Topic

World Models

All digests tagged World Models

Runway’s Bet Beyond Video: World Models, Robotics, and the Neural OS — Anastasis Germanidis thumbnail

· 1:38:06

Runway’s Bet Beyond Video: World Models, Robotics, and the Neural OS — Anastasis Germanidis

Runway is evolving beyond generative video to build 'Interface World Models' and general-purpose world simulators. The core thesis is that scaling video models is sufficient to learn physics and human dynamics, making them suitable for robotics and simulating complex software interfaces. The company emphasizes that the ultimate goal is a fully neural operating system, where the interface itself is generated by the model, rather than relying on traditional code like HTML/CSS.

Key takeaways

  1. World Models as the Endgame 1:30:20

    The ultimate goal is a fully neural operating system where the model delivers the application end-to-end, generating both the language model output and the rendered pixels/interface. This shifts the focus from content creation to general-purpose world simulation.

  2. Scaling Video Models for Physics 1:14:40

    The belief is that if scaling laws apply to language models (LLMs), they will also apply to video. By scaling up video models, the system will inherently learn to simulate physics, human actions, and dynamics, making the model a general simulator.

  3. Third-Person Video Data Advantage 1:25:50

    The most plentiful source of data for training robotics models is third-person video data (observing others perform tasks), which is far more abundant than teleoperation or egocentric data. Video pre-training allows models to generalize to new environments and tasks.

  4. The Importance of Counterfactual Generation 1:29:10

    A key difference between standard video models and true world models is the ability to generate counterfactuals—simulating 'what if' scenarios (e.g., scoring a goal vs. failing to score a goal). This is critical for robust robotics training.

Watch on YouTube Full article

One Operator, Many Drones: Inside Skydio's Autonomy Stack — Suchet Bargoti, Skydio thumbnail

· 20:48

One Operator, Many Drones: Inside Skydio's Autonomy Stack — Suchet Bargoti, Skydio

Skydio presented its full-stack autonomy solution, demonstrating how drones are evolving from hobbyist tools into critical infrastructure. The system enables large-scale, multi-agent orchestration, allowing a single operator to manage multiple drones performing diverse tasks (e.g., utility inspection, tracking stolen vehicles) across different geographical locations simultaneously. The core technical advancements involve splitting intelligence between the edge (on-drone actions) and the cloud (long-term planning, heavy lifting), utilizing World Models for global path planning, and employing Visual Language Models (VLMs) for agentic, rule-free object tracking and semantic reasoning.

Key takeaways

  1. Drones as Infrastructure 2:00

    Skydio is positioning its drones as critical infrastructure, with thousands of docks deployed across utilities, public safety, and construction sectors. This allows for continuous, reliable operation (day/night, rain/sunshine) and scales beyond the limitations of requiring a dedicated pilot for every incident.

  2. Full-Stack Autonomy Architecture 18:50

    The autonomy stack splits intelligence between the edge (for immediate actions) and the cloud (for heavy lifting and long-term planning). This architecture is designed to maintain high reliability (targeting 99.9999%) while managing vast amounts of data and complex decision-making.

  3. Agentic Orchestration

    The system moves beyond hand-coded rules by using agentic tools. A VLM can receive a high-level command (e.g., 'find a white Jeep') and autonomously access APIs and tools to command a drone's trajectory, enabling 'find and follow' without specific coding for every scenario.

Watch on YouTube Full article

Physical AI's Next Bottleneck Is Finding the Right Video — Rafael Levi, Bright Data thumbnail

· 16:32

Physical AI's Next Bottleneck Is Finding the Right Video — Rafael Levi, Bright Data

The primary bottleneck in advanced AI development is no longer the model architecture, but the availability and quality of training data. While Large Language Models (LLMs) have access to trillions of words, robotics training is severely limited, relying on small, often biased datasets (e.g., only about a million robot videos). The speaker, Rafael Levi of Bright Data, argues that the solution lies in leveraging the massive, real-world video data available on the web. Bright Data's approach, 'search first, collect second,' indexes billions of videos by specific actions, allowing users to retrieve precisely trimmed, ready-to-use video snippets via an API, drastically reducing data noise and collection waste compared to traditional methods.

Key takeaways

  1. The Data Bottleneck in Robotics 4:00

    For robotics, training data is scarce compared to text (trillions of words) or images (billions of labeled images), with current datasets limited to about a million videos of robots doing actions. (2:40)

  2. Bias in Staged Data 3:15

    Paying people to record staged actions (e.g., opening a door) results in biased data because human behavior is different when uninstructed versus when performing a task for recording. (3:15)

  3. The Web as a Solution 5:14

    The web holds billions of hours of real-world video showing natural physics, gravity, and cause-and-effect interactions, making it a massive training source for world models. (5:14)

  4. Bright Data's Indexing Approach 8:33

    Instead of downloading massive amounts of video and discarding the majority (e.g., NVIDIA discarding 96% of downloaded video), Bright Data indexes videos by specific actions, allowing users to search for precise clips (e.g., 'washing dishes,' 'folding a t-shirt'). (8:33)

Watch on YouTube Full article

When Will AI Make Me Scrambled Eggs? I Went To NVIDIA To Find Out. thumbnail

· 46:27

When Will AI Make Me Scrambled Eggs? I Went To NVIDIA To Find Out.

The video details the shift in AI from Large Language Models (LLMs) generating text to World Models (WMs) that generate physical actions and simulations. NVIDIA, through its Cosmos Lab, is building WMs to enable physical AI in complex domains like robotics, self-driving cars, and factory automation. The Cosmos 3 platform fuses world understanding, world simulation, and action capability into a single model, allowing developers to test and verify policies in a simulated environment before real-world deployment. The core architectural components include a reasoner, a generator, and an action module.

Key takeaways

  1. World Models vs. LLMs

    While LLMs are symbolic and semantic (dealing with text), World Models are designed to produce direct actions (e.g., issuing guidance to a robot arm) and model the physical world. WMs allow developers to simulate complex physical scenarios (like a factory floor) without needing to build thousands of physical prototypes.

  2. The Cosmos 3 Architecture 16:30

    The Cosmos platform integrates three key components into a single model: world understanding (interpreting the physical state), world simulation (predicting how the world changes), and action capability (generating physical commands). This unified approach is critical for physical AI applications.

  3. Scaling and Deployment 19:17

    WMs are designed to operate in real-time, necessitating models of different sizes (e.g., Super Nano and Nano). The architecture supports a mix of deployment environments—from embedded devices (like Jetson or Dig Spark) to powerful data centers—to balance performance and computational constraints.

  4. Verifiable Reward and Simulation

    A major advantage of WMs is the ability to perform policy verification in simulation. This allows engineers to test safety and performance (e.g., for self-driving cars) across thousands of edge cases, dramatically accelerating development velocity compared to physical testing.

Watch on YouTube Full article

SOTA Generative Media Panel — Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind thumbnail

· 56:59

SOTA Generative Media Panel — Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind

The session provided an overview of SOTA generative media models, highlighting new APIs like NanoBanana 2 Lite and Gemini Omni Flash. Key architectural discussions centered on the limitations of language as a sole intermediate representation for complex sensory data (taste, smell, skin tone). The consensus points toward a future requiring unified 'World Models' that integrate visual, temporal, and symbolic reasoning, moving beyond single-modality generation. Evaluation remains highly dependent on human judgment, making robust testing and field feedback critical.

Key takeaways

  1. New APIs Launched for Developers 0:15

    Google launched NanoBanana 2 Lite (the fastest/cheapest image model in the family) and the Gemini Omni Flash APIs. The Omni Flash API enables video generation and editing, priced similarly to V3 fast, making it accessible for developers [0:15-0:40].

  2. Generative Media Capabilities 2:00

    Models can now take diverse inputs (e.g., a storyboard of images, an audio track) to generate video. Furthermore, natural language processing allows for advanced video editing tasks like adding or removing elements from existing footage [1:20-3:00].

  3. The Limitation of Language as Representation 1:50

    Speakers argued that language is an insufficient intermediate representation for highly sensitive sensory data (e.g., taste, smell, skin tone). This suggests a need for more foundational representations, potentially including code or direct binary/latent space conditioning [1:50-2:30].

  4. Evaluation Challenges and Reward Hacking 0:35

    While human preference often favors AI output (e.g., sharper, more saturated images), this metric is unreliable for optimization. External testers have found 'reward hacking' artifacts, such as the model consistently adding wedding rings to hands [0:35-0:45].

Watch on YouTube Full article

Voice agents with Realtime Video — Sidney Primas, LemonSlice thumbnail

· 26:36

Voice agents with Realtime Video — Sidney Primas, LemonSlice

LemonSlice aims to break the Avatar Turing test by creating highly realistic, real-time video avatars. The core technical challenges addressed include achieving emotional expressiveness (requiring specialized audio encoders beyond monotone audiobook data), mitigating error accumulation over extended generation periods (e.g., 8+ hours), and optimizing for real-time performance. A major focus is on the 'model harness'—the complex orchestration of threads and queues across GPU/CPU to ensure uninterrupted, stutter-free video streaming at scale. The company also highlights cost parity, noting that generating high-resolution video costs comparably to running a voice model.

Key takeaways

  1. Real-Time Video Generation Challenges 17:05

    Generating real-time avatars requires training models with an attention mask that enforces looking only into the past, as future frames do not exist yet. Furthermore, speeding up generation involves collapsing many denoising steps (e.g., 30 steps) down to a single step.

  2. Error Accumulation Mitigation 20:22

    A significant technical hurdle is error accumulation, where errors introduced in previous video blocks compound over time. The company claims to have developed a novel method to generate very long videos with no noticeable error buildup.

  3. Model Harness and Cost Parity 23:54

    The most durable value lies in the 'model harness'—the orchestration of threads and queues across GPU and CPU to maintain real-time, stutter-free video. Surprisingly, serving this complex visual layer costs about the same as serving a voice model.

  4. Future Direction: Emotional Engine

    The next generation involves building an 'emotion engine' that predicts and controls emotional reactions and actions based on both audio input and text input, moving beyond current awkward interactions.

Watch on YouTube Full article