Topic

Latent Diffusion

All digests tagged Latent Diffusion

Runway’s Bet Beyond Video: World Models, Robotics, and the Neural OS — Anastasis Germanidis thumbnail

· 1:38:06

Runway’s Bet Beyond Video: World Models, Robotics, and the Neural OS — Anastasis Germanidis

Runway is evolving beyond generative video to build 'Interface World Models' and general-purpose world simulators. The core thesis is that scaling video models is sufficient to learn physics and human dynamics, making them suitable for robotics and simulating complex software interfaces. The company emphasizes that the ultimate goal is a fully neural operating system, where the interface itself is generated by the model, rather than relying on traditional code like HTML/CSS.

Key takeaways

  1. World Models as the Endgame 1:30:20

    The ultimate goal is a fully neural operating system where the model delivers the application end-to-end, generating both the language model output and the rendered pixels/interface. This shifts the focus from content creation to general-purpose world simulation.

  2. Scaling Video Models for Physics 1:14:40

    The belief is that if scaling laws apply to language models (LLMs), they will also apply to video. By scaling up video models, the system will inherently learn to simulate physics, human actions, and dynamics, making the model a general simulator.

  3. Third-Person Video Data Advantage 1:25:50

    The most plentiful source of data for training robotics models is third-person video data (observing others perform tasks), which is far more abundant than teleoperation or egocentric data. Video pre-training allows models to generalize to new environments and tasks.

  4. The Importance of Counterfactual Generation 1:29:10

    A key difference between standard video models and true world models is the ability to generate counterfactuals—simulating 'what if' scenarios (e.g., scoring a goal vs. failing to score a goal). This is critical for robust robotics training.

Watch on YouTube Full article