Latent Space

Recursive Language Models — Alex Zhang, MIT PhD

Published 2026-10-02 · Duration 1:43:25

Summary

The discussion explores the evolution of AI capability beyond simple frontier models, arguing that the design of the 'harness' (the surrounding program structure) is often more critical than the base Language Model (LM) itself. The focus is on Recursive Language Models (RLMs) and multi-agent swarms, which enable compositional generalization by allowing the model to treat its own sub-agents and tools as internal components. Key architectural shifts include context offloading, programmatic subagent calling, and moving from simple autoregressive decoding to complex, stateful, and highly structured workflows.

Download summary

Key takeaways

  1. Harness Design Drives Capability 27:52

    The core argument is that modern coding agents (like Claude Code, Codex, Pi, and Prime Agent) are structurally very similar. The true breakthrough lies in designing sophisticated harnesses that enable compositional generalization, allowing a model to solve diverse tasks (e.g., math vs. writing) using the same underlying strategy.

  2. RLMs Enable Compositional Generalization 30:00

    RLMs are presented as a primitive inductive bias that allows the model to maintain a central, persistent context (often stored on disk) while calling sub-agents. This structure allows the model to learn a core problem-solving strategy that is directly transferable across vastly different tasks and domains.

  3. The Importance of Specialized Benchmarking 1:43

    The field benefits from specialized benchmarks (like SWE-bench, Quiet-STaR, and KernelBench) because they force the development of novel, non-obvious techniques, pushing the boundaries of what is considered 'possible' in AI systems.

  4. Systemic Efficiency and Optimization 20:00

    For GPU kernels, optimization is not just about speed, but also memory and power consumption (pico/nano-joules). Furthermore, end-to-end models may require sacrificing speed in early operations to keep data within the cache for later, more critical operations (the 'fusion problem').

Technical details

  • Recursive Language Models (RLMs) 1720s

    RLMs are a harness design where the only tool is code, enabling programmatic subagent calling. They solve the problem of long context by offloading context to a central memory (e.g., disk), allowing the model to maintain a persistent state across complex, multi-step tasks.

  • Agent Swarms and Multi-Agent Systems 1900s

    These systems involve multiple sub-agents communicating over a shared context (like a message board or file system). The goal is to move beyond single-agent limitations by distributing the problem-solving effort, though coordination remains a bottleneck.

  • Context Offloading and Persistence 1720s

    Traditional harnesses often treat the context as a prompt (trajectory as a prompt). RLMs improve this by explicitly offloading context to a persistent memory store, ensuring the model can always reference its original state even after compaction or multiple agent runs.

  • GPU Kernel Optimization 1200s

    Optimization involves considering not just computational speed, but also memory bandwidth, power consumption, and the 'fusion' of operations to keep data within the cache hierarchy, which is crucial for end-to-end model performance.

Mentioned resources

  • GPU Mode (Community/Platform)
  • Prime Agent (Harness/System)
  • KernelBench (Benchmark)
  • SWE-bench (Benchmark)
  • Harvey (Legal AI) (Industry Example)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.