# Recursive Language Models — Alex Zhang, MIT PhD

## Executive summary

The discussion explores the evolution of AI capability beyond simple frontier models, arguing that the design of the 'harness' (the surrounding program structure) is often more critical than the base Language Model (LM) itself. The focus is on Recursive Language Models (RLMs) and multi-agent swarms, which enable compositional generalization by allowing the model to treat its own sub-agents and tools as internal components. Key architectural shifts include context offloading, programmatic subagent calling, and moving from simple autoregressive decoding to complex, stateful, and highly structured workflows.

## Key takeaways

- Harness Design Drives Capability: The core argument is that modern coding agents (like Claude Code, Codex, Pi, and Prime Agent) are structurally very similar. The true breakthrough lies in designing sophisticated harnesses that enable compositional generalization, allowing a model to solve diverse tasks (e.g., math vs. writing) using the same underlying strategy.
- RLMs Enable Compositional Generalization: RLMs are presented as a primitive inductive bias that allows the model to maintain a central, persistent context (often stored on disk) while calling sub-agents. This structure allows the model to learn a core problem-solving strategy that is directly transferable across vastly different tasks and domains.
- The Importance of Specialized Benchmarking: The field benefits from specialized benchmarks (like SWE-bench, Quiet-STaR, and KernelBench) because they force the development of novel, non-obvious techniques, pushing the boundaries of what is considered 'possible' in AI systems.
- Systemic Efficiency and Optimization: For GPU kernels, optimization is not just about speed, but also memory and power consumption (pico/nano-joules). Furthermore, end-to-end models may require sacrificing speed in early operations to keep data within the cache for later, more critical operations (the 'fusion problem').

## Technical details

- Recursive Language Models (RLMs): RLMs are a harness design where the only tool is code, enabling programmatic subagent calling. They solve the problem of long context by offloading context to a central memory (e.g., disk), allowing the model to maintain a persistent state across complex, multi-step tasks.
- Agent Swarms and Multi-Agent Systems: These systems involve multiple sub-agents communicating over a shared context (like a message board or file system). The goal is to move beyond single-agent limitations by distributing the problem-solving effort, though coordination remains a bottleneck.
- Context Offloading and Persistence: Traditional harnesses often treat the context as a prompt (trajectory as a prompt). RLMs improve this by explicitly offloading context to a persistent memory store, ensuring the model can always reference its original state even after compaction or multiple agent runs.
- GPU Kernel Optimization: Optimization involves considering not just computational speed, but also memory bandwidth, power consumption, and the 'fusion' of operations to keep data within the cache hierarchy, which is crucial for end-to-end model performance.

## Practical implications

- Focusing development efforts on improving the *architecture* (the harness) rather than solely on increasing model size or raw compute power.
- The ability to design highly opinionated, specialized harnesses is key to unlocking latent model capabilities and achieving compositional generalization.
- For build engineers, this suggests that complex, multi-step workflows should be modeled as stateful, persistent systems (like RLMs) rather than simple sequential API calls.
- The trend towards multi-agent swarms and programmatic tool calling suggests that future AI systems will be less like monolithic models and more like orchestrated pipelines.

## Topics

Recursive Language Models (RLMs), Agent Swarms, Harness Design, Compositional Generalization, GPU Kernel Optimization, Context Offloading, AI Architecture, GPU Mode, Prime Agent, KernelBench, SWE-bench, Harvey (Legal AI)

Source: https://www.youtube.com/watch?v=kog7mwsDqnk
