# Training Cognition SWE-1.7: Asynchronous RL at Global Scale

## Executive summary

This talk details the complex systems and algorithmic recipes required to train large coding models (e.g., SWE-1.7) at a global, massive scale. Key challenges addressed include optimizing the trade-off between computational cost and performance, maintaining training stability in asynchronous environments, and ensuring fault tolerance across distributed hardware. Solutions involve techniques like sparse weight update compression, decoupling training from inference, and implementing robust fault-tolerant infrastructure.

## Key takeaways

- Optimizing Cost vs. Performance: The objective function for RL training must optimize the trade-off between cost (computational effort) and performance (solved problems). This is achieved by applying a linear penalty to the reward function, where the optimal coefficient is determined by the slope of the tangent to the model's test-time scaling curve.
- Asynchronous RL for Scale: To maximize GPU throughput and avoid idle time caused by variable problem durations, the system uses Asynchronous RL. This allows the model weights to be updated while rollouts are continuously generated, meaning different tokens in a single trajectory can be generated by different model versions.
- Distributed Inference vs. Training: Inference engines can be distributed across multiple continents because each engine occupies a small, independent amount of GPUs. However, training GPUs require synchronizing gradients and must be collocated in a single, large cluster.

## Technical details

- Weight Update Compression and Synchronization: Weight updates in RL are extremely sparse element-wise. This property allows for compression (up to 99%) by only sending the list of changed bits (delta). This compressed delta can be pulled independently by distributed inference engines, allowing asynchronous weight updates without stopping the generation process.
- Fault Tolerance: Fault tolerance is implemented differently for training and inference. Inference engines are designed with individual fault tolerance units, allowing requests to be rerouted if a node fails. Training utilizes large clusters (e.g., Blackwell/GB200) and leverages spare nodes to auto-swap faulty nodes, ensuring long-term run stability.
- Training Stability and Mismatch Mitigation: In asynchronous RL, distributional mismatch can occur due to differences in weight versions, numerics, or sampling methods between training and inference. Stability is maintained by 'replaying' the exact quantization procedures (e.g., using FP8 on the RoPE part of the MLA layer) and sampling strategies (e.g., Top-P) used during inference during the training process.

## Practical implications

- The ability to decouple training and inference allows for highly efficient, globally distributed inference fleets.
- Weight update compression significantly reduces network bandwidth requirements for large-scale model deployment.
- The necessity of replaying inference-time quantization and sampling strategies during training is critical for achieving stable, high-performance RL runs.

## Topics

Reinforcement Learning (RL), Large Language Models (LLMs), Distributed Computing, Hardware Acceleration, Model Optimization, Frontier Code, Weights & Biases

Source: https://www.youtube.com/watch?v=GgZnWDcDxvE
