The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Cerebras CTO Sean Lie details the shift in AI infrastructure from merely increasing model size to achieving ultra-fast inference speeds. The core argument is that speed (throughput) itself enables entirely new classes of intelligent applications and agents. Cerebras showcased its CS4 wafer-scale architecture, demonstrating GPTOSS running at over 4,400 tokens per second (TPS). They previewed the next generation, CS5, which aims for up to 10,000 TPS on medium models and 5,000 TPS on frontier models. The industry is moving toward heterogeneous, disaggregated systems that integrate specialized components like OpenAI's Jalapeño chip to solve complex scaling challenges.
Key takeaways
-
Speed Defines Capability
0:35
What was once considered fast (100–200 TPS) is now viewed as 'batch mode.' Ultra-fast inference enables more reasoning loops and significantly more capable agents, transforming previously offline applications into real-time experiences.
-
CS4 Performance Milestone
3:05
The CS4 architecture is a modular platform designed to bring wafer scale to hyperscale. It provides significantly more power and interconnect bandwidth, demonstrated by running GPTOSS at over 4,400 TPS.
-
CS5 Roadmap (Preview)
7:56
The next generation CS5 platform is designed for multiple generations of products. It aims to push performance further: up to 10,000 TPS for medium models (e.g., Gemma) and up to 5,000 TPS for frontier models (e.g., Gemini/DeepSeek).
-
The Future is Heterogeneous & Disaggregated
11:20
Lie argues that the future of AI infrastructure requires integrating specialized components—such as prefill, attention, and KV cache loading—across different hardware types (e.g., wafer scale, SRAM-based chips like Jalapeño) to solve scaling challenges.