Latent Space

The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO

Published 2026-09-02 · Duration 44:03

Summary

Cerebras CTO Sean Lie details the shift in AI infrastructure from merely increasing model size to achieving ultra-fast inference speeds. The core argument is that speed (throughput) itself enables entirely new classes of intelligent applications and agents. Cerebras showcased its CS4 wafer-scale architecture, demonstrating GPTOSS running at over 4,400 tokens per second (TPS). They previewed the next generation, CS5, which aims for up to 10,000 TPS on medium models and 5,000 TPS on frontier models. The industry is moving toward heterogeneous, disaggregated systems that integrate specialized components like OpenAI's Jalapeño chip to solve complex scaling challenges.

Download summary

Key takeaways

  1. Speed Defines Capability 0:35

    What was once considered fast (100–200 TPS) is now viewed as 'batch mode.' Ultra-fast inference enables more reasoning loops and significantly more capable agents, transforming previously offline applications into real-time experiences.

  2. CS4 Performance Milestone 3:05

    The CS4 architecture is a modular platform designed to bring wafer scale to hyperscale. It provides significantly more power and interconnect bandwidth, demonstrated by running GPTOSS at over 4,400 TPS.

  3. CS5 Roadmap (Preview) 7:56

    The next generation CS5 platform is designed for multiple generations of products. It aims to push performance further: up to 10,000 TPS for medium models (e.g., Gemma) and up to 5,000 TPS for frontier models (e.g., Gemini/DeepSeek).

  4. The Future is Heterogeneous & Disaggregated 11:20

    Lie argues that the future of AI infrastructure requires integrating specialized components—such as prefill, attention, and KV cache loading—across different hardware types (e.g., wafer scale, SRAM-based chips like Jalapeño) to solve scaling challenges.

Technical details

  • Wafer Scale Architecture 195s

    Cerebras' modular platform is designed to bring wafer scale computing to hyperscale data centers. The architecture increases power delivery and interconnect bandwidth compared to previous generations.

  • Inference Throughput Benchmarks 236s

    CS4 demonstrated GPTOSS running at over 4,400 TPS. CS5 is projected to reach up to 10,000 TPS (medium models) and 5,000 TPS (frontier models).

  • AI-First Chip Design 760s

    OpenAI’s Jalapeño chip exemplifies an AI-first design methodology, which Lie argues is the future of the semiconductor industry. This approach enables significant gains in both throughput and latency.

  • System Scaling Challenges 850s

    The necessity for wafer scale integration stems from the massive memory requirements of frontier models (e.g., trillions of parameters), which cannot be contained on single chips, necessitating advanced packaging and interconnect solutions.

Mentioned resources

  • CS4 (Hardware Architecture)
  • CS5 (Hardware Architecture (Preview))
  • Jalapeño chip (AI Accelerator Chip)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.