Topic

Data Center Architecture

All digests tagged Data Center Architecture

Scaling in the Agentic Era with NVIDIA Vera Rubin NVL72 thumbnail

· 24:38

Scaling in the Agentic Era with NVIDIA Vera Rubin NVL72

The session details the infrastructure required to support the 'agentic era' of AI, which demands continuous, multi-step, and highly interactive workloads. CoreWeave leverages the NVIDIA Vera Rubin NVL72 platform to address critical scaling challenges related to memory, interactivity, and power. Key innovations include the development of specialized hardware management systems (Racky, Valve) and advanced automation (RLCC) to ensure optimal performance and reliability across the entire rack, enabling customers to move sophisticated AI systems from experimentation to reliable production scale.

Key takeaways

  1. Shift to Agentic AI Workloads 1:30

    Agentic systems require persistent state and continuous, multi-step reasoning, moving beyond simple chat interactions. This shift increases the demand for high-speed, low-latency infrastructure and necessitates managing complex, distributed operations (Timestamp: 0:01:30-0:02:10).

  2. NVL72 Performance Gains 3:50

    Comparing Vera Rubin NVL72 to previous generations (like GB200), benchmarks showed a significant improvement in efficiency, specifically increasing token throughput from 80,000 tokens/s/MW to 800,000 tokens/s/MW, demonstrating superior power efficiency for large-scale inference (Timestamp: 0:03:50-0:04:30).

  3. Comprehensive Rack Management 5:10

    The platform integrates specialized hardware components like Racky (patent pending rack manager) and Valve (for remote CDU telemetry) to provide centralized control over power, cooling, and leak detection, ensuring optimal component health and preventing catastrophic failures (Timestamp: 0:05:10-0:06:30).

  4. Software-Defined Infrastructure Control 6:30

    CoreWeave utilizes RLCC (Rack Lifecycle Controller) and Kubernetes operators to manage the entire rack as a single unit, addressing the complexity of distributed power, cooling, and high-speed interconnects (e.g., NVL link). This allows for faster, guaranteed deployment of complex compute resources (Timestamp: 0:06:30-0:07:50).

Watch on YouTube Full article