Weights & Biases

Scaling in the Agentic Era with NVIDIA Vera Rubin NVL72

Published 2026-10-07 · Duration 24:38

Summary

The session details the infrastructure required to support the 'agentic era' of AI, which demands continuous, multi-step, and highly interactive workloads. CoreWeave leverages the NVIDIA Vera Rubin NVL72 platform to address critical scaling challenges related to memory, interactivity, and power. Key innovations include the development of specialized hardware management systems (Racky, Valve) and advanced automation (RLCC) to ensure optimal performance and reliability across the entire rack, enabling customers to move sophisticated AI systems from experimentation to reliable production scale.

Download summary

Key takeaways

  1. Shift to Agentic AI Workloads 1:30

    Agentic systems require persistent state and continuous, multi-step reasoning, moving beyond simple chat interactions. This shift increases the demand for high-speed, low-latency infrastructure and necessitates managing complex, distributed operations (Timestamp: 0:01:30-0:02:10).

  2. NVL72 Performance Gains 3:50

    Comparing Vera Rubin NVL72 to previous generations (like GB200), benchmarks showed a significant improvement in efficiency, specifically increasing token throughput from 80,000 tokens/s/MW to 800,000 tokens/s/MW, demonstrating superior power efficiency for large-scale inference (Timestamp: 0:03:50-0:04:30).

  3. Comprehensive Rack Management 5:10

    The platform integrates specialized hardware components like Racky (patent pending rack manager) and Valve (for remote CDU telemetry) to provide centralized control over power, cooling, and leak detection, ensuring optimal component health and preventing catastrophic failures (Timestamp: 0:05:10-0:06:30).

  4. Software-Defined Infrastructure Control 6:30

    CoreWeave utilizes RLCC (Rack Lifecycle Controller) and Kubernetes operators to manage the entire rack as a single unit, addressing the complexity of distributed power, cooling, and high-speed interconnects (e.g., NVL link). This allows for faster, guaranteed deployment of complex compute resources (Timestamp: 0:06:30-0:07:50).

Technical details

  • Hardware Architecture 160s

    The Vera Rubin NVL72 platform integrates 36 CPUs and 72 GPUs within a high-bandwidth domain. The architecture is enhanced by new components including NVL link switches and CX9 backend Ethernet switches, making the entire rack function as a single, cohesive computer (Timestamp: 0:02:40-0:03:10).

  • Cooling and Power Management 320s

    The system transitioned from in-rack CDUs to remote CDUs, requiring the development of Valve (for telemetry like differential pressure and flow rate) and Racky (for centralized power and cooling control) to maintain per-rack observability (Timestamp: 0:05:20-0:06:00).

  • Software Orchestration 390s

    CoreWeave uses Kubernetes with custom operators and controllers (RLCC) to manage the rack's complex components (power shelves, CDUs, interconnects). This ensures the entire rack is validated and operational before deployment, treating the rack as the primary compute unit (Timestamp: 0:06:30-0:07:30).

  • Workload Placement and Observability 500s

    The Kubernetes service now exposes labels for NVLink domains and rack numbers, allowing users to enforce workload affinities. Observability is provided via 'Cabinet Wrangler' and 'Cabinet Virtualizer' dashboards, enabling monitoring of the entire rack or data center, not just individual nodes (Timestamp: 0:08:20-0:09:30).

Mentioned resources

  • NVIDIA Vera Rubin NVL72 (Compute Platform)
  • NVIDIA Spectrum 6 (Liquid Cooled Switch)
  • Racky McRrack Manager (Patent Pending Hardware)
  • Valve McValve Assembly (Patent Pending Hardware)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.