Weights & Biases

Building Reliable AI Infrastructure at Scale with Extreme Co-design

Published 2026-10-07 · Duration 40:38

Summary

The presentation details how building reliable, next-generation AI infrastructure requires 'extreme co-design' across compute, networking, storage, and software. NVIDIA's Vera Rubin platform is presented as an integrated AI factory, moving beyond single components to provide a scalable, resilient, and highly efficient ecosystem. Key advancements include the shift from traditional Generative AI to complex Agentic AI, which demands specialized CPUs (Vera CPU) and sophisticated, multi-plane networking (Spectrum-X Ethernet) to ensure high throughput and fault tolerance at massive scale.

Download summary

Key takeaways

  1. The Shift to Agentic AI 22:17

    AI workloads are evolving from simple question-answering (Generative AI) to complex Agentic AI, which involves multi-step reasoning, tool-calling, and continuous context management. This requires a platform that supports complex, multi-tasking workflows, not just single-pass inference.

  2. Co-Design for Reliability and Scale 17:30

    The Vera Rubin platform is an integrated AI factory designed for co-design across seven chips and five racks. This holistic approach ensures that components (compute, networking, storage) work together seamlessly, maximizing efficiency and minimizing downtime.

  3. Advanced Networking and Resilience 27:30

    The architecture utilizes Multi-plane topology and technologies like Spectrum-X Ethernet with co-packaged optics (CPO) to achieve massive scale (supporting 512,000 GPUs) while maintaining low latency and high resilience. Hardware-level fault detection and recovery can occur in milliseconds, preventing application timeouts.

  4. Efficiency through Quantization and Sparsity 20:20

    The platform supports advanced techniques like NVFP4 (4-bit quantization) combined with sparsity features, enabling models to achieve quality comparable to 16-bit models while significantly boosting performance and efficiency.

Technical details

  • AI Infrastructure Architecture 200s

    NVIDIA positions itself not just as a chip manufacturer, but as an AI infrastructure builder, defining a five-layer stack: Power, Chips/Hardware, Infrastructure, Models, and Applications. The entire stack is built upon the CUDA foundation.

  • Vera Rubin Platform Components 1050s

    The platform consists of seven chips across five racks, including the Vera Rubin NVL 72 (for training/inference), Groq LPX (for inference acceleration), Vera CPU (for tool-calling/RL workloads), Vera BlueField 4 STX (storage), and Spectrum-X Ethernet (networking).

  • Performance Metrics (100 MW Factory) 1280s

    A 100 MW Vera Rubin factory can achieve inference capacity of up to 2 Zettaflops (NVFB4) with fast memory up to 42 Petabytes, providing flexibility for training, inference, and post-training workloads.

  • CPU Performance for Agentic AI 1380s

    The Vera CPU is optimized for Agentic AI, offering double the single-thread performance of current accelerators. Key features include 88 cores on a single die and a second-generation fabric providing 3.4 Tb/s bandwidth, crucial for coordinating multiple agents and maintaining context.

  • Software and Deployment (DSX) 1850s

    DSX is the software platform for designing and deploying the entire AI factory. DSX Max LPS can reallocate unused power across racks, increasing GPU capacity by up to 40% without increasing power budget.

Mentioned resources

  • NVIDIA Vera Rubin (AI Platform)
  • CUDA (Software Platform)
  • NVLink (Interconnect Technology)
  • Spectrum-X Ethernet (Networking Hardware)
  • A100 / H100 (GPU Accelerators)
  • DSX (NVIDIA Software) (Design/Deployment Platform)

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.