Building Reliable AI Infrastructure at Scale with Extreme Co-design
The presentation details how building reliable, next-generation AI infrastructure requires 'extreme co-design' across compute, networking, storage, and software. NVIDIA's Vera Rubin platform is presented as an integrated AI factory, moving beyond single components to provide a scalable, resilient, and highly efficient ecosystem. Key advancements include the shift from traditional Generative AI to complex Agentic AI, which demands specialized CPUs (Vera CPU) and sophisticated, multi-plane networking (Spectrum-X Ethernet) to ensure high throughput and fault tolerance at massive scale.
Key takeaways
-
The Shift to Agentic AI
22:17
AI workloads are evolving from simple question-answering (Generative AI) to complex Agentic AI, which involves multi-step reasoning, tool-calling, and continuous context management. This requires a platform that supports complex, multi-tasking workflows, not just single-pass inference.
-
Co-Design for Reliability and Scale
17:30
The Vera Rubin platform is an integrated AI factory designed for co-design across seven chips and five racks. This holistic approach ensures that components (compute, networking, storage) work together seamlessly, maximizing efficiency and minimizing downtime.
-
Advanced Networking and Resilience
27:30
The architecture utilizes Multi-plane topology and technologies like Spectrum-X Ethernet with co-packaged optics (CPO) to achieve massive scale (supporting 512,000 GPUs) while maintaining low latency and high resilience. Hardware-level fault detection and recovery can occur in milliseconds, preventing application timeouts.
-
Efficiency through Quantization and Sparsity
20:20
The platform supports advanced techniques like NVFP4 (4-bit quantization) combined with sparsity features, enabling models to achieve quality comparable to 16-bit models while significantly boosting performance and efficiency.