AI Engineer

GPU Died. Training Didn't: Self-Healing Training at Scale — Crusoe

Published 2026-10-03 · Duration 16:54

Summary

The talk details the architecture of Crusoe's managed infrastructure, which combines Slurm and Kubernetes to provide robust, self-healing platforms for large-scale AI training. Since hardware failures are inevitable when running thousands of GPUs, the solution integrates Slurm's familiar job scheduling (gang scheduling, topology awareness) with Kubernetes' infrastructure resilience and observability. The key feature, AutoClusters, automatically detects critical hardware errors (e.g., XID 79 GPU failure), drains the node, replaces it, and seamlessly resumes the training job from the last checkpoint with minimal human intervention.

Download summary

Key takeaways

  1. Automated Remediation Flow 12:15

    AutoClusters automatically handles GPU node failures: the user is notified, the Slurm operator sets the node to down, the job is requeued, the unhealthy node is cordoned and replaced, and the training job resumes from the checkpoint. This process can complete in under 15 minutes.

  2. Seamless Integration 6:30

    By running Slurm on Kubernetes, platform teams avoid managing two separate infrastructure stacks. Slurm users maintain their familiar `sbatch` workflows, while the platform gains Kubernetes' built-in self-healing, load balancing, and autoscaling features.

  3. Dynamic Resource Allocation 8:00

    The unified infrastructure allows for dynamic resource shifting. GPUs can be reallocated between training clusters and inference services (or vice versa) as demand shifts, maximizing hardware utilization.

Technical details

  • Slurm's Strengths and Limitations 220s

    Slurm excels at high-performance job scheduling for multi-node training, offering features like gang scheduling and topology awareness. However, it is traditionally static and lacks modern dynamic resource management, comprehensive observability into failure causes (e.g., network flaps), and automated node health maintenance.

  • Managed Slurm on Kubernetes Architecture 390s

    The architecture uses a Crusoe Slurm Operator running on a Managed Kubernetes service (CMK). This allows Slurm to leverage Kubernetes' maturity, built-in self-healing, and observability while maintaining the familiar Slurm user experience. The system treats the GPU nodes as Kubernetes nodes first.

  • AutoClusters Remediation Process 560s

    Upon detecting a critical hardware error (e.g., XID 79), the process involves: 1) User notification; 2) Slurm operator setting the node to down and canceling jobs; 3) Kubernetes cordoning and draining the node; 4) Removing the unhealthy node and replacing it with a new healthy node; 5) The Slurm job restarting and resuming training from the checkpoint.

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.