GPU Died. Training Didn't: Self-Healing Training at Scale — Crusoe
The talk details the architecture of Crusoe's managed infrastructure, which combines Slurm and Kubernetes to provide robust, self-healing platforms for large-scale AI training. Since hardware failures are inevitable when running thousands of GPUs, the solution integrates Slurm's familiar job scheduling (gang scheduling, topology awareness) with Kubernetes' infrastructure resilience and observability. The key feature, AutoClusters, automatically detects critical hardware errors (e.g., XID 79 GPU failure), drains the node, replaces it, and seamlessly resumes the training job from the last checkpoint with minimal human intervention.
Key takeaways
-
Automated Remediation Flow
12:15
AutoClusters automatically handles GPU node failures: the user is notified, the Slurm operator sets the node to down, the job is requeued, the unhealthy node is cordoned and replaced, and the training job resumes from the checkpoint. This process can complete in under 15 minutes.
-
Seamless Integration
6:30
By running Slurm on Kubernetes, platform teams avoid managing two separate infrastructure stacks. Slurm users maintain their familiar `sbatch` workflows, while the platform gains Kubernetes' built-in self-healing, load balancing, and autoscaling features.
-
Dynamic Resource Allocation
8:00
The unified infrastructure allows for dynamic resource shifting. GPUs can be reallocated between training clusters and inference services (or vice versa) as demand shifts, maximizing hardware utilization.