# GPU Died. Training Didn't: Self-Healing Training at Scale — Crusoe

## Executive summary

The talk details the architecture of Crusoe's managed infrastructure, which combines Slurm and Kubernetes to provide robust, self-healing platforms for large-scale AI training. Since hardware failures are inevitable when running thousands of GPUs, the solution integrates Slurm's familiar job scheduling (gang scheduling, topology awareness) with Kubernetes' infrastructure resilience and observability. The key feature, AutoClusters, automatically detects critical hardware errors (e.g., XID 79 GPU failure), drains the node, replaces it, and seamlessly resumes the training job from the last checkpoint with minimal human intervention.

## Key takeaways

- Automated Remediation Flow: AutoClusters automatically handles GPU node failures: the user is notified, the Slurm operator sets the node to down, the job is requeued, the unhealthy node is cordoned and replaced, and the training job resumes from the checkpoint. This process can complete in under 15 minutes.
- Seamless Integration: By running Slurm on Kubernetes, platform teams avoid managing two separate infrastructure stacks. Slurm users maintain their familiar `sbatch` workflows, while the platform gains Kubernetes' built-in self-healing, load balancing, and autoscaling features.
- Dynamic Resource Allocation: The unified infrastructure allows for dynamic resource shifting. GPUs can be reallocated between training clusters and inference services (or vice versa) as demand shifts, maximizing hardware utilization.

## Technical details

- Slurm's Strengths and Limitations: Slurm excels at high-performance job scheduling for multi-node training, offering features like gang scheduling and topology awareness. However, it is traditionally static and lacks modern dynamic resource management, comprehensive observability into failure causes (e.g., network flaps), and automated node health maintenance.
- Managed Slurm on Kubernetes Architecture: The architecture uses a Crusoe Slurm Operator running on a Managed Kubernetes service (CMK). This allows Slurm to leverage Kubernetes' maturity, built-in self-healing, and observability while maintaining the familiar Slurm user experience. The system treats the GPU nodes as Kubernetes nodes first.
- AutoClusters Remediation Process: Upon detecting a critical hardware error (e.g., XID 79), the process involves: 1) User notification; 2) Slurm operator setting the node to down and canceling jobs; 3) Kubernetes cordoning and draining the node; 4) Removing the unhealthy node and replacing it with a new healthy node; 5) The Slurm job restarting and resuming training from the checkpoint.

## Practical implications

- Reduces operational burden by automating hardware failure remediation, eliminating the need for manual intervention at 3 a.m.
- Enables seamless scaling and resource reallocation between different workloads (e.g., training and inference) within a single cluster.
- Allows ML engineers to continue using familiar Slurm workflows while benefiting from the resilience and observability of Kubernetes.

## Topics

High-Performance Computing (HPC), Container Orchestration, Distributed Training, Infrastructure as Code, Fault Tolerance, Crusoe Cloud, Managed Kubernetes (CMK), Managed Slurm, Slinky, AutoClusters

Source: https://www.youtube.com/watch?v=bRGyYaE0lxI
