Keep It Secret, Keep It Safe: How IBM Deploys Sensitive Data and Workloads on CoreWeave
Summary
IBM Research details its complex journey to deploying highly sensitive, large-scale AI and deep learning workloads (including foundational models) on CoreWeave's infrastructure. The discussion highlights the immense challenges of scaling compute (H100, GB200) while maintaining stringent security and compliance standards required by large financial and research institutions. Key architectural focus areas include establishing robust identity controls (PKI, federated identity), managing multi-cluster resource allocation (Slurm, Kubernetes), and implementing advanced security primitives like Confidential Computing to ensure data integrity and control access at every layer.
Key takeaways
-
Scaling AI Compute and Infrastructure
20:00
The deployment process involved significant hardware scaling, moving from internal H100 clusters to large-scale GB200 deployments, requiring careful management of power delivery and cooling (row vs. in-rack cooling). The successful scaling to over 4,000 GPUs demonstrated a high level of operational maturity.
-
Identity and Security as the Central Challenge
23:20
Security is not an afterthought. The architecture emphasizes a 'federation first' approach, using systems like a global UID and GD namespace to maintain a single source of truth for identity across multiple clusters and services. This is critical for auditing and access control, especially when dealing with agentic and autonomous workloads.
-
Evolving Trust Boundaries
30:00
The definition of trust has shifted from the traditional shared responsibility model to requiring cryptographic proof. CoreWeave supports this by offering solutions like RKE (Restricted Key Encryption) where data can be encrypted with keys held solely by the customer, ensuring the provider cannot access the data.
-
Operationalizing Complex Workloads
33:20
The system must support diverse workloads—from pre-training and fine-tuning to inference and execution—often involving malicious prompt testing. This requires sophisticated sandboxing environments and tight control over ingress/egress rules to prevent lateral movement and unauthorized access.
Technical details
-
Compute Scaling and Hardware
600s
The deployment utilized advanced hardware like H100 and GB200 clusters. Scaling required addressing physical constraints, including power delivery and cooling methods (row vs. in-rack).
-
Identity and Access Management (IAM)
1400s
IBM implemented a global UID and GD namespace to unify identity across clusters. The process moved from simple SSH keys to modern, controlled methods like using short-lived keys (replacing 'bring your own key'). The system leverages Red Hat directory services as the source of truth.
-
Workload Orchestration and Control
1600s
Workloads are managed across multiple systems, including Slurm and Kubernetes. The architecture must coordinate these disparate systems, ensuring that user identity, Slurm accounting, and GPFS are all linked. The introduction of sandboxing is key for running potentially malicious or exploratory agentic workloads.
-
Security and Compliance
1800s
Security controls are paramount, driven by internal audits and CISO requirements. Key controls include: 1) Implementing granular ingress and egress rules; 2) Utilizing Confidential Computing/Trusted Execution Environments to protect the workload cryptographically; and 3) Establishing telemetry endpoints for auditing purposes (e.g., integrating CrowdStrike feeds).
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.