# Keep It Secret, Keep It Safe: How IBM Deploys Sensitive Data and Workloads on CoreWeave

## Executive summary

IBM Research details its complex journey to deploying highly sensitive, large-scale AI and deep learning workloads (including foundational models) on CoreWeave's infrastructure. The discussion highlights the immense challenges of scaling compute (H100, GB200) while maintaining stringent security and compliance standards required by large financial and research institutions. Key architectural focus areas include establishing robust identity controls (PKI, federated identity), managing multi-cluster resource allocation (Slurm, Kubernetes), and implementing advanced security primitives like Confidential Computing to ensure data integrity and control access at every layer.

## Key takeaways

- Scaling AI Compute and Infrastructure: The deployment process involved significant hardware scaling, moving from internal H100 clusters to large-scale GB200 deployments, requiring careful management of power delivery and cooling (row vs. in-rack cooling). The successful scaling to over 4,000 GPUs demonstrated a high level of operational maturity.
- Identity and Security as the Central Challenge: Security is not an afterthought. The architecture emphasizes a 'federation first' approach, using systems like a global UID and GD namespace to maintain a single source of truth for identity across multiple clusters and services. This is critical for auditing and access control, especially when dealing with agentic and autonomous workloads.
- Evolving Trust Boundaries: The definition of trust has shifted from the traditional shared responsibility model to requiring cryptographic proof. CoreWeave supports this by offering solutions like RKE (Restricted Key Encryption) where data can be encrypted with keys held solely by the customer, ensuring the provider cannot access the data.
- Operationalizing Complex Workloads: The system must support diverse workloads—from pre-training and fine-tuning to inference and execution—often involving malicious prompt testing. This requires sophisticated sandboxing environments and tight control over ingress/egress rules to prevent lateral movement and unauthorized access.

## Technical details

- Compute Scaling and Hardware: The deployment utilized advanced hardware like H100 and GB200 clusters. Scaling required addressing physical constraints, including power delivery and cooling methods (row vs. in-rack).
- Identity and Access Management (IAM): IBM implemented a global UID and GD namespace to unify identity across clusters. The process moved from simple SSH keys to modern, controlled methods like using short-lived keys (replacing 'bring your own key'). The system leverages Red Hat directory services as the source of truth.
- Workload Orchestration and Control: Workloads are managed across multiple systems, including Slurm and Kubernetes. The architecture must coordinate these disparate systems, ensuring that user identity, Slurm accounting, and GPFS are all linked. The introduction of sandboxing is key for running potentially malicious or exploratory agentic workloads.
- Security and Compliance: Security controls are paramount, driven by internal audits and CISO requirements. Key controls include: 1) Implementing granular ingress and egress rules; 2) Utilizing Confidential Computing/Trusted Execution Environments to protect the workload cryptographically; and 3) Establishing telemetry endpoints for auditing purposes (e.g., integrating CrowdStrike feeds).

## Practical implications

- For build engineers, the core lesson is that AI infrastructure security must be designed around a 'federation first' identity model, rather than relying on a single system of record.
- When planning large-scale compute deployments, anticipate the complexity of integrating disparate systems (e.g., Slurm, Kubernetes, storage) under a unified, auditable identity framework.
- The shift in trust requires moving beyond simple network segmentation and implementing cryptographic controls (like client-side encryption keys) to prove data ownership and protection.
- Designing for agentic workloads necessitates dedicated, highly controlled sandboxing environments to manage the risk of autonomous or malicious code execution.

## Topics

AI Infrastructure, High-Performance Computing (HPC), Cloud Security, Identity Management (PKI), Container Orchestration (Kubernetes), Deep Learning

Source: https://www.youtube.com/watch?v=2XdarmFGTFY
