Weights & Biases

Keep It Secret, Keep It Safe: How IBM Deploys Sensitive Data and Workloads on CoreWeave

Published 2026-10-09 · Duration 35:05

Summary

IBM Research details its complex journey to deploying highly sensitive, large-scale AI and deep learning workloads (including foundational models) on CoreWeave's infrastructure. The discussion highlights the immense challenges of scaling compute (H100, GB200) while maintaining stringent security and compliance standards required by large financial and research institutions. Key architectural focus areas include establishing robust identity controls (PKI, federated identity), managing multi-cluster resource allocation (Slurm, Kubernetes), and implementing advanced security primitives like Confidential Computing to ensure data integrity and control access at every layer.

Download summary

Key takeaways

  1. Scaling AI Compute and Infrastructure 20:00

    The deployment process involved significant hardware scaling, moving from internal H100 clusters to large-scale GB200 deployments, requiring careful management of power delivery and cooling (row vs. in-rack cooling). The successful scaling to over 4,000 GPUs demonstrated a high level of operational maturity.

  2. Identity and Security as the Central Challenge 23:20

    Security is not an afterthought. The architecture emphasizes a 'federation first' approach, using systems like a global UID and GD namespace to maintain a single source of truth for identity across multiple clusters and services. This is critical for auditing and access control, especially when dealing with agentic and autonomous workloads.

  3. Evolving Trust Boundaries 30:00

    The definition of trust has shifted from the traditional shared responsibility model to requiring cryptographic proof. CoreWeave supports this by offering solutions like RKE (Restricted Key Encryption) where data can be encrypted with keys held solely by the customer, ensuring the provider cannot access the data.

  4. Operationalizing Complex Workloads 33:20

    The system must support diverse workloads—from pre-training and fine-tuning to inference and execution—often involving malicious prompt testing. This requires sophisticated sandboxing environments and tight control over ingress/egress rules to prevent lateral movement and unauthorized access.

Technical details

  • Compute Scaling and Hardware 600s

    The deployment utilized advanced hardware like H100 and GB200 clusters. Scaling required addressing physical constraints, including power delivery and cooling methods (row vs. in-rack).

  • Identity and Access Management (IAM) 1400s

    IBM implemented a global UID and GD namespace to unify identity across clusters. The process moved from simple SSH keys to modern, controlled methods like using short-lived keys (replacing 'bring your own key'). The system leverages Red Hat directory services as the source of truth.

  • Workload Orchestration and Control 1600s

    Workloads are managed across multiple systems, including Slurm and Kubernetes. The architecture must coordinate these disparate systems, ensuring that user identity, Slurm accounting, and GPFS are all linked. The introduction of sandboxing is key for running potentially malicious or exploratory agentic workloads.

  • Security and Compliance 1800s

    Security controls are paramount, driven by internal audits and CISO requirements. Key controls include: 1) Implementing granular ingress and egress rules; 2) Utilizing Confidential Computing/Trusted Execution Environments to protect the workload cryptographically; and 3) Establishing telemetry endpoints for auditing purposes (e.g., integrating CrowdStrike feeds).

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.