Topic

Kubernetes

All digests tagged Kubernetes

Accelerating AI innovation thumbnail

· 8:40

Accelerating AI innovation

CoreWeave presents its purpose-built, end-to-end AI platform designed to manage complex AI workloads from training to inference. The stack integrates specialized foundational infrastructure (high-density GPU clusters, high-speed interconnects) with advanced software layers, including CoreWeave AI Object Storage (using LOTA) and CoreWeave Kubernetes Service (CKS). Operational control is unified through Mission Control™, providing observability and management across the entire lifecycle, from model development (W&B) to deployment.

Key takeaways

  1. Purpose-Built Infrastructure 1:50

    CoreWeave manages data centers optimized specifically for high-density GPUs and high-speed interconnects, scaling across multiple regions. The bare metal infrastructure is managed via Kubernetes, supplemented by services like Slurm for cluster management.

  2. Low-Latency Storage Solution 4:00

    The CoreWeave AI Object Storage solution utilizes the Local Object Transport Accelerator (LOTA). Each GPU node runs a LOTA proxy, which caches the object storage across the cluster, achieving throughput of up to 7 GB/s per GPU.

  3. Unified Operational Control 6:20

    Mission Control™ unifies security, observability, and talent services into a single pane of glass, providing real-time visibility into job state, cluster health, and performance signals, helping teams diagnose issues proactively.

  4. Accelerated Development Lifecycle 7:10

    The platform integrates W&B Models at the top layer, shortening the path to production by providing integrated experiment tracking, governance, and workflow tooling for model and agent development.

Watch on YouTube Full article

How to Secure & Run AI Agents with NVIDIA OpenShell thumbnail

· 5:30

How to Secure & Run AI Agents with NVIDIA OpenShell

NVIDIA OpenShell 0.1 provides a secure, governed runtime boundary for AI agents, preventing compromised agents from executing unauthorized actions. It allows developers to deploy agents—which are LLMs capable of tool calls and code execution—in a sandboxed environment where all components (tools, subagents, etc.) inherit consistent policy, credential, and audit controls from launch. The system supports live policy updates, delegated subagent controls, and multi-tenant deployments via Kubernetes.

Key takeaways

  1. Runtime Boundary Enforcement

    OpenShell creates a boundary around the agent that the agent cannot escape, similar to OS-level app restrictions. This prevents malicious actions, such as an agent attempting to upload private data to a public repository, even if compromised.

  2. Policy Granularity and Control 2:00

    Policies can be configured at the level of binaries, destinations, methods, and paths. The system supports reviewing and approving network requests in real-time, ensuring the agent only performs intended actions.

  3. Multi-Tenant and Production Deployment 4:00

    For large-scale production environments, OpenShell supports Kubernetes deployment (specifically OpenShift) and utilizes multi-tenant SDKs to manage sandboxes for multiple users and business units.

Watch on YouTube Full article

Operating Distributed Inference Systems at Scale — Nishant Gupta & Naman Ahuja, Meta thumbnail

· 19:51

Operating Distributed Inference Systems at Scale — Nishant Gupta & Naman Ahuja, Meta

Inference workloads are rapidly becoming foundational, hyperscale infrastructure, outpacing even major microservices. The complexity has shifted from optimizing models or kernels to mastering the orchestration layer—the 'control plane.' This requires treating a request as a distributed transaction, necessitating sophisticated scheduling and reliability mechanisms that account for seven axes (e.g., GPU generation, KV cache state, tenant priority). The core optimization metric must shift from 'cost per token' to 'cost per successful task.'

Key takeaways

  1. Inference as a Distributed Transaction 11:45

    Unlike traditional RPC calls, inference involves multiple network hops (gateway, router, scheduler, runtime) where any step can retry, time out, or fail. Reliability must therefore be managed by the control plane, which sees the entire workflow, especially when partial failures occur (e.g., after streaming 200 tokens).

  2. The Shift to Orchestration 3:50

    The value has moved from simple models to the orchestration layer. The system must manage complex interactions, such as a routing decision changing the cache hit rate, which subsequently affects batch composition and GPU utilization. This coupling is the new complexity.

  3. Optimization Metric: Cost per Successful Task 14:05

    The goal of optimization is not merely minimizing cost per token or per request. The critical metric is 'cost per successful task,' as this reflects the actual value delivered to the user and accounts for retries, failures, and operational overhead.

  4. The Need for a Dedicated Control Plane 17:00

    As inference scales, it requires its own dedicated control plane, analogous to how Kubernetes managed VMs. This plane must manage resources beyond CPU/Memory, including GPU, KV cache, and token limits, to make optimal scheduling and routing decisions.

Watch on YouTube Full article

Tethered: Our Agents Are Us — Shu Fang, Two Sigma thumbnail

· 21:10

Tethered: Our Agents Are Us — Shu Fang, Two Sigma

Two Sigma implemented a framework allowing every employee to run cloud agents using their own unique user identity, addressing the challenges of permissions drift and maintaining security in a highly regulated environment. The solution leverages existing Kubernetes infrastructure (dedicated namespaces per person) and introduces two critical guardrails: propagating a trace header for full action provenance, and utilizing Google's web grounding for enterprise—a restricted search index that eliminates external egress vulnerabilities while accepting a data freshness constraint of up to 24 hours.

Key takeaways

  1. Running Agents as User Identity 2:00

    By running agents with the user's exact identity, the system bypasses conventional constraints like permissions drift and licensing issues associated with separate machine identities. This capability was supported by pre-existing infrastructure: a Kubernetes namespace per individual in every region, where automated jobs already ran using the user's identity via a sidecar mounting mechanism.

  2. Ensuring Action Provenance (Attribution) 8:37

    To differentiate between actions taken by the human and those performed by the agent, a dedicated header is propagated throughout the system. This trace ID allows for full provenance tracking, enabling the replay of the entire chain of actions leading to an end result, which is superior to simple identity verification.

  3. Securing Web Access with Grounding 9:18

    To mitigate risks like exfiltration and prompt injection from open web access, the firm adopted Google's 'web grounding for enterprise.' This service provides search and fetch capabilities within the internal VPC network boundary, while blocking native tools (e.g., Brave web browser) to ensure all requests route through the controlled index.

Watch on YouTube Full article

KV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red Hat thumbnail

· 21:48

KV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red Hat

This talk details advanced strategies for optimizing LLM inference in complex agentic workloads, moving beyond the limitations of steady-state public benchmarks. The focus is on two critical levers: KV Cache-Aware Routing and Prefill/Decode (P/D) Disaggregation. Implementing these techniques—using frameworks like LLMD—significantly improves latency and throughput by managing volatile cache usage and separating compute-bound prefill from memory-bandwidth-hungry decode phases, particularly in the middle concurrency band.

Key takeaways

  1. Agentic Workloads vs. Benchmarks 5:29

    Real-world agentic workloads exhibit chaotic multi-turn interactions (up to 3,000 turns) with high cache hit rates (>90%) and massive input/output ratios (often >100:1), which standard public benchmarks fail to capture [0:00], [3:29].

  2. KV Cache Routing Optimization 10:28

    Implementing KV cache-aware routing (via Endpoint Picker) is a cost-effective optimization, as the token cost difference between cached and uncached tokens can be as high as 10x [5:12]. This helps solve Time to First Token (TTFT) issues.

  3. P/D Disaggregation Benefits 15:28

    Separating prefill and decode into independent, scalable pods prevents 'phase interference'—where a long prefill stalls token generation (decode)—leading to drastically reduced P99 Inter Token Latency (ITL) from ~900ms down to ~100ms [9:28], [10:46].

  4. Prerequisites for PD 17:26

    Effective P/D disaggregation requires an advanced, high-speed network fabric like RDMA or RoCE to facilitate the transfer of KV caches between prefill and decode workers [10:46]. If such a fabric is unavailable, aggregated serving may be preferable.

Watch on YouTube Full article

Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai thumbnail

· 16:55

Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai

Gabriel Jorge Menezes details the complex infrastructure required to train and serve Krea 2, a diffusion transformer model trained from scratch on thousands of GPUs. The system addresses challenges like silent failures at scale, GPU thermal throttling, and cross-node communication issues by implementing advanced monitoring (tensor core utilization, InfiniBand metrics). For serving, they built a robust architecture using Gang scheduling and Kubernetes features (virtual kubelet, taints/tolerations) to ensure training workloads can utilize the entire cluster while maintaining production uptime through seamless traffic flipping.

Key takeaways

  1. Metrics are essential for large-scale pre-training 9:50

    Do not rely on GPU utilization (which is 'a lie'). Instead, monitor tensor core utilization and collect custom metrics like InfiniBand/NVLink errors, as most failures relate to cross-node communication. [5:58], [6:48]

  2. Embrace failure for stability 4:18

    When scaling training runs, instead of debugging every crash, it is often more efficient to 'let it crash.' The system should be designed to recover and run successfully on the same nodes over extended periods. [4:18]

  3. Checkpointing must be extremely fast 8:29

    To make long training runs survivable, checkpoint aggressively against a high-speed filesystem capable of writing terabytes quickly (e.g., achieving >1 TB/30 seconds). [8:29]

  4. Decouple training and production workloads 11:01

    Use a system that allows high-priority training jobs to utilize the entire cluster while seamlessly migrating inference traffic (production) to external providers or other clusters, ensuring zero downtime. [11:01]

Watch on YouTube Full article