Weights & Biases

Dedicated Inference

Published 2026-09-28 · Duration 13:41

Summary

This session outlines the AI inference landscape, detailing three production-grade paths for deploying custom models at scale: Serverless Inference, Dedicated Inference, and Inference on CKS. For build engineers, the key takeaway is selecting the right level of control—from rapid, managed iteration (Serverless) to explicit infrastructure control (CKS)—while maintaining a consistent, OpenAI-compatible endpoint for seamless integration into complex agentic frameworks.

Download summary

Key takeaways

  1. The Inference Progression 0:15

    The deployment path moves from general-purpose, proprietary models (low control) to highly customized, controlled deployments that require explicit management of scale and cost.

  2. Serverless Inference 1:10

    Ideal for rapid model iteration and deployment, offering automatic scaling, observability, and native Weights & Biases tracing without requiring infrastructure management.

  3. Dedicated Inference 2:20

    Provides a controlled, production-ready path for custom models, allowing users to bring their own weights, select GPUs, and use open runtimes while CoreWeave handles operations and cost visibility.

  4. Inference on CKS 4:40

    Offers maximum control for ultra-high scale and latency-critical workloads. It is a fully managed service on CoreWeave Kubernetes Service (CKS), supporting runtimes like LLMD and Nvidia Dynamo for self-hosted, distributed inference.

Technical details

  • Multi-Agent Architecture 190s

    Complex agents (e.g., research assistants) require coordinating multiple model types (OSS, custom fine-tuned, multimodal) that must communicate through a common interface: OpenAI-compatible endpoints.

  • Dedicated Inference Mechanics 210s

    Deployment requires creating a **Gateway** (a single tenant isolated endpoint used for authentication, authorization, traffic splitting, and request routing). Multiple **Deployments** sit behind a Gateway, allowing explicit selection of runtime, version, GPU count, and min/max replicas.

  • Inference on CKS Deployment 350s

    Utilizes Kubernetes native control for distributed inference. Deployment involves applying YAML manifests (e.g., for LLMD or Nvidia Dynamo) to run models on dedicated single-tenant GPU nodes, providing infrastructure-level tuning and isolated boundaries.

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.