Dedicated Inference
Summary
This session outlines the AI inference landscape, detailing three production-grade paths for deploying custom models at scale: Serverless Inference, Dedicated Inference, and Inference on CKS. For build engineers, the key takeaway is selecting the right level of control—from rapid, managed iteration (Serverless) to explicit infrastructure control (CKS)—while maintaining a consistent, OpenAI-compatible endpoint for seamless integration into complex agentic frameworks.
Key takeaways
-
The Inference Progression
0:15
The deployment path moves from general-purpose, proprietary models (low control) to highly customized, controlled deployments that require explicit management of scale and cost.
-
Serverless Inference
1:10
Ideal for rapid model iteration and deployment, offering automatic scaling, observability, and native Weights & Biases tracing without requiring infrastructure management.
-
Dedicated Inference
2:20
Provides a controlled, production-ready path for custom models, allowing users to bring their own weights, select GPUs, and use open runtimes while CoreWeave handles operations and cost visibility.
-
Inference on CKS
4:40
Offers maximum control for ultra-high scale and latency-critical workloads. It is a fully managed service on CoreWeave Kubernetes Service (CKS), supporting runtimes like LLMD and Nvidia Dynamo for self-hosted, distributed inference.
Technical details
-
Multi-Agent Architecture
190s
Complex agents (e.g., research assistants) require coordinating multiple model types (OSS, custom fine-tuned, multimodal) that must communicate through a common interface: OpenAI-compatible endpoints.
-
Dedicated Inference Mechanics
210s
Deployment requires creating a **Gateway** (a single tenant isolated endpoint used for authentication, authorization, traffic splitting, and request routing). Multiple **Deployments** sit behind a Gateway, allowing explicit selection of runtime, version, GPU count, and min/max replicas.
-
Inference on CKS Deployment
350s
Utilizes Kubernetes native control for distributed inference. Deployment involves applying YAML manifests (e.g., for LLMD or Nvidia Dynamo) to run models on dedicated single-tenant GPU nodes, providing infrastructure-level tuning and isolated boundaries.
Mentioned resources
- CoreWeave Inference
- Weights & Biases
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.