# Dedicated Inference

## Executive summary

This session outlines the AI inference landscape, detailing three production-grade paths for deploying custom models at scale: Serverless Inference, Dedicated Inference, and Inference on CKS. For build engineers, the key takeaway is selecting the right level of control—from rapid, managed iteration (Serverless) to explicit infrastructure control (CKS)—while maintaining a consistent, OpenAI-compatible endpoint for seamless integration into complex agentic frameworks.

## Key takeaways

- The Inference Progression: The deployment path moves from general-purpose, proprietary models (low control) to highly customized, controlled deployments that require explicit management of scale and cost.
- Serverless Inference: Ideal for rapid model iteration and deployment, offering automatic scaling, observability, and native Weights & Biases tracing without requiring infrastructure management.
- Dedicated Inference: Provides a controlled, production-ready path for custom models, allowing users to bring their own weights, select GPUs, and use open runtimes while CoreWeave handles operations and cost visibility.
- Inference on CKS: Offers maximum control for ultra-high scale and latency-critical workloads. It is a fully managed service on CoreWeave Kubernetes Service (CKS), supporting runtimes like LLMD and Nvidia Dynamo for self-hosted, distributed inference.

## Technical details

- Multi-Agent Architecture: Complex agents (e.g., research assistants) require coordinating multiple model types (OSS, custom fine-tuned, multimodal) that must communicate through a common interface: OpenAI-compatible endpoints.
- Dedicated Inference Mechanics: Deployment requires creating a **Gateway** (a single tenant isolated endpoint used for authentication, authorization, traffic splitting, and request routing). Multiple **Deployments** sit behind a Gateway, allowing explicit selection of runtime, version, GPU count, and min/max replicas.
- Inference on CKS Deployment: Utilizes Kubernetes native control for distributed inference. Deployment involves applying YAML manifests (e.g., for LLMD or Nvidia Dynamo) to run models on dedicated single-tenant GPU nodes, providing infrastructure-level tuning and isolated boundaries.

## Practical implications

- Build engineers must select the inference path (Serverless, Dedicated, or CKS) based on the required level of governance, control, and scale (e.g., using Dedicated for predictable SLOs, or CKS for strict regulatory tenancy).
- The use of a common OpenAI-compatible endpoint simplifies the integration of diverse model types (OSS, custom, multimodal) into agentic frameworks.
- Understanding the Gateway/Deployment structure is crucial for building robust, multi-tenant inference services that require advanced traffic routing and authorization.

## Topics

AI Inference, Model Deployment, Scalability, Kubernetes (CKS), OpenAI API Compatibility, CoreWeave Inference, Weights & Biases

Source: https://www.youtube.com/watch?v=v3LOzZHPcFc
