What Makes Open Models Fast in Production — Sujee Maniyam, Nebius
The talk details the engineering challenges of deploying open Large Language Models (LLMs) in production. Nebius Token Factory positions itself as a full-stack AI cloud infrastructure that solves the complexity of the entire LLM lifecycle—from inference and data capture to post-training and deployment. Key technical optimizations discussed include cache-aware routing, speculative decoding, KV cache offloading, and separating prefill/decode stages to maximize performance and minimize cost, allowing companies to achieve self-hosting control with the simplicity of a managed service.
Key takeaways
-
The LLM Deployment Loop
7:04
A successful AI product requires managing the full loop: Inference $ ightarrow$ Data Collection $ ightarrow$ Post-training (fine-tuning/distillation) $ ightarrow$ Deployment. Token Factory manages this entire cycle, enabling continuous improvement and iteration.
-
Optimizing for Scale and Cost
19:06
Performance is optimized through hardware-level techniques like utilizing NVFP4 on the latest NVIDIA chips, and software-level optimizations such as cache-aware routing and disaggregating prefill and decode stages.
-
The Value of Managed Inference
1:40
Managed inference services offer the control and performance of self-hosting without the massive engineering overhead, allowing teams to focus on product development rather than infrastructure maintenance.