AI Engineer

What Makes Open Models Fast in Production — Sujee Maniyam, Nebius

Published 2026-10-03 · Duration 20:32

Summary

The talk details the engineering challenges of deploying open Large Language Models (LLMs) in production. Nebius Token Factory positions itself as a full-stack AI cloud infrastructure that solves the complexity of the entire LLM lifecycle—from inference and data capture to post-training and deployment. Key technical optimizations discussed include cache-aware routing, speculative decoding, KV cache offloading, and separating prefill/decode stages to maximize performance and minimize cost, allowing companies to achieve self-hosting control with the simplicity of a managed service.

Download summary

Key takeaways

  1. The LLM Deployment Loop 7:04

    A successful AI product requires managing the full loop: Inference $ ightarrow$ Data Collection $ ightarrow$ Post-training (fine-tuning/distillation) $ ightarrow$ Deployment. Token Factory manages this entire cycle, enabling continuous improvement and iteration.

  2. Optimizing for Scale and Cost 19:06

    Performance is optimized through hardware-level techniques like utilizing NVFP4 on the latest NVIDIA chips, and software-level optimizations such as cache-aware routing and disaggregating prefill and decode stages.

  3. The Value of Managed Inference 1:40

    Managed inference services offer the control and performance of self-hosting without the massive engineering overhead, allowing teams to focus on product development rather than infrastructure maintenance.

Technical details

  • Hardware and Infrastructure 140s

    Nebius operates its own data centers across the EU and US, utilizing bare metal capacity and specialized hardware like Blackwell Ultra AGXP300 and GB300. They are early adopters of Nvidia Rubin CP and Bluefield storage.

  • Cache-Aware Routing

    Unlike naive load balancing, cache-aware routing ensures that incoming requests are directed to GPUs that already possess relevant cached tokens, significantly improving speed and efficiency in LLM inference.

  • Speculative Decoding

    This technique uses a smaller, faster model (the 'draft model') to quickly generate candidate tokens, which are then verified by the larger, more accurate model. This dramatically improves generation speed.

  • KV Cache Management

    The Key-Value (KV) cache stores previously generated tokens to avoid redundant computation. To manage memory constraints, the system automatically implements KV cache offloading, saving the cache to regular memory and retrieving it when needed.

  • Prefill/Decode Disaggregation

    The two stages of LLM generation (context filling/prefill and token generation/decode) are separated onto different GPU sets. This allows the system to optimize for compute-intensive prefill and memory-intensive decode independently.

  • Quantization Sweet Spot

    Models are quantized (reducing precision) to improve efficiency, but the platform runs experiments to find the optimal 'sweet spot' that maximizes speed while minimizing performance degradation.

Mentioned resources

Channel & topics

Watch on YouTube · Back to latest

This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.