# What Makes Open Models Fast in Production — Sujee Maniyam, Nebius

## Executive summary

The talk details the engineering challenges of deploying open Large Language Models (LLMs) in production. Nebius Token Factory positions itself as a full-stack AI cloud infrastructure that solves the complexity of the entire LLM lifecycle—from inference and data capture to post-training and deployment. Key technical optimizations discussed include cache-aware routing, speculative decoding, KV cache offloading, and separating prefill/decode stages to maximize performance and minimize cost, allowing companies to achieve self-hosting control with the simplicity of a managed service.

## Key takeaways

- The LLM Deployment Loop: A successful AI product requires managing the full loop: Inference $ ightarrow$ Data Collection $ ightarrow$ Post-training (fine-tuning/distillation) $ ightarrow$ Deployment. Token Factory manages this entire cycle, enabling continuous improvement and iteration.
- Optimizing for Scale and Cost: Performance is optimized through hardware-level techniques like utilizing NVFP4 on the latest NVIDIA chips, and software-level optimizations such as cache-aware routing and disaggregating prefill and decode stages.
- The Value of Managed Inference: Managed inference services offer the control and performance of self-hosting without the massive engineering overhead, allowing teams to focus on product development rather than infrastructure maintenance.

## Technical details

- Hardware and Infrastructure: Nebius operates its own data centers across the EU and US, utilizing bare metal capacity and specialized hardware like Blackwell Ultra AGXP300 and GB300. They are early adopters of Nvidia Rubin CP and Bluefield storage.
- Cache-Aware Routing: Unlike naive load balancing, cache-aware routing ensures that incoming requests are directed to GPUs that already possess relevant cached tokens, significantly improving speed and efficiency in LLM inference.
- Speculative Decoding: This technique uses a smaller, faster model (the 'draft model') to quickly generate candidate tokens, which are then verified by the larger, more accurate model. This dramatically improves generation speed.
- KV Cache Management: The Key-Value (KV) cache stores previously generated tokens to avoid redundant computation. To manage memory constraints, the system automatically implements KV cache offloading, saving the cache to regular memory and retrieving it when needed.
- Prefill/Decode Disaggregation: The two stages of LLM generation (context filling/prefill and token generation/decode) are separated onto different GPU sets. This allows the system to optimize for compute-intensive prefill and memory-intensive decode independently.
- Quantization Sweet Spot: Models are quantized (reducing precision) to improve efficiency, but the platform runs experiments to find the optimal 'sweet spot' that maximizes speed while minimizing performance degradation.

## Practical implications

- Build teams can transition from complex, manual infrastructure management to a managed, full-stack platform, accelerating time-to-market for AI products.
- By optimizing the entire LLM lifecycle (inference, data, training, deployment), organizations can achieve significant cost savings and performance gains compared to stitching together disparate tools.
- The ability to train custom draft models using proprietary production data allows for deep model customization and reduced vendor lock-in.

## Topics

LLMs, AI Infrastructure, Model Optimization, Inference Engineering, Cloud Computing, Speculative Decoding, Quantization, Nebius Token Factory, Nebius

Source: https://www.youtube.com/watch?v=TRe1u7dHYiA
