Topic

Nebius

All digests tagged Nebius

What Makes Open Models Fast in Production — Sujee Maniyam, Nebius thumbnail

· 20:32

What Makes Open Models Fast in Production — Sujee Maniyam, Nebius

The talk details the engineering challenges of deploying open Large Language Models (LLMs) in production. Nebius Token Factory positions itself as a full-stack AI cloud infrastructure that solves the complexity of the entire LLM lifecycle—from inference and data capture to post-training and deployment. Key technical optimizations discussed include cache-aware routing, speculative decoding, KV cache offloading, and separating prefill/decode stages to maximize performance and minimize cost, allowing companies to achieve self-hosting control with the simplicity of a managed service.

Key takeaways

  1. The LLM Deployment Loop 7:04

    A successful AI product requires managing the full loop: Inference $ ightarrow$ Data Collection $ ightarrow$ Post-training (fine-tuning/distillation) $ ightarrow$ Deployment. Token Factory manages this entire cycle, enabling continuous improvement and iteration.

  2. Optimizing for Scale and Cost 19:06

    Performance is optimized through hardware-level techniques like utilizing NVFP4 on the latest NVIDIA chips, and software-level optimizations such as cache-aware routing and disaggregating prefill and decode stages.

  3. The Value of Managed Inference 1:40

    Managed inference services offer the control and performance of self-hosting without the massive engineering overhead, allowing teams to focus on product development rather than infrastructure maintenance.

Watch on YouTube Full article