Topic

NVIDIA Neotron

All digests tagged NVIDIA Neotron

How H-Company Optimizes VLM Serving for Computer Use Agents With NVIDIA Dynamo thumbnail

· 37:32

How H-Company Optimizes VLM Serving for Computer Use Agents With NVIDIA Dynamo

This session details how H-Company, a startup building computer-use agents, optimizes the serving of Vision Language Models (VLMs) using NVIDIA Dynamo. Computer-use agents, which can read screens and interact with desktop/web/mobile applications, present unique serving challenges compared to traditional chatbots due to the inclusion of screenshots and large context windows. The solution leverages Dynamo components—including the embedding cache, EP endpoint picker, and Model Express—to manage complex workloads, improve throughput, and enable faster autoscaling.

Key takeaways

  1. Computer-Use Agents (Holo/HoloTron) 2:30

    These agents go beyond text generation; they can read a screen and perform actions (clicking, typing) across real desktop, web, and mobile applications. H-Company's models, Holo and Holotron, are built by post-training open-source models (like Quen and Neotron) for these specific tasks.

  2. The Challenge of VLM Serving 5:30

    Unlike standard chat models, computer-use agents require passing screenshots, which significantly increases the context size (potentially 1-4k tokens per image). Furthermore, the history of screenshots (sliding window) invalidates the cache at each turn, requiring re-computation of the VLM.

  3. Dynamo's Role in Optimization 6:40

    NVIDIA Dynamo is used to address these scaling challenges by providing components for efficient request routing, caching, and model deployment, enabling production-scale service of these complex agentic workloads.

  4. Model Express for Fast Bootup 21:40

    Model Express is a Dynamo component that significantly speeds up VLM replica bootup (7x to 12x speedup). It uses NVIDIA GPU-to-GPU LDMA capability to pull weights directly from other replicas on the cluster, improving autoscaling speed and resource utilization.

Watch on YouTube Full article