How H-Company Optimizes VLM Serving for Computer Use Agents With NVIDIA Dynamo
Summary
This session details how H-Company, a startup building computer-use agents, optimizes the serving of Vision Language Models (VLMs) using NVIDIA Dynamo. Computer-use agents, which can read screens and interact with desktop/web/mobile applications, present unique serving challenges compared to traditional chatbots due to the inclusion of screenshots and large context windows. The solution leverages Dynamo components—including the embedding cache, EP endpoint picker, and Model Express—to manage complex workloads, improve throughput, and enable faster autoscaling.
Key takeaways
-
Computer-Use Agents (Holo/HoloTron)
2:30
These agents go beyond text generation; they can read a screen and perform actions (clicking, typing) across real desktop, web, and mobile applications. H-Company's models, Holo and Holotron, are built by post-training open-source models (like Quen and Neotron) for these specific tasks.
-
The Challenge of VLM Serving
5:30
Unlike standard chat models, computer-use agents require passing screenshots, which significantly increases the context size (potentially 1-4k tokens per image). Furthermore, the history of screenshots (sliding window) invalidates the cache at each turn, requiring re-computation of the VLM.
-
Dynamo's Role in Optimization
6:40
NVIDIA Dynamo is used to address these scaling challenges by providing components for efficient request routing, caching, and model deployment, enabling production-scale service of these complex agentic workloads.
-
Model Express for Fast Bootup
21:40
Model Express is a Dynamo component that significantly speeds up VLM replica bootup (7x to 12x speedup). It uses NVIDIA GPU-to-GPU LDMA capability to pull weights directly from other replicas on the cluster, improving autoscaling speed and resource utilization.
Technical details
-
Dynamo Embedding Cache
750s
To prevent redundant encoding, the embedding cache (residing within the VLM replica) maps a unique UUID attached to a screenshot to its already computed image encoding/embedding. This allows subsequent requests to only send the UUID, saving payload and computation time, especially with larger sliding windows (e.g., 10 screenshots).
-
Dynamo EP Endpoint Picker
1000s
The EP endpoint picker optimizes request routing by allowing the system to maintain state (e.g., sticky sessions) for a single client session across multiple requests, which is crucial for optimizing cache usage (like the embedding cache and prefix cache) within a single client trajectory.
-
Benchmarking with AIPF
550s
H-Company uses a custom dataset of realistic agentic traces for computer-use workloads. They benchmark infrastructure improvements using a CLI that allows specifying the dataset, the shape of the data (e.g., k screenshots in context), and the load shape (e.g., concurrency level).
-
GPU Resource Allocation Strategy
1650s
When optimizing for throughput (e.g., RL/synthetic data generation), the goal is to maximize performance at the high-throughput point on the interactivity vs. throughput curve. When optimizing for latency (e.g., production demo), the goal is to maximize throughput under a target latency constraint.
Mentioned resources
- Holo
- HoloTron
- Quen family of models
- NVIDIA Neotron
- Hollow 4 (27B dense, 35B OBoE)
- NVIDIA Dynamo
- AIPF
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.