# How H-Company Optimizes VLM Serving for Computer Use Agents With NVIDIA Dynamo

## Executive summary

This session details how H-Company, a startup building computer-use agents, optimizes the serving of Vision Language Models (VLMs) using NVIDIA Dynamo. Computer-use agents, which can read screens and interact with desktop/web/mobile applications, present unique serving challenges compared to traditional chatbots due to the inclusion of screenshots and large context windows. The solution leverages Dynamo components—including the embedding cache, EP endpoint picker, and Model Express—to manage complex workloads, improve throughput, and enable faster autoscaling.

## Key takeaways

- Computer-Use Agents (Holo/HoloTron): These agents go beyond text generation; they can read a screen and perform actions (clicking, typing) across real desktop, web, and mobile applications. H-Company's models, Holo and Holotron, are built by post-training open-source models (like Quen and Neotron) for these specific tasks.
- The Challenge of VLM Serving: Unlike standard chat models, computer-use agents require passing screenshots, which significantly increases the context size (potentially 1-4k tokens per image). Furthermore, the history of screenshots (sliding window) invalidates the cache at each turn, requiring re-computation of the VLM.
- Dynamo's Role in Optimization: NVIDIA Dynamo is used to address these scaling challenges by providing components for efficient request routing, caching, and model deployment, enabling production-scale service of these complex agentic workloads.
- Model Express for Fast Bootup: Model Express is a Dynamo component that significantly speeds up VLM replica bootup (7x to 12x speedup). It uses NVIDIA GPU-to-GPU LDMA capability to pull weights directly from other replicas on the cluster, improving autoscaling speed and resource utilization.

## Technical details

- Dynamo Embedding Cache: To prevent redundant encoding, the embedding cache (residing within the VLM replica) maps a unique UUID attached to a screenshot to its already computed image encoding/embedding. This allows subsequent requests to only send the UUID, saving payload and computation time, especially with larger sliding windows (e.g., 10 screenshots).
- Dynamo EP Endpoint Picker: The EP endpoint picker optimizes request routing by allowing the system to maintain state (e.g., sticky sessions) for a single client session across multiple requests, which is crucial for optimizing cache usage (like the embedding cache and prefix cache) within a single client trajectory.
- Benchmarking with AIPF: H-Company uses a custom dataset of realistic agentic traces for computer-use workloads. They benchmark infrastructure improvements using a CLI that allows specifying the dataset, the shape of the data (e.g., k screenshots in context), and the load shape (e.g., concurrency level).
- GPU Resource Allocation Strategy: When optimizing for throughput (e.g., RL/synthetic data generation), the goal is to maximize performance at the high-throughput point on the interactivity vs. throughput curve. When optimizing for latency (e.g., production demo), the goal is to maximize throughput under a target latency constraint.

## Practical implications

- Build engineers can significantly improve the scalability and resource efficiency of VLM serving by implementing specialized caching mechanisms (like UUID-based embedding caches) to handle large, multi-modal context windows.
- Utilizing Model Express drastically reduces the time required for autoscaling and replica bootup, allowing systems to absorb sudden load spikes without over-provisioning GPU capacity.
- The Dynamo EP endpoint picker provides granular control over request routing, enabling the optimization of latency and throughput by maintaining session state (sticky sessions) across multiple service calls.

## Topics

Vision Language Models (VLMs), Agentic Workloads, Inference Optimization, Distributed Systems, GPU Computing, Open Source AI, Holo, HoloTron, Quen family of models, NVIDIA Neotron, Hollow 4 (27B dense, 35B OBoE), NVIDIA Dynamo, AIPF

Source: https://www.youtube.com/watch?v=eXlgMVAm28I
