Build Visual AI Agents From a Prompt With NVIDIA Cosmos and VSS 3.3
This session details how to build scalable visual AI agents from a natural language prompt using NVIDIA Cosmos and the updated VSS 3.3 blueprint. VSS enables the combination of video search, live monitoring, and summarization into a single, deployable application. Key innovations include Adaptive Efficient Video Sampling (EVS) to reduce inference overhead and streaming NIM support for low-latency alerting. The architecture allows developers to build, customize, and deploy agents by integrating specialized components (e.g., fill level measurement) into a unified workflow.
Key takeaways
-
Building Agents from Prompts
2:30
The process allows developers to define a complex visual AI agent's requirements using a natural language prompt, which the coding agent (e.g., Astra in Codeex) uses to propose and deploy the necessary architecture (search, alerts, or summarization).
-
Adaptive Efficient Video Sampling (EVS)
4:00
EVS improves video processing efficiency by identifying and only passing tokens corresponding to temporal or spatial changes (deltas) to the Large Language Model (LLM), aiming to reduce latency by up to 60% and increase concurrent stream capacity.
-
Streaming and Low-Latency Alerts
4:40
The introduction of streaming NIM capability allows for rolling window processing, enabling alerts to be generated in under 300 milliseconds, moving beyond traditional video chunking methods.
-
Extensibility and Tool Integration
7:30
VSS supports building specialized capabilities (like measuring visible liquid fill levels) by integrating custom components, APIs, and tools directly into the agent's workflow, allowing the LLM to query and explain specific data points.
-
Scalability and Deployment
3:00
The platform provides reference architectures for both stored video and high-throughput streams, supporting deployment on various hardware, including edge devices like Jetson Orin and DGX systems.