Build Visual AI Agents From a Prompt With NVIDIA Cosmos and VSS 3.3
Summary
This session details how to build scalable visual AI agents from a natural language prompt using NVIDIA Cosmos and the updated VSS 3.3 blueprint. VSS enables the combination of video search, live monitoring, and summarization into a single, deployable application. Key innovations include Adaptive Efficient Video Sampling (EVS) to reduce inference overhead and streaming NIM support for low-latency alerting. The architecture allows developers to build, customize, and deploy agents by integrating specialized components (e.g., fill level measurement) into a unified workflow.
Key takeaways
-
Building Agents from Prompts
2:30
The process allows developers to define a complex visual AI agent's requirements using a natural language prompt, which the coding agent (e.g., Astra in Codeex) uses to propose and deploy the necessary architecture (search, alerts, or summarization).
-
Adaptive Efficient Video Sampling (EVS)
4:00
EVS improves video processing efficiency by identifying and only passing tokens corresponding to temporal or spatial changes (deltas) to the Large Language Model (LLM), aiming to reduce latency by up to 60% and increase concurrent stream capacity.
-
Streaming and Low-Latency Alerts
4:40
The introduction of streaming NIM capability allows for rolling window processing, enabling alerts to be generated in under 300 milliseconds, moving beyond traditional video chunking methods.
-
Extensibility and Tool Integration
7:30
VSS supports building specialized capabilities (like measuring visible liquid fill levels) by integrating custom components, APIs, and tools directly into the agent's workflow, allowing the LLM to query and explain specific data points.
-
Scalability and Deployment
3:00
The platform provides reference architectures for both stored video and high-throughput streams, supporting deployment on various hardware, including edge devices like Jetson Orin and DGX systems.
Technical details
-
VSS Architecture
190s
VSS provides a reference architecture that connects various models (VLMs, LLMs, embedding models, CV transformers) to handle large volumes of video data, supporting agentic deep search across embedding spaces, natural language, and dense captions.
-
Video Search Workflow
390s
Search involves three steps: 1) Retrieval (using the embedding model to find candidate clips), 2) Evaluation (Cosmos/VLM evaluates candidates against the requested visual condition), and 3) Verification (providing a verdict like 'confirmed' or 'rejected' with video evidence).
-
System Components
410s
The system utilizes a multi-stage process: VideoIO and storage (VSSs), Embedding Model (for searchable representations), ElasticSearch (for the search index), and the VLM/LLM (for evaluation and conversation).
-
Hardware Scaling
580s
For alert verification, a DGX Spark can process 5-6 steady streams (30 fps, 1080p) through a vision transformer pipeline, while pure VLM streaming is limited to 2-3 streams due to compute intensity.
Mentioned resources
Channel & topics
Watch on YouTube · Back to latest
This independent, AI-assisted summary is provided for commentary and informational purposes. It may contain errors or omit important context. Please watch the original video for the creator's complete presentation. Video, thumbnail, and related copyrights belong to their respective owners.