# Build Visual AI Agents From a Prompt With NVIDIA Cosmos and VSS 3.3

## Executive summary

This session details how to build scalable visual AI agents from a natural language prompt using NVIDIA Cosmos and the updated VSS 3.3 blueprint. VSS enables the combination of video search, live monitoring, and summarization into a single, deployable application. Key innovations include Adaptive Efficient Video Sampling (EVS) to reduce inference overhead and streaming NIM support for low-latency alerting. The architecture allows developers to build, customize, and deploy agents by integrating specialized components (e.g., fill level measurement) into a unified workflow.

## Key takeaways

- Building Agents from Prompts: The process allows developers to define a complex visual AI agent's requirements using a natural language prompt, which the coding agent (e.g., Astra in Codeex) uses to propose and deploy the necessary architecture (search, alerts, or summarization).
- Adaptive Efficient Video Sampling (EVS): EVS improves video processing efficiency by identifying and only passing tokens corresponding to temporal or spatial changes (deltas) to the Large Language Model (LLM), aiming to reduce latency by up to 60% and increase concurrent stream capacity.
- Streaming and Low-Latency Alerts: The introduction of streaming NIM capability allows for rolling window processing, enabling alerts to be generated in under 300 milliseconds, moving beyond traditional video chunking methods.
- Extensibility and Tool Integration: VSS supports building specialized capabilities (like measuring visible liquid fill levels) by integrating custom components, APIs, and tools directly into the agent's workflow, allowing the LLM to query and explain specific data points.
- Scalability and Deployment: The platform provides reference architectures for both stored video and high-throughput streams, supporting deployment on various hardware, including edge devices like Jetson Orin and DGX systems.

## Technical details

- VSS Architecture: VSS provides a reference architecture that connects various models (VLMs, LLMs, embedding models, CV transformers) to handle large volumes of video data, supporting agentic deep search across embedding spaces, natural language, and dense captions.
- Video Search Workflow: Search involves three steps: 1) Retrieval (using the embedding model to find candidate clips), 2) Evaluation (Cosmos/VLM evaluates candidates against the requested visual condition), and 3) Verification (providing a verdict like 'confirmed' or 'rejected' with video evidence).
- System Components: The system utilizes a multi-stage process: VideoIO and storage (VSSs), Embedding Model (for searchable representations), ElasticSearch (for the search index), and the VLM/LLM (for evaluation and conversation).
- Hardware Scaling: For alert verification, a DGX Spark can process 5-6 steady streams (30 fps, 1080p) through a vision transformer pipeline, while pure VLM streaming is limited to 2-3 streams due to compute intensity.

## Practical implications

- The platform significantly reduces development effort by allowing complex agent pipelines to be configured and deployed using natural language prompts.
- Developers can build highly specialized, multi-modal applications (e.g., combining video search with physical measurements) without manual integration.
- The architecture supports flexible deployment, allowing sensitive, latency-critical components (VLM/Vision Transformer) to run on the edge (e.g., Jetson Orin) while the LLM/reporting runs in the cloud.
- The modular nature of VSS allows for the integration of proprietary or specialized business logic (e.g., custom fill level calculation) as tools accessible to the AI agent.

## Topics

Visual AI Agents, Video Analytics, NVIDIA Cosmos, VSS Blueprint, Edge AI, LLM Integration, Video Processing, VSS 3.3 Tech Blog, VSS Skills, VSS Build, Brev Launchable

Source: https://www.youtube.com/watch?v=PQJKs1dyK7I
