NVIDIA Unveils Technology to Cut Video AI Development and Operational Costs with VSS 3.3
Key point
Adaptive EVS reduces VLM input tokens by 80% and improves concurrent stream throughput by 46%
Details
NVIDIA released version 3.3 of VSS (NVIDIA Metropolis Blueprint for Video Search and Summarization), a reference architecture for video search and summarization, introducing two key features that significantly reduce development efficiency and operational costs.
Build Vision Agent Skill
This feature enables developers to automatically generate validated profile-based Docker Compose deployments through coding agents (such as Claude Code and Codex). It adds or integrates only the necessary services based on requests to eliminate redundancy, and establishes a deployment plan after reviewing the architecture diagram. According to NVIDIA's measurements, recording-based builds were completed in under 30 minutes in an environment with two RTX PRO 6000 Blackwell GPUs.
Adaptive EVS (Adaptive Efficient Video Sampling)
This technology reduces VLM processing costs by dynamically removing tokens from areas of video frames that show no change. It uses patch-wise cosine similarity to compare against the previous frame, and through event-aware batching, it groups or discards clips based on the remaining token ratio.
Key Performance Metrics (Based on RTX PRO 6000 Blackwell, Cosmos 3 Super Reasoner FP8):
- VLM Input Token Reduction: When summarizing a 60-minute video, input tokens are reduced by 80%, and processing time is cut by approximately half.
- Increased Concurrent Throughput: The number of real-time VLM streams increased by 46%, from 13 to 19.
- Latency Improvement: Alert contextualization latency decreased by 17%, from 1,021ms to 844ms.
Application and Considerations:
- It is effective for tasks with many input frames and short outputs, such as detailed captioning, long-video summarization, and alert verification.
- Currently, only local vLLM-compatible models based on the Qwen3-VL architecture are supported; remote endpoints or Omni models are not supported.
- The feature is disabled by default, and restarting RT-VLM is required when changing settings.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.