Google Launches Agentic Video Understanding in Gemini, Cutting Tokens by 88%
Key point
Google introduced dynamic video analysis to Gemini models, reducing token consumption by up to 88% and increasing accuracy by 7%.
Details
Google DeepMind has launched agentic video understanding for the Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite models. Unlike previous methods that processed video at fixed frame rates, this feature enables active reasoning where the model dynamically searches and scans for target segments.
Efficiency and Performance Improvements
Based on standard benchmarks, applying agentic video understanding results in up to 88% reduction in token consumption, up to 66% cost savings in analysis, and up to 7% improvement in accuracy. It overcomes the limitations of static processing methods, providing efficient analysis for videos ranging from 10-minute guides to multi-hour long-form content.
Key Features and How It Works
Using Gemini's native video tools, it selectively retrieves only the necessary modalities such as frames, audio, and subtitles. This enables precise tasks such as:
- Sub-second moment retrieval: Accurately captures fleeting state changes or edit points that are easily missed at 1 FPS.
- Anomaly detection: Detects subtle anomalies by dynamically resampling suspicious time windows.
- Needle-in-a-haystack search: Answers complex questions within long videos without consuming millions of tokens.
Availability and Expansion
This feature can be enabled via the APIs in Google AI Studio and the Gemini Enterprise Agent Platform, available at existing token rates with no additional cost. It is scheduled to be sequentially applied to Flash models in the Gemini app and YouTube's visual question-answering feature in the future.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.