Google DeepMind
@GoogleDeepMind
Today, we're launching agentic video understanding across our latest models: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. This new capability improves accuracy while dramatically reducing token usage and costs for video analysis.
Unlike static processing, where the model ingests the video at a fixed frames-per-second rate, agentic video understanding pairs the model's core reasoning with native video tools to dynamically search, scan, and inspect target video segments across visual frames, audio, and transcripts. Across standard video analysis benchmarks, Gemini models with agentic video understanding reduce analysis costs by up to 66% and token consumption by up to 88%, while improving accuracy by up to 7%.
The feature is available today for video uploads and YouTube videos via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. To enable it, set processing to "agentic" in the API configuration. It will roll out to the Gemini app soon, and later power YouTube's Ask YouTube feature.