Google DeepMind

@GoogleDeepMind

Today, we're launching agentic video understanding across our latest models: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. This new capability improves accuracy while dramatically reducing token usage and costs for video analysis. Unlike static processing, where the model ingests the video at a fixed frames-per-second rate, agentic video understanding pairs the model's core reasoning with native video tools to dynamically search, scan, and inspect target video segments across visual frames, audio, and transcripts. Across standard video analysis benchmarks, Gemini models with agentic video understanding reduce analysis costs by up to 66% and token consumption by up to 88%, while improving accuracy by up to 7%. The feature is available today for video uploads and YouTube videos via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. To enable it, set processing to "agentic" in the API configuration. It will roll out to the Gemini app soon, and later power YouTube's Ask YouTube feature.
Open original#294511
  1. Research

    You Hold the Keys, I Run the Alarms: How Anthropic's New Architecture Unties Enterprise AI's Privacy Deadlock
  2. Research

    Epoch: ECI frontier sped up to 14 points/year with reasoning models
  3. Products

    Perplexity introduces hybrid compute on Mac