时间:2026年9月1日
地点:美国加利福尼亚州山景城(Google DeepMind)
人物:Rohan Doshi(DeepMind 高级产品经理)、Mario Lučić(DeepMind 研究总监)
事件详情:Google DeepMind 正式上线"agentic video understanding"(智能体视频理解)功能,覆盖 Gemini 3.7 Flash、Gemini 3.6 Flash 与 Gemini 3.5 Flash-Lite 三款最新模型。新功能把传统"按固定帧率(默认 1 FPS)逐帧采样"的静态处理方式,替换为模型主动判断"看哪里、以什么速度看、用哪条模态"的智能体循环,可在视觉帧、音频、转录文本三种通道间动态切换,按需加载相关片段。
背景:此前多模态模型分析长视频时必须按固定帧率逐帧采样,导致 token 消耗巨大且容易漏掉帧间细微变化。从 10 分钟教程到 90 分钟讲座、再到数小时录像,开发者长期陷于"高 token 换高细节"或"低细节换低 token"的两难。
影响:DeepMind 官方基准显示,agentic 视频理解最高可减少 88% 的 token 消耗、66% 的分析成本,同时准确率最高提升 7%。Gemini 3.7 Flash 在 1H-VideoQA 基准的精度-成本 Pareto 前沿上同时跑赢 GPT-5.6 Sol、GPT-5.6 Terra、Claude Opus 5 与 Grok 4.6。该功能还解锁了"亚秒级时刻检索""多小时视频大海捞针""动态帧率异常检测""动作与对象精确计数"四类新能力。
总结:DeepMind 把"agentic"思路从图像扩展到视频,让模型从被动消费者变成会查档的研究者。功能已通过 Gemini API 在 Google AI Studio 与 Gemini Enterprise Agent Platform 上线,沿用标准 token 定价、不另收功能费,未来数月还将随 Gemini App 与 YouTube 的"Ask YouTube"功能覆盖更多用户。
参考来源:
- https://deepmind.google/blog/introducing-agentic-video-in-gemini/
- https://agentictribune.com/article/20260901-google-rolls-out-agentic-video-understanding-in-gemini-flash-to-cut-token-use-and-costs
- https://www.startuphub.ai/ai-news/ai-research/2026/gemini-agentic-video-understanding-cuts-costs
- https://theaicronicle.com/en/news/tools/google-gemini-agentic-video-understanding
- https://nextcryptotrends.com/googles-gemini-models-launch-agentic-video-understanding
- https://www.orcarouter.ai/blog/gemini-agentic-video-understanding
- https://wpnews.pro/news/google-says-gemini-3-7-flash-is-at-the-frontier-of-accuracy-and-cost-at-agentic









