时间:2026年8月27日至28日
地点:加州山景城 / 新加坡
人物:Google DeepMind 团队、新加坡 AI 安全研究所、OpenMined、AVERI、MLCommons
事件:Google DeepMind 启动全球首个针对前沿 AI 模型的"双盲"评估试点,在加密"黑箱"中以 Gemini Flash Lite 模型对一组机密基准测试进行评估,确保评估方看不到模型权重、Google 也看不到测试题目。
背景:AI 模型若在训练阶段见过测试题,基准分数便不可信;过去的高敏感度外部评估只能二选一——评估机构交出测试提示词让厂商提前看到,或厂商交出模型权重承担 IP 泄露风险,最新案例是 Anthropic Fable 5 在 ARC-AGI 基准上的延期评估。
影响:DeepMind 采用 Google Cloud 机密计算产品 Confidential Space 进行密码学隔离,使独立机构能在不牺牲数据主权的前提下对前沿模型做严格测试,对网络安全和政府部门测试尤为关键,并有望成为行业模型监管新标准。
总结:Google 同时发布技术报告公开方法论与结果,呼应 OpenAI 2026 年 5 月的独立第三方评估路线图,推动前沿 AI 评估从合同约束升级为密码学约束。
参考来源:
1. The Decoder - https://the-decoder.com/ai-benchmarks-have-a-trust-problem-and-google-wants-to-fix-it/
2. Google DeepMind 官方博客 - https://deepmind.google/blog/piloting-the-worlds-first-double-blind-ai-evaluations
3. The AI Wrap - https://www.theaiwrap.com/story/ai-benchmarks-have-a-trust-problem-and-google-wants-to-fix-it
4. Brief News - https://www.brief.news/2026/08/27/google-deepmind-doubles-blind-ai-evaluation
5. AI Pulse Lab - https://aipulselab.tech/news/ai-benchmarks-have-a-trust-problem-and-google-wants-to-fix-it-79dc71
6. TokenFeed - https://tokenfeed.ai/ai-tests-get-a-blindfold-google-stops-models-from-cheating









