Piloting the world's first double-blind AI evaluations

Piloting the world’s first double-blind AI evaluations

试点全球首个双盲人工智能评估

August 27, 2026 | Responsibility & Safety | William Isaac, Sol Messing and Kristian Lum 2026年8月27日 | 责任与安全 | William Isaac, Sol Messing 和 Kristian Lum

Building trust in proprietary model benchmarks using cryptographically secure environments 利用加密安全环境建立对专有模型基准测试的信任

Imagine a student is set to take a high-stakes exam. If they accidentally peek at the test questions in advance, achieving a perfect score is influenced by this knowledge, making it a meaningless accomplishment. To truly measure what they know, they must have no visibility of the test questions until it’s time to take the exam. That is the exact challenge the industry faces when evaluating advanced AI models. If a model has already seen the test questions - a problem known as benchmark contamination - the results can only be trusted to an extent. 想象一下,一名学生即将参加一场至关重要的考试。如果他不小心提前看到了考题,那么他取得满分就受到了这种已知信息的影响,从而使这一成就变得毫无意义。为了真正衡量他们的知识水平,他们必须在考试开始前对考题一无所知。这正是行业在评估先进人工智能模型时所面临的挑战。如果模型已经“看过”考题——即所谓的“基准测试污染”问题——那么其评估结果的可信度就会大打折扣。

Today, we’re introducing the world’s first double-blind evaluation of a proprietary, frontier class AI model, which keeps external evaluations confined to a cryptographic “box” where they can’t be used by models later to optimize performance ahead of testing. We’re partnering with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, to test a Gemini Flash Lite model against confidential benchmarks in a privacy-preserving environment, increasing evaluation integrity. 今天,我们推出了全球首个针对专有前沿级人工智能模型的双盲评估。该评估将外部测试限制在一个加密的“盒子”中,确保模型无法在测试前利用这些题目来优化性能。我们正与新加坡人工智能安全研究所(AISI)、OpenMined、AVERI 和 MLCommons 合作,在一个保护隐私的环境中对 Gemini Flash Lite 模型进行机密基准测试,从而提高评估的完整性。

At Google, we assess our AI systems using a broad spectrum of evaluations throughout model development and deployment, but we don’t rely on internal testing alone. To identify potential blindspots, we work with a diverse group of external partners, including specialized research labs, civil society and national AI Safety and Security Institutes (AISIs), using their unique expertise to stress-test our models. 在谷歌,我们在模型开发和部署的全过程中使用广泛的评估手段来评估我们的人工智能系统,但我们并不单纯依赖内部测试。为了识别潜在的盲点,我们与多元化的外部合作伙伴开展合作,包括专业研究实验室、民间社会组织以及各国的人工智能安全研究所(AISIs),利用他们独特的专业知识对我们的模型进行压力测试。

As AI models become more capable, ensuring the model has not seen the test questions or prompts in advance is critical, as this can skew the results. Policymakers, researchers, and enterprises need to trust that AI benchmarks accurately reflect a model’s true capabilities and safety, but if models are able to “peek” at the evaluation questions in advance, it can artificially inflate scores and undermine this trust. 随着人工智能模型能力越来越强,确保模型没有提前接触过测试题目或提示词至关重要,因为这会扭曲评估结果。政策制定者、研究人员和企业需要相信人工智能基准测试能够准确反映模型的真实能力和安全性;但如果模型能够提前“窥视”评估题目,就会人为地提高分数,从而破坏这种信任。

Although zero-logging protocols and rigorous contractual safeguards have long kept external test prompts confidential, incorporating technical and cryptographic safeguards marks a major step forward in secure model evaluation. 尽管零日志协议和严格的合同保障措施长期以来一直确保外部测试提示词的机密性,但引入技术和加密保障措施标志着安全模型评估迈出了重要的一步。

How double-blind evaluations work 双盲评估的工作原理

Historically, high-stakes external evaluations required a tradeoff. Either evaluators handed over their testing prompts (risking the model provider seeing the test questions in advance), or the model provider handed over their model weights (risking their intellectual property). 从历史上看,高风险的外部评估往往需要权衡。要么评估者交出测试提示词(冒着模型提供商提前看到考题的风险),要么模型提供商交出模型权重(冒着知识产权泄露的风险)。

Double-blind evaluations eliminate this compromise. By using Confidential Space within Google Cloud’s Confidential Computing portfolio, we can cryptographically verify that both the external evaluation data and the proprietary model remain private to their respective owners. The evaluator cannot see the Gemini model weights, and Google cannot see the evaluator’s test prompts. 双盲评估消除了这种妥协。通过使用谷歌云机密计算产品组合中的“机密空间”(Confidential Space),我们可以通过加密方式验证外部评估数据和专有模型对各自的所有者保持私密。评估者无法看到 Gemini 的模型权重,谷歌也无法看到评估者的测试提示词。

A novel approach to building trust in model evaluations 建立模型评估信任的新方法

This cryptographic evidence helps prevent benchmark contamination and protects sensitive data. As models become more capable this becomes particularly important for highly sensitive evaluations, such as those used for cybersecurity or by government bodies. Double-blind evaluations unlock the ability for independent organizations to rigorously test advanced models without compromising data sovereignty or security. 这种加密证据有助于防止基准测试污染并保护敏感数据。随着模型能力不断增强,这一点对于高度敏感的评估(例如用于网络安全或政府机构的评估)尤为重要。双盲评估使独立组织能够在不损害数据主权或安全的前提下,对先进模型进行严格测试。

We hope this pilot establishes a new frontier for model oversight, helping the broader industry build safer, more reliable, and widely trusted AI systems. To learn more about our methodology and findings, read our technical report. 我们希望此次试点能为模型监管开辟新领域,帮助更广泛的行业构建更安全、更可靠且广受信任的人工智能系统。如需了解更多关于我们的方法和研究结果,请阅读我们的技术报告。