DeepMind has launched the world's first double-blind AI evaluations to ensure benchmark integrity. The initiative involves a Gemini Flash Lite model tested against confidential benchmarks in a privacy-preserving environment. By using cryptographic safeguards, the company aims to prevent benchmark contamination, where models might have seen test questions in advance, skewing results. This effort is part of a broader strategy to build trust in AI systems by ensuring evaluations reflect true model capabilities and safety. The pilot involves partnerships with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons. The approach aims to address the challenge of evaluating advanced AI models without compromising data security or intellectual property. Source: deepmind

Double-blind evaluations eliminate the traditional tradeoff between model providers and evaluators. Historically, either evaluators shared test prompts, risking model providers seeing questions in advance, or model providers shared model weights, risking intellectual property exposure. By using Confidential Space within Google Cloud’s Confidential Computing portfolio, DeepMind ensures that both the external evaluation data and the proprietary model remain private to their respective owners. The evaluator cannot see the Gemini model weights, and Google cannot see the evaluator’s test prompts. This cryptographic verification strengthens the integrity of the evaluation process and protects sensitive data. Source: deepmind

The initiative is part of DeepMind’s broader strategy to assess AI systems using a wide range of evaluations throughout development and deployment. The company does not rely solely on internal testing but collaborates with external partners, including specialized research labs and national AI Safety and Security Institutes (AISIs), to stress-test models. As AI models become more capable, ensuring they have not seen test questions in advance is critical, as this can artificially inflate scores and undermine trust. DeepMind emphasizes that this pilot represents a novel approach to building trust in model evaluations, particularly for highly sensitive applications like cybersecurity and government use. Source: deepmind