Google Deepmind has introduced a cryptographic method to address the trust issues in AI benchmarking by conducting a double-blind evaluation of its Gemini model. The initiative, in collaboration with the Singapore AI Safety Institute, aims to ensure that models cannot access test questions in advance, thereby making evaluation results more reliable. The pilot project involves testing a Gemini Flash Lite model against confidential benchmarks using a secure setup that prevents any leakage of test data or model weights.

The method is designed to resolve the longstanding dilemma in external evaluations, where either the model provider or the evaluator had to compromise data privacy. Previously, companies like Anthropic faced delays in evaluating models like Fable 5 due to a 30-day data retention policy. The new approach uses Google Cloud's Confidential Space to cryptographically verify the privacy of both the model and the test data, ensuring neither party can access the other’s information. This setup eliminates the need for zero-logging protocols and contractual safeguards alone, offering a more robust solution for secure model evaluation.

The initiative is particularly significant for sensitive applications such as cybersecurity and government testing, where data sovereignty and security are critical. Google claims the effort sets a new standard for model oversight and aims to help the industry build more reliable and trusted AI systems. The company has detailed the methodology and results in a technical report.

Source: thedecoder