
Google DeepMind has started what it describes as the first double-blind AI evaluation of a proprietary frontier model, using a cryptographic method that keeps external test questions hidden from the model and keeps the model's weights hidden from the evaluators. Announced on August 28, 2026, the pilot runs with the Singapore AI Safety Institute and other partners and tests a model from the Gemini Flash Lite line against confidential benchmarks. The aim is to give independent testers a way to measure advanced models without either side handing over sensitive assets.
What problem does the double-blind test solve?
Benchmark scores are only trustworthy when a model has not already seen the questions. If test items appear in training data, a high score reflects memorization rather than capability, an issue known as benchmark contamination. A double-blind AI evaluation removes that possibility by keeping the external tests locked in a cryptographic "box," so a model provider cannot later tune a model specifically to those questions.
According to Google DeepMind, sensitive external evaluations used to force a compromise. Evaluators either handed over their test prompts, letting the model provider see the questions in advance, or the provider handed over its model weights and put its intellectual property at risk. The company points to a delayed ARC-AGI evaluation involving an Anthropic model, where Anthropic's 30-day data retention policy for its strongest models complicated the process, as an example of that dilemma.
How does the cryptographic setup work?
The evaluation runs on Confidential Space, part of Google Cloud's confidential computing portfolio. The setup cryptographically verifies that both the external test data and the model stay private to their respective owners. The evaluator never sees the Gemini weights, and Google never sees the test prompts.
Google DeepMind says external prompts were previously protected through zero-logging protocols and contractual safeguards. Adding technical and cryptographic protection on top of those measures is what the company calls a step forward for secure model evaluation. For this first run, the model under test comes from the Gemini Flash Lite line and is measured against confidential benchmarks supplied by the external partners.
Where does this matter most?
The cryptographic proof is meant to prevent contamination and protect sensitive data at the same time. Google DeepMind says this matters most for highly sensitive evaluations, such as cybersecurity assessments or tests run by government agencies. In those settings, independent organizations can rigorously test advanced models while keeping data sovereignty and security intact.
The company frames the pilot with the Singapore AI Safety Institute as a template for how outside groups and model developers can work together without either party giving up private information. Google hopes the approach sets a standard for model oversight and helps the industry build more reliable and widely trusted AI systems. Full details on the methodology and results are laid out in a technical report.
What to watch next
The pilot is limited to one model from the Gemini Flash Lite line and a defined set of confidential benchmarks, so results from a single run will not settle how the method scales to larger models or broader test suites. The value of the design is that any organization with sensitive benchmarks could adopt the same cryptographic guarantees, verifying scores without exposing questions or weights. That combination gives evaluators a path to test frontier systems and publish results others can trust.
FAQ
What is a double-blind AI evaluation?
It is a test where the model provider never sees the evaluator's questions and the evaluator never sees the model's weights. Google DeepMind uses Confidential Space from Google Cloud to keep both the external test data and the model private to their owners.
Which model is being tested in the pilot?
Google is testing a model from the Gemini Flash Lite line against confidential benchmarks. The pilot runs with the Singapore AI Safety Institute and other partners.
Why does benchmark contamination matter?
If a model has already seen test questions during training, its scores reflect memorization rather than real capability. Keeping the questions locked in a cryptographic box prevents a provider from optimizing a model specifically for the test, so the results stay meaningful.
This article summarizes reporting from the-decoder.com. See our editorial disclaimer for how our articles are produced.
Run a free scan to see your AI Visibility Score, SEO rating, and local citation accuracy.