Editorial illustration: two sealed envelopes—dark weights and parchment prompts—enter a brass-locked glass evaluation chamber; only a thin slip of bounded metrics exits into empty cool space
Reciprocal secrecy: sealed weights and prompts meet inside an attested chamber; only bounded results leave., AI-generated editorial illustration, not a news photo

On 27 August 2026, Google DeepMind’s Responsibility & Safety blog published Piloting the world’s first double-blind AI evaluations, by William Isaac, Sol Messing and Kristian Lum. Fact (DeepMind): the lab describes what it calls the world’s first double-blind evaluation of a proprietary, frontier-class model, pairing a Gemini Flash Lite model with confidential benchmarks inside a cryptographically protected environment. Partners named on the post are the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons. Fact (technical report / AVERI): the model under test is Gemini 2.5 Flash Lite; the main reserved prompts come from MLCommons’ AILuminate corpus; AVERI handled encryption, decryption and grading; Singapore AISI supplied a separate private harmful-content set focused on Singapore’s context. Inference (labelled): if the architecture works as described, the durable object is not a new Gemini safety league table—scores were not published—but a reciprocal-secrecy pattern for closed-model audits.

[1][2]

What the pilot actually ran

Strip adjectives and the stack is concrete.

Hardware and cloud. The evaluation ran in Google Cloud Confidential Space on an A3 Confidential VM with Intel TDX host-memory encryption and an NVIDIA H100 Confidential GPU. DeepMind’s blog frames this as Confidential Space within Google Cloud’s Confidential Computing portfolio: the evaluator cannot see Gemini weights; Google cannot see the evaluator’s test prompts.

Software coordination. OpenMined’s PySyft coordinated submissions, approvals and result release. Parties verify remote attestation before releasing secrets; approved code runs; only bounded metrics leave; the enclave is meant to be decommissioned without leftover state.

Benchmark content. Per AVERI and the technical report, prompts were drawn from a reserve AILuminate set (never-before-used), covering domains including CBRNE hazards, cyberattacks, hate speech, self-harm and violent-crime elicitation. Singapore AISI’s parallel set targeted harmful-content elicitation in Singapore’s context. AVERI encrypted prompts and outputs with a key shared with no other party, then graded against AILuminate criteria.

What was not released. DeepMind’s announcement and the public write-ups describe architecture and hazard categories. They do not publish Gemini 2.5 Flash Lite’s public scores or a task-by-task breakdown. AVERI says it provided Google DeepMind a confidential findings report.

[1][2]

Reciprocal secrecy is the product

Independent evaluation of closed models has long forced a bad bargain. Either evaluators hand over prompts—risking contamination if the provider retains or later trains on them—or providers hand over weights, exposing intellectual property and dual-use assets. Contracts and zero-logging policies reduce risk; they still require trusting the counterparty’s systems and behaviour. Research cited in the technical report (including Xu et al., 2024 on leakage evidence and Schaeffer et al., 2026 on contamination inflation) is why “trust us not to peek” is no longer treated as enough for high-stakes sets.

Double-blind evaluation, as DeepMind and partners define it, relocates trust. Both sides send assets into a measured workload they inspect before release. Encryption limits casual observation of memory; attestation is meant to answer whether the intended stack is running before decryption keys or protected resources are released.

Inference (labelled): the strategically interesting claim is therefore about audit infrastructure, not about where Gemini ranks on AILuminate. Calling the run a “world’s first” for a proprietary frontier-class model is DeepMind’s competitive framing; the falsifiable engineering object is whether reserved prompts stayed out of Google’s training path and weights stayed out of evaluators’ hands under the stated attestation model.

A 2024 OpenMined collaboration with Anthropic and the UK AI Security Institute had already shown enclave evaluation with proxies (including a GPT-2 stand-in and a tiny sample). AVERI’s account stresses that this 2026 pilot moved from proxies to a production proprietary model and a real reserved safety corpus—mechanism validation at production scale, still not a public scoreboard.

[1][2]

What cryptography proves—and what it cannot

Fact vs claim. That the parties ran an attested Confidential Space workload with mutual asset secrecy is the core stated achievement. That this yields “trusted evaluation” in the everyday sense is a broader claim the cryptography alone cannot finish.

Attestation can support a strong procedural statement: assets were released to a measured workload both sides approved, and hardware encryption limited access while it ran. It does not prove the benchmark is complete, culturally adequate, well labelled, or predictive of deployment behaviour. A clean double-blind run can faithfully emit an unhelpful metric. Contamination control and test design are different problems.

Independence has dimensions. External organisations supplied and scored reserved prompts, which is independence of question ownership and analysis. Google still provided the model and hosted cloud technology. Those are not the same dimension. Buyers and regulators should keep them separate when reading “independent evaluation” headlines.

Output policy is part of the perimeter. Even a well-isolated computation can leak a benchmark through raw prompts, token traces, or verbose errors. The security story depends on pre-agreed bounded outputs—not only on encrypted RAM.

[1][2]

Steelman, gaps, and falsifiers

Steelmanning DeepMind and partners. Suppose high-stakes safety and cyber evaluations only remain diagnostic if prompts never enter provider training pipelines, and frontier labs will not ship weights to every auditor. Then building a reciprocal enclave pattern—mock interfaces, dual approval, allow-listed redactions via tools such as syft-restrict, ephemeral teardown—is the honest next step beyond NDAs. On that steelman, withholding public scores can be correct for a methodology pilot: publishing the architecture without burning the reserved set is the point. Procedural friction (legal agreements, code review), which the technical report says dominated over <5% compute overhead, is then a feature of seriousness rather than a failure of GPUs.

Gaps the primary materials themselves mark. The technical report acknowledges that not all proprietary inference code could be fully inspected or allow-listed; that individual Confidential Space builds are not independently reproducible (private signing keys enter the build); and that Google services are used to sign and verify the attestation report—placing Google in the verification path and increasing trust required in the model provider. MLCommons and secondary reporting also stress that technical secrecy does not replace stewardship, origination rules and legal controls over a benchmark’s lifetime.

Labels for editors. Facts: date and byline; partners; Gemini 2.5 Flash Lite; Confidential Space / A3 / TDX / H100; PySyft; AILuminate reserve domains; Singapore AISI parallel set; no public score table in the launch materials; documented limitations above. Claims: “world’s first” double-blind evaluation of a proprietary frontier-class model; major step beyond contracts/zero-logging; template for government and cyber-sensitive audits. Inferences: reciprocal-secrecy infrastructure is the durable news; public readers should not treat the pilot as a disclosed Gemini safety grade.

Falsifiers. The reading weakens if partners publish independently reproducible attestation artifacts with a non-Google verification path and bounded public metrics for the same reserved methodology; or if credible evidence shows reserved prompts later surfaced in Google training or logging. It strengthens if other cloud/hardware operators repeat the pattern without the provider sitting in the attestation chain.

[1][2]

What to watch in six months

By roughly early 2027, three checks matter more than another “world’s first” headline. First: do attestation artifacts become reproducible above the unavoidable vendor firmware layer, with verification that does not require Google to sit in the signing path? Second: do at least one non-Google cloud or hardware operator and one additional frontier lab complete a comparable reserved-set run, so the pattern is not a single-vendor demo? Third: do AISIs and auditors gain publication rules that protect evaluator independence when findings are unfavourable—while still keeping prompts sealed?

If those checks hold, double-blind evaluation will have earned its keep as a way to widen access to closed models without burning private exams. If not, the industry risks more secure ceremonies without more trustworthy public results. The box can keep the exam from the student and the student’s internals from the examiner. It cannot decide whether the exam was fair, or whether the institution will act on a bad grade.

[1][2]