Google DeepMind said this week it has begun testing a Gemini Flash Lite model inside a cryptographically sealed evaluation environment, structured so the model cannot access external test prompts before or during grading. DeepMind describes the pilot as a first-of-its-kind double-blind test of one of its own proprietary, frontier-class systems, a label that is the company’s own description of its project rather than a claim independently verified by a third party. Four outside groups are running the external tests: OpenMined, AVERI, MLCommons, and the Singapore AI Safety Institute.

The problem the pilot targets is real and widely acknowledged inside the industry: benchmark contamination. If a model’s training data, or the evaluation pipeline itself, has been exposed to test questions in advance, a high score measures memorization rather than capability. That gap is a major reason published leaderboard numbers have drifted from what developers actually experience when they put a model into production. DeepMind’s own post uses the analogy of a student who has already seen the exam questions.

The mechanism is what separates this from prior safeguards. DeepMind says labs have relied on zero-logging protocols and contractual confidentiality to keep test prompts secret, agreements that depend on trust rather than verification. The new setup adds a cryptographic layer: external evaluators submit benchmarks into a sealed environment where the model can generate answers, but the underlying test content stays inaccessible to anyone who might later use it to tune the system. DeepMind frames the design as increasing evaluation integrity rather than replacing internal testing, which it says remains ongoing alongside checks from research labs, civil society groups, and the national safety institutes now testing frontier systems.

The governance question sits underneath the cryptography. DeepMind is both the subject of the evaluation and the party that selected the evaluators, funded the pilot, and will decide how the results get used or disclosed. A sealed box prevents a model from peeking at questions, but it does not, on its own, guarantee that a negative result becomes public or that the benchmark’s difficulty was set independently of the lab being graded. The credibility of double-blind evaluation as an industry norm will depend on whether AISIs and outside groups can run it without a lab’s cooperation, not just with a lab’s participation.

For teams that cite vendor benchmark scores in procurement decisions, this pilot is worth tracking rather than treating as settled practice. If the Singapore AI Safety Institute or MLCommons publishes results (or a methodology) independent of DeepMind’s framing, that is the signal a cryptographic evaluation standard is becoming portable across labs rather than a one-off pilot on a single company’s smallest model.

Google DeepMind detailed the pilot in a blog post published August 27, 2026.