Google DeepMind said Aug. 27 it is piloting double-blind AI evaluations that hide a model’s identity from human graders and hide scoring rules from developers, aiming to curb the benchmark inflation spreading across AI model comparisons.

Key Takeaways

  • Google DeepMind announced on Aug. 27 it is piloting double-blind AI evaluations to curb benchmark inflation

  • In the pilot, human raters do not know which model produced which response during judging

  • DeepMind’s post does not publish rater agreement rates or inter-annotator reliability figures

  • A 300-task benchmark suite for visual reasoning models also addressed the same credibility gap this week

The Case Against Self-Graded AI Benchmarks

DeepMind, the AI research lab owned by Google, described the pilot as a move toward verifiable evaluation rather than self-reported scores. The lab said the approach marks the first attempt at running double-blind AI evaluations across AI systems at meaningful scale.

Most labs today publish benchmark scores chosen and run by the same teams building the models, leaving room for cherry-picked test sets and inflated results.

How Benchmark Gaming Became AI’s Credibility Problem

AI evaluation has followed a familiar pattern since 2025. Labs raced to publish leaderboard-topping scores as language models multiplied, and outside researchers repeatedly found that some numbers did not hold once tests moved past the exact data a model had already seen.

A 300-task benchmark suite built for visual reasoning models this week tackled a similar credibility gap, a sign the problem spans well beyond text-based chatbots.

Also Read: OpenAI’s 2026 Report Reveals a Real Gap After Hugging Face Hack

What Double-Blind AI Evaluations Actually Test

In double-blind AI evaluations, neither the human rater scoring an answer nor the lab that built the model knows which system produced which response during judging. The structure borrows from clinical trial methodology, where separating knowledge from judgment keeps expectation from skewing results.

The more pointed question is what the blinding actually measures: whether human raters score differently when they cannot anchor on a model’s reputation, and whether that gap is large enough to matter.

DeepMind’s post does not publish rater agreement rates or inter-annotator reliability figures that would let outside observers assess how much the identity-hiding step shifts scores in practice.

What Happens After The Pilot

DeepMind’s post does not name the outside evaluators taking part, specify how many models are included, or give a publication date for results. Whether other labs adopt a similar model depends on whether double-blind AI evaluations produce scores that diverge meaningfully from labs’ self-reported numbers.

A wide gap between blinded and self-reported results would validate the skepticism already building around benchmark claims industrywide.

Read Next: Anthropic Gives 3 Labs First Access to Claude Usage Data