AI article
Piloting the world's first double-blind AI evaluations
No preview is available. Read the original article for the full story.
deepmind | Aug 27, 2026 | William Isaac, Sol Messing, Kristian Lum
Automated excerpt
August 27, 2026 Responsibility & SafetyWilliam Isaac, Sol Messing and Kristian LumBuilding trust in proprietary model benchmarks using cryptographically secure environmentsImagine a student is set to take a high-stakes exam. We're partnering with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, to test a Gemini Flash Lite model against confidential benchmarks in a privacy-preserving environment, increasing evaluation integrity. At Google, we assess our AI systems using a broad spectrum of evaluations throughout model development and deployment, but we don’t rely on internal testing alone.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.