AI article

Piloting the world's first double-blind AI evaluations

No preview is available. Read the original article for the full story.

deepmind | Aug 27, 2026 | William Isaac, Sol Messing, Kristian Lum

Automated excerpt

August 27, 2026 Responsibility & SafetyWilliam Isaac, Sol Messing and Kristian LumBuilding trust in proprietary model benchmarks using cryptographically secure environmentsImagine a student is set to take a high-stakes exam. We're partnering with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, to test a Gemini Flash Lite model against confidential benchmarks in a privacy-preserving environment, increasing evaluation integrity. At Google, we assess our AI systems using a broad spectrum of evaluations throughout model development and deployment, but we don’t rely on internal testing alone.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news