Piloting the world's first double-blind AI evaluations
Google ● Covered by 2 sources
Google introduced the world’s first double-blind evaluation of a proprietary frontier AI model, using a cryptographic “box” so external evaluators’ benchmarks can’t be used for later optimization. On August 27, 2026, Google said it would test a Gemini Flash Lite model against confidential benchmarks with partners including the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons. This adds cryptographic safeguards to reduce benchmark contamination and improve trust in the measured model results.