Google DeepMind piloted a double-blind evaluation method for a Gemini frontier model using cryptographic and confidential-computing technology
Feature update Provisional 86% confidence first seen
Google DeepMind announced a pilot for double-blind evaluation of a proprietary Gemini model, designed to prevent evaluators from seeing the model weights and to prevent the model from accessing the test questions. The company plans to test Gemini Flash Lite against confidential benchmarks with partners including the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons using confidential GPU hardware and cryptographic safeguards to reduce benchmark contamination.