TLDRocket
Sign in

Google DeepMind piloted a double-blind evaluation method for a Gemini frontier model using cryptographic and confidential-computing technology

Feature update Provisional 86% confidence first seen

Google DeepMind announced a pilot for double-blind evaluation of a proprietary Gemini model, designed to prevent evaluators from seeing the model weights and to prevent the model from accessing the test questions. The company plans to test Gemini Flash Lite against confidential benchmarks with partners including the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons using confidential GPU hardware and cryptographic safeguards to reduce benchmark contamination.

Source coverage

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.