The Quest for Embedded Evaluators
Zvi (Don't Worry About the Vase) TheZvi ● Covered by 18 sources
Anthropic and OpenAI now want outside evaluators inside the lab. The hard part is finding reviewers who are independent, trusted, and not hired by the company they’re judging.
Based on reporting by Zvi (Don't Worry About the Vase), TheZvi — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Anthropic’s Dario Amodei has pushed the idea of “embedded evaluators”: outsiders placed inside a lab with employee-level access so they can bring an outside view and report on what’s going on. OpenAI has now echoed that idea, along with a mild call for international coordination. The concept sounds simple until you ask the obvious question: who gets to do the evaluating?
That problem sits at the center of the piece. If the labs pay, independence gets fuzzy. If governments refuse to fund or oversee the work, the pool of credible evaluators gets thin fast. The author argues that this is not a reason to skip auditing, just a reason to be honest about how awkward the setup is. The revolving door is real. So is the funding problem.
A public letter from Geoffrey Hinton, Stuart Russell, and Arvind Narayanan lays out a basic standard for credible embedded evaluators: they should be meaningfully independent, their pay should not depend on their findings, they should have no ownership or commercial ties, conflicts should be disclosed, and they should be protected from retaliation. They should also get access similar to highly privileged employees, aside from limits needed to protect data. That list gets a clear thumbs-up here.
The most concrete example so far is Anthropic’s partnership with Accenture. Anthropic says Faculty, Accenture’s specialist AI business, will lead the work, including model evaluation, red-teaming, alignment assessments, and testing safeguards. Anthropic and Accenture each expect to invest at least $1 billion over the next five years in building capacity for this, and Anthropic says it will fund Accenture directly. Anthropic also says it is talking with METR and other nonprofit evaluators about piloting parts of embedded evaluation using their own funding.
The author’s read is that the best answer is not one perfect evaluator, because there probably isn’t one. METR looks like the strongest pick for the expert slot; a more legible, boring organization would help balance that out. Accenture may be a weird fit on paper, but the broader point is that frontier labs probably need several kinds of oversight at once, with different strengths and different weaknesses. That is messy. It is also probably the only workable version of the idea right now.
My take — AI-written commentary, not fact-checked reporting
The clean-room fantasy is doing a lot of work here. If the only acceptable evaluator is pure, detached, and magically funded, then nothing gets audited and everyone gets to feel principled while the models keep shipping. The less glamorous answer is a mixed bench: experts, boring firms, and a lot more sunlight than labs usually enjoy.
Read more about this at: Zvi (Don't Worry About the Vase)