TLDRocket
18 September 2026
AI’s main plotline today wasn’t about bigger models—it was about who gets to measure them, and how. Anthropic is partnering with Accenture to stand up independent embedded evaluation for “frontier AI,” including red-teaming plus alignment and safeguards testing, with both companies expecting to invest at least $1 billion over five years to build evaluation capacity. The idea is simple but expensive: add evaluators with near-employee access inside labs to generate incident reporting and a more verifiable public account of safety, while keeping Anthropic as the accountable party. Dario Amodei’s parallel push to “pace the frontier” leans on the same mechanism—independent safety evaluators, coordinated across democratic countries—yet the hard part remains definition and enforcement, not evaluation tech.
Read the full briefing →