TLDRocket
Sign in

Open-world evaluations for measuring frontier AI capabilities

AI Snake Oil Sayash Kapoor

Researchers introduced open-world evaluations, a new method for testing AI capabilities in complex real-world tasks beyond standard benchmarks, and launched CRUX, a collaboration of 17 researchers that successfully tasked an AI agent with building and publishing an iOS app to the App Store. In CRUX's first experiment, the agent completed the task after two errors, one requiring manual intervention, with the entire process costing approximately $1,000. This approach aims to provide early warnings about emerging AI capabilities and identify blind spots in existing benchmarks before such abilities become widespread.

Why it matters

Introducing CRUX, a new project for evaluating AI on long, messy tasks

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.