TLDRocket
Sign in

Open-world evaluations for measuring frontier AI capabilities

AI as Normal Technology Sayash Kapoor

Researchers built CRUX, a project that tests AI by having it actually do things—like publish a real iOS app—instead of just acing benchmarks. Their AI agent got the app live with just one hiccup, which is a neat trick and a warning sign for app-store spam.

Based on reporting by AI as Normal Technology, Sayash Kapoor — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Benchmarks are running out of road. Once a test gets popular enough, it also gets precise enough to train against, and that's exactly what's happening across coding, browsing, and reasoning suites—SWE-Bench, ARC-AGI, Terminal Bench, you name it, they're all getting 'successor' versions because the originals saturated fast. A group of 17 researchers spanning academia, government, and industry decided the fix isn't a better benchmark. It's ditching the tidy sandbox altogether.

Their new effort, called CRUX, throws AI agents at messy, real tasks that can't be gamed the way a leaderboard can. No held-out test sets, no reward-hacking shortcuts—just an agent trying to do something a human would actually care about. First up: publishing a working app to Apple's App Store. That means writing code, sure, but also signing certificates, hosting a privacy policy, filling out Apple's review forms, and surviving the actual App Store review process, not a simulated one.

The agent pulled it off. Two mistakes total, one small enough to shrug off, one bad enough that a human had to step in—the agent lost track of its own credentials and, weirdly, invented a fake phone number for Apple's review team. The whole run cost about $1,000, though the researchers note the coding and submission itself only ran $25; nearly all the money went to babysitting the app's status afterward. The app is live right now on the App Store.

That's the impressive part. The uncomfortable part is what it implies at scale. If one lightly-supervised agent can get an app through Apple's gates for roughly the price of a nice dinner, nothing stops thousands of agents from doing the same thing simultaneously, flooding app stores with auto-generated junk. The team actually gave Apple a heads-up a month before publishing, which tells you they see this less as a cool demo and more as an early warning flare.

The bigger argument here is that sample-size-of-one experiments, however unscientific they look on paper, catch things averaged benchmark scores miss entirely—CAPTCHA walls that tank scores for reasons unrelated to real skill, or agents that pass automated tests but would get rejected by an actual human maintainer. CRUX plans to keep running these open-ended trials across new domains, including AI-driven research and governance tasks, treating each one less as a leaderboard entry and more as a canary in the coal mine.

My take — AI-written commentary, not fact-checked reporting

I've been saying for a year that benchmark charts are a lousy crystal ball, and this paper is basically proof: an $1,000 experiment with a sample size of one told us more about near-term risk than another point on METR's time-horizon graph ever could. The app-store-spam angle is the real headline, not the coding success — regulators and platform operators are always a year behind these capability jumps, and disclosing it to Apple ahead of publication is exactly the kind of quiet, unglamorous safety work that gets ignored in favor of splashier AGI debates.

Read more about this at: AI as Normal Technology

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.