TLDRocket
Sign in

Giving your AI a Job Interview

One Useful Thing Ethan Mollick

AI benchmarks like MMLU are getting shaky as a way to judge which model is actually good. Turns out picking the right AI is more like hiring a person than buying software.

Based on reporting by One Useful Thing, Ethan Mollick — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Ethan Mollick's latest piece pokes a hole in the AI industry's favorite bragging metric: the benchmark score. MMLU-Pro, one of the most cited tests, asks things like the mean cranial capacity of Homo erectus or which Cheap Trick album a 1979 live record shares its name with. Getting those right tells you almost nothing about whether a model can do your job. Add in the fact that many benchmarks leak into training data, that scores aren't calibrated so the jump from 84% to 85% might be trivial or brutal, and that some tests contain outright errors making a perfect score impossible, and you start to see why raw leaderboard rankings deserve some side-eye.

That said, Mollick isn't arguing benchmarks are useless. Stacked together, tests like AIME, GPQA, ARC-AGI and METR Long Tasks all trend the same exponential direction, and that composite signal does seem to track real capability gains showing up in medicine, finance and other fields. The catch is that almost all the rigorous benchmarks live in math, science, coding and reasoning. If what you need is writing quality, business judgment or empathetic advice, the data basically doesn't exist.

So people improvise. Mollick has his own party trick, asking every model to draw an otter on a plane, mirroring Simon Willison's pelican-on-a-bike test, and he runs models through poetry, JavaScript for starship control panels, and tiny fiction prompts to feel out their quirks. In one example, he asked four top models to write a single paragraph about someone rationing their final 47 words before death, holding a newborn. Claude 4.5 Sonnet nailed the tone, Gemini 2.5 Pro lost track of the word count entirely, GPT-5 Thinking went for wild metaphor at the cost of coherence, and Kimi K2 Thinking produced striking phrases wrapped around a story that didn't quite hold together. Fun, but subjective and hard to repeat fairly.

The more serious fix, he argues, is treating AI selection like hiring rather than procurement. OpenAI's GDPval study is his model example: experts averaging 14 years of experience built realistic four-to-seven-hour tasks across finance, law and retail, ran both AI and paid human experts through them, then had a separate blind panel grade the results. The findings were genuinely jagged, top models beat humans as financial advisors and software developers but lost badly to pharmacists, industrial engineers and real estate agents. Mollick also ran his own quirky test, pitching a guacamole drone delivery startup to multiple AIs ten times each. Grok and Copilot loved it, GPT-5 and Claude 4.5 were skeptical, and each model was oddly consistent with itself while wildly inconsistent with each other, revealing built-in risk appetites that could quietly bias thousands of downstream business decisions.

His conclusion is blunt: companies picking an AI based on a leaderboard score are essentially hiring a VP off their SAT results. If a model is going to advise hundreds of employees or handle thousands of tasks, it needs a real interview, built from your own realistic scenarios, run multiple times, judged by people who know the domain, and repeated every time a new model drops.

My take — AI-written commentary, not fact-checked reporting

This tracks with what I keep telling people who ask me which model is 'best', there's no such thing, only best-for-what. I'd add that the benchmark obsession is partly a marketing artifact, labs know journalists and investors love a leaderboard number, so the incentive to game or cherry-pick benchmarks isn't going away no matter how many flaws researchers point out. If you're deploying AI at any real scale and you're not running your own eval suite, you're basically flying blind with a nice PR chart taped to the windshield.

Read more about this at: One Useful Thing

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.