On Working with Wizards
One Useful Thing Ethan Mollick
Ethan Mollick says AI has stopped feeling like a coworker and started feeling like a wizard casting spells you can't watch. GPT-5 Pro caught an error in his own decade-old academic paper in under ten minutes — but he still can't explain how it found it.
Based on reporting by One Useful Thing, Ethan Mollick — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Ethan Mollick has spent years telling people to treat AI like a co-intelligence: a slightly erratic intern you correct, guide, and argue with. In his latest post, he admits that framing is starting to crack. The newest models, he says, aren't behaving like partners anymore. They're behaving like wizards — you make a vague request, something impressive appears, and you have almost no idea how it got made.
His evidence is pretty concrete. He fed his book and roughly 140 blog posts into NotebookLM and asked for a video summarizing what's happened in AI since he wrote Co-Intelligence. It nailed the MMLU scores, correctly cited a neurosurgery-exam study he'd practically forgotten citing himself, and only missed crediting his co-authors on the "jagged frontier" paper. Then he handed GPT-5 Pro his own decade-old job-market paper with instructions to critique and improve the methodology. Nine minutes and forty seconds later, it had run its own Monte Carlo simulations, re-checked his fixed-effects models, and flagged a genuine, previously unnoticed error linking two of his tables. Mollick verified it. The AI was right.
He pushed further with Claude 4.1 Opus, asking it to take a decades-old teaching spreadsheet about a desk manufacturer and rebuild it, formulas and all, as a cheese shop — same lesson, new business. A few minutes later he had a working Excel file that preserved the pedagogical point. Ask for a pitch deck next, and Claude produced a rough-but-solid deck, no factual howlers, just not launch-ready. Impressive, but also a reminder that these systems are jagged: brilliant at some things, mediocre at others, and you rarely know which until you try.
What unsettles Mollick isn't the quality of the output — it's the invisibility of the process. Claude reportedly fixed its own spreadsheet errors twice without being asked. GPT-5 Pro won't say exactly what it ran. NotebookLM gives almost no insight at all. Because these systems are trained with reinforcement learning to find their own paths to a goal, nobody — not Mollick, not the model's own summary logs — can fully reconstruct the steps. You're left checking facts you can verify and trusting the rest, which he compares to GPS or Netflix, except with much higher stakes: when GPS is wrong you hit a dead end immediately, but when an AI reworks your research, the error might be undetectable precisely because the tool has gotten good enough to hide it in competence.
His prescription is basically triage. Know when a task calls for a wizard versus a collaborator versus no AI at all. Get better at judging outputs rather than processes, since the process is increasingly locked away. And accept "good enough" as a real standard, because perfect verification, he argues, is quietly becoming impossible for a growing share of tasks — even as the tasks themselves keep getting more important.
My take — AI-written commentary, not fact-checked reporting
The uncomfortable part Mollick almost says but doesn't quite land on: if you can't verify the wizard's work, you're not actually gaining leverage, you're just outsourcing risk to a black box with excellent PR. Fine for a cheese-shop spreadsheet, less fine when the same opacity creeps into medicine, law, or infrastructure — and it will, because "good enough" has a way of becoming the only standard once it's the cheapest one.
Read more about this at: One Useful Thing