TLDRocket
Sign in

Addy Osmani on Building AI Evaluation Practices

X

Addy Osmani says the secret to good AI use isn't a prompt trick — it's actually reading the outputs and logging where they go wrong. Sounds boring, but that discipline is what separates people who trust AI blindly from those who actually get good results.

Based on reporting by X — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Addy Osmani has been making the rounds with a message that's less exciting than most AI advice but probably more useful: stop chasing magic prompts and start building an evaluation habit. His argument boils down to something almost old-fashioned. You get good at judging AI output the same way you get good at anything else — by doing it a lot, paying attention to what fails, and keeping notes.

That means actually reading what the model spits out instead of skimming it and moving on. It means writing down the specific ways an answer was wrong, not just filing it away as "AI messed up again." Osmani's framing treats this like a craft skill, closer to code review or editing than to typing clever instructions into a chat box. The people who get consistently strong results aren't the ones with the best prompt library. They're the ones who've built a mental (or literal) log of failure modes over months of use.

This matters because most public conversation about AI still fixates on model releases and benchmark scores, as if a bigger number on a leaderboard automatically translates into better real-world output. Osmani's point cuts against that. Two people with the same GPT-4 or Claude access can get wildly different results depending on whether they've developed any evaluation instinct at all. One treats every answer as gospel. The other has learned, through repeated exposure to mistakes, exactly where the model tends to hallucinate, oversimplify, or miss context.

There's also a quieter implication here for teams building products on top of these models. If judgment is a skill built through logged failures, then companies without a systematic process for tracking what goes wrong are essentially operating on vibes. Osmani's advice reads less like a productivity tip and more like an argument for treating AI evaluation as an actual discipline, with the same rigor you'd expect from QA testing, rather than something you pick up passively by using the tool a lot.

None of this is flashy, and that's sort of the point. In a space obsessed with the next model drop, someone arguing for slow, deliberate error-logging stands out precisely because it's unglamorous. But it's also the kind of advice that tends to be right long after this week's model announcement is forgotten.

My take — AI-written commentary, not fact-checked reporting

I've said it before and I'll keep saying it: the model matters less than the muscle you build around it, and Osmani is basically confirming that from the inside. Everyone wants a shortcut, but the people getting real value from AI right now are the ones treating it like a junior colleague whose work needs checking, not an oracle. That's not a sexy takeaway, but it's the one that actually compounds.

Read more about this at: X

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.