TLDRocket
Sign in

A new medical AI study found the same flaw in OpenEvidence, OpenAI, Anthropic, and Doximity

Fortune Lily Mae Lazarus

A new benchmark called NOHARM tested medical AI tools from OpenEvidence, Doximity, OpenAI, and Anthropic on real clinical cases. Doximity scored best, but all four models mostly failed by leaving stuff out, not by being wrong.

There's a quiet arms race happening in doctors' pockets, and it just got its first real report card. NOHARM, a new benchmark built by researchers at Stanford, Harvard, and the ARISE network, ran 1,100 real clinical cases through four medical AI tools and gathered about 13,000 physician annotations to grade them for patient harm. Doximity's Ask assistant came out ahead of OpenEvidence, OpenAI's GPT-5.6 Sol, and Anthropic's Claude Fable 5.

OpenEvidence isn't taking that quietly. CEO Daniel Nadler pushed back hard on the study's methodology, telling Fortune that a rigorous benchmark shouldn't let AI systems with perfect memories request re-tests, and pointing out, correctly, that NOHARM hasn't been peer-reviewed. It's a fair jab, especially given how much is riding on perception right now. OpenEvidence has gone from a $1 billion valuation in February to $12 billion this January, raising roughly $700 million along the way. Doximity, by contrast, sells its Ask tool through enterprise contracts with more than 150 health systems and posted $145.4 million in quarterly revenue this spring.

But the scoreboard isn't really the story here. The real finding is that 76.6% of harmful errors across every single model tested were omissions — things the AI simply left out — rather than outright factual mistakes. Eric Topol, the Scripps Research cardiologist who co-chairs Doximity's human-review layer PeerCheck, has built a career studying diagnostic error, and he says omissions are the kind of mistake medicine can least afford. He calls the current state of these tools an 'illusion of readiness,' a phrase that should probably get stitched onto a few pitch decks.

None of this is happening in a regulatory vacuum, either. The FDA loosened its rules on AI clinical decision-support tools in January, as long as a doctor can still trace the AI's reasoning. States went the other way, passing more than a dozen laws this year requiring human sign-off before any AI-assisted decision reaches a patient. And malpractice law hasn't caught up with either approach — nobody has cleanly settled who's on the hook when a model's suggestion goes sideways: the doctor, the hospital, or the company that built the tool.

That unresolved liability question is probably the thing to watch. Benchmarks like NOHARM will keep sparking methodology fights between well-funded rivals, but the courts and legislatures deciding who pays when an omission hurts someone will shape this market far more than any single leaderboard.

My take

Nobody should be surprised that the flashiest, best-funded medical AI tool isn't automatically the safest one — hype and clinical rigor have never tracked together, and a $12 billion valuation says nothing about whether a model knows what it doesn't know. The omission problem is the real story buried under the vendor spat, and until liability law catches up, doctors are the ones absorbing the risk that investors are pricing as upside.

Read more about this at: Fortune

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.