TLDRocket
Sign in

Model Evaluation

56 summarised stories about Model Evaluation, each linking back to the original source. Browse all topics →

+ Follow this topic

Thursday, 6 August 2026

A new medical AI study found the same flaw in OpenEvidence, OpenAI, Anthropic, and Doximity

Fortune 18

A Stanford-led benchmark called NOHARM tested AI systems from OpenEvidence, OpenAI, Anthropic, and Doximity on 1,100 real clinical cases and found a common flaw: all models frequently omit important information rather than stating falsehoods, with 76.6% of harmful errors being omissions. Doximity's Ask tool performed best in the study, though OpenEvidence disputed the methodology. The findings highlight that current medical AI systems maintain what cardiologist Eric Topol calls an "illusion of readiness," creating liability questions as regulators and hospitals decide who bears responsibility when AI suggestions prove wrong.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.