TLDRocket
6 August 2026
The divergence between what AI does in the lab and what it does in the world is crystallizing into a liability crisis. Stanford's NOHARM benchmark tested medical systems from OpenAI, Anthropic, and Doximity on 1,100 real clinical cases and found they all share the same critical flaw: they omit crucial information at alarming rates—76.6% of harmful errors were omissions rather than false statements. This creates what cardiologist Eric Topol calls an "illusion of readiness," a dangerous comfort that masks when AI systems fail silently rather than obviously. Doximity's Ask tool performed best, yet even that victory is conditional; the question of who bears legal responsibility when medical AI goes quiet remains unsettled as regulators and hospitals navigate uncharted territory.
Read the full briefing →