TLDRocket
Sign in

Sick and wrong: Ontario auditors find doctors' AI note takers routinely blow basic facts

The Register

Ontario's auditor checked the AI note-takers doctors use, and most flunked basic accuracy tests. Some invented symptoms and treatment suggestions patients never mentioned.

Based on reporting by The Register — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Ontario runs a program called AI Scribe, letting doctors and nurse practitioners use AI tools to transcribe and summarize patient visits instead of scribbling notes by hand. The province's Auditor General just checked how well 20 approved vendors actually perform, using simulated doctor-patient recordings reviewed by medical professionals against the AI-generated output. The results are rough.

Nine of the 20 systems fabricated information outright, including treatment suggestions that were never discussed in the recording. Evaluators found notes claiming no masses were found or that a patient seemed anxious, none of which came up in the actual conversation. Twelve systems inserted incorrect drug information. Seventeen missed key mental health details that patients had explicitly raised, and six either partially or fully dropped mental health issues from the notes entirely. This isn't a case of AI being clumsy with formatting or missing a date. These are the kinds of errors that could change how a patient gets treated.

Part of the problem traces back to how these tools got approved in the first place. The procurement scoring gave 30 percent of a vendor's total evaluation to whether it had a physical presence in Ontario, while medical note accuracy counted for just 4 percent. Bias controls got 2 percent. Privacy and risk assessments got 2 percent. SOC 2 Type 2 compliance, a common security certification, added another 4 points. Add it up and the things that actually determine whether a note is safe to put in someone's chart barely moved the needle compared to whether the company had a local office.

OntarioMD, the group that helped run the procurement and supports physicians adopting new tech, has told doctors to manually double-check their AI notes. Fine advice, except the report points out that none of the approved systems have a mandatory attestation step forcing that review to happen. So the safety net depends entirely on busy clinicians remembering to catch what the software got wrong, with no system-level backstop if they don't.

The Ministry told CBC that more than 5,000 physicians are already using AI Scribe and that no patient harm has been reported yet. The Register asked the Ministry directly whether it plans to act on the auditor's recommendations and hasn't heard back. Given that large language models have separately been shown to botch differential diagnoses in the majority of test cases, treating 'no harm reported yet' as reassurance feels like betting the house that nobody's noticed the wiring is bad.

My take — AI-written commentary, not fact-checked reporting

This is what happens when procurement checklists get written by people optimizing for economic development points instead of patient safety, and it's a preview of what's coming everywhere AI gets rubber-stamped into critical workflows without real accuracy audits. Ontario deserves credit for actually measuring this and publishing it, which is more transparency than most jurisdictions bother with, but 'no known harms reported' is a low bar when nobody's built a system to catch the harms in the first place.

Read more about this at: The Register

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.