TLDRocket
Sign in

Improving breast cancer screening workflows with machine learning

Google Research Covered by 2 sources

Google and the NHS tested AI as a second reader for breast cancer screening across two big studies. It caught more cancers and could cut human reading workload by nearly half.

Based on reporting by Google Research — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Britain's breast screening program is running on fumes. A 30% shortage of radiologists, projected to hit 40% by 2028, is straining a double-read system that has otherwise proven itself over decades. Into that gap steps Google Research, which teamed up with several NHS trusts on the AIMS study, publishing two papers in Nature Cancer that test whether an AI system can actually help rather than just look good on paper.

The first study threw the AI at nearly 116,000 real mammograms from five NHS screening services, each with its own quirks in how second readers and arbitration work. The results were striking: the AI system caught cancers at a rate of 9.33 per 1,000 women, compared to 7.54 for the original human first reader, and it flagged a quarter of the interval cancers that slipped through the standard double-read process undetected. It was especially good at spotting invasive cancers and at reading first-time screening mammograms, where it boosted sensitivity while cutting false positives. A follow-up deployment phase pushed the system live at 12 sites across two London screening services, processing over 9,200 cases with a median AI turnaround of 17.7 minutes versus more than two days for a human read. That live run also exposed a data drift problem — the system's training data didn't perfectly match current clinical patterns — which the researchers caught and corrected by recalibrating thresholds on the fly.

The second study is where things get more interesting, because it moves past raw accuracy and into how actual radiologists behave when an algorithm is looking over their shoulder. Twenty-two accredited readers arbitrated more than 8,700 cases pulled from 45,602 women's records, comparing the traditional two-human workflow against one where a single human reader's opinion was paired with the AI's call. The AI-enabled version held its own: statistically non-inferior sensitivity and specificity, but with an estimated 46% cut in the total number of human reads needed, and a 36-44% drop in the actual time radiologists spent reading. For a system buckling under staffing shortfalls, that's the kind of number that gets attention.

But the study also surfaced something uncomfortable. Human arbitrators overturned the AI's correct recall decisions on 93 cancer-positive cases, many of them the trickiest interval and next-round cancers to spot. That's not a small footnote — it means the biggest obstacle to squeezing value out of this technology might not be the AI's accuracy at all, but how much radiologists trust it when it disagrees with their gut. Google's researchers frame this as a call for better explainability tools, so clinicians understand why the AI flagged what it flagged instead of just seeing a number they can override.

None of this is a green light for AI to replace radiologists in the UK tomorrow. The researchers are careful to note that prospective clinical trials still need to prove the system works reliably in the wild, and thorny operational issues remain, like handling surges in arbitration cases and keeping the system properly calibrated to different patient populations over time. Still, the shape of the argument is coming into focus: not AI replacing radiologists, but AI absorbing enough of the workload that the shrinking pool of human experts can actually keep up.

My take — AI-written commentary, not fact-checked reporting

This is one of the more credible healthcare AI stories I've seen this year precisely because it doesn't oversell itself — the researchers surface the failure mode (93 missed cancers overruled by humans) instead of burying it, which is rare. The real story isn't detection accuracy, it's trust calibration: humans overriding a system that's often right is a much harder problem than training a better model, and it's the one nobody's solved yet.

Read more about this at: Google Research

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.