Research into how AI can help users understand skin conditions
Google Research
Google tested an AI skin-condition tool on thousands of people, and it nearly tripled their odds of naming what they saw. But knowing the name didn't mean people knew what to do next.
Based on reporting by Google Research — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Google Research just published two studies tackling a problem most of us have faced at 2am: staring at a weird spot on your arm and typing increasingly desperate phrases into a search bar. The company's answer is an AI tool that shows a scrollable set of matching skin conditions, complete with textbook photos and symptom notes, instead of the usual pile of search results and forum posts.
The numbers are the interesting part. In a JAMA Dermatology study of 2,345 participants who reviewed real, de-identified skin cases, people using the AI tool correctly named the condition 23% of the time, versus just 8% for those using standard web search. A rigged "Wizard of Oz" version, where the AI's suggestions always matched what a panel of dermatologists actually diagnosed, pushed accuracy to 36%. Even that ceiling isn't great, which says something about how genuinely hard visual diagnosis is, AI or not.
Where things get murkier is the next step. Knowing you might have palpable purpura doesn't tell you whether to book an urgent appointment or slap on some hydrocortisone and wait it out. The standard AI group showed no statistically significant improvement in choosing the right next step compared to people using plain search, and were slightly more likely than the control group to underestimate urgency, 30% versus 27%. Naming a condition, it turns out, is not the same skill as knowing what to do about it.
Google paired that survey with a smaller, real-world study run with Stanford's HEA3RT team and Santa Clara Family Health Plan, involving 110 people from a Medi-Cal-reliant community who used the app for their own actual skin concerns, in four different languages, then talked to a clinician right after. Correct self-diagnosis jumped 260% relative to baseline, though absolute accuracy stayed low, and people leaned heavily on matching their skin to textbook photos, which is a strong argument for stocking those libraries with a wider range of skin tones and severities. Clinicians called the app's predictions consistent with their own assessment 86% of the time and said it helped the conversation 92% of the time, mostly by giving patient and doctor a shared visual reference instead of a vague description.
What's notable here is Google resisting the urge to make the tool prescriptive. It matches images to conditions and leaves interpretation to the user, on purpose. That restraint is probably smart from a liability standpoint, but it's also exactly why the next-step numbers came in flat. A tool that names things well but stays silent on urgency risks giving people confidence without giving them judgment, and that gap is where the real research problem is sitting.
My take — AI-written commentary, not fact-checked reporting
This is a rare piece of AI-in-health research that admits AI can nail the pattern-matching and still fail the part that actually matters, telling someone whether to worry. I'd rather see that acknowledged than buried under a headline about accuracy gains, and it's a useful reminder that most consumer health AI is optimized for the easy half of the problem while quietly outsourcing the hard half, urgency and judgment, back to the user.
Read more about this at: Google Research