TLDRocket
Sign in

SensorFM: Towards a general intelligence and interface for wearable health data

Google Research

Google built an AI on a trillion minutes of Fitbit data to model human health. It beat custom-built models on 34 of 35 tasks without ever seeing a diagnosis.

Based on reporting by Google Research — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

For years, wearable health AI has worked one condition at a time. You wanted a heart-rate anomaly detector, you built one. You wanted a sleep-apnea flag, you built another. Each model needed its own labeled dataset, and labels in medicine are brutally expensive — confirmed diagnoses, lab draws, validated surveys, the stuff you can't retroactively conjure from six months of accelerometer readings. Google Research's new SensorFM tries a different bet: skip the labels almost entirely and just learn physiology itself, at a scale nobody has attempted before.

The numbers behind the training run are genuinely staggering. Google pulled de-identified data from five million consenting Fitbit and Pixel Watch users across more than 100 countries and all 50 U.S. states, spanning over 20 device models between September 2024 and September 2025. That adds up to more than two billion hours — over a trillion minutes — of minute-by-minute readings covering heart rate variability, blood oxygen, motion, skin temperature, and electrodermal activity. Rather than filling in the inevitable gaps caused by dead batteries or a watch left on the nightstand, SensorFM's training method, built on something called Adaptive and Inherited Masking, treats those missing stretches as just another kind of signal to learn from. That's a small architectural choice with a big practical payoff, since real-world wearable data is fragmented by nature.

The scaling results read like the kind of curve AI researchers dream about: pre-training loss dropped 31% between the smallest and largest model variants, and that improvement translated directly into better downstream health predictions — a 9% average AUC gain on classification and 21% on regression tasks. Tested against 35 discriminative tasks drawn from nearly 14,000 participants across cardiovascular, metabolic, mental health, sleep, demographic, and lifestyle categories, a frozen SensorFM encoder with nothing more than a simple linear layer on top beat hand-engineered supervised baselines on 34 of 35 tasks. It was especially strong on depression and anxiety — conditions that vary wildly between individuals and usually get lost in the noise of raw sensor data.

Google then handed the job of building task-specific prediction heads to a swarm of LLM agents, an approach they're calling an agentic classroom, where competing and collaborating models generated and refined over 30,000 candidate solutions. Those agent-built heads outperformed a plain linear probe on 16 of 20 classification tasks and 12 of 15 regression tasks, and unsurprisingly, better underlying LLMs produced better solutions — though weaker models closed some of the gap through collaboration.

The most consequential test, though, involved actual clinicians. Google plugged SensorFM's predictions into a personal health AI agent and had blinded doctors rate 31 real participant summaries across five dimensions, including justifiability and potential for harm. Grounding the agent in SensorFM's inferred metrics improved every rated dimension over a baseline using raw daily wearable data alone. And there was no statistically meaningful difference between summaries built on SensorFM's predictions versus actual ground-truth lab measurements — meaning the model's guesses were, for practical purposes, as good as the real thing.

My take — AI-written commentary, not fact-checked reporting

This is the kind of result that should make people uneasy in the exact same breath it impresses them: a single company now holds a model that can infer depression risk and metabolic health from wristband data on five million people who probably didn't picture their step counts ending up in a foundation model training run. The science is genuinely good — missingness-aware pretraining at this scale is a real advance, not marketing dressing. But Google publishing a research paper isn't the same as Google publishing weights, and until there's an open, auditable version of something like this, the EU and everyone else should treat 'trust us, clinicians reviewed it' as a start, not an answer.

Read more about this at: Google Research

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.