TLDRocket
Sign in

Transfer learning for genomic prediction in underrepresented populations

Google Research

Google Research found UK Biobank data helps Japanese PRS only when local data is tiny. Once BBJ grows, the European boost can start to hurt accuracy.

Based on reporting by Google Research — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Polygenic risk scores are supposed to turn piles of genetic variants into useful disease predictions. In practice, they’ve been much better trained for European populations than for everyone else, which is one reason they’re still not used much in clinics. Google Research looked at what happens when you try to move those models into Biobank Japan, and the answer is less tidy than the usual “more data is better.”

The team compared UK Biobank, with hundreds of thousands of European participants, against Biobank Japan, a deeply phenotyped cohort of nearly 200,000 Japanese individuals. They tested eight traits measured in both datasets: BMI, systolic and diastolic blood pressure, red blood cell count, white blood cell count, HDL, LDL, and blood glucose. Across those traits, the estimated fraction of variance explained by the genetic variants in UKB ranged from 0.07 to 0.28.

The basic pattern was clear. When the Japanese training set was very small, pulling in UKB data gave PRS models a real lift. But once the target sample size got large enough, local data started winning. For target-population models trained on 15,000 samples or more, the BBJ-only approach generally outperformed the cross-population version. And in traits that were more genetically conserved across populations, that crossover happened later, around 25,000 to 40,000 samples or more. For HDL, LDL and blood glucose, the break came earlier.

That matters because the source of the problem is not just sample size. The study says traits like lipids and blood glucose are more population-specific, so UKB data is farther out of distribution. In those cases, adding European samples can stop helping sooner than expected. The same is true even when the model is more sophisticated. PRS-CSx, which is meant to handle linkage disequilibrium differences across populations, needed more data than elastic net models and only caught up at much larger sample sizes.

The one place where mixing populations still paid off was meta-analysis, especially for the more population-specific traits. For HDL and LDL, and to a lesser extent blood glucose, combining UKB and BBJ GWAS results beat single-population discovery. The larger message is pretty blunt: cross-population transfer helps when local data is scarce, but it is not a permanent shortcut. For underrepresented populations, the real fix is still more local data, plus a model choice that matches the trait and the sample size.

My take — AI-written commentary, not fact-checked reporting

This is the annoying truth the AI crowd keeps skipping over: a bigger model is not the same thing as a better fit. Genetics punishes lazy transfer more brutally than language models do. If the data comes from the wrong population, the elegant wrapper just makes the mistake look expensive.

Read more about this at: Google Research

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.