TLDRocket
Sign in

Teaching Gemini to spot exploding stars with just a few examples

Google Research

Google taught Gemini to spot fake vs real exploding stars using just 15 examples per telescope, not millions of images. It matches specialist AI accuracy and actually explains its reasoning instead of just saying yes or no.

Based on reporting by Google Research — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Astronomy has a spam problem. Every night, sky surveys spit out millions of alerts flagging things that might be supernovae, and almost all of them are junk: satellite trails, cosmic ray hits, camera glitches dressed up as cosmic drama. For a decade the fix has been specialized convolutional neural networks trained on mountains of labeled images. They work, but they're mute. You get a real-or-bogus label and nothing else, which means astronomers either trust a black box or burn hours double-checking it by hand.

Google Research decided to test whether a general-purpose model could do better, and not by brute force. In a paper published in Nature Astronomy, the team fed Gemini just 15 annotated examples per survey — Pan-STARRS, MeerLICHT, and ATLAS — each one a trio of images showing a new observation, a reference shot, and the difference between them, plus a short expert note and an interest score. No massive training set, no months of curation. Just few-shot learning and a tight set of instructions.

The results landed at 93 percent average accuracy across all three surveys, roughly matching purpose-built CNNs that needed far more data to get there. But accuracy wasn't really the interesting part. Gemini also wrote out its reasoning for each classification and assigned an interest score to help astronomers decide what deserves follow-up telescope time. A panel of 12 professional astronomers graded 200 of these explanations on a 0-to-5 coherence scale and found them consistently aligned with expert logic, not just plausible-sounding filler.

The bigger surprise was that Gemini could flag its own uncertainty. Low self-rated coherence scores turned out to reliably predict wrong answers, which means the system can tell researchers when to double-check its work instead of quietly failing. Feed those flagged failures back into the prompt as new examples, and accuracy climbs — on MeerLICHT, from about 93.4 percent up to 96.7 percent, all without retraining anything.

That matters because the next generation of instruments, like the Vera C. Rubin Observatory, is expected to generate 10 million alerts a night, a volume no team of humans could manually sort through. Google's pitch here is that this few-shot, self-explaining approach could be ported to other instruments and other sciences with minimal setup, turning Gemini into something closer to a lab assistant that knows when to ask for help rather than a black box that never does.

My take — AI-written commentary, not fact-checked reporting

This is the most convincing case yet for treating big multimodal models as reasoning partners rather than pattern-matching machines, and the self-flagging uncertainty trick is the part everyone should be stealing for other domains, not the headline accuracy number. I'd rather trust a system that admits when it's confused than one that's marginally more accurate but silent about its failure modes, and if this scales to Rubin's ten-million-alert nights, it's a genuinely useful application rather than another chatbot demo dressed up as science.

Read more about this at: Google Research

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.