TLDRocket
Sign in

Teaching AI to see the world more like we do

Google DeepMind

DeepMind found AI vision models often see the world weirdly compared to humans, then fixed it. Their new method makes models understand concepts like we do, and it also makes them better at other tasks.

Based on reporting by Google DeepMind — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Google DeepMind's latest paper, published in Nature, tackles a problem that anyone who's used image-recognition tools has probably run into without realizing it: these systems don't organize what they see the way we do. A model might nail the make and model of two hundred different cars, then completely miss that a car and an airplane share something basic — they're both big metal vehicles. That gap between machine perception and human intuition is what researchers Andrew Lampinen and Klaus Greff, along with lead author Lukas Muttenthaler and a team of collaborators, set out to measure and correct.

Their method borrows a classic tool from cognitive science called the odd-one-out task: show three images, ask which one doesn't belong. Humans and AI models agree easily on some triplets — a birthday cake stands out next to a tapir and a sheep, no argument there. But feed both a starfish, a cat, and something else visually similar in texture or color, and models frequently latch onto superficial cues like background shading instead of the deeper category humans instinctively grasp. Plotted out, an unaligned model's internal map of the world looks like a mess, animals and furniture and food all tangled together with no discernible logic.

The fix required some cleverness because the best human-judgment dataset available, called THINGS, only contains a few thousand images — nowhere near enough to fine-tune a large vision model without it overfitting and forgetting everything else it knew. So the team built a three-stage pipeline. First, they trained a small adapter on top of a frozen, pretrained model (SigLIP-SO400M) using THINGS, creating a

Read more about this at: Google DeepMind

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.