TLDRocket
Sign in

Show Me Examples: Inferring Visual Concepts from Image Sets

Apple

Apple researchers built a test showing AI image models can't figure out what a set of pictures has in common and copy that idea onto a new photo. Turns out today's top vision-language models just guess or ignore the visual clues entirely, so Apple built a fix.

Based on reporting by Apple — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Ask a person to look at five photos of, say, red vintage cars and then apply that same 'red vintage car' vibe to a new image, and they'll manage it without much thought. Ask a state-of-the-art vision-language model to do the same thing, and it mostly falls apart. That's the blunt finding behind a new paper out of Apple's ML research group, written with LMU Munich, which introduces a benchmark called VICIS — Visual Concept Inference from Sets.

The setup is simple to describe and apparently very hard to solve. You give a model a small context set of images that all share some concept, plus a separate query image, and ask it to generate new images that honor the shared concept while staying true to the query. No text prompt spelling out what the concept is — the model has to infer it purely from looking at the examples. Researchers Nick Stracke, Kolja Bauer, Josh Susskind, Miguel Angel Bautista Martin and Björn Ommer found that today's leading VLMs frequently just ignore the context images altogether, or fall back on generic, biased outputs that have little to do with what was actually shown.

To fix this, the team built a training framework and matching architecture specifically designed to extract concept embeddings from a set of example images and carry that concept into a new generation. They tested it against synthetic datasets and against large-scale real-world data pulled from ImageNet and WordNet, and the custom model outperformed existing VLMs on accuracy and on output diversity — meaning it didn't just latch onto one narrow interpretation of the concept. Notably, it also generalized to concepts and image types it hadn't seen during training, including sketches, which suggests the model is picking up something closer to an abstract notion of 'concept' rather than memorizing surface patterns.

The timing lines up with a broader theme in Apple's recent research output. A related post from the same research page mentions work on identifying 'expert units' inside transformer models — individual neurons that reliably encode specific concepts — using a dataset of over 1,600 concepts derived from labeled sentence sets. Another nearby project tackles the classic sim-to-real gap by refining synthetic training images so they look more realistic, cutting down on the need for expensive manual labeling. Different problems, same underlying itch: getting models to understand structure and meaning without needing everything spelled out in text or hand-annotated data.

My take — AI-written commentary, not fact-checked reporting

This is the kind of gap that gets glossed over in the current hype cycle, where everyone's obsessed with chat-based reasoning and forgets that pure visual pattern-matching — the thing toddlers do effortlessly — is still a genuinely unsolved problem for these systems. I'd rather see Apple publish honest benchmarks showing where VLMs fail than another cherry-picked demo, and if this pushes the field toward models that actually reason from pixels instead of defaulting to text-shaped shortcuts, that's a real contribution, not just an incremental leaderboard bump.

Read more about this at: Apple

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.