TLDRocket
Sign in

AI Still Sees Like a Toddler

YouTube

AI can caption a photo but still can't reason about a wiring diagram or a floor plan. Elorian's CEO says vision models are stuck at toddler-level logic, and that's a bigger deal than it sounds.

Based on reporting by YouTube — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Andrew Dai runs Elorian, a startup betting that the next real bottleneck in AI isn't language, it's sight. Ask a modern model what's in a photo and it'll nail it: a dog on a couch, a red Honda, a cluttered desk. Ask it to trace which cord plugs into which port, or to figure out if a robot arm will clip a shelf mid-swing, and the wheels come off fast.

Dai's point is that description and reasoning are two different skills, and AI vision systems mastered the first while barely touching the second. A toddler can point at a cat and say cat long before she can explain how gears mesh or predict where a ball will land after bouncing off a wall. Current models are stuck in that same phase: fluent at naming, clumsy at inferring structure, sequence, and cause and effect from a static image.

That gap matters more than a parlor trick failing. Engineering drawings, satellite imagery, product schematics, floor plans, the guts of a server rack, these are all visual puzzles that demand spatial logic, not just object recognition. A model that can identify a resistor but not trace the circuit it's part of isn't much use to an electrical engineer. Same problem shows up in robotics, where a machine has to understand not just what it sees but what will happen next if it moves.

Elorian's pitch, as Dai frames it, is that whoever cracks visual reasoning first unlocks a wave of automation in fields that have mostly sat outside the generative AI boom: industrial design, remote sensing, physical robotics, even architecture. Text and code got the flashy early wins because language is linear and well-labeled. Vision is messier, spatial, and full of implicit relationships that don't show up neatly in training data.

So the interesting part isn't that AI struggles with tangled cords. It's that the struggle points to a specific, fixable gap between recognition and reasoning, and whoever closes it stands to matter a lot more than the next chatbot upgrade.

My take — AI-written commentary, not fact-checked reporting

I'll believe the 'toddler vision' framing is real progress when someone ships a model that can debug a messy breadboard photo without hand-holding, not another benchmark screenshot. Everyone's chasing bigger context windows and flashier demos while the actually useful problem, teaching machines to understand physical space, sits underfunded and unsexy. That's exactly the kind of boring, hard problem that tends to make someone very rich a few years late.

Read more about this at: YouTube

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.