TLDRocket
Sign in

Molmo learns to point and act

Allen Institute (AI2)

AI2 just shipped MolmoPoint and MolmoWeb, two open AI models that point at pixels and click through websites on their own. The web agent beats GPT-4o on some benchmarks despite being smaller and open source — that's the surprising bit.

Based on reporting by Allen Institute (AI2) — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Ai2 has spent the last few years building an open alternative to the big closed vision-language models, and this week's Molmo update shows why that bet is starting to pay dividends. Two new releases, MolmoPoint and MolmoWeb, tackle a problem that sounds simple but has quietly stumped a lot of labs: getting a model to point at exactly the right pixel and then act on it.

Pointing turns out to be deceptively hard. Most VLMs guess at coordinates as text, which is brittle and falls apart on cluttered screens full of tiny buttons. MolmoPoint, released in March, ditches that approach entirely. It picks a coarse region first, then narrows in on the exact spot, treating pointing as a cross-modal selection problem rather than a coordinate-guessing exercise. Chris Clark, who leads the Molmo research effort, says the jump in training efficiency and end-task accuracy surprised even his own team. The model now leads open-weight benchmarks on pointing, screen-element identification, and object tracking, and Ai2 released the datasets behind it — thousands of annotated screenshots and human-labeled tracks — so anyone can retrain their own version.

MolmoWeb takes that pointing skill and puts it to work in a browser. Instead of parsing HTML or accessibility trees, it looks at a screenshot, the same interface a human uses, and decides what to click or type next. Tanmay Gupta, who leads the project, frames this as a deliberate bet: screenshots survive website redesigns that would break code-based scrapers, and a single image is cheaper to process than thousands of lines of markup. The team started small, targeting just 20 websites with a handful of templated tasks, before scaling up through 2025 and into this year. Getting reliable evaluation was the harder problem — a single bad click early in a task can wreck everything downstream, so much of the work was forensic, tracing failures back through data generation and training.

The results are notable less for the round number of benchmarks won and more for who's already using the underlying tech. Harvard and Broad Institute researchers built an animal-behavior tracking agent on Molmo's pointing without retraining anything. A University of Edinburgh team folded it into an AI oversight framework where its descriptions help catch reasoning errors. Trento researchers used the open training pipeline to study how these models represent space. None of that happens with a closed API you can only prompt.

Ai2 is positioning this as one ecosystem rather than isolated demos — MolmoBot and MolmoSpaces for robotics, WildDet3D for 3D scene understanding, now pointing and web browsing joining the stack. Everything shares the same open weights, code, and data philosophy. Whether that adds up to something that outpaces the well-funded closed labs long-term is an open question, but the near-term evidence — beating GPT-4o-based agents with a smaller, fully public model — is a real point in favor of the approach.

My take — AI-written commentary, not fact-checked reporting

I've been skeptical of the constant drumbeat of 'open beats closed' claims, but this one has receipts: a smaller open model outperforming GPT-4o-based agents on real browsing tasks isn't marketing, it's a benchmark table. What I actually care about is the compounding effect — Harvard, Edinburgh, and Trento teams building on Molmo's weights within months of release is the whole argument for open models in one sentence, and no amount of API access to a closed frontier model replicates that. The EU keeps debating how to build 'sovereign AI'; funding and adopting infrastructure like this, rather than another sovereign cloud initiative nobody uses, would actually move the needle.

Read more about this at: Allen Institute (AI2)

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.