Molmo learns to point and act
Allen Institute (AI2)
Researchers at AI2 released MolmoPoint, a vision-language model that points by selecting directly from visual inputs rather than generating text coordinates, achieving state-of-the-art results on pointing and tracking benchmarks among open models. The model demonstrated significant improvements in training efficiency and end-task performance, particularly for high-resolution images and cluttered user interfaces. AI2 also released MolmoWeb, a suite of web agents that navigate websites and complete tasks using only screenshots and mouse/keyboard controls, outperforming comparable open models and matching larger proprietary systems like GPT-4o.
Why it matters
MolmoPoint and MolmoWeb extend the Molmo family from visual understanding to visual action, giving researchers open tools for models that can point, navigate, and interact with the world they see.