MolmoMotion: Language-guided 3D motion forecasting
Allen Institute (AI2)
Researchers released MolmoMotion, a model that predicts how 3D points on objects will move in the future based on video frames and language instructions describing actions. The model achieved 0.109 meters average displacement error on the PointMotionBench benchmark, outperforming existing forecasting methods. The model and accompanying datasets enable applications in robot manipulation planning and controllable video generation.
Why it matters
MolmoMotion is an open, language-guided 3D motion forecasting model that predicts how object points will move in the future, enabling stronger motion prediction for robotics, video generation, and other systems that need to reason about what happens next.