TLDRocket
Sign in

MoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization

Apple ML Research

Researchers developed MoMo, a framework that teaches robots to perform manipulation tasks with different execution styles by learning motion modes as reusable behavioral factors. The system uses spatiotemporal action tokenization and behavior cloning to enable robots to vary joint speed, acceleration, and approach pitch across six real-world tasks. This approach allows robots to generalize learned motion modes to new task combinations, enabling flexible skill execution beyond what was directly demonstrated.

Why it matters

To operate effectively across diverse contexts, robots must not only perform manipulation tasks accurately but also adapt how their actions unfold to the task, object, and interaction setting. We ask whether this execution-level variation can be learned as a reusable behavioral factor shared across tasks. We present MoMo, a two-stage imitation-learning framework consisting of a spatiotemporal action tokenizer and a behavior-cloning transformer that takes task and a continuous motion-mode condition as inputs. Across six real-robot manipulation tasks, varying this condition produces steady…

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.