MoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization
Apple ML Research
Researchers developed MoMo, a framework that teaches robots to perform manipulation tasks with different execution styles by learning motion modes as reusable behavioral factors. The system uses spatiotemporal action tokenization and behavior cloning to enable robots to vary joint speed, acceleration, and approach pitch across six real-world tasks. This approach allows robots to generalize learned motion modes to new task combinations, enabling flexible skill execution beyond what was directly demonstrated.
Why it matters
To operate effectively across diverse contexts, robots must not only perform manipulation tasks accurately but also adapt how their actions unfold to the task, object, and interaction setting. We ask whether this execution-level variation can be learned as a reusable behavioral factor shared across tasks. We present MoMo, a two-stage imitation-learning framework consisting of a spatiotemporal action tokenizer and a behavior-cloning transformer that takes task and a continuous motion-mode condition as inputs. Across six real-robot manipulation tasks, varying this condition produces steady…