MoMo: Controlling Robot Manipulation Motion Modes via Spatiotemporal Action Tokenization
Key point
We propose MoMo, a framework that allows robots to adjust their motion style according to the task environment through spatiotemporal action tokenization.
Details
For robots to operate effectively across diverse environments, they must not only perform tasks accurately but also adapt their motion style to the target object and interaction setting. MoMo is a two-stage Imitation Learning framework that learns such execution-level variability as a reusable behavior factor shared across multiple tasks.
MoMo consists of the following key components:
- Spatiotemporal Action Tokenizer: Performs spatiotemporal action tokenization
- Behavior-Cloning Transformer: Generates motion by taking the Task and a continuous Motion-mode condition as input
Through experiments on 6 real-world robot manipulation tasks, we demonstrate that simply changing the motion-mode condition allows the robot's motion style to be reliably adjusted.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.