Efficient Kinematics Generation via Long-Horizon Motion Embedding
Key point
It generates long motions using a motion embedding compressed 64x from large-scale trajectories.
Details
It collects large-scale trajectory data with a tracker model to learn a long-horizon motion embedding. This representation is a highly compressed space with 64x temporal compression applied, within which a conditional flow-matching model generates motion latents.
The generation is conditioned on text prompts and spatial pokes. It is designed to handle long-horizon motion far more efficiently without directly running full video synthesis, while producing realistic motion over long durations.
The researchers explain that this approach yields better motion distribution than state-of-the-art video models and dedicated task-specific approaches.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.