AllenAI Releases MolmoMotion Vision Model
·2026.06.21 13:26
Key point
AllenAI has unveiled MolmoMotion, a model that predicts the 3D trajectories of objects based on natural language instructions.
Details
AllenAI has released two new vision-language models in the MolmoMotion family. These models predict 3D point trajectories following natural language motion instructions, based on a short history of RGB observations.
Key features of the models include:
- How it works: Given user-specified 2D query points and their 3D history as input, the models predict future 3D positions (in the camera frame, in meters).
- Model configurations: Two versions are provided—one using a 3-frame history and another using a 1-frame history.
- Use cases: Useful for any visual AI application that needs to predict the future positions of objects based on past observation data.
The models have been released on Hugging Face.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.