Motus
AI Video CreationOpen Source AI & Machine Learning

Motus

Official code of Motus: A Unified Latent Action World Model

Open Source

About

Motus is a unified latent action world model that leverages existing pretrained models and rich, sharable motion information. Motus introduces a Mixture-of-Transformers (MoT) architecture to integrate three experts (understanding, action, and video generation) and adopts a UniDiffuser-style scheduler to enable flexible switching between different modeling modes (World Models, Vision-Language-Action Models, Inverse Dynamics Models, Video Generation Models, and Video-Action Joint Prediction Models). Motus further leverages optical flow to learn latent actions and adopts a three-phase training pipeline and six-layer data pyramid, thereby extracting pixel-level "delta action" and enabling large-scale action pretraining.

Open Source Health

Not enough history
Stars
1,277
Forks
74
License
Apache-2.0
Last commit
9 months ago
Python

Related Categories