Motion-O is a family of Qwen2.5-VL models fine-tuned for motion-aware trajectory reasoning in videos. This work is introduced in the paper Motion-o: Trajectory-Grounded Video Reasoning.
Fuente del modelo
Extracto de la fuente
Motion-O is a family of Qwen2.5-VL models fine-tuned for motion-aware trajectory reasoning in videos. This work is introduced in the paper Motion-o: Trajectory-Grounded Video Reasoning.
Fuentes
1 fuenteVerificado 6 ago
Artefactos del modelo
1 artefactoExtractos de fuentes
2 extractos--- base_model: - Qwen/Qwen2.5-VL-7B-Instruct datasets: - lmms-lab/Video-MME - OpenGVLab/MVBench - bishoygaloaa/Motion-o-MCoT-PLM-motion-keyframes license: mit pipeline_tag: video-text-to-text library_name: transformers --- # Motion-O: Motion-Aware Trajectory Reasoning for Video Motion-O is a family of Qwen2.5-VL models fine-tuned for **motion-aware trajectory reasoning** in videos. This work is introduced in the paper [Motion-o: Trajectory-Grounded Video Reasoning](https://huggingface.co/papers/2603.18856). The models learn to produce structured `<think>...</think>` chains with `<obj>`, `<box>`, `<t>`, and `<motion>` tags that describe object motion over time, and to answer a final question about the video. **Links:** - [Project Page](https://ostadabbas.github.io/motion-o.github.io/) - [GitHub Repository](https://github.com/ostadabbas/Motion-o) - [arXiv Paper](https://arxiv.org/abs/2603.18856) ### Available variants All variants live in this repository as subfolders: - **`(root)`** – `grpo_dense_t07_4737145/checkpoint-800/merged` - **Name**: Motion-O (no visual grounding) - **Description**: GRPO on STGR with motion-aware rewards; no explicit open-o3 visual grounding. - **`open-o3-mcot`** – `open-o3_grpo_v3_074917638/checkpoint-600/merged` - **Name**: Open-o3 + MCoT (with visual grounding) - **Description**: Open-o3-Video style model with multi-chain-of-thought and explicit visual grounding. - **`open-o3-mcot-no-vg`** – `open-o3_grpo_v2_4896760/checkpoint-1000/merged` - **Name**: Open-o3 + MCoT (no visual grounding) - **Description**: Same training recipe as above but without the additional visual-grounding objective. ### How to load ```python from transformers import AutoModelForCausalLM, AutoProcessor # 1) Motion-O (no visual grounding) – repo root model = AutoModelForCausalLM.from_pretrained( "bishoygaloaa/motion-o", torch_dtype="auto", ) processor = AutoProcessor.from_pretrained("bishoygaloaa/motion-o") # 2) Open-o3 + MCoT (with visual grounding) model_vg = AutoModelForCausalLM.from_pretrained( "bishoygaloaa/motion-o", subfolder="open-o3-mcot", torch_dtype="auto", ) processor_vg = AutoProcessor.from_pretrained( "bishoygaloaa/motion-o", subfolder="open-o3-mcot", ) # 3) Open-o3 + MCoT (no visual grounding) model_no_vg = AutoModelForCausalLM.from_pretrained( "bishoygaloaa/motion-o", subfolder="open-o3-mcot-no-vg", torch_dtype=...
Source context: 0 downloads · 2 likes · Pipeline video-text-to-text · Library transformers · Repo bishoygaloaa/motion-o