We evaluate LongVPO-Stage2-InternVL3-7B on various video understanding benchmarks, comparing it with the baseline InternVL3-8B.
Modellquelle
Quellenauszug
We evaluate LongVPO-Stage2-InternVL3-7B on various video understanding benchmarks, comparing it with the baseline InternVL3-8B.
Quellen
1 QuelleVerifiziert 6. Aug.
Modellartefakte
4 Artefaktemodel-00001-of-00004.safetensors
safetensors · 4,65 GB · SHA-256 04fb76b820ca…3cab · Hugging Face
Herunterladenmodel-00002-of-00004.safetensors
safetensors · 4,62 GB · SHA-256 7880f855dde7…35c0 · Hugging Face
HerunterladenQuellenauszüge
2 Auszügemodel-00003-of-00004.safetensors
safetensors · 4,47 GB · SHA-256 fcc11656bc04…ee09 · Hugging Face
model-00004-of-00004.safetensors
safetensors · 1,06 GB · SHA-256 629964a1fe42…4daf · Hugging Face
Herunterladen--- license: mit language: - en metrics: - accuracy base_model: - OpenGVLab/InternVL3-8B pipeline_tag: video-text-to-text library_name: transformers tags: - multimodal model-index: - name: LongVPO-Stage2-InternVL3-7B results: - task: type: multimodal dataset: name: MLVU type: mlvu metrics: - type: accuracy value: 76.4 name: accuracy verified: true - task: type: multimodal dataset: name: LongVideoBench type: longvideobench metrics: - type: accuracy value: 66.0 name: accuracy verified: true - task: type: multimodal dataset: name: LVBench type: lvbench metrics: - type: accuracy value: 53.6 name: accuracy verified: true - task: type: multimodal dataset: name: Video-MME w/o sub type: video-mme metrics: - type: accuracy value: 68.9 name: accuracy verified: true - task: type: multimodal dataset: name: Video-MME w/ sub type: video-mme metrics: - type: accuracy value: 74.0 name: accuracy verified: true - task: type: multimodal dataset: name: MVBench type: mvbench metrics: - type: accuracy value: 75.0 name: accuracy verified: true --- # LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization [\[📂 GitHub\]](https://github.com/MCG-NJU/LongVPO) [\[📜 Paper\]](https://arxiv.org/abs/2602.02341) ## 🏆 Performance We evaluate **LongVPO-Stage2-InternVL3-7B** on various video understanding benchmarks, comparing it with the baseline **InternVL3-8B**. LongVPO achieves **significant improvements** on long-video benchmarks (MLVU, LongVideoBench, LVBench, Video-MME) while maintaining competitive performance on short-video tasks (MVBench). The results demonstrate the progressive improvements from the baseline to Stage 1 (Anchored Cues) and finally to Stage 2 (Self-Reasoning). | Benchmark | Type | InternVL3-8B (Base) | **LongVPO-InternVL3-8B (Stage 1)** | **LongVPO-InternVL3-8B (Stage 2)** | | :--- | :---: | :---: | :---: | :---: | | **MLVU** | Long Video | 71.4 | 75.1 | **76.4** | | **LongVideoBench** | Long Video | 62.3 | 66.8 | **66.0** | | **LVBench** | Long Video | 48.8 | 52.4 | **53.6** | | **Video-MME** (w/o sub) | Long Video | 66.5 | 68.1 | **68.9** | | **Video-MME** (w/ sub) | Long Video | 72.5 | 74.0 | **74.0** | | **MVBench** | Short Video | 75.4 | 75.1 | 75.0 | ## 🚀 Quick Start > [!IMPORTANT] > **Please use `transformers>=4.37.2` to ensure the model works normall...
Source context: 7 downloads · 0 likes · Pipeline video-text-to-text · Library transformers · Repo MCG-NJU/LongVPO-Stage2-InternVL3-8B