A GPT-4V Level MLLM for Single Image, Multi Image and Video on Your Phone
Modellquelle
Quellenauszug
A GPT-4V Level MLLM for Single Image, Multi Image and Video on Your Phone
Quellen
1 QuelleVerifiziert 21. Aug.
Modellartefakte
4 Artefaktemodel-00001-of-00004.safetensors
safetensors · 4,54 GB · SHA-256 4510216c25d6…dcf7 · Hugging Face
Herunterladenmodel-00002-of-00004.safetensors
safetensors · 4,59 GB · SHA-256 ffd1b385bb70…f08a · Hugging Face
HerunterladenQuellenauszüge
2 Auszügemodel-00003-of-00004.safetensors
safetensors · 4,03 GB · SHA-256 7d8d8787819b…fdf8 · Hugging Face
Herunterladenmodel-00004-of-00004.safetensors
safetensors · 1,92 GB · SHA-256 e2463f690c7d…95f1 · Hugging Face
Herunterladen--- pipeline_tag: visual-question-answering datasets: - openbmb/RLAIF-V-Dataset --- <h1>A GPT-4V Level MLLM for Single Image, Multi Image and Video on Your Phone</h1> [GitHub](https://github.com/OpenBMB/MiniCPM-V) | [Demo](https://huggingface.co/spaces/openbmb/MiniCPM-V-2_6)</a> ## MiniCPM-V 2.6 **MiniCPM-V 2.6** is the latest and most capable model in the MiniCPM-V series. The model is built on SigLip-400M and Qwen2-7B with a total of 8B parameters. It exhibits a significant performance improvement over MiniCPM-Llama3-V 2.5, and introduces new features for multi-image and video understanding. Notable features of MiniCPM-V 2.6 include: - 🔥 **Leading Performance.** MiniCPM-V 2.6 achieves an average score of 65.2 on the latest version of OpenCompass, a comprehensive evaluation over 8 popular benchmarks. **With only 8B parameters, it surpasses widely used proprietary models like GPT-4o mini, GPT-4V, Gemini 1.5 Pro, and Claude 3.5 Sonnet** for single image understanding. - 🖼️ **Multi Image Understanding and In-context Learning.** MiniCPM-V 2.6 can also perform **conversation and reasoning over multiple images**. It achieves **state-of-the-art performance** on popular multi-image benchmarks such as Mantis-Eval, BLINK, Mathverse mv and Sciverse mv, and also shows promising in-context learning capability. - 🎬 **Video Understanding.** MiniCPM-V 2.6 can also **accept video inputs**, performing conversation and providing dense captions for spatial-temporal information. It outperforms **GPT-4V, Claude 3.5 Sonnet and LLaVA-NeXT-Video-34B** on Video-MME with/without subtitles. - 💪 **Strong OCR Capability and Others.** MiniCPM-V 2.6 can process images with any aspect ratio and up to 1.8 million pixels (e.g., 1344x1344). It achieves **state-of-the-art performance on OCRBench, surpassing proprietary models such as GPT-4o, GPT-4V, and Gemini 1.5 Pro**. Based on the the latest [RLAIF-V](https://github.com/RLHF-V/RLAIF-V/) and [VisCPM](https://github.com/OpenBMB/VisCPM) techniques, it features **trustworthy behaviors**, with significantly lower hallucination rates than GPT-4o and GPT-4V on Object HalBench, and supports **multilingual capabilities** on English, Chinese, German, French, Italian, Korean, etc. - 🚀 **Superior Efficiency.** In addition to its friendly size, MiniCPM-V 2.6 also shows **state-of-the-art token density** (i.e., number of pixels encod...
Source context: 14 downloads · 0 likes · Pipeline visual-question-answering · Repo lei-HuggingFace/MinCPM-V2_6_Level_Image_08162024