A Gemini 2.5 Flash Level MLLM for Vision, Speech, and Full-Duplex Mulitmodal Live Streaming on Your Phone
Fonte do modelo
Trecho da fonte
A Gemini 2.5 Flash Level MLLM for Vision, Speech, and Full-Duplex Mulitmodal Live Streaming on Your Phone
Fontes
1 fonteVerificado 26 de set.
Artefatos de modelo
8 artefatosTrechos de fonte
2 trechosassets/token2wav/hift.pt
pt · 79,5 MB · SHA-256 3386cc880324…9879 · Hugging Face
assets/token2wav/speech_tokenizer_v2_25hz.onnx
onnx · 473 MB · SHA-256 d43342aa1216…6f71 · Hugging Face
Baixarmodel-00001-of-00004.safetensors
safetensors · 4,91 GB · SHA-256 30c40b9a1038…e077 · Hugging Face
Baixarmodel-00002-of-00004.safetensors
safetensors · 4,94 GB · SHA-256 fe0faef420ac…bf96 · Hugging Face
Baixarmodel-00003-of-00004.safetensors
safetensors · 4,94 GB · SHA-256 5d0b20153f9b…7821 · Hugging Face
Baixarmodel-00004-of-00004.safetensors
safetensors · 2,67 GB · SHA-256 f61addf4747c…9b31 · Hugging Face
Baixar--- license: apache-2.0 pipeline_tag: any-to-any library_name: transformers tags: - minicpm-o - minicpm-v - multimodal - full-duplex --- A Gemini 2.5 Flash Level MLLM for Vision, Speech, and Full-Duplex Mulitmodal Live Streaming on Your Phone [GitHub](https://github.com/OpenBMB/MiniCPM-o) | [CookBook](https://github.com/OpenSQZ/MiniCPM-V-CookBook) | [Omni-modal Demo](https://openbmb.github.io/MiniCPM-o-Demo/) | [Vision-Language Demo](http://211.93.21.133:18121/) </br> [WeChat](https://github.com/OpenBMB/MiniCPM-o/blob/main/docs/wechat.md) | [Discord](https://discord.gg/N2RnxGdJ) | CaseBook([Audio](https://openbmb.github.io/minicpm-o-4_5/), [Omni Full-Duplex](https://openbmb.github.io/minicpm-o-4_5-omni/)) ## News > [!NOTE] > [2026.02.06] 🥳 🥳 🥳 We open-sourced a realtime web demo deployable on your own devices like Mac or GPU. [Try it now](#deploy-a-realtime-web-demo-on-your-own-device)! ## MiniCPM-o 4.5 **MiniCPM-o 4.5** is the latest and most capable model in the MiniCPM-o series. The model is built in an end-to-end fashion based on SigLip2, Whisper-medium, CosyVoice2, and Qwen3-8B with a total of 9B parameters. It exhibits a significant performance improvement, and introduces new features for full-duplex multimodal live streaming. Notable features of MiniCPM-o 4.5 include: - 🔥 **Leading Visual Capability.** MiniCPM-o 4.5 achieves an average score of 77.6 on OpenCompass, a comprehensive evaluation of 8 popular benchmarks. **With only 9B parameters, it surpasses widely used proprietary models like GPT-4o, Gemini 2.0 Pro, and approaches Gemini 2.5 Flash** for vision-language capabilities. It supports instruct and thinking modes in a single model, better covering efficiency and performance trade-offs in different user scenarios. - 🎙 **Strong Speech Capability.** MiniCPM-o 4.5 supports **bilingual real-time speech conversation with configurable voices** in English and Chinese. It features **more natural, expressive and stable speech conversation**. The model also allows for fun features such as **voice cloning and role play via a simple reference audio clip**, where the cloning performance surpasses strong TTS tools such as CosyVoice2. - 🎬 **New Full-Duplex and Proactive Multimodal Live Streaming Capability.** As a new feature, MiniCPM-o 4.5 can process real-time, continuous video and audio input streams simultaneously while generating conc...
Source context: 12 downloads · 0 likes · Pipeline any-to-any · Library transformers · Repo rycerzes/MiniCPM-o-4_5