qwen-research
Modellquelle
Quellenauszug
qwen-research
Quellen
1 QuelleVerifiziert 27. Sept.
Modellartefakte
4 Artefaktemodel-00001-of-00003.safetensors
safetensors · 4,65 GB · SHA-256 349972cebff4…49a6 · Hugging Face
Herunterladenmodel-00002-of-00003.safetensors
safetensors · 4,66 GB · SHA-256 b29a76dfefb3…6d8c · Hugging Face
HerunterladenQuellenauszüge
2 Auszügemodel-00003-of-00003.safetensors
safetensors · 1,84 GB · SHA-256 8f2138170a87…b170 · Hugging Face
--- license: other license_name: qwen-research license_link: LICENSE language: - en tags: - multimodal library_name: transformers pipeline_tag: any-to-any --- # Qwen2.5-Omni <a href="https://chat.qwen.ai/" target="_blank" style="margin: 2px;"> <img alt="Chat" src="https://img.shields.io/badge/%F0%9F%92%9C%EF%B8%8F%20Qwen%20Chat%20-536af5" style="display: inline-block; vertical-align: middle;"/> </a> ## Overview ### Introduction Qwen2.5-Omni is an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner. <p align="center"> <img src="https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen2.5-Omni/qwen_omni.png" width="80%"/> <p> ### Key Features * **Omni and Novel Architecture**: We propose Thinker-Talker architecture, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner. We propose a novel position embedding, named TMRoPE (Time-aligned Multimodal RoPE), to synchronize the timestamps of video inputs with audio. * **Real-Time Voice and Video Chat**: Architecture designed for fully real-time interactions, supporting chunked input and immediate output. * **Natural and Robust Speech Generation**: Surpassing many existing streaming and non-streaming alternatives, demonstrating superior robustness and naturalness in speech generation. * **Strong Performance Across Modalities**: Exhibiting exceptional performance across all modalities when benchmarked against similarly sized single-modality models. Qwen2.5-Omni outperforms the similarly sized Qwen2-Audio in audio capabilities and achieves comparable performance to Qwen2.5-VL-7B. * **Excellent End-to-End Speech Instruction Following**: Qwen2.5-Omni shows performance in end-to-end speech instruction following that rivals its effectiveness with text inputs, evidenced by benchmarks such as MMLU and GSM8K. ### Model Architecture <p align="center"> <img src="https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen2.5-Omni/overview.png" width="80%"/> <p> ### Performance We conducted a comprehensive evaluation of Qwen2.5-Omni, which demonstrates strong performance across all modalities when compared to similarly sized single-mo...
Source context: 8 downloads · 0 likes · Pipeline any-to-any · Library transformers · Repo Ares-Realm-Studios/Qwen2.5-Omni-3B