Voila: Voi ce- La nguage Foundation Models
Source du modèle
Extrait de la source
Voila: Voi ce- La nguage Foundation Models
Sources
1 sourceVérifié 5 sept.
Artefacts du modèle
4 artefactsmodel-00001-of-00004.safetensors
safetensors · 4,59 GB · SHA-256 1039359a8fa2…7c65 · Hugging Face
Téléchargermodel-00002-of-00004.safetensors
safetensors · 4,66 GB · SHA-256 cfd1aa2cee67…e107 · Hugging Face
TéléchargerExtraits de sources
2 extraitsmodel-00003-of-00004.safetensors
safetensors · 4,58 GB · SHA-256 6de5ed66d0dd…72a6 · Hugging Face
model-00004-of-00004.safetensors
safetensors · 1,37 GB · SHA-256 3228f78ac5cb…163b · Hugging Face
Télécharger--- library_name: transformers license: mit datasets: - maitrix-org/Voila-Benchmark - maitrix-org/Voila-million-voice language: - en - zh - fr - de - ja - ko base_model: - maitrix-org/Voila-base pipeline_tag: audio-to-audio tags: - audio - audio-language-model - speech-recognition - audio-understanding - text-to-speech - speech-conversation - end-to-end --- <p align="center"> <img src="https://voila.maitrix.org/static/images/logo.png" width="400"/><br/> <b>Voila: <span style="color:#ca00f9">Voi</span>ce-<span style="color:#ca00f9">La</span>nguage Foundation Models</b><br/><br/> 💜 <a href="https://voila.maitrix.org"><b>Project Page</b></a>    |    🖥️ <a href="https://github.com/maitrix-org/Voila">GitHub</a>    |   🤗 <a href="https://huggingface.co/collections/maitrix-org/voila-67e0d96962c19f221fc73fa5">Hugging Face</a>   |    📑 <a href="http://arxiv.org/abs/2505.02707">Paper</a>    |    🌐 <a href="https://huggingface.co/spaces/maitrix-org/Voila-demo">Online Demo</a>   |    🏠<a href="https://maitrix.org">Maitrix.org</a> </p> Voila is a new family of large voice-language foundation models aiming to lift human-AI interaction experiences to the next level. Breaking away from the constraints of traditional voice AI systems—high latency, loss of vocal nuances, and mechanical responses—Voila employs an innovative end-to-end model design and a novel hierarchical Transformer architecture. This approach enables real-time, autonomous, and rich voice interactions, with latency as low as 195 ms, surpassing average human response times. Combining advanced voice and language modeling, Voila offers customizable, persona-driven engagements and excels in a range of audio tasks from ASR and TTS to speech translation across six languages. With the online [web demo](https://huggingface.co/spaces/maitrix-org/Voila-demo), Voila invites you to explore a transformative, natural dialogue experience between human and AI. # ✨ Highlights - ⭐ High-fidelity, low-latency, real-time streaming audio processing - ⭐ Effective integration of voice and language modeling capabilities - ⭐ Millions of pre-built and custom voices, fast voice switching during conversation - ⭐ Unified model for various audio tasks # 🎥 Video Demo [](https://www.youtu...
Source context: 72 downloads · 54 likes · Pipeline audio-to-audio · Library transformers · Repo maitrix-org/Voila-chat