Chatterbox is a family of four state-of-the-art, open-source text-to-speech models by Resemble AI.
Fuente del modelo
Extracto de la fuente
Chatterbox is a family of four state-of-the-art, open-source text-to-speech models by Resemble AI.
Fuentes
1 fuenteVerificado 15 ago
Artefactos del modelo
5 artefactosExtractos de fuentes
2 extractoss3gen.safetensors
safetensors · 1008 MB · SHA-256 2b78103c6542…1d4e · Hugging Face
--- license: mit language: - en pipeline_tag: text-to-speech tags: - text-to-speech - speech - speech-generation - voice-cloning ---  # Chatterbox TTS <div style="display: flex; align-items: center; gap: 12px"> <a href="https://resemble-ai.github.io/chatterbox_turbo_demopage/"> <img src="https://img.shields.io/badge/listen-demo_samples-blue" alt="Listen to Demo Samples" /> </a> <a href="https://huggingface.co/spaces/ResembleAI/chatterbox-nano-demo"> <img src="https://huggingface.co/datasets/huggingface/badges/resolve/main/open-in-hf-spaces-sm.svg" alt="Open in HF Spaces" /> </a> <a href="https://podonos.com/resembleai/chatterbox"> <img src="https://static-public.podonos.com/badges/insight-on-pdns-sm-dark.svg" alt="Insight on Podos" /> </a> </div> <div style="display: flex; align-items: center; gap: 8px;"> <span style="font-style: italic;white-space: pre-wrap">Made with ❤️ by</span> <img width="100" alt="resemble-logo-horizontal" src="https://github.com/user-attachments/assets/35cf756b-3506-4943-9c72-c05ddfa4e525" /> </div> **Chatterbox** is a family of four state-of-the-art, open-source text-to-speech models by Resemble AI. We are excited to introduce **Chatterbox-Nano**, our most efficient model yet. Built on a streamlined 110M parameter architecture, **Nano** delivers high-quality speech on CPU - 3x faster than realtime on 8 cores. We have also distilled the speech-token-to-mel decoder, previously a bottleneck, reducing generation from 10 steps to just **one**, while retaining high-fidelity audio output. **Paralinguistic tags** are now native to both Nano and Turbo models, allowing you to use `[cough]`, `[laugh]`, `[chuckle]`, and more to add distinct realism. While Turbo was built primarily for low-latency voice agents, it excels at narration and creative workflows. If you like the model but need to scale or tune it for higher accuracy, check out our competitively priced TTS service (<a href="https://resemble.ai">link</a>). It delivers reliable performance with ultra-low latency of sub 200ms—ideal for production use in agents, applications, or interactive media. <img width="1200" height="600" alt="Podonos Turbo Eval" src="https://storage.googleapis.com/chatterbox-demo-samples/turbo/podonos_turbo.png" /> ### ⚡ Model Zoo Choose...
Source context: 0 downloads · 19 likes · Pipeline text-to-speech · Repo ResembleAI/chatterbox-nano