Chatterbox is a family of four state-of-the-art, open-source text-to-speech models by Resemble AI.
Model source
Source excerpt
Chatterbox is a family of four state-of-the-art, open-source text-to-speech models by Resemble AI.
Sources
1 sourceVerified Aug 15
Model artifacts
5 artifactsSource excerpts
2 excerptss3gen.safetensors
safetensors · 1,008 MB · SHA-256 2b78103c6542…1d4e · Hugging Face
--- license: mit language: - en pipeline_tag: text-to-speech tags: - text-to-speech - speech - speech-generation - voice-cloning ---  # Chatterbox TTS <div style="display: flex; align-items: center; gap: 12px"> <a href="https://resemble-ai.github.io/chatterbox_turbo_demopage/"> <img src="https://img.shields.io/badge/listen-demo_samples-blue" alt="Listen to Demo Samples" /> </a> <a href="https://huggingface.co/spaces/ResembleAI/chatterbox-nano-demo"> <img src="https://huggingface.co/datasets/huggingface/badges/resolve/main/open-in-hf-spaces-sm.svg" alt="Open in HF Spaces" /> </a> <a href="https://podonos.com/resembleai/chatterbox"> <img src="https://static-public.podonos.com/badges/insight-on-pdns-sm-dark.svg" alt="Insight on Podos" /> </a> </div> <div style="display: flex; align-items: center; gap: 8px;"> <span style="font-style: italic;white-space: pre-wrap">Made with ❤️ by</span> <img width="100" alt="resemble-logo-horizontal" src="https://github.com/user-attachments/assets/35cf756b-3506-4943-9c72-c05ddfa4e525" /> </div> **Chatterbox** is a family of four state-of-the-art, open-source text-to-speech models by Resemble AI. We are excited to introduce **Chatterbox-Nano**, our most efficient model yet. Built on a streamlined 110M parameter architecture, **Nano** delivers high-quality speech on CPU - 3x faster than realtime on 8 cores. We have also distilled the speech-token-to-mel decoder, previously a bottleneck, reducing generation from 10 steps to just **one**, while retaining high-fidelity audio output. **Paralinguistic tags** are now native to both Nano and Turbo models, allowing you to use `[cough]`, `[laugh]`, `[chuckle]`, and more to add distinct realism. While Turbo was built primarily for low-latency voice agents, it excels at narration and creative workflows. If you like the model but need to scale or tune it for higher accuracy, check out our competitively priced TTS service (<a href="https://resemble.ai">link</a>). It delivers reliable performance with ultra-low latency of sub 200ms—ideal for production use in agents, applications, or interactive media. <img width="1200" height="600" alt="Podonos Turbo Eval" src="https://storage.googleapis.com/chatterbox-demo-samples/turbo/podonos_turbo.png" /> ### ⚡ Model Zoo Choose...
Source context: 0 downloads · 19 likes · Pipeline text-to-speech · Repo ResembleAI/chatterbox-nano