WordVoice is a novel speech generation framework that transforms traditional implicit end-to-end TTS generation into an explicit, highly controllable paradigm. Built on top of the CosyVoice3 framework, it enables...
Fuente del modelo
Extracto de la fuente
WordVoice is a novel speech generation framework that transforms traditional implicit end-to-end TTS generation into an explicit, highly controllable paradigm. Built on top of the CosyVoice3 framework, it enables...
Fuentes
1 fuenteVerificado 3 ago
Artefactos del modelo
3 artefactosExtractos de fuentes
2 extractoswordvoice_llm_zh.pt
pt · 1,89 GB · SHA-256 3d88acbce69b…63eb · Hugging Face
--- base_model: - FunAudioLLM/Fun-CosyVoice3-0.5B-2512 datasets: - XXH333/WordVoice-5A language: - zh - en license: apache-2.0 pipeline_tag: text-to-speech --- # WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS WordVoice is a novel speech generation framework that transforms traditional implicit end-to-end TTS generation into an explicit, highly controllable paradigm. Built on top of the CosyVoice3 framework, it enables precise and decoupled word-level control over five acoustic dimensions. - **Paper:** [WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS](https://huggingface.co/papers/2607.06461) - **GitHub Repository:** [XXH333/WordVoice-main](https://github.com/XXH333/WordVoice-main) - **Project Page & Demo:** [WordVoice Demo](https://xxh333.github.io/wordvoice-demo/) - **Interactive Online Demo:** Try WordVoice on [Hugging Face Spaces](https://huggingface.co/spaces/XXH333/wordvoice-tts) *Thanks to the Hugging Face team for creating and transferring this Space to us.* --- ## ✨ Features ### 🎯 Explicit Word-Level Control Supports independent and decoupled control of five acoustic attributes for each input word: - ⏱️ **Duration**: Word-level pronunciation duration. - ⏸️ **Boundary**: 5-level pause classification (`b0`–`b4`). - 🔊 **Energy**: Word-level volume/loudness (`0`–`1`). - 🎵 **Pitch**: Word-level core fundamental frequency (`-1`–`1`). - 📈 **Tone**: 7 categories of prosodic morphologies (flat, rise, strong rise, fall, strong fall, peak, valley). ### 🧠 "Acoustic Thinking" Mechanism via Bound-Token Employs a `bound-token` (`<b>`) mechanism within the autoregressive (AR) language model. Before generating the speech tokens for a specific word, the model explicitly predicts its acoustic attributes, realizing an intelligent process of "planning prosody first, then generating sound." --- ## 🛠️ Quick Start ### Installation We recommend using **Conda** to manage your Python environment. ```bash conda create -n wordvoice python=3.10 -y conda activate wordvoice git clone https://github.com/XXH333/WordVoice-main.git cd WordVoice-main pip install -e . pip install num2words==0.5.14 x_transformers==2.11.24 ``` ### Download Model Weights Run the following script to automatically download the pre-trained weights and dependencies (such as CosyVoice3, MMS-FA, etc.): ``...
Source context: 0 downloads · 11 likes · Pipeline text-to-speech · Repo XXH333/WordVoice-base-0.5B