MeanVC is a lightweight and streaming zero-shot voice conversion system that enables real-time timbre transfer from any source speaker to any target speaker while preserving linguistic content. The system introduces...
Source du modèle
Extrait de la source
MeanVC is a lightweight and streaming zero-shot voice conversion system that enables real-time timbre transfer from any source speaker to any target speaker while preserving linguistic content. The system introduces...
Sources
1 sourceVérifié 29 août
Artefacts du modèle
1 artefactExtraits de sources
2 extraits--- language: - en - zh license: apache-2.0 pipeline_tag: audio-to-audio --- # MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows <div align="center"> [](https://arxiv.org/pdf/2510.08392) [](https://github.com/ASLP-lab/MeanVC) [](https://aslp-lab.github.io/MeanVC/) </div> **MeanVC** is a lightweight and streaming zero-shot voice conversion system that enables real-time timbre transfer from any source speaker to any target speaker while preserving linguistic content. The system introduces a diffusion transformer with a chunk-wise autoregressive denoising strategy and mean flows for efficient single-step inference.  ## ✨ Key Features - **🚀 Streaming Inference**: Real-time voice conversion with chunk-wise processing. - **⚡ Single-Step Generation**: Direct mapping from start to endpoint via mean flows for fast generation. - **🎯 Zero-Shot Capability**: Convert to any unseen target speaker without re-training. - **💾 Lightweight**: Significantly fewer parameters than existing methods. - **🔊 High Fidelity**: Superior speech quality and speaker similarity. ## 💻 Sample Usage ### 1. Environment Setup First, follow these steps to clone the repository and install the required environment. ```bash # Clone the repository and enter the directory git clone https://github.com/ASLP-lab/MeanVC.git cd MeanVC # Create and activate a Conda environment conda create -n meanvc python=3.11 -y conda activate meanvc # Install dependencies pip install -r requirements.txt ``` ### 2. Download Pre-trained Models Run the provided script to automatically download all necessary pre-trained models. ```bash python download_ckpt.py ``` This will download the main VC model, vocoder, and ASR model into the `src/ckpt/` directories. The speaker verification model (`wavlm_large_finetune.pth`) must be downloaded manually from Google Drive. Download the file from [this link](https://drive.google.com/file/d/1-aE1NfzpRCLxA4GUxX9ITI3F9LlbtEGP/view). Place the downloaded `wavlm_large_finetune.pth` file into the `src/runtime/speaker_verification/ckpt/` directory. ### 3. Real-Time Voice Conversion T...
Source context: 29 downloads · 25 likes · Pipeline audio-to-audio · Repo ASLP-lab/MeanVC