microsoft/VibeVoice-Realtime-0.5B with the missing acoustic encoder added — enabling voice cloning from your own audio.
Modellquelle
Quellenbeschreibung
microsoft/VibeVoice-Realtime-0.5B with the missing acoustic encoder added — enabling voice cloning from your own audio.
pip install "transformers==4.51.3" torch soundfile
pip install git+https://github.com/microsoft/VibeVoice
# get the scripts (the model itself downloads automatically on first run)
huggingface-cli download mohammed-bahumaish/vibevoice-realtime-0.5b-with-encoder \
make_voice_prompt.py run_tts.py --local-dir .
# 1) build a voice prompt from ~15-30s of reference audio
python make_voice_prompt.py \
--voice_wav my_voice.wav \
--transcript "exact transcript of the reference audio" \
--output my_voice.pt
# 2) speak anything in that voice
python run_tts.py \
--voice_pt my_voice.pt \
--text "Hello! This works with the stock Microsoft inference code." \
--output out.wav
Quellen
1 QuelleVerifiziert 19. Sept.
Modellartefakte
1 ArtefaktQuellenauszüge
2 AuszügeThe .pt files are drop-in compatible with Microsoft's own demos, like the
prebaked demo/voices/streaming_model/*.pt voices.
transformers must be 4.51.x — 5.x silently breaks the model.--transcript explicitly for best results (auto-transcription is English-only).MIT. Base model by Microsoft; its responsible-use guidelines apply — clone only voices you have the right to use.
--- license: mit library_name: transformers base_model: microsoft/VibeVoice-Realtime-0.5B base_model_relation: finetune language: - en pipeline_tag: text-to-speech tags: - vibevoice - voice-cloning - tts - streaming --- # VibeVoice-Realtime-0.5B — with encoder (voice cloning) [microsoft/VibeVoice-Realtime-0.5B](https://huggingface.co/microsoft/VibeVoice-Realtime-0.5B) with the missing acoustic encoder added — enabling voice cloning from your own audio. ## Usage ```bash pip install "transformers==4.51.3" torch soundfile pip install git+https://github.com/microsoft/VibeVoice # get the scripts (the model itself downloads automatically on first run) huggingface-cli download mohammed-bahumaish/vibevoice-realtime-0.5b-with-encoder \ make_voice_prompt.py run_tts.py --local-dir . # 1) build a voice prompt from ~15-30s of reference audio python make_voice_prompt.py \ --voice_wav my_voice.wav \ --transcript "exact transcript of the reference audio" \ --output my_voice.pt # 2) speak anything in that voice python run_tts.py \ --voice_pt my_voice.pt \ --text "Hello! This works with the stock Microsoft inference code." \ --output out.wav ``` The `.pt` files are drop-in compatible with Microsoft's own demos, like the prebaked `demo/voices/streaming_model/*.pt` voices. ## Tips - `transformers` must be **4.51.x** — 5.x silently breaks the model. - Use a true 24 kHz+ recording, ≥ 15 s, clean single speaker. - Pass `--transcript` explicitly for best results (auto-transcription is English-only). ## License MIT. Base model by [Microsoft](https://huggingface.co/microsoft/VibeVoice-Realtime-0.5B); its responsible-use guidelines apply — clone only voices you have the right to use.
Source context: 460 downloads · 4 likes · Pipeline text-to-speech · Library transformers · Repo mohammed-bahumaish/vibevoice-realtime-0.5b-with-encoder