GGUF / ggml conversions of Qwen/Qwen3-TTS-Tokenizer-12Hz for use with the qwen3-tts backend in CrispStrobe/CrispASR.
Fonte do modelo
Descrição da fonte
GGUF / ggml conversions of Qwen/Qwen3-TTS-Tokenizer-12Hz for use with the qwen3-tts backend in CrispStrobe/CrispASR.
Qwen3-TTS-Tokenizer-12Hz is the separate used by the Qwen3-TTS family:
Fontes
1 fonteVerificado 2 de ago.
Artefatos de modelo
2 artefatosTrechos de fonte
2 trechoszh en ja ko de fr ru pt es itThis repo contains the tokenizer / codec only. Use it together with the talker GGUF from cstr/qwen3-tts-0.6b-base-GGUF.
| File | Size | Notes |
|---|---|---|
qwen3-tts-tokenizer-12hz.gguf | 342 MB | F16 |
qwen3-tts-tokenizer-12hz-q8_0.gguf | 277 MB | Q8_0 codec quant |
./build/bin/crispasr \
--backend qwen3-tts \
-m qwen3-tts-12hz-0.6b-base.gguf \
--codec-model qwen3-tts-tokenizer-12hz.gguf \
--voice clone.wav \
--ref-text "Exact transcript of clone.wav" \
--tts "Hello there" \
--tts-output hello.wav
The tokenizer GGUF is used for:
ref_codeCurrent CrispASR validation status:
qwen3-tts-tokenizer-12hz.gguf
qwen3-tts-tokenizer-12hz-q8_0.gguf
For best fidelity, keep the tokenizer / codec at F16 even when quantising the talker. In current CrispASR testing, codec quantisation drifts earlier in the codec-encoder path than talker-only quantisation.
This repo may also publish lower-bit talker variants in the companion talker repo. If you use them, the safest pairing is still:
qwen3-tts-tokenizer-12hz.gguf kept at F16In other words: if you must choose where to keep precision, keep it in the tokenizer / codec first.
models/convert-qwen3-tts-tokenizer-to-gguf.py.src/qwen3_tts.cpp, sharing the same runtime as the qwen3-tts talker backend.Architecture and behaviour were checked against the official Qwen release:
Qwen/Qwen3-TTS-Tokenizer-12HzQwenLM/Qwen3-TTSQwen3-TTS Technical Reportcstr/qwen3-tts-0.6b-base-GGUFQwen/Qwen3-TTS-Tokenizer-12HzCrispStrobe/CrispASRApache-2.0, inherited from the base model.
--- license: apache-2.0 language: - zh - en - ja - ko - de - fr - ru - pt - es - it pipeline_tag: audio-to-audio tags: - audio - tts - speech - codec - ggml - gguf - crispasr - qwen3_tts_tokenizer_12hz library_name: ggml base_model: Qwen/Qwen3-TTS-Tokenizer-12Hz --- # Qwen3-TTS Tokenizer 12Hz — GGUF (CrispASR) GGUF / ggml conversions of [`Qwen/Qwen3-TTS-Tokenizer-12Hz`](https://huggingface.co/Qwen/Qwen3-TTS-Tokenizer-12Hz) for use with the `qwen3-tts` backend in **[CrispStrobe/CrispASR](https://github.com/CrispStrobe/CrispASR)**. Qwen3-TTS-Tokenizer-12Hz is the separate **speech tokenizer / codec** used by the Qwen3-TTS family: - 10 supported languages: `zh en ja ko de fr ru pt es it` - 12.5 Hz, 16-codebook speech representation - used for reference-audio encoding, voice-pack baking, and final waveform decode - Apache-2.0 licence This repo contains the **tokenizer / codec** only. Use it together with the talker GGUF from [`cstr/qwen3-tts-0.6b-base-GGUF`](https://huggingface.co/cstr/qwen3-tts-0.6b-base-GGUF). ## Files File | Size | Notes --- | --- | --- `qwen3-tts-tokenizer-12hz.gguf` | 342 MB | F16 `qwen3-tts-tokenizer-12hz-q8_0.gguf` | 277 MB | Q8_0 codec quant ## Quick Start ```bash ./build/bin/crispasr \ --backend qwen3-tts \ -m qwen3-tts-12hz-0.6b-base.gguf \ --codec-model qwen3-tts-tokenizer-12hz.gguf \ --voice clone.wav \ --ref-text "Exact transcript of clone.wav" \ --tts "Hello there" \ --tts-output hello.wav ``` The tokenizer GGUF is used for: - encoding reference audio into `ref_code` - baking / loading voice-pack GGUFs - decoding generated codes back into 24 kHz mono WAV ## Quantisation Notes Current CrispASR validation status: - `qwen3-tts-tokenizer-12hz.gguf` - reference baseline - `qwen3-tts-tokenizer-12hz-q8_0.gguf` - usable, but numerically less faithful than the F16 codec in strict diff tests For best fidelity, keep the tokenizer / codec at F16 even when quantising the talker. In current CrispASR testing, codec quantisation drifts earlier in the codec-encoder path than talker-only quantisation. This repo may also publish lower-bit talker variants in the companion talker repo. If you use them, the safest pairing is still: - quantised talker - `qwen3-tts-tokenizer-12hz.gguf` kept at F16 In other words: if you must choose where to keep precision, keep it in the tokenizer / codec first. ## How this...
Source context: 200 downloads · 0 likes · Pipeline audio-to-audio · Library ggml · Repo gto85/qwen3-tts-tokenizer-12hz-GGUF