GGUF / ggml conversion of Aratako/Irodori-TTS-500M-v3 for use with CrispStrobe/CrispASR.
Fonte do modelo
Descrição da fonte
GGUF / ggml conversion of Aratako/Irodori-TTS-500M-v3 for use with CrispStrobe/CrispASR.
Irodori-TTS is a 500M-param text-to-speech model with zero-shot voice cloning via DAC-VAE latent conditioning. Japanese-focused with a 102K-vocab sarashina2.2 tokenizer. 48 kHz mono output.
Fontes
1 fonteVerificado 15 de ago.
Artefatos de modelo
8 artefatosTrechos de fonte
2 trechosLicense: MIT (follows upstream Irodori-TTS license).
| File | Quant | Size | Notes |
|---|---|---|---|
irodori-tts-500m-v3-f16.gguf | F16 | ~1.9 GB | Reference quality |
irodori-tts-500m-v3-q4_k.gguf | Q4_K | ~852 MB | Recommended — fits 8 GB RAM |
irodori-tts-500m-v3-q8_0.gguf | Q8_0 | ~896 MB | Near-lossless |
irodori-tts-ref.gguf | F32 | ~4 KB | Reference activations for diff harness |
Text Input (Japanese / mixed)
│
sarashina2.2 Tokenize (102K vocab, BPE)
│
TextEncoder (14L, 1280d, 10 heads, RoPE + SwiGLU)
│── Each position: self-attention + gated residual + SwiGLU FFN
│
├── [Optional] ReferenceLatentEncoder (14L, 1280d)
│ └── DAC-VAE latent from reference audio → speaker conditioning
│
DiT Backbone (24L, 2048d, 16 heads)
│── LowRankAdaLN (rank=256) timestep conditioning
│── JointAttention: self-KV + text-context-KV + speaker-context-KV
│── Half-RoPE (first half of head_dim rotated, rest passthrough)
│── SwiGLU MLP (ratio 2.875)
│
Euler RF ODE Solver (40 steps, CFG)
│── noise → DAC-VAE latent sequence (32-dim continuous)
│
Semantic-DACVAE Decoder (48 kHz reconstruction)
│── 32-dim latent → Snake1d + ConvTranspose1d upsampling → PCM
│
Output: float32 mono @ 48 kHz
models/convert-irodori-tts-to-gguf.pyAratako.mit. This repository redistributes under the same terms; it grants no rights the upstream licence does not.irodori-tts-500m-v3-q4_k.gguf
gguf · 389 MB · SHA-256 188126c5813b…bc03 · Hugging Face
irodori-tts-ref.gguf
gguf · 50,3 KB · Hugging Face
--- license: mit language: - ja base_model: - Aratako/Irodori-TTS-500M-v3 pipeline_tag: text-to-speech tags: - tts - text-to-speech - irodori-tts - flow-matching - diffusion-transformer - dacvae - gguf - crispasr - japanese library_name: ggml --- # Irodori-TTS-500M-v3 — GGUF (ggml-quantised) GGUF / ggml conversion of [`Aratako/Irodori-TTS-500M-v3`](https://huggingface.co/Aratako/Irodori-TTS-500M-v3) for use with **[CrispStrobe/CrispASR](https://github.com/CrispStrobe/CrispASR)**. Irodori-TTS is a 500M-param **Rectified Flow Diffusion Transformer (RF-DiT)** text-to-speech model with zero-shot voice cloning via DAC-VAE latent conditioning. Japanese-focused with a 102K-vocab sarashina2.2 tokenizer. 48 kHz mono output. **License:** MIT (follows upstream Irodori-TTS license). ## Files | File | Quant | Size | Notes | |---|---|---:|---| | `irodori-tts-500m-v3-f16.gguf` | F16 | ~1.9 GB | Reference quality | | `irodori-tts-500m-v3-q4_k.gguf` | Q4_K | ~852 MB | **Recommended** — fits 8 GB RAM | | `irodori-tts-500m-v3-q8_0.gguf` | Q8_0 | ~896 MB | Near-lossless | | `irodori-tts-ref.gguf` | F32 | ~4 KB | Reference activations for diff harness | ## Architecture ``` Text Input (Japanese / mixed) │ sarashina2.2 Tokenize (102K vocab, BPE) │ TextEncoder (14L, 1280d, 10 heads, RoPE + SwiGLU) │── Each position: self-attention + gated residual + SwiGLU FFN │ ├── [Optional] ReferenceLatentEncoder (14L, 1280d) │ └── DAC-VAE latent from reference audio → speaker conditioning │ DiT Backbone (24L, 2048d, 16 heads) │── LowRankAdaLN (rank=256) timestep conditioning │── JointAttention: self-KV + text-context-KV + speaker-context-KV │── Half-RoPE (first half of head_dim rotated, rest passthrough) │── SwiGLU MLP (ratio 2.875) │ Euler RF ODE Solver (40 steps, CFG) │── noise → DAC-VAE latent sequence (32-dim continuous) │ Semantic-DACVAE Decoder (48 kHz reconstruction) │── 32-dim latent → Snake1d + ConvTranspose1d upsampling → PCM │ Output: float32 mono @ 48 kHz ``` ## Source model - **Upstream:** [Aratako/Irodori-TTS-500M-v3](https://huggingface.co/Aratako/Irodori-TTS-500M-v3) (safetensors, ~1.9 GB) - **Codec:** [Aratako/Semantic-DACVAE-Japanese-32dim](https://huggingface.co/Aratako/Semantic-DACVAE-Japanese-32dim) - **Tokenizer:** [sbintuitions/sarashina2.2-0.5b](https://huggingface.co/sbintuitions/sarashina2.2-0.5b) - **Code:** [Aratako/Irodori-TTS](https://g...
Source context: 947 downloads · 1 likes · Pipeline text-to-speech · Library ggml · Repo cstr/irodori-tts-GGUF