Pre-converted, pre-quantized MLX weights for Zyphra's ZONOS2 — an 8B-parameter Mixture-of-Experts autoregressive text-to-speech model — running natively on Apple Silicon.
Fonte do modelo
Trecho da fonte
Pre-converted, pre-quantized MLX weights for Zyphra's ZONOS2 — an 8B-parameter Mixture-of-Experts autoregressive text-to-speech model — running natively on Apple Silicon.
Fontes
1 fonteVerificado 13 de set.
Artefatos de modelo
9 artefatosTrechos de fonte
2 trechosbf16/zonos2-bf16.safetensors
safetensors · 14,3 GB · SHA-256 cd2fd4c9e867…26b4 · Hugging Face
int4/dac_44khz/model.safetensors
safetensors · 292 MB · SHA-256 6128ebff483a…913a · Hugging Face
Baixarint4/speaker_encoder/model.safetensors
safetensors · 22,9 MB · SHA-256 df60a638e7f4…ea48 · Hugging Face
Baixarint8/dac_44khz/model.safetensors
safetensors · 292 MB · SHA-256 6128ebff483a…913a · Hugging Face
Baixarint8/speaker_encoder/model.safetensors
safetensors · 22,9 MB · SHA-256 df60a638e7f4…ea48 · Hugging Face
Baixar--- license: apache-2.0 library_name: mlx pipeline_tag: text-to-speech base_model: drbaph/ZONOS2-BF16 tags: - mlx - apple-silicon - text-to-speech - tts - voice-cloning - zonos - zonos2 - moe language: - en --- # zonos2-mlx — ready-to-run MLX weights Pre-converted, pre-quantized [MLX](https://github.com/ml-explore/mlx) weights for Zyphra's [**ZONOS2**](https://github.com/Zyphra/ZONOS2) — an **8B-parameter Mixture-of-Experts** autoregressive text-to-speech model — running natively on Apple Silicon. **Download and run. No PyTorch in the inference path, no conversion step.** - 🧠 **Model:** 16-expert top-1 MoE AR trunk (layer 26 routes top-2) → DAC 44.1 kHz neural codec for the waveform, with an ECAPA-TDNN speaker encoder (+ LDA) for voice cloning from a short reference clip. - 🍎 **Runtime:** [`sb1992/mlx-zonos2`](https://github.com/sb1992/mlx-zonos2) — a clean-room MLX reimplementation of the inference runtime, gated per-stage against the original PyTorch model. - 📦 **This repo:** the weights only. Three precision tiers, each a **self-contained folder**. ## Tiers Each folder (`bf16/`, `int8/`, `int4/`) is **self-contained** — it bundles the quantized trunk plus the (tier-independent) DAC codec and ECAPA speaker encoder, so you download **one folder** and it just runs. | Folder | what's quantized | folder size | peak RAM | target Macs | |---|---|---|---|---| | `bf16/` | nothing (reference) | ~14 GB | ~44 GB | 64 GB | | `int8/` | attention/FFN/lm_head + experts int8; router/embeddings/norms bf16 | ~7.9 GB | ~13 GB | 32 GB | | `int4/` | attention/FFN/lm_head int8; experts gate/up int4, down int8; router/embeddings/norms bf16 | ~5.7 GB | ~10.6 GB | 16 GB | <sub>Folder size includes the bundled ~315 MB DAC codec + ECAPA speaker encoder (identical across tiers — Hugging Face Xet de-dups them, so they cost storage only once).</sub> The MoE experts (the bulk of the 8B) carry the int4; the **router/gate**, the **`lm_head`**, and the sensitive expert **`down`** projection stay int8/bf16 — the MoE-quant recipe that keeps the model intact. All three tiers produce **full, intelligible audio** — they're equal options, pick by the RAM you have. ## Quick start ```bash # 1. get the runtime git clone https://github.com/sb1992/mlx-zonos2.git cd mlx-zonos2 uv sync --extra oracle # `oracle` extra = torchaudio, for enrolling a voice from raw audio # 2. dow...
Source context: 0 downloads · 0 likes · Pipeline text-to-speech · Library mlx · Repo shraey/zonos2-mlx