4-bit (MLX) quantization of kyutai/hibiki-zero-3b-pytorch-bf16, a simultaneous speech-to-speech + speech-to-text translation model (FR / ES / PT / DE → EN) with voice transfer.
Source du modèle
Description de la source
4-bit (MLX) quantization of
kyutai/hibiki-zero-3b-pytorch-bf16,
a simultaneous speech-to-speech + speech-to-text translation model
(FR / ES / PT / DE → EN) with voice transfer.
The language model is quantized to , shrinking it from . On an Apple M4 Pro this runs at (16.5 tok/s @ 12.5 Hz), faster than the bf16 PyTorch/MPS path. The Mimi codec is kept separate in bf16.
Sources
1 sourceVérifié 7 août
Artefacts du modèle
2 artefactsmimi-pytorch-e351c8d8@125.safetensors
safetensors · 367 MB · SHA-256 09b782f06298…3f50 · Hugging Face
TéléchargerExtraits de sources
2 extraitsgroup_size=32 is intentional: stock moshi-mlx (and the moshi-swift iOS
loader) hardcode gs32 for .q4.safetensors, so a larger group size would not
load without patching the loader.
leon.wav, FR→EN)| Value | |
|---|---|
| LM weights | 5.8 GB bf16 → 2.2 GB q4 (578 layers quantized) |
| Speed | 16.5 tok/s ≈ 1.3× real-time — vs ~0.7× for the PyTorch/MPS path (~1.9× faster) |
| Quality | Coherent FR→EN — correctly translated the Léon Marchand / Paris 2024 Olympics commentary; minor q4 artifacts (e.g. one "Paris 1024" slip) |
| Audio out | 24 kHz wav ✅ |
| File | Description |
|---|---|
hibiki.q4.safetensors | 4-bit quantized LM (MLX) |
config.json | model config (needed to build the LmConfig) |
mimi-pytorch-e351c8d8@125.safetensors | Mimi codec (bf16) |
tokenizer_spm_48k_multi6_2.model | SentencePiece tokenizer |
mlx_hibiki_patch.py | runtime patches for moshi-mlx (required) |
verify_mlx_q4.py | example inference script |
Stock moshi-mlx (0.3.0) targets
moshi / older hibiki and misses three hibiki-zero deltas, so these weights will
not load without mlx_hibiki_patch.py:
hidden_scale (feedforward dim) and kv_repeat=2
instead of the hardcoded 4*dim / kv_repeat=1.kv_repeat=2).rope_concat == RoPE with interleave=False
(MLX traditional=False).pip install moshi-mlx
python - <<'PY'
import mlx_hibiki_patch # patches moshi_mlx for hibiki-zero — import first
from moshi_mlx import run_inference
import sys
sys.argv = [
"run_inference",
"--lm-config", "config.json",
"--moshi-weights", "hibiki.q4.safetensors",
"--mimi-weights", "mimi-pytorch-e351c8d8@125.safetensors",
"--tokenizer", "tokenizer_spm_48k_multi6_2.model",
"input_fr.wav", "output_en.wav",
]
run_inference.main()
PY
See verify_mlx_q4.py for a ready-to-run example.
Inherits CC BY-NC-SA 4.0 (non-commercial, share-alike) from the base model
kyutai/hibiki-zero-3b-pytorch-bf16.
--- license: cc-by-nc-sa-4.0 base_model: kyutai/hibiki-zero-3b-pytorch-bf16 tags: - mlx - speech-translation - speech-to-speech - quantized language: - en - fr - es - pt - de pipeline_tag: audio-to-audio --- # Hibiki-Zero 3B — MLX 4-bit 4-bit ([MLX](https://github.com/ml-explore/mlx)) quantization of [`kyutai/hibiki-zero-3b-pytorch-bf16`](https://huggingface.co/kyutai/hibiki-zero-3b-pytorch-bf16), a simultaneous speech-to-speech + speech-to-text translation model (FR / ES / PT / DE → EN) with voice transfer. The language model is quantized to **4-bit (group_size=32)**, shrinking it from **5.8 GB → 2.2 GB**. On an Apple M4 Pro this runs at **~1.3× real-time** (16.5 tok/s @ 12.5 Hz), faster than the bf16 PyTorch/MPS path. The Mimi codec is kept separate in bf16. `group_size=32` is intentional: stock `moshi-mlx` (and the moshi-swift iOS loader) hardcode gs32 for `.q4.safetensors`, so a larger group size would not load without patching the loader. ## Results (Apple M4 Pro, `leon.wav`, FR→EN) | | Value | |--------------|-------| | **LM weights** | 5.8 GB bf16 → **2.2 GB** q4 (578 layers quantized) | | **Speed** | 16.5 tok/s ≈ **1.3× real-time** — vs ~0.7× for the PyTorch/MPS path (~1.9× faster) | | **Quality** | Coherent FR→EN — correctly translated the Léon Marchand / Paris 2024 Olympics commentary; minor q4 artifacts (e.g. one "Paris 1024" slip) | | **Audio out** | 24 kHz wav ✅ | ## Files | File | Description | |------|-------------| | `hibiki.q4.safetensors` | 4-bit quantized LM (MLX) | | `config.json` | model config (needed to build the `LmConfig`) | | `mimi-pytorch-e351c8d8@125.safetensors` | Mimi codec (bf16) | | `tokenizer_spm_48k_multi6_2.model` | SentencePiece tokenizer | | `mlx_hibiki_patch.py` | runtime patches for `moshi-mlx` (**required**) | | `verify_mlx_q4.py` | example inference script | ## Why the patch is required Stock [`moshi-mlx`](https://pypi.org/project/moshi-mlx/) (0.3.0) targets moshi / older hibiki and misses three hibiki-zero deltas, so these weights will **not** load without `mlx_hibiki_patch.py`: 1. **config** — honour `hidden_scale` (feedforward dim) and `kv_repeat=2` instead of the hardcoded `4*dim` / `kv_repeat=1`. 2. **attention** — grouped-query attention in the forward pass (the main transformer uses `kv_repeat=2`). 3. **positional embedding** — `rope_concat` == RoPE with `interleave=False` (MLX `tradit...
Source context: 27 downloads · 0 likes · Pipeline audio-to-audio · Library mlx · Repo huybik/hibiki-zero-3b-mlx-q4