A 4-bit (Q40) GGUF quantization of Qwen3-ASR 1.7B, for use with llama.cpp's audio-multimodal path (llama-server / llama-mtmd-cli).
Fonte do modelo
Descrição da fonte
A 4-bit (Q4_0) GGUF quantization of Qwen3-ASR 1.7B, for use with llama.cpp's
audio-multimodal path (llama-server / llama-mtmd-cli).
Quantized with (llama.cpp) from the GGUF in . The audio projector () is the upstream projector from that same repo, unmodified.
Fontes
1 fonteVerificado 7 de ago.
Artefatos de modelo
2 artefatosTrechos de fonte
2 trechosllama-quantizemmproj| file | role | approx size | sha256 |
|---|---|---|---|
Qwen3-ASR-1.7B-Q4_0.gguf | LLM (4-bit) | 1.23 GB | adf2c487fcce3993bb30c9406748f5ef111a3a28b0a6c26f34223185011427c7 |
mmproj-Qwen3-ASR-1.7B-Q8_0.gguf | audio projector | 0.36 GB | 46c1d533af3f354ceb37ce855dbceff7da7fa7cf1e6a523df3b13440bd164c0d |
On Intel Arc / Vulkan integrated GPUs, the legacy Q4_0 format decodes faster than the
Q4_K_M K-quant (which is under-optimized on those backends) at essentially the same
size - a better speed/size trade-off there.
llama-server -m Qwen3-ASR-1.7B-Q4_0.gguf --mmproj mmproj-Qwen3-ASR-1.7B-Q8_0.gguf -c 1024 -ngl 99
License follows upstream Qwen3-ASR (Apache-2.0).
--- license: apache-2.0 base_model: ggml-org/Qwen3-ASR-1.7B-GGUF tags: - gguf - llama.cpp - automatic-speech-recognition - qwen3-asr --- # Qwen3-ASR 1.7B - Q4_0 GGUF A 4-bit (`Q4_0`) GGUF quantization of **Qwen3-ASR 1.7B**, for use with llama.cpp's audio-multimodal path (`llama-server` / `llama-mtmd-cli`). ## Provenance Quantized with `llama-quantize` (llama.cpp) from the **bf16** GGUF in [`ggml-org/Qwen3-ASR-1.7B-GGUF`](https://huggingface.co/ggml-org/Qwen3-ASR-1.7B-GGUF). The audio projector (`mmproj`) is the upstream **Q8_0** projector from that same repo, unmodified. ## Files | file | role | approx size | sha256 | |---|---|---|---| | `Qwen3-ASR-1.7B-Q4_0.gguf` | LLM (4-bit) | 1.23 GB | `adf2c487fcce3993bb30c9406748f5ef111a3a28b0a6c26f34223185011427c7` | | `mmproj-Qwen3-ASR-1.7B-Q8_0.gguf` | audio projector | 0.36 GB | `46c1d533af3f354ceb37ce855dbceff7da7fa7cf1e6a523df3b13440bd164c0d` | ## Why Q4_0 (rather than Q4_K_M) On Intel Arc / Vulkan integrated GPUs, the legacy `Q4_0` format decodes faster than the `Q4_K_M` K-quant (which is under-optimized on those backends) at essentially the same size - a better speed/size trade-off there. ## Usage ``` llama-server -m Qwen3-ASR-1.7B-Q4_0.gguf --mmproj mmproj-Qwen3-ASR-1.7B-Q8_0.gguf -c 1024 -ngl 99 ``` License follows upstream Qwen3-ASR (Apache-2.0).
Source context: 185 downloads · 0 likes · Pipeline automatic-speech-recognition · Repo getonit/Qwen3-ASR-1.7B-Q4_0-GGUF