Experimental low-memory GGUF quantizations of funasr/paraformer-zh. The vocabulary and CMVN tensors are embedded in each file.
Modellquelle
Quellenbeschreibung
Experimental low-memory GGUF quantizations of
funasr/paraformer-zh.
The vocabulary and CMVN tensors are embedded in each file.
Quellen
1 QuelleVerifiziert 15. Aug.
Modellartefakte
3 Artefakteparaformer-punctuation-int8.onnx
onnx · 72,0 MB · SHA-256 65a3fb9f5ad7…24b1 · Hugging Face
HerunterladenQuellenauszüge
2 Auszüge| File | Size | Quantization | Intended use |
|---|
paraformer-q5_0.gguf | 149.7 MiB | Q5_0 matrices, remaining tensors F32 | Recommended low-memory balance |
paraformer-q4_0.gguf | 124.3 MiB | Q4_0 matrices, remaining tensors F32 | Experimental minimum-memory option |
The original exporter only supports F32, F16, and Q8_0. These files were exported directly from the official FP32 checkpoint with an extended exporter; they were not requantized from Q8.
Measurements used the persistent CPU Paraformer GGUF runtime on an Intel Core i7-14700KF and a 30-second Mandarin/English mixed-language clip.
| Quantization | Resident peak working set | Recognition time |
|---|---|---|
| Q8_0 upstream baseline | 252.8 MiB | 996 ms |
| Q5_0 | 176.5 MiB | 1,201 ms |
| Q4_0 | 151.1 MiB | 1,009 ms |
Across 59 clips from the same recording, Q5_0 and Q4_0 each matched the Q8_0 token sequence exactly on 18 clips. Their aggregate token edit distance from Q8_0 was 8.8% and 10.1%, respectively. This is output drift relative to Q8, not a labeled word-error-rate measurement. Validate the quantization with your own language and acoustic conditions before replacing Q8 in quality-sensitive deployments.
These files use the custom paraformer GGUF architecture implemented by the
FunASR runtime/llama.cpp/paraformer runtime. They are not language models and
cannot be loaded by general-purpose llama.cpp applications.
The compatible runtime is FunASR's runtime/llama.cpp/paraformer
implementation. The extended export script used for these files is included
in this repository.
The export input was the official model.pt file:
size: 880502012 bytes
sha256: 5bba782a5e9196166233b9ab12ba04cadff9ef9212b4ff6153ed9290ff679025
Exported files:
paraformer-q5_0.gguf
size: 156967168 bytes
sha256: 1f2309eacd761c1f4184177c718321cd6ad3c07e7b17d6f796e5fb15565906ec
paraformer-q4_0.gguf
size: 130313216 bytes
sha256: 992562722aa2c4e88158245a0fc5be0e1338c580db808a97c6262f6864317584
The base model and the FunASR exporter/runtime are provided under Apache-2.0.
See LICENSE and the base model repository for details.
paraformer-q5_0.gguf
gguf · 150 MB · SHA-256 1f2309eacd76…06ec · Hugging Face
--- license: apache-2.0 language: - zh - en base_model: funasr/paraformer-zh pipeline_tag: automatic-speech-recognition tags: - paraformer - funasr - gguf - quantized --- # paraformer-zh GGUF Experimental low-memory GGUF quantizations of [`funasr/paraformer-zh`](https://huggingface.co/funasr/paraformer-zh). The vocabulary and CMVN tensors are embedded in each file. ## Files | File | Size | Quantization | Intended use | | --- | ---: | --- | --- | | `paraformer-q5_0.gguf` | 149.7 MiB | Q5_0 matrices, remaining tensors F32 | Recommended low-memory balance | | `paraformer-q4_0.gguf` | 124.3 MiB | Q4_0 matrices, remaining tensors F32 | Experimental minimum-memory option | The original exporter only supports F32, F16, and Q8_0. These files were exported directly from the official FP32 checkpoint with an extended exporter; they were not requantized from Q8. ## Measured behavior Measurements used the persistent CPU Paraformer GGUF runtime on an Intel Core i7-14700KF and a 30-second Mandarin/English mixed-language clip. | Quantization | Resident peak working set | Recognition time | | --- | ---: | ---: | | Q8_0 upstream baseline | 252.8 MiB | 996 ms | | Q5_0 | 176.5 MiB | 1,201 ms | | Q4_0 | 151.1 MiB | 1,009 ms | Across 59 clips from the same recording, Q5_0 and Q4_0 each matched the Q8_0 token sequence exactly on 18 clips. Their aggregate token edit distance from Q8_0 was 8.8% and 10.1%, respectively. This is output drift relative to Q8, not a labeled word-error-rate measurement. Validate the quantization with your own language and acoustic conditions before replacing Q8 in quality-sensitive deployments. ## Compatibility These files use the custom `paraformer` GGUF architecture implemented by the FunASR `runtime/llama.cpp/paraformer` runtime. They are not language models and cannot be loaded by general-purpose llama.cpp applications. The compatible runtime is FunASR's `runtime/llama.cpp/paraformer` implementation. The extended export script used for these files is included in this repository. ## Source integrity The export input was the official `model.pt` file: ```text size: 880502012 bytes sha256: 5bba782a5e9196166233b9ab12ba04cadff9ef9212b4ff6153ed9290ff679025 ``` Exported files: ```text paraformer-q5_0.gguf size: 156967168 bytes sha256: 1f2309eacd761c1f4184177c718321cd6ad3c07e7b17d6f796e5fb15565906ec paraformer-q4_0.gguf size: 1303132...
Source context: 46 downloads · 0 likes · Pipeline automatic-speech-recognition · Repo AIRsLight/paraformer-zh-GGUF