GGUF conversions of OpenMOSS-Team/MOSS-Transcribe-Diarize for use with transcribe.cpp.
Modellquelle
Quellenauszug
GGUF conversions of OpenMOSS-Team/MOSS-Transcribe-Diarize for use with transcribe.cpp.
Quellen
1 QuelleVerifiziert 30. Aug.
Modellartefakte
6 ArtefakteMOSS-Transcribe-Diarize-BF16.gguf
gguf · 1,70 GB · SHA-256 ccf9aa138e18…7712 · Hugging Face
HerunterladenMOSS-Transcribe-Diarize-F16.gguf
gguf · 1,71 GB · SHA-256 fee2e436f7a6…7f4b · Hugging Face
HerunterladenQuellenauszüge
2 AuszügeMOSS-Transcribe-Diarize-Q4_K_M.gguf
gguf · 589 MB · SHA-256 e371990f2c8f…0b8b · Hugging Face
MOSS-Transcribe-Diarize-Q5_K_M.gguf
gguf · 668 MB · SHA-256 52deaeff9312…473c · Hugging Face
HerunterladenMOSS-Transcribe-Diarize-Q6_K.gguf
gguf · 733 MB · SHA-256 f36e6edd7cae…2bbd · Hugging Face
HerunterladenMOSS-Transcribe-Diarize-Q8_0.gguf
gguf · 941 MB · SHA-256 64ec654dc6ff…5039 · Hugging Face
Herunterladen--- license: apache-2.0 base_model: OpenMOSS-Team/MOSS-Transcribe-Diarize base_model_relation: quantized library_name: transcribe.cpp pipeline_tag: automatic-speech-recognition language: - en - zh tags: - gguf - transcribe.cpp - asr - speech-to-text - moss - audio-llm - whisper-encoder - qwen3 - diarization transcribe_cpp: wer_librispeech_test_clean: bf16: 2.08 f16: 2.07 q8_0: 1.93 q6_k: 1.96 q5_k_m: 1.99 q4_k_m: 2.59 rtf_m4_max: metal: 28.1 cpu: 5.8 rtf_ryzen_4750u: vulkan: 3.0 cpu: 1.6 streaming: false translate: false lang_detect: false timestamps: segment --- # MOSS-Transcribe-Diarize: transcribe.cpp GGUF GGUF conversions of [OpenMOSS-Team/MOSS-Transcribe-Diarize](https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize) for use with [transcribe.cpp](https://github.com/handy-computer/transcribe.cpp). Ported from upstream commit [d7231bb](https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize/commit/d7231bb), pinned 2026-07-12. Validated against the MOSS author repo (OpenMOSS/MOSS-Transcribe-Diarize) reference at transcribe.cpp commit [3f5e15c](https://github.com/handy-computer/transcribe.cpp/tree/3f5e15c) on 2026-07-12. Offline English/Chinese speech-to-text with speaker diarization. A 0.9B audio-LLM: a Whisper-Medium encoder (24 layers, d_model=1024) feeds a 4x temporal merge + VQAdaptor bridge into a Qwen3-0.6B decoder (28 layers) via audio-token injection. Takes a 16 kHz mono WAV and emits transcript text in the canonical diarized format `[start][Sxx]text[end]`, where the speaker tags and segment timestamps are generated text, not special tokens. Not a streaming model. ## Downloads | Quantization | Download | Size | WER (LibriSpeech test-clean) | | --- | --- | ---: | ---: | | BF16 | [MOSS-Transcribe-Diarize-BF16.gguf](https://huggingface.co/handy-computer/MOSS-Transcribe-Diarize-gguf/resolve/main/MOSS-Transcribe-Diarize-BF16.gguf) | 1.83 GB | 2.08% | | F16 | [MOSS-Transcribe-Diarize-F16.gguf](https://huggingface.co/handy-computer/MOSS-Transcribe-Diarize-gguf/resolve/main/MOSS-Transcribe-Diarize-F16.gguf) | 1.83 GB | 2.07% | | Q8_0 | [MOSS-Transcribe-Diarize-Q8_0.gguf](https://huggingface.co/handy-computer/MOSS-Transcribe-Diarize-gguf/resolve/main/MOSS-Transcribe-Diarize-Q8_0.gguf) | 987 MB | 1.93% | | Q6_K | [MOSS-Transcribe-Diarize-Q6_K.gguf](https://huggingface.co/handy-computer/MOSS-Transcribe-Diarize-...
Source context: 2200 downloads · 0 likes · Pipeline automatic-speech-recognition · Library transcribe.cpp · Repo handy-computer/moss-transcribe-diarize-gguf