GGUF conversions of the CTC branch of nvidia/stthrfastconformerhybridlargepc for CrispASR. The upstream model is a hybrid transducer+CTC Croatian ASR release; the shared FastConformer encoder plus the auxiliary CTC...
Fonte do modelo
Descrição da fonte
GGUF conversions of the CTC branch of nvidia/stt_hr_fastconformer_hybrid_large_pc for CrispASR. The upstream model is a hybrid transducer+CTC Croatian ASR release; the shared FastConformer encoder plus the auxiliary CTC head are extracted here as a standalone CTC model (the RNNT prediction network and joint are dropped), giving a compact Croatian ASR model with punctuation + capitalisation.
Fontes
1 fonteVerificado 7 de set.
Artefatos de modelo
3 artefatosTrechos de fonte
2 trechos| Quant | Size | Description |
|---|---|---|
| F16 | 219 MB | Full precision |
| Q8_0 | 129 MB | 8-bit |
| Q4_K | 82 MB | 4-bit K-quant (recommended) |
17-layer NeMo FastConformer encoder + Conv1d CTC head. d_model=512, 8 heads, SentencePiece vocab, 80 log-mel features, ~115M params.
# Croatian ASR:
crispasr --backend fastconformer-ctc -m stt-hr-fastconformer-hybrid-ctc-large-q4_k.gguf -f audio.wav
# Forced alignment (word timestamps for known text, or re-timing an .srt):
crispasr --align-only -am stt-hr-fastconformer-hybrid-ctc-large-q4_k.gguf \
-f audio.wav --text-file subtitles.srt --align-output retimed.srt
All credit for the model goes to NVIDIA's NeMo team; this repository only repackages the CTC branch in GGUF form under the same CC-BY-4.0 license. Conversion: models/convert-stt-fastconformer-ctc-to-gguf.py in CrispASR.
nvidia.cc-by-4.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not.stt-hr-fastconformer-hybrid-ctc-large-q8_0.gguf
gguf · 112 MB · SHA-256 7a7e0d08d7e0…dd4d · Hugging Face
Baixar--- license: cc-by-4.0 base_model: nvidia/stt_hr_fastconformer_hybrid_large_pc language: - hr tags: - automatic-speech-recognition - forced-alignment - gguf - crispasr - fastconformer - ctc - nemo pipeline_tag: automatic-speech-recognition --- # stt-hr-fastconformer-hybrid-ctc-large-GGUF GGUF conversions of the **CTC branch** of [nvidia/stt_hr_fastconformer_hybrid_large_pc](https://huggingface.co/nvidia/stt_hr_fastconformer_hybrid_large_pc) for [CrispASR](https://github.com/CrispStrobe/CrispASR). The upstream model is a hybrid transducer+CTC Croatian ASR release; the shared FastConformer encoder plus the auxiliary CTC head are extracted here as a standalone CTC model (the RNNT prediction network and joint are dropped), giving a compact Croatian ASR **and forced-alignment** model with punctuation + capitalisation. | Quant | Size | Description | |---|---|---| | F16 | 219 MB | Full precision | | Q8_0 | 129 MB | 8-bit | | Q4_K | 82 MB | 4-bit K-quant (recommended) | ## Architecture 17-layer NeMo FastConformer encoder + Conv1d CTC head. d_model=512, 8 heads, SentencePiece vocab, 80 log-mel features, ~115M params. ## Usage ```bash # Croatian ASR: crispasr --backend fastconformer-ctc -m stt-hr-fastconformer-hybrid-ctc-large-q4_k.gguf -f audio.wav # Forced alignment (word timestamps for known text, or re-timing an .srt): crispasr --align-only -am stt-hr-fastconformer-hybrid-ctc-large-q4_k.gguf \ -f audio.wav --text-file subtitles.srt --align-output retimed.srt ``` ## Attribution All credit for the model goes to NVIDIA's NeMo team; this repository only repackages the CTC branch in GGUF form under the same CC-BY-4.0 license. Conversion: `models/convert-stt-fastconformer-ctc-to-gguf.py` in CrispASR. ## Provenance and EU AI Act Art. 53 note - **Upstream model:** [nvidia/stt_hr_fastconformer_hybrid_large_pc](https://huggingface.co/nvidia/stt_hr_fastconformer_hybrid_large_pc) — published by `nvidia`. - **Upstream licence:** `cc-by-4.0`. This repository redistributes under the same terms; it grants no rights the upstream licence does not. - **What was done here:** format conversion and/or quantisation only (GGUF). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs. - **Training data:** documented — where it is documented at all...
Source context: 350 downloads · 0 likes · Pipeline automatic-speech-recognition · Library nemo · Repo cstr/stt-hr-fastconformer-hybrid-ctc-large-GGUF