GGUF conversion of Silero's 95-language classifier for use with CrispASR.
Modellquelle
Quellenbeschreibung
GGUF conversion of Silero's 95-language classifier for use with CrispASR.
Quellen
1 QuelleVerifiziert 29. Aug.
Modellartefakte
1 ArtefaktQuellenauszüge
2 Auszüge| File | Type | Size | Notes |
|---|---|---|---|
silero-lid-lang95-f32.gguf | F32 | 16 MB | Full precision, recommended |
Quantized versions (Q8_0, Q5_0) were tested but break accuracy — the model is dominated by small Conv1d kernels where block quantization is destructive. At 16 MB F32, the model is already very small.
# Language detection pre-step for backends without native LID
crispasr --backend cohere -m cohere-transcribe-q5_0.gguf \
-f audio.wav -l auto \
--lid-backend silero --lid-model silero-lid-lang95-f32.gguf
# Standalone detection (via the C API)
silero_lid_context * ctx = silero_lid_init("silero-lid-lang95-f32.gguf", 4);
float conf;
const char * lang = silero_lid_detect(ctx, samples, n_samples, &conf);
// lang = "en, English"
silero_lid_free(ctx);
The model supports 95 languages across 58 language groups. Language detection works best on audio clips between 3-20 seconds of speech.
Converted from the ONNX export (lang_classifier_95.onnx) using models/convert-silero-lid-to-gguf.py from the CrispASR repo.
mit. This repository redistributes under the same terms; it grants no rights the upstream licence does not.--- license: mit tags: - language-identification - silero - gguf - crispasr - audio-classification language: - multilingual base_model: snakers4/silero-lang-classifier-95 pipeline_tag: audio-classification --- # Silero Language Classifier 95 — GGUF GGUF conversion of [Silero's 95-language classifier](https://github.com/snakers4/silero-vad) for use with [CrispASR](https://github.com/cstr/whisper.cpp/tree/integrated_cli). ## Model Details - **Architecture:** Learned STFT frontend + 8-stage MobileNet-style depthwise-separable conv encoder with interleaved post-norm transformers + attention-weighted pooling + 95-language / 58-group classifiers - **Parameters:** 507 tensors, ~4M parameters - **Input:** Raw 16 kHz mono PCM audio (best results on clips < 20 seconds) - **Output:** 95-language log-probabilities + 58-language-group log-probabilities - **License:** MIT (same as upstream Silero) ## Files | File | Type | Size | Notes | |------|------|------|-------| | `silero-lid-lang95-f32.gguf` | F32 | 16 MB | Full precision, recommended | Quantized versions (Q8_0, Q5_0) were tested but break accuracy — the model is dominated by small Conv1d kernels where block quantization is destructive. At 16 MB F32, the model is already very small. ## Usage with CrispASR ```bash # Language detection pre-step for backends without native LID crispasr --backend cohere -m cohere-transcribe-q5_0.gguf \ -f audio.wav -l auto \ --lid-backend silero --lid-model silero-lid-lang95-f32.gguf # Standalone detection (via the C API) silero_lid_context * ctx = silero_lid_init("silero-lid-lang95-f32.gguf", 4); float conf; const char * lang = silero_lid_detect(ctx, samples, n_samples, &conf); // lang = "en, English" silero_lid_free(ctx); ``` ## Supported Languages (95) The model supports 95 languages across 58 language groups. Language detection works best on audio clips between 3-20 seconds of speech. ## Conversion Converted from the ONNX export (`lang_classifier_95.onnx`) using `models/convert-silero-lid-to-gguf.py` from the CrispASR repo. ## Acknowledgements - [Silero Team](https://github.com/snakers4) for the original model - [CrispASR](https://github.com/cstr/whisper.cpp/tree/integrated_cli) for the native GGUF runtime ## Provenance and EU AI Act Art. 53 note - **Upstream model:** snakers4/silero-lang-classifier-95. - **Upstream licence:** `mit`. This repository redi...
Source context: 311 downloads · 2 likes · Pipeline audio-classification · Repo cstr/silero-lid-lang95-GGUF