Meta omniASRCTC1Bv2 exported for server-side ONNX Runtime (CUDA EP with CPU fallback). Exported by models/omnilingual-asr/export/convertonnx.py in soniqo/speech-models.
Model source
Source description
Meta omniASR_CTC_1B_v2 exported for server-side ONNX Runtime (CUDA EP with
CPU fallback). Exported by models/omnilingual-asr/export/convert_onnx.py in
soniqo/speech-models.
Sources
1 sourceVerified Aug 30
Model artifacts
1 artifactSource excerpts
2 excerpts| input | audio [1, S] f32 — 16 kHz waveform, z-scored by the caller |
| output | logits [1, T, 10288] — CTC logits, T = ceil(S / 320) |
The time axis is dynamic: cost tracks the real clip length. (The LiteRT and CoreML exports bake a fixed window instead, because on-device wants a static shape.)
The caller must z-score the waveform — the graph does not normalise internally. Feeding raw audio scores 67.7% WER against 11.6% normalised.
| int8 ONNX (CPU, 8 threads) | fp32 PyTorch (GPU) | |
|---|---|---|
| English | 11.64 | 11.54 |
| French | 14.69 | 14.69 |
| German | 8.22 | 8.03 |
| Arabic (ar_eg) | 19.02 | 18.91 |
int8 costs ~0.1 WER: CTC is a single encoder pass, so there is no recurrent state for quantization error to accumulate through. Throughput is 17-21x realtime on 8 CPU threads.
--- license: cc-by-nc-4.0 library_name: onnx tags: [automatic-speech-recognition, ctc, multilingual, onnx] --- # Omnilingual ASR CTC 1B — ONNX (int8) Meta `omniASR_CTC_1B_v2` exported for server-side ONNX Runtime (CUDA EP with CPU fallback). Exported by `models/omnilingual-asr/export/convert_onnx.py` in soniqo/speech-models. ## Graph | | | |---|---| | input | `audio` [1, S] f32 — 16 kHz waveform, **z-scored by the caller** | | output | `logits` [1, T, 10288] — CTC logits, T = ceil(S / 320) | The time axis is **dynamic**: cost tracks the real clip length. (The LiteRT and CoreML exports bake a fixed window instead, because on-device wants a static shape.) **The caller must z-score the waveform** — the graph does not normalise internally. Feeding raw audio scores 67.7% WER against 11.6% normalised. ## Quality (FLEURS, 50 utts/language, Whisper normalizer) | | int8 ONNX (CPU, 8 threads) | fp32 PyTorch (GPU) | |---|---|---| | English | 11.64 | 11.54 | | French | 14.69 | 14.69 | | German | 8.22 | 8.03 | | Arabic (ar_eg) | 19.02 | 18.91 | int8 costs ~0.1 WER: CTC is a single encoder pass, so there is no recurrent state for quantization error to accumulate through. Throughput is 17-21x realtime on 8 CPU threads.
Source context: 3 downloads · 0 likes · Pipeline automatic-speech-recognition · Library onnx · Repo soniqo/Omnilingual-ASR-CTC-1B-ONNX