> ⚠️ Non-commercial use only. This model is licensed CC BY-NC 4.0, > inherited from the original Meta AI checkpoint. Unlike most other onnx-asr > conversions in this collection (Apache-2.0 / CC0 / CC-BY), this model...
Modellquelle
Quellenauszug
⚠️ Non-commercial use only. This model is licensed CC BY-NC 4.0, > inherited from the original Meta AI checkpoint. Unlike most other onnx-asr > conversions in this collection (Apache-2.0 / CC0 / CC-BY), this model...
Quellen
1 QuelleVerifiziert 13. Sept.
Modellartefakte
1 ArtefaktQuellenauszüge
2 Auszüge--- license: cc-by-nc-4.0 language: - hu tags: - onnx - onnx-asr - wav2vec2 - automatic-speech-recognition base_model: - facebook/wav2vec2-base-10k-voxpopuli-ft-hu --- # wav2vec2-base-10k-voxpopuli-ft-hu-onnx (ONNX) > **⚠️ Non-commercial use only.** This model is licensed **CC BY-NC 4.0**, > inherited from the original Meta AI checkpoint. Unlike most other onnx-asr > conversions in this collection (Apache-2.0 / CC0 / CC-BY), this model may > **not** be used for commercial purposes. ONNX export of [facebook/wav2vec2-base-10k-voxpopuli-ft-hu](https://huggingface.co/facebook/wav2vec2-base-10k-voxpopuli-ft-hu), a Hungarian wav2vec2 CTC ASR model from Meta AI (FAIR) — VoxPopuli project, fine-tuned on the [VoxPopuli](https://github.com/facebookresearch/voxpopuli) European Parliament speech corpus — for use with [onnx-asr](https://github.com/istupakov/onnx-asr) (`wav2vec2-ctc` model type). Per-utterance zero-mean/unit-variance normalization is baked into the ONNX graph, masked by `input_lengths` for correct behavior with padded/batched input, so the model works with onnx-asr's plain `identity` preprocessor (raw 16kHz waveform in). ## Usage ```py import onnx_asr model = onnx_asr.load_model("OpenVoiceOS/wav2vec2-base-10k-voxpopuli-ft-hu-onnx") print(model.recognize("test.wav")) ``` ## Files * `model.onnx` / `model.onnx.data` — fp32 ONNX graph (inputs: `input_values` (batch, samples) float32, `input_lengths` (batch,) int64; output: `logprobs` (batch, frames, vocab) float32 log-softmax). * `vocab.txt` — CTC vocabulary in onnx-asr's `token id` format (word-delimiter -> `▁`, pad token -> `<blk>`). * `config.json` — `{"model_type": "wav2vec2-ctc", "subsampling_factor": 320}`. No int8 quantized variant is included yet -- `onnxruntime.quantization` does not currently support the `torch.onnx` dynamo-exported graph for this architecture. ## License **CC BY-NC 4.0 — non-commercial use only.** Inherited from the base model (`facebook/wav2vec2-base-10k-voxpopuli-ft-hu`, Meta AI / FAIR). See the [license text](https://creativecommons.org/licenses/by-nc/4.0/) for full terms.
Source context: 12 downloads · 0 likes · Pipeline automatic-speech-recognition · Repo OpenVoiceOS/wav2vec2-base-10k-voxpopuli-ft-hu-onnx