ONNX export of [openai/whisper-large-v3][base] in the [transformers.js][tjs] layout, with the cross-attention alignment heads baked into generationconfig.json so the decoder emits word-level timestamps...
Model source
Source description
ONNX export of in the layout, with the baked into so the decoder emits (). Packaged for the runtime ( on the execution provider).
Sources
1 sourceVerified Sep 6
Model artifacts
1 artifactSource excerpts
2 excerptsgeneration_config.jsonreturn_timestamps: 'word'packages/aionnx-community publishes such _timestamped variants for
turbo, small and tiny, but not for
large-v3 — hence this export. It is produced by scripts/onnx/whisper
in musetric-toolkit (optimum + the transformers.js converter).
| file | role |
|---|---|
encoder_model_q4.onnx | audio encoder (q4) |
decoder_model_merged_q4.onnx | decoder with merged KV-cache (q4) |
config.json, generation_config.json, preprocessor_config.json | model / feature-extractor config |
tokenizer.json, tokenizer_config.json, vocab.json, merges.txt, added_tokens.json, special_tokens_map.json, normalizer.json | tokenizer |
Only the q4 graphs are shipped; the runtime loads dtype: 'q4' for both the
encoder and the merged decoder.
import { pipeline } from '@huggingface/transformers';
const asr = await pipeline(
'automatic-speech-recognition',
'musetric/whisper-large-v3-onnx',
{
device: 'webgpu',
subfolder: '',
dtype: { encoder_model: 'q4', decoder_model_merged: 'q4' },
},
);
const out = await asr(audio, { return_timestamps: 'word', chunk_length_s: 30 });
Intended: client/edge speech-to-text with word timings via WebGPU through
@huggingface/transformers — e.g. lyric alignment in musetric.
Limitations:
shader-f16.musetric pipeline adds language
forcing and loop guards on top).Documented only as far as it is verifiable.
openai/whisper-large-v3
@ 06f233fe06e710322aca913c1bc4249a0d71fce1. Inference-only re-export in the
transformers.js layout; no fine-tuning.[layer, head] pairs correlated with word timing come
from hollance's gist, which covers tiny through large-v2. large-v3
is absent there, so its heads are taken from the
generation_config.json of the base model itself.transformers==4.42.4 / optimum==1.21.3 /
torch==2.4.1 (transformers 4.43 adds a cache_position decoder input that
@huggingface/transformers 4.2.0 does not feed).scripts/onnx/whisper in musetric-toolkit.Apache-2.0, following the model card the weights are downloaded from.
Upstream is inconsistent here and it is worth knowing: the
openai/whisper repository states that "Whisper's code and model weights are
released under the MIT License", while the Hugging Face model card these weights
are actually fetched from declares apache-2.0. This export follows the source
it downloads from; consult both before relying on either.
--- license: apache-2.0 library_name: transformers.js pipeline_tag: automatic-speech-recognition base_model: openai/whisper-large-v3 tags: - audio - automatic-speech-recognition - whisper - onnx - transformers.js - webgpu --- # Whisper large-v3 — word-timestamped ONNX (q4 / WebGPU) ONNX export of [**openai/whisper-large-v3**][base] in the [transformers.js][tjs] layout, with the **cross-attention alignment heads** baked into `generation_config.json` so the decoder emits **word-level timestamps** (`return_timestamps: 'word'`). Packaged for the [`musetric`][musetric] `packages/ai` runtime ([`@huggingface/transformers`][tjs] on the **WebGPU** execution provider). [onnx-community][oc] publishes such `_timestamped` variants for [turbo][oc-turbo], [small][oc-small] and [tiny][oc-tiny], but not for **large-v3** — hence this export. It is produced by [`scripts/onnx/whisper`][sc] in [musetric-toolkit][toolkit] (optimum + the [transformers.js converter][conv]). ## Files | file | role | |---|---| | `encoder_model_q4.onnx` | audio encoder (q4) | | `decoder_model_merged_q4.onnx` | decoder with merged KV-cache (q4) | | `config.json`, `generation_config.json`, `preprocessor_config.json` | model / feature-extractor config | | `tokenizer.json`, `tokenizer_config.json`, `vocab.json`, `merges.txt`, `added_tokens.json`, `special_tokens_map.json`, `normalizer.json` | tokenizer | Only the **q4** graphs are shipped; the runtime loads `dtype: 'q4'` for both the encoder and the merged decoder. ## How to use ```ts import { pipeline } from '@huggingface/transformers'; const asr = await pipeline( 'automatic-speech-recognition', 'musetric/whisper-large-v3-onnx', { device: 'webgpu', subfolder: '', dtype: { encoder_model: 'q4', decoder_model_merged: 'q4' }, }, ); const out = await asr(audio, { return_timestamps: 'word', chunk_length_s: 30 }); ``` ## Intended uses & limitations **Intended:** client/edge speech-to-text with word timings via WebGPU through [`@huggingface/transformers`][tjs] — e.g. lyric alignment in `musetric`. **Limitations:** - Requires a WebGPU adapter with `shader-f16`. - q4 weight quantization trades a little accuracy for size/speed; for non-English audio quality varies (the `musetric` pipeline adds language forcing and loop guards on top). - Word timestamps come from the whisper cross-attention heads, not a separate forced aligner. ## Sou...
Source context: 11 downloads · 0 likes · Pipeline automatic-speech-recognition · Library transformers.js · Repo musetric/whisper-large-v3-onnx