ONNX exports of LS-EEND, a long-form streaming end-to-end neural diarization model with online attractor extraction.
Fuente del modelo
Descripción de la fuente
ONNX exports of LS-EEND, a long-form streaming end-to-end neural diarization model with online attractor extraction.
This repository contains non-quantized ONNX step models for four LS-EEND variants:
AMICALLHOMEFuentes
1 fuenteVerificado 6 sept
Artefactos del modelo
4 artefactosExtractos de fuentes
2 extractosDIHARD IIDIHARD IIIThese models are intended for stateful streaming inference. Each package runs one LS-EEND step at a time with explicit recurrent/cache tensors, rather than processing an entire utterance in a single call.
Each variant directory contains:
*.onnx: the ONNX model*.json: metadata needed by the runtimeVariant directories:
AMI/CALLHOME/DIHARD II/DIHARD III/| Variant | Package | Configured max speakers | Model output capacity |
|---|---|---|---|
| AMI | AMI/ls_eend_ami_step.onnx | 4 | 6 |
| CALLHOME | CALLHOME/ls_eend_callhome_step.onnx | 7 | 9 |
| DIHARD II | DIHARD II/ls_eend_dih2_step.onnx | 10 | 12 |
| DIHARD III | DIHARD III/ls_eend_dih3_step.onnx | 10 | 12 |
The metadata JSON distinguishes between:
max_speakers: the dataset/config speaker setting from the LS-EEND infer YAMLmax_nspks: the exported model's full decode/output capacityAll four non-quantized exports in this repo use the same frontend settings:
8000 Hz200 samples80 samples102423710logmel23_cummn10 Hzfloat32These are step-wise streaming models. A runtime must maintain and feed the recurrent state tensors between calls:
enc_ret_kvenc_ret_scaleenc_conv_cachedec_ret_kvdec_ret_scaletop_bufferThe ONNX inputs and outputs follow the LS-EEND step export used by the reference Python and Swift runtimes.
Use these packages with a runtime that:
8 kHzingest/decode control inputs to handle the encoder delay and final tail flushThis repository is not a drop-in replacement for generic Hugging Face transformers inference. It is meant for custom ONNX runtimes, such as:
Setup the virtual environment:
# Create and activate virtual environment
python3.10 -m venv venv
source venv/bin/activate
# Install dependencies
pip install -r example/requirements.txt
To run the Python microphone inference script for the DIHARD III variant, run the following command:
python example/ls_eend_onnx_mic_gui.py --onnx-model DIHARD\ III/ls_eend_dih3_step.onnx
For the other variants, replace DIHARD\ III/ls_eend_dih3_step.onnx with the path to the desired variant:
# AMI variant
python example/ls_eend_onnx_mic_gui.py --onnx-model AMI/ls_eend_ami_step.onnx
# CALLHOME variant
python example/ls_eend_onnx_mic_gui.py --onnx-model CALLHOME/ls_eend_callhome_step.onnx
# DIHARD II variant
python example/ls_eend_onnx_mic_gui.py --onnx-model DIHARD\ II/ls_eend_dih2_step.onnx
Each variant ships a sidecar JSON with fields like:
{
"sample_rate": 8000,
"win_length": 200,
"hop_length": 80,
"n_fft": 1024,
"n_mels": 23,
"context_recp": 7,
"subsampling": 10,
"feat_type": "logmel23_cummn",
"frame_hz": 10.0,
"max_speakers": 10,
"max_nspks": 12
}
Check the variant-specific *.json file for the exact state tensor shapes and output dimensions.
These ONNX exports were produced from the LS-EEND code in the FS-EEND repository:
The export path is based on the LS-EEND ONNX step exporter and variant batch exporter in that project.
From the source project, the reported real-world diarization error rates are:
| Dataset | DER (%) |
|---|---|
| CALLHOME | 12.11 |
| DIHARD II | 27.58 |
| DIHARD III | 19.61 |
| AMI Dev | 20.97 |
| AMI Eval | 20.76 |
These numbers come from the upstream LS-EEND project README and reflect the original training/evaluation setup, not a Hugging Face evaluation pipeline.
The upstream LS-EEND model/codebase used for these ONNX exports is MIT-licensed, and this repository is published as MIT accordingly.
The underlying evaluation and fine-tuning datasets still have their own access and usage terms:
This repository redistributes ONNX exports of the LS-EEND model variants. Dataset licensing and access requirements remain governed by the original dataset providers.
If you use LS-EEND, cite the original paper:
@ARTICLE{11122273,
author={Liang, Di and Li, Xiaofei},
journal={IEEE Transactions on Audio, Speech and Language Processing},
title={LS-EEND: Long-Form Streaming End-to-End Neural Diarization With Online Attractor Extraction},
year={2025},
volume={33},
pages={3568-3581},
doi={10.1109/TASLPRO.2025.3597446}
}
DIHARD II/ls_eend_dih2_step.onnx
onnx · 42,4 MB · SHA-256 5df89a22ba87…8302 · Hugging Face
--- language: - en license: mit library_name: onnx pipeline_tag: audio-classification tags: - speaker-diarization - diarization - onnx - streaming - audio - ls-eend - eend pretty_name: LS-EEND ONNX Models model-index: - name: LS-EEND ONNX Models results: [] --- # LS-EEND ONNX Models ONNX exports of LS-EEND, a long-form streaming end-to-end neural diarization model with online attractor extraction. This repository contains non-quantized ONNX step models for four LS-EEND variants: - `AMI` - `CALLHOME` - `DIHARD II` - `DIHARD III` These models are intended for stateful streaming inference. Each package runs one LS-EEND step at a time with explicit recurrent/cache tensors, rather than processing an entire utterance in a single call. ## Included files Each variant directory contains: - `*.onnx`: the ONNX model - `*.json`: metadata needed by the runtime Variant directories: - `AMI/` - `CALLHOME/` - `DIHARD II/` - `DIHARD III/` ## Variants | Variant | Package | Configured max speakers | Model output capacity | | ---------- | ------------------------------------- | ----------------------: | --------------------: | | AMI | `AMI/ls_eend_ami_step.onnx` | 4 | 6 | | CALLHOME | `CALLHOME/ls_eend_callhome_step.onnx` | 7 | 9 | | DIHARD II | `DIHARD II/ls_eend_dih2_step.onnx` | 10 | 12 | | DIHARD III | `DIHARD III/ls_eend_dih3_step.onnx` | 10 | 12 | The metadata JSON distinguishes between: - `max_speakers`: the dataset/config speaker setting from the LS-EEND infer YAML - `max_nspks`: the exported model's full decode/output capacity ## Frontend and runtime assumptions All four non-quantized exports in this repo use the same frontend settings: - sample rate: `8000 Hz` - window length: `200` samples - hop length: `80` samples - FFT size: `1024` - mel bins: `23` - context receptive field: `7` - subsampling: `10` - feature type: `logmel23_cummn` - output frame rate: `10 Hz` - compute precision: `float32` These are step-wise streaming models. A runtime must maintain and feed the recurrent state tensors between calls: - `enc_ret_kv` - `enc_ret_scale` - `enc_conv_cache` - `dec_ret_kv` - `dec_ret_scale` - `top_buffer` The ONNX inputs and outputs follow the LS-EEND step export used by the reference Python and Swift runtimes. ## Intended usage Use these packages with a runtime that: 1. Resamples audio to mono `8 kHz` 2. Extracts LS-EEND features with the setti...
Source context: 13 downloads · 1 likes · Pipeline audio-classification · Library onnx · Repo GradientDescent2718/LS-EEND-ONNX