ECAPA-TDNN language identification model converted to MLX safetensors format for inference on Apple Silicon (Metal GPU).
Source du modèle
Description de la source
ECAPA-TDNN language identification model converted to MLX safetensors format for inference on Apple Silicon (Metal GPU).
Identifies 107 languages from audio. Original model: speechbrain/lang-id-voxlingua107-ecapa.
Sources
1 sourceVérifié 13 sept.
Artefacts du modèle
1 artefactecapa_tdnn_lid107.safetensors
safetensors · 81,2 MB · SHA-256 bae5627c78e9…dbc8 · Hugging Face
TéléchargerExtraits de sources
2 extraits| Framework | Latency | Russian | English |
|---|
Python MLX (with mx.compile()) | 15.7ms | 99.6% | 99.9% |
Swift MLX (with compile() + GPU mel) | 14.8ms | 99.5% | 99.9% |
| CoreML GPU | 17ms | 99.7% | 98.6% |
# Clone the benchmark repo
git clone https://github.com/beshkenadze/lid-bench
cd lid-bench/mlx
# Setup
uv venv && uv pip install mlx numpy soundfile safetensors
# Run
python ecapa_tdnn_lid.py path/to/audio.wav --benchmark
Full implementation: github.com/beshkenadze/lid-bench
.ckpt → Conv1d axis swap → safetensorsWeights were converted from the original SpeechBrain checkpoint:
[out, in, kernel] → [out, kernel, in] (MLX convention)Convert script: mlx/convert_ecapa_weights.py in the lid-bench repo.
--- license: apache-2.0 tags: - mlx - audio-classification - language-identification - ecapa-tdnn - apple-silicon language: - multilingual base_model: speechbrain/lang-id-voxlingua107-ecapa pipeline_tag: audio-classification --- # ECAPA-TDNN VoxLingua107 — MLX ECAPA-TDNN language identification model converted to MLX safetensors format for inference on Apple Silicon (Metal GPU). Identifies **107 languages** from audio. Original model: [speechbrain/lang-id-voxlingua107-ecapa](https://huggingface.co/speechbrain/lang-id-voxlingua107-ecapa). ## Performance (M1, 10s audio) | Framework | Latency | Russian | English | |---|---|---|---| | **Python MLX** (with `mx.compile()`) | **15.7ms** | 99.6% | 99.9% | | **Swift MLX** (with `compile()` + GPU mel) | **14.8ms** | 99.5% | 99.9% | | CoreML GPU | 17ms | 99.7% | 98.6% | ## Usage ```python # Clone the benchmark repo git clone https://github.com/beshkenadze/lid-bench cd lid-bench/mlx # Setup uv venv && uv pip install mlx numpy soundfile safetensors # Run python ecapa_tdnn_lid.py path/to/audio.wav --benchmark ``` Full implementation: [github.com/beshkenadze/lid-bench](https://github.com/beshkenadze/lid-bench) ## Model Details - **Architecture**: ECAPA-TDNN (channels=1024, embed_dim=256) - **Input**: Log-mel spectrogram (60 mels, periodic Hamming window, SpeechBrain-compatible) - **Output**: 107 language probabilities - **Parameters**: 20M - **Weight format**: safetensors (81.2 MB) - **Conversion**: SpeechBrain `.ckpt` → Conv1d axis swap → safetensors ## Weight Conversion Weights were converted from the original SpeechBrain checkpoint: - Conv1d axes swapped: `[out, in, kernel]` → `[out, kernel, in]` (MLX convention) - BatchNorm running stats preserved - No quantization applied Convert script: `mlx/convert_ecapa_weights.py` in the [lid-bench repo](https://github.com/beshkenadze/lid-bench).
Source context: 46 downloads · 0 likes · Pipeline audio-classification · Library mlx · Repo beshkenadze/lang-id-voxlingua107-ecapa-mlx