CoreML conversion of WeSpeaker ResNet34-LM for Apple Neural Engine.
Fuente del modelo
Descripción de la fuente
CoreML conversion of WeSpeaker ResNet34-LM for Apple Neural Engine.
Produces 256-dimensional L2-normalized speaker embeddings from audio.
Fuentes
1 fuenteVerificado 29 ago
Artefactos del modelo
3 artefactosExtractos de fuentes
2 extractos| Detail | Value |
|---|
| Architecture | ResNet34 with statistics pooling |
| Parameters | ~6.6M |
| Input | 80-bin log-mel spectrogram (16kHz) |
| Output | 256-dim L2-normalized speaker embedding |
| BatchNorm | Fused into Conv2d at conversion time |
let model = try await WeSpeakerModel.fromPretrained(backend: .coreML)
let embedding = model.embed(audio: samples, sampleRate: 16000)
let similarity = WeSpeakerModel.cosineSimilarity(embeddingA, embeddingB)
| Variant | Backend | Model ID |
|---|---|---|
| MLX | GPU | aufklarer/WeSpeaker-ResNet34-LM-MLX |
| CoreML | Neural Engine | aufklarer/WeSpeaker-ResNet34-LM-CoreML |
wespeaker.mlmodelc/weights/weight.bin
bin · 12,7 MB · SHA-256 6dba18a57a81…d688 · Hugging Face
Descargar--- license: mit tags: - coreml - speaker-embedding - speaker-verification - speaker-diarization - wespeaker - neural-engine base_model: pyannote/wespeaker-voxceleb-resnet34-LM pipeline_tag: audio-classification --- # WeSpeaker-ResNet34-LM — CoreML CoreML conversion of [WeSpeaker ResNet34-LM](https://huggingface.co/pyannote/wespeaker-voxceleb-resnet34-LM) for Apple Neural Engine. Produces 256-dimensional L2-normalized speaker embeddings from audio. ## Model Details | Detail | Value | |--------|-------| | Architecture | ResNet34 with statistics pooling | | Parameters | ~6.6M | | Input | 80-bin log-mel spectrogram (16kHz) | | Output | 256-dim L2-normalized speaker embedding | | BatchNorm | Fused into Conv2d at conversion time | ## Usage ```swift let model = try await WeSpeakerModel.fromPretrained(backend: .coreML) let embedding = model.embed(audio: samples, sampleRate: 16000) let similarity = WeSpeakerModel.cosineSimilarity(embeddingA, embeddingB) ``` ## Variants | Variant | Backend | Model ID | |---------|---------|----------| | MLX | GPU | [aufklarer/WeSpeaker-ResNet34-LM-MLX](https://huggingface.co/aufklarer/WeSpeaker-ResNet34-LM-MLX) | | **CoreML** | **Neural Engine** | **aufklarer/WeSpeaker-ResNet34-LM-CoreML** | ## Links - **Swift library**: [soniqo/speech-swift](https://github.com/soniqo/speech-swift) - **Original model**: [pyannote/wespeaker-voxceleb-resnet34-LM](https://huggingface.co/pyannote/wespeaker-voxceleb-resnet34-LM) --- --- - **Guide**: [soniqo.audio/guides/embed-speaker](https://soniqo.audio/guides/embed-speaker) - **Docs**: [soniqo.audio](https://soniqo.audio) - **GitHub**: [soniqo/speech-swift](https://github.com/soniqo/speech-swift)
Source context: 5733 downloads · 0 likes · Pipeline audio-classification · Repo aufklarer/WeSpeaker-ResNet34-LM-CoreML