Developed by: Cnam-LMSSC Model: EBEN(M=4,P=4,Q=4) (see publication in IEEE TASLP - arXiv link) Language: French License: MIT Training dataset: mix(speechclean+speechlessnoisy) of Cnam-LMSSC/vibravox (see VibraVox...
Model source
Source excerpt
Developed by: Cnam-LMSSC Model: EBEN(M=4,P=4,Q=4) (see publication in IEEE TASLP - arXiv link) Language: French License: MIT Training dataset: mix(speechclean+speechlessnoisy) of Cnam-LMSSC/vibravox (see VibraVox...
Sources
1 sourceVerified Sep 12
Model artifacts
1 artifactSource excerpts
2 excerpts--- language: fr license: mit library_name: transformers tags: - audio - audio-to-audio - speech datasets: - Cnam-LMSSC/vibravox model-index: - name: EBEN(M=4,P=4,Q=4) results: - task: type: speech-enhancement name: Bandwidth Extension dataset: name: Vibravox["forehead_accelerometer"] type: Cnam-LMSSC/vibravox args: fr metrics: - type: stoi value: 0.971 name: Test SQUIM-STOI, in-domain training - type: n-mos value: 4.20 name: Test Noresqa-MOS, in-domain training --- <p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/65302a613ecbe51d6a6ddcec/zhB1fh-c0pjlj-Tr4Vpmr.png" style="object-fit:contain; width:280px; height:280px;" > </p> # Model Card - **Developed by:** [Cnam-LMSSC](https://huggingface.co/Cnam-LMSSC) - **Model:** [EBEN(M=4,P=4,Q=4)](https://github.com/jhauret/vibravox/blob/main/vibravox/torch_modules/dnn/eben_generator.py) (see [publication in IEEE TASLP](https://ieeexplore.ieee.org/document/10244161) - [arXiv link](https://arxiv.org/abs/2303.10008)) - **Language:** French - **License:** MIT - **Training dataset:** mix(`speech_clean`+`speechless_noisy`) of [Cnam-LMSSC/vibravox](https://huggingface.co/datasets/Cnam-LMSSC/vibravox) (see [VibraVox paper on arXiV](https://arxiv.org/abs/2407.11828)) - **Testing dataset:** `speech_noisy` of [Cnam-LMSSC/vibravox](https://huggingface.co/datasets/Cnam-LMSSC/vibravox) - **Samplerate for usage:** 16kHz ## Overview This bandwidth extension model, trained on [Vibravox](https://huggingface.co/datasets/Cnam-LMSSC/vibravox) body conduction sensor data, enhances body-conducted speech audio by denoising and regenerating mid and high frequencies from low-frequency content. ## Disclaimer This model, trained for **a specific non-conventional speech sensor**, is intended to be used with **in-domain data**. Using it with other sensor data may lead to suboptimal performance. ## Link to BWE models trained on other body conducted sensors : The entry point to all EBEN models for Bandwidth Extension (BWE) is available at [https://huggingface.co/Cnam-LMSSC/vibravox_EBEN_models](https://huggingface.co/Cnam-LMSSC/vibravox_EBEN_models). ## Training procedure Detailed instructions for reproducing the experiments are available on the [jhauret/vibravox](https://github.com/jhauret/vibravox) Github repository. ## Inference script : ```python import torch, torchaudio from vibr...
Source context: 11 downloads · 0 likes · Pipeline audio-to-audio · Library transformers · Repo Cnam-LMSSC/EBEN_noisy_forehead_accelerometer