Real-time RE-USE is a unified, real-time universal speech enhancement framework that explicitly controls both algorithmic and computational latency within a single model. The proposed framework supports 30 distinct...
Model source
Source excerpt
Real-time RE-USE is a unified, real-time universal speech enhancement framework that explicitly controls both algorithmic and computational latency within a single model. The proposed framework supports 30 distinct...
Sources
1 sourceVerified Aug 2
Model artifacts
1 artifactSource excerpts
2 excerpts--- license: other track_downloads: true pipeline_tag: audio-to-audio library_name: mamba-ssm tags: - streaming speech-enhancement - speech-enhancement - universal speech enhancement - multiple input sampling rates - language-agnostic --- # **<span style="color:#76b900;"> 🚀+🤫 Real-time RE-USE: Real-time Multilingual Universal Speech Enhancement</span>** # Model Overview ## Description <p align="center"> <img src="./img/fig1.png" width="80%"/> <p> Real-time RE-USE is a unified, real-time universal speech enhancement framework that explicitly controls both **algorithmic** and **computational** latency within a single model. The proposed framework supports **30** distinct latency configurations while maintaining performance close to specialized models, making it easy to adapt to **different latency budgets**. <p align="center"> <img src="./img/fig2.png" width="80%"/> <p> For **Non Real-time applications**, check this model => [RE-USE](https://huggingface.co/nvidia/RE-USE) In universal speech enhancement, the goal is to restore the **quality** of diverse degraded speech while preserving **fidelity**, ensuring that all other factors remain unchanged, e.g., linguistic content, speaker identity, emotion, accent, and other paralinguistic attributes. Inspired by the **distortion–perception trade-off theory**, our proposed single model achieves a good balance between these two objectives and has the following desirable properties: - Robustness to **diverse degradations**, including additive noise, reverberation, clipping, bandwidth limitation, codec artifacts, packet loss and low-quality mics . - Support for **multiple input sampling rates**, including 8, 16, 22.05, 24, 32, 44.1, and 48 kHz. - Strong **language-agnostic** capability, enabling effective performance across different languages. This model is for research and development only. ## Usage Directly try our [**Gradio Interactive Demo**](https://huggingface.co/spaces/nvidia/Real-time_RE-USE) by uploading your noisy audio/video !! !! Note that this demo page uses offline inference, so you can easily check the audio quality not the latency. Please refer to `online_inference.py` and `online_inference.sh` for streaming inference. ## Environment Setup 1. (For **Mamba** setup)Pre-built Docker environments can be downloaded [here](https://github.com/RoyChao19477/SEMamba?tab=readme-ov-file#-docker-sup...
Source context: 569 downloads · 31 likes · Pipeline audio-to-audio · Library mamba-ssm · Repo nvidia/Real-time_RE-USE