Abstract: Speech bandwidth extension (BWE) refers to widening the frequency bandwidth range of speech signals, enhancing the speech quality towards brighter and fuller. This paper proposes a generative adversarial...
Fonte do modelo
Trecho da fonte
Abstract: Speech bandwidth extension (BWE) refers to widening the frequency bandwidth range of speech signals, enhancing the speech quality towards brighter and fuller. This paper proposes a generative adversarial...
Fontes
1 fonteVerificado 5 de set.
Artefatos de modelo
14 artefatosTrechos de fonte
2 trechosweights/16kto48k/g_16kto48k.ckpt
ckpt · 114 MB · SHA-256 83f088f3a3d6…9ae2 · Hugging Face
weights/16kto48k/g_16kto48k.safetensors
safetensors · 114 MB · SHA-256 89471f6d001a…50d9 · Hugging Face
Baixarweights/24kto48k/g_24kto48k.safetensors
safetensors · 114 MB · SHA-256 09af95615d75…62bf · Hugging Face
Baixarweights/2kto16k/g_2kto16k.safetensors
safetensors · 114 MB · SHA-256 c7c311909521…9b07 · Hugging Face
Baixarweights/4kto16k/g_4kto16k.safetensors
safetensors · 114 MB · SHA-256 7ab69cfb6f7a…728c · Hugging Face
Baixarweights/8kto16k/g_8kto16k.safetensors
safetensors · 114 MB · SHA-256 486e7e3d251b…f1c8 · Hugging Face
Baixarweights/8kto48k/g_8kto48k.safetensors
safetensors · 114 MB · SHA-256 e2c2c0bcec68…d121 · Hugging Face
Baixar--- license: mit datasets: - CSTR-Edinburgh/vctk language: - en pipeline_tag: audio-to-audio tags: - audio --- # Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction ### Ye-Xin Lu, Yang Ai, Hui-Peng Du, Zhen-Hua Ling **Abstract:** Speech bandwidth extension (BWE) refers to widening the frequency bandwidth range of speech signals, enhancing the speech quality towards brighter and fuller. This paper proposes a generative adversarial network (GAN) based BWE model with parallel prediction of Amplitude and Phase spectra, named AP-BWE, which achieves both high-quality and efficient wideband speech waveform generation. The proposed AP-BWE generator is entirely based on convolutional neural networks (CNNs). It features a dual-stream architecture with mutual interaction, where the amplitude stream and the phase stream communicate with each other and respectively extend the high-frequency components from the input narrowband amplitude and phase spectra. To improve the naturalness of the extended speech signals, we employ a multi-period discriminator at the waveform level and design a pair of multi-resolution amplitude and phase discriminators at the spectral level, respectively. Experimental results demonstrate that our proposed AP-BWE achieves state-of-the-art performance in terms of speech quality for BWE tasks targeting sampling rates of both 16 kHz and 48 kHz. In terms of generation efficiency, due to the all-convolutional architecture and all-frame-level operations, the proposed AP-BWE can generate 48 kHz waveform samples 292.3 times faster than real-time on a single RTX 4090 GPU and 18.1 times faster than real-time on a single CPU. Notably, to our knowledge, AP-BWE is the first to achieve the direct extension of the high-frequency phase spectrum, which is beneficial for improving the effectiveness of existing BWE methods. **We provide our implementation as open source in this repository. Audio samples can be found at the [demo website](http://yxlu-0102.github.io/AP-BWE).** ## License The code in this repository is licensed under the MIT License.<br> The weights for this model are stored in [Google Drive](https://drive.google.com/drive/folders/1IIYTf2zbJWzelu4IftKD6ooHloJ8mnZF?usp=share_link). You can access and use the weights under the same MIT License. Please refer to the [weights_LICENSE.txt](weight...
Source context: 0 downloads · 1 likes · Pipeline audio-to-audio · Repo rsxdalv/AP-BWE