This model was trained by wyz based on the universalsev1 recipe in espnet. More information can be found at
Modellquelle
Quellenauszug
This model was trained by wyz based on the universalsev1 recipe in espnet. More information can be found at
Quellen
1 QuelleVerifiziert 19. Sept.
Modellartefakte
2 Artefakteexp_vctk_dns20_whamr/enh_tfgridnet_xxtiny_raw/83epoch.pth
pth · 492 KB · SHA-256 6d0e66630d6d…23e8 · Hugging Face
Herunterladenexp_vctk_dns20_whamr/enh_tfgridnet_xxtiny_raw/valid.loss.best.pth
pth · 11 B · Hugging Face
Quellenauszüge
2 Auszüge--- tags: - espnet - audio - audio-to-audio language: en datasets: - VCTK_DEMAND - DNS2020 - WHAMR license: cc-by-4.0 --- ## ESPnet2 ENH model ### `wyz/vctk_dns2020_whamr_tfgridnet_xxtiny` This model was trained by wyz based on the universal_se_v1 recipe in [espnet](https://github.com/espnet/espnet/). More information can be found at https://github.com/Emrys365/se-scaling. ### Demo: How to use in ESPnet2 Follow the [ESPnet installation instructions](https://espnet.github.io/espnet/installation.html) if you haven't done that already. To use the model in the Python interface, you could use the following code: ```python import soundfile as sf from espnet2.bin.enh_inference import SeparateSpeech # For model downloading + loading model = SeparateSpeech.from_pretrained( model_tag="wyz/vctk_dns2020_whamr_tfgridnet_xxtiny", normalize_output_wav=True, device="cuda", ) # For loading a downloaded model # model = SeparateSpeech( # train_config="exp_vctk_dns20_whamr/enh_tfgridnet_xxtiny_raw/config.yaml", # model_file="exp_vctk_dns20_whamr/enh_tfgridnet_xxtiny_raw/xxxx.pth", # normalize_output_wav=True, # device="cuda", # ) audio, fs = sf.read("/path/to/noisy/utt1.flac") enhanced = model(audio[None, :], fs=fs)[0] ``` <!-- Generated by ./scripts/utils/show_enh_score.sh --> # RESULTS ## Environments - date: `Sun Mar 3 22:03:37 EST 2024` - python version: `3.8.16 (default, Mar 2 2023, 03:21:46) [GCC 11.2.0]` - espnet version: `espnet 202304` - pytorch version: `pytorch 2.0.1+cu118` - Git hash: `443028662106472c60fe8bd892cb277e5b488651` - Commit date: `Thu May 11 03:32:59 2023 +0000` ## enhanced_test_16k |dataset|PESQ_WB|STOI|SAR|SDR|SIR|SI_SNR|OVRL|SIG|BAK|P808_MOS| |---|---|---|---|---|---|---|---|---|---|---| |chime4_et05_real_isolated_6ch_track|1.18|52.77|-3.19|-3.19|0.00|-31.84|2.75|3.06|3.75|3.47| |chime4_et05_simu_isolated_6ch_track|1.42|81.92|8.23|8.23|0.00|1.62|2.63|3.00|3.64|3.13| |dns20_tt_synthetic_no_reverb|2.82|96.64|17.99|17.99|0.00|17.85|3.24|3.52|4.01|3.93| |reverb_et_real_8ch_multich|1.13|70.16|3.29|3.29|0.00|0.91|2.99|3.30|3.91|3.76| |reverb_et_simu_8ch_multich|2.12|92.45|10.33|10.33|0.00|-8.65|3.12|3.42|3.96|3.84| |whamr_tt_mix_single_reverb_max_16k|1.84|90.09|8.74|8.74|0.00|5.98|3.01|3.30|3.95|3.57| ## enhanced_test_48k |dataset|STOI|SAR|SDR|SIR|SI_SNR|OVRL|SIG|BAK|P808_MOS| |---|---|---|---|---|---|---|---|---|---| |vctk_noisy_tt_2sp...
Source context: 6 downloads · 0 likes · Pipeline audio-to-audio · Library espnet · Repo wyz/vctk_dns2020_whamr_tfgridnet_xxtiny