Official Repository of Paper: "Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios"(AAAI 2026)
Modellquelle
Quellenauszug
Official Repository of Paper: "Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios"(AAAI 2026)
Quellen
1 QuelleVerifiziert 3. Aug.
Modellartefakte
1 Artefaktutils/pretrain/250000_step_val_loss_0.50.pth
pth · 230 MB · SHA-256 c32005b52735…3653 · Hugging Face
HerunterladenQuellenauszüge
2 Auszüge--- tags: - singing - svc - speech - synthesis - aigc - super-resolution license: apache-2.0 pipeline_tag: audio-to-audio --- # HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios Official Repository of Paper: "Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios"(AAAI 2026) <div align="center"> <p> <img src="images/kon-new.gif" alt="HQ-SVC Logo" width="300"> </p> <a href="https://arxiv.org/abs/2511.08496"><img src="https://img.shields.io/badge/arXiv-2511.08496-b31b1b.svg?logo=arxiv&logoColor=white" alt="arXiv"></a> <a href="https://shawnpi233.github.io/HQ-SVC-demo"><img src="https://img.shields.io/badge/Demos-🌐-blue" alt="Demos"></a> <a href="https://huggingface.co/shawnpi/HQ-SVC"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Models%20-%20Access-orange" alt="Models Access"></a> <a href="https://github.com/ShawnPi233/HQ-SVC" target="_blank" rel="noopener noreferrer"> <img src="https://img.shields.io/badge/GitHub-Repository-blue?logo=github" alt="GitHub Repository"></a> </div> HQ-SVC is an efficient framework for high-quality zero-shot singing voice conversion (SVC) in low-resource scenarios. It achieves disentanglement of content and speaker features via a unified decoupled codec, and enhances synthesis quality through multi-feature fusion and progressive optimization. Unlike existing methods that demand large datasets or heavy computational resources, **HQ-SVC** unifies: - 🚀 Zero-shot conversion for unseen speakers without fine-tuning - ⚡ Low-resource training (single consumer-grade GPU, <80h data) - 🎧 Dual capabilities: high-quality singing voice conversion + voice super-resolution - 🎯 Superior naturalness and speaker similarity compared to SOTA methods ## 🗞 News - **[2025-11-08]** 🎉 Paper accepted by AAAI 2026 - **[2025-11-12]** 🎉 arXiv paper released - **[2025-11-12]** 🎉 Demo released - **[2025-12-24]** 🎉 Inference codes and pre-trained models released ## 📅 Release Plan - [x] arXiv preprint - [x] Online demo - [x] Inference codes - [x] Pre-trained models - [ ] Training codes ## ✨ New features - [ ] Singing style control - [ ] Improved quality ## 🎸 Try Inference ### 1. Download Codes and Environment(下载代码和环境) * Tested only on Linux platforms with CUDA >= 11.8 (仅在 Linux 平台、CUDA >= 11.8 的环境上测试通过) * Windows users can use WSL (Ubuntu) for deployment and ex...
Source context: 60 downloads · 10 likes · Pipeline audio-to-audio · Repo shawnpi/HQ-SVC