Official PyTorch Implementation of ICASSP 2026 paper "HEAR: Hierarchically Enhanced Aesthetic Representations for Multidimensional Music Evaluation"
Source du modèle
Extrait de la source
Official PyTorch Implementation of ICASSP 2026 paper "HEAR: Hierarchically Enhanced Aesthetic Representations for Multidimensional Music Evaluation"
Sources
1 sourceVérifié 19 sept.
Artefacts du modèle
2 artefactsExtraits de sources
2 extraits--- license: apache-2.0 pipeline_tag: audio-classification tags: - music - song - aesthetics - ASAE --- # **HEAR**: Hierarchically Enhanced Aesthetic Representations for Multidimensional Music Evaluation [**Paper**](https://arxiv.org/pdf/2511.18869) | [**Model**](https://huggingface.co/earlab/EAR_HEAR) <br> Official PyTorch Implementation of ICASSP 2026 paper "HEAR: Hierarchically Enhanced Aesthetic Representations for Multidimensional Music Evaluation" This repository contains the training and evaluation code for HEAR, a robust framework designed to address the challenges of multidimensional music aesthetic evaluation under limited labeled data.  ## 🌟 Key Features * **Excellent Performance**: Ranked 2nd/19 on Track 1 and 5th/17 on Track 2 in the [ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge](https://aslp-lab.github.io/Automatic-Song-Aesthetics-Evaluation-Challenge/). * **Robustness**: Synergizes Multi-Source Multi-Scale Representations and Hierarchical Augmentation to capture robust features under limited labeled data. * **Dual Capability**: Optimized for both exact score prediction and ranking (Top-Tier Song Identification). ## 📦 Installation Clone the repository and install dependencies: ``` git clone https://github.com:Eps-Acoustic-Revolution-Lab/EAR_HEAR.git git submodule update --init --recursive conda create -n hear python=3.10 -y conda activate hear pip install -r requirements.txt ``` ## 🚀 Quick Start ``` # Download pretrained model weights export HF_ENDPOINT=https://hf-mirror.com # For users in Mainland China, this is needed for HuggingFace downloads hf download earlab/EAR_HEAR --local-dir pretrained_models # Track 1: Single-Label Inference (Musicality) python inference.py \ --input_audio_path data_pipeline/origin_song_eval_dataset/mp3/0.mp3 \ --output_json_path output.json --model_path pretrained_models/track_1.pth \ --model_config_path config_track_1.yaml # Track 2: Multi-Label Inference (5 Dimensions) python inference.py \ --input_audio_path data_pipeline/origin_song_eval_dataset/mp3/0.mp3 \ --output_json_path output.json --model_path pretrained_models/track_2.pth \ --model_config_path config_track_2.yaml ``` ## 🎯 Training ### Step 1: Data Preparation First, prepare the dataset by running the data pipeline: ```bash cd data_pipeline bash run.sh ``` This script will: 1. **Download Dataset**: Do...
Source context: 0 downloads · 4 likes · Pipeline audio-classification · Repo earlab/EAR_HEAR