| Quantization | Perplexity | ± (Std. Err.) | ΔPPL vs Q80 | | Q40 | 10.6926 | 0.05953 | +0.2354 | | Q41 | 10.2493 | 0.05640 | -0.2079 | | MXFP4MOE | 10.4572 | 0.05838 | +0.0000 | | Q50 | 10.6152 | 0.05954 | +0.1580 |...
Source du modèle
Extrait de la source
| Quantization | Perplexity | ± (Std. Err.) | ΔPPL vs Q80 | | Q40 | 10.6926 | 0.05953 | +0.2354 | | Q41 | 10.2493 | 0.05640 | -0.2079 | | MXFP4MOE | 10.4572 | 0.05838 | +0.0000 | | Q50 | 10.6152 | 0.05954 | +0.1580 |...
Sources
1 sourceVérifié 7 août
Artefacts du modèle
16 artefactsExtraits de sources
2 extraitsmusic-flamingo-IQ4_NL.gguf
gguf · 4,15 GB · SHA-256 01dcda91d6e9…f47a · Hugging Face
--- license: other language: - en base_model: - nvidia/music-flamingo-hf pipeline_tag: audio-text-to-text library_name: transformers tags: - music/songs - music - music reasoning - music understanding datasets: - nvidia/MF-Skills --- # GGUF quantization perplexity: | Quantization | Perplexity | ± (Std. Err.) | ΔPPL vs Q8_0 | |---|---:|---:|---:| | Q4_0 | 10.6926 | 0.05953 | +0.2354 | | Q4_1 | 10.2493 | 0.05640 | -0.2079 | | MXFP4_MOE | 10.4572 | 0.05838 | +0.0000 | | Q5_0 | 10.6152 | 0.05954 | +0.1580 | | Q5_1 | 10.3985 | 0.05776 | -0.0587 | | Q2_K | 12.0840 | 0.06471 | +1.6268 | | Q3_K | 10.6406 | 0.05863 | +0.1834 | | Q3_K_S | 11.6617 | 0.06428 | +1.2045 | | Q3_K_M | 10.6406 | 0.05863 | +0.1834 | | Q3_K_L | 10.6608 | 0.05917 | +0.2036 | | IQ4_NL | 10.9692 | 0.06225 | +0.5120 | | IQ4_XS | 11.0113 | 0.06249 | +0.5541 | | Q6_K | 10.4604 | 0.05829 | +0.0032 | | Q8_0 | 10.4572 | 0.05838 | +0.0000 | *Perplexity tested using a random sample from the music bench captions dataset.* # LLAMACPP inference Download one of the quantizations of the text-text part, then download the mmproj audio processing model (in full BF16 precision). Ensure you are using the latest version of llama.cpp (or at least a release after the PR tag: b7593). Run the model specifying the path of both the `--mproj` and `--model`. # Model Overview <div align="center" style="display: flex; justify-content: center; align-items: center; text-align: center;"> <a href="https://github.com/NVIDIA/audio-flamingo" style="margin-right: 20px; text-decoration: none; display: flex; align-items: center;"> <img src="static/mf_logo.png" alt="Music Flamingo 🔥🚀🔥" width="120"> </a> </div> <div align="center" style="display: flex; justify-content: center; align-items: center; text-align: center;"> <h2> Music Flamingo: Scaling Music Understaning in Audio Language Models </h2> </div> <div align="center" style="display: flex; justify-content: center; margin-top: 10px;"> <a href="https://arxiv.org/abs/2511.10289"><img src="https://img.shields.io/badge/arXiv-2511.10289-AD1C18" style="margin-right: 5px;"></a> <a href="https://research.nvidia.com/labs/adlr/MF/"><img src="https://img.shields.io/badge/Demo page-228B22" style="margin-right: 5px;"></a> <a href="https://github.com/NVIDIA/audio-flamingo"><img src='https://img.shields.io/badge/Github-Audio Flamingo 3-9C276A' style="margin-right: 5px;"></a...
Source context: 273 downloads · 5 likes · Pipeline audio-text-to-text · Library transformers · Repo henry1477/music-flamingo-gguf