Quantized GGUFs of LongCat-Video-Avatar for ComfyUI + WanVideoWrapper
Modellquelle
Quellenbeschreibung
Quantized GGUFs of LongCat-Video-Avatar for ComfyUI + WanVideoWrapper
Original model Link: https://huggingface.co/meituan-longcat/LongCat-Video-Avatar
Watch us at Youtube: @VantageWithAI
Quellen
1 QuelleVerifiziert 16. Juli
Modellartefakte
1 Artefaktmulti/LongCat-Avatar-Multi_comfy-Q2_K.gguf
gguf · 7,74 GB · SHA-256 4e63875d0ca1…dd56 · Hugging Face
HerunterladenQuellenauszüge
3 AuszügeWe are excited to announce the release of LongCat-Video-Avatar, a unified model that delivers expressive and highly dynamic audio-driven character animation, supporting native tasks including Audio-Text-to-Video, Audio-Text-Image-to-Video, and Video Continuation with seamless compatibility for both single-stream and multi-stream audio inputs.
For more detail, please refer to the comprehensive LongCat-Video-Avatar Technical Report.
--> The following videos showcase example generations from our model and have been compressed for easier viewing.
Human evaluation on naturalness and realism of the synthesized videos. The benchmark EvalTalker [1] contains more than 400 testing samples with different difficulty levels for evaluating the single and multiple human video generation.
Reference:
[1] Zhou Y, Zhu X, Ren S, et al. EvalTalker: Learning to Evaluate Real-Portrait-Driven Multi-Subject Talking Humans[J]. arXiv preprint arXiv:2512.01340, 2025.
The model weights are released under the MIT License.
Any contributions to this repository are licensed under the MIT License, unless otherwise stated. This license does not grant any rights to use Meituan trademarks or patents.
See the LICENSE file for the full license text.
This model has not been specifically designed or comprehensively evaluated for every possible downstream application.
Developers should take into account the known limitations of large language models, including performance variations across different languages, and carefully assess accuracy, safety, and fairness before deploying the model in sensitive or high-risk scenarios. It is the responsibility of developers and downstream users to understand and comply with all applicable laws and regulations relevant to their use case, including but not limited to data protection, privacy, and content safety requirements.
Nothing in this Model Card should be interpreted as altering or restricting the terms of the MIT License under which the model is released.
We kindly encourage citation of our work if you find it useful.
@misc{meituanlongcatteam2025longcatvideoavatartechnicalreport,
title={LongCat-Video-Avatar Technical Report},
author={Meituan LongCat Team},
year={2025},
eprint={},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={},
}
We would like to thank the contributors to the Wan, UMT5-XXL, Diffusers and HuggingFace repositories, for their open research.
Please contact us at longcat-team@meituan.com or join our WeChat Group if you have any questions.
--- license: mit language: - en - zh library_name: diffusers tags: - audio-text-to-video - audio-image-text-to-video - audio-driven-video-continuation - diffusers - transformers - avatar - video-generation --- **Quantized GGUFs of LongCat-Video-Avatar for ComfyUI + WanVideoWrapper** **Original model Link:** [https://huggingface.co/meituan-longcat/LongCat-Video-Avatar](https://huggingface.co/meituan-longcat/LongCat-Video-Avatar) **Watch us at Youtube:** [@VantageWithAI](https://www.youtube.com/@vantagewithai) # LongCat-Video-Avatar <div align="center"> <img src="https://huggingface.co/meituan-longcat/LongCat-Video-Avatar/resolve/main/assets/longcat_video_avatar_logo.svg" width="45%" alt="LongCat-Video" /> </div> <hr> <div align="center" style="line-height: 1;"> <a href='https://meigen-ai.github.io/LongCat-Video-Avatar/'><img src='https://img.shields.io/badge/Project-Page-green'></a> <a href='https://github.com/meituan-longcat/LongCat-Video/blob/main/assets/LongCat-Video-Avatar-Tech-Report.pdf'><img src='https://img.shields.io/badge/Technique-Report-red'></a> <a href='https://huggingface.co/meituan-longcat/LongCat-Video-Avatar'><img src='https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Model-blue'></a> </div> <div align="center" style="line-height: 1;"> <a href='https://github.com/meituan-longcat/LongCat-Flash-Chat/blob/main/figures/wechat_official_accounts.png'><img src='https://img.shields.io/badge/WeChat-LongCat-brightgreen?logo=wechat&logoColor=white'></a> <a href='https://x.com/Meituan_LongCat'><img src='https://img.shields.io/badge/Twitter-LongCat-white?logo=x&logoColor=white'></a> </div> <div align="center" style="line-height: 1;"> <a href='LICENSE'><img src='https://img.shields.io/badge/License-MIT-f5de53?&color=f5de53'></a> </div> ## 🚀 Model Introduction We are excited to announce the release of LongCat-Video-Avatar, a unified model that delivers expressive and highly dynamic audio-driven character animation, supporting native tasks including Audio-Text-to-Video, Audio-Text-Image-to-Video, and Video Continuation with seamless compatibility for both single-stream and multi-stream audio inputs. ### Key Features - 🌟 **Support Multiple Generation Modes**: One unified model can be used for *audio-text-to-video (AT2V)* generation, *audio-text-image-to-video (ATI2V)* generation, and *Video Continuation*. - 🌟 **Natural Human D...
Source context: 508 downloads · 3 likes · Library diffusers · Repo Affcccbvdv2333-12/LongCat-Video-Avatar-ComfyUI-GGUF
Source context: 478 downloads · 3 likes · Library diffusers · Repo Affcccbvdv2333-12/LongCat-Video-Avatar-ComfyUI-GGUF