💜 Wan | 🖥️ GitHub | 🤗 Hugging Face | 🤖 ModelScope | 📑 Paper (Coming soon) | 📑 Blog | 💬 WeChat Group | 📖 Discord
Model source
Source excerpt
💜 Wan | 🖥️ GitHub | 🤗 Hugging Face | 🤖 ModelScope | 📑 Paper (Coming soon) | 📑 Blog | 💬 WeChat Group | 📖 Discord
Requirements
1 requirementSources
1 sourceVerified Aug 15
Model artifacts
8 artifactsSource excerpts
2 excerptsVAE · Hugging Face · Aiwithus/Wan2.1-T2V-1.3B-Diffusers · SHA-256 d6e524b3fffe
text_encoder/model-00003-of-00005.safetensors
safetensors · 4.63 GB · SHA-256 f93148bcc040…35c3 · Hugging Face
Downloadtext_encoder/model-00004-of-00005.safetensors
safetensors · 4.66 GB · SHA-256 a451792c739c…57c9 · Hugging Face
Downloadtext_encoder/model-00005-of-00005.safetensors
safetensors · 2.69 GB · SHA-256 7e76e18d2245…4eea · Hugging Face
Downloadtransformer/diffusion_pytorch_model-00001-of-00002.safetensors
safetensors · 4.66 GB · SHA-256 6d011927dbd2…3a24 · Hugging Face
Downloadtransformer/diffusion_pytorch_model-00002-of-00002.safetensors
safetensors · 646 MB · SHA-256 b92ec2309b1f…f872 · Hugging Face
Downloadvae/diffusion_pytorch_model.safetensors
safetensors · 484 MB · SHA-256 d6e524b3fffe…e793 · Hugging Face
Download--- license: apache-2.0 language: - en - zh pipeline_tag: text-to-video library_name: diffusers tags: - video - video-generation --- # Wan2.1 <p align="center"> <img src="assets/logo.png" width="400"/> <p> <p align="center"> 💜 <a href=""><b>Wan</b></a>    |    🖥️ <a href="https://github.com/Wan-Video/Wan2.1">GitHub</a>    |   🤗 <a href="https://huggingface.co/Wan-AI/">Hugging Face</a>   |   🤖 <a href="https://modelscope.cn/organization/Wan-AI">ModelScope</a>   |    📑 <a href="">Paper (Coming soon)</a>    |    📑 <a href="https://wanxai.com">Blog</a>    |   💬 <a href="https://gw.alicdn.com/imgextra/i2/O1CN01tqjWFi1ByuyehkTSB_!!6000000000015-0-tps-611-1279.jpg">WeChat Group</a>   |    📖 <a href="https://discord.gg/p5XbdQV7">Discord</a>   <br> ----- [**Wan: Open and Advanced Large-Scale Video Generative Models**]("#") <be> In this repository, we present **Wan2.1**, a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. **Wan2.1** offers these key features: - 👍 **SOTA Performance**: **Wan2.1** consistently outperforms existing open-source models and state-of-the-art commercial solutions across multiple benchmarks. - 👍 **Supports Consumer-grade GPUs**: The T2V-1.3B model requires only 8.19 GB VRAM, making it compatible with almost all consumer-grade GPUs. It can generate a 5-second 480P video on an RTX 4090 in about 4 minutes (without optimization techniques like quantization). Its performance is even comparable to some closed-source models. - 👍 **Multiple Tasks**: **Wan2.1** excels in Text-to-Video, Image-to-Video, Video Editing, Text-to-Image, and Video-to-Audio, advancing the field of video generation. - 👍 **Visual Text Generation**: **Wan2.1** is the first video model capable of generating both Chinese and English text, featuring robust text generation that enhances its practical applications. - 👍 **Powerful Video VAE**: **Wan-VAE** delivers exceptional efficiency and performance, encoding and decoding 1080P videos of any length while preserving temporal information, making it an ideal foundation for video and image generation. This repository hosts our T2V-1.3B model, a versatile solution for video generation that is compatible with nearly all consumer-grade GPUs. In...
Source context: 1 downloads · 0 likes · Pipeline text-to-video · Library diffusers · Repo Aiwithus/Wan2.1-T2V-1.3B-Diffusers