A image to video ComfyUI workflow with CogVideoX. Tested with CogvideoX Fun 1.1 and 1.5. Note that the motion lora does not work with the Fun 1.5 model. Just with the 1.1 one.
Profil d'exécution
Description de la source
A image to video ComfyUI workflow with CogVideoX. Tested with CogvideoX Fun 1.1 and 1.5. Note that the motion lora does not work with the Fun 1.5 model. Just with the 1.1 one.
The zipfile contains the json file, the starter image, and the png file from creation, which also contains the workflow.
This workflow also contains a CogVideoX motion lora for the camera movement. And you can also add further instructions in the prompt. CogVideoX relies at motion informations in text form.
It also has a very simple upscaling method implemented. I am at my journey to figure out a special upscaling workflow though. But for some it might still be useful. It is super fast compared to an upsampling by another ksampler.
CogVideo creation size is limited. The old version 1.1 is fixed to a 16:10 format. And 720x480 resolution. The new version 1.5 goes up to double size, but the motion lora that i use here does not work with it.
Commentaire généré par l’IA
Explication générée par l’IA à partir des détails de la source et de la configuration. Les suggestions sont clairement signalées.
Use this ComfyUI workflow to turn a starter image and prompt text with motion instructions into a CogVideoX video, with a simple upscaling path.
It includes a CogVideoX motion LoRA for camera movement and a simple upscaling method.
Input: a starter image and prompt text, which can include additional movement instructions.
Output: a generated video, with simple upscaling available in the workflow.
The archive contains the JSON workflow, the starter image, and a PNG from creation that also contains the workflow.
Expand collapsed Note nodes for model links, placement information, and other notes.
The workflow lists these model files or model IDs: alibaba-pai/CogVideoX-Fun-V1.1-5b-InP, microsoft/Florence-2-large, 4x-UltraSharp.pth, pytorch_lora_weights.safetensors, rife47.pth, and t5xxl_fp8_e4m3fn.safetensors.
Named ComfyUI nodes include CLIPLoader, CogVideoDecode, CogVideoImageEncodeFunInP, CogVideoLoraSelect, CogVideoSampler, CogVideoTextEncode, Crop (mtb), DF_Get_image_size, DownloadAndLoadCogVideoModel, DownloadAndLoadFlorence2Model, Fast Groups Bypasser (rgthree), Florence2Run, ۽.
Other named nodes include ImageScale, ImageUpscaleWithModel, Int Literal, JWImageMix, JWIntegerMul, LoadImage, RIFE VFI, ShowText|pysssss, TextCombinations, UpscaleModelLoader, VHS_VideoCombine, and VRAM_Debug.
Add movement instructions to the prompt; CogVideoX uses motion information in text form.
Suggestion · non vérifié
The listed size limits are 720x480 in 16:10 for Fun 1.1, while Fun 1.5 goes up to double size.
The included motion LoRA does not work with Fun 1.5; verify that you are using Fun 1.1 if you need that LoRA.
Suggestion · non vérifié
Before use, verify that pytorch_lora_weights.safetensors, VHS_VideoCombine, CogVideoLoraSelect, and RIFE VFI are available. The description reports that 8 GB was too low, says 12 GB should work, and says 12 GB was not tested.
Suggestion · non vérifié
Faut-il modifier cela ?
Connectez-vous pour envoyer une demande de modification.
Sources
1 sourceExtraits de sources
1 extraitSource context: 1839 downloads · Type Workflows · Base model CogVideoX
There are some collapsed Note nodes besides the important nodes. Click at them to expand them.
The Note nodes contain further informations. And in case of the models also links to the models, and where to put them.
This is my first upload at Civitai, so suggesstions and feedback is very welcome.
EDIT: The example video rendered in 8 minutes plus some overhead for preparation, without the upscaling with CogvideoX Fun version 1.1. at an 4060 TI. Version 1.5 renders doube as fast. But the motion lora does not work. Upscaling is another 4 minutes then if i remember right. I was not able to run CogvideoX at my old card with 8 Gb. That's too low. 12 gb should work. But i cannot test it. I render at a card with 16 gb at the moment.
And there is an explanation video at Youtube:
Estimation des besoins VRAM
Estimation indisponible
4,62 GB sur 2 de 8 fichiers de modèle. Total des fichiers du modèle + 25 % de surcharge de chargement + 2 Go de tampon d'exécution, arrondi à l'unité supérieure.
Exigences
Exigences 30Upscaler · 63.9 MB · PT · Unknown
alibaba-pai/CogVideoX-Fun-V1.1-5b-InP
Non résoluCheckpoint · Hugging Face · alibaba-pai/CogVideoX-Fun-V1.1-5b-InP
CLIPLoader
Non résoluText encoder · Unknown
ImageUpscaleWithModel
Non résoluUpscaler · Unknown
Vision model · Hugging Face · microsoft/Florence-2-large · Florence 2 Large
pytorch_lora_weights.safetensors
PossibleLoRA · Unknown
Text encoder · 4.56 GB · SAFETENSOR · fp8 · Unknown
UpscaleModelLoader
Non résoluUpscaler · Unknown
CogVideoDecode
Non résoluPack de nœud · Unknown
CogVideoImageEncodeFunInP
Non résoluPack de nœud · Unknown
CogVideoLoraSelect
Non résoluPack de nœud · Unknown
Pack de nœud · Unknown
CogVideoTextEncode
Non résoluPack de nœud · Unknown
Crop (mtb)
Non résoluPack de nœud · Unknown
DF_Get_image_size
Non résoluPack de nœud · Unknown
DownloadAndLoadCogVideoModel
Non résoluPack de nœud · Unknown
DownloadAndLoadFlorence2Model
Non résoluPack de nœud · Unknown
Fast Groups Bypasser (rgthree)
Non résoluPack de nœud · Unknown
Florence2Run
Non résoluPack de nœud · Unknown
ImageResizeKJ
Non résoluPack de nœud · Unknown
ImageScale
Non résoluPack de nœud · Unknown
Int Literal
Non résoluPack de nœud · Unknown
JWImageMix
Non résoluPack de nœud · Unknown
JWIntegerMul
Non résoluPack de nœud · Unknown
LoadImage
Non résoluPack de nœud · Unknown
RIFE VFI
Non résoluPack de nœud · Unknown
ShowText|pysssss
Non résoluPack de nœud · Unknown
TextCombinations
Non résoluPack de nœud · Unknown
VHS_VideoCombine
Non résoluPack de nœud · Unknown
VRAM_Debug
Non résoluPack de nœud · Unknown