A image to video ComfyUI workflow with CogVideoX. Tested with CogvideoX Fun 1.1 and 1.5. Note that the motion lora does not work with the Fun 1.5 model. Just with the 1.1 one.
Runtime profile
Source description
A image to video ComfyUI workflow with CogVideoX. Tested with CogvideoX Fun 1.1 and 1.5. Note that the motion lora does not work with the Fun 1.5 model. Just with the 1.1 one.
The zipfile contains the json file, the starter image, and the png file from creation, which also contains the workflow.
This workflow also contains a CogVideoX motion lora for the camera movement. And you can also add further instructions in the prompt. CogVideoX relies at motion informations in text form.
It also has a very simple upscaling method implemented. I am at my journey to figure out a special upscaling workflow though. But for some it might still be useful. It is super fast compared to an upsampling by another ksampler.
CogVideo creation size is limited. The old version 1.1 is fixed to a 16:10 format. And 720x480 resolution. The new version 1.5 goes up to double size, but the motion lora that i use here does not work with it.
AI-generated commentary
AI-generated explanation based on source and configuration details. Suggestions are clearly labeled.
Use this ComfyUI workflow to turn a starter image and prompt text with motion instructions into a CogVideoX video, with a simple upscaling path.
It includes a CogVideoX motion LoRA for camera movement and a simple upscaling method.
Input: a starter image and prompt text, which can include additional movement instructions.
Output: a generated video, with simple upscaling available in the workflow.
The archive contains the JSON workflow, the starter image, and a PNG from creation that also contains the workflow.
Expand collapsed Note nodes for model links, placement information, and other notes.
The workflow lists these model files or model IDs: alibaba-pai/CogVideoX-Fun-V1.1-5b-InP, microsoft/Florence-2-large, 4x-UltraSharp.pth, pytorch_lora_weights.safetensors, rife47.pth, and t5xxl_fp8_e4m3fn.safetensors.
Named ComfyUI nodes include CLIPLoader, CogVideoDecode, CogVideoImageEncodeFunInP, CogVideoLoraSelect, CogVideoSampler, CogVideoTextEncode, Crop (mtb), DF_Get_image_size, DownloadAndLoadCogVideoModel, DownloadAndLoadFlorence2Model, Fast Groups Bypasser (rgthree), Florence2Run, ۽.
Other named nodes include ImageScale, ImageUpscaleWithModel, Int Literal, JWImageMix, JWIntegerMul, LoadImage, RIFE VFI, ShowText|pysssss, TextCombinations, UpscaleModelLoader, VHS_VideoCombine, and VRAM_Debug.
Add movement instructions to the prompt; CogVideoX uses motion information in text form.
Suggestion · not verified
The listed size limits are 720x480 in 16:10 for Fun 1.1, while Fun 1.5 goes up to double size.
The included motion LoRA does not work with Fun 1.5; verify that you are using Fun 1.1 if you need that LoRA.
Suggestion · not verified
Before use, verify that pytorch_lora_weights.safetensors, VHS_VideoCombine, CogVideoLoraSelect, and RIFE VFI are available. The description reports that 8 GB was too low, says 12 GB should work, and says 12 GB was not tested.
Suggestion · not verified
Does this need to be edited?
Sign in to send an edit request.
Sources
1 sourceSource excerpts
1 excerptSource context: 1839 downloads · Type Workflows · Base model CogVideoX
There are some collapsed Note nodes besides the important nodes. Click at them to expand them.
The Note nodes contain further informations. And in case of the models also links to the models, and where to put them.
This is my first upload at Civitai, so suggesstions and feedback is very welcome.
EDIT: The example video rendered in 8 minutes plus some overhead for preparation, without the upscaling with CogvideoX Fun version 1.1. at an 4060 TI. Version 1.5 renders doube as fast. But the motion lora does not work. Upscaling is another 4 minutes then if i remember right. I was not able to run CogvideoX at my old card with 8 Gb. That's too low. 12 gb should work. But i cannot test it. I render at a card with 16 gb at the moment.
And there is an explanation video at Youtube:
Estimated VRAM requirement
Estimate unavailable
4.62 GB across 2 of 8 model files. Model file total + 25% loading overhead + 2 GB execution buffer, rounded up.
Requirements
30 requirementsUpscaler · 63.9 MB · PT · Unknown
alibaba-pai/CogVideoX-Fun-V1.1-5b-InP
Not resolvedCheckpoint · Hugging Face · alibaba-pai/CogVideoX-Fun-V1.1-5b-InP
CLIPLoader
Not resolvedText encoder · Unknown
ImageUpscaleWithModel
Not resolvedUpscaler · Unknown
Vision model · Hugging Face · microsoft/Florence-2-large · Florence 2 Large
pytorch_lora_weights.safetensors
PossibleLoRA · Unknown
Text encoder · 4.56 GB · SAFETENSOR · fp8 · Unknown
UpscaleModelLoader
Not resolvedUpscaler · Unknown
CogVideoDecode
Not resolvedNode pack · Unknown
CogVideoImageEncodeFunInP
Not resolvedNode pack · Unknown
CogVideoLoraSelect
Not resolvedNode pack · Unknown
Node pack · Unknown
CogVideoTextEncode
Not resolvedNode pack · Unknown
Crop (mtb)
Not resolvedNode pack · Unknown
DF_Get_image_size
Not resolvedNode pack · Unknown
DownloadAndLoadCogVideoModel
Not resolvedNode pack · Unknown
DownloadAndLoadFlorence2Model
Not resolvedNode pack · Unknown
Fast Groups Bypasser (rgthree)
Not resolvedNode pack · Unknown
Florence2Run
Not resolvedNode pack · Unknown
ImageResizeKJ
Not resolvedNode pack · Unknown
ImageScale
Not resolvedNode pack · Unknown
Int Literal
Not resolvedNode pack · Unknown
JWImageMix
Not resolvedNode pack · Unknown
JWIntegerMul
Not resolvedNode pack · Unknown
LoadImage
Not resolvedNode pack · Unknown
RIFE VFI
Not resolvedNode pack · Unknown
ShowText|pysssss
Not resolvedNode pack · Unknown
TextCombinations
Not resolvedNode pack · Unknown
VHS_VideoCombine
Not resolvedNode pack · Unknown
VRAM_Debug
Not resolvedNode pack · Unknown