A image to video ComfyUI workflow with CogVideoX. Tested with CogvideoX Fun 1.1 and 1.5. Note that the motion lora does not work with the Fun 1.5 model. Just with the 1.1 one.
Laufzeitprofil
Quellenbeschreibung
A image to video ComfyUI workflow with CogVideoX. Tested with CogvideoX Fun 1.1 and 1.5. Note that the motion lora does not work with the Fun 1.5 model. Just with the 1.1 one.
The zipfile contains the json file, the starter image, and the png file from creation, which also contains the workflow.
This workflow also contains a CogVideoX motion lora for the camera movement. And you can also add further instructions in the prompt. CogVideoX relies at motion informations in text form.
It also has a very simple upscaling method implemented. I am at my journey to figure out a special upscaling workflow though. But for some it might still be useful. It is super fast compared to an upsampling by another ksampler.
CogVideo creation size is limited. The old version 1.1 is fixed to a 16:10 format. And 720x480 resolution. The new version 1.5 goes up to double size, but the motion lora that i use here does not work with it.
KI-generierter Kommentar
KI-generierte Erklärung auf Grundlage von Quellen- und Konfigurationsdetails. Vorschläge sind klar gekennzeichnet.
Use this ComfyUI workflow to turn a starter image and prompt text with motion instructions into a CogVideoX video, with a simple upscaling path.
It includes a CogVideoX motion LoRA for camera movement and a simple upscaling method.
Input: a starter image and prompt text, which can include additional movement instructions.
Output: a generated video, with simple upscaling available in the workflow.
The archive contains the JSON workflow, the starter image, and a PNG from creation that also contains the workflow.
Expand collapsed Note nodes for model links, placement information, and other notes.
The workflow lists these model files or model IDs: alibaba-pai/CogVideoX-Fun-V1.1-5b-InP, microsoft/Florence-2-large, 4x-UltraSharp.pth, pytorch_lora_weights.safetensors, rife47.pth, and t5xxl_fp8_e4m3fn.safetensors.
Named ComfyUI nodes include CLIPLoader, CogVideoDecode, CogVideoImageEncodeFunInP, CogVideoLoraSelect, CogVideoSampler, CogVideoTextEncode, Crop (mtb), DF_Get_image_size, DownloadAndLoadCogVideoModel, DownloadAndLoadFlorence2Model, Fast Groups Bypasser (rgthree), Florence2Run, ۽.
Other named nodes include ImageScale, ImageUpscaleWithModel, Int Literal, JWImageMix, JWIntegerMul, LoadImage, RIFE VFI, ShowText|pysssss, TextCombinations, UpscaleModelLoader, VHS_VideoCombine, and VRAM_Debug.
Add movement instructions to the prompt; CogVideoX uses motion information in text form.
Vorschlag · nicht geprüft
The listed size limits are 720x480 in 16:10 for Fun 1.1, while Fun 1.5 goes up to double size.
The included motion LoRA does not work with Fun 1.5; verify that you are using Fun 1.1 if you need that LoRA.
Vorschlag · nicht geprüft
Before use, verify that pytorch_lora_weights.safetensors, VHS_VideoCombine, CogVideoLoraSelect, and RIFE VFI are available. The description reports that 8 GB was too low, says 12 GB should work, and says 12 GB was not tested.
Vorschlag · nicht geprüft
Muss das bearbeitet werden?
Melden Sie sich an, um eine Änderungsanfrage zu senden.
Quellen
1 QuelleQuellenauszüge
1 AuszugSource context: 1839 downloads · Type Workflows · Base model CogVideoX
There are some collapsed Note nodes besides the important nodes. Click at them to expand them.
The Note nodes contain further informations. And in case of the models also links to the models, and where to put them.
This is my first upload at Civitai, so suggesstions and feedback is very welcome.
EDIT: The example video rendered in 8 minutes plus some overhead for preparation, without the upscaling with CogvideoX Fun version 1.1. at an 4060 TI. Version 1.5 renders doube as fast. But the motion lora does not work. Upscaling is another 4 minutes then if i remember right. I was not able to run CogvideoX at my old card with 8 Gb. That's too low. 12 gb should work. But i cannot test it. I render at a card with 16 gb at the moment.
And there is an explanation video at Youtube:
Geschätzter VRAM Bedarf
Schätzung nicht verfügbar
4,62 GB über 2 von 8 Modell-Dateien. Gesamtmodell-Dateien + 25% Ladeaufwand + 2 GB Ausführungs-Puffer, aufgerundet.
Anforderungen
30 AnforderungenUpscaler · 63.9 MB · PT · Unknown
alibaba-pai/CogVideoX-Fun-V1.1-5b-InP
Nicht aufgelöstCheckpoint · Hugging Face · alibaba-pai/CogVideoX-Fun-V1.1-5b-InP
CLIPLoader
Nicht aufgelöstText encoder · Unknown
ImageUpscaleWithModel
Nicht aufgelöstUpscaler · Unknown
Vision model · Hugging Face · microsoft/Florence-2-large · Florence 2 Large
pytorch_lora_weights.safetensors
MöglichLoRA · Unknown
Text encoder · 4.56 GB · SAFETENSOR · fp8 · Unknown
UpscaleModelLoader
Nicht aufgelöstUpscaler · Unknown
CogVideoDecode
Nicht aufgelöstKnotenpaket · Unknown
CogVideoImageEncodeFunInP
Nicht aufgelöstKnotenpaket · Unknown
CogVideoLoraSelect
Nicht aufgelöstKnotenpaket · Unknown
Knotenpaket · Unknown
CogVideoTextEncode
Nicht aufgelöstKnotenpaket · Unknown
Crop (mtb)
Nicht aufgelöstKnotenpaket · Unknown
DF_Get_image_size
Nicht aufgelöstKnotenpaket · Unknown
DownloadAndLoadCogVideoModel
Nicht aufgelöstKnotenpaket · Unknown
DownloadAndLoadFlorence2Model
Nicht aufgelöstKnotenpaket · Unknown
Fast Groups Bypasser (rgthree)
Nicht aufgelöstKnotenpaket · Unknown
Florence2Run
Nicht aufgelöstKnotenpaket · Unknown
ImageResizeKJ
Nicht aufgelöstKnotenpaket · Unknown
ImageScale
Nicht aufgelöstKnotenpaket · Unknown
Int Literal
Nicht aufgelöstKnotenpaket · Unknown
JWImageMix
Nicht aufgelöstKnotenpaket · Unknown
JWIntegerMul
Nicht aufgelöstKnotenpaket · Unknown
LoadImage
Nicht aufgelöstKnotenpaket · Unknown
RIFE VFI
Nicht aufgelöstKnotenpaket · Unknown
ShowText|pysssss
Nicht aufgelöstKnotenpaket · Unknown
TextCombinations
Nicht aufgelöstKnotenpaket · Unknown
VHS_VideoCombine
Nicht aufgelöstKnotenpaket · Unknown
VRAM_Debug
Nicht aufgelöstKnotenpaket · Unknown