MiniMax H3 – Minimal I2V + Audio Workflow for ComfyUI
Perfil de execução
Descrição da fonte
MiniMax H3 – Minimal I2V + Audio Workflow for ComfyUI
A compact MiniMax H3 image-to-video workflow for ComfyUI with synchronized audio generation.
This package contains a tested workflow JSON , documentation, and direct model download references. Model weights are not included . Please download them from the original Comfy-Org repository.
Features
Image-to-video from a single starting image
Joint video + audio generation
124 frames at 24 fps (about 5.17 seconds)
res_multistep sampler
Fontes
1 fonteTrechos de fonte
2 trechosSource context: 389 downloads · Type Workflows · Base model MiniMax H3
20 steps
MiniMax H3 sigma shift:
Video: 12
Audio: 3
Standard ComfyUI VAE loaders
Uses the official Comfy-Org Audio VAE with latent statistics included
VideoHelperSuite output to H.264 MP4 with audio
Required model files
minimax_h3_fl2va_pruned_int8_convrot.safetensors
Direct download: Download diffusion model
Place in: ComfyUI/models/diffusion_models/ or your configured diffusion_models folder.
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
Direct download: Download text encoder
Place in: ComfyUI/models/text_encoders/ or your configured text_encoders folder.
minimax_h3_video_vae_fp16.safetensors
Direct download: Download video VAE
Place in: ComfyUI/models/vae/ or your configured vae folder.
minimax_h3_audio_vae_fp32.safetensors
Direct download: Download audio VAE
Place in: ComfyUI/models/vae/ or your configured vae folder.
Required custom node
ComfyUI-VideoHelperSuite
MiniMax H3 model / conditioning nodes used by this workflow are ComfyUI core nodes in a ComfyUI build with MiniMax H3 support .
Sample prompt – English speech
A woman looks directly at the camera and speaks naturally in English: "Hello! I'm testing the audio generation. Can you hear my voice clearly?" Her voice is clear and easy to understand. Natural synchronized English speech, accurate lip movements, and subtle room ambience.
Workflow settings
Resolution in included tested sample: 512 × 512
Length: 124 frames
Frame rate: 24 fps
Sampler: res_multistep
Scheduler: normal
Steps: 20
Denoise: 1.0
Video sigma shift: 12
Audio sigma shift: 3
You can change the prompt , source image , and resolution as needed.
Included files
MiniMax_H3_I2V_test.json
Tested 512×512 English speech example
Source image connected to first_frame
README.md
CIVITAI_DESCRIPTION.md
MODEL_DOWNLOADS.md
MANIFEST.txt
Important note
Use the official Comfy-Org Audio VAE listed above.
A differently packaged Audio VAE that lacks latents_mean / latents_std , or uses unfused weight_g / weight_v decoder weights, may decode to:
silence
noise
a constant tone
This package does not redistribute model weights. Please follow the licenses and terms of the original model repositories.
Estimativa de requisito VRAM
Estimativa indisponível
34,7 GB distribuídos em 3 de 4 arquivos de modelo. Total de arquivos do modelo + 25% de overhead de carregamento + 2 GB de buffer de execução, arredondado para cima.
Requisitos
Requisitos 5minimax_h3_audio_vae_fp32.safetensors
EncontradoVAE · 577 MB · SAFETENSORS · Hugging Face · Comfy-Org/MiniMax-H3
Checkpoint · 19.5 GB · SAFETENSORS · Hugging Face · Comfy-Org/MiniMax-H3
MiniMax-H3-video_vae_fp16.safetensors
Arquivo não verificadoVAE · Hugging Face · Comfy-Org/MiniMax-H3 · MiniMax H3
Text encoder · 14.6 GB · SAFETENSORS · Hugging Face · Comfy-Org/MiniMax-H3
Pacote de nós · Registry
Source context: 388 downloads · Type Workflows · Base model MiniMax H3