Package profile
README
Upscaling a 15-second video to about 2 megapixels (1080p) with this tiled sampler used roughly 75% of a 16 GB RTX 4060 Ti (~12 GB VRAM) and took 70 minutes.
A single ComfyUI node, MiniMax H3 Tiled Sampler (model/sampling/minimax), that
denoises a MiniMax H3 audio-video latent in overlapping spatial tiles.
H3 packs video and audio into one attention sequence, so its cost is driven by the video token count () and its trained canvas stops at 768x1344 pixels. Tiling the video stream keeps every tile inside that native field of view and ties peak VRAM to the tile size rather than to the canvas, which is what makes larger canvases practical.
Sources
3 sourcesVerified Sep 29
latent_t * h/2 * w/2Same controls as KSampler (seed, steps, cfg, sampler_name, scheduler,
denoise), plus:
| Input | Meaning |
|---|---|
tile_width / tile_height | Tile size in pixels, rounded down to 32. Defaults are the model's native canvas, 1344x768. |
tile_overlap | Overlap between neighbouring tiles in pixels. Larger values hide seams and add tiles. |
latent must be an H3 AV latent, from Empty MiniMax H3 AV Latent,
MiniMax H3 Image to Video or MiniMax H3 Reference to Video. When the canvas
fits in a single tile the node samples it whole, so it is a drop-in replacement
for KSampler in an H3 workflow.
Sampling itself is unchanged: the node clones the model and adds a
DIFFUSION_MODEL wrapper, then hands the latent to the stock sampler. Per model
call, the wrapper:
A PREPARE_SAMPLING wrapper shrinks the VRAM estimate to one tile's worth of
video cells, so the model is not offloaded for a sequence length that is never
built.
tile_overlap if you see seams.Nodes in this pack
1 nodeVerified Sep 29