ComfyUI nodes for vision-language models: Qwen3-VL, Moondream 3, Florence-2, SmolVLM2, InternVL, Gemma 3, MiniCPM-V. Plus open-vocabulary detection, SAM2/SAM3 segmentation, video temporal reasoning, GGUF via...
Package profile
README
ComfyUI nodes for vision-language models: Qwen3-VL, Moondream 3, Florence-2, SmolVLM2, InternVL, Gemma 3, MiniCPM-V. Plus open-vocabulary detection, SAM2/SAM3 segmentation, video temporal reasoning, GGUF via llama.cpp, and hosted LLM/VLM APIs.
Nodes in this pack
34 nodesSources
1 sourceSource excerpts
1 excerptSource context: Repo gokayfem/ComfyUI_VLM_nodes
AudioLDM2Node
Compact description unavailable for this node.
VLM Nodes/Audio · 2 inputs · 6 parameters