A simple implementation of QwenVL-series LLM model in ComfyUI.
Perfil do pacote
README
A comprehensive and modular QwenVL integration for ComfyUI, providing advanced vision-language capabilities with support for both HuggingFace Transformers and GGUF models. This extension consolidates features from multiple QwenVL implementations while introducing enhanced error handling, attention backend optimization, and a clean, maintainable codebase.
This project builds upon and consolidates features from multiple excellent QwenVL implementations:
Fontes
3 fontesTrechos de fonte
1 trechoSource context: Repo AkihaTatsu/ComfyUI-QwenVL-Utils
ComfyUI-QwenVL by 1038lab
ComfyUI_Qwen2-VL-Instruct by IuvenisSapiens
Qwen3.5 introduces unified thinking/instruct mode — a single model supports both deep chain-of-thought reasoning and direct instruction-following, controlled by the enable_thinking toggle in the node UI. No need for separate "-Instruct" and "-Thinking" model files.
| Model | Size | Architecture | VRAM (FP16) | VRAM (8-bit) | VRAM (4-bit) |
|---|---|---|---|---|---|
| Qwen3.5-9B | 9B | Dense (Hybrid Gated Delta Net) | ~20GB | ~12GB | ~7GB |
| Qwen3.5-27B | 27B | Dense (Hybrid Gated Delta Net) | ~56GB | ~30GB | ~18GB |
| Qwen3.5-35B-A3B | 35B total / 3B active | MoE (Hybrid Gated Delta Net) | ~10GB | ~6GB | ~4GB |
FP8 Pre-Quantized Qwen3.5 (40-series GPU recommended):
Qwen3.5-9B (FP8): ~12GB VRAMQwen3.5-27B (FP8): ~30GB VRAMQwen3.5-35B-A3B (FP8): ~6GB VRAMNote: Qwen3.5 uses a novel hybrid architecture combining Gated Delta Networks with sparse MoE. The 35B-A3B variant activates only ~3B parameters per token, making it very memory-efficient despite 35B total parameters.
| Model | Size | Features | VRAM (FP16) | VRAM (8-bit) | VRAM (4-bit) |
|---|---|---|---|---|---|
| Qwen3-VL-2B-Instruct | 2B | General VL | ~4GB | ~2.5GB | ~1.5GB |
| Qwen3-VL-2B-Thinking | 2B | CoT reasoning | ~4GB | ~2.5GB | ~1.5GB |
| Qwen3-VL-4B-Instruct | 4B | Balanced | ~6GB | ~3.5GB | ~2GB |
| Qwen3-VL-4B-Thinking | 4B | CoT reasoning | ~6GB | ~3.5GB | ~2GB |
| Qwen3-VL-8B-Instruct | 8B | High quality | ~12GB | ~7GB |
FP8 Pre-Quantized Models (40-series GPU recommended):
Qwen3-VL-2B-*-FP8: ~2.5GB VRAMQwen3-VL-4B-*-FP8: ~2.5GB VRAMQwen3-VL-8B-*-FP8: ~7.5GB VRAMQwen3-VL-32B-*-FP8: ~24GB VRAMAll GGUF models are sourced from unsloth for consistent quality and compatibility.
| Model | Source | Variants | Features |
|---|---|---|---|
| Qwen3.5 (Unified Thinking/Instruct) | |||
| Qwen3.5-9B-GGUF | unsloth | Q4_K_M, Q8_0, BF16 | Unified thinking + instruct |
| Qwen3.5-27B-GGUF | unsloth | Q4_K_M, Q8_0, BF16 | Unified thinking + instruct |
| Qwen3.5-35B-A3B-GGUF | unsloth | Q4_K_M, Q8_0, BF16 | MoE, unified thinking + instruct |
| Qwen3-VL | |||
| Qwen3-VL-2B-Instruct-GGUF | unsloth | Q4_K_M, Q8_0 | Instruct tuned |
| Qwen3-VL-4B-Instruct-GGUF | unsloth | Q4_K_M, Q8_0 | Instruct tuned |
Nós neste pacote
Nós 6Verificado 11 de set.
Verificado 11 de set.
| ~4.5GB |
| Qwen3-VL-8B-Thinking | 8B | Advanced CoT | ~12GB | ~7GB | ~4.5GB |
| Qwen3-VL-32B-Instruct | 32B | Best quality | ~28GB | ~14GB | ~8.5GB |
| Qwen3-VL-32B-Thinking | 32B | Complex reasoning | ~28GB | ~14GB | ~8.5GB |
| Qwen2.5-VL-3B-Instruct | 3B | Previous gen | ~6GB | ~3.5GB | ~2GB |
| Qwen2.5-VL-7B-Instruct | 7B | Previous gen | ~15GB | ~8.5GB | ~5GB |
| Qwen3-VL-8B-Instruct-GGUF | unsloth | Q4_K_M, Q8_0 | Instruct tuned |
| Qwen3-VL-4B-Thinking-GGUF | unsloth | Q4_K_M |