This repo packages the huihui-ai/Huihui-Qwen3-VL-4B-Instruct-abliterated Vision-Language Model in two ready-to-use formats for ComfyUI:
Source du modèle
Description de la source
This repo packages the huihui-ai/Huihui-Qwen3-VL-4B-Instruct-abliterated Vision-Language Model in two ready-to-use formats for ComfyUI:
Ressources connexes
1 connexionSources
1 sourceVérifié 28 juil.
Artefacts du modèle
1 artefactHuihui-Qwen3-VL-4B-Instruct-abliterated-fp8_scaled.safetensors
safetensors · FP8 · 4,88 GB · SHA-256 45fe15d359fb…b866 · Hugging Face
TéléchargerExtraits de sources
2 extraits| File | Size | Format | Use case |
|---|
Huihui-Qwen3-VL-4B-Instruct-abliterated.safetensors | 8.88 GiB | BF16 single safetensors | Maximum fidelity / training / full-precision workflows |
Huihui-Qwen3-VL-4B-Instruct-abliterated-fp8_scaled.safetensors | 5.24 GiB | FP8 (E4M3FN) per-tensor scaled | ComfyUI Qwen3-VL Text Encoder node — recommended |
huihui-ai/Huihui-Qwen3-VL-4B-Instruct-abliterated (apache-2.0)Qwen/Qwen3-VL-4B-Instruct*.safetensors, 8.88 GiB)The original upstream repo ships the weights split across two safetensors shards (model-00001-of-00002.safetensors + model-00002-of-00002.safetensors). They were merged into a single safetensors file using the original model.safetensors.index.json mapping. No weights modified.
bfloat16*-fp8_scaled.safetensors, 5.24 GiB)Per-tensor abs-max quantization to float8_e4m3fn for the 252 linear projections of the language model (q/k/v/o_proj + gate/up/down_proj across all layers). Embeddings, layer norms, biases and the entire visual encoder stay in BF16.
float8_e4m3fn weights + float32 per-tensor scale + uint8[64] comfy_quant marker (JSON: {"format": "float8_e4m3fn", "full_precision_matrix_mult": false})max(|w|) / 448 (E4M3FN max)qwen3vl_4b_fp8_scaled.safetensors)The FP8 conversion was done on GPU (NVIDIA RTX 3090) in ~4 seconds. Script is available on request.
*-fp8_scaled.safetensors into your ComfyUI models/text_encoders/ directory.Huihui-Qwen3-VL-4B-Instruct-abliterated-fp8_scaled.*.safetensors into models/text_encoders/.The upstream repo huihui-ai/Huihui-Qwen3-VL-4B-Instruct-abliterated has the matching tokenizer, processor and configs. For BF16 inference:
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
import torch
model = Qwen3VLForConditionalGeneration.from_pretrained(
"ahmed22xa/Huihui-Qwen3-VL-4B-Instruct-abliterated-comfy",
torch_dtype=torch.bfloat16,
device_map="auto",
)
processor = AutoProcessor.from_pretrained("huihui-ai/Huihui-Qwen3-VL-4B-Instruct-abliterated")
The FP8 file is not loadable with transformers.from_pretrained directly — it follows ComfyUI's per-tensor-FP8 layout with comfy_quant markers.
Qwen/Qwen3-VL-4B-Instruct).--- license: apache-2.0 base_model: huihui-ai/Huihui-Qwen3-VL-4B-Instruct-abliterated base_model_relation: quantized tags: - qwen3_vl - vision-language - abliterated - uncensored - safetensors - comfyui - fp8 - transformers pipeline_tag: image-text-to-text library_name: transformers --- # Huihui-Qwen3-VL-4B-Instruct-abliterated — ComfyUI Edition This repo packages the [huihui-ai/Huihui-Qwen3-VL-4B-Instruct-abliterated](https://huggingface.co/huihui-ai/Huihui-Qwen3-VL-4B-Instruct-abliterated) Vision-Language Model in two ready-to-use formats for **ComfyUI**: | File | Size | Format | Use case | |---|---|---|---| | `Huihui-Qwen3-VL-4B-Instruct-abliterated.safetensors` | 8.88 GiB | **BF16** single safetensors | Maximum fidelity / training / full-precision workflows | | `Huihui-Qwen3-VL-4B-Instruct-abliterated-fp8_scaled.safetensors` | 5.24 GiB | **FP8 (E4M3FN) per-tensor scaled** | ComfyUI Qwen3-VL Text Encoder node — recommended | ## Source / Provenance - **Base model:** [`huihui-ai/Huihui-Qwen3-VL-4B-Instruct-abliterated`](https://huggingface.co/huihui-ai/Huihui-Qwen3-VL-4B-Instruct-abliterated) (apache-2.0) - **Origin model:** [`Qwen/Qwen3-VL-4B-Instruct`](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct) - **Abliteration:** only the text part was processed (not the vision part), so the visual encoder behaves identically to the upstream model. ## What was done ### BF16 single-file (`*.safetensors`, 8.88 GiB) The original upstream repo ships the weights split across two safetensors shards (`model-00001-of-00002.safetensors` + `model-00002-of-00002.safetensors`). They were merged into a single safetensors file using the original `model.safetensors.index.json` mapping. No weights modified. - 713 tensors - dtype: `bfloat16` - Verified structurally identical to upstream (same key set) ### FP8 scaled (`*-fp8_scaled.safetensors`, 5.24 GiB) Per-tensor abs-max quantization to `float8_e4m3fn` for the 252 linear projections of the language model (q/k/v/o_proj + gate/up/down_proj across all layers). Embeddings, layer norms, biases and the entire visual encoder stay in BF16. - 1217 tensors (252 × 3 + 461 BF16) - Quantised layers: `float8_e4m3fn` weights + `float32` per-tensor scale + `uint8[64]` comfy_quant marker (JSON: `{"format": "float8_e4m3fn", "full_precision_matrix_mult": false}`) - Per-tensor scale = `max(|w|) / 448` (E4M3FN max) - Mean round-...
Source context: 0 downloads · 59 likes · Pipeline image-text-to-text · Library transformers · Repo ahmed22xa/Huihui-Qwen3-VL-4B-Instruct-abliterated-comfy