💬 Community: Join the Abliterlitics Discord for discussion, model releases and support.
Fuente del modelo
Descripción de la fuente
💬 Community: Join the Abliterlitics Discord for discussion, model releases and support.
ComfyUI checkpoints of the of . Five quantisations are included, covering everything from any modern GPU through to Blackwell, suitable as an uncensored text encoder for image-generation workflows including , which runs on a Qwen3VL-4B encoder.
Fuentes
1 fuenteVerificado 19 sept
Artefactos del modelo
6 artefactosExtractos de fuentes
3 extractosThis model is published across three repos. Pick the one that matches your runtime.
| Repo | Best for | Contents |
|---|---|---|
| Qwen3-VL-4b-Heretic | transformers, vLLM, HF Hub | bf16 weights with config, vision encoder preserved |
| Qwen3-VL-4b-Heretic-GGUF | llama.cpp, Ollama, LM Studio, ComfyUI-GGUF | GGUF quants from Q3_K_M up to F16 (text path) |
| Qwen3-VL-4b-Heretic-ComfyUI (this repo) | ComfyUI text encoder | bf16, fp8, int8, int4, nvfp4 and mxfp8 checkpoints |
The Docker setup, scripts and configs that produced these files are in Heretic Docker.
Several Heretic trials were run against Qwen3-VL-4B and all of them reach 100% HarmBench ASR, up from 30.8% on the base. This build was picked because, with safety tied, it wins on the tie-breakers:
| Base | Heretic | |
|---|---|---|
| HarmBench ASR | 30.8% | 100% |
| KL divergence (lower is better) | 0.0283 (lowest of the candidates) | |
| GSM8K | 78.62% | 77.18% (−1.83%, smallest drop) |
| MMLU | 69.58% | 69.61% (+0.03%) |
| Tensors changed | 54 (pure rank-1) |
See the full report for the comparison.
Benchmark deltas across the Heretic variants
| File | Format | Size | HW | Description |
|---|---|---|---|---|
qwen3-vl-4b-heretic.safetensors | bf16 | 8.3 GB | Any | Full precision |
qwen3-vl-4b-heretic_fp8_e4m3fn.safetensors | FP8 E4M3 | 4.5 GB | Ada+ | Per-tensor scaled, learned rounding |
qwen3-vl-4b-heretic_int8.safetensors | INT8 | 4.6 GB | Any | ConvRot row-wise, learned rounding |
qwen3-vl-4b-heretic_int4.safetensors | INT4 W4A4 | 2.7 GB | Any* | ConvRot Hadamard-rotated, learned rounding |
qwen3-vl-4b-heretic_nvfp4.safetensors | NVFP4 E2M1 | 2.9 GB | Blackwell† |
* INT4 W4A4 uses ComfyUI's convrot_w4a4 path (ComfyUI 0.30.0+). Layers with incompatible dimensions are kept in bf16.
† NVFP4 also runs on older GPUs through software dequantisation (tested on an RTX 4090). Native FP4 tensor cores need SM100+ (RTX 5090/5080).
Quality: All quantised variants use SVD-guided learned rounding (AdaRound via convert-to-quant), which optimises each weight's rounding direction to minimise output reconstruction error, noticeably higher fidelity than naive round-to-nearest quantisation.
ComfyUI/models/text_encoders/.Krea 2 is an image-generation model that runs on a Qwen3VL-4B text encoder. Because these checkpoints are the same Qwen3-VL-4B architecture (abliterated), you can drop one in as the encoder in a Krea 2 ComfyUI workflow to give it an uncensored text encoder. The fp8 checkpoint is the closest match to the stock qwen3vl_4b_fp8_scaled.safetensors; use int8 or bf16 if you want higher fidelity and have the VRAM.
FP8 (E4M3) is per-tensor scaled quantisation with learned rounding via convert-to-quant. It runs on Ada (RTX 4090) and newer, giving a good balance of size and quality.
INT8 (ConvRot row-wise) uses group-wise Hadamard rotation (ConvRot) to spread activation outliers, followed by per-row symmetric INT8 with learned rounding. For each weight tensor, an SVD-guided gradient descent loop (Prodigy optimiser) learns the rounding direction that minimises output error, which keeps INT8 quality near lossless. It runs on any modern GPU (Ampere+) with no Blackwell requirement.
INT4 (W4A4 ConvRot) is 4-bit signed INT with mandatory group-wise Hadamard rotation. The smallest variant at about a third of bf16. Layers whose dimensions are not divisible by the ConvRot group size are automatically kept in bf16. Uses ComfyUI's convrot_w4a4 path (ComfyUI 0.30.0+).
NVFP4 (E2M1) is 4-bit floating point with double quantisation (a per-tensor f32 scale plus a per-block FP8 scale, block size 16) and learned rounding. It is about three times smaller than bf16 and loads natively in ComfyUI without plugins. Blackwell GPUs (RTX 5090/5080, SM100+) use native FP4 tensor cores for best speed, while older GPUs fall back to software dequantisation (tested on an RTX 4090).
MXFP8 is microscaling FP8 (the OCP MX standard). It stores FP8 E4M3 data with E8M0 (power-of-two) per-block scales on 32-element blocks, with learned rounding. It needs SM100+ (Blackwell).
Produced with Heretic Docker, which wraps:
This model has had its safety alignment removed. It complies with harmful requests, including content related to violence, illegal activities and other harmful behaviour. Use it responsibly and in line with the laws and regulations that apply to you. The authors do not condone or encourage using this model for harmful purposes.
qwen3-vl-4b-heretic_int8.safetensors
safetensors · 4,50 GB · SHA-256 f70bc614fa54…a436 · Hugging Face
Descargarqwen3-vl-4b-heretic_mxfp8.safetensors
safetensors · 4,62 GB · SHA-256 2f2394363d66…8c42 · Hugging Face
Descargarqwen3-vl-4b-heretic_nvfp4.safetensors
safetensors · 2,85 GB · SHA-256 dd2ee4b686d3…d1c8 · Hugging Face
Descargarqwen3-vl-4b-heretic.safetensors
safetensors · 8,27 GB · SHA-256 3fba9f5a2059…f9ce · Hugging Face
Descargar--- license: apache-2.0 library_name: transformers pipeline_tag: image-text-to-text base_model: Qwen/Qwen3-VL-4B-Instruct base_model_relation: quantized language: - en tags: - abliteration - heretic - uncensored - qwen3-vl - qwen3 - vision-language - text-encoder - comfyui - fp8 - int8 - int4 - convrot - nvfp4 - mxfp8 - learned-rounding - blackwell - image-generation - krea - krea2 --- # Qwen3-VL-4B-Instruct Heretic (ComfyUI) 💬 **Community:** Join the [Abliterlitics Discord](https://discord.gg/AqmDnBjPvM) for discussion, model releases and support. ComfyUI checkpoints of the [Heretic abliteration](https://huggingface.co/DreamFast/Qwen3-VL-4b-Heretic) of [Qwen3-VL-4B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct). Five quantisations are included, covering everything from any modern GPU through to Blackwell, suitable as an uncensored text encoder for image-generation workflows including [Krea 2](https://docs.comfy.org/tutorials/image/krea/krea-2), which runs on a Qwen3VL-4B encoder. ## Available formats This model is published across three repos. Pick the one that matches your runtime. | Repo | Best for | Contents | |------|----------|----------| | [Qwen3-VL-4b-Heretic](https://huggingface.co/DreamFast/Qwen3-VL-4b-Heretic) | transformers, vLLM, HF Hub | bf16 weights with config, vision encoder preserved | | [Qwen3-VL-4b-Heretic-GGUF](https://huggingface.co/DreamFast/Qwen3-VL-4b-Heretic-GGUF) | llama.cpp, Ollama, LM Studio, ComfyUI-GGUF | GGUF quants from Q3_K_M up to F16 (text path) | | **[Qwen3-VL-4b-Heretic-ComfyUI](https://huggingface.co/DreamFast/Qwen3-VL-4b-Heretic-ComfyUI)** (this repo) | ComfyUI text encoder | bf16, fp8, int8, int4, nvfp4 and mxfp8 checkpoints | The Docker setup, scripts and configs that produced these files are in [Heretic Docker](https://github.com/dreamfast/heretic-docker). ## Why this variant? Several Heretic trials were run against Qwen3-VL-4B and all of them reach 100% HarmBench ASR, up from 30.8% on the base. This build was picked because, with safety tied, it wins on the tie-breakers: | | Base | **Heretic** | |---|------|------| | HarmBench ASR | 30.8% | **100%** | | KL divergence (lower is better) | | **0.0283** (lowest of the candidates) | | GSM8K | 78.62% | 77.18% (**−1.83%**, smallest drop) | | MMLU | 69.58% | 69.61% (+0.03%) | | Tensors changed | | 54 (pure rank-1) | See the [...
--- license: apache-2.0 library_name: transformers pipeline_tag: image-text-to-text base_model: Qwen/Qwen3-VL-4B-Instruct base_model_relation: quantized language: - en tags: - abliteration - heretic - uncensored - qwen3-vl - qwen3 - vision-language - text-encoder - comfyui - fp8 - int8 - convrot - nvfp4 - mxfp8 - blackwell - image-generation - krea - krea2 --- # Qwen3-VL-4B-Instruct Heretic (ComfyUI) 💬 **Community:** Join the [Abliterlitics Discord](https://discord.gg/AqmDnBjPvM) for discussion, model releases and support. ComfyUI checkpoints of the [Heretic abliteration](https://huggingface.co/DreamFast/Qwen3-VL-4b-Heretic) of [Qwen3-VL-4B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct). Five quantisations are included, covering everything from any modern GPU through to Blackwell, suitable as an uncensored text encoder for image-generation workflows including [Krea 2](https://docs.comfy.org/tutorials/image/krea/krea-2), which runs on a Qwen3VL-4B encoder. ## Available formats This model is published across three repos. Pick the one that matches your runtime. | Repo | Best for | Contents | |------|----------|----------| | [Qwen3-VL-4b-Heretic](https://huggingface.co/DreamFast/Qwen3-VL-4b-Heretic) | transformers, vLLM, HF Hub | bf16 weights with config, vision encoder preserved | | [Qwen3-VL-4b-Heretic-GGUF](https://huggingface.co/DreamFast/Qwen3-VL-4b-Heretic-GGUF) | llama.cpp, Ollama, LM Studio, ComfyUI-GGUF | GGUF quants from Q3_K_M up to F16 (text path) | | **[Qwen3-VL-4b-Heretic-ComfyUI](https://huggingface.co/DreamFast/Qwen3-VL-4b-Heretic-ComfyUI)** (this repo) | ComfyUI text encoder | bf16, fp8, int8, nvfp4 and mxfp8 checkpoints | The Docker setup, scripts and configs that produced these files are in [Heretic Docker](https://github.com/dreamfast/heretic-docker). ## Why this variant? Several Heretic trials were run against Qwen3-VL-4B and all of them reach 100% HarmBench ASR, up from 30.8% on the base. This build was picked because, with safety tied, it wins on the tie-breakers: | | Base | **Heretic** | |---|------|------| | HarmBench ASR | 30.8% | **100%** | | KL divergence (lower is better) | | **0.0283** (lowest of the candidates) | | GSM8K | 78.62% | 77.18% (**−1.83%**, smallest drop) | | MMLU | 69.58% | 69.61% (+0.03%) | | Tensors changed | | 54 (pure rank-1) | See the [full report](https://huggingface.c...
Source context: 0 downloads · 110 likes · Pipeline image-text-to-text · Library transformers · Repo DreamFast/Qwen3-VL-4b-Heretic-ComfyUI
| 4-bit float, double quantisation |
qwen3-vl-4b-heretic_mxfp8.safetensors | MXFP8 | 4.7 GB | Blackwell | Microscaling FP8, E8M0 block scales |