INT8 row-wise quantized versions of Z-Image Base and Z-Image Turbo, using ConvRot (Hadamard-rotation outlier suppression) for improved quantization fidelity, plus a matching INT8 ConvRot quantization of the Qwen 3 4B...
Modellquelle
Quellenauszug
INT8 row-wise quantized versions of Z-Image Base and Z-Image Turbo, using ConvRot (Hadamard-rotation outlier suppression) for improved quantization fidelity, plus a matching INT8 ConvRot quantization of the Qwen 3 4B...
Verwandte Ressourcen
1 VerbindungQuellen
1 QuelleVerifiziert 29. Sept.
Modellartefakte
1 Artefaktqwen_3_4b_int8_convrot.safetensors
safetensors · INT8 · 4,20 GB · SHA-256 9199e1ec6301…e400 · Hugging Face
HerunterladenQuellenauszüge
2 Auszüge--- license: apache-2.0 tags: - text-to-image - int8 - quantized - convrot - comfyui - z-image base_model: - Tongyi-MAI/Z-Image-Turbo - Tongyi-MAI/Z-Image --- # Z-Image Turbo + Base — INT8 ConvRot (with Qwen 3 4B Text Encoder)  INT8 row-wise quantized versions of **Z-Image Base** and **Z-Image Turbo**, using ConvRot (Hadamard-rotation outlier suppression) for improved quantization fidelity, plus a matching INT8 ConvRot quantization of the Qwen 3 4B text encoder. Converted with [convert_to_quant](https://github.com/silveroxides/convert_to_quant) for native ComfyUI compatibility. ## Files | File | Description | |---|---| | `z_image_int8_convrot.safetensors` | Z-Image Base, INT8 + ConvRot | | `z_image_turbo_int8_convrot.safetensors` | Z-Image Turbo, INT8 + ConvRot | | `qwen_3_4b_int8_convrot.safetensors` | Qwen 3 4B text encoder, INT8 + ConvRot | ## Why ConvRot + Row-Wise Scaling ConvRot applies a group-wise Hadamard rotation to suppress weight outliers before quantization, improving INT8 fidelity versus plain per-tensor or per-row quantization alone. Critically, **these conversions use `--scaling_mode row`, not `tensor`**. Tensor-wise scaling computes a single scale factor for an entire weight matrix; even a small number of outlier values forces that global scale to widen, coarsening quantization precision across the rest of the matrix. In testing, this combination (ConvRot + tensor-wise scaling) produced visibly fuzzy, detail-smoothed output. Switching to row-wise scaling — which computes an independent scale per row, isolating outliers to the rows that contain them — resolved this and produced output sharpness matching or exceeding plain INT8 row-wise quantization. If you encounter other ConvRot-quantized models with soft or "waxy" output, this scaling mode mismatch is the most likely culprit. ## Quantization Recipe ``` ctq -i <model>.safetensors -o <model>-int8-convrot.safetensors \ --int8 --scaling_mode row --simple --low-memory \ --convrot --convrot-group-size 64 \ --zimage --comfy_quant --save-quant-metadata ``` The Qwen 3 4B text encoder was converted with the same flags, omitting `--zimage` (no architecture-specific preset needed for this text encoder; verify its native hidden dimensions div...
Source context: 0 downloads · 24 likes · Pipeline text-to-image · Repo Winnougan/Z-Image-Base-Turbo-INT8-Convrot