fair-ai-public-license-1.0-sd
Model source
Source description
--bf16-unet command-line flag to work with this model. There will be differences in output with this flag enabled vs without the flag, even when using the normal 16-bit checkpoint. However, with the flag enabled, this DF11 model should produce identical outputs to the 16-bit model. Also, the original checkpoint weights are still required because of the Clip and VAE components.For more information (including how to compress models yourself), check out https://huggingface.co/DFloat11 and https://github.com/LeanModels/DFloat11
Sources
1 sourceVerified Aug 5
Model artifacts
1 artifactChenkinNoob-XL-V0.2-DF11.safetensors
safetensors · 3.41 GB · SHA-256 b18cd4655b36…7088 · Hugging Face
DownloadSource excerpts
3 excerptsSome SDXL-based checkpoints are actually distributed in BF16 format, instead of the more commonly used FP16, which makes them possible to be losslessly compressed, so I decided to give it a try. Currently, only the unet component (the main diffusion model) is compressed, but in SDXL it is the largest part of the pipeline anyway.
Unfortunately, the stock DF11 codebase does not support loading of compressed torch.nn.Conv2D tensors, even though they compress just fine, so I chose to skip compressing them, in order to avoid the need for users to manually patch the DFloat11 codebase. This makes the final compressed model ~200 MB larger than the expected size, but that is the price to pay for compatibility. Nevertheless, the reduction in VRAM footprint is still rather significant, from 5.14 GB to 3.66 GB, which should make the unet fit in 6 GB GPUs (assuming it has BF16 support).
Feel free to request for other models for compression as well (for either the diffusers library, ComfyUI, or any other model), although models that use architectures which are unfamiliar to me might be more difficult.
Install my own fork of the DF11 ComfyUI custom node: https://github.com/mingyi456/ComfyUI-DFloat11-Extended. After installing the DF11 custom node, simply replace the "Load Checkpoint" node of an existing SDXL workflow with the "Load Checkpoint with DF11 Unet" node. If you run into any issues, feel free to leave a comment.
diffusersRefer to this model instead.
This is the pattern_dict for compression:
pattern_dict_comfyui = {
r"time_embed" : (
"0",
"2",
),
r"label_emb.0" : (
"0",
"2",
),
r"input_blocks\.[12]\.0" : (
"emb_layers.1",
),
r"input_blocks\.4\.0" : (
"emb_layers.1",
),
r"input_blocks\.4\.1\.transformer_blocks\.\d+" : (
"attn1.to_q",
"attn1.to_k",
"attn1.to_v",
"attn1.to_out.0",
"ff.net.0.proj",
"ff.net.2",
"attn2.to_q",
"attn2.to_k",
"attn2.to_v",
"attn2.to_out.0",
),
r"input_blocks\.5\.0" : (
"emb_layers.1",
),
r"input_blocks\.5\.1\.transformer_blocks\.\d+" : (
"attn1.to_q",
"attn1.to_k",
"attn1.to_v",
"attn1.to_out.0",
"ff.net.0.proj",
"ff.net.2",
"attn2.to_q",
"attn2.to_k",
"attn2.to_v",
"attn2.to_out.0",
),
r"input_blocks\.7\.0" : (
"emb_layers.1",
),
r"input_blocks\.7\.1\.transformer_blocks\.\d+" : (
"attn1.to_q",
"attn1.to_k",
"attn1.to_v",
"attn1.to_out.0",
"ff.net.0.proj",
"ff.net.2",
"attn2.to_q",
"attn2.to_k",
"attn2.to_v",
"attn2.to_out.0",
),
r"input_blocks\.8\.0" : (
"emb_layers.1",
),
r"input_blocks\.8\.1\.transformer_blocks\.\d+" : (
"attn1.to_q",
"attn1.to_k",
"attn1.to_v",
"attn1.to_out.0",
"ff.net.0.proj",
"ff.net.2",
"attn2.to_q",
"attn2.to_k",
"attn2.to_v",
"attn2.to_out.0",
),
r"middle_block\.0" : (
"emb_layers.1",
),
r"middle_block\.1\.transformer_blocks\.\d+" : (
"attn1.to_q",
"attn1.to_k",
"attn1.to_v",
"attn1.to_out.0",
"ff.net.0.proj",
"ff.net.2",
"attn2.to_q",
"attn2.to_k",
"attn2.to_v",
"attn2.to_out.0",
),
r"middle_block\.2" : (
"emb_layers.1",
),
r"output_blocks\.[01]\.0" : (
"emb_layers.1",
),
r"output_blocks\.[01]\.1\.transformer_blocks\.\d+" : (
"attn1.to_q",
"attn1.to_k",
"attn1.to_v",
"attn1.to_out.0",
"ff.net.0.proj",
"ff.net.2",
"attn2.to_q",
"attn2.to_k",
"attn2.to_v",
"attn2.to_out.0",
),
r"output_blocks\.2\.0" : (
"emb_layers.1",
),
r"output_blocks\.2\.1\.transformer_blocks\.\d+" : (
"attn1.to_q",
"attn1.to_k",
"attn1.to_v",
"attn1.to_out.0",
"ff.net.0.proj",
"ff.net.2",
"attn2.to_q",
"attn2.to_k",
"attn2.to_v",
"attn2.to_out.0",
),
r"output_blocks\.3\.0" : (
"emb_layers.1",
),
r"output_blocks\.3\.1\.transformer_blocks\.\d+" : (
"attn1.to_q",
"attn1.to_k",
"attn1.to_v",
"attn1.to_out.0",
"ff.net.0.proj",
"ff.net.2",
"attn2.to_q",
"attn2.to_k",
"attn2.to_v",
"attn2.to_out.0",
),
r"output_blocks\.4\.0" : (
"emb_layers.1",
),
r"output_blocks\.4\.1\.transformer_blocks\.\d+" : (
"attn1.to_q",
"attn1.to_k",
"attn1.to_v",
"attn1.to_out.0",
"ff.net.0.proj",
"ff.net.2",
"attn2.to_q",
"attn2.to_k",
"attn2.to_v",
"attn2.to_out.0",
),
r"output_blocks\.5\.0" : (
"emb_layers.1",
),
r"output_blocks\.5\.1\.transformer_blocks\.\d+" : (
"attn1.to_q",
"attn1.to_k",
"attn1.to_v",
"attn1.to_out.0",
"ff.net.0.proj",
"ff.net.2",
"attn2.to_q",
"attn2.to_k",
"attn2.to_v",
"attn2.to_out.0",
),
r"output_blocks\.[678]\.0" : (
"emb_layers.1",
),
}
--- license: other license_name: fair-ai-public-license-1.0-sd license_link: https://freedevproject.org/faipl-1.0-sd/ language: - en base_model: - ChenkinNoob/ChenkinNoob-XL-V0.2 base_model_relation: quantized pipeline_tag: text-to-image tags: - comfyui - diffusion-single-file --- # Important: ComfyUI must be launched with the `--bf16-unet` command-line flag to work with this model. There will be differences in output with this flag enabled vs without the flag, even when using the normal 16-bit checkpoint. However, with the flag enabled, this DF11 model should produce identical outputs to the 16-bit model. Also, the original checkpoint weights are still required because of the Clip and VAE components. For more information (including how to compress models yourself), check out https://huggingface.co/DFloat11 and https://github.com/LeanModels/DFloat11 Some SDXL-based checkpoints are actually distributed in BF16 format, instead of the more commonly used FP16, which makes them possible to be losslessly compressed, so I decided to give it a try. Currently, only the unet component (the main diffusion model) is compressed, but in SDXL it is the largest part of the pipeline anyway. Unfortunately, the stock DF11 codebase does not support loading of compressed `torch.nn.Conv2D` tensors, even though they compress just fine, so I chose to skip compressing them, in order to avoid the need for users to manually patch the DFloat11 codebase. This makes the final compressed model ~200 MB larger than the expected size, but that is the price to pay for compatibility. Nevertheless, the reduction in VRAM footprint is still rather significant, from 5.14 GB to 3.66 GB, which should make the unet fit in 6 GB GPUs (assuming it has BF16 support). Feel free to request for other models for compression as well (for either the `diffusers` library, ComfyUI, or any other model), although models that use architectures which are unfamiliar to me might be more difficult. ### How to Use #### ComfyUI Install my own fork of the DF11 ComfyUI custom node: https://github.com/mingyi456/ComfyUI-DFloat11-Extended. After installing the DF11 custom node, simply replace the "Load Checkpoint" node of an existing SDXL workflow with the "Load Checkpoint with DF11 Unet" node. If you run into any issues, feel free to leave a comment. #### `diffusers` Refer to this [model](https://huggingface.co/ming...
Source context: 34 downloads · 0 likes · Pipeline text-to-image · Library diffusion-single-file · Repo mingyi456/ChenkinNoob-XL-V0.2-DF11-ComfyUI
Source context: 29 downloads · 0 likes · Pipeline text-to-image · Library diffusion-single-file · Repo mingyi456/ChenkinNoob-XL-V0.2-DF11-ComfyUI