A Control LoRA model trained on Stable Diffusion XL to control image generation through brightness/grayscale information. This model uses LoRA (Low-Rank Adaptation) combined with ControlNet architecture for efficient...
Fonte do modelo
Descrição da fonte
A Control LoRA model trained on Stable Diffusion XL to control image generation through brightness/grayscale information. This model uses LoRA (Low-Rank Adaptation) combined with ControlNet architecture for efficient control, providing an ultra-lightweight alternative to full ControlNet with excellent pattern preservation.
This Control LoRA enables brightness-based conditioning for SDXL image generation. By providing a grayscale image as input, you can control the brightness distribution and lighting structure while maintaining creative freedom through text prompts.
Fontes
1 fonteVerificado 29 de jul.
Artefatos de modelo
9 artefatosTrechos de fonte
2 trechosTrained on 100,000 samples from latentcat/grayscale_image_aesthetic_3M:
| Parameter | Value |
|---|---|
| Base Model | stabilityai/stable-diffusion-xl-base-1.0 |
| VAE | madebyollin/sdxl-vae-fp16-fix (improved stability) |
| Architecture | ControlLoRA v3 (~7M trainable parameters) |
| LoRA Rank | 16 |
| Extra Conv Rank | 64 (conv_in layer) |
| Training Resolution | 1024×1024 |
| Training Steps | 3,125 (1 epoch) |
| Batch Size | 8 per device |
| Gradient Accumulation | 4 (effective batch: 32) |
| Learning Rate | 1e-4 constant (no decay) |
| LR Warmup | 0 steps |
| Model | Parameters | Size | Training | Resolution |
|---|---|---|---|---|
| This Control LoRA | ~7M | ~24MB | 100k @ 1024 | 1024×1024 |
| ControlNet (SDXL) | ~700M | 4.7GB | 100k @ 512 | 512×512 |
| T2I Adapter (SDXL) | ~77M | 302MB | 100k @ 1024 | 1024×1024 |
| Flux Control LoRA | ~7M | 25MB | 10k @ 512 | 512×512 |
pip install diffusers transformers accelerate torch peft
# Install ControlLoRA v3
git clone https://github.com/HighCWu/control-lora-v3
import torch
import sys
sys.path.insert(0, '/path/to/control-lora-v3')
from pipeline_sdxl import StableDiffusionXLControlLoraV3Pipeline
from model import UNet2DConditionModelEx
from diffusers import AutoencoderKL
from PIL import Image
# Load improved VAE (same as used in training)
vae = AutoencoderKL.from_pretrained(
"madebyollin/sdxl-vae-fp16-fix",
torch_dtype=torch.float16,
)
# Load UNet with LoRA support
unet = UNet2DConditionModelEx.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
subfolder="unet",
torch_dtype=torch.bfloat16,
)
unet = unet.add_extra_conditions(["brightness"])
# Load SDXL Control LoRA pipeline with improved VAE
pipe = StableDiffusionXLControlLoraV3Pipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
vae=vae,
unet=unet,
torch_dtype=torch.bfloat16,
)
# Load Control LoRA weights
pipe.load_lora_weights("Oysiyl/controlnet-lora-brightness-sdxl", adapter_name="brightness")
pipe.to("cuda")
# Load grayscale/brightness control image
control_image = Image.open("path/to/grayscale_image.png")
control_image = control_image.resize((1024, 1024))
# Generate image
prompt = "a beautiful garden scene with colorful flowers and butterflies, highly detailed, professional photography, vibrant colors"
image = pipe(
prompt=prompt,
image=control_image,
num_inference_steps=30,
guidance_scale=7.5,
extra_condition_scale=1.0, # Controls conditioning strength
height=1024,
width=1024,
).images[0]
image.save("output.png")
The extra_condition_scale parameter controls how strongly the brightness map influences generation:
# Subtle control (scale 0.5-0.7)
image = pipe(
prompt=prompt,
image=control_image,
extra_condition_scale=0.5,
...
).images[0]
# Balanced control (scale 1.0-1.5) - Recommended for artistic QR codes
image = pipe(
prompt=prompt,
image=control_image,
extra_condition_scale=1.0,
...
).images[0]
# Strong control (scale 1.5-2.0)
image = pipe(
prompt=prompt,
image=control_image,
extra_condition_scale=1.5,
...
).images[0]
import qrcode
from PIL import Image
# Generate QR code
qr = qrcode.QRCode(
version=1,
error_correction=qrcode.constants.ERROR_CORRECT_H,
box_size=10,
border=4
)
qr.add_data("https://your-url.com")
qr.make(fit=True)
qr_image = qr.make_image(fill_color="black", back_color="white")
qr_image = qr_image.resize((1024, 1024), Image.LANCZOS).convert("RGB")
# Generate artistic QR code (scale 1.0-1.5 works best)
image = pipe(
prompt="a beautiful garden with colorful flowers and butterflies, highly detailed, professional photography",
image=qr_image,
num_inference_steps=30,
guidance_scale=7.5,
extra_condition_scale=1.0,
height=1024,
width=1024,
).images[0]
image.save("artistic_qr.png")
The model includes intermediate checkpoints from throughout training:
# Early checkpoint (25% - 25,000 samples)
pipe.load_lora_weights("Oysiyl/controlnet-lora-brightness-sdxl",
adapter_name="brightness",
subfolder="checkpoint-781")
# Mid checkpoint (50% - 50,000 samples)
pipe.load_lora_weights("Oysiyl/controlnet-lora-brightness-sdxl",
adapter_name="brightness",
subfolder="checkpoint-1562")
# Late checkpoint (75% - 75,000 samples)
pipe.load_lora_weights("Oysiyl/controlnet-lora-brightness-sdxl",
adapter_name="brightness",
subfolder="checkpoint-2343")
# Near-final checkpoint (99% - 99,968 samples)
pipe.load_lora_weights("Oysiyl/controlnet-lora-brightness-sdxl",
adapter_name="brightness",
subfolder="checkpoint-3124")
# Final model (100,000 samples, main branch - recommended)
pipe.load_lora_weights("Oysiyl/controlnet-lora-brightness-sdxl",
adapter_name="brightness")
The extra_condition_scale parameter controls how strongly the brightness map influences generation:
| Scale | Behavior | Best For |
|---|---|---|
| 0.5-0.7 | Subtle artistic integration with hints of pattern | Natural images, soft lighting hints |
| 0.7-1.0 | Light control - visible structure with artistic freedom | Artistic images, creative reinterpretation |
| 1.0-1.5 | 🔥 Balanced control | Artistic QR codes, watermarks (recommended) |
| 1.5-2.0 | Strong control - clear patterns with artistic overlay | Geometric patterns, structured designs |
| 2.0+ | Maximum control - dominant patterns | Strong brightness maps, technical applications |
| Metric | ControlNet (SDXL) | This Control LoRA | Advantage |
|---|---|---|---|
| Parameters | ~700M | ~7M | 100x smaller |
| Model Size | 4.7GB | 24MB | 196x smaller |
| Load Time | ~5-10 seconds | <1 second | 10x faster loading |
| Storage (w/ checkpoints) | ~18.8GB | ~490MB | 38x less storage |
| Training Time | ~3 hours | 4.5 hours | Comparable |
| Pattern Preservation @ Scale 1.0 | Excellent | Excellent | Comparable quality |
| Metric | T2I Adapter (SDXL) | This Control LoRA | Advantage |
|---|---|---|---|
| Parameters | ~77M | ~7M | 11x smaller |
| Model Size | 302MB | 24MB | 12.6x smaller |
| Training Samples | 100k | 100k | Matched data |
| Architecture | Separate adapter | Integrated LoRA | Simpler loading |
The model includes checkpoints from throughout training:
Each comparison shows QR input + all 7 conditioning scales (0.25, 0.5, 0.7, 0.75, 1.0, 1.25, 1.5) for a specific checkpoint:
Checkpoint-781 Scale Progression
Checkpoint-1562 Scale Progression
Checkpoint-2343 Scale Progression
Checkpoint-3124 Scale Progression
Final Model Scale Progression
All checkpoints show consistent, high-quality performance across scales. The progression analysis reveals:
Early Checkpoint (781 steps, 25k samples):
Mid Checkpoints (1562-2343 steps, 50k-75k samples):
Final Model (3125 steps, 100k samples):
checkpoint-234/pytorch_lora_weights.safetensors
safetensors · 23,5 MB · SHA-256 5078514e6ad4…48d7 · Hugging Face
Baixarcheckpoint-2343/pytorch_lora_weights.safetensors
safetensors · 93,8 MB · SHA-256 042ecc91bab4…6719 · Hugging Face
Baixarcheckpoint-312/pytorch_lora_weights.safetensors
safetensors · 23,5 MB · SHA-256 2b9dc11a212d…130c · Hugging Face
Baixarcheckpoint-3124/pytorch_lora_weights.safetensors
safetensors · 93,8 MB · SHA-256 dab6d1c27f02…da68 · Hugging Face
Baixarcheckpoint-78/pytorch_lora_weights.safetensors
safetensors · 23,5 MB · SHA-256 23b178b5c75e…ef61 · Hugging Face
Baixarcheckpoint-781/pytorch_lora_weights.safetensors
safetensors · 93,8 MB · SHA-256 87e8a9dfe06d…ecad · Hugging Face
Baixarpytorch_lora_weights.safetensors
safetensors · 93,8 MB · SHA-256 b006bb0d4a8f…ed2e · Hugging Face
Baixar--- license: apache-2.0 base_model: stabilityai/stable-diffusion-xl-base-1.0 tags: - stable-diffusion-xl - sdxl - text-to-image - diffusers - lora - control - controlnet - control-lora - brightness - grayscale - template:sd-lora widget: - text: "a beautiful garden scene with colorful flowers and butterflies, highly detailed, professional photography, vibrant colors" output: url: "https://huggingface.co/Oysiyl/controlnet-lora-brightness-sdxl/resolve/main/examples/example.png" inference: true --- # ControlNet LoRA SDXL - Brightness Control (100k @ 1024×1024) A Control LoRA model trained on Stable Diffusion XL to control image generation through brightness/grayscale information. This model uses **LoRA (Low-Rank Adaptation)** combined with ControlNet architecture for efficient control, providing an **ultra-lightweight alternative** to full ControlNet with excellent pattern preservation. ## Model Description This Control LoRA enables brightness-based conditioning for SDXL image generation. By providing a grayscale image as input, you can control the brightness distribution and lighting structure while maintaining creative freedom through text prompts. ### Key Features: - 🎨 **Excellent brightness and pattern control** across multiple scales (0.5-2.0) - 🚀 **196x smaller than full ControlNet**: ~24MB vs ~4.7GB - ⚡ **Ultra-fast loading**: LoRA weights load in <1 second - 💡 **Flexible scale control**: Adjustable conditioning scale from 0.5 to 2.0+ - 🔄 **Compatible with ControlLoRA v3**: Uses the efficient ControlLoRA v3 architecture - 📦 **Minimal storage**: All checkpoints + final model = ~490MB total - 🖼️ **Native SDXL resolution**: Trained at 1024×1024 - 🎯 **Production-scale training**: 100,000 samples with PiSSA initialization ### Intended Uses: - **Artistic QR code generation** (scale 1.0-1.5 recommended) - Image recoloring and colorization - Lighting control in text-to-image generation - Brightness-based pattern integration - Watermark and subtle pattern embedding - Photo enhancement and stylization ## Training Details ### Training Data Trained on 100,000 samples from `latentcat/grayscale_image_aesthetic_3M`: - High-quality aesthetic images - Paired with grayscale/brightness versions - Native resolution: 1024×1024 (SDXL native) ### Training Configuration | Parameter | Value | |-----------|-------| | **...
Source context: 26 downloads · 0 likes · Pipeline text-to-image · Library diffusers · Repo Oysiyl/controlnet-lora-brightness-sdxl
| Empty Prompts | 20% (improved from 10k's 10%) |
| Init Method | PiSSA (niter=4) for faster convergence |
| Mixed Precision | BF16 |
| Hardware | NVIDIA H100 80GB |
| Training Time | ~4.5 hours |
| Final Loss | ~0.05 |
| Fixed architecture |
| Adjustable weights |
| More versatile |