FLUXNATION
Fused CUDA kernel for FLUX.1 image generation — RTX 4090 optimized. 9.3 seconds for 20 steps at 1024×1024. No other kernel does what this does. Three technologies fused into one ComfyUI custom node: - FP8 fused CUDA kernel — replaces ComfyUI's SingleStreamBlock forward pass with a single torch.ops call. Modulation, L1 GEMM, QKV split, RoPE, L2 GEMM, gate+residual+bias all fused. Zero ctypes overhead. - Spike attention — neuromorphic block-sparse Triton attention. After step 10/20, scores attention blocks by dot-product similarity and only computes the top 45%. Saves 90% attention FLOPs on spike steps. - Step caching — alternates between computing spike attention and replaying the cached output. Zero attention compute on every other spike step. No fallbacks. If it breaks, it crashes — no silent quality degradation. - Windows (tested) / Linux (should work) - RTX 4090 (SM 8.9) — other Ampere/Ada cards may work but untested - Python 3.10 - CUDA 12.x - PyTorch 2.x with CUDA - ComfyUI - Visual Studio 2022 (for building the extension) bash git clone https://github.com/ULT7RA/cuda-kernels.git Copy FLUXNATION/customnodes/FLUXNATION/ into your ComfyUI custom nodes folder: ComfyUI/customnodes/FLUXNATION/ From the FLUXNATION/ folder: bat buildext.bat This requires: - Visual Studio 2022 (with C++ workload) - CUDA Toolkit 12.x - Python 3.10 The build installs fluxext.cp310-winamd64.pyd into your Python site-packages automatically. Set these environment variables before launching: bat set FLUXSPIKE=1 set FLUXSPIKESWITCH=0.50 set FLUXSPIKECAP=0.45 set FLUXSPIKETAU=0.05 set FLUXSPIKECACHE=…
FLUXNATION · 1 Eingabe · 0 Parameter · 1 Ausgabe