Repository navigation
feat(sdxl): run SDXL attention on SageAttention when enabled - #9663
Draft
Pfannkuchensack wants to merge 24 commits into
Draft
Pfannkuchensack wants to merge 24 commits into
Pfannkuchensack wants to merge 24 commits into
Conversation
New `attention_backend: auto | sage` setting; with `sage`, eligible PyTorch SDPA calls inside a denoising scope run on SageAttention 2.2 and every other call keeps SDPA unchanged. Each GPU's first SageAttention call is checked against SDPA and a failing GPU falls back for the session; no model family enters the scope yet. Adds the SageAttention docs page and the regenerated config schema.
With `attention_backend: sage`, FLUX.1 denoising, including ControlNets inside the loop, enters the SageAttention scope. RTX 4090, FLUX.1 dev FP8: 16.1 s → 14.0 s at 1024², 84.2 s → 54.8 s at 2048², peak VRAM unchanged; a blind review judged no image worse.
With `attention_backend: sage`, FLUX.2 denoising enters the SageAttention scope. RTX 4090, FLUX.2 Klein 9B FP8: 15.6 s → 15.1 s at 1024², 27.1 s → 23.1 s at 2048², peak VRAM unchanged; a blind review judged no image worse.
With `attention_backend: sage`, Z-Image denoising enters the SageAttention scope; Z-Image Control masks its attention and keeps SDPA. RTX 4090, Z-Image FP8: 10.4 s → 10.0 s at 1024², 33.7 s → 25.9 s at 2048², peak VRAM unchanged; a blind review judged no image worse.
With `attention_backend: sage`, Qwen-Image denoising enters the SageAttention scope. RTX 4090, Qwen-Image 2512 GGUF Q4_K_M: 78.8 s → 73.5 s at 1024², 208.9 s → 150.3 s at 2048², peak VRAM unchanged; a blind review judged no image worse.
With `attention_backend: sage`, Krea-2 denoising enters the SageAttention scope unless INVOKE_KREA2_SDPA_BACKEND pins a PyTorch kernel. RTX 4090, Krea-2 Turbo GGUF Q4_K_M: 11.5 s → 11.0 s at 1024², 42.7 s → 36.2 s at 2048², peak VRAM unchanged; a blind review judged no image worse.
With `attention_backend: sage`, SDXL denoising and tiled multi-diffusion enter the SageAttention scope; SD1.5 and SD2 do not. RTX 4090: 4.4 s → 4.2 s at 1024², 19.6 s → 16.1 s at 2048², peak VRAM unchanged; a blind review judged no image worse.
This was referenced Oct 3, 2026
…ent SageAttention failures sm86 now runs the Triton kernel upstream's sageattn picks there, and sm87 gets none; the Windows build's CUDA routing for both is unexplained and unmeasured here. Out-of-memory and Triton cache errors fall back to SDPA for that call (three retire the device), and the first-call check runs per precision and head size. The kernel runs with the tensors' device current, so a single-device install on cuda:1 uses SageAttention too.
An out-of-memory error now moves only the rest of that generation to SDPA, logged once, instead of counting toward retiring the device; out-of-memory is recognised by the shared is_oom_error. Only Triton cache failures in a row retire a device, and a success resets the count.
# Conflicts: # invokeai/frontend/api/openapi.json
# Conflicts: # invokeai/frontend/api/openapi.json # tests/test_config.py
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Note
Stacked PR 6 of 8: merge after #9662. It targets
main, so until the PRs before it merge, its diff also shows their commits and the merges that carry them forward. This PR's own change is one commit,1968eb9c58.Summary
With
attention_backend: sage, SDXL denoising and SDXL tiled multi-diffusion enter the SageAttention scope from #9657, so eligible attention calls run on SageAttention, including ControlNets that run inside the loop. SD1.5 and SD2 do not enter it: SD1.5's head sizes (40, 80, 160) never qualify, and SD2 is unmeasured. Withattention_type: sliced, SDXL's attention does not go through SDPA, so nothing changes there. The defaultautois unchanged. The docs page lists SDXL under Supported models and notes the sliced case.Related Issues / Discussions
Stacked on #9657 (SageAttention backend) and #9662, the PR before it.
QA Instructions
SDXL, model loaded, three arms per seed as described in #9657:
sageattn()'s default K smoothing; the selection feat(attention): add opt-in SageAttention backend for diffusion models #9657 ships is faster per call.To try it: install SageAttention as the docs page from #9657 describes, set
attention_backend: sage, restart and generate with SDXL. The first generation logsSageAttention ... serves diffusion-model attention on cuda:0 (...)with the GPU and kernel name, and withlog_level: debugevery generation logs how many calls ran on SageAttention and why the others did not. Withattention_backend: autothe output is bit-identical tomain.Checks: ruff clean;
pytest tests/app/invocationspasses with the whole stack applied (2090 passed). No call-site test was added: this change only enters the scope, whose behavior #9657 tests.Review
Resolved: the scope first opened for every UNet base, including the unmeasured SD2; it now opens for SDXL only.
Compatibility / Rollout
Merge after #9662. No API or persisted-state change; only
attention_backend: sagebehaves differently.Checklist
What's Newcopy (if doing a release after this PR)