This repository was archived by the owner on Oct 4, 2026. It is now read-only.
Repository navigation
chore(deps): move the cpu and cuda extras to torch 2.13.0 (cu130) - #369
Merged
Merged
Conversation
Pin cpu/cuda to 2.13.0+cpu/+cu130 (macOS and linux-aarch64 stay on 2.7.1), hold the no-extra lock to 2.13.0 and take xformers from the cu130 index. Draw seeded noise in float32 by default (new `noise_dtype` setting), since torch changed its half-precision CPU random numbers. Log a startup error for NVIDIA drivers older than the build's CUDA and GPUs below its lowest architecture, and document the new GPU floor.
Pfannkuchensack
requested review from
JPPhoto,
blessedcoolant and
lstein
as code owners
September 29, 2026 02:24
2 of 5 tasks
Collaborator
|
conflictssssssss |
Carries the branch's regenerated API contracts to their moved location under frontend/api.
Pfannkuchensack
enabled auto-merge
October 1, 2026 21:06
Pfannkuchensack
disabled auto-merge
October 1, 2026 21:06
…-2.13 Keeps the CUDA or Intel XPU requirement and drops the CUDA 12.x mention from the FP8 Storage page, and regenerates the API contracts.
…-2.13 Regenerates the API contracts with both noise_dtype and the ROCm memory settings.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The
cpuandcudaextras were still on torch 2.7.1 (cu128) whilerocmandxpuwere already on 2.13.0. This moves them to 2.13.0+cpu / 2.13.0+cu130, so every backend on Windows and Linux x86_64 now runs the same torch. cu128 stops at torch 2.11, which is why the CUDA build moves to cu130. macOS (capped below 2.8) and linux-aarch64 stay on 2.7.1.The bump broke three things. This PR fixes all three:
randn/randfor tensors of 16+ elements; float32 is unchanged. SD1.5/SDXL (noise node), SD3, FLUX.1, FLUX.2, CogView4 and the Z-Image seed variance enhancer drew their seeded noise in half precision. The same seed gave a different image after the bump, and a different image on macOS (2.7.1) than on Windows/Linux. A new settingnoise_dtype(defaultfloat32, optionalfloat16) now controls the draw, throughTorchDevice.choose_noise_dtype. After the fix, 2.7.1 and 2.13 give the same composition: 31–36 dB instead of 11–12 dB.memory_efficient_attentioncall raised. InvokeAI prefers xformers on compute capability ≤ 7 when it is installed, and the launcher installs it for 20xx cards. xformers now comes from PyTorch's cu130 index, where it is built for CUDA 13.Before → after, RTX 4090, Windows, warm run, same code,
fp8_computeon, model cache emptied before each case:_scaled_mm)This supports the claim "no warm regression for models that fit in 24 GB on this machine", not more (one warm sample per case).
Related Issues / Discussions
QA Instructions
Run on Windows with an RTX 4090 and driver 610.47 (reports CUDA 13.3), using a venv synced from the lock (
uv sync --locked --extra cuda --extra test):uv lock --locked(uv 0.11.28),python scripts/check_pins.pyandscripts/check_aarch64_lock.pypass.uv exportper extra × platform: cpu/cuda/xpu/rocm and the no-extra resolution give 2.13.0 on win32 and linux x86_64. darwin and linux-aarch64 give 2.7.1.torch.cuda.get_arch_list():sm_75, sm_80, sm_86, sm_90, sm_100, sm_120.pytest -n 6gives 11,674 passed, 184 skipped, 9 xfailed. This was run withHF_ENDPOINTunset: on this machine a local HF proxy makes 9 download/install tests fail, and they all pass without it._scaled_mm, allocator, attention probe, nvfp4, bitsandbytes, Krea-2 SDPA, FLUX.2 working memory.test_sdpa_scopeconfirms torch'ssdpa_kernelstill leaks across threads in 2.13.-m slow: 21 passed. The 3 errors intests/backend/ip_adapter/test_ip_adapter.pycome from a missingmodel_installerfixture and happen on main too.tests/backend/util/test_noise_dtype.pyandtests/app/util/test_cuda_build_compatibility.py.fp8_compute(fp8 storage path) 31.2 dB.scripts/calibrate_flux2_working_memory.py,calibrate_qwen_vae_working_memory.py) on both versions: every estimate is still an upper bound on the path CUDA actually dispatches.memory_efficient_attentionruns; max abs difference vs SDPA is 0.0. Its build metadata lists CUDA 13.0 and arch 7.5–12.1.uv pip compile --torch-backend cu130also resolves xformers fromwhl/cu130, so manual installs get the right build.pnpm buildandcheck-deploy-outputpass.docs/src/generated/settings.json,openapi.jsonandschema.tsare regenerated.Not tested:
--torch-backend=cu126fallback for older GPUs. The docs say so.Review
Resolved:
torch.randdraw was also affected; it now draws throughchoose_noise_dtype.noise_dtypedescription overstated whatfloat16restores.sm_90a).Remaining risks and limitations:
normal_fillpath; it was measured on x86 only.test_flux2_working_memory.py::…forced_math_forward[1-4096-512]measures 0 B when earlier CUDA tests ran in the same process. It fails the same way on 2.7.1.Compatibility / Rollout
noise_dtype: float16keeps the previous images. A FAQ entry explains this.noise_dtype(float32|float16, defaultfloat32) is additive and needs no migration.openapi.jsonandschema.tsare regenerated.pins.json. The cuda URLs moved to cu130. They are only read by launcher installs of releases before 6.14; v7 syncs--frozenfrom the lock.[tool.uv] constraint-dependenciesholds win32/linux x86_64 resolutions to torch 2.13.0. It is not part of the published metadata, so installing the released package with--torch-backendkeeps the open range.uv pip installinside a checkout is held to it.Checklist
What's Newcopy (if doing a release after this PR)