Skip to content
This repository was archived by the owner on Oct 4, 2026. It is now read-only.

chore(deps): move the cpu and cuda extras to torch 2.13.0 (cu130) - #369

Merged
Pfannkuchensack merged 4 commits into
upstream-mergefrom
chore/torch-2.13
Oct 1, 2026
Merged

Pfannkuchensack merged 4 commits into
upstream-mergefrom
chore/torch-2.13

Conversation

@Pfannkuchensack

Copy link
Copy Markdown
Member

Summary

The cpu and cuda extras were still on torch 2.7.1 (cu128) while rocm and xpu were already on 2.13.0. This moves them to 2.13.0+cpu / 2.13.0+cu130, so every backend on Windows and Linux x86_64 now runs the same torch. cu128 stops at torch 2.11, which is why the CUDA build moves to cu130. macOS (capped below 2.8) and linux-aarch64 stay on 2.7.1.

The bump broke three things. This PR fixes all three:

  • Seeds changed. torch 2.13 changed its CPU float16/bfloat16 randn/rand for tensors of 16+ elements; float32 is unchanged. SD1.5/SDXL (noise node), SD3, FLUX.1, FLUX.2, CogView4 and the Z-Image seed variance enhancer drew their seeded noise in half precision. The same seed gave a different image after the bump, and a different image on macOS (2.7.1) than on Windows/Linux. A new setting noise_dtype (default float32, optional float16) now controls the draw, through TorchDevice.choose_noise_dtype. After the fix, 2.7.1 and 2.13 give the same composition: 31–36 dB instead of 11–12 dB.
  • xformers stopped working. PyPI's xformers 0.0.35 is built for torch 2.10+cu128 and cannot load its extension next to cu130 torch, so every memory_efficient_attention call raised. InvokeAI prefers xformers on compute capability ≤ 7 when it is installed, and the launcher installs it for 20xx cards. xformers now comes from PyTorch's cu130 index, where it is built for CUDA 13.
  • Dropped GPUs and old drivers failed silently. cu130 needs compute capability 7.5+ (Turing and newer) and an R580+ driver. With an older driver, torch has no CUDA and Invoke quietly runs on the CPU. On Maxwell/Pascal/Volta, every generation fails with "no kernel image". A new startup check logs an error for both cases, and the docs name the floor, the symptoms and the ways out.

Before → after, RTX 4090, Windows, warm run, same code, fp8_compute on, model cache emptied before each case:

Model (1024², Wan 512²) 2.7.1 2.13
SDXL, 25 steps 5.0 s 4.5 s
FLUX.1 dev int8, 20 steps 20.7 s 18.7 s
FLUX.2 Klein 4B fp8, 4 steps (_scaled_mm) 2.5 s 2.0 s
Z-Image Turbo nvfp4, 8 steps 6.5 s 6.6 s
Krea-2 Turbo nvfp4, 8 steps 12.1 s 11.1 s
Wan 2.2 TI2V-5B GGUF Q4_K_M, 20 steps 12.1 s 10.6 s
RealESRGAN 4× upscale 4.0 s 4.0 s

This supports the claim "no warm regression for models that fit in 24 GB on this machine", not more (one warm sample per case).

Related Issues / Discussions

QA Instructions

Run on Windows with an RTX 4090 and driver 610.47 (reports CUDA 13.3), using a venv synced from the lock (uv sync --locked --extra cuda --extra test):

  • Lock and pins: uv lock --locked (uv 0.11.28), python scripts/check_pins.py and scripts/check_aarch64_lock.py pass.
    • uv export per extra × platform: cpu/cuda/xpu/rocm and the no-extra resolution give 2.13.0 on win32 and linux x86_64. darwin and linux-aarch64 give 2.7.1.
    • Only torch, torchvision, triton, the NVIDIA/CUDA runtime wheels and the xformers source moved.
    • torch.cuda.get_arch_list(): sm_75, sm_80, sm_86, sm_90, sm_100, sm_120.
  • Full suite: pytest -n 6 gives 11,674 passed, 184 skipped, 9 xfailed. This was run with HF_ENDPOINT unset: on this machine a local HF proxy makes 9 download/install tests fail, and they all pass without it.
  • CUDA-dependent tests on the 4090 pass: fp8 _scaled_mm, allocator, attention probe, nvfp4, bitsandbytes, Krea-2 SDPA, FLUX.2 working memory. test_sdpa_scope confirms torch's sdpa_kernel still leaks across threads in 2.13.
  • -m slow: 21 passed. The 3 errors in tests/backend/ip_adapter/test_ip_adapter.py come from a missing model_installer fixture and happen on main too.
  • New tests on both torch 2.7.1 and 2.13: tests/backend/util/test_noise_dtype.py and tests/app/util/test_cuda_build_compatibility.py.
    • Golden float32 values, identical on both versions.
    • Mutation runs: ignoring the setting fails 11 tests; always drawing in float32 fails 6.
  • E2E before/after (table above): all images inspected and clean.
    • With the final code, 2.7.1 vs 2.13 give SDXL 35.8 dB, FLUX.1 32.6 dB, upscale 34.3 dB, and Klein fp8 without fp8_compute (fp8 storage path) 31.2 dB.
    • On 2.13 the float32 SDXL image is bit-identical to the former float16-draw image.
  • Working-memory calibrators (scripts/calibrate_flux2_working_memory.py, calibrate_qwen_vae_working_memory.py) on both versions: every estimate is still an upper bound on the path CUDA actually dispatches.
  • xformers from the cu130 index (synced from the lock): memory_efficient_attention runs; max abs difference vs SDPA is 0.0. Its build metadata lists CUDA 13.0 and arch 7.5–12.1. uv pip compile --torch-backend cu130 also resolves xformers from whl/cu130, so manual installs get the right build.
  • Startup check on this machine: driver probe 13030, no error logged.
  • Docs: pnpm build and check-deploy-output pass. docs/src/generated/settings.json, openapi.json and schema.ts are regenerated.

Not tested:

  • Linux, and the Docker cuda image.
  • Turing, Pascal or older cards, and drivers older than R580. The startup check is covered with faked device queries only.
  • The --torch-backend=cu126 fallback for older GPUs. The docs say so.
  • NF4 (no model installed here).
  • Qwen-Image timings are not comparable: its ~29 GiB transformer does not fit in 24 GB, and the offload-bound runs swing about 4× in both directions on both versions. It served as a functional check only.

Review

Resolved:

  • The Z-Image seed variance enhancer's CPU bf16 torch.rand draw was also affected; it now draws through choose_noise_dtype.
  • PyPI xformers was incompatible with cu130.
  • There was no guidance or failure message for dropped GPUs and old drivers.
  • The noise_dtype description overstated what float16 restores.
  • A golden-value test now catches a future change to torch's float32 draw.
  • Stale aarch64 comments were corrected, and the startup check now parses suffixed arch entries (sm_90a).

Remaining risks and limitations:

  • Qwen-Image VAE CUDA constants. Measured decode peaks are 2770 on 2.13 (2734 on 2.7.1) against the shipped 2900 (4.7 % headroom). Encode peaks are 1557 against 1600 on both versions (2.8 %, same as before this PR). ROCm on Windows: keep VRAM resident and bound attention and VAE memory #304 rewrites exactly those lines, so the adjustment follows after it.
  • Golden-value test on macOS. Its tolerance (1e-4) allows for the ARM normal_fill path; it was measured on x86 only.
  • Pre-existing, unrelated flake. test_flux2_working_memory.py::…forced_math_forward[1-4096-512] measures 0 B when earlier CUDA tests ran in the same process. It fails the same way on 2.7.1.

Compatibility / Rollout

  • GPU floor. Nvidia GPUs now need compute capability 7.5+ and a driver from the R580 series or newer. This covers Docker hosts running the cuda image as well. Maxwell, Pascal and Volta (GTX 9xx/10xx, Titan V, P40/P100/V100) are no longer supported by the cuda extra; the release notes should say so.
  • Seeds. On Windows and Linux, images made before this update cannot be reproduced exactly for SD1.5/SDXL, SD3, FLUX.1, FLUX.2, CogView4 and Z-Image seed variance. On macOS, noise_dtype: float16 keeps the previous images. A FAQ entry explains this.
  • New config setting. noise_dtype (float32 | float16, default float32) is additive and needs no migration. openapi.json and schema.ts are regenerated.
  • pins.json. The cuda URLs moved to cu130. They are only read by launcher installs of releases before 6.14; v7 syncs --frozen from the lock.
  • Lock. [tool.uv] constraint-dependencies holds win32/linux x86_64 resolutions to torch 2.13.0. It is not part of the published metadata, so installing the released package with --torch-backend keeps the open range. uv pip install inside a checkout is held to it.

Checklist

  • The PR has a short but descriptive title, suitable for a changelog
  • Meaningful regression coverage added / updated where needed; obsolete tests/code removed
  • Persisted-state and API changes include required migrations / compatibility validation
  • Relevant performance/efficiency opportunities considered; material claims have evidence
  • Material review findings resolved and relevant checks rerun
  • Documentation added / updated (if applicable)
  • Updated What's New copy (if doing a release after this PR)

Pin cpu/cuda to 2.13.0+cpu/+cu130 (macOS and linux-aarch64 stay on 2.7.1), hold the no-extra lock to 2.13.0 and take xformers from the cu130 index.
Draw seeded noise in float32 by default (new `noise_dtype` setting), since torch changed its half-precision CPU random numbers.
Log a startup error for NVIDIA drivers older than the build's CUDA and GPUs below its lowest architecture, and document the new GPU floor.
@joshistoast

Copy link
Copy Markdown
Collaborator

conflictssssssss

Carries the branch's regenerated API contracts to their moved location under frontend/api.
@Pfannkuchensack
Pfannkuchensack changed the base branch from main to upstream-merge October 1, 2026 21:06
…-2.13

Keeps the CUDA or Intel XPU requirement and drops the CUDA 12.x mention from the FP8 Storage page, and regenerates the API contracts.
…-2.13

Regenerates the API contracts with both noise_dtype and the ROCm memory settings.
@Pfannkuchensack
Pfannkuchensack merged commit 19310a2 into upstream-merge Oct 1, 2026
13 checks passed
@Pfannkuchensack
Pfannkuchensack deleted the chore/torch-2.13 branch October 1, 2026 21:25
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants