Skip to content

feat(deps): support ROCm on Windows through the rocm extra - #9606

Closed
0xDELUXA wants to merge 1 commit into
invoke-ai:mainfrom
0xDELUXA:windows-rocm
Closed

0xDELUXA wants to merge 1 commit into
invoke-ai:mainfrom
0xDELUXA:windows-rocm

Conversation

@0xDELUXA

Copy link
Copy Markdown
Contributor

Summary

Feature: invokeai[rocm] now installs on Windows. The rocm extra previously resolved only on Linux, so AMD GPUs on Windows fell back to the CPU.

On Windows the extra takes torch 2.13 and torchvision 0.28 from AMD's stable ROCm 10 index (https://stable.repo.amd.com/rocm/whl-next/), with GPU kernels for the RDNA3, RDNA3.5 and RDNA4 targets AMD supports on Windows. That index also mirrors common PyPI packages, so it is explicit, and the packages that exist only there are listed in the extra so their index source applies (uv applies sources to direct requirements only). AMD's Linux torch wheels require a Linux-only triton build that the Windows wheel does not use, so a dependency-metadata entry restates the Windows wheel's dependencies. The base torch requirement is split into linux and win32 lines so uv forks them, since the rocm extra now uses a different index on each.

bitsandbytes is floored at 0.50.2 on Windows, the first release that ships a ROCm binary there. Because uv resolves a single Windows fork, this also moves the Windows cpu and cuda installs from 0.49.2 to 0.50.2. Linux and macOS resolutions are unchanged for every extra and Python version.

TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL now defaults to 1 at startup. Without it, torch offers no fused SDPA kernel on GPUs AOTriton marks experimental, RDNA4 among them (on Windows and Linux), and falls back to the math backend. On an RX 9060 XT a 4096-token attention call takes 2.2 ms with flash and 36.9 ms on the math backend. An explicitly set value is respected.

Docs cover the Windows requirements and a manual uv pip install.

Related Issues / Discussions

Related: #9530, #9321

The launcher still routes Windows AMD GPUs to the CPU backend (hasWindowsAmdGpu in gpu-detection.ts), so a follow-up there is needed before launcher installs pick this up.

QA Instructions

Tested on Windows 11 with an RX 9060 XT (gfx1200):

  • uv sync --frozen --extra rocm installs torch 2.13.0+rocm10.0.0, and torch.cuda.is_available() is true.
  • Installed SD1.5 through the model manager and generated 512x512 images at 20 steps, 1.5 s per image after warmup.
  • bitsandbytes loads its ROCm binary; NF4 and int8 Linear layers produce correct output on the GPU.
  • pytest tests/test_check_pins.py tests/backend/pid passes, including the CUDA-marked PiD tests.
  • The manual uv pip install command from the docs installs the released package with ROCm torch.
  • Compared uv export per extra (none, cpu, cuda, rocm, xpu, each with and without xformers) against main on linux x86_64/aarch64, macOS arm64/x86_64 and Windows for Python 3.11 and 3.12: only the Windows changes described above differ.

Merge Plan

uv.lock is regenerated; if another dependency change lands first, rebase and rerun uv lock.

Checklist

  • The PR has a short but descriptive title, suitable for a changelog
  • Tests added / updated (if applicable)
  • ❗Changes to a redux slice have a corresponding migration
  • Documentation added / updated (if applicable)
  • Updated What's New copy (if doing a release after this PR)

@github-actions github-actions Bot added python PRs that change python files Root docs PRs that change docs python-deps PRs that change python dependencies labels Sep 28, 2026
@0xDELUXA

Copy link
Copy Markdown
Contributor Author

I'm not sure whether this PR should stay here or be opened against https://github.com/invoke-ai/InvokeAI-7.

Could the maintainers let me know whether they'd prefer this for v6 or v7?

It looks like even Windows on ARM64 support is planned for v7, so Windows ROCm support could potentially be added there as well. There are also a couple of related PRs over there: invoke-ai#304 and invoke-ai#369.

@lstein

lstein commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

I'm not sure whether this PR should stay here or be opened against https://github.com/invoke-ai/InvokeAI-7.

Could the maintainers let me know whether they'd prefer this for v6 or v7?

It looks like even Windows on ARM64 support is planned for v7, so Windows ROCm support could potentially be added there as well. There are also a couple of related PRs over there: invoke-ai#304 and invoke-ai#369.

Version 7 is about to be pulled into this repository. Your PR will be automatically retargeted to the v7 backend.

@lstein lstein self-assigned this Sep 29, 2026
@0xDELUXA

Copy link
Copy Markdown
Contributor Author

Version 7 is about to be pulled into this repository. Your PR will be automatically retargeted to the v7 backend.

Got it, thanks!

@Pfannkuchensack

Copy link
Copy Markdown
Member

i am waiting that invoke-ai#304 gets merged then i will add the rcom 10 stuff from you into invoke-ai#369 if thats ok for you @0xDELUXA ?

@0xDELUXA

Copy link
Copy Markdown
Contributor Author

i am waiting that invoke-ai#304 gets merged then i will add the rcom 10 stuff from you into invoke-ai#369 if thats ok for you @0xDELUXA ?

@Pfannkuchensack Sounds good, go ahead. Bringing the changes into invoke-ai#369 after invoke-ai#304 makes sense. A co-authored-by trailer on the commit that carries the changes over would be appreciated.

One part of #9606 needs to come along with the deps: run_app.py sets TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1 via os.environ.setdefault before torch is imported. On gfx1101, gfx1102, gfx1150, gfx1151 and gfx1200, PyTorch only enables the AOTriton flash and memory-efficient SDPA backends with that variable set, otherwise every SDPA call falls back to the math backend. InvokeAI-7 does not set it anywhere, so the "no fused SDPA kernel exists on gfx1200" finding in #304 matches a process without the flag rather than the hardware.

With the flag on an RX 9060 XT (gfx1200, torch 2.13.0+rocm10.0.0), a Z-Image-sized call at 1024px (30 heads, head_dim 128, 4608 tokens, bf16) dispatches to FLASH_ATTENTION: 8.8 ms per call and 68.5 MiB of extra peak memory, instead of materializing a 2.37 GiB fp32 score matrix. Output matches an fp32 math reference to 5e-4. Calls with head_dim above 256, like the VAE's 512-wide heads, have no fused kernel either way, so the chunked math path in #304 still matters for those.

@0xDELUXA

0xDELUXA commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

Closing as superseded by #9643.

@0xDELUXA 0xDELUXA closed this Oct 2, 2026
lstein pushed a commit that referenced this pull request Oct 10, 2026
The rocm extra moves from PyTorch's 2.13.0+rocm7.2 (Linux only) to AMD's 2.13.0+rocm10.0.0 for Linux
x86_64 and Windows. The lock carries the GPU kernel packages of every target AMD builds, so the launcher
can leave out the ones a machine does not need. Linux loses gfx900/gfx906, which ROCm 10 dropped.
bitsandbytes moves to 0.50.2, the first release with a Windows ROCm binary.

A new setting, rocm_aotriton_experimental (auto/on/off), sets TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL
before anything runs attention. auto turns AOTriton's experimental fused kernels on only for a ROCm 10
build on GPUs where they were measured correct (gfx1200 so far); under torch 2.12+rocm7.14 the same
switch made every fused call fail on that GPU.

Builds on the dependency and documentation changes of #9606.

Co-authored-by: 0xDELUXA <djernovevo@gmail.com>
lstein added a commit that referenced this pull request Oct 10, 2026
## Summary

The `rocm` extra moves from PyTorch's 2.13.0+rocm7.2, which exists for
Linux only, to AMD's **2.13.0+rocm10.0.0 for Linux x86_64 and Windows**.
Until now, `uv sync --extra rocm` on Windows installed PyPI's CPU torch.
It now installs ROCm, and the Windows ROCm path becomes a supported
install.

This carries over the dependency and documentation changes of #9606,
with @0xDELUXA as co-author. Three things changed on the way.

**Fused attention: a switch instead of a fixed
`TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1`.**
- On several RDNA3/3.5/4 GPUs, PyTorch runs its fused (AOTriton)
attention only with that variable set. Otherwise every attention call
falls back to the math kernel.
- Setting it unconditionally is not safe. Under torch 2.12+rocm7.14.1 it
made every flash and memory-efficient call on an RX 9060 XT fail with
`hipErrorInvalidValue`. Under 2.13.0+rocm10.0.0 every fused case matched
the math reference.
- A new setting, `rocm_aotriton_experimental` (`auto`/`on`/`off`),
decides. It is applied before anything runs attention. `auto` turns the
kernels on only for a ROCm 10 build, and only when every generation GPU
is one they were measured on: so far gfx1200. `on` and `off` work for
any build and GPU, and an exported variable always wins. The startup log
names the effective value.

**Integrated GPUs.** HIP enumerates the Radeon graphics of a Ryzen CPU
next to the discrete card. On the test machine (RX 9060 XT plus the
gfx1036 iGPU of a Ryzen 7900, Windows 10), the iGPU caused three
problems:
- `generation_devices: auto` started a second worker on it.
- Any kernel launched on it, even `torch.randn`, crashed the process
with an access violation.
- While HIP can merely see it, a process's first context on the discrete
card shows 17.9 GiB of shared GPU memory, and its first kernel launch
another 17.9 GiB. This happens with every ROCm build tried and without
any Invoke code.
- It is neither RAM nor commit charge: commit grew by 1.7 GiB, free RAM
fell by 0.8 GiB.
- But Task Manager shows 36 GiB of shared GPU memory, and the paging
check from #304 reads the same counter. It warned after every
generation, and the residency correction of the pre-tiling decision
dropped out.

The fix has two parts:
- `generation_devices: auto` and `device: auto` leave integrated GPUs
out on CUDA/ROCm too. Until now that filter existed for XPU only.
- On Windows ROCm, a short child process (1.4–2.4 s at startup) asks HIP
which devices are integrated, and `HIP_VISIBLE_DEVICES` hides them
before torch initializes HIP. This happens whenever no visibility
variable is set, whatever `device` and `generation_devices` say, so the
settings UI, the saved config and the runtime all count `cuda:N` among
the discrete GPUs. Naming an integrated GPU next to a discrete one is
refused with a clear error, since every kernel on it crashed the
process. The parent takes the child's answer as soon as it is printed
and ends the child; after 60 s without an answer it logs a warning with
the reason and leaves every GPU visible.

**Device packages.** AMD ships the GPU kernels as one package per target
and library. The lock carries all of them for both platforms
(gfx101x–gfx120x, gfx908/90a; on Linux also gfx942/950/1250). The
launcher can then leave out the targets a machine does not have with `uv
sync --no-install-package`. ROCm 10 has no kernels for gfx900/gfx906, so
Linux loses Vega 10/20 cards; the docs name the way out.

RX 9060 XT (gfx1200), Windows 10, venv synced from this lock, no
`HIP_VISIBLE_DEVICES` set. Warm runs, 1024², the rest from the first
run:

| | before | after |
|---|---|---|
| Z-Image Turbo nvfp4, 8 steps | 35.7 / 35.8 s (math kernel, `off`) |
16.7 / 17.2 s (`auto`) |
| Krea-2 Turbo fp8, 8 steps | ~60 s (#304, math kernel) | 32.1 / 32.1 s
|
| shared GPU memory held by the process (PDH) | 36.1 GiB | 0.26 GiB |

## Related Issues / Discussions

- Supersedes #9606 (Windows ROCm through the `rocm` extra); its changes
are carried over in the first commit, co-authored by @0xDELUXA.
- Builds on #304 (ROCm on Windows: VRAM residency, bounded attention and
VAE memory) and #369 (torch 2.13 for cpu/cuda).
- Follow-up in invoke-ai/launcher: route AMD on Windows to the `rocm`
extra, and leave out device packages the machine does not need.

## QA Instructions

**Lock and pins**
- `uv lock --locked` with uv 0.11.28, `scripts/check_pins.py` and
`scripts/check_aarch64_lock.py` pass.
- `uv export` per extra against `main`:
- only torch/torchvision/triton for `rocm` move, plus the 78 device
packages, `rocm`, `rocm-sdk-core`, `rocm-sdk-libraries` and
`rocm-bootstrap`;
- bitsandbytes goes to 0.50.2 for `rocm` on Windows and Linux x86_64:
AMD's torch reports HIP 7.15, for which 0.49 looks for a `rocm85`
library that does not exist. uv's fork preference also moves `cpu`,
`cuda` and the no-extra resolution to 0.50.2 on Windows and Linux
x86_64. 0.50 logs the int8 bf16→fp16 cast notice on every matmul;
`silence_int8_cast_notice()` filters it for the FLUX T5 and Qwen-Image
int8 paths;
  - `xpu` is unchanged.
- Two metadata overrides keep the lock independent of the machine that
writes it:
- The Linux torch wheel requires AMD's triton without a marker, and uv
reads one wheel's metadata for both platforms.
- `rocm` is an sdist whose build asks the machine's GPU which kernel
package to require.

**Installs from the lock**
- Windows: `uv sync --locked --extra rocm --extra test` installs. torch
2.13.0+rocm10.0.0 sees both GPUs and reports 22 arch targets.
- Linux (WSL, no GPU): `uv sync --locked --extra rocm` installs. `import
torch, triton` works, and `get_arch_list()` lists all 25 targets.
  - The venv is 20 GB, against 17 GB for the previous rocm7.2 lock.

**Tests**
- New: `tests/app/util/test_rocm_aotriton.py`,
`tests/app/util/test_rocm_integrated_gpu.py`, and CUDA/ROCm iGPU cases
in `tests/backend/util/test_devices.py`.
- Passing on the CUDA venv (RTX 4090, torch 2.13.0+cu130):
- `tests/backend/util`, `tests/app/util`,
`tests/backend/model_manager/load/model_cache`,
`tests/app/services/config`, `tests/app/services/session_processor`,
`tests/test_check_pins.py`;
  - the bitsandbytes NF4/LLM.int8 tests with bitsandbytes 0.50.2.
- Passing on the ROCm venv: `tests/backend/util`, `tests/app/util`,
`tests/app/routers/test_app_info.py`.
- `test_get_generation_devices_auto_expands_to_all_cuda` read the real
iGPU and is now hardware-independent.
-
`tests/backend/model_manager/load/model_cache/test_shared_weights_gpu.py`
crashes on the ROCm venv when the iGPU is visible: it computes on
`cuda:1`, which is the iGPU there. This is the crash described above,
not a regression.
- `-m slow` ROCm hardware tests, with and without the AOTriton variable:
6 pass, and the two canaries fail as they are meant to on this GPU:
- The Wan conv3d decomposition is 2× slower than native conv3d here (89
vs 42 ms).
  - With the variable, the fused kernels are correct for 512-wide heads.
  - Both workarounds stay (see Review).
- `ruff check`/`ruff format --check` are clean. `openapi.json`,
`schema.ts` and `docs/src/generated/settings.json` are regenerated.

**E2E on the RX 9060 XT** (table above):
- Images inspected, clean. With `auto`, PSNR against the rocm7.14.1
math-kernel reference is 42.2/32.1 dB, the same as in the earlier
measurement; `off` gives 57.8 dB.
- The startup log shows the iGPU hidden and the kernels on. No paging
warning after the fix; before it, one after every generation.

**Not tested:**
- Linux with an AMD GPU: @lstein, could you run it on the W7900?
- The Docker ROCm image build.
- GPUs other than gfx1200, and Windows 11 (AMD lists Windows 11; Windows
10 worked here).
- The launcher path.
- The full test suite locally (CI runs it).

## Review

Remaining risks and limitations:
- **AMD's index** serves no hashes and redirects per package. It must
stay explicit because it also mirrors PyPI packages. CI's `uv lock
--locked` now needs it reachable.
- **Size:**
- A plain `uv sync --extra rocm` installs the kernels of every target,
about 4.3 GiB extra download on Windows and 8.3 GiB on Linux.
- The Docker ROCm image grows accordingly. Its build job has a 90-minute
timeout, and the image build has not been run.
- **Two ROCm workarounds stay, measured as unnecessary on gfx1200/ROCm
10:**
- The wide-head math guard: with the variable, the fused kernels compute
512-wide heads correctly here, 37.9 ms instead of 93 ms per VAE
mid-block call.
  - The Wan conv3d decomposition: 2× slower here.
- Both were measured necessary on a W7900 under rocm7.2, and RDNA3 under
ROCm 10 is unmeasured, so dropping them is left for a follow-up with
that measurement.
- **`auto` covers gfx1200 only.** gfx1101, gfx1102, gfx1150 and gfx1151
need the variable too (per #9606), but have not been measured here; on
them, `on` is an opt-in.
- **An AMD APU used on its own** (e.g. Ryzen AI Max) keeps the model
cache's discrete-GPU memory model. That case is unmeasured.

## Compatibility / Rollout

- **New setting:** `rocm_aotriton_experimental` (`auto` | `on` | `off`,
default `auto`). It is additive, needs no migration, and the API
contracts are regenerated.
- **`HIP_VISIBLE_DEVICES` on Windows ROCm** is now set at startup when
an iGPU sits next to a discrete GPU, unless the user set it,
`CUDA_VISIBLE_DEVICES` or `ROCR_VISIBLE_DEVICES`. Explicit `cuda:N`
values then count the discrete GPUs only. Anyone who ran `main` with a
self-installed Windows ROCm torch and an explicit `cuda:N` in HIP's full
numbering gets renumbered once: an index past the end fails at startup
with a clear error, and an iGPU index is refused, but with two discrete
GPUs behind an iGPU, `cuda:1` would silently select the other card.
- **`pins.json`:** `linux.rocm` points at AMD's index. It is read only
by launcher installs of releases before 6.14, so legacy Windows installs
still get no ROCm entry.
- **GPU support on Linux:** gfx900/gfx906 (Radeon VII, Vega 56/64,
MI25/MI50/MI60) are no longer covered by the `rocm` extra. The docs
point to the previous release or `--torch-backend=rocm7.2`.

## Checklist

- [x] _The PR has a short but descriptive title, suitable for a
changelog_
- [x] _Meaningful regression coverage added / updated where needed;
obsolete tests/code removed_
- [x] _Persisted-state and API changes include required migrations /
compatibility validation_
- [x] _Relevant performance/efficiency opportunities considered;
material claims have evidence_
- [x] _Material review findings resolved and relevant checks rerun_
- [x] _Documentation added / updated (if applicable)_
- [ ] _Updated `What's New` copy (if doing a release after this PR)_
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

docs PRs that change docs python PRs that change python files python-deps PRs that change python dependencies Root

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants