Repository navigation
Conversation
|
I'm not sure whether this PR should stay here or be opened against https://github.com/invoke-ai/InvokeAI-7. Could the maintainers let me know whether they'd prefer this for v6 or v7? It looks like even Windows on ARM64 support is planned for v7, so Windows ROCm support could potentially be added there as well. There are also a couple of related PRs over there: invoke-ai#304 and invoke-ai#369. |
Version 7 is about to be pulled into this repository. Your PR will be automatically retargeted to the v7 backend. |
Got it, thanks! |
|
i am waiting that invoke-ai#304 gets merged then i will add the rcom 10 stuff from you into invoke-ai#369 if thats ok for you @0xDELUXA ? |
@Pfannkuchensack Sounds good, go ahead. Bringing the changes into invoke-ai#369 after invoke-ai#304 makes sense. A co-authored-by trailer on the commit that carries the changes over would be appreciated. One part of #9606 needs to come along with the deps: With the flag on an RX 9060 XT ( |
|
Closing as superseded by #9643. |
The rocm extra moves from PyTorch's 2.13.0+rocm7.2 (Linux only) to AMD's 2.13.0+rocm10.0.0 for Linux x86_64 and Windows. The lock carries the GPU kernel packages of every target AMD builds, so the launcher can leave out the ones a machine does not need. Linux loses gfx900/gfx906, which ROCm 10 dropped. bitsandbytes moves to 0.50.2, the first release with a Windows ROCm binary. A new setting, rocm_aotriton_experimental (auto/on/off), sets TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL before anything runs attention. auto turns AOTriton's experimental fused kernels on only for a ROCm 10 build on GPUs where they were measured correct (gfx1200 so far); under torch 2.12+rocm7.14 the same switch made every fused call fail on that GPU. Builds on the dependency and documentation changes of #9606. Co-authored-by: 0xDELUXA <djernovevo@gmail.com>
## Summary The `rocm` extra moves from PyTorch's 2.13.0+rocm7.2, which exists for Linux only, to AMD's **2.13.0+rocm10.0.0 for Linux x86_64 and Windows**. Until now, `uv sync --extra rocm` on Windows installed PyPI's CPU torch. It now installs ROCm, and the Windows ROCm path becomes a supported install. This carries over the dependency and documentation changes of #9606, with @0xDELUXA as co-author. Three things changed on the way. **Fused attention: a switch instead of a fixed `TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1`.** - On several RDNA3/3.5/4 GPUs, PyTorch runs its fused (AOTriton) attention only with that variable set. Otherwise every attention call falls back to the math kernel. - Setting it unconditionally is not safe. Under torch 2.12+rocm7.14.1 it made every flash and memory-efficient call on an RX 9060 XT fail with `hipErrorInvalidValue`. Under 2.13.0+rocm10.0.0 every fused case matched the math reference. - A new setting, `rocm_aotriton_experimental` (`auto`/`on`/`off`), decides. It is applied before anything runs attention. `auto` turns the kernels on only for a ROCm 10 build, and only when every generation GPU is one they were measured on: so far gfx1200. `on` and `off` work for any build and GPU, and an exported variable always wins. The startup log names the effective value. **Integrated GPUs.** HIP enumerates the Radeon graphics of a Ryzen CPU next to the discrete card. On the test machine (RX 9060 XT plus the gfx1036 iGPU of a Ryzen 7900, Windows 10), the iGPU caused three problems: - `generation_devices: auto` started a second worker on it. - Any kernel launched on it, even `torch.randn`, crashed the process with an access violation. - While HIP can merely see it, a process's first context on the discrete card shows 17.9 GiB of shared GPU memory, and its first kernel launch another 17.9 GiB. This happens with every ROCm build tried and without any Invoke code. - It is neither RAM nor commit charge: commit grew by 1.7 GiB, free RAM fell by 0.8 GiB. - But Task Manager shows 36 GiB of shared GPU memory, and the paging check from #304 reads the same counter. It warned after every generation, and the residency correction of the pre-tiling decision dropped out. The fix has two parts: - `generation_devices: auto` and `device: auto` leave integrated GPUs out on CUDA/ROCm too. Until now that filter existed for XPU only. - On Windows ROCm, a short child process (1.4–2.4 s at startup) asks HIP which devices are integrated, and `HIP_VISIBLE_DEVICES` hides them before torch initializes HIP. This happens whenever no visibility variable is set, whatever `device` and `generation_devices` say, so the settings UI, the saved config and the runtime all count `cuda:N` among the discrete GPUs. Naming an integrated GPU next to a discrete one is refused with a clear error, since every kernel on it crashed the process. The parent takes the child's answer as soon as it is printed and ends the child; after 60 s without an answer it logs a warning with the reason and leaves every GPU visible. **Device packages.** AMD ships the GPU kernels as one package per target and library. The lock carries all of them for both platforms (gfx101x–gfx120x, gfx908/90a; on Linux also gfx942/950/1250). The launcher can then leave out the targets a machine does not have with `uv sync --no-install-package`. ROCm 10 has no kernels for gfx900/gfx906, so Linux loses Vega 10/20 cards; the docs name the way out. RX 9060 XT (gfx1200), Windows 10, venv synced from this lock, no `HIP_VISIBLE_DEVICES` set. Warm runs, 1024², the rest from the first run: | | before | after | |---|---|---| | Z-Image Turbo nvfp4, 8 steps | 35.7 / 35.8 s (math kernel, `off`) | 16.7 / 17.2 s (`auto`) | | Krea-2 Turbo fp8, 8 steps | ~60 s (#304, math kernel) | 32.1 / 32.1 s | | shared GPU memory held by the process (PDH) | 36.1 GiB | 0.26 GiB | ## Related Issues / Discussions - Supersedes #9606 (Windows ROCm through the `rocm` extra); its changes are carried over in the first commit, co-authored by @0xDELUXA. - Builds on #304 (ROCm on Windows: VRAM residency, bounded attention and VAE memory) and #369 (torch 2.13 for cpu/cuda). - Follow-up in invoke-ai/launcher: route AMD on Windows to the `rocm` extra, and leave out device packages the machine does not need. ## QA Instructions **Lock and pins** - `uv lock --locked` with uv 0.11.28, `scripts/check_pins.py` and `scripts/check_aarch64_lock.py` pass. - `uv export` per extra against `main`: - only torch/torchvision/triton for `rocm` move, plus the 78 device packages, `rocm`, `rocm-sdk-core`, `rocm-sdk-libraries` and `rocm-bootstrap`; - bitsandbytes goes to 0.50.2 for `rocm` on Windows and Linux x86_64: AMD's torch reports HIP 7.15, for which 0.49 looks for a `rocm85` library that does not exist. uv's fork preference also moves `cpu`, `cuda` and the no-extra resolution to 0.50.2 on Windows and Linux x86_64. 0.50 logs the int8 bf16→fp16 cast notice on every matmul; `silence_int8_cast_notice()` filters it for the FLUX T5 and Qwen-Image int8 paths; - `xpu` is unchanged. - Two metadata overrides keep the lock independent of the machine that writes it: - The Linux torch wheel requires AMD's triton without a marker, and uv reads one wheel's metadata for both platforms. - `rocm` is an sdist whose build asks the machine's GPU which kernel package to require. **Installs from the lock** - Windows: `uv sync --locked --extra rocm --extra test` installs. torch 2.13.0+rocm10.0.0 sees both GPUs and reports 22 arch targets. - Linux (WSL, no GPU): `uv sync --locked --extra rocm` installs. `import torch, triton` works, and `get_arch_list()` lists all 25 targets. - The venv is 20 GB, against 17 GB for the previous rocm7.2 lock. **Tests** - New: `tests/app/util/test_rocm_aotriton.py`, `tests/app/util/test_rocm_integrated_gpu.py`, and CUDA/ROCm iGPU cases in `tests/backend/util/test_devices.py`. - Passing on the CUDA venv (RTX 4090, torch 2.13.0+cu130): - `tests/backend/util`, `tests/app/util`, `tests/backend/model_manager/load/model_cache`, `tests/app/services/config`, `tests/app/services/session_processor`, `tests/test_check_pins.py`; - the bitsandbytes NF4/LLM.int8 tests with bitsandbytes 0.50.2. - Passing on the ROCm venv: `tests/backend/util`, `tests/app/util`, `tests/app/routers/test_app_info.py`. - `test_get_generation_devices_auto_expands_to_all_cuda` read the real iGPU and is now hardware-independent. - `tests/backend/model_manager/load/model_cache/test_shared_weights_gpu.py` crashes on the ROCm venv when the iGPU is visible: it computes on `cuda:1`, which is the iGPU there. This is the crash described above, not a regression. - `-m slow` ROCm hardware tests, with and without the AOTriton variable: 6 pass, and the two canaries fail as they are meant to on this GPU: - The Wan conv3d decomposition is 2× slower than native conv3d here (89 vs 42 ms). - With the variable, the fused kernels are correct for 512-wide heads. - Both workarounds stay (see Review). - `ruff check`/`ruff format --check` are clean. `openapi.json`, `schema.ts` and `docs/src/generated/settings.json` are regenerated. **E2E on the RX 9060 XT** (table above): - Images inspected, clean. With `auto`, PSNR against the rocm7.14.1 math-kernel reference is 42.2/32.1 dB, the same as in the earlier measurement; `off` gives 57.8 dB. - The startup log shows the iGPU hidden and the kernels on. No paging warning after the fix; before it, one after every generation. **Not tested:** - Linux with an AMD GPU: @lstein, could you run it on the W7900? - The Docker ROCm image build. - GPUs other than gfx1200, and Windows 11 (AMD lists Windows 11; Windows 10 worked here). - The launcher path. - The full test suite locally (CI runs it). ## Review Remaining risks and limitations: - **AMD's index** serves no hashes and redirects per package. It must stay explicit because it also mirrors PyPI packages. CI's `uv lock --locked` now needs it reachable. - **Size:** - A plain `uv sync --extra rocm` installs the kernels of every target, about 4.3 GiB extra download on Windows and 8.3 GiB on Linux. - The Docker ROCm image grows accordingly. Its build job has a 90-minute timeout, and the image build has not been run. - **Two ROCm workarounds stay, measured as unnecessary on gfx1200/ROCm 10:** - The wide-head math guard: with the variable, the fused kernels compute 512-wide heads correctly here, 37.9 ms instead of 93 ms per VAE mid-block call. - The Wan conv3d decomposition: 2× slower here. - Both were measured necessary on a W7900 under rocm7.2, and RDNA3 under ROCm 10 is unmeasured, so dropping them is left for a follow-up with that measurement. - **`auto` covers gfx1200 only.** gfx1101, gfx1102, gfx1150 and gfx1151 need the variable too (per #9606), but have not been measured here; on them, `on` is an opt-in. - **An AMD APU used on its own** (e.g. Ryzen AI Max) keeps the model cache's discrete-GPU memory model. That case is unmeasured. ## Compatibility / Rollout - **New setting:** `rocm_aotriton_experimental` (`auto` | `on` | `off`, default `auto`). It is additive, needs no migration, and the API contracts are regenerated. - **`HIP_VISIBLE_DEVICES` on Windows ROCm** is now set at startup when an iGPU sits next to a discrete GPU, unless the user set it, `CUDA_VISIBLE_DEVICES` or `ROCR_VISIBLE_DEVICES`. Explicit `cuda:N` values then count the discrete GPUs only. Anyone who ran `main` with a self-installed Windows ROCm torch and an explicit `cuda:N` in HIP's full numbering gets renumbered once: an index past the end fails at startup with a clear error, and an iGPU index is refused, but with two discrete GPUs behind an iGPU, `cuda:1` would silently select the other card. - **`pins.json`:** `linux.rocm` points at AMD's index. It is read only by launcher installs of releases before 6.14, so legacy Windows installs still get no ROCm entry. - **GPU support on Linux:** gfx900/gfx906 (Radeon VII, Vega 56/64, MI25/MI50/MI60) are no longer covered by the `rocm` extra. The docs point to the previous release or `--torch-backend=rocm7.2`. ## Checklist - [x] _The PR has a short but descriptive title, suitable for a changelog_ - [x] _Meaningful regression coverage added / updated where needed; obsolete tests/code removed_ - [x] _Persisted-state and API changes include required migrations / compatibility validation_ - [x] _Relevant performance/efficiency opportunities considered; material claims have evidence_ - [x] _Material review findings resolved and relevant checks rerun_ - [x] _Documentation added / updated (if applicable)_ - [ ] _Updated `What's New` copy (if doing a release after this PR)_
Summary
Feature:
invokeai[rocm]now installs on Windows. Therocmextra previously resolved only on Linux, so AMD GPUs on Windows fell back to the CPU.On Windows the extra takes torch 2.13 and torchvision 0.28 from AMD's stable ROCm 10 index (
https://stable.repo.amd.com/rocm/whl-next/), with GPU kernels for the RDNA3, RDNA3.5 and RDNA4 targets AMD supports on Windows. That index also mirrors common PyPI packages, so it is explicit, and the packages that exist only there are listed in the extra so their index source applies (uv applies sources to direct requirements only). AMD's Linux torch wheels require a Linux-only triton build that the Windows wheel does not use, so adependency-metadataentry restates the Windows wheel's dependencies. The base torch requirement is split into linux and win32 lines so uv forks them, since the rocm extra now uses a different index on each.bitsandbytesis floored at 0.50.2 on Windows, the first release that ships a ROCm binary there. Because uv resolves a single Windows fork, this also moves the Windows cpu and cuda installs from 0.49.2 to 0.50.2. Linux and macOS resolutions are unchanged for every extra and Python version.TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTALnow defaults to 1 at startup. Without it, torch offers no fused SDPA kernel on GPUs AOTriton marks experimental, RDNA4 among them (on Windows and Linux), and falls back to the math backend. On an RX 9060 XT a 4096-token attention call takes 2.2 ms with flash and 36.9 ms on the math backend. An explicitly set value is respected.Docs cover the Windows requirements and a manual
uv pipinstall.Related Issues / Discussions
Related: #9530, #9321
The launcher still routes Windows AMD GPUs to the CPU backend (
hasWindowsAmdGpuingpu-detection.ts), so a follow-up there is needed before launcher installs pick this up.QA Instructions
Tested on Windows 11 with an RX 9060 XT (gfx1200):
uv sync --frozen --extra rocminstallstorch 2.13.0+rocm10.0.0, andtorch.cuda.is_available()is true.Linearlayers produce correct output on the GPU.pytest tests/test_check_pins.py tests/backend/pidpasses, including the CUDA-marked PiD tests.uv pip installcommand from the docs installs the released package with ROCm torch.uv exportper extra (none, cpu, cuda, rocm, xpu, each with and without xformers) againstmainon linux x86_64/aarch64, macOS arm64/x86_64 and Windows for Python 3.11 and 3.12: only the Windows changes described above differ.Merge Plan
uv.lockis regenerated; if another dependency change lands first, rebase and rerunuv lock.Checklist
What's Newcopy (if doing a release after this PR)