Repository navigation
feat(deps): ROCm 10 on Linux and Windows from AMD's index - #9643
Pfannkuchensack wants to merge 12 commits into
Conversation
The rocm extra moves from PyTorch's 2.13.0+rocm7.2 (Linux only) to AMD's 2.13.0+rocm10.0.0 for Linux x86_64 and Windows. The lock carries the GPU kernel packages of every target AMD builds, so the launcher can leave out the ones a machine does not need. Linux loses gfx900/gfx906, which ROCm 10 dropped. bitsandbytes moves to 0.50.2, the first release with a Windows ROCm binary. A new setting, rocm_aotriton_experimental (auto/on/off), sets TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL before anything runs attention. auto turns AOTriton's experimental fused kernels on only for a ROCm 10 build on GPUs where they were measured correct (gfx1200 so far); under torch 2.12+rocm7.14 the same switch made every fused call fail on that GPU. Builds on the dependency and documentation changes of invoke-ai#9606. Co-authored-by: 0xDELUXA <djernovevo@gmail.com>
…n CUDA/ROCm ROCm enumerates the Radeon graphics of a Ryzen CPU next to a discrete card (measured: gfx1036 beside an RX 9060 XT on Windows), so `auto` handed the iGPU half of the queue. Integrated GPUs were only filtered on XPU. device_is_integrated now also answers for CUDA/HIP devices from their properties, and `device: auto` no longer defaults to an iGPU that is enumerated before the discrete card.
…ializes it While HIP on Windows can see the Ryzen iGPU, a process's first context on the discrete card shows 17.9 GiB of shared GPU memory and its first kernel launch another 17.9 GiB, with every ROCm build tried and without any Invoke code. It is neither RAM nor commit charge, but Task Manager shows 36 GiB of shared GPU memory, and Invoke's paging check read it as memory pushed out of VRAM: its warning fired after every generation. A short child process now asks HIP which devices are integrated (reading device properties creates no context), and HIP_VISIBLE_DEVICES hides them. Only on Windows ROCm, only when no visibility variable is set, and only when device and generation_devices are both auto, so explicit cuda:N indices keep their meaning.
lstein
left a comment
There was a problem hiding this comment.
Thanks for this — it's a big step for AMD users, and the lock work (the two dependency-metadata overrides, the per-target device packages) is careful. I tested it on Linux with 2× Radeon PRO W7900 (gfx1100), venv synced from this lock with the non-gfx1100 device packages left out via --no-install-package. That install path works, and the gfx1100 rocBLAS/hipBLASLt kernels are all present.
One blocker for Linux ROCm, plus a few smaller issues.
Blocking
1. bitsandbytes is broken on Linux ROCm
pyproject.toml:171 pins bitsandbytes>=0.50.2 for win32 only, so the Linux rocm extra stays on 0.49.2. AMD's ROCm 10 torch reports torch.version.hip == '7.15.26333' (same on both wheels). bitsandbytes 0.49.2 builds the library name as major * 10 + minor, so 7.15 becomes 85 and it looks for libbitsandbytes_rocm85.so, which exists in no release. It then falls back to its stand-in library, and every native call raises.
Reproduced on the W7900. At startup:
bitsandbytes library load error: Configured ROCm binary not found at .../bitsandbytes/libbitsandbytes_rocm85.so
and a FLUX.1 NF4 transformer + t5_bnb_int8_quantized_encoder generation fails:
Detected PyTorch ROCm version: 7.15
Available pre-compiled versions: 6.2 6.3 6.4 7.0 7.1 7.2
...
Native code method attempted to call: lib.cint8_vector_quant()
On main this works, because torch 2.13.0+rocm7.2 reports HIP 7.2 and rocm72.so ships.
0.50.2 joins the digits (f"{major}{minor}" → 715) and ships libbitsandbytes_rocm715.so for Linux as well as Windows. With bitsandbytes==0.50.2 installed in the same venv, the same FLUX NF4 + int8-T5 generation completes and the image is fine.
Fix: widen the marker to sys_platform == 'win32' or (sys_platform == 'linux' and platform_machine == 'x86_64'). A Linux install on a machine without a GPU can't catch this, because torch.cuda is unavailable there and bitsandbytes never picks a GPU library.
Should fix
2. 0.50.2 floods the log with the int8 cast warning
0.50.2 moved MatMul8bitLt: inputs will be cast from torch.bfloat16 to float16 during quantization from warnings.warn to logger.warning on the bitsandbytes.autograd._functions logger. The existing warnings.filterwarnings in invokeai/backend/quantization/bnb_llm_int8.py:14 (and the Qwen-Image encoder's catch_warnings block) no longer catches it, so it prints on every int8 matmul. I saw a wall of these during the generation above.
As the PR body notes, the lock also moves cpu, cuda and the no-extra install to 0.50.2 on Windows and Linux x86_64, so CUDA users of the int8 T5 get the same flood. A logging.Filter on that logger next to the existing warnings filter would cover both versions:
class _SuppressInt8CastWarning(logging.Filter):
def filter(self, record: logging.LogRecord) -> bool:
return not record.getMessage().startswith("MatMul8bitLt: inputs will be cast")
logging.getLogger("bitsandbytes.autograd._functions").addFilter(_SuppressInt8CastWarning())3. Hiding the iGPU only while the config is auto lets device indices shift
hide_integrated_gpus_on_rocm_windows only sets HIP_VISIBLE_DEVICES while device and generation_devices are both auto (rocm_integrated_gpu.py:87). But the settings UI lists, validates and saves devices in the hidden numbering: get_generation_device_options (app_info.py:221) enumerates torch.cuda after hiding, and update_runtime_config (app_info.py:261) validates against it, then writes invokeai.yaml.
Example: HIP order is [discrete A, iGPU, discrete B]. With the iGPU hidden, the UI offers cuda:0 = A and cuda:1 = B. The user selects both and saves. On restart generation_devices is no longer auto, so nothing is hidden, and cuda:1 is now the iGPU. Per the PR description, any kernel there crashes the process with an access violation. With the iGPU enumerated first, picking [cuda:0] lands on it. With the iGPU last, nothing crashes, but the 36 GiB phantom shared memory and the paging warning come back.
Options: hide the iGPU whenever it sits next to a discrete GPU, regardless of device/generation_devices (it can't run kernels on Windows ROCm anyway, per the PR's own measurement), or keep the UI and the saved values in the full HIP numbering. I couldn't test this one; the W7900 box has no iGPU.
4. The probe can stall startup for 2 minutes
_PROBE_TIMEOUT_SECONDS = 120 (rocm_integrated_gpu.py:33). The PR itself notes a HIP process hanging on exit under Windows (the CLIP case under 7.14.1). If the child prints its answer and then hangs in teardown, startup waits the full 120 s and falls back with only a debug log (:96). Suggest ending the probe with sys.stdout.flush(); os._exit(0), a much shorter timeout, and a warning on the fallback.
Minor
torch.version.rocmexists on these wheels ('10.0.0') and would be a cleaner source forrocm_major_versionthan parsing the version label. It also covers locally built torch.- AMD's artifacts carry no hashes in the lock, including the
rocmsdist, which runs build code at install time. You've already called this out; noting it as an accepted risk.
W7900 results (Linux, gfx1100)
-
Correctness: fp32/fp16/bf16 matmul on both cards match the CPU reference to the same error as rocm7.2.
-
SDXL, warm:
Build Denoise VAE decode rocm7.2 (default, rocBLAS) 8.91 / 9.12 s 0.51 s ROCm 10, TORCH_BLAS_PREFER_HIPBLASLT=0(rocBLAS)8.93 s 0.54 s ROCm 10, default (hipBLASLt) 7.84 s 0.77 s ROCm 10 is as fast as or faster than rocm7.2 here. AMD's build defaults to hipBLASLt, where PyTorch's rocm7.2 build used rocBLAS. A plain
a @ bbenchmark makes hipBLASLt look ~30% slower at 8192³ on gfx1100, but end to end it's ~12% faster for SDXL denoising, so the default is the right choice. The VAE decode is ~0.2 s slower with hipBLASLt (single run). -
First run: the first cold SDXL run on ROCm 10 had ~4 s extra in denoise. That's probably MIOpen's per-version kernel cache starting empty, so it's likely a one-time cost after install.
-
AOTriton
auto: leaves the experimental kernels off on gfx1100, as intended, and the startup log says so clearly. -
I did not run the
-m slowhardware canaries (wide-head guard, Wan conv3d decomposition) under ROCm 10, so the RDNA3 measurement for the follow-up is still open.
|
I run ROCm 10 in production on Linux / RX 7900 XT (gfx1100) with a self-built image and measured large gains (numbers in Additional Content). Measured evidence (Krea-2 Turbo fp8, 1024², 8 steps)
Application process (steps we ran, Linux / RX 7900 XT / gfx1100)
|
# Conflicts: # invokeai/frontend/api/openapi.json # tests/backend/util/test_devices.py
AMD's ROCm 10 torch reports HIP 7.15; bitsandbytes 0.49 then looks for libbitsandbytes_rocm85.so, which no release ships, and every native call fails. 0.50.2 names the library rocm715 and ships it for Linux as well as Windows.
…s on every matmul bitsandbytes 0.50 reports the bf16-to-fp16 cast through its logger instead of warnings, past the existing filter, once per int8 matmul. One silence_int8_cast_notice() now filters both forms, for the FLUX T5 int8 path and the Qwen-Image int8 encoder.
…d bound the probe The settings UI numbers devices after the iGPU is hidden, so hiding it only while device and generation_devices were auto renumbered saved choices on the next start. Naming an integrated GPU beside a discrete one under Windows ROCm is now refused with a reason instead of crashing at the first kernel. The probe takes the answer as soon as the child prints it, gives up after 60 s, and warns with the reason when it cannot answer.
The +rocmN label of the torch version stays the fallback for builds that leave torch.version.rocm unset.
|
@lstein Thanks for testing on the W7900. All points are addressed in a493101..513e04a:
|
| logger.info( | ||
| f"ROCm ({gpus}): AOTriton's experimental fused attention kernels are off ({reason}). On a GPU that " | ||
| "AOTriton marks experimental, attention then runs on the slower math kernel; set " | ||
| "rocm_aotriton_experimental: on to try the fused kernels." |
There was a problem hiding this comment.
P2 Quote the YAML value in the opt-in instruction
A user following this startup message and adding rocm_aotriton_experimental: on to invokeai.yaml cannot restart InvokeAI. load_and_migrate_config() uses yaml.safe_load(), which converts unquoted on and off to booleans, while the new setting accepts only the strings auto, on, and off. Reproducing the parse and validation produces a literal_error for True, which the config loader turns into a startup failure. Recommend rocm_aotriton_experimental: "on" (and quote "off" in corresponding guidance), or normalize YAML booleans before validating this field.
There was a problem hiding this comment.
Good catch, thanks. Fixed in 6a0048a. Instead of only quoting the hint, the field now accepts YAML's booleans: an unquoted on/off (read as True/False by yaml.safe_load) maps to "on"/"off" before validation. Both the log line's advice and the natural way of writing the setting work. tests/test_config.py loads real YAML files with unquoted on, off, quoted "on" and auto through load_and_migrate_config.
YAML 1.1 reads them as booleans, so `rocm_aotriton_experimental: on`, as the startup log suggests, failed validation and stopped InvokeAI from starting.
…ing bitsandbytes The Qwen-Image int8 encoder reached the filter through bnb_llm_int8, which imports bitsandbytes; macOS has none, so its int8 load tests failed there. The filter now lives in bnb_cast_notice, which needs only logging and warnings, and its test runs on every platform.
Summary
The
rocmextra moves from PyTorch's 2.13.0+rocm7.2, which exists for Linux only, to AMD's 2.13.0+rocm10.0.0 for Linux x86_64 and Windows. Until now,uv sync --extra rocmon Windows installed PyPI's CPU torch. It now installs ROCm, and the Windows ROCm path becomes a supported install.This carries over the dependency and documentation changes of #9606, with @0xDELUXA as co-author. Three things changed on the way.
Fused attention: a switch instead of a fixed
TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1.hipErrorInvalidValue. Under 2.13.0+rocm10.0.0 every fused case matched the math reference.rocm_aotriton_experimental(auto/on/off), decides. It is applied before anything runs attention.autoturns the kernels on only for a ROCm 10 build, and only when every generation GPU is one they were measured on: so far gfx1200.onandoffwork for any build and GPU, and an exported variable always wins. The startup log names the effective value.Integrated GPUs. HIP enumerates the Radeon graphics of a Ryzen CPU next to the discrete card. On the test machine (RX 9060 XT plus the gfx1036 iGPU of a Ryzen 7900, Windows 10), the iGPU caused three problems:
generation_devices: autostarted a second worker on it.torch.randn, crashed the process with an access violation.The fix has two parts:
generation_devices: autoanddevice: autoleave integrated GPUs out on CUDA/ROCm too. Until now that filter existed for XPU only.HIP_VISIBLE_DEVICEShides them before torch initializes HIP. This happens whenever no visibility variable is set, whateverdeviceandgeneration_devicessay, so the settings UI, the saved config and the runtime all countcuda:Namong the discrete GPUs. Naming an integrated GPU next to a discrete one is refused with a clear error, since every kernel on it crashed the process. The parent takes the child's answer as soon as it is printed and ends the child; after 60 s without an answer it logs a warning with the reason and leaves every GPU visible.Device packages. AMD ships the GPU kernels as one package per target and library. The lock carries all of them for both platforms (gfx101x–gfx120x, gfx908/90a; on Linux also gfx942/950/1250). The launcher can then leave out the targets a machine does not have with
uv sync --no-install-package. ROCm 10 has no kernels for gfx900/gfx906, so Linux loses Vega 10/20 cards; the docs name the way out.RX 9060 XT (gfx1200), Windows 10, venv synced from this lock, no
HIP_VISIBLE_DEVICESset. Warm runs, 1024², the rest from the first run:off)auto)Related Issues / Discussions
rocmextra); its changes are carried over in the first commit, co-authored by @0xDELUXA.rocmextra, and leave out device packages the machine does not need.QA Instructions
Lock and pins
uv lock --lockedwith uv 0.11.28,scripts/check_pins.pyandscripts/check_aarch64_lock.pypass.uv exportper extra againstmain:rocmmove, plus the 78 device packages,rocm,rocm-sdk-core,rocm-sdk-librariesandrocm-bootstrap;rocmon Windows and Linux x86_64: AMD's torch reports HIP 7.15, for which 0.49 looks for arocm85library that does not exist. uv's fork preference also movescpu,cudaand the no-extra resolution to 0.50.2 on Windows and Linux x86_64. 0.50 logs the int8 bf16→fp16 cast notice on every matmul;silence_int8_cast_notice()filters it for the FLUX T5 and Qwen-Image int8 paths;xpuis unchanged.rocmis an sdist whose build asks the machine's GPU which kernel package to require.Installs from the lock
uv sync --locked --extra rocm --extra testinstalls. torch 2.13.0+rocm10.0.0 sees both GPUs and reports 22 arch targets.uv sync --locked --extra rocminstalls.import torch, tritonworks, andget_arch_list()lists all 25 targets.Tests
tests/app/util/test_rocm_aotriton.py,tests/app/util/test_rocm_integrated_gpu.py, and CUDA/ROCm iGPU cases intests/backend/util/test_devices.py.tests/backend/util,tests/app/util,tests/backend/model_manager/load/model_cache,tests/app/services/config,tests/app/services/session_processor,tests/test_check_pins.py;tests/backend/util,tests/app/util,tests/app/routers/test_app_info.py.test_get_generation_devices_auto_expands_to_all_cudaread the real iGPU and is now hardware-independent.tests/backend/model_manager/load/model_cache/test_shared_weights_gpu.pycrashes on the ROCm venv when the iGPU is visible: it computes oncuda:1, which is the iGPU there. This is the crash described above, not a regression.-m slowROCm hardware tests, with and without the AOTriton variable: 6 pass, and the two canaries fail as they are meant to on this GPU:ruff check/ruff format --checkare clean.openapi.json,schema.tsanddocs/src/generated/settings.jsonare regenerated.E2E on the RX 9060 XT (table above):
auto, PSNR against the rocm7.14.1 math-kernel reference is 42.2/32.1 dB, the same as in the earlier measurement;offgives 57.8 dB.Not tested:
Review
Remaining risks and limitations:
uv lock --lockednow needs it reachable.uv sync --extra rocminstalls the kernels of every target, about 4.3 GiB extra download on Windows and 8.3 GiB on Linux.autocovers gfx1200 only. gfx1101, gfx1102, gfx1150 and gfx1151 need the variable too (per feat(deps): support ROCm on Windows through the rocm extra #9606), but have not been measured here; on them,onis an opt-in.Compatibility / Rollout
rocm_aotriton_experimental(auto|on|off, defaultauto). It is additive, needs no migration, and the API contracts are regenerated.HIP_VISIBLE_DEVICESon Windows ROCm is now set at startup when an iGPU sits next to a discrete GPU, unless the user set it,CUDA_VISIBLE_DEVICESorROCR_VISIBLE_DEVICES. Explicitcuda:Nvalues then count the discrete GPUs only. Anyone who ranmainwith a self-installed Windows ROCm torch and an explicitcuda:Nin HIP's full numbering gets renumbered once: an index past the end fails at startup with a clear error, and an iGPU index is refused, but with two discrete GPUs behind an iGPU,cuda:1would silently select the other card.pins.json:linux.rocmpoints at AMD's index. It is read only by launcher installs of releases before 6.14, so legacy Windows installs still get no ROCm entry.rocmextra. The docs point to the previous release or--torch-backend=rocm7.2.Checklist
What's Newcopy (if doing a release after this PR)