You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[enhancement]: Replacement ROCM 7 to ROCM 10 #9692
Move the rocm extra from PyTorch's rocm7.2 index to AMD's ROCm 10 wheels (torch==2.13.0+rocm10.0.0 and friends), i.e. land #9643 — and additionally build and publish the rocm variant of ghcr.io/invoke-ai/invokeai against ROCm 10, because that is the exact step #9643 lists as "Not tested: The Docker ROCm image build" and "Linux with an AMD GPU".
We run ROCm 10 in production on Linux / RX 7900 XT (gfx1100) with a self-built image and measured large gains (numbers in Additional Content). We would like this to become the supported, officially published path instead of a per-user custom build.
Per-user custom images (what we do today). Works — we built and deployed one — but it is outside the supported path: no official main-rocm tag on ROCm 10, no CI coverage, and every user has to re-derive the recipe.
torch-2.13.0+rocm10.0.0 (cp312, manylinux x86_64), matching torchvision / triton, plus per-architecture device packages (gfx101x–gfx120x). There is no official download.pytorch.org/whl/rocm10.* index yet, so AMD's index must be the source and must stay explicit = true (it mirrors PyPI packages too).
/dev/kfd and /dev/dri passed into the container; the container user in the video (GID 44) and render groups (our host has render at GID 992 — the entrypoint's RENDER_GROUP_ID handling covers this); seccomp=unconfined as in docker/README.md.
A shared MIOpen cache volume so gfx kernels are reused across restarts.
Env deltas: on gfx1100, ROCm 10 supports the ISA natively — HSA_OVERRIDE_GFX_VERSION is not required; PYTORCH_ROCM_ARCH=gfx1100 still applies for source builds; run_app.py already defaults MIOPEN_FIND_MODE=FAST.
Application process (steps we ran, Linux / RX 7900 XT / gfx1100)
→ 2.13.0+rocm10.0.0, HIP 7.15.x (this is the HIP runtime version, not the ROCm release — beware any gate keyed on this string), device gfx1100 reported as (11, 0).
Known caveats we hit (worth deciding on before merge)
rocm_aotriton_experimental: auto keeps fused attention off on gfx1100 — fusedcorrectness under ROCm 10 is measured on gfx1200 only (the same gap feat(deps): ROCm 10 on Linux and Windows from AMD's index #9643 itself flags;RDNA3 under ROCm 10 is unmeasured). Everything we ran used the math kernel and producedcorrect images; a follow-up measurement on a W7900 would let the wide-head math guardand the Wan conv3d decomposition be retired.
ROCm 10 ships no kernels for gfx900/gfx906 → Radeon VII / Vega 56/64 fall out of therocm extra; docs should point them at --torch-backend=rocm7.2 or a previous release.
There is no ROCm hardware CI (python-tests.yml matrix = CPU/macos/windows), so thenumerical/device behaviour of this bump rests on community hardware.
bitsandbytes has no ROCm-specific wheel (PyPI 0.49.2 → 0.50.2 as resolved).
We are happy to provide the compose/env recipe and the raw timings as a starting point forthe ROCm 10 image build job in build-container.yml.
Is there an existing issue for this?
Contact Details
No response
What should this feature add?
Move the
rocmextra from PyTorch'srocm7.2index to AMD's ROCm 10 wheels (torch==2.13.0+rocm10.0.0and friends), i.e. land #9643 — and additionally build and publish therocmvariant ofghcr.io/invoke-ai/invokeaiagainst ROCm 10, because that is the exact step #9643 lists as "Not tested: The Docker ROCm image build" and "Linux with an AMD GPU".We run ROCm 10 in production on Linux / RX 7900 XT (gfx1100) with a self-built image and measured large gains (numbers in Additional Content). We would like this to become the supported, officially published path instead of a per-user custom build.
Alternatives
rocm7.2(status quo). Cost: RDNA3 users run the math attention kernel unless they export an experimental env var by hand; the pathological first-load copies from mmap-backed safetensors ([bug]: First load of a model is extremely slow on ROCm — copies to VRAM from mmap-backed safetensors run at ~2 MB/s #9487) remain; no fused-attention speed-up that feat(deps): ROCm 10 on Linux and Windows from AMD's index #9643 measured (~35 s to 17 s, ~60 s to 32 s per image on gfx1200).main-rocmtag on ROCm 10, no CI coverage, and every user has to re-derive the recipe.Additional Content
Preparation checklist
Wheels exist today.
https://stable.repo.amd.com/rocm/whl-next/servestorch-2.13.0+rocm10.0.0(cp312, manylinux x86_64), matchingtorchvision/triton,plus per-architecture device packages (gfx101x–gfx120x). There is no official
download.pytorch.org/whl/rocm10.*index yet, so AMD's index must be the source and muststay
explicit = true(it mirrors PyPI packages too).Repo-side changes (mirrors what #9643 does):
pins.json—linux.rocmpoints at the AMD index.pyproject.tomlrocmextra —torch==2.13.0+rocm10.0.0,torchvision==0.28.0+rocm10.0.0, updatedtriton-rocm; update the[[tool.uv.index]] torch-rocmURL; regenerateuv lockwith uv 0.11.28.scripts/check_pins.pyandscripts/check_aarch64_lock.pymust pass —uv-lock-checks.ymlenforces pins.json ↔ pyproject sync, and the 6.3 → 7.1 drift([bug]: Release 6.13.5rc1 not installing ROCm 7.1 as expected #9328) is exactly what that gate exists for.
Budget size: a full
--extra rocmvenv is ~20 GB (was ~17 GB at rocm7.2); thebuild-container.ymlrocm job has a 90-minute timeout that needs headroom.Docker side — the Dockerfile needs no structural change.
library/ubuntu:24.04baseARG GPU_DRIVER=rocm+uv sync --extra $GPU_DRIVER --frozenalready works: the ROCmuserspace ships inside AMD's torch wheel. Keep the existing ROCm layers (
amdgpu.idssymlink,
groupadd render).Host/runtime prerequisites:
/dev/kfdand/dev/dripassed into the container; the container user in thevideo(GID 44) and
rendergroups (our host hasrenderat GID 992 — the entrypoint'sRENDER_GROUP_IDhandling covers this);seccomp=unconfinedas indocker/README.md.A shared MIOpen cache volume so gfx kernels are reused across restarts.
Env deltas: on gfx1100, ROCm 10 supports the ISA natively —
HSA_OVERRIDE_GFX_VERSIONis not required;
PYTORCH_ROCM_ARCH=gfx1100still applies for source builds;run_app.pyalready defaultsMIOPEN_FIND_MODE=FAST.Application process (steps we ran, Linux / RX 7900 XT / gfx1100)
docker build --build-arg GPU_DRIVER=rocm .from the ROCm 10 branch (we used the headcommit of feat(deps): ROCm 10 on Linux and Windows from AMD's index #9643 plus the [bug]: First load of a model is extremely slow on ROCm — copies to VRAM from mmap-backed safetensors run at ~2 MB/s #9487 clone fix ported into
tensor_aliases.py).Push to our registry, deploy with the existing
invokeai-rocmcompose shape(
runtime: amd,profiles: [rocm]) adjusted to the host's group IDs.Startup verification inside the container:
python -c "import torch; print(torch.version, torch.version.hip, torch.cuda.get_device_name(0))"→
2.13.0+rocm10.0.0, HIP 7.15.x (this is the HIP runtime version, not the ROCmrelease — beware any gate keyed on this string), device
gfx1100reported as(11, 0).Generate via the queue API. First-load behaviour with the [bug]: First load of a model is extremely slow on ROCm — copies to VRAM from mmap-backed safetensors run at ~2 MB/s #9487 fix: the 7.67 GB
mmap-backed CLIP/VAE encoder reached VRAM in 5.64 s (~1.36 GB/s) — no longer the
~2 MB/s pathological path described in [bug]: First load of a model is extremely slow on ROCm — copies to VRAM from mmap-backed safetensors run at ~2 MB/s #9487.
Measured evidence (Krea-2 Turbo fp8, 1024², 8 steps)
Known caveats we hit (worth deciding on before merge)
rocm_aotriton_experimental: autokeeps fused attention off on gfx1100 — fused correctness under ROCm 10 is measured on gfx1200 only (the same gap feat(deps): ROCm 10 on Linux and Windows from AMD's index #9643 itself flags; RDNA3 under ROCm 10 is unmeasured). Everything we ran used the math kernel and produced correct images; a follow-up measurement on a W7900 would let the wide-head math guard and the Wan conv3d decomposition be retired.ROCm 10 ships no kernels for gfx900/gfx906 → Radeon VII / Vega 56/64 fall out of the
rocmextra; docs should point them at--torch-backend=rocm7.2or a previous release.There is no ROCm hardware CI (
python-tests.ymlmatrix = CPU/macos/windows), so the numerical/device behaviour of this bump rests on community hardware.bitsandbytes has no ROCm-specific wheel (PyPI 0.49.2 → 0.50.2 as resolved).
We are happy to provide the compose/env recipe and the raw timings as a starting point for the ROCm 10 image build job in
build-container.yml.