Skip to content

[enhancement]: Replacement ROCM 7 to ROCM 10 #9692

Description

@baramofme

Is there an existing issue for this?

  • I have searched the existing issues

Contact Details

No response

What should this feature add?

Move the rocm extra from PyTorch's rocm7.2 index to AMD's ROCm 10 wheels (torch==2.13.0+rocm10.0.0 and friends), i.e. land #9643 — and additionally build and publish the rocm variant of ghcr.io/invoke-ai/invokeai against ROCm 10, because that is the exact step #9643 lists as "Not tested: The Docker ROCm image build" and "Linux with an AMD GPU".

We run ROCm 10 in production on Linux / RX 7900 XT (gfx1100) with a self-built image and measured large gains (numbers in Additional Content). We would like this to become the supported, officially published path instead of a per-user custom build.

Alternatives

Additional Content

Preparation checklist

Wheels exist today. https://stable.repo.amd.com/rocm/whl-next/ serves

torch-2.13.0+rocm10.0.0 (cp312, manylinux x86_64), matching torchvision / triton,
plus per-architecture device packages (gfx101x–gfx120x). There is no official
download.pytorch.org/whl/rocm10.* index yet, so AMD's index must be the source and must
stay explicit = true (it mirrors PyPI packages too).

Repo-side changes (mirrors what #9643 does):

  1. pins.json — linux.rocm points at the AMD index.

  2. pyproject.toml rocm extra — torch==2.13.0+rocm10.0.0,
    torchvision==0.28.0+rocm10.0.0, updated triton-rocm; update the
    [[tool.uv.index]] torch-rocm URL; regenerate uv lock with uv 0.11.28.

  3. scripts/check_pins.py and scripts/check_aarch64_lock.py must pass —
    uv-lock-checks.yml enforces pins.json ↔ pyproject sync, and the 6.3 → 7.1 drift
    ([bug]: Release 6.13.5rc1 not installing ROCm 7.1 as expected #9328) is exactly what that gate exists for.

  4. Budget size: a full --extra rocm venv is ~20 GB (was ~17 GB at rocm7.2); the
    build-container.yml rocm job has a 90-minute timeout that needs headroom.

Docker side — the Dockerfile needs no structural change. library/ubuntu:24.04 base

  • ARG GPU_DRIVER=rocm + uv sync --extra $GPU_DRIVER --frozen already works: the ROCm
    userspace ships inside AMD's torch wheel. Keep the existing ROCm layers (amdgpu.ids
    symlink, groupadd render).

Host/runtime prerequisites:

  • /dev/kfd and /dev/dri passed into the container; the container user in the video
    (GID 44) and render groups (our host has render at GID 992 — the entrypoint's
    RENDER_GROUP_ID handling covers this); seccomp=unconfined as in docker/README.md.

  • A shared MIOpen cache volume so gfx kernels are reused across restarts.

  • Env deltas: on gfx1100, ROCm 10 supports the ISA natively — HSA_OVERRIDE_GFX_VERSION
    is not required; PYTORCH_ROCM_ARCH=gfx1100 still applies for source builds;
    run_app.py already defaults MIOPEN_FIND_MODE=FAST.

Application process (steps we ran, Linux / RX 7900 XT / gfx1100)

  1. docker build --build-arg GPU_DRIVER=rocm . from the ROCm 10 branch (we used the head
    commit of feat(deps): ROCm 10 on Linux and Windows from AMD's index #9643 plus the [bug]: First load of a model is extremely slow on ROCm — copies to VRAM from mmap-backed safetensors run at ~2 MB/s #9487 clone fix ported into tensor_aliases.py).

  2. Push to our registry, deploy with the existing invokeai-rocm compose shape
    (runtime: amd, profiles: [rocm]) adjusted to the host's group IDs.

  3. Startup verification inside the container:

    python -c "import torch; print(torch.version, torch.version.hip, torch.cuda.get_device_name(0))"

    → 2.13.0+rocm10.0.0, HIP 7.15.x (this is the HIP runtime version, not the ROCm
    release — beware any gate keyed on this string), device gfx1100 reported as (11, 0).

  4. Generate via the queue API. First-load behaviour with the [bug]: First load of a model is extremely slow on ROCm — copies to VRAM from mmap-backed safetensors run at ~2 MB/s #9487 fix: the 7.67 GB
    mmap-backed CLIP/VAE encoder reached VRAM in 5.64 s (~1.36 GB/s) — no longer the
    ~2 MB/s pathological path described in [bug]: First load of a model is extremely slow on ROCm — copies to VRAM from mmap-backed safetensors run at ~2 MB/s #9487.

Measured evidence (Krea-2 Turbo fp8, 1024², 8 steps)

stage rocm7.2 (main-rocm) rocm 10 (this build)
cold first model load, encoder 18.68 s 3.50 s
cold first denoise pass 72.43 s 20.17 s
warm denoise — 16.06 s
end-to-end per image (warm) — ~31–40 s

Known caveats we hit (worth deciding on before merge)

  • rocm_aotriton_experimental: auto keeps fused attention off on gfx1100 — fused correctness under ROCm 10 is measured on gfx1200 only (the same gap feat(deps): ROCm 10 on Linux and Windows from AMD's index #9643 itself flags; RDNA3 under ROCm 10 is unmeasured). Everything we ran used the math kernel and produced correct images; a follow-up measurement on a W7900 would let the wide-head math guard and the Wan conv3d decomposition be retired.

  • ROCm 10 ships no kernels for gfx900/gfx906 → Radeon VII / Vega 56/64 fall out of the rocm extra; docs should point them at --torch-backend=rocm7.2 or a previous release.

  • There is no ROCm hardware CI (python-tests.yml matrix = CPU/macos/windows), so the numerical/device behaviour of this bump rests on community hardware.

  • bitsandbytes has no ROCm-specific wheel (PyPI 0.49.2 → 0.50.2 as resolved).

We are happy to provide the compose/env recipe and the raw timings as a starting point for the ROCm 10 image build job in build-container.yml.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions