Skip to content

feat(deps): ROCm 10 on Linux and Windows from AMD's index - #9643

Open
Pfannkuchensack wants to merge 12 commits into
invoke-ai:mainfrom
Pfannkuchensack:feat/rocm10
Open

Pfannkuchensack wants to merge 12 commits into
invoke-ai:mainfrom
Pfannkuchensack:feat/rocm10

Conversation

@Pfannkuchensack

@Pfannkuchensack Pfannkuchensack commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

Summary

The rocm extra moves from PyTorch's 2.13.0+rocm7.2, which exists for Linux only, to AMD's 2.13.0+rocm10.0.0 for Linux x86_64 and Windows. Until now, uv sync --extra rocm on Windows installed PyPI's CPU torch. It now installs ROCm, and the Windows ROCm path becomes a supported install.

This carries over the dependency and documentation changes of #9606, with @0xDELUXA as co-author. Three things changed on the way.

Fused attention: a switch instead of a fixed TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1.

  • On several RDNA3/3.5/4 GPUs, PyTorch runs its fused (AOTriton) attention only with that variable set. Otherwise every attention call falls back to the math kernel.
  • Setting it unconditionally is not safe. Under torch 2.12+rocm7.14.1 it made every flash and memory-efficient call on an RX 9060 XT fail with hipErrorInvalidValue. Under 2.13.0+rocm10.0.0 every fused case matched the math reference.
  • A new setting, rocm_aotriton_experimental (auto/on/off), decides. It is applied before anything runs attention. auto turns the kernels on only for a ROCm 10 build, and only when every generation GPU is one they were measured on: so far gfx1200. on and off work for any build and GPU, and an exported variable always wins. The startup log names the effective value.

Integrated GPUs. HIP enumerates the Radeon graphics of a Ryzen CPU next to the discrete card. On the test machine (RX 9060 XT plus the gfx1036 iGPU of a Ryzen 7900, Windows 10), the iGPU caused three problems:

  • generation_devices: auto started a second worker on it.
  • Any kernel launched on it, even torch.randn, crashed the process with an access violation.
  • While HIP can merely see it, a process's first context on the discrete card shows 17.9 GiB of shared GPU memory, and its first kernel launch another 17.9 GiB. This happens with every ROCm build tried and without any Invoke code.
    • It is neither RAM nor commit charge: commit grew by 1.7 GiB, free RAM fell by 0.8 GiB.
    • But Task Manager shows 36 GiB of shared GPU memory, and the paging check from #304 reads the same counter. It warned after every generation, and the residency correction of the pre-tiling decision dropped out.

The fix has two parts:

  • generation_devices: auto and device: auto leave integrated GPUs out on CUDA/ROCm too. Until now that filter existed for XPU only.
  • On Windows ROCm, a short child process (1.4–2.4 s at startup) asks HIP which devices are integrated, and HIP_VISIBLE_DEVICES hides them before torch initializes HIP. This happens whenever no visibility variable is set, whatever device and generation_devices say, so the settings UI, the saved config and the runtime all count cuda:N among the discrete GPUs. Naming an integrated GPU next to a discrete one is refused with a clear error, since every kernel on it crashed the process. The parent takes the child's answer as soon as it is printed and ends the child; after 60 s without an answer it logs a warning with the reason and leaves every GPU visible.

Device packages. AMD ships the GPU kernels as one package per target and library. The lock carries all of them for both platforms (gfx101x–gfx120x, gfx908/90a; on Linux also gfx942/950/1250). The launcher can then leave out the targets a machine does not have with uv sync --no-install-package. ROCm 10 has no kernels for gfx900/gfx906, so Linux loses Vega 10/20 cards; the docs name the way out.

RX 9060 XT (gfx1200), Windows 10, venv synced from this lock, no HIP_VISIBLE_DEVICES set. Warm runs, 1024², the rest from the first run:

before after
Z-Image Turbo nvfp4, 8 steps 35.7 / 35.8 s (math kernel, off) 16.7 / 17.2 s (auto)
Krea-2 Turbo fp8, 8 steps ~60 s (#304, math kernel) 32.1 / 32.1 s
shared GPU memory held by the process (PDH) 36.1 GiB 0.26 GiB

Related Issues / Discussions

QA Instructions

Lock and pins

  • uv lock --locked with uv 0.11.28, scripts/check_pins.py and scripts/check_aarch64_lock.py pass.
  • uv export per extra against main:
    • only torch/torchvision/triton for rocm move, plus the 78 device packages, rocm, rocm-sdk-core, rocm-sdk-libraries and rocm-bootstrap;
    • bitsandbytes goes to 0.50.2 for rocm on Windows and Linux x86_64: AMD's torch reports HIP 7.15, for which 0.49 looks for a rocm85 library that does not exist. uv's fork preference also moves cpu, cuda and the no-extra resolution to 0.50.2 on Windows and Linux x86_64. 0.50 logs the int8 bf16→fp16 cast notice on every matmul; silence_int8_cast_notice() filters it for the FLUX T5 and Qwen-Image int8 paths;
    • xpu is unchanged.
  • Two metadata overrides keep the lock independent of the machine that writes it:
    • The Linux torch wheel requires AMD's triton without a marker, and uv reads one wheel's metadata for both platforms.
    • rocm is an sdist whose build asks the machine's GPU which kernel package to require.

Installs from the lock

  • Windows: uv sync --locked --extra rocm --extra test installs. torch 2.13.0+rocm10.0.0 sees both GPUs and reports 22 arch targets.
  • Linux (WSL, no GPU): uv sync --locked --extra rocm installs. import torch, triton works, and get_arch_list() lists all 25 targets.
    • The venv is 20 GB, against 17 GB for the previous rocm7.2 lock.

Tests

  • New: tests/app/util/test_rocm_aotriton.py, tests/app/util/test_rocm_integrated_gpu.py, and CUDA/ROCm iGPU cases in tests/backend/util/test_devices.py.
  • Passing on the CUDA venv (RTX 4090, torch 2.13.0+cu130):
    • tests/backend/util, tests/app/util, tests/backend/model_manager/load/model_cache, tests/app/services/config, tests/app/services/session_processor, tests/test_check_pins.py;
    • the bitsandbytes NF4/LLM.int8 tests with bitsandbytes 0.50.2.
  • Passing on the ROCm venv: tests/backend/util, tests/app/util, tests/app/routers/test_app_info.py.
    • test_get_generation_devices_auto_expands_to_all_cuda read the real iGPU and is now hardware-independent.
  • tests/backend/model_manager/load/model_cache/test_shared_weights_gpu.py crashes on the ROCm venv when the iGPU is visible: it computes on cuda:1, which is the iGPU there. This is the crash described above, not a regression.
  • -m slow ROCm hardware tests, with and without the AOTriton variable: 6 pass, and the two canaries fail as they are meant to on this GPU:
    • The Wan conv3d decomposition is 2× slower than native conv3d here (89 vs 42 ms).
    • With the variable, the fused kernels are correct for 512-wide heads.
    • Both workarounds stay (see Review).
  • ruff check/ruff format --check are clean. openapi.json, schema.ts and docs/src/generated/settings.json are regenerated.

E2E on the RX 9060 XT (table above):

  • Images inspected, clean. With auto, PSNR against the rocm7.14.1 math-kernel reference is 42.2/32.1 dB, the same as in the earlier measurement; off gives 57.8 dB.
  • The startup log shows the iGPU hidden and the kernels on. No paging warning after the fix; before it, one after every generation.

Not tested:

  • Linux with an AMD GPU: @lstein, could you run it on the W7900?
  • The Docker ROCm image build.
  • GPUs other than gfx1200, and Windows 11 (AMD lists Windows 11; Windows 10 worked here).
  • The launcher path.
  • The full test suite locally (CI runs it).

Review

Remaining risks and limitations:

  • AMD's index serves no hashes and redirects per package. It must stay explicit because it also mirrors PyPI packages. CI's uv lock --locked now needs it reachable.
  • Size:
    • A plain uv sync --extra rocm installs the kernels of every target, about 4.3 GiB extra download on Windows and 8.3 GiB on Linux.
    • The Docker ROCm image grows accordingly. Its build job has a 90-minute timeout, and the image build has not been run.
  • Two ROCm workarounds stay, measured as unnecessary on gfx1200/ROCm 10:
    • The wide-head math guard: with the variable, the fused kernels compute 512-wide heads correctly here, 37.9 ms instead of 93 ms per VAE mid-block call.
    • The Wan conv3d decomposition: 2× slower here.
    • Both were measured necessary on a W7900 under rocm7.2, and RDNA3 under ROCm 10 is unmeasured, so dropping them is left for a follow-up with that measurement.
  • auto covers gfx1200 only. gfx1101, gfx1102, gfx1150 and gfx1151 need the variable too (per feat(deps): support ROCm on Windows through the rocm extra #9606), but have not been measured here; on them, on is an opt-in.
  • An AMD APU used on its own (e.g. Ryzen AI Max) keeps the model cache's discrete-GPU memory model. That case is unmeasured.

Compatibility / Rollout

  • New setting: rocm_aotriton_experimental (auto | on | off, default auto). It is additive, needs no migration, and the API contracts are regenerated.
  • HIP_VISIBLE_DEVICES on Windows ROCm is now set at startup when an iGPU sits next to a discrete GPU, unless the user set it, CUDA_VISIBLE_DEVICES or ROCR_VISIBLE_DEVICES. Explicit cuda:N values then count the discrete GPUs only. Anyone who ran main with a self-installed Windows ROCm torch and an explicit cuda:N in HIP's full numbering gets renumbered once: an index past the end fails at startup with a clear error, and an iGPU index is refused, but with two discrete GPUs behind an iGPU, cuda:1 would silently select the other card.
  • pins.json: linux.rocm points at AMD's index. It is read only by launcher installs of releases before 6.14, so legacy Windows installs still get no ROCm entry.
  • GPU support on Linux: gfx900/gfx906 (Radeon VII, Vega 56/64, MI25/MI50/MI60) are no longer covered by the rocm extra. The docs point to the previous release or --torch-backend=rocm7.2.

Checklist

  • The PR has a short but descriptive title, suitable for a changelog
  • Meaningful regression coverage added / updated where needed; obsolete tests/code removed
  • Persisted-state and API changes include required migrations / compatibility validation
  • Relevant performance/efficiency opportunities considered; material claims have evidence
  • Material review findings resolved and relevant checks rerun
  • Documentation added / updated (if applicable)
  • Updated What's New copy (if doing a release after this PR)

Pfannkuchensack and others added 3 commits October 2, 2026 03:52
The rocm extra moves from PyTorch's 2.13.0+rocm7.2 (Linux only) to AMD's 2.13.0+rocm10.0.0 for Linux
x86_64 and Windows. The lock carries the GPU kernel packages of every target AMD builds, so the launcher
can leave out the ones a machine does not need. Linux loses gfx900/gfx906, which ROCm 10 dropped.
bitsandbytes moves to 0.50.2, the first release with a Windows ROCm binary.

A new setting, rocm_aotriton_experimental (auto/on/off), sets TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL
before anything runs attention. auto turns AOTriton's experimental fused kernels on only for a ROCm 10
build on GPUs where they were measured correct (gfx1200 so far); under torch 2.12+rocm7.14 the same
switch made every fused call fail on that GPU.

Builds on the dependency and documentation changes of invoke-ai#9606.

Co-authored-by: 0xDELUXA <djernovevo@gmail.com>
…n CUDA/ROCm

ROCm enumerates the Radeon graphics of a Ryzen CPU next to a discrete card (measured: gfx1036 beside an
RX 9060 XT on Windows), so `auto` handed the iGPU half of the queue. Integrated GPUs were only filtered
on XPU. device_is_integrated now also answers for CUDA/HIP devices from their properties, and
`device: auto` no longer defaults to an iGPU that is enumerated before the discrete card.
…ializes it

While HIP on Windows can see the Ryzen iGPU, a process's first context on the discrete card shows 17.9 GiB
of shared GPU memory and its first kernel launch another 17.9 GiB, with every ROCm build tried and without
any Invoke code. It is neither RAM nor commit charge, but Task Manager shows 36 GiB of shared GPU memory,
and Invoke's paging check read it as memory pushed out of VRAM: its warning fired after every generation.

A short child process now asks HIP which devices are integrated (reading device properties creates no
context), and HIP_VISIBLE_DEVICES hides them. Only on Windows ROCm, only when no visibility variable is
set, and only when device and generation_devices are both auto, so explicit cuda:N indices keep their
meaning.
@github-actions github-actions Bot added python PRs that change python files Root backend PRs that change backend files services PRs that change app services python-tests PRs that change python tests docs PRs that change docs python-deps PRs that change python dependencies labels Oct 2, 2026

@lstein lstein left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this — it's a big step for AMD users, and the lock work (the two dependency-metadata overrides, the per-target device packages) is careful. I tested it on Linux with 2× Radeon PRO W7900 (gfx1100), venv synced from this lock with the non-gfx1100 device packages left out via --no-install-package. That install path works, and the gfx1100 rocBLAS/hipBLASLt kernels are all present.

One blocker for Linux ROCm, plus a few smaller issues.

Blocking

1. bitsandbytes is broken on Linux ROCm

pyproject.toml:171 pins bitsandbytes>=0.50.2 for win32 only, so the Linux rocm extra stays on 0.49.2. AMD's ROCm 10 torch reports torch.version.hip == '7.15.26333' (same on both wheels). bitsandbytes 0.49.2 builds the library name as major * 10 + minor, so 7.15 becomes 85 and it looks for libbitsandbytes_rocm85.so, which exists in no release. It then falls back to its stand-in library, and every native call raises.

Reproduced on the W7900. At startup:

bitsandbytes library load error: Configured ROCm binary not found at .../bitsandbytes/libbitsandbytes_rocm85.so

and a FLUX.1 NF4 transformer + t5_bnb_int8_quantized_encoder generation fails:

Detected PyTorch ROCm version:   7.15
Available pre-compiled versions: 6.2 6.3 6.4 7.0 7.1 7.2
...
Native code method attempted to call: lib.cint8_vector_quant()

On main this works, because torch 2.13.0+rocm7.2 reports HIP 7.2 and rocm72.so ships.

0.50.2 joins the digits (f"{major}{minor}" → 715) and ships libbitsandbytes_rocm715.so for Linux as well as Windows. With bitsandbytes==0.50.2 installed in the same venv, the same FLUX NF4 + int8-T5 generation completes and the image is fine.

Fix: widen the marker to sys_platform == 'win32' or (sys_platform == 'linux' and platform_machine == 'x86_64'). A Linux install on a machine without a GPU can't catch this, because torch.cuda is unavailable there and bitsandbytes never picks a GPU library.

Should fix

2. 0.50.2 floods the log with the int8 cast warning

0.50.2 moved MatMul8bitLt: inputs will be cast from torch.bfloat16 to float16 during quantization from warnings.warn to logger.warning on the bitsandbytes.autograd._functions logger. The existing warnings.filterwarnings in invokeai/backend/quantization/bnb_llm_int8.py:14 (and the Qwen-Image encoder's catch_warnings block) no longer catches it, so it prints on every int8 matmul. I saw a wall of these during the generation above.

As the PR body notes, the lock also moves cpu, cuda and the no-extra install to 0.50.2 on Windows and Linux x86_64, so CUDA users of the int8 T5 get the same flood. A logging.Filter on that logger next to the existing warnings filter would cover both versions:

class _SuppressInt8CastWarning(logging.Filter):
    def filter(self, record: logging.LogRecord) -> bool:
        return not record.getMessage().startswith("MatMul8bitLt: inputs will be cast")

logging.getLogger("bitsandbytes.autograd._functions").addFilter(_SuppressInt8CastWarning())

3. Hiding the iGPU only while the config is auto lets device indices shift

hide_integrated_gpus_on_rocm_windows only sets HIP_VISIBLE_DEVICES while device and generation_devices are both auto (rocm_integrated_gpu.py:87). But the settings UI lists, validates and saves devices in the hidden numbering: get_generation_device_options (app_info.py:221) enumerates torch.cuda after hiding, and update_runtime_config (app_info.py:261) validates against it, then writes invokeai.yaml.

Example: HIP order is [discrete A, iGPU, discrete B]. With the iGPU hidden, the UI offers cuda:0 = A and cuda:1 = B. The user selects both and saves. On restart generation_devices is no longer auto, so nothing is hidden, and cuda:1 is now the iGPU. Per the PR description, any kernel there crashes the process with an access violation. With the iGPU enumerated first, picking [cuda:0] lands on it. With the iGPU last, nothing crashes, but the 36 GiB phantom shared memory and the paging warning come back.

Options: hide the iGPU whenever it sits next to a discrete GPU, regardless of device/generation_devices (it can't run kernels on Windows ROCm anyway, per the PR's own measurement), or keep the UI and the saved values in the full HIP numbering. I couldn't test this one; the W7900 box has no iGPU.

4. The probe can stall startup for 2 minutes

_PROBE_TIMEOUT_SECONDS = 120 (rocm_integrated_gpu.py:33). The PR itself notes a HIP process hanging on exit under Windows (the CLIP case under 7.14.1). If the child prints its answer and then hangs in teardown, startup waits the full 120 s and falls back with only a debug log (:96). Suggest ending the probe with sys.stdout.flush(); os._exit(0), a much shorter timeout, and a warning on the fallback.

Minor

  • torch.version.rocm exists on these wheels ('10.0.0') and would be a cleaner source for rocm_major_version than parsing the version label. It also covers locally built torch.
  • AMD's artifacts carry no hashes in the lock, including the rocm sdist, which runs build code at install time. You've already called this out; noting it as an accepted risk.

W7900 results (Linux, gfx1100)

  • Correctness: fp32/fp16/bf16 matmul on both cards match the CPU reference to the same error as rocm7.2.

  • SDXL, warm:

    Build Denoise VAE decode
    rocm7.2 (default, rocBLAS) 8.91 / 9.12 s 0.51 s
    ROCm 10, TORCH_BLAS_PREFER_HIPBLASLT=0 (rocBLAS) 8.93 s 0.54 s
    ROCm 10, default (hipBLASLt) 7.84 s 0.77 s

    ROCm 10 is as fast as or faster than rocm7.2 here. AMD's build defaults to hipBLASLt, where PyTorch's rocm7.2 build used rocBLAS. A plain a @ b benchmark makes hipBLASLt look ~30% slower at 8192³ on gfx1100, but end to end it's ~12% faster for SDXL denoising, so the default is the right choice. The VAE decode is ~0.2 s slower with hipBLASLt (single run).

  • First run: the first cold SDXL run on ROCm 10 had ~4 s extra in denoise. That's probably MIOpen's per-version kernel cache starting empty, so it's likely a one-time cost after install.

  • AOTriton auto: leaves the experimental kernels off on gfx1100, as intended, and the startup log says so clearly.

  • I did not run the -m slow hardware canaries (wide-head guard, Wan conv3d decomposition) under ROCm 10, so the RDNA3 measurement for the follow-up is still open.

@baramofme

Copy link
Copy Markdown

I run ROCm 10 in production on Linux / RX 7900 XT (gfx1100) with a self-built image and measured large gains (numbers in Additional Content).

Measured evidence (Krea-2 Turbo fp8, 1024², 8 steps)

stage rocm7.2 (main-rocm) rocm 10 (this build)
cold first model load, encoder 18.68 s 3.50 s
cold first denoise pass 72.43 s 20.17 s
warm denoise — 16.06 s
end-to-end per image (warm) — ~31–40 s

Application process (steps we ran, Linux / RX 7900 XT / gfx1100)

  1. docker build --build-arg GPU_DRIVER=rocm . from the ROCm 10 branch (we used the head
    commit of feat(deps): ROCm 10 on Linux and Windows from AMD's index #9643 plus the [bug]: First load of a model is extremely slow on ROCm — copies to VRAM from mmap-backed safetensors run at ~2 MB/s #9487 clone fix ported into tensor_aliases.py).

  2. Push to our registry, deploy with the existing invokeai-rocm compose shape
    (runtime: amd, profiles: [rocm]) adjusted to the host's group IDs.

  3. Startup verification inside the container:

    python -c "import torch; print(torch.__version__, torch.version.hip, torch.cuda.get_device_name(0))"
    

    → 2.13.0+rocm10.0.0, HIP 7.15.x (this is the HIP runtime version, not the ROCm
    release — beware any gate keyed on this string), device gfx1100 reported as (11, 0).

  4. Generate via the queue API. First-load behaviour with the [bug]: First load of a model is extremely slow on ROCm — copies to VRAM from mmap-backed safetensors run at ~2 MB/s #9487 fix: the 7.67 GB
    mmap-backed CLIP/VAE encoder reached VRAM in 5.64 s (~1.36 GB/s) — no longer the
    ~2 MB/s pathological path described in [bug]: First load of a model is extremely slow on ROCm — copies to VRAM from mmap-backed safetensors run at ~2 MB/s #9487.

# Conflicts:
#	invokeai/frontend/api/openapi.json
#	tests/backend/util/test_devices.py
AMD's ROCm 10 torch reports HIP 7.15; bitsandbytes 0.49 then looks for libbitsandbytes_rocm85.so, which no release ships, and every native call fails.
0.50.2 names the library rocm715 and ships it for Linux as well as Windows.
…s on every matmul

bitsandbytes 0.50 reports the bf16-to-fp16 cast through its logger instead of warnings, past the existing filter, once per int8 matmul.
One silence_int8_cast_notice() now filters both forms, for the FLUX T5 int8 path and the Qwen-Image int8 encoder.
…d bound the probe

The settings UI numbers devices after the iGPU is hidden, so hiding it only while device and generation_devices were auto renumbered saved choices on the next start.
Naming an integrated GPU beside a discrete one under Windows ROCm is now refused with a reason instead of crashing at the first kernel.
The probe takes the answer as soon as the child prints it, gives up after 60 s, and warns with the reason when it cannot answer.
The +rocmN label of the torch version stays the fallback for builds that leave torch.version.rocm unset.
@github-actions github-actions Bot added the invocations PRs that change invocations label Oct 9, 2026
@Pfannkuchensack

Copy link
Copy Markdown
Member Author

@lstein Thanks for testing on the W7900. All points are addressed in a493101..513e04a:

  1. bitsandbytes: the marker now covers Linux x86_64, and the lock is regenerated. Only bitsandbytes' marker changes, so Linux rocm resolves 0.50.2. I can't test Linux ROCm here; could you rerun the FLUX NF4 + int8 T5 case?
  2. Cast notice: silence_int8_cast_notice() in bnb_llm_int8.py filters both the 0.49 warning and the 0.50 log record. The Qwen-Image int8 encoder uses it too; its old catch_warnings only wrapped the load. With the real bitsandbytes 0.50.2 on CUDA, every bf16 int8 matmul logged the notice without the filter, and none does with it.
  3. iGPU numbering: the iGPU is now hidden on Windows ROCm whatever device and generation_devices say, unless a *_VISIBLE_DEVICES variable is set. The settings UI, the saved config and the runtime now share one numbering. As a safety net, naming an integrated GPU next to a discrete one on Windows ROCm is refused with a clear error instead of crashing at the first kernel. I checked this on an RX 9060 XT with a Ryzen iGPU and generation_devices: ["cuda:0"]. Anyone who ran main with a self-installed Windows ROCm torch and an explicit cuda:N gets renumbered once; the PR description now says so.
  4. Probe: the parent takes the answer line as soon as it's printed and ends the child itself, so a teardown hang no longer delays startup. The child also exits via os._exit. The timeout is 60 s; it took 1.4–2.4 s warm on the 9060 XT. A shorter limit buys little: when HIP hangs while enumerating devices, which I saw after a driver fault, the main process hangs too as soon as it imports torch. A failure now logs a warning with the reason.
  • Minor: torch.version.rocm is read first, with the +rocmN label as fallback. The missing lock hashes stay an accepted risk.

logger.info(
f"ROCm ({gpus}): AOTriton's experimental fused attention kernels are off ({reason}). On a GPU that "
"AOTriton marks experimental, attention then runs on the slower math kernel; set "
"rocm_aotriton_experimental: on to try the fused kernels."

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Quote the YAML value in the opt-in instruction

A user following this startup message and adding rocm_aotriton_experimental: on to invokeai.yaml cannot restart InvokeAI. load_and_migrate_config() uses yaml.safe_load(), which converts unquoted on and off to booleans, while the new setting accepts only the strings auto, on, and off. Reproducing the parse and validation produces a literal_error for True, which the config loader turns into a startup failure. Recommend rocm_aotriton_experimental: "on" (and quote "off" in corresponding guidance), or normalize YAML booleans before validating this field.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, thanks. Fixed in 6a0048a. Instead of only quoting the hint, the field now accepts YAML's booleans: an unquoted on/off (read as True/False by yaml.safe_load) maps to "on"/"off" before validation. Both the log line's advice and the natural way of writing the setting work. tests/test_config.py loads real YAML files with unquoted on, off, quoted "on" and auto through load_and_migrate_config.

YAML 1.1 reads them as booleans, so `rocm_aotriton_experimental: on`, as the startup log suggests, failed validation and stopped InvokeAI from starting.
@Pfannkuchensack
Pfannkuchensack requested a review from lstein October 9, 2026 04:53
…ing bitsandbytes

The Qwen-Image int8 encoder reached the filter through bnb_llm_int8, which imports bitsandbytes; macOS has none, so its int8 load tests failed there.
The filter now lives in bnb_cast_notice, which needs only logging and warnings, and its test runs on every platform.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backend PRs that change backend files docs PRs that change docs invocations PRs that change invocations python PRs that change python files python-deps PRs that change python dependencies python-tests PRs that change python tests Root services PRs that change app services

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants