Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ sidebar:
order: 2
---

As of v5.6.0, Invoke has a low-VRAM mode. It works on systems with dedicated GPUs (Nvidia GPUs on Windows/Linux and AMD GPUs on Linux).
As of v5.6.0, Invoke has a low-VRAM mode. It works on systems with dedicated GPUs (Nvidia and AMD GPUs on Windows/Linux).

This allows you to generate even if your GPU doesn't have enough VRAM to hold full models. Most users should be able to run even the beefiest models - like the ~24GB unquantised FLUX dev model.

Expand Down Expand Up @@ -255,7 +255,7 @@ It is strongly suggested to disable this feature:

### AMD GPUs on Windows (ROCm)

Invoke does not officially support ROCm on Windows (see the [System Requirements](/start-here/system-requirements/)). If you run a ROCm build of PyTorch there anyway, Windows behaves as if the Nvidia sysmem fallback were always on: an allocation that does not fit is moved into much slower system memory instead of failing, and it stays there. Task Manager shows it as shared GPU memory. Invoke works around this where it can:
With a ROCm build of PyTorch (see the [System Requirements](/start-here/system-requirements/)), Windows behaves as if the Nvidia sysmem fallback were always on: an allocation that does not fit is moved into much slower system memory instead of failing, and it stays there. Task Manager shows it as shared GPU memory. Invoke works around this where it can:

- The model cache plans against the video-memory budget Windows grants Invoke, not against the card's total VRAM.
- The PyTorch allocator defaults to `expandable_segments:True` (see [PyTorch CUDA allocator config](#pytorch-cuda-allocator-config)), so fragmented VRAM does not force allocations out.
Expand Down
15 changes: 14 additions & 1 deletion docs/src/content/docs/start-here/manual.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -111,6 +111,13 @@ The following commands vary depending on the version of Invoke being installed a
--torch-backend=xpu
```
</TabItem>
<TabItem label="AMD GPU">
Do not use a torch backend. AMD publishes its ROCm wheels on its own index, with the GPU kernels in a separate package per GPU target. In the next step, pin them for your GPU's target and point `uv` at that index instead:
```sh
uv pip install <PACKAGE_SPECIFIER>==<VERSION> "torch[device-<TARGET>]==2.13.0+rocm10.0.0" "torchvision[device-<TARGET>]==0.28.0+rocm10.0.0" --index https://stable.repo.amd.com/rocm/whl-next/ --index-strategy unsafe-best-match --python 3.12 --python-preference only-managed --force-reinstall
```
`<TARGET>` is `gfx1201` for the RX 9070 series and Radeon AI PRO R9700, `gfx1200` for the RX 9060 series, `gfx1100` for the RX 7900 series, `gfx1101` for the RX 7700 and 7800 series, `gfx1102` for the RX 7600 series, and `gfx1151` for Ryzen AI Max. For other GPUs, see AMD's [GPU target table](https://github.com/ROCm/TheRock/blob/main/RELEASES.md#gfx-target-lookup-table).
</TabItem>
<TabItem label="Other / no GPU">
Do not use a torch backend.
</TabItem>
Expand All @@ -133,7 +140,13 @@ The following commands vary depending on the version of Invoke being installed a
On ARM64 (`aarch64`, e.g. Raspberry Pi 5), do **not** use a torch backend. The default PyPI wheels are CPU-only on ARM64 and work out of the box.
</TabItem>
<TabItem label="Linux + AMD GPU">
Use:
Do not use a torch backend. Install AMD's ROCm 10 wheels from AMD's index, as the launcher does, with the GPU kernels pinned for your GPU's target:
```sh
uv pip install <PACKAGE_SPECIFIER>==<VERSION> "torch[device-<TARGET>]==2.13.0+rocm10.0.0" "torchvision[device-<TARGET>]==0.28.0+rocm10.0.0" --index https://stable.repo.amd.com/rocm/whl-next/ --index-strategy unsafe-best-match --python 3.12 --python-preference only-managed --force-reinstall
```
`<TARGET>` is `gfx1201` for the RX 9070 series and Radeon AI PRO R9700, `gfx1200` for the RX 9060 series, `gfx1100` for the RX 7900 series, `gfx1101` for the RX 7700 and 7800 series, `gfx1102` for the RX 7600 series, and `gfx1151` for Ryzen AI Max. For other GPUs, see AMD's [GPU target table](https://github.com/ROCm/TheRock/blob/main/RELEASES.md#gfx-target-lookup-table).

ROCm 10 has no kernels for Vega-based cards (gfx900/gfx906). For those, use PyTorch's ROCm 7.2 build instead:
```sh
--torch-backend=rocm7.2
```
Expand Down
47 changes: 39 additions & 8 deletions docs/src/content/docs/start-here/system-requirements.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ The requirements below are rough guidelines for best performance. GPUs with less

- All Apple Silicon (M1, M2, etc) Macs work, but 16GB+ memory is recommended.
- Nvidia GPUs need compute capability 7.5 or newer (GTX 16xx, RTX 20xx and everything after) and a driver from the R580 series or newer. Maxwell, Pascal and Volta cards (GTX 9xx/10xx, Titan V, Tesla P40/P100/V100) are not supported by Invoke's CUDA build: stay on the previous Invoke release, or do a [manual install](../manual) with `--torch-backend=cu126`, whose PyTorch build still includes these GPUs (not tested by the Invoke team). With a driver that is too old, Invoke starts on the CPU and logs an error saying so; on an unsupported GPU it logs an error at startup and generation fails.
- AMD GPUs are supported on Linux only. The VRAM requirements are the same as Nvidia GPUs.
- AMD GPUs are supported on Windows and Linux `x86_64` through ROCm 10 (see [AMD](#amd) below for the GPUs and drivers). On Linux, Vega-based cards (gfx900/gfx906) are no longer supported. The VRAM requirements are the same as Nvidia GPUs.
- Intel Arc GPUs (Alchemist, Battlemage and newer) are supported on Windows and Linux `x86_64`. The VRAM requirements are the same as Nvidia GPUs.
- Linux ARM64 (`aarch64`) devices — e.g. Raspberry Pi 5, other SBCs, ARM servers — are supported in CPU-only mode. Local generation is slow without a GPU, but API-backed models (e.g. GPT Image, Gemini) work well.

Expand Down Expand Up @@ -145,8 +145,15 @@ If, after restarting the app, this doesn't improve your performance, either rest

### AMD

:::tip[Linux Only]{icon="linux"}
AMD GPUs are supported on Linux only, due to ROCm (the AMD equivalent of CUDA) support being Linux only.
Invoke installs AMD's ROCm 10 build of PyTorch on Windows and Linux. It brings the ROCm runtime with it, so
there is no separate ROCm or HIP SDK to install, but the GPU needs an up-to-date driver: [AMD Software: Adrenalin
Edition] on Windows, AMD's kernel driver on Linux (see [Linux](#linux) below). AMD's [ROCm compatibility matrix] lists the GPUs and operating
systems ROCm 10 supports. AMD lists Windows 11; the Invoke team has run it on Windows 10 with an RX 9060 XT.

:::caution[Vega-based cards on Linux]
ROCm 10 has no kernels for gfx900 and gfx906 (Radeon RX Vega 56/64, Radeon VII, Instinct MI25/MI50/MI60).
Stay on the previous Invoke release, or do a [manual install](../manual) with `--torch-backend=rocm7.2`,
whose PyTorch build still includes them (not tested by the Invoke team).
:::

:::caution[Bumps Ahead]
Expand All @@ -160,11 +167,33 @@ If, after restarting the app, this doesn't improve your performance, either rest
(for example `MIOPEN_FIND_MODE=NORMAL` to let MIOpen re-tune its kernel database).
:::

:::note[Fused attention on newer Radeon GPUs]
On some GPUs - among them the RX 7600/7700/7800 series, Ryzen AI 300/Max and the RX 9060 series -
PyTorch uses its fused attention kernels only when `TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL` is set.
Without them, attention runs on the slower math kernel, which also needs far more memory. The
`rocm_aotriton_experimental` setting decides. Its default, `auto`, turns the kernels on where they
were measured correct with ROCm 10: so far the RX 9060 series (gfx1200), where Z-Image at 1024px
takes 17 s instead of 36 s. Set it to `on` to try them on another GPU, and back to `auto` if images
come out wrong or generation fails. The startup log says whether they are on.
:::

:::note[Integrated graphics]
ROCm lists the Radeon graphics built into Ryzen CPUs as a GPU of its own. When a discrete GPU is
present, Invoke leaves the integrated one out of `generation_devices: auto`. On Windows it also hides
it from ROCm at startup (by setting `HIP_VISIBLE_DEVICES`) unless you set `HIP_VISIBLE_DEVICES`,
`CUDA_VISIBLE_DEVICES` or `ROCR_VISIBLE_DEVICES` yourself: while ROCm can see the integrated GPU,
Task Manager shows tens of GB of shared GPU memory for Invoke that is not really in use, and nothing
can run on it. Device names such as `cuda:1` in `device`, `generation_devices` and the settings then
count the discrete GPUs only, and naming the integrated GPU is refused with an error. Set
`HIP_VISIBLE_DEVICES` yourself to choose the devices.
:::

:::note[Wide-head attention runs on the math kernel]
On ROCm builds Invoke routes attention whose head dimension exceeds 256 to PyTorch's math kernel.
The fused kernels return wrong output there on torch 2.13 / ROCm 7.2 (measured on a W7900), which
showed up as black or noisy images from the image VAEs - their mid-block attention is one 512-wide
head. The math kernel builds the whole attention matrix (about 4 GB for a 1024px VAE decode, 21 GB
head. With ROCm 10 they were correct on an RX 9060 XT, but RDNA3 cards have not been measured
again, so the threshold stays. The math kernel builds the whole attention matrix (about 4 GB for a 1024px VAE decode, 21 GB
at 1536px), so Invoke computes it in pieces of at most 1 GB. The result can differ from a single
math call in the last bit of a bf16 value, and each such call is roughly 10% slower. Export
`INVOKE_ROCM_FUSED_SDPA_MAX_HEAD_DIM` to raise the threshold on a build whose fused kernels are
Expand All @@ -187,11 +216,11 @@ If, after restarting the app, this doesn't improve your performance, either rest
`rocprof`/ROCTracer-based kernel profiling, which this disables.
:::

Run `rocm-smi` on your system's command line verify that drivers and ROCm are installed. If this command fails, or doesn't report versions, you will need to install them.

Go to the [ROCm Documentation] and carefully follow the instructions for your system to get everything installed.
#### Linux

Confirm that `rocm-smi` displays driver and CUDA versions after installation.
The ROCm libraries come with Invoke's PyTorch packages, so ROCm itself does not need to be installed. The GPU needs
AMD's `amdgpu` kernel driver; on the distributions in AMD's [ROCm compatibility matrix], the driver that comes with
the kernel is enough. For other distributions, follow the [ROCm Documentation] to install the driver.

#### Linux - via Docker Container

Expand All @@ -209,6 +238,8 @@ An alternative to installing ROCm locally is to use a [ROCm docker container] to
[Intel Arc & Iris Xe Graphics driver]: https://www.intel.com/content/www/us/en/download/785597/intel-arc-iris-xe-graphics-windows.html
[Intel client GPU install guide]: https://dgpu-docs.intel.com/driver/client/overview.html
[ROCm Documentation]: https://rocmdocs.amd.com
[ROCm compatibility matrix]: https://rocm.docs.amd.com/en/docs-10.0.0/compatibility/compatibility-matrix.html
[AMD Software: Adrenalin Edition]: https://www.amd.com/en/support/download/drivers.html
[ROCm docker container]: https://rocmdocs.amd.com/en/latest/Deep_learning/Deep_learning.html#docker-containers
[ROCm/ROCm#6522]: https://github.com/ROCm/ROCm/issues/6522
[ROCm/TheRock#7051]: https://github.com/ROCm/TheRock/issues/7051
17 changes: 16 additions & 1 deletion docs/src/generated/settings.json
Original file line number Diff line number Diff line change
Expand Up @@ -661,7 +661,7 @@
{
"category": "DEVICE",
"default": "auto",
"description": "Devices to use for parallel generation. `auto` (the default) uses every available GPU, running one generation session per GPU concurrently and distributing jobs fairly across users \u2014 unless the legacy `device` setting is pinned to a specific device, in which case `auto` uses only that device (preserving configs that pinned `device` before multi-GPU support existed). Provide an explicit list (e.g. `[cuda:0, cuda:1]`) to use specific devices regardless of `device`, or a single-device list (e.g. `[cuda:0]`) to run serially. On systems without a GPU, `auto` resolves to the single `cpu`/`mps` device.<br>Valid values: `auto`, or a list whose entries are each `cpu`, `cuda`, `mps`, `xpu`, `cuda:N`, or `xpu:N` (where N is a device number)",
"description": "Devices to use for parallel generation. `auto` (the default) uses every available GPU, except an integrated GPU next to a discrete one, running one generation session per GPU concurrently and distributing jobs fairly across users \u2014 unless the legacy `device` setting is pinned to a specific device, in which case `auto` uses only that device (preserving configs that pinned `device` before multi-GPU support existed). Provide an explicit list (e.g. `[cuda:0, cuda:1]`) to use specific devices regardless of `device`, or a single-device list (e.g. `[cuda:0]`) to run serially. On systems without a GPU, `auto` resolves to the single `cpu`/`mps` device.<br>Valid values: `auto`, or a list whose entries are each `cpu`, `cuda`, `mps`, `xpu`, `cuda:N`, or `xpu:N` (where N is a device number)",
"env_var": "INVOKEAI_GENERATION_DEVICES",
"literal_values": [],
"name": "generation_devices",
Expand Down Expand Up @@ -696,6 +696,21 @@
"type": "typing.Literal['auto', 'float16', 'bfloat16', 'float32']",
"validation": {}
},
{
"category": "DEVICE",
"default": "auto",
"description": "Use AOTriton's fused (flash and memory-efficient) attention kernels on AMD GPUs that PyTorch marks experimental for them, such as the RX 7600/7700/7800 series, Ryzen AI 300/Max and the RX 9060 series, by setting TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL. Without them, attention on those GPUs runs on the slower math kernel, which also needs far more memory. `auto` turns them on only with a ROCm 10 build of PyTorch, and only when every GPU used for generation is one they were measured correct on (so far gfx1200, the RX 9060 series); `on` and `off` decide for any build and GPU. A TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL value already set in the environment takes precedence. Has no effect on other GPUs.",
"env_var": "INVOKEAI_ROCM_AOTRITON_EXPERIMENTAL",
"literal_values": [
"auto",
"on",
"off"
],
"name": "rocm_aotriton_experimental",
"required": false,
"type": "typing.Literal['auto', 'on', 'off']",
"validation": {}
},
{
"category": "GENERATION",
"default": false,
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@
from invokeai.app.services.shared.invocation_context import InvocationContext
from invokeai.backend.model_manager.load.model_cache.model_cache import MB, MODEL_LOAD_LOCK
from invokeai.backend.model_manager.load.model_util import calc_model_size_by_fs
from invokeai.backend.quantization.bnb_cast_notice import silence_int8_cast_notice
from invokeai.backend.quantization.dequantizing_linear import peak_dequant_transient_bytes
from invokeai.backend.qwen2_5_vl.qwen2_5_vl_assets import (
load_bundled_qwen2_5_vl_preprocessor_config_dict,
Expand Down Expand Up @@ -395,8 +396,6 @@ def _load_quantized_encoder(self, context: InvocationContext) -> _LoadedEncoder:
layers to the CPU, which BnB int8 refuses with "Some modules are dispatched on
the CPU or the disk" (issue #9147).
"""
import warnings

from transformers import BitsAndBytesConfig, Qwen2_5_VLConfig, Qwen2_5_VLForConditionalGeneration

encoder_config = context.models.get_config(self.qwen_vl_encoder.text_encoder)
Expand All @@ -415,6 +414,7 @@ def _load_quantized_encoder(self, context: InvocationContext) -> _LoadedEncoder:
)
else: # int8
bnb_config = BitsAndBytesConfig(load_in_8bit=True)
silence_int8_cast_notice()

# Load onto this worker's execution device, never `device_map="auto"`: "auto" sizes its plan from whatever
# VRAM is free *right now* and quietly spills to the CPU when the cached models fill the card, and it shards
Expand Down Expand Up @@ -464,9 +464,7 @@ def _load_quantized_encoder(self, context: InvocationContext) -> _LoadedEncoder:
model_config = Qwen2_5_VLConfig.from_pretrained(str(encoder_path), local_files_only=True)
state_dict = _read_checkpoint(encoder_path)

with MODEL_LOAD_LOCK.write_lock(), warnings.catch_warnings():
# BnB int8 internally casts bfloat16→float16; the warning is harmless
warnings.filterwarnings("ignore", message="MatMul8bitLt.*cast.*float16")
with MODEL_LOAD_LOCK.write_lock():
text_encoder = Qwen2_5_VLForConditionalGeneration.from_pretrained(
None if state_dict is not None else str(encoder_path),
config=model_config,
Expand Down
13 changes: 13 additions & 0 deletions invokeai/app/run_app.py
Original file line number Diff line number Diff line change
Expand Up @@ -69,13 +69,26 @@ def run_app() -> None:

logger = InvokeAILogger.get_logger(config=app_config)

# Before torch is imported anywhere: HIP reads its device list once, when torch initializes it, and an explicit
# allocator configuration below imports torch. Runs a short child process on Windows ROCm.
from invokeai.app.util.rocm_integrated_gpu import hide_integrated_gpus_on_rocm_windows

hide_integrated_gpus_on_rocm_windows(logger)

# Configure the torch CUDA memory allocator.
# NOTE: It is important that this happens before torch is imported.
if app_config.pytorch_cuda_alloc_conf:
configure_torch_cuda_allocator(app_config.pytorch_cuda_alloc_conf, logger)
else:
apply_rocm_windows_allocator_default(logger)

# Decide on AOTriton's experimental attention kernels before anything can run attention: torch reads the switch
# once, at its first fused-kernel check, and the server already runs attention while starting up. This imports
# torch, so it comes after the allocator configuration.
from invokeai.app.util.rocm_aotriton import apply_rocm_aotriton_setting

apply_rocm_aotriton_setting(app_config.rocm_aotriton_experimental, app_config.generation_devices, logger)

# This import must happen after configure_torch_cuda_allocator() is called, because the module imports torch.
from invokeai.app.invocations.baseinvocation import InvocationRegistry
from invokeai.app.invocations.load_custom_nodes import load_custom_nodes
Expand Down
Loading
Loading