Skip to content
This repository was archived by the owner on Oct 4, 2026. It is now read-only.
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 2 additions & 4 deletions docker/Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -73,10 +73,8 @@ WORKDIR ${INVOKEAI_SRC}
# Install project's dependencies as a separate layer so they aren't rebuilt every commit.
# bind-mount instead of copy to defer adding sources to the image until next layer.
#
# NOTE: on arm64 (aarch64), torch/torchvision resolve to PyPI CPU wheels regardless of
# GPU_DRIVER — the WHL indexes carry aarch64 torch wheels (+cpu/+cu128) but no aarch64
# torchvision wheels, so the pinned pairs only apply on x86_64. Use GPU_DRIVER=cpu on arm64.
# CUDA-on-arm64 (e.g. Grace/GH200) would need a torchvision-specific workaround; see pyproject.toml.
# NOTE: on arm64 (aarch64), torch/torchvision resolve to the PyPI CPU wheels (2.7.1) regardless of
# GPU_DRIVER -- the pinned WHL-index pairs only apply on x86_64 (see pyproject.toml). Use GPU_DRIVER=cpu on arm64.
# x86_64/CUDA is the default
RUN --mount=type=cache,target=/root/.cache/uv \
--mount=type=bind,source=pyproject.toml,target=pyproject.toml \
Expand Down
2 changes: 1 addition & 1 deletion docs/src/content/docs/configuration/02.docker.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ import SystemRequirementsLink from '@components/SystemRequirmentsLink.astro'

## TL;DR

Ensure your Docker setup is able to use your GPU. Then:
Ensure your Docker setup is able to use your GPU. The CUDA image needs an Nvidia driver from the R580 series or newer on the host and a GPU with compute capability 7.5 or newer (GTX 16xx, RTX 20xx and later). Then:

```bash
docker run --runtime=nvidia --gpus=all --publish 9090:9090 ghcr.io/invoke-ai/invokeai
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ FP8 Storage is meant for **full precision** checkpoints (FP16 / BF16 / FP32). It
## Requirements

- **A CUDA or Intel XPU device.** Before first use InvokeAI runs a small FP8 test on the device. If the device, driver or PyTorch build cannot store and upcast `float8_e4m3fn` (some older AMD ROCm GPUs, for example), the test fails, InvokeAI logs a warning and loads the model without FP8. CPU and MPS are never used for FP8 Storage.
- **CUDA 12.x and recent PyTorch.** The `float8_e4m3fn` dtype was added in PyTorch 2.1 — InvokeAI's bundled versions satisfy this.
- **Recent PyTorch.** The `float8_e4m3fn` dtype was added in PyTorch 2.1 — InvokeAI's bundled versions satisfy this.

There is no hardware requirement for FP8 *compute* — InvokeAI casts back to FP16/BF16 for math. This means FP8 Storage works on GPUs that do not natively support FP8 matmul (e.g. RTX 30-series), at a small per-step throughput cost.

Expand Down
8 changes: 4 additions & 4 deletions docs/src/content/docs/start-here/manual.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -87,7 +87,7 @@ The following commands vary depending on the version of Invoke being installed a

6. Determine the package specifier to use when installing. This is a performance optimization.

- If you have an Nvidia 20xx series GPU or older, use `invokeai[xformers]`.
- If you have an Nvidia 16xx or 20xx series GPU, use `invokeai[xformers]`.
- If you have an Nvidia 30xx series GPU or newer, or do not have an Nvidia GPU, use `invokeai`.

7. Determine the torch backend to use for installation, if any. This is necessary to get the right version of torch installed. This is achieved by using [UV's built in torch support.](https://docs.astral.sh/uv/guides/integration/pytorch/#automatic-backend-selection)
Expand All @@ -102,7 +102,7 @@ The following commands vary depending on the version of Invoke being installed a
<TabItem label="Nvidia GPU">
Use:
```sh
--torch-backend=cu128
--torch-backend=cu130
```
</TabItem>
<TabItem label="Intel Arc GPU">
Expand All @@ -122,15 +122,15 @@ The following commands vary depending on the version of Invoke being installed a
<TabItem label="Linux + Nvidia GPU">
Use:
```sh
--torch-backend=cu128
--torch-backend=cu130
```
</TabItem>
<TabItem label="Linux + no GPU">
On `x86_64`, use:
```sh
--torch-backend=cpu
```
On ARM64 (`aarch64`, e.g. Raspberry Pi 5), do **not** use a torch backend — PyTorch's `cpu` index has no ARM64 torchvision wheels. The default PyPI wheels are CPU-only on ARM64 and work out of the box.
On ARM64 (`aarch64`, e.g. Raspberry Pi 5), do **not** use a torch backend. The default PyPI wheels are CPU-only on ARM64 and work out of the box.
</TabItem>
<TabItem label="Linux + AMD GPU">
Use:
Expand Down
5 changes: 4 additions & 1 deletion docs/src/content/docs/start-here/system-requirements.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,7 @@ Hardware requirements vary significantly depending on model and image output siz
The requirements below are rough guidelines for best performance. GPUs with less VRAM typically still work, if a bit slower. Follow the [Low VRAM Guide] to optimize performance.

- All Apple Silicon (M1, M2, etc) Macs work, but 16GB+ memory is recommended.
- Nvidia GPUs need compute capability 7.5 or newer (GTX 16xx, RTX 20xx and everything after) and a driver from the R580 series or newer. Maxwell, Pascal and Volta cards (GTX 9xx/10xx, Titan V, Tesla P40/P100/V100) are not supported by Invoke's CUDA build: stay on the previous Invoke release, or do a [manual install](../manual) with `--torch-backend=cu126`, whose PyTorch build still includes these GPUs (not tested by the Invoke team). With a driver that is too old, Invoke starts on the CPU and logs an error saying so; on an unsupported GPU it logs an error at startup and generation fails.
- AMD GPUs are supported on Linux only. The VRAM requirements are the same as Nvidia GPUs.
- Intel Arc GPUs (Alchemist, Battlemage and newer) are supported on Windows and Linux `x86_64`. The VRAM requirements are the same as Nvidia GPUs.
- Linux ARM64 (`aarch64`) devices — e.g. Raspberry Pi 5, other SBCs, ARM servers — are supported in CPU-only mode. Local generation is slow without a GPU, but API-backed models (e.g. GPT Image, Gemini) work well.
Expand All @@ -24,7 +25,7 @@ The requirements below are rough guidelines for best performance. GPUs with less

| Model Family | Best resolution | GPU (series) | VRAM (min) | RAM (min) | Notes |
|---|---:|---|---:|---:|---|
| SD1.5 | 512x512 | Nvidia 10xx+ | 4GB | 8GB | |
| SD1.5 | 512x512 | Nvidia 16xx+ | 4GB | 8GB | |
| SDXL | 1024x1024 | Nvidia 20xx+ | 8GB | 16GB | |
| FLUX.1 | 1024x1024 | Nvidia 20xx+ | 10GB | 32GB | Download totals: NF4 ~12GB, int8 ~18GB, bf16 ~33GB. The int8 transformer is 11.5GB resident on any supported GPU, so it streams on a 10GB card |
| FLUX.2 Klein 4B | 1024x1024 | Nvidia 30xx+ | 12GB | 16GB | FP8 works with 8GB+; Diffusers + encoder |
Expand Down Expand Up @@ -114,6 +115,8 @@ Go to the [CUDA Toolkit Downloads] and carefully follow the instructions for you

Confirm that `nvidia-smi` displays driver and CUDA versions after installation.

Invoke's CUDA build (PyTorch 2.13 with CUDA 13.0) needs an Nvidia driver from the R580 series or newer. See [Hardware](#hardware) for the supported GPUs.

#### Linux - via Nvidia Container Runtime

An alternative to installing CUDA locally is to use the [Nvidia Container Runtime] to run the application in a container.
Expand Down
4 changes: 4 additions & 0 deletions docs/src/content/docs/troubleshooting/faq.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -59,6 +59,10 @@ Firefox does not allow Invoke to directly access the clipboard by default. As a

Most example images with prompts that you'll find on the internet have been generated using different software, so you can't expect to get identical results. In order to reproduce an image, you need to replicate the exact settings and processing steps, including (but not limited to) the model, the positive and negative prompts, the seed, the sampler, the exact image size, any upscaling steps, etc.

## Same seed, different image after an update

Invoke now draws its seeded noise in `float32`, because PyTorch 2.13 changed its half-precision random numbers. For SD1.5, SDXL, SD3, FLUX.1, FLUX.2, CogView4 and Z-Image's seed variance enhancer, a seed therefore gives a different image than it did before this update. On Windows and Linux the older images cannot be reproduced exactly. On macOS, set `noise_dtype: float16` in the [configuration file][configuration docs] to keep the images your seeds gave before.

## Invalid configuration file

Everything seems to install ok, you get a `ValidationError` when starting up the app.
Expand Down
14 changes: 14 additions & 0 deletions docs/src/generated/settings.json
Original file line number Diff line number Diff line change
Expand Up @@ -652,6 +652,20 @@
"type": "<class 'bool'>",
"validation": {}
},
{
"category": "GENERATION",
"default": "float32",
"description": "The dtype seeded noise is drawn in for SD1.5/SDXL, SD3, FLUX.1, FLUX.2, CogView4 and the Z-Image seed variance enhancer. `float32` draws the same noise on every platform. `float16` is the half-precision draw of earlier versions: on macOS it keeps the images your seeds gave before this update; on Windows and Linux both values give (nearly) the same images, and images made before the update cannot be reproduced there.",
"env_var": "INVOKEAI_NOISE_DTYPE",
"literal_values": [
"float32",
"float16"
],
"name": "noise_dtype",
"required": false,
"type": "typing.Literal['float32', 'float16']",
"validation": {}
},
{
"category": "GENERATION",
"default": false,
Expand Down
2 changes: 1 addition & 1 deletion invokeai/app/invocations/cogview4/cogview4_denoise.py
Original file line number Diff line number Diff line change
Expand Up @@ -140,7 +140,7 @@ def _get_noise(
) -> torch.Tensor:
# We always generate noise on the same device and dtype then cast to ensure consistency across devices/dtypes.
rand_device = "cpu"
rand_dtype = torch.float16
rand_dtype = TorchDevice.choose_noise_dtype(torch.float16)

return torch.randn(
batch_size,
Expand Down
88 changes: 13 additions & 75 deletions invokeai/app/invocations/latent_noise.py
Original file line number Diff line number Diff line change
Expand Up @@ -58,79 +58,17 @@ def generate_noise_tensor(
dtype: torch.dtype,
use_cpu: bool = True,
) -> torch.Tensor:
validate_noise_dimensions(noise_type, width, height)
shape = get_expected_noise_shape(noise_type, width, height)
rand_device = "cpu" if use_cpu else device.type
rand_dtype = TorchDevice.choose_torch_dtype(device=device)

if noise_type == "SD":
return torch.randn(
1,
4,
height // LATENT_SCALE_FACTOR,
width // LATENT_SCALE_FACTOR,
dtype=rand_dtype,
device=rand_device,
generator=torch.Generator(device=rand_device).manual_seed(seed),
).to("cpu")
if noise_type == "FLUX":
return torch.randn(
1,
16,
height // LATENT_SCALE_FACTOR,
width // LATENT_SCALE_FACTOR,
device=rand_device,
dtype=rand_dtype,
generator=torch.Generator(device=rand_device).manual_seed(seed),
).to("cpu")
if noise_type == "FLUX.2":
return torch.randn(
1,
32,
height // LATENT_SCALE_FACTOR,
width // LATENT_SCALE_FACTOR,
device=rand_device,
dtype=rand_dtype,
generator=torch.Generator(device=rand_device).manual_seed(seed),
).to("cpu")
if noise_type == "SD3":
return torch.randn(
1,
16,
height // LATENT_SCALE_FACTOR,
width // LATENT_SCALE_FACTOR,
device=rand_device,
dtype=rand_dtype,
generator=torch.Generator(device=rand_device).manual_seed(seed),
).to("cpu")
if noise_type == "CogView4":
return torch.randn(
1,
16,
height // LATENT_SCALE_FACTOR,
width // LATENT_SCALE_FACTOR,
device=rand_device,
dtype=rand_dtype,
generator=torch.Generator(device=rand_device).manual_seed(seed),
).to("cpu")
if noise_type == "Z-Image":
return torch.randn(
1,
16,
height // LATENT_SCALE_FACTOR,
width // LATENT_SCALE_FACTOR,
device=rand_device,
dtype=torch.float32,
generator=torch.Generator(device=rand_device).manual_seed(seed),
).to("cpu")
if noise_type == "Anima":
return torch.randn(
1,
16,
1,
height // LATENT_SCALE_FACTOR,
width // LATENT_SCALE_FACTOR,
device=rand_device,
dtype=torch.float32,
generator=torch.Generator(device=rand_device).manual_seed(seed),
).to("cpu")
raise ValueError(f"Unsupported noise type: {noise_type}")
if noise_type in ("Z-Image", "Anima"):
output_dtype = rand_dtype = torch.float32
else:
# The tensor keeps the device precision it always had; only the dtype it is drawn in follows `noise_dtype`.
output_dtype = dtype
rand_dtype = TorchDevice.choose_noise_dtype(dtype)
return torch.randn(
*shape,
device=rand_device,
dtype=rand_dtype,
generator=torch.Generator(device=rand_device).manual_seed(seed),
).to(device="cpu", dtype=output_dtype)
2 changes: 1 addition & 1 deletion invokeai/app/invocations/sd3/sd3_denoise.py
Original file line number Diff line number Diff line change
Expand Up @@ -159,7 +159,7 @@ def _get_noise(
) -> torch.Tensor:
# We always generate noise on the same device and dtype then cast to ensure consistency across devices/dtypes.
rand_device = "cpu"
rand_dtype = torch.float16
rand_dtype = TorchDevice.choose_noise_dtype(torch.float16)

return torch.randn(
num_samples,
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@
ConditioningFieldData,
ZImageConditioningInfo,
)
from invokeai.backend.util.devices import TorchDevice


@invocation(
Expand Down Expand Up @@ -82,8 +83,11 @@ def invoke(self, context: InvocationContext) -> ZImageConditioningOutput:
generator = torch.Generator(device=prompt_embeds.device)
generator.manual_seed(self.seed)
noise = torch.rand(
prompt_embeds.shape, generator=generator, device=prompt_embeds.device, dtype=prompt_embeds.dtype
)
prompt_embeds.shape,
generator=generator,
device=prompt_embeds.device,
dtype=TorchDevice.choose_noise_dtype(prompt_embeds.dtype),
).to(prompt_embeds.dtype)
noise = noise * 2 - 1 # Scale to [-1, 1)
noise = noise * actual_strength

Expand Down
2 changes: 2 additions & 0 deletions invokeai/app/run_app.py
Original file line number Diff line number Diff line change
Expand Up @@ -87,6 +87,7 @@ def run_app() -> None:
# Import from startup_utils here to avoid importing torch before configure_torch_cuda_allocator() is called.
from invokeai.app.util.startup_utils import (
apply_monkeypatches,
check_cuda_build_compatibility,
check_cudnn,
enable_dev_reload,
find_open_port,
Expand All @@ -105,6 +106,7 @@ def run_app() -> None:
apply_monkeypatches()
register_mime_types()
check_cudnn(logger)
check_cuda_build_compatibility(logger)
# Fail here rather than inside a generation: the value is read per generation, so a typo would
# otherwise surface as a failed queue item minutes after the server came up.
resolve_krea2_sdpa_backends()
Expand Down
3 changes: 3 additions & 0 deletions invokeai/app/services/config/config_default.py
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@
LEGACY_INIT_FILE = Path("invokeai.init")
PRECISION = Literal["auto", "float16", "bfloat16", "float32"]
ATTENTION_TYPE = Literal["auto", "normal", "xformers", "sliced", "torch-sdp"]
NOISE_DTYPE = Literal["float32", "float16"]
ATTENTION_SLICE_SIZE = Literal["auto", "balanced", "max", 1, 2, 3, 4, 5, 6, 7, 8]
LOG_FORMAT = Literal["plain", "color", "syslog", "legacy"]
LOG_LEVEL = Literal["debug", "info", "warning", "error", "critical"]
Expand Down Expand Up @@ -120,6 +121,7 @@ class InvokeAIAppConfig(BaseSettings):
device: Preferred execution device. `auto` will choose the device depending on the hardware platform and the installed torch capabilities.<br>Valid values: `auto`, `cpu`, `cuda`, `mps`, `xpu`, `cuda:N`, `xpu:N` (where N is a device number)
precision: Floating point precision. `float16` will consume half the memory of `float32` but produce slightly lower-quality images. The `auto` setting will guess the proper precision based on your video card and operating system.<br>Valid values: `auto`, `float16`, `bfloat16`, `float32`
sequential_guidance: Whether to calculate guidance in serial instead of in parallel, lowering memory requirements.
noise_dtype: The dtype seeded noise is drawn in for SD1.5/SDXL, SD3, FLUX.1, FLUX.2, CogView4 and the Z-Image seed variance enhancer. `float32` draws the same noise on every platform. `float16` is the half-precision draw of earlier versions: on macOS it keeps the images your seeds gave before this update; on Windows and Linux both values give (nearly) the same images, and images made before the update cannot be reproduced there.<br>Valid values: `float32`, `float16`
wan_memory_optimization: Enable experimental Wan memory optimizations at the cost of slower generation.
pid_memory_optimization: Enable experimental PiD decode memory optimizations. Roughly halves the peak activation memory of a PiD decode; in exchange the decoded image changes slightly, because neither the chunked pixel pathway nor the float32 sampler intermediates are bit-exact with the default path.
attention_type: Attention type.<br>Valid values: `auto`, `normal`, `xformers`, `sliced`, `torch-sdp`
Expand Down Expand Up @@ -245,6 +247,7 @@ class InvokeAIAppConfig(BaseSettings):

# GENERATION
sequential_guidance: bool = Field(default=False, description="Whether to calculate guidance in serial instead of in parallel, lowering memory requirements.")
noise_dtype: NOISE_DTYPE = Field(default="float32", description="The dtype seeded noise is drawn in for SD1.5/SDXL, SD3, FLUX.1, FLUX.2, CogView4 and the Z-Image seed variance enhancer. `float32` draws the same noise on every platform. `float16` is the half-precision draw of earlier versions: on macOS it keeps the images your seeds gave before this update; on Windows and Linux both values give (nearly) the same images, and images made before the update cannot be reproduced there.")
wan_memory_optimization: bool = Field(default=False, description="Enable experimental Wan memory optimizations at the cost of slower generation.")
pid_memory_optimization: bool = Field(default=False, description="Enable experimental PiD decode memory optimizations. Roughly halves the peak activation memory of a PiD decode; in exchange the decoded image changes slightly, because neither the chunked pixel pathway nor the float32 sampler intermediates are bit-exact with the default path.")
attention_type: ATTENTION_TYPE = Field(default="auto", description="Attention type.")
Expand Down
Loading
Loading