Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
---
title: Model Format Support
description: Which file formats and quantizations InvokeAI loads for every model family, text encoder, VAE and adapter.
lastUpdated: 2026-10-10
lastUpdated: 2026-10-11
sidebar:
order: 3
---
Expand Down Expand Up @@ -54,7 +54,7 @@ Ministral 3B encoder. Others are recognized by their layer names and fail, or mi

| Encoder | Used by | Folder | Single file (bf16) | fp8 scaled | int8 | nvfp4 | GGUF | SDNQ |
|---|---|---|---|---|---|---|---|---|
| T5-XXL | FLUX.1, SD 3 | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ (and bitsandbytes int8) |
| T5-XXL | FLUX.1, SD 3 | ✓ | ✓ | ✓ | R | R | ✓ | ✓ (and bitsandbytes int8) |
| CLIP-L / CLIP-G | SD, SDXL, FLUX.1, SD 3 | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | CLIP-L inside a FLUX.1 SDNQ pipeline |
| Qwen3 0.6B / 4B / 8B | Z-Image, FLUX.2 Klein, Anima | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Qwen3-VL 4B / 8B | Krea-2, Ideogram 4 | ✓ | ✓ | ✓ | ✓ | ✗ | ✓ | ✗ |
Expand All @@ -67,6 +67,12 @@ Ministral 3B encoder. Others are recognized by their layer names and fail, or mi
| Gemma-4 12B | LTX-2 | ✓ (bf16 or int8) | ✗ | L | in the folder | L | ✗ | ✗ |
| UMT5 | Wan 2.2 | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |

The single-file T5-XXL builds are ComfyUI's `t5xxl_fp16`, `t5xxl_fp8_e4m3fn` and `t5xxl_fp8_e4m3fn_scaled`. The
fp8 builds keep their weights in fp8 on a GPU that can hold them, about 4.9 GB instead of about 9.5 GB. Prefer
`t5xxl_fp8_e4m3fn_scaled` over `t5xxl_fp8_e4m3fn`: the unscaled build stores every tensor in fp8, its layer norms
included, and its prompt embeddings stray measurably further from the full-precision encoder. Wan's UMT5 file looks
similar but is not a T5-XXL and is not recognized as one.

The Qwen3-VL GGUF encoder is the language model only; a separate `mmproj` vision file is refused, and split
multi-part GGUFs have to be joined before installing. The Gemma-4 encoder installs only as a folder: its
`config.json` and tokenizer files next to exactly one weight file, bf16 or int8. A standalone `.safetensors` file is
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -71,6 +71,9 @@ components. Installing a starter model installs them as dependencies:
| T5 encoder | bfloat16 (~9.5 GB), bitsandbytes int8 (~5 GB), GGUF Q6_K (~3.9 GB, near-lossless), GGUF Q3_K_S (~2.1 GB, lower quality) |
| CLIP embed | `clip-vit-large-patch14` (~250 MB) |

You can also install the T5 encoder ComfyUI uses as a single file (`t5xxl_fp16`, or `t5xxl_fp8_e4m3fn_scaled` for
about half the VRAM) and select it here; see [Model Format Support](/users-guide/models/model-format-support/#text-encoders).

Select them in the **Components** section next to the model. The SDNQ pipeline carries its own encoders
and VAE, so its component pickers stay empty.

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,9 @@ you can replace any of them with a model you have installed separately:
- **CLIP G**
- **VAE**: only SD3 VAEs are offered

The T5 encoder can also be one of the single files ComfyUI uses: `t5xxl_fp16`, or `t5xxl_fp8_e4m3fn_scaled` for
about half the VRAM.

## Generation settings

When you select an SD3 model, InvokeAI applies these defaults, which come from the SD3.5 Medium
Expand Down
3 changes: 2 additions & 1 deletion invokeai/app/invocations/text_encoder/flux_text_encoder.py
Original file line number Diff line number Diff line change
Expand Up @@ -99,7 +99,8 @@ def _t5_encode(self, context: InvocationContext) -> torch.Tensor:
# Determine if the model is quantized.
# If the model is quantized, then we need to apply the LoRA weights as sidecar layers. This results in
# slower inference than direct patching, but is agnostic to the quantization format.
if t5_encoder_config.format in [ModelFormat.T5Encoder, ModelFormat.Diffusers]:
# A single-file fp8 build may keep its Linear weights in float8; the patcher sidecars those itself.
if t5_encoder_config.format in [ModelFormat.T5Encoder, ModelFormat.Diffusers, ModelFormat.Checkpoint]:
model_is_quantized = False
elif t5_encoder_config.format in [
ModelFormat.BnbQuantizedLlmInt8b,
Expand Down
2 changes: 2 additions & 0 deletions invokeai/backend/model_manager/configs/factory.py
Original file line number Diff line number Diff line change
Expand Up @@ -160,6 +160,7 @@
)
from invokeai.backend.model_manager.configs.t5_encoder import (
T5Encoder_BnBLLMint8_Config,
T5Encoder_Checkpoint_Config,
T5Encoder_GGUF_Config,
T5Encoder_SDNQ_Config,
T5Encoder_T5Encoder_Config,
Expand Down Expand Up @@ -527,6 +528,7 @@ def has_model_export(module: Any, name: Any, expected_bases: tuple[type, ...]) -
Annotated[T5Encoder_BnBLLMint8_Config, T5Encoder_BnBLLMint8_Config.get_tag()],
Annotated[T5Encoder_SDNQ_Config, T5Encoder_SDNQ_Config.get_tag()],
Annotated[T5Encoder_GGUF_Config, T5Encoder_GGUF_Config.get_tag()],
Annotated[T5Encoder_Checkpoint_Config, T5Encoder_Checkpoint_Config.get_tag()],
# Qwen3-VL Encoder (Qwen3-VL multimodal encoder for Krea-2) - checked BEFORE the text-only Qwen3
# encoder so single-file VL checkpoints (which also carry generic model.layers.* keys) are not
# misclassified as the Z-Image Qwen3 encoder. The VL probe requires the visual tower.
Expand Down
40 changes: 40 additions & 0 deletions invokeai/backend/model_manager/configs/identification_utils.py
Original file line number Diff line number Diff line change
Expand Up @@ -184,6 +184,46 @@ def state_dict_has_any_keys_ending_with(state_dict: dict[str | int, Any], suffix
return any(any(key.endswith(suffix) for suffix in _suffixes) for key in state_dict.keys() if isinstance(key, str))


def raise_if_quantized_beyond_fp8(
mod: ModelOnDisk, state_dict: dict[str | int, Any], what: str, supported_note: str
) -> None:
"""Refuse at install a single file quantized in a scheme a loader that only reads (scaled) fp8 cannot decode.

ComfyUI redistributes text encoders as int8 (convrot), nvfp4 and mxfp8 builds next to the fp16 and fp8
ones. Recognising such a file and registering it would leave a model that fails at its first use, so it is
an `InvalidMatchError`, worded for the caller's model. Identification has every dtype on the meta device but
no bytes, so the schemes are read from the weights' dtypes, nvfp4's second scale, and the header's
declarations (a header parse plus one seek per `.comfy_quant` marker).
"""
import torch

if state_dict_has_any_keys_ending_with(state_dict, ".weight_scale_2"):
raise InvalidMatchError(f"This {what} is nvfp4-quantized, which is not supported. {supported_note}")

integer_weights = sorted(
key
for key, value in state_dict.items()
if isinstance(key, str)
and key.endswith(".weight")
and getattr(value, "dtype", None) in (torch.int8, torch.uint8)
)
if integer_weights:
raise InvalidMatchError(
f"This {what} carries {len(integer_weights)} integer-quantized weight(s) (e.g. '{integer_weights[0]}'), "
f"which are not supported. {supported_note}"
)

# Lazy: the loader package imports the config factory, which imports the configs calling this.
from invokeai.backend.model_manager.load.model_loaders._single_file_guards import (
reject_formats_declared_in_the_header,
)

try:
reject_formats_declared_in_the_header(mod.path, what, None, {"float8_e4m3fn"}, supported_note)
except ValueError as e:
raise InvalidMatchError(str(e)) from e


def common_config_paths(path: Path) -> set[Path]:
"""Returns common config file paths for models stored in directories."""
return {path / "config.json", path / "model_index.json"}
Expand Down
66 changes: 66 additions & 0 deletions invokeai/backend/model_manager/configs/t5_encoder.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,18 +6,24 @@

from invokeai.backend.model_manager.configs.base import Checkpoint_Config_Base, Config_Base
from invokeai.backend.model_manager.configs.identification_utils import (
InvalidMatchError,
NotAMatchError,
raise_for_class_name,
raise_for_override_fields,
raise_if_not_dir,
raise_if_not_file,
raise_if_quantized_beyond_fp8,
state_dict_has_any_keys_ending_with,
state_dict_has_any_keys_starting_with,
)
from invokeai.backend.model_manager.model_on_disk import ModelOnDisk
from invokeai.backend.model_manager.taxonomy import BaseModelType, ModelFormat, ModelType
from invokeai.backend.quantization.gguf.ggml_tensor import GGMLTensor
from invokeai.backend.quantization.sdnq.detection import folder_has_sdnq_keys
from invokeai.backend.t5.t5_tokenizer import T5_VOCAB_SIZE

# The width of T5 v1.1 XXL, the encoder FLUX.1 and SD 3 condition on.
T5_XXL_D_MODEL = 4096


def _safetensors_dir_has_sdnq_keys(directory) -> bool:
Expand Down Expand Up @@ -231,3 +237,63 @@ def raise_if_doesnt_look_like_gguf_quantized(cls, mod: ModelOnDisk) -> None:
has_ggml = any(isinstance(v, GGMLTensor) for v in mod.load_state_dict().values())
if not has_ggml:
raise NotAMatchError("state dict does not look like GGUF quantized")


class T5Encoder_Checkpoint_Config(Checkpoint_Config_Base, Config_Base):
"""Configuration for a T5 encoder in a single safetensors file with transformers key names.

This is how ComfyUI distributes T5-XXL (``t5xxl_fp16``, ``t5xxl_fp8_e4m3fn`` and
``t5xxl_fp8_e4m3fn_scaled``): the ``T5EncoderModel`` state dict as is, with no config and no tokenizer.
"""

base: Literal[BaseModelType.Any] = Field(default=BaseModelType.Any)
type: Literal[ModelType.T5Encoder] = Field(default=ModelType.T5Encoder)
format: Literal[ModelFormat.Checkpoint] = Field(default=ModelFormat.Checkpoint)
cpu_only: bool | None = Field(default=None, description="Whether this model should run on CPU only")

@classmethod
def from_model_on_disk(cls, mod: ModelOnDisk, override_fields: dict[str, Any]) -> Self:
raise_if_not_file(mod)

raise_for_override_fields(cls, override_fields)

if mod.path.suffix != ".safetensors":
raise NotAMatchError("not a safetensors file")

state_dict = mod.load_state_dict()
cls._raise_if_not_t5_encoder(state_dict)
cls._raise_if_not_t5_xxl(state_dict)
raise_if_quantized_beyond_fp8(
mod, state_dict, "T5 encoder", "Use the fp16 or fp8 (scaled) build of T5-XXL instead."
)

return cls(**override_fields)

@classmethod
def _raise_if_not_t5_encoder(cls, state_dict: dict[str | int, Any]) -> None:
if "encoder.block.0.layer.0.SelfAttention.q.weight" not in state_dict or not (
"shared.weight" in state_dict or "encoder.embed_tokens.weight" in state_dict
):
raise NotAMatchError("state dict does not look like a transformers T5 encoder")
# The first shard of a sharded export has block 0 and the embedding too; only a whole encoder ends in its norm.
if "encoder.final_layer_norm.weight" not in state_dict:
raise NotAMatchError("state dict has no encoder.final_layer_norm; not a complete T5 encoder")
if state_dict_has_any_keys_starting_with(state_dict, "decoder."):
raise NotAMatchError("state dict carries a T5 decoder; only encoder-only files are supported")
# UMT5 (Wan's text encoder) shares T5's key names but gives every block its own position bias,
# where T5 has one in the first block only.
if "encoder.block.1.layer.0.SelfAttention.relative_attention_bias.weight" in state_dict:
raise NotAMatchError("state dict looks like UMT5 (a relative attention bias in every block)")
# T5 v1.0 has a single, ungated input projection; v1.1 (what FLUX.1 and SD 3 use) is gated.
if "encoder.block.0.layer.1.DenseReluDense.wi_0.weight" not in state_dict:
raise NotAMatchError("state dict does not have T5 v1.1's gated feed-forward")

@classmethod
def _raise_if_not_t5_xxl(cls, state_dict: dict[str | int, Any]) -> None:
embedding = state_dict.get("shared.weight", state_dict.get("encoder.embed_tokens.weight"))
vocab_size, d_model = (int(x) for x in getattr(embedding, "shape", (0, 0)))
if d_model != T5_XXL_D_MODEL or vocab_size != T5_VOCAB_SIZE:
raise InvalidMatchError(
f"this T5 encoder has width {d_model} and vocabulary {vocab_size}, but FLUX.1 and SD 3 need "
f"T5 v1.1 XXL (width {T5_XXL_D_MODEL}, vocabulary {T5_VOCAB_SIZE})."
)
Loading
Loading