Skip to content

feat(model-manager): load ComfyUI single-file T5-XXL encoders - #9799

Open
Pfannkuchensack wants to merge 2 commits into
invoke-ai:mainfrom
Pfannkuchensack:feat/single-file-text-encoders
Open

Pfannkuchensack wants to merge 2 commits into
invoke-ai:mainfrom
Pfannkuchensack:feat/single-file-text-encoders

Conversation

@Pfannkuchensack

@Pfannkuchensack Pfannkuchensack commented Oct 11, 2026 •

Copy link
Copy Markdown
Member

Summary

ComfyUI distributes T5-XXL as bare safetensors files: t5xxl_fp16, t5xxl_fp8_e4m3fn and t5xxl_fp8_e4m3fn_scaled. InvokeAI loaded T5 only as a Diffusers folder, GGUF, SDNQ or bnb int8, so these files installed as unknown models. They now install as T5 encoders and can be selected for FLUX.1 and SD 3.

Identification

A new config, T5Encoder_Checkpoint_Config (tag t5_encoder.checkpoint.any), reads only the safetensors header. It requires:

  • transformers key names
  • an encoder-only layout with a final norm
  • T5 v1.1's gated feed-forward
  • the XXL width and vocabulary

Files that are not recognized:

  • UMT5, Wan's encoder, which has a position bias in every block
  • encoder-decoder checkpoints
  • T5 v1.0
  • the first shard of a sharded export

Files that are refused at install, with the reason shown:

  • narrower T5 variants
  • int8, nvfp4 and header-declared formats other than fp8 (mxfp8, for example)

Loading

A new loader, T5EncoderSingleFileLoader, works like this:

  • Tokenizer: the bundled T5 tokenizer, so nothing is fetched from the Hub.
  • Architecture: built from tensor shapes.
  • Token embedding: the copy ComfyUI writes twice is dropped and tied.
  • fp8 builds on a device that can store float8: the Linear weights stay in float8 with their scales, at 4.56 GiB resident instead of 8.87 GiB. They are dequantized per forward and never go through the fp8 matmul, even with fp8_compute on. Quantizing T5's activations measurably hurts (see QA).
  • Everywhere else: scales are folded and weights are cast to the bf16-safe dtype.
  • RAM: reserved with reserve_for_load before anything is widened. The tokenizer submodel is sized 0, so loading it no longer evicts models to make room for the whole file.

Shared T5 code

infer_t5_encoder_config and make_t5_feed_forward_safe move from the GGUF T5 loader into invokeai/backend/t5/t5_encoder.py. The feed-forward workaround now also keeps activations out of a float8 wo. GGUF behaviour is unchanged.

FLUX text encoder

The FLUX text encoder accepts the new format. LayerPatcher already routes LoRA patches on float8 layers through sidecars.

Related Issues / Discussions

Part of #9736 (the T5 half). CLIP-L / CLIP-G single files are #9800.

QA Instructions

Automated (Windows, CPU):

  • uv run --no-sync pytest tests/model_identification tests/backend/architectures/test_variants.py tests/backend/model_manager/configs tests/backend/t5 tests/backend/model_manager/load tests/app/invocations -n 4: 4335 passed, 158 skipped, 6 xfailed.
  • New tests:
    • Identification through the whole config factory, from header-only files written from the captured ComfyUI headers. Covers fp16, raw fp8, scaled fp8, UMT5, decoder, first shard, narrow T5, int8, nvfp4 and mxfp8, plus a round trip of the stored record to its loader.
    • Loader tests with a tiny T5 written in each ComfyUI layout. Covers folded vs kept fp8, the tie, the scale and full-precision matmul.
  • ruff check / ruff format --check clean. make frontend-openapi frontend-typegen regenerated; pnpm -C invokeai/frontend/api lint test passes.

Manual (RTX 4090, CUDA, Windows):

  • Installs: all three files install as T5 Encoder / Checkpoint. They show 4.67 GB VRAM for the fp8 builds in the model cache log (kept 168 weight(s) in fp8).

  • Prompt embeddings: compared against the bf16 Diffusers T5 folder. Three prompts at 512 tokens; cosine per valid token:

    File mean worst token
    t5xxl_fp16 0.9997 0.991
    t5xxl_fp8_e4m3fn_scaled 0.996 0.874
    t5xxl_fp8_e4m3fn 0.934 -0.23

    The raw file stores its norms in fp8 too, so this loss is the file's own; it is documented. The bundled tokenizer gives the same ids as the folder's.

  • fp8 matmul on the scaled build: worst/mean token cosine 0.459 / 0.984, against 0.874 / 0.996 with storage only. This is why kept T5 weights never take the fp8 matmul. With fp8_compute on, the result is now identical to storage only.

  • FLUX.1 dev (int8), 768², 20 steps, same seed, each run twice without the invocation cache: reproducible. PSNR vs the folder encoder: fp16 28.1 dB, scaled 20.2 dB, raw 19.1 dB. All images are coherent.

  • SD3.5 Medium, 1024², 28 steps, same seed: reproducible. PSNR vs the pipeline's own fp32 T5: fp16 31.0 dB, scaled 26.7 dB. The bf16 T5 folder reaches 26.3 dB.

  • Encode working memory: +0.57 GiB in every case. Dequantizing per forward needs no extra estimate.

Review

Material findings resolved:

  • With fp8_compute on, kept T5 Linears ran the fp8 matmul; it was measured and is now disabled for this loader.
  • The first shard of a sharded export installed as T5; identification now requires the final norm.
  • A file carrying only encoder.embed_tokens with a scale would have lost the scale; the key and any scale are now renamed together.
  • An unreachable meta-parameter sweep was removed.
  • A stored-record round trip and loader-registry test was added.

Remaining limitations:

  • Only CUDA was exercised on hardware. CPU loads fold and are unit-tested. ROCm, MPS and XPU were not run.
  • The raw fp8 build is inherently lossy; the docs recommend the scaled build.

Compatibility / Rollout

  • Additive: a new config class and tag, with openapi.json / schema.ts regenerated.
  • No migration: no existing records carry the new tag.
  • The GGUF T5 loader keeps its behaviour; the helpers it used moved without change.
  • webv2 needs no change, since its T5 pickers filter by type only.

Checklist

  • The PR has a short but descriptive title, suitable for a changelog
  • Meaningful regression coverage added / updated where needed; obsolete tests/code removed
  • Persisted-state and API changes include required migrations / compatibility validation
  • Relevant performance/efficiency opportunities considered; material claims have evidence
  • Material review findings resolved and relevant checks rerun
  • Documentation added / updated (if applicable)
  • Updated What's New copy (if doing a release after this PR)

Identify t5xxl_fp16 / t5xxl_fp8_e4m3fn(_scaled) as T5 encoder checkpoints; refuse UMT5, non-XXL and int8/nvfp4/mxfp8 builds at install.
fp8 builds keep their Linear weights in fp8 on capable GPUs (~4.9 instead of ~9.5 GB), always dequantized per forward.
Share T5 config inference and the feed-forward dtype fix with the GGUF loader.
@github-actions github-actions Bot added python PRs that change python files invocations PRs that change invocations backend PRs that change backend files python-tests PRs that change python tests docs PRs that change docs labels Oct 11, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backend PRs that change backend files docs PRs that change docs invocations PRs that change invocations python PRs that change python files python-tests PRs that change python tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant