Repository navigation
feat(model-manager): load ComfyUI single-file T5-XXL encoders - #9799
Open
Pfannkuchensack wants to merge 2 commits into
Open
Pfannkuchensack wants to merge 2 commits into
Pfannkuchensack wants to merge 2 commits into
Conversation
Identify t5xxl_fp16 / t5xxl_fp8_e4m3fn(_scaled) as T5 encoder checkpoints; refuse UMT5, non-XXL and int8/nvfp4/mxfp8 builds at install. fp8 builds keep their Linear weights in fp8 on capable GPUs (~4.9 instead of ~9.5 GB), always dequantized per forward. Share T5 config inference and the feed-forward dtype fix with the GGUF loader.
6 of 7 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
ComfyUI distributes T5-XXL as bare safetensors files:
t5xxl_fp16,t5xxl_fp8_e4m3fnandt5xxl_fp8_e4m3fn_scaled. InvokeAI loaded T5 only as a Diffusers folder, GGUF, SDNQ or bnb int8, so these files installed as unknown models. They now install as T5 encoders and can be selected for FLUX.1 and SD 3.Identification
A new config,
T5Encoder_Checkpoint_Config(tagt5_encoder.checkpoint.any), reads only the safetensors header. It requires:Files that are not recognized:
Files that are refused at install, with the reason shown:
Loading
A new loader,
T5EncoderSingleFileLoader, works like this:fp8_computeon. Quantizing T5's activations measurably hurts (see QA).reserve_for_loadbefore anything is widened. The tokenizer submodel is sized 0, so loading it no longer evicts models to make room for the whole file.Shared T5 code
infer_t5_encoder_configandmake_t5_feed_forward_safemove from the GGUF T5 loader intoinvokeai/backend/t5/t5_encoder.py. The feed-forward workaround now also keeps activations out of a float8wo. GGUF behaviour is unchanged.FLUX text encoder
The FLUX text encoder accepts the new format. LayerPatcher already routes LoRA patches on float8 layers through sidecars.
Related Issues / Discussions
Part of #9736 (the T5 half). CLIP-L / CLIP-G single files are #9800.
QA Instructions
Automated (Windows, CPU):
uv run --no-sync pytest tests/model_identification tests/backend/architectures/test_variants.py tests/backend/model_manager/configs tests/backend/t5 tests/backend/model_manager/load tests/app/invocations -n 4: 4335 passed, 158 skipped, 6 xfailed.ruff check/ruff format --checkclean.make frontend-openapi frontend-typegenregenerated;pnpm -C invokeai/frontend/api lint testpasses.Manual (RTX 4090, CUDA, Windows):
Installs: all three files install as T5 Encoder / Checkpoint. They show 4.67 GB VRAM for the fp8 builds in the model cache log (
kept 168 weight(s) in fp8).Prompt embeddings: compared against the bf16 Diffusers T5 folder. Three prompts at 512 tokens; cosine per valid token:
t5xxl_fp16t5xxl_fp8_e4m3fn_scaledt5xxl_fp8_e4m3fnThe raw file stores its norms in fp8 too, so this loss is the file's own; it is documented. The bundled tokenizer gives the same ids as the folder's.
fp8 matmul on the scaled build: worst/mean token cosine 0.459 / 0.984, against 0.874 / 0.996 with storage only. This is why kept T5 weights never take the fp8 matmul. With
fp8_computeon, the result is now identical to storage only.FLUX.1 dev (int8), 768², 20 steps, same seed, each run twice without the invocation cache: reproducible. PSNR vs the folder encoder: fp16 28.1 dB, scaled 20.2 dB, raw 19.1 dB. All images are coherent.
SD3.5 Medium, 1024², 28 steps, same seed: reproducible. PSNR vs the pipeline's own fp32 T5: fp16 31.0 dB, scaled 26.7 dB. The bf16 T5 folder reaches 26.3 dB.
Encode working memory: +0.57 GiB in every case. Dequantizing per forward needs no extra estimate.
Review
Material findings resolved:
fp8_computeon, kept T5 Linears ran the fp8 matmul; it was measured and is now disabled for this loader.encoder.embed_tokenswith a scale would have lost the scale; the key and any scale are now renamed together.Remaining limitations:
Compatibility / Rollout
openapi.json/schema.tsregenerated.Checklist
What's Newcopy (if doing a release after this PR)