Repository navigation
feat(model-manager): load ComfyUI single-file CLIP-L/CLIP-G encoders - #9800
Draft
Pfannkuchensack wants to merge 1 commit into
Draft
Pfannkuchensack wants to merge 1 commit into
Pfannkuchensack wants to merge 1 commit into
Conversation
Identify clip_l / clip_g as CLIP Embed checkpoints with a bundled tokenizer and configs; refuse Long-CLIP, quantized and projection-less CLIP-G files. Fix SD 3 with separately selected CLIP encoders (CLIP-G submodel, pooled output) and load Diffusers CLIP-G with its projection. Restrict FLUX.1's CLIP selection to CLIP-L in the loader node and webv2.
6 of 7 tasks
Member
Author
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Summary
ComfyUI distributes the CLIP text towers FLUX.1 and SD 3 use as bare safetensors files,
clip_landclip_g. InvokeAI loaded CLIP only as a Diffusers folder, so these installed as unknown models. They now install as CLIP Embed models (variants L and G). Along the way, two existing SD 3 bugs with separately selected CLIP encoders are fixed.Identification
New configs:
CLIPEmbed_Checkpoint_L_Config/_G_Config(tagsclip_embed.checkpoint.any.large/.gigantic). The variant comes from width and depth, matched against vendored configs. The record discriminator now includes the variant for checkpoint CLIP records as it already did for Diffusers ones.Files that are not recognized:
Files that are refused at install, with the reason shown:
text_projection, which SD 3 needsLoading
A new loader,
CLIPSingleFileLoader, works like this:CLIPTextModel, which FLUX.1 requires. Thetext_model.prefix is stripped, because transformers ≥5.6 flattened that class.CLIPTextModelWithProjection.openai/clip-vit-large-patch14's tokenizer (MIT) and the two text configs are vendored ininvokeai/backend/clip/, gzip-compressed like the Qwen tokenizers. It gives the same ids as the original.SD 3 fixes
TextEncoder2, which no CLIP loader provides, so every such generation failed. It is now requested asTextEncoder.output[0]. That islast_hidden_statefor a plainCLIPTextModel, the class a separately selected CLIP-L loads as. It is nowtext_embedsforCLIPTextModelWithProjection, elsepooler_output. The second is exact for CLIP-L, because SD 3's CLIP-L projection is the identity (max abs diff 0.0 in the SD 3 Medium Diffusers weights).FLUX.1 takes CLIP-L only
Installing ComfyUI's
text_encodersfolder now brings in a CLIP-G as well, so the FLUX CLIP selection is restricted to CLIP-L:ui_model_variant)A CLIP-G wired into a workflow gets a clear error instead of an
AssertionError.Related Issues / Discussions
Closes #9736 together with the T5 single-file PR (#9799).
QA Instructions
Automated (Windows, CPU):
uv run --no-sync pytest tests/model_identification tests/backend/architectures/test_variants.py tests/backend/model_manager/configs tests/backend/model_manager/load tests/app/invocations tests/backend/util/test_bundled_tokenizer.py -n 4: 4336 passed, 158 skipped, 6 xfailed.CLIPTextModel/CLIPTextModelWithProjection.pnpm -C invokeai/frontend/webv2 lintpasses (format, oxlint, tsc, architecture).pnpm -C invokeai/frontend/webv2 test: 10068 passed, 3 failed. The three are in untouched files (queue/ui/modelCacheStore.test.ts,workbench/image-map/*) and fail only under a German Windows locale, which formats1,500as1.500.ruffclean.make frontend-openapi frontend-typegenregenerated;pnpm -C invokeai/frontend/api lintpasses.Manual (RTX 4090, CUDA, Windows):
clip_l,clip_g, or both produce bit-identical images to the pipeline's own encoders (PSNR ∞, max diff 0). The log confirms the single files loaded asCLIPTextModel/CLIPTextModelWithProjection. Before this PR the same selection failed for CLIP-G.clip_lagainst the bf16 Diffusers CLIP-L folder gives 33.2 dB (fp16 vs bf16 weights). Reproducible.clip_l/clip_gas Checkpoint with their variant.Screenshots: attach Model Manager detail (clip_g), SD 3 components with clip_l/clip_g, FLUX CLIP picker.
Review
Material findings resolved:
Remaining limitations:
ui_model_variant, as before. The FLUX text encoder now names the problem instead.Compatibility / Rollout
variantfor checkpoint CLIP records. No such records existed before, and Diffusers CLIP records parse exactly as before.openapi.json/schema.tsregenerated, additively (new config classes,ui_model_varianton the FLUX loader's CLIP field).invokeai.backend.clip.Checklist
What's Newcopy (if doing a release after this PR)