Skip to content

feat(model-manager): load ComfyUI single-file CLIP-L/CLIP-G encoders - #9800

Draft
Pfannkuchensack wants to merge 1 commit into
invoke-ai:mainfrom
Pfannkuchensack:feat/single-file-clip-encoders
Draft

Pfannkuchensack wants to merge 1 commit into
invoke-ai:mainfrom
Pfannkuchensack:feat/single-file-clip-encoders

Conversation

@Pfannkuchensack

Copy link
Copy Markdown
Member

Summary

ComfyUI distributes the CLIP text towers FLUX.1 and SD 3 use as bare safetensors files, clip_l and clip_g. InvokeAI loaded CLIP only as a Diffusers folder, so these installed as unknown models. They now install as CLIP Embed models (variants L and G). Along the way, two existing SD 3 bugs with separately selected CLIP encoders are fixed.

Identification

New configs: CLIPEmbed_Checkpoint_L_Config / _G_Config (tags clip_embed.checkpoint.any.large / .gigantic). The variant comes from width and depth, matched against vendored configs. The record discriminator now includes the variant for checkpoint CLIP records as it already did for Diffusers ones.

Files that are not recognized:

  • full CLIP models with a vision tower
  • other text-tower sizes

Files that are refused at install, with the reason shown:

  • Long-CLIP (a longer position table)
  • quantized weights
  • a CLIP-G without text_projection, which SD 3 needs

Loading

A new loader, CLIPSingleFileLoader, works like this:

  • CLIP-L is built as CLIPTextModel, which FLUX.1 requires. The text_model. prefix is stripped, because transformers ≥5.6 flattened that class.
  • CLIP-G is built as CLIPTextModelWithProjection.
  • Tokenizer: openai/clip-vit-large-patch14's tokenizer (MIT) and the two text configs are vendored in invokeai/backend/clip/, gzip-compressed like the Qwen tokenizers. It gives the same ids as the original.

SD 3 fixes

  • A separately selected CLIP-G was requested as TextEncoder2, which no CLIP loader provides, so every such generation failed. It is now requested as TextEncoder.
  • The pooled output was output[0]. That is last_hidden_state for a plain CLIPTextModel, the class a separately selected CLIP-L loads as. It is now text_embeds for CLIPTextModelWithProjection, else pooler_output. The second is exact for CLIP-L, because SD 3's CLIP-L projection is the identity (max abs diff 0.0 in the SD 3 Medium Diffusers weights).
  • A Diffusers CLIP-G folder now loads with its projection instead of dropping it. A folder without one is refused instead of getting a randomly initialised projection.

FLUX.1 takes CLIP-L only

Installing ComfyUI's text_encoders folder now brings in a CLIP-G as well, so the FLUX CLIP selection is restricted to CLIP-L:

  • the loader node field (ui_model_variant)
  • the webv2 Generate slot
  • the Upscale auto-pick and picker

A CLIP-G wired into a workflow gets a clear error instead of an AssertionError.

Related Issues / Discussions

Closes #9736 together with the T5 single-file PR (#9799).

QA Instructions

Automated (Windows, CPU):

  • uv run --no-sync pytest tests/model_identification tests/backend/architectures/test_variants.py tests/backend/model_manager/configs tests/backend/model_manager/load tests/app/invocations tests/backend/util/test_bundled_tokenizer.py -n 4: 4336 passed, 158 skipped, 6 xfailed.
  • New tests:
    • Identification through the whole config factory from header-only files written from the captured ComfyUI headers. Covers L, G, extras, vision tower, OpenCLIP-H size, Long-CLIP, a G without projection, quantized weights, and the stored-record round trip per variant.
    • Single-file loader with tiny models.
    • Diffusers loader class per variant, including the missing-projection refusal.
    • SD 3 pooled output with tiny real CLIPTextModel / CLIPTextModelWithProjection.
    • The SD 3 model loader's submodel choice.
  • pnpm -C invokeai/frontend/webv2 lint passes (format, oxlint, tsc, architecture).
  • pnpm -C invokeai/frontend/webv2 test: 10068 passed, 3 failed. The three are in untouched files (queue/ui/modelCacheStore.test.ts, workbench/image-map/*) and fail only under a German Windows locale, which formats 1,500 as 1.500.
  • New webv2 tests cover the FLUX Generate slot and the Upscale auto-pick/sync rejecting a CLIP-G; both were mutation-checked.
  • ruff clean. make frontend-openapi frontend-typegen regenerated; pnpm -C invokeai/frontend/api lint passes.

Manual (RTX 4090, CUDA, Windows):

  • SD3.5 Medium, 1024², 28 steps, same seed, each run twice without the invocation cache: separately selected clip_l, clip_g, or both produce bit-identical images to the pipeline's own encoders (PSNR ∞, max diff 0). The log confirms the single files loaded as CLIPTextModel / CLIPTextModelWithProjection. Before this PR the same selection failed for CLIP-G.
  • FLUX.1 dev (int8), 768², 20 steps: clip_l against the bf16 Diffusers CLIP-L folder gives 33.2 dB (fp16 vs bf16 weights). Reproducible.
  • webv2 in the browser:
    • The Model Manager lists clip_l / clip_g as Checkpoint with their variant.
    • The SD 3 CLIP L / CLIP G pickers offer the matching file.
    • A generation with both started from the UI succeeded.
    • The FLUX CLIP Embed picker lists only CLIP-L models.

Screenshots: attach Model Manager detail (clip_g), SD 3 components with clip_l/clip_g, FLUX CLIP picker.

Review

Material findings resolved:

  • Long-CLIP-L installed as CLIP-L and failed at generation; it is now refused at install.
  • FLUX.1 could be given a CLIP-G, which this PR made easy to install; it is now restricted to CLIP-L in backend and webv2.
  • A Diffusers CLIP-G without a projection would have been initialised randomly; it is now refused.
  • The SD 3 pooled-output change and the Diffusers CLIP-G class had no failing-on-regression tests; both are now added.

Remaining limitations:

  • The workflow editor still ignores ui_model_variant, as before. The FLUX text encoder now names the problem instead.
  • Only CUDA was exercised on hardware.

Compatibility / Rollout

  • Discriminator: it now includes variant for checkpoint CLIP records. No such records existed before, and Diffusers CLIP records parse exactly as before.
  • Generated contracts: openapi.json / schema.ts regenerated, additively (new config classes, ui_model_variant on the FLUX loader's CLIP field).
  • Packaging: the new vendored files are covered by a package-data entry for invokeai.backend.clip.
  • Stored webv2 selections: they carry the model's variant. An Upscale selection that lacks one is dropped and re-picked.

Checklist

  • The PR has a short but descriptive title, suitable for a changelog
  • Meaningful regression coverage added / updated where needed; obsolete tests/code removed
  • Persisted-state and API changes include required migrations / compatibility validation
  • Relevant performance/efficiency opportunities considered; material claims have evidence
  • Material review findings resolved and relevant checks rerun
  • Documentation added / updated (if applicable)
  • Updated What's New copy (if doing a release after this PR)

Identify clip_l / clip_g as CLIP Embed checkpoints with a bundled tokenizer and configs; refuse Long-CLIP, quantized and projection-less CLIP-G files.
Fix SD 3 with separately selected CLIP encoders (CLIP-G submodel, pooled output) and load Diffusers CLIP-G with its projection.
Restrict FLUX.1's CLIP selection to CLIP-L in the loader node and webv2.
@Pfannkuchensack

Copy link
Copy Markdown
Member Author
mm_clip_g sd3_result flux_clip_picker

@github-actions github-actions Bot added python PRs that change python files Root invocations PRs that change invocations backend PRs that change backend files frontend PRs that change frontend files python-tests PRs that change python tests docs PRs that change docs python-deps PRs that change python dependencies labels Oct 11, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backend PRs that change backend files docs PRs that change docs frontend PRs that change frontend files invocations PRs that change invocations python PRs that change python files python-deps PRs that change python dependencies python-tests PRs that change python tests Root

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[enhancement]: T5-XXL and CLIP text encoders as single files

1 participant