Skip to content

feat(ernie-image): load GGUF transformers - #9725

Merged
lstein merged 6 commits into
invoke-ai:mainfrom
Pfannkuchensack:feat/ernie-image-gguf
Oct 10, 2026
Merged

lstein merged 6 commits into
invoke-ai:mainfrom
Pfannkuchensack:feat/ernie-image-gguf

Conversation

@Pfannkuchensack

Copy link
Copy Markdown
Member

Stacked on #9709, which is stacked on #9708. Until those merge, this PR also shows their commits. The change here is the last commit, feat(ernie-image): load GGUF transformers.

Summary

ERNIE-Image already loads Diffusers pipelines and Comfy-Org's safetensors single files. GGUF transformers did not register. This PR adds them.

  • Identification. Main_GGUF_ErnieImage_Config identifies ERNIE GGUFs by the same four-key fingerprint as the safetensors config. unsloth's GGUFs (ERNIE-Image and ERNIE-Image Turbo, the most downloaded) carry the diffusers key layout unchanged: all 409 tensors match ErnieImageTransformer2DModel, checked programmatically against the header. Their general.architecture names another model, though: wan for unsloth, flux for vantagewithai. Identification therefore rests on the keys alone, and the tests run the whole factory over both headers.
  • Loader. ErnieImageGGUFModel builds the diffusers transformer and loads the GGUF strictly.
    • Quantized Linears stay packed.
    • Everything the forward reads outside a Linear matmul is dequantized once at load, as is everything that is not actually quantized. That covers the RMSNorms, the F32 biases, the BF16 stem weights, and the 4-D BF16 patch convolution, which aten.convolution cannot run packed.
    • The helper doing this, unpack_ggml_at_load, moves from the Ideogram 4 loader (feat(ideogram4): load community GGUF transformers #9708) into quantization/gguf/loaders.py, since ERNIE is its second caller.
  • Node and denoise. The ERNIE loader node treats GGUF as a single file. The denoise passes the GGUF dequantization transient from fix(model-cache): reserve the GGUF dequantization transient in every node #9709. webv2 needs no change, because its ERNIE component slots already key on format !== 'diffusers'.
  • Starters. unsloth ERNIE-Image and ERNIE-Image Turbo at Q4_K_M and Q8_0, each pinned to a commit. Each installs the Ministral 3B encoder and the FLUX.2 VAE.

Measured on an RTX 4090 with ERNIE-Image Turbo, 1024², 8 steps, warm (3 runs after a warm-up):

Time per image GPU peak Transformer resident
Diffusers pipeline 11.5–11.7 s 23.6 GB 15.3 GB
GGUF Q4_K_M 11.0–11.1 s 17.8 GB 4.8 GB

The GGUF image is visually indistinguishable from the Diffusers one, with the text rendered cleanly. The record received the Turbo defaults (8 steps, CFG 1.0) from its name.

Related Issues / Discussions

Closes #9499. Ideogram 4 GGUF is #9708; ERNIE-Image GGUF is this PR.

QA Instructions

  1. Install the starter ERNIE-Image Turbo (GGUF, Q4_K_M). It also pulls the Ministral 3B encoder and the FLUX.2 VAE. The transformer installs as gguf_quantized with 8 steps and CFG 1.0.
  2. Select it in Generate. The Components section asks for the Mistral encoder and the VAE, as it does for the safetensors single file. Invoke and check that the image is clean.

Checks run:

  • uv run pytest tests/app/invocations tests/backend/quantization tests/backend/model_manager -n 4: 4585 passed.
  • New coverage:
    • Factory identification over real tiny GGUFs with wan and flux headers, including the Turbo/base defaults.
    • A safetensors file is not claimed as GGUF.
    • The forward matches a dense reference for Q8_0, Q5_1 and Q4_0, using the published BF16/F32 layout including the 4-D conv. It fails with a convolution dispatch error when the at-load unpacking is removed.
    • The absolute load reservation.
    • The denoise passes the transient.
    • The loader node accepts GGUF as a single file.
  • ruff check and ruff format --check. openapi.json and schema.ts regenerated; API package lint passes.

Review

Material findings, resolved:

  • The test fixture had the BF16 layout backwards: it made text_proj BF16 and the conv F32. The real files keep the 4-D conv in BF16 and quantize text_proj. The fixture now follows the published headers, so the 4-D BF16 unpack path is under test.
  • Starter copy called Q4_K_M "the smallest build". The repos also hold Q2_K/Q3_K. Sizes now use decimal GB consistently.

Out of scope:

  • GGUF Ministral 3B encoder. It is still refused by the Mistral GGUF config, which builds a MistralModel without Ministral's YaRN RoPE. The encoder stays the 7.7 GB safetensors file.
  • Wan's own at-load unwrap. _unwrap_unquantized_to_compute_dtype covers F32/F16 only. Switching it to unpack_ggml_at_load is a follow-up.

Compatibility / Rollout

No migration; this adds a config class and a loader. The generated openapi.json and schema.ts gain Main_GGUF_ErnieImage_Config. Existing Diffusers and safetensors installs behave as before.

Checklist

  • The PR has a short but descriptive title, suitable for a changelog
  • Meaningful regression coverage added / updated where needed; obsolete tests/code removed
  • Persisted-state and API changes include required migrations / compatibility validation
  • Relevant performance/efficiency opportunities considered; material claims have evidence
  • Material review findings resolved and relevant checks rerun
  • Documentation added / updated (if applicable)
  • Updated What's New copy (if doing a release after this PR)

@github-actions github-actions Bot added python PRs that change python files invocations PRs that change invocations backend PRs that change backend files python-tests PRs that change python tests docs PRs that change docs labels Oct 9, 2026
Add Main_GGUF_ErnieImage_Config and a GGUF loader; identification rests on the keys, since the published files name another architecture in their header.
Move unpack_ggml_at_load into the GGUF package as its second caller; the ERNIE node accepts GGUF as a single file and the denoise reserves its dequantization transient.
Add unsloth Q4_K_M and Q8_0 starters for ERNIE-Image and ERNIE-Image Turbo, pinned to their commits.

@lstein lstein left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. I checked this against the published files as well as the tests. I range-fetched the headers of the pinned unsloth Turbo Q4_K_M and Q8_0 GGUFs:

  • At the real geometry, the 409 tensors match ErnieImageTransformer2DModel key for key and shape for shape, so the strict load holds.
  • Nothing quantized sits outside an nn.Linear, so the per-module rule in unpack_ggml_at_load covers every tensor the forward reads outside a matmul.
  • Running the whole factory over the real key set, with the real general.architecture: wan metadata, yields exactly one match, Main_GGUF_ErnieImage_Config, with the right Turbo and base defaults.
  • All four starter URLs resolve at their pinned commits (5.02 GB and 8.69 GB).

Locally, the new tests and tests/backend/model_manager plus tests/backend/quantization pass (2257 passed). Two non-blocking nits are inline.

Comment thread invokeai/app/invocations/ernie_image/ernie_image_model_loader.py
Comment thread tests/backend/model_manager/load/test_ernie_image_gguf_loader.py Outdated
lstein and others added 3 commits October 10, 2026 14:52
Only a Diffusers folder is expected to carry scheduler/, so single files of any format no longer warn.
The GGUF forward test compared all-zero outputs (diffusers zero-inits adaLN and the final layer); it now fills
those in and also covers Q4_K, the starters' type, on a 256-wide geometry.
lstein added a commit that referenced this pull request Oct 10, 2026
## Summary

Which file format loads for which model was spread over the Model
Families prose and the individual family pages, and went stale whenever
a format landed (Ideogram 4 and ERNIE-Image GGUF most recently). This
adds one **Model Format Support** page under Users Guide → Models with
four tables:

- **Main models.** Diffusers, bf16 single file, fp8 scaled, int8, nvfp4,
MXFP8, GGUF, SDNQ and NF4 for all 14 families.
- **Text encoders.** T5, CLIP, Qwen3, Qwen3-VL (4B/8B and the truncated
32B), Qwen3.5, Qwen2.5-VL, Mistral, Ministral, Gemma-2, Gemma-4 and
UMT5.
- **VAEs.** Single file and Diffusers, plus which family uses which VAE.
- **LoRAs and adapters per family.** LoRA formats, ControlNet,
IP-/T2I-Adapter, Redux, PiD.

Cells are ✓ (loads), ✗ (not supported), **R** (refused at install with a
message) or – (no such build). Model Families links to the page from its
quantized-formats section. The new-model integration checklist asks to
update it.

The tables reflect `main` as of #9708/#9709. ERNIE-Image GGUF is still ✗
here; #9725 flips that cell when it merges.

## Related Issues / Discussions

Follow-up to #9708, #9709 and #9725.

## QA Instructions

- `pnpm -C docs build`: completes.
- Every internal link on the new page resolves to a built page (16
checked against `docs/dist`), and the `#quantized-formats` anchor
exists.
- The cells come from the loader registrations, the config classes and
their explicit refusals in `invokeai/backend/model_manager/`.

## Review

Docs only. No material findings.

## Checklist

- [x] _The PR has a short but descriptive title, suitable for a
changelog_
- [ ] _Meaningful regression coverage added / updated where needed;
obsolete tests/code removed_
- [ ] _Persisted-state and API changes include required migrations /
compatibility validation_
- [ ] _Relevant performance/efficiency opportunities considered;
material claims have evidence_
- [x] _Material review findings resolved and relevant checks rerun_
- [x] _Documentation added / updated (if applicable)_
- [ ] _Updated `What's New` copy (if doing a release after this PR)_
@lstein
lstein enabled auto-merge October 10, 2026 22:12
@lstein
lstein merged commit 0dc9c17 into invoke-ai:main Oct 10, 2026
17 checks passed
@Pfannkuchensack
Pfannkuchensack deleted the feat/ernie-image-gguf branch October 11, 2026 00:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

7.0.0 backend PRs that change backend files docs PRs that change docs invocations PRs that change invocations python PRs that change python files python-tests PRs that change python tests

Projects

Status: 7.0 Theme: Modular Design

Development

Successfully merging this pull request may close these issues.

[bug]: Ernie and Ideogram safetensors + GGUF support

2 participants