Repository navigation
feat(ernie-image): load GGUF transformers - #9725
Merged
Merged
Conversation
Add Main_GGUF_ErnieImage_Config and a GGUF loader; identification rests on the keys, since the published files name another architecture in their header. Move unpack_ggml_at_load into the GGUF package as its second caller; the ERNIE node accepts GGUF as a single file and the denoise reserves its dequantization transient. Add unsloth Q4_K_M and Q8_0 starters for ERNIE-Image and ERNIE-Image Turbo, pinned to their commits.
3 of 7 tasks
Pfannkuchensack
force-pushed
the
feat/ernie-image-gguf
branch
from
October 9, 2026 21:11
5bef7f4 to
d481dbe
Compare
This was referenced Oct 9, 2026
lstein
approved these changes
Oct 10, 2026
lstein
left a comment
Collaborator
There was a problem hiding this comment.
Approving. I checked this against the published files as well as the tests. I range-fetched the headers of the pinned unsloth Turbo Q4_K_M and Q8_0 GGUFs:
- At the real geometry, the 409 tensors match
ErnieImageTransformer2DModelkey for key and shape for shape, so the strict load holds. - Nothing quantized sits outside an
nn.Linear, so the per-module rule inunpack_ggml_at_loadcovers every tensor the forward reads outside a matmul. - Running the whole factory over the real key set, with the real
general.architecture: wanmetadata, yields exactly one match,Main_GGUF_ErnieImage_Config, with the right Turbo and base defaults. - All four starter URLs resolve at their pinned commits (5.02 GB and 8.69 GB).
Locally, the new tests and tests/backend/model_manager plus tests/backend/quantization pass (2257 passed). Two non-blocking nits are inline.
Only a Diffusers folder is expected to carry scheduler/, so single files of any format no longer warn. The GGUF forward test compared all-zero outputs (diffusers zero-inits adaLN and the final layer); it now fills those in and also covers Q4_K, the starters' type, on a 256-wide geometry.
lstein
added a commit
that referenced
this pull request
Oct 10, 2026
## Summary Which file format loads for which model was spread over the Model Families prose and the individual family pages, and went stale whenever a format landed (Ideogram 4 and ERNIE-Image GGUF most recently). This adds one **Model Format Support** page under Users Guide → Models with four tables: - **Main models.** Diffusers, bf16 single file, fp8 scaled, int8, nvfp4, MXFP8, GGUF, SDNQ and NF4 for all 14 families. - **Text encoders.** T5, CLIP, Qwen3, Qwen3-VL (4B/8B and the truncated 32B), Qwen3.5, Qwen2.5-VL, Mistral, Ministral, Gemma-2, Gemma-4 and UMT5. - **VAEs.** Single file and Diffusers, plus which family uses which VAE. - **LoRAs and adapters per family.** LoRA formats, ControlNet, IP-/T2I-Adapter, Redux, PiD. Cells are ✓ (loads), ✗ (not supported), **R** (refused at install with a message) or – (no such build). Model Families links to the page from its quantized-formats section. The new-model integration checklist asks to update it. The tables reflect `main` as of #9708/#9709. ERNIE-Image GGUF is still ✗ here; #9725 flips that cell when it merges. ## Related Issues / Discussions Follow-up to #9708, #9709 and #9725. ## QA Instructions - `pnpm -C docs build`: completes. - Every internal link on the new page resolves to a built page (16 checked against `docs/dist`), and the `#quantized-formats` anchor exists. - The cells come from the loader registrations, the config classes and their explicit refusals in `invokeai/backend/model_manager/`. ## Review Docs only. No material findings. ## Checklist - [x] _The PR has a short but descriptive title, suitable for a changelog_ - [ ] _Meaningful regression coverage added / updated where needed; obsolete tests/code removed_ - [ ] _Persisted-state and API changes include required migrations / compatibility validation_ - [ ] _Relevant performance/efficiency opportunities considered; material claims have evidence_ - [x] _Material review findings resolved and relevant checks rerun_ - [x] _Documentation added / updated (if applicable)_ - [ ] _Updated `What's New` copy (if doing a release after this PR)_
lstein
enabled auto-merge
October 10, 2026 22:12
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
ERNIE-Image already loads Diffusers pipelines and Comfy-Org's safetensors single files. GGUF transformers did not register. This PR adds them.
Main_GGUF_ErnieImage_Configidentifies ERNIE GGUFs by the same four-key fingerprint as the safetensors config. unsloth's GGUFs (ERNIE-Image and ERNIE-Image Turbo, the most downloaded) carry the diffusers key layout unchanged: all 409 tensors matchErnieImageTransformer2DModel, checked programmatically against the header. Theirgeneral.architecturenames another model, though:wanfor unsloth,fluxfor vantagewithai. Identification therefore rests on the keys alone, and the tests run the whole factory over both headers.ErnieImageGGUFModelbuilds the diffusers transformer and loads the GGUF strictly.aten.convolutioncannot run packed.unpack_ggml_at_load, moves from the Ideogram 4 loader (feat(ideogram4): load community GGUF transformers #9708) intoquantization/gguf/loaders.py, since ERNIE is its second caller.format !== 'diffusers'.Measured on an RTX 4090 with ERNIE-Image Turbo, 1024², 8 steps, warm (3 runs after a warm-up):
The GGUF image is visually indistinguishable from the Diffusers one, with the text rendered cleanly. The record received the Turbo defaults (8 steps, CFG 1.0) from its name.
Related Issues / Discussions
Closes #9499. Ideogram 4 GGUF is #9708; ERNIE-Image GGUF is this PR.
QA Instructions
gguf_quantizedwith 8 steps and CFG 1.0.Checks run:
uv run pytest tests/app/invocations tests/backend/quantization tests/backend/model_manager -n 4: 4585 passed.wanandfluxheaders, including the Turbo/base defaults.ruff checkandruff format --check.openapi.jsonandschema.tsregenerated; API package lint passes.Review
Material findings, resolved:
text_projBF16 and the conv F32. The real files keep the 4-D conv in BF16 and quantizetext_proj. The fixture now follows the published headers, so the 4-D BF16 unpack path is under test.Out of scope:
MistralModelwithout Ministral's YaRN RoPE. The encoder stays the 7.7 GB safetensors file._unwrap_unquantized_to_compute_dtypecovers F32/F16 only. Switching it tounpack_ggml_at_loadis a follow-up.Compatibility / Rollout
No migration; this adds a config class and a loader. The generated
openapi.jsonandschema.tsgainMain_GGUF_ErnieImage_Config. Existing Diffusers and safetensors installs behave as before.Checklist
What's Newcopy (if doing a release after this PR)