Skip to content

docs(models): add a format support matrix for every model type - #9726

Open
Pfannkuchensack wants to merge 6 commits into
invoke-ai:mainfrom
Pfannkuchensack:docs/model-format-matrix
Open

Pfannkuchensack wants to merge 6 commits into
invoke-ai:mainfrom
Pfannkuchensack:docs/model-format-matrix

Conversation

@Pfannkuchensack

Copy link
Copy Markdown
Member

Summary

Which file format loads for which model was spread over the Model Families prose and the individual family pages, and went stale whenever a format landed (Ideogram 4 and ERNIE-Image GGUF most recently). This adds one Model Format Support page under Users Guide → Models with four tables:

  • Main models. Diffusers, bf16 single file, fp8 scaled, int8, nvfp4, MXFP8, GGUF, SDNQ and NF4 for all 14 families.
  • Text encoders. T5, CLIP, Qwen3, Qwen3-VL (4B/8B and the truncated 32B), Qwen3.5, Qwen2.5-VL, Mistral, Ministral, Gemma-2, Gemma-4 and UMT5.
  • VAEs. Single file and Diffusers, plus which family uses which VAE.
  • LoRAs and adapters per family. LoRA formats, ControlNet, IP-/T2I-Adapter, Redux, PiD.

Cells are ✓ (loads), ✗ (not supported), R (refused at install with a message) or – (no such build). Model Families links to the page from its quantized-formats section. The new-model integration checklist asks to update it.

The tables reflect main as of #9708/#9709. ERNIE-Image GGUF is still ✗ here; #9725 flips that cell when it merges.

Related Issues / Discussions

Follow-up to #9708, #9709 and #9725.

QA Instructions

  • pnpm -C docs build: completes.
  • Every internal link on the new page resolves to a built page (16 checked against docs/dist), and the #quantized-formats anchor exists.
  • The cells come from the loader registrations, the config classes and their explicit refusals in invokeai/backend/model_manager/.

Review

Docs only. No material findings.

Checklist

  • The PR has a short but descriptive title, suitable for a changelog
  • Meaningful regression coverage added / updated where needed; obsolete tests/code removed
  • Persisted-state and API changes include required migrations / compatibility validation
  • Relevant performance/efficiency opportunities considered; material claims have evidence
  • Material review findings resolved and relevant checks rerun
  • Documentation added / updated (if applicable)
  • Updated What's New copy (if doing a release after this PR)

New Model Format Support page lists which formats load, are refused or do not exist for every main model, text encoder, VAE and adapter family.
Model Families links to it from its quantized-formats section, and the new-model checklist asks to keep it current.

@joshistoast joshistoast left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Mark standalone Gemma-4 files as unsupported. The matrix advertises single-file bf16 support, but Gemma-4 requires a folder containing configuration, tokenizer, and weight files. Clarify that bf16 and int8 weights require those accompanying files.

[P2] Distinguish installation refusal from loading refusal. The legend says R formats are refused during installation, but some—including ERNIE-Image quantization and Qwen3.5 cases—are rejected only when loading for generation. Remove the installation-specific wording or distinguish these stages.

…mat matrix

Mark formats that install but fail on load as L, keep R for install-time refusals, and mark unrecognized files ✗.
Gemma-4 installs only as a folder with its config and tokenizer; quantized VAEs are not refused everywhere.
@Pfannkuchensack

Copy link
Copy Markdown
Member Author

Thanks, both points were right. Fixed in 73172b0.

Gemma-4: it is only recognized as a folder: config.json and the tokenizer files next to exactly one unsharded weight file, bf16 or int8. The row now reads Folder "✓ (bf16 or int8)", Single file ✗, and a note below the table says that a standalone .safetensors is not recognized.

Install vs. load: I went through every R cell against the configs and loaders. The legend now separates three outcomes:

  • R: the install fails with a message. This is now only Ideogram 4 nvfp4 in the table, plus LTX-2.0/2.3 files and Q8_CR GGUFs in the notes.
  • L: the model installs, then fails with a message when it is loaded for generation. This applies to the ERNIE-Image quantizations, Wan int8/nvfp4/MXFP8, LTX-2 and MiniMax H3 fp8/nvfp4, Qwen3.5, and the int8/nvfp4/fp8 encoder cells.
  • ✗: the file is not recognized, so it installs as an unknown model and the reason only shows up in the server log. This covers the LTX-2 Diffusers layout, the LTX-2, MiniMax H3 and Ministral 3B GGUFs, and Wan 2.1/Animate/S2V/Fun-Control/VACE.

While checking this I also corrected the VAE note. Quantized VAEs are refused when they load for SD 3, FLUX.2, Qwen-Image, Wan/Anima and Ideogram 4. The FLUX.1 and SD 1.x/SDXL single-file VAEs do not check at all.

@joshistoast joshistoast left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Include the supported FLUX.2 dev NF4 pipeline. Starter Models exposes FLUX.2 [dev] (Diffusers, NF4) from diffusers/FLUX.2-dev-bnb-4bit, and the existing family page documents it. Mark NF4 as supported inside the Diffusers folder, consistent with the Ideogram 4 row.

@Pfannkuchensack

Copy link
Copy Markdown
Member Author

Right, thanks. FLUX.2 [dev]'s NF4 Diffusers pipeline (diffusers/FLUX.2-dev-bnb-4bit) loads through from_pretrained, which applies the bitsandbytes quantization_config itself. The FLUX.2 NF4 cell now reads "inside the Diffusers folder (FLUX.2 [dev])", consistent with the Ideogram 4 row. Ideogram 4 and FLUX.2 [dev] are the only families with an NF4 Diffusers starter, so no other row changes. Fixed in e8e2318.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

docs PRs that change docs

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants