Skip to content

perf(model-manager): identify GGUF models from the header alone - #9816

Open
Pfannkuchensack wants to merge 1 commit into
invoke-ai:mainfrom
Pfannkuchensack:perf/gguf-header-identification
Open

Pfannkuchensack wants to merge 1 commit into
invoke-ai:mainfrom
Pfannkuchensack:perf/gguf-header-identification

Conversation

@Pfannkuchensack

@Pfannkuchensack Pfannkuchensack commented Oct 11, 2026 •

Copy link
Copy Markdown
Member

Summary

Installing a GGUF model no longer reads its tensors into RAM. Identification probes each file with every candidate config, and for GGUF that meant gguf_sd_loader, which copies every tensor out of the file. Memory peaked at about the file size, plus as much again in the working set. A 23.6 GB transformer therefore needed ~22 GiB of free RAM just to be identified, and a machine with less RAM could not install it at all.

Identification reads names, shapes and quantization types, never values, and the GGUF header holds all of them. ModelOnDisk.load_state_dict() now builds the same GGMLTensors from the header via the new gguf_header_state_dict, the way the safetensors branch already does with meta tensors:

  • The packed shape and storage dtype match the full read.
  • The logical shape, including ComfyUI's comfy.gguf.orig_shape, comes from the same _logical_shape helper the full read uses.
  • The storage is on the meta device.

gguf_sd_loader's q8_cr="ignore" mode existed only for identification and is removed.

Before / after, measured on real files (peak private memory of one identification; time with a warm file cache):

File Before After
LTX-2.5 transformer Q8_0 (23.6 GB) 22.1 GiB, 17.6 s 0.05 GiB, 7.3 s
Ideogram 4 Q5_K (6.4 GB) 6.0 GiB, 5.0 s 0.03 GiB, 2.3 s

What time remains is the whole-file hash.

Related Issues / Discussions

Found while adding LTX-2 GGUF support (#9814, #9734). #9814's LTX-2 docs say installing a GGUF reads the whole file into RAM. Whichever PR merges second drops that sentence.

QA Instructions

Audit. Nothing reads tensor values during identification. The configs read only .shape, tensor_shape, dtypes and the GGMLTensor type. The other ModelOnDisk users are the reidentify route (which runs the same identification) and a migration that only reads the file size.

Real files. I identified 15 local GGUFs both ways, header and full read, and every file produced the identical config:

Families covered Peak memory: header / full read
FLUX.2 (dev, Klein), Krea-2, Ideogram 4, ERNIE-Image, Z-Image, Qwen-Image, Wan 2.2; encoders Qwen3, Qwen3-VL, T5, Gemma-2, Mistral; a FLUX VAE; a 12B LLM 0.01–0.5 GiB / ≈ file size

The 0.5 GiB cases are encoders, whose tokenizer metadata is parsed either way.

Tests.

  • New: per tensor, the header state dict matches the full read for F32, F16, BF16, Q8_0, Q4_K and an orig_shape tensor: same type, same logical and packed shapes, same storage dtype, and meta storage. Mutation probes on the packed shape and on orig_shape turn it red.
  • The Q8_CR identification test now runs through the header reader.
  • A new test checks the routing: ModelOnDisk.load_state_dict() returns meta-backed tensors for a GGUF. Reverting the branch to the full read turns it red.
  • pytest -n 4 tests/backend/quantization tests/backend/model_manager tests/model_identification: 2616 passed, 156 skipped, 1 xfailed.
  • ruff check and format: clean.

Review

No blocking findings. Resolved:

  • A test now pins the routing at the owning interface: ModelOnDisk.load_state_dict() returns meta-backed tensors.
  • The docstring is precise about Q8_CR layers, which come back as raw I8 codes without markers.
  • A stale test name is fixed.

Left as follow-ups, both outside this change:

  • WrappedGGUFReader.close() runs gc.collect(), about 0.19 s of a 0.2 s header read, and it frees nothing because the memmap is released by refcount on return.
  • model_util.read_checkpoint_meta has no callers and a comment that no longer matches.

Checklist

  • The PR has a short but descriptive title, suitable for a changelog
  • Meaningful regression coverage added / updated where needed; obsolete tests/code removed
  • Persisted-state and API changes include required migrations / compatibility validation
  • Relevant performance/efficiency opportunities considered; material claims have evidence
  • Material review findings resolved and relevant checks rerun
  • Documentation added / updated (if applicable)
  • Updated What's New copy (if doing a release after this PR)

View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Build identification's GGUF state dict as meta-backed GGMLTensors from the header instead of copying every tensor.
Peak RAM per probe drops from about the file size to megabytes; drop gguf_sd_loader's identification-only mode.
@github-actions github-actions Bot added python PRs that change python files backend PRs that change backend files python-tests PRs that change python tests labels Oct 11, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backend PRs that change backend files python PRs that change python files python-tests PRs that change python tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants