Skip to content

feat(ltx2): load LTX-2.5 transformer GGUF builds - #9814

Open
Pfannkuchensack wants to merge 1 commit into
invoke-ai:mainfrom
Pfannkuchensack:feat/ltx2-gguf-transformer
Open

Pfannkuchensack wants to merge 1 commit into
invoke-ai:mainfrom
Pfannkuchensack:feat/ltx2-gguf-transformer

Conversation

@Pfannkuchensack

@Pfannkuchensack Pfannkuchensack commented Oct 11, 2026 •

Copy link
Copy Markdown
Member

Summary

Community GGUF builds of the LTX-2.5 transformer now install and run. Until now every LTX-2 GGUF installed as an unknown model, with the reason only in the server log.

  • Install: a new Main_GGUF_LTX2_Config claims LTX-2.5 transformer GGUFs from their keys. That covers bare or model.diffusion_model.-prefixed keys under any general.architecture. The generation comes from the header's model_version when there is one, and from the structure otherwise. Dev or Distilled is read from the file name.
  • Refusals:
    • LTX-2.3 GGUFs (unsloth, QuantStack) are refused with a message naming the generation. Safetensors transformers without a header version get the same treatment: a transformer with video feed-forward biases is now dated "2.0/2.3" and refused, instead of "cannot tell" → unknown.
    • ComfyUI-GGUF Q8_CR files are refused as such. The factory now checks DECODES_GGUF_Q8_CR before parsing the markers, so molbal's build no longer fails on "names 'keyframes_abs_pos_embedding', not a weight".
  • Load: a new LTX2GGUFModel keeps the quantized Linears packed as GGMLTensor. Everything read outside a matmul, and everything unquantized, is unpacked once at load. The text connectors that every build bundles are skipped before their data is copied, through a new keep filter on gguf_sd_loader; the component folder supplies them, as it does for the official files. Key renaming and the meta build are shared with the safetensors path.
  • UI: webv2 accepts gguf_quantized LTX-2 mains, which run like the single files: component folder plus Gemma-4 encoder.
  • Starters: vantagewithai Dev and Distilled Q4_K_M are added as separate starters, not in the bundle.

Related Issues / Discussions

Closes #9728

QA Instructions

Automated checks

  • uv tool run ruff@0.11.2 check and format --check on all 17 changed Python files: clean.
  • pytest -n 4 tests/backend/model_manager tests/backend/quantization tests/app/invocations/test_ltx2_nodes.py tests/app/invocations/test_text_encoders_with_packed_layers.py tests/model_identification: 2758 passed, 156 skipped, 1 xfailed.
  • webv2 (src/features/video/core): vitest 508 passed; lint (format, oxlint, tsc --noEmit, architecture) passed.
  • The OpenAPI artifacts are regenerated, and identical to a fresh generation.
  • Not run: the docs Astro build (no docs node_modules locally). The docs change is prose only.

New tests

  • Loader tests over real tiny GGUFs packed like the published releases (Q8_0, Q4_0, Q4_K; ruygar-style quantized stem/adaLN; molbal-style BF16 and quantized keyframe embedding; prefixed keys). They compare a float32 forward against a dense model of the dequantized values, check the reservation exactly, and check the bf16 compute dtype.
  • Identification tests, including two stripped real headers in tests/model_identification: vantagewithai distilled Q4_K_M installs; unsloth LTX-2.3 dev Q4_K_M is refused. The identification harness gains an optional expected_refusal field.
  • Mutation probes catch each of these: dropping the keep filter, hard-coding float32, leaving quantized non-Linears packed, the old dating, and a disabled refusal branch.

E2E on an RTX 4090

Setup: distilled, 8 steps, same seed. A/B runs are interleaved, with 1 warm-up and 3 runs each, and the model cache is emptied before every run.

1248×704×121, device_working_mem_gb: 5 s/step (runs 1–3) resident peak VRAM (driver)
int8 single file 5.99 / 5.80 / 5.78 94 % of 18.1 GiB 22.2 GB
GGUF Q4_K_M 6.10 / 6.15 / 6.10 100 % of 12.5 GiB 18.3–18.5 GB
GGUF Q8_0 5.45 / 5.44 / 5.46 89 % of 19.2 GiB 22.1 GB
  • 896×512×49: int8 1.65 / 1.63 / 1.64, Q4_K_M 1.87 / 1.87 / 1.89, Q8_0 1.28 / 1.28 / 1.29 s/step. Unpacking K-quants costs the same every step, so the gap is larger on short sequences.
  • Quality, at 100 % residency: every variant is bit-identical from run to run. Against int8, Q8_0 reaches PSNR 28.7 dB and audio SNR 12.6 dB, with the same composition. Q4_K_M renders the same scene framed differently (PSNR 15.8 dB).
  • Real-file weight check: InvokeAI's Q4_K / Q5_K / Q6_K kernels match gguf-py's reader bit for bit on the real file. Q4_K_M differs from Q8_0 by 7–10 % / 3.9 % / 2.1–2.6 % rel. RMS (cos ≥ 0.995), the expected quantization error.
  • RAM:
    • Loading peaks at the held model + ~1 GB private memory; the bundled connectors are never copied.
    • Installing reads the whole file once. The private peak is file size + ~1.6 GB (17.3 GB for Q4_K_M, 25.2 GB for Q8_0). The working set briefly reaches about twice that, because of file cache.
    • A header-only probe (follow-up) would cut install to ~2 MB. It was measured on a 6.4 GB GGUF: +12.9 GB → +2 MB, 17 s → 0.03 s.
  • Browser (webv2):
    • The GGUF record shows the LTX-2 / GGUF badges, its variant and its default settings.
    • It is selectable in the Video panel.
    • Without components, Invoke is disabled with the reason shown at the slot.
    • Recall All on a GGUF video restores that GGUF.
    • A Dev Q4_K_M render (30 steps, CFG 3) produced clean video with sound.

Review

Material findings resolved:

  • The GGUF load follows the shared extra-keys policy ([enhancement]: unify unexpected/missing key handling across single-file model loaders #9437) and its completeness check.
  • The keep semantics are simplified: Q8_CR is refused file-wide.
  • Identification is pinned against real published headers.
  • The fixtures follow the published layout, with a dedicated test for molbal's quantized keyframe embedding.
  • The production compute dtype is tested.
  • The docs include recovery steps for GGUFs installed as unknown models before this change.

Remaining limitations:

  • Header metadata is not kept by the stripped fixtures (pre-existing stripper behaviour), so the real-file fixtures exercise structural dating only; header dating is unit-tested.
  • Installation still reads the whole GGUF (pre-existing for all GGUF families; follow-up).
  • The ERNIE-Image GGUF cell in the format matrix is corrected from ✗ to ✓ in passing. ERNIE GGUF already loaded on main.

Compatibility / Rollout

  • A new config class joins AnyModelConfig; openapi.json and schema.ts are regenerated. There is no migration.
  • LTX-2 GGUFs installed before this change remain unknown records. Re-identify model converts them; the docs say so.

Checklist

  • The PR has a short but descriptive title, suitable for a changelog
  • Meaningful regression coverage added / updated where needed; obsolete tests/code removed
  • Persisted-state and API changes include required migrations / compatibility validation
  • Relevant performance/efficiency opportunities considered; material claims have evidence
  • Material review findings resolved and relevant checks rerun
  • Documentation added / updated (if applicable)
  • Updated What's New copy (if doing a release after this PR)

View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Add Main_GGUF_LTX2_Config and LTX2GGUFModel: Linears stay packed, the bundled connectors are skipped at read.
Refuse LTX-2.3 GGUFs and ComfyUI-GGUF Q8_CR files with a reason; pin both against real stripped headers.
Add the vantagewithai Dev/Distilled Q4_K_M starters, webv2 support and measured docs.
@github-actions github-actions Bot added python PRs that change python files invocations PRs that change invocations backend PRs that change backend files frontend PRs that change frontend files python-tests PRs that change python tests docs PRs that change docs labels Oct 11, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backend PRs that change backend files docs PRs that change docs frontend PRs that change frontend files invocations PRs that change invocations python PRs that change python files python-tests PRs that change python tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[enhancement]: LTX-2 GGUF transformers

2 participants