Is there an existing issue for this?
What should this feature add?
Load community GGUF builds of the MiniMax H3 transformer. Today installation is refused in
invokeai/backend/model_manager/configs/main.py with "GGUF-quantized MiniMax H3 checkpoints are not supported yet";
only Diffusers folders, bf16 single files and Comfy int8_tensorwise single files load.
Alternatives
Use the int8 single file (already supported) or the bf16 file.
Additional Content
The full matrix of what loads today is in #9726 (Model Format Support page).
Is there an existing issue for this?
What should this feature add?
Load community GGUF builds of the MiniMax H3 transformer. Today installation is refused in
invokeai/backend/model_manager/configs/main.pywith "GGUF-quantized MiniMax H3 checkpoints are not supported yet";only Diffusers folders, bf16 single files and Comfy
int8_tensorwisesingle files load.Alternatives
Use the int8 single file (already supported) or the bf16 file.
Additional Content
unsloth/MiniMax-H3-GGUF(~1.7M downloads),Abiray/MiniMax-H3-GGUF,Abiray/MiniMax-H3-Pruned-GGUF,realrebelai/MiniMax-H3_GGUFs.(Add single file support for Krea 2 huggingface/diffusers#14914). InvokeAI pins
diffusers==0.40.0and vendors the MiniMax-H3 code, so the bumpand its key mapping are worth evaluating here, both for the single file and for the GGUF key layout.
Main_GGUF_*_Configfingerprinted by keys (GGUF headers often name another architecture), a loader that keeps quantized Linears packed and unpacks the rest withunpack_ggml_at_load, and the denoise node'speak_dequant_transient_bytes, which already counts GGUF layers (fix(model-cache): reserve the GGUF dequantization transient in every node #9709).The full matrix of what loads today is in #9726 (Model Format Support page).