Is there an existing issue for this?
What should this feature add?
Load quantized builds of the Qwen3.5 4B encoder Anima 3.8B uses. Today only the bf16 single file loads:
scaled fp8, int8 and nvfp4 are refused by reject_quantized_side_channel (load/model_loaders/qwen3_5_encoder.py),
and a GGUF file does not match any config.
Alternatives
Use the bf16 encoder.
Additional Content
The Qwen3 encoder loader already handles fp8, int8, nvfp4, GGUF and SDNQ and is the natural template.
The full matrix of what loads today is in #9726 (Model Format Support page).
Is there an existing issue for this?
What should this feature add?
Load quantized builds of the Qwen3.5 4B encoder Anima 3.8B uses. Today only the bf16 single file loads:
scaled fp8, int8 and nvfp4 are refused by
reject_quantized_side_channel(load/model_loaders/qwen3_5_encoder.py),and a GGUF file does not match any config.
Alternatives
Use the bf16 encoder.
Additional Content
The Qwen3 encoder loader already handles fp8, int8, nvfp4, GGUF and SDNQ and is the natural template.
The full matrix of what loads today is in #9726 (Model Format Support page).