Repository navigation
fix(qwen4): exclude PLE n-gram base state from PEFT saves - #220
Merged
Merged
Conversation
hjh0119
approved these changes
Oct 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
When saving a Qwen4Exp LoRA adapter with
peft_format=True,save_weights()callsexport_weights(..., _is_saving=True). The PLE export condition then evaluatesskip_ngram_statetoFalse, so the adapter receives the base model'slayer_multipliers,ngram_heads_offsets,ngram_heads_vocab_sizes, and n-gramembedding shards (plus the scale when emitted). These are not LoRA factors and
can cause adapter consumers to reject the checkpoint, as well as unnecessarily
exporting a very large PLE table.
Make PEFT exports skip the base n-gram state regardless of
_is_saving.Full-model saves still include it, and online full-model exports retain the
existing CPU-offload behavior. PLE projections continue through the existing
PEFT-aware
_set_state_dictpath. The adapter import fix added in #204 remainsunchanged.
Validation:
mainat35bc42c070e68dddf17c2043a0412a9be420cc95(checked on 2026-10-09); the same save issue is present in
release/1.6at73d6807a11acd8caf039a27459e3f4c377327a32andswift-dev-v5at99bf4022a3d20338d683b7b96513331bd32ab9bc.python -m pytest -q tests/test_qwen4_exp_peft_export.py: 5 tests and 22 subtestspass after the fix. The original source produces six expected failures in the adapter-save scenarios (and no test errors).
adapter saving on a PP stage without the PLE layer, full-model saves, online
full-model exports, and adapter/full-model imports.
tensor operations, the table iterator, and PP collectives are mocked. They
verify whether base buffers are exported and table export is scheduled.
change addresses PLE base-state leakage; other LoRA module mapping issues
remain outside its scope.