Skip to content

fix(qwen4): exclude PLE n-gram base state from PEFT saves - #220

Merged
hjh0119 merged 1 commit into
modelscope:mainfrom
0KEAHA:fix/qwen4-ple-peft-save
Oct 9, 2026
Merged

hjh0119 merged 1 commit into
modelscope:mainfrom
0KEAHA:fix/qwen4-ple-peft-save

Conversation

@0KEAHA

@0KEAHA 0KEAHA commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

When saving a Qwen4Exp LoRA adapter with peft_format=True, save_weights() calls
export_weights(..., _is_saving=True). The PLE export condition then evaluates
skip_ngram_state to False, so the adapter receives the base model's
layer_multipliers, ngram_heads_offsets, ngram_heads_vocab_sizes, and n-gram
embedding shards (plus the scale when emitted). These are not LoRA factors and
can cause adapter consumers to reject the checkpoint, as well as unnecessarily
exporting a very large PLE table.

Make PEFT exports skip the base n-gram state regardless of _is_saving.
Full-model saves still include it, and online full-model exports retain the
existing CPU-offload behavior. PLE projections continue through the existing
PEFT-aware _set_state_dict path. The adapter import fix added in #204 remains
unchanged.

Validation:

  • Tested against upstream main at 35bc42c070e68dddf17c2043a0412a9be420cc95
    (checked on 2026-10-09); the same save issue is present in release/1.6 at
    73d6807a11acd8caf039a27459e3f4c377327a32 and swift-dev-v5 at
    99bf4022a3d20338d683b7b96513331bd32ab9bc.
  • python -m pytest -q tests/test_qwen4_exp_peft_export.py: 5 tests and 22 subtests
    pass after the fix. The original source produces six expected failures in the adapter-save scenarios (and no test errors).
  • Covers adapter export/save with CPU offload on/off and simulated PP1/PP2,
    adapter saving on a PP stage without the PLE layer, full-model saves, online
    full-model exports, and adapter/full-model imports.
  • The tests compile the actual production method and buffer names from source;
    tensor operations, the table iterator, and PP collectives are mocked. They
    verify whether base buffers are exported and table export is scheduled.
  • Flake8, isort, YAPF, and diff whitespace checks pass on the changed files.
  • No multi-GPU checkpoint round trip or real vLLM loading has been run. This
    change addresses PLE base-state leakage; other LoRA module mapping issues
    remain outside its scope.

@hjh0119
hjh0119 merged commit 356caec into modelscope:main Oct 9, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants