Repository navigation
Conversation
Absorbed MLA called the base checkpoint helper with hidden states and the compressed query in the value and attention-mask positions, and passed position_ids, which that helper does not accept. Selective core-attention recompute then either raised or ran DSA without x and qr. The non-checkpoint path already passed those tensors as keywords. Checkpoint the same keyword call. Tensor inputs go through tensor_parallel.checkpoint; packed sequence params stay closed over, matching the base helper. MLA does the same for x and qr, using core_attention_extra_kwargs when the installed mcore accepts it and a local checkpoint on 0.16, which drops unknown kwargs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Selective core-attention recompute for absorbed MLA called the base checkpoint helper as
(q, kv, hidden_states, q_compressed, attention_mask, v_up_weight, position_ids=...). That helper's parameters are(query, key, value, attention_mask, ...), and it does not acceptposition_ids. The call either raisedTypeErroror ran DSA with hidden states bound asvalueand the compressed query bound as the mask. The non-checkpoint path already passedx,qr, andup_v_weightas keywords.Absorbed MLA now checkpoints that same keyword call. Tensor inputs go through
tensor_parallel.checkpoint.packed_seq_paramsandattn_mask_typeare closed over, which is what the base helper does.MLA's own checkpoint path dropped
xandqrentirely. It now forwards them throughcore_attention_extra_kwargswhen the installed Megatron accepts that argument, and uses a local checkpoint on 0.16, whose helper ignores unknown keyword arguments. The rotary injection previously applied only on the non-checkpoint path is applied before both paths.Checked with
python3 -m pytest tests/test_absorbed_dsa_checkpoint.py tests/test_split_cp_inputs.py(2 passed). The absorbed test stubs Megatron and checks that recompute callscore_attentionwithvalue=None,x=hidden_states, andqr=q_compressed.flake8is clean on the changed files.