Skip to content

fix: pass DSA hidden states through selective attention recompute - #216

Open
shiaho777 wants to merge 1 commit into
modelscope:mainfrom
shiaho777:fix/dsa-checkpoint-extra-args
Open

shiaho777 wants to merge 1 commit into
modelscope:mainfrom
shiaho777:fix/dsa-checkpoint-extra-args

Conversation

@shiaho777

Copy link
Copy Markdown
Contributor

Selective core-attention recompute for absorbed MLA called the base checkpoint helper as (q, kv, hidden_states, q_compressed, attention_mask, v_up_weight, position_ids=...). That helper's parameters are (query, key, value, attention_mask, ...), and it does not accept position_ids. The call either raised TypeError or ran DSA with hidden states bound as value and the compressed query bound as the mask. The non-checkpoint path already passed x, qr, and up_v_weight as keywords.

Absorbed MLA now checkpoints that same keyword call. Tensor inputs go through tensor_parallel.checkpoint. packed_seq_params and attn_mask_type are closed over, which is what the base helper does.

MLA's own checkpoint path dropped x and qr entirely. It now forwards them through core_attention_extra_kwargs when the installed Megatron accepts that argument, and uses a local checkpoint on 0.16, whose helper ignores unknown keyword arguments. The rotary injection previously applied only on the non-checkpoint path is applied before both paths.

Checked with python3 -m pytest tests/test_absorbed_dsa_checkpoint.py tests/test_split_cp_inputs.py (2 passed). The absorbed test stubs Megatron and checks that recompute calls core_attention with value=None, x=hidden_states, and qr=q_compressed. flake8 is clean on the changed files.

Absorbed MLA called the base checkpoint helper with hidden states and the
compressed query in the value and attention-mask positions, and passed
position_ids, which that helper does not accept. Selective core-attention
recompute then either raised or ran DSA without x and qr. The non-checkpoint
path already passed those tensors as keywords.

Checkpoint the same keyword call. Tensor inputs go through
tensor_parallel.checkpoint; packed sequence params stay closed over, matching
the base helper. MLA does the same for x and qr, using
core_attention_extra_kwargs when the installed mcore accepts it and a local
checkpoint on 0.16, which drops unknown kwargs.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant