Skip to content

fix: do not install the QSA sparse kernel off CUDA - #217

Open
shiaho777 wants to merge 1 commit into
modelscope:mainfrom
shiaho777:fix/qsa-kernel-cuda-only
Open

shiaho777 wants to merge 1 commit into
modelscope:mainfrom
shiaho777:fix/qsa-kernel-cuda-only

Conversation

@shiaho777

Copy link
Copy Markdown
Contributor

qsa_sparse_supported() warned that the Triton kernel is only tested on CUDA, then returned True. On Ascend, Triton still imports, so Qwen4 training with packing or context parallelism compiles the CUDA tile sizes and aborts in the compiler with a UB/Cc overflow.

The kernel is no longer installed when CUDA is unavailable. With context parallel size 1 and packing off, QSA still uses the bool-mask path. Packing and context parallelism keep the explicit QSA_SPARSE_KERNEL=0 fallback to full attention. That fallback is not the sparse result, so it stays opt-in. The raised error now says a non-CUDA device is one reason the kernel is absent.

Checked with python3 -m pytest tests/test_qsa_sparse_supported.py (3 passed): non-CUDA returns false, a power-of-two head on CUDA stays enabled, and a non-power-of-two head stays disabled. flake8 is clean on the changed files.

qsa_sparse_supported() warned that the Triton kernel is only tested on CUDA,
then returned True anyway. On Ascend, Triton still imports, so Qwen4 training
with packing or context parallelism compiles the CUDA tile sizes and aborts
with a UB/Cc overflow. The kernel is now left uninstalled on non-CUDA devices.

CP == 1 without packing still uses the bool-mask path. Packing and CP raise
until QSA_SPARSE_KERNEL=0, which is the full-attention fallback. That fallback
is not the sparse result, so it stays opt-in.
@hjh0119

hjh0119 commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator

Thanks for contribution. Have you tested this on NPUs?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants