Skip to content

fix: place collectives and new modules on the process accelerator - #218

Open
shiaho777 wants to merge 1 commit into
modelscope:mainfrom
shiaho777:fix/accelerator-device
Open

shiaho777 wants to merge 1 commit into
modelscope:mainfrom
shiaho777:fix/accelerator-device

Conversation

@shiaho777

Copy link
Copy Markdown
Contributor

Checkpoint broadcast, pipeline-parallel flag reductions, the PLE export all-reduce, DSpark parameter placement, and the Hugging Face vision tower used device='cuda' or torch.cuda.current_device(). On Ascend those calls target a CUDA device that is not present, so export and model construction fail before an NPU kernel runs. A CPU-only process raises the same current_device() error. The chunked broadcast also treated any non-CUDA tensor as needing a copy to CUDA, which drops an NPU tensor onto the wrong device.

accelerator_device() returns the current NPU device, otherwise the current CUDA device, otherwise CPU or MPS. CUDA still uses the current CUDA device, which is what device='cuda' already selected.

Checked with python3 -m pytest tests/test_accelerator_device.py (2 passed) on a machine without CUDA: the helper returns CPU and does not call torch.cuda.current_device(). flake8 is clean on the changed files.

Checkpoint broadcast, PP flag reductions, the PLE export all-reduce, DSpark
parameters, and the Hugging Face vision tower all used device='cuda' or
torch.cuda.current_device(). On Ascend that targets a CUDA device that is not
there, so export and model construction fail before any NPU kernel runs. A
CPU-only process hits the same current_device() error.

accelerator_device() returns the current NPU or CUDA device, otherwise CPU or
MPS. CUDA still uses the current CUDA device, which is what device='cuda'
already selected.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant