Skip to content

Add the fp8 version of fused-moe - #10

Draft
sfc-gh-reyazda wants to merge 6 commits into
arctic-hao-0504from
fp8-fused-moe
Draft

Add the fp8 version of fused-moe#10
sfc-gh-reyazda wants to merge 6 commits into
arctic-hao-0504from
fp8-fused-moe

Conversation

@sfc-gh-reyazda

Copy link
Copy Markdown
Collaborator

This PR adds the fp8 version of the fused-moe for fusing the dequantization with the moe-gemm.
The initial performance evaluation shows between 15% to 33% speedup running inference with #token between 1 to 8K.
To run benchmark with this branch, please use this branch of deepspeed.

  • TODO: check the performance for different batch sizes on the right environment.
  • TODO: check the accuracy before releasing this feature.

@sfc-gh-reyazda
sfc-gh-reyazda changed the base branch from arctic to arctic-hao-0504 May 8, 2024 19:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant