Skip to content

Add the fp8-quantized GeMM for dense linear layers - #18

Open
sfc-gh-reyazda wants to merge 9 commits into
pp-stagingfrom
fp8-fused-gemm
Open

Add the fp8-quantized GeMM for dense linear layers#18
sfc-gh-reyazda wants to merge 9 commits into
pp-stagingfrom
fp8-fused-gemm

Conversation

@sfc-gh-reyazda

Copy link
Copy Markdown
Collaborator

No description provided.

"deepspeedfp quantizer.") from err
reduce_dim = -1
if transposed:
orig_shape = (orig[:-2]+(orig[-1],orig[-2]))

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

orig should be orig_shape?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants