Skip to content

update new template to support weight cache and batch - #36

Merged
HaibaraAiChan merged 1 commit into
ai-decentralized:mainfrom
TomekWei:template_new
Nov 30, 2025
Merged

update new template to support weight cache and batch#36
HaibaraAiChan merged 1 commit into
ai-decentralized:mainfrom
TomekWei:template_new

Conversation

@TomekWei

Copy link
Copy Markdown
Contributor

I have refactored the original template based on the current main branch, and it now supports some of the system features that XU and Charlie previously added or modified, such as multi-batch size and caching. I tested the LLaMA model and the outputs are correct. With a batch size of 4, the throughput is 7.01 tokens/sec/sequence, and the effective throughput is 27.18 tokens/sec. I’m not sure whether this matches Charlie’s results exactly but the muti-batch seems works correctly.

@HaibaraAiChan
HaibaraAiChan merged commit 0794a52 into ai-decentralized:main Nov 30, 2025
JiuChen0 referenced this pull request in JiuChen0/BloomBee Mar 22, 2026
Co-authored-by: TomekWei <twei11@illinois.edu>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants