Skip to content

Pull requests: lvyufeng/PocketLLM

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Optimize GQA attention with online softmax (1.46× prefill speedup)
#176 opened Sep 13, 2026 by lvyufeng Owner Loading…
6 tasks done
Add Phase 3.5 performance validation benchmarks
#127 opened Sep 3, 2026 by lvyufeng Owner Loading…
5 of 7 tasks
Fix bench_qwen_gqa_prefill comparing every mode against itself
#87 opened Aug 28, 2026 by lvyufeng Owner Loading…
Add an all-reduce bandwidth bench at prefill chunk sizes
#86 opened Aug 28, 2026 by lvyufeng Owner Loading…
Add Qwen drafter acceptance and speedup comparison
#79 opened Aug 25, 2026 by lvyufeng Owner Loading…
ProTip! Mix and match filters to narrow down what you’re looking for.