You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Decode on every measured host is a bandwidth wall: expert bytes ÷ effective GB/s is the largest bucket on the critical path (6×5090 forensics in #431: ~47 ms/token of CPU-tier expert sweep at hit 100%). Bytes are speed. The quantization ablations (#81, #255, #347) found the lever:
container
quality (MMLU-class Δ)
bytes vs int4
shipped per-row int4
−9.3pp
1.00×
uniform int3-g64
−7.5pp
0.75×
int3-g64-e8-rot
−5.9pp
0.75×
A 3-bit E8-lattice container measures better than the shipped 4-bit while carrying 25% fewer bytes — that is simultaneously: −25% on the CPU-tier wall, +33% experts per GB of VRAM (the whole routed set of GLM-5.2 drops to ~280 GB), and a quality gain over what users run today.
But those numbers are simulated: tools/quant_ablation.py quantizes to the nearest E8 point and stores floats. There is no deployable index encoding, no container format, no decode kernel. This issue is the plan to build them.
The pedigree (so nobody thinks this is exotic)
E8 codebooks are QuIP#'s core trick (ICML'24) and are mass-deployed in llama.cpp's IQ2/IQ3 quant family (E8-derived grids + sign-bit factoring, MIT). Our additions are the rate-scaled ball (#347 — the radius must grow with the bit budget), the g64 outer scale, and the target: a tiered MoE expert stream where bytes convert to tok/s at a measured rate.
Design: fmt=5 (E8 grouped container)
Block encoding — decision point, two candidates:
(a) IQ3-style subset codebook (adopt llama.cpp's proven construction, MIT): small LUT (256–512 entries) that fits in SIMD registers / GPU shared memory, signs factored out by parity. Deployable today, kernels are a known quantity.
(b) our rate-scaled ball with structured coordinate coding: matches the ablation exactly but needs entropy-shaped index construction (a 3-bpw E8 ball has ~10⁷ points — naive indexing doesn't close). Research-grade.
Plan: (a) first, re-validated through the ablation harness (swap _quant_e8's ball for the candidate codebook, OLMoE A/B in hours, then GLM n=200). (b) stays a research fork if (a) leaves quality on the table.
Container layout: per tensor — packed block indices + sign words, [O, ng] fp32 outer scales (g64, same shape discipline as fmt=4), rotation seeds (rotations regenerate deterministically from seed — nothing stored beyond a per-tensor seed). Rider tensors (.qs shape) keep the existing shard/index machinery.
Converter: port _e8_nearest/rotation from the ablation into convert_fp8_to_int4.py with index emission; --ebits 3 --e8 on the existing mixed-precision flag surface (MTP head stays int8, io/shared tensors keep their existing overrides).
Placement arithmetic (the payoff): ~280 GB total → on a 6×5090/251 GB host, VRAM 176 GB holds ~12.4k experts (+34%) and RAM covers the rest twice over; on 128 GB hosts the pinned share rises ~33%.
Validation ladder (each step gates the next)
Codebook A/B in the ablation harness (OLMoE, torch, hours) — candidate (a) vs the published ball numbers.
GLM full convert + n=200 quality three-suite vs the int4 baseline (baseline run in flight).
Speed A/B on the 6×5090 host: same protocol as this week's forensics — P0-EXEC GB/s fingerprint, usage-snapshot control, warm-segment rates.
I'll take the workstream; kernels and converter in increments, same discipline as the #431 series. Input welcome on the codebook decision (1) — especially from anyone who has run the llama.cpp IQ3 grids up close.
Why now
Decode on every measured host is a bandwidth wall: expert bytes ÷ effective GB/s is the largest bucket on the critical path (6×5090 forensics in #431: ~47 ms/token of CPU-tier expert sweep at hit 100%). Bytes are speed. The quantization ablations (#81, #255, #347) found the lever:
A 3-bit E8-lattice container measures better than the shipped 4-bit while carrying 25% fewer bytes — that is simultaneously: −25% on the CPU-tier wall, +33% experts per GB of VRAM (the whole routed set of GLM-5.2 drops to ~280 GB), and a quality gain over what users run today.
But those numbers are simulated:
tools/quant_ablation.pyquantizes to the nearest E8 point and stores floats. There is no deployable index encoding, no container format, no decode kernel. This issue is the plan to build them.The pedigree (so nobody thinks this is exotic)
E8 codebooks are QuIP#'s core trick (ICML'24) and are mass-deployed in llama.cpp's IQ2/IQ3 quant family (E8-derived grids + sign-bit factoring, MIT). Our additions are the rate-scaled ball (#347 — the radius must grow with the bit budget), the g64 outer scale, and the target: a tiered MoE expert stream where bytes convert to tok/s at a measured rate.
Design:
fmt=5(E8 grouped container)_quant_e8's ball for the candidate codebook, OLMoE A/B in hours, then GLM n=200). (b) stays a research fork if (a) leaves quality on the table.[O, ng]fp32 outer scales (g64, same shape discipline as fmt=4), rotation seeds (rotations regenerate deterministically from seed — nothing stored beyond a per-tensor seed). Rider tensors (.qsshape) keep the existing shard/index machinery.QTgains fmt=5;matmul_qtroutes to new CPU kernels (AVX2/AVX-512: LUT gather + sign-apply + FMA, group scale exactly asmatmul_i4_grouped); CUDA rides the cuda: grouped-int4 (fmt=4) support in the expert-group kernels — opens the GPU tier to g64 and E8 containers (#334) #451GroupDescpath — the per-group scale plumbing is already merged-ready, the g4 kernels get e8 siblings with the LUT in shared memory._e8_nearest/rotationfrom the ablation intoconvert_fp8_to_int4.pywith index emission;--ebits 3 --e8on the existing mixed-precision flag surface (MTP head stays int8, io/shared tensors keep their existing overrides).Validation ladder (each step gates the next)
I'll take the workstream; kernels and converter in increments, same discipline as the #431 series. Input welcome on the codebook decision (1) — especially from anyone who has run the llama.cpp IQ3 grids up close.