Skip to content

fmt=5: a deployable E8-lattice grouped container — int3 that beats shipped int4 quality at 25% fewer bytes (design + plan) #452

Description

@ZacharyZcR

Why now

Decode on every measured host is a bandwidth wall: expert bytes ÷ effective GB/s is the largest bucket on the critical path (6×5090 forensics in #431: ~47 ms/token of CPU-tier expert sweep at hit 100%). Bytes are speed. The quantization ablations (#81, #255, #347) found the lever:

container quality (MMLU-class Δ) bytes vs int4
shipped per-row int4 −9.3pp 1.00×
uniform int3-g64 −7.5pp 0.75×
int3-g64-e8-rot −5.9pp 0.75×

A 3-bit E8-lattice container measures better than the shipped 4-bit while carrying 25% fewer bytes — that is simultaneously: −25% on the CPU-tier wall, +33% experts per GB of VRAM (the whole routed set of GLM-5.2 drops to ~280 GB), and a quality gain over what users run today.

But those numbers are simulated: tools/quant_ablation.py quantizes to the nearest E8 point and stores floats. There is no deployable index encoding, no container format, no decode kernel. This issue is the plan to build them.

The pedigree (so nobody thinks this is exotic)

E8 codebooks are QuIP#'s core trick (ICML'24) and are mass-deployed in llama.cpp's IQ2/IQ3 quant family (E8-derived grids + sign-bit factoring, MIT). Our additions are the rate-scaled ball (#347 — the radius must grow with the bit budget), the g64 outer scale, and the target: a tiered MoE expert stream where bytes convert to tok/s at a measured rate.

Design: fmt=5 (E8 grouped container)

  1. Block encoding — decision point, two candidates:
    • (a) IQ3-style subset codebook (adopt llama.cpp's proven construction, MIT): small LUT (256–512 entries) that fits in SIMD registers / GPU shared memory, signs factored out by parity. Deployable today, kernels are a known quantity.
    • (b) our rate-scaled ball with structured coordinate coding: matches the ablation exactly but needs entropy-shaped index construction (a 3-bpw E8 ball has ~10⁷ points — naive indexing doesn't close). Research-grade.
    • Plan: (a) first, re-validated through the ablation harness (swap _quant_e8's ball for the candidate codebook, OLMoE A/B in hours, then GLM n=200). (b) stays a research fork if (a) leaves quality on the table.
  2. Container layout: per tensor — packed block indices + sign words, [O, ng] fp32 outer scales (g64, same shape discipline as fmt=4), rotation seeds (rotations regenerate deterministically from seed — nothing stored beyond a per-tensor seed). Rider tensors (.qs shape) keep the existing shard/index machinery.
  3. Engine dispatch: QT gains fmt=5; matmul_qt routes to new CPU kernels (AVX2/AVX-512: LUT gather + sign-apply + FMA, group scale exactly as matmul_i4_grouped); CUDA rides the cuda: grouped-int4 (fmt=4) support in the expert-group kernels — opens the GPU tier to g64 and E8 containers (#334) #451 GroupDesc path — the per-group scale plumbing is already merged-ready, the g4 kernels get e8 siblings with the LUT in shared memory.
  4. Converter: port _e8_nearest/rotation from the ablation into convert_fp8_to_int4.py with index emission; --ebits 3 --e8 on the existing mixed-precision flag surface (MTP head stays int8, io/shared tensors keep their existing overrides).
  5. Placement arithmetic (the payoff): ~280 GB total → on a 6×5090/251 GB host, VRAM 176 GB holds ~12.4k experts (+34%) and RAM covers the rest twice over; on 128 GB hosts the pinned share rises ~33%.

Validation ladder (each step gates the next)

  1. Codebook A/B in the ablation harness (OLMoE, torch, hours) — candidate (a) vs the published ball numbers.
  2. Round-trip exactness oracle: encode→decode == the torch dequant, bit-for-bit (the fmt=4 oracle pattern, cuda+engine: full fmt=4 (grouped int4 gs=64) support + diagnostic harness #298).
  3. CPU kernel oracle vs reference; CUDA oracle via the cuda: grouped-int4 (fmt=4) support in the expert-group kernels — opens the GPU tier to g64 and E8 containers (#334) #451 test pattern.
  4. GLM full convert + n=200 quality three-suite vs the int4 baseline (baseline run in flight).
  5. Speed A/B on the 6×5090 host: same protocol as this week's forensics — P0-EXEC GB/s fingerprint, usage-snapshot control, warm-segment rates.

I'll take the workstream; kernels and converter in increments, same discipline as the #431 series. Input welcome on the codebook decision (1) — especially from anyone who has run the llama.cpp IQ3 grids up close.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    qualityQualità del modello / quantizzazione

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions