Summary
"A cache that learns" is one of the first things the README promises, and routing history is what feeds it (PIN=auto, AUTOPIN, REPIN, the .coli_pairs coupling prefetch, the expert atlas in the dashboard). That machinery is per-engine today, and the four engines disagree:
| Engine |
ROUTE_TRACE |
writes .coli_usage |
reads it for PIN |
colibri.c (GLM-5.2) |
✅ |
✅ sparse text triples |
✅ |
inkling.c |
❌ |
✅ dense binary IKU1 |
✅ (own format only) |
olmoe.c |
❌ |
❌ |
❌ |
kimi_k3.c (#676) |
❌ |
❌ |
❌ |
So: routing traces exist on exactly one engine, and the usage file is written in two mutually unreadable formats under the same filename.
I hit this while instrumenting K3 to collect routing data, and rather than send a K3-only patch I would rather propose the shared facility — if that shape appeals, I am happy to implement it.
The format collision, precisely
colibri.c writes sparse text through stats_dump_q() (telemetry.h):
FILE *f=fopen(tmp,"w"); /* text */
... fprintf(f,"%d %d %u\n", i, e, m->eusage[i][e]); /* layer expert count, non-zero only */
inkling.c:798 writes a dense binary block with its own magic:
FILE *f = fopen(tp, "wb"); /* binary */
uint32_t hdr[3] = { 0x31554B49u /* "IKU1" */, (uint32_t)c->n_layers, (uint32_t)E };
fwrite(hdr, 4, 3, f);
for (int i = 0; i < c->n_layers; i++) fwrite(m->eusage[i] ? ... , 4, E, f);
Both default to <snap>/.coli_usage and both honour PIN=<path>. The comment above inkling.c's writer says "same contract as glm's .coli_usage" — it is a different contract.
What that costs today:
- Point
inkling at a snapshot where GLM has run: the magic check at inkling.c:750 fails and it declines cleanly — "[pin] …: not an inkling usage file, ignoring" — so the user does get told, but pinning is off and the history they accumulated is unusable.
- The other direction is quieter. GLM's
usage_load is while(fscanf(f,"%d %d %u",...)==3): handed inkling's binary file it parses zero triples and returns 0 with no message, which is indistinguishable from "no history yet".
- Share a model directory between the two engines and whichever ran last has overwritten the other's history, in a format the other cannot read.
Why the fix is not "share the types"
telemetry.h already exists but is bound to GLM's structures — its own header says so:
Include after Model/Cfg/QT/ESlot/shards and st.h are defined; requires qt_bytes(), now_s(), rss_gb(), edisk_s(), and the g_cuda_* globals.
kimi_k3.c cannot include it: its Model/Cfg are different types with different fields, which is exactly why #676 is a separate engine. And #667 proposes unifying those types while #676 argues the architectures are incompatible — this proposal deliberately takes no side there. Routing telemetry does not need shared types. It needs a facility that speaks integers.
Proposal: c/route_trace.h
Dependency-free, no Model/Cfg, no st.h. Every engine already knows its layer index, the ids it selected and their gates — that is the entire input.
/* Engine-agnostic routing telemetry. Include anywhere; owns its own counters. */
static void rt_init(int n_layers, int n_experts); /* reads ROUTE_TRACE= and the usage path */
static void rt_route(int layer, int row, /* one call per routing decision: */
const int *ids, const float *gates, int k); /* traces AND counts */
static void rt_save(void); /* atomic write of the usage file */
static int rt_load(const char *path); /* read history back, for PIN/AUTOPIN */
static const uint32_t *rt_counts(int layer); /* what PIN ranking consumes */
Per-engine cost after that: rt_init at startup, one rt_route in the router loop, rt_save where the engine already persists, rt_load where PIN is handled. For kimi_k3.c that is four call sites; olmoe.c likewise.
Format contract
The wire format should stay exactly what colibri.c emits today, because the ecosystem is already built on it — PIN=auto, tools/route_pairs.py, the dashboard's expert atlas, and any user's saved histories:
- trace line:
<call> <row> <layer> <id>:<gate%.4f> ... (colibri.c:3028-3031)
- usage line:
<layer> <expert> <count>, sparse, non-zero only
inkling.c's binary format is more compact for a dense history, so the sensible migration is: the shared writer emits text, and inkling.c's reader keeps sniffing the IKU1 magic for one release so existing histories still load. Happy to do it the other way round if you prefer the binary format as canonical — the point is one format, not which one.
Validation
The refactor of colibri.c is the only risky part, and it is provable rather than arguable:
ROUTE_TRACE=/tmp/a on the current build, ROUTE_TRACE=/tmp/b after the refactor, same prompt and seed → cmp /tmp/a /tmp/b must be silent.
- Same for
.coli_usage after an identical run.
- Oracle:
SNAP=./glm_tiny TF=1 ./colibri 64 16 16 → 32/32, and greedy → 20/20.
make check, clean build, 0 warnings, CUDA=1 and portable.
If (1) and (2) are byte-identical then PIN=auto, REPIN, .coli_pairs and the dashboard cannot tell the refactor happened.
Suggested phasing
Each phase stands alone; if only A lands, the collision is documented and nothing is worse.
What this is not
Offer
I have K3 routing telemetry working already (that is how I found this), and an H100 box with GLM-5.2 resident, the oracle fixture built, and tooling that consumes both formats. If the shape above looks right I can send phase A with the byte-identical diff evidence, and phase B behind it.
Summary
"A cache that learns" is one of the first things the README promises, and routing history is what feeds it (
PIN=auto,AUTOPIN,REPIN, the.coli_pairscoupling prefetch, the expert atlas in the dashboard). That machinery is per-engine today, and the four engines disagree:ROUTE_TRACE.coli_usagePINcolibri.c(GLM-5.2)inkling.cIKU1olmoe.ckimi_k3.c(#676)So: routing traces exist on exactly one engine, and the usage file is written in two mutually unreadable formats under the same filename.
I hit this while instrumenting K3 to collect routing data, and rather than send a K3-only patch I would rather propose the shared facility — if that shape appeals, I am happy to implement it.
The format collision, precisely
colibri.cwrites sparse text throughstats_dump_q()(telemetry.h):inkling.c:798writes a dense binary block with its own magic:Both default to
<snap>/.coli_usageand both honourPIN=<path>. The comment aboveinkling.c's writer says "same contract as glm's.coli_usage" — it is a different contract.What that costs today:
inklingat a snapshot where GLM has run: the magic check atinkling.c:750fails and it declines cleanly — "[pin] …: not an inkling usage file, ignoring" — so the user does get told, but pinning is off and the history they accumulated is unusable.usage_loadiswhile(fscanf(f,"%d %d %u",...)==3): handedinkling's binary file it parses zero triples and returns 0 with no message, which is indistinguishable from "no history yet".Why the fix is not "share the types"
telemetry.halready exists but is bound to GLM's structures — its own header says so:kimi_k3.ccannot include it: itsModel/Cfgare different types with different fields, which is exactly why #676 is a separate engine. And #667 proposes unifying those types while #676 argues the architectures are incompatible — this proposal deliberately takes no side there. Routing telemetry does not need shared types. It needs a facility that speaks integers.Proposal:
c/route_trace.hDependency-free, no
Model/Cfg, nost.h. Every engine already knows its layer index, the ids it selected and their gates — that is the entire input.Per-engine cost after that:
rt_initat startup, onert_routein the router loop,rt_savewhere the engine already persists,rt_loadwherePINis handled. Forkimi_k3.cthat is four call sites;olmoe.clikewise.Format contract
The wire format should stay exactly what
colibri.cemits today, because the ecosystem is already built on it —PIN=auto,tools/route_pairs.py, the dashboard's expert atlas, and any user's saved histories:<call> <row> <layer> <id>:<gate%.4f> ...(colibri.c:3028-3031)<layer> <expert> <count>, sparse, non-zero onlyinkling.c's binary format is more compact for a dense history, so the sensible migration is: the shared writer emits text, andinkling.c's reader keeps sniffing theIKU1magic for one release so existing histories still load. Happy to do it the other way round if you prefer the binary format as canonical — the point is one format, not which one.Validation
The refactor of
colibri.cis the only risky part, and it is provable rather than arguable:ROUTE_TRACE=/tmp/aon the current build,ROUTE_TRACE=/tmp/bafter the refactor, same prompt and seed →cmp /tmp/a /tmp/bmust be silent..coli_usageafter an identical run.SNAP=./glm_tiny TF=1 ./colibri 64 16 16→ 32/32, and greedy → 20/20.make check, clean build, 0 warnings,CUDA=1andportable.If (1) and (2) are byte-identical then
PIN=auto,REPIN,.coli_pairsand the dashboard cannot tell the refactor happened.Suggested phasing
route_trace.h+colibri.crefactored onto it, with the byte-identical evidence above. No new behaviour, no new env vars. This is the PR that needs your review.kimi_k3.c— proves the facility is actually engine-agnostic rather than just claimed to be. Best landed after Kimi K3 engine: native-MXFP4 expert streaming, KDA + gated NoPE MLA + AttnRes + LatentMoE (#658) #676, or offered to @steve-m for inclusion.olmoe.candinkling.c, plus theIKU1transition read.Each phase stands alone; if only A lands, the collision is documented and nothing is worse.
What this is not
stdio/stdlib/stdint.Offer
I have K3 routing telemetry working already (that is how I found this), and an H100 box with GLM-5.2 resident, the oracle fixture built, and tooling that consumes both formats. If the shape above looks right I can send phase A with the byte-identical diff evidence, and phase B behind it.