Skip to content

The learning cache is engine-specific: .coli_usage has two incompatible writers and two engines that cannot produce it at all #700

Description

@terrizoaguimor

Summary

"A cache that learns" is one of the first things the README promises, and routing history is what feeds it (PIN=auto, AUTOPIN, REPIN, the .coli_pairs coupling prefetch, the expert atlas in the dashboard). That machinery is per-engine today, and the four engines disagree:

Engine ROUTE_TRACE writes .coli_usage reads it for PIN
colibri.c (GLM-5.2) ✅ ✅ sparse text triples ✅
inkling.c ❌ ✅ dense binary IKU1 ✅ (own format only)
olmoe.c ❌ ❌ ❌
kimi_k3.c (#676) ❌ ❌ ❌

So: routing traces exist on exactly one engine, and the usage file is written in two mutually unreadable formats under the same filename.

I hit this while instrumenting K3 to collect routing data, and rather than send a K3-only patch I would rather propose the shared facility — if that shape appeals, I am happy to implement it.

The format collision, precisely

colibri.c writes sparse text through stats_dump_q() (telemetry.h):

FILE *f=fopen(tmp,"w");                       /* text */
... fprintf(f,"%d %d %u\n", i, e, m->eusage[i][e]);   /* layer expert count, non-zero only */

inkling.c:798 writes a dense binary block with its own magic:

FILE *f = fopen(tp, "wb");                    /* binary */
uint32_t hdr[3] = { 0x31554B49u /* "IKU1" */, (uint32_t)c->n_layers, (uint32_t)E };
fwrite(hdr, 4, 3, f);
for (int i = 0; i < c->n_layers; i++) fwrite(m->eusage[i] ? ... , 4, E, f);

Both default to <snap>/.coli_usage and both honour PIN=<path>. The comment above inkling.c's writer says "same contract as glm's .coli_usage" — it is a different contract.

What that costs today:

  1. Point inkling at a snapshot where GLM has run: the magic check at inkling.c:750 fails and it declines cleanly — "[pin] …: not an inkling usage file, ignoring" — so the user does get told, but pinning is off and the history they accumulated is unusable.
  2. The other direction is quieter. GLM's usage_load is while(fscanf(f,"%d %d %u",...)==3): handed inkling's binary file it parses zero triples and returns 0 with no message, which is indistinguishable from "no history yet".
  3. Share a model directory between the two engines and whichever ran last has overwritten the other's history, in a format the other cannot read.

Why the fix is not "share the types"

telemetry.h already exists but is bound to GLM's structures — its own header says so:

Include after Model/Cfg/QT/ESlot/shards and st.h are defined; requires qt_bytes(), now_s(), rss_gb(), edisk_s(), and the g_cuda_* globals.

kimi_k3.c cannot include it: its Model/Cfg are different types with different fields, which is exactly why #676 is a separate engine. And #667 proposes unifying those types while #676 argues the architectures are incompatible — this proposal deliberately takes no side there. Routing telemetry does not need shared types. It needs a facility that speaks integers.

Proposal: c/route_trace.h

Dependency-free, no Model/Cfg, no st.h. Every engine already knows its layer index, the ids it selected and their gates — that is the entire input.

/* Engine-agnostic routing telemetry. Include anywhere; owns its own counters. */
static void rt_init(int n_layers, int n_experts);   /* reads ROUTE_TRACE= and the usage path */
static void rt_route(int layer, int row,            /* one call per routing decision:      */
                     const int *ids, const float *gates, int k);   /* traces AND counts    */
static void rt_save(void);                          /* atomic write of the usage file       */
static int  rt_load(const char *path);              /* read history back, for PIN/AUTOPIN   */
static const uint32_t *rt_counts(int layer);        /* what PIN ranking consumes            */

Per-engine cost after that: rt_init at startup, one rt_route in the router loop, rt_save where the engine already persists, rt_load where PIN is handled. For kimi_k3.c that is four call sites; olmoe.c likewise.

Format contract

The wire format should stay exactly what colibri.c emits today, because the ecosystem is already built on it — PIN=auto, tools/route_pairs.py, the dashboard's expert atlas, and any user's saved histories:

  • trace line: <call> <row> <layer> <id>:<gate%.4f> ... (colibri.c:3028-3031)
  • usage line: <layer> <expert> <count>, sparse, non-zero only

inkling.c's binary format is more compact for a dense history, so the sensible migration is: the shared writer emits text, and inkling.c's reader keeps sniffing the IKU1 magic for one release so existing histories still load. Happy to do it the other way round if you prefer the binary format as canonical — the point is one format, not which one.

Validation

The refactor of colibri.c is the only risky part, and it is provable rather than arguable:

  1. ROUTE_TRACE=/tmp/a on the current build, ROUTE_TRACE=/tmp/b after the refactor, same prompt and seed → cmp /tmp/a /tmp/b must be silent.
  2. Same for .coli_usage after an identical run.
  3. Oracle: SNAP=./glm_tiny TF=1 ./colibri 64 16 16 → 32/32, and greedy → 20/20.
  4. make check, clean build, 0 warnings, CUDA=1 and portable.

If (1) and (2) are byte-identical then PIN=auto, REPIN, .coli_pairs and the dashboard cannot tell the refactor happened.

Suggested phasing

Each phase stands alone; if only A lands, the collision is documented and nothing is worse.

What this is not

Offer

I have K3 routing telemetry working already (that is how I found this), and an H100 box with GLM-5.2 resident, the oracle fixture built, and tooling that consumes both formats. If the shape above looks right I can send phase A with the byte-identical diff evidence, and phase B behind it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    discussionProposta / discussione aperta, non un taskenhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions