TL;DR: measured routing on OLMoE-1B-7B-0924 (520 prompts, 5 domains + Wikipedia en/ru, ~50k tokens): the 32 most popular experts of a layer cover only 75% of routed top-8 slots; globally the top-256 of 1024 (layer, expert) pairs cover 52%. The popularity tail is flat - a pinned/learned hot-store buys much less on this model than a Zipf-heavy one.
Numbers
- Per-layer popularity of the top-32/64 experts: mean 0.752, min 0.681 of slots
- Global (layer, expert) pairs: top-64/1024 = 21.4%, top-128 = 34.7%, top-256 = 51.9%
- Unique experts touched per token: mean 56.3 / 64 (p10 = 54)
- Adjacent-token union Jaccard: 0.83 (temporal locality is high even though popularity is flat)
Method
Router top-8 indices collected per token per layer via HF forward hooks (output_router_logits) on 520 prompts (template domains: code/english/math/prose/russian + Wikipedia summaries en/ru). Collection script and aggregates available; happy to share the npz.
Why it matters
A hot-store sized from popularity curves assumes a steep head. On OLMoE the head is shallow: pinning the top-25% of experts still leaves ~48% of traffic unpinned. Suggest gating hot-store-style policies on a measured popularity skew check (or a flag) for OLMoE-like models. This is meant as a held-out datapoint for the hypothesis table, not a critique - the policy may still win on 256+expert models with steeper routing.
TL;DR: measured routing on OLMoE-1B-7B-0924 (520 prompts, 5 domains + Wikipedia en/ru, ~50k tokens): the 32 most popular experts of a layer cover only 75% of routed top-8 slots; globally the top-256 of 1024 (layer, expert) pairs cover 52%. The popularity tail is flat - a pinned/learned hot-store buys much less on this model than a Zipf-heavy one.
Numbers
Method
Router top-8 indices collected per token per layer via HF forward hooks (
output_router_logits) on 520 prompts (template domains: code/english/math/prose/russian + Wikipedia summaries en/ru). Collection script and aggregates available; happy to share the npz.Why it matters
A hot-store sized from popularity curves assumes a steep head. On OLMoE the head is shallow: pinning the top-25% of experts still leaves ~48% of traffic unpinned. Suggest gating hot-store-style policies on a measured popularity skew check (or a flag) for OLMoE-like models. This is meant as a held-out datapoint for the hypothesis table, not a critique - the policy may still win on 256+expert models with steeper routing.