Skip to content

OLMoE has no hot expert zone (counter-datapoint for pinned hot-store policies) #864

Description

@outtodata

TL;DR: measured routing on OLMoE-1B-7B-0924 (520 prompts, 5 domains + Wikipedia en/ru, ~50k tokens): the 32 most popular experts of a layer cover only 75% of routed top-8 slots; globally the top-256 of 1024 (layer, expert) pairs cover 52%. The popularity tail is flat - a pinned/learned hot-store buys much less on this model than a Zipf-heavy one.

Numbers

  • Per-layer popularity of the top-32/64 experts: mean 0.752, min 0.681 of slots
  • Global (layer, expert) pairs: top-64/1024 = 21.4%, top-128 = 34.7%, top-256 = 51.9%
  • Unique experts touched per token: mean 56.3 / 64 (p10 = 54)
  • Adjacent-token union Jaccard: 0.83 (temporal locality is high even though popularity is flat)

Method

Router top-8 indices collected per token per layer via HF forward hooks (output_router_logits) on 520 prompts (template domains: code/english/math/prose/russian + Wikipedia summaries en/ru). Collection script and aggregates available; happy to share the npz.

Why it matters

A hot-store sized from popularity curves assumes a steep head. On OLMoE the head is shallow: pinning the top-25% of experts still leaves ~48% of traffic unpinned. Suggest gating hot-store-style policies on a measured popularity skew check (or a flag) for OLMoE-like models. This is meant as a held-out datapoint for the hypothesis table, not a critique - the policy may still win on 256+expert models with steeper routing.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    benchmarkDatapoint di misurazione hardwarediscussionProposta / discussione aperta, non un taskperformanceVelocità / tok-s / ottimizzazioni

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions