Skip to content

fix(benchmark): preserve leaderboard scores on rate limits - #182

Merged
Andyyyy64 merged 4 commits into
Andyyyy64:mainfrom
hannibal-lee:fix/ollb-rate-limit
Oct 3, 2026
Merged

Andyyyy64 merged 4 commits into
Andyyyy64:mainfrom
hannibal-lee:fix/ollb-rate-limit

Conversation

@hannibal-lee

@hannibal-lee hannibal-lee commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

What

  • Honor numeric and HTTP-date Retry-After values in the shared bounded retry helper.
  • Give the archived Open LLM Leaderboard rows fetch source-specific retry settings.
  • Preserve already fetched leaderboard pages when a later page remains rate-limited after all retries.
  • Add direct coverage for the 46-page production fallback, exhausted 429s, first-page failures, and the no-pyarrow path.

Why

The archived dataset currently has 4,576 rows, while the datasets-server rows endpoint caps each response at 100 rows. Normal installations do not include the development-only pyarrow dependency, so they use 46 sequential requests instead of the one-request parquet path.

When the request at offset 3,600 remained rate-limited, the source discarded all scores collected from the first 36 pages. This keeps the change within the source fetch behavior: no mirror, runtime dependency, cache format, or ranking changes.

Refs #180. Issue #180 will stay open until the published version is verified.

Testing

  • Tests pass (pytest) - 555 passed
  • New tests added
  • Tested on real hardware (not hardware-related)

Notes

A first-page 429 still fails the source normally. Partial results are returned only when at least one page succeeded, and a warning reports the failed offset and retained score count.

Partial results use the existing 24-hour cache. The fetch warning goes to stderr and does not repeat when reading the cache. Use --refresh to fetch the scores again.

@Andyyyy64

Copy link
Copy Markdown
Owner

Thanks for picking up #180 as well. The explanation of why a normal install makes 46 requests helps. I'll review the retry timing and what happens when collection stops partway through, including how the retained scores reach the cache and ranking.

@hannibal-lee

hannibal-lee commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

One flag on the cache half of your review: the partial OLLB result is cached exactly like a full one — {cached_at, scores}, 24h TTL, no completeness marker — so the warning line is the only thing that distinguishes them.

I kept it that way to stay inside the "no cache format change" scope from #180. I can instead skip the cache write when the source reports partial (tradeoff: it redoes the 46-page walk), or surface a partial marker in the summary. Both are internal-only — say which you prefer.

@Andyyyy64
Andyyyy64 merged commit 0043d82 into Andyyyy64:main Oct 3, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants