fix(benchmark): preserve leaderboard scores on rate limits - #182
Merged
Merged
Conversation
Owner
|
Thanks for picking up #180 as well. The explanation of why a normal install makes 46 requests helps. I'll review the retry timing and what happens when collection stops partway through, including how the retained scores reach the cache and ranking. |
Contributor
Author
|
One flag on the cache half of your review: the partial OLLB result is cached exactly like a full one — I kept it that way to stay inside the "no cache format change" scope from #180. I can instead skip the cache write when the source reports partial (tradeoff: it redoes the 46-page walk), or surface a partial marker in the summary. Both are internal-only — say which you prefer. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Retry-Aftervalues in the shared bounded retry helper.Why
The archived dataset currently has 4,576 rows, while the datasets-server rows endpoint caps each response at 100 rows. Normal installations do not include the development-only
pyarrowdependency, so they use 46 sequential requests instead of the one-request parquet path.When the request at offset 3,600 remained rate-limited, the source discarded all scores collected from the first 36 pages. This keeps the change within the source fetch behavior: no mirror, runtime dependency, cache format, or ranking changes.
Refs #180. Issue #180 will stay open until the published version is verified.
Testing
pytest) - 555 passedNotes
A first-page 429 still fails the source normally. Partial results are returned only when at least one page succeeded, and a warning reports the failed offset and retained score count.
Partial results use the existing 24-hour cache. The fetch warning goes to stderr and does not repeat when reading the cache. Use
--refreshto fetch the scores again.