test(tokenless): stop the late-wake re-list scenario flaking on a loaded host - #3315
Conversation
…ded host Scenario 14 of tests/test-claude-code-detect-retry.sh pins a real contract -- a backoff the deadline clamp shortened still buys no attempt once the host wakes it past the window -- but it ran that contract under a 1s re-list window, and $SECONDS counts whole seconds. The re-list loop reads the clock twice: once to decide the window is still open, then again, after two awk spawns, to clamp the backoff against what is left of it. On a host that spends that single second on the initial `plugin list`, the loop is left either nothing to clamp or no iteration at all, and detect.sh -- which behaved correctly -- is reported as broken. Observed on an unmodified checkout of main (a9a3719), running the exact command CI runs: FAIL: the 2s backoff should have been clamped to the 1s window, requested 0.000s Reproduced deterministically by making the stubbed registry list slow, which is what a loaded runner does: STUB_SLOWDOWN=1.2 bash <scenario 14 only> FAIL: expected exactly one backoff before the window closed (saw 0 sleeps) Give the scenario the 3s window scenario 12 already runs on and the same knobs (a 30s backoff under a 60s ceiling, so the clamp always bites), and keep only the strict invariant that holds whichever way the host behaved: exactly one `plugin list`. A clamped sleep of `left` plus 3s of lateness lands at least 3s past a deadline whose slack is 1s, so the wake-up rule has to deny the attempt. The sleep count and the requested delay become bounds, matching scenario 13's style. The scenario still fails when the contract breaks. Mutating detect.sh locally: - probe_attempt_allowed() always allowing -> "must start no new list (saw 2 calls)" - probe_delay() clamping against the backoff ceiling instead of the window -> "should have been clamped to the 3s window (requested 4.000s)" detect.sh is unchanged; this is a test-only fix.
ikunkun-sys
left a comment
There was a problem hiding this comment.
Reviewed at 67e766f. No blocking findings.
This test-only change addresses a real false failure: the initial plugin query can consume the old one-second window, making the required positive sleep or exact sleep count invalid even when detect.sh correctly refuses further work. The wider window and bounded sleep assertions retain the strict single-query invariant. Production detect.sh is unchanged.
Independent validation:
- bash -n and the full test-claude-code-detect-retry.sh suite passed; git diff --check passed.
- In extracted scenario-14 scratch harnesses, a 1.2-second initial-query delay failed the parent test with zero sleeps and passed the revised test. A 3.5-second initial-query delay also passed the revised test.
- A scratch mutation making probe_attempt_allowed() always succeed caused the revised scenario to fail on two plugin-list calls, confirming it still detects the late-wake regression when that path executes.
- Current Test tokenless CI passed; no existing review threads.
Non-blocking coverage tradeoff: allowing zero sleeps means a run that exhausts the entire window during its initial query does not exercise the post-sleep guard. This is acceptable for avoiding the false failure, but the upper-bound assertion alone does not prove the clamp executed on every run. Validation used injected delays, not a genuinely contended runner.
What
Scenario 14 of
src/tokenless/tests/test-claude-code-detect-retry.sh— "a clamped backoff that wakes late buys no attempt" — pins a real contract, but it ran that contract under a 1s re-list window while$SECONDScounts whole seconds. The re-list loop indetect.shreads that clock twice: once to decide the window is still open, then again, after twoawkspawns, to clamp the backoff against what is left of it. A host that spends that single second on the initialplugin listleaves the loop either nothing to clamp or no iteration at all, anddetect.sh— which behaved correctly — gets reported as broken.Reproduction (before)
On an unmodified checkout of
main(a9a3719), running the exact command CI runs:It is a race rather than a hard failure — the same suite passed 3/3 runs immediately afterwards — so it was also reproduced deterministically by slowing the stubbed registry list, which is what a loaded runner does. Scenario 14 was extracted into a scratch harness for that (harness deleted; not part of this PR):
Change
Test-only.
adapters/tokenless/claude-code/scripts/detect.shis untouched.run_detect 3 30 5 3 60: a 30s backoff under a 60s ceiling, so the clamp always bites), late wake 2s → 3s.plugin list. A clamped sleep ofleftplus 3s of lateness lands at least 3s past a deadline whose slack isDETECT_PROBE_DEADLINE_SLACK=1, so the wake-up rule has to deny the attempt whichever way the host behaved before it.-le 1,<= 3), matching the style scenario 13 already uses for the same reason. The upper bound is what proves the clamp was exercised: unclamped, the request would be 30.000s.Test report
Environment: Linux x86_64, GNU bash 4.4.20, GNU Make 4.2.1, GNU awk, python3 3.8.17, rustc/cargo 1.96.0, jq 1.6, shellcheck 0.10.0. Base:
main@a9a3719.Before/after under simulated load (scratch harness running scenario 14 alone, stubbed
claude pluginlist slowed bySTUB_SLOWDOWN):STUB_SLOWDOWNFull suite:
bash src/tokenless/tests/test-claude-code-detect-retry.sh— 3/3 runs pass, 11.4s each (was 8.4s; the 3s is the longer late-wake sleep).Mutation checks — the scenario must still fail when the contract breaks; each mutation was reverted afterwards:
detect.shprobe_attempt_allowed()always returns 0probe_delay()clamps against the backoff ceiling (4) instead of the windowWider target:
make -C src/tokenless test-integration— hook golden parity OK (2 tests), claude-code detect retry OK, toon cleanup OK, then the Python suites:test_compress_response_hook39 skipped,test_codex_response_diagnostics4 OK,test_hermes_lifecycle13 OK,test_compress_schema_hook18 skipped. It stops attests/test_hook_contract.pywithTypeError: unsupported operand type(s) for |: 'type' and 'NoneType'— that file annotates with PEP 604str | None, which needs Python >= 3.10, and the local interpreter is 3.8.17 while CI pins 3.11. Pre-existing local toolchain limit, unrelated to this change (this PR touches one shell test file).Lint/format:
bash -nclean,shellcheck -S warning tests/test-claude-code-detect-retry.shclean.cargo fmt --all -- --checkandcargo clippy --workspace --all-targets -- -D warningsclean (no Rust touched).Rust sanity on the same checkout:
cargo test --workspace -- --test-threads=1— 811 passed / 0 failed / 3 ignored.Merge cleanliness:
git merge-treeagainst the heads of every open tokenless PR (#2322, #2452, #2530, #2532, #2877, #3242) — no conflicts.Not run: the
Test tokenlessCI job itself (no self-hosted runner locally), and the scenario under a genuinely contended runner rather than a stubbed delay.