What
tests/test_at_scale_ingestion_benchmark.py::TestRunIngestionBenchmark::test_checkpoint_summary_is_present_and_self_consistent fails intermittently at a low rate. Surfaced while verifying the DB lease protocol branch (PR #269).
Evidence, and its limits
| observation |
result |
| full-suite run on the lease branch |
1 failure |
| second full-suite run, same tree |
passed |
| 12 isolated runs, lease branch |
1 failure |
| 12 isolated runs, master worktree |
0 failures |
| 20 further isolated runs, lease branch |
0 failures |
| 8 runs under 8 busy loops |
0 failures |
Roughly 1 failure in 48 runs on the branch, 0 in 12 on master.
Attribution is unproven in both directions. 1-vs-0 over those sample sizes is not significant, and I never captured the assertion text — so which of the four assertions failed is unknown. Do not record this as pre-existing, and do not record it as caused by the lease protocol. Both claims are currently unsupported.
What was ruled out
checkpoints == 0 is not the cause. Measured the distribution directly over 10 clean benchmark runs: 3-4 checkpoints, never 0. _CheckpointPolicy._budget_elapsed returns True when nothing has been measured yet ("checkpoint once to seed d"), so the first gated checkpoint always fires provided _db_checkpoint_gated is called at all.
Not load-sensitive in the way #261's matcher test was — 0/8 under 8 busy loops on an 8-core box.
Suggested first step
Capture the failure rather than sampling for it. Run with -l --tb=long in a loop until it reproduces, and record which assertion failed and the full checkpoint_summary dict at that moment. The remaining candidates are the realised_duty == approx(total_seconds / elapsed_seconds) self-consistency check and summary is not None.
Found during PR #269. Not a blocker for it.
What
tests/test_at_scale_ingestion_benchmark.py::TestRunIngestionBenchmark::test_checkpoint_summary_is_present_and_self_consistentfails intermittently at a low rate. Surfaced while verifying the DB lease protocol branch (PR #269).Evidence, and its limits
Roughly 1 failure in 48 runs on the branch, 0 in 12 on master.
Attribution is unproven in both directions. 1-vs-0 over those sample sizes is not significant, and I never captured the assertion text — so which of the four assertions failed is unknown. Do not record this as pre-existing, and do not record it as caused by the lease protocol. Both claims are currently unsupported.
What was ruled out
checkpoints == 0is not the cause. Measured the distribution directly over 10 clean benchmark runs: 3-4 checkpoints, never 0._CheckpointPolicy._budget_elapsedreturnsTruewhen nothing has been measured yet ("checkpoint once to seed d"), so the first gated checkpoint always fires provided_db_checkpoint_gatedis called at all.Not load-sensitive in the way #261's matcher test was — 0/8 under 8 busy loops on an 8-core box.
Suggested first step
Capture the failure rather than sampling for it. Run with
-l --tb=longin a loop until it reproduces, and record which assertion failed and the fullcheckpoint_summarydict at that moment. The remaining candidates are therealised_duty == approx(total_seconds / elapsed_seconds)self-consistency check andsummary is not None.Found during PR #269. Not a blocker for it.