keep contending for job locks after losing them - #311
Merged
Merged
Conversation
The indexer and the percolator took their job lock with a single RunWithContext. A lost lock ended the run, and the indexer stayed registered with the coordinator, so the replica never contended for the lock again and reconciliation saw nothing to fix. Each Postgres blip took one more replica out, until stage had no indexer at all. Both now run under joblock.Run, which re-acquires the lock whenever the job returns. The percolator re-reads its persisted position on every acquisition, since another replica may have advanced it.
danijelvukoje
approved these changes
Sep 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stage stopped indexing for ~80 minutes today, and the root cause is here rather than in Postgres.
Indexer.Runand the percolator took their job lock with a singleRunWithContext. When a replica lost the lock, the indexer loop returnedniland loggedindexer has stopped, but the indexer stayed inc.indexers. So reconciliation saw an enabled set that already had an indexer and did nothing, and that replica never contended for the lock again. Each Postgres blip took one more replica out:elephant-db.postgrescost-4jt4rits lock.-ghvngits lock. No replica was left contending, and a rollout restart brought indexing back.The percolator had the same flaw, but it only stopped when
percolateEventsreturnednil, so it survived by luck.Changes:
joblock.Run, which re-acquires the lock whenever the job returns and recovers panics.pg_job_lock_*lands in the test registry like everything else.TestIndexerSurvivesLockLossdeletes the indexer's lock row out from under it, then requires the lock to be taken again and a document to be indexed. It fails on the old code with the same log sequence as stage.acquire-lockinelephant_indexer_percolator_lifecycle_totalnow counts acquisitions rather than attempts.docs/observability.mdsays so.The reason a slow Postgres cost the lock in the first place was in elephantine: a ping that commits after its client-side timeout left the in-memory iteration behind, and the next ping read as lost (ELE-1606). elephantine v0.30.2 fixes that, and this PR bumps to it, so a slow ping no longer costs the lock. Losing the lock for any other reason now costs a handover rather than the indexer.
elephantine goes from v0.29.1 to v0.30.2 for its job lock fixes.
NewHTTPClientIntrumentationin the test setup is renamed toNewHTTPClientInstrumentation, as in #310.