Skip to content

keep contending for job locks after losing them - #311

Merged
hugowetterberg merged 3 commits into
mainfrom
feature/job-lock-reacquire
Sep 28, 2026
Merged

hugowetterberg merged 3 commits into
mainfrom
feature/job-lock-reacquire

Conversation

@hugowetterberg

@hugowetterberg hugowetterberg commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Stage stopped indexing for ~80 minutes today, and the root cause is here rather than in Postgres.

Indexer.Run and the percolator took their job lock with a single RunWithContext. When a replica lost the lock, the indexer loop returned nil and logged indexer has stopped, but the indexer stayed in c.indexers. So reconciliation saw an enabled set that already had an indexer and did nothing, and that replica never contended for the lock again. Each Postgres blip took one more replica out:

  • 2026-09-23 19:12 UTC: a DNS failure resolving elephant-db.postgres cost -4jt4r its lock.
  • 2026-09-24 12:10 UTC: a second blip cost -ghvng its lock. No replica was left contending, and a rollout restart brought indexing back.

The percolator had the same flaw, but it only stopped when percolateEvents returned nil, so it survived by luck.

Changes:

  • Both jobs now run under joblock.Run, which re-acquires the lock whenever the job returns and recovers panics.
  • The percolator re-reads its persisted position on every acquisition instead of resuming from its own in-memory one, since another replica may have advanced it.
  • Both locks get the injected metrics registerer, so pg_job_lock_* lands in the test registry like everything else.
  • TestIndexerSurvivesLockLoss deletes the indexer's lock row out from under it, then requires the lock to be taken again and a document to be indexed. It fails on the old code with the same log sequence as stage.
  • acquire-lock in elephant_indexer_percolator_lifecycle_total now counts acquisitions rather than attempts. docs/observability.md says so.

The reason a slow Postgres cost the lock in the first place was in elephantine: a ping that commits after its client-side timeout left the in-memory iteration behind, and the next ping read as lost (ELE-1606). elephantine v0.30.2 fixes that, and this PR bumps to it, so a slow ping no longer costs the lock. Losing the lock for any other reason now costs a handover rather than the indexer.

elephantine goes from v0.29.1 to v0.30.2 for its job lock fixes. NewHTTPClientIntrumentation in the test setup is renamed to NewHTTPClientInstrumentation, as in #310.

The indexer and the percolator took their job lock with a single
RunWithContext. A lost lock ended the run, and the indexer stayed
registered with the coordinator, so the replica never contended for
the lock again and reconciliation saw nothing to fix. Each Postgres
blip took one more replica out, until stage had no indexer at all.

Both now run under joblock.Run, which re-acquires the lock whenever
the job returns. The percolator re-reads its persisted position on
every acquisition, since another replica may have advanced it.
@hugowetterberg
hugowetterberg merged commit 9cca3a0 into main Sep 28, 2026
5 checks passed
@hugowetterberg
hugowetterberg deleted the feature/job-lock-reacquire branch September 28, 2026 14:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants