entity-extraction-worker: rows stranded in processing are never reclaimed
Summary
If an invocation dies after claiming rows but before finishing them, those rows stay
status = 'processing' forever. Nothing in the worker picks them back up, so the queue
silently stops draining while every surface-level signal says the run succeeded.
claimQueueItems() (index.ts:493) selects candidates with:
and claims them with the same guard. Both the select and the atomic update are
pending-only, so a row that reached processing and never advanced is invisible to
every subsequent invocation.
The file already anticipates this. The comment at index.ts:58-62 notes that a stalled
upstream can "consume the Edge Function's 150s wall-clock and leave claimed rows in
'processing' with no status update." That's accurate, and FETCH_TIMEOUT_MS reduces how
often it happens, but nothing recovers the rows once it does.
What I hit
Backfilled 781 thoughts and drained the queue with repeated ?limit=50 calls. At round
32 the function returned:
{ "code": "WORKER_RESOURCE_LIMIT", "message": "Function failed due to not having enough
compute resources (please check logs)" }
That looks like an isolate memory ceiling after roughly 30 back-to-back invocations,
which is plausibly its own issue, but the relevant part is the aftermath. I re-ran, the
loop reported processed 0 and exited cleanly, and every round in the final run showed
failed: 0. Everything looked finished.
The queue said otherwise:
status | count
------------+-------
complete | 744
processing | 28
pending | 8
failed | 1
28 rows claimed by the crashed invocation, permanently stuck. I only found them because
I ran the status query manually. The worker's own output gave no indication.
Why it's worth fixing
The failure is silent in both directions. The response body reports failed: 0 because
those rows never reached the error path, and processed: 0 on a later run reads as "queue
is empty" rather than "queue is blocked." Anyone monitoring the worker's return values
alone will conclude the backfill completed.
It also compounds. Each crash strands another batch, and since the auto-enqueue trigger
keeps adding new thoughts, the queue looks alive and healthy while an ever-growing set of
rows sits unreachable.
Not the same as #314
#314 covers INVOCATION_BUDGET_MS being hardcoded and non-overridable. That constant is
what produces the truncated: wall_clock_budget responses, so the two touch adjacent
code, but they're different failures. #314 is about not being able to tune when the
worker stops early; this is about rows being unrecoverable when it stops unexpectedly.
Raising or exposing the budget would not reclaim a stranded row.
Suggested fix
The straightforward version is a staleness reclaim at the top of claimQueueItems(),
before the pending select:
const staleCutoff = new Date(Date.now() - STALE_CLAIM_MS).toISOString();
await supabase
.from("entity_extraction_queue")
.update({ status: "pending", started_at: null, worker_version: null })
.eq("status", "processing")
.lt("started_at", staleCutoff);
started_at and worker_version are already written at claim time (index.ts:511-513),
so the data needed is present. A cutoff comfortably above the 150s platform wall-clock,
say 10 minutes, can't race a live invocation.
Alternatives if that's the wrong shape for the project: expose the reclaim behind a query
parameter so it's opt-in, or document the recovery SQL in the README's troubleshooting
section so people at least know to look. Right now neither the README nor the response
body mentions the state exists.
Manual recovery, for anyone who lands here from a search:
UPDATE entity_extraction_queue
SET status = 'pending', started_at = NULL, worker_version = NULL
WHERE status = 'processing';
Happy to open a PR if the reclaim-on-claim approach is the one you'd want.
Environment
- Supabase Edge Functions
- OpenRouter,
anthropic/claude-haiku-4-5
- 781 queued thoughts, UUID
thoughts.id
- 28 rows stranded by a single crash
entity-extraction-worker: rows stranded in
processingare never reclaimedSummary
If an invocation dies after claiming rows but before finishing them, those rows stay
status = 'processing'forever. Nothing in the worker picks them back up, so the queuesilently stops draining while every surface-level signal says the run succeeded.
claimQueueItems()(index.ts:493) selects candidates with:and claims them with the same guard. Both the select and the atomic update are
pending-only, so a row that reachedprocessingand never advanced is invisible toevery subsequent invocation.
The file already anticipates this. The comment at
index.ts:58-62notes that a stalledupstream can "consume the Edge Function's 150s wall-clock and leave claimed rows in
'processing' with no status update." That's accurate, and
FETCH_TIMEOUT_MSreduces howoften it happens, but nothing recovers the rows once it does.
What I hit
Backfilled 781 thoughts and drained the queue with repeated
?limit=50calls. At round32 the function returned:
That looks like an isolate memory ceiling after roughly 30 back-to-back invocations,
which is plausibly its own issue, but the relevant part is the aftermath. I re-ran, the
loop reported
processed 0and exited cleanly, and every round in the final run showedfailed: 0. Everything looked finished.The queue said otherwise:
28 rows claimed by the crashed invocation, permanently stuck. I only found them because
I ran the status query manually. The worker's own output gave no indication.
Why it's worth fixing
The failure is silent in both directions. The response body reports
failed: 0becausethose rows never reached the error path, and
processed: 0on a later run reads as "queueis empty" rather than "queue is blocked." Anyone monitoring the worker's return values
alone will conclude the backfill completed.
It also compounds. Each crash strands another batch, and since the auto-enqueue trigger
keeps adding new thoughts, the queue looks alive and healthy while an ever-growing set of
rows sits unreachable.
Not the same as #314
#314 covers
INVOCATION_BUDGET_MSbeing hardcoded and non-overridable. That constant iswhat produces the
truncated: wall_clock_budgetresponses, so the two touch adjacentcode, but they're different failures. #314 is about not being able to tune when the
worker stops early; this is about rows being unrecoverable when it stops unexpectedly.
Raising or exposing the budget would not reclaim a stranded row.
Suggested fix
The straightforward version is a staleness reclaim at the top of
claimQueueItems(),before the pending select:
started_atandworker_versionare already written at claim time (index.ts:511-513),so the data needed is present. A cutoff comfortably above the 150s platform wall-clock,
say 10 minutes, can't race a live invocation.
Alternatives if that's the wrong shape for the project: expose the reclaim behind a query
parameter so it's opt-in, or document the recovery SQL in the README's troubleshooting
section so people at least know to look. Right now neither the README nor the response
body mentions the state exists.
Manual recovery, for anyone who lands here from a search:
Happy to open a PR if the reclaim-on-claim approach is the one you'd want.
Environment
anthropic/claude-haiku-4-5thoughts.id