Skip to content

entity-extraction-worker: rows stranded in processing are never reclaimed #467

Description

@sjgold

entity-extraction-worker: rows stranded in processing are never reclaimed

Summary

If an invocation dies after claiming rows but before finishing them, those rows stay
status = 'processing' forever. Nothing in the worker picks them back up, so the queue
silently stops draining while every surface-level signal says the run succeeded.

claimQueueItems() (index.ts:493) selects candidates with:

.eq("status", "pending")

and claims them with the same guard. Both the select and the atomic update are
pending-only, so a row that reached processing and never advanced is invisible to
every subsequent invocation.

The file already anticipates this. The comment at index.ts:58-62 notes that a stalled
upstream can "consume the Edge Function's 150s wall-clock and leave claimed rows in
'processing' with no status update." That's accurate, and FETCH_TIMEOUT_MS reduces how
often it happens, but nothing recovers the rows once it does.

What I hit

Backfilled 781 thoughts and drained the queue with repeated ?limit=50 calls. At round
32 the function returned:

{ "code": "WORKER_RESOURCE_LIMIT", "message": "Function failed due to not having enough
compute resources (please check logs)" }

That looks like an isolate memory ceiling after roughly 30 back-to-back invocations,
which is plausibly its own issue, but the relevant part is the aftermath. I re-ran, the
loop reported processed 0 and exited cleanly, and every round in the final run showed
failed: 0. Everything looked finished.

The queue said otherwise:

status      | count
------------+-------
complete    |   744
processing  |    28
pending     |     8
failed      |     1

28 rows claimed by the crashed invocation, permanently stuck. I only found them because
I ran the status query manually. The worker's own output gave no indication.

Why it's worth fixing

The failure is silent in both directions. The response body reports failed: 0 because
those rows never reached the error path, and processed: 0 on a later run reads as "queue
is empty" rather than "queue is blocked." Anyone monitoring the worker's return values
alone will conclude the backfill completed.

It also compounds. Each crash strands another batch, and since the auto-enqueue trigger
keeps adding new thoughts, the queue looks alive and healthy while an ever-growing set of
rows sits unreachable.

Not the same as #314

#314 covers INVOCATION_BUDGET_MS being hardcoded and non-overridable. That constant is
what produces the truncated: wall_clock_budget responses, so the two touch adjacent
code, but they're different failures. #314 is about not being able to tune when the
worker stops early; this is about rows being unrecoverable when it stops unexpectedly.
Raising or exposing the budget would not reclaim a stranded row.

Suggested fix

The straightforward version is a staleness reclaim at the top of claimQueueItems(),
before the pending select:

const staleCutoff = new Date(Date.now() - STALE_CLAIM_MS).toISOString();
await supabase
  .from("entity_extraction_queue")
  .update({ status: "pending", started_at: null, worker_version: null })
  .eq("status", "processing")
  .lt("started_at", staleCutoff);

started_at and worker_version are already written at claim time (index.ts:511-513),
so the data needed is present. A cutoff comfortably above the 150s platform wall-clock,
say 10 minutes, can't race a live invocation.

Alternatives if that's the wrong shape for the project: expose the reclaim behind a query
parameter so it's opt-in, or document the recovery SQL in the README's troubleshooting
section so people at least know to look. Right now neither the README nor the response
body mentions the state exists.

Manual recovery, for anyone who lands here from a search:

UPDATE entity_extraction_queue
SET status = 'pending', started_at = NULL, worker_version = NULL
WHERE status = 'processing';

Happy to open a PR if the reclaim-on-claim approach is the one you'd want.

Environment

  • Supabase Edge Functions
  • OpenRouter, anthropic/claude-haiku-4-5
  • 781 queued thoughts, UUID thoughts.id
  • 28 rows stranded by a single crash

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions