Skip to content

About

Self-contained task queue on Postgres: SKIP LOCKED claims, leases with fencing tokens, retry with jittered backoff, DLQ, worker harness, React dashboard

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Task Queue

A small, self-contained task queue with a worker harness and a minimal React UI, built from primitives on Postgres (pg driver, no queue library). Producers enqueue typed tasks (llm, js, http); worker processes claim them under concurrency, run the matching handler, and ack/nack; failures retry with backoff and exhausted tasks land in a DLQ that a human can requeue.

Design rationale and trade-offs are in DESIGN.md.

Requirements

  • Node 20+
  • Docker (for Postgres, via docker-compose.yml)

Run it locally

npm install && npm run dev

One command: starts Postgres (Docker), applies the schema, then runs the API (:3001), one worker for the default queue, and the React UI (http://localhost:5173) side by side.

The parts also run individually — e.g. to demonstrate several workers in parallel:

npm run api        # Postgres + migrations + API only
npm run worker     # one worker; QUEUE=<name> WORKER_ID=<id> to vary
npm run web        # the React UI

To see the whole lifecycle live (success, terminal failure, retries with backoff into the DLQ, a delayed task, a dedupe collapse), seed the demo set in a second terminal and watch the UI:

npm run demo

Quick end-to-end check with the API on :3001:

# enqueue a js task
curl -sX POST localhost:3001/queues/default/tasks \
  -H 'content-type: application/json' \
  -d '{"type":"js","payload":{"script":"40 + 2"}}'

# the worker bundled in `npm run dev` picks it up within a second
curl -s localhost:3001/queues/default/stats     # {"ready":0,"inFlight":0,"dlq":0}

Run the tests

docker compose up -d     # tests need Postgres (npm run dev also starts it)
npm test                 # Vitest: 53 tests
npm run typecheck        # tsc --noEmit

The required integration test is test/concurrent-claim.test.ts: 50 mixed tasks, 4 concurrent worker loops, asserting at-least-once and no double-claim.

HTTP API

Method Path Body
POST /queues/:name/tasks { type, payload, delay?, dedupeKey? }
POST /queues/:name/claim { workerId, max }
POST /tasks/:id/ack { leaseToken, result? }
POST /tasks/:id/nack { leaseToken, reason, terminal? }
POST /tasks/:id/extend { leaseToken }
POST /tasks/:id/requeue —
GET /tasks/:id —
GET /queues —
GET /queues/:name/stats —
GET /queues/:name/dlq?limit=&offset= —

Errors are { error: { code, message } }. Codes: 400 validation, 404 not-found, 409 state conflict (wrong_status, stale_lease), 500 internal.

What's implemented

  • Claim under concurrency — FOR UPDATE SKIP LOCKED; no two workers ever hold the same task at once (proven by the integration test).
  • At-least-once delivery with lease-based recovery — a dead worker's task is reclaimed automatically once its lease expires (lazily, inside the claim query — no background process).
  • Fencing tokens — a rotating lease_token rejects a stale worker's late ack/nack, including the same-worker-reclaim case.
  • Retry + backoff — exponential with full jitter, configurable per queue; terminal errors skip straight to the DLQ.
  • DLQ + requeue — DLQ is a task status; requeue resets attempts to 0.
  • Idempotent enqueue — dedupeKey with a fixed time window (SQS-style), race-safe via a task_dedupes upsert.
  • Three handlers — js (node:vm, sandboxed, timed out), http (fetch with scheme allowlist, timeout, Idempotency-Key), llm (stub behind an interface, driven by payload._sim).
  • Worker harness — register handlers, claim loop, dispatch, heartbeat.
  • React UI — live queue counts, DLQ drill-in, per-row requeue.

What isn't (deliberately)

  • Exactly-once delivery. Not offered — see below.
  • Real js isolation. node:vm contains accidents and runaway loops, not a determined attacker. Real isolation (a child process / isolated-vm) is noted in DESIGN.md as future work.
  • Full SSRF protection for http tasks (post-DNS private-range blocking).
  • Priority queues, pub/sub fan-out, multi-tenancy, auth beyond a worker id — all explicit non-goals.

Why not exactly-once

A worker can die after it performs a side effect but before it acks. The queue cannot tell a dead worker from a slow one — both are silent — so when the lease expires it hands the task out again, and the effect may run twice. No recovery strategy avoids this; the only alternative is to never retry, which loses tasks. So the queue guarantees at-least-once, and correctness for side-effecting tasks is a shared contract: the handler's effect must be idempotent at the target (e.g. an idempotency key derived from the stable task id). The http handler sends Idempotency-Key automatically to make that concrete. Full reasoning is in DESIGN.md.

About

Self-contained task queue on Postgres: SKIP LOCKED claims, leases with fencing tokens, retry with jittered backoff, DLQ, worker harness, React dashboard

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages