Conversation
…y Slack paging
pg_cron's job_run_details records 'succeeded' when the HTTP request was
QUEUED by pg_net — not when the Edge Function returned 2xx. A scheduled
function can fail for weeks while cron shows green (observed live: a run
returning {"succeeded":0,"failed":2} logged as succeeded).
This recipe closes the gap:
- ops_http_post_logged captures the pg_net request id of every scheduled
call; ops_harvest_responses copies outcomes before pg_net purges them;
ops_cron_http_failures joins them back deterministically.
- brain-health-monitor Edge Function (hourly): harvest -> critical checks
(cron failures, entity-queue poison, missing embeddings >1h, and — when
editorial-policy hygiene.sql is installed — fingerprint/entity-dedup
regressions) -> Slack pages on breach only, deduped by a 24h per-alert
cooldown -> weekly all-green heartbeat (dead-man switch: monitor death
is distinguishable from health) -> snapshot row for trends.
- Opt-in backup-receipt staleness check; bounded retention purge.
- Every optional dependency is guarded: on a stock install the extra
checks skip with a console warning instead of failing.
Battle-tested pattern: running in production, where the forced-failure
drill (documented in schedule.sql) proved the pager end-to-end.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S9FUbMiGenJMFguwQsQYZa
OB1 PR Gate✅ Folder structure — All files are in allowed directories Result: All 15 checks passed! Ready for human review. Post-Merge TasksThese don't block merge — they're reminders for admins after this PR lands.
|
What
New recipe: an hourly ops probe that makes cron→function outcomes observable and pages Slack only when something is actually wrong.
The problem it solves
cron.job_run_detailsrecordssucceededwhen pg_net queued the HTTP request — not when your Edge Function returned 2xx. A scheduled function can fail every run for weeks while cron shows green. Observed live: a run whose function returned{"succeeded":0,"failed":2}was loggedsucceeded.How
schema.sql—ops_http_post_logged()wrapper stores the pg_net request id of every scheduled call;ops_harvest_responses()copies outcomes out ofnet._http_responsebefore pg_net purges it;ops_cron_http_failuresjoins them back deterministically (bad status /"ok":falsebody / no response after grace). Plus alert-state (24h per-alert page cooldown), health snapshots, and a bounded retention purge.monitor/index.ts— hourly Edge Function: harvest → checks → page on breach → weekly all-green heartbeat (dead-man switch: breach pages within the hour, health heartbeats weekly, missing heartbeat = the monitor itself is dead — every state has a distinguishable signal) → snapshot.hygiene.sql— pairs with [recipes] editorial-policy: add mechanical hygiene block to the auditor #439), and backup staleness (opt-in via a_backup_receiptrow) all skip cleanly when their dependency is absent. A stock install still gets cron-failure + embedding-gap coverage.schedule.sqlincludes the migration template for existing jobs and a documented forced-failure drill — prove the pager works, don't trust theory.Provenance
Running in production for a live brain; the drill paged correctly against a manufactured 404, and the heartbeat proved the green path. Vault-header auth throughout (no
?key=URLs — pairs with #437).🤖 Generated with Claude Code