OpenRig version: 0.5.17 (f137f11). The probe is unchanged on main (packages/daemon/src/domain/native-resume-probe.ts:124 and :230).
OS, Node and tmux: macOS 26.5.2 (arm64), Node 22.23.1 (daemon), tmux 3.7c
Harnesses involved: Claude Code 2.1.283
What happened
After a daemon restart, the restore of an 11-seat rig ran claude --resume for every Claude seat. Four of them (Opus and Sonnet, all with "bypass permissions on") failed with:
Harness launch failed: Claude resume failed: timed out waiting for Claude to become active
restore.completed recorded those nodes as failed and startup_status became attention_required. The Claude processes came up fine a few seconds later and have been working normally since (runtime-hook activity running/idle, panes mid-turn, queue rows progressing). But rig ps / the TUI keep showing them as attention_required / "needs you", and nothing can clear it:
$ rig seat clear-attention orch1-lead@rigname
not_demonstrably_responsive — Uncleared attention class restore_outcome: pane_not_usable:
Pane state is inconclusive (awaiting_runtime); reconciliation requires resumed.
The same result with --reason "<operator attestation>". Seats whose pane happened to match the TUI heuristic during the restore window came up ready; the four that were still loading a large session did not.
Cause
-
Claude Code 2.1.283 sets its process title to its version string, so tmux #{pane_current_command} reports 2.1.283, not claude:
$ tmux display -p -t orch1-lead@rigname '#{pane_current_command}'
2.1.283
-
assessNativeResumeProbe (claude-code branch) only returns resumed when paneCommand === "claude" (line 124) or looksLikeClaudeTui(paneContent) (line 230), which needs a ❯ prompt plus the text "Claude Code v" or "accept edits on". A seat running "bypass permissions on" (or "auto mode on") shows neither once the welcome banner scrolls off. Everything else falls through to inconclusive / awaiting_runtime.
-
Reproduced by calling the shipped module against live panes:
orch1-lead@rigname cmd=2.1.283 {"status":"inconclusive","code":"awaiting_runtime",...}
↳ same pane content, cmd='claude' → resumed
data1-analyst@rigname cmd=2.1.283 {"status":"inconclusive","code":"awaiting_runtime",...}
↳ same pane content, cmd='claude' → resumed
(The pane-identity reconciler's pane_process layer verifies these same seats as verified, so process lineage already knows the pane is Claude.)
-
Knock-on effects, all from the same probe:
claude-resume times out at restore → node failed + startup_status=attention_required.
ContextMonitor.normalizeStartupStatus re-probes on every poll with the same function → never flips back to ready.
SeatAttentionReconciler.clearAttention delegates a restore.completed-sourced outcome to reconcileNodeRuntimeTruth, which also requires resumed; the --reason attestation branch is never reached for that class, so there is no operator path at all.
deriveNodeLifecycleState ranks the stale startup/restore verdict above live agentActivity, so a seat that is demonstrably working renders as "needs you".
Expected
A seat whose runtime hook is reporting running/idle under the current occupant generation, or whose pane process lineage verifies as Claude, should count as resumed, and an operator attestation should be able to clear a restore-derived attention flag (in the spirit of #7).
Suggested fix
native-resume-probe.ts: for claude-code, treat a pane command matching ^\d+\.\d+\.\d+$ as the Claude process (or use the same process-lineage check the identity reconciler uses), and extend looksLikeClaudeTui to accept the "bypass permissions on" / "auto mode on" footers.
seat-attention-reconciler.ts: let fresh positive hook activity or --reason clear the restore_outcome class, as they already can for startup_status.
Side note (possibly its own issue)
Running rig reconcile-session <seat> --no-launch on a live managed seat as a workaround registered a new occupant tenure (kind=adopt) while the running process still carries the previous OPENRIG_OCCUPANT_GENERATION in its environment. Every later hook event is now judged generation_mismatch, so that seat's agentActivity is permanently unknown until it is relaunched. Refusing adoption of a seat that already has a live managed tenure, or carrying the tenure forward, would avoid that.
Related: #7 (startup readiness is one-shot), #79 (same class of probe/TUI drift on Codex 0.153).
OpenRig version: 0.5.17 (f137f11). The probe is unchanged on
main(packages/daemon/src/domain/native-resume-probe.ts:124and:230).OS, Node and tmux: macOS 26.5.2 (arm64), Node 22.23.1 (daemon), tmux 3.7c
Harnesses involved: Claude Code 2.1.283
What happened
After a daemon restart, the restore of an 11-seat rig ran
claude --resumefor every Claude seat. Four of them (Opus and Sonnet, all with "bypass permissions on") failed with:restore.completedrecorded those nodes asfailedandstartup_statusbecameattention_required. The Claude processes came up fine a few seconds later and have been working normally since (runtime-hook activityrunning/idle, panes mid-turn, queue rows progressing). Butrig ps/ the TUI keep showing them asattention_required/ "needs you", and nothing can clear it:The same result with
--reason "<operator attestation>". Seats whose pane happened to match the TUI heuristic during the restore window came upready; the four that were still loading a large session did not.Cause
Claude Code 2.1.283 sets its process title to its version string, so tmux
#{pane_current_command}reports2.1.283, notclaude:assessNativeResumeProbe(claude-code branch) only returnsresumedwhenpaneCommand === "claude"(line 124) orlooksLikeClaudeTui(paneContent)(line 230), which needs a❯prompt plus the text "Claude Code v" or "accept edits on". A seat running "bypass permissions on" (or "auto mode on") shows neither once the welcome banner scrolls off. Everything else falls through toinconclusive / awaiting_runtime.Reproduced by calling the shipped module against live panes:
(The pane-identity reconciler's
pane_processlayer verifies these same seats asverified, so process lineage already knows the pane is Claude.)Knock-on effects, all from the same probe:
claude-resumetimes out at restore → nodefailed+startup_status=attention_required.ContextMonitor.normalizeStartupStatusre-probes on every poll with the same function → never flips back toready.SeatAttentionReconciler.clearAttentiondelegates arestore.completed-sourced outcome toreconcileNodeRuntimeTruth, which also requiresresumed; the--reasonattestation branch is never reached for that class, so there is no operator path at all.deriveNodeLifecycleStateranks the stale startup/restore verdict above liveagentActivity, so a seat that is demonstrably working renders as "needs you".Expected
A seat whose runtime hook is reporting
running/idleunder the current occupant generation, or whose pane process lineage verifies as Claude, should count asresumed, and an operator attestation should be able to clear a restore-derived attention flag (in the spirit of #7).Suggested fix
native-resume-probe.ts: forclaude-code, treat a pane command matching^\d+\.\d+\.\d+$as the Claude process (or use the same process-lineage check the identity reconciler uses), and extendlooksLikeClaudeTuito accept the "bypass permissions on" / "auto mode on" footers.seat-attention-reconciler.ts: let fresh positive hook activity or--reasonclear therestore_outcomeclass, as they already can forstartup_status.Side note (possibly its own issue)
Running
rig reconcile-session <seat> --no-launchon a live managed seat as a workaround registered a new occupant tenure (kind=adopt) while the running process still carries the previousOPENRIG_OCCUPANT_GENERATIONin its environment. Every later hook event is now judgedgeneration_mismatch, so that seat'sagentActivityis permanentlyunknownuntil it is relaunched. Refusing adoption of a seat that already has a live managed tenure, or carrying the tenure forward, would avoid that.Related: #7 (startup readiness is one-shot), #79 (same class of probe/TUI drift on Codex 0.153).