Skip to content

🤖 fix: keep the thinking replay receipt when a step-0 retry fails or a turn rebuilds - #5994

Merged
ThomasK33 merged 3 commits into
mainfrom
fix/thinking-replay-receipt-survives-retry
Oct 10, 2026
Merged

ThomasK33 merged 3 commits into
mainfrom
fix/thinking-replay-receipt-survives-retry

Conversation

@ThomasK33

@ThomasK33 ThomasK33 commented Oct 10, 2026 •

Copy link
Copy Markdown
Member

Summary

The Anthropic thinking-repair receipt (MuxMetadata.anthropicThinkingReplay: "off") now survives a repaired turn that produced no output, and same-turn rebuilds keep the removed thinking out. Before this change, both cases sent the removed thinking again, and the next request paid one more signature 400 and repair.

Fixes #5886

Background

After a thinking-signature 400, Xum strips Anthropic thinking for one retry and writes the receipt, so later turns in the context segment keep that thinking out (preserved thinking, #5841 stack). #5886 listed two gaps:

  1. Receipt lost before any output. If the repair fires at step 0 and the retry also fails (or the app crashes) before a part streams, commitPartial deleted the empty assistant row, and the receipt went with it.
  2. Same-turn rebuilds. The step-0 thinking rebuild (rebuildFirstStepForThinkingLevel) and the refusal model fallback (prepare()) rebuild messages from history, which does not hold the current row's receipt yet.

While writing the failing test, I found a third gap of the same kind: the empty-row filter (filterEmptyAssistantMessages, inside prepareProviderRequestMessages) ran before the receipt check, so the check could not see an empty receipt row even when it was kept.

Implementation

  • HistoryService.commitPartial: when nothing commits and the partial carries the receipt, add only the receipt to the stored row (parts unchanged) instead of deleting it. The message-ID, sequence and orphan checks stay as before. An ordinary empty failed turn without a receipt is still deleted. The kept row has no parts, so it never reaches the provider and does not show as a reply.
  • prepareMessagesForProvider takes an optional replayReceiptMessages: the rows searched for the receipt. All three callers pass the rows from before the empty-row filter (activeContextMessages, the same source the Sonnet 5.5 effort pin uses): assemblePromptPayload, the continuous compaction journal prefix rebuild (rebuildContinuousPrefix), and the continuous compaction summary.
  • StreamManager: the step-0 thinking rebuild and the fallback prepare() result are stripped with stripReasoningReplay(..., "anthropic") when the live turn holds the receipt. There is one source of truth: a StepMessageTracker getter that reads streamInfo.initialMetadata.

Validation

Tests use the real HistoryService and the real Anthropic SDK with scripted HTTP responses:

  • A failed step-0 retry, then commitPartial: the next turn sends no earlier thinking but keeps the earlier answer.
  • A crash-shaped partial (receipt, no error metadata): same.
  • A crash after the receipt already reached the row, before the errored partial was deleted: the next commit keeps the row and its receipt.
  • A fallback over a consumed continuous-compaction prefix sends none of that prefix's thinking (pins the existing 🤖 Map the "off" thinking level to between_tools on Claude Sonnet 5.5 #5086 prefix-swap guard).
  • Control: an empty failed turn without a receipt still leaves no row, and thinking still replays.
  • A step-0 thinking rebuild after the repair sends no thinking (tool loop kept).
  • A model fallback after the repair sends no thinking (tool call kept).
  • Continuous compaction: rebuildContinuousPrefix with an empty receipt row in the kept tail equals assemblePromptPayload (the existing parity rule) and drops the earlier thinking.

Each new behavior test was red before its fix. Removing each fix piece on its own fails exactly its tests.

The continuous compaction summary path has no dedicated test: exercising it needs a full AIService, agent resolution and config. It gets the same one-line change as the journal path, which the parity test covers.

Risks

Low. Receipt detection reads the pre-filter rows instead of the filtered rows (same scan, Anthropic only). commitPartial rewrites the active epoch once more, only when a repaired turn ends with no committed output. If the receipt is wrongly kept, Anthropic thinking replay stops for the rest of the context segment, as it already does after any repair. That costs prompt-cache reuse of the thinking blocks, not correctness.


Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high • Cost: $65.94

…a turn rebuilds

The thinking-repair receipt (anthropicThinkingReplay "off") now survives a turn
without output: commitPartial adds it to the stored empty row instead of deleting
the row, and receipt detection reads the rows from before the empty-row filter.
Same-turn rebuilds (step-0 thinking rebuild, refusal model fallback) strip
Anthropic thinking when the live turn already holds the receipt.

Fixes #5886

_Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking: `high` • Cost: `$59.35`_
The journal prefix rebuild and the continuous compaction summary also run the
empty-row filter before the receipt check. Pass them the pre-filter rows too, so
all three provider-message callers read the receipt the same way (#5886).

_Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking: `high` • Cost: `$63.48`_
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-10T00:57:54.652111Z 2978801 New commits
🔒 Security Review ✅ Completed 2026-10-10T00:55:33.753067Z 2978801 New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 012d2cb5e8

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/node/services/historyService.ts
Comment thread src/node/services/streamManager.ts
Comment thread src/node/services/continuousCompactionSummary.ts
A crash after commitPartial wrote the receipt to the row, but before it deleted
the errored partial, made the next commit delete that row (Codex round 1). Pin
also that a fallback over a consumed continuous-compaction prefix sends none of
the prefix's thinking (the existing #5086 prefix-swap guard strips it).

_Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking: `high` • Cost: `$68.61`_
@ThomasK33
ThomasK33 added this pull request to the merge queue Oct 10, 2026
Merged via the queue into main with commit a26e428 Oct 10, 2026
37 checks passed
@ThomasK33
ThomasK33 deleted the fix/thinking-replay-receipt-survives-retry branch October 10, 2026 01:38
yermakoffivan pushed a commit to yermakoffivan/mux that referenced this pull request Oct 10, 2026
… thinking blocks (coder#6004)

## Summary

When an Anthropic server tool (web search) runs between two thinking
blocks, Xum now turns thinking replay off for the rest of the context
segment, using the existing replay receipt from coder#5994. Before this
change, the next turn replayed the later thinking block against a prefix
that Xum never sends again. On accounts that enforce the prefix check,
that request gets a 400 and pays one signature repair. Elsewhere the API
drops the block.

Part of coder#5887. The issue stays open: full native replay of server tools
is a separate decision (see "Not in this PR").

## Background

Preserved thinking binds each replayed thinking block to everything
before it. In one step the API returns `thinking, server_tool_use,
web_search_tool_result (encrypted_content), thinking, tool_use`. The SDK
replays exactly that within the turn.

Xum stores the server tool differently, so history cannot rebuild it:

- The persisted `dynamic-tool` part has no `providerExecuted` (the
`tool-call` case in `StreamManager` never copies it, and
`MuxToolPartBase` has no such field). `providerExecuted` exists only on
the `tool-call-end` event.
- `stripEncryptedContent` removes `encryptedContent` from the result
before storage.

The next turn therefore sends the server tool as a client
`tool_use`/`tool_result` pair. The thinking after it is bound to the
native blocks, so it cannot be valid.

## Implementation

`StreamManager` only (+42 lines). Nothing new is persisted.

- The `tool-call` stream case records the call ID in an in-memory set
when `providerExecuted === true` on the Anthropic Messages wire
(`isAnthropicMessagesModel`).
- When `appendPartAndEmit` stores a reasoning part and a recorded server
tool's part is in the row, it sets
`initialMetadata.anthropicThinkingReplay = "off"`. The receipt rides on
the partial write that every reasoning append already schedules. The
existing receipt readers (`messagePipeline`, `commitPartial`, the
continuous-compaction rebuild) then strip Anthropic thinking for the
rest of the segment.
- The helper checks the parts, not only the ID set, so a retry that
drops parts also drops the trigger. On a miss it clears the set, so the
parts scan runs at most once per server-tool call, not once per
reasoning delta. Today every retry path keeps existing parts, so
reaching that stale-ID case takes a test-only reset. The bound holds
anyway.

## Validation

All tests use the real `HistoryService` and the real `@ai-sdk/anthropic`
/ `@ai-sdk/openai` SDKs with scripted SSE. The next turn is built
through `assemblePromptPayload`.

- A server tool between thinking blocks: the committed row carries the
receipt. The next request body has no `thinking`/`redacted_thinking`
blocks. Every `tool_use` is answered by the next message, and no
`tool_result` is left over. The answer text is kept.
- Receipt timing and crash: the first partial on disk that holds the
bound thinking already has the receipt. That exact partial, committed
through `commitPartial` and reloaded, gives the same clean request.
- Continuous compaction: `rebuildContinuousPrefix` over the committed
row shape matches `assemblePromptPayload`, carries no signed thinking,
and keeps both tool pairs.
- Controls: a client-tool-only turn writes no receipt and replays every
block as returned. A server tool with no thinking after it writes no
receipt. A provider-executed tool before reasoning on the OpenAI
Responses wire writes no receipt.
- Scan cost: a server-tool part dropped by a reset, then 50 reasoning
deltas, gives at most one parts scan (50 before the fix).

Each behavior test failed before its fix. Mutations confirmed the guard
tests: without the receipt, without the Anthropic guard, and without the
stale-ID drop, exactly the matching test fails.

## Risks and known limits

- **Tradeoff:** after such a turn, Anthropic thinking replay stops for
the rest of the context segment. That costs prompt-cache reuse of the
thinking blocks, not correctness. The preserved-thinking docs list
removing every thinking block as a valid edit.
- **Client-pair limit:** with thinking dropped, the request still sends
the server tool as a client `tool_use`/`tool_result` pair named
`web_search`. That is the shape Xum has sent after every web search turn
since that code shipped. This PR neither introduces nor worsens it, and
no issue reports a rejection. The scripted fixtures prove the strip, not
upstream acceptance of the client-pair shape.
- Regression surface: Anthropic turns that use server tools. Other
providers and client tools are unchanged (covered by the controls
above).

## Not in this PR

Full native replay (persisting `providerExecuted` and `encryptedContent`
so history sends `server_tool_use` + `web_search_tool_result` in the
original order) is being prototyped and measured locally (history size
and input tokens), for a user decision. coder#5887 stays open for it and for
the text-first reorder audit.

---

_Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking:
`high` • Cost: `$86.63`_

<!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=high
costs=86.63 -->
yermakoffivan pushed a commit to yermakoffivan/mux that referenced this pull request Oct 10, 2026
… in the retained tail (coder#6008)

## Summary

The continuous compaction summary now honors the Anthropic
thinking-repair receipt (`MuxMetadata.anthropicThinkingReplay: "off"`)
when the receipt sits in the retained tail or exists only on the running
turn. Before this change, the summarizer looked for the receipt only in
the cut head, so it could send signed thinking that the receipt says
must stay out. The headless summary path has no reasoning-rejection
retry, so that request fails on the signature 400.

Fixes coder#5996

## Background

After a thinking-signature repair (coder#5994) or a server tool between
thinking blocks (coder#6004), Xum writes the receipt, and later requests in
the context segment send no Anthropic thinking.
`summarizeContinuousCompaction` passed `replayReceiptMessages` built
from the head only. When the receipt was on a tail row, or only in the
running turn's in-memory metadata (its partial not yet written), the
summary request still carried the head's signed thinking.

## Implementation

- `ContinuousCompactor`: `deps.summarize` gets a 4th argument,
`receiptRows`. It is the raw segment snapshot (`rows`, the same rows the
cut is taken from, model-hidden rows included), so an empty receipt-only
row counts too. `readSnapshot` copies the live turn's receipt onto the
live row next to `stepStartPartIndices`.
- `summarizeContinuousCompaction` takes a required `receiptRows` and
passes it as `replayReceiptMessages`. It is required so that a caller
cannot silently fall back to the head alone. `continuous.ts` forwards
it.
- `StreamManager.getStreamInfo` and the session host type expose
`initialMetadata.anthropicThinkingReplay`. This is one more field on the
snapshot object that `getStreamInfo` already builds.

The journal prefix rebuild (`rebuildContinuousPrefix`) is unchanged. It
already reads `prefixSourceRows`, which include the tail, and coder#5994's
parity test covers it.

## Validation

Each test failed before the fix:

- Compactor: a receipt on a tail row reaches `summarize`'s
`receiptRows`, while the head is only the older rows.
- Compactor: a receipt held only by the live stream snapshot reaches
`receiptRows`.
- Summary path, with an Anthropic pinned model: with only the head as
receipt rows, the head's signed thinking is sent (control). With the
head plus a tail receipt row, no thinking is sent and the answer text
stays.
- `StreamManager`: during a real turn, `getStreamInfo` reports no
receipt before the server-tool reasoning and `"off"` after it.

Mutation checks: removing each piece on its own fails exactly its test.
Dropping the forwarding in `continuous.ts` fails `tsc`.

The summary-path test uses the file's existing `MockLanguageModelV3`
seam, so it asserts on the SDK-level prompt, not on an Anthropic HTTP
body.

## Risks

Low. The receipt check reads the whole segment, which is already in
memory, instead of the head. If a receipt is found wrongly, the
summarizer sends no thinking. That costs prompt-cache reuse, not
correctness.

---

_Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking:
`high` • Cost: `$96.87`_

<!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=high
costs=96.87 -->
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

🤖 fix: keep the thinking replay receipt when a step-0 retry fails or a turn rebuilds

1 participant