Repository navigation
🤖 fix: keep the thinking replay receipt when a step-0 retry fails or a turn rebuilds - #5994
Merged
Merged
Conversation
…a turn rebuilds The thinking-repair receipt (anthropicThinkingReplay "off") now survives a turn without output: commitPartial adds it to the stored empty row instead of deleting the row, and receipt detection reads the rows from before the empty-row filter. Same-turn rebuilds (step-0 thinking rebuild, refusal model fallback) strip Anthropic thinking when the live turn already holds the receipt. Fixes #5886 _Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking: `high` • Cost: `$59.35`_
The journal prefix rebuild and the continuous compaction summary also run the empty-row filter before the receipt check. Pass them the pre-filter rows too, so all three provider-message callers read the receipt the same way (#5886). _Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking: `high` • Cost: `$63.48`_
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 012d2cb5e8
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
A crash after commitPartial wrote the receipt to the row, but before it deleted the errored partial, made the next commit delete that row (Codex round 1). Pin also that a fallback over a consumed continuous-compaction prefix sends none of the prefix's thinking (the existing #5086 prefix-swap guard strips it). _Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking: `high` • Cost: `$68.61`_
This was referenced Oct 10, 2026
yermakoffivan
pushed a commit
to yermakoffivan/mux
that referenced
this pull request
Oct 10, 2026
… thinking blocks (coder#6004) ## Summary When an Anthropic server tool (web search) runs between two thinking blocks, Xum now turns thinking replay off for the rest of the context segment, using the existing replay receipt from coder#5994. Before this change, the next turn replayed the later thinking block against a prefix that Xum never sends again. On accounts that enforce the prefix check, that request gets a 400 and pays one signature repair. Elsewhere the API drops the block. Part of coder#5887. The issue stays open: full native replay of server tools is a separate decision (see "Not in this PR"). ## Background Preserved thinking binds each replayed thinking block to everything before it. In one step the API returns `thinking, server_tool_use, web_search_tool_result (encrypted_content), thinking, tool_use`. The SDK replays exactly that within the turn. Xum stores the server tool differently, so history cannot rebuild it: - The persisted `dynamic-tool` part has no `providerExecuted` (the `tool-call` case in `StreamManager` never copies it, and `MuxToolPartBase` has no such field). `providerExecuted` exists only on the `tool-call-end` event. - `stripEncryptedContent` removes `encryptedContent` from the result before storage. The next turn therefore sends the server tool as a client `tool_use`/`tool_result` pair. The thinking after it is bound to the native blocks, so it cannot be valid. ## Implementation `StreamManager` only (+42 lines). Nothing new is persisted. - The `tool-call` stream case records the call ID in an in-memory set when `providerExecuted === true` on the Anthropic Messages wire (`isAnthropicMessagesModel`). - When `appendPartAndEmit` stores a reasoning part and a recorded server tool's part is in the row, it sets `initialMetadata.anthropicThinkingReplay = "off"`. The receipt rides on the partial write that every reasoning append already schedules. The existing receipt readers (`messagePipeline`, `commitPartial`, the continuous-compaction rebuild) then strip Anthropic thinking for the rest of the segment. - The helper checks the parts, not only the ID set, so a retry that drops parts also drops the trigger. On a miss it clears the set, so the parts scan runs at most once per server-tool call, not once per reasoning delta. Today every retry path keeps existing parts, so reaching that stale-ID case takes a test-only reset. The bound holds anyway. ## Validation All tests use the real `HistoryService` and the real `@ai-sdk/anthropic` / `@ai-sdk/openai` SDKs with scripted SSE. The next turn is built through `assemblePromptPayload`. - A server tool between thinking blocks: the committed row carries the receipt. The next request body has no `thinking`/`redacted_thinking` blocks. Every `tool_use` is answered by the next message, and no `tool_result` is left over. The answer text is kept. - Receipt timing and crash: the first partial on disk that holds the bound thinking already has the receipt. That exact partial, committed through `commitPartial` and reloaded, gives the same clean request. - Continuous compaction: `rebuildContinuousPrefix` over the committed row shape matches `assemblePromptPayload`, carries no signed thinking, and keeps both tool pairs. - Controls: a client-tool-only turn writes no receipt and replays every block as returned. A server tool with no thinking after it writes no receipt. A provider-executed tool before reasoning on the OpenAI Responses wire writes no receipt. - Scan cost: a server-tool part dropped by a reset, then 50 reasoning deltas, gives at most one parts scan (50 before the fix). Each behavior test failed before its fix. Mutations confirmed the guard tests: without the receipt, without the Anthropic guard, and without the stale-ID drop, exactly the matching test fails. ## Risks and known limits - **Tradeoff:** after such a turn, Anthropic thinking replay stops for the rest of the context segment. That costs prompt-cache reuse of the thinking blocks, not correctness. The preserved-thinking docs list removing every thinking block as a valid edit. - **Client-pair limit:** with thinking dropped, the request still sends the server tool as a client `tool_use`/`tool_result` pair named `web_search`. That is the shape Xum has sent after every web search turn since that code shipped. This PR neither introduces nor worsens it, and no issue reports a rejection. The scripted fixtures prove the strip, not upstream acceptance of the client-pair shape. - Regression surface: Anthropic turns that use server tools. Other providers and client tools are unchanged (covered by the controls above). ## Not in this PR Full native replay (persisting `providerExecuted` and `encryptedContent` so history sends `server_tool_use` + `web_search_tool_result` in the original order) is being prototyped and measured locally (history size and input tokens), for a user decision. coder#5887 stays open for it and for the text-first reorder audit. --- _Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking: `high` • Cost: `$86.63`_ <!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=high costs=86.63 -->
yermakoffivan
pushed a commit
to yermakoffivan/mux
that referenced
this pull request
Oct 10, 2026
… in the retained tail (coder#6008) ## Summary The continuous compaction summary now honors the Anthropic thinking-repair receipt (`MuxMetadata.anthropicThinkingReplay: "off"`) when the receipt sits in the retained tail or exists only on the running turn. Before this change, the summarizer looked for the receipt only in the cut head, so it could send signed thinking that the receipt says must stay out. The headless summary path has no reasoning-rejection retry, so that request fails on the signature 400. Fixes coder#5996 ## Background After a thinking-signature repair (coder#5994) or a server tool between thinking blocks (coder#6004), Xum writes the receipt, and later requests in the context segment send no Anthropic thinking. `summarizeContinuousCompaction` passed `replayReceiptMessages` built from the head only. When the receipt was on a tail row, or only in the running turn's in-memory metadata (its partial not yet written), the summary request still carried the head's signed thinking. ## Implementation - `ContinuousCompactor`: `deps.summarize` gets a 4th argument, `receiptRows`. It is the raw segment snapshot (`rows`, the same rows the cut is taken from, model-hidden rows included), so an empty receipt-only row counts too. `readSnapshot` copies the live turn's receipt onto the live row next to `stepStartPartIndices`. - `summarizeContinuousCompaction` takes a required `receiptRows` and passes it as `replayReceiptMessages`. It is required so that a caller cannot silently fall back to the head alone. `continuous.ts` forwards it. - `StreamManager.getStreamInfo` and the session host type expose `initialMetadata.anthropicThinkingReplay`. This is one more field on the snapshot object that `getStreamInfo` already builds. The journal prefix rebuild (`rebuildContinuousPrefix`) is unchanged. It already reads `prefixSourceRows`, which include the tail, and coder#5994's parity test covers it. ## Validation Each test failed before the fix: - Compactor: a receipt on a tail row reaches `summarize`'s `receiptRows`, while the head is only the older rows. - Compactor: a receipt held only by the live stream snapshot reaches `receiptRows`. - Summary path, with an Anthropic pinned model: with only the head as receipt rows, the head's signed thinking is sent (control). With the head plus a tail receipt row, no thinking is sent and the answer text stays. - `StreamManager`: during a real turn, `getStreamInfo` reports no receipt before the server-tool reasoning and `"off"` after it. Mutation checks: removing each piece on its own fails exactly its test. Dropping the forwarding in `continuous.ts` fails `tsc`. The summary-path test uses the file's existing `MockLanguageModelV3` seam, so it asserts on the SDK-level prompt, not on an Anthropic HTTP body. ## Risks Low. The receipt check reads the whole segment, which is already in memory, instead of the head. If a receipt is found wrongly, the summarizer sends no thinking. That costs prompt-cache reuse, not correctness. --- _Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking: `high` • Cost: `$96.87`_ <!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=high costs=96.87 -->
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The Anthropic thinking-repair receipt (
MuxMetadata.anthropicThinkingReplay: "off") now survives a repaired turn that produced no output, and same-turn rebuilds keep the removed thinking out. Before this change, both cases sent the removed thinking again, and the next request paid one more signature 400 and repair.Fixes #5886
Background
After a thinking-signature 400, Xum strips Anthropic thinking for one retry and writes the receipt, so later turns in the context segment keep that thinking out (preserved thinking, #5841 stack). #5886 listed two gaps:
commitPartialdeleted the empty assistant row, and the receipt went with it.rebuildFirstStepForThinkingLevel) and the refusal model fallback (prepare()) rebuild messages from history, which does not hold the current row's receipt yet.While writing the failing test, I found a third gap of the same kind: the empty-row filter (
filterEmptyAssistantMessages, insideprepareProviderRequestMessages) ran before the receipt check, so the check could not see an empty receipt row even when it was kept.Implementation
HistoryService.commitPartial: when nothing commits and the partial carries the receipt, add only the receipt to the stored row (parts unchanged) instead of deleting it. The message-ID, sequence and orphan checks stay as before. An ordinary empty failed turn without a receipt is still deleted. The kept row has no parts, so it never reaches the provider and does not show as a reply.prepareMessagesForProvidertakes an optionalreplayReceiptMessages: the rows searched for the receipt. All three callers pass the rows from before the empty-row filter (activeContextMessages, the same source the Sonnet 5.5 effort pin uses):assemblePromptPayload, the continuous compaction journal prefix rebuild (rebuildContinuousPrefix), and the continuous compaction summary.StreamManager: the step-0 thinking rebuild and the fallbackprepare()result are stripped withstripReasoningReplay(..., "anthropic")when the live turn holds the receipt. There is one source of truth: aStepMessageTrackergetter that readsstreamInfo.initialMetadata.Validation
Tests use the real
HistoryServiceand the real Anthropic SDK with scripted HTTP responses:commitPartial: the next turn sends no earlier thinking but keeps the earlier answer.rebuildContinuousPrefixwith an empty receipt row in the kept tail equalsassemblePromptPayload(the existing parity rule) and drops the earlier thinking.Each new behavior test was red before its fix. Removing each fix piece on its own fails exactly its tests.
The continuous compaction summary path has no dedicated test: exercising it needs a full
AIService, agent resolution and config. It gets the same one-line change as the journal path, which the parity test covers.Risks
Low. Receipt detection reads the pre-filter rows instead of the filtered rows (same scan, Anthropic only).
commitPartialrewrites the active epoch once more, only when a repaired turn ends with no committed output. If the receipt is wrongly kept, Anthropic thinking replay stops for the rest of the context segment, as it already does after any repair. That costs prompt-cache reuse of the thinking blocks, not correctness.Generated with
xum• Model:anthropic:claude-opus-5-5• Thinking:high• Cost:$65.94