Repository navigation
🤖 fix: turn Anthropic thinking replay off after a server tool between thinking blocks - #6004
Merged
Merged
Conversation
…hinking blocks (#5887 containment) Xum stores Anthropic server tools as client tool calls without their encrypted results, so history replays them as a tool_use/tool_result pair. Thinking after such a tool is bound to a prefix Xum never sends again. Write the existing thinking-replay receipt when reasoning follows a server tool in the row.
…ests (#5887) A retry that drops the server-tool part left its ID recorded, so every later reasoning delta scanned all parts again. Clear the IDs on the first miss. Adds the scan-count test and the non-Anthropic (OpenAI Responses) guard test.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 78da2be585
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
This was referenced Oct 10, 2026
Merged
yermakoffivan
pushed a commit
to yermakoffivan/mux
that referenced
this pull request
Oct 10, 2026
… in the retained tail (coder#6008) ## Summary The continuous compaction summary now honors the Anthropic thinking-repair receipt (`MuxMetadata.anthropicThinkingReplay: "off"`) when the receipt sits in the retained tail or exists only on the running turn. Before this change, the summarizer looked for the receipt only in the cut head, so it could send signed thinking that the receipt says must stay out. The headless summary path has no reasoning-rejection retry, so that request fails on the signature 400. Fixes coder#5996 ## Background After a thinking-signature repair (coder#5994) or a server tool between thinking blocks (coder#6004), Xum writes the receipt, and later requests in the context segment send no Anthropic thinking. `summarizeContinuousCompaction` passed `replayReceiptMessages` built from the head only. When the receipt was on a tail row, or only in the running turn's in-memory metadata (its partial not yet written), the summary request still carried the head's signed thinking. ## Implementation - `ContinuousCompactor`: `deps.summarize` gets a 4th argument, `receiptRows`. It is the raw segment snapshot (`rows`, the same rows the cut is taken from, model-hidden rows included), so an empty receipt-only row counts too. `readSnapshot` copies the live turn's receipt onto the live row next to `stepStartPartIndices`. - `summarizeContinuousCompaction` takes a required `receiptRows` and passes it as `replayReceiptMessages`. It is required so that a caller cannot silently fall back to the head alone. `continuous.ts` forwards it. - `StreamManager.getStreamInfo` and the session host type expose `initialMetadata.anthropicThinkingReplay`. This is one more field on the snapshot object that `getStreamInfo` already builds. The journal prefix rebuild (`rebuildContinuousPrefix`) is unchanged. It already reads `prefixSourceRows`, which include the tail, and coder#5994's parity test covers it. ## Validation Each test failed before the fix: - Compactor: a receipt on a tail row reaches `summarize`'s `receiptRows`, while the head is only the older rows. - Compactor: a receipt held only by the live stream snapshot reaches `receiptRows`. - Summary path, with an Anthropic pinned model: with only the head as receipt rows, the head's signed thinking is sent (control). With the head plus a tail receipt row, no thinking is sent and the answer text stays. - `StreamManager`: during a real turn, `getStreamInfo` reports no receipt before the server-tool reasoning and `"off"` after it. Mutation checks: removing each piece on its own fails exactly its test. Dropping the forwarding in `continuous.ts` fails `tsc`. The summary-path test uses the file's existing `MockLanguageModelV3` seam, so it asserts on the SDK-level prompt, not on an Anthropic HTTP body. ## Risks Low. The receipt check reads the whole segment, which is already in memory, instead of the head. If a receipt is found wrongly, the summarizer sends no thinking. That costs prompt-cache reuse, not correctness. --- _Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking: `high` • Cost: `$96.87`_ <!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=high costs=96.87 -->
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
When an Anthropic server tool (web search) runs between two thinking blocks, Xum now turns thinking replay off for the rest of the context segment, using the existing replay receipt from #5994. Before this change, the next turn replayed the later thinking block against a prefix that Xum never sends again. On accounts that enforce the prefix check, that request gets a 400 and pays one signature repair. Elsewhere the API drops the block.
Part of #5887. The issue stays open: full native replay of server tools is a separate decision (see "Not in this PR").
Background
Preserved thinking binds each replayed thinking block to everything before it. In one step the API returns
thinking, server_tool_use, web_search_tool_result (encrypted_content), thinking, tool_use. The SDK replays exactly that within the turn.Xum stores the server tool differently, so history cannot rebuild it:
dynamic-toolpart has noproviderExecuted(thetool-callcase inStreamManagernever copies it, andMuxToolPartBasehas no such field).providerExecutedexists only on thetool-call-endevent.stripEncryptedContentremovesencryptedContentfrom the result before storage.The next turn therefore sends the server tool as a client
tool_use/tool_resultpair. The thinking after it is bound to the native blocks, so it cannot be valid.Implementation
StreamManageronly (+42 lines). Nothing new is persisted.tool-callstream case records the call ID in an in-memory set whenproviderExecuted === trueon the Anthropic Messages wire (isAnthropicMessagesModel).appendPartAndEmitstores a reasoning part and a recorded server tool's part is in the row, it setsinitialMetadata.anthropicThinkingReplay = "off". The receipt rides on the partial write that every reasoning append already schedules. The existing receipt readers (messagePipeline,commitPartial, the continuous-compaction rebuild) then strip Anthropic thinking for the rest of the segment.Validation
All tests use the real
HistoryServiceand the real@ai-sdk/anthropic/@ai-sdk/openaiSDKs with scripted SSE. The next turn is built throughassemblePromptPayload.thinking/redacted_thinkingblocks. Everytool_useis answered by the next message, and notool_resultis left over. The answer text is kept.commitPartialand reloaded, gives the same clean request.rebuildContinuousPrefixover the committed row shape matchesassemblePromptPayload, carries no signed thinking, and keeps both tool pairs.Each behavior test failed before its fix. Mutations confirmed the guard tests: without the receipt, without the Anthropic guard, and without the stale-ID drop, exactly the matching test fails.
Risks and known limits
tool_use/tool_resultpair namedweb_search. That is the shape Xum has sent after every web search turn since that code shipped. This PR neither introduces nor worsens it, and no issue reports a rejection. The scripted fixtures prove the strip, not upstream acceptance of the client-pair shape.Not in this PR
Full native replay (persisting
providerExecutedandencryptedContentso history sendsserver_tool_use+web_search_tool_resultin the original order) is being prototyped and measured locally (history size and input tokens), for a user decision. #5887 stays open for it and for the text-first reorder audit.Generated with
xum• Model:anthropic:claude-opus-5-5• Thinking:high• Cost:$86.63