Skip to content

🤖 fix: turn Anthropic thinking replay off after a server tool between thinking blocks - #6004

Merged
ThomasK33 merged 3 commits into
mainfrom
fix/anthropic-thinking-order-server-tools
Oct 10, 2026
Merged

ThomasK33 merged 3 commits into
mainfrom
fix/anthropic-thinking-order-server-tools

Conversation

@ThomasK33

@ThomasK33 ThomasK33 commented Oct 10, 2026 •

Copy link
Copy Markdown
Member

Summary

When an Anthropic server tool (web search) runs between two thinking blocks, Xum now turns thinking replay off for the rest of the context segment, using the existing replay receipt from #5994. Before this change, the next turn replayed the later thinking block against a prefix that Xum never sends again. On accounts that enforce the prefix check, that request gets a 400 and pays one signature repair. Elsewhere the API drops the block.

Part of #5887. The issue stays open: full native replay of server tools is a separate decision (see "Not in this PR").

Background

Preserved thinking binds each replayed thinking block to everything before it. In one step the API returns thinking, server_tool_use, web_search_tool_result (encrypted_content), thinking, tool_use. The SDK replays exactly that within the turn.

Xum stores the server tool differently, so history cannot rebuild it:

  • The persisted dynamic-tool part has no providerExecuted (the tool-call case in StreamManager never copies it, and MuxToolPartBase has no such field). providerExecuted exists only on the tool-call-end event.
  • stripEncryptedContent removes encryptedContent from the result before storage.

The next turn therefore sends the server tool as a client tool_use/tool_result pair. The thinking after it is bound to the native blocks, so it cannot be valid.

Implementation

StreamManager only (+42 lines). Nothing new is persisted.

  • The tool-call stream case records the call ID in an in-memory set when providerExecuted === true on the Anthropic Messages wire (isAnthropicMessagesModel).
  • When appendPartAndEmit stores a reasoning part and a recorded server tool's part is in the row, it sets initialMetadata.anthropicThinkingReplay = "off". The receipt rides on the partial write that every reasoning append already schedules. The existing receipt readers (messagePipeline, commitPartial, the continuous-compaction rebuild) then strip Anthropic thinking for the rest of the segment.
  • The helper checks the parts, not only the ID set, so a retry that drops parts also drops the trigger. On a miss it clears the set, so the parts scan runs at most once per server-tool call, not once per reasoning delta. Today every retry path keeps existing parts, so reaching that stale-ID case takes a test-only reset. The bound holds anyway.

Validation

All tests use the real HistoryService and the real @ai-sdk/anthropic / @ai-sdk/openai SDKs with scripted SSE. The next turn is built through assemblePromptPayload.

  • A server tool between thinking blocks: the committed row carries the receipt. The next request body has no thinking/redacted_thinking blocks. Every tool_use is answered by the next message, and no tool_result is left over. The answer text is kept.
  • Receipt timing and crash: the first partial on disk that holds the bound thinking already has the receipt. That exact partial, committed through commitPartial and reloaded, gives the same clean request.
  • Continuous compaction: rebuildContinuousPrefix over the committed row shape matches assemblePromptPayload, carries no signed thinking, and keeps both tool pairs.
  • Controls: a client-tool-only turn writes no receipt and replays every block as returned. A server tool with no thinking after it writes no receipt. A provider-executed tool before reasoning on the OpenAI Responses wire writes no receipt.
  • Scan cost: a server-tool part dropped by a reset, then 50 reasoning deltas, gives at most one parts scan (50 before the fix).

Each behavior test failed before its fix. Mutations confirmed the guard tests: without the receipt, without the Anthropic guard, and without the stale-ID drop, exactly the matching test fails.

Risks and known limits

  • Tradeoff: after such a turn, Anthropic thinking replay stops for the rest of the context segment. That costs prompt-cache reuse of the thinking blocks, not correctness. The preserved-thinking docs list removing every thinking block as a valid edit.
  • Client-pair limit: with thinking dropped, the request still sends the server tool as a client tool_use/tool_result pair named web_search. That is the shape Xum has sent after every web search turn since that code shipped. This PR neither introduces nor worsens it, and no issue reports a rejection. The scripted fixtures prove the strip, not upstream acceptance of the client-pair shape.
  • Regression surface: Anthropic turns that use server tools. Other providers and client tools are unchanged (covered by the controls above).

Not in this PR

Full native replay (persisting providerExecuted and encryptedContent so history sends server_tool_use + web_search_tool_result in the original order) is being prototyped and measured locally (history size and input tokens), for a user decision. #5887 stays open for it and for the text-first reorder audit.


Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high • Cost: $86.63

…hinking blocks (#5887 containment)

Xum stores Anthropic server tools as client tool calls without their encrypted
results, so history replays them as a tool_use/tool_result pair. Thinking after
such a tool is bound to a prefix Xum never sends again. Write the existing
thinking-replay receipt when reasoning follows a server tool in the row.
…ests (#5887)

A retry that drops the server-tool part left its ID recorded, so every later
reasoning delta scanned all parts again. Clear the IDs on the first miss.
Adds the scan-count test and the non-Anthropic (OpenAI Responses) guard test.
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-10T02:29:56.249006Z b157785 New commits
🔒 Security Review ✅ Completed 2026-10-10T02:30:08.994361Z b157785 New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 78da2be585

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/node/services/streamManager.ts
@ThomasK33
ThomasK33 added this pull request to the merge queue Oct 10, 2026
Merged via the queue into main with commit 8b2c699 Oct 10, 2026
35 checks passed
@ThomasK33
ThomasK33 deleted the fix/anthropic-thinking-order-server-tools branch October 10, 2026 02:51
yermakoffivan pushed a commit to yermakoffivan/mux that referenced this pull request Oct 10, 2026
… in the retained tail (coder#6008)

## Summary

The continuous compaction summary now honors the Anthropic
thinking-repair receipt (`MuxMetadata.anthropicThinkingReplay: "off"`)
when the receipt sits in the retained tail or exists only on the running
turn. Before this change, the summarizer looked for the receipt only in
the cut head, so it could send signed thinking that the receipt says
must stay out. The headless summary path has no reasoning-rejection
retry, so that request fails on the signature 400.

Fixes coder#5996

## Background

After a thinking-signature repair (coder#5994) or a server tool between
thinking blocks (coder#6004), Xum writes the receipt, and later requests in
the context segment send no Anthropic thinking.
`summarizeContinuousCompaction` passed `replayReceiptMessages` built
from the head only. When the receipt was on a tail row, or only in the
running turn's in-memory metadata (its partial not yet written), the
summary request still carried the head's signed thinking.

## Implementation

- `ContinuousCompactor`: `deps.summarize` gets a 4th argument,
`receiptRows`. It is the raw segment snapshot (`rows`, the same rows the
cut is taken from, model-hidden rows included), so an empty receipt-only
row counts too. `readSnapshot` copies the live turn's receipt onto the
live row next to `stepStartPartIndices`.
- `summarizeContinuousCompaction` takes a required `receiptRows` and
passes it as `replayReceiptMessages`. It is required so that a caller
cannot silently fall back to the head alone. `continuous.ts` forwards
it.
- `StreamManager.getStreamInfo` and the session host type expose
`initialMetadata.anthropicThinkingReplay`. This is one more field on the
snapshot object that `getStreamInfo` already builds.

The journal prefix rebuild (`rebuildContinuousPrefix`) is unchanged. It
already reads `prefixSourceRows`, which include the tail, and coder#5994's
parity test covers it.

## Validation

Each test failed before the fix:

- Compactor: a receipt on a tail row reaches `summarize`'s
`receiptRows`, while the head is only the older rows.
- Compactor: a receipt held only by the live stream snapshot reaches
`receiptRows`.
- Summary path, with an Anthropic pinned model: with only the head as
receipt rows, the head's signed thinking is sent (control). With the
head plus a tail receipt row, no thinking is sent and the answer text
stays.
- `StreamManager`: during a real turn, `getStreamInfo` reports no
receipt before the server-tool reasoning and `"off"` after it.

Mutation checks: removing each piece on its own fails exactly its test.
Dropping the forwarding in `continuous.ts` fails `tsc`.

The summary-path test uses the file's existing `MockLanguageModelV3`
seam, so it asserts on the SDK-level prompt, not on an Anthropic HTTP
body.

## Risks

Low. The receipt check reads the whole segment, which is already in
memory, instead of the head. If a receipt is found wrongly, the
summarizer sends no thinking. That costs prompt-cache reuse, not
correctness.

---

_Generated with `xum` • Model: `anthropic:claude-opus-5-5` • Thinking:
`high` • Cost: `$96.87`_

<!-- mux-attribution: model=anthropic:claude-opus-5-5 thinking=high
costs=96.87 -->
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant