Skip to content

🤖 fix: replay Anthropic web search natively next to its thinking blocks - #6027

Draft
ThomasK33 wants to merge 6 commits into
mainfrom
fix/anthropic-native-server-tool-replay
Draft

ThomasK33 wants to merge 6 commits into
mainfrom
fix/anthropic-native-server-tool-replay

Conversation

@ThomasK33

@ThomasK33 ThomasK33 commented Oct 10, 2026 •

Copy link
Copy Markdown
Member

Summary

Xum now replays a successful Anthropic web_search from history in its original position: server_tool_use and web_search_tool_result (with encrypted_content), next to the signed thinking blocks around them. Before this change, history replayed the search as a client tool_use/tool_result pair. The #6004 receipt then had to drop all thinking for the rest of the context segment, and the cached prefix changed. With native replay, later requests keep the thinking and resend the same prefix.

Part of #5887. The issue stays open for the text-first reorder audit.

Background

Preserved thinking binds each thinking block to everything before it. In one step the API returns thinking, server_tool_use, web_search_tool_result, thinking, .... The SDK replays exactly that within the turn. Xum's stored part lost two things: the providerExecuted flag, and the result's encryptedContent (removed by stripEncryptedContent). So the next turn could only send the client pair. #6004 contained this by writing the thinking-replay receipt. This PR keeps the native identity, so a successful search no longer needs the receipt.

Implementation

  • Storage (StreamManager): the tool-call part of an Anthropic Messages server tool stores providerExecuted: true. The result keeps its ciphertext only when the stored call part carries the flag. toStoredServerToolPart (in the new src/common/utils/messages/anthropicNativeServerTools.ts) decides the stored form when the result arrives. The part keeps the flag only when history can replay it natively (isNativeAnthropicReplayable). Null titles are stored as null, because the SDK stream omits them and its replay schema requires them.
  • Receipt trigger: the 🤖 fix: turn Anthropic thinking replay off after a server tool between thinking blocks #6004 trigger now counts only server tools that are not native-replayable, using the same predicate. The scan bound is unchanged.
  • Request projection (messagePipeline.ts): projectAnthropicServerTools keeps replayable parts native only when providerForMessages === "anthropic". Every other provider-executed part is demoted to the client pair without ciphertext. If a demoted part has reasoning after it in the same row, the request sends no Anthropic thinking. That flag is a pure function of the rows, so every request in the segment decides the same way. Rows with no server tool are returned as they are, with no new allocation.
  • Splitter, reorder and validator (modelMessageTransform.ts): splitMixedContentMessages keeps provider-executed calls inside their message. ensureAnthropicThinkingBeforeToolCalls keeps the API's order for a message that opens with a provider-executed call: the API itself started that response with server_tool_use, and moving the thinking in front of the search breaks the thinking's binding. validateAnthropicCompliance skips provider-executed calls, because their result rides in the same message.
  • UI: WebSearchToolCall hides encryptedContent in its expanded JSON. tool-call-end carries the stored output, so dropped ciphertext never crosses IPC.
  • stripEncryptedContent moved from src/node to src/common (git mv) so the common projection can use it.

What still replays as the client pair (today's behavior, plus the receipt)

  1. A web search with an error result, and every other server tool (web_fetch, code_execution, ...).
  2. A search that would take the row's total encryptedContent above ANTHROPIC_NATIVE_SERVER_TOOL_MAX_ROW_CIPHERTEXT_CHARS (256 KiB summed over every native search in the row, src/constants/anthropicServerTools.ts), or any single result whose fields the 12,000-character sanitizer would rewrite. The partial holding it is rewritten on every throttled write for the rest of the turn, the row crosses IPC, and chat.jsonl readers skip rows over 1 MiB.
  3. A search that Claude called in parallel with a client tool. The response ends after both calls, and the search result opens the next step. One stored part cannot replay the call and the result at their two positions.
  4. Requests on non-Anthropic wires (for example a model switch to OpenAI): the part is demoted and its ciphertext stripped.
  5. Old rows: rows stored before this change have no flag and keep today's shape. A flagged row that lost its ciphertext is demoted, and the request sends no thinking.

Known limits

  • Results with long ciphertext keep the 🤖 fix: turn Anthropic thinking replay off after a server tool between thinking blocks #6004 containment. Native replay sends a result back unchanged, and every build still runs the 12,000-character provider-output sanitizer on stored rows. So a search where any single result has more than 12,000 characters of encryptedContent, or any field the sanitizer would rewrite, is stored as the client pair. The receipt then drops the thinking after it, and the prompt cache does not hold past that search. Option A preserves thinking and the cache only for searches that fit. In L1, the 9 real results were 752 to 2,864 characters each, so 0 of 9 were demoted by the 12,000-character limit.
  • The continuous-compaction summary keeps the client pair until P1. The summary request declares no tools, and I have not proven that the API accepts undeclared native server-tool blocks. P1 is the free count_tokens probe for that, and it needs real ciphertext from L1. The summary is headless and one-shot, so this costs no turn cache.
  • Turning web search off mid-segment is undecided until P1. If P1 shows that undeclared native blocks are rejected, the main path needs a decision for that case.
  • Live checks (Opus 5.5 through the gateway, block_binding: drop_block):
  • Citations are not stored, and text blocks are merged. Thinking that follows cited text in the same message would still mismatch. In the live runs the text came after the thinking, so it did not matter. Fixing that needs citation storage, a separate change.

Validation

All tests drive scripted SSE through the real Anthropic (and OpenAI) SDKs, with a real HistoryService, and read the request body the SDK sends.

  • T1 native: the next turn's request equals the SDK's in-turn reference. No receipt is written. The stored part has the flag and the ciphertext. A null-title variant replays without throwing.
  • T2 cache content over 3 turns: system and tools are equal across requests. Each request's messages equal the previous request's messages, plus the previous reply exactly as the API returned it, plus the new user turn. cache_control is compared separately: system and tools keep the same breakpoints, and the message breakpoint sits only on the newest block.
  • T3 fallbacks: the OpenAI wire gets the client pair without ciphertext. A flagged row without ciphertext gets the client pair and no thinking. The summary gets the client pair. The prefix rebuild keeps native, in parity with assemblePromptPayload.
  • T4 old rows: a client-pair row from before this change keeps today's request shape.
  • T5 non-native: a failed search between thinking blocks writes the receipt and is stored without the flag. A search whose result opens the next step (parallel client call) is stored as the client pair, writes the receipt, and its tool-call-end event carries no ciphertext. The existing 🤖 fix: turn Anthropic thinking replay off after a server tool between thinking blocks #6004 and 🤖 fix: continuous compaction summary misses a replay receipt in the retained tail #5996 tests now use a non-native search, so they still guard the receipt path.
  • Ciphertext limit: unit tests on both sides of the limit (the sum is per call, across results).
  • Red first: on main's production code, the native, null-title, T2 and lost-ciphertext tests fail.
  • Mutations: I removed each piece on its own, and only the matching tests failed. The pieces: store flag, keep ciphertext, receipt predicate, wire gate, derived strip, splitter, ciphertext strip on demote, summary flag, stored demotion, stored null title, ciphertext limit, split-step demotion, stored output on the event, and the validator skip.
  • Downgrade probe: I reverted messagePipeline.ts and modelMessageTransform.ts to main and fed them a new-format native row. The request builds without throwing. The old splitter keeps native blocks when no client tool follows them, and splits or drops them when one does: degraded, not bricked. Rows that fail the predicate are stored without the flag, so an older build's SDK never sees a flagged result it cannot validate.

Risks

Medium, Anthropic only. The change touches the stored shape of server-tool parts and the provider request for every Anthropic request that contains one. If the API rejects a native replay that the scripted fixtures accept, the next request in that segment fails. I have not verified that the existing thinking-signature repair recovers from that error. L1 is the check for this. Other providers and client tools are unchanged: their parts carry no flag, and the projection returns their rows as they are.


Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high • Cost: $115.89

…#5887 option A)

Store providerExecuted and the encrypted results for successful Anthropic
web_search calls, and replay server_tool_use + web_search_tool_result in place,
so later requests keep the signed thinking after a search and the prompt cache
holds. Other server tools and other providers keep the client pair; the #6004
receipt stays the fallback for them.
…he client pair (#5887)

Perf-owner condition: a web search whose summed encryptedContent exceeds
ANTHROPIC_NATIVE_SERVER_TOOL_MAX_CIPHERTEXT_CHARS (256 KiB) is stored as the
pre-#5887 client pair without ciphertext, and the receipt keeps the thinking
after it out.

A search called in parallel with a client tool gets its result at the start
of the next step. One stored part cannot replay the call and the result at
their two positions, so it is stored as the client pair too.

tool-call-end now carries the stored output, so dropped ciphertext does not
cross IPC. validateAnthropicCompliance skips provider-executed calls, whose
result rides in the same message. The projection skips rows with no server
tool without allocating.
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-10T08:49:38.617963Z 9ca04a8 New commits
🔒 Security Review ✅ Completed 2026-10-10T08:45:44.077558Z 9ca04a8 New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 7c546df6ea

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/node/services/messagePipeline.ts
Comment thread src/common/utils/messages/anthropicNativeServerTools.ts

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9410c83879

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/browser/utils/messages/applyToolOutputRedaction.ts Outdated
Comment thread src/common/utils/messages/anthropicNativeServerTools.ts Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 800c93b634

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/node/services/streamManager.ts
Comment thread src/common/utils/messages/anthropicNativeServerTools.ts Outdated
@ThomasK33

Copy link
Copy Markdown
Member Author

Paused at the review cap. The 6 Codex reviews (3 code-plus-security pairs) are used up. The last pair found 2 blockers. Their fix (1f6881e659, local, not yet pushed) narrows native replay to results that the 12,000-character provider-output sanitizer leaves unchanged, and requires pageAge. This PR waits for a decision on one final review pair on that head before anything else happens. If that pair finds another valid blocker in this mechanism, the work stops here, and the #6004 containment stays the behavior.


Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high • Cost: $125.74

@ThomasK33

Copy link
Copy Markdown
Member Author

Final review pair. One extra code-plus-security review pair was granted for this PR (reviews 7 and 8), on the pushed head. It is an extension, not a budget reset. If this pair finds any valid blocker, the PR is parked with no further repair cycle, and the #6004 containment stays the behavior.

This head carries the round-3 fixes and the reorder fix found by the live in-message check (#5887 issuecomment-6095755816).


Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high • Cost: $130.73

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9ca04a885c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +22 to +26
if (part.providerExecuted !== true || part.toolName !== "web_search") return false;
if (part.state !== "output-available") return false;
// The SDK validates every native result against a schema that requires encryptedContent and
// a string-or-null title: one missing field throws, and no request is sent at all.
return Array.isArray(part.output) && part.output.every(isReplayableWebSearchResult);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Demote native searches with malformed inputs

When a persisted providerExecuted search has a corrupted non-object input, sanitizeToolInputs() rewrites it to {}, but this predicate validates only the output and still replays the part natively. Any later signed thinking is then sent against an altered server_tool_use prefix, causing Anthropic signature validation to reject subsequent turns; validate the input before replay or demote the search so the existing reasoning-stripping recovery runs.

AGENTS.md reference: AGENTS.md:L128-L130

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Valid. This is the blocker that parks this PR, so it is not fixed here.

  • isNativeAnthropicReplayable checks only the output. A stored providerExecuted search with a non-object input passes it, and sanitizeToolInputs (src/browser/utils/messages/sanitizeToolInput.ts) then rewrites that input to {}. So the request sends an altered server_tool_use, and the signed thinking after it no longer matches.
  • Trigger: only corrupted stored input. The stream stores the object the API returned, and in the live re-check the replayed server_tool_use was identical to the raw response block.
  • Impact: on block_binding: drop_block routes, the API drops that thinking (degraded, no 400). On other routes, the request gets one 400, then the existing signature repair writes the replay receipt and the workspace continues.
  • Fix for a future attempt: require a plain-object input in the predicate, so such a part is demoted and its thinking stripped (about 1 line).

Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high • Cost: $133.33

@ThomasK33

Copy link
Copy Markdown
Member Author

Parked. The final granted review pair (reviews 7 and 8) on 9ca04a885c found a valid blocker: a stored native search with a corrupted, non-object input still replays natively, and its input is rewritten before sending (thread on anthropicNativeServerTools.ts). Under the stop rule set for this pair, this PR is not merged and gets no further repair cycle here. The #6004 containment stays the behavior.

The branch stays at 9ca04a885c as a draft. What this PR proved, and what a future attempt needs, is summarized on #5887.


Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high • Cost: $133.33

@ThomasK33
ThomasK33 marked this pull request as draft October 10, 2026 08:53

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant