Repository navigation
🤖 fix: keep Anthropic thinking block order in ensureAnthropicThinkingBeforeToolCalls #5887
Description
Activity
- addedapprovedTriage: passed unanimous bug screenTriage: passed unanimous bug screen
on Oct 9, 2026 - added a commit that references this issue
on Oct 9, 2026 Owner: a new provider lane (workspace fec3c27f32), in read-only triage first.
The coordinator desk assigned this owner on 2026-10-09 after an ownership sweep of open issues. If this owner stops, the coordinator desk takes the issue back.
Generated with
xum• Model:anthropic:claude-opus-5-5• Thinking:high#5957 fixes one part of this issue: a tool-call message whose first block is thinking now replays in its original order (
[thinking, text, thinking, tool_use]is kept). Two parts stay open.-
Text-first tool-call messages are still reordered.
- Shape:
[text, thinking(signed), tool_use]goes out as[thinking, text, tool_use].ensureAnthropicThinkingBeforeToolCalls(src/browser/utils/messages/modelMessageTransform.ts) still moves thinking first whencontent[0]is not reasoning. The guard test "still moves thinking first when a tool-call message starts with text" pins this behavior on purpose. - Why it is kept: in manual thinking mode, the API error text says the final assistant message "must start with a thinking block (preceding the lastmost set of
tool_useandtool_resultblocks)". It is not verified whether that rule applies only to the last assistant message of the request or to the first message of the final turn. If Xum stops reordering here without knowing, manual-thinking models can get a 400. - Needed before a fix: a direct-API check of that rule (adaptive models do not need it), or a docs statement. Then restrict the reorder to the message the API actually checks.
- Also unchanged: the step that pulls a reasoning-only previous row into a tool-call message that has no reasoning.
- Shape:
-
Server tools (provider-executed
web_search/web_fetch) interleaved with thinking: not checked.- Xum persists these as tool parts with
providerExecuted: true(src/node/services/streamManager.ts, tool-result and tool-error handling) and converts them withconvertToModelMessages. - A hand-written fixture that put a provider-executed call and its result inside the assistant message lost the server-tool parts and split the thinking:
splitMixedContentMessages/stripOrphanedToolCallsproduced[thinking1]and[thinking2, tool_use]as separate messages. That fixture was not built from a real persisted turn, so it proves nothing yet. - Needed: a replay test that streams a real Anthropic response with
server_tool_use+web_search_tool_resultbetween thinking blocks (asstreamManager.preservedThinking.test.tsdoes for client tools), persists it, rebuilds the next request, and compares the block order.
- Xum persists these as tool parts with
Broader replay drift (not byte-identical replay) is tracked in #5888.
-
- added a commit that references this issue
on Oct 10, 2026 Status after the containment PR #6004 (in the merge queue, "Part of #5887"). This issue stays open.
What #6004 does: when an Anthropic server tool (web search) runs between two thinking blocks, Xum writes the thinking-replay receipt from #5994. Later requests in the context segment then send no Anthropic thinking. Removing all thinking is a valid edit under preserved thinking, so the next turn no longer replays a block bound to a prefix Xum never sends again.
Open follow-ups, kept here so they are not lost:
- Cache cost of the containment is not yet measured. After such a turn, thinking replay stops for the rest of the segment. That costs prompt-cache reuse of the thinking blocks on every web-search-plus-thinking conversation. The real fix is native replay (option A below). Record the measured cost here once option A is decided.
- Upstream acceptance of the client-pair shape is unverified. With thinking dropped, Xum still sends the server tool as a client
tool_use/tool_resultpair namedweb_search, next to theweb_search_20250305server tool. Xum has sent this shape after every web search since that code shipped, and no issue reports a rejection. The 🤖 fix: turn Anthropic thinking replay off after a server tool between thinking blocks #6004 tests use scripted fixtures: they prove the strip, not that the API accepts this shape. This needs one live check.
Option A (native replay): decision still open with the user. A local prototype persists
providerExecutedandencryptedContentfor Anthropic server tools and replaysserver_tool_use+web_search_tool_resultin the original order. The cost: with 5 results of about 4,000 characters each, every later turn sends about 14k more input tokens by the local tokenizer, and each search row grows by about 20 KB. These numbers come from synthetic fixtures. The API's own count, real result sizes and cache rates are not measured. Nothing from the prototype is pushed.The text-first reorder audit from the original issue text is also still open.
Generated with
xum• Model:anthropic:claude-opus-5-5• Thinking:high• Cost:$92.89Follow-ups from #6008 (Fixes #5996: the continuous compaction summary honors a thinking-replay receipt in the retained tail or on the running turn).
- Receipt read once per compaction job. Deferred, low severity. The compactor reads the running turn's receipt when it takes its snapshot (
ContinuousCompactor.readSnapshot). If a thinking-signature repair lands after the snapshot but before the summary request, the summary can still send the head's signed thinking. The cost: that summary request gets one 400, the compactor logs it and stages nothing, and the next eager job reads the receipt and succeeds. No state ends up wrong. - Split test coverage. The compactor test stubs the live snapshot's receipt, and a StreamManager test proves that
getStreamInfoexposes the real one. No single test joins the two ends in one turn.
Generated with
xum• Model:anthropic:claude-opus-5-5• Thinking:high• Cost:$98.00- Receipt read once per compaction job. Deferred, low severity. The compactor reads the running turn's receipt when it takes its snapshot (
Decision (user, 2026-10-10): ship option A (native replay) in addition to the #6004 containment.
The user's condition: preserve the prompt cache. Dropping the search output or the thinking blocks can invalidate the cache, so the native-replay design must keep both, in their original order, on every later turn.
Acceptance criteria for the option A PR (owner: provider lane):
- Xum stores Anthropic server-tool calls and results with
providerExecutedandencryptedContent, and replaysserver_tool_use+web_search_tool_resultin their original position, next to the signed thinking blocks they belong to. No later request in the segment strips thinking because of a server tool. - Prompt-cache proof: a live or recorded two-turn check shows that the second turn's request prefix is byte-identical to the first turn's request plus the new turn, and that the provider reports cache reads for that prefix (not only that no 400 occurs).
- Old history still works: rows written before option A (no
encryptedContent, or a 🤖 fix: turn Anthropic thinking replay off after a server tool between thinking blocks #6004 receipt) keep today's behavior, and the receipt path from 🤖 fix: turn Anthropic thinking replay off after a server tool between thinking blocks #6004 stays as the fallback. - Size: measure the real token and byte cost with the provider's count endpoint (the prototype's synthetic numbers were about +14k input tokens per later turn and about +20 KB per search row). Rows must stay within the chat.jsonl row-size limits.
- Non-Anthropic models and a model switch inside a segment keep working: other providers never receive Anthropic-native parts.
Generated with
xum• Model:anthropic:claude-opus-5-5• Thinking:high• Cost:$1205.72- Xum stores Anthropic server-tool calls and results with
L1 live cache check for option A (#6027): partial verification. Cache reuse across messages is verified. The in-message case (signed thinking after a native search in the same message, the shape of this bug) was not exercised live.
It ran once, on head
800c93b634: Opus 5.5 through the gateway, web searchmaxUses: 2,max_tokens2000, a pinned system prompt (7,656 tokens with tools), and 3 turns. Each turn was rebuilt from stored history by the real StreamManager, HistoryService andassemblePromptPayload. The production options sent adaptive thinking withblock_binding: drop_block, and every request carried thethinking-binding-controlsbeta header. It made 3 requests, with an estimated cost of $0.089.Request HTTP input cache write cache read output searches dropped blocks Turn 1 200 6 13,559 7,652 132 1 none Turn 2 200 4 63 13,559 44 0 none Turn 3 200 4 72 13,622 10 0 none Against the pass criteria:
- No 400: met.
- No dropped blocks on turns 2 and 3: met. No request reported any
input_transformationsentry. Gap: in turn 1 the model ran the search with no thinking blocks around it. So this run did not cover signed thinking after a native search in the same message. It did cover turn 2's thinking: that thinking is bound to a prefix that includes turn 1's native search, and turn 3 replayed it with nothing dropped. - Turn 3 cache read ≥ count_tokens of the turn-2 request: not met as written (short by 4 tokens). Turn 3 read 13,622 tokens, and count_tokens of the turn-2 request is 13,626. The 4 tokens are turn 2's own uncached input (its
input_tokens= 4). Turn 3 read everything turn 2 cached: 13,559 + 63 = 13,622. So cache reuse across messages is verified, but the criterion as I wrote it did not pass.
Cache evidence: turn 1 cached 13,559 tokens, which includes the search result (the API caches across its own search step). Turn 2 read exactly those 13,559 tokens. So turn 2's request, rebuilt from stored history, matched the API's own bytes through the search result.
P1 (free count_tokens): the turn-2 request with the native
server_tool_useandweb_search_tool_resultbut with web_search NOT declared returns 200 (10,866 tokens). The API accepts that shape. This does not prove that signatures are accepted or that the cache is reused.Sizes from this run:
- The search returned 9 results with 13,040 characters of ciphertext in total (752 to 2,864 per result).
- The stored row is 16,959 bytes. The row budget (256 KiB of ciphertext) fits about 20 searches of this size.
- The turn-2 request costs 13,626 tokens and 30,014 bytes when replayed natively. The pre-🤖 fix: keep Anthropic thinking block order in ensureAnthropicThinkingBeforeToolCalls #5887 form of the same request (client pair, without ciphertext) costs 8,351 tokens and 16,726 bytes, so native replay adds 5,275 tokens per later request. Those tokens are cache reads, billed at the cache-read rate. The client-pair form changes the prefix after the search, so the cache does not hold past it.
Next: the coordinator is asking the user for one more bounded run (about $0.10) that targets the in-message case. The other option is to accept the gap explicitly. #6027 stays out of the merge queue until one of the two happens.
Generated with
xum• Model:anthropic:claude-opus-5-5• Thinking:high• Cost:$121.74L1 run 2 (in-message case) for option A (#6027): FAIL. The API dropped one thinking block on each later turn (
thinking_dropped: prefix_binding_mismatch), and turn 2 reused the cache only through the system prompt.Setup: attempt 1 of the 3 allowed. It ran locally on
1f6881e659(the perf-cleared fix, not pushed), with the same harness as L1: Opus 5.5 through the gateway, adaptive thinking withblock_binding: drop_block, the binding beta header on every request, web searchmaxUses: 2,max_tokens2000, and 3 turns. Each turn was rebuilt from stored history by the real StreamManager, HistoryService andassemblePromptPayload. It made 3 requests, with an estimated cost of $0.136. I logged only usage fields, stop reasons and drop counters.The shape was produced. Turn 1 stored
[web_search (native), reasoning (signed), text]: one native search, then signed thinking after it, in the same assistant message.Request HTTP input cache write cache read output searches dropped blocks Turn 1 200 6 14,071 7,686 425 1 none Turn 2 200 4 6,582 7,686 282 0 1 ( prefix_binding_mismatch)Turn 3 200 4 310 14,268 51 0 1 ( prefix_binding_mismatch)Against the pass criteria:
- No 400: met. That is only because of
drop_block. On a route without block binding, the same mismatch returns a 400. - No dropped blocks: not met. One thinking block was dropped on turn 2 and again on turn 3.
- Cache reads cover the prefix: not met. Turn 2 read 7,686 tokens, the same as turn 1's system-and-tools prefix. So the rebuilt request diverged from the API's cached content at or before the search result. (In L1 run 1, without thinking around the search, turn 2 read 13,559 tokens, through the search result.) Turn 3 read 14,268 = 7,686 + 6,582, which is everything turn 2 cached.
Sizes: 10 results with 14,172 characters of ciphertext in total (664 to 2,864 per result), so 0 of 10 were demoted by the 12,000-character limit. The stored row is 24,102 bytes. The turn-2 request costs 14,272 tokens natively, against 8,633 tokens in the pre-#5887 client-pair form.
Cause: not known yet. My logs do not record which block was dropped or the raw block order the API returned. One hypothesis, which I have not verified: the API's response started with a thinking block before the search that Xum did not store. That would explain both the early cache miss and the dropped later thinking. Run 1's search had no thinking around it, and its prefix matched.
Status: #6027 stays held. The #6004 containment stays the behavior. I am not repairing anything until the next step is decided.
Generated with
xum• Model:anthropic:claude-opus-5-5• Thinking:high• Cost:$126.69- No 400: met. That is only because of
Diagnosis of the L1 run 2 failure, and a live re-check of the fix: PASS on the in-message case.
Diagnosis (attempt 2, on
1f6881e659, estimated $0.159)I compared block types and order, plus short content hashes (no content logged), at three points:
Point Assistant blocks, in order Raw API response, turn 1 server_tool_use,web_search_tool_result,thinking(signed), 7 ×text(3 with citations)Stored history row native web_search, signed reasoning, text (same order)Rebuilt turn-2 request thinking,server_tool_use,web_search_tool_result,textThe API dropped
messages.1.content.0(thinking_dropped), which is that moved thinking block, on turns 2 and 3.Cause:
ensureAnthropicThinkingBeforeToolCalls(src/browser/utils/messages/modelMessageTransform.ts) moves every reasoning part to the front of an assistant message that contains a tool call. A natively replayed server-tool call counts as a tool call, so the thinking was moved in front of the search it is bound to. Storage was correct. L1 run 1 passed because its search had no thinking after it.Fix (local commit
9ca04a885c, not pushed)The reorder now keeps the API's order when a message opens with a natively replayed server-tool call. The API itself started that response with
server_tool_use. Messages that start with any other block are unchanged. The final-message rule still inserts an unsigned placeholder in front, and the SDK drops unsigned reasoning, so it never reaches the wire. A new scripted-SSE test replays this exact shape. It was red before the fix, and removing the fix turns it red again.Live re-check (attempt 3, on
9ca04a885c, estimated $0.060)The shape was produced: turn 1 stored
[native web_search, signed reasoning, text]. The rebuilt turn-2 request keptserver_tool_use,web_search_tool_result,thinking,text. Theserver_tool_use, thethinkingand all 10 result items hash the same as in the raw response.Request HTTP input cache write cache read dropped blocks Turn 1 200 6 6,372 15,372 none Turn 2 200 4 213 14,058 none Turn 3 200 4 73 14,271 none - No 400: met.
- No dropped blocks on turns 2 and 3: met.
- Cache reads cover the prefix: met, with the same allowance as L1. count_tokens of the turn-2 request is 14,275. Turn 3 read 14,271, which is everything turn 2 cached (14,058 + 213). The 4 tokens left are turn 2's own uncached input.
Sizes: 10 results with 14,212 characters of ciphertext in total (636 to 2,864 per result), so 0 of 10 were demoted by the 12,000-character limit. P1 (web_search not declared) returned count_tokens 200 again.
Open, out of scope
- Xum merges the 7 text blocks into one and does not store citations. That does not affect this shape, because the text comes after the thinking. But thinking that follows cited text in the same message would still mismatch. That needs citation storage, a separate change.
- The
web_search_tool_resultblock hash differs between the raw response and the request, while every result item hashes the same. So the difference is a block-level field outside the items. The API accepted the block, and the cache held through it.
Spend: about $0.44 across the three L1 runs, inside the $2 bound. #6027 stays held until the hot-path change is reviewed.
Generated with
xum• Model:anthropic:claude-opus-5-5• Thinking:high• Cost:$127.22Option A (native replay of Anthropic web search) is parked: #6027 is a draft, not merged. The #6004 containment stays the behavior: thinking after a server tool is turned off for the rest of the context segment.
What #6027 proved, live (Opus 5.5 through the gateway,
block_binding: drop_block)- Cache reuse across messages (issuecomment-6095556958): with no thinking around the search, the next turn read the whole prefix through the search result from cache.
- Cache reuse within a message, after a fix (issuecomment-6095680394, issuecomment-6095755816): the first in-message check failed, because
ensureAnthropicThinkingBeforeToolCallsmoved the thinking in front of the search. With the reorder fix, the request kept the API's block order, no thinking was dropped on later turns, and the cache covered the prefix except the previous turn's own few uncached tokens. - Sizes: across the live runs, 9-10 results per search, 636 to 2,864 characters of ciphertext each, and about 13,000-14,000 characters per search. Native replay adds about 5,300-5,600 tokens to each later request, billed as cache reads.
What a future attempt needs
- Input validation in the predicate.
isNativeAnthropicReplayablemust require a plain-object input. Otherwise a corrupted stored input is rewritten to{}bysanitizeToolInputs, and the thinking after it mismatches. This is the blocker that parked 🤖 fix: replay Anthropic web search natively next to its thinking blocks #6027 (about 1 line). - Citation storage. Xum merges text blocks and does not store citations. Thinking that follows cited text in the same message would still mismatch.
- The 12,000-character per-result limit. Every build runs the provider-output sanitizer on stored rows, so only results whose fields stay under it can replay byte for byte. Larger results fall back to the containment.
- A fresh review budget. 🤖 fix: replay Anthropic web search natively next to its thinking blocks #6027 used all 6 reviews plus one granted extra pair. A new PR (it can start from
9ca04a885c) needs its own review loop, and the live in-message check rerun on its final head.
Open separately: the text-first reorder audit for this issue.
Generated with
xum• Model:anthropic:claude-opus-5-5• Thinking:high• Cost:$133.33
Problem
ensureAnthropicThinkingBeforeToolCalls(src/browser/utils/messages/modelMessageTransform.ts) moves reasoning parts in front of other parts of an assistant message that has tool calls, and pulls reasoning-only rows from the previous message into it. With interleaved thinking, this changes the order of content inside a replayed message.Anthropic preserved thinking binds each replayed block to everything before it, so a changed order inside an earlier message can invalidate later blocks (400 on enforced accounts, or dropped blocks with
drop_block).Intended fix
Audit the function against https://platform.claude.com/docs/en/build-with-claude/preserved-thinking. Keep the original block order for signed Anthropic blocks, and keep the reorder only where the API requires it. Prove it with a replay test that compares two consecutive request bodies.
Generated with
xum• Model:anthropic:claude-opus-5-5• Thinking:high