Skip to content

fix(openai): preserve reasoning replay and check stream completion - #407

Open
kevinle128 wants to merge 6 commits into
charmbracelet:mainfrom
kevinle128:fix/openai-reasoning-replay-stream-completion
Open

kevinle128 wants to merge 6 commits into
charmbracelet:mainfrom
kevinle128:fix/openai-reasoning-replay-stream-completion

Conversation

@kevinle128

@kevinle128 kevinle128 commented Oct 1, 2026 •

Copy link
Copy Markdown

Closes #406

Summary

  • Replay encrypted Responses reasoning inline when store=false, once per item ID, with the original summary.
    Use the completed streaming item for final encrypted content and summary metadata.
    Stored requests keep their current behavior.
  • Mark reasoning metadata as Finalized when it comes from the completed output item (Generate output or response.output_item.done).
    Stateless replay skips unfinalized metadata, so partial content from response.output_item.added, or metadata persisted before this field existed, is not sent back.
    The field name and JSON tag (finalized,omitempty) match fix: replay OpenAI Responses reasoning from encrypted content聽coder/fantasy#63, so persisted metadata stays compatible.
  • Add WithLanguageModelRequireFinishReason() so callers can reject Chat Completions streams that close without a finish reason before completed tool calls are published.
    The option is off by default to preserve compatible-provider behavior.
  • Anthropic: keep replayed tool argument numbers exact, including integers above 2^53.
    Input values are decoded as json.RawMessage, because the SDK encoder writes json.Number as a string.
  • Re-record the four OpenAI summary-thinking cassettes against the live API, and update the two Azure cassette request bodies offline.

Before this change, stateless follow-up requests lost encrypted reasoning, and streaming metadata could retain missing or partial encrypted content.
A stream with valid tool arguments could also finish without a terminal reason.
After this change, stateless requests retain the completed reasoning data, and callers can require the terminal reason with the new option.

Validation

  • Focused OpenAI and Anthropic tests pass. The new tests fail when the Finalized gate is removed, and the Anthropic test fails against both the float64 and json.Number decoders.
  • OpenAI cassettes (gpt-5, o4-mini, generate and streaming) re-recorded live: the follow-up requests with inline encrypted reasoning return 200.
  • go test ./... -count=1 -timeout=30m
  • go build ./...
  • go vet ./...
  • git diff --check

The two Azure cassettes were edited offline; I don't have Azure access to re-record them.
golangci-lint was not run locally.

kevinle128 and others added 3 commits October 5, 2026 14:51
Mark reasoning metadata as finalized when it comes from the completed
output item in Generate or from response.output_item.done in Stream.
Stateless replay now skips unfinalized metadata, so partial content from
the added event, or metadata persisted before this field existed, is not
sent back to the API.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Record the OpenAI gpt-5 and o4-mini summary thinking cassettes against
the live API, so the follow-up requests that replay encrypted reasoning
inline are accepted by OpenAI.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Decode replayed tool call input into json.RawMessage values. The SDK
encoder writes json.Number as a string, so the earlier UseNumber decode
sent {"a":"2"} for {"a":2}. Raw values keep each number exact,
including integers above 2^53.

The test now marshals through the SDK params, which is the encoder that
writes the request body.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kevinle128

Copy link
Copy Markdown
Author

A question for the maintainers about the store=true path.

This PR replays encrypted reasoning inline only when store=false. With store=true, reasoning parts are still skipped, as before. coder#63 fixes the same bug and replays finalized reasoning inline regardless of store, so reasoning models keep their reasoning between tool steps and turns even when the stored item is not referenced.

I kept the store=true behavior unchanged here to avoid changing requests for existing stored-response callers. Would you prefer to replay finalized reasoning inline for store=true as well, in this PR or in a follow-up? I'm happy to do either.

The Finalized gate in this PR uses the same field name and JSON tag as coder#63, so metadata persisted by either side stays compatible.

@kevinle128

Copy link
Copy Markdown
Author

The follow-up Responses replay metadata, request overrides, error details, and echoed service-tier changes moved to #412. This PR is restored to its original six commits.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

OpenAI replay drops encrypted reasoning and cannot require stream completion

1 participant