Skip to content

fix(openai): preserve Responses replay metadata and error details - #412

Open
kevinle128 wants to merge 11 commits into
charmbracelet:mainfrom
kevinle128:feat/openai-responses-replay-metadata
Open

kevinle128 wants to merge 11 commits into
charmbracelet:mainfrom
kevinle128:feat/openai-responses-replay-metadata

Conversation

@kevinle128

Copy link
Copy Markdown

Closes #411

Depends on #407. Until #407 merges, this PR also shows its six commits; review the last five.

Summary

  • Preserve assistant message item IDs and phase in Generate and text stream metadata.
    With store:false, replay text as an output-message item when an ID is present, and group text parts that share the same ID.
  • Preserve function-call item IDs in Generate and all function-call stream parts.
    Replay an item ID only when it starts with fc_ and storage is disabled.
    Keep ToolCallID as call_id and preserve reasoning/function-call order.
  • Add Responses ExtraBody, applied last with the same semantics as openaicompat.
    A boolean store override also controls prompt conversion and PreviousResponseID validation.
  • Preserve error code, type, message, and actual HTTP status through a typed error, including errors after stream data has arrived.
    Preserve terminal response status and the raw incomplete_details.reason without changing existing normalized finish reasons or retry rules.
  • Expose the terminal echoed service_tier for pricing.
    A missing echo remains empty; the requested or initial tier is not used as a fallback.

New public types and fields:

  • openai.ResponsesTextMetadata: ItemID string, Phase string.
    Registered as openai.TypeResponsesTextMetadata (openai.responses.text_metadata).
  • openai.ResponsesToolCallMetadata: ItemID string.
    Registered as openai.TypeResponsesToolCallMetadata (openai.responses.tool_call_metadata).
  • openai.ResponsesProviderOptions.ExtraBody map[string]any.
  • openai.ResponsesError: Code string, Type string, and embedded *fantasy.ProviderError, which exposes Message and StatusCode.
    errors.As can inspect the Responses error, provider error, and SDK cause when present.
    Error strings contain no request dump or Authorization header; the embedded provider error is excluded from Responses error JSON.
  • openai.ResponsesProviderMetadata.ResponseStatus string, RawFinishReason string, and ServiceTier openai.ServiceTier.

Before this change, stateless text and function-call replay lost item metadata, callers could not set untyped Responses fields, stream errors lost classification details, and finish metadata omitted the raw reason and actual service tier.
After this change, callers can store and restore replay metadata, apply request overrides, inspect errors, and read the terminal reason and tier.
Metadata-free text and store:true replay retain their existing behavior.
Generate, Stream, GenerateObject, and StreamObject share the error and finish metadata handling.

Validation

  • Offline local HTTP tests cover text and function-call capture/replay, metadata JSON storage, message grouping, item-prefix checks, typed-nil metadata, and reasoning/function-call order.
  • Offline tests cover extra-body overrides, storage conflicts, and previous-response validation.
  • Offline tests cover mid-stream errors, HTTP/payload status conflicts, safe error serialization, unknown incomplete reasons, and terminal service-tier precedence.
  • go test ./providers/openai/... -count=1
  • go test ./... -count=1 -timeout=30m
  • go build ./...
  • golangci-lint run (0 issues)
  • VCR request expectations regenerated offline in 24 existing Responses fixtures, using 30 function-call IDs from earlier recorded responses with matching call_id values.
    Recorded response sections are byte-identical; no cassette was re-recorded.

These checks used Go 1.27.0 and golangci-lint 2.12.2, cached dependencies, local HTTP servers, and existing VCR recordings.
No live provider test was run.
Offline tests prove the request shapes and metadata paths but do not prove whether live OpenAI target models require message IDs, phase, or function-call item IDs.
Live service-tier pricing has not been verified.

kevinle128 and others added 11 commits October 1, 2026 16:20
Mark reasoning metadata as finalized when it comes from the completed
output item in Generate or from response.output_item.done in Stream.
Stateless replay now skips unfinalized metadata, so partial content from
the added event, or metadata persisted before this field existed, is not
sent back to the API.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Record the OpenAI gpt-5 and o4-mini summary thinking cassettes against
the live API, so the follow-up requests that replay encrypted reasoning
inline are accepted by OpenAI.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Decode replayed tool call input into json.RawMessage values. The SDK
encoder writes json.Number as a string, so the earlier UseNumber decode
sent {"a":"2"} for {"a":2}. Raw values keep each number exact,
including integers above 2^53.

The test now marshals through the SDK params, which is the encoder that
writes the request body.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

OpenAI Responses drops replay metadata, error details, and echoed service tier

1 participant