Skip to content

Copilot harness: choose the wire API from AWF endpoint data, fail fast on a model/endpoint mismatch, and report mismatch errors clearly #67485

Description

@SivaKesava1

Summary

pelikhan asked for better diagnostics for model misconfiguration in general, after a customer report (Slack, 9–10 Oct). A workflow with model: auto failed mid-run with:

400 Cannot translate Copilot request feature 'tools[custom]' between Responses and Chat Completions. Routing model "gpt-5.6-luna-utility" to /responses is incompatible

On older AWF versions, the same mismatch shows as Unsupported Responses custom tool '(unknown)'.

This issue proposes three changes to the Copilot harness:

  1. Choose the wire API from AWF endpoint data instead of guessing it from the model name.
  2. Fail before the CLI starts when the model and wire API don't match.
  3. When a mismatch still happens mid-run, classify it as a model misconfiguration, don't retry it, and report the cause and fix.

Cause

When AWF routing hasn't selected a model, actions/setup/js/copilot_harness.cjs does two things:

  1. applyCopilotModelAliasResolution resolves the alias. In the customer's run, auto became gpt-5.6-luna-utility.
  2. applyCopilotWireAPI picks COPILOT_PROVIDER_WIRE_API from the model name, trying these sources in order:
    • The catalog wire_api, looked up with getCatalogModelEntry in awf_reflect.cjs. That lookup includes the -utility → base-model fallback from Infer Copilot wire API for utility model variants #64177.
    • A gpt-5+ name regex.
    • If neither matches, it sets nothing, so the CLI uses its default, /chat/completions.

The Copilot CLI uses this one wire API for the whole session, including sub-agents. If the guess is wrong, startup still succeeds. The run fails later, when AWF's api-proxy can't translate a request.

The customer was on v0.89.21, which predates #64177 and the gpt-5+ rule. Even so, the name-based approach has needed a per-model fix four times: #60054, #62893, #63117 (fixed by #64177), and #64965. Today gh-aw's own PR Code Quality Reviewer hit the same class of failure in a sub-agent (#67460): a claude-haiku-5.5 sub-agent ran under a GPT main model and failed with Cannot translate … 'include'. That is tracked upstream in github/copilot-cli#5103.

The data needed to avoid the guess already exists:

  • On the AWF-routed path, applyCopilotRoutingSelection uses the endpoint AWF selected. awf_model_routing.cjs checks that endpoint against the model's routing_models[].supported_endpoints, using getAWFRoutingModel, getAWFRoutingModelAmbiguityError, and candidate_metadata_complete.
  • pi_models_json.cjs already rejects an API that isn't listed in supported_endpoints.
  • The alias and name path doesn't use this data.

The mid-run 400 is also retried. HTTP_400_RESPONSE_ERROR_PATTERN in copilot_harness.cjs (mirrored in agent_error_patterns.cjs) doesn't match these messages, so they fall through to the generic retry path.

Reproduction

Tested in a private sandbox on main at 540e47a:

  • model: gpt-5.6-luna with engine.env setting COPILOT_PROVIDER_WIRE_API: completions fails with:

    400 Cannot translate Copilot request feature 'tools[custom]' between Responses and Chat Completions. Routing model "gpt-5.6-luna" to /responses is incompatible: this request needs /chat/completions to preserve 'tools[custom]'…

  • The harness retries the failure. The log shows [copilot-harness] attempt 1: outputTail=… followed by more attempts.
  • model: auto resolved to claude-sonnet-5 in that org and the run worked.

Plan

1. Choose or verify the wire API from AWF data

  • After applyCopilotModelAliasResolution, look up the resolved model in the saved /reflect data: the configured Copilot endpoint's routing_models entry and its supported_endpoints. Reuse getAWFRoutingModel and getAWFRoutingModelAmbiguityError from awf_model_routing.cjs instead of adding a second parser.
  • In applyCopilotWireAPI, set COPILOT_PROVIDER_WIRE_API to a supported endpoint that the Copilot CLI can use: /responses maps to responses, and /chat/completions maps to completions. When a model supports both, keep the current preference (catalog wire_api, then the name rule) so models that work today keep the same behavior.
  • Keep explicit COPILOT_PROVIDER_WIRE_API overrides from engine.env, but check them against supported_endpoints as well (see step 2).
  • Use getCatalogModelEntry, the -utility fallback, and the gpt-5+ regex only when /reflect has no usable metadata for the model. That covers AWF disabled, an older AWF, or candidate_metadata_complete not being true. Log which source chose the wire API.

2. Fail before spawning the CLI

  • In the copilot_harness.cjs startup path (where applyCopilotModelAliasResolution and applyCopilotWireAPI run, before the CLI is spawned), exit with a non-zero status and don't start the CLI in either of these cases:
    • The resolved model doesn't support any wire API the CLI can use.
    • An explicit COPILOT_PROVIDER_WIRE_API override isn't in the model's supported_endpoints.
  • The error message should include:
    • the configured model, and the resolved model when the configured one was an alias
    • the configured wire API and whether it came from an override or was inferred
    • the model's supported endpoints
    • how to fix it: pin a compatible model, remove the COPILOT_PROVIDER_WIRE_API override, or upgrade gh-aw
  • Record this failure in the unified session (see step 3), so it appears in the step summary, the failure issue, and audit, and not only in the log.
  • Sub-agents: the harness doesn't read declared sub-agent models today. Read the model: of custom agents in the checked-out .github/agents/ directory. If an agent's model doesn't support the main session's wire API in its supported_endpoints, emit a warning, not a failure. The warning should name the agent, its model, and both endpoints, and link [aw] PR Code Quality Reviewer had a request rejected #67460 and BYOK: sub-agents always use the session's wire API, so a sub-agent on a model from the other family fails with a 400 copilot-cli#5103. Only warn when /reflect has endpoint data for both models.

3. Classify mid-run mismatches as model misconfiguration

  • Add a model/endpoint-mismatch category to harness_error_patterns.cjs, next to the model_policy_violation entry. It should match these AWF 400 messages:

    • Cannot translate Copilot request feature
    • Unsupported Responses custom tool
    • model_endpoint_incompatible
    • not accessible via the … endpoint

    When the message includes the model name or endpoint, capture them.

  • Return the new category from detectNonRetryableHarnessGuard in harness_retry_guard.cjs. In the copilot_harness.cjs retry loop, check it before the generic HTTP 400 branch so these errors are never retried, including the fresh-run retry after --continue. Add the same patterns to agent_error_patterns.cjs so detect_agent_errors.cjs classifies them too.

  • Unified session: follow pelikhan's rule that audit and logs read from the unified session. Steps:

    • Have the harness write a small JSON record under the agent directory, the same way recordAWFModelRoutingOutcome writes agent/awf-routing-outcome.json.
    • Register that record in the source table in unified_session.cjs as a new event type in usage/aw_session.jsonl.
    • Update types/unified_session.d.ts, docs/public/schemas/unified-session.schema.json, and docs/src/content/docs/specs/unified-agent-session-specification.md for the new event type.
    • Use the same record for the pre-spawn failure (step 2) and for the mid-run 400.
  • Step summary: render the event in unified_session_render.cjs with the cause and the fix.

  • Agent-failure issue, in handle_agent_failure.cjs:

    • Add a context builder next to buildModelNotSupportedErrorContext and buildHTTP400ResponseErrorContext. It should read the cause from the unified session, not by scanning raw logs again.
    • Add its placeholder to actions/setup/md/agent_failure_issue.md and agent_failure_comment.md.
    • Add the category to buildFailureMatchCategories so repeated failures are grouped into one issue.
  • Audit: have gh aw audit read the new event from aw_session.jsonl, next to readSessionModelRouting in pkg/cli/model_routing_session.go, and report the model, the endpoints, and the fix.

4. Tests

  • copilot_harness.test.cjs:
    • A model that only supports /responses, combined with a completions override, fails before spawn. The message names the model, both endpoints, and the fix.
    • auto resolving to a -utility variant picks responses from /reflect data.
    • A model missing from the models.json catalog but present in /reflect gets the correct endpoint.
    • Without /reflect metadata, the current name-based behavior doesn't change.
    • The sub-agent warning fires for a declared model from the other family, and doesn't fire for a model from the same family.
  • harness_retry_guard.test.cjs: each of the four 400 messages is classified, and other 400s aren't.
  • detect_agent_errors.test.cjs: each of the four messages is classified.
  • Retry loop: each message stops the run without a retry, including after --continue.
  • unified_session.test.cjs, unified_session_render.test.cjs, handle_agent_failure.test.cjs, and the pkg/cli audit tests: the event is recorded, rendered in the step summary, included in the failure issue, and reported by audit.

Acceptance criteria

  • When /reflect shows that a model and wire API don't match, the run fails before the CLI starts, with a message that explains how to fix it.
  • A mismatch that happens mid-run isn't retried. Its cause and fix appear in the step summary, the agent-failure issue, and gh aw audit.
  • A new model doesn't need a per-model name rule when AWF advertises its endpoints.

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions