Skip to content

Copilot sub-agents fail with 400 when their model needs a different API than the main model (Claude under GPT, GPT under Claude) #66280

Description

@SivaKesava1

Summary

The Copilot engine runs inside AWF with a BYOK provider pointed at the API proxy. In that mode,
Copilot CLI sends every request in the session through one wire API, responses or
completions, and that includes sub-agents. Copilot's API doesn't serve Claude models on
/responses, and it doesn't serve GPT-5 family models such as gpt-5.4-mini on
/chat/completions. So an inline or imported sub-agent whose model: is from the other family
fails with an upstream 400 on every call:

  • a Claude sub-agent under a GPT-5 main model;
  • a GPT-5 sub-agent under a Claude main model.

This happens with and without model routing, in the default CLI mode. The inline sub-agents
reference uses model: claude-haiku-4.5 as its example, and that fails under any GPT-5 main
model.

#66234 makes this more visible under model routing. Declared sub-agent models are now admitted
by the request policy, but the main model is chosen per run. A workflow with a fixed-model
sub-agent therefore works on some routed arms and fails on others, and the author can't tell
which in advance.

Evidence

The test repository runs gh-aw main 3fd44f8dc1 (compiled with --gh-aw-ref) and AWF 0.28.35
images on GitHub-hosted runners.

  • The workflow has two inline sub-agents: file-summarizer, with model: claude-haiku-4.5, and
    quick-checker, with a GPT model.
  • The task tells the agent to delegate one question to each and not to answer them itself.
Main model Routing Claude sub-agent (claude-haiku-4.5) GPT sub-agent (gpt-5.4-mini) Run result
gpt-5.6-luna routed, allowed-models: [gpt-5.6-luna] 400 on /responses, 2 attempts 200 on /responses (through the small alias) Success; the agent fell back to a general-purpose sub-agent
gpt-5.6-luna none (model:) 400 on /responses 200 (through small) Success; the agent reran the task on the main model
claude-sonnet-5 none (model:) 200 on /chat/completions 400 on /chat/completions, 3 attempts Failure: the agent reported the sub-agent as a missing tool
claude-sonnet-5 routed, allowed-models: [claude-sonnet-5] 200 on /chat/completions 400 on /chat/completions, 4 attempts Success, without the GPT sub-agent's answer

Upstream error messages:

  • 400 model claude-haiku-4.5 does not support Responses API.
  • 400 model "gpt-5.4-mini" is not accessible via the /chat/completions endpoint

These are recorded in each run's artifacts:

  • the Copilot session events, as subagent.failed with the sub-agent name and model;
  • the API proxy's upstream-errors.jsonl;
  • for routed runs, model-routing.jsonl.

They aren't policy rejections: the request policy admitted these models (no 403
model_policy_violation).

The harness logs show where the session's wire API came from:

  • GPT main, not routed: COPILOT_PROVIDER_WIRE_API already set to responses — skipping auto-configure.
  • Routed: inference routing: mode=awf-routed model=gpt-5.6-luna effort=high wire_api=responses,
    and ... model=claude-sonnet-5 effort=medium wire_api=completions for the Claude arm.
  • Claude main, not routed: no wire API is set, and Copilot CLI uses completions for the session.

Cause

  • With AWF, the agent always runs Copilot CLI in BYOK mode: COPILOT_PROVIDER_BASE_URL points
    at the API proxy, for credential isolation. A BYOK provider has a single wire API, set by
    COPILOT_PROVIDER_WIRE_API or Copilot CLI's default.
  • Three places set it from the main model only:
    • AWF's copilot-credential-env.ts sets responses whenever COPILOT_MODEL is GPT-5 family;
    • gh-aw's applyCopilotWireAPI in copilot_harness.cjs does the same when it's unset;
    • under routing, applyCopilotRoutingSelection sets it from the selected arm's endpoint.
  • Sub-agents spawned by the task tool reuse the session's provider, so their requests use the
    main model's wire API, whatever their own model needs.
  • gh-aw already decides the wire API from the model name: from the catalogue's wire_api in
    models.json when present, otherwise from the GPT-5 family rule in applyCopilotWireAPI. So
    the mismatch can be detected at compile time.

Proposed plan

  1. Give each model its own wire API.
    • Configure the agent's BYOK provider so that each model uses the wire API it needs.
    • One way: two providers pointing at the same API proxy, one for responses and one for
      completions, with each model (main, sub-agent, and routed candidate) mapped to the
      provider that matches its catalogue wire_api.
    • The SDK driver path already builds a multi-provider config from AWF's reflect data
      (resolveMultiProviderFromReflect). Check whether CLI mode can use the same mechanism, and
      whether Copilot CLI supports a per-model provider in CLI mode.
    • If Copilot CLI can't do this, raise it with the Copilot CLI team. Use steps 2 and 3 in the
      meantime.
  2. Warn at compile time. When a sub-agent's model needs a different wire API from the main
    model, emit a compiler warning naming the sub-agent and both models.
    • Under routing, compare against every routing candidate and warn when any candidate differs.
      That routed arm would break the sub-agent.
    • Use the catalogue's wire_api, falling back to the GPT-5 family rule the harness already
      uses.
  3. Document the constraint in the inline sub-agents reference until step 1 lands.
    • Its claude-haiku-4.5 example fails under any GPT-5 main model.
    • State that a sub-agent's model must use the same API family as the main model, or (under
      routing) as every allowed model.
  4. Tests.
    • Harness: a session with a GPT main model and a Claude sub-agent model resolves a provider
      with the right wire API for each.
    • Compiler: warnings for a Claude sub-agent under a GPT model, for a GPT sub-agent under a
      Claude model, and for a routed workflow whose candidates span both families.
    • The existing smoke test "Smoke Copilot Sub Agents" covers one direction in SDK mode only.
      Add a CLI-mode case.

Open questions

  • SDK mode. "Smoke Copilot Sub Agents" passes daily with copilot-sdk: true, a
    gpt-5.3-codex main model, and a claude-haiku-4.5 sub-agent, so SDK mode may already handle
    this. We couldn't test SDK mode in our repository: the SDK driver exited with No GitHub OAuth token or Copilot HMAC key provided, because the workflow authenticates with github.token and
    copilot-requests: write and has no Copilot token. Is SDK mode the intended path for
    mixed-family sub-agents, and does it work under routing?
  • Ownership of the default. AWF sets the GPT-5 default in the agent environment, and gh-aw
    sets it again in the harness. Whichever fix is chosen should leave one owner for the wire API
    decision.

Notes

Activity

  1. locked and limited conversation to collaborators on Oct 6, 2026
  2. unlocked this conversation on Oct 6, 2026
  3. SivaKesava1 commented on Oct 6, 2026

    @SivaKesava1
    CollaboratorAuthor
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions