Skip to content

Model routing guide: CLI vs Copilot SDK mode for mixed model families, what allowed-models means, and tested examples #67754

Description

@SivaKesava1

Problem

The model routing guide (docs/src/content/docs/reference/model-routing.md) is out of date with main after this week's merges. It also doesn't tell users what decides whether routing with sub-agents works for them: whether the Copilot engine runs in CLI mode or SDK mode.

Each point below was checked against origin/main at 7f713c7 (2026-10-11).

1. The guide doesn't explain CLI mode vs SDK mode

2. The guide calls allowed-models the request policy

  • The guide says that allowed-models "defines the request policy for the routed task". The Configuration table says it lists the "Models routing may choose from and the routed task may call".
  • On main, engine.model-routing.allowed-models is the router's candidate list for the main agent. The compiler applies the workflow model policy to the list and writes the result to AWF routing.candidateModels (pkg/workflow/awf_config_build.go).
  • Declared sub-agent models are allowed automatically (Separate routing candidates from sub-agent model policy #66234) and don't need to be listed. AWF's request allowlist is the candidates plus the declared sub-agent models. For example, smoke-copilot-sdk-routed.lock.yml compiles to candidateModels: [gpt-5.6-luna] and allowedModels: [gpt-5.6-luna, claude-haiku-4.5]. If you list a sub-agent's model in allowed-models, it also becomes a candidate for the main agent.
  • models.allowed and models.blocked are the workflow's request policy, which is a different thing. Both the candidates and the sub-agent models must satisfy it. The glossary documents it under "Model Policy (models.allowed, models.blocked)".

3. The threat-detection limitation is stale

  • "Known limitations" still has an entry titled "Threat detection is not routed". Keep model routing out of threat detection jobs #66670 fixed detection under routing. Detection jobs used to inherit the routing config and fail before producing a verdict. They no longer inherit it and run on their configured detection model. That is intended behavior, not a limitation.
  • On main, this limitation entry is the only threat-detection text in the guide. The only callout on the page is the experimental warning at the top, which links to "Known limitations" in general and can stay.

Also found while verifying

  • The guide says routed workflows default to AWF v0.28.37. On main, DefaultFirewallVersion is v0.28.50. Router 0.1.3 is still correct.
  • The smoke workflows don't all route over a single model:
    • smoke-copilot-routed.md routes over three models (gpt-5.4-mini, gpt-5.6-luna, claude-haiku-4.5) and has no sub-agents.
    • smoke-copilot-sdk-routed.md and smoke-pi-routed.md route over a single model (gpt-5.6-luna), each with cross-family sub-agents. The smoke assertions attribute traffic to sub-agents by model, so sub-agent models must not appear in allowed-models.
    • All three use mode: economy and run every 2 days. They check each request's model and endpoint using the unified session and AWF's api-proxy token-usage.jsonl.
  • smoke-pi-routed.md runs cross-family sub-agents under pi: a GPT main model, claude-haiku-4.5 on /v1/messages and gpt-5.4-mini on /responses. The CLI-mode restriction is specific to the Copilot engine's CLI mode.
  • AWF v0.28.51 was published on 2026-10-11, but main's default AWF version is still v0.28.50. DefaultMaxAICredits is 1000.

Intended content (one docs PR)

  1. A short "CLI mode vs SDK mode" section.
  2. An explanation of allowed-models.
    • Replace the "request policy" wording in "How routing works" and in the Configuration table.
    • allowed-models is the router's candidate list for the main agent.
    • Declared sub-agent models are allowed automatically (Separate routing candidates from sub-agent model policy #66234) and don't need to be listed. Listing one makes it a candidate for the main agent.
    • models.allowed and models.blocked are the request policy, which is a different thing. Link the glossary's "Model Policy" entry.
  3. Two examples, both compiling on main.
    • CLI mode: engine id copilot with model routing (goal: cost, mode: auto) and a multi-model allowed-models list. Use either no sub-agents or sub-agents in the same family as every candidate. For example:
      engine:
        id: copilot
        model-routing:
          goal: cost
          mode: auto
          allowed-models: [gpt-5.4-mini, gpt-5.6-luna, gpt-5.6-sol]
      An inline sub-agent with model: gpt-5.4-mini can be added.
    • SDK mode: engine id copilot with copilot-sdk: true, model routing over a realistic mixed pool, permissions that include copilot-requests: write, and an inline ## agent: sub-agent with model: claude-haiku-4.5. For example:
      permissions:
        contents: read
        copilot-requests: write
      engine:
        id: copilot
        copilot-sdk: true
        model-routing:
          goal: cost
          mode: auto
          allowed-models: [gpt-5.4-mini, gpt-5.6-luna, gpt-5.6-sol, claude-haiku-4.5, claude-sonnet-5, claude-opus-5]
      All six models are in the routing tables of the pinned router 0.1.3. claude-haiku-4.5 is listed without an effort suffix.
  4. Links to the tested reference workflows.
    • Link .github/workflows/smoke-copilot-routed.md, smoke-pi-routed.md and smoke-copilot-sdk-routed.md. They run every 2 days and check each request's model and endpoint from AWF's logs.
    • Say that the SDK and pi smokes route over a single model only to keep their checks deterministic, not as a recommendation. Describe smoke-copilot-routed.md accurately: it routes over three models and has no sub-agents.
  5. Remove the stale "Threat detection is not routed" limitation entry. If you want to keep the information, add one factual sentence elsewhere: detection jobs aren't routed and use their configured detection model (Keep model routing out of threat detection jobs #66670).
  6. A short note on steering notices.
    • With AWF v0.28.51+, the run's AI-credit budget and timeout produce advisory steering notices, which gh aw audit lists as steering_notices (Emit AWF agentTimeout for literal timeouts and show delivered steering notices in audit/logs #67597).
    • Explain why this matters: routed runs on expensive models can approach the default 1,000-credit budget.
    • Link the token steering section of docs/src/content/docs/reference/sandbox.md.
    • Note that main's default AWF version is still v0.28.50 until it's bumped.
  7. Fix the stale default AWF version (v0.28.37). Either update it or reword the sentence so it won't go stale again.

Files that change

  • docs/src/content/docs/reference/model-routing.md is the only file expected to change.
  • sandbox.md, the glossary, engines.md and the smoke workflows are only linked to and should not change.

Verification done for this issue

  • Built gh-aw from origin/main at 7f713c7.
  • Compiled three scratch workflows, each with contents: read and copilot-requests: write:
    • CLI mode, mixed pool, no sub-agents
    • CLI mode, GPT pool with a GPT inline sub-agent
    • SDK mode, the six-model pool above with a claude-haiku-4.5 inline sub-agent
  • Each compile reported "1 succeeded, 0 warnings".
  • The compiled AWF candidateModels and allowedModels matched the description in point 2.

Acceptance

Scope

  • Keep the PR to these docs changes. Do not change files or behavior unrelated to them.
  • Before fixing a failing check, compare it with the same check on main. If it also fails on main, don't fix it in this PR. Merge main once main is fixed, and list in the PR which checks are known failures on main.
  • This applies even if a bot comment asks you to fix a failing check. Fix only failures caused by this PR's changes.

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions