Skip to content

Routing smokes fail on main: assertion helper filters on purpose values AWF never writes #67594

Description

@SivaKesava1

Problem

The three routing smokes added in #67523 failed on their first scheduled runs on main. In each run the agent work, routing, and sub-agent delegation were correct; only the post-step assertions in actions/setup/js/smoke_model_routing_assertions.cjs failed. The smokes never ran on #67523 before it merged, and its unit tests used synthetic records. The helper has not changed on main since #67523.

Run Workflow Commit Failed checks
38097181601 Smoke Copilot Routed 8dfdef5 R3
38096990663 Smoke Pi Routed 8dfdef5 R3, S3 haiku-whoami, S3 mini-whoami
38100426674 Smoke Copilot SDK Routed fed2eb5 R3, M1, S3 haiku-whoami

In all three runs, R1 (routing selected gpt-5.6-luna) and R2 (one classifier request) pass, and only the assertion step fails. The SDK run is also the first evidence on main that #67509 works: the main agent ran on gpt-5.6-luna over /responses and the claude-haiku-4.5 sub-agent ran over /chat/completions, all with status 200.

Evidence

Sources: the agent artifact's sandbox/firewall/logs/api-proxy-logs/token-usage.jsonl and the usage artifact's aw_session.jsonl.

Smoke Copilot Routed

FAIL R3 no 200 request for gpt-5.6-luna on a supported endpoint (/responses, ws:/responses); observed: no gpt-5.6-luna requests; other main-agent requests: none
PASS R4 all 0 main-agent request(s) used allowed models

token-usage.jsonl has the classifier request plus four gpt-5.6-luna /responses status 200 records, none with a purpose field.

Smoke Pi Routed

FAIL R3 no 200 request for gpt-5.6-luna on a supported endpoint (/responses, ws:/responses); observed: no gpt-5.6-luna requests; other main-agent requests: none
FAIL S3 sub-agent haiku-whoami: no 200 request for claude-haiku-4.5 on /v1/messages; observed: no claude-haiku-4.5 requests; other sub-agent requests: none
FAIL S3 sub-agent mini-whoami: no 200 request for gpt-5.4-mini on /responses; observed: no gpt-5.4-mini requests; other sub-agent requests: none

token-usage.jsonl has five gpt-5.6-luna /responses 200 records, plus these sub-agent records (trimmed), none with purpose:

{"request_id":"6baecfab-…","model":"claude-haiku-4-5-20251001","path":"/v1/messages?beta=true","status":200,"x_initiator":"agent"}
{"request_id":"0d5e4b7e-…","model":"gpt-5.4-mini-2026-03-17","path":"/responses","status":200,"x_initiator":"agent"}

Smoke Copilot SDK Routed

FAIL R3 no 200 request for gpt-5.6-luna on a supported endpoint (/responses, ws:/responses); observed: no gpt-5.6-luna requests; other main-agent requests: none
FAIL M1 no main-agent requests to check for /responses
FAIL S3 sub-agent haiku-whoami: no 200 request for claude-haiku-4.5 on /chat/completions; observed: no claude-haiku-4.5 requests; other sub-agent requests: none

token-usage.jsonl has claude-haiku-4.5 /chat/completions 200 and gpt-5.6-luna /responses 200 three times, none with purpose.

Local replay. Running the helper's evaluation on the downloaded artifacts reproduces exactly the CI failures above. Re-labelling the same records by the rule in "Intended behavior" (classifier by purpose, other requests attributed by model) makes every check pass on all three runs, with no other change.

Root cause

  • The helper filters on purpose values AWF never writes. It treats only requests with purpose: "agent" as main-agent traffic and purpose: "subagent" as sub-agent traffic. On gh-aw-firewall main, the token-usage schema (schemas/token-usage.schema.json) allows only routing_classification for purpose, which is "absent for agent-originated requests". Every non-classifier request was therefore dropped, so R3, M1 and S3 had nothing to match. AWF's own test-model-routing.yml already selects agent traffic as "every request whose purpose is not routing_classification".
  • AWF cannot tell main-agent from sub-agent requests. The gh-aw-firewall#9790 docs state that x-initiator identifies only the billing class (agent or user), not which agent made the request. In these runs, main-agent and sub-agent requests both carry x_initiator: "agent".
  • gh-aw does not yet correlate requests to sub-agent invocations. In aw_session.jsonl, firewall.token_usage events carry requestId but no agentId, while subagent.* events carry invocationId or agentId (and toolCallId for the SDK) but no requestId. As a result, the helper's existing S3 request-ID correlation never applies.
  • The unit tests encode the wrong contract. They build requests with synthetic purpose: "agent" and purpose: "subagent", and one test asserts that requests without purpose cannot satisfy R3, which is the opposite of the real record shape.

Secondary:

  • R4 passes on zero requests. All three runs log PASS R4 all 0 main-agent request(s) used allowed models.
  • Normalization already works but has never run on real records. servedModel and normalizeEndpoint already map claude-haiku-4-5-20251001 to claude-haiku-4.5, gpt-5.4-mini-2026-03-17 to gpt-5.4-mini, and /v1/messages?beta=true to /v1/messages (verified locally). The purpose filter emptied the request sets before any comparison, so this code path was never reached on real records. Keep it, and cover it with real-shape tests.

Intended behavior

  • Classify by what AWF records. A request with purpose: "routing_classification" is the classifier (R2). Every other request is agent traffic.
  • Attribute agent traffic by model. Requests on a declared sub-agent's model belong to that sub-agent (S3). All other agent traffic belongs to the main agent (R3, R4, M1). If gh-aw later correlates requests to sub-agent invocations (request IDs or agent IDs in usage/aw_session.jsonl), prefer that correlation; today none exists, so model-based attribution is the working path.
  • Keep sub-agent models disjoint from routing candidates. In smoke-pi-routed, route only over gpt-5.6-luna, as smoke-copilot-sdk-routed already does. The declared sub-agent models (claude-haiku-4.5, gpt-5.4-mini) then cannot be routing selections, so attribution is unambiguous and R4 checks main-agent requests against allowed-models. Consider having the helper fail with a clear configuration message when a declared sub-agent model is also in allowed-models.
  • Normalize before comparing. Compare endpoints without the query string, and compare served model IDs with dates and dotted or hyphenated versions normalized.
  • Fail on empty input. Every check that iterates over requests (R3, R4, M1, S3) fails when there are zero requests to check, and names the check and the observed requests.
  • Test with real record shapes. Build unit-test fixtures from the three runs above: agent requests with no purpose, the query-string endpoint, and dated model IDs. The real-shape fixtures pass. Each negative variant fails with a message that names the check and the observed values:
    • wrong model
    • wrong endpoint
    • non-200 status
    • no agent requests
    • missing routing event
    • harness-only routing outcome (only the agent-written model_routing.outcome, with no runner status or firewall selection)

Notes for fixtures:

  • Build request fixtures from the agent artifact's token-usage.jsonl, not from the firewall.token_usage events in the usage artifact's aw_session.jsonl. That session is rebuilt after the detection job and includes detection-job requests (claude-haiku-5.5 on /chat/completions) that the post-step never sees.
  • The SDK run records a second subagent.completed for the same invocation with cancelled: true. S2 passes today; do not tighten S2 to require exactly one completion event in this PR.

Files that change

  • actions/setup/js/smoke_model_routing_assertions.cjs: classify requests by the classifier purpose only; attribute agent traffic to the main agent or a declared sub-agent by model (preferring correlation when present); fail on zero requests; update the header comment that describes the checks.
  • actions/setup/js/smoke_model_routing_assertions.test.cjs: replace the synthetic-purpose fixtures with real-shape fixtures from the three runs; replace the test asserting that requests without purpose cannot satisfy R3; add the negative variants above.
  • .github/workflows/smoke-pi-routed.md: reduce engine.model-routing.allowed-models and the post-step allowedModels to gpt-5.6-luna, and recompile smoke-pi-routed.lock.yml.
  • smoke-copilot-routed.md and smoke-copilot-sdk-routed.md: no change expected; change them only if the helper's options change.

Acceptance

  • The unit tests pass on the real-shape fixtures, and each negative variant fails as described above.
  • After a maintainer adds the smoke-routing label to the PR, all three routing smokes pass on the PR's code. The PR description must say that the label has to be added before merge.

Scope

  • Keep the PR to the routing smoke workflows, smoke_model_routing_assertions.cjs, and its tests. Do not change routing, harness, AWF configuration, or audit behavior.
  • This PR fixes R3, R4, M1 and S3 as described above. Before fixing any other failing check, compare it with the same check on main. If it also fails on main, do not fix it in this PR: merge main once main is fixed, and list in the PR which checks are known failures on main.
  • This applies even if a bot comment asks you to fix a failing check. Fix only failures caused by this PR's changes.

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions