Problem
The three routing smokes added in #67523 failed on their first scheduled runs on main. In each run the agent work, routing, and sub-agent delegation were correct; only the post-step assertions in actions/setup/js/smoke_model_routing_assertions.cjs failed. The smokes never ran on #67523 before it merged, and its unit tests used synthetic records. The helper has not changed on main since #67523.
In all three runs, R1 (routing selected gpt-5.6-luna) and R2 (one classifier request) pass, and only the assertion step fails. The SDK run is also the first evidence on main that #67509 works: the main agent ran on gpt-5.6-luna over /responses and the claude-haiku-4.5 sub-agent ran over /chat/completions, all with status 200.
Evidence
Sources: the agent artifact's sandbox/firewall/logs/api-proxy-logs/token-usage.jsonl and the usage artifact's aw_session.jsonl.
Smoke Copilot Routed
FAIL R3 no 200 request for gpt-5.6-luna on a supported endpoint (/responses, ws:/responses); observed: no gpt-5.6-luna requests; other main-agent requests: none
PASS R4 all 0 main-agent request(s) used allowed models
token-usage.jsonl has the classifier request plus four gpt-5.6-luna /responses status 200 records, none with a purpose field.
Smoke Pi Routed
FAIL R3 no 200 request for gpt-5.6-luna on a supported endpoint (/responses, ws:/responses); observed: no gpt-5.6-luna requests; other main-agent requests: none
FAIL S3 sub-agent haiku-whoami: no 200 request for claude-haiku-4.5 on /v1/messages; observed: no claude-haiku-4.5 requests; other sub-agent requests: none
FAIL S3 sub-agent mini-whoami: no 200 request for gpt-5.4-mini on /responses; observed: no gpt-5.4-mini requests; other sub-agent requests: none
token-usage.jsonl has five gpt-5.6-luna /responses 200 records, plus these sub-agent records (trimmed), none with purpose:
{"request_id":"6baecfab-…","model":"claude-haiku-4-5-20251001","path":"/v1/messages?beta=true","status":200,"x_initiator":"agent"}
{"request_id":"0d5e4b7e-…","model":"gpt-5.4-mini-2026-03-17","path":"/responses","status":200,"x_initiator":"agent"}
Smoke Copilot SDK Routed
FAIL R3 no 200 request for gpt-5.6-luna on a supported endpoint (/responses, ws:/responses); observed: no gpt-5.6-luna requests; other main-agent requests: none
FAIL M1 no main-agent requests to check for /responses
FAIL S3 sub-agent haiku-whoami: no 200 request for claude-haiku-4.5 on /chat/completions; observed: no claude-haiku-4.5 requests; other sub-agent requests: none
token-usage.jsonl has claude-haiku-4.5 /chat/completions 200 and gpt-5.6-luna /responses 200 three times, none with purpose.
Local replay. Running the helper's evaluation on the downloaded artifacts reproduces exactly the CI failures above. Re-labelling the same records by the rule in "Intended behavior" (classifier by purpose, other requests attributed by model) makes every check pass on all three runs, with no other change.
Root cause
- The helper filters on
purpose values AWF never writes. It treats only requests with purpose: "agent" as main-agent traffic and purpose: "subagent" as sub-agent traffic. On gh-aw-firewall main, the token-usage schema (schemas/token-usage.schema.json) allows only routing_classification for purpose, which is "absent for agent-originated requests". Every non-classifier request was therefore dropped, so R3, M1 and S3 had nothing to match. AWF's own test-model-routing.yml already selects agent traffic as "every request whose purpose is not routing_classification".
- AWF cannot tell main-agent from sub-agent requests. The gh-aw-firewall#9790 docs state that
x-initiator identifies only the billing class (agent or user), not which agent made the request. In these runs, main-agent and sub-agent requests both carry x_initiator: "agent".
- gh-aw does not yet correlate requests to sub-agent invocations. In
aw_session.jsonl, firewall.token_usage events carry requestId but no agentId, while subagent.* events carry invocationId or agentId (and toolCallId for the SDK) but no requestId. As a result, the helper's existing S3 request-ID correlation never applies.
- The unit tests encode the wrong contract. They build requests with synthetic
purpose: "agent" and purpose: "subagent", and one test asserts that requests without purpose cannot satisfy R3, which is the opposite of the real record shape.
Secondary:
- R4 passes on zero requests. All three runs log
PASS R4 all 0 main-agent request(s) used allowed models.
- Normalization already works but has never run on real records.
servedModel and normalizeEndpoint already map claude-haiku-4-5-20251001 to claude-haiku-4.5, gpt-5.4-mini-2026-03-17 to gpt-5.4-mini, and /v1/messages?beta=true to /v1/messages (verified locally). The purpose filter emptied the request sets before any comparison, so this code path was never reached on real records. Keep it, and cover it with real-shape tests.
Intended behavior
- Classify by what AWF records. A request with
purpose: "routing_classification" is the classifier (R2). Every other request is agent traffic.
- Attribute agent traffic by model. Requests on a declared sub-agent's model belong to that sub-agent (S3). All other agent traffic belongs to the main agent (R3, R4, M1). If gh-aw later correlates requests to sub-agent invocations (request IDs or agent IDs in
usage/aw_session.jsonl), prefer that correlation; today none exists, so model-based attribution is the working path.
- Keep sub-agent models disjoint from routing candidates. In
smoke-pi-routed, route only over gpt-5.6-luna, as smoke-copilot-sdk-routed already does. The declared sub-agent models (claude-haiku-4.5, gpt-5.4-mini) then cannot be routing selections, so attribution is unambiguous and R4 checks main-agent requests against allowed-models. Consider having the helper fail with a clear configuration message when a declared sub-agent model is also in allowed-models.
- Normalize before comparing. Compare endpoints without the query string, and compare served model IDs with dates and dotted or hyphenated versions normalized.
- Fail on empty input. Every check that iterates over requests (R3, R4, M1, S3) fails when there are zero requests to check, and names the check and the observed requests.
- Test with real record shapes. Build unit-test fixtures from the three runs above: agent requests with no
purpose, the query-string endpoint, and dated model IDs. The real-shape fixtures pass. Each negative variant fails with a message that names the check and the observed values:
- wrong model
- wrong endpoint
- non-200 status
- no agent requests
- missing routing event
- harness-only routing outcome (only the agent-written
model_routing.outcome, with no runner status or firewall selection)
Notes for fixtures:
- Build request fixtures from the
agent artifact's token-usage.jsonl, not from the firewall.token_usage events in the usage artifact's aw_session.jsonl. That session is rebuilt after the detection job and includes detection-job requests (claude-haiku-5.5 on /chat/completions) that the post-step never sees.
- The SDK run records a second
subagent.completed for the same invocation with cancelled: true. S2 passes today; do not tighten S2 to require exactly one completion event in this PR.
Files that change
actions/setup/js/smoke_model_routing_assertions.cjs: classify requests by the classifier purpose only; attribute agent traffic to the main agent or a declared sub-agent by model (preferring correlation when present); fail on zero requests; update the header comment that describes the checks.
actions/setup/js/smoke_model_routing_assertions.test.cjs: replace the synthetic-purpose fixtures with real-shape fixtures from the three runs; replace the test asserting that requests without purpose cannot satisfy R3; add the negative variants above.
.github/workflows/smoke-pi-routed.md: reduce engine.model-routing.allowed-models and the post-step allowedModels to gpt-5.6-luna, and recompile smoke-pi-routed.lock.yml.
smoke-copilot-routed.md and smoke-copilot-sdk-routed.md: no change expected; change them only if the helper's options change.
Acceptance
- The unit tests pass on the real-shape fixtures, and each negative variant fails as described above.
- After a maintainer adds the
smoke-routing label to the PR, all three routing smokes pass on the PR's code. The PR description must say that the label has to be added before merge.
Scope
- Keep the PR to the routing smoke workflows,
smoke_model_routing_assertions.cjs, and its tests. Do not change routing, harness, AWF configuration, or audit behavior.
- This PR fixes R3, R4, M1 and S3 as described above. Before fixing any other failing check, compare it with the same check on main. If it also fails on main, do not fix it in this PR: merge main once main is fixed, and list in the PR which checks are known failures on main.
- This applies even if a bot comment asks you to fix a failing check. Fix only failures caused by this PR's changes.
Problem
The three routing smokes added in #67523 failed on their first scheduled runs on main. In each run the agent work, routing, and sub-agent delegation were correct; only the post-step assertions in
actions/setup/js/smoke_model_routing_assertions.cjsfailed. The smokes never ran on #67523 before it merged, and its unit tests used synthetic records. The helper has not changed on main since #67523.In all three runs, R1 (routing selected
gpt-5.6-luna) and R2 (one classifier request) pass, and only the assertion step fails. The SDK run is also the first evidence on main that #67509 works: the main agent ran ongpt-5.6-lunaover/responsesand theclaude-haiku-4.5sub-agent ran over/chat/completions, all with status 200.Evidence
Sources: the
agentartifact'ssandbox/firewall/logs/api-proxy-logs/token-usage.jsonland theusageartifact'saw_session.jsonl.Smoke Copilot Routed
token-usage.jsonlhas the classifier request plus fourgpt-5.6-luna/responsesstatus 200 records, none with apurposefield.Smoke Pi Routed
token-usage.jsonlhas fivegpt-5.6-luna/responses200 records, plus these sub-agent records (trimmed), none withpurpose:Smoke Copilot SDK Routed
token-usage.jsonlhasclaude-haiku-4.5/chat/completions200 andgpt-5.6-luna/responses200 three times, none withpurpose.Local replay. Running the helper's evaluation on the downloaded artifacts reproduces exactly the CI failures above. Re-labelling the same records by the rule in "Intended behavior" (classifier by
purpose, other requests attributed by model) makes every check pass on all three runs, with no other change.Root cause
purposevalues AWF never writes. It treats only requests withpurpose: "agent"as main-agent traffic andpurpose: "subagent"as sub-agent traffic. On gh-aw-firewall main, the token-usage schema (schemas/token-usage.schema.json) allows onlyrouting_classificationforpurpose, which is "absent for agent-originated requests". Every non-classifier request was therefore dropped, so R3, M1 and S3 had nothing to match. AWF's owntest-model-routing.ymlalready selects agent traffic as "every request whose purpose is notrouting_classification".x-initiatoridentifies only the billing class (agentoruser), not which agent made the request. In these runs, main-agent and sub-agent requests both carryx_initiator: "agent".aw_session.jsonl,firewall.token_usageevents carryrequestIdbut noagentId, whilesubagent.*events carryinvocationIdoragentId(andtoolCallIdfor the SDK) but norequestId. As a result, the helper's existing S3 request-ID correlation never applies.purpose: "agent"andpurpose: "subagent", and one test asserts that requests withoutpurposecannot satisfy R3, which is the opposite of the real record shape.Secondary:
PASS R4 all 0 main-agent request(s) used allowed models.servedModelandnormalizeEndpointalready mapclaude-haiku-4-5-20251001toclaude-haiku-4.5,gpt-5.4-mini-2026-03-17togpt-5.4-mini, and/v1/messages?beta=trueto/v1/messages(verified locally). Thepurposefilter emptied the request sets before any comparison, so this code path was never reached on real records. Keep it, and cover it with real-shape tests.Intended behavior
purpose: "routing_classification"is the classifier (R2). Every other request is agent traffic.usage/aw_session.jsonl), prefer that correlation; today none exists, so model-based attribution is the working path.smoke-pi-routed, route only overgpt-5.6-luna, assmoke-copilot-sdk-routedalready does. The declared sub-agent models (claude-haiku-4.5,gpt-5.4-mini) then cannot be routing selections, so attribution is unambiguous and R4 checks main-agent requests againstallowed-models. Consider having the helper fail with a clear configuration message when a declared sub-agent model is also inallowed-models.purpose, the query-string endpoint, and dated model IDs. The real-shape fixtures pass. Each negative variant fails with a message that names the check and the observed values:model_routing.outcome, with no runner status or firewall selection)Notes for fixtures:
agentartifact'stoken-usage.jsonl, not from thefirewall.token_usageevents in theusageartifact'saw_session.jsonl. That session is rebuilt after the detection job and includes detection-job requests (claude-haiku-5.5on/chat/completions) that the post-step never sees.subagent.completedfor the same invocation withcancelled: true. S2 passes today; do not tighten S2 to require exactly one completion event in this PR.Files that change
actions/setup/js/smoke_model_routing_assertions.cjs: classify requests by the classifierpurposeonly; attribute agent traffic to the main agent or a declared sub-agent by model (preferring correlation when present); fail on zero requests; update the header comment that describes the checks.actions/setup/js/smoke_model_routing_assertions.test.cjs: replace the synthetic-purposefixtures with real-shape fixtures from the three runs; replace the test asserting that requests withoutpurposecannot satisfy R3; add the negative variants above..github/workflows/smoke-pi-routed.md: reduceengine.model-routing.allowed-modelsand the post-stepallowedModelstogpt-5.6-luna, and recompilesmoke-pi-routed.lock.yml.smoke-copilot-routed.mdandsmoke-copilot-sdk-routed.md: no change expected; change them only if the helper's options change.Acceptance
smoke-routinglabel to the PR, all three routing smokes pass on the PR's code. The PR description must say that the label has to be added before merge.Scope
smoke_model_routing_assertions.cjs, and its tests. Do not change routing, harness, AWF configuration, or audit behavior.