You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
agentic-workflows skill: add a checklist for model and engine misconfiguration #67486
The /agentic-workflows skill doesn't diagnose model and engine misconfiguration well. In a customer case (Slack, 9–10 Oct), a workflow with model: auto failed with:
Running /agentic-workflows on that failure suggested switching to model: gpt-4.1. That works around the problem without fixing it. pelikhan: "Did you try to fix this with our skill? If it can't handle we should improve that."
The same thread also mixed up the version fields:
Setting tried
What it actually is
What happened
engine.version: 1.94.0
The Copilot CLI version, which is 1.0.x
Failed at "Install GitHub Copilot CLI" with bare curl 404s
sandbox.agent.version: 1.94.0
The AWF release version
Failed at "Install AWF binary"
engine.copilot.version
Not a field
—
The real causes were:
An old gh-aw. v0.89.21 is still the only non-prerelease, so gh extension install and gh extension upgrade install it. It predates the wire-API inference from Infer Copilot wire API for utility model variants #64177. The newest prerelease is v0.91.7.
A wire-API mismatch between the resolved model and the endpoint the Copilot CLI used.
This issue proposes adding a "Model and engine misconfiguration" checklist that the skill uses when it sees these symptoms.
Plan
Where the guidance goes
.github/skills/agentic-workflows/SKILL.md already sends debugging to .github/aw/debug-agentic-workflow.md. Add a routing line so these symptoms go straight to the new checklist:
AWF 400s that mention models or endpoints
model: auto failures
install-step 404s after a version pin
questions about which version field to set
In .github/aw/debug-agentic-workflow.md, add a "Model and engine misconfiguration" section. Link to it from "Collect Existing Evidence" and "Identify the First Failing Boundary".
Add matching, shorter guidance to the root debug.md, which external agents load from the lock-file header link. It can link to the full section.
What the checklist covers
1. Check the gh-aw version first.
Find the version the workflow was compiled with:
compiler_version in the gh-aw-metadata lock-file header
cli_version in aw_info.json
Compare it with the newest gh-aw release, including prereleases. Explain that gh extension install and gh extension upgrade only install the latest non-prerelease, so a user can be several releases behind without knowing it. Explain how to install a specific prerelease tag.
If the version predates the relevant fix, recommend upgrading and recompiling before trying anything else.
2. Explain how the Copilot harness picks the model and wire API.
How an alias like auto resolves to a concrete model.
How COPILOT_PROVIDER_WIRE_API is chosen, in this order:
an explicit engine.env override
the catalog wire_api
the -utility → base-model fallback
the gpt-5+ name rule
the CLI default, /chat/completions
That the CLI uses one wire API for the whole session, including sub-agents.
How to read these AWF 400s as a model/endpoint mismatch, not a transient failure or a prompt problem:
Cannot translate Copilot request feature
Unsupported Responses custom tool
model_policy_violation
not accessible via the … endpoint
3. Recommend fixes in this order.
Upgrade gh-aw and recompile.
Pin a model that supports the endpoint the run needs.
Remove a conflicting COPILOT_PROVIDER_WIRE_API override.
Recommend switching to an older model only when none of these work. Say why that's the last resort: it trades capability for a workaround and leaves the real misconfiguration in place.
4. Explain the version fields.
engine.version is the agent CLI version. For Copilot, that's the Copilot CLI, which is 1.0.x.
sandbox.agent.version is the AWF release version, in vX.Y.Z form, and must match a GitHub release of gh-aw-firewall.
There's no engine.copilot.version.
Map each failing install step to the field that controls it: "Install GitHub Copilot CLI" 404s point to engine.version, and "Install AWF binary" failures point to sandbox.agent.version. In most cases the fix is to remove the pin and use the compiled default.
5. List the artifacts to read.
agent-stdio.log: the [copilot-harness] lines for model alias resolution and the chosen COPILOT_PROVIDER_WIRE_API.
sandbox/firewall/logs/api-proxy-logs/token-usage.jsonl: the model and path for each request, which shows which endpoint each call actually used.
aw_info.json: model, requested_model, cli_version, version, and awf_version.
gh aw audit <run>: the combined view.
Regression coverage
Add a skill-level eval or worked example for the customer case, so a future change to the skill or debug guide can't bring back the "switch to gpt-4.1" advice. This skill has no eval harness today. The nearest precedents are per-skill tests/ folders such as .github/skills/operational-value-designer/tests/.
Fixture: a redacted lock-file header with an old compiler_version, model: auto, the [copilot-harness] alias and wire-API lines, and the AWF 400 text.
Expected diagnosis: upgrade gh-aw and recompile, then pin a model that supports the needed endpoint. It shouldn't recommend an older model first.
Second case: the misplaced engine.version and sandbox.agent.version pins. The expected diagnosis identifies the correct field for each, or recommends removing the pin.
Acceptance criteria
Debugging the customer's run with /agentic-workflows points to the outdated gh-aw and the wire-API mismatch, and recommends upgrading and pinning a compatible model before downgrading the model.
The skill explains engine.version and sandbox.agent.version correctly and doesn't suggest engine.copilot.version.
.github/aw/debug-agentic-workflow.md and the root debug.md carry the same checklist, and SKILL.md routes these symptoms to it.
The eval or example fails if the guidance regresses.
Summary
The
/agentic-workflowsskill doesn't diagnose model and engine misconfiguration well. In a customer case (Slack, 9–10 Oct), a workflow withmodel: autofailed with:Running
/agentic-workflowson that failure suggested switching tomodel: gpt-4.1. That works around the problem without fixing it. pelikhan: "Did you try to fix this with our skill? If it can't handle we should improve that."The same thread also mixed up the version fields:
engine.version: 1.94.0sandbox.agent.version: 1.94.0engine.copilot.versionThe real causes were:
gh extension installandgh extension upgradeinstall it. It predates the wire-API inference from Infer Copilot wire API for utility model variants #64177. The newest prerelease is v0.91.7.This issue proposes adding a "Model and engine misconfiguration" checklist that the skill uses when it sees these symptoms.
Plan
Where the guidance goes
.github/skills/agentic-workflows/SKILL.mdalready sends debugging to.github/aw/debug-agentic-workflow.md. Add a routing line so these symptoms go straight to the new checklist:model: autofailures.github/aw/debug-agentic-workflow.md, add a "Model and engine misconfiguration" section. Link to it from "Collect Existing Evidence" and "Identify the First Failing Boundary".debug.md, which external agents load from the lock-file header link. It can link to the full section.What the checklist covers
1. Check the gh-aw version first.
compiler_versionin thegh-aw-metadatalock-file headercli_versioninaw_info.jsongh extension installandgh extension upgradeonly install the latest non-prerelease, so a user can be several releases behind without knowing it. Explain how to install a specific prerelease tag.2. Explain how the Copilot harness picks the model and wire API.
autoresolves to a concrete model.COPILOT_PROVIDER_WIRE_APIis chosen, in this order:engine.envoverridewire_api-utility→ base-model fallbackgpt-5+name rule/chat/completionsCannot translate Copilot request featureUnsupported Responses custom toolmodel_policy_violationnot accessible via the … endpoint3. Recommend fixes in this order.
COPILOT_PROVIDER_WIRE_APIoverride.Recommend switching to an older model only when none of these work. Say why that's the last resort: it trades capability for a workaround and leaves the real misconfiguration in place.
4. Explain the version fields.
engine.versionis the agent CLI version. For Copilot, that's the Copilot CLI, which is 1.0.x.sandbox.agent.versionis the AWF release version, invX.Y.Zform, and must match a GitHub release of gh-aw-firewall.engine.copilot.version.engine.version, and "Install AWF binary" failures point tosandbox.agent.version. In most cases the fix is to remove the pin and use the compiled default.5. List the artifacts to read.
agent-stdio.log: the[copilot-harness]lines for model alias resolution and the chosenCOPILOT_PROVIDER_WIRE_API.sandbox/firewall/logs/api-proxy-logs/token-usage.jsonl: the model and path for each request, which shows which endpoint each call actually used.aw_info.json:model,requested_model,cli_version,version, andawf_version.gh aw audit <run>: the combined view.Regression coverage
Add a skill-level eval or worked example for the customer case, so a future change to the skill or debug guide can't bring back the "switch to gpt-4.1" advice. This skill has no eval harness today. The nearest precedents are per-skill
tests/folders such as.github/skills/operational-value-designer/tests/.compiler_version,model: auto, the[copilot-harness]alias and wire-API lines, and the AWF 400 text.engine.versionandsandbox.agent.versionpins. The expected diagnosis identifies the correct field for each, or recommends removing the pin.Acceptance criteria
/agentic-workflowspoints to the outdated gh-aw and the wire-API mismatch, and recommends upgrading and pinning a compatible model before downgrading the model.engine.versionandsandbox.agent.versioncorrectly and doesn't suggestengine.copilot.version..github/aw/debug-agentic-workflow.mdand the rootdebug.mdcarry the same checklist, andSKILL.mdroutes these symptoms to it.