Skip to content

agentic-workflows skill: add a checklist for model and engine misconfiguration #67486

Description

@SivaKesava1

Summary

The /agentic-workflows skill doesn't diagnose model and engine misconfiguration well. In a customer case (Slack, 9–10 Oct), a workflow with model: auto failed with:

400 Cannot translate Copilot request feature 'tools[custom]' … gpt-5.6-luna-utility … /responses is incompatible

Running /agentic-workflows on that failure suggested switching to model: gpt-4.1. That works around the problem without fixing it. pelikhan: "Did you try to fix this with our skill? If it can't handle we should improve that."

The same thread also mixed up the version fields:

Setting tried What it actually is What happened
engine.version: 1.94.0 The Copilot CLI version, which is 1.0.x Failed at "Install GitHub Copilot CLI" with bare curl 404s
sandbox.agent.version: 1.94.0 The AWF release version Failed at "Install AWF binary"
engine.copilot.version Not a field —

The real causes were:

  1. An old gh-aw. v0.89.21 is still the only non-prerelease, so gh extension install and gh extension upgrade install it. It predates the wire-API inference from Infer Copilot wire API for utility model variants #64177. The newest prerelease is v0.91.7.
  2. A wire-API mismatch between the resolved model and the endpoint the Copilot CLI used.

This issue proposes adding a "Model and engine misconfiguration" checklist that the skill uses when it sees these symptoms.

Plan

Where the guidance goes

  • .github/skills/agentic-workflows/SKILL.md already sends debugging to .github/aw/debug-agentic-workflow.md. Add a routing line so these symptoms go straight to the new checklist:
    • AWF 400s that mention models or endpoints
    • model: auto failures
    • install-step 404s after a version pin
    • questions about which version field to set
  • In .github/aw/debug-agentic-workflow.md, add a "Model and engine misconfiguration" section. Link to it from "Collect Existing Evidence" and "Identify the First Failing Boundary".
  • Add matching, shorter guidance to the root debug.md, which external agents load from the lock-file header link. It can link to the full section.

What the checklist covers

1. Check the gh-aw version first.

  • Find the version the workflow was compiled with:
    • compiler_version in the gh-aw-metadata lock-file header
    • cli_version in aw_info.json
  • Compare it with the newest gh-aw release, including prereleases. Explain that gh extension install and gh extension upgrade only install the latest non-prerelease, so a user can be several releases behind without knowing it. Explain how to install a specific prerelease tag.
  • If the version predates the relevant fix, recommend upgrading and recompiling before trying anything else.

2. Explain how the Copilot harness picks the model and wire API.

  • How an alias like auto resolves to a concrete model.
  • How COPILOT_PROVIDER_WIRE_API is chosen, in this order:
    • an explicit engine.env override
    • the catalog wire_api
    • the -utility → base-model fallback
    • the gpt-5+ name rule
    • the CLI default, /chat/completions
  • That the CLI uses one wire API for the whole session, including sub-agents.
  • How to read these AWF 400s as a model/endpoint mismatch, not a transient failure or a prompt problem:
    • Cannot translate Copilot request feature
    • Unsupported Responses custom tool
    • model_policy_violation
    • not accessible via the … endpoint

3. Recommend fixes in this order.

  1. Upgrade gh-aw and recompile.
  2. Pin a model that supports the endpoint the run needs.
  3. Remove a conflicting COPILOT_PROVIDER_WIRE_API override.
  4. For sub-agents, use a model from the main model's family. See [aw] PR Code Quality Reviewer had a request rejected #67460 and BYOK: sub-agents always use the session's wire API, so a sub-agent on a model from the other family fails with a 400 copilot-cli#5103.

Recommend switching to an older model only when none of these work. Say why that's the last resort: it trades capability for a workaround and leaves the real misconfiguration in place.

4. Explain the version fields.

  • engine.version is the agent CLI version. For Copilot, that's the Copilot CLI, which is 1.0.x.
  • sandbox.agent.version is the AWF release version, in vX.Y.Z form, and must match a GitHub release of gh-aw-firewall.
  • There's no engine.copilot.version.
  • Map each failing install step to the field that controls it: "Install GitHub Copilot CLI" 404s point to engine.version, and "Install AWF binary" failures point to sandbox.agent.version. In most cases the fix is to remove the pin and use the compiled default.

5. List the artifacts to read.

  • agent-stdio.log: the [copilot-harness] lines for model alias resolution and the chosen COPILOT_PROVIDER_WIRE_API.
  • sandbox/firewall/logs/api-proxy-logs/token-usage.jsonl: the model and path for each request, which shows which endpoint each call actually used.
  • aw_info.json: model, requested_model, cli_version, version, and awf_version.
  • gh aw audit <run>: the combined view.

Regression coverage

Add a skill-level eval or worked example for the customer case, so a future change to the skill or debug guide can't bring back the "switch to gpt-4.1" advice. This skill has no eval harness today. The nearest precedents are per-skill tests/ folders such as .github/skills/operational-value-designer/tests/.

  • Fixture: a redacted lock-file header with an old compiler_version, model: auto, the [copilot-harness] alias and wire-API lines, and the AWF 400 text.
  • Expected diagnosis: upgrade gh-aw and recompile, then pin a model that supports the needed endpoint. It shouldn't recommend an older model first.
  • Second case: the misplaced engine.version and sandbox.agent.version pins. The expected diagnosis identifies the correct field for each, or recommends removing the pin.

Acceptance criteria

  • Debugging the customer's run with /agentic-workflows points to the outdated gh-aw and the wire-API mismatch, and recommends upgrading and pinning a compatible model before downgrading the model.
  • The skill explains engine.version and sandbox.agent.version correctly and doesn't suggest engine.copilot.version.
  • .github/aw/debug-agentic-workflow.md and the root debug.md carry the same checklist, and SKILL.md routes these symptoms to it.
  • The eval or example fails if the guidance regresses.

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions