Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/outcome-collector.lock.yml

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

24 changes: 15 additions & 9 deletions .github/workflows/outcome-collector.md
Original file line number Diff line number Diff line change
Expand Up @@ -97,12 +97,15 @@ The summary JSON produced by the pre-agent step includes:
| `total_outcomes` | int | Actionable items evaluated (excludes noops) |
| `accepted` | int | Items kept/merged/resolved |
| `rejected` | int | Items undone/dismissed/removed |
| `ignored` | int | Items with no observable follow-up within the window |
| `ignored` | int | Actions that reached a state without the expected effect |
Comment thread
pelikhan marked this conversation as resolved.
| `pending` | int | Items not yet at a terminal state |
| `unknown` | int | Missing execution evidence, unsupported evaluators, or existence-only checks |
| `errors` | int | API or evaluation failures |
| `lifecycle` | int | Bot lifecycle closures, including retained close actions |
| `noop` | int | Non-actionable items (noops, missing_tool, etc.) |
| `accepted_strong` | int | Accepted with strong evidence (merged, completed, approved) |
| `accepted_medium` | int | Accepted with medium evidence (engagement, retention) |
| `accepted_weak` | int | Accepted with weak evidence (object still exists) |
| `accepted_weak` | int | Legacy weak acceptance; existence-only checks now count as unknown |
| `fallback_exists_only_count` | int | Items evaluated using only the generic existence fallback — a data quality signal |
| `acceptance_rate` | float | `accepted / (accepted + rejected)` |
| `waste_rate` | float | `rejected / total_outcomes` |
Expand Down Expand Up @@ -159,8 +162,8 @@ List concrete actions the team should take based on the data directly under the
2. **Stuck pending items** — List any items pending >48 hours or any workflow classified as 🔴 stuck. These need human review or the workflow needs a timeout.
3. **Underdefined workflows** — Any workflow classified as ⚪ underdefined needs clearer acceptance/rejection criteria or a dedicated evaluator. The outcome model for that workflow is not yet mature.
4. **Low zero-touch workflows** — Workflows where accepted items always need human edits indicate the agent's output quality needs improvement.
5. **High ignored rate** — If ignored items exceed 30% of total outcomes, the workflow may be producing outputs that nobody engages with; consider refining targeting or output type.
6. **Data quality: fallback evaluations** — If `fallback_exists_only_count` > 20% of total outcomes, many items were evaluated with only a generic existence check (weak signal). This means the acceptance numbers may be overstated; note this in the report.
5. **High ignored rate** — If ignored items exceed 30% of total outcomes, inspect actions that completed without their expected effect (for example neutral/skipped dispatch conclusions or ready-for-review actions closed without a qualifying review). Recommend fixing those action-specific paths, not inferring lack of engagement.
6. **Data quality: fallback evaluations** — If `fallback_exists_only_count` > 20% of total outcomes, many items have only generic existence evidence and remain unknown. Note this coverage gap; do not count them as accepted.

**Lifecycle health classification** — assign one label per workflow based on its outcome history:

Expand Down Expand Up @@ -193,27 +196,30 @@ Place all detailed metrics, numeric breakdowns, evidence quality, and trends ins
| Accepted | {accepted} / {total_outcomes} | — |
| — strong evidence | {accepted_strong} | merged, completed, approved |
| — medium evidence | {accepted_medium} | engaged, retained |
| — weak evidence | {accepted_weak} | existence only |
| — weak evidence | {accepted_weak} | legacy weak acceptance |
| Rejected | {rejected} | — |
| Ignored | {ignored} | no observable follow-up |
| Ignored | {ignored} | expected effect not observed |
| Zero-touch | {zero_touch} / {accepted} | — |
| Pending | {pending} | — |
| Unknown | {unknown} | insufficient evidence |
| Evaluation errors | {errors} | verification failed |
| Lifecycle | {lifecycle} | bot closures, not rejection |
| Runs checked | {runs_checked} | — |

### Per-Workflow Breakdown

For each workflow with outcomes, show a mini-scorecard:

| Workflow | Accepted | Rejected | Ignored | Pending | Acceptance | Zero-touch |
|----------|----------|----------|---------|---------|------------|------------|
| Workflow | Accepted | Rejected | Ignored | Pending | Unknown | Errors | Lifecycle | Acceptance | Zero-touch |
|----------|----------|----------|---------|---------|---------|--------|-----------|------------|------------|

Sort by waste rate descending (worst first).

### Evidence Quality

If `fallback_exists_only_count` > 0, include this note:

> ⚠️ **{fallback_exists_only_count} item(s)** were evaluated using only a generic existence check (signal: `target_exists_only`). These contribute to `accepted_weak` and may overstate acceptance. Dedicated evaluators for `add_reviewer`, `submit_pull_request_review`, `update_issue`, `update_pull_request`, and other types provide stronger evidence.
> ⚠️ **{fallback_exists_only_count} item(s)** have only generic existence evidence (signal: `target_exists_only`). They count as `unknown`, not accepted. Dedicated action-specific evaluators require execution evidence and provide stronger attribution.

### Trend Signal

Expand Down
6 changes: 6 additions & 0 deletions actions/setup/js/emit_outcome_spans.cjs
Original file line number Diff line number Diff line change
Expand Up @@ -191,6 +191,9 @@ async function main() {
if (reactionsPositive !== null) attributes.push(buildAttr("gh-aw.outcome.reactions_positive", reactionsPositive));
if (reactionsNegative !== null) attributes.push(buildAttr("gh-aw.outcome.reactions_negative", reactionsNegative));
if (comments !== null) attributes.push(buildAttr("gh-aw.outcome.comments", comments));
for (const field of ["human_comments", "human_reviews", "human_edits"]) {
if (typeof eval_[field] === "number") attributes.push(buildAttr(`gh-aw.outcome.${field}`, eval_[field]));
}
if (zeroTouch) attributes.push(buildAttr("gh-aw.outcome.zero_touch", true));

// Map normalized outcome_status to OTLP status: accepted=OK, rejected=ERROR, all others=UNSET
Expand Down Expand Up @@ -225,6 +228,9 @@ async function main() {
buildAttr("gh-aw.outcome.rejected", getSummaryNumber("rejected", 0)),
buildAttr("gh-aw.outcome.ignored", getSummaryNumber("ignored", 0)),
buildAttr("gh-aw.outcome.pending", getSummaryNumber("pending", 0)),
buildAttr("gh-aw.outcome.unknown", getSummaryNumber("unknown", 0)),
buildAttr("gh-aw.outcome.errors", getSummaryNumber("errors", 0)),
buildAttr("gh-aw.outcome.lifecycle", getSummaryNumber("lifecycle", 0)),
buildAttr("gh-aw.outcome.noop", getSummaryNumber("noop", 0)),
buildAttr("gh-aw.outcome.accepted_strong", getSummaryNumber("accepted_strong", 0)),
buildAttr("gh-aw.outcome.accepted_medium", getSummaryNumber("accepted_medium", 0)),
Expand Down
12 changes: 12 additions & 0 deletions actions/setup/js/emit_outcome_spans.test.cjs
Original file line number Diff line number Diff line change
Expand Up @@ -195,6 +195,9 @@ describe("emit_outcome_spans.cjs", () => {
rejected: 1,
ignored: 0,
pending: 0,
unknown: 1,
errors: 2,
lifecycle: 3,
noop: 0,
accepted_strong: 1,
accepted_medium: 0,
Expand Down Expand Up @@ -231,6 +234,9 @@ describe("emit_outcome_spans.cjs", () => {
reactions_positive: 4,
reactions_negative: 1,
comments: 0,
human_comments: 0,
human_reviews: 0,
human_edits: 0,
zero_touch: true,
}),
JSON.stringify({
Expand Down Expand Up @@ -302,6 +308,9 @@ describe("emit_outcome_spans.cjs", () => {
expect(summarySpan.attributes).toContainEqual({ key: "gh-aw.outcome.date", value: "2026-05-13" });
expect(summarySpan.attributes).toContainEqual({ key: "gh-aw.outcome.zero_touch_count", value: 1 });
expect(summarySpan.attributes).toContainEqual({ key: "gh-aw.outcome.accepted_strong", value: 1 });
expect(summarySpan.attributes).toContainEqual({ key: "gh-aw.outcome.unknown", value: 1 });
expect(summarySpan.attributes).toContainEqual({ key: "gh-aw.outcome.errors", value: 2 });
expect(summarySpan.attributes).toContainEqual({ key: "gh-aw.outcome.lifecycle", value: 3 });
expect(summarySpan.attributes).toContainEqual({ key: "gh-aw.outcome.fallback_exists_only_count", value: 1 });
expect(spans[1].attributes).toContainEqual({ key: "gh-aw.exporter.name", value: "outcome-collector" });
expect(spans[1].attributes).toContainEqual({ key: "gh-aw.outcome.url", value: "https://github.com/github/gh-aw/issues/1" });
Expand All @@ -318,6 +327,9 @@ describe("emit_outcome_spans.cjs", () => {
expect(spans[1].attributes).toContainEqual({ key: "gh-aw.outcome.reactions_positive", value: 4 });
expect(spans[1].attributes).toContainEqual({ key: "gh-aw.outcome.reactions_negative", value: 1 });
expect(spans[1].attributes).toContainEqual({ key: "gh-aw.outcome.comments", value: 0 });
expect(spans[1].attributes).toContainEqual({ key: "gh-aw.outcome.human_comments", value: 0 });
expect(spans[1].attributes).toContainEqual({ key: "gh-aw.outcome.human_reviews", value: 0 });
expect(spans[1].attributes).toContainEqual({ key: "gh-aw.outcome.human_edits", value: 0 });
expect(spans[1].attributes).toContainEqual({ key: "gh-aw.outcome.zero_touch", value: true });
expect(spans[2].attributes.find(attr => attr.key === "gh-aw.outcome.review_comments")).toBeUndefined();
expect(spans[2].attributes.find(attr => attr.key === "gh-aw.outcome.changed_files")).toBeUndefined();
Expand Down
Loading
Loading