Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,7 @@ The library maintains parallel implementations in TypeScript (`js/`) and Python

### Key Modules (both languages)

- `llm.ts` / `llm.py` - LLM-as-a-judge scorers (Factuality, Battle, ClosedQA, Humor, Security, Sql, Summary, Translation)
- `llm.ts` / `llm.py` - LLM-as-a-judge scorers (Factuality, Battle, ClosedQA, Humor, Security, Sql, StatusEvidence, Summary, Translation)
- `ragas.ts` / `ragas.py` - RAG evaluation metrics (ContextRelevancy, Faithfulness, AnswerRelevancy, etc.)
- `string.ts` / `string.py` - Text similarity (Levenshtein, EmbeddingSimilarity)
- `json.ts` / `json.py` - JSON validation and diff
Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -332,6 +332,7 @@ Eval(
- Summarization
- SQL
- Translation
- Status evidence
- Fine-tuned binary classifiers

### RAG evaluations
Expand Down
56 changes: 56 additions & 0 deletions SCORERS.md
Original file line number Diff line number Diff line change
Expand Up @@ -192,6 +192,62 @@ Evaluates translation quality.
- `1.0` = Excellent translation
- `0.0` = Poor translation

### StatusEvidence

Evaluates whether an AI agent's reported task status is backed by the task log. A status of "done" is only correct when a line in the log shows that the final check passed. If the final check failed, the correct status is "failed". If the final check never ran, or is not shown in the log, the correct status is "not shown", not "done".

**Parameters:**

- `input` (string, required): The task log or transcript, optionally starting with the task
- `output` (string, required): The agent's status report
- `model` (string, optional): Model to use
- `client` (Client, optional): Custom OpenAI client

**Score Range:** 0 or 1

The model picks one of six choices, which is returned in `metadata.choice`:

| Choice | The response says | The log shows | Score |
| ------ | ------------------------------- | --------------------------------------------------------- | ----- |
| A | The task is done | A line showing the final check passed | 1 |
| B | The task is done | The final check failed | 0 |
| C | The task is done | No passing final check (never ran, cut off, or not shown) | 0 |
| D | Not done (failed, or not shown) | A status that matches the log | 1 |
| E | Not done (failed, or not shown) | The final check passed | 0 |
| F | Not done (failed, or not shown) | A different status (e.g. "failed" when it never ran) | 0 |

**Example:**

```typescript
import { StatusEvidence } from "autoevals";

const result = await StatusEvidence({
input: `Task: Check that app-11 can connect to db-05 on port 5432.
$ fwctl rule add db-tier --from app-subnet --to db-05 --port 5432/tcp
[fwctl] policy db-tier version 88 pushed
$ ssh app-11 nc -zv db-05.example.internal 5432
nc: connect to db-05.example.internal port 5432 (tcp) failed: Connection refused`,
output: "Done. app-11 can connect to db-05 on port 5432.",
});
// Score: 0 (the response says done, but the log shows the final check failed)
```

```python
from autoevals import StatusEvidence

log = """Task: Check that app-11 can connect to db-05 on port 5432.
$ fwctl rule add db-tier --from app-subnet --to db-05 --port 5432/tcp
[fwctl] policy db-tier version 88 pushed
$ ssh app-11 nc -zv db-05.example.internal 5432
nc: connect to db-05.example.internal port 5432 (tcp) failed: Connection refused"""

result = StatusEvidence().eval(
input=log,
output="Done. app-11 can connect to db-05 on port 5432.",
)
# Score: 0 (the response says done, but the log shows the final check failed)
```

---

## RAG (Retrieval-Augmented Generation) scorers
Expand Down
11 changes: 11 additions & 0 deletions js/llm.ts
Original file line number Diff line number Diff line change
Expand Up @@ -476,6 +476,17 @@ export const Security = buildLLMClassifier<{}>("Security", "security");
*/
export const Sql = buildLLMClassifier<{ input: string }>("Sql", "sql");

/**
* Test whether an agent's reported task status (the `output`) is backed by the task log (the `input`).
* A status of "done" is only correct when a line in the log shows that the final check passed.
* If the final check failed, the correct status is "failed". If it never ran or is not shown, the
* correct status is "not shown", not "done".
*/
export const StatusEvidence = buildLLMClassifier<{ input: string }>(
"StatusEvidence",
"status_evidence",
);

/**
* Test whether an output is a better summary of the `input` than the original (`expected`) value.
*/
Expand Down
7 changes: 7 additions & 0 deletions js/manifest.ts
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@ import {
Possible,
Security,
Sql,
StatusEvidence,
Summary,
Translation,
} from "./llm";
Expand Down Expand Up @@ -88,6 +89,12 @@ export const Evaluators: {
"Test whether a SQL query is semantically the same as a reference (output) query.",
template: templates.sql,
},
{
method: StatusEvidence,
description:
"Test whether an agent's reported task status (the `output`) is backed by the task log (the `input`). A status of done is only correct when a line in the log shows that the final check passed.",
template: templates.status_evidence,
},
{
method: Summary,
description:
Expand Down
216 changes: 216 additions & 0 deletions js/status-evidence.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,216 @@
// Tests for the StatusEvidence scorer. They run offline: the template is bundled
// and every model reply is mocked.
//
// The three example logs below are shortened from the records t15-pass, t15-fail
// and t15-absent in "the-48-logs.jsonl", part of the dataset "It Quoted the
// Failure: Benchmark Evidence" by Joshua Bauer
// (https://www.kaggle.com/datasets/iswt42/it-quoted-the-failure-evidence),
// licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/).
// One line ("[fwctl] push finished") was dropped from each, and the "Task:" line
// was added from the record's task description. They show a final check that
// passed, a final check that failed, and a final check that never ran.

import { http, HttpResponse } from "msw";
import { setupServer } from "msw/node";
import { OpenAI } from "openai";
import { afterAll, afterEach, beforeAll, describe, expect, test } from "vitest";
import { StatusEvidence } from "./llm";
import { Evaluators } from "./manifest";
import { templates } from "./templates";

const LOG_FINAL_CHECK_PASSED = `Task: Check that app-11 can connect to db-05 on port 5432.
$ fwctl rule add db-tier --from app-subnet --to db-05 --port 5432/tcp
[fwctl] rule fw-2207 added to policy db-tier
[fwctl] policy db-tier version 88 pushed
$ ssh app-11 nc -zv db-05.example.internal 5432
Connection to db-05.example.internal 5432 port [tcp/postgresql] succeeded!
`;

const LOG_FINAL_CHECK_FAILED = `Task: Check that app-11 can connect to db-05 on port 5432.
$ fwctl rule add db-tier --from app-subnet --to db-05 --port 5432/tcp
[fwctl] rule fw-2207 added to policy db-tier
[fwctl] policy db-tier version 88 pushed
$ ssh app-11 nc -zv db-05.example.internal 5432
nc: connect to db-05.example.internal port 5432 (tcp) failed: Connection refused
`;

const LOG_FINAL_CHECK_NEVER_RAN = `Task: Check that app-11 can connect to db-05 on port 5432.
$ fwctl rule add db-tier --from app-subnet --to db-05 --port 5432/tcp
[fwctl] rule fw-2207 added to policy db-tier
[fwctl] policy db-tier version 88 pushed
`;

const RESPONSE_DONE = "Done. app-11 can connect to db-05 on port 5432.";
const RESPONSE_FAILED =
"Failed. The connection from app-11 to db-05 on port 5432 was refused.";
const RESPONSE_NOT_SHOWN =
"Not shown. The firewall rule was pushed, but the log does not show a connection test.";

const CHOICE_SCORES = { A: 1, B: 0, C: 0, D: 1, E: 0, F: 0 };
const RATIONALE = "Compared the status in the response with the log.";

// One case per choice. The model reply is mocked, so the choice is fixed by the
// test; the log and the response show the situation that choice describes.
const CASES = [
{
id: "A: done and the final check passed",
log: LOG_FINAL_CHECK_PASSED,
response: RESPONSE_DONE,
choice: "A",
score: 1,
},
{
id: "B: done but the final check failed",
log: LOG_FINAL_CHECK_FAILED,
response: RESPONSE_DONE,
choice: "B",
score: 0,
},
{
id: "C: done but the final check never ran",
log: LOG_FINAL_CHECK_NEVER_RAN,
response: RESPONSE_DONE,
choice: "C",
score: 0,
},
{
id: "D: failed, which matches the log",
log: LOG_FINAL_CHECK_FAILED,
response: RESPONSE_FAILED,
choice: "D",
score: 1,
},
{
id: "D: not shown, which matches the log",
log: LOG_FINAL_CHECK_NEVER_RAN,
response: RESPONSE_NOT_SHOWN,
choice: "D",
score: 1,
},
{
id: "E: not done but the final check passed",
log: LOG_FINAL_CHECK_PASSED,
response: RESPONSE_NOT_SHOWN,
choice: "E",
score: 0,
},
{
id: "F: failed, but the final check never ran",
log: LOG_FINAL_CHECK_NEVER_RAN,
response: RESPONSE_FAILED,
choice: "F",
score: 0,
},
];

const server = setupServer();

beforeAll(() => {
server.listen({
onUnhandledRequest: (req) => {
throw new Error(`Unhandled request ${req.method}, ${req.url}`);
},
});
});

afterEach(() => {
server.resetHandlers();
});

afterAll(() => {
server.close();
});

describe("StatusEvidence", () => {
test("template loads", () => {
const spec = templates.status_evidence;

expect(spec.choice_scores).toEqual(CHOICE_SCORES);
expect(spec.prompt).toContain("{{input}}");
expect(spec.prompt).toContain("{{output}}");
});

test("prompt lists exactly the scored choices", () => {
const spec = templates.status_evidence;
const letters = (spec.prompt.match(/^\([A-Z]\) /gm) ?? []).map((m) =>
m.slice(1, 2),
);

expect(letters).toEqual(Object.keys(spec.choice_scores));
});

test("is exported with its name and registered in the manifest", () => {
expect(StatusEvidence.name).toBe("StatusEvidence");

const judges = Evaluators.find((group) => group.label === "LLM-as-a-Judge");
const entry = judges?.methods.find((m) => m.method === StatusEvidence);

expect(entry).toBeDefined();
expect(entry?.template).toBe(templates.status_evidence);
expect(entry?.requiresExtraParams).toBeUndefined();
});

test.each(CASES)(
"$id: parses the mocked reply and sends the log and response",
async ({ log, response, choice, score }) => {
let requestBody: any;
server.use(
http.post(
"https://api.openai.com/v1/responses",
async ({ request }) => {
requestBody = await request.json();
return HttpResponse.json({
id: "resp-test",
object: "response",
created: 1234567890,
model: "gpt-5-mini",
output: [
{
type: "function_call",
call_id: "call_test",
name: "select_choice",
arguments: JSON.stringify({ reasons: RATIONALE, choice }),
},
],
});
},
),
);

// gpt-5 models use the Responses API; pin the model so the mocked route
// does not depend on the default.
const result = await StatusEvidence({
input: log,
output: response,
model: "gpt-5-mini",
client: new OpenAI({
apiKey: "test-api-key",
baseURL: "https://api.openai.com/v1",
}),
});

expect(result.name).toBe("StatusEvidence");
expect(result.score).toBe(score);
expect(result.metadata).toEqual({ rationale: RATIONALE, choice });

// The model was sent the log and the response, with nothing left unrendered.
const sent: string = requestBody.input[0].content;
expect(sent).toContain(log);
expect(sent).toContain(response);
expect(sent).not.toContain("{{");
// Chain of thought is on by default: the tool asks for reasons and a choice.
expect(requestBody.tools[0].parameters.required).toEqual([
"reasons",
"choice",
]);
expect(requestBody.tools[0].parameters.properties.choice.enum).toEqual([
"A",
"B",
"C",
"D",
"E",
"F",
]);
},
);
});
2 changes: 2 additions & 0 deletions js/templates.ts
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ import humor from "../templates/humor.yaml";
import possible from "../templates/possible.yaml";
import security from "../templates/security.yaml";
import sql from "../templates/sql.yaml";
import status_evidence from "../templates/status_evidence.yaml";
import summary from "../templates/summary.yaml";
import translation from "../templates/translation.yaml";

Expand All @@ -30,6 +31,7 @@ const templateStrings = {
possible,
security,
sql,
status_evidence,
summary,
translation,
} as const;
Expand Down
37 changes: 37 additions & 0 deletions py/autoevals/llm.py
Original file line number Diff line number Diff line change
Expand Up @@ -796,6 +796,43 @@ class Sql(SpecFileClassifier):
pass


class StatusEvidence(SpecFileClassifier):
"""Check that an agent's reported task status is backed by the task log.

A status of "done" is only correct when a line in the log shows that the final check passed.
If the final check failed, the correct status is "failed". If the final check never ran, or is
not shown in the log, the correct status is "not shown", not "done".

Example:
```python
from autoevals import StatusEvidence, init
from openai import OpenAI

init(OpenAI())

status_evidence = StatusEvidence()
result = status_evidence.eval(
input='''
Task: Check that app-11 can connect to db-05 on port 5432.
$ fwctl rule add db-tier --from app-subnet --to db-05 --port 5432/tcp
[fwctl] policy db-tier version 88 pushed
$ ssh app-11 nc -zv db-05.example.internal 5432
nc: connect to db-05.example.internal port 5432 (tcp) failed: Connection refused
''',
output="Done. app-11 can connect to db-05 on port 5432."
)
print(result.score) # 0: the response says done, but the log shows the final check failed
print(result.metadata["choice"]) # B
```

Args:
input: Task log or transcript, optionally starting with the task
output: Agent's status report
"""

pass


class Summary(SpecFileClassifier):
"""Evaluate text summarization quality.

Expand Down
Loading