Repository navigation
Add StatusEvidence scorer for agent status reports - #233
Open
Joshua Bauer (ISWT42) wants to merge 2 commits into
Open
Joshua Bauer (ISWT42) wants to merge 2 commits into
Joshua Bauer (ISWT42) wants to merge 2 commits into
Conversation
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This adds
StatusEvidence, a new LLM-as-a-judge scorer in both the Python and the TypeScript package. It checks whether an AI agent's reported status for a task is backed by the task log.inputis the task log or transcript (it can start with the task).outputis the agent's status report.The rule is simple. A status of "done" is only correct when a line in the log shows that the final check passed. If the final check failed, the correct status is "failed". If the final check never ran, was cut off, or is not shown, the correct status is "not shown", which is not done.
The judge picks one of six choices:
Why
Agents often end a run with "done" or "all tests pass". Sometimes the log shows that the last check failed. Sometimes that check never ran at all. These are two different mistakes, and comparing the report to an expected answer does not catch either one, because the evidence is in the log. This scorer reads the report and the log together.
What changed
templates/status_evidence.yaml: the prompt and the choice scores, laid out likefactuality.yaml(lettered options, chain-of-thought on by default).StatusEvidenceinpy/autoevals/llm.py, aSpecFileClassifierlikeSecurityandSql. It is exported through the existingfrom .llm import *.js/templates.ts,StatusEvidenceinjs/llm.ts, and an entry injs/manifest.ts.SCORERS.mdsection with an example in each language, plus the lists inREADME.mdandAGENTS.md.py/autoevals/test_status_evidence.pyandjs/status-evidence.test.ts.No new dependencies and no new infrastructure.
How it was tested
The new tests run offline. They use the same mocking as the existing tests (respx in Python, msw in TypeScript), so they need no API keys. They check that:
The three example logs (final check passed, failed, never ran) are short versions of logs from the public dataset below. Each test file credits the source in a comment.
Everything was run locally with no API keys:
pre-commit run --all-files(black, ruff, codespell, prettier, whitespace) passes, andpnpm run buildsucceeds.test_litellm.py, because the optional extra was not installed). They fail the same way onmain.The tests do not show how accurately a real model picks among the six options, because every model reply is mocked. The scorer has not been run against a live model yet.
To run just the new tests:
uv run --extra dev pytest py/autoevals/test_status_evidence.py pnpm run test -- js/status-evidence.test.tsWhere the idea comes from
The idea comes from my public benchmark, "It Quoted the Failure: Two Kinds of False 'Done'" (Joshua Bauer, 2026): https://dev.to/iswt42/it-quoted-the-failure-two-kinds-of-false-done-1ago. The logs and data are published as a dataset under CC BY 4.0: https://www.kaggle.com/datasets/iswt42/it-quoted-the-failure-evidence.
I prepared this with AI assistance.