Skip to content

fix(halueval): read Yes/No verdicts as whole words, not substrings - #1800

Open
Zhuoxi2000 wants to merge 1 commit into
modelscope:mainfrom
Zhuoxi2000:fix-halu-eval-yes-no-substring
Open

Zhuoxi2000 wants to merge 1 commit into
modelscope:mainfrom
Zhuoxi2000:fix-halu-eval-yes-no-substring

Conversation

@Zhuoxi2000

Copy link
Copy Markdown

What

HaluEvalAdapter.match_score credited a sample when reference in filtered_prediction.strip().upper(). For a NO label, any word that contains "no" (not, now, note, know, knowledge, ...) turned a Yes reply into a correct answer. A reply naming both verdicts (Yes and No: ...) was credited for either label, and a reply with no verdict (I do not know.) was credited for No. aggregate_scores derives the predicted class from acc, so precision, recall, f1_score and yes_ratio were affected too.

This PR applies the rule from the official evaluator (RUCAIBox/HaluEval evaluation/evaluate.py): a reply is correct only when it contains exactly one verdict and that verdict equals the label. Replies containing both verdicts or neither are incorrect.

How

  • New extract_verdict() returns 'YES' / 'NO' when the reply contains exactly one verdict and None otherwise. match_score credits the sample only when that verdict equals the reference.
  • Verdicts are matched as whole words, either as the prompt writes them (Yes / No) or in capitals (YES / NO). The official evaluator uses a case-sensitive substring check ("No" in ans), so it treats Note ... and Nothing ... as a "No". Matching whole words avoids that. Lowercase prose such as "there is no mention of ..." after a Yes is still read as a single Yes verdict, as in the official script. This is stricter than the reference about word boundaries and otherwise follows its exactly-one-verdict rule. It does not reproduce the reference byte for byte. One visible change: a bare lowercase yes / no was credited before and is not now, which matches the official script and the prompt (The answer you give MUST be "Yes" or "No"). If you prefer case-insensitive matching, I can switch, at the cost of reading lowercase prose such as "no mention" as a second verdict.
  • aggregate_scores is unchanged. It derives the prediction from acc and is correct once acc is.
  • HaluEval evaluation_version is bumped from v1.0 (default) to v1.1 because the scoring semantics change, following the Evaluation versioning rules in AGENTS.md. BenchmarkMeta.description is unchanged, so no _meta / docs regeneration is needed.

Tests

New tests/benchmark/test_halu_eval_scoring.py (14 cases): Yes replies containing not / now / Note with a No label, a both-verdict reply under each label, a reply with no verdict, plain verdicts, and single verdicts followed by an explanation (including YES, NO. and lowercase no in prose).

# before (main @ cfb9af3, test file only)
python -m pytest tests/benchmark/test_halu_eval_scoring.py -q
6 failed, 8 passed

# after
python -m pytest tests/benchmark/test_halu_eval_scoring.py -q
14 passed

python -m pytest tests/api/test_evaluation_versioning.py tests/api/test_benchmark_registry_contract.py tests/api/judge/test_gates.py -q
48 passed

pre-commit run --files evalscope/benchmarks/halu_eval/halu_eval_adapter.py tests/benchmark/test_halu_eval_scoring.py   # all passed
lint-imports   # 5 kept, 0 broken

The same reference in filtered_prediction.strip().upper() pattern exists in pope, hallusion_bench and drivelology_binary. Their prompts ask for a bare YES/NO, so they are less exposed. This PR is intentionally scoped to HaluEval, and I can follow up on those if wanted.

AI assistance: this change was drafted with an AI coding assistant (Claude) and verified locally with the tests above.

HaluEvalAdapter.match_score credited a sample when the label appeared
anywhere in reference in prediction.upper(). For a "No" label any word
containing "no" (not, now, note, know, knowledge, another, ...) made a
"Yes" reply count as correct, and a reply naming both verdicts was
credited for either label. aggregate_scores rebuilds the predicted class
from acc, so precision, recall, f1_score and yes_ratio inherited the
error.

Follow the official HaluEval evaluator (RUCAIBox/HaluEval
evaluation/evaluate.py): a reply is correct only when it contains exactly
one of the verdicts "Yes" / "No" and that verdict equals the label;
replies with both or neither are incorrect. Verdicts are matched as whole
words, as written in the prompt or in capitals, so prose like "not" or
"there is no mention" is not read as a verdict. Bump HaluEval
evaluation_version to v1.1 because the scoring semantics change.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant