Repository navigation
fix(halueval): read Yes/No verdicts as whole words, not substrings - #1800
Open
Zhuoxi2000 wants to merge 1 commit into
Open
Zhuoxi2000 wants to merge 1 commit into
Zhuoxi2000 wants to merge 1 commit into
Conversation
HaluEvalAdapter.match_score credited a sample when the label appeared anywhere in reference in prediction.upper(). For a "No" label any word containing "no" (not, now, note, know, knowledge, another, ...) made a "Yes" reply count as correct, and a reply naming both verdicts was credited for either label. aggregate_scores rebuilds the predicted class from acc, so precision, recall, f1_score and yes_ratio inherited the error. Follow the official HaluEval evaluator (RUCAIBox/HaluEval evaluation/evaluate.py): a reply is correct only when it contains exactly one of the verdicts "Yes" / "No" and that verdict equals the label; replies with both or neither are incorrect. Verdicts are matched as whole words, as written in the prompt or in capitals, so prose like "not" or "there is no mention" is not read as a verdict. Bump HaluEval evaluation_version to v1.1 because the scoring semantics change.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
HaluEvalAdapter.match_scorecredited a sample whenreference in filtered_prediction.strip().upper(). For aNOlabel, any word that contains "no" (not,now,note,know,knowledge, ...) turned aYesreply into a correct answer. A reply naming both verdicts (Yes and No: ...) was credited for either label, and a reply with no verdict (I do not know.) was credited forNo.aggregate_scoresderives the predicted class fromacc, soprecision,recall,f1_scoreandyes_ratiowere affected too.This PR applies the rule from the official evaluator (RUCAIBox/HaluEval
evaluation/evaluate.py): a reply is correct only when it contains exactly one verdict and that verdict equals the label. Replies containing both verdicts or neither are incorrect.How
extract_verdict()returns'YES'/'NO'when the reply contains exactly one verdict andNoneotherwise.match_scorecredits the sample only when that verdict equals the reference.Yes/No) or in capitals (YES/NO). The official evaluator uses a case-sensitive substring check ("No" in ans), so it treatsNote ...andNothing ...as a "No". Matching whole words avoids that. Lowercase prose such as "there is no mention of ..." after aYesis still read as a singleYesverdict, as in the official script. This is stricter than the reference about word boundaries and otherwise follows its exactly-one-verdict rule. It does not reproduce the reference byte for byte. One visible change: a bare lowercaseyes/nowas credited before and is not now, which matches the official script and the prompt (The answer you give MUST be "Yes" or "No"). If you prefer case-insensitive matching, I can switch, at the cost of reading lowercase prose such as "no mention" as a second verdict.aggregate_scoresis unchanged. It derives the prediction fromaccand is correct onceaccis.evaluation_versionis bumped fromv1.0(default) tov1.1because the scoring semantics change, following the Evaluation versioning rules in AGENTS.md.BenchmarkMeta.descriptionis unchanged, so no_meta/ docs regeneration is needed.Tests
New
tests/benchmark/test_halu_eval_scoring.py(14 cases):Yesreplies containingnot/now/Notewith aNolabel, a both-verdict reply under each label, a reply with no verdict, plain verdicts, and single verdicts followed by an explanation (includingYES,NO.and lowercasenoin prose).The same
reference in filtered_prediction.strip().upper()pattern exists inpope,hallusion_benchanddrivelology_binary. Their prompts ask for a bare YES/NO, so they are less exposed. This PR is intentionally scoped to HaluEval, and I can follow up on those if wanted.AI assistance: this change was drafted with an AI coding assistant (Claude) and verified locally with the tests above.