Keep a mixed document's parse metrics when the provider emits no layout - #153
Merged
boyang-zhang1 merged 2 commits intoSep 14, 2026
Conversation
boyang-zhang1
force-pushed
the
boyang/multi-task-eval-hardening
branch
from
September 14, 2026 15:56
7384bae to
a80e773
Compare
boyang-zhang1
force-pushed
the
boyang/multi-task-eval-hardening
branch
2 times, most recently
from
September 14, 2026 16:27
b0f3ffc to
745ecd5
Compare
A document carrying both layout and parse rules routes to `_evaluate_multi_task`, which resolved the layout adapter outside its error boundary. A markdown-only parse provider has nothing to adapt, so the call raised, `_evaluate_single` turned the whole document into `success=False, metrics=[]`, and `_aggregate_metrics` dropped every metric of a failed result — including the parse rules the provider had just passed — before `_is_infra_failure` zero-padded the document on top of that. Catch only that condition (the default adapter's "not LayoutOutput" `ValueError`) and leave the layout half unscored for the document, so the parse half still counts. A real adapter bug still propagates and still fails the result, which is what `_is_infra_failure` exists to sort out. An adapter that *matched* but found no elements is a different case and is now scored rather than failed: the evaluator runs against the empty prediction set and returns a genuine 0 across the whole metric set, with the denominators it builds itself — localization and classification are separate checks, so an element count would under-count them. That branch used to append "Could not extract layout from PARSE output" and fail the result. Two related losses on the same path: - The split parse test case was rebuilt with `expected_markdown=None` and default table settings, so a mixed document was scored without the markdown GT that drives text similarity/TEDS/GriTS and with the title-strip and TRM fallback reset. It now inherits `expected_markdown`, `allow_splitting_ambiguous_merged_tables`, `trm_unsupported`, `max_top_title_rows` and `metadata` from the document. - `metadata` was read only from `LayoutDetectionTestCase`, silently dropping `source_dataset` when the mixed rules arrive on a `ParseTestCase`. It is read by attribute now, and a `LayoutDetectionTestCase`'s own `source_dataset` field is preferred over the metadata copy. The layout evaluator emits `rule_pass_rate` as an alias of the `layout_rule_pass_rate` beside it, so a mixed result carried two entries under one name: the document contributed two samples to `avg_*`, pooled unrelated denominators into `micro_*`, and the CSV and detailed report kept whichever came last. Drop the alias when the parse half already emitted that name — the parse entry carries the `rule_results` payload the report renders, and its value is a graduated score rather than `passed / total`, so the two cannot be combined arithmetically. No shipped dataset carries a mixed document today, so none of this moves current leaderboard numbers.
boyang-zhang1
force-pushed
the
boyang/multi-task-eval-hardening
branch
from
September 14, 2026 17:01
745ecd5 to
e133626
Compare
boyang-zhang1
changed the base branch from
main
to
fix/131-layout-order-rules-dropped
September 14, 2026 17:01
When no layout adapter matches a mixed document's inference result, the layout half was left unscored. One perfect document plus one markdown-only document then aggregated to avg_AP50 = 1.0, because the missing document dropped out of the layout denominators. That is the opposite verdict from the pure-layout path, where _is_infra_failure counts "not LayoutOutput" as a genuine provider zero. Score an empty LayoutOutput instead, tagged with a new LayoutDetectionModel sentinel NONE. The evaluator emits the full metric set with its own denominators, so the document counts as 0 in AP/F1/layout_rule_pass_rate while its parse half is untouched. Also move the test stub adapters off LLAMAPARSE: that label mapper lazily imports llama_cloud, which is only installed with the runners extra, so the empty-layout test failed under a plain dev install. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
boyang-zhang1
added this pull request to stack #155
September 14, 2026 17:19
logan-markewich
approved these changes
Sep 14, 2026
boyang-zhang1
merged commit Sep 14, 2026
8a84d81
into
fix/131-layout-order-rules-dropped
4 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #133 (its commit is the first commit here, authorship preserved). Groundwork for #154. A document that carries both layout rules and parse rules routes to
_evaluate_multi_task, and that path has three losses that only become reachable once such a document exists — which is exactly what #133 enables. This lands them first so #133 can go in without moving anyone's score in a way nobody expects.No shipped dataset carries a mixed document today (
data/layout.jsonlis 100%type: "layout", and the group key is(category, pdf)), so this changes no current leaderboard number. Verified by loadingdata/anddata/test/before and after: identical.1. A markdown-only provider scored 0 on a document it partly passed
create_layout_adapter_for_resultsat outside the try block. For any provider with no layout output (openai, anthropic, gemini, reducto…) it raises,_evaluate_singleturns the document intosuccess=False, metrics=[], and_aggregate_metricsdrops every metric of a failed result — including the parse rules that had just passed — before_is_infra_failurezero-pads it.Reproduced on
mainwith one layout rule + onepresentrule and a markdown-only parse result:Only the default adapter's
"not LayoutOutput"ValueErroris caught — a real adapter bug still propagates and still fails the result, which is what_is_infra_failureexists to sort out (covered by a test).When no adapter matches, the provider emitted nothing the layout task can score, so the layout half scores a genuine 0 — the same verdict
_is_infra_failurereaches for a pure-layout document. An emptyLayoutOutputtagged with a newLayoutDetectionModel.NONEsentinel is run through the evaluator, which emits the full metric set with its own denominators, so the document stays in the AP/F1/layout_rule_pass_ratedenominators. Leaving it unscored (the first revision of this PR) let one perfect document plus one markdown-only document aggregate toavg_AP50 = 1.0.When an adapter matched but found no elements, that is scoreable and is now scored instead of failed. The evaluator runs against the empty prediction set and returns a genuine 0 across the whole metric set, with the denominators it builds itself:
That branch previously appended
"Could not extract layout from PARSE output"and failed the whole result.2. The split parse test case lost the document's parse configuration
It was rebuilt with
expected_markdown=Noneand default table settings, so a mixed document was scored without the markdown GT that drives text similarity / TEDS / GriTS, and with the title-strip and TRM fallback reset — a document flaggedtrm_unsupported: truewas scored with the metric it was flagged unreliable for. It now inheritsexpected_markdown,allow_splitting_ambiguous_merged_tables,trm_unsupported,max_top_title_rowsandmetadata.Separately,
metadatawas read only fromLayoutDetectionTestCase, sosource_datasetwas silentlyNonewhenever the mixed rules arrived on aParseTestCase— which is the case the sidecar loader already produces. It is read by attribute now, and aLayoutDetectionTestCase's ownsource_datasetfield is preferred over the metadata copy.3.
rule_pass_ratewas emitted twice into one resultThe layout evaluator emits
rule_pass_rateas an alias of thelayout_rule_pass_rateright beside it, andRuleBasedMetricemits the same name. A mixed result carried both, so one document contributed two samples toavg_*, pooled unrelated denominators intomicro_*, and the CSV export and detailed report — which flatten with{m.metric_name: m.value}— kept whichever came last, making the per-example view disagree with the rule detail pane under it.The layout alias is dropped when the parse half already emitted that name. Not folded arithmetically: the parse entry carries the
rule_resultspayload the report renders, and its value is a graduated score rather thanpassed / total, so combining them by counts would erase partial credit and drop the rule details.layout_rule_pass_ratekeeps the layout half's number and subcounts.Tests
tests/parse_bench/evaluation/test_multi_task_evaluation.py, 10 tests: the markdown-only case for both test-case classes, the two-document aggregate that must not read 100% when one document has no layout, the empty-layout-output case scoring a real zero with the evaluator's own denominators, a real adapter bug still failing the result, config andsource_datasetinheritance from either carrier, and the alias-drop rules.ruff check,ruff format --checkand the evaluation + test_cases + analysis suites (2596 tests) are clean.