Skip to content

Keep a mixed document's parse metrics when the provider emits no layout - #153

Merged
boyang-zhang1 merged 2 commits into
fix/131-layout-order-rules-droppedfrom
boyang/multi-task-eval-hardening
Sep 14, 2026
Merged

boyang-zhang1 merged 2 commits into
fix/131-layout-order-rules-droppedfrom
boyang/multi-task-eval-hardening

Conversation

@boyang-zhang1

@boyang-zhang1 boyang-zhang1 commented Sep 14, 2026 •

Copy link
Copy Markdown
Member

Stacked on #133 (its commit is the first commit here, authorship preserved). Groundwork for #154. A document that carries both layout rules and parse rules routes to _evaluate_multi_task, and that path has three losses that only become reachable once such a document exists — which is exactly what #133 enables. This lands them first so #133 can go in without moving anyone's score in a way nobody expects.

No shipped dataset carries a mixed document today (data/layout.jsonl is 100% type: "layout", and the group key is (category, pdf)), so this changes no current leaderboard number. Verified by loading data/ and data/test/ before and after: identical.

1. A markdown-only provider scored 0 on a document it partly passed

create_layout_adapter_for_result sat outside the try block. For any provider with no layout output (openai, anthropic, gemini, reducto…) it raises, _evaluate_single turns the document into success=False, metrics=[], and _aggregate_metrics drops every metric of a failed result — including the parse rules that had just passed — before _is_infra_failure zero-pads it.

Reproduced on main with one layout rule + one present rule and a markdown-only parse result:

main: rules log "1/1 passed", then success=False n_metrics=0
      error='Evaluation error: Inference output is not LayoutOutput and no provider adapter matched.'
this: success=True  metrics=[('rule_pass_rate', 1.0), ('rule_present_pass_rate', 1.0)]

Only the default adapter's "not LayoutOutput" ValueError is caught — a real adapter bug still propagates and still fails the result, which is what _is_infra_failure exists to sort out (covered by a test).

When no adapter matches, the provider emitted nothing the layout task can score, so the layout half scores a genuine 0 — the same verdict _is_infra_failure reaches for a pure-layout document. An empty LayoutOutput tagged with a new LayoutDetectionModel.NONE sentinel is run through the evaluator, which emits the full metric set with its own denominators, so the document stays in the AP/F1/layout_rule_pass_rate denominators. Leaving it unscored (the first revision of this PR) let one perfect document plus one markdown-only document aggregate to avg_AP50 = 1.0.

When an adapter matched but found no elements, that is scoreable and is now scored instead of failed. The evaluator runs against the empty prediction set and returns a genuine 0 across the whole metric set, with the denominators it builds itself:

rule_pass_rate 1.0 {passed: 1, total: 1}     <- parse half intact
layout_rule_pass_rate 0.0 {passed: 0, total: 2}   <- localization + classification
AP50 0.0   mean_f1 0.0   num_ground_truth 1.0   unmatched_gt_elements 1.0

That branch previously appended "Could not extract layout from PARSE output" and failed the whole result.

2. The split parse test case lost the document's parse configuration

It was rebuilt with expected_markdown=None and default table settings, so a mixed document was scored without the markdown GT that drives text similarity / TEDS / GriTS, and with the title-strip and TRM fallback reset — a document flagged trm_unsupported: true was scored with the metric it was flagged unreliable for. It now inherits expected_markdown, allow_splitting_ambiguous_merged_tables, trm_unsupported, max_top_title_rows and metadata.

Separately, metadata was read only from LayoutDetectionTestCase, so source_dataset was silently None whenever the mixed rules arrived on a ParseTestCase — which is the case the sidecar loader already produces. It is read by attribute now, and a LayoutDetectionTestCase's own source_dataset field is preferred over the metadata copy.

3. rule_pass_rate was emitted twice into one result

The layout evaluator emits rule_pass_rate as an alias of the layout_rule_pass_rate right beside it, and RuleBasedMetric emits the same name. A mixed result carried both, so one document contributed two samples to avg_*, pooled unrelated denominators into micro_*, and the CSV export and detailed report — which flatten with {m.metric_name: m.value} — kept whichever came last, making the per-example view disagree with the rule detail pane under it.

The layout alias is dropped when the parse half already emitted that name. Not folded arithmetically: the parse entry carries the rule_results payload the report renders, and its value is a graduated score rather than passed / total, so combining them by counts would erase partial credit and drop the rule details. layout_rule_pass_rate keeps the layout half's number and subcounts.

Tests

tests/parse_bench/evaluation/test_multi_task_evaluation.py, 10 tests: the markdown-only case for both test-case classes, the two-document aggregate that must not read 100% when one document has no layout, the empty-layout-output case scoring a real zero with the evaluator's own denominators, a real adapter bug still failing the result, config and source_dataset inheritance from either carrier, and the alias-drop rules. ruff check, ruff format --check and the evaluation + test_cases + analysis suites (2596 tests) are clean.

@boyang-zhang1
boyang-zhang1 force-pushed the boyang/multi-task-eval-hardening branch from 7384bae to a80e773 Compare September 14, 2026 15:56
@boyang-zhang1
boyang-zhang1 force-pushed the boyang/multi-task-eval-hardening branch 2 times, most recently from b0f3ffc to 745ecd5 Compare September 14, 2026 16:27
A document carrying both layout and parse rules routes to
`_evaluate_multi_task`, which resolved the layout adapter outside its error
boundary. A markdown-only parse provider has nothing to adapt, so the call
raised, `_evaluate_single` turned the whole document into
`success=False, metrics=[]`, and `_aggregate_metrics` dropped every metric of a
failed result — including the parse rules the provider had just passed — before
`_is_infra_failure` zero-padded the document on top of that.

Catch only that condition (the default adapter's "not LayoutOutput"
`ValueError`) and leave the layout half unscored for the document, so the parse
half still counts. A real adapter bug still propagates and still fails the
result, which is what `_is_infra_failure` exists to sort out.

An adapter that *matched* but found no elements is a different case and is now
scored rather than failed: the evaluator runs against the empty prediction set
and returns a genuine 0 across the whole metric set, with the denominators it
builds itself — localization and classification are separate checks, so an
element count would under-count them. That branch used to append "Could not
extract layout from PARSE output" and fail the result.

Two related losses on the same path:

- The split parse test case was rebuilt with `expected_markdown=None` and
  default table settings, so a mixed document was scored without the markdown
  GT that drives text similarity/TEDS/GriTS and with the title-strip and TRM
  fallback reset. It now inherits `expected_markdown`,
  `allow_splitting_ambiguous_merged_tables`, `trm_unsupported`,
  `max_top_title_rows` and `metadata` from the document.
- `metadata` was read only from `LayoutDetectionTestCase`, silently dropping
  `source_dataset` when the mixed rules arrive on a `ParseTestCase`. It is read
  by attribute now, and a `LayoutDetectionTestCase`'s own `source_dataset`
  field is preferred over the metadata copy.

The layout evaluator emits `rule_pass_rate` as an alias of the
`layout_rule_pass_rate` beside it, so a mixed result carried two entries under
one name: the document contributed two samples to `avg_*`, pooled unrelated
denominators into `micro_*`, and the CSV and detailed report kept whichever came
last. Drop the alias when the parse half already emitted that name — the parse
entry carries the `rule_results` payload the report renders, and its value is a
graduated score rather than `passed / total`, so the two cannot be combined
arithmetically.

No shipped dataset carries a mixed document today, so none of this moves
current leaderboard numbers.
@boyang-zhang1
boyang-zhang1 force-pushed the boyang/multi-task-eval-hardening branch from 745ecd5 to e133626 Compare September 14, 2026 17:01
@boyang-zhang1
boyang-zhang1 changed the base branch from main to fix/131-layout-order-rules-dropped September 14, 2026 17:01
When no layout adapter matches a mixed document's inference result, the
layout half was left unscored. One perfect document plus one markdown-only
document then aggregated to avg_AP50 = 1.0, because the missing document
dropped out of the layout denominators. That is the opposite verdict from
the pure-layout path, where _is_infra_failure counts "not LayoutOutput" as a
genuine provider zero.

Score an empty LayoutOutput instead, tagged with a new LayoutDetectionModel
sentinel NONE. The evaluator emits the full metric set with its own
denominators, so the document counts as 0 in AP/F1/layout_rule_pass_rate
while its parse half is untouched.

Also move the test stub adapters off LLAMAPARSE: that label mapper lazily
imports llama_cloud, which is only installed with the runners extra, so the
empty-layout test failed under a plain dev install.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@boyang-zhang1
boyang-zhang1 added this pull request to stack #155 September 14, 2026 17:19
@boyang-zhang1
boyang-zhang1 merged commit 8a84d81 into fix/131-layout-order-rules-dropped Sep 14, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants