Skip to content

perf(parquet): skip Bloom reads for fully matched row groups - #25854

Merged
xudong963 merged 1 commit into
apache:mainfrom
xudong963:xudong963/skip-bloom-on-fully-matched-row-groups
Sep 30, 2026
Merged

xudong963 merged 1 commit into
apache:mainfrom
xudong963:xudong963/skip-bloom-on-fully-matched-row-groups

Conversation

@xudong963

@xudong963 xudong963 commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Which issue does this PR close?

No linked issue. This is a follow-up to the fully matched Parquet row-group work in PR #23696.

Rationale for this change

When row-group statistics prove that every row satisfies a scan predicate, a Bloom filter cannot prune that group. Reading its Bloom filters still adds object-store I/O, which can be especially costly for remote files.

What changes are included in this PR?

  • Skip Bloom filter reads and predicate evaluation for fully matched row groups.
  • Avoid creating a Bloom reader when every surviving row group is fully matched.
  • Keep the Bloom pruning matched metric accounting for skipped groups.

What is the testing strategy for this PR?

The new fully_matched_row_groups_skip_bloom_filter_reads test verifies that a partially matched group still reads Bloom filters, a fully matched group reduces bytes_scanned, and an all-fully-matched file reads zero Bloom bytes during open. It also checks that the returned rows are unchanged.

Are there any user-facing changes?

No API or query-result changes. Scans avoid unnecessary Bloom filter reads for fully matched row groups.

@github-actions github-actions Bot added the datasource Changes to the datasource crate label Sep 29, 2026
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 82.57%. Comparing base (1bd2ad3) to head (85e9608).
⚠️ Report is 13 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main   #25854      +/-   ##
==========================================
+ Coverage   82.51%   82.57%   +0.05%     
==========================================
  Files        1141     1142       +1     
  Lines      439774   440946    +1172     
  Branches   439774   440946    +1172     
==========================================
+ Hits       362896   364123    +1227     
+ Misses      54939    54822     -117     
- Partials    21939    22001      +62     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@xudong963

Copy link
Copy Markdown
Member Author

Thanks @zhuqi-lucas

@xudong963
xudong963 added this pull request to the merge queue Sep 30, 2026
Merged via the queue into apache:main with commit f427fb7 Sep 30, 2026
43 checks passed
@xudong963
xudong963 deleted the xudong963/skip-bloom-on-fully-matched-row-groups branch September 30, 2026 06:35
martin-g pushed a commit to martin-g/datafusion that referenced this pull request Sep 30, 2026
…25854)

## Which issue does this PR close?

No linked issue. This is a follow-up to the fully matched Parquet
row-group work in PR apache#23696.

## Rationale for this change

When row-group statistics prove that every row satisfies a scan
predicate, a Bloom filter cannot prune that group. Reading its Bloom
filters still adds object-store I/O, which can be especially costly for
remote files.

## What changes are included in this PR?

- Skip Bloom filter reads and predicate evaluation for fully matched row
groups.
- Avoid creating a Bloom reader when every surviving row group is fully
matched.
- Keep the Bloom pruning matched metric accounting for skipped groups.

## What is the testing strategy for this PR?

The new `fully_matched_row_groups_skip_bloom_filter_reads` test verifies
that a partially matched group still reads Bloom filters, a fully
matched group reduces `bytes_scanned`, and an all-fully-matched file
reads zero Bloom bytes during open. It also checks that the returned
rows are unchanged.


## Are there any user-facing changes?

No API or query-result changes. Scans avoid unnecessary Bloom filter
reads for fully matched row groups.
@rgbuilds

Copy link
Copy Markdown

Thanks for this optimization. In row_group_filter.rs, I noticed this branch records a Bloom match even though Bloom reads and evaluation are skipped:

if self.access_plan.is_fully_matched(idx) {
    metrics.row_groups_pruned_bloom_filter.add_matched(1);
    continue;
}

This overlaps with #25822 (addressing #18355), which proposes counting only actual Bloom evaluation outcomes. I’ll incorporate this new path and extend the test to cover its metric accounting when updating that PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

datasource Changes to the datasource crate v56.0.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants