Skip to content

introduce optional rle reads from parquet - #24227

Merged
kumarUjjawal merged 10 commits into
apache:mainfrom
Rich-T-kid:rich-T-kid/introduce-rle-parquet-flag
Oct 5, 2026
Merged

kumarUjjawal merged 10 commits into
apache:mainfrom
Rich-T-kid:rich-T-kid/introduce-rle-parquet-flag

Conversation

@Rich-T-kid

@Rich-T-kid Rich-T-kid commented Aug 10, 2026 •

Copy link
Copy Markdown
Contributor

a large portion of this PR is test!

Which issue does this PR close?

Rationale for this change

When DataFusion reads a parquet file with dictionary-encoded string or binary columns, it currently decodes the dictionary and returns plain Utf8/Binary arrays, discarding the encoding. For low-cardinality columns (status, country, category, etc.) this doesn't take full advantage of the compacted format parquet gives the engine Preserving the dictionary encoding reduces memory usage and can improve aggregation performance on these columns.

What changes are included in this PR?

Adds datafusion.execution.parquet.enable_rle_to_dictionary with a default of false.

When enabled for inferred-schema Parquet tables, DataFusion inspects Parquet footer metadata and promotes top-level string/binary columns with dictionary pages to Arrow dictionary types. Mixed dictionary/plain files are normalized before schema merge when the value types are compatible. At scan time, the parquet opener passes the promoted schema to arrow-rs so those columns can be read as dictionary arrays directly.

Tables with a user-supplied schema are not promoted because DataFusion does not use footer metadata to infer their schema.

Are these changes tested?

yes.

  • datafusion/sqllogictest/test_files/parquet_rle_to_dictionary.slt

  • datafusion/datasource-parquet/src/schema_coercion.rs

    • uniform_dict_schemas_respects_value_type_families checks that mixed file schemas are normalized only across compatible string/binary families.
    • rle_schema_coercion_respects_dictionary_value_type checks scan-time coercion into dictionary types, including incompatible cases and the flag-off path.
  • datafusion/datasource-parquet/src/opener/mod.rs

    • test_rle_binary_column_promotion verifies the opener passes a promoted Dictionary(Int32, Binary) schema to arrow-rs so binary columns can be read as dictionary arrays directly.

Are there any user-facing changes?

New session config option: SET datafusion.execution.parquet.enable_rle_to_dictionary = true. Default is false so existing behavior is unchanged.

@github-actions github-actions Bot added sqllogictest SQL Logic Tests (.slt) common Related to common crate datasource Changes to the datasource crate labels Aug 10, 2026
Comment thread datafusion/datasource-parquet/src/opener/mod.rs Outdated
Comment thread datafusion/common/src/config.rs Outdated
@github-actions github-actions Bot added the auto detected api change Auto detected API change label Aug 10, 2026
@Rich-T-kid
Rich-T-kid marked this pull request as ready for review August 10, 2026 15:45
@github-actions github-actions Bot added proto Related to proto crate documentation Improvements or additions to documentation labels Aug 10, 2026
@codecov-commenter

codecov-commenter commented Aug 10, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 95.87302% with 26 lines in your changes missing coverage. Please review.
✅ Project coverage is 82.68%. Comparing base (5741773) to head (aceff34).
⚠️ Report is 1 commits behind head on main.

Files with missing lines Patch % Lines
...tafusion/datasource-parquet/src/schema_coercion.rs 97.23% 6 Missing and 7 partials ⚠️
datafusion/proto-common/src/generated/pbjson.rs 26.66% 8 Missing and 3 partials ⚠️
datafusion/datasource-parquet/src/metadata.rs 96.29% 2 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main   #24227      +/-   ##
==========================================
+ Coverage   82.66%   82.68%   +0.01%     
==========================================
  Files        1147     1147              
  Lines      446522   447137     +615     
  Branches   446522   447137     +615     
==========================================
+ Hits       369134   369713     +579     
- Misses      54988    55009      +21     
- Partials    22400    22415      +15     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

@adriangb pinging you since you seem interested in parquet related speed ups 👍

@adriangb

Copy link
Copy Markdown
Contributor

If I understand correctly the goal is to evaluate filters during filter pushdown against dictionary / RLE encoded columns? We can't propagate these dynamic type changes to the rest of the query plan / scan. Is that right?

@Rich-T-kid

Rich-T-kid commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor Author

@adriangb

If I understand correctly the goal is to evaluate filters during filter pushdown against dictionary / RLE encoded columns?

no not exactly. The goal of this PR is to keep RLE parquet columns in their compacted form by materializing them as dictionary arrays instead of regular strings.

We can't propagate these dynamic type changes to the rest of the query plan / scan. Is that right?

exactly! This is why it needs to be done as early in the plan as possible. we inspect DFParquetMetadata::fetch_schema to see if any RLE columns exist, if so change the plan type from utf8/binary to dict<_,utf8/binary>. Thanks to some plumbing in arrow-rs it will handle the conversions correctly returning a dictionary array, this can cause a 60x perf boost

@Rich-T-kid
Rich-T-kid force-pushed the rich-T-kid/introduce-rle-parquet-flag branch from 0479f9c to 6e1cbc1 Compare August 13, 2026 15:10
@adriangb

Copy link
Copy Markdown
Contributor

no not exactly. The goal of this PR is to keep RLE parquet columns in their compacted form by materializing them as dictionary arrays instead of regular strings.

Why only RLE and not dictionaries as well? How does this compare to / relate to the schema_force_view_types option?

It also looks like this goes through infer_schema right? A lot of code paths never touch that (CREATE EXTERNAL TABLE (a VARCHAR), any custom table providers, etc.).

I'd be more interested in seeing something at the parquet scan level that was able to e.g. optimize how row filters are applied by applying them to the dictionary instead of expanding into Utf8View. That would be applicable to all DataFusion users.

@Rich-T-kid
Rich-T-kid force-pushed the rich-T-kid/introduce-rle-parquet-flag branch from 6a4e89d to d1273bf Compare August 13, 2026 15:48
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

Why only RLE and not dictionaries as well? How does this compare to / relate to the schema_force_view_types option?

my bad when I say RLE i'm referring to RLE_DICTIONARY

It also looks like this goes through infer_schema right? A lot of code paths never touch that (CREATE EXTERNAL TABLE (a VARCHAR), any custom table providers, etc.).

I'd be more interested in seeing something at the parquet scan level that was able to e.g. optimize how row filters are applied by applying them to the dictionary instead of expanding into Utf8View. That would be applicable to all DataFusion users.

I agree, ill update the PR to target all parquet scans.

@github-actions github-actions Bot added the core Core DataFusion crate label Aug 13, 2026
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

ideally we surface columns that are physically RLE_DICTIONARY-encoded in the parquet file as Arrow Dictionary(Int32, Utf8) arrays rather than decoding them back to plain Utf8.

To know whether a specific column is RLE_DICTIONARY-encoded you need to read the parquet file footer. For the infer_schema path this happens at table registration, but for explicit schemas (CREATE EXTERNAL TABLE (col VARCHAR)) and direct ParquetSource construction no footer is ever read during planning, so per-column encoding information isn't available for all paths.

Downstream physical operators (FilterExec, AggregateExec) are compiled against the scan's declared output schema during physical planning, before any files are opened. If the scan declares Utf8 but produces Dictionary(Int32, Utf8) at execution time that's a type mismatch.

So when the flag is enabled we promote all string/binary columns to dict at planning time, not just the ones that are actually RLE-encoded, because that's the only way to guarantee schema consistency across all parquet scan paths without introducing file I/O into the planning stage.

I feel like i'm missing something here. if we could take a peak at the parquets metadata before physical planning and change the schema for all operators from the point forward that would be perfect. Im not sure this is currently possible

@Rich-T-kid
Rich-T-kid marked this pull request as draft August 13, 2026 19:53
@Rich-T-kid
Rich-T-kid force-pushed the rich-T-kid/introduce-rle-parquet-flag branch from 717d31a to fc360e8 Compare August 14, 2026 17:27
@Rich-T-kid

Rich-T-kid commented Aug 14, 2026 •

Copy link
Copy Markdown
Contributor Author

@adriangb When the flag is on, DataFusion promotes string and binary columns that are physically RLE_DICTIONARY encoded in the parquet file to Dictionary(Int32, Utf8) or Dictionary(Int32, Binary) at schema inference time. This applies to all parquet scans regardless of how the table was registered.

the PR is ready for review

@Rich-T-kid
Rich-T-kid marked this pull request as ready for review August 14, 2026 17:36
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

How does this compare to / relate to the schema_force_view_types option?

utf8 columns become dict<_,utf8 so the default string type for parquet columns utf8View is replaced.

@Rich-T-kid
Rich-T-kid force-pushed the rich-T-kid/introduce-rle-parquet-flag branch 6 times, most recently from bc5011f to 6899579 Compare August 15, 2026 18:45
@kumarUjjawal

Copy link
Copy Markdown
Contributor

Thank you @Rich-T-kid for moving this forward. Since this is a larger change I will need second pair of eyes on this before this can be approved. Let's see if anyone has time to take a look. Can you share the work in discord that might bring some people over.

@Rich-T-kid
Rich-T-kid force-pushed the rich-T-kid/introduce-rle-parquet-flag branch from cb3a578 to 3c359b4 Compare September 27, 2026 21:36
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

@kumarUjjawal seems we didn't get much of a response. @alamb & @adriangb have some context on what were doing, they may have input

cc @asolimando for visibility

Comment thread datafusion/datasource-parquet/src/schema_coercion.rs Outdated
Comment thread datafusion/datasource-parquet/src/schema_coercion.rs Outdated
…ssion in schema coercion

- Drop rle_column_allowlist from ParquetFormat and DFParquetMetadata: no callers
  existed and new public API is hard to remove post-release
- Remove Dictionary arm from transform_schema_to_view and transform_binary_to_string:
  those arms fired unconditionally regardless of enable_rle_to_dictionary, causing
  pre-existing dict columns to get wrong value types when the flag was off
- Remove the unit test that covered the now-deleted transform_binary_to_string dict arm
- Drop Utf8View/BinaryView arms added to common_dictionary_value_type: the parquet
  reader never produces view types in the physical file schema so they were dead code
- Clean up parquet_rle_to_dictionary.slt: remove redundant SET and spurious RESET
…w in SLT

- common_dictionary_key_type: add Int64 arm before catch-all so UInt64+Int64
  selects UInt64 (wider) rather than silently narrowing to Int64
- update test expectation from Int64 to UInt64 for that case
- SLT high-cardinality test: write rle.parquet via arrow_cast to
  Dictionary(Int8, Utf8) so the file embeds an Int8-keyed Arrow schema;
  add a standalone check confirming the round-trip; this exercises the
  actual Dict(Int8)+plain-130-values overflow path that the previous
  plain-Utf8 write never triggered
…type_coercions_with_rle

Replace the duplicate full implementation with a call to the existing
coerce_fields_by_name helper, extended with an enable_rle flag threaded
through coerce_data_type and coerce_child. This restores correct
positional map-entry coercion via coerce_map_entries and avoids
allocating a Vec for every field when nothing changes.
@Rich-T-kid
Rich-T-kid force-pushed the rich-T-kid/introduce-rle-parquet-flag branch from ae6b03b to d514497 Compare September 28, 2026 07:31
@kumarUjjawal

Copy link
Copy Markdown
Contributor

Lets keep this open for few days so others can comment if interested.

@alamb

alamb commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Is the goal of this PR to improve performance for some usecases? If so, do we have any benchmark results that show the performance improves?

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

Is the goal of this PR to improve performance for some usecases? If so, do we have any benchmark results that show the performance improves?

@alamb yes & no. The end goal is faster queries & providing users with more control over how their data is read. This initial PR solves the second problem. As for the query speedup, I would have expected a larger speedup for certain queries. As we suspected, and as @yinli-systems pointed out here, there is a cutoff for the usefulness of dictionaries. I made a follow-up issue to address this and bridge the performance gaps.

On a slightly unrelated note, it will be much easier to track performance changes and gains with this PR, as we can just run the regular benchmark suite with this flag enabled. This means a faster iteration cycle.

I plan on going through the list of optimizations pointed out in the issue and then running the benchmarks with the flag enabled vs disabled. for any other discrepancies with performance ill do in in depth profile to see where we may be missing optimizations

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

I can rebase this before we merge

@kumarUjjawal

Copy link
Copy Markdown
Contributor

I can rebase this before we merge

Please resolve the conflicts.

…e-rle-parquet-flag

# Conflicts:
#	datafusion/common/src/config.rs
#	datafusion/common/src/file_options/parquet_writer.rs
#	datafusion/datasource-parquet/src/opener/mod.rs
#	datafusion/datasource-parquet/src/source.rs
#	datafusion/proto-common/proto/datafusion_common.proto
#	datafusion/proto-common/src/generated/pbjson.rs
#	datafusion/proto-common/src/generated/prost.rs
#	datafusion/proto-models/src/generated/datafusion_proto_common.rs
#	docs/source/user-guide/configs.md
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

I can rebase this before we merge

Please resolve the conflicts.

@kumarUjjawal fixed. should be good to go now 🚀

@kumarUjjawal

Copy link
Copy Markdown
Contributor

Thank you @Rich-T-kid

@kumarUjjawal
kumarUjjawal added this pull request to the merge queue Oct 5, 2026
Merged via the queue into apache:main with commit 5e5d79d Oct 5, 2026
43 checks passed
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

thanks @kumarUjjawal 🙇

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

auto detected api change Auto detected API change common Related to common crate datasource Changes to the datasource crate documentation Improvements or additions to documentation proto Related to proto crate sqllogictest SQL Logic Tests (.slt) v56.0.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Allow users to enable dictionary column reads from parquet files

7 participants