Repository navigation
55.1 upgrade - #180
Closed
gene-bordegaray wants to merge 14 commits into
Closed
55.1 upgrade#180gene-bordegaray wants to merge 14 commits into
gene-bordegaray wants to merge 14 commits into
Conversation
…ize (apache#24384) (apache#24529) ## Which issue does this PR close? - Part of apache#24462 - Backport of apache#24384 to `branch-55` (for 55.1.0, tracked in apache#24462). - Fixes apache#24383 ## Rationale for this change `UnnestExec` emitted exactly one output batch per input batch, however many rows the unnesting produced, never consulting `datafusion.execution.batch_size`. This let downstream operators receive arbitrarily large batches and made peak memory scale with input batch size times list length instead of `batch_size`. No API change, so it fits the backport criteria. ## What changes are included in this PR? Clean cherry-pick of apache#24384 (commit 6187d47). No adaptation required. ## Are these changes tested? Yes. Carries the original regression coverage, all tests pass. ## Are there any user-facing changes? `UnnestExec` output batches now respect `datafusion.execution.batch_size` as an upper bound. No API changes. Co-authored-by: Andy Grove <agrove@apache.org>
…s adaptation (apache#24125) (apache#24530) ## Which issue does this PR close? - Part of apache#24462 - Backport of apache#24125 to `branch-55` (for 55.1.0, tracked in apache#24462). - Fixes apache#24109 ## Rationale for this change With `datafusion.execution.parquet.pushdown_filters = true`, a predicate on a struct field was reported as fully handled by the scan whenever the file needed schema adaptation, so `FilterExec` was removed from the plan and the predicate was silently dropped — returning every row instead of the filtered set. This is a correctness bug (wrong results), not specific to 55.0.0, so it fits the backport criteria. ## What changes are included in this PR? Cherry-pick of apache#24125 (commit 40c208e). Git's recursive merge auto-resolved surrounding context differences in `datafusion/physical-expr-adapter/src/schema_rewriter.rs` and `datafusion/sqllogictest/test_files/parquet_nested_schema_pruning.slt`; no manual conflict resolution or adaptation of the fix itself was required. One follow-up commit adapts a test expectation: `branch-55` doesn't have apache#24130/apache#24315, which taught nested schema pruning to union the leaves needed by mixed whole-column + field-access reads (e.g. `select s, s['y'] from narrow`). Without that optimization, the mixed-access case falls back to reading every physical leaf, so `bytes_scanned` is `219` here instead of the `146` the original PR's test expects on `main`. This is a pre-existing difference in pruning capability, not a correctness regression from this fix. ## Are these changes tested? Yes. Carries the original regression coverage, all tests pass. ## Are there any user-facing changes? `WHERE s['field'] = ...` predicates on struct columns now filter correctly when Parquet filter pushdown requires schema adaptation. No API changes. --------- Co-authored-by: Adrian Garcia Badaracco <1755071+adriangb@users.noreply.github.com> Co-authored-by: Claude <noreply@anthropic.com>
…te_physical_plan (apache#24690) (apache#24694) This PR is a cherry-pick of apache#24690 onto `branch-55`. See the original PR for a description of the issue. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…uction (apache#24723) (apache#24752) This is a back port of apache#24723 onto `branch-55` to support `datafusion-python` upgrade to 55.1.0. The details can be found in the linked PR. This is needed for apache/datafusion-python#1677 Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
… to planned schema in aggregation (apache#24394) (apache#24699) This is a back port of apache#24394 for `branch-55`. The original issue is apache#24069 Note that the original PR was marked as an auto-detected api change, but I believe that is a false positive. --------- Co-authored-by: Patrick Ribbsaeter <patrickswedish@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
… and new_default (apache#24582) (apache#24876) This is a back port of apache#24582 into `branch-55` for inclusion in 55.1.0. Please see the original PR for details. Co-authored-by: Unik Dahal <61407386+unikdahal@users.noreply.github.com>
This is a small update to the Cargo.lock file to update - chacha20 v0.10.0 -> v0.10.2 - h2 v0.4.13 -> v0.4.19 The first was due to a yanked crate. The second was due to a security vulnerability, RUSTSEC-2026-0258. The yanked crate is a warning in cargo audit and will auto-resolve on building to the newer 0.10.2, and the security vulnerability is already updated in `main`.
…asts (apache#23169) (apache#24875) This PR is a backport of apache#23169 onto `branch-55`. I made one update the `use` statement in `planner.rs`. Co-authored-by: Dewey Dunnington <dewey@wherobots.com> Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
…ache#24670) (apache#24992) Backports apache#24670 to branch-55. Co-authored-by: Tim Saucer <timsaucer@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This PR updates the version number and changelog for the 55.1.0 release.
- Relates to apache#22989 but does not close it. More types are coming. - Support the Arrow type `Dictionary` for `approx_distinct` - Enable `HLLAccumulator` and `HllGroupsAccumulator` to support `Dictionary` - Tests Yes Yes, `approx_distinct` supports now `Dictionary` but no breaking changes.
…r-defined types (apache#25015) ## Which issue does this PR close? N/A ## Rationale for this change `consume_user_defined_type` maps a Substrait user-defined type straight to a bare `DataType`, so a consumer has no way to mark the resulting `Field` for a downstream reader to recover semantic type lost in that mapping. For example, a UDT representing JSON is commonly mapped to plain `Utf8`, which is then indistinguishable from a real string column once it reaches a consumer that only sees the Arrow schema (e.g. over Arrow Flight, with no access to the original plan). ## What changes are included in this PR? - Adds `consume_user_defined_type_metadata`, a second, additive extension point on `SubstraitConsumer` defaulting to `Ok(None)`. - Applies its result to the `Field` built in `from_substrait_struct_type`, via `Field::with_metadata`. - Existing consumers are unaffected: the default keeps today's behavior exactly (no metadata attached). ## Are these changes tested? Yes, two new unit tests: a consumer that supplies metadata for a `UserDefined` field gets it attached to the field (data type unchanged), and a consumer returning `None` (the default) attaches nothing. ## Are there any user-facing changes? No behavior change for existing consumers. This is a new, optional trait method with a default implementation. (cherry picked from commit d682553)
…2026/09/datafusion-55.1-datadog-patches # Conflicts: # datafusion/expr/src/expr_schema.rs # datafusion/physical-plan/src/projection.rs
Author
|
Closing because this revision is maintained as the standalone branch |
gene-bordegaray
deleted the
gene.bordegaray/2026/09/datafusion-55.1-datadog-patches
branch
September 9, 2026 16:38
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## branch-55 #180 +/- ##
=============================================
+ Coverage 81.15% 81.22% +0.06%
=============================================
Files 1110 1110
Lines 386654 388726 +2072
Branches 386654 388726 +2072
=============================================
+ Hits 313798 315739 +1941
- Misses 54365 54434 +69
- Partials 18491 18553 +62 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
55.1 upgrade