Repository navigation
Conversation
## Which issue does this PR close? - Closes apache#24970. ## Rationale for this change `INSERT` currently aligns source expressions with target columns by storage `DataType` only. If both fields use the same storage type but the target is an Arrow extension type, the cast is skipped and the DML projection loses the target extension metadata. A table sink can then receive a schema that does not describe the target logical type. ## What changes are included in this PR? After storage-type coercion, compare the resulting field metadata with the target table field. When they differ, use the existing field-aware `Cast` representation to enforce the target metadata. Existing casts are replaced rather than nested. The regression uses an unannotated `FixedSizeBinary(16)` source and an `arrow.uuid` target so the behavior is independent of any particular downstream consumer. Existing parameterized UUID INSERT snapshots are updated to show the target field-aware cast. ## What is the testing strategy for this PR? - Added `plan_insert_preserves_target_extension_metadata`. - Ran all 575 `datafusion-sql` integration tests. - Ran `cargo clippy -p datafusion-sql --all-targets --all-features -- -D warnings`. - Ran `cargo fmt --all -- --check`. ## Are there any user-facing changes? Yes. SQL `INSERT` plans now preserve target Arrow field metadata, including extension metadata. There are no public API changes.
## Which issue does this PR close? Related to apache#18083. Full support for NULL partition values remains covered by that issue. ## Rationale for this change Partitioned writes read partition values without checking their validity. A NULL integer partition key can be written under `p=0`, so reading the output turns NULL into zero and silently changes the data. For example: ```sql CREATE TABLE src(p INT, v INT) AS VALUES (NULL, 1), (0, 2); COPY src TO '/tmp/null_partition/' STORED AS PARQUET PARTITIONED BY (p); ``` Both rows can end up in the zero partition. This PR rejects the write with `NULL values are not supported for partition column 'p'`. ## What changes are included in this PR? - Check logical nulls before extracting partition keys, including NULL dictionary values referenced by valid dictionary indices. - Continue accepting nullable columns when the selected rows contain no NULL values. - Add focused unit tests and a `copy.slt` regression covering rejection and a filtered successful write. ## What is the testing strategy for this PR? - The rejection regressions fail without the production fix and pass with it. - Focused partition-key tests: 5 passed with `RUST_BACKTRACE=1`. - `cargo clippy --all-targets --all-features -- -D warnings` passed. - The extended workspace Rust suites passed after excluding four host memory-measurement runners that also fail on current main in this environment. - The SQL logic harness passed all 513 files, including encrypted Parquet coverage. - `cargo fmt --all -- --check` and `git diff --check` passed. ## Are there any user-facing changes? Partitioned writes containing NULL partition keys now return an execution error. Callers can filter or replace these values before writing. Validation happens per batch, so files written by earlier batches can remain after an error.
## Which issue does this PR close? Related to apache#24258. This covers runtime operands, floating-point literals, and volatile expressions in compound IN-list rewrites. The NULL-semantics changes in that PR remain separate. ## Rationale for this change The optimizer performs set intersection and difference using expression equality. Different expressions can evaluate to the same value, so these rewrites can remove matching rows even when every column is non-nullable: ```sql CREATE TABLE in_list_values(x INT NOT NULL, a INT NOT NULL, b INT NOT NULL) AS VALUES (1, 1, 1); SELECT x FROM in_list_values WHERE x IN (a, 2, 10, 11) AND x IN (b, 5, 12, 13); ``` The query should return `1`; the rewrite produces an empty intersection and returns no rows. Structural comparison also distinguishes positive and negative floating-point zero. Deduplicating volatile operands can change how many times they are evaluated. ## What changes are included in this PR? - Restrict intersection and difference rewrites to non-null, non-floating literals. - Guard compound rewrites against volatile inputs and list items. - Preserve valid union rewrites for nonvolatile runtime expressions. - Add unit regressions for runtime values, signed zero, mixed integer coercion, and volatility, plus runtime-value cases in `predicates.slt`. ## What is the testing strategy for this PR? - The new regressions fail without the production fix and pass with it. - Focused optimizer tests: 6 passed. - `cargo clippy --all-targets --all-features -- -D warnings` passed. - The extended workspace Rust suites passed after excluding four host memory-measurement runners that also fail on current main in this environment. - The SQL logic harness passed all 513 files, including encrypted Parquet coverage. - `cargo fmt --all -- --check` and `git diff --check` passed. ## Are there any user-facing changes? Affected compound IN predicates retain matching rows and preserve volatile expression evaluations. No public API changes. ## Downsides Queries excluded by the new guards may perform more comparisons because the optimizer retains their original predicates.
## Which issue does this PR close? No linked issue. The regression below reproduces the problem on current upstream. ## Rationale for this change Grouped CORR uses raw sums whose subtraction loses precision when values have a large common offset. It can return a negative correlation for a column correlated with itself: ```sql SELECT g, corr(x, x) FROM (VALUES (1, 1000000000.0), (1, 1000000007.0), (1, 1000000015.0), (2, 1.0) ) AS t(g, x) GROUP BY g; ``` Group 1 can return `-1` instead of `1`. ## What changes are included in this PR? - Use paired Welford updates and centered moments for grouped correlation. - Merge partial states using weighted mean differences, applying weights before multiplying deltas to avoid avoidable intermediate overflow. - Match the scalar CORR state layout across update, merge, singleton conversion, and preserving/draining emission. - Add regressions for large offsets, multiple batch sizes, scalar/grouped state interchange, and large-range merges, plus a grouped self-correlation case in `aggregate.slt`. ## What is the testing strategy for this PR? - The large-offset and scalar/grouped state compatibility regressions fail without the production fix and pass with it. - Focused correlation tests: 7 passed. - `cargo clippy --all-targets --all-features -- -D warnings` passed. - The extended workspace Rust suites passed after excluding four host memory-measurement runners that also fail on current main in this environment. - The SQL logic harness passed all 513 files, including encrypted Parquet coverage. - `cargo fmt --all -- --check` and `git diff --check` passed. - H2O groupby q09 returned the same 6,320,797 rows on every run. Across 20 release-mode runs, current main averaged 479.5 ms and this PR averaged 373.7 ms. Excluding each process's first run, the median was 333.9 ms on main and 312.7 ms on this PR, so the stability fix showed no regression in this environment. ## Are there any user-facing changes? Grouped CORR produces stable results for the covered large-offset inputs. Floating-point results can differ in their final bits. SQL signatures and output types remain unchanged. ## Downsides Centered updates add arithmetic, including per-row divisions. The grouped intermediate state representation changes from raw sums to centered moments, so partial states produced by the old implementation cannot be mixed with the new representation.
## Which issue does this PR close? Closes a planner gap where `ORDER BY ALL` only accepted raw column projections and was silently ignored for set-operation outputs. ## Rationale for this change `ORDER BY ALL` means sorting by every output column in select-list order. Expanding it to the equivalent 1-based ordinal keys supports computed expressions, aliases, aggregate outputs, wildcard-expanded projections, and set-operation outputs through one planner path. The ordinal expansion also sorts already-projected columns rather than evaluating computed expressions again. Physical execution remains on DataFusion's existing vectorized `SortExec`; this adds no row-wise conversion or custom physical operator. ## What changes are included? - expand `OrderByKind::All` to ordinal `OrderByExpr`s from the output width - provide output width for non-`SELECT` set expressions instead of dropping `ORDER BY ALL` - add focused execution coverage for computed expressions, null-order options, set operations, and aggregate output ## Are these changes tested? - `cargo +1.95.0 test -p datafusion-sqllogictest --test sqllogictests -- order_by_all.slt --test-threads 1` - `cargo +1.95.0 test -p datafusion-sql --test sql_integration -- --test-threads 8` (591 passed) - `cargo +1.95.0 clippy -p datafusion-sql --all-targets -- -D warnings` - `cargo +1.95.0 fmt --all -- --check` Snowflake dialect parsing support is proposed independently in apache/datafusion-sqlparser-rs#2502; this planner change is generic to every dialect that already emits `OrderByKind::All`.
## Which issue does this PR close? - Closes apache#23920 ## Rationale for this change just more uses of the new property and comment on some non strictly order preserving for why not ## What changes are included in this PR? marked if something is order preserving or not, for those who are not I've added a comment why not, since some expressions might be confusing why not ## Are these changes tested? no ## Are there any user-facing changes? not api ones
## Which issue does this PR close? - Part of apache#21048. The issue stays open for the other checklist items. ## Rationale for this change `dev/rust_lint.sh` is the local mirror of the CI lint jobs, but it does not run the Markdown link check. CI runs `ci/scripts/markdown_link_check.sh` in `dev.yml`. A developer can push a broken internal link after a clean local lint run and then see the failure only in CI. This PR adds the existing checker to the local suite. ## What changes are included in this PR? - `dev/rust_lint.sh` runs `ci/scripts/markdown_link_check.sh` as a read-only step, after the workflow install check and before the Rust documentation build. The `--write` and `--allow-dirty` flags never reach it. - The runner loads `LYCHEE_VERSION` from `ci/scripts/utils/tool_versions.sh` and installs that version when `lychee` is missing. An installed `lychee` is used as is, which is the runner's existing policy for other tools. - `ci/scripts/markdown_link_check.sh` becomes executable, because the runner invokes each registered script directly. The script body, `lychee.toml`, the file selection, and the GitHub workflow do not change. - `docs/source/contributor-guide/testing.md` documents the new behavior and keeps the standalone instructions. ## What is the testing strategy for this PR? - A disposable fixture with stubbed steps and tools: the checker runs once with no arguments in check mode and in both write modes, a missing `lychee` triggers exactly the pinned install command, a present `lychee` triggers no install, and a checker failure stops the suite before the later steps. - A disposable clone with an injected broken internal link: the standalone checker and `./dev/rust_lint.sh` both fail with the same exit code and name the file. After the fix, the full suite passes. - The real checker and the full `./dev/rust_lint.sh` pass on this branch with `lychee` 0.23.0. ## Are there any user-facing changes? No. Developers get the link check in the local lint suite. CI is unchanged. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
## Which issue does this PR close? Closes apache#25239. Follows up on apache#21907 and [Arrow apache#6485](apache/arrow-rs#6485). ## Rationale for this change An absent Parquet `null_count` is unknown. Treating it as zero can discard NULL rows during `NULLS FIRST` TopK and `IS NULL` pruning, and let metadata-only `COUNT(column)` return an incorrect result. ## What changes are included in this PR? Preserve missing null counts in the shared row-group statistics adapter and file statistics. Remove the per-caller compatibility flag so static pruning, runtime pruning, and fully matched proofs use the same interpretation. Add an upgrade guide covering affected writers, plan changes, and migration. ## What is the testing strategy for this PR? Regressions cover missing and explicit-zero counts, static pruning, file-statistics precision, `COUNT`, and TopK with ASC/DESC, NULLS FIRST/LAST, and dynamic pruning enabled/disabled. The affected cases fail without the fix. Extended workspace tests, strict all-target/all-feature Clippy, and the full lint suite pass. ## Are there any user-facing changes? Correct results for Parquet files with missing null counts. Affected files can lose `IS NULL` pruning, metadata-only `COUNT(column)`, and `NULLS FIRST` TopK pruning. Sort pushdown can also retain a full `SortExec` for `ORDER BY` on nullable columns, even when files are sorted and non-overlapping. This includes older files written with parquet-rs before 53.1.0, used by DataFusion releases before 42.1.0. The DataFusion 56 upgrade guide explains how to identify and rewrite affected files. Explicit counts retain their existing behavior. No public API changes. --------- Signed-off-by: peterxcli <peterxcli@gmail.com>
## Which issue does this PR close? - N/A ## Rationale for this change Periodic update to latest stable Rust toolchain. ## What changes are included in this PR? * Bump pinned toolchain * Bump `depcheck` sub-workspace toolchain * Update a few references in docs * Remove `from_iter_instead_of_collect` from workspace clippy allow list (lint was removed in 1.98) * Resolve new clippy warnings that appear with clippy 1.98 ## What is the testing strategy for this PR? Compiles cleanly, all clippy / lint checks pass with no new warnings. ## Are there any user-facing changes? No.
…chain (apache#25286) ## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes apache#123` indicates that this PR will close issue apache#123. --> - Closes #. ## Rationale for this change <!-- Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed. Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes. Please explain the problem you are trying to solve in terms of the user-visible behavior, rather than the implementation. For example, "The code in `foo.rs` doesn't handle nulls" is a symptom of the implementation. "COUNT(DISTINCT) returns wrong results when the column contains nulls" is the user-visible problem. --> `datafusion::physical_plan::joins::chain` is a public module with no public member. See https://docs.rs/datafusion/latest/datafusion/physical_plan/joins/chain/index.html I think it was made public by mistake, this PR revert this module to private. ## What changes are included in this PR? <!-- There is no need to duplicate the description in the issue here, but it is sometimes worth providing a summary of the individual changes in this PR. --> ## What is the testing strategy for this PR? <!-- We typically require tests for all PRs in order to: 1. Prevent the code from being accidentally broken by subsequent changes 2. Serve as another way to document the expected behavior of the code Briefly describe how this PR is tested, and point to the specific tests you added. For example: 'This new feature is covered by the `sqllogictest` cases added in `foo.slt`'. If this PR does not add tests, explain why. For example, if the change is already covered by existing tests, please mention it. You should also check the `codecov` bot reply on this PR to confirm the changed code is exercised. --> ## Are there any user-facing changes? <!-- If there are user-facing changes then we may require documentation to be updated before approving the PR. If there are any breaking changes to public APIs, please add the `api change` label. -->
## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes apache#123` indicates that this PR will close issue apache#123. --> - Closes apache#24318 ## Rationale for this change <!-- Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed. Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes. Please explain the problem you are trying to solve in terms of the user-visible behavior, rather than the implementation. For example, "The code in `foo.rs` doesn't handle nulls" is a symptom of the implementation. "COUNT(DISTINCT) returns wrong results when the column contains nulls" is the user-visible problem. --> See issue ## What changes are included in this PR? <!-- There is no need to duplicate the description in the issue here, but it is sometimes worth providing a summary of the individual changes in this PR. --> ## What is the testing strategy for this PR? <!-- We typically require tests for all PRs in order to: 1. Prevent the code from being accidentally broken by subsequent changes 2. Serve as another way to document the expected behavior of the code Briefly describe how this PR is tested, and point to the specific tests you added. For example: 'This new feature is covered by the `sqllogictest` cases added in `foo.slt`'. If this PR does not add tests, explain why. For example, if the change is already covered by existing tests, please mention it. You should also check the `codecov` bot reply on this PR to confirm the changed code is exercised. --> Existing tests. ## Are there any user-facing changes? <!-- If there are user-facing changes then we may require documentation to be updated before approving the PR. If there are any breaking changes to public APIs, please add the `api change` label. --> No.
## Which issue does this PR close? - Part of apache#318. - Umbrella PR: apache#23738. - All required ASOF feature layers (apache#23828 through apache#23832) are merged. ## Rationale for this change This is the final core benchmark layer of the ASOF JOIN stack. It uses the standard SQL benchmark framework so the workloads share the existing runner, result format, and Criterion mode. ## What changes are included in this PR? - Add an `asof_join` SQL benchmark suite, runnable with `./bench.sh run asof_join` or `benchmark_runner`. - Add seven workloads covering relative input sizes, available ordering, equality-group cardinality and skew, match direction, and payload width. - Describe each workload and its explicit ASOF match semantics in the query. - Generate pre-sorted Parquet inputs under `DATA_DIR/asof_join` with `./bench.sh data asof_join`. The benchmark load hook only registers them with `WITH ORDER`, keeping data generation and input sorting outside the measured query. - Assert that every workload plans to `AsOfJoinExec`. ## Are these changes tested? Yes: - `cargo fmt --all` - `./ci/scripts/doc_prettier_check.sh --write --allow-dirty` - `cargo clippy --all-targets --all-features -- -D warnings` - `cargo test -p datafusion-benchmarks --lib --bins` - `DATA_DIR=/tmp/asof_join CARGO_COMMAND='cargo run' ./bench.sh data asof_join` - `DATA_DIR=/tmp/asof_join CARGO_COMMAND='cargo run' ./bench.sh run asof_join 7` - `cargo run -p datafusion-benchmarks --bin benchmark_runner -- asof_join --iterations 1 --partitions 4 --path /tmp/asof_join` - `cargo run -p datafusion-benchmarks --bin benchmark_runner -- asof_join --query 7 --partitions 4 --path /tmp/asof_join --criterion` The pre-sorted workload's physical plan contains `AsOfJoinExec` without an input `SortExec`. These smoke runs validate the suite and are not presented as performance claims. ## Are there any user-facing changes? This adds ASOF workloads to the repository benchmark tooling. It does not change query semantics or runtime behavior. This is the final core item tracked by apache#23738. It does not depend on the optional floating-point follow-up apache#24375.
…pache#25193) - ATM, the Parquet file schema type coercion applies view transforms for example only on top-level fields, but not on nested fields. - For some Parquet schemas, where we request a view transform for a nested field, this is translated to an unmodified schema passed to the reader. Then, when we try to apply predicates on the actual data on that nested field, the actual predicate execution needs a cast, because the schemas are different - the predicate refers to a string view, but the actual data schema refers to a string ## Which issue does this PR close? - Closes apache#25192 ## Rationale for this change - Rationale is explained in apache#25192 - Also, I don't see any reason (explicit comments) why this function would not work on deeply nested schemas ## What changes are included in this PR? - Modification of the `apply_file_schema_type_coercions` function to work on nested fields - Respective tests for the functionality ## What is the testing strategy for this PR? - All tests pass - Newly added unit tests pass ## Are there any user-facing changes? - Unsure ? Probably not ? The documentation in the function does NOT explicitly say the current behavior.
## Which issue does this PR close? - Closes apache#23022. ## Rationale for this change Predicate subqueries such as `EXISTS`, `IN`, and `NOT IN` are valid scalar expressions in a `SELECT` list, but DataFusion currently only decorrelates them when they appear in filters. Projection queries therefore reach physical planning with an unsupported logical subquery expression. This continues the work from the now-closed stale apache#23039 and covers multiple and correlated projection subqueries as well as SQL null semantics. ## What changes are included in this PR? - decorrelate `EXISTS`, `NOT EXISTS`, `IN`, and `NOT IN` expressions inside projections into mark joins - materialize projected `IN`/`NOT IN` values with SQL three-valued logic by tracking matches, nulls, and empty subqueries - preserve projection output names and avoid duplicate columns in correlated subquery projections ## What is the testing strategy for this PR? Optimizer plan tests cover projected `IN` rewriting. New SQLLogicTest cases cover uncorrelated and correlated `EXISTS`/`IN`, multiple subqueries in one projection, empty inputs, null inputs, and `NOT IN` three-valued logic. Locally validated with: - `cargo test -p datafusion-optimizer decorrelate_predicate_subquery` (53 passed) - `cargo test -p datafusion-sqllogictest --test sqllogictests -- subquery_projection` - `cargo clippy -p datafusion-optimizer --all-targets --all-features -- -D warnings` - `cargo fmt --all -- --check` ## Are there any user-facing changes? Yes. Predicate subqueries can now be evaluated as projected scalar expressions. There are no public API changes.
## Which issue does this PR close? Fixes the [security audit failure](https://github.com/apache/datafusion/actions/runs/34914826729/job/104210077720) caused by [RUSTSEC-2026-0285](https://rustsec.org/advisories/RUSTSEC-2026-0285.html). ## Rationale for this change The locked rustls 0.23.39 accepts TLS 1.3 handshake messages across encryption level boundaries and now fails the security audit. Rustls 0.23.45 fixes the vulnerability. ## What changes are included in this PR? Update rustls to 0.23.45 in Cargo.lock, together with its required aws-lc-rs, aws-lc-sys, and rustls-webpki dependency updates. ## What is the testing strategy for this PR? The existing security audit reproduces the vulnerability with the original lockfile and passes with the updated lockfile: ```sh cargo audit --ignore RUSTSEC-2026-0194 --ignore RUSTSEC-2026-0195 ``` The workspace extended tests also pass: ```sh RUST_BACKTRACE=1 cargo test --profile ci \ --exclude datafusion-examples --exclude datafusion-benchmarks --exclude datafusion-cli \ --workspace --lib --tests --bins \ --features avro,json,backtrace,extended_tests,recursive_protection,parquet_encryption ``` Also passed `cargo fmt --all`, `cargo clippy --all-targets --all-features -- -D warnings`, and `./dev/rust_lint.sh`. ## Are there any user-facing changes? Builds using the checked-in lockfile use the patched TLS implementation. No DataFusion API changes.
…ng (apache#25261) ## Which issue does this PR close? - Closes apache#25260. ## Rationale for this change A malformed `null_regex` panics the query task instead of returning an error. `CsvFormat::infer_schema_from_stream` compiled the pattern with ```rust let regex = Regex::new(null_regex.as_str()) .expect("Unable to parse CSV null regex."); ``` so any pattern the `regex` crate rejects aborted the task. It is reachable straight from SQL, through a `CREATE EXTERNAL TABLE` that leaves its columns to schema inference: ```sql CREATE EXTERNAL TABLE t STORED AS CSV LOCATION 'data.csv' OPTIONS ('format.has_header' 'true', 'format.null_regex' '('); ``` ``` task 9 panicked with message "Unable to parse CSV null regex.: Syntax( regex parse error: ( ^ error: unclosed group )" ``` An invalid regex is a bad option value, not an internal invariant, so it should come back as an error naming the pattern. ## What changes are included in this PR? The regex is compiled once, before the per-chunk loop, and a failure is propagated with `exec_datafusion_err!` rather than panicking. Hoisting it also stops the pattern being recompiled for every chunk of the inference stream. The error text matches the one used on the read side in apache#25254, so the same bad option reads the same whichever path hits it first. ## What is the testing strategy for this PR? A `statement error Unable to parse CSV null regex` case in `datafusion/sqllogictest/test_files/csv_files.slt`. I checked it is not vacuous: with this change reverted, that case fails with the panic quoted above rather than an error, so the test reproduces the bug. `cargo fmt --check`, `cargo clippy --all-targets` and the crate's unit tests are clean. ## Are there any user-facing changes? An invalid `null_regex` now produces an error and leaves the session usable, where it previously panicked the task. No API changes. Independent of apache#25254 — that one is the read path in `source.rs`, this is the inference path in `file_format.rs` — but they touch the same crate, so whichever lands second may want a trivial rebase. Co-authored-by: Prajwal Narayana <sprajwalln@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
## Which issue does this PR close? None; found by sweeping the scalar functions with edge-case arguments. ## Rationale for this change `array_resize` with a negative size raises an internal error, which tells the caller they hit a DataFusion bug and asks them to open a report: ```console > SELECT array_resize(make_array(1,2,3), -1, 0); Internal error: array_resize: failed to convert size to usize. This issue was likely caused by a bug in DataFusion's code. Please help us to resolve this by filing a bug report in our issue tracker: https://github.com/apache/datafusion/issues ``` The size is an argument to the query, so a negative value is the caller's input rather than a broken internal invariant. Sending someone to the issue tracker for their own argument is misleading, and it hides what was actually wrong. The behaviour was already meant to be an error. `array_resize.slt` has expected this call to fail since before this change, but with a bare `query error`, so the internal error satisfied it and went unnoticed. ## What changes are included in this PR? Reject a negative size with an execution error naming the value, which matches how the neighbouring maximum-size check already reports bad input: ```console > SELECT array_resize(make_array(1,2,3), -1, 0); Execution error: array_resize: size must not be negative, got -1 ``` Valid input is untouched: `0` still returns `[]`, `5` still grows to `[1, 2, 3, 9, 9]`, and a NULL size still returns NULL. ## Are these changes tested? Yes. The slt file gains the negative case for both `List` and `LargeList`, plus `i64::MIN` for the boundary, each pinned to the new message so a regression back to an internal error fails the suite. ```console $ cargo test -p datafusion-sqllogictest --test sqllogictests -- array Progress: 53/53 files completed (100%) $ cargo test -p datafusion-functions-nested --lib resize test result: ok. 4 passed; 0 failed ``` The pre-existing `query error` case for `-5` still passes, so the contract that a negative size fails is unchanged. ## Are there any user-facing changes? The error for a negative size changes from an internal error to an execution error with a clear message. No valid call changes behaviour, and no API changes. --------- Signed-off-by: 1fanwang <1fannnw@gmail.com> Co-authored-by: Neil Conway <neil.conway@gmail.com> Co-authored-by: Jeffrey Vo <jeffrey.vo.australia@gmail.com>
…4979) ## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes apache#123` indicates that this PR will close issue apache#123. --> - Closes #. ## Rationale for this change <!-- Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed. Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes. Please explain the problem you are trying to solve in terms of the user-visible behavior, rather than the implementation. For example, "The code in `foo.rs` doesn't handle nulls" is a symptom of the implementation. "COUNT(DISTINCT) returns wrong results when the column contains nulls" is the user-visible problem. --> `PartialSortStream` concatenates each incoming batch with the buffered incomplete prefix before looking for a completed prefix boundary. When one prefix spans multiple batches, this repeatedly copies rows accumulated from earlier batches. When a batch contains multiple completed prefix groups, the completed region is also sorted using the full ordering even though the prefix ordering is already satisfied. This adds unnecessary copying and sorting work, especially for wide batches and workloads containing multiple prefix groups per batch. ## What changes are included in this PR? <!-- There is no need to duplicate the description in the issue here, but it is sometimes worth providing a summary of the individual changes in this PR. --> - Keep the unfinished trailing prefix as a list of `RecordBatch` slices instead of concatenating it with every incoming batch. - Detect completed prefix groups within each batch and compare prefix expression results directly across batch boundaries. - Sort completed groups only by the remaining suffix ordering. - Concatenate a single completed prefix at most once. - For multiple completed prefixes, evaluate suffix expressions once per source batch and materialize the final output with one interleave operation. - Preserve fetch handling, input release, empty-batch handling, and zero-column batch semantics. ## What is the testing strategy for this PR? <!-- We typically require tests for all PRs in order to: 1. Prevent the code from being accidentally broken by subsequent changes 2. Serve as another way to document the expected behavior of the code Briefly describe how this PR is tested, and point to the specific tests you added. For example: 'This new feature is covered by the `sqllogictest` cases added in `foo.slt`'. If this PR does not add tests, explain why. For example, if the change is already covered by existing tests, please mention it. You should also check the `codecov` bot reply on this PR to confirm the changed code is exercised. --> The implementation is covered by the existing `PartialSortExec` unit tests and the additional SQL logic test cases in `group_by.slt`. ## Are there any user-facing changes? <!-- If there are user-facing changes then we may require documentation to be updated before approving the PR. If there are any breaking changes to public APIs, please add the `api change` label. --> No. This is an internal performance improvement. ## Benchmark ``` group main optimized ----- ---- --------- partial_sort/rows_per_prefix/100 3.10 12.1±0.13ms ? ?/sec 1.00 3.9±0.05ms ? ?/sec partial_sort/rows_per_prefix/1000 4.17 13.2±0.07ms ? ?/sec 1.00 3.2±0.14ms ? ?/sec partial_sort/rows_per_prefix/10000 1.30 2.8±0.02ms ? ?/sec 1.00 2.1±0.07ms ? ?/sec partial_sort/rows_per_prefix/20000 1.41 3.0±0.22ms ? ?/sec 1.00 2.1±0.15ms ? ?/sec partial_sort/rows_per_prefix/5000 3.85 12.4±0.27ms ? ?/sec 1.00 3.2±0.03ms ? ?/sec partial_sort/rows_per_prefix/8192 1.56 2.8±0.05ms ? ?/sec 1.00 1780.8±34.70µs ? ?/sec ```
…() (apache#25306) ## Which issue does this PR close? - Closes apache#25305. ## Rationale for this change `named_struct(...)` and `struct(...)` never produce a NULL row: `invoke_with_args` builds the output `StructArray` with no null buffer. Both functions nevertheless reported a nullable return field, and `ExprSchemable::nullable` for a scalar function reads that field. As a result the simplifier rule that rewrites `a IS NOT NULL` to `true` for a non-nullable `a` never fired on a struct constructor. The user-visible cost is lost pruning. A guard such as `WHERE s IS NOT NULL` on a struct built by a view cannot remove any row, but it keeps the whole struct expression alive, so projection pushdown can no longer prune the scan to the fields that are read: ```sql SET datafusion.explain.format = 'indent'; CREATE TABLE t (a INT, b INT, c INT) AS VALUES (1, 2, 3), (NULL, 5, 6); CREATE VIEW v AS SELECT named_struct('a', a, 'b', b, 'c', c) AS s FROM t; EXPLAIN SELECT s['b'] FROM v WHERE s IS NOT NULL; ``` Before: ``` logical_plan Projection: __datafusion_extracted_1 AS v.s[b] SubqueryAlias: v Projection: __datafusion_extracted_1 Filter: named_struct(Utf8("a"), t.a, Utf8("b"), t.b, Utf8("c"), t.c) IS NOT NULL Projection: t.b AS __datafusion_extracted_1, t.a, t.b, t.c TableScan: t projection=[a, b, c] ``` After: ``` logical_plan Projection: __datafusion_extracted_1 AS v.s[b] SubqueryAlias: v Projection: t.b AS __datafusion_extracted_1 TableScan: t projection=[b] ``` This matches the plan the same query already gets without the guard, and lets the existing `get_field(named_struct(...), 'f')` simplification (apache#22239) remove the constructor. ## What changes are included in this PR? - `datafusion/functions/src/core/named_struct.rs`: `return_field_from_args` now reports a non-nullable field instead of a nullable one. - `datafusion/functions/src/core/struct.rs`: adds a `return_field_from_args` override that reports a non-nullable field. The default implementation it previously used made the field nullable. Both are a one-line nullability correction; the inner struct fields stay nullable, since an individual member value can still be NULL. ## What is the testing strategy for this PR? - New `sqllogictest` cases at the end of `datafusion/sqllogictest/test_files/struct.slt` covering the folded scalar results (`named_struct(...) IS NOT NULL`, `struct(...) IS NULL`), the `EXPLAIN` for `WHERE s IS NOT NULL` over a view showing `TableScan: t projection=[b]`, the values that query returns, and the `CASE WHEN named_struct(...) IS NOT NULL` fold. - Existing `datafusion-functions` unit tests: `cargo test -p datafusion-functions --lib`, 356 passed. - The full `sqllogictest` suite: all 518 files pass, with no other expectation changes needed. ## Are there any user-facing changes? Yes, but no API change. - The schema nullability of the output of `named_struct()` and `struct()` changes from nullable to non-nullable. This now matches the array those functions actually build. - `IS NULL` and `IS NOT NULL` applied directly to one of these constructors constant-fold to `false` and `true`. Query results are unchanged; plans get smaller and can prune more.
## Which issue does this PR close? - closes apache#11900 ## Rationale for this change When a query has `ORDER BY <cols> LIMIT N` on top of an outer join and all sort columns come from the preserved side, DataFusion currently runs the full join first, then sorts and limits. We can push a copy of the `Sort(fetch=N)` to the preserved input, reducing the number of rows entering the join. **Before:** ``` Sort: t1.b ASC, fetch=3 Left Join: t1.a = t2.a Scan: t1 ← scans ALL rows Scan: t2 ``` **After:** ``` Sort: t1.b ASC, fetch=3 Left Join: t1.a = t2.a Sort: t1.b ASC, fetch=3 ← pushed down Scan: t1 ← only top-3 rows enter join Scan: t2 ``` ## What changes are included in this PR? A new logical optimizer rule `PushDownTopKThroughJoin` that: 1. Matches `Sort` with `fetch = Some(N)` (TopK) 2. Looks through an optional `Projection` to find a `Join` 3. Checks join type is LEFT or RIGHT with no non-equijoin filter 4. Verifies all sort expression columns come from the preserved side 5. Inserts a copy of the `Sort(fetch=N)` on the preserved child 6. Keeps the top-level sort for correctness ## Are these changes tested? Yes through UT ## Are there any user-facing changes? No API changes. --------- Co-authored-by: Subham Singhal <subhamsinghal@Subhams-MacBook-Air.local> Co-authored-by: Dmitrii Blaginin <dmitrii@blaginin.me> Co-authored-by: Kumar Ujjawal <ujjawalpathak6@gmail.com>
…pache#25259) ## Which issue does this PR close? - Closes apache#25244. - Closes apache#25262. - Closes apache#25263. - Closes apache#25264. - Closes apache#21246. - Closes apache#25296. ## Rationale for this change Enabling `datafusion.optimizer.enable_join_dynamic_filter_pushdown` can silently drop matching rows when the probe side of a join contains several columns with the same name, for example `a.id` and `b.id` from a nested join. Physical filter pushdown resolved a pushed filter's columns in the child schema by name, so a predicate on the second `id` column was rewritten to the first `id` column, which holds a different value. The same name-based mapping also affects TopK dynamic filters under the default configuration (apache#25296). For a nested join of orders and payments that both expose `amount`, `ORDER BY p.amount LIMIT 1` can push the payment threshold onto `orders.amount`. Later matching orders are incorrectly pruned, returning `(100, 30)` instead of `(300, 10)`. Every operator that forwards parent filters had the same weakness, in slightly different forms: - `HashJoinExec` computed which output columns belong to each side by position, but then resolved the child column by name. - `FilterExec`, `SortExec`, `RepartitionExec`, `CoalesceBatchesExec`, `UnionExec` and similar nodes used the generic name lookup even though their output positions equal their input positions. - `FilterExec` with an embedded projection resolved unconsumed parent filters back into input coordinates by name. - `ProjectionExec` looked up output aliases by name to find the expression to substitute, so two outputs aliased `id` both mapped to the first. - `AggregateExec` restricted pushdown to grouping positions but resolved the input column by name. The join-filter regression queries return one row with dynamic filtering disabled and zero rows with it enabled across these shapes. The `FilterExec`, `ProjectionExec` and `AggregateExec` cases were found while auditing the remaining name-based paths for apache#25244; they share the root cause, so this PR fixes them together. The TopK case in apache#25296 is fixed by the same positional mapping. ## What changes are included in this PR? Built-in operators now remap filter columns by position. The deprecated public API retains its historical name-based resolution for compatibility. - `FilterRemapper` uses a `ColumnMapping` enum: `Identity` preserves positions and checks that names match; `Explicit` uses a caller-supplied parent-output to child-input mapping and allows names to differ. - `ChildFilterDescription::from_child` now maps by position. New `ChildFilterDescription::from_child_with_column_mapping` takes an explicit `HashMap<usize, usize>`. - `HashJoinExec` builds the explicit mapping from its `column_indices` and output projection. For semi joins, output join keys on the emitted side are mapped to the paired key on the other side, which also supports differently named keys. - `FilterExec` maps through its embedded projection in both pushdown phases and when folding unconsumed parent filters back into its predicate. - `ProjectionExec` substitutes the expression at each output position instead of looking the alias up in the output schema. - `AggregateExec` maps each grouping output position to the input column that grouping expression reads; non-column grouping expressions are not forwarded, as before. - `ChildFilterDescription::from_child_with_allowed_indices` is kept as a deprecated wrapper that preserves name-based resolution to the first matching child field. It translates allowed parent column references into an explicit positional mapping; new callers should supply positions directly to avoid ambiguous duplicate names. ## What is the testing strategy for this PR? - `dynamic_filter_pushdown_config.slt` gains four regression queries over two small Parquet tables, run once with join dynamic filtering disabled and once enabled, asserting identical rows: nested joins with `RepartitionExec` between them (the query from apache#25244), a `FilterExec` with an embedded projection, a `ProjectionExec` with duplicate aliases, and an `AggregateExec` grouping on same-named columns. All four return zero rows on `main` with dynamic filtering enabled. - `dynamic_filter_pushdown_config.slt` also covers apache#25296 using nested joins of orders, payments and customers, with one row per batch and one row per orders Parquet row group. The query returns `(300, 10)` with TopK dynamic filtering disabled, enabled, and enabled together with Parquet filter pushdown. On `main` at `85d4cbb0a9`, both enabled cases reproduce the incorrect `(100, 30)` result; all three pass on this branch. - `datafusion/core/tests/physical_optimizer/filter_pushdown.rs` gains focused tests that call `gather_filters_for_pushdown` directly on `RepartitionExec`, `HashJoinExec` (duplicate child columns with and without projection, both semi join directions with differently named keys, and a semi join whose key is not a plain column), `FilterExec` with a projection in both phases, `ProjectionExec` with duplicate aliases, and `AggregateExec` with reordered same-named grouping columns. Further tests cover the deprecated `from_child_with_allowed_indices` wrapper preserving the first name match, accepting allowed parent indices outside the child schema, and rejecting unresolvable names, the identity mapping rejecting a column whose name differs from the child field at that position, and `FilterExec` rejecting an unconsumed parent filter outside its projection. Passed locally on the updated implementation: - `cargo fmt --all` - `./ci/scripts/doc_prettier_check.sh --write --allow-dirty` - `cargo test --profile ci -p datafusion --test core_integration physical_optimizer::filter_pushdown` — 70 tests passed. - `cargo test --profile ci --test sqllogictests -- dynamic_filter_pushdown_config.slt` — passed. - The isolated apache#25296 regression fails on `main` in both TopK-enabled configurations with the expected wrong-result mismatch. ## Are there any user-facing changes? Queries that push join or TopK dynamic filters through operators with duplicate column names now return the correct rows, including the default-configuration TopK query in apache#25296. API change in `datafusion-physical-plan`: `ChildFilterDescription::from_child_with_allowed_indices` is deprecated in favour of `from_child` and the new `from_child_with_column_mapping`; the deprecated function preserves its previous name-based behavior. Callers should migrate to explicit positions because duplicate names make name resolution ambiguous. `ChildFilterDescription::from_child` also resolves by position, which only affects callers that used it on a node whose output positions differ from its child's.
…#25198) ## Which issue does this PR close? [Related to apache#24822. No issue closed.](apache#25185) ## Rationale for this change The benchmarks in this file build a fresh values array per batch and cap cardinality at the batch size. Neither shape matches a real plan: `take` and `filter` clone the values `Arc` and rewrite only the keys, so batches downstream of a repartition share one values array, and because it is never compacted its cardinality can far exceed the rows in a batch. Per-batch work that scales with dictionary cardinality is invisible under the current parameters. ## What changes are included in this PR? Adds `bench_shared_values_arc`: 32 batches of 8192 rows sharing one values array, at cardinalities 8192, 100k and 500k. ## What is the testing strategy for this PR? Benchmark-only. Ran the bench locally and `cargo clippy --all-targets -D warnings`. ## Are there any user-facing changes? No
…in keys (inner and semi joins) (apache#25255) ## Which issue does this PR close? No dedicated issue yet. Related: apache#18290, apache#19858, apache#7955 (dynamic filter pushdown follow-ups). ## Rationale for this change ### What "transferring a filter across join keys" means For an inner or semi join `ON a.k = b.k`, every output row has `a.k = b.k`. So a filter that only touches `a.k` (say `a.k IN (1, 3)`) is also true of `b.k` for every row that can match. Today `HashJoinExec` routes a parent filter only to the side that owns its columns, so that filter reaches `a` and never `b`. With this PR the join also pushes `b.k IN (1, 3)` to `b`. Static predicates already get this at the logical level (equality inference in `push_down_filter`). Dynamic filters do not exist there, and they are the ones that matter: the dynamic filter of a join above is a parent filter for the joins below it. Before, with `dim` (small, filtered) joined to `mid`, and `mid` joined to `fact`: ```text HashJoinExec on=[(d_key, m_key)] build: dim, probe: (mid ⋈ fact) ├── DataSourceExec dim └── HashJoinExec on=[(m_key, f_key)] build: mid, probe: fact ├── DataSourceExec mid predicate=DynamicFilter[d_key from dim] <- direct: m_key is a mid column └── DataSourceExec fact predicate=DynamicFilter[m_key from mid] <- only the lower join's own filter ``` After: ```text HashJoinExec on=[(d_key, m_key)] ├── DataSourceExec dim └── HashJoinExec on=[(m_key, f_key)] ├── DataSourceExec mid predicate=DynamicFilter[d_key from dim] └── DataSourceExec fact predicate=DynamicFilter[m_key from mid] AND DynamicFilter[d_key from dim, rewritten over f_key] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ transferred across m_key = f_key ``` The transferred copy is the same `DynamicFilterPhysicalExpr` with its key columns remapped, so it is populated when `dim` finishes building and prunes `fact` at the scan (file / row-group / page pruning always, row-level with `pushdown_filters=true`). Because `fact` is now pruned before the lower join builds, the lower join's own filter tightens too. Why the lower join's own filter is not enough on its own: - it is a hash-table lookup once the lower build side exceeds `hash_join_inlist_pushdown_max_size` / `_max_distinct_values`, which cannot prune files, row groups or pages, while the transferred filter carries the small table's min/max bounds and IN list; - it is absent when the lower join creates none (Partitioned mode without routing, `preserve_file_partitions`, null-aware joins with nullable build keys, a build-side scan that does not accept filters); - it can only be as tight as its own build input, whereas the transferred filter prunes that input first (see TPC-H Q5 below, where `customer` had no dynamic filter at all before). ### Semi-join routing bug fixed on the way The same mechanism replaces the semi-join special case in `gather_filters_for_pushdown`. That code added the key's *output index* to the non-output side's allowed set and then let `FilterRemapper::try_remap` map the column to that side by *name*. Two failure modes followed when the keys were named differently: - the name does not exist on the other side: the remap fails and nothing is pushed there (a missed optimisation); - the name exists on the other side but is not the key: the filter is pushed onto that unrelated column, and because a child accepted it the `FilterExec` above the join is removed. That is a wrong-result bug. Example: `LeftSemi` on `left.k = right.j` where `right` also has a non-key column `k`. A parent filter `k = 'x'` reached the right scan as `k@2 = 'x'`; with this PR it is `j@0 = 'x'`. The `RightSemi` branch had the mirror-image bug. The existing `test_hashjoin_parent_filter_pushdown_semi_anti_join` did not catch either, because both keys in it are named `k`. The transfer rewrites over the key *expression* on the other side, so column names no longer matter. ### Benchmarks **TL;DR:** JOB total 3 % faster with row-level parquet pushdown (16a 2.3x, 16c/16d 1.6x, 33a/33c 1.25x), TPC-H SF10 Q5 1.3x, everything else within noise, no reproducible regression. Details below. M4 Pro (12 cores / 24 GB), release binaries built in separate target dirs, sides alternated per round, per-query minimum over all iterations. `default` = stock parquet config (dynamic filters prune files / row groups / pages only); `pushdown` = `datafusion.execution.parquet.pushdown_filters=true`. | suite | mode | geomean ratio | total time | |---|---|---|---| | JOB (113 q), 2 x 5 iter | default | 0.997 | 51.07 s -> 50.95 s | | JOB (113 q), 2 x 5 iter | pushdown | 0.982 | 26.48 s -> 25.66 s | | TPC-H SF10 (22 q), 2 x 3 iter | default | 1.004 | 5.19 s -> 5.24 s (no query outside 5 %) | | TPC-H SF10 (22 q), 2 x 3 iter | pushdown | 0.996 | 6.49 s -> 6.42 s | Queries outside +/-5 % (pushdown mode, ms): | query | before | after | ratio | |---|---|---|---| | JOB 16a | 723 | 321 | 0.44 | | JOB 16d | 722 | 439 | 0.61 | | JOB 16c | 719 | 464 | 0.64 | | JOB 33a | 134 | 105 | 0.78 | | JOB 33c | 143 | 116 | 0.81 | | TPC-H Q5 | 416 | 324 | 0.78 | | JOB 18a / 17e / 17f | 208-329 | 221-349 | 1.06-1.07 (within run-to-run spread) | | JOB 16b | 874 | 990 | 1.13 (not reproducible: identical scan predicates on both binaries; interleaved re-timing gives before 922-1642 ms vs after 874-1144 ms) | Where the time goes, from per-scan `EXPLAIN ANALYZE` metrics in pushdown mode. JOB 16a: the top join builds on the filtered `title` (68 K rows) and its dynamic filter lands on `ci.movie_id`, the build-side key of the join below. Transferred across that join's keys it now also reaches two scans: | scan | dynamic filters before -> after | output rows before -> after | |---|---|---| | movie_companies | 1 -> 2 | 1.15 M -> 7.6 K | | cast_info | 2 -> 3 | 6.38 M -> 45 K | TPC-H SF10 Q5: the 5-row `nation` filter reaches `customer` through the key equivalence, which previously received no dynamic filter, and the pruning cascades down the join chain: | scan | dynamic filters before -> after | output rows before -> after | |---|---|---| | customer | 0 -> 1 | 1.50 M -> 300 K | | orders | 1 -> 1 | 2.28 M -> 457 K | | lineitem | 1 -> 1 | 9.10 M -> 1.83 M | Cost side: the transferred copy is rewritten over the target side's key expression, so where the join key is a `CAST` it pays a cast per row. The bounds builder also emits duplicated bound pairs when several key pairs share one column (pre-existing, visible in the before plans too, just more often now). ## What changes are included in this PR? - `HashJoinExec::gather_filters_for_pushdown`: after the plain column-based routing, every parent filter whose columns are all plain `Column` join keys of one side is rewritten over the other side's key expressions and marked supported for that child. Inner, LeftSemi and RightSemi only: outer, anti and mark joins also emit unmatched rows, so the transferred filter would not be exact there and `if_any` could wrongly drop the parent filter. Outer joins would need "prune-only" semantics and are left as a follow-up. - A `DynamicFilterPhysicalExpr` is rewritten through `with_new_children`, so the transferred copy shares the original's state and keeps tracking the build side. - Removed the name-based semi-join routing, now covered by the general transfer. This fixes the wrong-column pushdown described above. - A transfer only targets a side that `lr_is_preserved` permits, the same gate as the plain column routing (a no-op for the three permitted join types today, but it couples the two functions). Join keys are `debug_assert`ed to have equal data types, since the planner coerces them and `try_new` does not check. ## What is the testing strategy for this PR? - `test_hashjoin_parent_filter_transferred_across_join_keys`: key names differ; key filters land on both scans, a non-key filter stays on its side. - `test_hashjoin_parent_filter_transfer_semi_join_different_key_names` and its `RightSemi` variant: the case the old name-based routing missed (nothing reached the non-output scan). - `test_hashjoin_parent_filter_transfer_semi_join_key_name_shadowed_by_non_key`: regression test for the wrong-column pushdown, `LeftSemi` and `RightSemi`. - `test_hashjoin_parent_filter_transfer_uses_first_on_pair`: `ON k = x AND k = y` transfers over the first pair. - `test_hashjoin_parent_filter_transfer_cast_key_with_projection`: a `CAST` key with a projection that puts the other side's key at the same output index as the cast's inner column. The rewrite must not descend into the substituted expression, or it would loop forever. - `test_hashjoin_dynamic_filter_transferred_through_nested_join`: an upper join's dynamic filter reaches the lower probe scan. The lower build scan rejects filters, so the lower join's own filter still lists all four keys and the two pruned rows are attributable to the transfer alone (checked via scan metrics). - `join_dynamic_filter_transfer.slt`: SQL plan shape and results. - Two existing snapshots changed only by a build-key filter now also appearing on the probe scan. Full sqllogictest suite, physical-plan join unit tests, `cargo fmt` and `clippy -D warnings` pass. ## Are there any user-facing changes? No new configuration. `EXPLAIN` may now show a parent or dynamic filter on both scans of an inner or semi join where it previously appeared on one.
…_fanout (apache#25272) ## Which issue does this PR close? - Closes apache#25077 . ## Rationale for this change `EXPLAIN ANALYZE` reports `probe_hit_rate` and `avg_fanout` of `HashJoinExec` too low whenever a probe batch produces more than `batch_size` output rows, e.g. when the build side has many duplicate keys. Both metrics are defined per probe row, but their counters were updated once per output chunk. ## What changes are included in this PR? - `probe_hit_rate` total: added only on the first chunk of each probe batch (`state.offset == (0, None)`), the same first-chunk check that `process_probe_batch` already uses for correlated null-aware `LeftMark`. - `probe_hit_rate` part and `avg_fanout` total: a new `ProcessProbeBatchState::matched_probe_idx` records the last probe index counted so far. `count_new_matched_probe_rows` does not count a chunk's first index again if it equals that value. Empty chunks leave it unchanged. - Existing state is not reused: `joined_probe_idx` is tracked after the join filter, while these metrics are defined before it, and the lookup `offset` can point at a probe row that has not been counted yet. ## What is the testing strategy for this PR? Two new tests in `hash_join/exec.rs`, each run over `hash_join_exec_configs` (batch sizes 8192/10/5/2/1, perfect hash join on and off): - `join_probe_metrics_count_each_probe_row_once`: duplicate build keys, so one probe row's matches are split across chunks. Fails without this change for batch sizes 5, 2 and 1. - `join_probe_metrics_count_probe_row_starting_new_chunk`: unique build keys, so chunks split between probe rows. It guards against deduplicating by the lookup offset: comparing against `offset.0` (on non-first chunks) instead makes this test fail for batch sizes 2 and 1, while the first test still passes. ## Are there any user-facing changes? `probe_hit_rate` and `avg_fanout` shown by `EXPLAIN ANALYZE` are now correct when a probe batch is processed in several chunks. No public API changes.
…he#25185) ## Which issue does this PR close? - Closes #. ## Rationale for this change `vectorized_append` is the path taken when we have non streaming hash aggregations (when there is no ordering in the `group by`), right now it re hashed the entire dictionary values on every batch unlike append_val which that only does this if the values array was already seen before. ## What changes are included in this PR? `vectorized_append` now reuses the cached value hashes when a batch carries the same dictionary values array as the previous one (this is the common case of a repartition or filter), instead of re hashing the whole dictionary and re-initialising the map on every batch. This extends the existing `Arc::ptr_eq` cache from apache#24418, which only covered the scalar `append_val` path. ## What is the testing strategy for this PR? All dict tests pass, I have a PR for a benchmark that exercises this path -> apache#25198 ## Are there any user-facing changes? no, this is a pure perf PR
## Which issue does this PR close? Part of apache#25148. As suggested by @comphead apache#25231 (review) ## Rationale for this change The DataFusion and Substrait feature checks wait for the workspace check, but start on separate runners and restore a cache intended for test builds. Sharing the workspace check's outputs can avoid repeating compatible compilation work within the same workflow run, including on a cold PR run. ## What changes are included in this PR? - Upload the workspace build outputs and Cargo registry once from `linux-build-lib`, then restore them in the DataFusion and Substrait feature-check jobs. - Include tracked source files in the archive to preserve timestamps, and use tar archives to preserve Cargo's hidden fingerprint files and executable permissions. - Exclude generated incremental compiler caches from the Cargo registry archive. - Keep the producer's dependency cache, retain artifacts for one day, support reruns, and run the checks normally if artifacts cannot be uploaded or downloaded. - Preserve every existing check command, runner, trigger, job name, and job dependency. ## What is the testing strategy for this PR? - The updated workspace producer and both feature-check jobs pass in [PR CI](https://github.com/apache/datafusion/actions/runs/34743091547). The affected jobs also pass actionlint. - [Repacking the original artifact](https://github.com/kumarUjjawal/datafusion/actions/runs/34742695510) with the same compression reduced it from 2.34 GB to 577 MB (75%), while preserving the workspace archive. Excluded paths were checked against the dependency source manifests. - [Three controlled comparisons per consumer](https://github.com/kumarUjjawal/datafusion/actions/runs/34743010860) use the same source, toolchain, container, and physical runner for each baseline/artifact pair, clear build data between variants, and reverse execution order in one sample. All twelve consumer executions pass. - The two consumers together save 59–81 seconds, including restoration. Packaging/upload measured separately takes 19–23 seconds; using its 19-second median gives an estimated net saving of 40–62 runner-seconds per run. These fork measurements do not establish AWS cost savings or faster overall PR feedback. ## Are there any user-facing changes? No.
## Which issue does this PR close? Closes apache#25047. ## Rationale for this change `ordered_aggregate_spill.slt` can fail when spill-merge buffers leave too little memory for aggregate replay while another partition retains aggregate state in the shared greedy pool. Limit merge fan-in to two to reduce this contention while preserving the 600 KiB memory limit, two partitions, and existing query results. ## What changes are included in this PR? - Set and reset the SQL test's spill merge fan-in. - Add controlled two-partition tests for grouped results, spilling, and reservation cleanup on completion, input error, and cancellation. - Check spilled rows without depending on the formatted byte unit, and remove exact spill-count claims from comments. ## What is the testing strategy for this PR? Local validation: - 320 SQL test executions (16 concurrent copies per run) and 50 repetitions of the focused tests. - Full extended workspace tests, including all 518 SQL test files. - Workspace Clippy with all targets and features, formatting, and submission checks. ## Are there any user-facing changes? No. This changes tests only; production execution and defaults are unchanged.
Bumps [rstest](https://github.com/la10736/rstest) from 0.26.1 to 0.27.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/la10736/rstest/releases">rstest's releases</a>.</em></p> <blockquote> <h2>v0.27.0</h2> <h2>[0.27.0] 2026/9/6</h2> <h3>Changed</h3> <ul> <li>Bump msrv to 1.85.0 both for <code>rstest</code> and <code>rstest_reuse</code></li> <li>Disabled default features of <code>futures-util</code></li> </ul> <h3>Added</h3> <ul> <li>Doc comments before <code>#[values(...)]</code> entries can be used to override the generated matrix test names (both for the legacy <code>arg => [..]</code> syntax and the new attribute form). See <a href="https://redirect.github.com/la10736/rstest/pull/321">#321</a> thanks to <a href="https://github.com/orhun"><code>@orhun</code></a>.</li> </ul> <h3>Fixed</h3> <ul> <li>Use fully-qualified <code>core</code> import. See <a href="https://redirect.github.com/la10736/rstest/pull/336">#336</a>.</li> <li>Fix <code>mut</code> arguments failing to compile with <code>#[trace]</code>. See <a href="https://redirect.github.com/la10736/rstest/pull/345">#345</a> thanks to <a href="https://github.com/super-cooper"><code>@super-cooper</code></a>.</li> <li>Fix compilation under bazel by upgrading proc-macro-crate to 3.4.0.</li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/la10736/rstest/blob/master/CHANGELOG.md">rstest's changelog</a>.</em></p> <blockquote> <h2>[0.27.0] 2026/9/6</h2> <h3>Changed</h3> <ul> <li>Bump msrv to 1.85.0 both for <code>rstest</code> and <code>rstest_reuse</code></li> <li>Disabled default features of <code>futures-util</code></li> </ul> <h3>Added</h3> <ul> <li>Doc comments before <code>#[values(...)]</code> entries can be used to override the generated matrix test names (both for the legacy <code>arg => [..]</code> syntax and the new attribute form). See <a href="https://redirect.github.com/la10736/rstest/pull/321">#321</a> thanks to <a href="https://github.com/orhun"><code>@orhun</code></a>.</li> </ul> <h3>Fixed</h3> <ul> <li>Use fully-qualified <code>core</code> import. See <a href="https://redirect.github.com/la10736/rstest/pull/336">#336</a>.</li> <li>Fix <code>mut</code> arguments failing to compile with <code>#[trace]</code>. See <a href="https://redirect.github.com/la10736/rstest/pull/345">#345</a> thanks to <a href="https://github.com/super-cooper"><code>@super-cooper</code></a>.</li> <li>Fix compilation under bazel by upgrading proc-macro-crate to 3.4.0.</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/la10736/rstest/commit/59cd3d1919fc471922297606b8819af4e5bfd472"><code>59cd3d1</code></a> Release 0.27.0</li> <li><a href="https://github.com/la10736/rstest/commit/0821ebc65d0c8434f0d9b7773a23043549313426"><code>0821ebc</code></a> test: Add E2E test for mut arguments with #[trace]</li> <li><a href="https://github.com/la10736/rstest/commit/509ffef32538324bd75d0addf49ee7e0296bff81"><code>509ffef</code></a> fix: <code>mut</code> arguments with <code>#[trace]</code></li> <li><a href="https://github.com/la10736/rstest/commit/043437d63956f1c9e3f6284c4463169cb484cd3b"><code>043437d</code></a> fix: Resolve clippy warnings and truncate long test project names</li> <li><a href="https://github.com/la10736/rstest/commit/9aa8d1a3c51627ec87d0620abda9707488a6f8fe"><code>9aa8d1a</code></a> fix: Suppress nightly cargo lints in test scaffolding</li> <li><a href="https://github.com/la10736/rstest/commit/3d3c76c4fb7215aae92008e390b7fb618ecd5522"><code>3d3c76c</code></a> fix: Bump rstest_test MSRV to 1.85 and mark as unpublished</li> <li><a href="https://github.com/la10736/rstest/commit/d9ae990e323b6910d39364d94876a86b003b4729"><code>d9ae990</code></a> chore: Add changelog entry</li> <li><a href="https://github.com/la10736/rstest/commit/05d4b1a817c1ee4684e1cff728d805f3a647fc0b"><code>05d4b1a</code></a> fix: Use fully-qualified core import</li> <li><a href="https://github.com/la10736/rstest/commit/6da56a112e65b96685e06be41f8d98a84a6e7452"><code>6da56a1</code></a> Bump msrv to 1.85 also for rstest_reuse (<a href="https://redirect.github.com/la10736/rstest/issues/342">#342</a>)</li> <li><a href="https://github.com/la10736/rstest/commit/1e9963bbc1a0feff27afdd1c6cec25d00c96d607"><code>1e9963b</code></a> Add CLAUDE.md for Claude Code guidance (<a href="https://redirect.github.com/la10736/rstest/issues/340">#340</a>)</li> <li>Additional commits viewable in <a href="https://github.com/la10736/rstest/compare/v0.26.1...v0.27.0">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps dirs from 6.0.0 to 7.0.0. [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…he#25322) Bumps the all-other-cargo-deps group with 3 updates: [uuid](https://github.com/uuid-rs/uuid), [async-compression](https://github.com/Nullus157/async-compression) and [toml](https://github.com/toml-rs/toml). Updates `uuid` from 1.26.0 to 1.26.1 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/uuid-rs/uuid/releases">uuid's releases</a>.</em></p> <blockquote> <h2>v1.26.1</h2> <h2>What's Changed</h2> <ul> <li>Seat the v7 counter below the version nibble by <a href="https://github.com/lenamonj"><code>@lenamonj</code></a> in <a href="https://redirect.github.com/uuid-rs/uuid/pull/907">uuid-rs/uuid#907</a></li> <li>Don't panic in overflowing Timestamp to SystemTime conversion by <a href="https://github.com/KodrAus"><code>@KodrAus</code></a> in <a href="https://redirect.github.com/uuid-rs/uuid/pull/909">uuid-rs/uuid#909</a></li> <li>Prepare for 1.26.1 release by <a href="https://github.com/KodrAus"><code>@KodrAus</code></a> in <a href="https://redirect.github.com/uuid-rs/uuid/pull/910">uuid-rs/uuid#910</a></li> </ul> <h2>New Contributors</h2> <ul> <li><a href="https://github.com/lenamonj"><code>@lenamonj</code></a> made their first contribution in <a href="https://redirect.github.com/uuid-rs/uuid/pull/907">uuid-rs/uuid#907</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/uuid-rs/uuid/compare/v1.26.0...v1.26.1">https://github.com/uuid-rs/uuid/compare/v1.26.0...v1.26.1</a></p> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/uuid-rs/uuid/commit/9f927126c89892ddfed6cd2f92df16852f3f9aa6"><code>9f92712</code></a> Merge pull request <a href="https://redirect.github.com/uuid-rs/uuid/issues/910">#910</a> from uuid-rs/cargo/v1.26.1</li> <li><a href="https://github.com/uuid-rs/uuid/commit/d4df8f0cd9f461b4ef493254420052ffa5ce6277"><code>d4df8f0</code></a> prepare for 1.26.1 release</li> <li><a href="https://github.com/uuid-rs/uuid/commit/5613f2357c1fc06afc5fffd98e96da2ccf25a608"><code>5613f23</code></a> Merge pull request <a href="https://redirect.github.com/uuid-rs/uuid/issues/909">#909</a> from uuid-rs/fix/ts-conversion-overflow</li> <li><a href="https://github.com/uuid-rs/uuid/commit/fda00eba938383d242bad33143c8af73227f2a2c"><code>fda00eb</code></a> don't panic in overflowing Timestamp to SystemTime conversion</li> <li><a href="https://github.com/uuid-rs/uuid/commit/c82e88ca184e4ab83ce6b4ac0be33d32a0b9c3c4"><code>c82e88c</code></a> Merge pull request <a href="https://redirect.github.com/uuid-rs/uuid/issues/907">#907</a> from lenamonj/v7-counter-placement</li> <li><a href="https://github.com/uuid-rs/uuid/commit/ac065a6c17389a67dc6e98ef5f9a2e1d7ad4a670"><code>ac065a6</code></a> Align the counter diagram</li> <li><a href="https://github.com/uuid-rs/uuid/commit/34ec10208d813672928c299146dfe4c18cedcec7"><code>34ec102</code></a> Seat the v7 counter below the version nibble</li> <li>See full diff in <a href="https://github.com/uuid-rs/uuid/compare/v1.26.0...v1.26.1">compare view</a></li> </ul> </details> <br /> Updates `async-compression` from 0.4.44 to 0.4.46 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/Nullus157/async-compression/releases">async-compression's releases</a>.</em></p> <blockquote> <h2>async-compression-v0.4.46</h2> <h3>Other</h3> <ul> <li>Return UnexpectedEof when deflate64 does not progress (<a href="https://redirect.github.com/Nullus157/async-compression/pull/483">#483</a>)</li> </ul> <h2>async-compression-v0.4.45</h2> <h3>Other</h3> <ul> <li><em>(deps)</em> update zstd requirement from 0.13 to 0.14 (<a href="https://redirect.github.com/Nullus157/async-compression/pull/481">#481</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/Nullus157/async-compression/commit/3f2a60041bdb51c087166893504b35f7d007e0d1"><code>3f2a600</code></a> chore: release (<a href="https://redirect.github.com/Nullus157/async-compression/issues/484">#484</a>)</li> <li><a href="https://github.com/Nullus157/async-compression/commit/b22cd694cf6cb2f4753c8cae0bfdebd4b78b6fe3"><code>b22cd69</code></a> Return UnexpectedEof when deflate64 does not progress (<a href="https://redirect.github.com/Nullus157/async-compression/issues/483">#483</a>)</li> <li><a href="https://github.com/Nullus157/async-compression/commit/d8c1b476bda99a05c4cc67af1ef45c2dc2405926"><code>d8c1b47</code></a> chore: release (<a href="https://redirect.github.com/Nullus157/async-compression/issues/482">#482</a>)</li> <li><a href="https://github.com/Nullus157/async-compression/commit/9ffc96c79755cb24d80c1e15507d4ae6d795e9de"><code>9ffc96c</code></a> chore(deps): update zstd requirement from 0.13 to 0.14 (<a href="https://redirect.github.com/Nullus157/async-compression/issues/481">#481</a>)</li> <li>See full diff in <a href="https://github.com/Nullus157/async-compression/compare/async-compression-v0.4.44...async-compression-v0.4.46">compare view</a></li> </ul> </details> <br /> Updates `toml` from 1.1.5+spec-1.1.0 to 1.1.6+spec-1.1.0 <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/toml-rs/toml/commit/572c005d80cca5f7bd163805c2f33ba0a5207b6d"><code>572c005</code></a> chore: Release</li> <li><a href="https://github.com/toml-rs/toml/commit/66d0c53e1ec7cfc61f98c30b92f72e2ad1df4536"><code>66d0c53</code></a> docs: Update changelog</li> <li><a href="https://github.com/toml-rs/toml/commit/07af4e7644d34d2d784bcd4f570c0619ee1626b2"><code>07af4e7</code></a> perf: Reduce allocations in toml_edit parsing and dumping (<a href="https://redirect.github.com/toml-rs/toml/issues/1215">#1215</a>)</li> <li><a href="https://github.com/toml-rs/toml/commit/0ff90dbd7d7931b995f91cac661a12818cfd7ab7"><code>0ff90db</code></a> perf(display): Write encoded strings directly</li> <li><a href="https://github.com/toml-rs/toml/commit/efb25365450e5512efa76457fcd53b0bb1ee9c41"><code>efb2536</code></a> perf(display): Move generated representation strings</li> <li><a href="https://github.com/toml-rs/toml/commit/075c444fa05f359cce53dda0f1382a07ad72776e"><code>075c444</code></a> refactor(display): Consolidate key-path encoding</li> <li><a href="https://github.com/toml-rs/toml/commit/b48f33869f1024949945157c0ab540a114676fc0"><code>b48f338</code></a> perf(display): Borrow table keys during document output</li> <li><a href="https://github.com/toml-rs/toml/commit/c6c1de348f6e276a58ce57ffdcdeb51065698709"><code>c6c1de3</code></a> perf(parser): Move completed table header keys</li> <li><a href="https://github.com/toml-rs/toml/commit/ec904630a54a9857babc7eab99c39d4bdc8173c4"><code>ec90463</code></a> perf(parser): Borrow input when creating editable documents</li> <li><a href="https://github.com/toml-rs/toml/commit/d76a48af2f5bb783969295cf7dbbfb447f324ca1"><code>d76a48a</code></a> test: Benchmark rendering generated keys and values</li> <li>Additional commits viewable in <a href="https://github.com/toml-rs/toml/compare/toml-v1.1.5...toml-v1.1.6">compare view</a></li> </ul> </details> <br /> Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore <dependency name> major version` will close this group update PR and stop Dependabot creating any more for the specific dependency's major version (unless you unignore this specific dependency's major version or upgrade to it yourself) - `@dependabot ignore <dependency name> minor version` will close this group update PR and stop Dependabot creating any more for the specific dependency's minor version (unless you unignore this specific dependency's minor version or upgrade to it yourself) - `@dependabot ignore <dependency name>` will close this group update PR and stop Dependabot creating any more for the specific dependency (unless you unignore this specific dependency or upgrade to it yourself) - `@dependabot unignore <dependency name>` will remove all of the ignore conditions of the specified dependency - `@dependabot unignore <dependency name> <ignore condition>` will remove the ignore condition of the specified dependency and ignore conditions </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…ter (apache#24924) ## Rationale for this change Allows to roughly limit the size of parquet files when the number of rows is a hard-to-use estimator. This is especially useful when a column contains potentially large blobs. ## What changes are included in this PR? - Adds the config parameter - Add a new `AsyncWrite` implementation created by `ObjectWriterBuilder`. It sits between the optional compression and the `BufWriter` implementation, counting the number of effectively flushed bytes in a `Arc<AtomicUsize>`. - The `Arc<AtomicU64>` is used in the demuxer logic to get an estimation of the written size. This estimation lags because writes happens asynchronously but this avoid having to block on encoding and compression. - Use `ObjectWriterBuilder` when creating the single file `AsyncArrowWriter`
## Which issue does this PR close? - Closes apache#24719. ## Rationale for this change The statistics runner previously captured SQL parser settings only once. A successful `SET` in one query file therefore changed the session but did not affect parsing of later files, which made ordered query suites behave inconsistently with session state. ## What changes are included in this PR? - Read the current session's parser options immediately before parsing each query file. - Extract file processing into a focused helper without changing the existing same-file pre-parsing behavior. - Document that parser-setting changes still do not affect later statements in the same already-parsed file. - Add ordered-file regressions for MySQL dialect propagation and parser recursion-limit propagation. ## Are these changes tested? - `cargo test -p datafusion-benchmarks --lib statistics::tests -- --nocapture` (7 passed) - `cargo test -p datafusion-benchmarks --lib 'statistics::tests::refreshes_' -- --nocapture` after rebasing (2 passed) - `cargo check -p datafusion-benchmarks --lib` - `cargo clippy -p datafusion-benchmarks --lib -- -D warnings` - `cargo fmt --all -- --check` - Ablation: both new regressions fail when parsing uses stale default settings. The broader `cargo clippy -p datafusion-benchmarks --all-targets --all-features -- -D warnings` could not complete locally because the optional `snmalloc` target requires `cmake`, which is unavailable in this environment. It produced no Rust diagnostics before that build-tool failure. ## Are there any user-facing changes? Yes. Parser settings changed by an earlier statistics query file now apply when later query files are parsed. There are no public API changes. This draft was prepared with AI assistance and has not yet received human review. I understand the implementation end-to-end; there are no known design assumptions beyond the documented same-file parsing limitation. --------- Signed-off-by: Elton Chang <tchang52@ucsc.edu>
## Which issue does this PR close? - Closes apache#25280. ## Rationale for this change Adds a release version label to merged PRs so it's easy to determine which version a PR will land/has landed. ## What changes are included in this PR? - New workflow. - Updated the `dev/update_datafusion_versions.py` script to automatically update the version in the workflow. ## What is the testing strategy for this PR? This PR. ## Are there any user-facing changes? No.
…ache#25777) ## Which issue does this PR close? - Closes apache#25578. ## Rationale for this change See the issue and apache#14305. ## What changes are included in this PR? - New CI job that checks if `datafusion-examples` uses subcrates directly which are already exported by the main `datafusion` crate. - Updated the examples to only use the main `datafusion` crate. ## What is the testing strategy for this PR? This PR. ## Are there any user-facing changes? Examples now do not use subcrates directly without needing to.
## Which issue does this PR close? - Related to apache#17098. ## Rationale for this change Once the TopK heap is full, most input batches are rejected outright by the threshold filter. `insert_batch` still evaluated every sort expression over the whole batch before checking the filter, then evaluated the same expressions again inside the filter, and cloned the batch even when nothing from it was kept. With an expression sort key that doubles the per-batch cost for no output. ## What changes are included in this PR? - The threshold filter runs before the sort keys are evaluated; fully rejected batches skip sort-key evaluation. - The batch moves into the heap entry instead of being cloned. The last row's common prefix is encoded before the move so early completion still compares against the same boundary at the same point. - The per-batch timer clones one metric instead of the whole baseline set, here and in `PartitionedTopK`, `PartitionedTopKRank`, and `PartitionedTopKDenseRank`. `SELECT l_orderkey, l_extendedprice FROM lineitem ORDER BY l_extendedprice + 0 DESC LIMIT 1` on TPC-H SF1, TopK `elapsed_compute`, median of five: 152 ms before, 94 ms after. `sort-tpch --limit 10` shows no change. One nuance: a sort-key expression that errors only on rows the threshold already rejects no longer raises, because those batches skip sort-key evaluation. The filter evaluates the same expression over the batch first, so this can only differ for a secondary key under short-circuit evaluation. ## What is the testing strategy for this PR? New unit tests in `datafusion/physical-plan/src/topk`: `test_batch_rejected_by_threshold_leaves_heap_untouched` and `test_expression_sort_key_matches_column_sort_key`. Existing TopK tests and the `topk`, `window` and `limit` sqllogictest files pass unchanged. ## Are there any user-facing changes? No.
## Which issue does this PR close? Closes apache#25582. ## Rationale for this change Consumers reporting metrics for one partition currently clone all registered metrics before filtering. This becomes increasingly expensive when a physical plan is retained across many partition executions. ## What changes are included in this PR? - Add `MetricsSet::for_partition` with indexed partition selection. - Record snapshot boundaries so `metrics()` does not need to clone every metric handle. Build and update the shared partition index only when selecting a partition. - Preserve registration order and existing iterator APIs through lazy materialization. ## What is the testing strategy for this PR? Unit and integration tests cover partition selection, snapshot isolation, concurrent registration, and FFI round trips. ## Are there any user-facing changes? Callers can select a partition with `plan.metrics().map(|m| m.for_partition(partition))`. Unpartitioned metrics remain available in the full set; unknown partitions produce an empty set. The existing `ExecutionPlan` API and FFI layout remain unchanged. Partition selection does not eliminate work already performed by a metrics provider, such as FFI transport.
…he#25845) ## Which issue does this PR close? - N/A. ## Rationale for this change Follow up to apache#25776. Adds the missing checkout to the release version workflow. ## What changes are included in this PR? - Added the missing checkout. - Fix the script fmt. ## What is the testing strategy for this PR? N/A. ## Are there any user-facing changes? No.
…apache#25749) ## Which issue does this PR close? - Closes apache#25717. ## Rationale for this change A transparent wrapper `ExecutionPlan` (one that implements `downcast_delegate`, such as `datafusion-tracing`'s `InstrumentedExec`) around a `ForeignExecutionPlan` is silently dropped when it is passed back across the FFI boundary. For example, installing the tracing instrumentation rule via FFI leaves the plan un-instrumented. Also, when such a wrapper is the child of a foreign plan, it does not receive the runtime handle. ## What changes are included in this PR? - `FFI_ExecutionPlan::new`: the "already foreign" shortcut now checks the concrete type via `Any` instead of the delegate-aware `downcast_ref`, so it only fires for a real `ForeignExecutionPlan`. - `pass_runtime_to_children`: the child check also uses `Any`, so a local wrapper child of a foreign parent is re-wrapped with the runtime. - `plan_is_foreign` intentionally stays delegate-aware, since a delegating wrapper typically forwards `replace_children` to the foreign plan and its children still need the runtime. A comment explains this. ## What is the testing strategy for this PR? - Added an optional `downcast_delegate` mode to the existing `EmptyExec` test plan (`with_downcast_delegate`). - Added `test_ffi_execution_plan_delegating_wrapper`, which checks that the wrapper survives `FFI_ExecutionPlan::new` and that a wrapper child of a foreign parent gets the runtime. Reverting either fix on its own makes the test fail. ## Are there any user-facing changes? No API changes. Wrapper plans that implement `downcast_delegate` are now kept when they cross the FFI boundary instead of being discarded. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
## Which issue does this PR close? Follow-up to [the review of apache#25383](apache#25383 (review)), which suggested reducing intermediate aggregate spill work. ## Rationale for this change An intermediate aggregate merge can rewrite rows that could remain on disk until final replay. With eight uniform full-batch runs and an admitted fan-in of seven, merging six inputs leaves three for the final pass. The ordinary selection merges seven and leaves two, unnecessarily rewriting one run. The final pass must still fit if another partition consumes available memory during the intermediate write. The retained seven-input buffer reservation can cover three final inputs plus their replay headroom; it cannot guarantee seven final inputs with that headroom. ## What changes are included in this PR? Eligible spill-only intermediate merges select only enough inputs to leave a final pass whose expected buffers and replay headroom fit the existing buffer reservation. An actually trimmed pass retains that reservation in the merge builder through intermediate EOF and spill completion, transferring it to the next ordinary admission. Sizing acquires no additional pool memory and preserves read-ahead and output batching. Passes that are not trimmed keep ordinary reservation ownership. Only the newer aggregate implementation (`AggregateSpill`, enabled by default) opts in. Legacy aggregation keeps its original selection. Sizing requires equal recorded maximum batch memory and batch-size limits across admitted and pending runs. Every original run must contain a full target-size batch; short original runs disable sizing for their replay. Intermediate and split spill writes also check actual largest-batch row counts against their output limit. These checks use private metadata; public APIs are unchanged. The sizing budget assumes the produced run is no wider than the guarded inputs. Variable-width output can still require further intermediate work. Final admission and its normal headroom release remain unchanged. ## What is the testing strategy for this PR? Head: `66569072850d91476b042010906ac6461ebbbd08`. Comparison base: `991fd23dd0be0046af5945b8c3905612859b657a`. The regression tests cover competing memory consumers at replay-headroom release and intermediate EOF, both beneficial and unchanged merge selections, short batches, and heterogeneous input budgets. They check exact ordered results, intermediate rewrite counts, actual competing reservations, memory and spill-file cleanup, and reservation ownership when the intermediate stream and builder are dropped. CI runs for this head: [push Rust run](https://github.com/apache/datafusion/actions/runs/36461591867) and [pull-request Rust run](https://github.com/apache/datafusion/actions/runs/36461597627). ### Performance validation The updated head has not been rebenchmarked against `991fd23dd0be0046af5945b8c3905612859b657a`. ## Are there any user-facing changes? Eligible spilling aggregations can perform fewer intermediate row writes. Results and configuration remain unchanged. Short runs and selections without sufficient retained replay capacity use the ordinary selection.
…try_new()` (apache#25786) ## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes apache#123` indicates that this PR will close issue apache#123. --> - Closes #. ## Rationale for this change <!-- Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed. Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes. Please explain the problem you are trying to solve in terms of the user-visible behavior, rather than the implementation. For example, "The code in `foo.rs` doesn't handle nulls" is a symptom of the implementation. "COUNT(DISTINCT) returns wrong results when the column contains nulls" is the user-visible problem. --> If both group by expression and aggregate expression is empty, such aggregate is not valid. ``` AggregateExec { group_by_expr: vec![], aggr_expr: vec![], .... } ``` Now `AggregateExec::try_new()` with such invalid input would succeed, this PR adds a validation for it. (this case is not possible from SQL interface, but it can be built through raw `AggregateExec` API) ## What changes are included in this PR? <!-- There is no need to duplicate the description in the issue here, but it is sometimes worth providing a summary of the individual changes in this PR. --> 1. Error for this case 2. Update several optimizer UTs to avoid building such plan ## What is the testing strategy for this PR? <!-- We typically require tests for all PRs in order to: 1. Prevent the code from being accidentally broken by subsequent changes 3. Serve as another way to document the expected behavior of the code Briefly describe how this PR is tested, and point to the specific tests you added. For example: 'This new feature is covered by the `sqllogictest` cases added in `foo.slt`'. If this PR does not add tests, explain why. For example, if the change is already covered by existing tests, please mention it. You should also check the `codecov` bot reply on this PR to confirm the changed code is exercised. --> UT (this case it not able to test from SQL) ## Are there any user-facing changes? <!-- If there are user-facing changes then we may require documentation to be updated before approving the PR. If there are any breaking changes to public APIs, please add the `api change` label. --> no
…apache#25549) ## Which issue does this PR close? - Closes apache#25548. ## Rationale for this change When we have a `Filter` next to a `Projection`, we might not be able to unparse them at the same level if the predicate depends on columns aliased by the projection. For those cases, we need to add a subquery in between. ## What changes are included in this PR? - Iff any column of a `Filter` is aliased by a `Projection` directly below it, a subquery is added with `derive`. ## What is the testing strategy for this PR? Unit tests. ## Are there any user-facing changes? No, this only adds the subquery when it is necessary, so existing tests remain unchanged.
…c filters (apache#25579) The fix prevents one aggregate’s filter from discarding rows that another aggregate still needs in dynamic filters. ### Example: Query: `MIN(a), MAX(a), MIN(b), MAX(b)` ``` First batch: (1, NULL), (8, NULL) | v a has bounds [1, 8] b has no bounds yet | +---------+----------+ | | Before After | | a < 1 OR a > 8 filter = true | | v v Next batch: (2, 5), (3, 7) | | Both rows rejected Both rows accepted | | v v Result: Now b has bounds [5, 7] 1 8 NULL NULL | v Filtering can start: a < 1 OR a > 8 OR b < 5 OR b > 7 | v Correct result: 1 8 5 7 ``` The latter is correct 1. According to [Postgres](https://www.postgresql.org/docs/current/functions-aggregate.html): > - MIN: “Computes the minimum of the non-null input values.” > - MAX: “Computes the maximum of the non-null input values.” 2. According to [DataFusion](https://github.com/apache/datafusion/blob/6ff484a96d7eca4b1bf4a263d7190ada5af2343c/datafusion/sqllogictest/test_files/aggregate.slt#L5490-L5500) Executing `1, NULL, 3, 4, 5` expects `MIN = 1, MAX = 5` --------- Co-authored-by: Gabriel <45515538+gabotechs@users.noreply.github.com>
…rgo-deps group (apache#25869) Bumps the all-other-cargo-deps group with 1 update: [thiserror](https://github.com/dtolnay/thiserror). Updates `thiserror` from 2.0.20 to 2.0.21 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/dtolnay/thiserror/releases">thiserror's releases</a>.</em></p> <blockquote> <h2>2.0.21</h2> <ul> <li>Fix parsing of generic unit variants in display expressions (<a href="https://redirect.github.com/dtolnay/thiserror/issues/459">#459</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/dtolnay/thiserror/commit/b1827ee06f81a7d7f954e676771e19a97f9f8e3b"><code>b1827ee</code></a> Release 2.0.21</li> <li><a href="https://github.com/dtolnay/thiserror/commit/58037b575a2a5543d0d54a49360554325b49f55a"><code>58037b5</code></a> Merge pull request <a href="https://redirect.github.com/dtolnay/thiserror/issues/459">#459</a> from dtolnay/turbofish</li> <li><a href="https://github.com/dtolnay/thiserror/commit/f82a0cf8f06c71263c2caa26e759f337d2778153"><code>f82a0cf</code></a> Keep track of nested turbofish depth</li> <li><a href="https://github.com/dtolnay/thiserror/commit/72ea49262d6412ccbc08f1696dda8626409e657d"><code>72ea492</code></a> Raise required compiler to Rust 1.77</li> <li><a href="https://github.com/dtolnay/thiserror/commit/72eea0d4ddb17ff9fa873aa74469de850021b22e"><code>72eea0d</code></a> Resolve io_other_error clippy lint in tests</li> <li><a href="https://github.com/dtolnay/thiserror/commit/07f09a2ec934df58508ce4fa5e7fbf27cca552c2"><code>07f09a2</code></a> Raise required compiler to Rust 1.74</li> <li><a href="https://github.com/dtolnay/thiserror/commit/2715388e8cc5787b3643251f7a7a637e89376688"><code>2715388</code></a> Update ui test suite to nightly-2026-09-22</li> <li><a href="https://github.com/dtolnay/thiserror/commit/5a306c7d0a8588caaaaa6a7567aeb25c1c10719b"><code>5a306c7</code></a> Update ui test suite to nightly-2026-09-05</li> <li><a href="https://github.com/dtolnay/thiserror/commit/ef9383b37d7c96b01e10df35d3e9e313c7b98b5a"><code>ef9383b</code></a> Update ui test suite to nightly-2026-08-22</li> <li><a href="https://github.com/dtolnay/thiserror/commit/8336b8407fbfe69177c504cdbc379db20cf6f131"><code>8336b84</code></a> Update ui tests for version 2.0.20</li> <li>See full diff in <a href="https://github.com/dtolnay/thiserror/compare/2.0.20...2.0.21">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore <dependency name> major version` will close this group update PR and stop Dependabot creating any more for the specific dependency's major version (unless you unignore this specific dependency's major version or upgrade to it yourself) - `@dependabot ignore <dependency name> minor version` will close this group update PR and stop Dependabot creating any more for the specific dependency's minor version (unless you unignore this specific dependency's minor version or upgrade to it yourself) - `@dependabot ignore <dependency name>` will close this group update PR and stop Dependabot creating any more for the specific dependency (unless you unignore this specific dependency or upgrade to it yourself) - `@dependabot unignore <dependency name>` will remove all of the ignore conditions of the specified dependency - `@dependabot unignore <dependency name> <ignore condition>` will remove the ignore condition of the specified dependency and ignore conditions </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
) Bumps [astral-sh/setup-uv](https://github.com/astral-sh/setup-uv) from 10.1.0 to 10.2.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/astral-sh/setup-uv/releases">astral-sh/setup-uv's releases</a>.</em></p> <blockquote> <h2>v10.2.0 🌈 Disable automatic cache saves for merge queues</h2> <h2>Changes</h2> <p>This release contains the known-checksum of the most recent uv releases and also disabled the uploading(saving) of the cache when in a merge queue since theses caches would almost never be used.</p> <h2>🚀 Enhancements</h2> <ul> <li>Disable automatic cache saves for merge queues <a href="https://github.com/eifinger"><code>@eifinger</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1056">#1056</a>)</li> </ul> <h2>🧰 Maintenance</h2> <ul> <li>chore: update known checksums for 0.12.17 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1058">#1058</a>)</li> <li>chore: update known checksums for 0.12.16 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1057">#1057</a>)</li> <li>chore: update known checksums for 0.12.15 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1054">#1054</a>)</li> <li>chore: update known checksums for 0.12.14 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1053">#1053</a>)</li> <li>chore: update known checksums for 0.12.13 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1045">#1045</a>)</li> </ul> <h2>📚 Documentation</h2> <ul> <li>docs: update version references to v10.1.0 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1044">#1044</a>)</li> </ul> <h2>⬆️ Dependency updates</h2> <ul> <li>chore(deps): roll up Dependabot updates <a href="https://github.com/eifinger"><code>@eifinger</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1059">#1059</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/astral-sh/setup-uv/commit/c18668ad3cf93ea998bef934396af7bb5c839dc7"><code>c18668a</code></a> chore(deps): roll up Dependabot updates (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1059">#1059</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/ffe14763056ca34ecd158146a9fc7e144c8a2753"><code>ffe1476</code></a> chore: update known checksums for 0.12.17 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1058">#1058</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/f5548c55522a1db0af3c84f2af3d058bc9bc2de2"><code>f5548c5</code></a> chore: update known checksums for 0.12.16 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1057">#1057</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/a761a4e9afd7b2f353ae020bd6d3a3af34c6c4d5"><code>a761a4e</code></a> Disable automatic cache saves for merge queues (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1056">#1056</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/3377a30666f438759955882b3eba6a92b2b29b12"><code>3377a30</code></a> chore: update known checksums for 0.12.15 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1054">#1054</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/dfb5f386776afcea37f271f3b656318d9949b9f1"><code>dfb5f38</code></a> chore: update known checksums for 0.12.14 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1053">#1053</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/45c121f982720f3bdf236c2ac06f126ca211949c"><code>45c121f</code></a> chore: update known checksums for 0.12.13 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1045">#1045</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/8073452fd4b566e886f04f731bcdbd8332b0b779"><code>8073452</code></a> docs: update version references to v10.1.0 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/1044">#1044</a>)</li> <li>See full diff in <a href="https://github.com/astral-sh/setup-uv/compare/bec219d24cd3e171d82865faccec33120bb574f4...c18668ad3cf93ea998bef934396af7bb5c839dc7">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…che#25867) Bumps [taiki-e/install-action](https://github.com/taiki-e/install-action) from 2.87.16 to 2.87.21. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/taiki-e/install-action/releases">taiki-e/install-action's releases</a>.</em></p> <blockquote> <h2>2.87.21</h2> <ul> <li> <p>Update <code>wasm-bindgen@latest</code> to 0.2.129.</p> </li> <li> <p>Update <code>protoc-gen-connect-openapi@latest</code> to 0.27.3.</p> </li> <li> <p>Update <code>wasmtime@latest</code> to 49.0.1.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.12.19.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.9.13.</p> </li> <li> <p>Update <code>kingfisher@latest</code> to 2.7.0.</p> </li> </ul> <h2>2.87.20</h2> <ul> <li> <p>Update <code>uv@latest</code> to 0.12.18.</p> </li> <li> <p>Update <code>cargo-shear@latest</code> to 1.14.0.</p> </li> </ul> <h2>2.87.19</h2> <ul> <li> <p>Update <code>wasmtime@latest</code> to 49.0.0.</p> </li> <li> <p>Update <code>cargo-shear@latest</code> to 1.13.5.</p> </li> <li> <p>Update <code>cargo-nextest@latest</code> to 0.9.146.</p> </li> </ul> <h2>2.87.18</h2> <ul> <li> <p>Update <code>oxfmt@latest</code> to 1.84.0.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.9.12.</p> </li> <li> <p>Update <code>kache@latest</code> to 0.26.3.</p> </li> <li> <p>Update <code>cargo-tarpaulin@latest</code> to 0.37.4.</p> </li> <li> <p>Update <code>cargo-rdme@latest</code> to 2.2.3.</p> </li> </ul> <h2>2.87.17</h2> <ul> <li> <p>Update <code>uv@latest</code> to 0.12.17.</p> </li> <li> <p>Update <code>release-plz@latest</code> to 0.3.169.</p> </li> <li> <p>Update <code>kingfisher@latest</code> to 2.5.0.</p> </li> <li> <p>Update <code>kache@latest</code> to 0.25.0.</p> </li> <li> <p>Update <code>git-cliff@latest</code> to 2.14.2.</p> </li> <li> <p>Update <code>cargo-tarpaulin@latest</code> to 0.37.3.</p> </li> <li> <p>Update <code>cargo-leptos@latest</code> to 0.3.9.</p> </li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/taiki-e/install-action/blob/main/CHANGELOG.md">taiki-e/install-action's changelog</a>.</em></p> <blockquote> <h1>Changelog</h1> <p>All notable changes to this project will be documented in this file.</p> <p>This project adheres to <a href="https://semver.org">Semantic Versioning</a>.</p> <!-- raw HTML omitted --> <h2>[Unreleased]</h2> <h2>[2.87.21] - 2026-09-26</h2> <ul> <li> <p>Update <code>wasm-bindgen@latest</code> to 0.2.129.</p> </li> <li> <p>Update <code>protoc-gen-connect-openapi@latest</code> to 0.27.3.</p> </li> <li> <p>Update <code>wasmtime@latest</code> to 49.0.1.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.12.19.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.9.13.</p> </li> <li> <p>Update <code>kingfisher@latest</code> to 2.7.0.</p> </li> </ul> <h2>[2.87.20] - 2026-09-24</h2> <ul> <li> <p>Update <code>uv@latest</code> to 0.12.18.</p> </li> <li> <p>Update <code>cargo-shear@latest</code> to 1.14.0.</p> </li> </ul> <h2>[2.87.19] - 2026-09-23</h2> <ul> <li> <p>Update <code>wasmtime@latest</code> to 49.0.0.</p> </li> <li> <p>Update <code>cargo-shear@latest</code> to 1.13.5.</p> </li> <li> <p>Update <code>cargo-nextest@latest</code> to 0.9.146.</p> </li> </ul> <h2>[2.87.18] - 2026-09-22</h2> <ul> <li> <p>Update <code>oxfmt@latest</code> to 1.84.0.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.9.12.</p> </li> <li> <p>Update <code>kache@latest</code> to 0.26.3.</p> </li> <li> <p>Update <code>cargo-tarpaulin@latest</code> to 0.37.4.</p> </li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/taiki-e/install-action/commit/4cef1412cce204788f482e778a0b9187f9626a29"><code>4cef141</code></a> Release 2.87.21</li> <li><a href="https://github.com/taiki-e/install-action/commit/133135e7fd1570baff8964a044bcf0d2147fc57a"><code>133135e</code></a> Update <code>wasm-bindgen@latest</code> to 0.2.129</li> <li><a href="https://github.com/taiki-e/install-action/commit/8589241798155a8c0c86c7c0ae4f642a90a698e1"><code>8589241</code></a> Update <code>protoc-gen-connect-openapi@latest</code> to 0.27.3</li> <li><a href="https://github.com/taiki-e/install-action/commit/95e5bbc4a58633e12ebefd904f798c1eb76c18f7"><code>95e5bbc</code></a> Update wasmtime manifest</li> <li><a href="https://github.com/taiki-e/install-action/commit/48ce11d4f97676e64a76a3dc4b83a4f59245b702"><code>48ce11d</code></a> Update <code>wasmtime@latest</code> to 49.0.1</li> <li><a href="https://github.com/taiki-e/install-action/commit/368e6617a899743920fa751bb5d414b6f0352cc0"><code>368e661</code></a> Update wasm-bindgen manifest</li> <li><a href="https://github.com/taiki-e/install-action/commit/57f22a457e27d10a13bfcaee43e03dff1c20dedd"><code>57f22a4</code></a> Update <code>uv@latest</code> to 0.12.19</li> <li><a href="https://github.com/taiki-e/install-action/commit/888634e2727c1021a94ebfac9a2d266109a366b8"><code>888634e</code></a> Update typos manifest</li> <li><a href="https://github.com/taiki-e/install-action/commit/3bcab01fb9790f533b19647e60396514a9d08f96"><code>3bcab01</code></a> Update protoc-gen-connect-openapi manifest</li> <li><a href="https://github.com/taiki-e/install-action/commit/947e90eb285641d8325cc2d55ae24bd51d5c6b8a"><code>947e90e</code></a> Update <code>mise@latest</code> to 2026.9.13</li> <li>Additional commits viewable in <a href="https://github.com/taiki-e/install-action/compare/9114bf4d891761788c546334fd37538eae1bf8b3...4cef1412cce204788f482e778a0b9187f9626a29">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…updates (apache#25866) Bumps the codeql-actions group with 2 updates in the / directory: [github/codeql-action/init](https://github.com/github/codeql-action) and [github/codeql-action/analyze](https://github.com/github/codeql-action). Updates `github/codeql-action/init` from 4.38.1 to 4.38.2 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/releases">github/codeql-action/init's releases</a>.</em></p> <blockquote> <h2>v4.38.2</h2> <ul> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.27.1">2.27.1</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4160">#4160</a></li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/blob/main/CHANGELOG.md">github/codeql-action/init's changelog</a>.</em></p> <blockquote> <h1>CodeQL Action Changelog</h1> <p>See the <a href="https://github.com/github/codeql-action/releases">releases page</a> for the relevant changes to the CodeQL CLI and language packs.</p> <h2>[UNRELEASED]</h2> <p>No user facing changes.</p> <h2>4.38.2 - 24 Sept 2026</h2> <ul> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.27.1">2.27.1</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4160">#4160</a></li> </ul> <h2>4.38.1 - 18 Sept 2026</h2> <ul> <li>The CodeQL Action now has experimental support for CodeQL releases for which per-language bundles are available. Per-language bundles support analysis for a single language and are therefore smaller than the combined bundles that allow analysis for all supported languages. As a result, per-language bundles take up less space on disk and are faster to download. We expect to roll this change out to everyone in the coming weeks. <a href="https://redirect.github.com/github/codeql-action/pull/4146">#4146</a></li> </ul> <h2>4.38.0 - 09 Sept 2026</h2> <ul> <li>On GitHub-hosted runners, the CodeQL Action now deletes unused CodeQL bundles from the toolcache before downloading a different bundle, which frees up disk space for the analysis. We expect to roll this change out to everyone in September. <a href="https://redirect.github.com/github/codeql-action/pull/4124">#4124</a></li> <li>The CodeQL Action now supports CodeQL releases that are compatible with Linux Arm64 and downloads the native <code>linux-arm64</code> CodeQL bundle when available. <a href="https://redirect.github.com/github/codeql-action/pull/4072">#4072</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.27.0">2.27.0</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4129">#4129</a></li> </ul> <h2>4.37.9 - 26 Aug 2026</h2> <ul> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.4">2.26.4</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4106">#4106</a></li> </ul> <h2>4.37.8 - 21 Aug 2026</h2> <p>No user facing changes.</p> <h2>4.37.7 - 13 Aug 2026</h2> <ul> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.3">2.26.3</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4085">#4085</a></li> </ul> <h2>4.37.6 - 04 Aug 2026</h2> <ul> <li>Changed the default filepath for the new remote file address format that was introduced in CodeQL Action 4.37.0 / 3.37.0 to <code>.github/codeql-config.yml</code> to align it with the suggested path that is used elsewhere. <a href="https://redirect.github.com/github/codeql-action/pull/4070">#4070</a></li> </ul> <h2>4.37.5 - 03 Aug 2026</h2> <ul> <li>Fixed a bug where a network error while streaming the download of the CodeQL bundle could terminate the <code>init</code> Action instead of falling back to downloading the bundle before extracting it. <a href="https://redirect.github.com/github/codeql-action/pull/4061">#4061</a></li> </ul> <h2>4.37.4 - 29 Jul 2026</h2> <ul> <li>This version of the CodeQL Action adds support for the <code>tools</code> input for the <code>codeql-action/init</code> step to be specified using a <code>github-codeql-tools</code> <a href="https://docs.github.com/en/organizations/managing-organization-settings/managing-custom-properties-for-repositories-in-your-organization">repository property</a>. This feature will gradually be rolled out following the release of this version. Once rolled out, this allows for the CodeQL CLI version that is used in GitHub-managed workflows, such as Default Setup, to be set to a custom value. For example, customers who run into issues with rate limits when a new CodeQL CLI version is released can set the value to <code>toolcache</code> to always use the CodeQL CLI version that is available in the runner toolcache. For Advanced Setup workflows, the value provided for <code>tools</code> in the workflow definition always takes precedence unless the value of the repository property starts with <code>!</code>. <a href="https://redirect.github.com/github/codeql-action/pull/4037">#4037</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.2">2.26.2</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4051">#4051</a></li> </ul> <h2>4.37.3 - 22 Jul 2026</h2> <p>No user facing changes.</p> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/github/codeql-action/commit/2892aa5e19bbd11bc0cff5427e3b750a04d9e3c2"><code>2892aa5</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4168">#4168</a> from github/update-v4.38.2-a6ef2c96f</li> <li><a href="https://github.com/github/codeql-action/commit/8ad03a333eb88de8ad6833eda208d0fc51a9c571"><code>8ad03a3</code></a> Trigger workflows</li> <li><a href="https://github.com/github/codeql-action/commit/98af865db5041cee73c7185896319367f8c0adf2"><code>98af865</code></a> Update changelog for v4.38.2</li> <li><a href="https://github.com/github/codeql-action/commit/a6ef2c96fc0e37d0b44fb2bd0b32db4bcb89ae24"><code>a6ef2c9</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4156">#4156</a> from github/mario-campos/fix-validate-cmd</li> <li><a href="https://github.com/github/codeql-action/commit/1ef28a1b7603ca158fd774d1ab328cbd6a40b84b"><code>1ef28a1</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4166">#4166</a> from github/dependabot/github_actions/dot-github/wor...</li> <li><a href="https://github.com/github/codeql-action/commit/26cb08bab0037de74cc66ad9ec0dca31d6d9e8a7"><code>26cb08b</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4163">#4163</a> from github/mbg/fix-getCommitOid-stubs</li> <li><a href="https://github.com/github/codeql-action/commit/f035ce3a985a1223a9f59fb719542598640160b2"><code>f035ce3</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4165">#4165</a> from github/dependabot/npm_and_yarn/npm-minor-8eaed9...</li> <li><a href="https://github.com/github/codeql-action/commit/5e4e2550b48d7f3de205c9d752eb5176bf07f6d9"><code>5e4e255</code></a> Rebuild</li> <li><a href="https://github.com/github/codeql-action/commit/b13f5f47d5398d0fb982942ced6fdfc4e3951804"><code>b13f5f4</code></a> Bump ruby/setup-ruby</li> <li><a href="https://github.com/github/codeql-action/commit/c87fe5756c0c0bcd5e0005d2169945cfee9a232f"><code>c87fe57</code></a> Rebuild</li> <li>Additional commits viewable in <a href="https://github.com/github/codeql-action/compare/1c5b675653bb5c22dbe9b12b556ec555138e09fd...2892aa5e19bbd11bc0cff5427e3b750a04d9e3c2">compare view</a></li> </ul> </details> <br /> Updates `github/codeql-action/analyze` from 4.38.1 to 4.38.2 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/releases">github/codeql-action/analyze's releases</a>.</em></p> <blockquote> <h2>v4.38.2</h2> <ul> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.27.1">2.27.1</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4160">#4160</a></li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/blob/main/CHANGELOG.md">github/codeql-action/analyze's changelog</a>.</em></p> <blockquote> <h1>CodeQL Action Changelog</h1> <p>See the <a href="https://github.com/github/codeql-action/releases">releases page</a> for the relevant changes to the CodeQL CLI and language packs.</p> <h2>[UNRELEASED]</h2> <p>No user facing changes.</p> <h2>4.38.2 - 24 Sept 2026</h2> <ul> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.27.1">2.27.1</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4160">#4160</a></li> </ul> <h2>4.38.1 - 18 Sept 2026</h2> <ul> <li>The CodeQL Action now has experimental support for CodeQL releases for which per-language bundles are available. Per-language bundles support analysis for a single language and are therefore smaller than the combined bundles that allow analysis for all supported languages. As a result, per-language bundles take up less space on disk and are faster to download. We expect to roll this change out to everyone in the coming weeks. <a href="https://redirect.github.com/github/codeql-action/pull/4146">#4146</a></li> </ul> <h2>4.38.0 - 09 Sept 2026</h2> <ul> <li>On GitHub-hosted runners, the CodeQL Action now deletes unused CodeQL bundles from the toolcache before downloading a different bundle, which frees up disk space for the analysis. We expect to roll this change out to everyone in September. <a href="https://redirect.github.com/github/codeql-action/pull/4124">#4124</a></li> <li>The CodeQL Action now supports CodeQL releases that are compatible with Linux Arm64 and downloads the native <code>linux-arm64</code> CodeQL bundle when available. <a href="https://redirect.github.com/github/codeql-action/pull/4072">#4072</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.27.0">2.27.0</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4129">#4129</a></li> </ul> <h2>4.37.9 - 26 Aug 2026</h2> <ul> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.4">2.26.4</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4106">#4106</a></li> </ul> <h2>4.37.8 - 21 Aug 2026</h2> <p>No user facing changes.</p> <h2>4.37.7 - 13 Aug 2026</h2> <ul> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.3">2.26.3</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4085">#4085</a></li> </ul> <h2>4.37.6 - 04 Aug 2026</h2> <ul> <li>Changed the default filepath for the new remote file address format that was introduced in CodeQL Action 4.37.0 / 3.37.0 to <code>.github/codeql-config.yml</code> to align it with the suggested path that is used elsewhere. <a href="https://redirect.github.com/github/codeql-action/pull/4070">#4070</a></li> </ul> <h2>4.37.5 - 03 Aug 2026</h2> <ul> <li>Fixed a bug where a network error while streaming the download of the CodeQL bundle could terminate the <code>init</code> Action instead of falling back to downloading the bundle before extracting it. <a href="https://redirect.github.com/github/codeql-action/pull/4061">#4061</a></li> </ul> <h2>4.37.4 - 29 Jul 2026</h2> <ul> <li>This version of the CodeQL Action adds support for the <code>tools</code> input for the <code>codeql-action/init</code> step to be specified using a <code>github-codeql-tools</code> <a href="https://docs.github.com/en/organizations/managing-organization-settings/managing-custom-properties-for-repositories-in-your-organization">repository property</a>. This feature will gradually be rolled out following the release of this version. Once rolled out, this allows for the CodeQL CLI version that is used in GitHub-managed workflows, such as Default Setup, to be set to a custom value. For example, customers who run into issues with rate limits when a new CodeQL CLI version is released can set the value to <code>toolcache</code> to always use the CodeQL CLI version that is available in the runner toolcache. For Advanced Setup workflows, the value provided for <code>tools</code> in the workflow definition always takes precedence unless the value of the repository property starts with <code>!</code>. <a href="https://redirect.github.com/github/codeql-action/pull/4037">#4037</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.2">2.26.2</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4051">#4051</a></li> </ul> <h2>4.37.3 - 22 Jul 2026</h2> <p>No user facing changes.</p> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/github/codeql-action/commit/2892aa5e19bbd11bc0cff5427e3b750a04d9e3c2"><code>2892aa5</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4168">#4168</a> from github/update-v4.38.2-a6ef2c96f</li> <li><a href="https://github.com/github/codeql-action/commit/8ad03a333eb88de8ad6833eda208d0fc51a9c571"><code>8ad03a3</code></a> Trigger workflows</li> <li><a href="https://github.com/github/codeql-action/commit/98af865db5041cee73c7185896319367f8c0adf2"><code>98af865</code></a> Update changelog for v4.38.2</li> <li><a href="https://github.com/github/codeql-action/commit/a6ef2c96fc0e37d0b44fb2bd0b32db4bcb89ae24"><code>a6ef2c9</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4156">#4156</a> from github/mario-campos/fix-validate-cmd</li> <li><a href="https://github.com/github/codeql-action/commit/1ef28a1b7603ca158fd774d1ab328cbd6a40b84b"><code>1ef28a1</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4166">#4166</a> from github/dependabot/github_actions/dot-github/wor...</li> <li><a href="https://github.com/github/codeql-action/commit/26cb08bab0037de74cc66ad9ec0dca31d6d9e8a7"><code>26cb08b</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4163">#4163</a> from github/mbg/fix-getCommitOid-stubs</li> <li><a href="https://github.com/github/codeql-action/commit/f035ce3a985a1223a9f59fb719542598640160b2"><code>f035ce3</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4165">#4165</a> from github/dependabot/npm_and_yarn/npm-minor-8eaed9...</li> <li><a href="https://github.com/github/codeql-action/commit/5e4e2550b48d7f3de205c9d752eb5176bf07f6d9"><code>5e4e255</code></a> Rebuild</li> <li><a href="https://github.com/github/codeql-action/commit/b13f5f47d5398d0fb982942ced6fdfc4e3951804"><code>b13f5f4</code></a> Bump ruby/setup-ruby</li> <li><a href="https://github.com/github/codeql-action/commit/c87fe5756c0c0bcd5e0005d2169945cfee9a232f"><code>c87fe57</code></a> Rebuild</li> <li>Additional commits viewable in <a href="https://github.com/github/codeql-action/compare/1c5b675653bb5c22dbe9b12b556ec555138e09fd...2892aa5e19bbd11bc0cff5427e3b750a04d9e3c2">compare view</a></li> </ul> </details> <br /> Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore <dependency name> major version` will close this group update PR and stop Dependabot creating any more for the specific dependency's major version (unless you unignore this specific dependency's major version or upgrade to it yourself) - `@dependabot ignore <dependency name> minor version` will close this group update PR and stop Dependabot creating any more for the specific dependency's minor version (unless you unignore this specific dependency's minor version or upgrade to it yourself) - `@dependabot ignore <dependency name>` will close this group update PR and stop Dependabot creating any more for the specific dependency (unless you unignore this specific dependency or upgrade to it yourself) - `@dependabot unignore <dependency name>` will remove all of the ignore conditions of the specified dependency - `@dependabot unignore <dependency name> <ignore condition>` will remove the ignore condition of the specified dependency and ignore conditions </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
## Which issue does this PR close? Related to apache#25620. Split from apache#25648 following [the review request](apache#25648 (comment)). This PR can be reviewed and merged independently. ## Rationale for this change An equality filter can narrow a column's bounds to one value without proving that any row contains that value. With input bounds [1, 100] and actual values 1 and 100, `a = 5` returns no rows, but currently reports an exact distinct count of one when its estimated row count is positive. ## What changes are included in this PR? Report `Inexact(1)` for a singleton domain when the filtered row count is an estimate or is unknown. Preserve exact zero for an exactly empty output and exact one when an exact positive row count proves that the singleton is present. Add an execution regression for the absent value and update affected unit-test and Parquet EXPLAIN expectations. ## What is the testing strategy for this PR? Passed on this independent branch: - `cargo test --profile ci -p datafusion-physical-plan --lib` (2,364 tests) - `cargo clippy --profile ci -p datafusion-physical-plan --all-targets --all-features -- -D warnings` - `cargo test --profile ci -p datafusion-sqllogictest --test sqllogictests -- parquet_statistics.slt` - `cargo fmt --all` and `git diff --check` Benchmarks were not rerun for this split. This extracts the singleton-precision implementation already present in apache#25648. ## Are there any user-facing changes? EXPLAIN reports an estimated distinct count when a matching row is not proven. No public API changes.
## Which issue does this PR close? N/A ## Rationale for this change The introduction page invites projects using DataFusion to add a link. ## What changes are included in this PR? Adds [narwhals-datafusion](https://github.com/s5dsn-eqee/narwhals-datafusion), a plugin that runs the Narwhals dataframe API on DataFusion, to the Integrations list. ## Are these changes tested? Docs only. ## Are there any user-facing changes? No. Co-authored-by: Neil Conway <neil.conway@gmail.com>
## Which issue does this PR close? - Closes apache#25796 ## Rationale for this change `array_agg(x ORDER BY y)`, `string_agg(x, d ORDER BY y)`, `first_value`, `last_value`, and any other aggregate whose ordering is written inside the argument list, lost that ordering when unparsed. Only aggregates using `WITHIN GROUP` consumed `params.order_by`; every other aggregate emitted an empty clause list. The unparsed SQL stays valid, so it runs without error and quietly returns different results. Any tool that forwards unparsed SQL to another engine (for example `datafusion-federation`) gets wrong answers with no warning. ## What changes are included in this PR? `Expr::AggregateFunction` unparsing now decides the spelling of the ordering once and fills the matching AST slot: - ordered-set aggregates (`supports_within_group_clause()`: `percentile_cont` and `approx_percentile_cont*`) keep emitting `WITHIN GROUP (ORDER BY ..)`; - every other aggregate with a non-empty `order_by` now emits `ast::FunctionArgumentClause::OrderBy` inside the argument list; - aggregates without ordering are unchanged. The `clauses` list is keyed off whether a `WITHIN GROUP` clause is actually emitted, so the two slots can never both end up empty. ## What is the testing strategy for this PR? - `roundtrip_ordered_aggregate_order_by` round-trips `last_value`, `first_value`, `array_agg` and `string_agg` from the issue, plus `DISTINCT` and a multi-key ordering. - `ordered_aggregate_order_by_survives_replanning` re-parses and re-plans the emitted SQL and asserts the plan is unchanged, so a frozen-but-wrong snapshot cannot hide a dropped ordering. It covers both spellings. Both tests fail before the change and pass after it. ## Are there any user-facing changes? Unparsed SQL now preserves the ordering of ordered aggregates instead of silently dropping it. No API change. ## AI assistance This change was prepared with AI assistance. Per the project's AI-assisted contributions policy, the following assumptions are called out explicitly so reviewers can check them rather than take them on trust: - The `clauses` slot is keyed off whether a `WITHIN GROUP` clause is actually emitted, rather than off `supports_within_group_clause()` alone. A function that reports within-group support but reaches the argument-list path therefore cannot end up with both slots empty. - `f(x ORDER BY y)` is emitted unconditionally, with no dialect hook. MySQL cannot express these functions anyway (`ARRAY_AGG` is `JSON_ARRAYAGG` there), so nothing that previously worked regresses, but a dialect guard is deliberately not part of this PR. - `null_treatment` at the same site remains hardcoded to `None`, so `IGNORE NULLS` is still dropped (see apache#25462). Kept out of scope rather than mixing two fixes into one PR. - sqlparser parses argument-list clauses in a fixed order (`WHERE` -> `IGNORE|RESPECT NULLS` -> `ORDER BY` -> `LIMIT`). Emitting only `ORDER BY` is safe today, but that order must be respected if `null_treatment` is emitted later. - The re-planning test compares `display_indent()` output rather than structural plan equality. It catches a dropped or relocated ordering but is not a general plan-equivalence check. --------- Co-authored-by: Neil Conway <neil.conway@gmail.com>
…fusion/wasmtest/datafusion-wasm-app (apache#25879) Bumps [webpack-dev-middleware](https://github.com/webpack/webpack-dev-middleware) from 8.0.4 to 8.3.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/webpack/webpack-dev-middleware/releases">webpack-dev-middleware's releases</a>.</em></p> <blockquote> <h2>v8.3.0</h2> <h3>Minor Changes</h3> <ul> <li> <p>Added a <code>hot</code> option that enables hot module replacement, replacing the need for <code>webpack-hot-middleware</code>. Pass <code>hot: true</code> to enable with defaults, or <code>hot: { path, heartbeat, progress, statsOptions }</code> to customize. The client runtime is served by the middleware itself. (by <a href="https://github.com/bjohansebas"><code>@bjohansebas</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2370">#2370</a>)</p> </li> <li> <p>Take the diagnostics a hot payload carries from the <code>stats</code> option, so one setting governs what a build reports in the terminal and in the browser: <code>stats: "errors-only"</code> keeps warnings out of both, and <code>stats: false</code> keeps errors and warnings out of both, the client's error overlay included — reach for the client's <code>?logging=</code> or <code>?overlay=</code> to quiet the browser alone. <code>hot.statsOptions</code> is deprecated and will be removed in the next major release; its <code>hash</code>, <code>timings</code> and <code>children</code> keys are now ignored, because they could leave a payload without the hash the client compares, or carry a child compilation's hash instead, which stopped updates applying and forced a full page reload on every rebuild. (by <a href="https://github.com/alexander-akait"><code>@alexander-akait</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2392">#2392</a>)</p> </li> </ul> <h3>Patch Changes</h3> <ul> <li> <p>Fixed a crash when calling <code>invalidate()</code> in plugin mode (<code>isPlugin = true</code>). Since the host (webpack-cli, webpack-dev-server, etc.) owns <code>compiler.watch()</code>, the middleware now invalidates the host's <code>watching</code> instead (each child compiler's one for a <code>MultiCompiler</code> on webpack < 5.109). When nothing is watching it logs a warning and completes the callback, as <code>close()</code> does, rather than leaving <code>invalidate(callback)</code> waiting on a build that never runs. (by <a href="https://github.com/bjohansebas"><code>@bjohansebas</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2378">#2378</a>)</p> </li> <li> <p>Reject with <code>403 Forbidden</code> the requests whose resolved filename falls outside <code>outputPath</code> (<a href="https://github.com/webpack/webpack-dev-middleware/security/advisories/GHSA-g84c-rxfj-3j2c">GHSA-g84c-rxfj-3j2c</a>). With a <code>publicPath</code> without a trailing slash, a sibling path sharing its prefix (<code>/assets../secret</code>) escaped the output root once the prefix was stripped and joined. (by <a href="https://github.com/bjohansebas"><code>@bjohansebas</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2404">#2404</a>)</p> </li> <li> <p>Update the changelog generator to the <code>@changesets/get-github-info</code> 1.0 API. (by <a href="https://github.com/alexander-akait"><code>@alexander-akait</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2396">#2396</a>)</p> </li> <li> <p>Update dependencies. (by <a href="https://github.com/alexander-akait"><code>@alexander-akait</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2394">#2394</a>)</p> </li> </ul> <h2>v8.2.0</h2> <h3>Minor Changes</h3> <ul> <li>Added a <code>hot</code> option that enables hot module replacement, replacing the need for <code>webpack-hot-middleware</code>. Pass <code>hot: true</code> to enable with defaults, or <code>hot: { path, heartbeat, progress, statsOptions }</code> to customize. The client runtime ships with the package and is added as a webpack entry. (by <a href="https://github.com/bjohansebas"><code>@bjohansebas</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2322">#2322</a>)</li> </ul> <h2>v8.1.1</h2> <h3>Patch Changes</h3> <ul> <li>Fixed a crash when calling <code>close()</code> in plugin mode (<code>isPlugin = true</code>). Since the host (webpack-cli, webpack-dev-server, etc.) owns <code>compiler.watch()</code>, the middleware has no <code>watching</code> of its own to close, so <code>close()</code> now just calls the callback instead of throwing. (by <a href="https://github.com/bjohansebas"><code>@bjohansebas</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2347">#2347</a>)</li> </ul> <h2>v8.1.0</h2> <h3>Minor Changes</h3> <ul> <li>Reuse an already active <code>MultiCompiler</code> watching session instead of starting a duplicate one (requires webpack >= 5.109). (by <a href="https://github.com/bjohansebas"><code>@bjohansebas</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2371">#2371</a>)</li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/webpack/webpack-dev-middleware/blob/main/CHANGELOG.md">webpack-dev-middleware's changelog</a>.</em></p> <blockquote> <h2>8.3.0</h2> <h3>Minor Changes</h3> <ul> <li> <p>Added a <code>hot</code> option that enables hot module replacement, replacing the need for <code>webpack-hot-middleware</code>. Pass <code>hot: true</code> to enable with defaults, or <code>hot: { path, heartbeat, progress, statsOptions }</code> to customize. The client runtime is served by the middleware itself. (by <a href="https://github.com/bjohansebas"><code>@bjohansebas</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2370">#2370</a>)</p> </li> <li> <p>Take the diagnostics a hot payload carries from the <code>stats</code> option, so one setting governs what a build reports in the terminal and in the browser: <code>stats: "errors-only"</code> keeps warnings out of both, and <code>stats: false</code> keeps errors and warnings out of both, the client's error overlay included — reach for the client's <code>?logging=</code> or <code>?overlay=</code> to quiet the browser alone. <code>hot.statsOptions</code> is deprecated and will be removed in the next major release; its <code>hash</code>, <code>timings</code> and <code>children</code> keys are now ignored, because they could leave a payload without the hash the client compares, or carry a child compilation's hash instead, which stopped updates applying and forced a full page reload on every rebuild. (by <a href="https://github.com/alexander-akait"><code>@alexander-akait</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2392">#2392</a>)</p> </li> </ul> <h3>Patch Changes</h3> <ul> <li> <p>Fixed a crash when calling <code>invalidate()</code> in plugin mode (<code>isPlugin = true</code>). Since the host (webpack-cli, webpack-dev-server, etc.) owns <code>compiler.watch()</code>, the middleware now invalidates the host's <code>watching</code> instead (each child compiler's one for a <code>MultiCompiler</code> on webpack < 5.109). When nothing is watching it logs a warning and completes the callback, as <code>close()</code> does, rather than leaving <code>invalidate(callback)</code> waiting on a build that never runs. (by <a href="https://github.com/bjohansebas"><code>@bjohansebas</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2378">#2378</a>)</p> </li> <li> <p>Reject with <code>403 Forbidden</code> the requests whose resolved filename falls outside <code>outputPath</code> (<a href="https://github.com/webpack/webpack-dev-middleware/security/advisories/GHSA-g84c-rxfj-3j2c">GHSA-g84c-rxfj-3j2c</a>). With a <code>publicPath</code> without a trailing slash, a sibling path sharing its prefix (<code>/assets../secret</code>) escaped the output root once the prefix was stripped and joined. (by <a href="https://github.com/bjohansebas"><code>@bjohansebas</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2404">#2404</a>)</p> </li> <li> <p>Update the changelog generator to the <code>@changesets/get-github-info</code> 1.0 API. (by <a href="https://github.com/alexander-akait"><code>@alexander-akait</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2396">#2396</a>)</p> </li> <li> <p>Update dependencies. (by <a href="https://github.com/alexander-akait"><code>@alexander-akait</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2394">#2394</a>)</p> </li> </ul> <h2>8.2.0</h2> <h3>Minor Changes</h3> <ul> <li>Added a <code>hot</code> option that enables hot module replacement, replacing the need for <code>webpack-hot-middleware</code>. Pass <code>hot: true</code> to enable with defaults, or <code>hot: { path, heartbeat, progress, statsOptions }</code> to customize. The client runtime ships with the package and is added as a webpack entry. (by <a href="https://github.com/bjohansebas"><code>@bjohansebas</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2322">#2322</a>)</li> </ul> <h2>8.1.1</h2> <h3>Patch Changes</h3> <ul> <li>Fixed a crash when calling <code>close()</code> in plugin mode (<code>isPlugin = true</code>). Since the host (webpack-cli, webpack-dev-server, etc.) owns <code>compiler.watch()</code>, the middleware has no <code>watching</code> of its own to close, so <code>close()</code> now just calls the callback instead of throwing. (by <a href="https://github.com/bjohansebas"><code>@bjohansebas</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2347">#2347</a>)</li> </ul> <h2>8.1.0</h2> <h3>Minor Changes</h3> <ul> <li>Reuse an already active <code>MultiCompiler</code> watching session instead of starting a duplicate one (requires webpack >= 5.109). (by <a href="https://github.com/bjohansebas"><code>@bjohansebas</code></a> in <a href="https://redirect.github.com/webpack/webpack-dev-middleware/pull/2371">#2371</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/webpack/webpack-dev-middleware/commit/1abbf0898f1ca6cb1f8cc42b7469c7c0c45a5787"><code>1abbf08</code></a> chore(release): new release (<a href="https://redirect.github.com/webpack/webpack-dev-middleware/issues/2393">#2393</a>)</li> <li><a href="https://github.com/webpack/webpack-dev-middleware/commit/ef479c9e4d35fa1fda3aff09b51bded65025e065"><code>ef479c9</code></a> feat: add missing changelog (<a href="https://redirect.github.com/webpack/webpack-dev-middleware/issues/2404">#2404</a>)</li> <li><a href="https://github.com/webpack/webpack-dev-middleware/commit/13344e78b26d88fe9b7c8381a3860e248a52359b"><code>13344e7</code></a> Merge commit from fork</li> <li><a href="https://github.com/webpack/webpack-dev-middleware/commit/e58a70ba9c36f6acc10617f338b73269f951de35"><code>e58a70b</code></a> test(hot): reach 100% of the client, off the deprecated statsOptions (<a href="https://redirect.github.com/webpack/webpack-dev-middleware/issues/2402">#2402</a>)</li> <li><a href="https://github.com/webpack/webpack-dev-middleware/commit/17e689e3290f6d0add35443ca049ea9b906d001b"><code>17e689e</code></a> chore(deps): bump the dependencies group across 1 directory with 2 updates (#...</li> <li><a href="https://github.com/webpack/webpack-dev-middleware/commit/d81681815b998fc4a13260c9462ec8ecdfdabf93"><code>d816818</code></a> fix(plugin): resolve crash on invalidate() in plugin mode (<a href="https://redirect.github.com/webpack/webpack-dev-middleware/issues/2378">#2378</a>)</li> <li><a href="https://github.com/webpack/webpack-dev-middleware/commit/09a9095d9e88cf727626423233659c366b984940"><code>09a9095</code></a> chore(deps): bump changesets/action (<a href="https://redirect.github.com/webpack/webpack-dev-middleware/issues/2384">#2384</a>)</li> <li><a href="https://github.com/webpack/webpack-dev-middleware/commit/4c1dcfcf02ca827d8802fda395e393fa154cbcd1"><code>4c1dcfc</code></a> test: rewrite to e2e (<a href="https://redirect.github.com/webpack/webpack-dev-middleware/issues/2370">#2370</a>)</li> <li><a href="https://github.com/webpack/webpack-dev-middleware/commit/9133f47afea2673d1e2097263ecfad11e4929c90"><code>9133f47</code></a> build: update changesets tooling (<a href="https://redirect.github.com/webpack/webpack-dev-middleware/issues/2396">#2396</a>)</li> <li><a href="https://github.com/webpack/webpack-dev-middleware/commit/a337174ccb37d33c2559d11c13929c445a29da80"><code>a337174</code></a> build: update dependencies (<a href="https://redirect.github.com/webpack/webpack-dev-middleware/issues/2394">#2394</a>)</li> <li>Additional commits viewable in <a href="https://github.com/webpack/webpack-dev-middleware/compare/v8.0.4...v8.3.0">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) You can disable automated security fix PRs for this repo from the [Security Alerts page](https://github.com/apache/datafusion/network/alerts). </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…robin (apache#25809) ## Which issue does this PR close? - No issue. Found while benchmarking apache#24086. ## Rationale for this change With `SIMULATE_LATENCY=true`, a benchmark result can change when a PR changes the number of object store requests in *other* queries, or off the critical path of the same query. `LatencyObjectStore` gives each request the next entry of a 20-entry latency table, with one counter for the whole process: ```rust let idx = self.get_counter.fetch_add(1, Ordering::Relaxed) % GET_LATENCIES_MS.len(); ``` So the latency of a request depends on how many requests came before it, in all earlier queries of the run. The simulator is deterministic, so the same wrong result comes back on every rerun. Example from apache#24086 (TPC-DS SF1, `SIMULATE_LATENCY=true`): the bot reported Q27 1.55x slower (141 → 218 ms) and Q44 1.38x slower. Both queries have only two serial round trips on the critical path. Q44 makes identical requests with and without the change. Local runs, median of 9, ms: | latency model | Q27 base | Q27 PR | Q44 base | Q44 PR | |---|---|---|---|---| | fixed 50 ms per request | 111 | 111 | 114 | 121 | | fixed 100 ms per request | 216 | 211 | – | – | | current round-robin table | 238 | 288 | 212 | 211 | A sweep over the 20 possible start positions of the counter (100 iterations each) gives Q27 medians of 239 ms (base) and 241 ms (PR). The min-of-5 that the bot reports depends on the start position: the base build gets a lucky minimum at 11 of 20 positions, the PR build at 5. ## What changes are included in this PR? `LatencyObjectStore` draws each GET and LIST latency independently at random from the same 20-entry distributions. The distributions do not change. | approach | latency of a request depends on | result | |---|---|---| | shared counter (before) | the number of earlier requests in the process | a change in one query moves latencies in other queries | | hash of the request (not chosen) | the request's path and byte range | a request shape gets the same latency in every iteration, so builds with different request shapes do not average out | | independent random draw (this PR) | nothing | each request samples the same distribution, and iterations average out | ## What is the testing strategy for this PR? This change affects only the benchmark harness. - With this change, Q27 over 99 iterations gives base median 217 ms and PR median 224 ms (mean 223.5 ± 7.5 vs 217.6 ± 6.9), so there is no difference, as with a fixed latency. - Queries that do read data keep their differences: TPC-DS Q3 703 → 353 ms, Q7 1078 → 526 ms, TPC-H Q1 1006 → 209 ms, Q6 1862 → 218 ms (median of 5, same two builds). - `cargo test -p datafusion-benchmarks --lib` and clippy pass. There is no new unit test: the property ("no dependency on earlier requests") is structural, since the store has no state left. ## Are there any user-facing changes? No. Benchmark runs with `SIMULATE_LATENCY=true` are no longer bit-for-bit reproducible between runs. Queries with few serial round trips have more spread per iteration (Q27 SD about 70 ms), so compare them with more iterations or with means. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…ections (apache#25338) ## Which issue does this PR close? - Closes apache#25341. This PR supersedes apache#21363 by @crm26. It reuses the single mark join approach from that PR, and @crm26 is a co-author of the main commit. It also builds on apache#24972, which added the decorrelation of `IN` subqueries in a projection. ## Rationale for this change An `IN` or `NOT IN` subquery in a SELECT list gives correct results today, but the plan is quadratic. The optimizer builds three mark joins for each subquery. Two of them have no join predicate, so they run as nested loop joins over all outer rows and all inner rows. `COALESCE` duplicates the expression, so that shape gets six mark joins. This PR keeps one hash mark join for each subquery. The join is null-aware when a key can be NULL. A non-equality correlation (Q6 below) stays a residual filter of that join, which the null-aware hash join applies since apache#25339. | Shape | Query | main | this PR | Result rows identical | | --- | --- | --- | --- | --- | | Q1 | Bare `IN` in the SELECT list | 12.893 s | 0.005 s | yes | | Q2 | `COALESCE((x IN (...))::boolean, false)` | 36.430 s | 0.014 s | yes | | Q3 | Correlated `IN` with an equality predicate | 0.278 s | 0.008 s | yes | | Q4 | Two `IN` subqueries in separate columns | 20.904 s | 0.011 s | yes | | Q5 | Correlated `EXISTS` | 0.005 s | 0.003 s | yes | | Q6 | Correlated `IN` with a non-equality predicate | 2.489 s | 0.018 s | yes | The table is one run of the script below after the rebase onto `main` at `39ca2c74a9`, `datafusion-cli` built with `--profile release-nonlto` and the default `target_partitions`. Q6 in 3 more interleaved runs: `main` 2.57 s to 3.18 s, this PR 0.022 s to 0.058 s. These shapes are the `projection_subquery` benchmark suite from apache#25346. <details> <summary>Reproduction with datafusion-cli: script, plans and timings for each shape</summary> Run the script with `datafusion-cli -f mre.sql`. It creates two tables with 200000 rows each, then for each shape it prints `EXPLAIN` and runs the query. `datafusion-cli` prints the elapsed time after each statement. The result rows of every query are identical between the two builds. The plans and times for Q1 to Q5 below are from the first version of this PR, with `main` at 22651d2 and `target_partitions = 4`. The plans of this PR for those shapes did not change with the rebase. Q6 was captured again after the rebase. ```sql SET datafusion.execution.target_partitions = 4; CREATE TABLE outer_t AS SELECT CAST(v AS INT) AS id, CAST(v % 1000 AS INT) AS z FROM (SELECT unnest(generate_series(1, 200000)) AS v); CREATE TABLE inner_t AS SELECT CASE WHEN v % 97 = 0 THEN NULL ELSE CAST(v * 2 AS INT) END AS id, CAST(v % 1000 AS INT) AS z FROM (SELECT unnest(generate_series(1, 200000)) AS v); SELECT 'Q1' AS shape; EXPLAIN SELECT id, id IN (SELECT id FROM inner_t) AS m FROM outer_t; SELECT count(*) FILTER (WHERE m), count(*) FILTER (WHERE m IS NULL) FROM (SELECT id, id IN (SELECT id FROM inner_t) AS m FROM outer_t); SELECT 'Q2' AS shape; EXPLAIN SELECT id, COALESCE((id IN (SELECT id FROM inner_t))::boolean, false) AS m FROM outer_t; SELECT count(*) FILTER (WHERE m) FROM (SELECT id, COALESCE((id IN (SELECT id FROM inner_t))::boolean, false) AS m FROM outer_t); SELECT 'Q3' AS shape; EXPLAIN SELECT id, id IN (SELECT i.id FROM inner_t i WHERE i.z = o.z) AS m FROM outer_t o; SELECT count(*) FILTER (WHERE m), count(*) FILTER (WHERE m IS NULL) FROM (SELECT id, id IN (SELECT i.id FROM inner_t i WHERE i.z = o.z) AS m FROM outer_t o); SELECT 'Q4' AS shape; EXPLAIN SELECT id IN (SELECT id FROM inner_t) AS a, id IN (SELECT id FROM inner_t WHERE z < 500) AS b FROM outer_t; SELECT count(*) FILTER (WHERE a), count(*) FILTER (WHERE b) FROM (SELECT id IN (SELECT id FROM inner_t) AS a, id IN (SELECT id FROM inner_t WHERE z < 500) AS b FROM outer_t); SELECT 'Q5' AS shape; EXPLAIN SELECT id, EXISTS (SELECT 1 FROM inner_t i WHERE i.id = o.id) AS e FROM outer_t o; SELECT count(*) FILTER (WHERE e) FROM (SELECT id, EXISTS (SELECT 1 FROM inner_t i WHERE i.id = o.id) AS e FROM outer_t o); SELECT 'Q6' AS shape; EXPLAIN SELECT id, id IN (SELECT i.id FROM inner_t i WHERE i.z < o.z) AS m FROM outer_t o; SELECT count(*) FILTER (WHERE m), count(*) FILTER (WHERE m IS NULL) FROM (SELECT id, id IN (SELECT i.id FROM inner_t i WHERE i.z < o.z) AS m FROM outer_t o); ``` <details> <summary>Q1: Bare `IN` in the SELECT list, plans on main and on this PR</summary> **main**, query time 190.730 s ``` logical_plan Projection: outer_t.id, __correlated_sq_1.mark IS NOT DISTINCT FROM Boolean(true) OR (__correlated_sq_2.mark OR outer_t.id IS NULL AND __correlated_sq_3.mark) IS NOT DISTINCT FROM Boolean(true) AND __correlated_sq_1.mark IS DISTINCT FROM Boolean(true) AND Boolean(NULL) AS m LeftMark Join: LeftMark Join: LeftMark Join: outer_t.id = __correlated_sq_1.id null_aware TableScan: outer_t projection=[id] SubqueryAlias: __correlated_sq_1 TableScan: inner_t projection=[id] SubqueryAlias: __correlated_sq_2 Filter: inner_t.id IS NULL TableScan: inner_t projection=[id] SubqueryAlias: __correlated_sq_3 TableScan: inner_t projection=[id] physical_plan ProjectionExec: expr=[id@0 as id, mark@1 IS NOT DISTINCT FROM true OR (mark@2 OR id@0 IS NULL AND mark@3) IS NOT DISTINCT FROM true AND mark@1 IS DISTINCT FROM true AND NULL as m] NestedLoopJoinExec: join_type=LeftMark CoalescePartitionsExec NestedLoopJoinExec: join_type=RightMark CoalescePartitionsExec FilterExec: id@0 IS NULL DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0)], null_aware CoalescePartitionsExec DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] ``` **this PR**, query time 0.034 s ``` logical_plan Projection: outer_t.id, __correlated_sq_1.mark AS m LeftMark Join: outer_t.id = __correlated_sq_1.id null_aware TableScan: outer_t projection=[id] SubqueryAlias: __correlated_sq_1 TableScan: inner_t projection=[id] physical_plan ProjectionExec: expr=[id@0 as id, mark@1 as m] HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0)], null_aware CoalescePartitionsExec DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] ``` </details> <details> <summary>Q2: `COALESCE((x IN (...))::boolean, false)`, plans on main and on this PR</summary> **main**, query time 351.579 s ``` logical_plan Projection: outer_t.id, CAST(__correlated_sq_1.mark IS NOT DISTINCT FROM Boolean(true) OR (__correlated_sq_2.mark OR outer_t.id IS NULL AND __correlated_sq_3.mark) IS NOT DISTINCT FROM Boolean(true) AND __correlated_sq_1.mark IS DISTINCT FROM Boolean(true) AND Boolean(NULL) AS Boolean) IS NOT NULL AND CAST(__correlated_sq_4.mark IS NOT DISTINCT FROM Boolean(true) OR (__correlated_sq_5.mark OR outer_t.id IS NULL AND __correlated_sq_6.mark) IS NOT DISTINCT FROM Boolean(true) AND __correlated_sq_4.mark IS DISTINCT FROM Boolean(true) AND Boolean(NULL) AS Boolean) AS m LeftMark Join: LeftMark Join: LeftMark Join: outer_t.id = __correlated_sq_4.id null_aware LeftMark Join: LeftMark Join: LeftMark Join: outer_t.id = __correlated_sq_1.id null_aware TableScan: outer_t projection=[id] SubqueryAlias: __correlated_sq_1 TableScan: inner_t projection=[id] SubqueryAlias: __correlated_sq_2 Filter: inner_t.id IS NULL TableScan: inner_t projection=[id] SubqueryAlias: __correlated_sq_3 TableScan: inner_t projection=[id] SubqueryAlias: __correlated_sq_4 TableScan: inner_t projection=[id] SubqueryAlias: __correlated_sq_5 Filter: inner_t.id IS NULL TableScan: inner_t projection=[id] SubqueryAlias: __correlated_sq_6 TableScan: inner_t projection=[id] physical_plan ProjectionExec: expr=[id@0 as id, mark@1 IS NOT DISTINCT FROM true OR (mark@2 OR id@0 IS NULL AND mark@3) IS NOT DISTINCT FROM true AND mark@1 IS DISTINCT FROM true AND NULL IS NOT NULL AND (mark@4 IS NOT DISTINCT FROM true OR (mark@5 OR id@0 IS NULL AND mark@6) IS NOT DISTINCT FROM true AND mark@4 IS DISTINCT FROM true AND NULL) as m] NestedLoopJoinExec: join_type=LeftMark CoalescePartitionsExec NestedLoopJoinExec: join_type=RightMark CoalescePartitionsExec FilterExec: id@0 IS NULL DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0)], null_aware CoalescePartitionsExec NestedLoopJoinExec: join_type=LeftMark CoalescePartitionsExec NestedLoopJoinExec: join_type=RightMark CoalescePartitionsExec FilterExec: id@0 IS NULL DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0)], null_aware CoalescePartitionsExec DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] ``` **this PR**, query time 0.054 s ``` logical_plan Projection: outer_t.id, CAST(__correlated_sq_1.mark AS Boolean) IS NOT NULL AND CAST(__correlated_sq_2.mark AS Boolean) AS m LeftMark Join: outer_t.id = __correlated_sq_2.id null_aware LeftMark Join: outer_t.id = __correlated_sq_1.id null_aware TableScan: outer_t projection=[id] SubqueryAlias: __correlated_sq_1 TableScan: inner_t projection=[id] SubqueryAlias: __correlated_sq_2 TableScan: inner_t projection=[id] physical_plan ProjectionExec: expr=[id@0 as id, mark@1 IS NOT NULL AND mark@2 as m] HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0)], null_aware CoalescePartitionsExec HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0)], null_aware CoalescePartitionsExec DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] ``` </details> <details> <summary>Q3: Correlated `IN` with an equality predicate, plans on main and on this PR</summary> **main**, query time 1.372 s ``` logical_plan Projection: o.id, __correlated_sq_1.mark IS NOT DISTINCT FROM Boolean(true) OR (__correlated_sq_2.mark OR o.id IS NULL AND __correlated_sq_3.mark) IS NOT DISTINCT FROM Boolean(true) AND __correlated_sq_1.mark IS DISTINCT FROM Boolean(true) AND Boolean(NULL) AS m LeftMark Join: o.z = __correlated_sq_3.z LeftMark Join: o.z = __correlated_sq_2.z LeftMark Join: o.id = __correlated_sq_1.id, o.z = __correlated_sq_1.z null_aware SubqueryAlias: o TableScan: outer_t projection=[id, z] SubqueryAlias: __correlated_sq_1 SubqueryAlias: i TableScan: inner_t projection=[id, z] SubqueryAlias: __correlated_sq_2 SubqueryAlias: i Projection: inner_t.z Filter: inner_t.id IS NULL TableScan: inner_t projection=[id, z] SubqueryAlias: __correlated_sq_3 SubqueryAlias: i TableScan: inner_t projection=[z] physical_plan ProjectionExec: expr=[id@0 as id, mark@1 IS NOT DISTINCT FROM true OR (mark@2 OR id@0 IS NULL AND mark@3) IS NOT DISTINCT FROM true AND mark@1 IS DISTINCT FROM true AND NULL as m] HashJoinExec: mode=CollectLeft, join_type=RightMark, on=[(z@0, z@1)], projection=[id@0, mark@2, mark@3, mark@4] CoalescePartitionsExec DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] HashJoinExec: mode=CollectLeft, join_type=RightMark, on=[(z@0, z@1)] CoalescePartitionsExec FilterExec: id@0 IS NULL, projection=[z@1] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0), (z@1, z@1)], null_aware CoalescePartitionsExec DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] ``` **this PR**, query time 0.091 s ``` logical_plan Projection: o.id, __correlated_sq_1.mark AS m LeftMark Join: o.id = __correlated_sq_1.id, o.z = __correlated_sq_1.z null_aware SubqueryAlias: o TableScan: outer_t projection=[id, z] SubqueryAlias: __correlated_sq_1 SubqueryAlias: i TableScan: inner_t projection=[id, z] physical_plan ProjectionExec: expr=[id@0 as id, mark@1 as m] HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0), (z@1, z@1)], projection=[id@0, mark@2], null_aware CoalescePartitionsExec DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] ``` </details> <details> <summary>Q4: Two `IN` subqueries in separate columns, plans on main and on this PR</summary> **main**, query time 296.676 s ``` logical_plan Projection: __correlated_sq_1.mark IS NOT DISTINCT FROM Boolean(true) OR (__correlated_sq_2.mark OR outer_t.id IS NULL AND __correlated_sq_3.mark) IS NOT DISTINCT FROM Boolean(true) AND __correlated_sq_1.mark IS DISTINCT FROM Boolean(true) AND Boolean(NULL) AS a, __correlated_sq_4.mark IS NOT DISTINCT FROM Boolean(true) OR (__correlated_sq_5.mark OR outer_t.id IS NULL AND __correlated_sq_6.mark) IS NOT DISTINCT FROM Boolean(true) AND __correlated_sq_4.mark IS DISTINCT FROM Boolean(true) AND Boolean(NULL) AS b LeftMark Join: LeftMark Join: LeftMark Join: outer_t.id = __correlated_sq_4.id null_aware LeftMark Join: LeftMark Join: LeftMark Join: outer_t.id = __correlated_sq_1.id null_aware TableScan: outer_t projection=[id] SubqueryAlias: __correlated_sq_1 TableScan: inner_t projection=[id] SubqueryAlias: __correlated_sq_2 Filter: inner_t.id IS NULL TableScan: inner_t projection=[id] SubqueryAlias: __correlated_sq_3 TableScan: inner_t projection=[id] SubqueryAlias: __correlated_sq_4 Projection: inner_t.id Filter: inner_t.z < Int32(500) TableScan: inner_t projection=[id, z] SubqueryAlias: __correlated_sq_5 Projection: inner_t.id Filter: inner_t.z < Int32(500) AND inner_t.id IS NULL TableScan: inner_t projection=[id, z] SubqueryAlias: __correlated_sq_6 Projection: inner_t.id Filter: inner_t.z < Int32(500) TableScan: inner_t projection=[id, z] physical_plan ProjectionExec: expr=[mark@1 IS NOT DISTINCT FROM true OR (mark@2 OR id@0 IS NULL AND mark@3) IS NOT DISTINCT FROM true AND mark@1 IS DISTINCT FROM true AND NULL as a, mark@4 IS NOT DISTINCT FROM true OR (mark@5 OR id@0 IS NULL AND mark@6) IS NOT DISTINCT FROM true AND mark@4 IS DISTINCT FROM true AND NULL as b] NestedLoopJoinExec: join_type=LeftMark CoalescePartitionsExec NestedLoopJoinExec: join_type=RightMark CoalescePartitionsExec FilterExec: z@1 < 500 AND id@0 IS NULL, projection=[id@0] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0)], null_aware CoalescePartitionsExec NestedLoopJoinExec: join_type=LeftMark CoalescePartitionsExec NestedLoopJoinExec: join_type=RightMark CoalescePartitionsExec FilterExec: id@0 IS NULL DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0)], null_aware CoalescePartitionsExec DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] FilterExec: z@1 < 500, projection=[id@0] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] FilterExec: z@1 < 500, projection=[id@0] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] ``` **this PR**, query time 0.063 s ``` logical_plan Projection: __correlated_sq_1.mark AS a, __correlated_sq_2.mark AS b LeftMark Join: outer_t.id = __correlated_sq_2.id null_aware LeftMark Join: outer_t.id = __correlated_sq_1.id null_aware TableScan: outer_t projection=[id] SubqueryAlias: __correlated_sq_1 TableScan: inner_t projection=[id] SubqueryAlias: __correlated_sq_2 Projection: inner_t.id Filter: inner_t.z < Int32(500) TableScan: inner_t projection=[id, z] physical_plan ProjectionExec: expr=[mark@0 as a, mark@1 as b] HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0)], projection=[mark@1, mark@2], null_aware CoalescePartitionsExec HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0)], null_aware CoalescePartitionsExec DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] FilterExec: z@1 < 500, projection=[id@0] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] ``` </details> <details> <summary>Q5: Correlated `EXISTS`, plans on main and on this PR</summary> **main**, query time 0.021 s ``` logical_plan Projection: o.id, __correlated_sq_1.mark AS e LeftMark Join: o.id = __correlated_sq_1.id SubqueryAlias: o TableScan: outer_t projection=[id] SubqueryAlias: __correlated_sq_1 SubqueryAlias: i TableScan: inner_t projection=[id] physical_plan ProjectionExec: expr=[id@0 as id, mark@1 as e] HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0)] CoalescePartitionsExec DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] ``` **this PR**, query time 0.021 s ``` logical_plan Projection: o.id, __correlated_sq_1.mark AS e LeftMark Join: o.id = __correlated_sq_1.id SubqueryAlias: o TableScan: outer_t projection=[id] SubqueryAlias: __correlated_sq_1 SubqueryAlias: i TableScan: inner_t projection=[id] physical_plan ProjectionExec: expr=[id@0 as id, mark@1 as e] HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0)] CoalescePartitionsExec DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] DataSourceExec: partitions=4, partition_sizes=[7, 6, 6, 6] ``` </details> <details> <summary>Q6: Correlated `IN` with a non-equality predicate, plans on main and on this PR</summary> **main** (`39ca2c74a9`), query time 2.489 s ``` logical_plan Projection: o.id, CASE WHEN __correlated_sq_1.mark THEN Boolean(true) WHEN __correlated_sq_2.mark OR o.id IS NULL AND __correlated_sq_3.mark THEN Boolean(NULL) ELSE Boolean(false) END AS m LeftMark Join: Filter: __correlated_sq_3.z < o.z LeftMark Join: Filter: __correlated_sq_2.z < o.z LeftMark Join: o.id = __correlated_sq_1.id Filter: __correlated_sq_1.z < o.z null_aware SubqueryAlias: o TableScan: outer_t projection=[id, z] SubqueryAlias: __correlated_sq_1 SubqueryAlias: i TableScan: inner_t projection=[id, z] SubqueryAlias: __correlated_sq_2 SubqueryAlias: i Projection: inner_t.z Filter: inner_t.id IS NULL TableScan: inner_t projection=[id, z] SubqueryAlias: __correlated_sq_3 SubqueryAlias: i TableScan: inner_t projection=[z] physical_plan ProjectionExec: expr=[id@0 as id, CASE WHEN mark@1 THEN true WHEN mark@2 OR id@0 IS NULL AND mark@3 THEN NULL ELSE false END as m] NestedLoopJoinExec: join_type=LeftMark, filter=z@1 < z@0, projection=[id@0, mark@2, mark@3, mark@4] CoalescePartitionsExec NestedLoopJoinExec: join_type=RightMark, filter=z@1 < z@0 CoalescePartitionsExec FilterExec: id@0 IS NULL, projection=[z@1] DataSourceExec: partitions=12, partition_sizes=[3, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2] HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0)], filter=z@1 < z@0, null_aware CoalescePartitionsExec DataSourceExec: partitions=12, partition_sizes=[3, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2] DataSourceExec: partitions=12, partition_sizes=[3, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2] DataSourceExec: partitions=12, partition_sizes=[3, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2] ``` **this PR**, query time 0.018 s ``` logical_plan Projection: o.id, __correlated_sq_1.mark AS m LeftMark Join: o.id = __correlated_sq_1.id Filter: __correlated_sq_1.z < o.z null_aware SubqueryAlias: o TableScan: outer_t projection=[id, z] SubqueryAlias: __correlated_sq_1 SubqueryAlias: i TableScan: inner_t projection=[id, z] physical_plan ProjectionExec: expr=[id@0 as id, mark@1 as m] HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0)], filter=z@1 < z@0, projection=[id@0, mark@2], null_aware CoalescePartitionsExec DataSourceExec: partitions=12, partition_sizes=[3, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2] DataSourceExec: partitions=12, partition_sizes=[3, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2] ``` </details> </details> ## What changes are included in this PR? 1. `build_join` now reports whether the mark column of the `LeftMark` join is exact under three-valued logic. It is exact when no side of the `IN` predicate can be NULL in scope, or when the join is null-aware and the `IN` equality is its first key. A residual non-equality filter does not change this, because the null-aware hash join applies it when it decides whether a NULL makes the mark UNKNOWN (apache#25339). `in_subquery_value_mark_join` builds this join first. If the mark is exact, it returns the mark column alone, or `NOT mark` for `NOT IN`. The three join materialization stays for the other cases. 2. The rule decides null-awareness from one question: can the `IN` value or the subquery output be NULL inside the scope of an outer row? Only then can `IN` be UNKNOWN. - The test uses the key expressions, not the columns they reference. A key such as `NULLIF(id, 1)` or `TRY_CAST(s AS INT)` can be NULL over a `NOT NULL` column. - `PullUpCorrelatedExpr` now records every correlated conjunct, before it drops the conjunct that repeats the `IN` predicate. A conjunct such as `y = x`, `y > x` or `y IS NOT NULL` is never TRUE for a NULL `y`, so it keeps a NULL `y` out of the scope. The conjuncts keep their outer references, so a subquery column with the same qualified name as an outer column is not mistaken for the outer value. - A correlation key that is NULL only empties the scope, so it never makes the join null-aware. - The recorded conjuncts are cleared when the pull up passes an outer join, a union or a grouping set, because such a node can put a NULL back into the column. With `ROLLUP`, the grand-total row is a NULL for every outer row. 3. A `NOT IN` filter uses the same scope-aware nullability for its null-aware `LeftAnti` join. If the `IN` equality does not become the first key of that join, the rule gives up and the `NOT IN` becomes the mark joins, which give the correct result. 4. The regression test for apache#24574 in `projection_pushdown.slt` now uses a correlated subquery with `LIMIT 1`. The subquery then still reaches `ExtractLeafExpressions`, and the alias generator still starts at 2 inside it. The earlier version of the test let the subquery be flattened, so it no longer exercised the scan inside a subquery path. 5. New sqllogictest cases in `subquery_projection.slt`: - `EXPLAIN` guards: one hash mark join for each subquery, one null-aware mark join with a residual filter, a null-aware `LeftAnti` join with two keys, and with one key and a residual filter. - NULL semantics: `IN` and `NOT IN` with a NULL in the subquery, a NULL outer value, `NOT` inside `CASE`, a projection over an aggregate, the `COALESCE` and cast shape, and the residual filter shape. - Nullable key expressions over `NOT NULL` columns: a `NULLIF` value, a `TRY_CAST` value and a `NULLIF` subquery output, in a projection and in a `WHERE` clause. - A correlation that repeats the `IN` predicate, and a subquery column with the same qualified name as the outer value. - The same correlation below a `ROLLUP`, as a projected `IN` with an `EXPLAIN` guard and as a `NOT IN` filter. The expected results agree with DuckDB 1.5.2, and the earlier cases also with PostgreSQL. 6. The `projection_subquery` benchmark docs no longer describe q07 as the nested-loop control. ## DuckDB comparison For an absolute reference, the same seven queries in `datafusion-cli` and in DuckDB 1.5.2, on the same two tables, release build, Apple M4 Pro. `main` is the merge-base `64871d9` and "this PR" is `3af87370c1`. The later commits of this PR change the correlated and `NOT IN` filter paths, and, after the rebase onto apache#25339, the plan of q07. q07 is measured again below. The tables are the ones the suite builds, at its default of 30,000 rows each. Each number is the median of 7 runs. One round runs one session per engine, and the rounds interleave the three engines. | Query | `main` | this PR | DuckDB | | --- | --- | --- | --- | | q01 bare `IN` | 533 ms | 2 ms | 2 ms | | q02 `COALESCE` over `IN` | 1352 ms | 3 ms | 2 ms | | q03 correlated `IN`, equality | 14 ms | 3 ms | 3 ms | | q04 two `IN` columns | 1026 ms | 3 ms | 2 ms | | q05 bare `NOT IN` | 646 ms | 1 ms | 2 ms | | q06 correlated `EXISTS` | 4 ms | 2 ms | 2 ms | | q07 correlated `IN`, non-equality | 117 ms | 93 ms | 458 ms | All three engines give the same result for all seven queries, and the counts are the ones in the suite's checked-in result files. For the four quadratic shapes, q01, q02, q04 and q05, `main` is 270x to 680x slower than DuckDB. This PR puts them at DuckDB's cost. These times are at the 1 ms resolution of both command line tools, so the exact factor is approximate; the size of the difference is not. q07 now uses one null-aware mark join with the `<` correlation as a residual filter. Measured again after the rebase, median of 7 runs at 30,000 rows: `main` at `39ca2c74a9` 78 ms, this PR 3 ms, DuckDB 404 ms. The counts are the same in all three. A second run at 100,000 rows per table gives the same picture: `main` takes 2741 ms to 6355 ms for q01, q02, q04 and q05, this PR takes 2 ms to 4 ms and DuckDB 3 ms to 4 ms. All three engines again give the same counts. <details><summary>How to run it</summary> The queries are the suite's own `benchmarks/sql_benchmarks/projection_subquery/queries/q0*.sql`, with an alias added to the derived table. The load SQL needs one change for DuckDB, because `generate_series` gives a column of that name there and a column named `value` in DataFusion: ```sql CREATE TABLE outer_t AS SELECT CAST(value AS INT) AS id, CAST(value % 1000 AS INT) AS z FROM generate_series(1, 30000) AS g(value); CREATE TABLE inner_t AS SELECT CASE WHEN value % 97 = 0 THEN NULL ELSE CAST(value * 2 AS INT) END AS id, CAST(value % 1000 AS INT) AS z FROM generate_series(1, 30000) AS g(value); ``` `datafusion-cli` reports the time of each statement as `Elapsed`, and `duckdb` reports it as `Run Time (s): real` after `.timer on`. </details> ## What is the testing strategy for this PR? - Unit tests in `datafusion/optimizer/src/decorrelate_predicate_subquery.rs`. The updated snapshots show a single mark join. New tests cover a residual filter, `NOT IN`, nullable key expressions, a correlation that repeats the `IN` predicate, and the same correlation below a `ROLLUP`. The last one fails without the clearing in item 2. - The new sqllogictest cases in `subquery_projection.slt` listed above. - The full sqllogictest suite, the optimizer crate and workspace clippy are green locally. - The benchmark numbers above. ## Are there any user-facing changes? Plans for projected `IN` and `NOT IN` subqueries are much faster. There are no changes to any public API, and no documentation change is needed. Some query results change, and each change is a correction. DuckDB 1.5.2 gives the new results: - An `IN` or `NOT IN` whose key is an expression that can be NULL over a `NOT NULL` column, such as `NULLIF(id, 1)`. For a projected `IN`, `main` returns `false` instead of `NULL`. For a `NOT IN` filter, `main` keeps rows that it must drop. The correlated form of that is apache#25347. - A correlated `NOT IN` whose correlation repeats the `IN` predicate, `x NOT IN (SELECT y FROM t WHERE y = x)`. `main` drops rows that it must keep (apache#25480). - A projected `IN` whose correlation repeats the `IN` predicate below a `ROLLUP`, `k IN (SELECT i.k FROM i WHERE i.k = o.k GROUP BY ROLLUP(i.k))`. `main` fails to plan it since apache#25529. It now gives NULL for a miss, as DuckDB does. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts: # datafusion/execution/src/config.rs
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue does this PR close?
This PR builds on apache#25887 so that should be merged first.
Rationale for this change
To be completed
What changes are included in this PR?
LogicalPlan::DescribeTablehas been made been renamed toLogicalPlan::Showand extended to support all show statements.Handling of
showstatements has been moved fromsql::statementtocore::physical_planner.The implementations themselves have been decoupled from the
information_schematables. This allowsSHOWstatements to work even if the information schema provider is not enabled.A new
common::information_schemamodule has been added to avoid string based duplication of the information_schema schema. In other words, duplicate magic strings have been removed.Are these changes tested?
Covered by SLTs
Are there any user-facing changes?
Yes
LogicalPlan::DescribeTablehas been replaced withLogicalPlan::ShowSHOWstatements is no longer tied to theinformation_schemaschema being accessible in the cataloginformation_schema.viewsno longer contains base tables