Repository navigation
Conversation
…ossible (apache#23702) ## Which issue does this PR close? N/A ## Rationale for this change `SortPreservingMergeStream` is a little complex, so add some guiding comments and make it as textbook-like as possible ## What changes are included in this PR? Added comments, reorder code While this was done this also fixed couple of bugs due to how it work: 1. leftover drain was not counted in the `elapsed_compute` 2. `limit(0)` returns 0 rows and not 1 ## Are these changes tested? Existing tests ## Are there any user-facing changes? Not API ones. `limit(0)` now returns 0 rows
## Which issue does this PR close? - Closes apache#21428. ## Rationale for this change This PR adds `BuildHasher`-based variants for `hash_utils` so callers can compute row hashes with a caller-provided hash builder instead of always using DataFusion's default `RandomState`. The main constraint is performance: `with_hashes` is a hot path, especially for string, dictionary, and nested array hashing. A previous version in apache#21429 caused measurable regressions in the default `RandomState` path, for example `large_utf8: single, no nulls` regressed from roughly `26.7us` to `36.3us`, and `large_utf8: multiple, no nulls` from roughly `112us` to `127us`. This version keeps the default path performance-oriented by avoiding a fully generic `BuildHasher` rewrite of the existing hot loops. ## What changes are included in this PR? This PR adds: - `with_hashes_with_hasher` - `create_hashes_with_hasher` - custom-hasher implementations for primitive, string, binary, byte-view, dictionary, and nested arrays - tests covering custom hashers, multi-column hashing, and dictionary equivalence The implementation intentionally uses a hybrid design: - Default `RandomState` leaf hot paths remain specialized. - Custom `BuildHasher` leaf paths live separately in `hash_utils/build_hasher.rs`. - Nested/structural logic is shared through an internal child-hashing adapter, so struct/list/map/union/run/dictionary behavior does not need to be broadly duplicated. The trade-off is that there is still some duplication for primitive/string/binary leaf loops. That duplication is intentional: those are the hottest loops, and keeping them separate prevents the existing `RandomState` path from becoming generic over `BuildHasher` or being perturbed by the custom-hasher implementation. ## Are these changes tested? Yes. --------- Co-authored-by: Dmitrii Blaginin <dmitrii@blaginin.me> Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
…t whether they keep the same ordering of the input (apache#23807) ## Which issue does this PR close? - Closes apache#23798 Related to: - apache#16217 ## Rationale for this change To be able to keep the same sorting order allowing for more optimizations Now comet own `cast` implementation can be recognized as not modifying sort order in the same cases that datafusion cast does. ## What changes are included in this PR? added `strictly_order_preserving` property to `ExprProperties` + varius other places and replaced the hard coded logic for cast (`substitute_cast_ordering`) about keeping input order with more generic approach that now any expression can implement and have the same advantage of sort elimination also marked from_unixtime as keeping ordering to show a case of this optimization ## Are these changes tested? yes ## Are there any user-facing changes? yes, breaking change, added `strictly_order_preserving` property to `ExprProperties` and to `FFI_ExprProperties` this property means that given expression `f` and 2 values from the input column `a` and `b` the following variants are kept: 1. `a.cmp(b) == f(a).cmp(f(b))` 2. nulls maps to nulls Example of satisfying expression: `cast(col_a as BIGINT)` where `col_a` is `INT` it is keeping the properties Example of not satisfying: `floor` - floor can not `array_repeat(my_col, 2)` which might look like at first glance as keeping the property as well but in fact it does not. the reason is that `array_repeat(null, 2)` will output list of 2 nulls which breaks the 2nd property that nulls must be kept as nulls how to migrate: Option 1 - keeping the old behavior (safest but least performant) set `strictly_order_preserving` to false Option 2 - Using the new optimization that this opens up: set `strictly_order_preserving` only if the expression keep both variants
## Which issue does this PR close? Part of apache#22330. ## Rationale for this change `ForeignScalarUDF` inherits the default `preserves_lex_ordering`, so producer overrides are lost across the FFI boundary. ## What changes are included in this PR? - Forward `preserves_lex_ordering` through `FFI_ScalarUDF`. - Reuse the existing placement UDF for unit and dynamic-library coverage. ## Are these changes tested? - `cargo test -p datafusion-ffi --features integration-tests` - `cargo clippy --all-targets --all-features -- -D warnings` ## Are there any user-facing changes? The FFI ABI changes. Foreign libraries must rebuild against the new DataFusion version. --------- Signed-off-by: Amogh Ramesh <ramogh2404@gmail.com>
…tonic deques (apache#23826) (apache#23827) ## Which issue does this PR close? - Closes apache#23826 ### What changes are included in this PR? This PR optimizes sliding window `MIN`/`MAX` aggregate functions using a **Sequence-Numbered Monotonic Deque** instead of a Two-Stack Queue. **Revised Design:** - We store `(sequence_number, value)` pairs in a single `VecDeque`, kept in strictly monotonic order. - **`push(val)`**: Evicts dominated elements from the back, then pushes the new value with the current `push_seq` and increments `push_seq`. - **`pop()`**: Increments `pop_seq`. If the front elements sequence number equals the old `pop_seq`, it is expired and popped. - **Benefits**: No secondary FIFO queue (lower memory) and no `clone()` overhead. ### Are these changes tested? Yes, existing tests pass. Added tests for empty-window `pop()` and duplicate-heavy scenarios. ### Are there any user-facing changes? **API Change**: `MovingMin` and `MovingMax` were changed to `pub(crate)` visibility, and their `pop()` methods now return `()` instead of returning a value. Performance is significantly improved (2x-3.5x throughput). --------- Co-authored-by: Pavan <pavan@Lakshmis-MacBook-Air.local>
## Which issue does this PR close? NA ## Rationale for this change Adds [datapress](https://docs.datap-rs.org) to the list of known users in the documentation. ## What changes are included in this PR? NA ## Are these changes tested? NA ## Are there any user-facing changes? NA
## Which issue does this PR close? - Closes apache#11748 ## Rationale for this change `EliminateGroupByConstant` removes GROUP BY expressions that are constants. If all of the GROUP BY expressions are constants, eliminating all of them results in converting a grouped aggregate into an ungrouped (global) aggregate. This changes the semantics of the query: a grouped aggregate query on an empty input returns zero rows, whereas an ungrouped aggregate query produces a single row. ## What changes are included in this PR? `EliminateGroupByConstant` now declines to eliminate constant GROUP BY expressions, if doing so would result in removing all of the grouping expressions. ## Are these changes tested? Yes. Sqllogictest reproducer for the original issue, updated unit test and optimizer SLT expectations. ## Are there any user-facing changes? Queries with all-constant GROUP BY may now return fewer (correct) rows.
…pache#23874) ## Which issue does this PR close? - Closes apache#23872 ## Rationale for this change `SlidingMinAccumulator::update_batch` skipped NULL values. This meant that if all the non-NULL values in a window frame were retracted, the window frame would not be empty but the `MovingMin` data structure by the `SlidingMinAccumulator` would not contain any values. This resulted in incorrectly returning a stale non-NULL value for a sliding `min()` over a window frame consisting of only NULL values. Along the way, optimize the min and max sliding window accumulators to make them both more efficient and more symmetric with one another. In the original coding, `SlidingMaxAccumulator` included NULL values but `SlidingMinAccumulator` omitted them, in part because omitting NULLs made the original `retract_batch` implementation more expensive. This PR optimizes `retract_batch`, so we can now use the same scheme for both the min and max sliding accumulators: * Omit NULLs on `update_batch` (this improves on the prior behavior of `max`) * Efficiently account for NULLs in `retract_batch` (this improves on the prior behavior of `min`) * Ensure correct results for all-NULL window frames (this fixes the prior bug in `min`). ## What changes are included in this PR? * Fix bug in `min()` over all-NULL window frames * Optimize `SlidingMinAccumulator::retract_batch` (avoid materializing values just to count NULLs) * Optimize `SlidingMinAccumulator::update_batch` (omit NULLs), also improving symmetry with `min` * Optimize both accumulators to stop caching the current `min` / `max`; this saves a few clones, but perhaps more importantly it is simpler and avoids the risk of inconsistency between the cached value and the underlying `MovingMin` / `MovingMax` data structure ## Are these changes tested? Yes, new tests added. ## Are there any user-facing changes? No, aside from the bug fix.
## Which issue does this PR close?
- N/A
## Rationale for this change
```
$ cargo test -p datafusion-functions-aggregate
[...]
warning: associated function `from_parts` is never used
--> datafusion/physical-expr/src/expressions/dynamic_filters/mod.rs:464:8
[...]
```
Fix this by appropriately gating the compilation of `from_parts`.
## What changes are included in this PR?
* Squelch unused code warning
## Are these changes tested?
Yes.
## Are there any user-facing changes?
No.
…w retract (apache#23913) ## Which issue does this PR close? - Closes apache#23912. ## Rationale for this change `DistinctPercentileContAccumulator` reused the shared set-based `GenericDistinctBuffer`, which doesn't fit it: (1) the buffer asserts a single input column but `percentile_cont` passes two (value + percentile), so every `percentile_cont(DISTINCT ...)` panicked; (2) the buffer is a plain `HashSet` with no multiplicity, so sliding-window `retract_batch` dropped a value while duplicates were still in the frame. ## What changes are included in this PR? - Replace the shared buffer in this accumulator with a per-accumulator `HashMap<Hashable, usize>` count map: `update_batch` reads only the value column and increments; `retract_batch` decrements and removes a key only at zero; `state`/`merge_batch` keep the same List state shape. Other `GenericDistinctBuffer` users are untouched. - Regression tests in `aggregate.slt` for the plain distinct query and the sliding-window duplicate-retract case. ## Are these changes tested? Yes — new regression tests; the full `aggregate.slt` suite passes. ## Are there any user-facing changes? `percentile_cont(DISTINCT ...)` now works instead of panicking, and returns correct results in sliding windows.
## Which issue does this PR close? - Closes apache#23717. ## Rationale for this change Spark and Java format negative numeric values with parentheses when the `(` flag is present. The decimal formatting path always emitted a minus sign because it ignored `negative_in_parentheses`, while the floating-point path already handled the flag. This made `format_string` inconsistent across numeric input types. ## What changes are included in this PR? - Add a closing suffix when a negative decimal uses parentheses formatting. - Include the suffix when calculating width, left alignment, and zero padding. - Add regression coverage for grouped negative decimals with and without an explicit width. ## Are these changes tested? Yes. The following checks pass locally: - `cargo fmt --all -- --check` - `cargo clippy --all-targets --all-features -- -D warnings` - `cargo test -p datafusion-spark --lib` - The full workspace test command from the contributor guide with `avro,json,backtrace,extended_tests,recursive_protection,parquet_encryption` enabled ## Are there any user-facing changes? Yes. Spark-compatible formatting of negative decimal values now honors the parentheses flag. There are no public API or breaking changes.
## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes #123` indicates that this PR will close issue #123. --> - Closes #. ## Rationale for this change <!-- Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed. Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes. --> Precedes apache#23720. While displaying a plan with metrics, allow filtering them by name, not just by category or type. This is very useful when writing snapshot tests where only a specific metrics needs to be asserted, without other unrelated metrics polluting the snapshot assertion. Regardless of what happens with apache#23720, I think this PR is still worth it, as it allows creating some really nice `insta` tests asserting runtime properties reliably by just cherry picking the runtime metrics relevant for that specific test. This is relevant not only within DataFusion codebase, but also for other people's codebases using DataFusion and relying on `insta` for snapshot testing, See an example of this here: https://github.com/apache/datafusion/pull/23720/changes#diff-8281405e117428077c19c07afbbe59fed303a3460096e958f761d34f0899b618 ## What changes are included in this PR? <!-- There is no need to duplicate the description in the issue here but it is sometimes worth providing a summary of the individual changes in this PR. --> While displaying metrics, allows filtering them by name ## Are these changes tested? <!-- We typically require tests for all PRs in order to: 1. Prevent the code from being accidentally broken by subsequent changes 2. Serve as another way to document the expected behavior of the code If tests are not included in your PR, please explain why (for example, are they covered by existing tests)? --> Yes, by a new small test. ## Are there any user-facing changes? <!-- If there are user-facing changes then we may require documentation to be updated before approving the PR. --> <!-- If there are any breaking changes to public APIs, please add the `api change` label. --> People using the metrics display API will be able to filter metrics by name.
## Which issue does this PR close? N/A ## Rationale for this change I saw that `empty` implementation is inefficient. I wrote review comments for more why inefficient. there are some optimizations that do not require benchmark since they are obvious once you understand, this is one of them ## What changes are included in this PR? rewrote the function to be fast ## Are these changes tested? existing tests ## Are there any user-facing changes? no
Bumps the codeql-actions group with 2 updates: [github/codeql-action/init](https://github.com/github/codeql-action) and [github/codeql-action/analyze](https://github.com/github/codeql-action). Updates `github/codeql-action/init` from 4.37.1 to 4.37.3 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/releases">github/codeql-action/init's releases</a>.</em></p> <blockquote> <h2>v4.37.3</h2> <p>No user facing changes.</p> <h2>v4.37.2</h2> <ul> <li>The new address format for the <code>config-file</code> input that was introduced in CodeQL Action 4.37.0 is now enabled by default. In addition to the format described there, the <code>remote=</code> prefix can now be used to explicitly indicate that the input refers to a remote file. All previous input formats continue to be accepted as well. <a href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li> <li>The CodeQL Action can now make use of <a href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured private registries</a> in Default Setup to retrieve CodeQL configuration files from remote repositories that require authentication. This will allow customers to store their CodeQL configuration in a single repository that can then be referenced by Default Setup workflows in other repositories. We expect to roll this and other, related changes out to everyone in July. <a href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/blob/main/CHANGELOG.md">github/codeql-action/init's changelog</a>.</em></p> <blockquote> <h1>CodeQL Action Changelog</h1> <p>See the <a href="https://github.com/github/codeql-action/releases">releases page</a> for the relevant changes to the CodeQL CLI and language packs.</p> <h2>[UNRELEASED]</h2> <ul> <li>This version of the CodeQL Action adds support for the <code>tools</code> input for the <code>codeql-action/init</code> step to be specified using a <code>github-codeql-tools</code> <a href="https://docs.github.com/en/organizations/managing-organization-settings/managing-custom-properties-for-repositories-in-your-organization">repository property</a>. This feature will gradually be rolled out following the release of this version. Once rolled out, this allows for the CodeQL CLI version that is used in GitHub-managed workflows, such as Default Setup, to be set to a custom value. For example, customers who run into issues with rate limits when a new CodeQL CLI version is released can set the value to <code>toolcache</code> to always use the CodeQL CLI version that is available in the runner toolcache. For Advanced Setup workflows, the value provided for <code>tools</code> in the workflow definition always takes precedence unless the value of the repository property starts with <code>!</code>. <a href="https://redirect.github.com/github/codeql-action/pull/4037">#4037</a></li> </ul> <h2>4.37.3 - 22 Jul 2026</h2> <p>No user facing changes.</p> <h2>4.37.2 - 21 Jul 2026</h2> <ul> <li>The new address format for the <code>config-file</code> input that was introduced in CodeQL Action 4.37.0 is now enabled by default. In addition to the format described there, the <code>remote=</code> prefix can now be used to explicitly indicate that the input refers to a remote file. All previous input formats continue to be accepted as well. <a href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li> <li>The CodeQL Action can now make use of <a href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured private registries</a> in Default Setup to retrieve CodeQL configuration files from remote repositories that require authentication. This will allow customers to store their CodeQL configuration in a single repository that can then be referenced by Default Setup workflows in other repositories. We expect to roll this and other, related changes out to everyone in July. <a href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li> </ul> <h2>4.37.1 - 16 Jul 2026</h2> <ul> <li><em>Upcoming breaking change</em>: Add a deprecation warning for customers using CodeQL version 2.20.6 and earlier. These versions of CodeQL were discontinued on 1 July 2026 alongside GitHub Enterprise Server 3.16, and will be unsupported by the next minor release of the CodeQL Action. <a href="https://redirect.github.com/github/codeql-action/pull/3956">#3956</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.1">2.26.1</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4019">#4019</a></li> </ul> <h2>4.37.0 - 08 Jul 2026</h2> <ul> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.0">2.26.0</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3995">#3995</a></li> <li>In addition to the existing input format, the <code>config-file</code> input for the <code>codeql-action/init</code> step will soon support a new <code>[owner/]repo[@ref][:path]</code> format. All components except the repository name are optional. If omitted, <code>owner</code> defaults to the same owner as the repository the analysis is running for, <code>ref</code> to <code>main</code>, and <code>path</code> to <code>.github/codeql-action.yaml</code>. Support for this format ships in this version of the CodeQL Action, but will only be enabled over the coming weeks. <a href="https://redirect.github.com/github/codeql-action/pull/3973">#3973</a></li> </ul> <h2>4.36.3 - 01 Jul 2026</h2> <p>No user facing changes.</p> <h2>4.36.2 - 04 Jun 2026</h2> <ul> <li>Cache CodeQL CLI version information across Actions steps. <a href="https://redirect.github.com/github/codeql-action/pull/3943">#3943</a></li> <li>Reduce requests while waiting for analysis processing by using exponential backoff when polling SARIF processing status. <a href="https://redirect.github.com/github/codeql-action/pull/3937">#3937</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.6">2.25.6</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3948">#3948</a></li> </ul> <h2>4.36.1 - 02 Jun 2026</h2> <p>No user facing changes.</p> <h2>4.36.0 - 22 May 2026</h2> <ul> <li><em>Breaking change</em>: Bump the minimum required CodeQL bundle version to 2.19.4. <a href="https://redirect.github.com/github/codeql-action/pull/3894">#3894</a></li> <li>Add support for SHA-256 Git object IDs. <a href="https://redirect.github.com/github/codeql-action/pull/3893">#3893</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.5">2.25.5</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3926">#3926</a></li> </ul> <h2>4.35.5 - 15 May 2026</h2> <ul> <li>We have improved how the JavaScript bundles for the CodeQL Action are generated to avoid duplication across bundles and reduce the size of the repository by around 70%. This should have no effect on the runtime behaviour of the CodeQL Action. <a href="https://redirect.github.com/github/codeql-action/pull/3899">#3899</a></li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/github/codeql-action/commit/e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81"><code>e4fba86</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4031">#4031</a> from github/update-v4.37.3-72f6a9da0</li> <li><a href="https://github.com/github/codeql-action/commit/fb50ab5d62a274adf3ef3e22cfe750ae87a0ede7"><code>fb50ab5</code></a> Update changelog for v4.37.3</li> <li><a href="https://github.com/github/codeql-action/commit/72f6a9da0def52d9193d6a758f0378b65091f8d1"><code>72f6a9d</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4030">#4030</a> from github/mbg/fix/no-proxy</li> <li><a href="https://github.com/github/codeql-action/commit/3b5ee58597653d9cc6785f3f1277f796d81f3646"><code>3b5ee58</code></a> Use default <code>request</code> options instead of <code>undefined</code></li> <li><a href="https://github.com/github/codeql-action/commit/bfb6be4b5ecd3650f02f530571453e8c64ef0778"><code>bfb6be4</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4028">#4028</a> from github/mergeback/v4.37.2-to-main-e0647621</li> <li><a href="https://github.com/github/codeql-action/commit/526ab84f9858816d9cf5f7b9df4dd5e2235f0eba"><code>526ab84</code></a> Rebuild</li> <li><a href="https://github.com/github/codeql-action/commit/d6217b9b8c14166e4851db94c11155d03bd13c07"><code>d6217b9</code></a> Update changelog and version after v4.37.2</li> <li><a href="https://github.com/github/codeql-action/commit/e0647621c2984b5ed2f768cb892365bf2a616ad1"><code>e064762</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4027">#4027</a> from github/update-v4.37.2-385bcdc5a</li> <li><a href="https://github.com/github/codeql-action/commit/e0faed839190caa67a5cd42f1cc16246028ca3df"><code>e0faed8</code></a> Add a couple of change notes</li> <li><a href="https://github.com/github/codeql-action/commit/73aad0eaa9df172668665a150d17b8bc5a650c20"><code>73aad0e</code></a> Update changelog for v4.37.2</li> <li>Additional commits viewable in <a href="https://github.com/github/codeql-action/compare/7188fc363630916deb702c7fdcf4e481b751f97a...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81">compare view</a></li> </ul> </details> <br /> Updates `github/codeql-action/analyze` from 4.37.1 to 4.37.3 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/releases">github/codeql-action/analyze's releases</a>.</em></p> <blockquote> <h2>v4.37.3</h2> <p>No user facing changes.</p> <h2>v4.37.2</h2> <ul> <li>The new address format for the <code>config-file</code> input that was introduced in CodeQL Action 4.37.0 is now enabled by default. In addition to the format described there, the <code>remote=</code> prefix can now be used to explicitly indicate that the input refers to a remote file. All previous input formats continue to be accepted as well. <a href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li> <li>The CodeQL Action can now make use of <a href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured private registries</a> in Default Setup to retrieve CodeQL configuration files from remote repositories that require authentication. This will allow customers to store their CodeQL configuration in a single repository that can then be referenced by Default Setup workflows in other repositories. We expect to roll this and other, related changes out to everyone in July. <a href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/github/codeql-action/blob/main/CHANGELOG.md">github/codeql-action/analyze's changelog</a>.</em></p> <blockquote> <h1>CodeQL Action Changelog</h1> <p>See the <a href="https://github.com/github/codeql-action/releases">releases page</a> for the relevant changes to the CodeQL CLI and language packs.</p> <h2>[UNRELEASED]</h2> <ul> <li>This version of the CodeQL Action adds support for the <code>tools</code> input for the <code>codeql-action/init</code> step to be specified using a <code>github-codeql-tools</code> <a href="https://docs.github.com/en/organizations/managing-organization-settings/managing-custom-properties-for-repositories-in-your-organization">repository property</a>. This feature will gradually be rolled out following the release of this version. Once rolled out, this allows for the CodeQL CLI version that is used in GitHub-managed workflows, such as Default Setup, to be set to a custom value. For example, customers who run into issues with rate limits when a new CodeQL CLI version is released can set the value to <code>toolcache</code> to always use the CodeQL CLI version that is available in the runner toolcache. For Advanced Setup workflows, the value provided for <code>tools</code> in the workflow definition always takes precedence unless the value of the repository property starts with <code>!</code>. <a href="https://redirect.github.com/github/codeql-action/pull/4037">#4037</a></li> </ul> <h2>4.37.3 - 22 Jul 2026</h2> <p>No user facing changes.</p> <h2>4.37.2 - 21 Jul 2026</h2> <ul> <li>The new address format for the <code>config-file</code> input that was introduced in CodeQL Action 4.37.0 is now enabled by default. In addition to the format described there, the <code>remote=</code> prefix can now be used to explicitly indicate that the input refers to a remote file. All previous input formats continue to be accepted as well. <a href="https://redirect.github.com/github/codeql-action/pull/4023">#4023</a></li> <li>The CodeQL Action can now make use of <a href="https://docs.github.com/en/code-security/how-tos/secure-at-scale/configure-organization-security/manage-usage-and-access/giving-org-access-private-registries">configured private registries</a> in Default Setup to retrieve CodeQL configuration files from remote repositories that require authentication. This will allow customers to store their CodeQL configuration in a single repository that can then be referenced by Default Setup workflows in other repositories. We expect to roll this and other, related changes out to everyone in July. <a href="https://redirect.github.com/github/codeql-action/pull/4007">#4007</a></li> </ul> <h2>4.37.1 - 16 Jul 2026</h2> <ul> <li><em>Upcoming breaking change</em>: Add a deprecation warning for customers using CodeQL version 2.20.6 and earlier. These versions of CodeQL were discontinued on 1 July 2026 alongside GitHub Enterprise Server 3.16, and will be unsupported by the next minor release of the CodeQL Action. <a href="https://redirect.github.com/github/codeql-action/pull/3956">#3956</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.1">2.26.1</a>. <a href="https://redirect.github.com/github/codeql-action/pull/4019">#4019</a></li> </ul> <h2>4.37.0 - 08 Jul 2026</h2> <ul> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.26.0">2.26.0</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3995">#3995</a></li> <li>In addition to the existing input format, the <code>config-file</code> input for the <code>codeql-action/init</code> step will soon support a new <code>[owner/]repo[@ref][:path]</code> format. All components except the repository name are optional. If omitted, <code>owner</code> defaults to the same owner as the repository the analysis is running for, <code>ref</code> to <code>main</code>, and <code>path</code> to <code>.github/codeql-action.yaml</code>. Support for this format ships in this version of the CodeQL Action, but will only be enabled over the coming weeks. <a href="https://redirect.github.com/github/codeql-action/pull/3973">#3973</a></li> </ul> <h2>4.36.3 - 01 Jul 2026</h2> <p>No user facing changes.</p> <h2>4.36.2 - 04 Jun 2026</h2> <ul> <li>Cache CodeQL CLI version information across Actions steps. <a href="https://redirect.github.com/github/codeql-action/pull/3943">#3943</a></li> <li>Reduce requests while waiting for analysis processing by using exponential backoff when polling SARIF processing status. <a href="https://redirect.github.com/github/codeql-action/pull/3937">#3937</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.6">2.25.6</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3948">#3948</a></li> </ul> <h2>4.36.1 - 02 Jun 2026</h2> <p>No user facing changes.</p> <h2>4.36.0 - 22 May 2026</h2> <ul> <li><em>Breaking change</em>: Bump the minimum required CodeQL bundle version to 2.19.4. <a href="https://redirect.github.com/github/codeql-action/pull/3894">#3894</a></li> <li>Add support for SHA-256 Git object IDs. <a href="https://redirect.github.com/github/codeql-action/pull/3893">#3893</a></li> <li>Update default CodeQL bundle version to <a href="https://github.com/github/codeql-action/releases/tag/codeql-bundle-v2.25.5">2.25.5</a>. <a href="https://redirect.github.com/github/codeql-action/pull/3926">#3926</a></li> </ul> <h2>4.35.5 - 15 May 2026</h2> <ul> <li>We have improved how the JavaScript bundles for the CodeQL Action are generated to avoid duplication across bundles and reduce the size of the repository by around 70%. This should have no effect on the runtime behaviour of the CodeQL Action. <a href="https://redirect.github.com/github/codeql-action/pull/3899">#3899</a></li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/github/codeql-action/commit/e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81"><code>e4fba86</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4031">#4031</a> from github/update-v4.37.3-72f6a9da0</li> <li><a href="https://github.com/github/codeql-action/commit/fb50ab5d62a274adf3ef3e22cfe750ae87a0ede7"><code>fb50ab5</code></a> Update changelog for v4.37.3</li> <li><a href="https://github.com/github/codeql-action/commit/72f6a9da0def52d9193d6a758f0378b65091f8d1"><code>72f6a9d</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4030">#4030</a> from github/mbg/fix/no-proxy</li> <li><a href="https://github.com/github/codeql-action/commit/3b5ee58597653d9cc6785f3f1277f796d81f3646"><code>3b5ee58</code></a> Use default <code>request</code> options instead of <code>undefined</code></li> <li><a href="https://github.com/github/codeql-action/commit/bfb6be4b5ecd3650f02f530571453e8c64ef0778"><code>bfb6be4</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4028">#4028</a> from github/mergeback/v4.37.2-to-main-e0647621</li> <li><a href="https://github.com/github/codeql-action/commit/526ab84f9858816d9cf5f7b9df4dd5e2235f0eba"><code>526ab84</code></a> Rebuild</li> <li><a href="https://github.com/github/codeql-action/commit/d6217b9b8c14166e4851db94c11155d03bd13c07"><code>d6217b9</code></a> Update changelog and version after v4.37.2</li> <li><a href="https://github.com/github/codeql-action/commit/e0647621c2984b5ed2f768cb892365bf2a616ad1"><code>e064762</code></a> Merge pull request <a href="https://redirect.github.com/github/codeql-action/issues/4027">#4027</a> from github/update-v4.37.2-385bcdc5a</li> <li><a href="https://github.com/github/codeql-action/commit/e0faed839190caa67a5cd42f1cc16246028ca3df"><code>e0faed8</code></a> Add a couple of change notes</li> <li><a href="https://github.com/github/codeql-action/commit/73aad0eaa9df172668665a150d17b8bc5a650c20"><code>73aad0e</code></a> Update changelog for v4.37.2</li> <li>Additional commits viewable in <a href="https://github.com/github/codeql-action/compare/7188fc363630916deb702c7fdcf4e481b751f97a...e4fba868fa4b1b91e1fdab776edc8cfbe6e9fb81">compare view</a></li> </ul> </details> <br /> Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore <dependency name> major version` will close this group update PR and stop Dependabot creating any more for the specific dependency's major version (unless you unignore this specific dependency's major version or upgrade to it yourself) - `@dependabot ignore <dependency name> minor version` will close this group update PR and stop Dependabot creating any more for the specific dependency's minor version (unless you unignore this specific dependency's minor version or upgrade to it yourself) - `@dependabot ignore <dependency name>` will close this group update PR and stop Dependabot creating any more for the specific dependency (unless you unignore this specific dependency or upgrade to it yourself) - `@dependabot unignore <dependency name>` will remove all of the ignore conditions of the specified dependency - `@dependabot unignore <dependency name> <ignore condition>` will remove the ignore condition of the specified dependency and ignore conditions </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…e#23941) Bumps [taiki-e/install-action](https://github.com/taiki-e/install-action) from 2.84.0 to 2.85.2. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/taiki-e/install-action/releases">taiki-e/install-action's releases</a>.</em></p> <blockquote> <h2>2.85.2</h2> <ul> <li> <p>Update <code>prek@latest</code> to 0.4.11.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.7.13.</p> </li> <li> <p>Update <code>kingfisher@latest</code> to 1.109.0.</p> </li> </ul> <h2>2.85.1</h2> <ul> <li> <p>Update <code>vacuum@latest</code> to 0.30.0.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.11.32.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.7.12.</p> </li> <li> <p>Update <code>cyclonedx@latest</code> to 0.33.1.</p> </li> <li> <p>Update <code>cargo-neat@latest</code> to 0.5.2.</p> </li> </ul> <h2>2.85.0</h2> <ul> <li> <p>Support <code>wild</code> (alias: <code>wild-linker</code>). (<a href="https://redirect.github.com/taiki-e/install-action/pull/1949">#1949</a>)</p> </li> <li> <p>Support <code>bpf-linker</code>. (<a href="https://redirect.github.com/taiki-e/install-action/pull/1950">#1950</a>)</p> </li> <li> <p>Support <code>rafn</code>. (<a href="https://redirect.github.com/taiki-e/install-action/pull/1935">#1935</a>, thanks <a href="https://github.com/DarkWanderer"><code>@DarkWanderer</code></a>)</p> </li> <li> <p>Update <code>cargo-neat@latest</code> to 0.5.1.</p> </li> <li> <p>Update <code>zizmor@latest</code> to 1.28.0.</p> </li> <li> <p>Update <code>wasmtime@latest</code> to 47.0.2.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.11.31.</p> </li> <li> <p>Update <code>syft@latest</code> to 1.49.0.</p> </li> </ul> <h2>2.84.1</h2> <ul> <li> <p>Update <code>wasmtime@latest</code> to 47.0.1.</p> </li> <li> <p>Update <code>wasm-tools@latest</code> to 1.254.0.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.11.30.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.7.11.</p> </li> <li> <p>Update <code>cargo-neat@latest</code> to 0.5.0.</p> </li> <li> <p>Update <code>cargo-crap@latest</code> to 0.3.1.</p> </li> <li> <p>Update <code>biome@latest</code> to 2.5.5.</p> </li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/taiki-e/install-action/blob/main/CHANGELOG.md">taiki-e/install-action's changelog</a>.</em></p> <blockquote> <h1>Changelog</h1> <p>All notable changes to this project will be documented in this file.</p> <p>This project adheres to <a href="https://semver.org">Semantic Versioning</a>.</p> <!-- raw HTML omitted --> <h2>[Unreleased]</h2> <h2>[2.85.2] - 2026-07-26</h2> <ul> <li> <p>Update <code>prek@latest</code> to 0.4.11.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.7.13.</p> </li> <li> <p>Update <code>kingfisher@latest</code> to 1.109.0.</p> </li> </ul> <h2>[2.85.1] - 2026-07-25</h2> <ul> <li> <p>Update <code>vacuum@latest</code> to 0.30.0.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.11.32.</p> </li> <li> <p>Update <code>mise@latest</code> to 2026.7.12.</p> </li> <li> <p>Update <code>cyclonedx@latest</code> to 0.33.1.</p> </li> <li> <p>Update <code>cargo-neat@latest</code> to 0.5.2.</p> </li> </ul> <h2>[2.85.0] - 2026-07-23</h2> <ul> <li> <p>Support <code>wild</code> (alias: <code>wild-linker</code>). (<a href="https://redirect.github.com/taiki-e/install-action/pull/1949">#1949</a>)</p> </li> <li> <p>Support <code>bpf-linker</code>. (<a href="https://redirect.github.com/taiki-e/install-action/pull/1950">#1950</a>)</p> </li> <li> <p>Support <code>rafn</code>. (<a href="https://redirect.github.com/taiki-e/install-action/pull/1935">#1935</a>, thanks <a href="https://github.com/DarkWanderer"><code>@DarkWanderer</code></a>)</p> </li> <li> <p>Update <code>cargo-neat@latest</code> to 0.5.1.</p> </li> <li> <p>Update <code>zizmor@latest</code> to 1.28.0.</p> </li> <li> <p>Update <code>wasmtime@latest</code> to 47.0.2.</p> </li> <li> <p>Update <code>uv@latest</code> to 0.11.31.</p> </li> <li> <p>Update <code>syft@latest</code> to 1.49.0.</p> </li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/taiki-e/install-action/commit/41049aa56687c35e0afa74eed4f09cec4f9afabf"><code>41049aa</code></a> Release 2.85.2</li> <li><a href="https://github.com/taiki-e/install-action/commit/dfcf36552b1b743e910b9b9b6b114a08e56a1e5e"><code>dfcf365</code></a> Update <code>prek@latest</code> to 0.4.11</li> <li><a href="https://github.com/taiki-e/install-action/commit/eea03ccfa855b965a301aa52424915fddc32f855"><code>eea03cc</code></a> Update <code>mise@latest</code> to 2026.7.13</li> <li><a href="https://github.com/taiki-e/install-action/commit/81ca2feb84748ed3fa514707ee1004f4f3ad715a"><code>81ca2fe</code></a> Update martin manifest</li> <li><a href="https://github.com/taiki-e/install-action/commit/cc90ed04bc1a258ea90a83789f24b759a77c9671"><code>cc90ed0</code></a> Update <code>kingfisher@latest</code> to 1.109.0</li> <li><a href="https://github.com/taiki-e/install-action/commit/55639a3362f508fda5154d3451072205f5c32ca6"><code>55639a3</code></a> Update cargo-shear manifest</li> <li><a href="https://github.com/taiki-e/install-action/commit/3d7d7cd5ac7f994c1892ae0c06165095b9139094"><code>3d7d7cd</code></a> Release 2.85.1</li> <li><a href="https://github.com/taiki-e/install-action/commit/d09ccb4fe2105ac9ead6376f7a423ff88defeb57"><code>d09ccb4</code></a> Update <code>vacuum@latest</code> to 0.30.0</li> <li><a href="https://github.com/taiki-e/install-action/commit/ac43dee1a92e482c0e244bdb0bb126908159e245"><code>ac43dee</code></a> Update <code>uv@latest</code> to 0.11.32</li> <li><a href="https://github.com/taiki-e/install-action/commit/49b16979f38f30e44b1f68af9115be9a6ffc5215"><code>49b1697</code></a> Update prek manifest</li> <li>Additional commits viewable in <a href="https://github.com/taiki-e/install-action/compare/a6b2e2dcd845ddd7f509ce4f3ed3d922b80cc5d9...41049aa56687c35e0afa74eed4f09cec4f9afabf">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [actions/stale](https://github.com/actions/stale) from 10.4.0 to 11.0.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/actions/stale/releases">actions/stale's releases</a>.</em></p> <blockquote> <h2>v11.0.0</h2> <h2>What's Changed</h2> <h3>Enhancement</h3> <ul> <li>Migrate to ESM and update dependencies by <a href="https://github-grid.enterprise.slack.com/team/U08CVLQ4JKE"><code>@chiranjib-swain</code></a> in <a href="https://redirect.github.com/actions/stale/pull/1350">actions/stale#1350</a></li> </ul> <h3>Dependency Update</h3> <ul> <li>Override brace-expansion to 5.0.8 to address 24 high-severity dependency vulnerabilities by <a href="https://github.com/dependabot"><code>@dependabot</code></a> in <a href="https://redirect.github.com/actions/stale/pull/1351">actions/stale#1351</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/actions/stale/compare/v10...v11.0.0">https://github.com/actions/stale/compare/v10...v11.0.0</a></p> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/actions/stale/commit/4391f3da665fdf50b6810c1a66712fb9ba21aa93"><code>4391f3d</code></a> Fix 24 high severity vulnerabilities by overriding brace-expansion to 5.0.8 (...</li> <li><a href="https://github.com/actions/stale/commit/eaf9131fae5eafd0c31a64ebe3a2e183266fec48"><code>eaf9131</code></a> refactor: update imports to use ES module syntax and improve test structure (...</li> <li>See full diff in <a href="https://github.com/actions/stale/compare/1e223db275d687790206a7acac4d1a11bd6fe629...4391f3da665fdf50b6810c1a66712fb9ba21aa93">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [base64](https://github.com/marshallpierce/rust-base64) from 0.22.1 to 0.23.0. <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/marshallpierce/rust-base64/blob/master/RELEASE-NOTES.md">base64's changelog</a>.</em></p> <blockquote> <h1>0.23.0</h1> <ul> <li>Added more consts for preconfigured configs and engines</li> <li>Make DecodeError::InvalidLastSymbol more clear by including the decoded value</li> <li>Added SIMD-accelerated engines behind the default-on <code>simd-unsafe</code> feature: <code>Simd</code> picks the best instruction set at runtime (AVX2 on <code>x86_64</code>, NEON on <code>aarch64</code>) and falls back to the scalar <code>GeneralPurpose</code> engine, while <code>Avx2</code> and <code>Neon</code> target one instruction set with no runtime detection and work in <code>no_std</code>. The engines support the standard and URL-safe alphabets.</li> <li>Update MSRV to 1.71.0</li> <li>Add support for custom padding symbols</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/marshallpierce/rust-base64/commit/9e9220a4166f628de7c8803289e120ae1e944f78"><code>9e9220a</code></a> v0.23.0</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/870326ec592eebde9d6bfe4c5d8130c591273e9c"><code>870326e</code></a> Merge pull request <a href="https://redirect.github.com/marshallpierce/rust-base64/issues/306">#306</a> from marshallpierce/mp/trailing-bits-docs</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/fbec5f1050f9fc16e6a826ebabaa2b7b0644bd67"><code>fbec5f1</code></a> Document no trailing trailing bits</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/0a23549968f059b53cf39e96eba8f46779f322a7"><code>0a23549</code></a> Merge pull request <a href="https://redirect.github.com/marshallpierce/rust-base64/issues/305">#305</a> from marshallpierce/mp/edition-2021</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/f10b7e20614135aa61289140683fc93e5a45d338"><code>f10b7e2</code></a> Update deps & edition</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/9d21a598860645cb6290940e7a43033bc43ebd74"><code>9d21a59</code></a> Merge pull request <a href="https://redirect.github.com/marshallpierce/rust-base64/issues/304">#304</a> from marshallpierce/mp/custom-padding-rebase</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/f70bad2caaa85350b95d988bfb9a0997e824bfd8"><code>f70bad2</code></a> Support custom padding symbols</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/684d79cd3deb8dfd5323619634c75bc0ff6edfd9"><code>684d79c</code></a> Merge pull request <a href="https://redirect.github.com/marshallpierce/rust-base64/issues/301">#301</a> from marshallpierce/mp/simd-gardening</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/5bf66f2646c6fd99e1b18bf4e2bb0a47d34e1eaa"><code>5bf66f2</code></a> Merge pull request <a href="https://redirect.github.com/marshallpierce/rust-base64/issues/284">#284</a> from AbeZbm/add-tests</li> <li><a href="https://github.com/marshallpierce/rust-base64/commit/d3831cfbf7dafe226a8383450c3410e2f67c826a"><code>d3831cf</code></a> Followups to SIMD work</li> <li>Additional commits viewable in <a href="https://github.com/marshallpierce/rust-base64/compare/v0.22.1...v0.23.0">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [astral-sh/setup-uv](https://github.com/astral-sh/setup-uv) from 8.3.2 to 9.0.0. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/astral-sh/setup-uv/releases">astral-sh/setup-uv's releases</a>.</em></p> <blockquote> <h2>v9.0.0 🌈 Change <code>prune-cache</code> default to <code>false</code></h2> <h2>Changes</h2> <p>This release disables the default cache cache pruning to ease the load on the PyPi infrastructure. Since users might experience more GitHub Actions cache usage which might result in higher costs this is marked as a breaking change. To read more on why we did this (now) you can read the detailed analysis and reasoning in <a href="https://redirect.github.com/astral-sh/setup-uv/issues/967">#967</a></p> <p>Besides this big breaking change we also have a small bugfix while building caches for linux distributions that behave a big different than the "big ones" and a speed up in version resolution by only reading the version manifest until a matching version is found saving runtime and network bandwith.</p> <h2>🚨 Breaking changes</h2> <ul> <li>Change <code>prune-cache</code> default to <code>false</code> <a href="https://github.com/charliermarsh"><code>@charliermarsh</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/967">#967</a>)</li> </ul> <h2>🐛 Bug fixes</h2> <ul> <li>fix: fall back to distribution ID when os-release has no version field <a href="https://github.com/cxzhong"><code>@cxzhong</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/961">#961</a>)</li> </ul> <h2>🚀 Enhancements</h2> <ul> <li>Speed up version client by partial response reads <a href="https://github.com/eifinger"><code>@eifinger</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/807">#807</a>)</li> </ul> <h2>🧰 Maintenance</h2> <ul> <li>chore: update known checksums for 0.11.30 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/968">#968</a>)</li> <li>chore: update known checksums for 0.11.29 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/960">#960</a>)</li> </ul> <h2>📚 Documentation</h2> <ul> <li>docs: update version references to v8.3.2 @<a href="https://github.com/apps/github-actions">github-actions[bot]</a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/949">#949</a>)</li> </ul> <h2>⬆️ Dependency updates</h2> <ul> <li>chore(deps): roll up Dependabot updates <a href="https://github.com/eifinger"><code>@eifinger</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/970">#970</a>)</li> <li>chore(deps): roll up Dependabot updates <a href="https://github.com/eifinger"><code>@eifinger</code></a> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/962">#962</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/astral-sh/setup-uv/commit/c771a70e6277c0a99b617c7a806ffedaca235ff9"><code>c771a70</code></a> chore(deps): roll up Dependabot updates (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/970">#970</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/2f537ca87c1ffa233ca2a1b84815388e3e42d845"><code>2f537ca</code></a> chore: update known checksums for 0.11.30 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/968">#968</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/2269552d547df6f50e57442326930d30d943afe3"><code>2269552</code></a> Speed up version client by partial response reads (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/807">#807</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/47a7f4fb2e900d6c33a5b5f231fa21dbfaeba52f"><code>47a7f4f</code></a> Change <code>prune-cache</code> default to <code>false</code> (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/967">#967</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/71966eff34a27b0a62ed4b9f6f6e383e071b1bb5"><code>71966ef</code></a> chore(deps): roll up Dependabot updates (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/962">#962</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/f12b1f0a84bd6dc2331b36b2bbdbb1d1e617dbcc"><code>f12b1f0</code></a> fix: fall back to distribution ID when os-release has no version field (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/961">#961</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/ecd24dd710f2fb0dca1693a67af11fc4a5c5ec84"><code>ecd24dd</code></a> chore: update known checksums for 0.11.29 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/960">#960</a>)</li> <li><a href="https://github.com/astral-sh/setup-uv/commit/6a191366842ac1502ba6c07e9b5acd5c2d9d8db3"><code>6a19136</code></a> docs: update version references to v8.3.2 (<a href="https://redirect.github.com/astral-sh/setup-uv/issues/949">#949</a>)</li> <li>See full diff in <a href="https://github.com/astral-sh/setup-uv/compare/11f9893b081a58869d3b5fccaea48c9e9e46f990...c771a70e6277c0a99b617c7a806ffedaca235ff9">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…fy to be textbook like as possible (apache#23761) ## Which issue does this PR close? N/A ## Rationale for this change SortMergeJoin bitwise stream implementation is very complex and hard to understand while on paper it should be pretty simple. the reason for that is we have to store state between polls (we had `boundary`) and handle the case where both can get `Poll::Pending` from child and `Poll::Ready` from child which further complicate the code ## What changes are included in this PR? 1. Move to async generators 2. Rewrote main loop to be textbook like as possible (this was entirely written by Claude Fable, sorry, I tried manually but the code was too complex to hold in my head 😅 ) ## Are these changes tested? existing tests ## Are there any user-facing changes? The join_time now includes the time to read from the async spill stream between pending which is arguable more correct since this time is part of the operator, although long waits between pending calls will be counted in the op `join_time` while the alternative is not counting the read from file and decoding...
…lator (apache#23946) ## Rationale for this change Follow-up to post-merge review feedback from @neilconway on apache#23913. ## What changes are included in this PR? - Use the fast foldhash `RandomState` for the distinct-value count map instead of the standard library's default SipHash (the shared `GenericDistinctBuffer` already does this; the merged fix regressed to SipHash on this hot path). - Use `estimate_memory_size` in `size()` instead of a hand-rolled capacity calculation. - Adopt the null-free fast path in `update_batch`/`retract_batch` (skip per-element validity checks when the input has no nulls), mirroring `GenericDistinctBuffer`. - Add `ORDER BY` to the sliding-window regression test so its row order is deterministic. ## Are these changes tested? Yes — existing percentile unit tests and the full `aggregate.slt` pass. ## Are there any user-facing changes? No.
## Which issue does this PR close? - Related to apache/datafusion-comet#4131. ## Rationale for this change `TrivialLastValueAccumulator::merge_batch` filtered out unset states but then read row `0` from the remaining value array. As a result, merging multiple `LAST_VALUE` partial states returned the first valid state instead of the last one. ## What changes are included in this PR? - Select the final valid state row when merging unordered `LAST_VALUE` states. - Extend the existing merge test to verify the evaluated result. ## Are these changes tested? Yes: ```shell cargo test -p datafusion-functions-aggregate test_first_last_state_after_merge
Bumps [syn](https://github.com/dtolnay/syn) from 2.0.119 to 3.0.2. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/dtolnay/syn/releases">syn's releases</a>.</em></p> <blockquote> <h2>3.0.2</h2> <ul> <li>Add <a href="https://docs.rs/syn/3/syn/struct.Error.html#method.new_range"><code>Error::new_range(start..end, "msg")</code></a> (<a href="https://redirect.github.com/dtolnay/syn/issues/2068">#2068</a>, <a href="https://redirect.github.com/dtolnay/syn/issues/2070">#2070</a>)</li> <li>Add <a href="https://docs.rs/syn/3/syn/buffer/struct.Cursor.html#method.prev_span"><code>Cursor::prev_span</code></a></li> </ul> <h2>3.0.1</h2> <ul> <li>Parse const traits (<a href="https://redirect.github.com/dtolnay/syn/issues/2056">#2056</a>, <a href="https://redirect.github.com/dtolnay/syn/issues/2057">#2057</a>, <a href="https://redirect.github.com/dtolnay/syn/issues/2058">#2058</a>, <a href="https://redirect.github.com/dtolnay/syn/issues/2063">#2063</a>, <a href="https://redirect.github.com/dtolnay/syn/issues/2064">#2064</a>)</li> <li>Parse unsafe binder types (<a href="https://redirect.github.com/dtolnay/syn/issues/2065">#2065</a>)</li> <li>Parse impl restrictions (<a href="https://redirect.github.com/dtolnay/syn/issues/2066">#2066</a>)</li> </ul> <h2>3.0.0</h2> <p>This release contains adjustments to the syntax tree to account for ongoing Rust language development from the 3 years since syn 2.0.0 and to anticipate some in-flight Rust language RFCs.</p> <p>These include: default values in fields, pinned type sugar, raw lifetimes, generator blocks and functions, unnamed enum variants, attributes in tuple types and tuple patterns, named arguments in parenthesized generic argument lists, lightweight clones, const traits, const function pointers, mutability restricted fields, supertrait auto implementation, final associated functions, trait implementability restrictions, const blocks in path arguments, item-level const blocks, return type notation, never patterns, function delegation, mutable by-reference bindings, in-place initialization, field projections, explicitly dyn-compatible traits, view types, file-level frontmatter, generic const arguments, guard patterns, lazy type aliases, explicitly safe foreign items, super let, unsafe fields, pattern types, heterogeneous try-blocks, function contracts, async function trait bounds, static closure coroutine syntax, unsafe binder types, move expressions, for-await loops, and postfix keywords.</p> <!-- raw HTML omitted --> <!-- raw HTML omitted --> <h1>Breaking changes</h1> <h2>Modifiers</h2> <p>To reserve more room for language evolution, there are 10 new non-exhaustive structs in the syntax tree having the following commonality:</p> <ul> <li> <p>Name ending in <code>Modifiers</code>. {<code>BlockModifiers</code>, <code>ClosureModifiers</code>, <code>ConstModifiers</code>, <code>FieldModifiers</code>, <code>FnModifiers</code>, <code>ImplModifiers</code>, <code>LocalModifiers</code>, <code>TraitBoundModifiers</code>, <code>TraitModifiers</code>, <code>TypeModifiers</code>}</p> </li> <li> <p>Each implements <code>Default</code>. The default value is guaranteed to comprise no tokens.</p> </li> <li> <p>Non-exhaustive. Can only be instantiated by Syn's parser or by creating and then mutating <code>▁▁Modifiers::default()</code>.</p> </li> <li> <p>Does not implement <code>Parse</code>. When parsing, they are parsed by the enclosing syntax tree node.</p> </li> <li> <p>Does not implement <code>ToTokens</code>. In some cases the syntax that these nodes might hold in the future is not necessarily contiguous tokens.</p> </li> <li> <p>Provides <code>.require_empty() -> Result<()></code> which returns a meaningfully spanned error if the modifiers are different from the empty default. This enables a caller to reject syntax it does not recognize without knowing what that syntax may be.</p> </li> </ul> <h2>Types</h2> <ul> <li> <p><code>Type::BareFn</code> has been renamed to <code>Type::FnPtr</code> to mirror the compiler's terminology. Together with this, <code>BareVariadic</code> is renamed to <code>FnPtrVariadic</code>.</p> </li> <li> <p>The mutually exclusive <code>const_token</code> and <code>mutability</code> fields of <code>Type::Ptr</code> have been unified into an enum of type <code>PointerMutability</code>, which was already previously used by <code>Expr::RawAddr</code>.</p> </li> <li> <p>Every <code>Type</code> variant now holds attributes, which can represent the attributes of element types inside a tuple type, or attributes for a function return type.</p> </li> <li> <p><code>BareFnArg</code> is renamed to <code>NamedArg</code> and is used in <code>ParenthesizedGenericArguments</code>, in addition to the existing use in <code>Type::FnPtr</code>.</p> </li> </ul> <h2>Expressions</h2> <ul> <li>In <code>Expr::Closure</code>, the fields <code>or1_token</code> and <code>or2_token</code> have been renamed to <code>inputs_begin</code> and <code>inputs_end</code> to indicate the beginning and ending <code>|</code> token of the closure inputs.</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/dtolnay/syn/commit/88ee7be2197e61d6e84f0bff38eb2fe57998a765"><code>88ee7be</code></a> Release 3.0.2</li> <li><a href="https://github.com/dtolnay/syn/commit/587bc203a0e975e8261a65791203e8912edb1c42"><code>587bc20</code></a> Merge pull request <a href="https://redirect.github.com/dtolnay/syn/issues/2070">#2070</a> from dtolnay/emptyrange</li> <li><a href="https://github.com/dtolnay/syn/commit/96801f717ed1ae7943182afe1fd4c5857c50428b"><code>96801f7</code></a> Allow Error::new_range at empty cursor range</li> <li><a href="https://github.com/dtolnay/syn/commit/9dc16c94cf0468dc4e4b1f6011cfd4040246c06c"><code>9dc16c9</code></a> Merge pull request <a href="https://redirect.github.com/dtolnay/syn/issues/2069">#2069</a> from dtolnay/prevspan</li> <li><a href="https://github.com/dtolnay/syn/commit/1db76b7ba3f4e16430e6b2c96f160c05160e056e"><code>1db76b7</code></a> Align on using impl trait across all Error constructors</li> <li><a href="https://github.com/dtolnay/syn/commit/bfa1ebf336f5fae0a0a22a89aeb991e2cd96c65a"><code>bfa1ebf</code></a> Make Cursor::prev_span public</li> <li><a href="https://github.com/dtolnay/syn/commit/c6ac5e57086963d9fe549ea1576d062f30b1ae32"><code>c6ac5e5</code></a> Merge pull request <a href="https://redirect.github.com/dtolnay/syn/issues/2068">#2068</a> from dtolnay/newrange</li> <li><a href="https://github.com/dtolnay/syn/commit/143645404be46202362a576650f2e1f8d56ae086"><code>1436454</code></a> Add Error::new_range constructor taking Range<Cursor></li> <li><a href="https://github.com/dtolnay/syn/commit/123c1485eaa5cb6db1034711136a0d65b17e7f0a"><code>123c148</code></a> Release 3.0.1</li> <li><a href="https://github.com/dtolnay/syn/commit/bc11dddf65ebc8a1398f5ad9d0e1a1b4e9a1e9cf"><code>bc11ddd</code></a> Merge pull request <a href="https://redirect.github.com/dtolnay/syn/issues/2067">#2067</a> from dtolnay/fastpeek</li> <li>Additional commits viewable in <a href="https://github.com/dtolnay/syn/compare/2.0.119...3.0.2">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Part of apache#23929 ## Rationale for this change Spark's `elt` errors on out-of-range indices when ANSI is on, and returns `NULL` when off. DataFusion previously always returned `NULL` (there was a `TODO` for this in the source). ## What changes are included in this PR? - `elt.rs`: read `enable_ansi_mode` from `ScalarFunctionArgs::config_options`; raise `"The index N is out of bounds. The array has M elements."` in ANSI mode, keep `NULL` behavior otherwise. `NULL` index always returns `NULL`. - `elt.slt`: add missing ANSI-off cases (index 0, negative, large, `NULL` index, `NULL` value, vectorized) and a full ANSI-on section with `statement error` assertions. ## Are these changes tested? Yes. New Rust unit tests and sqllogictest cases cover both modes. ## Are there any user-facing changes? With `enable_ansi_mode = true`, `elt` on invalid indices now errors instead of returning `NULL`. Default behavior unchanged. No API changes.
Leftover from arrow, but not needed anyway - see apache/arrow#45526
…mas on the column-wise path (apache#23523) ## Which issue does this PR close? - Part of apache#22715 (nested type coverage in `GroupValuesColumn` EPIC) - Alternative to apache#23128 (per-type approach) — implements the direction @alamb proposed in <apache#23128 (comment)> - Step toward the terminal goal of retiring `GroupValuesRows` entirely (apache#23404) ## Rationale for this change Today `GroupValuesColumn` is **all-or-nothing**: a single nested column in the GROUP BY key (\`Struct\`, \`List\`, \`FixedSizeList\`, …) makes \`supported_schema\` return \`false\` and drops the *entire* aggregation onto the row-wise \`GroupValuesRows\` fallback — even when every other column would have qualified for the column-wise fast path. For a \`GROUP BY int_col, struct_col\` shape, the \`int_col\` pays the row-encoded storage cost for no reason. ## What changes are included in this PR? Add \`RowsGroupColumn\`: a generic \`GroupColumn\` backed by a single-field \`RowConverter\`, wired in as the nested-type dispatch arm of \`group_column_supported_type\` / \`make_group_column\`. Native columns keep their type-specialized builders; the nested column pays row-encoding only for its one column. Gated to \`data_type.is_nested()\` so intentionally excluded scalar types (Float16, Decimal256) stay on \`GroupValuesRows\` and the \`group_column_supported_type\` ⇔ \`make_group_column\` invariant holds. ## Impact Memory, measured with 4000 groups of \`8 × Int64 + 1 × FixedSizeList<Int64, 4>\` in \`mixed_schema_column_path_uses_less_memory_than_rows_fallback\`: | | Bytes | vs baseline | |-------------------------------------------------------|----------|-------------| | \`GroupValuesRows\` (today's fallback) | 1096 KB | 100% | | \`GroupValuesColumn\` + \`RowsGroupColumn\` fallback | 594 KB | **54.2%** | Speed: not benchmarked as a headline result — the wins come from native columns keeping their type-specialized \`equal_to\`/\`append_val\` fast paths instead of falling back to byte-encoded row comparisons. ## Are these changes tested? Yes: - Unit tests inside \`row_backed\`: FSL / Struct roundtrip, \`take_n\`, \`supports_type\` matches \`RowConverter::supports_fields\`. - \`mixed_schema_column_path_uses_less_memory_than_rows_fallback\` (mod.rs): the 54.2% memory claim + identical group assignment vs \`GroupValuesRows\`. - \`nested_float_edge_cases_match_rows_fallback\`: nested \`-0.0\` / \`NaN\` produce the same groupings as \`GroupValuesRows\` (the correctness invariant to watch, since hashing runs on the raw column and equality runs on the row bytes). - \`multi_batch_and_emit_first_matches_rows_fallback\`: multi-batch streaming intern + \`EmitTo::First\` + \`take_n\`. All 39 tests in \`aggregates::group_values\` pass. ## Are there any user-facing changes? No — internal aggregation representation only. Same query results, lower memory footprint on mixed-schema GROUP BY keys. ## Follow-ups (out of scope) - Add coverage for any type \`RowConverter\` cannot encode (currently arrow-rs 59.x handles Map fine; \`supports_type\` delegates to \`RowConverter::supports_fields\` so it auto-tracks upstream). - Retire \`GroupValuesRows\` entirely once coverage is complete (apache#23404).
## Which issue does this PR close? - Closes apache#23715. ## Rationale for this change Queries using `array_agg(DISTINCT col)` were significantly slower than expected. Profiling revealed that `DistinctArrayAggAccumulator::update_batch` was allocating a heap-owned `String` on every single input row — even for rows whose value was already present in the accumulator. For a typical low-cardinality workload (e.g. a column of ~25 database names across thousands of rows), this meant paying the full allocation cost for every duplicate, which dominated the runtime. ## What changes are included in this PR? This PR applies the same deduplication strategy already used by `AggregateExec` for `GROUP BY`: duplicate rows now cost only a hash probe with no heap allocation, and new distinct values are appended to a single shared buffer rather than allocated individually. The fix applies to all column types, not just strings, and `retract_batch` support (required for sliding window frames such as `ROWS BETWEEN N PRECEDING AND CURRENT ROW`) is fully preserved. ## Are these changes tested? Four unit tests were added to `DistinctArrayAggAccumulator` — one each for `Utf8`, `Int64`, `Float64`, and `Date32` — to pin the deduplication contract across the most common column types and serve as a regression guard for future changes. The existing sliding window sqllogictest suite (`array_agg_sliding_window.slt`) covers `retract_batch` correctness end-to-end and passes unchanged. Two `update_batch` micro-benchmarks were added to measure the before/after on realistic data: one with low cardinality (~25 distinct database names in 8 192 rows, modelling the common production case) and one with high cardinality (~7 800 distinct values, modelling the worst case where almost every row is new). Results on an 8 192-row batch: | Benchmark | Before | After | Speedup | |---|---|---|---| | Low cardinality (~25 distinct DB names) | 648.6 µs | 189.4 µs | **3.42×** | | High cardinality (~7 800 distinct values) | 1078.3 µs | 269.4 µs | **4.00×** | ## Are there any user-facing changes? No user-facing changes No breaking changes to public APIs --------- Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
…pache#23735) ## Which issue does this PR close? - Closes apache#23732 ## Rationale for this change Follow-up to apache#23704, which made the `TypeCoercion` analyzer coerce `SIMILAR TO` operands to a common string type. The SQL planner still rejects `LargeUtf8` and `Utf8View` **patterns** before the analyzer ever runs: ```sql SELECT s FROM test, patterns WHERE s SIMILAR TO arrow_cast(pat, 'Utf8View'); -- error: Invalid pattern in SIMILAR TO expression ``` Now that the analyzer coerces both operands, this planner restriction is unnecessary (and inconsistent with `LIKE`, whose planner path has no such check). ## What changes are included in this PR? - `datafusion/sql/src/expr/mod.rs`: the pattern-type check in `sql_similarto_to_expr` now accepts `Utf8` / `LargeUtf8` / `Utf8View` / `Null` instead of only `Utf8` / `Null`. - `datafusion/sqllogictest/test_files/type_coercion.slt`: regression tests with non-literal `Utf8View` and `LargeUtf8` patterns (plus a `NOT SIMILAR TO` case), next to the existing apache#22886 regression tests. The planner relaxation and tests were originally part of apache#22887; this PR lands the remaining piece of it. ## Are these changes tested? Yes — new sqllogictest cases in `type_coercion.slt`. The existing planner rejection of non-string patterns (`SELECT 'a' SIMILAR TO 1`) still errors as before. ## Are there any user-facing changes? Yes: `SIMILAR TO` queries with non-literal `LargeUtf8` / `Utf8View` patterns that previously failed during SQL planning with `Invalid pattern in SIMILAR TO expression` now plan and execute successfully. No breaking API changes.
) ## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes #123` indicates that this PR will close issue #123. --> - Closes apache#13363 ## Rationale for this change <!-- Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed. Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes. --> Fill in some missing support for utf8view (and date64) in function type coercion. Main affected functions are UDFs that use exact/uniform type signature; for example in `date_bin` it couldnt coerce utf8view to its input types (even though utf8 worked) ## What changes are included in this PR? <!-- There is no need to duplicate the description in the issue here but it is sometimes worth providing a summary of the individual changes in this PR. --> Add support for casting to/from utf8view in some arms. ## Are these changes tested? <!-- We typically require tests for all PRs in order to: 1. Prevent the code from being accidentally broken by subsequent changes 2. Serve as another way to document the expected behavior of the code If tests are not included in your PR, please explain why (for example, are they covered by existing tests)? --> Yes ## Are there any user-facing changes? <!-- If there are user-facing changes then we may require documentation to be updated before approving the PR. --> No <!-- If there are any breaking changes to public APIs, please add the `api change` label. -->
…edian` (apache#23954) ## Which issue does this PR close? - Closes apache#23953 ## Rationale for this change This PR makes three improvements to non-distinct `PercentileContAccumulator` and `MedianAccumulator`, inspired by recent work on `percentile_cont(DISTINCT)` (apache#23946): 1. Switch from SipHash to foldhash for internal hash maps 2. Add a null-free fast path to `update_batch` and `retract_batch` 3. Raise an error if we attempt to retract an unknown value in `retract_batch`, rather than silently ignoring it. Benchmarks: percentile_cont no_nulls window=256 -61.1% (249.0 -> 96.6 us) percentile_cont with_nulls window=256 -59.4% (232.7 -> 94.5 us) percentile_cont no_nulls window=4096 -50.9% (756.3 -> 370.4 us) percentile_cont with_nulls window=4096 -48.2% (552.2 -> 286.2 us) percentile_cont no_nulls window=16384 -47.3% (2.311 -> 1.218 ms) percentile_cont with_nulls window=16384 -42.0% (1.465 -> 0.851 ms) median no_nulls window=256 -62.8% (247.0 -> 92.0 us) median with_nulls window=256 -59.8% (228.8 -> 92.0 us) median no_nulls window=4096 -40.0% (411.4 -> 246.6 us) median with_nulls window=4096 -39.4% (364.4 -> 221.0 us) median no_nulls window=16384 -31.5% (1.107 -> 0.759 ms) median with_nulls window=16384 -32.9% (0.947 -> 0.637 ms) ## What changes are included in this PR? See above. ## Are these changes tested? Yes: existing tests pass, new tests added for the null-free fast path and the improved error handling. ## Are there any user-facing changes? No.
…pache#23956) ## Which issue does this PR close? - Closes apache#23955. ## Rationale for this change `PrimitiveDistinctCountAccumulator::update_batch` does a per-element validity check on every row even when the array has no nulls. Its own `merge_batch` already inserts wholesale from the values buffer, so the same fast path applies to `update_batch`. Inspired by the null-free fast paths added to `percentile_cont`/`median` in apache#23954. ## What changes are included in this PR? - When `null_count() == 0`, insert directly from the values buffer (`self.values.extend(arr.values().iter().copied())`) instead of the per-element `Option` check. - Unit test asserting the fast and general paths agree (a dense array and a null-containing array with the same non-null values yield the same distinct count). ## Benchmarks `count_distinct` benchmark, null-free 8192-row batches: ``` count_distinct i64 80% distinct -67.0% (54.9 -> 18.1 us) count_distinct i64 99% distinct -66.0% (54.9 -> 18.7 us) count_distinct u32 80% distinct -66.9% (53.9 -> 17.8 us) count_distinct u32 99% distinct -65.3% (53.3 -> 18.5 us) count_distinct i32 80% distinct -66.9% (54.1 -> 17.9 us) count_distinct i32 99% distinct -66.2% (53.2 -> 18.0 us) ``` The bitmap-backed cases (u8/i8/u16/i16) and the grouped-accumulator cases don't use this path and are unchanged, as expected. ## Are these changes tested? Yes — new unit test; existing count-distinct tests pass. ## Are there any user-facing changes? No.
…#24817) <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes #123` indicates that this PR will close issue #123. --> - Closes apache#24816. <!-- Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed. Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes. Please explain the problem you are trying to solve in terms of the user-visible behavior, rather than the implementation. For example, "The code in `foo.rs` doesn't handle nulls" is a symptom of the implementation. "COUNT(DISTINCT) returns wrong results when the column contains nulls" is the user-visible problem. --> Aggregate dynamic filtering currently ignores unsupported expressions while still generating a shared scan predicate from supported aggregates. This can prune rows required by the unsupported aggregate and produce incorrect results. <!-- There is no need to duplicate the description in the issue here, but it is sometimes worth providing a summary of the individual changes in this PR. --> - Disable aggregate dynamic filtering when any aggregate expression is unsupported. - Document that aggregate dynamic filtering requires every aggregate expression to be supported. - Update the existing sqllogictest to verify that no dynamic filter is produced for mixed supported and unsupported expressions. <!-- We typically require tests for all PRs in order to: 1. Prevent the code from being accidentally broken by subsequent changes 2. Serve as another way to document the expected behavior of the code Briefly describe how this PR is tested, and point to the specific tests you added. For example: 'This new feature is covered by the `sqllogictest` cases added in `foo.slt`'. If this PR does not add tests, explain why. For example, if the change is already covered by existing tests, please mention it. You should also check the `codecov` bot reply on this PR to confirm the changed code is exercised. --> Yes. Updated the existing sqllogictest case in `push_down_filter_regression.slt`. <!-- If there are user-facing changes then we may require documentation to be updated before approving the PR. If there are any breaking changes to public APIs, please add the `api change` label. --> Yes. This fixes potentially incorrect query results. No API changes. (cherry picked from commit b3cb365) Conflicts resolved when applying to spiceai-54 (DataFusion 54.1): - datafusion/physical-plan/src/aggregates/mod.rs: the base spells the cols_for_dynamic_filter check as debug_assert!(a == b); applied only the upstream rename (supported_accumulators_info -> accumulator_dyn_filter_info) to that line. - datafusion/sqllogictest/test_files/push_down_filter_regression.slt: the base carries an older version of the mixed-expressions section; replaced it with the upstream section verbatim. The upstream aggregate_stream.rs hunks apply to aggregates/no_grouping.rs, its name on 54.1. (cherry picked from commit 6a4a698)
…pache#24428) ChildFilterDescription::from_child (and from_child_with_allowed_indices) built a FilterRemapper before checking whether parent_filters was empty. FilterRemapper::new indexes every column of the child's schema into a HashMap, and with no filters to remap that index is never used: remap_filters returns an empty description regardless of what the remapper contains. Each plan node produces one ChildFilterDescription per child, so a union pays this cost once per child. On a 250,000-column scan a sampling profile put FilterRemapper::new at 794 of 20,373 samples, almost entirely HashMap insertion and hashing. Return ChildFilterDescription::empty() directly when parent_filters is empty, before the remapper is built. ## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes #123` indicates that this PR will close issue #123. --> Do I need an issue for this change, it feels like a really minor fix? Co-authored-by: kosiew <kosiew@gmail.com> (cherry picked from commit 6cc3442) (cherry picked from commit 2c0284f)
) ## Which issue does this PR close? <!-- We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes #123` indicates that this PR will close issue #123. --> - Part of apache#24246 . ## Rationale for this change Bitwise XOR simplifications cancelled repeated nullable operands, which could incorrectly produce a non-NULL result when the operand was NULL. For example: ```sql SELECT i XOR i, (i XOR 7) XOR i, i XOR (7 XOR i) FROM (VALUES (NULL::INT)) AS t(i); ``` These expressions should all return NULL, but simplification could replace them with 0, 7, and 7. <!-- Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed. Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes. Please explain the problem you are trying to solve in terms of the user-visible behavior, rather than the implementation. For example, "The code in `foo.rs` doesn't handle nulls" is a symptom of the implementation. "COUNT(DISTINCT) returns wrong results when the column contains nulls" is the user-visible problem. --> ## What changes are included in this PR? Only cancel repeated XOR operands when the removed operand is non-nullable. <!-- There is no need to duplicate the description in the issue here, but it is sometimes worth providing a summary of the individual changes in this PR. --> ## Are these changes tested? <!-- We typically require tests for all PRs in order to: 1. Prevent the code from being accidentally broken by subsequent changes 2. Serve as another way to document the expected behavior of the code If tests are not included in your PR, please explain why (for example, are they covered by existing tests)? --> Yes. Added sqllogictests covering results. ## Are there any user-facing changes? Yes. Bitwise XOR expressions with repeated nullable operands now correctly preserve NULLs. <!-- If there are user-facing changes then we may require documentation to be updated before approving the PR. If there are any breaking changes to public APIs, please add the `api change` label. --> (cherry picked from commit 50cdbde) (cherry picked from commit 7d9f743)
## Which issue does this PR close? - Closes apache#24379. ## Rationale for this change `simplify_regex_expr` rewrites `col ~ '.*'` to `col IS NOT NULL`. For a NULL input that returns `false`, but `NULL ~ '.*'` is `NULL` under three-valued logic — so the rewrite produces wrong results in a projection context: ```sql SELECT s, s ~ '.*' FROM (VALUES (CAST(NULL AS VARCHAR)), ('x')) t(s); -- NULL row currently returns `false`; it should be NULL ``` The `!~` (`RegexNotMatch`) branch of the same rule is already NULL-aware (`col IS NULL AND NULL`); only the `~` branch dropped the NULL. ## What changes are included in this PR? - Rewrite `col ~ '.*'` to `col IS NOT NULL OR NULL` — `true` for a non-NULL string, `NULL` for a NULL input. - In a WHERE filter both FALSE and NULL reject the row, so filter *results* are unchanged; only the plan text and projection-context values differ. Existing filter-plan expectations in `simplify_expr.slt` and the `test_simplify_regex_special_cases` unit test are updated accordingly, and a projection regression test is added. ## Are these changes tested? Yes. - New projection regression test in `simplify_expr.slt` asserting `col ~ '.*'` returns `true`/`true`/`NULL` for `'foo'`/`''`/`NULL`. - Updated the two filter-context plan expectations (logical + physical) that previously encoded the `IS NOT NULL` rewrite. - `simplify_expr.slt`, the `regexp/*` SLTs, and the optimizer simplify unit tests all pass. ## Are there any user-facing changes? `col ~ '.*'` in a projection now returns `NULL` for a NULL input instead of `false`, matching SQL semantics. No API changes. (cherry picked from commit 16b08db) (cherry picked from commit b90cb07)
## Which issue does this PR close? - No issue has been filed. ## Rationale for this change Anonymous nested projections that reuse the same output name can return wrong results. For example, each `i + 1 AS i` layer must be evaluated independently, but the projection optimizer could drop one layer when two consecutive projection expression vectors were structurally equal. Structural equality does not imply that a projection is safe to elide: repeated computations such as `i + 1 AS i` have the same expression shape but must still be evaluated twice. ## What changes are included in this PR? - Remove the structural-equality fast path that directly elided one of two consecutive projections. - Keep the existing iterative whole-chain merge, so deep projection chains still collapse within one optimizer rule invocation. - Continue using the normal projection rewrite path, which composes repeated expressions and preserves aliases and field metadata. - Add focused regressions for repeated non-idempotent projections, one-pass collapse of a 12-level chain, and metadata-bearing aliases. - Add an execution-level SQLLogicTest with six anonymous `i + 1 AS i` layers under both `max_passes = 1` and the default optimizer configuration. ## What is the testing strategy for this PR? The focused unit tests verify that: - two structurally equal `i + 1 AS i` projections retain both additions; - a 12-level chain preserves all 12 additions and collapses to one `Projection` with `max_passes = 1`; - a metadata-bearing `Alias(Column)` is merged without losing field metadata. The SQLLogicTest executes a six-level anonymous projection chain against a temporary table with both one optimizer pass and the default pass count. Verified with: ```text cargo fmt --all --check cargo test -p datafusion-optimizer optimize_projections # 58 passed; 0 failed cargo test -p datafusion-optimizer --test optimizer_integration # 26 passed; 0 failed cargo test --profile ci -p datafusion-sqllogictest --test sqllogictests -- projection.slt # 1/1 files completed; 0 failures cargo clippy -p datafusion-optimizer --all-targets --all-features -- -D warnings # passed git diff --check # passed ``` ## Are there any user-facing changes? Yes. Deep anonymous nested projections that reuse an output name now preserve every projection expression and return the correct result. There are no public API or configuration changes. --------- Signed-off-by: discord9 <discord9@163.com> Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> (cherry picked from commit 7a1f468) (cherry picked from commit 5b73507)
…che#24997) - Closes apache#25055. `COUNT` does not depend on input order, but it inherited the default `AggregateOrderSensitivity::HardRequirement`. As detailed in apache#25055, this could pass ordering keys to count accumulators as additional arguments, causing incorrect results or a panic, and could introduce unnecessary sorting. Override `Count::order_sensitivity` to return `AggregateOrderSensitivity::Insensitive`. This makes aggregate ordering keys ineffective for physical planning: they are excluded from accumulator inputs and no `SortExec` is required solely for `COUNT`'s `ORDER BY`. Regression tests in `aggregate.slt` cover grouped, multi-argument, and non-grouped counts. The data includes both a null ordering key, which must not affect the count, and a null counted argument, which must still be excluded. A bare global `COUNT(a ORDER BY b)` over the in-memory test table is folded by `AggregateStatistics` to `num_rows - null_count(a)`, so `CountAccumulator` never runs. The test uses `a + 0`, which preserves `a`'s nullness but prevents that fold, ensuring the regression exercises the accumulator. An `EXPLAIN` assertion verifies that the count does not require a sort. `COUNT(... ORDER BY ...)` returns the correct count without panicking. Nulls in counted arguments retain their usual behavior; nulls in ordering keys no longer incorrectly exclude rows. (cherry picked from commit 224cc56) Conflict resolved when applying to spiceai-54 (DataFusion 54.1): - datafusion/sqllogictest/test_files/aggregate.slt: upstream appends this commit's COUNT ... ORDER BY block after other blocks that are not on 54.1, including the nested-aggregate block. Added only this commit's 43 lines, at the end of the 54.1 file. (cherry picked from commit e72572d)
## Which issue does this PR close? - Closes apache#24513. ## Rationale for this change A scalar subquery that returns zero rows evaluates to `NULL`, even when its projected expression is non-nullable. DataFusion currently derives `Expr::ScalarSubquery` nullability from the subquery's output field. This can produce a non-nullable output schema containing `NULL`, and can cause `SimplifyExpressions` to incorrectly fold predicates such as: ```sql SELECT (SELECT 1 WHERE FALSE) IS NULL; ``` to `false`. This change deliberately marks all scalar subqueries nullable, including those guaranteed to return exactly one row (such as an ungrouped aggregate like `(SELECT count(*) FROM t)`). This is a conservative trade-off that gives up some nullability precision for correctness, and is consistent with how PostgreSQL treats scalar subqueries. A possible follow-up refinement is a `LogicalPlan::min_rows()` lower bound (mirroring the existing `max_rows()`), which would let uncorrelated scalar subqueries provably returning at least one row keep their projected field's nullability. ## What changes are included in this PR? Scalar subqueries are conservatively marked nullable in logical expression schema derivation and physical expression planning. The projected field's data type, name, and metadata are preserved. ## Are these changes tested? Yes. Unit tests cover logical schema derivation and expression simplification, and SQLLogicTests cover both zero-row execution and `IS NULL` correctness. The full workspace test suite and Clippy with warnings denied pass. ## Are there any user-facing changes? Yes. Zero-row scalar subqueries with non-nullable projections now return `NULL` without a schema validation error, and `IS NULL` predicates produce the correct result. In addition, output schemas containing scalar subqueries now always mark those fields as nullable, even for subqueries that can never produce `NULL`. Downstream consumers that inspect schema nullability can observe this change. No public APIs change. --- AI usage: Created with Claude Code and Opus 5. I have reviewed the code and made modifications where it made sense. --------- Signed-off-by: Fredrik Fornwall <fredrik@fornwall.net> (cherry picked from commit 4798476) (cherry picked from commit e9dc1dd)
…o few rows (apache#24958) ## Which issue does this PR close? - Closes #. ## Rationale for this change `ORDER BY` applied on top of a `LIMIT ... OFFSET` subquery returns too few rows. ```sql SELECT * FROM (SELECT a FROM t1 LIMIT 4 OFFSET 3) ORDER BY a; ``` `EnforceSorting` pushes the sort below the `GlobalLimitExec` and converts it into a TopK. The TopK's fetch was seeded from `GlobalLimitExec::fetch()`, which is the limit's *output* row count and ignores `skip`. So the TopK kept only 4 rows, the limit then skipped 3 of them, and the query returned a single row instead of 4. The same seeding happens in two places in `sort_pushdown.rs`: when a root's direct child is a `GlobalLimitExec` (`assign_initial_requirements`) and when the pushdown descends through one (`pushdown_sorts_helper`). Both had the bug. ## What changes are included in this PR? - Add an `input_fetch` helper in `sort_pushdown.rs` that returns `skip + fetch` for `GlobalLimitExec` and plain `fetch()` for every other operator. - Use it at both sites where the pushed-down fetch is derived from the plan node, so the TopK below a `GlobalLimitExec` retains enough rows for the limit to skip and still return `fetch` rows. ## What is the testing strategy for this PR? - `limit.slt`: new `EXPLAIN` + result test for `LIMIT 4 OFFSET 3` under `ORDER BY`, showing `TopK(fetch=7)` below `GlobalLimitExec: skip=3, fetch=4` and the correct 4 rows. - `ensure_requirements.rs`: two unit tests, one per code path (`test_sort_pushed_below_limit_with_skip_keeps_skip_rows` and `test_parent_ordering_over_limit_with_skip_keeps_skip_rows`), asserting the pushed sort has `fetch=15` for `skip=5, fetch=10`. ## Are there any user-facing changes? Queries with a sort above `LIMIT ... OFFSET` now return the correct number of rows. No API changes. (cherry picked from commit 16ace4f)
…25227) ## Which issue does this PR close? - Closes apache#25221. ## Rationale for this change `MIN`/`MAX` over a cast can return incorrect results when aggregate optimization replaces the scan with converted column statistics. Casting the original endpoints is only valid if they remain extrema in the target domain. For example, a Parquet string column containing `('1', '100', '2')` has string extrema `'1'` and `'2'`. Previously, `MIN(CAST(a AS INT)), MAX(CAST(a AS INT))` could return `1, 2` from those statistics instead of the correct `1, 100`. The same issue affects `BIGINT`. ## What changes are included in this PR? - Restrict CAST statistics propagation to conversions proven safe for the source type or its exact integer bounds. Unsupported or unsafe conversions return unknown column statistics, so aggregates evaluate the data. - Preserve the existing same-type passthrough and supported lossless widening conversions. - Retain integer narrowing and signed/unsigned conversions when both original bounds are exact and non-null, and both converted endpoints are non-null. Successful endpoint conversions establish that the entire integer range fits in the target type; overflow errors or null endpoints do not qualify. - Extend `CastExpr::check_bigger_cast` to recognize `Int32 ↔ Date32`, and clarify that its contract includes lossless, strictly order-preserving conversions that preserve nulls. Arrow reinterprets the same `i32` values as days since the epoch, so these conversions preserve both statistics and ordering properties. This also retains the statistics-based `MIN`/`MAX` optimization for ClickBench's `UInt16 → Int32 → Date32` projection. ## What is the testing strategy for this PR? - Add projection unit tests for non-monotonic string-to-integer casts and integer narrowing with safe, overflowing, and inexact bounds. Removing the statistics guard makes the new string-cast regression test fail. - Add end-to-end Parquet regression cases in `parquet_statistics.slt` for `INT` and `BIGINT`, both expecting `1, 100`. - Add a bidirectional `Int32 ↔ Date32` test covering `i32::MIN`, `i32::MAX`, negative values, zero, NULL, and strict ordering properties. - Run the related projection and cast unit tests, plus the ClickBench and Parquet statistics SLT files. The existing ClickBench plan expectation passes unchanged, confirming that its aggregate still folds to constants. Local formatting, Clippy, license, spelling, workflow, and documentation checks also pass. ## Are there any user-facing changes? Affected casts now produce correct `MIN`/`MAX` results. Conversions whose safety is not established may require a scan instead of using statistics-based aggregate optimization. Supported safe integer conversions and `Int32 ↔ Date32` retain that optimization. (cherry picked from commit ccfe704)
…pache#25259) - Closes apache#25244. - Closes apache#25262. - Closes apache#25263. - Closes apache#25264. - Closes apache#21246. - Closes apache#25296. Enabling `datafusion.optimizer.enable_join_dynamic_filter_pushdown` can silently drop matching rows when the probe side of a join contains several columns with the same name, for example `a.id` and `b.id` from a nested join. Physical filter pushdown resolved a pushed filter's columns in the child schema by name, so a predicate on the second `id` column was rewritten to the first `id` column, which holds a different value. The same name-based mapping also affects TopK dynamic filters under the default configuration (apache#25296). For a nested join of orders and payments that both expose `amount`, `ORDER BY p.amount LIMIT 1` can push the payment threshold onto `orders.amount`. Later matching orders are incorrectly pruned, returning `(100, 30)` instead of `(300, 10)`. Every operator that forwards parent filters had the same weakness, in slightly different forms: - `HashJoinExec` computed which output columns belong to each side by position, but then resolved the child column by name. - `FilterExec`, `SortExec`, `RepartitionExec`, `CoalesceBatchesExec`, `UnionExec` and similar nodes used the generic name lookup even though their output positions equal their input positions. - `FilterExec` with an embedded projection resolved unconsumed parent filters back into input coordinates by name. - `ProjectionExec` looked up output aliases by name to find the expression to substitute, so two outputs aliased `id` both mapped to the first. - `AggregateExec` restricted pushdown to grouping positions but resolved the input column by name. The join-filter regression queries return one row with dynamic filtering disabled and zero rows with it enabled across these shapes. The `FilterExec`, `ProjectionExec` and `AggregateExec` cases were found while auditing the remaining name-based paths for apache#25244; they share the root cause, so this PR fixes them together. The TopK case in apache#25296 is fixed by the same positional mapping. Built-in operators now remap filter columns by position. The deprecated public API retains its historical name-based resolution for compatibility. - `FilterRemapper` uses a `ColumnMapping` enum: `Identity` preserves positions and checks that names match; `Explicit` uses a caller-supplied parent-output to child-input mapping and allows names to differ. - `ChildFilterDescription::from_child` now maps by position. New `ChildFilterDescription::from_child_with_column_mapping` takes an explicit `HashMap<usize, usize>`. - `HashJoinExec` builds the explicit mapping from its `column_indices` and output projection. For semi joins, output join keys on the emitted side are mapped to the paired key on the other side, which also supports differently named keys. - `FilterExec` maps through its embedded projection in both pushdown phases and when folding unconsumed parent filters back into its predicate. - `ProjectionExec` substitutes the expression at each output position instead of looking the alias up in the output schema. - `AggregateExec` maps each grouping output position to the input column that grouping expression reads; non-column grouping expressions are not forwarded, as before. - `ChildFilterDescription::from_child_with_allowed_indices` is kept as a deprecated wrapper that preserves name-based resolution to the first matching child field. It translates allowed parent column references into an explicit positional mapping; new callers should supply positions directly to avoid ambiguous duplicate names. - `dynamic_filter_pushdown_config.slt` gains four regression queries over two small Parquet tables, run once with join dynamic filtering disabled and once enabled, asserting identical rows: nested joins with `RepartitionExec` between them (the query from apache#25244), a `FilterExec` with an embedded projection, a `ProjectionExec` with duplicate aliases, and an `AggregateExec` grouping on same-named columns. All four return zero rows on `main` with dynamic filtering enabled. - `dynamic_filter_pushdown_config.slt` also covers apache#25296 using nested joins of orders, payments and customers, with one row per batch and one row per orders Parquet row group. The query returns `(300, 10)` with TopK dynamic filtering disabled, enabled, and enabled together with Parquet filter pushdown. On `main` at `85d4cbb0a9`, both enabled cases reproduce the incorrect `(100, 30)` result; all three pass on this branch. - `datafusion/core/tests/physical_optimizer/filter_pushdown.rs` gains focused tests that call `gather_filters_for_pushdown` directly on `RepartitionExec`, `HashJoinExec` (duplicate child columns with and without projection, both semi join directions with differently named keys, and a semi join whose key is not a plain column), `FilterExec` with a projection in both phases, `ProjectionExec` with duplicate aliases, and `AggregateExec` with reordered same-named grouping columns. Further tests cover the deprecated `from_child_with_allowed_indices` wrapper preserving the first name match, accepting allowed parent indices outside the child schema, and rejecting unresolvable names, the identity mapping rejecting a column whose name differs from the child field at that position, and `FilterExec` rejecting an unconsumed parent filter outside its projection. Passed locally on the updated implementation: - `cargo fmt --all` - `./ci/scripts/doc_prettier_check.sh --write --allow-dirty` - `cargo test --profile ci -p datafusion --test core_integration physical_optimizer::filter_pushdown` — 70 tests passed. - `cargo test --profile ci --test sqllogictests -- dynamic_filter_pushdown_config.slt` — passed. - The isolated apache#25296 regression fails on `main` in both TopK-enabled configurations with the expected wrong-result mismatch. Queries that push join or TopK dynamic filters through operators with duplicate column names now return the correct rows, including the default-configuration TopK query in apache#25296. API change in `datafusion-physical-plan`: `ChildFilterDescription::from_child_with_allowed_indices` is deprecated in favour of `from_child` and the new `from_child_with_column_mapping`; the deprecated function preserves its previous name-based behavior. Callers should migrate to explicit positions because duplicate names make name resolution ambiguous. `ChildFilterDescription::from_child` also resolves by position, which only affects callers that used it on a node whose output positions differ from its child's. (cherry picked from commit b376290)
…ubquery (apache#25348) - Closes apache#25340. A follow-up issue covers the remaining `NOT IN` problems. See **Follow-up work** at the end. `3 NOT IN (1, NULL)` is UNKNOWN. A `WHERE` clause must remove the row. DataFusion kept every row. There was no error and no warning. ```sql CREATE TABLE t1(id INT) AS VALUES (1), (2); CREATE TABLE t2(id INT) AS VALUES (1), (NULL); SELECT id FROM t1 WHERE 3 NOT IN (SELECT id FROM t2) ORDER BY id; ``` | | result | |---|---| | DataFusion (before) | `1, 2` ❌ | | DataFusion (after) | *(no rows)* ✅ | | DuckDB 1.5.2 | *(no rows)* | | PostgreSQL 17.11 | *(no rows)* | The subquery gives `{1, NULL}`. The value `3` is not `NULL`, but `3` is also not known to be absent. Therefore the answer is UNKNOWN for every row of `t1`. The same expression in a `SELECT` list was already correct. Only the `WHERE` clause was wrong. The value `3` holds no column. Therefore it could not become a join key. It stayed as a join filter. Two problems followed. ``` BEFORE AFTER WHERE 3 NOT IN (SELECT id FROM t2) WHERE 3 NOT IN (SELECT id FROM t2) │ │ ▼ ▼ anti join, no key anti join, key = (3, t2.id) filter: 3 = t2.id no filter │ │ ├─ (1) the filter moves into the ├─ (1) nothing to move │ subquery: WHERE t2.id = 3 │ │ the NULL row disappears │ │ │ └─ (2) no key, so the join becomes a └─ (2) the join has a key, so it nested loop join, which cannot stays a hash join, which do null-aware work. The flag is does null-aware work dropped without an error correctly │ │ ▼ ▼ 1, 2 ❌ (no rows) ✅ ``` Three changes. **1. Make the constant a join key** (`datafusion/optimizer/src/decorrelate_predicate_subquery.rs`) Add the constant to the outer side as a column. The comparison then becomes a true equality of two columns, so the existing null-aware hash join does the work. ``` Before: LeftAnti Join: Filter: Int64(3) = __correlated_sq_1.id null_aware → NestedLoopJoinExec (the null_aware flag is lost) After: LeftAnti Join: __correlated_sq_1_value = __correlated_sq_1.id null_aware Projection: t1.id, Int64(3) AS __correlated_sq_1_value → HashJoinExec ... null_aware ``` This applies only to uncorrelated subqueries. A correlated subquery needs a second key, and a null-aware anti join accepts only one key. **2. Keep the filter out of the subquery** (`datafusion/optimizer/src/push_down_filter.rs`) Do not move a predicate into the subquery side of a null-aware join. The NULLs must reach the join. `infer_join_predicates` has the same rule already. **3. Report an error instead of a wrong answer** (`datafusion/core/src/physical_planner.rs`) Only a hash join can do null-aware work, and it needs a key. If a null-aware join has no key, report an error. Do not build a nested loop join that gives wrong results without a warning. New `sqllogictest` cases in `datafusion/sqllogictest/test_files/null_aware_anti_join.slt` and `null_aware_mark_join.slt`. They cover: - the four queries of the issue, and its controls; - a subquery with no NULL, which must keep all rows; - the `OR` and `IS NULL` forms, which use a mark join; - a user column with the same name as the new column, which must not be ambiguous; - the one shape that is not supported, which must report an error. New unit tests in `decorrelate_predicate_subquery.rs` and `push_down_filter.rs` hold the new plans. Results: | check | result | |---|---| | full `sqllogictest` suite | pass, and **no plan changes** anywhere else | | `datafusion-optimizer` unit tests | 853 pass | | extended workspace suite | 10,784 tests pass, 0 fail, 66 crates | | `cargo fmt --all` | clean | | `cargo clippy` | clean for the changed crates | The extended workspace figure comes from an earlier run of the same commit on a slightly older base. The full `sqllogictest` suite and the optimizer tests were re-run after the rebase onto current `main`. Yes. Two. **1. `<constant> NOT IN (<subquery>)` in a `WHERE` clause now gives correct results.** This is the fix. **2. One shape now reports an error.** A constant with a correlation that is not an equality: ```sql SELECT id FROM t1 WHERE 3 NOT IN (SELECT id FROM t2 WHERE t2.g > t1.g); ``` ``` Error during planning: null_aware LeftAnti join requires equi-join keys, but the join has none ``` This query gave wrong results before. No correct result is lost. An error is better than a wrong answer that a user cannot see. There are no API changes. This PR makes every **uncorrelated** `NOT IN` correct. **Correlated** `NOT IN` still has problems. They come from the join operator, not from the code this PR changes. ``` NOT IN (subquery) │ ├── no outer column ──────────────▶ CORRECT (this PR) │ └── reads an outer column ├── equality condition ───────▶ WRONG or ERROR (follow-up, parts 1 and 2) └── other condition ──────────▶ WRONG (follow-up, part 3) ``` A test matrix of 18 query shapes gives these totals: | | wrong or error | |---|---| | before this PR | 13 of 18 | | after this PR | 10 of 18 | The 10 remaining shapes are all correlated. They fall into three parts: 1. A null-aware anti join accepts only one key, but a correlated `NOT IN` needs two. 2. A constant does not become a key when the subquery is correlated. 3. A null-aware join looks for NULLs only in the key, not in a leftover filter. Part 3 must come first, as an error. A prototype showed that part 2 alone turns a clear error into a silent wrong answer. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VCeMPyJNgAiFpGXaz5CxF3 Co-authored-by: Claude <noreply@anthropic.com> (cherry picked from commit edc936f)
…als in a filter
A correlated `NOT IN` in anti join position was decorrelated into a null-aware
anti join. On this branch the null-aware hash join takes a single key and checks
the subquery's NULLs over every build row, so after equi-join extraction:
- a constant `NOT IN` with an equality correlation got the correlation column
as its only key, and the NULL rules applied to it: a NULL correlation key
dropped rows although the subquery held no NULL;
- a column `NOT IN` with an equality correlation had two keys and failed to
plan ("null_aware anti join only supports single column join key");
- with a non-equality correlation, a subquery NULL that the correlation
excludes still made every outer row UNKNOWN;
- a constant `NOT IN` with a non-equality correlation had no key and failed to
plan (the check apache#25348 added).
In a filter, `x NOT IN (SELECT y ... WHERE c)` keeps a row exactly when no
subquery row satisfying `c` has `x IS NULL OR y IS NULL OR x = y`. Build that
filter and a plain anti join for a correlated `NOT IN`; an uncorrelated one
keeps the null-aware hash join, which is correct for it. With non-nullable
columns the disjunction simplifies to `x = y`, which stays an equi-join key.
When the pull-up drops a subquery filter as a duplicate of the `IN` equality
(`... WHERE y = x`), the subquery kept only values equal to `x`, none NULL, so
the equality alone decides: `x = y` and the correlation, not null-aware. The
null-aware plan got this wrong when no other correlation was left, treating the
subquery as uncorrelated. `PullUpCorrelatedExpr` now records the drop.
The `NOT EXISTS` form is not used when the pull-up grouped a scalar aggregate by
the correlation, which loses the row the aggregate returns over no input, nor
for a volatile `x`, which the filter would evaluate more than once. Those keep
the null-aware plan, which does not support the aggregate shape.
Upstream evaluates these shapes in the null-aware join executor instead
(apache#25339, apache#25560), which builds on the null-aware mark joins of
the branch moves to a release that has them.
`not_in_correlated.slt` checks 9 shapes over 9 datasets against SQLite's
answers (46 of the 81 queries were wrong or failed to plan before this change,
0 after), the two dropped-duplicate shapes, and that the scalar aggregate shapes
still fail to plan instead of answering from missing rows. The apache#25348 case that
expected the planning error now expects the row SQL returns.
(cherry picked from commit 0fcd17a)
…epeated, gated by dialect (refs spiceai/spiceai#12751) (#227) (cherry picked from commit 8ba22a6)
…on it returns (#230) JoinType::RightMark returns a record for each record from the right input, so like RightSemi and RightAnti its probe side is join.right. It was missing from swaps_join_inputs, so the join arm built the outer query from join.left and the EXISTS body from join.right — probe and build reversed. The mark placeholder's qualifier followed from the same confusion: it read join.left for RightMark, which happened to be the build side only because the swap was not being applied. After the swap the build side is right_plan for both mark joins, which is the plan build_exists_subquery is already given, so the special case goes. Refs spiceai/spiceai#13022 Co-authored-by: claudespice <270518434+claudespice@users.noreply.github.com> (cherry picked from commit ec92649)
…ojection binds (refs spiceai/spiceai#13493) (#232) * fix(unparser): refuse an EXISTS build-side key that only the build projection binds (refs spiceai/spiceai#13493) build_exists_subquery replaces the build side's projection with SELECT 1. A build-half join key that only that projection bound - an output qualified with a relation the build side does not scan, or named something the relation it is qualified with does not have - then has nothing left to bind to inside the body, and binds outward instead: on the reported plan to the outer query, where both halves name the outer row and the comparison is a tautology, so a semi join keeps every probe row and an anti join drops them all, in valid SQL. The capture guard tests only the correlated half of each pair on purpose, so this asks the opposite question of the other half, and only where the plan proves it: a qualifier the emitted FROM does not introduce, or one naming a scan or alias that has no such column. An alias pushed down onto a scan answers to the scan's columns, its renames having been lifted into the select list SELECT 1 replaces. An unqualified key is kept: the body's exposed names include the erased outputs, so that shape needs the emitted FROM (spiceai/spiceai#13469). * fix(unparser): read an unpushed alias's derived table by its own schema The pushdown leaves an alias over a bare scan as written, so the projection becomes the derived table's own and the rename it makes is what the alias answers to. Mirror that arm's conditions instead of assuming every alias over a scan is pushed down, which refused a key that binds. Pinned by the unpushed-alias shape. * refactor(unparser): share the column check across the build-half walk's arms Quality-only: the scan and alias arms each set the flag, exposed the schema's columns and looked the name up; the match now yields the matching relation's schema and the check runs once after it. No assertion or snapshot changed. * fix(unparser): read an inline alias over a join by the relation it renames SQL has no syntax for naming a join, so an alias over one that the emitter does not wrap in a derived table renames the primary relation in place and leaves the rest named as they were. The alias's schema requalifies every output of the join, so reading it kept a build-half key that names a column of a relation beside the renamed one, which the emitted body cannot bind. schema_an_alias_answers_to now walks to the relation the emitter lands the alias on - derived table, pushed-down scan, or the primary relation of an inline join - and answers with that relation's columns. Pinned in both directions. * fix(unparser): a projection on an inline alias's primary side is the table the alias renames The alias arm takes the select list before it walks the join, so a projection met on the way is emitted as a derived table of its own and it is that table relation.alias renames. Walking through it to the scan below read the scan's columns and refused a key naming the projection's rename, which binds. Pinned. --------- Co-authored-by: grokspice <grokspice@users.noreply.github.com> (cherry picked from commit f20e0f3)
…input is a join (refs spiceai/spiceai#12593) (#231) * fix(unparser): keep a FULL JOIN input's scan filters scoped when the input is a join (refs spiceai/spiceai#12593) A FULL JOIN preserves both inputs, so a predicate from either one has no clause of the enclosing query: ON lets the rows it rejects reappear as unmatched, and WHERE discards the other input's unmatched rows. The join already isolates a bare scan's filters in a derived table, but an input that is itself a join is walked by the nested Join arm, which routes its scans' filters to the shared WHERE as it would anywhere else. The FULL JOIN now marks the shared SelectBuilder while its inputs are walked. A join nested inside such an input derives each filtered scan the way the FULL JOIN derives its own, keeps an inner join's subquery conjunct in ON rather than moving it to WHERE, and a predicate that no scan applies (a Filter above the nested join, a semi or anti join's EXISTS) is refused rather than emitted where it would drop rows. After the walk the FULL JOIN checks that its inputs contributed nothing to WHERE, so any path the rules miss is an internal error, not wrong rows. * fix(unparser): leave the predicate above a FULL JOIN in place while its inputs are walked Taking it out changed what the walk saw: a row-limited input read an empty WHERE and put its LIMIT on the enclosing query, and a mark join had nothing to substitute its EXISTS into. Count the predicates added through SelectBuilder::selection instead, so a contribution from an input is detected without moving the predicate that was already there. The RIGHT JOIN hoist puts its predicate back through restore_selection, which does not count. * refactor(unparser): read the FULL JOIN mark once, shorten the new tests Quality-only: the join arm reads the enclosing-FULL-JOIN mark into one local, the new tests import the join types they use, and two clones whose value was never used again are gone. No snapshot or assertion changed. * fix(unparser): do not hoist a nested RIGHT JOIN's predicate below a FULL JOIN Below a FULL JOIN's input the marked walk keeps every predicate in its own scope, so the RIGHT JOIN hoist has nothing to relocate — and setting the enclosing predicate aside for its walk hid it from a row-limited left input, which then put its LIMIT on the enclosing query, and from a mark join, which had nothing to substitute into. Pinned by the nested RIGHT JOIN shape added to the existing test. * fix(unparser): derive a filtered scan behind a folded projection below a FULL JOIN A FULL JOIN input shaped Projection(TableScan with filters) is not a simple scan to the join walk, so the projection is folded into the shared SELECT and the scan is reached with its filters still on it. The pushdown rewrite then materialized them as a Filter, which the marked Filter arm refused. The scan now keeps its filters (and fetch) in a derived table of its own, the way the join derives a bare filtered input; the folded projection still names the scan, which the derived table is aliased as. Pinned directly under the FULL JOIN and inside a nested join. * fix(unparser): derive a marked scan's filters only once the select list is taken It is the folded projection that lists the scan's columns; without one the pushdown rewrite is what lists them, so a filtered scan reached some other way below a FULL JOIN keeps the rewrite and the refusal, rather than a derived table whose columns nothing selects. Pinned by a limited filtered scan with no projection above the join. * fix(unparser): keep a marked scan's filters scoped through a plain alias A filtered scan below a FULL JOIN keeps its filters in a derived table of its own, but only when the join walk reaches the scan directly. Put a plain alias between the folded projection and the scan and the pushdown rewrite runs first, materializing the filters as a `Filter` over the aliased scan — which the `Filter` arm refuses, so the alias alone turned a supported plan into `NotImplemented`. Take the same derived table for the aliased scan, named by the alias the enclosing scope already uses. The filters are rebased onto that alias first: they still name the scan's own table, which the alias shadows inside the derived table, so `a.id` under `FROM a AS s` would bind to nothing. A filter holding a subquery keeps the refusal. `TableAliasRewriter` rewrites ordinary column references but does not descend into a subquery's plan, so a reference the subquery makes to the scan's table would outlive the alias and be emitted unbound — worse than the refusal it replaced. * fix(unparser): refuse an aliased scan with a subquery filter on a join side reached directly The join arm peels a `SubqueryAlias(TableScan)` input to a simple scan before recursing, so the `SubqueryAlias` arm's refusal of a subquery filter under a shadowing alias never ran for a side reached directly; the FULL JOIN scope path then derived `FROM a AS s WHERE EXISTS (... a.id ...)`, where `a` binds to nothing. The same refusal now runs in the join arm, for either side, and the un-projected shapes are pinned. * fix(unparser): name the aliased-scan subquery refusal for what it is The scan does apply the predicate; what cannot be done is rebasing a filter that holds a subquery onto the alias, since the rewriter does not descend into the subquery's plan. Both paths that decline the derived table — the join arm for a side reached directly, the `SubqueryAlias` arm for one reached through a projection — now say so, instead of falling through to the generic FULL JOIN predicate refusal. * fix(unparser): refuse a mark join as a FULL JOIN input The `EXISTS` that replaces a mark column is never NULL, but the mark is on a row the FULL JOIN null-extends, so `NOT x.mark` or `x.mark IS NULL` above the join would keep or drop that row differently from the plan. A mark join as a FULL JOIN input is refused with the semi and anti joins; the three predicate shapes are pinned. --------- Co-authored-by: grokspice <grokspice@users.noreply.github.com> Co-authored-by: claudespice <270518434+claudespice@users.noreply.github.com> (cherry picked from commit f9bd47d)
…ified DFSchema The 54 patch constructed TableAliasRewriter from the source's arrow schema; 55 takes a DFSchema and a rewrite_unqualified flag. Mirrors the pushdown rewrite it references.
- import GlobalLimitExec at module level for input_fetch (upstream had it from apache#24932) - define the nullable_scalar_mark_scan helper the apache#25348 tests use - build partition bounds with MinMaxColumnBounds, and name the HashJoinExec accumulator - pass NullEquality to estimate_join_cardinality in the join statistics tests, which did not compile against the 5-argument signature
It tests null-aware mark joins from apache#21585, which this line does not have, so a constant IN / NOT IN that decorrelates to a mark join (inside an OR, or under IS NULL) is not fixed here. The output of 'SELECT ... WHERE (id NOT IN (subquery with a NULL)) IS NULL' is identical before and after this branch's patches: no rows, where SQL returns the rows whose id is not in the set.
…g only the backports' own hunks The earlier resolutions of these two files took the 54 version whole, which dropped test content that exists only on 55. Restore the 55 files and apply the test changes of apache#24997 and apache#24817 from upstream.
9 of 58 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Brings the
spiceai-54line to DataFusion 55.1.0 and mergesspiceai-54back in (-s ours), re-applying what 55.1.0 does not contain.spiceai-54 commits since the branch was cut
#24958,#25227,#25259,#25348taken from upstream main)NOT IN→NOT EXISTSrewritenull_aware_mark_join.sltTests (same runs on the untouched df55 tip for comparison)
EXPLAINtext drift (accumulator=MinMaxLeftAccumulator); 0 new failures.not_in_correlated.sltpasses.datafusion-physical-planunit: the same 9 fail on both (tests did not compile at the tip; fixed the compile errors).core_integration physical_optimizer: the same 41 fail on both.(id NOT IN (subquery with NULL)) IS NULLand... OR ...return wrong rows (fixed upstream by feat: Support null aware hash mark-joins apache/datafusion#21585).Patch survival
fetchpushed into a join input keeps its own scope (fork PR #197)crates/data_components/src/federation.rs::a_fetch_pushed_into_a_join_input_survives_unparsingFilterabove aLimitkeeps the limit scoped (fork PR #198)…::a_filter_above_a_limit_keeps_the_limit_scopedORDER BYkept out of a derived table when the sort key is computed (fork PR #191)…::a_computed_sort_key_keeps_order_by_at_the_top_level…::a_stacked_aggregate_keeps_its_inner_group_by,…::a_grouped_stacked_aggregate_binds_its_outer_clauses_through_the_derived_scopeEXISTSbuild side is scoped, so its limit selects rows (fork PR #201)…::a_bounded_exists_build_side_is_scoped_outside_the_correlation,…::an_offset_only_exists_build_side_is_scoped_outside_the_correlation,…EXISTSbound that cannot be scoped, rather than emit wrong rows (fork PR #205)…::a_bounded_exists_refuses_a_correlation_naming_two_build_inputs…::a_derived_tables_unnamed_outputs_are_namedEXISTScorrelation the emittedFROMwould capture, at any bound (fork PR #207)…::an_unbounded_exists_refuses_a_correlation_shadowed_by_its_build_relation,…ProjectionemitsSELECT 1…::an_empty_projection_does_not_unparse_to_an_empty_select_listAT TIME ZONEfaithfully unparsed (fork PR #160), and suppressed for fixed-offset timezones on DuckDB……::a_timezone_survives_unparsing_except_where_the_engine_cannot_resolve_itFLOAT64notDOUBLE, timestamp literal format,date_field_extract_style/interval_style……::bigquery_names_the_float_type_the_way_bigquery_does,…::bigquery_attaches_a_timestamp_offset_to_the_time,…DISTINCTquantifier for distinct unionscrates/runtime-datafusion/src/dialect/bigquery.rs::the_wrapper_emits_valid_bigquery_set_and_window_syntax; real-engine guard:…crates/runtime-datafusion/src/dialect/bigquery.rs::the_wrapper_emits_valid_bigquery_set_and_window_syntax; real-engine guard:…Dialect::range_window_default_nulls_first, and the leadingIS NULLkey a dialect that reports one…crates/runtime-datafusion/src/dialect/bigquery.rs::the_wrapper_forwards_every_bigquery_specific_rendering(the RANGE-window arm), and in the fork…DISTINCT ONoutput alias does not capture its own key (fork PR #209)crates/data_components/src/federation.rs::a_distinct_on_output_is_not_named_over_its_own_key, with…Dialect::group_by_matches_select_subexpressions, and the aggregate scope a dialect that answers…crates/data_components/src/federation.rs::a_projection_wrapping_a_grouping_expression_keeps_the_aggregate_scoped, which reads the flag, so losing…ON— toWHEREon an inner join, and into the…plan_to_sql.rs::test_a_subquery_in_an_inner_join_predicate_moves_to_whereand…return_field_from_argscrates/runtime-datafusion/src/dialect/bigquery.rs::the_wrapper_forwards_every_bigquery_specific_rendering(the truncated-date-comparison arm); in…Dialect::string_to_date_to_sql, and theBigQueryrendering that parses text before narrowing it to a…crates/runtime-datafusion/src/dialect/bigquery.rs::the_wrapper_forwards_every_bigquery_specific_rendering(the text-parsed-into-a-date arm)crates/runtime-datafusion/src/dialect/bigquery.rs::the_wrapper_forwards_every_bigquery_specific_rendering(the text-compared-with-a-timestamp arm)crates/runtime-datafusion/src/dialect/bigquery.rs::a_recursive_cte_renders_through_the_wrapper_only_where_it_is_supportedand…Dialect::supports_recursive_cte, and theWITH RECURSIVErendering behind it (#219)crates/runtime-datafusion/src/dialect/bigquery.rs::a_recursive_cte_renders_through_the_wrapper_only_where_it_is_supported, which asserts both the…Dialect::string_to_timestamp_to_sql, and theBigQueryrendering that parses text into a…crates/runtime-datafusion/src/dialect/bigquery.rs::the_wrapper_forwards_every_bigquery_specific_rendering(the text-parsed-into-a-timestamp arm,…FILTER (WHERE …)on an aggregate is rewritten, and only for aggregates the rewriting is…crates/runtime-datafusion/src/dialect/bigquery.rs::the_wrapper_forwards_every_bigquery_specific_rendering(the filtered row-count, filtered-sum,…LEAD/LAGwhile aggregate frames are kept (fork PR #216)crates/runtime-datafusion/src/dialect/bigquery.rs::the_wrapper_forwards_every_bigquery_specific_rendering(theLEAD-frame arm, with a…crates/runtime-datafusion/src/dialect/bigquery.rs::the_wrapper_forwards_every_bigquery_specific_rendering(the instant-vs-date and…Dialect::decimal_type_to_sql, and the unparameterizedNUMERIC/BIGNUMERICBigQuery selects through…crates/runtime-datafusion/src/dialect/bigquery.rs::the_wrapper_forwards_every_bigquery_specific_rendering(the decimal, integer-width,…plan_to_sql.rs::test_a_sort_over_a_scoped_aggregate_names_the_scope_not_its_grouping_expr_location,_last_modified,_size) onListingOptions/FileScanConfig, and their…crates/data-connector-api/src/listing/connector.rs(metadata-column tests)ListingOptions(with_object_versioning_type), forwarded throughDFParquetMetadata…crates/data-connector-api/src/listing/connector.rs::a_versioned_parquet_read_pins_every_request_to_one_object_version,…crates/data-connector-api/src/listing/connector.rs::a_second_reader_for_the_same_file_keeps_the_version_the_first_one_pinnedfor the factory, which…Expr::infer_placeholder_types, incl.CASE,LIMIT/OFFSETInt64, name/metadata…crates/runtime/src/datafusion/query.rs::every_shape_the_fork_patches_cover_infers_its_parameter_type,…DATETIMEnotTIMESTAMP, a…crates/runtime-datafusion/src/dialect/bigquery.rs::the_wrapper_forwards_every_bigquery_specific_rendering(the four#212arms:DATE_DIFF,…crates/runtime/src/datafusion/builder.rs::tests::the_built_session_concatenates_an_untyped_nullcrates/runtime/src/flight/flightsql/statement_substrait_plan.rs::tests::decode_plan_executes_a_varchar_literal;…extractenum arguments lower todate_partcast to the plan's declared output type (fork PR #220)tools/substrait-compliance/src/mode_a.rs::enum_function_argument_lowers_to_date_part,::registered_extract_udf_takes_precedence_over_date_part(a…crates/data_components/src/federation.rs::a_projected_scan_under_an_alias_names_the_output_its_scope_references(a pushed-down scan projection…tools/substrait-compliance/src/mode_a.rs::correlated_subquery_over_the_same_table_keeps_its_rows(same-table correlated EXISTS over three rows…DIV(fork PR #222)crates/runtime-datafusion/src/dialect/bigquery.rs::the_wrapper_forwards_every_bigquery_specific_rendering(the integer cohort hours arm)crates/runtime-datafusion/src/dialect/bigquery.rs::filtered_recursive_join_inputs_keep_their_qualified_columnsand…crates/runtime-datafusion/src/dialect/bigquery.rs::recursive_column_list_survives_a_join_alias; real-engine guard:…datafusion/physical-optimizer/src/eager_aggregation.rs, ~3000 lines,…CollectLeftAccumulatorseam onHashJoinExeccrates/cayenneVerified from the
spiceai/spiceaiworkspace pinned at this head:cargo check --workspace(lint-gate features plus the data-connector features) passes, and the guard tests named below were run withcargo test --lib.