Is your feature request related to a problem or challenge?
Proposed by @2010YOUY01 in #23599 (comment).
sqllogictest runs under exactly one physical optimizer configuration: the default rule list. That makes a whole class of plan shapes untestable end-to-end, because the shape only arises under a different rule order.
The concrete case that surfaced this: WindowTopN sits at position 117 in the default list and ProjectionPushdown at 142. A FilterExec carrying an embedded projection therefore never reaches WindowTopN from a stock pipeline — but it does in downstream pipelines that re-run projection pushdown earlier. #21596 asks WindowTopN to handle that shape, and its premise ("this happens when ProjectionPushdown runs before WindowTopN") is not true of the default list.
This leaves us with a bad choice. Either a rule grows logic for shapes no slt can produce — untested in practice, and an open-ended maintenance burden as "handle every possible shape" is not a realistic bar for an early-phase rule — or downstream engines carry local patches for orderings that upstream never exercises.
The underlying tension, in @2010YOUY01's words: is the public API the ordered default rule list, or is it each individual rule? Today we rely on an ordering nobody wrote down.
Describe the solution you'd like
Run the slt corpus under a second, named optimizer pipeline, so an alternative rule order is a first-class thing upstream exercises rather than an assumption downstream engines rely on.
Most of the machinery already exists. # configMatrix: (#24493) re-runs a single slt file across a config sweep and attributes failures to the combination that produced them:
# configMatrix: datafusion.optimizer.prefer_hash_join=true,false
# configMatrix: datafusion.execution.batch_size=1,2,100,8192
Five files use it today (sort_merge_join_matrix.slt, piecewise_merge_join_matrix.slt, mark_join_matrix.slt, ...). What it sweeps is ConfigOptions key/value pairs, not rule lists.
So the missing piece is a config option that selects a pipeline variant, e.g.
datafusion.optimizer.pipeline = default | reoptimize
where reoptimize is the default list plus a few re-optimization passes (a second ProjectionPushdown, etc.) — chosen to be closer to what downstream engines actually run. prefer_hash_join is precedent for a config option that changes the physical plan. Files then opt in with:
# configMatrix: datafusion.optimizer.pipeline=default,reoptimize
and the existing matrix machinery handles the rest.
Two things this buys, which are worth keeping separate:
- Results agree. Sweeping the corpus proves the alternative pipeline computes the same answers everywhere. This is the part that scales, and the part that makes rule-level shape handling maintainable instead of speculative.
- What capability does it actually buy? This needs its own plan assertions — a sweep can't show it. Note the existing matrix files carry almost no
EXPLAIN (sort_merge_join_matrix.slt has one), because a different rule list produces different plans and slt has no per-configuration expected output. So plan-shape coverage stays in targeted tests, and the sweep covers correctness.
Describe alternatives you've considered
- Per-rule unit tests that construct the shape by hand. This is what's done today. It's a circular dependency — upstream UTs added to guard downstream use cases — and it can't catch interactions between rules.
- A custom optimizer pipeline running a single query in a Rust test. Weaker than
slt: it covers one query, and nothing keeps it in sync as the rule list evolves.
- Enumerating all orderings of the rule list. Combinatorially hopeless, and most orderings are not shapes anyone runs.
Additional context
Is your feature request related to a problem or challenge?
Proposed by @2010YOUY01 in #23599 (comment).
sqllogictestruns under exactly one physical optimizer configuration: the default rule list. That makes a whole class of plan shapes untestable end-to-end, because the shape only arises under a different rule order.The concrete case that surfaced this:
WindowTopNsits at position 117 in the default list andProjectionPushdownat 142. AFilterExeccarrying an embedded projection therefore never reachesWindowTopNfrom a stock pipeline — but it does in downstream pipelines that re-run projection pushdown earlier. #21596 asksWindowTopNto handle that shape, and its premise ("this happens whenProjectionPushdownruns beforeWindowTopN") is not true of the default list.This leaves us with a bad choice. Either a rule grows logic for shapes no
sltcan produce — untested in practice, and an open-ended maintenance burden as "handle every possible shape" is not a realistic bar for an early-phase rule — or downstream engines carry local patches for orderings that upstream never exercises.The underlying tension, in @2010YOUY01's words: is the public API the ordered default rule list, or is it each individual rule? Today we rely on an ordering nobody wrote down.
Describe the solution you'd like
Run the
sltcorpus under a second, named optimizer pipeline, so an alternative rule order is a first-class thing upstream exercises rather than an assumption downstream engines rely on.Most of the machinery already exists.
# configMatrix:(#24493) re-runs a singlesltfile across a config sweep and attributes failures to the combination that produced them:Five files use it today (
sort_merge_join_matrix.slt,piecewise_merge_join_matrix.slt,mark_join_matrix.slt, ...). What it sweeps isConfigOptionskey/value pairs, not rule lists.So the missing piece is a config option that selects a pipeline variant, e.g.
where
reoptimizeis the default list plus a few re-optimization passes (a secondProjectionPushdown, etc.) — chosen to be closer to what downstream engines actually run.prefer_hash_joinis precedent for a config option that changes the physical plan. Files then opt in with:and the existing matrix machinery handles the rest.
Two things this buys, which are worth keeping separate:
EXPLAIN(sort_merge_join_matrix.slthas one), because a different rule list produces different plans andslthas no per-configuration expected output. So plan-shape coverage stays in targeted tests, and the sweep covers correctness.Describe alternatives you've considered
slt: it covers one query, and nothing keeps it in sync as the rule list evolves.Additional context
WindowTopN+ embedded projections, the case that surfaced this. Blocked on having a way to test it.configMatrixrunner this would build on.