Skip to content

Discussion: Define the Contract for Custom Physical Optimizer Pipelines #26171

Description

@2010YOUY01

Motivation

The physical optimizer's contract is currently underspecified. In particular, it's unclear how much flexibility downstream projects should have when composing built-in optimizer rules, and what guarantees DataFusion should provide for such custom pipelines.

One possible interpretation is that each built-in rule must support arbitrary pipeline compositions. However, this could introduce substantial implementation complexity to support plan shapes that never occur in the default pipeline.

I'd like to reach an agreement on the intended contract here, then document it more clearly in a follow-up PR.

This discussion explores:

  1. Most restrictive: Treat the default optimizer pipeline as a public API.
  2. Most permissive: Treat every built-in rule as an independently supported public API.
  3. Proposed middle ground: Allow custom pipelines while separating correctness guarantees from optimization coverage.

Note that 1 and 2 are just thought experiments to illustrate the trade-offs. I don't think either should be adopted.

1. Most restrictive: The default pipeline is the contract

See the [#26107](https://github.com/apache/datafusion/pull/26107/changes) draft for a more detailed version.

Under this interpretation, the default rule sequence defines the supported behavior, and downstream projects/extensions should adapt to the default optimizer pipeline.

This gives built-in rules the strongest maintainability guarantees, but limits downstream flexibility.

2. Most permissive: Every built-in rule is independently supported

Under this interpretation, built-in rules are expected to handle all relevant plan shapes, regardless of where they appear in a custom pipeline. This could accumulate unnecessary maintenance overhead for the following reasons:

Progressive refinement in the optimizer pipeline

The default optimizer pipeline is a progressive refinement process, where the possible plan shapes vary across stages.

For example:

  • At rule 1, a hash join might have only one canonical representation.
  • At rule 20, the same logical join might have five possible physical representations.

A rule that rewrites joins could be much simpler when placed in the earlier stage.

Requiring every rule to handle every possible representation, regardless of its position, can unnecessarily bloat the implementation in the earlier stages.

An unintended scenario

Consider a built-in rule that identifies a join subtree and converts it into a join graph for join reordering.

In the default pipeline, earlier rules establish a canonical representation, allowing this rule to use a concise pattern-matching implementation.

Later, a downstream project introduces a custom pipeline that produces many equivalent but non-canonical join shapes. It then requests that the built-in rule be extended to recognize 10 additional shapes.

For example, these plans are semantically equivalent for inner joins:

SELECT *
FROM t1, t2, t3
WHERE t1.c1 = t2.c1 AND t1.c2 = t2.c2;
-- Shape 1: Canonical join conditions

Join
  Join(on=[(t1.c1, t2.c1), (t1.c2, t2.c2)])
    Scan(t1)
    Scan(t2)
  Scan(t3)

-- Shape 2: Join filter

Join
  Join(filter=(t1.c1 = t2.c1) AND (t1.c2 = t2.c2))
    Scan(t1)
    Scan(t2)
  Scan(t3)

-- Shape 3: Filter immediately above inner join

Join
  Filter((t1.c1 = t2.c1) AND (t1.c2 = t2.c2))
    Join
      Scan(t1)
      Scan(t2)
  Scan(t3)

-- Shape 4: Filter above entire join subtree

Filter((t1.c1 = t2.c1) AND (t1.c2 = t2.c2))
  Join
    Join
      Scan(t1)
      Scan(t2)
    Scan(t3)

... and more

-- Splitting the AND conditions and placing them at
-- different levels can also produce equivalent plans.

Accommodating this request has several implications:

  • Upstream complexity driven by downstream requirements: The implementation and supported plan shapes of built-in rules become dependent on arbitrary downstream pipeline choices.
  • Testing gaps: Most end-to-end optimizer tests use sqllogictest with EXPLAIN to verify the final optimized plan produced by the default pipeline. Additional rule behavior required only by custom pipelines would not be covered by these tests.
  • Accumulating complexity: Over time, built-in rules acquire increasingly complicated logic to support plan shapes that the default pipeline never produces.

This kind of extension is already fairly common practice.

The difficult question is: When is extending a built-in rule justified?

In this example, the downstream project could canonicalize its plans before invoking the join-graph conversion rule.

Without a clear contract, we risk turning a downstream-specific workaround into a permanent upstream maintenance obligation.

3. Proposed middle ground: Separate correctness from optimization coverage

This is my current thinking, and I'd appreciate feedback on whether it strikes a reasonable balance.

First, distinguish two guarantees:

  • Plan validity / correctness: Applying built-in optimizer rules preserves query semantics and produces a valid executable plan, without internal errors or broken plans, even when the optimizer pipeline is arbitrarily recomposed.
    • This should have relatively low maintenance overhead, since most optimizer rules follow a "whitelist" pattern: if a plan shape matches, the rule transforms it; otherwise, it leaves the plan unchanged.
  • Optimization coverage: An optimizer rule recognizes and optimizes the plan shapes it is intended to handle.

Consider:

# Default pipeline
rule1
rule2
rule3

# Custom pipeline
rule3
custom_rule1
rule1

If we restrict the contract to guarantee validity, but not optimization coverage:

  • The custom pipeline is guaranteed to be runnable.
  • However, it might miss some optimization opportunities. For example, in the default pipeline, rule1 might only encounter certain canonical plan shapes, while the custom pipeline could introduce more diverse plan variations at rule1.

My proposed contract is:

  1. Custom pipelines should preserve correctness, but are not guaranteed equivalent optimization coverage.

    By default, extension pipelines can expect correctness under recomposition, and add unit tests for built-in rules to guard against regressions.

    If they need additional optimization coverage, they should first try to accommodate it within the extension. If this becomes a common requirement across downstream users, see the next point.

  2. Additional optimization coverage should be justified and tested as an upstream feature.

    When downstream projects want to extend a built-in rule's coverage, they should demonstrate the value of that extension rather than simply requiring upstream to accommodate their custom pipeline.

    One practical approach would be to:

    • Add a representative custom pipeline to DataFusion's test infrastructure.
    • Run all sqllogictest cases through that pipeline to verify the additional optimization coverage end-to-end.
    • Explain why supporting this pipeline brings general value, rather than merely moving a downstream-specific hack upstream.

I'd be interested to hear whether this sounds like a reasonable middle ground, or if there are important downstream use cases it would fail to support.

Activity

  1. 2010YOUY01 commented on Oct 10, 2026

    @2010YOUY01
    ContributorAuthor

    This kind of extension is already fairly common practice.

    Here are two in-the-wild examples of downstream extension requirements. In the context of this issue, addressing (1) seems uncontroversial, while (2) may require further justification.

    1. Downstream requirement to avoid internal errors

    2. Downstream requirement for broader optimization coverage

  2. alamb commented on Oct 11, 2026

    @alamb
    Contributor

    If we restrict the contract to guarantee validity, but not optimization coverage:

    • The custom pipeline is guaranteed to be runnable.
    • However, it might miss some optimization opportunities. For example, in the default pipeline, rule1 might only encounter certain canonical plan shapes, while the custom pipeline could introduce more diverse plan variations at rule1.

    I agree with this description

    Custom pipelines should preserve correctness, but are not guaranteed equivalent optimization coverage.

    Yes

    Additional optimization coverage should be justified and tested as an upstream feature.

    Yes

    Later, a downstream project introduces a custom pipeline that produces many equivalent but non-canonical join shapes. It then requests that the built-in rule be extended to recognize 10 additional shapes.

    I feel like this is the core challenge that we need to address -- how to decide to push back (or accept) code to handle shapes that the built in pipeline never covers.

    Here are two in-the-wild examples of downstream extension requirements. In the context of this issue, addressing (1) seems uncontroversial, while (2) may require further justification.

    1. Downstream requirement to avoid internal errors

    In my mind, this is not really a new requirement on the optimizer, it is a requirement from HashJoinExec itself (which was being triggered by the optimizer). I do think the requirements/semantics of individual ExecutionPlan nodes could be much better

    1. Downstream requirement for broader optimization coverage

    The current PR as proposed I don't think is a request broader optimization coverage -- it looks to me like a bugfix, but it also looks like it has changed significantly since when @zhuqi-lucas first proposed it

    I do see the code / request for extending WindowExec pushdown when FilterExec has an embedded projection (which the default optimizer rule order never exposes)

  3. alamb commented on Oct 11, 2026

    @alamb
    Contributor

    So TLDR is I agree that your proposal (option 3)

    Separate correctness from optimization coverage

    Is a good idea and what I thought (clearly not correctly) was already the case

  4. alamb commented on Oct 11, 2026

    @alamb
    Contributor

    Perhaps @zhuqi-lucas @xudong963 @Dandandan @kumarUjjawal @jayzhan211 and @kosiew and others may have some ideas on this one

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions