Skip to content

Proposal: a relation conformance corpus — machine-checkable relation derivation and execution semantics in the spec repo #1164

Description

@nielspardon

Note

This is long — about 3,200 words, call it a 10–15 minute read. More than half of it is about how this proposal relates to work that already exists (substrait-validator's test corpus and the various Substrait text formats), rather than about the proposal itself. If you're short on time, the TL;DR below and What I'd like feedback on at the end are enough to react to; everything between them is reference material, safe to skim by heading. Nothing here is decided — it's a proposal looking for holes.

TL;DR

  • Problem. Relation semantics — field/type derivation, emit/remap, join nullability, set-op unification, decimal precision propagation, aggregate output typing — are specified in prose (site/docs/relations/*.md) and structurally in proto/substrait/algebra.proto. Nothing in this repo can be run to check whether an implementation matches them the way the spec intends.
  • Proposal. A spec-owned corpus of minimal, single-behavior plans, each paired with its expected output schema and expected rows, as protobuf-JSON manifests wrapping a plain substrait.Plan — loadable by any engine using the protobuf library it already has, with no new parser in any language.
  • What's machine-checked. CI independently recomputes each expected schema, so that half is correct by construction. Expected rows are author-asserted plus an automatic schema-conformance check — the same trust model the function .test files already run on, for the same reason (tests must land before engines can adopt the feature).
  • Who runs it. Consumers, in their own harness, gated on their dialects/ capability file so cases an engine can't support are skipped, not failed. This repo never executes rows.
  • This is not a new text format. Substrait already has four incompatible ones. The corpus is protobuf-JSON; text formats are treated as optional projections for review, and the comparison below explains why none of them can carry it today.
  • The ask. Is a non-normative substrait.test proto in this repo acceptable in principle, and is the trust model sound? Plus specific questions for the authors of substrait-explain, textplan and substrait-validator.

Where this comes from

This grew out of a discussion at a recent community sync. Several people pointed me at adjacent work — DataDog's substrait-explain, the textplan / .splan grammar, and substrait-validator's test corpus — so a good part of what follows is about how this fits alongside those rather than duplicating them. I've assessed those projects from the outside, so corrections from their authors are very welcome; where I've read a grammar or a test file I've cited the path so the claim is checkable.

The gap

Function testing here is comparatively well developed, though not finished: an ANTLR-based .test DSL (grammar/FuncTestCaseParser.g4) drives 1,297 cases across 132 .test files in tests/cases/, consumed by the harness in tests/coverage/, which enforces coverage against a golden tests/baseline.json, requires every case to match a declared function, and checks that .test return-type nullability agrees with the extension YAMLs.

Two caveats are worth stating plainly, because the same constraints apply to anything proposed here. Coverage is partial — baseline.json currently records 277 of 530 function variants covered, so a little over half. And the expected values in a .test case are author-asserted: nothing in CI evaluates them, so a case can encode a wrong expectation and still pass. I mention both not to undersell the function tests, which are the most developed testing asset the project has, but because a relation corpus will start out equally incomplete and will rely on author-asserted values for part of what it checks. Better to design around that openly than to imply a corpus arrives complete and infallible.

For whole plans and relations there is no equivalent. The plan-adjacent assets are:

  • site/examples/proto-textformat/ — strict-parse validation of three Expression sub-messages (Lambda, LambdaInvocation, FieldReference); never Rel / RelRoot / Plan.
  • site/docs/tutorial/final_plan.json — a documentation artifact, not a test.
  • dialects/tests/*.yaml — capability descriptors, schema-validated, never executed.

Existing conformance efforts — IBM/substrait-compliance and substrait-io/consumer-testing — work at the level of whole benchmark queries (TPC-H / TPC-DS, end to end). That's valuable, and complementary, but it doesn't isolate: a wrong result in a 200-node TPC-DS plan doesn't tell you which relation derived the wrong schema. The missing layer is unit-sized.

What a test case would be

  • A non-normative substrait.test.RelationTestCase protobuf message that embeds the spec protos, serialized as protobuf-JSON — one self-contained manifest per case.
  • Input schema and input rows ride inside the plan via ReadRel.VirtualTable (base_schema + literal rows), so a case is a single file with no sidecar data.
  • Each plan is the smallest plan that exercises one behavior, so a failure localizes to one relation.

Sketch — field layout would be settled in review:

syntax = "proto3";
package substrait.test;                 // non-normative tooling namespace

import "substrait/plan.proto";
import "substrait/type.proto";
import "substrait/algebra.proto";

message RelationTestCase {
  string name = 1;
  string description = 2;
  repeated string behaviors = 3;               // "emit_remap", "join_nullability", …
  substrait.Plan         plan            = 4;  // self-contained (VirtualTable input)
  substrait.NamedStruct  expected_schema = 5;
  ExpectedData           expected_rows   = 6;
  CompareConfig          compare         = 7;  // row order, float tolerance
  CapabilityRequirements requires        = 8;  // for dialect gating
}

Row comparison is per-case with defaults that match how engines actually behave: multiset (unordered) unless the plan ends in a terminal SortRel/TopNRel, float types compared with a tolerance, everything else exact.

Reusing the spec's own protos rather than inventing a case format means zero new parsers for consumers, free code generation and packaging, buf breaking coverage, and structural validation for free.

Trust model: what is machine-checked, and what isn't

This is the part I'd most like scrutiny on.

The schema half can be correct by construction. CI strict-parses each case, then independently recomputes the output schema from the input schema plus the relation and asserts it equals expected_schema. That requires a small derivation implementation in this repo, scoped deliberately to schema/type derivation only — not an execution engine.

The row half cannot be gated the same way, and I don't think it can be. There is a genuine catch-22: a new spec feature and its tests must land before the spec is released, but engines can only adopt that feature after release. So no pre-merge engine cross-check is possible, and the corpus has to act as the forcing function rather than a follower. What's automatic is a structural check that rows conform to expected_schema (types, arity, nullability); value correctness is human-reviewed, then validated in the field by real consumers, whose divergences show up as bug reports against either the engine or the corpus.

Worth noting that this is the trust model the function .test files already run on — their expected values are author-asserted too — so it's an established practice here rather than a new concession. The difference is that the schema half of each relation case doesn't depend on it: that part is independently recomputed in CI, which is strictly stronger than the existing function-test precedent.

Consumers execute; this repo does not. The spec repo's CI validates well-formedness and schema derivation. Engines run the corpus in their own harness and self-check both schema and rows. That's what lets it scale to engines this repo could never host.

Coverage and engine variance

First tranche = wherever consumers actually disagree today: emit/remap (RelCommon), join nullability (LEFT/RIGHT/OUTER/INNER × left-then-right concatenation), Project (passthrough plus computed-expression types, including decimal precision growth), Aggregate (grouping keys ++ measures, measure output types), Set (type/nullability unification), Read (base_schema plus projection/masking).

Two-tier coverage catalog, CI-enforced, mirroring the baseline.json pattern: every relation type in algebra.proto should have at least one case (breadth), plus a curated behavior checklist so depth gaps stay visible in review. As with baseline.json, the gate is a ratchet rather than a completeness requirement — it stops coverage regressing while it grows, and the corpus will be visibly incomplete for a long time, exactly as function coverage is today at 277 of 530 variants.

Dialect-gated execution. Cases are authored once at spec-canonical limits (including boundaries — P=1, P=38, scale/length extremes). At run time a consumer supplies its dialect file — the existing dialects/ capability schema already encodes supported_precision_range, per-type max_precision / max_scale / max_length, supported_types, supported_relations — and the harness runs a case only if the engine supports every type, parameter and relation it uses and its derived output. Unsupported cases are skipped, not failed. Derivation results stay universal; the dialect only governs applicability.

Home, versioning, distribution

Proposed home is this repo (tests/relations/, paralleling tests/cases/), versioned in lockstep with the protos — the derivation rules are the spec, so the tests that pin them shouldn't drift from it.

Cases would be emitted by a deterministic builder with a per-PR git diff --exit-code guard, the same drift protection already used for the ANTLR parsers.

Checked-in fixtures would be version-agnostic (Plan.version omitted), with substrait-packaging stamping the real SUBSTRAIT_VERSION during its existing subtree copy. That keeps per-release churn at zero here while published plans stay valid and correctly versioned.

Distribution reuses the existing streams: substrait.test bindings ride substrait-protobuf, corpus manifests ride substrait-extensions (which already carries the function .test files). No new packages or publish workflows.

How this relates to existing work

None of the projects below is a competing conformance corpus — but two have pieces I'd like to borrow, and one has something I'd like to offer back.

substrait-validator's testcase format — the closest prior art

This is the nearest existing thing to the derivation half of this proposal, and it deserves to be stated plainly rather than glossed over. substrait-validator/tests/tests/relations/ already has 96 cases across per-relation subdirectories (aggregate, common, cross, extensions, fetch, filter, join, project, read, root, set, sort) — the same layout this proposal arrived at independently — and 60 of them assert a derived type. They hit first-tranche behaviors directly: relations/join/left.yaml pins LEFT-join nullability as STRUCT<string, i32, fp32?, boolean?>, relations/common/emit-basic.yaml pins an outputMapping remap, and the set cases pin set-op nullability unification.

So the gap isn't that nobody has written derived schemas down. It's that what exists is a regression suite for one implementation rather than a portable spec artifact:

  • The assertion target is validator behavior. A green run says the validator still agrees with itself; it gives DuckDB or Isthmus no way to ask "do I derive this?"
  • Assertions are in-band: __test keys are interleaved into the plan, so the plan value isn't a substrait.Plan — runner.py must preprocess it into a private intermediate first. The envelope proposed here is the deliberate inverse: a wrapper message around a pristine substrait.Plan, so a manifest strict-parses with a vanilla protobuf library in any language.
  • No rows and no execution, and no capability gating for engines with narrower parameterized-type limits.
  • Version coupling: the released validator (v0.1.4) targets spec 0.57.x, with 0.87.x on unreleased main, while the spec is at 0.99.0.
  • It's diagnostic-centric — 82 of its 158 cases assert diagnostics. That's its strength, and it's precisely the negative/must-reject space this proposal defers to a later phase.

One thing I'd like to borrow from it. Its type: assertions use NSTRUCT<a: string, b: i32?> — and that syntax is a production in this repo's own grammar (grammar/SubstraitType.g4, grammar/SubstraitLexer.g4), the same grammar the function .test files use, with a generated parser already checked in. A one-line NSTRUCT<…> is far more reviewable than a NamedStruct proto-JSON block. So: keep expected_schema as the normative NamedStruct, and add an optional, redundant expected_schema_string that CI verifies against it. The redundancy is the point — a readable view that can't drift, because CI compares the two.

One thing I'd like to offer back. The validator is a second, independent implementation of relation derivation. It could consume the positive corpus as "must produce no diagnostic above info" cases: extra regression breadth for it, and an independent derivation cross-check for the corpus. Because it lives outside this repo and lags releases, that has to stay advisory and out-of-band rather than a merge gate here. And when negative cases arrive, its diagnostic corpus is the obvious thing to mine.

The text formats

Worth saying up front: the spec already advertises a text form it doesn't specify. site/docs/index.md lists "the text version of the Substrait plan" among Substrait's benefits, and site/docs/types/type_parsing.md links the type grammar into substrait-cpp's SubstraitPlanParser.g4, noting that "the grammar also supports an entire language for representing plans as text." So there's a text format acknowledged by reference, from another repo, with no normative status — and, partly as a result, four independent implementations:

Implementation Language Direction Status
substrait-cpp textplan (.splan) C++ (ANTLR) text ↔ protobuf in-tree, with a written language reference
substrait-io/substrait-textplan Rust (same grammar) + C/C++/Python/Go FFI text ↔ protobuf created 2025-11-24, no tagged release yet
DataDog/substrait-explain Rust text ↔ protobuf v0.6.0, actively released
substrait-java SubstraitStringify Java protobuf → text debug utility; completion tracked in substrait-java#302

I looked at each to see whether the corpus could be encoded in one instead of protobuf-JSON. My conclusion is no — but for reasons specific to this use case, not judgments about the projects.

Per-format detail (capabilities relevant to this corpus)

DataDog/substrait-explain — Apache-2.0 Rust crate and CLI; compact, EXPLAIN-like indented syntax with $index field references; tracks the spec via the substrait crate (0.63, post-URN). It is the only one of the three with inline virtual-table row syntax (Read:Virtual[(1, 'alice'), (2, 'bob') => id:i64, name:string]), which is exactly the mechanism this proposal uses for input data, and it makes emit/remap explicit (Read[t +> a:i64, b:string, c:i64 |> $1, $0]). Its === Version section is optional, which suits version-agnostic fixtures. Gaps for this purpose: its README lists Set and Window as planned but not yet implemented, and decimal literals aren't implemented — those are two of the six first-tranche targets. Its DESIGN.md is also explicit that it performs no validation or type checking and normalizes toward canonical output, which is right for a debugging tool but means it can't reliably pin the fine proto distinctions a conformance case exists to pin (e.g. an unset emit_kind versus an explicit Direct).

At the sync it was mentioned that DataDog runs an internal relation test suite built on this format, similar in spirit to this proposal — I'd be glad to be corrected on the details. If those cases could be shared, the plan half converts mechanically (substrait-explain convert -f text -t json), and the assertions — the genuinely valuable part — would come across as data rather than as a format. That's plausibly the fastest route to a substantial first tranche.

textplan / .splan (substrait-cpp src/substrait/textplan/, plus the Rust port at substrait-io/substrait-textplan) — the oldest and most formally specified of the three, with a real language reference at substrait-cpp/src/substrait/textplan/README.md. Structurally different from the others: named declarative blocks with the dataflow in a separate pipelines { read -> filter -> root; } section, and field references resolved by name through a symbol table rather than by index. On several axes it's the strongest candidate — substrait-io-owned, bindings for C/C++/Python/Go, a literal syntax richer than the alternatives (decimal literals such as -123_decimal<3,0>, plus interval, binary, UUID, list/map/struct literals), and relation coverage that includes Set.

Two things block it for this use, and both are grammar-level rather than maturity:

  1. No syntax for virtual-table rows. In both implementations the production is VIRTUAL_TABLE id LEFTBRACE RIGHTBRACE — an empty body (SubstraitPlanParser.g4, line 194 in substrait-cpp / 195 in substrait-textplan). There's no way to write the inline literal rows this proposal uses for input data.
  2. No Plan.version. plan_detail (line 23 in both) admits only pipelines, relation, root_relation, schema_definition, source_definition and extensionspace — so a .splan can't round-trip a required proto field.

Both look like additive grammar changes rather than redesigns. I'd be happy to raise them as feature requests: this corpus is a concrete consumer that would want them, which may be useful roadmap input regardless of whether the corpus ever renders to .splan. Separately, the Rust port is early — no tagged release or published package yet, pinned to substrait 0.61 (pre-URN, with the bump tracked in its issues 7 and 16), complex-literal support still catching up with its own language reference (issues 19–21, 23), and a build that currently needs a specific beta ANTLR4-Rust JAR plus a JRE and protoc.

substrait-java SubstraitStringify — a RelVisitor producing Spark/Calcite-style explain output. Its own Javadoc describes it as being for debug and development only and "not a replacement for any canonical form," it's one-way with no parser, and it lives in an examples module. Its tracking issue (substrait-java#302) frames the open question well: complete it, "or choose/implement an alternative string encoding."

Where that leaves the format question. Protobuf-JSON stays the encoding for the corpus, because the corpus has to load in every consumer language with no new parser, and because no text format currently expresses both inline virtual-table rows and Plan.version. But text formats are genuinely better for review — a four-node plan is ~70 lines of proto-JSON — so I'd document a converter one-liner for reviewers, and keep any checked-in text renderings non-authoritative and out of CI.

I want to be clear that this proposal shouldn't be the thing that decides a canonical text format for Substrait. That's a real and separate question, it's already half-asked in substrait-java#302 and by the spec's own docs, and it deserves its own thread.

What I'd like feedback on

  1. The proto. Is a non-normative substrait.test package in this repo acceptable in principle? It buys zero new parsers, free bindings via existing packaging, and buf breaking protection — but it is a new proto in the spec repo, which I assume needs explicit buy-in.
  2. The row-gating catch-22. Author-asserted expected_rows plus an automatic schema-conformance check is the best I could come up with, given that tests must land before release and engines adopt after. Is there a better answer?
  3. The in-repo deriver. Writing a schema-only derivation implementation here is real new code, and substrait-validator already implements derivation. Is correct-by-construction in-repo worth it, or should this lean on the validator as an out-of-band cross-check instead?
  4. Scope of the first tranche — is emit/remap → join nullability → project (incl. decimal) → aggregate → set → read the right priority order? It's chosen as where I believe engines diverge most; I'd rather be corrected early.
  5. expected_schema_string — worth adding as a CI-verified redundant NSTRUCT<…> mirror for reviewability, or unnecessary complexity?
  6. Dialect gating — reusing the dialects/ capability schema for skip-vs-fail decisions seems natural. Does that match how the dialect files are intended to be used?
  7. To the authors of the adjacent projects — are the capability statements above accurate? And is there interest in (a) sharing/translating existing cases, (b) the two textplan grammar additions, (c) the validator consuming the corpus as zero-diagnostic regression cases?

Happy to bring any of this to a sync, and happy to start with a single trivial passthrough case to stand up the harness before authoring breadth.

Activity

  1. julianhyde commented on Aug 4, 2026

    @julianhyde

    I needed to create a similar corpus in Morel; in fact, I just wrote a blog post about how it allowed us to translate Morel to Go.

    For Morel, what worked well was a format (.smli) that was idempotent (i.e. results mixed in with queries) and mergeable, where we could test only the aspects we cared about (e.g. metadata, or query plan, or query results), and having a reference implementation (the Java implementation of Morel).

    My advice to you:

    • I know you don't want to devise a new format, but choose the format carefully. If it isn't easy to edit and merge (e.g. JSON) you'll burn a lot of effort on tooling.
    • Consider adding a reference implementation of Substrait. It will make everything more concrete. It should be simplistic, inefficient, and only work on small data sets (no more than 100 rows), so that no one considers using it in production. Claude could probably write it in an hour. Then your proposed "corpus" then becomes the test suite of the reference implementation, as opposed to a stuffy artifact of a standards organization.

    A bonus of having a reference implementation: people implementing Substrait will point their agents at the reference implementation and get quick feedback on whether their implementation is correct.

  2. benbellick commented on Aug 5, 2026

    @benbellick
    Member

    Major ➕ on the need for a format to specify the semantics of relations. We internally have a tool at Datadog we are using for testing plans based off of the text format backed by substrait-explain (link). I don't have a timeline on if/when the format itself will be made OSS.

    In my opinion, the most important quality of such a test suite is readability. We need to make sure that it is easy for maintainers to quickly glance at tests proposed and recognize that the test adequately captures the intended semantics. I understand the aversion to relying on a new format. However, relying on a JSON representation of full substrait plans would be challenging. I know that I would not be confident looking at any non-trivial substrait plan and determining from inspection alone that the plan is legitimate / semantically-correct.

    Considering that there is prior art for writing tests for functions in a custom format, it seems reasonable to have a similar format for relational plans. Making the test format a new protobuf feels optimized for the machines rather than for us. (I have the obvious bias of working at DataDog, but I think substrait-explain is a great format for the plans, considering it is the most actively maintained of the ones suggested and is sufficiently compact to be readable as tests.)

    Nonetheless, something in this direction is a good idea and having the community settle on something that helps us all would be great :)


    By the way, I do think the idea of having a reference implementation is an interesting approach. However, I worry that a reference implementation unintentionally codifies semantics we didn't intend, especially if we try and have AI generate something quickly. Starting from tests means we codify the behavior we know and deliberately leave unspecified the behavior we don't. Taking the reference implementation approach feels like codifying all of the behavior at once, but I think we will find that there are lots of footguns we haven't yet anticipated.

  3. wackywendell commented on Aug 5, 2026

    @wackywendell
    Contributor

    👋 Chiming in, as the main author of substrait-explain, to answer some of the above questions related to that repo/format!

    are the capability statements above accurate?

    Some clarifications:

    • Set operations are now supported
    • Window functions are in process - a supporting PR was merged today
    • Decimal literals are not yet supported, but could be; contributions welcome, or if this was a significant blocker, I would do it.
    • Emit syntax exists, but is only implemented so far for Read relations; others are in progress.
    • It has an extensive grammar specification, with test cases built into the grammar spec
      • The test cases in the spec are not data-based - they simply convert text to protobuf and back, and check conversion fidelity.
      • The tests use Rust's doctests setup; GRAMMAR.md is imported into the Rust crate as Rust docs, and then cargo test runs the tests defined in it.

    Overall, in terms of maturity: as @benbellick mentioned, we're currently making this a test format within DataDog, so that it is both human- and machine-readable/writeable, and maps very cleanly and deterministically to Substrait. As the above examples show, its not quite there, but reasonably close, and quite actively developed. AI has been used, but all PRs are carefully human-reviewed to ensure clear, maintainable code and thoughtful syntax choices. It's open-source and available on crates.io.

    In terms of the data-based test cases, many of our internal uses use internal extensions, and would not be suitable for open-sourcing; but there are also a bunch that we've built that are cross-engine, and some folks here have even found DataFusion bugs using them. I'm mainly working on the substrait-explain grammar / parser; I'll defer to @benbellick on the test cases maturity.

  4. julianhyde commented on Aug 5, 2026

    @julianhyde

    However, I worry that a reference implementation unintentionally codifies semantics we didn't intend, especially if we try and have AI generate something quickly.

    That would have been my concern a few months ago, but I feel differently having used Claude extensively. Strictly, and in theory, the reference implementation needs to pass the corpus of tests and its behavior is undefined for any other case. But in practice, the agent will do the sensible thing (especially if a human provides some guidance at the start of the process) and will produce a parser with a grammar that tends to match the specification, a type inference process that matches the specification, and execution rules that look like denotational semantics. And therefore the reference implementation will tend to do the right thing almost all of the time. And, as I mentioned, agents can test their hypotheses against the reference implementation and help you write the tests correctly almost all of the time.

    There will be occasions when the reference implementation differs from tests that people propose adding to the corpus. Those PRs will be a forcing function for the humans to discuss what they intended/want.

  5. julianhyde commented on Aug 5, 2026

    @julianhyde

    Regarding format, take a look at Morel's smli. It designed for humans to read and write, and machines can merge it very easily. Because it is concise, encourages comments, and includes output, reading an .smli file is very much like reading a specification.

    Because it is an expression language, you can include only the output that is salient for a particular test: usually the output and/or the derived type, sometimes parse/validation errors, sometimes the plan. The output matcher is resilient to changes in output order and whitespace, and this reduces false-negatives and the need to update the output.

  6. alexandrefimov commented on Sep 2, 2026

    @alexandrefimov
    Contributor

    One area the (4) list does not cover is type and nullability derivation on consume, with no join involved: a case that reads a base_schema and asserts the output type the consumer derives from it. apache/datafusion#22105 was about nullability, substrait-io/substrait-java#1121 about temporal literals, both fixed by now, and in each the consumer derived something other than what the producer declared. Those cases are the cheapest I have written: a single read carrying ReadRel.projection found the mask ignored in Spark, Isthmus and substrait-python.

  7. alexandrefimov commented on Sep 2, 2026

    @alexandrefimov
    Contributor

    On the priority order in (4): across 95 cases and ten implementations, counting how many implementations disagree with the expectations I wrote from the spec, aggregate is the widest at five, decimal is four, join, set and masking are level at three, and emit is one.

    Aggregate disagrees in both directions: DataFusion widens a grouping key present in every set, substrait-python and substrait-go keep both keys required where one set omits them, the validator does not parse relation-level grouping expressions at all, and DuckDB binds sum:i64 to its own HUGEINT over the declared output type.

    Decimal: dividing dec(10,2) by dec(5,1) gives (21,8) in the spec, (15,6) in DataFusion, DOUBLE in DuckDB and (16,7) in Acero, while Spark keeps the spec's precision and scale and returns the result nullable. The formula is the spec's and the arithmetic is mine, so it is worth a second pair of eyes.

    Join: substrait-python and the validator concatenate the inputs instead of applying the join type, and substrait-go gets JoinRel right but not HashJoinRel or MergeJoinRel.

    set: DataFusion takes the first input's nullability on intersection where the spec derives it from all inputs, the validator does the same for every set operation including union, and Isthmus applies the union pattern to intersection and minus as well; in rows, DataFusion returns five where the spec gives three for INTERSECTION_MULTISET_ALL and none at all for MINUS_PRIMARY_ALL.

    Masking through ReadRel.projection belongs beside RelCommon.emit rather than under Read: the projection is ignored in Isthmus, substrait-python and Spark, emit only in DuckDB, which ignores the indices in output_mapping on six relation types with no error.

    @wackywendell, on your offer in (7): decimal literals would unblock the most of this.

  8. alexandrefimov commented on Sep 6, 2026

    @alexandrefimov
    Contributor

    For (3): a matching schema does not show that the consumer computed it. Swapping every declared output_type in the corpus for a false one moves substrait-java's reported schema on eleven cases, substrait-python's and the validator's on ten each.

    decimal_divide is the clearest of those: the spec's formula gives dec(21,8), the plan declares dec(21,8), substrait-java reports it, and declaring dec(20,2) instead makes substrait-java report dec(20,2).

    DuckDB moves on none of them, and substrait-go's producer API returns dec(21,8) for that division and dec(38,17) for dec(30,20) * dec(30,20) without being handed a return type at all - so the calculation exists; a corpus of declared plans just does not reach it.

    The validator cannot supply it either: its pinned implementation keeps the supplied return type, and checking it against the function definition is unimplemented.

    That makes (3) worth doing: a deriver in the spec repo that computes the return type from the argument types and the function's extension definition, then checks the plan's output_type against it. Its own tests should carry wrong declarations, so that a deriver which merely copies the declaration fails them.

  9. vbarua commented on Sep 9, 2026

    @vbarua
    Member

    Meta

    This discussion has gotten a little verbose due to agentic assistance. I've attempted to summarize the core of the discussion at a high-level so that we can have a human-centric design discussion about the test format itself before getting too deep in the weeds of implementation and planning every possible test we could want to write. Lets figure out the foundation of the house we're building before picking the tiles for the kitchen back splash.

    Please write to be read by your fellow contributors, and not agents. A little editing never hurt anyone :slight-smile:

    Problem

    Relation semantics can be fairly complicated and there is no existing mechanism of capturing behaviours in the spec that could be used for mechanical verification. Effectively, the textual spec is all we have for relations, and that is not always clear.

    Core Solution Idea

    The spec should have, alongside the textual definitions for relations, test vectors for those relations. These vectors can be used to test:

    • The language libraries in our ecosystem.
    • Engines consuming Substrait

    Beyond improving the quality of systems consuming and producing Substrait, these vectors would also help in documenting the expected behaviours of these relations. The primary properties to verify for relations would be:

    1. Expected output schemas
    2. Expected rows

    These properties would be encoded into tests by the authors of the test, and ideally would be written alongside the definition of any new relations. Each test would exercise a specific behaviour of the relation.

    This does introduce a potential bit of a catch-22, in that tests will exist for a relation before they can be integrated into any system, because the relation won't be available until we first publish it.

    Niels' suggestion is to introduce a RelationTestCase message that wraps a full Substrait plan (the test input), alongside other fields that capture the expectations on said plan.

    Prior Art

    There is some prior art in this space in various places that Niels' has helpfully documented.

    • Functions Test Cases: there is a custom text format in the spec repo to helps capture function behaviours.
    • substrait-validator: contains plans for various relations and expressions with type assertions on output schemas. These tests are not executable.
    • text formats: there are a number of text formats (textplan, substrait-textplan, substrait-explain, SubstraitStringify) with varying degrees of support for portions of the specification.

    Additional Suggestions

    Julian Hyde

    Julian made two suggestions:

    1. Consider the choice of format carefully in terms of editing and merging tests to avoid burning efforts on tooling.
    2. Consider a reference implementation for Substrait.

    Ben Bellick

    Ben echoed Julian's first point, specifically with regard to readability to make it easy to understand what tests are verifying.

  10. vbarua commented on Sep 9, 2026

    @vbarua
    Member

    To the things you requested feedback on @nielspardon, I'm in broad agreement with the following:

    • We should have some kind of machine-readable test format (1). Protobuf would be a reasonable choice for this.
    • Tests should live in the core repo, and ideally be defined alongside relations as they are added, with the caveat that these tests should be easily readable (and thus reviewable) by humans.
    • Having an expected schema string (when possible) make sense to me. (5)
    • Dialects can 100% be used to gate tests (6). Another way to think about this is that entities should only be added to a dialect for a system if all the tests for those entities pass.

    I don't think that we need an in-repo schema deriver for relations (3) IF we have a human-readable format. That can be handled as part of adding tests for the relation by the author. I also think it's fine to land tests before they are integrated (2), especially if we think of them as part of the spec documentation. The spec is still the source of truth, we're just adding a mechanism that helps us be even clearer about what that is, and easier for other systems to check their behaviours against.

    I don't feel strongly about the tranche priority (4). The hard part IMO is defining the format, once that's available writing new tests is straightforward.

    Test Format Suggestions

    The two properties that I think are useful for a test format are that:

    1. Tests should be easily loadable into consuming systems to make them easy to integrate.
    2. Tests should be easily read/written by humans.

    Niels' suggestion satisfies 1, but makes 2 challenging. However, I think that there is a fairly nice solution to this, which is that we standardize on a single machine-readable format and use that as a target for the human readable format(s).

    Case Studies in Test Generation

    The following examples are actually things that we do internally (that we really should prioritize for open-sourcing).

    Function Test Case Format

    Consider an existing function test like

    add(30000::i32, 30000::i32) = 60000::i32
    

    This is super easy to read, but its a pain to integrate into any consuming system because you need to parse the test format in each consumer. This was a real pain point that we ran into when testing DataFusion, especially because there is no test case parser in Rust.

    Instead, we pre-processed the tests into a Substrait plan like

    Root[result]
      Project[add(30000:i32, 30000:i32):i32?]
        Read:Virtual[
          - (0:i32)                             // Single Dummy Row
          - => dummy:i32]
    

    and then embedded the expectations alongside the plan like:

          result:
            type:
                i32:
                    nullability: NULLABILITY_REQUIRED
            expected_return_value:
                i32: 60000
    

    We use YAML as our machine-readable format, embedding the expectations into YAML and then converting the protobuf Substrait plan into protoyaml.

    This was easy to run against DataFusion because you effectively hand it a full protobuf Substrait plan (which is what it expects already), and then do a bit of processing to check the assertions.

    Relation Test Case

    Our intern this summer wrote a test case format that looks a bit like:

    # Simple Filter Rel
    == Query
    === Extensions
    URNs:
      @  1: extension:io.substrait:functions_comparison
    Functions:
      # 9 @  1: gt:any_any
    
    === Plan
    Root[id, name, age]
      Filter[gt($2, 30:i32):boolean => $0, $1, $2]
        Read:Virtual[(1:i32, 'Alice':string, 31:i32), (2:i32, 'Bob':string, 25:i32) => id:i32, name:string, age:i32]
    
    == Results
    === Plan
    Root[id, name, age]
      Read:Virtual[(1:i32, 'Alice':string, 31:i32) => id:i32, name:string, age:i32]
    

    Which we then compile into two machine readable Substrait plans (an input plan and a result plan). We assume that the system we are testing can handle basic VirtualTables for this.

    Bonus: Function Property Testing

    Another fun thing we do is generate tests based on function definitions. Consider the humble add function

      -
        name: "add"
        description: "Add two values."
        impls:
          - args:
              - name: x
                value: i32
              - name: y
                value: i32
    

    Without knowledge of its implementation, there are still number of useful properties that we can derive tests from:

    1. It has nullability mode MIRROR, so we can generate inputs with all possible nullability combinations and make sure that any consuming system matches the defined nullability handling behaviour.
    2. This invocation takes two arguments i32 arguments, so we can check that invocations with 1 i32 argument, or more two i32 argument result in failures in the consumers.
    3. We can check that invoking add:i32_i32 with non-i32 arguments fails in the consumer.

    Takeaways

    Internally we effectively utilize two human readable formats (the function test format + our substrait-explain based relation format) but we compile them down to a shared machine-readable format. I actually think this is a fairly nice setup, because not every (human) readable format will work for every type of test, and if we need to write a particularly complicated tests that we can't express we can drop into the machine-readable format (though in practice we just update substrait-explain with the features we need).

  11. alexandrefimov commented on Sep 9, 2026

    @alexandrefimov
    Contributor

    @vbarua, thanks for the summary.

    On (3): a case carrying a declared output_type does not show whether the consumer computed that type or took it from the plan, however readable the case is. Swapping the declaration tells the two apart, and substrait-java's reported schema follows it on eleven cases, substrait-python's and the validator's on ten each.

    If you open-source the relation format, I can convert the decimal cases to it and see whether a wrong output_type is visible there.

  12. vbarua commented on Sep 9, 2026

    @vbarua
    Member

    @alexandrefimov

    On (3): a case carrying a declared output_type does not show whether the consumer computed that type or took it from the plan

    I don't see how have an in-repo deriver addresses this, as it doesn't affect how the consumer handles the output type. As I understand it, the in-repo deriver is effectively an automated way to compute the expected output_type. Your swap test hints at something that could be used here though, which is writing tests with intentionally incorrect output type computations that we would expect to fail. If a library consumes then as is without issues, it let's us know it's not checking them.

  13. alexandrefimov commented on Sep 17, 2026

    @alexandrefimov
    Contributor

    Agreed on the deriver. I ran the test you describe: 27 relation plans with their declared output_type fields replaced and nothing else changed. None of the nine implementations the corpus covers rejected any of them. Results.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions