Repository navigation
Conversation
|
run benchmarks |
|
run benchmarks |
|
run benchmarks |
|
run benchmarks |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark clickbench_partitioned
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark tpch
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark tpcds
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"
SIMULATE_LATENCY: "true"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark tpcds
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark clickbench_partitioned
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"
SIMULATE_LATENCY: "true"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark tpch
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"CPU Details (lscpu)Details
Resource Usagetpch — base (merge-base)
tpch — branch
File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark tpch
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"
SIMULATE_LATENCY: "true"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark tpcds
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"CPU Details (lscpu)Details
Resource Usagetpcds — base (merge-base)
tpcds — branch
File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark clickbench_partitioned
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "pipelined"
SIMULATE_LATENCY: "true"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark clickbench_partitioned
env:
DATAFUSION_RUNTIME_MEMORY_LIMIT: "4G"
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"
SIMULATE_LATENCY: "true"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark tpcds
env:
DATAFUSION_RUNTIME_MEMORY_LIMIT: "4G"
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"
SIMULATE_LATENCY: "true"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark tpch
env:
DATAFUSION_RUNTIME_MEMORY_LIMIT: "4G"
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"
SIMULATE_LATENCY: "true"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark clickbench_partitioned
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"CPU Details (lscpu)Details
Resource Usageclickbench_partitioned — base (merge-base)
clickbench_partitioned — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark tpcds
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"
SIMULATE_LATENCY: "true"CPU Details (lscpu)Details
Resource Usagetpcds — base (merge-base)
tpcds — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark tpch
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"
SIMULATE_LATENCY: "true"CPU Details (lscpu)Details
Resource Usagetpch — base (merge-base)
tpch — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark tpch
env:
DATAFUSION_RUNTIME_MEMORY_LIMIT: "4G"
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"
SIMULATE_LATENCY: "true"CPU Details (lscpu)Details
Memory Pool PeaksPeak Base:
Pool accounting vs. process RSS Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.
Resource Usagetpch — base (merge-base)
tpch — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark clickbench_partitioned
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"
SIMULATE_LATENCY: "true"CPU Details (lscpu)Details
Resource Usageclickbench_partitioned — base (merge-base)
clickbench_partitioned — branch
File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark tpch
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "pipelined"
SIMULATE_LATENCY: "true"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark tpcds
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "pipelined"
SIMULATE_LATENCY: "true"Results will be posted here when complete File an issue against this benchmark runner |
|
run benchmarks |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark clickbench_partitioned
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "pipelined"
SIMULATE_LATENCY: "true"CPU Details (lscpu)Details
Resource Usageclickbench_partitioned — base (merge-base)
clickbench_partitioned — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (5ff7af6) to 39d5064 (merge-base) diff Run configurationrun benchmark tpcds
env:
DATAFUSION_RUNTIME_MEMORY_LIMIT: "4G"
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"
SIMULATE_LATENCY: "true"CPU Details (lscpu)Details
Memory Pool PeaksPeak Base:
Pool accounting vs. process RSS Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.
Resource Usagetpcds — base (merge-base)
tpcds — branch
File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (bcbbda1) to 39d5064 (merge-base) diff Run configurationrun benchmark clickbench_partitioned
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"
SIMULATE_LATENCY: "true"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (bcbbda1) to 39d5064 (merge-base) diff Run configurationrun benchmark tpcds
env:
DF_FETCH_BUDGET: "104857600"
DF_FETCH_POLICY: "streaming"
SIMULATE_LATENCY: "true"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark tpch
env:
DATAFUSION_EXECUTION_PARQUET_PUSHDOWN_FILTERS: "true"
SIMULATE_LATENCY: "true"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark tpch
env:
SIMULATE_LATENCY: "true"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"CPU Details (lscpu)Details
Resource Usagetpch — base (merge-base)
tpch — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark tpch
env:
DATAFUSION_EXECUTION_PARQUET_PUSHDOWN_FILTERS: "true"
SIMULATE_LATENCY: "true"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"CPU Details (lscpu)Details
Resource Usagetpch — base (merge-base)
tpch — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark tpch
env:
DATAFUSION_EXECUTION_PARQUET_PUSHDOWN_FILTERS: "true"
DATAFUSION_RUNTIME_MEMORY_LIMIT: "1G"
SIMULATE_LATENCY: "true"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"CPU Details (lscpu)Details
Memory Pool PeaksPeak Base:
Pool accounting vs. process RSS Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.
Resource Usagetpch — base (merge-base)
tpch — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark tpch10
env:
SIMULATE_LATENCY: "true"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"CPU Details (lscpu)Details
Resource Usagetpch10 — base (merge-base)
tpch10 — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark tpcds
env:
SIMULATE_LATENCY: "true"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"CPU Details (lscpu)Details
Resource Usagetpcds — base (merge-base)
tpcds — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark clickbench_partitioned
env:
SIMULATE_LATENCY: "true"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"CPU Details (lscpu)Details
Resource Usageclickbench_partitioned — base (merge-base)
clickbench_partitioned — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark clickbench_partitioned
env:
DATAFUSION_EXECUTION_PARQUET_PUSHDOWN_FILTERS: "true"
SIMULATE_LATENCY: "true"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"CPU Details (lscpu)Details
Resource Usageclickbench_partitioned — base (merge-base)
clickbench_partitioned — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark tpcds
env:
DATAFUSION_EXECUTION_PARQUET_PUSHDOWN_FILTERS: "true"
SIMULATE_LATENCY: "true"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"CPU Details (lscpu)Details
Resource Usagetpcds — base (merge-base)
tpcds — branch
File an issue against this benchmark runner |
Benchmark summary:
|
| suite | query total base → branch | speedup | previous | faster / slower / same | peak memory base → branch |
|---|---|---|---|---|---|
| tpch_sf1 | 17.02 s → 9.27 s | 1.84x | 1.75x | 21 / 0 / 1 | 890 MiB → 1.0 GiB |
| tpch_sf10 | 113.68 s → 15.71 s | 7.24x | 7.03x | 22 / 0 / 0 | 3.6 GiB → 4.4 GiB |
| tpcds_sf1 | 55.62 s → 56.43 s | 0.99x | 1.00x | 37 / 46 / 16 | 1.1 GiB → 1.0 GiB |
| clickbench_partitioned | 84.45 s → 42.64 s | 1.98x | 1.96x | 38 / 1 / 4 | 16.4 GiB → 17.0 GiB |
With SIMULATE_LATENCY, pushdown_filters=true
| suite | query total base → branch | speedup | previous | faster / slower / same | peak memory base → branch |
|---|---|---|---|---|---|
| tpch_sf1 | 38.58 s → 9.52 s | 4.05x | 3.73x | 22 / 0 / 0 | 677 MiB → 713 MiB |
| tpcds_sf1 | 107.74 s → 62.91 s | 1.71x | 1.72x | 94 / 5 / 0 | 1.2 GiB → 1.1 GiB |
| clickbench_partitioned | 105.25 s → 39.09 s | 2.69x | 2.75x | 41 / 0 / 2 | 9.5 GiB → 11.0 GiB |
tpch_sf1, DATAFUSION_RUNTIME_MEMORY_LIMIT=1G |
39.53 s → 9.63 s | 4.10x | 3.82x | 22 / 0 / 0 | 674 MiB → 847 MiB |
Notes
- The new arrow-rs pin does not change the results. All speedups are within a few percent of
4a9f276. The tpch pushdown runs are a little faster (4.05x and 4.10x, previously 3.73x and 3.82x). - tpcds without pushdown is flat, as before (0.99x). The query counts changed from 9 / 12 / 78 to 37 / 46 / 16 because the random latency draws make each query more variable: the standard deviation on the base side is 30–60% of the mean for many queries. The mean-based total is also flat (68.4 s → 69.3 s).
- tpcds Q27 is no longer slower in the no-pushdown run: its mean is 1.12x faster. It is 1.15x slower with pushdown (189 → 219 ms), previously 1.55x.
- tpcds Q1 and Q44 are slower in both tpcds runs: Q1 is 1.29x slower with pushdown and 1.56x without, and Q44 is 1.31x slower with pushdown and 1.33x (mean) without. These are small queries (100–250 ms). I will profile them.
- The 1 GB pool run has no failures.
Runs: pushdown + latency, latency, 1 GB pool.
|
run benchmark tpch tpch10 tpcds clickbench_partitioned env:
SIMULATE_LATENCY: "false"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600" |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark tpch
env:
SIMULATE_LATENCY: "false"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark tpch10
env:
SIMULATE_LATENCY: "false"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark tpcds
env:
SIMULATE_LATENCY: "false"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark clickbench_partitioned
env:
SIMULATE_LATENCY: "false"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark tpch
env:
SIMULATE_LATENCY: "false"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"CPU Details (lscpu)Details
Resource Usagetpch — base (merge-base)
tpch — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark tpcds
env:
SIMULATE_LATENCY: "false"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"CPU Details (lscpu)Details
Resource Usagetpcds — base (merge-base)
tpcds — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark tpch10
env:
SIMULATE_LATENCY: "false"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"CPU Details (lscpu)Details
Resource Usagetpch10 — base (merge-base)
tpch10 — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing claude/datafusion-rowgroup-buffering-ec9abc (02429ca) to 953518c (merge-base) diff Run configurationrun benchmark clickbench_partitioned
env:
SIMULATE_LATENCY: "false"
changed:
env:
DATAFUSION_EXECUTION_PARQUET_READ_AHEAD_BYTES: "104857600"CPU Details (lscpu)Details
Resource Usageclickbench_partitioned — base (merge-base)
clickbench_partitioned — branch
File an issue against this benchmark runner |
…robin (apache#25809) ## Which issue does this PR close? - No issue. Found while benchmarking apache#24086. ## Rationale for this change With `SIMULATE_LATENCY=true`, a benchmark result can change when a PR changes the number of object store requests in *other* queries, or off the critical path of the same query. `LatencyObjectStore` gives each request the next entry of a 20-entry latency table, with one counter for the whole process: ```rust let idx = self.get_counter.fetch_add(1, Ordering::Relaxed) % GET_LATENCIES_MS.len(); ``` So the latency of a request depends on how many requests came before it, in all earlier queries of the run. The simulator is deterministic, so the same wrong result comes back on every rerun. Example from apache#24086 (TPC-DS SF1, `SIMULATE_LATENCY=true`): the bot reported Q27 1.55x slower (141 → 218 ms) and Q44 1.38x slower. Both queries have only two serial round trips on the critical path. Q44 makes identical requests with and without the change. Local runs, median of 9, ms: | latency model | Q27 base | Q27 PR | Q44 base | Q44 PR | |---|---|---|---|---| | fixed 50 ms per request | 111 | 111 | 114 | 121 | | fixed 100 ms per request | 216 | 211 | – | – | | current round-robin table | 238 | 288 | 212 | 211 | A sweep over the 20 possible start positions of the counter (100 iterations each) gives Q27 medians of 239 ms (base) and 241 ms (PR). The min-of-5 that the bot reports depends on the start position: the base build gets a lucky minimum at 11 of 20 positions, the PR build at 5. ## What changes are included in this PR? `LatencyObjectStore` draws each GET and LIST latency independently at random from the same 20-entry distributions. The distributions do not change. | approach | latency of a request depends on | result | |---|---|---| | shared counter (before) | the number of earlier requests in the process | a change in one query moves latencies in other queries | | hash of the request (not chosen) | the request's path and byte range | a request shape gets the same latency in every iteration, so builds with different request shapes do not average out | | independent random draw (this PR) | nothing | each request samples the same distribution, and iterations average out | ## What is the testing strategy for this PR? This change affects only the benchmark harness. - With this change, Q27 over 99 iterations gives base median 217 ms and PR median 224 ms (mean 223.5 ± 7.5 vs 217.6 ± 6.9), so there is no difference, as with a fixed latency. - Queries that do read data keep their differences: TPC-DS Q3 703 → 353 ms, Q7 1078 → 526 ms, TPC-H Q1 1006 → 209 ms, Q6 1862 → 218 ms (median of 5, same two builds). - `cargo test -p datafusion-benchmarks --lib` and clippy pass. There is no new unit test: the property ("no dependency on earlier requests") is structural, since the store has no state left. ## Are there any user-facing changes? No. Benchmark runs with `SIMULATE_LATENCY=true` are no longer bit-for-bit reproducible between runs. Queries with few serial round trips have more spread per iteration (Q27 SD about 70 ms), so compare them with more iterations or with means. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
EXPERIMENT: pin every arrow crate to pydantic/arrow-rs claude/push-decoder-batch-granular-scan-plan-60.0.0: the 60.0.0 release plus batch-granular decoding in ParquetPushDecoder (apache/arrow-rs#6946) and ParquetPushDecoder::scan_plan (apache/arrow-rs#10555). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
02429ca to
5b7407e
Compare
… read-ahead Opt-in with datafusion.execution.parquet.read_ahead_bytes (default unset). If set, the scan builds the same ParquetPushDecoder as the default path with FetchGranularity::Batch and drives it with try_decode. ReadAhead fetches ranges from scan_plan() in the background within that many bytes. datafusion.execution.parquet.read_ahead_conditional (default false) also reads ahead ranges that a pushed-down filter can make unnecessary. Filtered scans, runtime row-group pruning and the per-row-group filter toggle use the same path as main. Squashed from the history in branch claude/datafusion-rowgroup-buffering-history (pydantic/datafusion). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ranges Skipping ranges that a pushed-down filter can make unnecessary made filtered scans on object storage 2-16x slower. Read-ahead now fetches them, bounded by the read-ahead window. Proto field 40 is reserved. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each read-ahead stream registers a `ParquetReadAhead[partition]` consumer. The reservation follows the bytes the decoder holds plus the bytes in flight. Bytes the decoder asks for are always reserved (grow). Speculative read-ahead takes only what the pool can grant (try_grow); ranges that do not fit stay pending. `FileSource::create_morselizer_with_context` gives the source the scan's `TaskContext`. The default delegates to `create_morselizer`. `ParquetSource` uses it to get the memory pool. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
5b7407e to
d6d54fd
Compare
…time A scan with `read_ahead_bytes` set had the same metrics as a scan without it, so EXPLAIN ANALYZE could not show that read-ahead was active or whether it kept ahead of the decoder. - `read_ahead_bytes_fetched`: bytes fetched before the decoder asked for them. - `read_ahead_wait_time`: time the decoder waited for byte ranges that it needed. Both are registered only for a file that is read with read-ahead.
|
Possible follow-ups (prototypes, not ready for review):
|
Which issue does this PR close?
ParquetPushDecoder, with the configured row selection policy: feat(parquet): predicate cache, page release and selection policy in the batch-granular push decoder arrow-rs#11240 (stacked on feat(parquet): let ParquetPushDecoder decode a batch as soon as its pages are pushed arrow-rs#11238, which starts from the test in test(parquet): show push decoder needs a whole row group before decoding a batch arrow-rs#11218)ParquetPushDecoder::scan_plan(feat(parquet): add ParquetPushDecoder::scan_plan for read-ahead planning arrow-rs#10555)Rationale for this change
Today a Parquet scan buffers a whole row group before it decodes anything: the push decoder's
NeedsDatadoes not resolve until every projected byte of the row group has arrived. Under object-store latency this has two effects:This PR schedules I/O at batch granularity. It uses the normal push-decoder path and adds only a read-ahead layer on top.
What changes are included in this PR?
One new Parquet read option:
datafusion.execution.parquet.read_ahead_bytesNonemain.The intended use: leave it unset for local disks (no speculative reads), set it for object storage.
ParquetPushDecoderas the default path, withFetchGranularity::Batch, and drives it withtry_decode.NeedsDataasks only for the pages of the next batch, and the decoder returns the batch as soon as they are pushed. The decoder releases pages when their rows are emitted.ReadAheadfetches ahead from the decoder's plan. It readsscan_plan()in decode order and fetches in the background while decoder-held bytes plus in-flight bytes stay within the window.NeedsDatais always fetched. A blocking fetch is filled with read-ahead up to the window, except the first fetch of a file whose plan is larger than the window, so the first batch comes early.pushdown_filters, the decoder evaluates predicates one window ofbatch_sizerows at a time. Read-ahead also fetchesconditionalranges (bytes that a predicate can make unnecessary), in plan order. Without this, each window waits for one round trip per predicate plus one: with simulated latency this made TPC-H 0.50x and TPC-DS 0.14x. With it, TPC-H is 3.79x and TPC-DS 1.74x, with no slower queries. The worst case reads about the bytes thatmainreads without pushdown. This matches what DuckDB, ClickHouse (remote) and Arrow C++pre_bufferdo by default.ParquetReadAhead[partition]reservation that equals the bytes the decoder holds plus the bytes in flight. Bytes thatNeedsDatarequests usegrow: they can exceed the pool limit, so the scan never blocks or fails on them. Read-ahead usestry_grow: it fetches only what the pool can give, and tries again at the next decode step. With the default unbounded pool, behavior does not change.FileSourcegetscreate_morselizer_with_context(default: callscreate_morselizer) so that the Parquet source can get the pool from theTaskContext.main. Read-ahead skips row groups that the dynamic pruner has already pruned. Fetched bytes of row groups that are pruned later are dropped.Changes in this revision
ParquetRecordBatchReaderover a customChunkReader(SharedBuffers)FetchGranularity::Batchmaintry_decode.DF_FETCH_POLICY,DF_FETCH_BUDGET,DF_FETCH_CONDITIONALenvironment variablesread_ahead_bytesconfig option (also in proto)MemoryPoolreservation: required bytesgrow, read-aheadtry_growtarget_partitions × windowThe pinned arrow-rs branch is
pydantic/arrow-rsclaude/push-decoder-batch-granular-scan-plan-60.0.0(commit6bea7c8969): the 60.0.0 release, plus the test in apache/arrow-rs#11218, plus the selection-policy commit08d956bca6(first in apache/arrow-rs#11241, now part of apache/arrow-rs#11240), plus apache/arrow-rs#10555 at0d91bca5e1, plus the upstream fix apache/arrow-rs#11182 (a test in apache/arrow-rs#10555 needs it). Benchmarks againstmaintherefore measure only these changes. This pin is why the PR cannot merge as-is.This PR is two commits on
main: the pin, then the streaming code. The earlier revisions (sync reader,pipelinedandbatchedpolicies) are in branchclaude/datafusion-rowgroup-buffering-history.What this does not do
GreedyMemoryPoolbefore another operator asks for them. A headroom rule (for example, stop read-ahead above X% of the pool) is a follow-up that needs measurements.get_byte_ranges.Are these changes tested?
read_ahead_bytesforced to 65536 for the whole suite (a local default change, not committed), 522 of 523 pass. The one difference isinformation_schema.slt, which prints the changed default. The small window forces many read-ahead waves and releases.pushdown_filters=true(forced on with a local patch, not committed) for 18 filter-heavy files: the default and streaming policies fail on the same statements, allEXPLAINdifferences caused by pushdown itself. The--completeoutput is byte-identical between the two policies.datafusion-datasource-parquetunit tests: all pass with read-ahead off and with it forced on. The proto round-trip tests forParquetOptionspass.opener::test::read_ahead): a scan with a 64 KiB pool is correct and returns all reserved bytes; the reservation stays within the window plus one fetch and returns to 0; a filtered scan reads conditional ranges ahead with a large pool and reads exactly the baseline bytes with a 1-byte pool. A unit test covers the accounting helper.NeedsDatarequests.main.Earlier results
Previous revision (
16c5ba3, sync reader), adriangbot GKE vs merge-base, 100 MB window. Full tables are in the benchmark summary.SIMULATE_LATENCYNew runs for this revision are requested below.
Are there any user-facing changes?
Yes: one new configuration option, off by default, and one new
FileSourcemethod with a default implementation. With the defaults, behavior is the same asmain. The[patch.crates-io]pin means this PR cannot merge in its current form.🤖 Generated with Claude Code