Skip to content

feat: blocked agg with dynamic flat and blocked - #25877

Open
rluvaton wants to merge 13 commits into
apache:mainfrom
rluvaton:blocked-agg-poc
Open

rluvaton wants to merge 13 commits into
apache:mainfrom
rluvaton:blocked-agg-poc

Conversation

@rluvaton

@rluvaton rluvaton commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Difference with my other PR:

That PR goes BlockedGroupsAccumulator first which means that it only interact with blocked and any non blocked are wrapped in adapters
the reason for that is:

  • Easier deprecation since there are no usages of the GroupsAccumulator in the API
  • Cleaner code since you only support one api
  • It have performance problems until the migration is complete since most of the code will use the slow adapter

this PR however does not go that way, it support both blocked and flat without adapter,
the reason for that is to keep good performance for unsupported cases while also use blocked when possible

Important note

Currently it sort every block and spill it which will create more spill files and it reduce performance by increasing batch size merge degree on spill path, the fix is to have sort kernel that get a vector of arrays to sort and return them sorted, I did not add it in this PR to make it easier to review, but will add in later pr,

I already note that this is needed in the related issue comment with my findings
#24704 (comment)

Which issue does this PR close?

Part of:

Rationale for this change

See issue

What changes are included in this PR?

It contain BlockedGroupsAccumulator trait, helper BlockedVec, support count in blocked so you see the example usage, change the non-ordered aggregate (skipped ordered to make this pr smaller) to work with either blocked or flat, implement BlockedGroupValues for primitive so you will see how it is being used

What is the testing strategy for this PR?

added tests + existing

Are there any user-facing changes?

yes, but not breaking ones

@github-actions github-actions Bot added logical-expr Logical plan and expressions physical-expr Changes to the physical-expr crates functions Changes to functions implementation physical-plan Changes to the physical-plan crate labels Sep 29, 2026
@rluvaton

Copy link
Copy Markdown
Member Author

run benchmarks

@rluvaton

Copy link
Copy Markdown
Member Author

run benchmarks h2o_medium external_aggr

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5892367849-2905-xpccx 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (a97e6b6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark clickbench_partitioned

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5892407262-2909-pj8dp 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (a97e6b6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark external_aggr

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5892367849-2906-czvg4 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (a97e6b6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark tpcds

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5892407262-2908-gtc7b 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (a97e6b6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark h2o_medium

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5892367849-2907-bdd7t 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (a97e6b6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark tpch

Results will be posted here when complete


File an issue against this benchmark runner

/// otherwise the whole aggregation uses [`GroupsAccumulator`].
///
/// [`GroupsAccumulator`]: crate::groups_accumulator::GroupsAccumulator
pub trait BlockedGroupsAccumulator: Send + Any {

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will add the rest of the functions (evaluate/state preserving) in later pr so it will be easier to review

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (a97e6b6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark tpch
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark tpch_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃     HEAD ┃ blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │ 42.99 ms │        42.35 ms │     no change │
│ QQuery 2  │ 19.91 ms │        19.87 ms │     no change │
│ QQuery 3  │ 30.04 ms │        29.94 ms │     no change │
│ QQuery 4  │ 18.47 ms │        18.10 ms │     no change │
│ QQuery 5  │ 37.54 ms │        37.13 ms │     no change │
│ QQuery 6  │ 17.05 ms │        17.09 ms │     no change │
│ QQuery 7  │ 44.69 ms │        44.75 ms │     no change │
│ QQuery 8  │ 42.83 ms │        42.36 ms │     no change │
│ QQuery 9  │ 52.75 ms │        51.57 ms │     no change │
│ QQuery 10 │ 43.57 ms │        42.49 ms │     no change │
│ QQuery 11 │ 14.78 ms │        14.55 ms │     no change │
│ QQuery 12 │ 22.61 ms │        22.26 ms │     no change │
│ QQuery 13 │ 45.04 ms │        42.23 ms │ +1.07x faster │
│ QQuery 14 │ 25.83 ms │        25.98 ms │     no change │
│ QQuery 15 │ 32.99 ms │        32.78 ms │     no change │
│ QQuery 16 │ 14.84 ms │        14.79 ms │     no change │
│ QQuery 17 │ 82.63 ms │        72.58 ms │ +1.14x faster │
│ QQuery 18 │ 64.66 ms │        66.16 ms │     no change │
│ QQuery 19 │ 33.74 ms │        33.61 ms │     no change │
│ QQuery 20 │ 34.49 ms │        34.01 ms │     no change │
│ QQuery 21 │ 61.99 ms │        61.37 ms │     no change │
│ QQuery 22 │ 15.24 ms │        14.79 ms │     no change │
└───────────┴──────────┴─────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━┓
┃ Benchmark Summary              ┃          ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━┩
│ Total Time (HEAD)              │ 798.67ms │
│ Total Time (blocked-agg-poc)   │ 780.76ms │
│ Average Time (HEAD)            │  36.30ms │
│ Average Time (blocked-agg-poc) │  35.49ms │
│ Queries Faster                 │        2 │
│ Queries Slower                 │        0 │
│ Queries with No Change         │       20 │
│ Queries with Failure           │        0 │
└────────────────────────────────┴──────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark tpch_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                           HEAD ┃                blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │ 42.99 / 45.81 ±2.08 / 49.11 ms │ 42.35 / 43.48 ±1.05 / 45.07 ms │ +1.05x faster │
│ QQuery 2  │ 19.91 / 20.42 ±0.38 / 20.95 ms │ 19.87 / 20.11 ±0.19 / 20.43 ms │     no change │
│ QQuery 3  │ 30.04 / 30.90 ±1.43 / 33.75 ms │ 29.94 / 30.08 ±0.14 / 30.33 ms │     no change │
│ QQuery 4  │ 18.47 / 18.97 ±0.55 / 20.02 ms │ 18.10 / 18.59 ±0.67 / 19.92 ms │     no change │
│ QQuery 5  │ 37.54 / 38.10 ±0.35 / 38.57 ms │ 37.13 / 37.47 ±0.27 / 37.84 ms │     no change │
│ QQuery 6  │ 17.05 / 17.16 ±0.08 / 17.24 ms │ 17.09 / 17.77 ±0.90 / 19.50 ms │     no change │
│ QQuery 7  │ 44.69 / 48.84 ±2.32 / 51.71 ms │ 44.75 / 45.65 ±0.83 / 47.00 ms │ +1.07x faster │
│ QQuery 8  │ 42.83 / 43.09 ±0.21 / 43.47 ms │ 42.36 / 42.65 ±0.17 / 42.89 ms │     no change │
│ QQuery 9  │ 52.75 / 53.36 ±0.49 / 54.17 ms │ 51.57 / 52.64 ±0.76 / 53.90 ms │     no change │
│ QQuery 10 │ 43.57 / 44.35 ±0.78 / 45.85 ms │ 42.49 / 43.47 ±1.12 / 45.20 ms │     no change │
│ QQuery 11 │ 14.78 / 14.91 ±0.14 / 15.16 ms │ 14.55 / 15.20 ±0.60 / 16.34 ms │     no change │
│ QQuery 12 │ 22.61 / 23.07 ±0.38 / 23.57 ms │ 22.26 / 22.56 ±0.27 / 23.08 ms │     no change │
│ QQuery 13 │ 45.04 / 46.09 ±1.26 / 48.42 ms │ 42.23 / 43.82 ±2.57 / 48.95 ms │     no change │
│ QQuery 14 │ 25.83 / 26.03 ±0.13 / 26.24 ms │ 25.98 / 26.19 ±0.14 / 26.33 ms │     no change │
│ QQuery 15 │ 32.99 / 33.50 ±0.66 / 34.78 ms │ 32.78 / 35.32 ±1.62 / 37.12 ms │  1.05x slower │
│ QQuery 16 │ 14.84 / 15.03 ±0.14 / 15.22 ms │ 14.79 / 15.05 ±0.16 / 15.27 ms │     no change │
│ QQuery 17 │ 82.63 / 84.15 ±1.30 / 85.73 ms │ 72.58 / 73.71 ±1.48 / 76.62 ms │ +1.14x faster │
│ QQuery 18 │ 64.66 / 66.12 ±0.85 / 67.24 ms │ 66.16 / 67.04 ±0.77 / 68.40 ms │     no change │
│ QQuery 19 │ 33.74 / 34.54 ±1.20 / 36.88 ms │ 33.61 / 34.25 ±1.19 / 36.63 ms │     no change │
│ QQuery 20 │ 34.49 / 35.11 ±0.64 / 36.34 ms │ 34.01 / 34.94 ±0.58 / 35.61 ms │     no change │
│ QQuery 21 │ 61.99 / 63.33 ±1.47 / 65.76 ms │ 61.37 / 62.99 ±1.18 / 64.35 ms │     no change │
│ QQuery 22 │ 15.24 / 15.45 ±0.27 / 15.97 ms │ 14.79 / 15.15 ±0.23 / 15.36 ms │     no change │
└───────────┴────────────────────────────────┴────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━┓
┃ Benchmark Summary              ┃          ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━┩
│ Total Time (HEAD)              │ 818.33ms │
│ Total Time (blocked-agg-poc)   │ 798.12ms │
│ Average Time (HEAD)            │  37.20ms │
│ Average Time (blocked-agg-poc) │  36.28ms │
│ Queries Faster                 │        3 │
│ Queries Slower                 │        1 │
│ Queries with No Change         │       18 │
│ Queries with Failure           │        0 │
└────────────────────────────────┴──────────┘

Resource Usage

tpch — base (merge-base)

Metric Value
Wall time 5.0s
Peak memory 1.1 GiB
Avg memory 634.9 MiB
CPU user 22.9s
CPU sys 2.0s
Peak spill 0 B

tpch — branch

Metric Value
Wall time 5.0s
Peak memory 1.3 GiB
Avg memory 672.7 MiB
CPU user 22.5s
CPU sys 2.0s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (a97e6b6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark tpcds
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark tpcds_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Query     ┃       HEAD ┃ blocked-agg-poc ┃    Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ QQuery 1  │    5.52 ms │         5.57 ms │ no change │
│ QQuery 2  │   80.50 ms │        79.98 ms │ no change │
│ QQuery 3  │   28.91 ms │        28.84 ms │ no change │
│ QQuery 4  │  444.05 ms │       443.04 ms │ no change │
│ QQuery 5  │   50.94 ms │        50.71 ms │ no change │
│ QQuery 6  │   35.55 ms │        35.70 ms │ no change │
│ QQuery 7  │   74.51 ms │        74.72 ms │ no change │
│ QQuery 8  │   36.50 ms │        36.21 ms │ no change │
│ QQuery 9  │   52.96 ms │        53.10 ms │ no change │
│ QQuery 10 │   62.28 ms │        61.39 ms │ no change │
│ QQuery 11 │  269.48 ms │       274.31 ms │ no change │
│ QQuery 12 │   28.58 ms │        28.94 ms │ no change │
│ QQuery 13 │  117.23 ms │       116.71 ms │ no change │
│ QQuery 14 │  392.65 ms │       391.81 ms │ no change │
│ QQuery 15 │   54.01 ms │        53.41 ms │ no change │
│ QQuery 16 │    6.43 ms │         6.36 ms │ no change │
│ QQuery 17 │   78.70 ms │        77.72 ms │ no change │
│ QQuery 18 │  103.98 ms │       102.61 ms │ no change │
│ QQuery 19 │   41.85 ms │        41.62 ms │ no change │
│ QQuery 20 │   36.12 ms │        35.72 ms │ no change │
│ QQuery 21 │   17.00 ms │        17.04 ms │ no change │
│ QQuery 22 │   63.77 ms │        63.40 ms │ no change │
│ QQuery 23 │  307.17 ms │       304.77 ms │ no change │
│ QQuery 24 │  193.67 ms │       192.84 ms │ no change │
│ QQuery 25 │  106.59 ms │       106.96 ms │ no change │
│ QQuery 26 │   48.60 ms │        48.05 ms │ no change │
│ QQuery 27 │    5.89 ms │         5.90 ms │ no change │
│ QQuery 28 │   56.84 ms │        56.40 ms │ no change │
│ QQuery 29 │   94.86 ms │        94.04 ms │ no change │
│ QQuery 30 │   33.06 ms │        32.45 ms │ no change │
│ QQuery 31 │  108.47 ms │       108.92 ms │ no change │
│ QQuery 32 │   20.26 ms │        20.51 ms │ no change │
│ QQuery 33 │   37.52 ms │        37.56 ms │ no change │
│ QQuery 34 │   10.35 ms │         9.94 ms │ no change │
│ QQuery 35 │   71.93 ms │        72.36 ms │ no change │
│ QQuery 36 │    5.70 ms │         5.59 ms │ no change │
│ QQuery 37 │    6.84 ms │         6.76 ms │ no change │
│ QQuery 38 │   61.05 ms │        61.93 ms │ no change │
│ QQuery 39 │   73.65 ms │        73.86 ms │ no change │
│ QQuery 40 │   23.26 ms │        22.66 ms │ no change │
│ QQuery 41 │   11.24 ms │        10.99 ms │ no change │
│ QQuery 42 │   23.89 ms │        24.23 ms │ no change │
│ QQuery 43 │    4.88 ms │         4.71 ms │ no change │
│ QQuery 44 │    8.84 ms │         8.58 ms │ no change │
│ QQuery 45 │   36.81 ms │        36.62 ms │ no change │
│ QQuery 46 │   11.88 ms │        11.68 ms │ no change │
│ QQuery 47 │  219.96 ms │       221.79 ms │ no change │
│ QQuery 48 │   95.77 ms │        94.26 ms │ no change │
│ QQuery 49 │   70.00 ms │        70.61 ms │ no change │
│ QQuery 50 │   57.70 ms │        57.42 ms │ no change │
│ QQuery 51 │   91.94 ms │        92.26 ms │ no change │
│ QQuery 52 │   24.08 ms │        23.81 ms │ no change │
│ QQuery 53 │   29.04 ms │        28.85 ms │ no change │
│ QQuery 54 │   54.49 ms │        54.03 ms │ no change │
│ QQuery 55 │   23.65 ms │        23.69 ms │ no change │
│ QQuery 56 │   38.60 ms │        39.26 ms │ no change │
│ QQuery 57 │  167.98 ms │       170.41 ms │ no change │
│ QQuery 58 │  109.22 ms │       108.03 ms │ no change │
│ QQuery 59 │  115.57 ms │       115.15 ms │ no change │
│ QQuery 60 │   39.78 ms │        38.81 ms │ no change │
│ QQuery 61 │   11.60 ms │        11.47 ms │ no change │
│ QQuery 62 │   44.13 ms │        44.01 ms │ no change │
│ QQuery 63 │   29.41 ms │        29.44 ms │ no change │
│ QQuery 64 │  358.58 ms │       357.06 ms │ no change │
│ QQuery 65 │  128.89 ms │       131.26 ms │ no change │
│ QQuery 66 │   76.61 ms │        77.42 ms │ no change │
│ QQuery 67 │  247.33 ms │       247.78 ms │ no change │
│ QQuery 68 │   11.85 ms │        11.80 ms │ no change │
│ QQuery 69 │   56.72 ms │        56.17 ms │ no change │
│ QQuery 70 │  104.54 ms │       103.64 ms │ no change │
│ QQuery 71 │   35.63 ms │        35.90 ms │ no change │
│ QQuery 72 │ 1728.85 ms │      1769.04 ms │ no change │
│ QQuery 73 │    9.87 ms │         9.91 ms │ no change │
│ QQuery 74 │  158.61 ms │       159.59 ms │ no change │
│ QQuery 75 │  139.49 ms │       139.99 ms │ no change │
│ QQuery 76 │   34.51 ms │        34.21 ms │ no change │
│ QQuery 77 │   60.91 ms │        60.44 ms │ no change │
│ QQuery 78 │  159.89 ms │       161.85 ms │ no change │
│ QQuery 79 │   65.58 ms │        66.32 ms │ no change │
│ QQuery 80 │   95.81 ms │        96.46 ms │ no change │
│ QQuery 81 │   26.08 ms │        26.80 ms │ no change │
│ QQuery 82 │   16.35 ms │        16.49 ms │ no change │
│ QQuery 83 │   33.27 ms │        33.43 ms │ no change │
│ QQuery 84 │   29.43 ms │        29.26 ms │ no change │
│ QQuery 85 │  102.19 ms │       102.55 ms │ no change │
│ QQuery 86 │   25.64 ms │        26.05 ms │ no change │
│ QQuery 87 │   62.37 ms │        62.01 ms │ no change │
│ QQuery 88 │   61.09 ms │        60.90 ms │ no change │
│ QQuery 89 │   35.01 ms │        34.81 ms │ no change │
│ QQuery 90 │   16.44 ms │        16.28 ms │ no change │
│ QQuery 91 │   44.49 ms │        44.46 ms │ no change │
│ QQuery 92 │   29.37 ms │        29.00 ms │ no change │
│ QQuery 93 │   48.60 ms │        48.87 ms │ no change │
│ QQuery 94 │   38.50 ms │        38.24 ms │ no change │
│ QQuery 95 │   79.75 ms │        79.26 ms │ no change │
│ QQuery 96 │   23.62 ms │        23.74 ms │ no change │
│ QQuery 97 │   50.50 ms │        50.97 ms │ no change │
│ QQuery 98 │   43.17 ms │        42.58 ms │ no change │
│ QQuery 99 │   65.39 ms │        65.67 ms │ no change │
└───────────┴────────────┴─────────────────┴───────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary              ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 8972.56ms │
│ Total Time (blocked-agg-poc)   │ 9010.71ms │
│ Average Time (HEAD)            │   90.63ms │
│ Average Time (blocked-agg-poc) │   91.02ms │
│ Queries Faster                 │         0 │
│ Queries Slower                 │         0 │
│ Queries with No Change         │        99 │
│ Queries with Failure           │         0 │
└────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark tpcds_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                  HEAD ┃                       blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │           5.52 / 6.15 ±1.12 / 8.38 ms │           5.57 / 6.20 ±1.02 / 8.24 ms │     no change │
│ QQuery 2  │        80.50 / 80.70 ±0.20 / 81.06 ms │        79.98 / 80.31 ±0.27 / 80.72 ms │     no change │
│ QQuery 3  │        28.91 / 29.34 ±0.31 / 29.79 ms │        28.84 / 29.10 ±0.22 / 29.44 ms │     no change │
│ QQuery 4  │     444.05 / 447.08 ±3.49 / 453.15 ms │     443.04 / 446.56 ±2.70 / 449.69 ms │     no change │
│ QQuery 5  │        50.94 / 52.67 ±2.00 / 56.54 ms │        50.71 / 51.41 ±0.43 / 51.97 ms │     no change │
│ QQuery 6  │        35.55 / 35.92 ±0.27 / 36.30 ms │        35.70 / 36.04 ±0.22 / 36.30 ms │     no change │
│ QQuery 7  │        74.51 / 75.58 ±0.76 / 76.80 ms │        74.72 / 74.90 ±0.19 / 75.25 ms │     no change │
│ QQuery 8  │        36.50 / 36.75 ±0.20 / 37.08 ms │        36.21 / 36.50 ±0.22 / 36.75 ms │     no change │
│ QQuery 9  │        52.96 / 54.69 ±1.93 / 58.22 ms │        53.10 / 54.27 ±0.85 / 55.30 ms │     no change │
│ QQuery 10 │        62.28 / 62.60 ±0.28 / 63.02 ms │        61.39 / 61.63 ±0.19 / 61.92 ms │     no change │
│ QQuery 11 │     269.48 / 273.65 ±3.00 / 277.15 ms │     274.31 / 277.74 ±2.13 / 280.62 ms │     no change │
│ QQuery 12 │        28.58 / 29.20 ±0.47 / 29.88 ms │        28.94 / 29.21 ±0.23 / 29.57 ms │     no change │
│ QQuery 13 │     117.23 / 118.14 ±0.59 / 118.93 ms │     116.71 / 117.46 ±0.61 / 118.57 ms │     no change │
│ QQuery 14 │     392.65 / 394.90 ±1.81 / 397.59 ms │     391.81 / 396.68 ±4.26 / 402.63 ms │     no change │
│ QQuery 15 │        54.01 / 55.95 ±3.08 / 62.10 ms │        53.41 / 54.41 ±0.79 / 55.43 ms │     no change │
│ QQuery 16 │           6.43 / 6.62 ±0.21 / 7.04 ms │           6.36 / 6.57 ±0.21 / 6.97 ms │     no change │
│ QQuery 17 │        78.70 / 79.69 ±1.18 / 82.00 ms │        77.72 / 78.38 ±0.54 / 79.29 ms │     no change │
│ QQuery 18 │     103.98 / 104.86 ±0.92 / 106.40 ms │     102.61 / 104.41 ±1.64 / 107.43 ms │     no change │
│ QQuery 19 │        41.85 / 42.14 ±0.26 / 42.55 ms │        41.62 / 41.97 ±0.48 / 42.89 ms │     no change │
│ QQuery 20 │        36.12 / 36.69 ±0.73 / 38.12 ms │        35.72 / 36.47 ±0.71 / 37.75 ms │     no change │
│ QQuery 21 │        17.00 / 17.32 ±0.20 / 17.54 ms │        17.04 / 17.25 ±0.19 / 17.50 ms │     no change │
│ QQuery 22 │        63.77 / 64.23 ±0.80 / 65.84 ms │        63.40 / 64.53 ±1.33 / 67.11 ms │     no change │
│ QQuery 23 │     307.17 / 311.16 ±2.38 / 314.43 ms │     304.77 / 311.32 ±5.64 / 321.57 ms │     no change │
│ QQuery 24 │     193.67 / 197.49 ±3.38 / 202.42 ms │     192.84 / 198.16 ±6.93 / 211.66 ms │     no change │
│ QQuery 25 │     106.59 / 107.75 ±1.36 / 110.42 ms │     106.96 / 108.81 ±1.32 / 110.70 ms │     no change │
│ QQuery 26 │        48.60 / 50.78 ±3.32 / 57.36 ms │        48.05 / 48.32 ±0.19 / 48.61 ms │     no change │
│ QQuery 27 │           5.89 / 6.15 ±0.18 / 6.42 ms │           5.90 / 6.09 ±0.19 / 6.45 ms │     no change │
│ QQuery 28 │        56.84 / 60.26 ±1.72 / 61.42 ms │        56.40 / 59.61 ±1.63 / 60.77 ms │     no change │
│ QQuery 29 │       94.86 / 98.34 ±4.66 / 107.52 ms │       94.04 / 96.63 ±3.08 / 102.11 ms │     no change │
│ QQuery 30 │        33.06 / 33.78 ±0.59 / 34.70 ms │        32.45 / 32.91 ±0.24 / 33.12 ms │     no change │
│ QQuery 31 │     108.47 / 109.25 ±0.52 / 109.80 ms │     108.92 / 109.58 ±0.36 / 109.95 ms │     no change │
│ QQuery 32 │        20.26 / 20.49 ±0.17 / 20.72 ms │        20.51 / 20.68 ±0.13 / 20.92 ms │     no change │
│ QQuery 33 │        37.52 / 37.95 ±0.29 / 38.43 ms │        37.56 / 38.94 ±2.27 / 43.47 ms │     no change │
│ QQuery 34 │        10.35 / 10.86 ±0.27 / 11.13 ms │         9.94 / 10.26 ±0.23 / 10.57 ms │ +1.06x faster │
│ QQuery 35 │        71.93 / 74.03 ±1.09 / 75.04 ms │        72.36 / 72.86 ±0.37 / 73.32 ms │     no change │
│ QQuery 36 │           5.70 / 5.91 ±0.22 / 6.34 ms │           5.59 / 5.77 ±0.22 / 6.20 ms │     no change │
│ QQuery 37 │           6.84 / 6.97 ±0.08 / 7.08 ms │           6.76 / 6.95 ±0.19 / 7.24 ms │     no change │
│ QQuery 38 │        61.05 / 61.74 ±0.37 / 62.14 ms │        61.93 / 62.21 ±0.27 / 62.59 ms │     no change │
│ QQuery 39 │        73.65 / 75.18 ±1.72 / 78.53 ms │        73.86 / 74.99 ±1.50 / 77.88 ms │     no change │
│ QQuery 40 │        23.26 / 24.32 ±0.69 / 25.01 ms │        22.66 / 23.50 ±0.51 / 24.27 ms │     no change │
│ QQuery 41 │        11.24 / 11.71 ±0.65 / 12.96 ms │        10.99 / 11.24 ±0.18 / 11.54 ms │     no change │
│ QQuery 42 │        23.89 / 24.18 ±0.27 / 24.63 ms │        24.23 / 25.13 ±1.09 / 27.23 ms │     no change │
│ QQuery 43 │           4.88 / 5.01 ±0.21 / 5.41 ms │           4.71 / 4.92 ±0.23 / 5.36 ms │     no change │
│ QQuery 44 │           8.84 / 8.95 ±0.12 / 9.19 ms │           8.58 / 8.79 ±0.16 / 9.05 ms │     no change │
│ QQuery 45 │        36.81 / 37.58 ±0.62 / 38.52 ms │        36.62 / 37.03 ±0.52 / 38.02 ms │     no change │
│ QQuery 46 │        11.88 / 12.22 ±0.29 / 12.67 ms │        11.68 / 12.00 ±0.24 / 12.40 ms │     no change │
│ QQuery 47 │     219.96 / 225.15 ±5.67 / 235.56 ms │     221.79 / 226.00 ±3.02 / 229.51 ms │     no change │
│ QQuery 48 │       95.77 / 98.28 ±3.65 / 105.39 ms │        94.26 / 94.94 ±0.49 / 95.74 ms │     no change │
│ QQuery 49 │        70.00 / 71.14 ±0.83 / 72.37 ms │        70.61 / 70.76 ±0.14 / 71.03 ms │     no change │
│ QQuery 50 │        57.70 / 61.96 ±7.68 / 77.31 ms │        57.42 / 59.94 ±4.44 / 68.78 ms │     no change │
│ QQuery 51 │        91.94 / 94.05 ±1.38 / 95.49 ms │        92.26 / 93.32 ±0.83 / 94.68 ms │     no change │
│ QQuery 52 │        24.08 / 24.41 ±0.25 / 24.70 ms │        23.81 / 24.20 ±0.26 / 24.59 ms │     no change │
│ QQuery 53 │        29.04 / 29.78 ±0.56 / 30.51 ms │        28.85 / 29.15 ±0.19 / 29.43 ms │     no change │
│ QQuery 54 │        54.49 / 58.13 ±5.74 / 69.57 ms │        54.03 / 56.35 ±3.19 / 62.68 ms │     no change │
│ QQuery 55 │        23.65 / 24.55 ±0.47 / 24.99 ms │        23.69 / 24.36 ±0.67 / 25.53 ms │     no change │
│ QQuery 56 │        38.60 / 39.09 ±0.30 / 39.41 ms │        39.26 / 39.43 ±0.15 / 39.63 ms │     no change │
│ QQuery 57 │     167.98 / 172.50 ±5.75 / 183.81 ms │     170.41 / 172.71 ±2.08 / 175.92 ms │     no change │
│ QQuery 58 │     109.22 / 110.18 ±0.75 / 111.07 ms │     108.03 / 109.86 ±1.06 / 111.30 ms │     no change │
│ QQuery 59 │     115.57 / 116.96 ±1.76 / 120.44 ms │     115.15 / 116.50 ±2.29 / 121.07 ms │     no change │
│ QQuery 60 │        39.78 / 40.26 ±0.53 / 41.22 ms │        38.81 / 39.56 ±0.72 / 40.87 ms │     no change │
│ QQuery 61 │        11.60 / 11.82 ±0.13 / 11.99 ms │        11.47 / 11.64 ±0.15 / 11.90 ms │     no change │
│ QQuery 62 │        44.13 / 44.43 ±0.27 / 44.88 ms │        44.01 / 44.17 ±0.19 / 44.41 ms │     no change │
│ QQuery 63 │        29.41 / 32.31 ±5.23 / 42.75 ms │        29.44 / 31.09 ±3.03 / 37.14 ms │     no change │
│ QQuery 64 │     358.58 / 364.30 ±5.06 / 371.03 ms │     357.06 / 362.56 ±3.70 / 366.17 ms │     no change │
│ QQuery 65 │     128.89 / 131.66 ±2.06 / 134.06 ms │     131.26 / 135.45 ±2.57 / 139.05 ms │     no change │
│ QQuery 66 │        76.61 / 76.93 ±0.42 / 77.74 ms │        77.42 / 77.96 ±0.37 / 78.52 ms │     no change │
│ QQuery 67 │     247.33 / 251.83 ±2.87 / 255.97 ms │     247.78 / 253.70 ±4.97 / 260.19 ms │     no change │
│ QQuery 68 │        11.85 / 12.18 ±0.28 / 12.61 ms │        11.80 / 12.05 ±0.21 / 12.37 ms │     no change │
│ QQuery 69 │        56.72 / 57.60 ±0.57 / 58.30 ms │        56.17 / 56.73 ±0.47 / 57.31 ms │     no change │
│ QQuery 70 │     104.54 / 107.87 ±4.72 / 117.24 ms │     103.64 / 108.71 ±9.19 / 127.08 ms │     no change │
│ QQuery 71 │        35.63 / 36.46 ±0.80 / 37.49 ms │        35.90 / 36.58 ±0.44 / 37.06 ms │     no change │
│ QQuery 72 │ 1728.85 / 1839.48 ±56.58 / 1880.34 ms │ 1769.04 / 1797.19 ±31.79 / 1858.01 ms │     no change │
│ QQuery 73 │         9.87 / 10.22 ±0.37 / 10.78 ms │         9.91 / 10.23 ±0.35 / 10.90 ms │     no change │
│ QQuery 74 │     158.61 / 160.86 ±2.94 / 166.61 ms │     159.59 / 162.42 ±2.09 / 165.64 ms │     no change │
│ QQuery 75 │     139.49 / 140.64 ±0.76 / 141.44 ms │     139.99 / 141.60 ±1.24 / 143.36 ms │     no change │
│ QQuery 76 │        34.51 / 36.68 ±3.34 / 43.29 ms │        34.21 / 35.36 ±1.54 / 38.38 ms │     no change │
│ QQuery 77 │        60.91 / 61.69 ±0.66 / 62.83 ms │        60.44 / 61.89 ±1.12 / 63.53 ms │     no change │
│ QQuery 78 │     159.89 / 163.34 ±2.55 / 166.97 ms │     161.85 / 169.01 ±7.41 / 180.78 ms │     no change │
│ QQuery 79 │        65.58 / 66.09 ±0.42 / 66.85 ms │        66.32 / 67.10 ±0.61 / 68.20 ms │     no change │
│ QQuery 80 │       95.81 / 98.62 ±3.49 / 105.42 ms │        96.46 / 97.02 ±0.54 / 97.91 ms │     no change │
│ QQuery 81 │        26.08 / 26.53 ±0.29 / 26.96 ms │        26.80 / 27.68 ±1.34 / 30.31 ms │     no change │
│ QQuery 82 │        16.35 / 16.60 ±0.22 / 16.90 ms │        16.49 / 16.72 ±0.26 / 17.19 ms │     no change │
│ QQuery 83 │        33.27 / 34.30 ±1.04 / 36.23 ms │        33.43 / 33.81 ±0.23 / 34.16 ms │     no change │
│ QQuery 84 │        29.43 / 29.65 ±0.31 / 30.26 ms │        29.26 / 30.56 ±1.88 / 34.26 ms │     no change │
│ QQuery 85 │     102.19 / 106.53 ±6.83 / 120.15 ms │     102.55 / 104.85 ±3.53 / 111.85 ms │     no change │
│ QQuery 86 │        25.64 / 26.08 ±0.29 / 26.52 ms │        26.05 / 26.30 ±0.23 / 26.71 ms │     no change │
│ QQuery 87 │        62.37 / 63.22 ±1.29 / 65.77 ms │        62.01 / 63.00 ±0.88 / 64.33 ms │     no change │
│ QQuery 88 │        61.09 / 61.28 ±0.32 / 61.91 ms │        60.90 / 61.35 ±0.34 / 61.86 ms │     no change │
│ QQuery 89 │        35.01 / 37.24 ±3.24 / 43.69 ms │        34.81 / 35.32 ±0.53 / 36.14 ms │ +1.05x faster │
│ QQuery 90 │        16.44 / 16.83 ±0.24 / 17.07 ms │        16.28 / 18.51 ±3.73 / 25.95 ms │  1.10x slower │
│ QQuery 91 │        44.49 / 45.76 ±1.28 / 48.22 ms │        44.46 / 45.02 ±0.32 / 45.33 ms │     no change │
│ QQuery 92 │        29.37 / 29.95 ±0.41 / 30.48 ms │        29.00 / 29.46 ±0.30 / 29.92 ms │     no change │
│ QQuery 93 │        48.60 / 49.75 ±0.72 / 50.77 ms │        48.87 / 50.43 ±2.04 / 54.41 ms │     no change │
│ QQuery 94 │        38.50 / 38.64 ±0.10 / 38.75 ms │        38.24 / 38.53 ±0.22 / 38.85 ms │     no change │
│ QQuery 95 │        79.75 / 80.90 ±0.90 / 82.15 ms │        79.26 / 80.44 ±0.70 / 81.38 ms │     no change │
│ QQuery 96 │        23.62 / 24.00 ±0.20 / 24.18 ms │        23.74 / 23.91 ±0.16 / 24.11 ms │     no change │
│ QQuery 97 │        50.50 / 52.28 ±1.43 / 54.24 ms │        50.97 / 52.51 ±0.99 / 53.67 ms │     no change │
│ QQuery 98 │        43.17 / 43.46 ±0.21 / 43.73 ms │        42.58 / 43.46 ±0.83 / 44.71 ms │     no change │
│ QQuery 99 │        65.39 / 68.28 ±4.14 / 76.45 ms │        65.67 / 67.80 ±3.30 / 74.36 ms │     no change │
└───────────┴───────────────────────────────────────┴───────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary              ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 9219.66ms │
│ Total Time (blocked-agg-poc)   │ 9167.90ms │
│ Average Time (HEAD)            │   93.13ms │
│ Average Time (blocked-agg-poc) │   92.61ms │
│ Queries Faster                 │         2 │
│ Queries Slower                 │         1 │
│ Queries with No Change         │        96 │
│ Queries with Failure           │         0 │
└────────────────────────────────┴───────────┘

Resource Usage

tpcds — base (merge-base)

Metric Value
Wall time 50.0s
Peak memory 2.0 GiB
Avg memory 1.4 GiB
CPU user 203.2s
CPU sys 5.8s
Peak spill 0 B

tpcds — branch

Metric Value
Wall time 50.0s
Peak memory 1.9 GiB
Avg memory 1.3 GiB
CPU user 203.2s
CPU sys 5.8s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (a97e6b6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark external_aggr
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark external_aggr.json
--------------------
┏━━━━━━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query        ┃      HEAD ┃ blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ Q1(64.0 MB)  │  53.69 ms │        53.59 ms │     no change │
│ Q1(32.0 MB)  │  52.16 ms │        50.59 ms │     no change │
│ Q1(16.0 MB)  │  54.43 ms │        50.94 ms │ +1.07x faster │
│ Q2(512.0 MB) │ 283.82 ms │       268.89 ms │ +1.06x faster │
│ Q2(256.0 MB) │ 270.61 ms │       271.44 ms │     no change │
│ Q2(128.0 MB) │ 241.07 ms │       243.47 ms │     no change │
│ Q2(64.0 MB)  │ 239.51 ms │       236.93 ms │     no change │
│ Q2(32.0 MB)  │ 301.20 ms │       296.85 ms │     no change │
└──────────────┴───────────┴─────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary              ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 1496.49ms │
│ Total Time (blocked-agg-poc)   │ 1472.69ms │
│ Average Time (HEAD)            │  187.06ms │
│ Average Time (blocked-agg-poc) │  184.09ms │
│ Queries Faster                 │         2 │
│ Queries Slower                 │         0 │
│ Queries with No Change         │         6 │
│ Queries with Failure           │         0 │
└────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark external_aggr.json
--------------------
┏━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query        ┃                               HEAD ┃                    blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ Q1(64.0 MB)  │     53.69 / 56.68 ±2.73 / 61.42 ms │     53.59 / 56.66 ±4.22 / 65.00 ms │     no change │
│ Q1(32.0 MB)  │     52.16 / 53.31 ±0.91 / 54.58 ms │     50.59 / 51.52 ±0.76 / 52.75 ms │     no change │
│ Q1(16.0 MB)  │     54.43 / 55.48 ±1.27 / 57.98 ms │     50.94 / 52.52 ±1.42 / 54.86 ms │ +1.06x faster │
│ Q2(512.0 MB) │  283.82 / 291.74 ±8.64 / 305.87 ms │ 268.89 / 285.79 ±10.81 / 298.43 ms │     no change │
│ Q2(256.0 MB) │ 270.61 / 281.54 ±11.22 / 303.05 ms │  271.44 / 281.30 ±9.30 / 296.90 ms │     no change │
│ Q2(128.0 MB) │  241.07 / 248.13 ±8.79 / 265.46 ms │  243.47 / 245.21 ±1.15 / 247.07 ms │     no change │
│ Q2(64.0 MB)  │ 239.51 / 246.79 ±11.36 / 269.37 ms │  236.93 / 239.62 ±3.51 / 246.47 ms │     no change │
│ Q2(32.0 MB)  │  301.20 / 306.02 ±3.34 / 309.40 ms │  296.85 / 301.13 ±3.94 / 307.56 ms │     no change │
└──────────────┴────────────────────────────────────┴────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary              ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 1539.70ms │
│ Total Time (blocked-agg-poc)   │ 1513.76ms │
│ Average Time (HEAD)            │  192.46ms │
│ Average Time (blocked-agg-poc) │  189.22ms │
│ Queries Faster                 │         1 │
│ Queries Slower                 │         0 │
│ Queries with No Change         │         7 │
│ Queries with Failure           │         0 │
└────────────────────────────────┴───────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only for benchmarks that write a results JSON (the dfbench suites, and the suites bench.sh runs through the Criterion SQL harness since #25644), and only under a memory limit: DATAFUSION_RUNTIME_MEMORY_LIMIT, or one the suite sets itself.

Base: dbf62c1 (merge-base) | Changed: blocked-agg-poc

external_aggr

Query Base Changed Change
1(64.0 MB) 36.2 MiB 36.8 MiB +1.6%
1(32.0 MB) 18.6 MiB 18.3 MiB -2.0%
1(16.0 MB) 10.2 MiB 9.6 MiB -6.0%
2(512.0 MB) 138.2 MiB 139.1 MiB +0.6%
2(256.0 MB) 97.5 MiB 98.0 MiB +0.5%
2(128.0 MB) 49.5 MiB 49.2 MiB -0.7%
2(64.0 MB) 32.5 MiB 32.5 MiB +0.0%
2(32.0 MB) 21.0 MiB 17.5 MiB -16.7%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
external_aggr base (dbf62c1 (merge-base)) 138.2 MiB 440.2 MiB 302.0 MiB 3.2×
external_aggr changed (blocked-agg-poc) 139.1 MiB 460.8 MiB 321.7 MiB 3.3×
Resource Usage

external_aggr — base (merge-base)

Metric Value
Wall time 105.0s
Peak memory 440.2 MiB
Avg memory 28.6 MiB
CPU user 25.7s
CPU sys 3.8s
Peak spill 93.1 MiB

external_aggr — branch

Metric Value
Wall time 110.0s
Peak memory 460.8 MiB
Avg memory 25.5 MiB
CPU user 21.9s
CPU sys 3.3s
Peak spill 93.1 MiB

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (a97e6b6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark clickbench_partitioned
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark clickbench_partitioned.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃       HEAD ┃ blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 0  │    1.27 ms │         1.22 ms │     no change │
│ QQuery 1  │   12.00 ms │        11.71 ms │     no change │
│ QQuery 2  │   37.45 ms │        37.24 ms │     no change │
│ QQuery 3  │   32.81 ms │        31.78 ms │     no change │
│ QQuery 4  │  253.84 ms │       211.47 ms │ +1.20x faster │
│ QQuery 5  │  286.93 ms │       272.15 ms │ +1.05x faster │
│ QQuery 6  │    1.36 ms │         1.30 ms │     no change │
│ QQuery 7  │   13.55 ms │        13.13 ms │     no change │
│ QQuery 8  │  344.87 ms │       329.21 ms │     no change │
│ QQuery 9  │  482.98 ms │       463.71 ms │     no change │
│ QQuery 10 │   66.52 ms │        63.81 ms │     no change │
│ QQuery 11 │   77.30 ms │        75.63 ms │     no change │
│ QQuery 12 │  268.83 ms │       261.30 ms │     no change │
│ QQuery 13 │  369.69 ms │       368.82 ms │     no change │
│ QQuery 14 │  287.80 ms │       278.22 ms │     no change │
│ QQuery 15 │  291.85 ms │       269.59 ms │ +1.08x faster │
│ QQuery 16 │  630.81 ms │       623.22 ms │     no change │
│ QQuery 17 │  633.26 ms │       631.04 ms │     no change │
│ QQuery 18 │ 1269.63 ms │      1282.06 ms │     no change │
│ QQuery 19 │   27.88 ms │        27.59 ms │     no change │
│ QQuery 20 │  518.07 ms │       524.74 ms │     no change │
│ QQuery 21 │  511.27 ms │       515.00 ms │     no change │
│ QQuery 22 │ 1018.51 ms │      1008.17 ms │     no change │
│ QQuery 23 │ 3151.41 ms │      3320.65 ms │  1.05x slower │
│ QQuery 24 │   42.23 ms │        41.61 ms │     no change │
│ QQuery 25 │  105.47 ms │       105.17 ms │     no change │
│ QQuery 26 │   41.36 ms │        42.06 ms │     no change │
│ QQuery 27 │  514.56 ms │       509.90 ms │     no change │
│ QQuery 28 │ 2884.77 ms │      2890.05 ms │     no change │
│ QQuery 29 │   41.33 ms │        41.21 ms │     no change │
│ QQuery 30 │  306.32 ms │       306.83 ms │     no change │
│ QQuery 31 │  283.64 ms │       286.77 ms │     no change │
│ QQuery 32 │ 1071.07 ms │      1144.63 ms │  1.07x slower │
│ QQuery 33 │ 1541.63 ms │      1550.05 ms │     no change │
│ QQuery 34 │ 1577.34 ms │      1582.82 ms │     no change │
│ QQuery 35 │  303.30 ms │       256.90 ms │ +1.18x faster │
│ QQuery 36 │   67.57 ms │        66.70 ms │     no change │
│ QQuery 37 │   35.01 ms │        35.63 ms │     no change │
│ QQuery 38 │   43.11 ms │        44.28 ms │     no change │
│ QQuery 39 │  141.71 ms │       147.44 ms │     no change │
│ QQuery 40 │   14.11 ms │        13.97 ms │     no change │
│ QQuery 41 │   13.56 ms │        13.46 ms │     no change │
│ QQuery 42 │   13.16 ms │        13.15 ms │     no change │
└───────────┴────────────┴─────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 19631.13ms │
│ Total Time (blocked-agg-poc)   │ 19715.37ms │
│ Average Time (HEAD)            │   456.54ms │
│ Average Time (blocked-agg-poc) │   458.50ms │
│ Queries Faster                 │          4 │
│ Queries Slower                 │          2 │
│ Queries with No Change         │         37 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark clickbench_partitioned.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                   HEAD ┃                       blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 0  │           1.27 / 4.21 ±5.70 / 15.60 ms │          1.22 / 3.95 ±5.37 / 14.70 ms │ +1.07x faster │
│ QQuery 1  │         12.00 / 12.27 ±0.23 / 12.56 ms │        11.71 / 11.93 ±0.13 / 12.11 ms │     no change │
│ QQuery 2  │         37.45 / 37.68 ±0.23 / 38.11 ms │        37.24 / 37.40 ±0.13 / 37.56 ms │     no change │
│ QQuery 3  │         32.81 / 33.51 ±0.53 / 34.33 ms │        31.78 / 31.94 ±0.13 / 32.10 ms │     no change │
│ QQuery 4  │      253.84 / 260.30 ±3.75 / 265.57 ms │     211.47 / 214.97 ±2.71 / 219.67 ms │ +1.21x faster │
│ QQuery 5  │      286.93 / 293.00 ±3.17 / 296.24 ms │     272.15 / 275.42 ±2.13 / 278.48 ms │ +1.06x faster │
│ QQuery 6  │            1.36 / 1.52 ±0.25 / 2.01 ms │           1.30 / 1.45 ±0.23 / 1.90 ms │     no change │
│ QQuery 7  │         13.55 / 13.63 ±0.06 / 13.71 ms │        13.13 / 13.25 ±0.18 / 13.60 ms │     no change │
│ QQuery 8  │      344.87 / 349.73 ±2.83 / 353.51 ms │     329.21 / 332.05 ±1.68 / 333.72 ms │ +1.05x faster │
│ QQuery 9  │     482.98 / 501.44 ±13.21 / 522.99 ms │    463.71 / 477.38 ±14.83 / 503.96 ms │     no change │
│ QQuery 10 │         66.52 / 67.85 ±1.04 / 69.45 ms │        63.81 / 66.52 ±4.43 / 75.35 ms │     no change │
│ QQuery 11 │         77.30 / 77.93 ±0.40 / 78.50 ms │        75.63 / 77.65 ±1.92 / 80.34 ms │     no change │
│ QQuery 12 │      268.83 / 273.60 ±3.84 / 279.48 ms │    261.30 / 274.45 ±11.45 / 293.31 ms │     no change │
│ QQuery 13 │     369.69 / 388.19 ±15.74 / 415.02 ms │    368.82 / 387.04 ±14.07 / 411.50 ms │     no change │
│ QQuery 14 │      287.80 / 293.46 ±4.85 / 300.01 ms │     278.22 / 285.73 ±5.02 / 293.06 ms │     no change │
│ QQuery 15 │     291.85 / 303.57 ±12.64 / 326.59 ms │    269.59 / 286.89 ±16.14 / 312.89 ms │ +1.06x faster │
│ QQuery 16 │      630.81 / 640.94 ±8.07 / 651.09 ms │    623.22 / 642.20 ±14.50 / 659.64 ms │     no change │
│ QQuery 17 │      633.26 / 644.55 ±9.20 / 661.30 ms │    631.04 / 644.97 ±10.62 / 662.17 ms │     no change │
│ QQuery 18 │  1269.63 / 1306.89 ±23.54 / 1344.03 ms │ 1282.06 / 1311.94 ±24.85 / 1343.66 ms │     no change │
│ QQuery 19 │        27.88 / 35.36 ±11.67 / 58.22 ms │        27.59 / 27.85 ±0.16 / 28.01 ms │ +1.27x faster │
│ QQuery 20 │     518.07 / 532.70 ±16.72 / 563.07 ms │     524.74 / 532.78 ±5.61 / 538.87 ms │     no change │
│ QQuery 21 │      511.27 / 522.37 ±9.76 / 540.29 ms │     515.00 / 521.54 ±4.77 / 529.68 ms │     no change │
│ QQuery 22 │   1018.51 / 1029.65 ±7.52 / 1041.18 ms │ 1008.17 / 1027.39 ±26.58 / 1080.03 ms │     no change │
│ QQuery 23 │ 3151.41 / 3353.75 ±156.08 / 3617.83 ms │ 3320.65 / 3413.51 ±85.37 / 3569.15 ms │     no change │
│ QQuery 24 │       42.23 / 71.20 ±53.81 / 178.74 ms │     41.61 / 119.30 ±85.79 / 275.80 ms │  1.68x slower │
│ QQuery 25 │     105.47 / 118.04 ±20.46 / 158.82 ms │    105.17 / 136.89 ±60.47 / 257.79 ms │  1.16x slower │
│ QQuery 26 │         41.36 / 44.55 ±6.14 / 56.82 ms │        42.06 / 44.48 ±3.52 / 51.45 ms │     no change │
│ QQuery 27 │      514.56 / 521.52 ±4.71 / 527.45 ms │     509.90 / 524.13 ±9.26 / 532.89 ms │     no change │
│ QQuery 28 │  2884.77 / 2920.39 ±26.06 / 2963.68 ms │ 2890.05 / 2923.39 ±18.42 / 2942.02 ms │     no change │
│ QQuery 29 │         41.33 / 42.64 ±1.76 / 46.04 ms │       41.21 / 50.77 ±18.27 / 87.30 ms │  1.19x slower │
│ QQuery 30 │      306.32 / 313.49 ±6.64 / 323.18 ms │     306.83 / 312.85 ±4.19 / 317.11 ms │     no change │
│ QQuery 31 │      283.64 / 295.59 ±6.89 / 304.29 ms │     286.77 / 294.85 ±8.86 / 310.48 ms │     no change │
│ QQuery 32 │  1071.07 / 1141.02 ±47.24 / 1195.94 ms │ 1144.63 / 1228.52 ±84.97 / 1357.04 ms │  1.08x slower │
│ QQuery 33 │  1541.63 / 1615.72 ±44.67 / 1654.67 ms │ 1550.05 / 1602.36 ±35.06 / 1649.28 ms │     no change │
│ QQuery 34 │  1577.34 / 1646.41 ±49.66 / 1707.99 ms │ 1582.82 / 1668.73 ±56.90 / 1755.55 ms │     no change │
│ QQuery 35 │     303.30 / 346.68 ±54.37 / 450.31 ms │    256.90 / 308.16 ±56.64 / 388.27 ms │ +1.13x faster │
│ QQuery 36 │         67.57 / 74.42 ±4.36 / 79.17 ms │        66.70 / 76.37 ±9.12 / 93.53 ms │     no change │
│ QQuery 37 │         35.01 / 40.31 ±8.97 / 58.18 ms │        35.63 / 39.72 ±3.38 / 44.54 ms │     no change │
│ QQuery 38 │         43.11 / 45.12 ±3.14 / 51.38 ms │        44.28 / 46.23 ±1.82 / 49.39 ms │     no change │
│ QQuery 39 │      141.71 / 148.49 ±4.55 / 154.74 ms │     147.44 / 156.93 ±6.85 / 167.91 ms │  1.06x slower │
│ QQuery 40 │         14.11 / 14.89 ±1.03 / 16.88 ms │        13.97 / 14.53 ±0.61 / 15.51 ms │     no change │
│ QQuery 41 │         13.56 / 16.62 ±4.13 / 24.27 ms │        13.46 / 14.79 ±2.12 / 19.01 ms │ +1.12x faster │
│ QQuery 42 │         13.16 / 13.37 ±0.20 / 13.69 ms │        13.15 / 14.57 ±2.24 / 19.00 ms │  1.09x slower │
└───────────┴────────────────────────────────────────┴───────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 20408.57ms │
│ Total Time (blocked-agg-poc)   │ 20477.78ms │
│ Average Time (HEAD)            │   474.62ms │
│ Average Time (blocked-agg-poc) │   476.23ms │
│ Queries Faster                 │          8 │
│ Queries Slower                 │          6 │
│ Queries with No Change         │         29 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Resource Usage

clickbench_partitioned — base (merge-base)

Metric Value
Wall time 105.0s
Peak memory 16.8 GiB
Avg memory 6.0 GiB
CPU user 1008.1s
CPU sys 99.0s
Peak spill 0 B

clickbench_partitioned — branch

Metric Value
Wall time 105.0s
Peak memory 18.2 GiB
Avg memory 6.0 GiB
CPU user 995.4s
CPU sys 104.8s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (a97e6b6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark h2o_medium
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark h2o.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃       HEAD ┃ blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │ 1126.24 ms │      1095.87 ms │     no change │
│ QQuery 2  │ 2513.15 ms │      2428.53 ms │     no change │
│ QQuery 3  │ 2222.87 ms │      2107.16 ms │ +1.05x faster │
│ QQuery 4  │ 1429.42 ms │      1400.91 ms │     no change │
│ QQuery 5  │ 2051.58 ms │      1854.61 ms │ +1.11x faster │
│ QQuery 6  │ 1733.23 ms │      1657.23 ms │     no change │
│ QQuery 7  │ 2027.95 ms │      1916.29 ms │ +1.06x faster │
│ QQuery 8  │ 3615.92 ms │      3549.73 ms │     no change │
│ QQuery 9  │ 3038.90 ms │      2864.21 ms │ +1.06x faster │
│ QQuery 10 │ 3076.17 ms │      3142.83 ms │     no change │
└───────────┴────────────┴─────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 22835.44ms │
│ Total Time (blocked-agg-poc)   │ 22017.38ms │
│ Average Time (HEAD)            │  2283.54ms │
│ Average Time (blocked-agg-poc) │  2201.74ms │
│ Queries Faster                 │          4 │
│ Queries Slower                 │          0 │
│ Queries with No Change         │          6 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark h2o.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                  HEAD ┃                       blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │ 1126.24 / 1136.37 ±13.20 / 1155.01 ms │  1095.87 / 1098.26 ±1.75 / 1100.00 ms │     no change │
│ QQuery 2  │ 2513.15 / 2591.00 ±57.34 / 2649.59 ms │ 2428.53 / 2475.57 ±45.49 / 2537.09 ms │     no change │
│ QQuery 3  │ 2222.87 / 2240.76 ±16.81 / 2263.26 ms │ 2107.16 / 2130.47 ±27.25 / 2168.70 ms │     no change │
│ QQuery 4  │  1429.42 / 1430.44 ±0.75 / 1431.20 ms │  1400.91 / 1409.50 ±9.30 / 1422.42 ms │     no change │
│ QQuery 5  │ 2051.58 / 2093.72 ±30.26 / 2121.27 ms │ 1854.61 / 1872.24 ±12.75 / 1884.36 ms │ +1.12x faster │
│ QQuery 6  │ 1733.23 / 1765.87 ±30.64 / 1806.87 ms │  1657.23 / 1664.45 ±8.45 / 1676.30 ms │ +1.06x faster │
│ QQuery 7  │ 2027.95 / 2081.36 ±40.35 / 2125.47 ms │ 1916.29 / 1946.01 ±28.33 / 1984.14 ms │ +1.07x faster │
│ QQuery 8  │ 3615.92 / 3672.43 ±45.15 / 3726.41 ms │ 3549.73 / 3581.40 ±31.82 / 3624.93 ms │     no change │
│ QQuery 9  │ 3038.90 / 3052.10 ±16.53 / 3075.41 ms │ 2864.21 / 2921.88 ±56.18 / 2998.04 ms │     no change │
│ QQuery 10 │ 3076.17 / 3135.40 ±44.68 / 3184.08 ms │ 3142.83 / 3194.07 ±37.15 / 3229.75 ms │     no change │
└───────────┴───────────────────────────────────────┴───────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 23199.45ms │
│ Total Time (blocked-agg-poc)   │ 22293.85ms │
│ Average Time (HEAD)            │  2319.95ms │
│ Average Time (blocked-agg-poc) │  2229.38ms │
│ Queries Faster                 │          3 │
│ Queries Slower                 │          0 │
│ Queries with No Change         │          7 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Resource Usage

h2o_medium — base (merge-base)

Metric Value
Wall time 75.0s
Peak memory 12.7 GiB
Avg memory 3.0 GiB
CPU user 725.8s
CPU sys 65.1s
Peak spill 0 B

h2o_medium — branch

Metric Value
Wall time 70.0s
Peak memory 11.9 GiB
Avg memory 3.0 GiB
CPU user 696.1s
CPU sys 60.2s
Peak spill 0 B

File an issue against this benchmark runner

@rluvaton

Copy link
Copy Markdown
Member Author

run benchmarks

env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5892964298-2910-6kklh 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (c160407) to dbf62c1 (merge-base) diff

Run configuration
run benchmark clickbench_partitioned
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5892964298-2911-grsf2 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (c160407) to dbf62c1 (merge-base) diff

Run configuration
run benchmark tpcds
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5892964298-2912-4thxg 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (c160407) to dbf62c1 (merge-base) diff

Run configuration
run benchmark tpch
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (ce4e7c6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark tpcds
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark tpcds_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query     ┃       HEAD ┃ blocked-agg-poc ┃       Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ QQuery 1  │    5.79 ms │         5.69 ms │    no change │
│ QQuery 2  │   81.46 ms │        81.37 ms │    no change │
│ QQuery 3  │   28.89 ms │        28.90 ms │    no change │
│ QQuery 4  │  477.14 ms │       475.98 ms │    no change │
│ QQuery 5  │   51.55 ms │        51.73 ms │    no change │
│ QQuery 6  │   36.09 ms │        36.46 ms │    no change │
│ QQuery 7  │   74.65 ms │        74.32 ms │    no change │
│ QQuery 8  │   36.57 ms │        36.59 ms │    no change │
│ QQuery 9  │   52.49 ms │        52.58 ms │    no change │
│ QQuery 10 │   62.83 ms │        61.96 ms │    no change │
│ QQuery 11 │  302.82 ms │       304.68 ms │    no change │
│ QQuery 12 │   29.38 ms │        29.85 ms │    no change │
│ QQuery 13 │  118.10 ms │       117.67 ms │    no change │
│ QQuery 14 │  397.74 ms │       398.40 ms │    no change │
│ QQuery 15 │   58.08 ms │        57.94 ms │    no change │
│ QQuery 16 │    6.80 ms │         6.98 ms │    no change │
│ QQuery 17 │   79.37 ms │        78.93 ms │    no change │
│ QQuery 18 │  105.67 ms │       106.84 ms │    no change │
│ QQuery 19 │   41.93 ms │        42.40 ms │    no change │
│ QQuery 20 │   36.25 ms │        36.08 ms │    no change │
│ QQuery 21 │   17.52 ms │        17.77 ms │    no change │
│ QQuery 22 │   65.95 ms │        66.13 ms │    no change │
│ QQuery 23 │  324.15 ms │       317.13 ms │    no change │
│ QQuery 24 │  201.13 ms │       201.42 ms │    no change │
│ QQuery 25 │  108.27 ms │       109.05 ms │    no change │
│ QQuery 26 │   49.03 ms │        48.64 ms │    no change │
│ QQuery 27 │    6.60 ms │         6.28 ms │    no change │
│ QQuery 28 │   61.34 ms │        60.82 ms │    no change │
│ QQuery 29 │   95.25 ms │        95.91 ms │    no change │
│ QQuery 30 │   33.89 ms │        33.59 ms │    no change │
│ QQuery 31 │  110.73 ms │       111.81 ms │    no change │
│ QQuery 32 │   21.36 ms │        21.08 ms │    no change │
│ QQuery 33 │   38.30 ms │        38.39 ms │    no change │
│ QQuery 34 │   10.78 ms │        10.55 ms │    no change │
│ QQuery 35 │   75.29 ms │        75.48 ms │    no change │
│ QQuery 36 │    5.84 ms │         6.14 ms │ 1.05x slower │
│ QQuery 37 │    7.06 ms │         7.20 ms │    no change │
│ QQuery 38 │   64.03 ms │        64.44 ms │    no change │
│ QQuery 39 │   75.58 ms │        75.87 ms │    no change │
│ QQuery 40 │   24.20 ms │        24.16 ms │    no change │
│ QQuery 41 │   11.73 ms │        11.84 ms │    no change │
│ QQuery 42 │   23.86 ms │        23.99 ms │    no change │
│ QQuery 43 │    5.07 ms │         4.97 ms │    no change │
│ QQuery 44 │    9.31 ms │         9.06 ms │    no change │
│ QQuery 45 │   38.58 ms │        39.70 ms │    no change │
│ QQuery 46 │   12.55 ms │        12.03 ms │    no change │
│ QQuery 47 │  237.80 ms │       241.48 ms │    no change │
│ QQuery 48 │   95.69 ms │        95.29 ms │    no change │
│ QQuery 49 │   70.44 ms │        71.51 ms │    no change │
│ QQuery 50 │   58.74 ms │        58.99 ms │    no change │
│ QQuery 51 │   93.95 ms │        92.81 ms │    no change │
│ QQuery 52 │   23.95 ms │        24.62 ms │    no change │
│ QQuery 53 │   29.40 ms │        29.74 ms │    no change │
│ QQuery 54 │   54.59 ms │        55.17 ms │    no change │
│ QQuery 55 │   23.85 ms │        24.00 ms │    no change │
│ QQuery 56 │   39.01 ms │        39.62 ms │    no change │
│ QQuery 57 │  170.93 ms │       171.46 ms │    no change │
│ QQuery 58 │  110.01 ms │       111.36 ms │    no change │
│ QQuery 59 │  116.23 ms │       116.52 ms │    no change │
│ QQuery 60 │   39.35 ms │        39.43 ms │    no change │
│ QQuery 61 │   11.96 ms │        11.99 ms │    no change │
│ QQuery 62 │   45.02 ms │        44.97 ms │    no change │
│ QQuery 63 │   29.34 ms │        29.74 ms │    no change │
│ QQuery 64 │  369.09 ms │       364.37 ms │    no change │
│ QQuery 65 │  130.06 ms │       132.94 ms │    no change │
│ QQuery 66 │   78.16 ms │        79.46 ms │    no change │
│ QQuery 67 │  261.73 ms │       258.55 ms │    no change │
│ QQuery 68 │   12.48 ms │        12.30 ms │    no change │
│ QQuery 69 │   57.73 ms │        56.67 ms │    no change │
│ QQuery 70 │  105.23 ms │       105.78 ms │    no change │
│ QQuery 71 │   35.24 ms │        35.49 ms │    no change │
│ QQuery 72 │ 1764.75 ms │      1831.07 ms │    no change │
│ QQuery 73 │   10.57 ms │        10.89 ms │    no change │
│ QQuery 74 │  174.65 ms │       177.16 ms │    no change │
│ QQuery 75 │  142.30 ms │       142.43 ms │    no change │
│ QQuery 76 │   34.65 ms │        34.58 ms │    no change │
│ QQuery 77 │   61.63 ms │        61.43 ms │    no change │
│ QQuery 78 │  165.09 ms │       169.13 ms │    no change │
│ QQuery 79 │   67.58 ms │        67.24 ms │    no change │
│ QQuery 80 │   97.99 ms │        99.19 ms │    no change │
│ QQuery 81 │   27.34 ms │        26.88 ms │    no change │
│ QQuery 82 │   16.92 ms │        16.96 ms │    no change │
│ QQuery 83 │   34.29 ms │        34.50 ms │    no change │
│ QQuery 84 │   29.87 ms │        30.06 ms │    no change │
│ QQuery 85 │  104.76 ms │       103.17 ms │    no change │
│ QQuery 86 │   25.95 ms │        26.29 ms │    no change │
│ QQuery 87 │   64.92 ms │        64.33 ms │    no change │
│ QQuery 88 │   62.31 ms │        61.27 ms │    no change │
│ QQuery 89 │   35.03 ms │        36.14 ms │    no change │
│ QQuery 90 │   17.26 ms │        17.32 ms │    no change │
│ QQuery 91 │   45.32 ms │        45.49 ms │    no change │
│ QQuery 92 │   30.05 ms │        30.38 ms │    no change │
│ QQuery 93 │   48.93 ms │        49.67 ms │    no change │
│ QQuery 94 │   39.44 ms │        39.80 ms │    no change │
│ QQuery 95 │   81.70 ms │        81.10 ms │    no change │
│ QQuery 96 │   24.20 ms │        24.19 ms │    no change │
│ QQuery 97 │   51.34 ms │        51.64 ms │    no change │
│ QQuery 98 │   42.92 ms │        43.56 ms │    no change │
│ QQuery 99 │   66.58 ms │        67.04 ms │    no change │
└───────────┴────────────┴─────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary              ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 9249.35ms │
│ Total Time (blocked-agg-poc)   │ 9326.03ms │
│ Average Time (HEAD)            │   93.43ms │
│ Average Time (blocked-agg-poc) │   94.20ms │
│ Queries Faster                 │         0 │
│ Queries Slower                 │         1 │
│ Queries with No Change         │        98 │
│ Queries with Failure           │         0 │
└────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark tpcds_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query     ┃                                   HEAD ┃                       blocked-agg-poc ┃       Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ QQuery 1  │            5.79 / 6.40 ±0.94 / 8.26 ms │           5.69 / 6.33 ±1.00 / 8.31 ms │    no change │
│ QQuery 2  │         81.46 / 81.86 ±0.31 / 82.35 ms │        81.37 / 81.58 ±0.17 / 81.78 ms │    no change │
│ QQuery 3  │         28.89 / 29.22 ±0.18 / 29.39 ms │        28.90 / 29.56 ±0.88 / 31.28 ms │    no change │
│ QQuery 4  │      477.14 / 480.44 ±1.94 / 482.90 ms │     475.98 / 482.75 ±4.24 / 488.80 ms │    no change │
│ QQuery 5  │         51.55 / 51.96 ±0.34 / 52.48 ms │        51.73 / 54.74 ±5.01 / 64.75 ms │ 1.05x slower │
│ QQuery 6  │         36.09 / 36.67 ±0.52 / 37.49 ms │        36.46 / 37.12 ±0.51 / 37.92 ms │    no change │
│ QQuery 7  │         74.65 / 75.47 ±0.43 / 75.85 ms │        74.32 / 74.65 ±0.22 / 74.94 ms │    no change │
│ QQuery 8  │         36.57 / 36.80 ±0.20 / 37.09 ms │        36.59 / 36.93 ±0.30 / 37.38 ms │    no change │
│ QQuery 9  │         52.49 / 54.57 ±1.37 / 56.29 ms │        52.58 / 56.18 ±3.36 / 62.39 ms │    no change │
│ QQuery 10 │         62.83 / 63.08 ±0.27 / 63.55 ms │        61.96 / 62.39 ±0.23 / 62.61 ms │    no change │
│ QQuery 11 │      302.82 / 307.31 ±2.94 / 310.56 ms │     304.68 / 306.62 ±1.36 / 308.07 ms │    no change │
│ QQuery 12 │         29.38 / 30.05 ±0.37 / 30.51 ms │        29.85 / 30.41 ±0.46 / 31.20 ms │    no change │
│ QQuery 13 │      118.10 / 118.69 ±0.39 / 119.15 ms │     117.67 / 118.62 ±0.58 / 119.42 ms │    no change │
│ QQuery 14 │      397.74 / 402.21 ±2.54 / 404.86 ms │     398.40 / 402.26 ±3.19 / 407.02 ms │    no change │
│ QQuery 15 │         58.08 / 60.45 ±2.76 / 65.34 ms │        57.94 / 58.34 ±0.48 / 59.07 ms │    no change │
│ QQuery 16 │            6.80 / 6.91 ±0.15 / 7.21 ms │           6.98 / 7.17 ±0.12 / 7.34 ms │    no change │
│ QQuery 17 │         79.37 / 80.47 ±1.21 / 82.68 ms │        78.93 / 81.13 ±2.19 / 85.24 ms │    no change │
│ QQuery 18 │      105.67 / 107.70 ±2.98 / 113.62 ms │     106.84 / 108.04 ±0.95 / 109.63 ms │    no change │
│ QQuery 19 │         41.93 / 42.45 ±0.28 / 42.78 ms │        42.40 / 42.66 ±0.21 / 42.98 ms │    no change │
│ QQuery 20 │         36.25 / 37.03 ±0.92 / 38.32 ms │        36.08 / 36.54 ±0.59 / 37.69 ms │    no change │
│ QQuery 21 │         17.52 / 17.79 ±0.34 / 18.46 ms │        17.77 / 17.96 ±0.13 / 18.17 ms │    no change │
│ QQuery 22 │         65.95 / 66.57 ±0.42 / 67.03 ms │        66.13 / 67.26 ±1.17 / 69.34 ms │    no change │
│ QQuery 23 │      324.15 / 329.15 ±3.76 / 335.38 ms │     317.13 / 324.57 ±4.55 / 331.40 ms │    no change │
│ QQuery 24 │      201.13 / 205.95 ±5.68 / 216.98 ms │     201.42 / 205.68 ±4.63 / 214.09 ms │    no change │
│ QQuery 25 │      108.27 / 111.09 ±3.72 / 118.18 ms │     109.05 / 110.23 ±1.03 / 112.00 ms │    no change │
│ QQuery 26 │         49.03 / 49.49 ±0.35 / 49.94 ms │        48.64 / 48.89 ±0.27 / 49.32 ms │    no change │
│ QQuery 27 │            6.60 / 6.72 ±0.09 / 6.88 ms │           6.28 / 6.48 ±0.21 / 6.88 ms │    no change │
│ QQuery 28 │         61.34 / 61.64 ±0.20 / 61.85 ms │        60.82 / 61.86 ±0.72 / 63.02 ms │    no change │
│ QQuery 29 │        95.25 / 98.30 ±2.79 / 102.77 ms │      95.91 / 100.84 ±9.24 / 119.32 ms │    no change │
│ QQuery 30 │         33.89 / 34.18 ±0.25 / 34.49 ms │        33.59 / 34.24 ±0.37 / 34.61 ms │    no change │
│ QQuery 31 │      110.73 / 111.61 ±0.96 / 112.86 ms │     111.81 / 112.29 ±0.25 / 112.50 ms │    no change │
│ QQuery 32 │         21.36 / 23.20 ±1.95 / 26.04 ms │        21.08 / 23.69 ±3.21 / 29.53 ms │    no change │
│ QQuery 33 │         38.30 / 38.95 ±0.50 / 39.72 ms │        38.39 / 38.66 ±0.29 / 39.17 ms │    no change │
│ QQuery 34 │         10.78 / 11.01 ±0.13 / 11.18 ms │        10.55 / 10.88 ±0.32 / 11.48 ms │    no change │
│ QQuery 35 │         75.29 / 76.08 ±0.68 / 77.32 ms │        75.48 / 76.45 ±0.65 / 77.22 ms │    no change │
│ QQuery 36 │            5.84 / 6.08 ±0.17 / 6.36 ms │           6.14 / 6.24 ±0.10 / 6.42 ms │    no change │
│ QQuery 37 │            7.06 / 7.29 ±0.13 / 7.47 ms │           7.20 / 7.29 ±0.05 / 7.35 ms │    no change │
│ QQuery 38 │         64.03 / 64.77 ±0.93 / 66.56 ms │        64.44 / 65.60 ±0.67 / 66.49 ms │    no change │
│ QQuery 39 │         75.58 / 76.31 ±0.73 / 77.46 ms │        75.87 / 76.43 ±0.60 / 77.46 ms │    no change │
│ QQuery 40 │         24.20 / 24.39 ±0.16 / 24.55 ms │        24.16 / 24.71 ±0.42 / 25.28 ms │    no change │
│ QQuery 41 │         11.73 / 11.89 ±0.11 / 12.03 ms │        11.84 / 11.89 ±0.04 / 11.97 ms │    no change │
│ QQuery 42 │         23.86 / 24.27 ±0.32 / 24.79 ms │        23.99 / 24.16 ±0.20 / 24.54 ms │    no change │
│ QQuery 43 │            5.07 / 5.28 ±0.15 / 5.49 ms │           4.97 / 5.08 ±0.12 / 5.29 ms │    no change │
│ QQuery 44 │            9.31 / 9.36 ±0.08 / 9.53 ms │           9.06 / 9.16 ±0.11 / 9.32 ms │    no change │
│ QQuery 45 │         38.58 / 40.32 ±1.31 / 42.65 ms │        39.70 / 42.38 ±2.89 / 48.00 ms │ 1.05x slower │
│ QQuery 46 │         12.55 / 13.24 ±0.79 / 14.78 ms │        12.03 / 12.60 ±0.47 / 13.47 ms │    no change │
│ QQuery 47 │      237.80 / 242.32 ±4.42 / 250.32 ms │     241.48 / 247.44 ±5.15 / 256.53 ms │    no change │
│ QQuery 48 │         95.69 / 96.28 ±0.38 / 96.73 ms │        95.29 / 96.01 ±0.46 / 96.57 ms │    no change │
│ QQuery 49 │         70.44 / 72.59 ±3.63 / 79.84 ms │        71.51 / 74.67 ±4.68 / 83.97 ms │    no change │
│ QQuery 50 │         58.74 / 59.40 ±0.59 / 60.45 ms │        58.99 / 59.58 ±0.49 / 60.25 ms │    no change │
│ QQuery 51 │         93.95 / 94.99 ±0.56 / 95.54 ms │        92.81 / 94.64 ±1.86 / 98.16 ms │    no change │
│ QQuery 52 │         23.95 / 25.47 ±2.15 / 29.74 ms │        24.62 / 25.78 ±1.33 / 27.48 ms │    no change │
│ QQuery 53 │         29.40 / 30.88 ±2.43 / 35.73 ms │        29.74 / 29.87 ±0.12 / 30.10 ms │    no change │
│ QQuery 54 │         54.59 / 54.75 ±0.18 / 55.10 ms │        55.17 / 56.19 ±0.60 / 56.92 ms │    no change │
│ QQuery 55 │         23.85 / 24.19 ±0.25 / 24.54 ms │        24.00 / 24.33 ±0.26 / 24.67 ms │    no change │
│ QQuery 56 │         39.01 / 39.63 ±0.45 / 40.36 ms │        39.62 / 39.84 ±0.24 / 40.27 ms │    no change │
│ QQuery 57 │      170.93 / 173.18 ±2.02 / 176.86 ms │     171.46 / 174.32 ±2.27 / 178.02 ms │    no change │
│ QQuery 58 │      110.01 / 111.50 ±1.43 / 113.83 ms │     111.36 / 112.60 ±1.43 / 114.93 ms │    no change │
│ QQuery 59 │      116.23 / 116.98 ±0.67 / 118.11 ms │     116.52 / 117.46 ±0.61 / 118.32 ms │    no change │
│ QQuery 60 │         39.35 / 39.72 ±0.23 / 40.01 ms │        39.43 / 39.71 ±0.24 / 40.05 ms │    no change │
│ QQuery 61 │         11.96 / 12.95 ±1.26 / 15.42 ms │        11.99 / 13.44 ±2.50 / 18.42 ms │    no change │
│ QQuery 62 │         45.02 / 46.62 ±2.26 / 51.10 ms │        44.97 / 45.43 ±0.36 / 45.96 ms │    no change │
│ QQuery 63 │         29.34 / 30.10 ±0.39 / 30.46 ms │        29.74 / 30.03 ±0.33 / 30.57 ms │    no change │
│ QQuery 64 │      369.09 / 372.97 ±3.12 / 377.11 ms │     364.37 / 370.65 ±4.83 / 377.24 ms │    no change │
│ QQuery 65 │      130.06 / 131.68 ±1.69 / 134.73 ms │     132.94 / 135.41 ±2.50 / 140.11 ms │    no change │
│ QQuery 66 │         78.16 / 79.00 ±0.87 / 80.60 ms │        79.46 / 81.44 ±2.69 / 86.79 ms │    no change │
│ QQuery 67 │      261.73 / 265.47 ±5.14 / 275.41 ms │     258.55 / 265.14 ±4.55 / 272.78 ms │    no change │
│ QQuery 68 │         12.48 / 12.59 ±0.06 / 12.64 ms │        12.30 / 12.43 ±0.09 / 12.53 ms │    no change │
│ QQuery 69 │         57.73 / 58.62 ±1.46 / 61.53 ms │        56.67 / 59.80 ±4.38 / 68.34 ms │    no change │
│ QQuery 70 │      105.23 / 107.56 ±2.76 / 112.89 ms │     105.78 / 106.87 ±1.08 / 108.86 ms │    no change │
│ QQuery 71 │         35.24 / 35.54 ±0.21 / 35.85 ms │        35.49 / 35.99 ±0.37 / 36.51 ms │    no change │
│ QQuery 72 │ 1764.75 / 1920.11 ±101.67 / 2054.98 ms │ 1831.07 / 1913.26 ±75.74 / 2042.56 ms │    no change │
│ QQuery 73 │         10.57 / 11.14 ±0.53 / 12.10 ms │        10.89 / 11.23 ±0.32 / 11.71 ms │    no change │
│ QQuery 74 │      174.65 / 178.47 ±3.17 / 184.31 ms │     177.16 / 179.70 ±1.51 / 181.88 ms │    no change │
│ QQuery 75 │      142.30 / 143.27 ±1.21 / 145.36 ms │     142.43 / 143.39 ±0.65 / 144.30 ms │    no change │
│ QQuery 76 │         34.65 / 35.67 ±1.11 / 37.63 ms │        34.58 / 37.57 ±4.70 / 46.94 ms │ 1.05x slower │
│ QQuery 77 │         61.63 / 62.69 ±0.99 / 64.49 ms │        61.43 / 62.04 ±0.47 / 62.69 ms │    no change │
│ QQuery 78 │      165.09 / 169.39 ±5.28 / 178.37 ms │     169.13 / 171.67 ±2.60 / 175.29 ms │    no change │
│ QQuery 79 │         67.58 / 68.72 ±0.94 / 70.24 ms │        67.24 / 67.53 ±0.26 / 68.02 ms │    no change │
│ QQuery 80 │       97.99 / 102.41 ±5.63 / 113.38 ms │      99.19 / 102.09 ±3.50 / 108.87 ms │    no change │
│ QQuery 81 │         27.34 / 27.93 ±0.40 / 28.47 ms │        26.88 / 27.29 ±0.27 / 27.65 ms │    no change │
│ QQuery 82 │         16.92 / 17.27 ±0.28 / 17.72 ms │        16.96 / 17.64 ±1.01 / 19.65 ms │    no change │
│ QQuery 83 │         34.29 / 34.76 ±0.32 / 35.07 ms │        34.50 / 34.89 ±0.28 / 35.21 ms │    no change │
│ QQuery 84 │         29.87 / 30.43 ±0.38 / 30.97 ms │        30.06 / 30.19 ±0.11 / 30.36 ms │    no change │
│ QQuery 85 │      104.76 / 107.58 ±3.08 / 113.13 ms │     103.17 / 109.02 ±4.58 / 117.12 ms │    no change │
│ QQuery 86 │         25.95 / 26.32 ±0.27 / 26.61 ms │        26.29 / 26.57 ±0.25 / 26.94 ms │    no change │
│ QQuery 87 │         64.92 / 65.71 ±0.51 / 66.27 ms │        64.33 / 65.06 ±0.76 / 66.43 ms │    no change │
│ QQuery 88 │         62.31 / 62.65 ±0.27 / 63.07 ms │        61.27 / 62.25 ±0.63 / 62.99 ms │    no change │
│ QQuery 89 │         35.03 / 36.94 ±2.13 / 40.98 ms │        36.14 / 39.12 ±5.47 / 50.06 ms │ 1.06x slower │
│ QQuery 90 │         17.26 / 17.45 ±0.14 / 17.61 ms │        17.32 / 17.68 ±0.24 / 17.93 ms │    no change │
│ QQuery 91 │         45.32 / 45.84 ±0.37 / 46.32 ms │        45.49 / 45.93 ±0.34 / 46.47 ms │    no change │
│ QQuery 92 │         30.05 / 30.33 ±0.21 / 30.58 ms │        30.38 / 30.75 ±0.34 / 31.22 ms │    no change │
│ QQuery 93 │         48.93 / 49.68 ±0.70 / 51.00 ms │        49.67 / 50.39 ±0.38 / 50.73 ms │    no change │
│ QQuery 94 │         39.44 / 40.45 ±1.54 / 43.50 ms │        39.80 / 40.54 ±1.10 / 42.71 ms │    no change │
│ QQuery 95 │         81.70 / 83.99 ±1.89 / 87.38 ms │        81.10 / 82.67 ±0.88 / 83.42 ms │    no change │
│ QQuery 96 │         24.20 / 24.37 ±0.21 / 24.76 ms │        24.19 / 24.28 ±0.08 / 24.43 ms │    no change │
│ QQuery 97 │         51.34 / 52.63 ±0.97 / 53.86 ms │        51.64 / 52.58 ±0.62 / 53.51 ms │    no change │
│ QQuery 98 │         42.92 / 43.96 ±0.74 / 44.74 ms │        43.56 / 45.93 ±1.83 / 49.12 ms │    no change │
│ QQuery 99 │         66.58 / 68.47 ±2.42 / 72.88 ms │        67.04 / 67.37 ±0.33 / 67.97 ms │    no change │
└───────────┴────────────────────────────────────────┴───────────────────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary              ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 9528.30ms │
│ Total Time (blocked-agg-poc)   │ 9547.22ms │
│ Average Time (HEAD)            │   96.25ms │
│ Average Time (blocked-agg-poc) │   96.44ms │
│ Queries Faster                 │         0 │
│ Queries Slower                 │         4 │
│ Queries with No Change         │        95 │
│ Queries with Failure           │         0 │
└────────────────────────────────┴───────────┘

Resource Usage

tpcds — base (merge-base)

Metric Value
Wall time 50.0s
Peak memory 1.8 GiB
Avg memory 1.3 GiB
CPU user 206.0s
CPU sys 5.4s
Peak spill 0 B

tpcds — branch

Metric Value
Wall time 50.0s
Peak memory 1.9 GiB
Avg memory 1.3 GiB
CPU user 207.0s
CPU sys 5.5s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (ce4e7c6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark external_aggr
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark external_aggr.json
--------------------
┏━━━━━━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query        ┃      HEAD ┃ blocked-agg-poc ┃       Change ┃
┡━━━━━━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ Q1(64.0 MB)  │  53.34 ms │        60.84 ms │ 1.14x slower │
│ Q1(32.0 MB)  │  48.83 ms │        55.23 ms │ 1.13x slower │
│ Q1(16.0 MB)  │  51.39 ms │        53.85 ms │    no change │
│ Q2(512.0 MB) │ 271.52 ms │       284.48 ms │    no change │
│ Q2(256.0 MB) │ 269.52 ms │       270.24 ms │    no change │
│ Q2(128.0 MB) │ 239.55 ms │       240.04 ms │    no change │
│ Q2(64.0 MB)  │ 237.48 ms │       235.05 ms │    no change │
│ Q2(32.0 MB)  │ 297.83 ms │       294.83 ms │    no change │
└──────────────┴───────────┴─────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary              ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 1469.47ms │
│ Total Time (blocked-agg-poc)   │ 1494.56ms │
│ Average Time (HEAD)            │  183.68ms │
│ Average Time (blocked-agg-poc) │  186.82ms │
│ Queries Faster                 │         0 │
│ Queries Slower                 │         2 │
│ Queries with No Change         │         6 │
│ Queries with Failure           │         0 │
└────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark external_aggr.json
--------------------
┏━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query        ┃                               HEAD ┃                   blocked-agg-poc ┃       Change ┃
┡━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ Q1(64.0 MB)  │     53.34 / 56.73 ±2.64 / 60.85 ms │    60.84 / 66.51 ±8.01 / 82.29 ms │ 1.17x slower │
│ Q1(32.0 MB)  │     48.83 / 51.17 ±1.40 / 52.84 ms │    55.23 / 55.99 ±0.93 / 57.83 ms │ 1.09x slower │
│ Q1(16.0 MB)  │     51.39 / 54.12 ±1.65 / 56.43 ms │    53.85 / 56.48 ±1.34 / 57.63 ms │    no change │
│ Q2(512.0 MB) │ 271.52 / 288.89 ±10.13 / 300.54 ms │ 284.48 / 290.12 ±4.01 / 296.42 ms │    no change │
│ Q2(256.0 MB) │ 269.52 / 283.19 ±14.66 / 311.00 ms │ 270.24 / 279.08 ±8.14 / 293.65 ms │    no change │
│ Q2(128.0 MB) │  239.55 / 243.25 ±2.88 / 247.81 ms │ 240.04 / 244.08 ±5.76 / 255.49 ms │    no change │
│ Q2(64.0 MB)  │  237.48 / 238.24 ±0.61 / 239.26 ms │ 235.05 / 238.99 ±4.06 / 246.81 ms │    no change │
│ Q2(32.0 MB)  │  297.83 / 302.24 ±3.60 / 308.38 ms │ 294.83 / 300.34 ±3.38 / 303.84 ms │    no change │
└──────────────┴────────────────────────────────────┴───────────────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary              ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 1517.84ms │
│ Total Time (blocked-agg-poc)   │ 1531.58ms │
│ Average Time (HEAD)            │  189.73ms │
│ Average Time (blocked-agg-poc) │  191.45ms │
│ Queries Faster                 │         0 │
│ Queries Slower                 │         2 │
│ Queries with No Change         │         6 │
│ Queries with Failure           │         0 │
└────────────────────────────────┴───────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only for benchmarks that write a results JSON (the dfbench suites, and the suites bench.sh runs through the Criterion SQL harness since #25644), and only under a memory limit: DATAFUSION_RUNTIME_MEMORY_LIMIT, or one the suite sets itself.

Base: dbf62c1 (merge-base) | Changed: blocked-agg-poc

external_aggr

Query Base Changed Change
1(64.0 MB) 36.8 MiB 38.8 MiB +5.4%
1(32.0 MB) 17.8 MiB 19.8 MiB +11.7%
1(16.0 MB) 9.6 MiB 10.0 MiB +4.3%
2(512.0 MB) 136.4 MiB 137.4 MiB +0.7%
2(256.0 MB) 97.3 MiB 98.3 MiB +1.0%
2(128.0 MB) 49.1 MiB 49.1 MiB -0.0%
2(64.0 MB) 32.5 MiB 32.5 MiB +0.0%
2(32.0 MB) 17.5 MiB 18.0 MiB +2.9%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
external_aggr base (dbf62c1 (merge-base)) 136.4 MiB 454.0 MiB 317.6 MiB 3.3×
external_aggr changed (blocked-agg-poc) 137.4 MiB 420.4 MiB 283.0 MiB 3.1×
Resource Usage

external_aggr — base (merge-base)

Metric Value
Wall time 110.0s
Peak memory 454.0 MiB
Avg memory 29.0 MiB
CPU user 26.0s
CPU sys 3.7s
Peak spill 93.1 MiB

external_aggr — branch

Metric Value
Wall time 105.0s
Peak memory 420.4 MiB
Avg memory 32.4 MiB
CPU user 26.0s
CPU sys 3.9s
Peak spill 85.5 MiB

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (ce4e7c6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark clickbench_partitioned
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark clickbench_partitioned.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃       HEAD ┃ blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 0  │    1.27 ms │         1.26 ms │     no change │
│ QQuery 1  │   11.96 ms │        11.80 ms │     no change │
│ QQuery 2  │   36.82 ms │        36.84 ms │     no change │
│ QQuery 3  │   31.61 ms │        31.20 ms │     no change │
│ QQuery 4  │  238.38 ms │       213.21 ms │ +1.12x faster │
│ QQuery 5  │  277.79 ms │       272.84 ms │     no change │
│ QQuery 6  │    1.29 ms │         1.28 ms │     no change │
│ QQuery 7  │   13.13 ms │        12.87 ms │     no change │
│ QQuery 8  │  339.37 ms │       331.43 ms │     no change │
│ QQuery 9  │  475.37 ms │       471.69 ms │     no change │
│ QQuery 10 │   66.07 ms │        63.90 ms │     no change │
│ QQuery 11 │   75.62 ms │        74.61 ms │     no change │
│ QQuery 12 │  265.14 ms │       267.92 ms │     no change │
│ QQuery 13 │  353.90 ms │       368.52 ms │     no change │
│ QQuery 14 │  282.45 ms │       282.38 ms │     no change │
│ QQuery 15 │  287.78 ms │       272.68 ms │ +1.06x faster │
│ QQuery 16 │  633.55 ms │       638.35 ms │     no change │
│ QQuery 17 │  632.66 ms │       645.93 ms │     no change │
│ QQuery 18 │ 1289.31 ms │      1283.42 ms │     no change │
│ QQuery 19 │   27.53 ms │        27.31 ms │     no change │
│ QQuery 20 │  514.24 ms │       520.49 ms │     no change │
│ QQuery 21 │  504.67 ms │       513.69 ms │     no change │
│ QQuery 22 │ 1008.48 ms │      1008.06 ms │     no change │
│ QQuery 23 │ 3033.74 ms │      3075.59 ms │     no change │
│ QQuery 24 │   40.48 ms │        40.40 ms │     no change │
│ QQuery 25 │  106.69 ms │       106.86 ms │     no change │
│ QQuery 26 │   41.28 ms │        41.17 ms │     no change │
│ QQuery 27 │  509.23 ms │       512.44 ms │     no change │
│ QQuery 28 │ 2902.99 ms │      2899.61 ms │     no change │
│ QQuery 29 │   41.55 ms │        41.41 ms │     no change │
│ QQuery 30 │  305.42 ms │       302.80 ms │     no change │
│ QQuery 31 │  273.41 ms │       273.52 ms │     no change │
│ QQuery 32 │ 1110.17 ms │      1059.30 ms │     no change │
│ QQuery 33 │ 1589.71 ms │      1545.51 ms │     no change │
│ QQuery 34 │ 1556.20 ms │      1595.56 ms │     no change │
│ QQuery 35 │  305.05 ms │       258.86 ms │ +1.18x faster │
│ QQuery 36 │   65.49 ms │        67.51 ms │     no change │
│ QQuery 37 │   36.73 ms │        36.33 ms │     no change │
│ QQuery 38 │   40.91 ms │        44.04 ms │  1.08x slower │
│ QQuery 39 │  149.06 ms │       159.56 ms │  1.07x slower │
│ QQuery 40 │   14.63 ms │        14.49 ms │     no change │
│ QQuery 41 │   13.70 ms │        13.96 ms │     no change │
│ QQuery 42 │   13.28 ms │        13.60 ms │     no change │
└───────────┴────────────┴─────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 19518.14ms │
│ Total Time (blocked-agg-poc)   │ 19454.18ms │
│ Average Time (HEAD)            │   453.91ms │
│ Average Time (blocked-agg-poc) │   452.42ms │
│ Queries Faster                 │          3 │
│ Queries Slower                 │          2 │
│ Queries with No Change         │         38 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark clickbench_partitioned.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                   HEAD ┃                        blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 0  │           1.27 / 4.15 ±5.62 / 15.39 ms │           1.26 / 4.17 ±5.68 / 15.52 ms │     no change │
│ QQuery 1  │         11.96 / 12.27 ±0.34 / 12.93 ms │         11.80 / 11.98 ±0.17 / 12.27 ms │     no change │
│ QQuery 2  │         36.82 / 37.16 ±0.26 / 37.53 ms │         36.84 / 37.21 ±0.39 / 37.92 ms │     no change │
│ QQuery 3  │         31.61 / 32.58 ±0.60 / 33.23 ms │         31.20 / 31.56 ±0.47 / 32.46 ms │     no change │
│ QQuery 4  │      238.38 / 244.51 ±4.07 / 250.28 ms │      213.21 / 215.22 ±1.71 / 218.35 ms │ +1.14x faster │
│ QQuery 5  │      277.79 / 279.20 ±0.81 / 280.25 ms │      272.84 / 275.54 ±1.83 / 278.04 ms │     no change │
│ QQuery 6  │            1.29 / 1.46 ±0.24 / 1.93 ms │            1.28 / 1.44 ±0.23 / 1.88 ms │     no change │
│ QQuery 7  │         13.13 / 13.35 ±0.14 / 13.53 ms │         12.87 / 13.00 ±0.08 / 13.10 ms │     no change │
│ QQuery 8  │      339.37 / 341.66 ±2.15 / 345.58 ms │      331.43 / 334.33 ±2.08 / 337.94 ms │     no change │
│ QQuery 9  │      475.37 / 482.96 ±4.25 / 487.63 ms │      471.69 / 480.77 ±6.48 / 489.35 ms │     no change │
│ QQuery 10 │         66.07 / 70.91 ±6.53 / 83.24 ms │         63.90 / 64.69 ±0.47 / 65.33 ms │ +1.10x faster │
│ QQuery 11 │         75.62 / 76.43 ±0.58 / 76.99 ms │         74.61 / 75.22 ±0.86 / 76.88 ms │     no change │
│ QQuery 12 │      265.14 / 273.19 ±5.56 / 281.52 ms │      267.92 / 274.96 ±9.71 / 294.21 ms │     no change │
│ QQuery 13 │     353.90 / 365.86 ±11.67 / 382.42 ms │      368.52 / 378.49 ±6.27 / 387.08 ms │     no change │
│ QQuery 14 │      282.45 / 288.40 ±6.71 / 299.99 ms │     282.38 / 300.77 ±17.61 / 329.50 ms │     no change │
│ QQuery 15 │      287.78 / 295.28 ±8.57 / 311.02 ms │      272.68 / 277.08 ±3.89 / 283.09 ms │ +1.07x faster │
│ QQuery 16 │      633.55 / 638.35 ±3.97 / 643.50 ms │      638.35 / 642.87 ±4.04 / 650.24 ms │     no change │
│ QQuery 17 │      632.66 / 639.34 ±5.53 / 645.77 ms │      645.93 / 652.82 ±5.47 / 660.77 ms │     no change │
│ QQuery 18 │  1289.31 / 1305.29 ±10.97 / 1323.41 ms │  1283.42 / 1314.62 ±22.83 / 1348.12 ms │     no change │
│ QQuery 19 │         27.53 / 28.03 ±0.37 / 28.63 ms │         27.31 / 27.62 ±0.33 / 28.21 ms │     no change │
│ QQuery 20 │      514.24 / 526.11 ±9.32 / 537.47 ms │     520.49 / 533.94 ±13.14 / 556.87 ms │     no change │
│ QQuery 21 │      504.67 / 517.62 ±8.91 / 528.61 ms │     513.69 / 523.23 ±14.11 / 551.02 ms │     no change │
│ QQuery 22 │   1008.48 / 1015.80 ±4.46 / 1022.36 ms │   1008.06 / 1020.02 ±9.87 / 1034.82 ms │     no change │
│ QQuery 23 │ 3033.74 / 3259.42 ±128.45 / 3400.97 ms │ 3075.59 / 3349.29 ±201.11 / 3612.85 ms │     no change │
│ QQuery 24 │        40.48 / 57.83 ±21.42 / 95.30 ms │         40.40 / 45.09 ±7.44 / 59.82 ms │ +1.28x faster │
│ QQuery 25 │    106.69 / 195.58 ±170.85 / 537.17 ms │    106.86 / 185.03 ±137.67 / 458.84 ms │ +1.06x faster │
│ QQuery 26 │         41.28 / 42.79 ±1.63 / 45.63 ms │         41.17 / 42.02 ±0.75 / 42.92 ms │     no change │
│ QQuery 27 │     509.23 / 520.65 ±10.81 / 539.84 ms │      512.44 / 524.72 ±7.48 / 533.00 ms │     no change │
│ QQuery 28 │  2902.99 / 2932.65 ±30.19 / 2970.02 ms │  2899.61 / 2923.07 ±21.99 / 2963.57 ms │     no change │
│ QQuery 29 │       41.55 / 68.38 ±44.54 / 156.32 ms │         41.41 / 41.71 ±0.21 / 41.96 ms │ +1.64x faster │
│ QQuery 30 │     305.42 / 314.50 ±11.22 / 334.36 ms │     302.80 / 321.99 ±14.38 / 344.77 ms │     no change │
│ QQuery 31 │     273.41 / 300.22 ±34.06 / 367.08 ms │     273.52 / 286.18 ±11.08 / 305.01 ms │     no change │
│ QQuery 32 │  1110.17 / 1146.51 ±34.86 / 1209.16 ms │ 1059.30 / 1196.59 ±153.90 / 1491.44 ms │     no change │
│ QQuery 33 │  1589.71 / 1632.69 ±37.71 / 1688.58 ms │  1545.51 / 1622.21 ±46.52 / 1674.81 ms │     no change │
│ QQuery 34 │  1556.20 / 1619.31 ±55.91 / 1719.34 ms │  1595.56 / 1651.77 ±61.05 / 1767.90 ms │     no change │
│ QQuery 35 │     305.05 / 355.21 ±70.15 / 489.54 ms │     258.86 / 343.39 ±98.71 / 489.55 ms │     no change │
│ QQuery 36 │         65.49 / 69.11 ±2.85 / 73.26 ms │         67.51 / 72.47 ±3.55 / 78.00 ms │     no change │
│ QQuery 37 │         36.73 / 44.70 ±6.96 / 55.68 ms │         36.33 / 39.89 ±4.40 / 48.57 ms │ +1.12x faster │
│ QQuery 38 │         40.91 / 47.77 ±6.41 / 59.95 ms │         44.04 / 48.66 ±4.56 / 56.32 ms │     no change │
│ QQuery 39 │      149.06 / 151.13 ±2.40 / 154.44 ms │      159.56 / 163.86 ±5.92 / 175.45 ms │  1.08x slower │
│ QQuery 40 │         14.63 / 17.09 ±4.26 / 25.58 ms │         14.49 / 14.98 ±0.52 / 15.97 ms │ +1.14x faster │
│ QQuery 41 │         13.70 / 14.19 ±0.42 / 14.76 ms │         13.96 / 14.13 ±0.11 / 14.28 ms │     no change │
│ QQuery 42 │         13.28 / 13.48 ±0.12 / 13.61 ms │         13.60 / 20.45 ±9.06 / 36.60 ms │  1.52x slower │
└───────────┴────────────────────────────────────────┴────────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 20293.17ms │
│ Total Time (blocked-agg-poc)   │ 20399.03ms │
│ Average Time (HEAD)            │   471.93ms │
│ Average Time (blocked-agg-poc) │   474.40ms │
│ Queries Faster                 │          8 │
│ Queries Slower                 │          2 │
│ Queries with No Change         │         33 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Resource Usage

clickbench_partitioned — base (merge-base)

Metric Value
Wall time 105.0s
Peak memory 18.4 GiB
Avg memory 6.2 GiB
CPU user 1003.3s
CPU sys 92.8s
Peak spill 0 B

clickbench_partitioned — branch

Metric Value
Wall time 105.0s
Peak memory 17.4 GiB
Avg memory 6.0 GiB
CPU user 1000.4s
CPU sys 96.7s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (ce4e7c6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark clickbench_partitioned
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "16G"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark clickbench_partitioned.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃       HEAD ┃ blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 0  │    1.24 ms │         1.24 ms │     no change │
│ QQuery 1  │   11.79 ms │        12.21 ms │     no change │
│ QQuery 2  │   37.01 ms │        37.36 ms │     no change │
│ QQuery 3  │   31.98 ms │        31.43 ms │     no change │
│ QQuery 4  │  235.15 ms │       214.24 ms │ +1.10x faster │
│ QQuery 5  │  273.31 ms │       273.41 ms │     no change │
│ QQuery 6  │    1.29 ms │         1.32 ms │     no change │
│ QQuery 7  │   13.00 ms │        12.96 ms │     no change │
│ QQuery 8  │  332.04 ms │       329.71 ms │     no change │
│ QQuery 9  │  480.37 ms │       467.22 ms │     no change │
│ QQuery 10 │   67.35 ms │        63.87 ms │ +1.05x faster │
│ QQuery 11 │   76.49 ms │        74.82 ms │     no change │
│ QQuery 12 │  265.90 ms │       295.25 ms │  1.11x slower │
│ QQuery 13 │  370.53 ms │       369.00 ms │     no change │
│ QQuery 14 │  279.17 ms │       281.92 ms │     no change │
│ QQuery 15 │  283.06 ms │       277.03 ms │     no change │
│ QQuery 16 │  627.64 ms │       640.31 ms │     no change │
│ QQuery 17 │  628.39 ms │       642.90 ms │     no change │
│ QQuery 18 │ 1269.46 ms │      1343.58 ms │  1.06x slower │
│ QQuery 19 │   27.36 ms │        28.14 ms │     no change │
│ QQuery 20 │  518.99 ms │       527.09 ms │     no change │
│ QQuery 21 │  513.18 ms │       525.30 ms │     no change │
│ QQuery 22 │  996.03 ms │      1011.04 ms │     no change │
│ QQuery 23 │ 3197.46 ms │      3096.23 ms │     no change │
│ QQuery 24 │   42.60 ms │        40.93 ms │     no change │
│ QQuery 25 │  109.38 ms │       106.21 ms │     no change │
│ QQuery 26 │   42.16 ms │        42.19 ms │     no change │
│ QQuery 27 │  522.12 ms │       526.96 ms │     no change │
│ QQuery 28 │ 2931.88 ms │      2947.53 ms │     no change │
│ QQuery 29 │   43.31 ms │        43.12 ms │     no change │
│ QQuery 30 │  329.04 ms │       313.31 ms │     no change │
│ QQuery 31 │  294.85 ms │       284.91 ms │     no change │
│ QQuery 32 │ 4028.76 ms │      4861.97 ms │  1.21x slower │
│ QQuery 33 │ 1552.06 ms │      1681.54 ms │  1.08x slower │
│ QQuery 34 │ 1589.01 ms │      1585.25 ms │     no change │
│ QQuery 35 │  308.66 ms │       257.35 ms │ +1.20x faster │
│ QQuery 36 │   66.03 ms │        72.32 ms │  1.10x slower │
│ QQuery 37 │   37.97 ms │        38.11 ms │     no change │
│ QQuery 38 │   42.67 ms │        42.19 ms │     no change │
│ QQuery 39 │  142.51 ms │       150.06 ms │  1.05x slower │
│ QQuery 40 │   15.19 ms │        15.81 ms │     no change │
│ QQuery 41 │   14.41 ms │        14.89 ms │     no change │
│ QQuery 42 │   13.30 ms │        14.09 ms │  1.06x slower │
└───────────┴────────────┴─────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 22664.11ms │
│ Total Time (blocked-agg-poc)   │ 23596.30ms │
│ Average Time (HEAD)            │   527.07ms │
│ Average Time (blocked-agg-poc) │   548.75ms │
│ Queries Faster                 │          3 │
│ Queries Slower                 │          7 │
│ Queries with No Change         │         33 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark clickbench_partitioned.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                   HEAD ┃                        blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 0  │           1.24 / 4.08 ±5.55 / 15.17 ms │           1.24 / 4.08 ±5.55 / 15.18 ms │     no change │
│ QQuery 1  │         11.79 / 11.91 ±0.10 / 12.06 ms │         12.21 / 12.32 ±0.14 / 12.57 ms │     no change │
│ QQuery 2  │         37.01 / 37.26 ±0.27 / 37.62 ms │         37.36 / 37.78 ±0.41 / 38.52 ms │     no change │
│ QQuery 3  │         31.98 / 32.36 ±0.38 / 33.09 ms │         31.43 / 32.02 ±0.43 / 32.59 ms │     no change │
│ QQuery 4  │      235.15 / 241.44 ±3.89 / 247.33 ms │      214.24 / 215.80 ±0.99 / 216.97 ms │ +1.12x faster │
│ QQuery 5  │      273.31 / 277.70 ±3.86 / 283.70 ms │      273.41 / 276.43 ±2.27 / 279.14 ms │     no change │
│ QQuery 6  │            1.29 / 1.45 ±0.24 / 1.92 ms │            1.32 / 1.48 ±0.24 / 1.94 ms │     no change │
│ QQuery 7  │         13.00 / 13.20 ±0.17 / 13.40 ms │         12.96 / 13.16 ±0.15 / 13.34 ms │     no change │
│ QQuery 8  │      332.04 / 338.55 ±6.13 / 349.58 ms │      329.71 / 332.73 ±2.29 / 336.66 ms │     no change │
│ QQuery 9  │     480.37 / 504.37 ±14.84 / 523.06 ms │      467.22 / 477.25 ±9.12 / 489.42 ms │ +1.06x faster │
│ QQuery 10 │         67.35 / 68.77 ±1.17 / 70.81 ms │         63.87 / 65.28 ±1.05 / 66.69 ms │ +1.05x faster │
│ QQuery 11 │         76.49 / 77.55 ±0.83 / 78.52 ms │         74.82 / 76.31 ±1.06 / 78.10 ms │     no change │
│ QQuery 12 │      265.90 / 275.70 ±9.73 / 293.48 ms │      295.25 / 300.09 ±3.11 / 304.59 ms │  1.09x slower │
│ QQuery 13 │     370.53 / 394.48 ±17.35 / 415.20 ms │     369.00 / 393.76 ±14.99 / 407.42 ms │     no change │
│ QQuery 14 │      279.17 / 287.51 ±8.60 / 303.38 ms │      281.92 / 289.58 ±4.72 / 296.09 ms │     no change │
│ QQuery 15 │      283.06 / 290.88 ±5.87 / 299.42 ms │     277.03 / 305.19 ±31.82 / 366.47 ms │     no change │
│ QQuery 16 │      627.64 / 634.26 ±5.31 / 641.30 ms │     640.31 / 657.18 ±11.47 / 669.36 ms │     no change │
│ QQuery 17 │      628.39 / 640.29 ±6.63 / 646.52 ms │      642.90 / 655.12 ±8.86 / 665.64 ms │     no change │
│ QQuery 18 │  1269.46 / 1298.33 ±31.57 / 1357.39 ms │  1343.58 / 1376.52 ±28.48 / 1421.87 ms │  1.06x slower │
│ QQuery 19 │         27.36 / 27.78 ±0.43 / 28.56 ms │        28.14 / 43.32 ±19.61 / 76.58 ms │  1.56x slower │
│ QQuery 20 │     518.99 / 527.67 ±11.67 / 550.27 ms │      527.09 / 536.96 ±9.16 / 552.05 ms │     no change │
│ QQuery 21 │      513.18 / 517.43 ±3.67 / 522.91 ms │      525.30 / 534.73 ±7.38 / 541.99 ms │     no change │
│ QQuery 22 │    996.03 / 1012.29 ±8.96 / 1019.64 ms │  1011.04 / 1019.64 ±10.39 / 1040.03 ms │     no change │
│ QQuery 23 │ 3197.46 / 3326.56 ±124.68 / 3553.22 ms │ 3096.23 / 3219.99 ±125.56 / 3446.26 ms │     no change │
│ QQuery 24 │     42.60 / 122.08 ±157.60 / 437.27 ms │     40.93 / 198.44 ±215.03 / 589.55 ms │  1.63x slower │
│ QQuery 25 │     109.38 / 128.90 ±26.47 / 177.99 ms │      106.21 / 108.95 ±3.81 / 116.37 ms │ +1.18x faster │
│ QQuery 26 │         42.16 / 43.13 ±1.11 / 44.99 ms │         42.19 / 43.43 ±1.17 / 44.91 ms │     no change │
│ QQuery 27 │      522.12 / 531.95 ±9.87 / 549.99 ms │     526.96 / 542.62 ±16.03 / 573.07 ms │     no change │
│ QQuery 28 │  2931.88 / 2991.02 ±42.25 / 3060.62 ms │  2947.53 / 2995.78 ±36.66 / 3051.65 ms │     no change │
│ QQuery 29 │         43.31 / 51.43 ±9.90 / 65.29 ms │        43.12 / 57.73 ±13.18 / 77.82 ms │  1.12x slower │
│ QQuery 30 │      329.04 / 337.26 ±7.19 / 350.65 ms │     313.31 / 327.97 ±13.87 / 352.80 ms │     no change │
│ QQuery 31 │      294.85 / 301.70 ±6.67 / 311.44 ms │     284.91 / 312.29 ±25.92 / 359.89 ms │     no change │
│ QQuery 32 │  4028.76 / 4152.22 ±79.46 / 4277.79 ms │ 4861.97 / 5111.50 ±160.73 / 5312.64 ms │  1.23x slower │
│ QQuery 33 │  1552.06 / 1589.98 ±51.73 / 1690.92 ms │  1681.54 / 1723.59 ±51.37 / 1820.97 ms │  1.08x slower │
│ QQuery 34 │ 1589.01 / 1681.03 ±100.89 / 1854.88 ms │  1585.25 / 1659.02 ±50.46 / 1719.97 ms │     no change │
│ QQuery 35 │     308.66 / 351.83 ±65.48 / 478.94 ms │     257.35 / 309.68 ±84.39 / 477.86 ms │ +1.14x faster │
│ QQuery 36 │         66.03 / 72.78 ±4.92 / 81.16 ms │         72.32 / 75.12 ±2.38 / 77.81 ms │     no change │
│ QQuery 37 │         37.97 / 43.45 ±5.50 / 54.02 ms │         38.11 / 39.05 ±0.64 / 39.93 ms │ +1.11x faster │
│ QQuery 38 │         42.67 / 48.99 ±5.38 / 55.77 ms │         42.19 / 48.62 ±5.39 / 56.84 ms │     no change │
│ QQuery 39 │     142.51 / 155.24 ±10.60 / 174.13 ms │     150.06 / 170.51 ±11.58 / 184.11 ms │  1.10x slower │
│ QQuery 40 │         15.19 / 16.59 ±2.17 / 20.89 ms │         15.81 / 19.46 ±5.55 / 30.51 ms │  1.17x slower │
│ QQuery 41 │         14.41 / 16.76 ±3.81 / 24.35 ms │         14.89 / 16.78 ±2.94 / 22.62 ms │     no change │
│ QQuery 42 │         13.30 / 13.56 ±0.22 / 13.96 ms │         14.09 / 14.38 ±0.18 / 14.58 ms │  1.06x slower │
└───────────┴────────────────────────────────────────┴────────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 23491.73ms │
│ Total Time (blocked-agg-poc)   │ 24651.61ms │
│ Average Time (HEAD)            │   546.32ms │
│ Average Time (blocked-agg-poc) │   573.29ms │
│ Queries Faster                 │          6 │
│ Queries Slower                 │         10 │
│ Queries with No Change         │         27 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only for benchmarks that write a results JSON (the dfbench suites, and the suites bench.sh runs through the Criterion SQL harness since #25644), and only under a memory limit: DATAFUSION_RUNTIME_MEMORY_LIMIT, or one the suite sets itself.

Base: dbf62c1 (merge-base) | Changed: blocked-agg-poc

clickbench_partitioned

Query Base Changed Change
Query 0 0 B 0 B 0.0%
Query 1 104 B 104 B +0.0%
Query 2 936 B 936 B +0.0%
Query 3 312 B 312 B +0.0%
Query 4 770.9 MiB 539.1 MiB -30.1%
Query 5 1.2 GiB 1.2 GiB -0.0%
Query 6 0 B 0 B 0.0%
Query 7 60.0 MiB 40.2 MiB -32.9%
Query 8 944.3 MiB 951.7 MiB +0.8%
Query 9 593.9 MiB 594.3 MiB +0.1%
Query 10 109.5 MiB 116.9 MiB +6.8%
Query 11 121.0 MiB 113.7 MiB -6.0%
Query 12 1.4 GiB 1.2 GiB -8.1%
Query 13 1.6 GiB 1.5 GiB -3.5%
Query 14 1.3 GiB 1.2 GiB -6.3%
Query 15 1.1 GiB 693.3 MiB -41.1%
Query 16 3.4 GiB 3.1 GiB -9.3%
Query 17 3.4 GiB 3.1 GiB -9.0%
Query 18 7.8 GiB 7.5 GiB -3.5%
Query 19 0 B 0 B 0.0%
Query 20 104 B 104 B +0.0%
Query 21 3.3 MiB 3.3 MiB +0.9%
Query 22 3.2 MiB 3.7 MiB +13.2%
Query 23 3.6 GiB 3.6 GiB -0.4%
Query 24 59.4 MiB 59.0 MiB -0.8%
Query 25 167.8 MiB 165.1 MiB -1.6%
Query 26 58.2 MiB 60.7 MiB +4.3%
Query 27 2.2 MiB 2.4 MiB +10.0%
Query 28 1.7 GiB 1.7 GiB -1.7%
Query 29 624 B 624 B +0.0%
Query 30 706.3 MiB 605.6 MiB -14.3%
Query 31 1.5 GiB 1.4 GiB -8.6%
Query 32 3.6 GiB 8.7 GiB +142.3%
Query 33 7.1 GiB 6.9 GiB -3.6%
Query 34 7.0 GiB 6.6 GiB -5.4%
Query 35 604.4 MiB 505.2 MiB -16.4%
Query 36 122.1 MiB 104.5 MiB -14.4%
Query 37 6.4 MiB 6.3 MiB -2.8%
Query 38 5.2 MiB 5.0 MiB -2.8%
Query 39 297.8 MiB 289.4 MiB -2.8%
Query 40 1.7 MiB 2.2 MiB +28.2%
Query 41 3.1 MiB 3.1 MiB +0.3%
Query 42 1.8 MiB 1.6 MiB -6.2%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
clickbench_partitioned base (dbf62c1 (merge-base)) 7.8 GiB 18.0 GiB 10.2 GiB 2.3×
clickbench_partitioned changed (blocked-agg-poc) 8.7 GiB 17.3 GiB 8.7 GiB 2.0×
Resource Usage

clickbench_partitioned — base (merge-base)

Metric Value
Wall time 120.0s
Peak memory 18.0 GiB
Avg memory 6.0 GiB
CPU user 1153.8s
CPU sys 118.6s
Peak spill 4.2 GiB

clickbench_partitioned — branch

Metric Value
Wall time 125.0s
Peak memory 17.3 GiB
Avg memory 6.5 GiB
CPU user 1188.6s
CPU sys 144.6s
Peak spill 4.1 GiB

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (ce4e7c6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark h2o_medium
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark h2o.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃       HEAD ┃ blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │ 1102.29 ms │      1019.02 ms │ +1.08x faster │
│ QQuery 2  │ 2382.16 ms │      2273.07 ms │     no change │
│ QQuery 3  │ 2057.92 ms │      1999.92 ms │     no change │
│ QQuery 4  │ 1402.55 ms │      1320.82 ms │ +1.06x faster │
│ QQuery 5  │ 1955.39 ms │      1739.38 ms │ +1.12x faster │
│ QQuery 6  │ 1648.58 ms │      1565.41 ms │ +1.05x faster │
│ QQuery 7  │ 1948.56 ms │      1838.75 ms │ +1.06x faster │
│ QQuery 8  │ 3555.70 ms │      3464.23 ms │     no change │
│ QQuery 9  │ 2905.58 ms │      2742.51 ms │ +1.06x faster │
│ QQuery 10 │ 3064.26 ms │      2942.79 ms │     no change │
└───────────┴────────────┴─────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 22022.99ms │
│ Total Time (blocked-agg-poc)   │ 20905.92ms │
│ Average Time (HEAD)            │  2202.30ms │
│ Average Time (blocked-agg-poc) │  2090.59ms │
│ Queries Faster                 │          6 │
│ Queries Slower                 │          0 │
│ Queries with No Change         │          4 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark h2o.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                  HEAD ┃                       blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │ 1102.29 / 1116.06 ±11.64 / 1130.77 ms │  1019.02 / 1020.74 ±1.25 / 1021.96 ms │ +1.09x faster │
│ QQuery 2  │ 2382.16 / 2423.91 ±29.91 / 2450.66 ms │ 2273.07 / 2319.07 ±32.60 / 2344.69 ms │     no change │
│ QQuery 3  │ 2057.92 / 2083.58 ±22.52 / 2112.74 ms │ 1999.92 / 2025.06 ±18.10 / 2041.82 ms │     no change │
│ QQuery 4  │  1402.55 / 1407.34 ±3.74 / 1411.67 ms │  1320.82 / 1322.54 ±2.22 / 1325.68 ms │ +1.06x faster │
│ QQuery 5  │ 1955.39 / 1974.08 ±14.32 / 1990.19 ms │ 1739.38 / 1756.64 ±14.59 / 1775.06 ms │ +1.12x faster │
│ QQuery 6  │  1648.58 / 1652.82 ±3.14 / 1656.08 ms │  1565.41 / 1571.01 ±7.92 / 1582.21 ms │     no change │
│ QQuery 7  │  1948.56 / 1953.17 ±3.27 / 1955.82 ms │ 1838.75 / 1874.38 ±44.99 / 1937.85 ms │     no change │
│ QQuery 8  │ 3555.70 / 3573.43 ±23.66 / 3606.87 ms │ 3464.23 / 3530.02 ±52.75 / 3593.38 ms │     no change │
│ QQuery 9  │  2905.58 / 2909.22 ±2.59 / 2911.39 ms │ 2742.51 / 2788.37 ±36.50 / 2831.82 ms │     no change │
│ QQuery 10 │ 3064.26 / 3094.24 ±34.67 / 3142.82 ms │ 2942.79 / 2976.51 ±23.84 / 2993.47 ms │     no change │
└───────────┴───────────────────────────────────────┴───────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 22187.83ms │
│ Total Time (blocked-agg-poc)   │ 21184.35ms │
│ Average Time (HEAD)            │  2218.78ms │
│ Average Time (blocked-agg-poc) │  2118.43ms │
│ Queries Faster                 │          3 │
│ Queries Slower                 │          0 │
│ Queries with No Change         │          7 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Resource Usage

h2o_medium — base (merge-base)

Metric Value
Wall time 70.0s
Peak memory 9.4 GiB
Avg memory 2.8 GiB
CPU user 701.5s
CPU sys 58.1s
Peak spill 0 B

h2o_medium — branch

Metric Value
Wall time 65.0s
Peak memory 12.9 GiB
Avg memory 3.1 GiB
CPU user 666.4s
CPU sys 56.8s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (ce4e7c6) to dbf62c1 (merge-base) diff

Run configuration
run benchmark h2o_medium
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "16G"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark h2o.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃        HEAD ┃ blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │  1096.90 ms │      1014.36 ms │ +1.08x faster │
│ QQuery 2  │  2372.88 ms │      2261.76 ms │     no change │
│ QQuery 3  │  2065.27 ms │      1933.76 ms │ +1.07x faster │
│ QQuery 4  │  1402.89 ms │      1314.86 ms │ +1.07x faster │
│ QQuery 5  │  1913.20 ms │      1700.84 ms │ +1.12x faster │
│ QQuery 6  │  1604.72 ms │      1530.04 ms │     no change │
│ QQuery 7  │  1875.12 ms │      1811.52 ms │     no change │
│ QQuery 8  │  3534.64 ms │      3348.05 ms │ +1.06x faster │
│ QQuery 9  │  7458.93 ms │      7341.11 ms │     no change │
│ QQuery 10 │ 12667.67 ms │      9856.57 ms │ +1.29x faster │
└───────────┴─────────────┴─────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 35992.22ms │
│ Total Time (blocked-agg-poc)   │ 32112.86ms │
│ Average Time (HEAD)            │  3599.22ms │
│ Average Time (blocked-agg-poc) │  3211.29ms │
│ Queries Faster                 │          6 │
│ Queries Slower                 │          0 │
│ Queries with No Change         │          4 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark h2o.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                      HEAD ┃                          blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │     1096.90 / 1107.82 ±13.77 / 1127.24 ms │     1014.36 / 1015.66 ±1.08 / 1017.01 ms │ +1.09x faster │
│ QQuery 2  │     2372.88 / 2392.88 ±20.73 / 2421.45 ms │    2261.76 / 2297.91 ±28.63 / 2331.76 ms │     no change │
│ QQuery 3  │      2065.27 / 2066.30 ±0.76 / 2067.07 ms │    1933.76 / 1943.24 ±12.70 / 1961.19 ms │ +1.06x faster │
│ QQuery 4  │      1402.89 / 1405.69 ±2.90 / 1409.68 ms │     1314.86 / 1316.09 ±0.96 / 1317.21 ms │ +1.07x faster │
│ QQuery 5  │     1913.20 / 1931.61 ±13.84 / 1946.59 ms │    1700.84 / 1721.31 ±18.25 / 1745.16 ms │ +1.12x faster │
│ QQuery 6  │     1604.72 / 1621.15 ±12.87 / 1636.16 ms │     1530.04 / 1538.59 ±6.33 / 1545.16 ms │ +1.05x faster │
│ QQuery 7  │     1875.12 / 1888.33 ±10.66 / 1901.22 ms │    1811.52 / 1821.06 ±11.53 / 1837.28 ms │     no change │
│ QQuery 8  │     3534.64 / 3602.51 ±55.44 / 3670.43 ms │    3348.05 / 3410.01 ±70.52 / 3508.66 ms │ +1.06x faster │
│ QQuery 9  │     7458.93 / 7474.68 ±17.04 / 7498.36 ms │    7341.11 / 7415.76 ±59.99 / 7488.00 ms │     no change │
│ QQuery 10 │ 12667.67 / 12943.76 ±230.68 / 13232.31 ms │ 9856.57 / 10256.71 ±356.82 / 10723.04 ms │ +1.26x faster │
└───────────┴───────────────────────────────────────────┴──────────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 36434.73ms │
│ Total Time (blocked-agg-poc)   │ 32736.34ms │
│ Average Time (HEAD)            │  3643.47ms │
│ Average Time (blocked-agg-poc) │  3273.63ms │
│ Queries Faster                 │          7 │
│ Queries Slower                 │          0 │
│ Queries with No Change         │          3 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only for benchmarks that write a results JSON (the dfbench suites, and the suites bench.sh runs through the Criterion SQL harness since #25644), and only under a memory limit: DATAFUSION_RUNTIME_MEMORY_LIMIT, or one the suite sets itself.

Base: dbf62c1 (merge-base) | Changed: blocked-agg-poc

h2o_medium — h2o

Query Base Changed Change
Query 1 1.7 MiB 1.7 MiB +0.0%
Query 2 4.2 GiB 4.2 GiB -0.0%
Query 3 138.1 MiB 133.1 MiB -3.6%
Query 4 4.4 MiB 3.9 MiB -10.0%
Query 5 92.6 MiB 98.3 MiB +6.2%
Query 6 1.1 GiB 1.0 GiB -2.9%
Query 7 117.6 MiB 117.7 MiB +0.2%
Query 8 3.9 GiB 3.9 GiB +0.0%
Query 9 4.7 GiB 4.7 GiB +0.0%
Query 10 4.8 GiB 9.3 GiB +94.8%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
h2o_medium base (dbf62c1 (merge-base)) 4.8 GiB 10.5 GiB 5.7 GiB 2.2×
h2o_medium changed (blocked-agg-poc) 9.3 GiB 15.4 GiB 6.1 GiB 1.7×
Resource Usage

h2o_medium — base (merge-base)

Metric Value
Wall time 115.0s
Peak memory 10.5 GiB
Avg memory 3.5 GiB
CPU user 1166.6s
CPU sys 93.8s
Peak spill 7.3 GiB

h2o_medium — branch

Metric Value
Wall time 100.0s
Peak memory 15.4 GiB
Avg memory 4.0 GiB
CPU user 998.9s
CPU sys 111.1s
Peak spill 7.2 GiB

File an issue against this benchmark runner

@rluvaton

Copy link
Copy Markdown
Member Author

run benchmarks

@rluvaton

Copy link
Copy Markdown
Member Author

run benchmarks h2o_medium external_aggr

@rluvaton

Copy link
Copy Markdown
Member Author

run benchmarks tpch10 h2o_medium clickbench_partitioned

env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "16G"

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-standard-32 (12 vCPU / 65 GiB) | Linux bench-c5914812338-2971-ndgd8 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  32
On-line CPU(s) list:                     0-31
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     32
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               2 MiB (32 instances)
L1i cache:                               2 MiB (32 instances)
L2 cache:                                64 MiB (32 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-31
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark tpch

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-standard-32 (12 vCPU / 65 GiB) | Linux bench-c5914812843-2967-mnvbh 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  32
On-line CPU(s) list:                     0-31
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     32
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               2 MiB (32 instances)
L1i cache:                               2 MiB (32 instances)
L2 cache:                                64 MiB (32 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-31
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark h2o_medium

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-standard-32 (12 vCPU / 65 GiB) | Linux bench-c5914812843-2968-fj9m6 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  32
On-line CPU(s) list:                     0-31
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     32
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               2 MiB (32 instances)
L1i cache:                               2 MiB (32 instances)
L2 cache:                                64 MiB (32 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-31
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark external_aggr

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-standard-32 (12 vCPU / 65 GiB) | Linux bench-c5914812338-2970-46swz 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  32
On-line CPU(s) list:                     0-31
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     32
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               2 MiB (32 instances)
L1i cache:                               2 MiB (32 instances)
L2 cache:                                64 MiB (32 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-31
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark tpcds

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-standard-32 (12 vCPU / 65 GiB) | Linux bench-c5914812338-2969-chffw 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  32
On-line CPU(s) list:                     0-31
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     32
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               2 MiB (32 instances)
L1i cache:                               2 MiB (32 instances)
L2 cache:                                64 MiB (32 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-31
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark clickbench_partitioned

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-standard-32 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark tpch
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  32
On-line CPU(s) list:                     0-31
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     32
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               2 MiB (32 instances)
L1i cache:                               2 MiB (32 instances)
L2 cache:                                64 MiB (32 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-31
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark tpch_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query     ┃     HEAD ┃ blocked-agg-poc ┃       Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ QQuery 1  │ 40.50 ms │        41.81 ms │    no change │
│ QQuery 2  │ 18.71 ms │        19.19 ms │    no change │
│ QQuery 3  │ 28.27 ms │        29.28 ms │    no change │
│ QQuery 4  │ 17.39 ms │        17.76 ms │    no change │
│ QQuery 5  │ 35.91 ms │        37.41 ms │    no change │
│ QQuery 6  │ 16.28 ms │        16.74 ms │    no change │
│ QQuery 7  │ 41.31 ms │        44.86 ms │ 1.09x slower │
│ QQuery 8  │ 40.23 ms │        42.18 ms │    no change │
│ QQuery 9  │ 48.65 ms │        51.49 ms │ 1.06x slower │
│ QQuery 10 │ 40.36 ms │        42.44 ms │ 1.05x slower │
│ QQuery 11 │ 13.48 ms │        14.38 ms │ 1.07x slower │
│ QQuery 12 │ 20.59 ms │        21.57 ms │    no change │
│ QQuery 13 │ 40.30 ms │        42.29 ms │    no change │
│ QQuery 14 │ 24.50 ms │        25.39 ms │    no change │
│ QQuery 15 │ 30.37 ms │        32.27 ms │ 1.06x slower │
│ QQuery 16 │ 13.99 ms │        14.72 ms │ 1.05x slower │
│ QQuery 17 │ 74.19 ms │        72.43 ms │    no change │
│ QQuery 18 │ 61.64 ms │        64.57 ms │    no change │
│ QQuery 19 │ 31.46 ms │        32.74 ms │    no change │
│ QQuery 20 │ 31.83 ms │        33.00 ms │    no change │
│ QQuery 21 │ 55.21 ms │        56.98 ms │    no change │
│ QQuery 22 │ 14.28 ms │        14.40 ms │    no change │
└───────────┴──────────┴─────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━┓
┃ Benchmark Summary              ┃          ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━┩
│ Total Time (HEAD)              │ 739.44ms │
│ Total Time (blocked-agg-poc)   │ 767.89ms │
│ Average Time (HEAD)            │  33.61ms │
│ Average Time (blocked-agg-poc) │  34.90ms │
│ Queries Faster                 │        0 │
│ Queries Slower                 │        6 │
│ Queries with No Change         │       16 │
│ Queries with Failure           │        0 │
└────────────────────────────────┴──────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark tpch_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query     ┃                           HEAD ┃                blocked-agg-poc ┃       Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ QQuery 1  │ 40.50 / 41.37 ±1.05 / 43.38 ms │ 41.81 / 42.80 ±1.26 / 45.28 ms │    no change │
│ QQuery 2  │ 18.71 / 18.93 ±0.25 / 19.40 ms │ 19.19 / 19.55 ±0.30 / 20.09 ms │    no change │
│ QQuery 3  │ 28.27 / 28.45 ±0.14 / 28.70 ms │ 29.28 / 29.40 ±0.09 / 29.51 ms │    no change │
│ QQuery 4  │ 17.39 / 17.48 ±0.08 / 17.58 ms │ 17.76 / 17.99 ±0.13 / 18.14 ms │    no change │
│ QQuery 5  │ 35.91 / 36.11 ±0.17 / 36.34 ms │ 37.41 / 37.82 ±0.26 / 38.12 ms │    no change │
│ QQuery 6  │ 16.28 / 16.39 ±0.13 / 16.64 ms │ 16.74 / 16.87 ±0.08 / 16.98 ms │    no change │
│ QQuery 7  │ 41.31 / 42.53 ±1.05 / 44.14 ms │ 44.86 / 45.61 ±0.75 / 46.90 ms │ 1.07x slower │
│ QQuery 8  │ 40.23 / 40.97 ±0.48 / 41.58 ms │ 42.18 / 42.41 ±0.17 / 42.62 ms │    no change │
│ QQuery 9  │ 48.65 / 49.72 ±0.82 / 50.51 ms │ 51.49 / 51.80 ±0.29 / 52.24 ms │    no change │
│ QQuery 10 │ 40.36 / 40.46 ±0.06 / 40.53 ms │ 42.44 / 43.17 ±0.58 / 44.14 ms │ 1.07x slower │
│ QQuery 11 │ 13.48 / 13.66 ±0.17 / 13.96 ms │ 14.38 / 14.52 ±0.09 / 14.67 ms │ 1.06x slower │
│ QQuery 12 │ 20.59 / 21.00 ±0.29 / 21.38 ms │ 21.57 / 21.83 ±0.20 / 22.11 ms │    no change │
│ QQuery 13 │ 40.30 / 41.86 ±1.82 / 44.67 ms │ 42.29 / 43.84 ±1.70 / 47.01 ms │    no change │
│ QQuery 14 │ 24.50 / 24.64 ±0.12 / 24.86 ms │ 25.39 / 25.63 ±0.21 / 25.91 ms │    no change │
│ QQuery 15 │ 30.37 / 30.92 ±0.33 / 31.35 ms │ 32.27 / 32.47 ±0.19 / 32.74 ms │ 1.05x slower │
│ QQuery 16 │ 13.99 / 14.16 ±0.12 / 14.29 ms │ 14.72 / 14.90 ±0.14 / 15.13 ms │ 1.05x slower │
│ QQuery 17 │ 74.19 / 74.74 ±0.45 / 75.40 ms │ 72.43 / 73.05 ±0.57 / 74.12 ms │    no change │
│ QQuery 18 │ 61.64 / 63.61 ±1.69 / 66.33 ms │ 64.57 / 66.71 ±2.35 / 71.07 ms │    no change │
│ QQuery 19 │ 31.46 / 31.89 ±0.37 / 32.58 ms │ 32.74 / 33.58 ±1.15 / 35.87 ms │ 1.05x slower │
│ QQuery 20 │ 31.83 / 32.25 ±0.41 / 32.92 ms │ 33.00 / 33.72 ±0.71 / 35.00 ms │    no change │
│ QQuery 21 │ 55.21 / 56.33 ±1.09 / 58.29 ms │ 56.98 / 57.61 ±0.40 / 58.13 ms │    no change │
│ QQuery 22 │ 14.28 / 14.35 ±0.05 / 14.43 ms │ 14.40 / 14.73 ±0.18 / 14.88 ms │    no change │
└───────────┴────────────────────────────────┴────────────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━┓
┃ Benchmark Summary              ┃          ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━┩
│ Total Time (HEAD)              │ 751.81ms │
│ Total Time (blocked-agg-poc)   │ 780.02ms │
│ Average Time (HEAD)            │  34.17ms │
│ Average Time (blocked-agg-poc) │  35.46ms │
│ Queries Faster                 │        0 │
│ Queries Slower                 │        6 │
│ Queries with No Change         │       16 │
│ Queries with Failure           │        0 │
└────────────────────────────────┴──────────┘

Resource Usage

tpch — base (merge-base)

Metric Value
Wall time 5.0s
Peak memory 1.3 GiB
Avg memory 503.4 MiB
CPU user 21.2s
CPU sys 1.7s
Peak spill 0 B

tpch — branch

Metric Value
Wall time 5.0s
Peak memory 1.3 GiB
Avg memory 502.4 MiB
CPU user 22.2s
CPU sys 1.8s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-standard-32 (12 vCPU / 65 GiB) | Linux bench-c5914813281-2972-9s49r 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  32
On-line CPU(s) list:                     0-31
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     32
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               2 MiB (32 instances)
L1i cache:                               2 MiB (32 instances)
L2 cache:                                64 MiB (32 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-31
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark tpch10
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "16G"

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-standard-32 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark tpcds
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  32
On-line CPU(s) list:                     0-31
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     32
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               2 MiB (32 instances)
L1i cache:                               2 MiB (32 instances)
L2 cache:                                64 MiB (32 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-31
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark tpcds_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Query     ┃       HEAD ┃ blocked-agg-poc ┃    Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ QQuery 1  │    5.71 ms │         5.55 ms │ no change │
│ QQuery 2  │   80.77 ms │        80.32 ms │ no change │
│ QQuery 3  │   28.79 ms │        29.10 ms │ no change │
│ QQuery 4  │  455.88 ms │       458.28 ms │ no change │
│ QQuery 5  │   50.53 ms │        51.09 ms │ no change │
│ QQuery 6  │   35.26 ms │        35.35 ms │ no change │
│ QQuery 7  │   73.63 ms │        73.55 ms │ no change │
│ QQuery 8  │   35.92 ms │        36.29 ms │ no change │
│ QQuery 9  │   49.66 ms │        50.72 ms │ no change │
│ QQuery 10 │   61.47 ms │        60.90 ms │ no change │
│ QQuery 11 │  286.46 ms │       285.77 ms │ no change │
│ QQuery 12 │   28.96 ms │        28.88 ms │ no change │
│ QQuery 13 │  117.72 ms │       116.99 ms │ no change │
│ QQuery 14 │  398.56 ms │       393.09 ms │ no change │
│ QQuery 15 │   54.46 ms │        54.64 ms │ no change │
│ QQuery 16 │    6.50 ms │         6.35 ms │ no change │
│ QQuery 17 │   77.64 ms │        78.98 ms │ no change │
│ QQuery 18 │  104.63 ms │       104.12 ms │ no change │
│ QQuery 19 │   41.30 ms │        41.59 ms │ no change │
│ QQuery 20 │   36.04 ms │        35.50 ms │ no change │
│ QQuery 21 │   16.66 ms │        16.96 ms │ no change │
│ QQuery 22 │   63.62 ms │        64.62 ms │ no change │
│ QQuery 23 │  313.15 ms │       312.19 ms │ no change │
│ QQuery 24 │  195.71 ms │       198.20 ms │ no change │
│ QQuery 25 │  107.64 ms │       107.97 ms │ no change │
│ QQuery 26 │   48.07 ms │        47.77 ms │ no change │
│ QQuery 27 │    6.06 ms │         6.08 ms │ no change │
│ QQuery 28 │   60.89 ms │        61.09 ms │ no change │
│ QQuery 29 │   96.34 ms │        94.54 ms │ no change │
│ QQuery 30 │   32.58 ms │        32.65 ms │ no change │
│ QQuery 31 │  109.15 ms │       109.36 ms │ no change │
│ QQuery 32 │   20.52 ms │        20.51 ms │ no change │
│ QQuery 33 │   37.50 ms │        37.36 ms │ no change │
│ QQuery 34 │   10.24 ms │        10.21 ms │ no change │
│ QQuery 35 │   72.89 ms │        72.91 ms │ no change │
│ QQuery 36 │    5.72 ms │         5.71 ms │ no change │
│ QQuery 37 │    6.84 ms │         6.92 ms │ no change │
│ QQuery 38 │   61.49 ms │        63.22 ms │ no change │
│ QQuery 39 │   73.47 ms │        74.29 ms │ no change │
│ QQuery 40 │   24.34 ms │        23.74 ms │ no change │
│ QQuery 41 │   11.39 ms │        11.23 ms │ no change │
│ QQuery 42 │   23.83 ms │        23.87 ms │ no change │
│ QQuery 43 │    4.83 ms │         4.76 ms │ no change │
│ QQuery 44 │    9.11 ms │         8.82 ms │ no change │
│ QQuery 45 │   37.07 ms │        38.72 ms │ no change │
│ QQuery 46 │   12.08 ms │        11.83 ms │ no change │
│ QQuery 47 │  229.09 ms │       228.86 ms │ no change │
│ QQuery 48 │   94.50 ms │        95.78 ms │ no change │
│ QQuery 49 │   70.36 ms │        70.05 ms │ no change │
│ QQuery 50 │   57.91 ms │        57.70 ms │ no change │
│ QQuery 51 │   94.03 ms │        94.95 ms │ no change │
│ QQuery 52 │   24.01 ms │        24.65 ms │ no change │
│ QQuery 53 │   29.00 ms │        29.31 ms │ no change │
│ QQuery 54 │   54.33 ms │        54.14 ms │ no change │
│ QQuery 55 │   24.03 ms │        23.99 ms │ no change │
│ QQuery 56 │   38.75 ms │        38.88 ms │ no change │
│ QQuery 57 │  169.04 ms │       171.93 ms │ no change │
│ QQuery 58 │  109.36 ms │       109.07 ms │ no change │
│ QQuery 59 │  115.10 ms │       116.43 ms │ no change │
│ QQuery 60 │   38.87 ms │        38.78 ms │ no change │
│ QQuery 61 │   11.84 ms │        11.57 ms │ no change │
│ QQuery 62 │   44.14 ms │        44.15 ms │ no change │
│ QQuery 63 │   29.26 ms │        29.20 ms │ no change │
│ QQuery 64 │  365.87 ms │       362.78 ms │ no change │
│ QQuery 65 │  127.82 ms │       130.87 ms │ no change │
│ QQuery 66 │   76.91 ms │        76.82 ms │ no change │
│ QQuery 67 │  255.11 ms │       254.75 ms │ no change │
│ QQuery 68 │   12.09 ms │        11.84 ms │ no change │
│ QQuery 69 │   55.49 ms │        55.72 ms │ no change │
│ QQuery 70 │  104.37 ms │       103.54 ms │ no change │
│ QQuery 71 │   35.13 ms │        35.26 ms │ no change │
│ QQuery 72 │ 1780.57 ms │      1790.74 ms │ no change │
│ QQuery 73 │   10.08 ms │        10.03 ms │ no change │
│ QQuery 74 │  166.25 ms │       164.08 ms │ no change │
│ QQuery 75 │  140.70 ms │       140.12 ms │ no change │
│ QQuery 76 │   34.18 ms │        34.11 ms │ no change │
│ QQuery 77 │   60.57 ms │        60.09 ms │ no change │
│ QQuery 78 │  163.54 ms │       163.21 ms │ no change │
│ QQuery 79 │   66.17 ms │        66.13 ms │ no change │
│ QQuery 80 │   94.91 ms │        96.51 ms │ no change │
│ QQuery 81 │   26.04 ms │        26.12 ms │ no change │
│ QQuery 82 │   16.50 ms │        16.53 ms │ no change │
│ QQuery 83 │   33.48 ms │        34.01 ms │ no change │
│ QQuery 84 │   29.57 ms │        29.65 ms │ no change │
│ QQuery 85 │  102.91 ms │       103.32 ms │ no change │
│ QQuery 86 │   25.57 ms │        26.28 ms │ no change │
│ QQuery 87 │   62.13 ms │        64.41 ms │ no change │
│ QQuery 88 │   60.66 ms │        60.49 ms │ no change │
│ QQuery 89 │   34.71 ms │        35.52 ms │ no change │
│ QQuery 90 │   16.91 ms │        16.96 ms │ no change │
│ QQuery 91 │   44.31 ms │        44.32 ms │ no change │
│ QQuery 92 │   29.38 ms │        29.43 ms │ no change │
│ QQuery 93 │   48.89 ms │        48.96 ms │ no change │
│ QQuery 94 │   38.25 ms │        38.44 ms │ no change │
│ QQuery 95 │   79.58 ms │        80.48 ms │ no change │
│ QQuery 96 │   23.85 ms │        23.94 ms │ no change │
│ QQuery 97 │   51.36 ms │        50.99 ms │ no change │
│ QQuery 98 │   42.58 ms │        42.50 ms │ no change │
│ QQuery 99 │   65.13 ms │        65.57 ms │ no change │
└───────────┴────────────┴─────────────────┴───────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary              ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 9106.26ms │
│ Total Time (blocked-agg-poc)   │ 9126.56ms │
│ Average Time (HEAD)            │   91.98ms │
│ Average Time (blocked-agg-poc) │   92.19ms │
│ Queries Faster                 │         0 │
│ Queries Slower                 │         0 │
│ Queries with No Change         │        99 │
│ Queries with Failure           │         0 │
└────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark tpcds_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                  HEAD ┃                       blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │           5.71 / 6.35 ±0.97 / 8.27 ms │           5.55 / 6.17 ±1.01 / 8.18 ms │     no change │
│ QQuery 2  │        80.77 / 80.92 ±0.14 / 81.09 ms │        80.32 / 80.78 ±0.29 / 81.18 ms │     no change │
│ QQuery 3  │        28.79 / 29.15 ±0.22 / 29.47 ms │        29.10 / 29.21 ±0.08 / 29.33 ms │     no change │
│ QQuery 4  │     455.88 / 457.90 ±1.19 / 459.09 ms │     458.28 / 462.70 ±3.52 / 469.05 ms │     no change │
│ QQuery 5  │        50.53 / 51.18 ±0.50 / 51.97 ms │        51.09 / 53.08 ±2.77 / 58.38 ms │     no change │
│ QQuery 6  │        35.26 / 35.54 ±0.20 / 35.79 ms │        35.35 / 35.57 ±0.23 / 35.97 ms │     no change │
│ QQuery 7  │        73.63 / 74.08 ±0.29 / 74.43 ms │        73.55 / 74.86 ±0.92 / 76.28 ms │     no change │
│ QQuery 8  │        35.92 / 36.12 ±0.19 / 36.46 ms │        36.29 / 36.51 ±0.14 / 36.68 ms │     no change │
│ QQuery 9  │        49.66 / 52.69 ±2.26 / 55.56 ms │        50.72 / 53.78 ±4.02 / 61.49 ms │     no change │
│ QQuery 10 │        61.47 / 61.77 ±0.29 / 62.31 ms │        60.90 / 61.46 ±0.45 / 61.93 ms │     no change │
│ QQuery 11 │     286.46 / 290.40 ±6.35 / 303.03 ms │     285.77 / 288.35 ±1.81 / 291.23 ms │     no change │
│ QQuery 12 │        28.96 / 30.28 ±2.34 / 34.95 ms │        28.88 / 30.41 ±2.02 / 34.34 ms │     no change │
│ QQuery 13 │     117.72 / 118.49 ±0.59 / 119.48 ms │     116.99 / 117.95 ±0.93 / 119.34 ms │     no change │
│ QQuery 14 │     398.56 / 402.11 ±2.12 / 404.52 ms │     393.09 / 398.02 ±4.08 / 403.56 ms │     no change │
│ QQuery 15 │        54.46 / 55.13 ±0.51 / 55.94 ms │        54.64 / 55.19 ±0.47 / 55.93 ms │     no change │
│ QQuery 16 │           6.50 / 6.67 ±0.17 / 6.99 ms │           6.35 / 6.54 ±0.19 / 6.89 ms │     no change │
│ QQuery 17 │        77.64 / 80.72 ±3.84 / 88.27 ms │        78.98 / 79.93 ±1.15 / 82.16 ms │     no change │
│ QQuery 18 │     104.63 / 105.23 ±0.66 / 106.37 ms │     104.12 / 105.69 ±1.01 / 107.30 ms │     no change │
│ QQuery 19 │        41.30 / 43.01 ±2.80 / 48.59 ms │        41.59 / 42.82 ±1.42 / 45.53 ms │     no change │
│ QQuery 20 │        36.04 / 36.26 ±0.26 / 36.76 ms │        35.50 / 36.17 ±0.52 / 36.75 ms │     no change │
│ QQuery 21 │        16.66 / 17.05 ±0.21 / 17.28 ms │        16.96 / 17.22 ±0.27 / 17.73 ms │     no change │
│ QQuery 22 │        63.62 / 64.46 ±0.62 / 65.46 ms │        64.62 / 65.70 ±0.69 / 66.75 ms │     no change │
│ QQuery 23 │     313.15 / 314.99 ±1.36 / 317.07 ms │     312.19 / 314.19 ±1.32 / 315.60 ms │     no change │
│ QQuery 24 │     195.71 / 201.48 ±5.15 / 208.98 ms │     198.20 / 204.86 ±5.54 / 214.15 ms │     no change │
│ QQuery 25 │     107.64 / 108.30 ±0.52 / 109.02 ms │     107.97 / 108.35 ±0.36 / 108.98 ms │     no change │
│ QQuery 26 │        48.07 / 48.51 ±0.25 / 48.83 ms │        47.77 / 48.28 ±0.52 / 49.08 ms │     no change │
│ QQuery 27 │           6.06 / 6.20 ±0.11 / 6.39 ms │           6.08 / 6.22 ±0.13 / 6.45 ms │     no change │
│ QQuery 28 │        60.89 / 61.11 ±0.28 / 61.65 ms │        61.09 / 61.40 ±0.29 / 61.93 ms │     no change │
│ QQuery 29 │       96.34 / 98.49 ±2.39 / 103.03 ms │        94.54 / 95.98 ±0.91 / 97.31 ms │     no change │
│ QQuery 30 │        32.58 / 32.90 ±0.27 / 33.27 ms │        32.65 / 33.82 ±1.29 / 36.29 ms │     no change │
│ QQuery 31 │     109.15 / 110.98 ±1.80 / 113.68 ms │     109.36 / 110.84 ±1.53 / 113.46 ms │     no change │
│ QQuery 32 │        20.52 / 20.68 ±0.09 / 20.81 ms │        20.51 / 20.71 ±0.19 / 20.96 ms │     no change │
│ QQuery 33 │        37.50 / 37.75 ±0.34 / 38.42 ms │        37.36 / 37.67 ±0.32 / 38.24 ms │     no change │
│ QQuery 34 │        10.24 / 10.88 ±0.39 / 11.28 ms │        10.21 / 10.37 ±0.12 / 10.57 ms │     no change │
│ QQuery 35 │        72.89 / 73.72 ±0.72 / 74.70 ms │        72.91 / 74.20 ±1.85 / 77.85 ms │     no change │
│ QQuery 36 │           5.72 / 5.88 ±0.23 / 6.34 ms │           5.71 / 5.86 ±0.18 / 6.18 ms │     no change │
│ QQuery 37 │           6.84 / 6.95 ±0.07 / 7.02 ms │           6.92 / 7.09 ±0.11 / 7.22 ms │     no change │
│ QQuery 38 │        61.49 / 62.43 ±0.69 / 63.42 ms │        63.22 / 63.64 ±0.30 / 64.00 ms │     no change │
│ QQuery 39 │        73.47 / 74.22 ±0.59 / 74.95 ms │        74.29 / 75.43 ±1.60 / 78.57 ms │     no change │
│ QQuery 40 │        24.34 / 24.93 ±0.74 / 26.39 ms │        23.74 / 24.25 ±0.39 / 24.89 ms │     no change │
│ QQuery 41 │        11.39 / 11.51 ±0.10 / 11.70 ms │        11.23 / 11.40 ±0.12 / 11.60 ms │     no change │
│ QQuery 42 │        23.83 / 23.97 ±0.14 / 24.15 ms │        23.87 / 24.23 ±0.52 / 25.23 ms │     no change │
│ QQuery 43 │           4.83 / 4.99 ±0.17 / 5.32 ms │           4.76 / 4.90 ±0.17 / 5.24 ms │     no change │
│ QQuery 44 │           9.11 / 9.18 ±0.06 / 9.26 ms │           8.82 / 9.00 ±0.12 / 9.19 ms │     no change │
│ QQuery 45 │        37.07 / 39.38 ±2.26 / 43.64 ms │        38.72 / 39.83 ±2.05 / 43.93 ms │     no change │
│ QQuery 46 │        12.08 / 12.30 ±0.16 / 12.50 ms │        11.83 / 11.89 ±0.07 / 12.01 ms │     no change │
│ QQuery 47 │     229.09 / 232.94 ±2.67 / 237.04 ms │     228.86 / 231.20 ±2.99 / 236.97 ms │     no change │
│ QQuery 48 │        94.50 / 94.91 ±0.28 / 95.21 ms │        95.78 / 98.75 ±1.50 / 99.74 ms │     no change │
│ QQuery 49 │        70.36 / 72.58 ±2.99 / 78.41 ms │        70.05 / 71.33 ±1.39 / 74.02 ms │     no change │
│ QQuery 50 │        57.91 / 58.19 ±0.27 / 58.61 ms │        57.70 / 58.30 ±0.65 / 59.52 ms │     no change │
│ QQuery 51 │        94.03 / 96.17 ±1.59 / 98.43 ms │        94.95 / 96.21 ±1.16 / 98.22 ms │     no change │
│ QQuery 52 │        24.01 / 24.16 ±0.12 / 24.38 ms │        24.65 / 26.60 ±3.23 / 33.03 ms │  1.10x slower │
│ QQuery 53 │        29.00 / 30.74 ±3.14 / 37.02 ms │        29.31 / 29.84 ±0.28 / 30.09 ms │     no change │
│ QQuery 54 │        54.33 / 55.16 ±0.48 / 55.71 ms │        54.14 / 54.84 ±0.56 / 55.70 ms │     no change │
│ QQuery 55 │        24.03 / 25.21 ±0.97 / 26.17 ms │        23.99 / 24.12 ±0.14 / 24.33 ms │     no change │
│ QQuery 56 │        38.75 / 39.10 ±0.41 / 39.89 ms │        38.88 / 39.46 ±0.51 / 40.36 ms │     no change │
│ QQuery 57 │     169.04 / 171.69 ±2.75 / 176.60 ms │     171.93 / 174.47 ±2.48 / 178.67 ms │     no change │
│ QQuery 58 │     109.36 / 110.77 ±1.49 / 112.92 ms │     109.07 / 110.94 ±1.49 / 113.15 ms │     no change │
│ QQuery 59 │     115.10 / 116.18 ±1.04 / 118.12 ms │     116.43 / 117.02 ±0.45 / 117.69 ms │     no change │
│ QQuery 60 │        38.87 / 39.58 ±0.44 / 40.00 ms │        38.78 / 39.69 ±0.49 / 40.24 ms │     no change │
│ QQuery 61 │        11.84 / 11.89 ±0.03 / 11.94 ms │        11.57 / 11.70 ±0.08 / 11.79 ms │     no change │
│ QQuery 62 │        44.14 / 45.77 ±2.89 / 51.54 ms │        44.15 / 44.42 ±0.22 / 44.83 ms │     no change │
│ QQuery 63 │        29.26 / 29.57 ±0.24 / 29.93 ms │        29.20 / 29.53 ±0.17 / 29.69 ms │     no change │
│ QQuery 64 │     365.87 / 369.81 ±3.31 / 375.45 ms │     362.78 / 366.44 ±2.89 / 371.15 ms │     no change │
│ QQuery 65 │     127.82 / 131.89 ±2.53 / 135.73 ms │     130.87 / 133.95 ±2.24 / 136.87 ms │     no change │
│ QQuery 66 │        76.91 / 80.03 ±5.39 / 90.77 ms │        76.82 / 79.50 ±4.55 / 88.56 ms │     no change │
│ QQuery 67 │     255.11 / 263.72 ±8.28 / 275.60 ms │     254.75 / 260.54 ±5.62 / 270.35 ms │     no change │
│ QQuery 68 │        12.09 / 12.17 ±0.06 / 12.27 ms │        11.84 / 12.02 ±0.19 / 12.38 ms │     no change │
│ QQuery 69 │        55.49 / 55.99 ±0.51 / 56.96 ms │        55.72 / 55.90 ±0.13 / 56.05 ms │     no change │
│ QQuery 70 │     104.37 / 108.38 ±7.46 / 123.28 ms │     103.54 / 107.91 ±5.70 / 119.12 ms │     no change │
│ QQuery 71 │        35.13 / 35.77 ±0.49 / 36.46 ms │        35.26 / 35.83 ±0.62 / 36.81 ms │     no change │
│ QQuery 72 │ 1780.57 / 1837.39 ±37.41 / 1883.70 ms │ 1790.74 / 1824.35 ±20.86 / 1851.14 ms │     no change │
│ QQuery 73 │        10.08 / 10.60 ±0.45 / 11.12 ms │        10.03 / 10.18 ±0.12 / 10.33 ms │     no change │
│ QQuery 74 │     166.25 / 168.45 ±1.23 / 169.72 ms │     164.08 / 169.37 ±4.28 / 175.26 ms │     no change │
│ QQuery 75 │     140.70 / 143.46 ±4.42 / 152.25 ms │     140.12 / 143.95 ±5.45 / 154.72 ms │     no change │
│ QQuery 76 │        34.18 / 34.96 ±0.55 / 35.85 ms │        34.11 / 35.25 ±1.21 / 37.55 ms │     no change │
│ QQuery 77 │        60.57 / 60.99 ±0.39 / 61.74 ms │        60.09 / 60.93 ±0.60 / 61.69 ms │     no change │
│ QQuery 78 │     163.54 / 170.91 ±8.77 / 187.42 ms │     163.21 / 167.38 ±3.61 / 173.13 ms │     no change │
│ QQuery 79 │        66.17 / 66.79 ±0.83 / 68.40 ms │        66.13 / 66.44 ±0.21 / 66.72 ms │     no change │
│ QQuery 80 │      94.91 / 100.05 ±7.32 / 114.57 ms │       96.51 / 98.22 ±1.31 / 100.16 ms │     no change │
│ QQuery 81 │        26.04 / 26.37 ±0.28 / 26.78 ms │        26.12 / 26.41 ±0.21 / 26.73 ms │     no change │
│ QQuery 82 │        16.50 / 16.76 ±0.23 / 17.03 ms │        16.53 / 16.96 ±0.28 / 17.39 ms │     no change │
│ QQuery 83 │        33.48 / 33.73 ±0.23 / 34.15 ms │        34.01 / 34.30 ±0.24 / 34.74 ms │     no change │
│ QQuery 84 │        29.57 / 31.63 ±3.89 / 39.40 ms │        29.65 / 29.87 ±0.19 / 30.20 ms │ +1.06x faster │
│ QQuery 85 │     102.91 / 104.91 ±1.94 / 108.08 ms │     103.32 / 108.21 ±6.49 / 121.01 ms │     no change │
│ QQuery 86 │        25.57 / 25.99 ±0.40 / 26.60 ms │        26.28 / 26.52 ±0.21 / 26.81 ms │     no change │
│ QQuery 87 │        62.13 / 62.45 ±0.23 / 62.74 ms │        64.41 / 64.89 ±0.52 / 65.71 ms │     no change │
│ QQuery 88 │        60.66 / 61.72 ±1.45 / 64.54 ms │        60.49 / 63.42 ±4.96 / 73.29 ms │     no change │
│ QQuery 89 │        34.71 / 35.75 ±0.77 / 37.07 ms │        35.52 / 37.26 ±1.89 / 40.91 ms │     no change │
│ QQuery 90 │        16.91 / 17.02 ±0.10 / 17.18 ms │        16.96 / 17.05 ±0.09 / 17.22 ms │     no change │
│ QQuery 91 │        44.31 / 44.69 ±0.35 / 45.16 ms │        44.32 / 44.96 ±0.33 / 45.30 ms │     no change │
│ QQuery 92 │        29.38 / 29.93 ±0.30 / 30.19 ms │        29.43 / 29.91 ±0.49 / 30.66 ms │     no change │
│ QQuery 93 │        48.89 / 49.49 ±0.46 / 50.17 ms │        48.96 / 49.87 ±0.56 / 50.54 ms │     no change │
│ QQuery 94 │        38.25 / 38.59 ±0.34 / 39.10 ms │        38.44 / 38.92 ±0.42 / 39.44 ms │     no change │
│ QQuery 95 │        79.58 / 80.35 ±0.78 / 81.76 ms │        80.48 / 81.92 ±1.58 / 84.86 ms │     no change │
│ QQuery 96 │        23.85 / 23.93 ±0.08 / 24.05 ms │        23.94 / 24.16 ±0.26 / 24.51 ms │     no change │
│ QQuery 97 │        51.36 / 53.24 ±1.73 / 56.47 ms │        50.99 / 52.31 ±1.38 / 54.74 ms │     no change │
│ QQuery 98 │        42.58 / 43.14 ±0.30 / 43.40 ms │        42.50 / 43.42 ±0.76 / 44.41 ms │     no change │
│ QQuery 99 │        65.13 / 66.13 ±0.99 / 67.96 ms │        65.57 / 66.24 ±0.84 / 67.79 ms │     no change │
└───────────┴───────────────────────────────────────┴───────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary              ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 9289.02ms │
│ Total Time (blocked-agg-poc)   │ 9285.72ms │
│ Average Time (HEAD)            │   93.83ms │
│ Average Time (blocked-agg-poc) │   93.80ms │
│ Queries Faster                 │         1 │
│ Queries Slower                 │         1 │
│ Queries with No Change         │        97 │
│ Queries with Failure           │         0 │
└────────────────────────────────┴───────────┘

Resource Usage

tpcds — base (merge-base)

Metric Value
Wall time 50.0s
Peak memory 1.7 GiB
Avg memory 1.2 GiB
CPU user 203.4s
CPU sys 5.9s
Peak spill 0 B

tpcds — branch

Metric Value
Wall time 50.0s
Peak memory 1.6 GiB
Avg memory 1.2 GiB
CPU user 202.6s
CPU sys 5.8s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-standard-32 (12 vCPU / 65 GiB) | Linux bench-c5914813281-2973-4h8nr 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  32
On-line CPU(s) list:                     0-31
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     32
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               2 MiB (32 instances)
L1i cache:                               2 MiB (32 instances)
L2 cache:                                64 MiB (32 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-31
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark h2o_medium
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "16G"

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-standard-32 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark external_aggr
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  32
On-line CPU(s) list:                     0-31
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     32
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               2 MiB (32 instances)
L1i cache:                               2 MiB (32 instances)
L2 cache:                                64 MiB (32 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-31
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark external_aggr.json
--------------------
┏━━━━━━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query        ┃      HEAD ┃ blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ Q1(64.0 MB)  │  53.59 ms │        54.43 ms │     no change │
│ Q1(32.0 MB)  │  51.64 ms │        46.20 ms │ +1.12x faster │
│ Q1(16.0 MB)  │  54.42 ms │        48.70 ms │ +1.12x faster │
│ Q2(512.0 MB) │ 282.48 ms │       283.19 ms │     no change │
│ Q2(256.0 MB) │ 276.11 ms │       276.52 ms │     no change │
│ Q2(128.0 MB) │ 247.25 ms │       243.96 ms │     no change │
│ Q2(64.0 MB)  │ 241.45 ms │       241.93 ms │     no change │
│ Q2(32.0 MB)  │ 306.02 ms │       303.84 ms │     no change │
└──────────────┴───────────┴─────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary              ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 1512.96ms │
│ Total Time (blocked-agg-poc)   │ 1498.78ms │
│ Average Time (HEAD)            │  189.12ms │
│ Average Time (blocked-agg-poc) │  187.35ms │
│ Queries Faster                 │         2 │
│ Queries Slower                 │         0 │
│ Queries with No Change         │         6 │
│ Queries with Failure           │         0 │
└────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark external_aggr.json
--------------------
┏━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query        ┃                               HEAD ┃                    blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ Q1(64.0 MB)  │     53.59 / 57.55 ±4.12 / 65.21 ms │     54.43 / 58.98 ±4.98 / 68.49 ms │     no change │
│ Q1(32.0 MB)  │     51.64 / 53.08 ±1.15 / 54.69 ms │     46.20 / 49.08 ±2.49 / 52.40 ms │ +1.08x faster │
│ Q1(16.0 MB)  │     54.42 / 55.58 ±1.18 / 57.50 ms │     48.70 / 51.61 ±2.24 / 54.49 ms │ +1.08x faster │
│ Q2(512.0 MB) │  282.48 / 294.54 ±6.67 / 301.08 ms │  283.19 / 293.49 ±6.42 / 302.88 ms │     no change │
│ Q2(256.0 MB) │  276.11 / 282.57 ±3.91 / 287.90 ms │ 276.52 / 287.30 ±10.48 / 306.07 ms │     no change │
│ Q2(128.0 MB) │ 247.25 / 258.43 ±13.72 / 279.32 ms │  243.96 / 248.41 ±5.37 / 258.36 ms │     no change │
│ Q2(64.0 MB)  │  241.45 / 243.75 ±1.56 / 245.36 ms │ 241.93 / 249.58 ±11.37 / 272.08 ms │     no change │
│ Q2(32.0 MB)  │  306.02 / 308.00 ±1.94 / 310.89 ms │  303.84 / 307.55 ±2.02 / 309.09 ms │     no change │
└──────────────┴────────────────────────────────────┴────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary              ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 1553.51ms │
│ Total Time (blocked-agg-poc)   │ 1546.00ms │
│ Average Time (HEAD)            │  194.19ms │
│ Average Time (blocked-agg-poc) │  193.25ms │
│ Queries Faster                 │         2 │
│ Queries Slower                 │         0 │
│ Queries with No Change         │         6 │
│ Queries with Failure           │         0 │
└────────────────────────────────┴───────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only for benchmarks that write a results JSON (the dfbench suites, and the suites bench.sh runs through the Criterion SQL harness since #25644), and only under a memory limit: DATAFUSION_RUNTIME_MEMORY_LIMIT, or one the suite sets itself.

Base: 6924c6a (merge-base) | Changed: blocked-agg-poc

external_aggr

Query Base Changed Change
1(64.0 MB) 36.8 MiB 28.4 MiB -22.7%
1(32.0 MB) 18.6 MiB 17.6 MiB -5.5%
1(16.0 MB) 9.6 MiB 10.3 MiB +8.0%
2(512.0 MB) 135.7 MiB 145.0 MiB +6.8%
2(256.0 MB) 97.4 MiB 97.2 MiB -0.1%
2(128.0 MB) 49.1 MiB 49.0 MiB -0.3%
2(64.0 MB) 32.5 MiB 34.4 MiB +5.9%
2(32.0 MB) 21.0 MiB 17.5 MiB -16.7%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
external_aggr base (6924c6a (merge-base)) 135.7 MiB 454.9 MiB 319.1 MiB 3.4×
external_aggr changed (blocked-agg-poc) 145.0 MiB 419.6 MiB 274.6 MiB 2.9×
Resource Usage

external_aggr — base (merge-base)

Metric Value
Wall time 105.0s
Peak memory 454.9 MiB
Avg memory 30.8 MiB
CPU user 26.0s
CPU sys 3.9s
Peak spill 83.9 MiB

external_aggr — branch

Metric Value
Wall time 110.0s
Peak memory 419.6 MiB
Avg memory 30.5 MiB
CPU user 25.7s
CPU sys 4.0s
Peak spill 91.3 MiB

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5914813281-2974-n9vq2 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark clickbench_partitioned
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "16G"

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-standard-32 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark clickbench_partitioned
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  32
On-line CPU(s) list:                     0-31
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     32
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               2 MiB (32 instances)
L1i cache:                               2 MiB (32 instances)
L2 cache:                                64 MiB (32 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-31
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark clickbench_partitioned.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃       HEAD ┃ blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 0  │    1.23 ms │         1.23 ms │     no change │
│ QQuery 1  │   11.71 ms │        11.74 ms │     no change │
│ QQuery 2  │   36.76 ms │        36.50 ms │     no change │
│ QQuery 3  │   31.22 ms │        31.15 ms │     no change │
│ QQuery 4  │  244.67 ms │       217.40 ms │ +1.13x faster │
│ QQuery 5  │  279.47 ms │       282.26 ms │     no change │
│ QQuery 6  │    1.30 ms │         1.32 ms │     no change │
│ QQuery 7  │   13.17 ms │        13.17 ms │     no change │
│ QQuery 8  │  343.02 ms │       341.35 ms │     no change │
│ QQuery 9  │  483.64 ms │       476.22 ms │     no change │
│ QQuery 10 │   64.40 ms │        64.64 ms │     no change │
│ QQuery 11 │   74.92 ms │        74.68 ms │     no change │
│ QQuery 12 │  269.79 ms │       274.93 ms │     no change │
│ QQuery 13 │  373.82 ms │       368.19 ms │     no change │
│ QQuery 14 │  286.70 ms │       289.31 ms │     no change │
│ QQuery 15 │  294.92 ms │       277.61 ms │ +1.06x faster │
│ QQuery 16 │  636.72 ms │       652.15 ms │     no change │
│ QQuery 17 │  636.82 ms │       649.34 ms │     no change │
│ QQuery 18 │ 1287.36 ms │      1308.64 ms │     no change │
│ QQuery 19 │   27.83 ms │        27.64 ms │     no change │
│ QQuery 20 │  526.77 ms │       522.78 ms │     no change │
│ QQuery 21 │  520.27 ms │       512.46 ms │     no change │
│ QQuery 22 │ 1042.86 ms │      1014.08 ms │     no change │
│ QQuery 23 │ 3152.13 ms │      3105.29 ms │     no change │
│ QQuery 24 │   40.84 ms │        42.20 ms │     no change │
│ QQuery 25 │  105.47 ms │       106.17 ms │     no change │
│ QQuery 26 │   41.60 ms │        41.46 ms │     no change │
│ QQuery 27 │  519.39 ms │       525.48 ms │     no change │
│ QQuery 28 │ 2857.67 ms │      2912.86 ms │     no change │
│ QQuery 29 │   41.70 ms │        41.81 ms │     no change │
│ QQuery 30 │  315.54 ms │       312.04 ms │     no change │
│ QQuery 31 │  285.11 ms │       286.82 ms │     no change │
│ QQuery 32 │ 1121.62 ms │      1113.99 ms │     no change │
│ QQuery 33 │ 1613.26 ms │      1651.70 ms │     no change │
│ QQuery 34 │ 1638.70 ms │      1660.21 ms │     no change │
│ QQuery 35 │  308.45 ms │       264.85 ms │ +1.16x faster │
│ QQuery 36 │   69.84 ms │        70.23 ms │     no change │
│ QQuery 37 │   35.10 ms │        36.45 ms │     no change │
│ QQuery 38 │   40.66 ms │        44.06 ms │  1.08x slower │
│ QQuery 39 │  154.61 ms │       144.60 ms │ +1.07x faster │
│ QQuery 40 │   14.30 ms │        14.75 ms │     no change │
│ QQuery 41 │   14.16 ms │        13.93 ms │     no change │
│ QQuery 42 │   13.66 ms │        13.55 ms │     no change │
└───────────┴────────────┴─────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 19873.17ms │
│ Total Time (blocked-agg-poc)   │ 19851.23ms │
│ Average Time (HEAD)            │   462.17ms │
│ Average Time (blocked-agg-poc) │   461.66ms │
│ Queries Faster                 │          4 │
│ Queries Slower                 │          1 │
│ Queries with No Change         │         38 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark clickbench_partitioned.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                  HEAD ┃                        blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 0  │          1.23 / 4.13 ±5.65 / 15.43 ms │           1.23 / 4.15 ±5.66 / 15.47 ms │     no change │
│ QQuery 1  │        11.71 / 11.93 ±0.12 / 12.02 ms │         11.74 / 11.98 ±0.14 / 12.13 ms │     no change │
│ QQuery 2  │        36.76 / 36.98 ±0.21 / 37.33 ms │         36.50 / 36.83 ±0.20 / 37.14 ms │     no change │
│ QQuery 3  │        31.22 / 31.72 ±0.56 / 32.79 ms │         31.15 / 31.27 ±0.12 / 31.47 ms │     no change │
│ QQuery 4  │     244.67 / 246.73 ±1.98 / 249.72 ms │      217.40 / 219.58 ±1.90 / 222.14 ms │ +1.12x faster │
│ QQuery 5  │     279.47 / 281.96 ±2.19 / 285.78 ms │      282.26 / 283.79 ±1.75 / 286.64 ms │     no change │
│ QQuery 6  │           1.30 / 1.44 ±0.21 / 1.86 ms │            1.32 / 1.44 ±0.21 / 1.85 ms │     no change │
│ QQuery 7  │        13.17 / 13.27 ±0.11 / 13.41 ms │         13.17 / 13.31 ±0.15 / 13.60 ms │     no change │
│ QQuery 8  │     343.02 / 345.22 ±1.33 / 346.87 ms │      341.35 / 343.51 ±1.15 / 344.71 ms │     no change │
│ QQuery 9  │    483.64 / 494.82 ±11.16 / 512.71 ms │     476.22 / 494.23 ±10.94 / 506.80 ms │     no change │
│ QQuery 10 │        64.40 / 65.01 ±0.53 / 65.61 ms │         64.64 / 65.96 ±1.36 / 68.42 ms │     no change │
│ QQuery 11 │        74.92 / 75.47 ±0.42 / 76.17 ms │         74.68 / 75.50 ±0.88 / 77.19 ms │     no change │
│ QQuery 12 │     269.79 / 276.05 ±4.13 / 280.55 ms │      274.93 / 280.82 ±5.62 / 290.45 ms │     no change │
│ QQuery 13 │    373.82 / 389.46 ±13.66 / 411.72 ms │     368.19 / 390.42 ±20.36 / 424.02 ms │     no change │
│ QQuery 14 │     286.70 / 290.17 ±3.24 / 295.72 ms │      289.31 / 296.59 ±8.19 / 312.18 ms │     no change │
│ QQuery 15 │     294.92 / 301.97 ±5.02 / 309.90 ms │      277.61 / 283.75 ±6.00 / 294.70 ms │ +1.06x faster │
│ QQuery 16 │     636.72 / 647.51 ±5.70 / 653.16 ms │      652.15 / 659.52 ±5.47 / 667.68 ms │     no change │
│ QQuery 17 │    636.82 / 659.91 ±21.90 / 701.42 ms │      649.34 / 662.24 ±8.41 / 672.69 ms │     no change │
│ QQuery 18 │ 1287.36 / 1328.23 ±27.61 / 1362.62 ms │  1308.64 / 1341.96 ±36.88 / 1408.98 ms │     no change │
│ QQuery 19 │        27.83 / 28.11 ±0.32 / 28.72 ms │         27.64 / 31.34 ±6.46 / 44.19 ms │  1.11x slower │
│ QQuery 20 │    526.77 / 536.21 ±10.66 / 550.35 ms │      522.78 / 532.55 ±8.41 / 543.25 ms │     no change │
│ QQuery 21 │     520.27 / 523.42 ±2.47 / 527.72 ms │      512.46 / 518.45 ±5.76 / 528.69 ms │     no change │
│ QQuery 22 │ 1042.86 / 1089.37 ±40.81 / 1144.09 ms │  1014.08 / 1027.51 ±12.57 / 1047.16 ms │ +1.06x faster │
│ QQuery 23 │ 3152.13 / 3277.24 ±90.79 / 3406.34 ms │ 3105.29 / 3339.87 ±193.40 / 3656.66 ms │     no change │
│ QQuery 24 │    40.84 / 121.84 ±130.38 / 378.10 ms │     42.20 / 124.69 ±153.80 / 431.96 ms │     no change │
│ QQuery 25 │     105.47 / 111.06 ±6.91 / 123.66 ms │      106.17 / 107.81 ±1.96 / 111.45 ms │     no change │
│ QQuery 26 │        41.60 / 42.79 ±1.84 / 46.45 ms │         41.46 / 42.10 ±0.47 / 42.77 ms │     no change │
│ QQuery 27 │     519.39 / 526.27 ±6.19 / 535.11 ms │     525.48 / 549.93 ±25.16 / 597.99 ms │     no change │
│ QQuery 28 │ 2857.67 / 2886.23 ±15.17 / 2902.28 ms │  2912.86 / 2958.78 ±29.83 / 2992.56 ms │     no change │
│ QQuery 29 │      41.70 / 57.99 ±32.11 / 122.21 ms │       41.81 / 54.41 ±23.10 / 100.56 ms │ +1.07x faster │
│ QQuery 30 │     315.54 / 323.52 ±6.99 / 335.86 ms │      312.04 / 318.71 ±3.57 / 322.82 ms │     no change │
│ QQuery 31 │    285.11 / 302.73 ±12.70 / 324.09 ms │      286.82 / 297.64 ±8.31 / 311.93 ms │     no change │
│ QQuery 32 │ 1121.62 / 1153.03 ±34.95 / 1207.84 ms │  1113.99 / 1194.27 ±77.82 / 1304.91 ms │     no change │
│ QQuery 33 │ 1613.26 / 1667.06 ±41.16 / 1734.13 ms │  1651.70 / 1722.63 ±40.79 / 1761.79 ms │     no change │
│ QQuery 34 │ 1638.70 / 1663.17 ±17.75 / 1692.69 ms │  1660.21 / 1679.52 ±12.12 / 1695.32 ms │     no change │
│ QQuery 35 │    308.45 / 346.11 ±34.15 / 393.65 ms │    264.85 / 379.78 ±174.26 / 724.81 ms │  1.10x slower │
│ QQuery 36 │        69.84 / 77.59 ±4.41 / 82.77 ms │         70.23 / 78.21 ±9.68 / 96.85 ms │     no change │
│ QQuery 37 │        35.10 / 35.80 ±0.64 / 36.90 ms │         36.45 / 37.84 ±1.01 / 39.35 ms │  1.06x slower │
│ QQuery 38 │        40.66 / 47.02 ±4.92 / 55.00 ms │         44.06 / 46.64 ±3.55 / 53.61 ms │     no change │
│ QQuery 39 │     154.61 / 164.67 ±6.01 / 169.74 ms │      144.60 / 154.88 ±8.42 / 164.57 ms │ +1.06x faster │
│ QQuery 40 │        14.30 / 14.74 ±0.28 / 15.00 ms │         14.75 / 15.88 ±1.90 / 19.65 ms │  1.08x slower │
│ QQuery 41 │        14.16 / 14.48 ±0.32 / 15.07 ms │         13.93 / 14.36 ±0.27 / 14.75 ms │     no change │
│ QQuery 42 │        13.66 / 15.23 ±2.62 / 20.45 ms │         13.55 / 15.67 ±4.04 / 23.75 ms │     no change │
└───────────┴───────────────────────────────────────┴────────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 20527.66ms │
│ Total Time (blocked-agg-poc)   │ 20740.32ms │
│ Average Time (HEAD)            │   477.39ms │
│ Average Time (blocked-agg-poc) │   482.33ms │
│ Queries Faster                 │          5 │
│ Queries Slower                 │          4 │
│ Queries with No Change         │         34 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Resource Usage

clickbench_partitioned — base (merge-base)

Metric Value
Wall time 105.0s
Peak memory 18.1 GiB
Avg memory 5.9 GiB
CPU user 1012.4s
CPU sys 100.9s
Peak spill 0 B

clickbench_partitioned — branch

Metric Value
Wall time 105.0s
Peak memory 18.1 GiB
Avg memory 5.9 GiB
CPU user 1015.6s
CPU sys 104.1s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-standard-32 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark h2o_medium
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  32
On-line CPU(s) list:                     0-31
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     32
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               2 MiB (32 instances)
L1i cache:                               2 MiB (32 instances)
L2 cache:                                64 MiB (32 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-31
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark h2o.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Query     ┃       HEAD ┃ blocked-agg-poc ┃    Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ QQuery 1  │ 1097.19 ms │      1097.39 ms │ no change │
│ QQuery 2  │ 2355.81 ms │      2416.74 ms │ no change │
│ QQuery 3  │ 2047.24 ms │      2044.84 ms │ no change │
│ QQuery 4  │ 1401.17 ms │      1399.03 ms │ no change │
│ QQuery 5  │ 1922.47 ms │      1829.24 ms │ no change │
│ QQuery 6  │ 1652.43 ms │      1643.00 ms │ no change │
│ QQuery 7  │ 1927.19 ms │      1949.25 ms │ no change │
│ QQuery 8  │ 3534.72 ms │      3525.32 ms │ no change │
│ QQuery 9  │ 2852.33 ms │      2918.40 ms │ no change │
│ QQuery 10 │ 3073.10 ms │      3039.39 ms │ no change │
└───────────┴────────────┴─────────────────┴───────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 21863.66ms │
│ Total Time (blocked-agg-poc)   │ 21862.61ms │
│ Average Time (HEAD)            │  2186.37ms │
│ Average Time (blocked-agg-poc) │  2186.26ms │
│ Queries Faster                 │          0 │
│ Queries Slower                 │          0 │
│ Queries with No Change         │         10 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark h2o.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                  HEAD ┃                       blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │ 1097.19 / 1106.55 ±12.46 / 1124.16 ms │  1097.39 / 1098.44 ±0.86 / 1099.49 ms │     no change │
│ QQuery 2  │ 2355.81 / 2428.57 ±55.16 / 2489.32 ms │ 2416.74 / 2456.98 ±34.02 / 2499.94 ms │     no change │
│ QQuery 3  │ 2047.24 / 2065.30 ±12.92 / 2076.80 ms │ 2044.84 / 2109.42 ±46.00 / 2148.49 ms │     no change │
│ QQuery 4  │  1401.17 / 1402.43 ±1.42 / 1404.40 ms │  1399.03 / 1400.29 ±0.89 / 1400.94 ms │     no change │
│ QQuery 5  │ 1922.47 / 1940.01 ±12.44 / 1949.93 ms │  1829.24 / 1839.55 ±9.46 / 1852.09 ms │ +1.05x faster │
│ QQuery 6  │  1652.43 / 1655.84 ±4.54 / 1662.26 ms │ 1643.00 / 1658.05 ±12.40 / 1673.37 ms │     no change │
│ QQuery 7  │ 1927.19 / 1942.52 ±12.08 / 1956.70 ms │ 1949.25 / 1970.54 ±16.23 / 1988.63 ms │     no change │
│ QQuery 8  │ 3534.72 / 3598.24 ±82.99 / 3715.46 ms │ 3525.32 / 3599.61 ±66.22 / 3686.13 ms │     no change │
│ QQuery 9  │ 2852.33 / 2897.44 ±34.76 / 2936.92 ms │ 2918.40 / 2932.82 ±12.37 / 2948.60 ms │     no change │
│ QQuery 10 │ 3073.10 / 3102.28 ±39.72 / 3158.44 ms │ 3039.39 / 3082.20 ±53.56 / 3157.72 ms │     no change │
└───────────┴───────────────────────────────────────┴───────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 22139.18ms │
│ Total Time (blocked-agg-poc)   │ 22147.90ms │
│ Average Time (HEAD)            │  2213.92ms │
│ Average Time (blocked-agg-poc) │  2214.79ms │
│ Queries Faster                 │          1 │
│ Queries Slower                 │          0 │
│ Queries with No Change         │          9 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Resource Usage

h2o_medium — base (merge-base)

Metric Value
Wall time 70.0s
Peak memory 10.5 GiB
Avg memory 2.9 GiB
CPU user 691.0s
CPU sys 58.8s
Peak spill 0 B

h2o_medium — branch

Metric Value
Wall time 70.0s
Peak memory 10.0 GiB
Avg memory 3.0 GiB
CPU user 688.9s
CPU sys 58.3s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-standard-32 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark tpch10
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "16G"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  32
On-line CPU(s) list:                     0-31
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     32
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               2 MiB (32 instances)
L1i cache:                               2 MiB (32 instances)
L2 cache:                                64 MiB (32 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-31
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark tpch_sf10.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃      HEAD ┃ blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │ 340.11 ms │       330.96 ms │     no change │
│ QQuery 2  │  93.54 ms │        89.71 ms │     no change │
│ QQuery 3  │ 222.37 ms │       215.91 ms │     no change │
│ QQuery 4  │ 119.44 ms │       115.17 ms │     no change │
│ QQuery 5  │ 351.54 ms │       350.23 ms │     no change │
│ QQuery 6  │ 128.03 ms │       124.88 ms │     no change │
│ QQuery 7  │ 464.04 ms │       443.28 ms │     no change │
│ QQuery 8  │ 360.02 ms │       353.55 ms │     no change │
│ QQuery 9  │ 506.27 ms │       511.14 ms │     no change │
│ QQuery 10 │ 284.96 ms │       307.70 ms │  1.08x slower │
│ QQuery 11 │  62.59 ms │        62.27 ms │     no change │
│ QQuery 12 │ 149.26 ms │       153.38 ms │     no change │
│ QQuery 13 │ 312.66 ms │       287.93 ms │ +1.09x faster │
│ QQuery 14 │ 167.82 ms │       168.01 ms │     no change │
│ QQuery 15 │ 298.11 ms │       292.37 ms │     no change │
│ QQuery 16 │  63.99 ms │        63.77 ms │     no change │
│ QQuery 17 │ 581.08 ms │       509.25 ms │ +1.14x faster │
│ QQuery 18 │ 735.12 ms │       791.49 ms │  1.08x slower │
│ QQuery 19 │ 232.39 ms │       238.90 ms │     no change │
│ QQuery 20 │ 276.89 ms │       269.98 ms │     no change │
│ QQuery 21 │ 664.30 ms │       673.05 ms │     no change │
│ QQuery 22 │  59.09 ms │        59.43 ms │     no change │
└───────────┴───────────┴─────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary              ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 6473.59ms │
│ Total Time (blocked-agg-poc)   │ 6412.37ms │
│ Average Time (HEAD)            │  294.25ms │
│ Average Time (blocked-agg-poc) │  291.47ms │
│ Queries Faster                 │         2 │
│ Queries Slower                 │         2 │
│ Queries with No Change         │        18 │
│ Queries with Failure           │         0 │
└────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark tpch_sf10.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                               HEAD ┃                    blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │  340.11 / 341.32 ±0.66 / 341.93 ms │  330.96 / 333.04 ±1.29 / 334.87 ms │     no change │
│ QQuery 2  │     93.54 / 96.86 ±2.56 / 99.92 ms │     89.71 / 91.37 ±1.15 / 92.64 ms │ +1.06x faster │
│ QQuery 3  │  222.37 / 228.57 ±4.60 / 236.60 ms │  215.91 / 219.79 ±3.40 / 224.72 ms │     no change │
│ QQuery 4  │  119.44 / 120.14 ±0.80 / 121.70 ms │  115.17 / 116.33 ±1.63 / 119.55 ms │     no change │
│ QQuery 5  │ 351.54 / 364.33 ±11.22 / 383.80 ms │  350.23 / 357.50 ±5.55 / 365.80 ms │     no change │
│ QQuery 6  │  128.03 / 130.60 ±3.13 / 136.65 ms │  124.88 / 126.84 ±2.30 / 131.36 ms │     no change │
│ QQuery 7  │ 464.04 / 488.40 ±20.53 / 520.54 ms │ 443.28 / 458.13 ±10.09 / 473.94 ms │ +1.07x faster │
│ QQuery 8  │  360.02 / 370.34 ±6.84 / 378.01 ms │  353.55 / 360.27 ±6.52 / 370.24 ms │     no change │
│ QQuery 9  │  506.27 / 522.53 ±9.05 / 532.38 ms │ 511.14 / 531.69 ±15.29 / 551.45 ms │     no change │
│ QQuery 10 │  284.96 / 296.10 ±7.65 / 306.54 ms │ 307.70 / 323.01 ±14.46 / 347.97 ms │  1.09x slower │
│ QQuery 11 │     62.59 / 66.36 ±6.84 / 80.03 ms │     62.27 / 62.83 ±0.42 / 63.58 ms │ +1.06x faster │
│ QQuery 12 │ 149.26 / 157.85 ±10.26 / 178.02 ms │ 153.38 / 163.35 ±12.38 / 183.85 ms │     no change │
│ QQuery 13 │  312.66 / 320.27 ±6.68 / 328.71 ms │ 287.93 / 297.28 ±10.04 / 315.30 ms │ +1.08x faster │
│ QQuery 14 │  167.82 / 169.92 ±3.41 / 176.70 ms │  168.01 / 174.15 ±7.37 / 185.89 ms │     no change │
│ QQuery 15 │  298.11 / 300.94 ±3.08 / 305.64 ms │  292.37 / 296.81 ±5.17 / 306.57 ms │     no change │
│ QQuery 16 │     63.99 / 67.59 ±3.23 / 72.62 ms │     63.77 / 64.91 ±1.64 / 68.13 ms │     no change │
│ QQuery 17 │ 581.08 / 595.05 ±11.74 / 615.65 ms │ 509.25 / 533.42 ±22.02 / 573.72 ms │ +1.12x faster │
│ QQuery 18 │ 735.12 / 747.95 ±11.60 / 767.62 ms │  791.49 / 797.12 ±3.83 / 802.88 ms │  1.07x slower │
│ QQuery 19 │ 232.39 / 250.71 ±14.06 / 266.08 ms │  238.90 / 250.13 ±9.75 / 262.06 ms │     no change │
│ QQuery 20 │  276.89 / 285.68 ±7.08 / 297.03 ms │  269.98 / 279.52 ±6.36 / 289.85 ms │     no change │
│ QQuery 21 │ 664.30 / 674.48 ±11.53 / 696.90 ms │  673.05 / 677.58 ±6.35 / 689.81 ms │     no change │
│ QQuery 22 │     59.09 / 62.61 ±4.06 / 70.29 ms │     59.43 / 63.01 ±2.46 / 66.41 ms │     no change │
└───────────┴────────────────────────────────────┴────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary              ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 6658.58ms │
│ Total Time (blocked-agg-poc)   │ 6578.11ms │
│ Average Time (HEAD)            │  302.66ms │
│ Average Time (blocked-agg-poc) │  299.00ms │
│ Queries Faster                 │         5 │
│ Queries Slower                 │         2 │
│ Queries with No Change         │        15 │
│ Queries with Failure           │         0 │
└────────────────────────────────┴───────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only for benchmarks that write a results JSON (the dfbench suites, and the suites bench.sh runs through the Criterion SQL harness since #25644), and only under a memory limit: DATAFUSION_RUNTIME_MEMORY_LIMIT, or one the suite sets itself.

Base: 6924c6a (merge-base) | Changed: blocked-agg-poc

tpch10 — tpch_sf10

Query Base Changed Change
Query 1 30.0 MiB 31.4 MiB +4.6%
Query 2 113.2 MiB 105.5 MiB -6.8%
Query 3 110.5 MiB 108.7 MiB -1.7%
Query 4 42.5 MiB 42.5 MiB +0.0%
Query 5 135.3 MiB 135.3 MiB +0.0%
Query 6 832 B 832 B +0.0%
Query 7 965.5 MiB 965.4 MiB -0.0%
Query 8 374.2 MiB 374.2 MiB -0.0%
Query 9 1.3 GiB 1.3 GiB +0.0%
Query 10 482.3 MiB 482.0 MiB -0.1%
Query 11 41.7 MiB 58.1 MiB +39.4%
Query 12 29.8 MiB 30.2 MiB +1.3%
Query 13 126.7 MiB 123.4 MiB -2.6%
Query 14 86.9 MiB 84.6 MiB -2.7%
Query 15 69.6 MiB 66.0 MiB -5.2%
Query 16 179.2 MiB 180.0 MiB +0.4%
Query 17 163.4 MiB 157.7 MiB -3.5%
Query 18 3.7 GiB 3.7 GiB +0.0%
Query 19 9.5 MiB 9.9 MiB +4.6%
Query 20 311.3 MiB 308.2 MiB -1.0%
Query 21 408.9 MiB 431.2 MiB +5.5%
Query 22 28.5 MiB 40.2 MiB +41.2%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
tpch10 base (6924c6a (merge-base)) 3.7 GiB 4.4 GiB 720.3 MiB 1.2×
tpch10 changed (blocked-agg-poc) 3.7 GiB 5.1 GiB 1.4 GiB 1.4×
Resource Usage

tpch10 — base (merge-base)

Metric Value
Wall time 35.0s
Peak memory 4.4 GiB
Avg memory 1.4 GiB
CPU user 340.2s
CPU sys 20.8s
Peak spill 0 B

tpch10 — branch

Metric Value
Wall time 35.0s
Peak memory 5.1 GiB
Avg memory 1.5 GiB
CPU user 335.7s
CPU sys 20.0s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark clickbench_partitioned
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "16G"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark clickbench_partitioned.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃       HEAD ┃ blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 0  │    1.27 ms │         1.29 ms │     no change │
│ QQuery 1  │   12.08 ms │        11.88 ms │     no change │
│ QQuery 2  │   37.85 ms │        37.40 ms │     no change │
│ QQuery 3  │   32.10 ms │        32.17 ms │     no change │
│ QQuery 4  │  250.89 ms │       217.78 ms │ +1.15x faster │
│ QQuery 5  │  282.25 ms │       282.38 ms │     no change │
│ QQuery 6  │    1.29 ms │         1.28 ms │     no change │
│ QQuery 7  │   13.41 ms │        13.17 ms │     no change │
│ QQuery 8  │  352.17 ms │       339.47 ms │     no change │
│ QQuery 9  │  484.67 ms │       489.61 ms │     no change │
│ QQuery 10 │   65.40 ms │        66.02 ms │     no change │
│ QQuery 11 │   76.41 ms │        76.57 ms │     no change │
│ QQuery 12 │  277.42 ms │       272.08 ms │     no change │
│ QQuery 13 │  388.82 ms │       376.16 ms │     no change │
│ QQuery 14 │  297.13 ms │       290.20 ms │     no change │
│ QQuery 15 │  309.13 ms │       278.92 ms │ +1.11x faster │
│ QQuery 16 │  653.32 ms │       643.04 ms │     no change │
│ QQuery 17 │  663.77 ms │       662.34 ms │     no change │
│ QQuery 18 │ 1329.56 ms │      1334.51 ms │     no change │
│ QQuery 19 │   28.17 ms │        27.86 ms │     no change │
│ QQuery 20 │  522.74 ms │       518.52 ms │     no change │
│ QQuery 21 │  523.06 ms │       516.20 ms │     no change │
│ QQuery 22 │ 1009.95 ms │      1030.59 ms │     no change │
│ QQuery 23 │ 3136.63 ms │      3118.14 ms │     no change │
│ QQuery 24 │   40.98 ms │        40.79 ms │     no change │
│ QQuery 25 │  107.59 ms │       106.90 ms │     no change │
│ QQuery 26 │   42.58 ms │        41.69 ms │     no change │
│ QQuery 27 │  529.18 ms │       517.74 ms │     no change │
│ QQuery 28 │ 2874.17 ms │      2920.48 ms │     no change │
│ QQuery 29 │   42.94 ms │        42.51 ms │     no change │
│ QQuery 30 │  317.57 ms │       322.48 ms │     no change │
│ QQuery 31 │  288.12 ms │       289.71 ms │     no change │
│ QQuery 32 │ 4105.70 ms │      4075.75 ms │     no change │
│ QQuery 33 │ 1576.64 ms │      1600.06 ms │     no change │
│ QQuery 34 │ 1600.05 ms │      1677.87 ms │     no change │
│ QQuery 35 │  325.26 ms │       272.52 ms │ +1.19x faster │
│ QQuery 36 │   68.28 ms │        71.72 ms │  1.05x slower │
│ QQuery 37 │   37.36 ms │        36.66 ms │     no change │
│ QQuery 38 │   44.29 ms │        43.67 ms │     no change │
│ QQuery 39 │  140.17 ms │       154.30 ms │  1.10x slower │
│ QQuery 40 │   14.84 ms │        15.37 ms │     no change │
│ QQuery 41 │   14.56 ms │        14.43 ms │     no change │
│ QQuery 42 │   14.27 ms │        13.65 ms │     no change │
└───────────┴────────────┴─────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 22934.06ms │
│ Total Time (blocked-agg-poc)   │ 22895.91ms │
│ Average Time (HEAD)            │   533.35ms │
│ Average Time (blocked-agg-poc) │   532.46ms │
│ Queries Faster                 │          3 │
│ Queries Slower                 │          2 │
│ Queries with No Change         │         38 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark clickbench_partitioned.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                   HEAD ┃                        blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 0  │           1.27 / 4.23 ±5.77 / 15.77 ms │           1.29 / 4.22 ±5.72 / 15.66 ms │     no change │
│ QQuery 1  │         12.08 / 12.36 ±0.15 / 12.49 ms │         11.88 / 12.30 ±0.32 / 12.70 ms │     no change │
│ QQuery 2  │         37.85 / 38.27 ±0.30 / 38.64 ms │         37.40 / 37.61 ±0.13 / 37.78 ms │     no change │
│ QQuery 3  │         32.10 / 33.11 ±0.72 / 34.25 ms │         32.17 / 32.68 ±0.43 / 33.31 ms │     no change │
│ QQuery 4  │      250.89 / 253.57 ±2.68 / 258.53 ms │      217.78 / 223.63 ±3.52 / 228.17 ms │ +1.13x faster │
│ QQuery 5  │      282.25 / 287.56 ±2.90 / 290.46 ms │      282.38 / 285.00 ±2.68 / 289.98 ms │     no change │
│ QQuery 6  │            1.29 / 1.44 ±0.24 / 1.90 ms │            1.28 / 1.44 ±0.23 / 1.89 ms │     no change │
│ QQuery 7  │         13.41 / 14.64 ±2.13 / 18.89 ms │         13.17 / 13.34 ±0.12 / 13.49 ms │ +1.10x faster │
│ QQuery 8  │      352.17 / 357.29 ±4.61 / 363.59 ms │      339.47 / 345.78 ±4.76 / 353.48 ms │     no change │
│ QQuery 9  │      484.67 / 499.53 ±7.79 / 506.76 ms │      489.61 / 493.34 ±2.95 / 497.78 ms │     no change │
│ QQuery 10 │         65.40 / 66.59 ±1.06 / 67.97 ms │         66.02 / 66.55 ±0.60 / 67.49 ms │     no change │
│ QQuery 11 │         76.41 / 79.93 ±5.43 / 90.75 ms │         76.57 / 77.33 ±0.88 / 78.77 ms │     no change │
│ QQuery 12 │      277.42 / 283.48 ±4.38 / 289.47 ms │     272.08 / 283.32 ±12.20 / 306.54 ms │     no change │
│ QQuery 13 │     388.82 / 401.73 ±10.35 / 417.94 ms │     376.16 / 390.43 ±13.09 / 409.74 ms │     no change │
│ QQuery 14 │      297.13 / 303.19 ±3.91 / 307.42 ms │      290.20 / 295.49 ±4.19 / 301.55 ms │     no change │
│ QQuery 15 │      309.13 / 315.24 ±5.06 / 321.11 ms │      278.92 / 284.92 ±6.23 / 294.11 ms │ +1.11x faster │
│ QQuery 16 │      653.32 / 665.33 ±8.43 / 676.46 ms │     643.04 / 665.07 ±13.41 / 683.93 ms │     no change │
│ QQuery 17 │      663.77 / 672.92 ±7.60 / 683.54 ms │     662.34 / 677.98 ±12.65 / 691.13 ms │     no change │
│ QQuery 18 │  1329.56 / 1369.94 ±33.23 / 1411.06 ms │  1334.51 / 1349.67 ±14.01 / 1373.64 ms │     no change │
│ QQuery 19 │         28.17 / 28.41 ±0.16 / 28.60 ms │        27.86 / 33.67 ±10.44 / 54.52 ms │  1.19x slower │
│ QQuery 20 │      522.74 / 530.34 ±7.74 / 545.19 ms │      518.52 / 526.33 ±5.36 / 531.83 ms │     no change │
│ QQuery 21 │      523.06 / 530.16 ±5.11 / 536.17 ms │     516.20 / 531.59 ±11.57 / 549.03 ms │     no change │
│ QQuery 22 │  1009.95 / 1021.41 ±12.94 / 1045.28 ms │  1030.59 / 1052.61 ±12.83 / 1067.81 ms │     no change │
│ QQuery 23 │ 3136.63 / 3314.10 ±111.38 / 3467.91 ms │ 3118.14 / 3349.54 ±160.55 / 3622.69 ms │     no change │
│ QQuery 24 │     40.98 / 139.32 ±107.33 / 324.67 ms │       40.79 / 81.17 ±79.54 / 240.25 ms │ +1.72x faster │
│ QQuery 25 │     107.59 / 130.55 ±41.11 / 212.45 ms │     106.90 / 130.40 ±45.37 / 221.12 ms │     no change │
│ QQuery 26 │         42.58 / 45.16 ±4.47 / 54.08 ms │         41.69 / 42.02 ±0.47 / 42.89 ms │ +1.07x faster │
│ QQuery 27 │      529.18 / 534.35 ±5.15 / 541.15 ms │     517.74 / 532.20 ±10.68 / 549.64 ms │     no change │
│ QQuery 28 │  2874.17 / 2899.17 ±19.03 / 2919.24 ms │  2920.48 / 2974.20 ±31.30 / 3006.42 ms │     no change │
│ QQuery 29 │         42.94 / 47.45 ±4.43 / 54.45 ms │        42.51 / 50.29 ±10.06 / 67.13 ms │  1.06x slower │
│ QQuery 30 │      317.57 / 325.08 ±4.45 / 328.87 ms │     322.48 / 343.24 ±20.46 / 379.62 ms │  1.06x slower │
│ QQuery 31 │      288.12 / 294.15 ±3.94 / 298.96 ms │      289.71 / 294.64 ±4.55 / 302.55 ms │     no change │
│ QQuery 32 │  4105.70 / 4222.49 ±79.07 / 4299.69 ms │  4075.75 / 4124.50 ±30.91 / 4162.68 ms │     no change │
│ QQuery 33 │  1576.64 / 1668.65 ±56.99 / 1753.89 ms │  1600.06 / 1694.97 ±55.94 / 1757.34 ms │     no change │
│ QQuery 34 │  1600.05 / 1700.59 ±98.35 / 1885.74 ms │  1677.87 / 1730.46 ±50.20 / 1815.99 ms │     no change │
│ QQuery 35 │     325.26 / 379.00 ±70.62 / 505.12 ms │    272.52 / 379.83 ±201.76 / 783.29 ms │     no change │
│ QQuery 36 │         68.28 / 77.30 ±7.43 / 86.89 ms │         71.72 / 76.54 ±5.13 / 84.31 ms │     no change │
│ QQuery 37 │        37.36 / 47.40 ±15.66 / 78.35 ms │         36.66 / 40.98 ±7.15 / 55.26 ms │ +1.16x faster │
│ QQuery 38 │         44.29 / 48.23 ±2.93 / 51.97 ms │        43.67 / 50.84 ±10.40 / 71.50 ms │  1.05x slower │
│ QQuery 39 │      140.17 / 153.18 ±7.78 / 161.72 ms │      154.30 / 159.80 ±4.98 / 168.19 ms │     no change │
│ QQuery 40 │         14.84 / 15.47 ±0.59 / 16.48 ms │        15.37 / 24.65 ±12.20 / 45.88 ms │  1.59x slower │
│ QQuery 41 │         14.56 / 14.78 ±0.30 / 15.36 ms │         14.43 / 14.87 ±0.39 / 15.41 ms │     no change │
│ QQuery 42 │         14.27 / 17.14 ±3.60 / 23.18 ms │         13.65 / 15.11 ±2.10 / 19.24 ms │ +1.13x faster │
└───────────┴────────────────────────────────────────┴────────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 23844.23ms │
│ Total Time (blocked-agg-poc)   │ 23794.57ms │
│ Average Time (HEAD)            │   554.52ms │
│ Average Time (blocked-agg-poc) │   553.36ms │
│ Queries Faster                 │          7 │
│ Queries Slower                 │          5 │
│ Queries with No Change         │         31 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only for benchmarks that write a results JSON (the dfbench suites, and the suites bench.sh runs through the Criterion SQL harness since #25644), and only under a memory limit: DATAFUSION_RUNTIME_MEMORY_LIMIT, or one the suite sets itself.

Base: 6924c6a (merge-base) | Changed: blocked-agg-poc

clickbench_partitioned

Query Base Changed Change
Query 0 0 B 0 B 0.0%
Query 1 104 B 104 B +0.0%
Query 2 936 B 936 B +0.0%
Query 3 312 B 312 B +0.0%
Query 4 754.7 MiB 538.9 MiB -28.6%
Query 5 1.2 GiB 1.2 GiB -3.0%
Query 6 0 B 0 B 0.0%
Query 7 60.0 MiB 60.0 MiB -0.0%
Query 8 967.0 MiB 946.6 MiB -2.1%
Query 9 593.9 MiB 594.5 MiB +0.1%
Query 10 109.1 MiB 110.1 MiB +0.9%
Query 11 117.8 MiB 130.0 MiB +10.4%
Query 12 1.4 GiB 1.3 GiB -7.4%
Query 13 1.6 GiB 1.6 GiB -0.8%
Query 14 1.3 GiB 1.2 GiB -3.0%
Query 15 1.1 GiB 682.2 MiB -40.5%
Query 16 3.4 GiB 3.1 GiB -8.4%
Query 17 3.4 GiB 3.1 GiB -9.0%
Query 18 8.1 GiB 7.8 GiB -3.6%
Query 19 0 B 0 B 0.0%
Query 20 104 B 104 B +0.0%
Query 21 3.7 MiB 4.0 MiB +8.3%
Query 22 4.0 MiB 3.5 MiB -14.3%
Query 23 2.1 GiB 2.3 GiB +8.4%
Query 24 58.9 MiB 57.7 MiB -2.1%
Query 25 170.4 MiB 162.7 MiB -4.5%
Query 26 58.9 MiB 58.0 MiB -1.5%
Query 27 2.5 MiB 2.2 MiB -9.2%
Query 28 1.8 GiB 1.7 GiB -6.4%
Query 29 624 B 624 B +0.0%
Query 30 722.1 MiB 649.2 MiB -10.1%
Query 31 1.5 GiB 1.3 GiB -7.7%
Query 32 3.6 GiB 5.3 GiB +49.6%
Query 33 7.1 GiB 7.1 GiB +0.9%
Query 34 7.3 GiB 7.1 GiB -2.6%
Query 35 598.9 MiB 505.0 MiB -15.7%
Query 36 114.7 MiB 109.2 MiB -4.9%
Query 37 7.3 MiB 7.3 MiB +0.0%
Query 38 5.2 MiB 5.6 MiB +9.1%
Query 39 296.9 MiB 288.0 MiB -3.0%
Query 40 2.0 MiB 2.0 MiB +0.5%
Query 41 3.1 MiB 3.1 MiB +0.3%
Query 42 1.6 MiB 1.6 MiB -5.5%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
clickbench_partitioned base (6924c6a (merge-base)) 8.1 GiB 18.2 GiB 10.1 GiB 2.3×
clickbench_partitioned changed (blocked-agg-poc) 7.8 GiB 17.5 GiB 9.7 GiB 2.3×
Resource Usage

clickbench_partitioned — base (merge-base)

Metric Value
Wall time 120.0s
Peak memory 18.2 GiB
Avg memory 5.9 GiB
CPU user 1176.0s
CPU sys 120.7s
Peak spill 4.2 GiB

clickbench_partitioned — branch

Metric Value
Wall time 120.0s
Peak memory 17.5 GiB
Avg memory 6.2 GiB
CPU user 1180.1s
CPU sys 119.5s
Peak spill 4.2 GiB

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-standard-32 (12 vCPU / 65 GiB)

Comparing blocked-agg-poc (51fca43) to 6924c6a (merge-base) diff

Run configuration
run benchmark h2o_medium
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "16G"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  32
On-line CPU(s) list:                     0-31
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     32
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               2 MiB (32 instances)
L1i cache:                               2 MiB (32 instances)
L2 cache:                                64 MiB (32 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-31
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark h2o.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃        HEAD ┃ blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │  1108.10 ms │      1109.01 ms │     no change │
│ QQuery 2  │  2431.03 ms │      2439.80 ms │     no change │
│ QQuery 3  │  2116.69 ms │      2117.93 ms │     no change │
│ QQuery 4  │  1408.44 ms │      1408.12 ms │     no change │
│ QQuery 5  │  1967.67 ms │      1854.29 ms │ +1.06x faster │
│ QQuery 6  │  1691.12 ms │      1677.60 ms │     no change │
│ QQuery 7  │  1975.61 ms │      1955.09 ms │     no change │
│ QQuery 8  │  3581.48 ms │      3479.01 ms │     no change │
│ QQuery 9  │  7956.83 ms │      7877.45 ms │     no change │
│ QQuery 10 │ 13681.53 ms │     14420.01 ms │  1.05x slower │
└───────────┴─────────────┴─────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 37918.50ms │
│ Total Time (blocked-agg-poc)   │ 38338.30ms │
│ Average Time (HEAD)            │  3791.85ms │
│ Average Time (blocked-agg-poc) │  3833.83ms │
│ Queries Faster                 │          1 │
│ Queries Slower                 │          1 │
│ Queries with No Change         │          8 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and blocked-agg-poc
--------------------
Benchmark h2o.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                      HEAD ┃                           blocked-agg-poc ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │     1108.10 / 1116.66 ±11.57 / 1133.02 ms │      1109.01 / 1109.47 ±0.63 / 1110.36 ms │     no change │
│ QQuery 2  │     2431.03 / 2481.76 ±37.93 / 2522.21 ms │     2439.80 / 2464.30 ±22.83 / 2494.77 ms │     no change │
│ QQuery 3  │     2116.69 / 2139.91 ±20.10 / 2165.72 ms │      2117.93 / 2127.60 ±7.41 / 2135.92 ms │     no change │
│ QQuery 4  │      1408.44 / 1412.54 ±4.01 / 1417.98 ms │      1408.12 / 1409.41 ±1.80 / 1411.95 ms │     no change │
│ QQuery 5  │     1967.67 / 2026.46 ±41.78 / 2060.96 ms │     1854.29 / 1864.35 ±13.51 / 1883.45 ms │ +1.09x faster │
│ QQuery 6  │      1691.12 / 1700.51 ±9.28 / 1713.15 ms │      1677.60 / 1686.58 ±9.35 / 1699.47 ms │     no change │
│ QQuery 7  │     1975.61 / 2006.13 ±29.79 / 2046.54 ms │     1955.09 / 1995.93 ±29.00 / 2019.55 ms │     no change │
│ QQuery 8  │     3581.48 / 3675.76 ±75.09 / 3765.24 ms │    3479.01 / 3679.55 ±255.61 / 4040.28 ms │     no change │
│ QQuery 9  │     7956.83 / 8041.27 ±82.46 / 8153.14 ms │     7877.45 / 7919.47 ±41.15 / 7975.35 ms │     no change │
│ QQuery 10 │ 13681.53 / 14075.23 ±280.41 / 14313.21 ms │ 14420.01 / 14570.33 ±115.92 / 14702.15 ms │     no change │
└───────────┴───────────────────────────────────────────┴───────────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary              ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)              │ 38676.25ms │
│ Total Time (blocked-agg-poc)   │ 38826.99ms │
│ Average Time (HEAD)            │  3867.62ms │
│ Average Time (blocked-agg-poc) │  3882.70ms │
│ Queries Faster                 │          1 │
│ Queries Slower                 │          0 │
│ Queries with No Change         │          9 │
│ Queries with Failure           │          0 │
└────────────────────────────────┴────────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only for benchmarks that write a results JSON (the dfbench suites, and the suites bench.sh runs through the Criterion SQL harness since #25644), and only under a memory limit: DATAFUSION_RUNTIME_MEMORY_LIMIT, or one the suite sets itself.

Base: 6924c6a (merge-base) | Changed: blocked-agg-poc

h2o_medium — h2o

Query Base Changed Change
Query 1 1.9 MiB 1.9 MiB +0.0%
Query 2 4.2 GiB 4.2 GiB +0.0%
Query 3 141.1 MiB 134.5 MiB -4.7%
Query 4 4.4 MiB 3.6 MiB -19.9%
Query 5 98.2 MiB 96.9 MiB -1.3%
Query 6 1.0 GiB 1.0 GiB +0.1%
Query 7 117.2 MiB 126.2 MiB +7.7%
Query 8 3.9 GiB 3.9 GiB +0.0%
Query 9 4.7 GiB 4.7 GiB -0.0%
Query 10 4.8 GiB 4.7 GiB -1.1%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
h2o_medium base (6924c6a (merge-base)) 4.8 GiB 12.9 GiB 8.1 GiB 2.7×
h2o_medium changed (blocked-agg-poc) 4.7 GiB 12.4 GiB 7.7 GiB 2.6×
Resource Usage

h2o_medium — base (merge-base)

Metric Value
Wall time 120.0s
Peak memory 12.9 GiB
Avg memory 3.8 GiB
CPU user 1231.2s
CPU sys 101.2s
Peak spill 7.3 GiB

h2o_medium — branch

Metric Value
Wall time 120.0s
Peak memory 12.4 GiB
Avg memory 4.2 GiB
CPU user 1238.6s
CPU sys 104.1s
Peak spill 7.3 GiB

File an issue against this benchmark runner

@rluvaton

rluvaton commented Sep 30, 2026 •

Copy link
Copy Markdown
Member Author

@jayzhan211

Thanks @rluvaton , I left some suggestions

  • Test for multi-block spilling. The table tests use MIN_BLOCK_SIZE + 10 groups but never spill, and the existing spill tests stay under 2^18 groups per table, so sort_and_spill_batches never sees more than one batch. A SingleHashAggregateStream (or Final) test with a memory limit and more than MIN_BLOCK_SIZE groups, compared against an unlimited run, would cover it.

changed to block size without 2^18, so it is now covered

  • Emit modes that can't be reached yet. materialize_batches is only called with EmitTo::All (common.rs:348, common.rs:416), so BlockedEmitTo::NextBlock/First, BlockedVec::take_first, the map-renumbering arms in blocked/primitive.rs:174-215, and FIXED_BLOCK_SIZE are unused. Would you consider moving them to the ordered-stream PR that needs them? It would make this one easier to review. Fine to leave if you prefer.

FIXED_BLOCK_SIZE is for nested/dictionary support so I won't have to introduce breaking changes

first and next block will be used by ordered but also outside of datafusion so they are valuable from the start, and we can avoid adding breaking change later

@jayzhan211 jayzhan211 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I left some comments and also the test failed


// Make sure we can hold on the hash tables and all the batches that need to be emitted (except the first one)
// if we can't hold it than we can't do anything about it.
self.reservation.try_resize(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

try_resize(hash_table.memory_size() + pending_memory)? dropped the fallback main had: when holding the materialized state doesn't fit, main reserved only the table and emitted unreserved. With blocked storage (a blocked count is enough, even with flat keys), partial aggregation now fails with ResourcesExhausted. 4 existing tests fail here and pass on the merge-base 6924c6a: both aggregate_grouping_sets_*_with_spill (Failed to allocate additional 872.0 B ... pool_size: 500.0 B), test_partial_hash_stream_emits_whole_batch_when_held_batch_does_not_fit, and test_partial_hash_stream_accounts_held_batch_on_memory_pressure_while_slicing.

Fix (with it, both grouping-sets tests pass locally):

-                        self.reservation.try_resize(
-                            hash_table.memory_size()
-                              // The batches that need to be held while emitting each one
-                              + pending_memory,
-                        )?;
+                        let hold_pending = match self
+                            .reservation
+                            .try_resize(hash_table.memory_size() + pending_memory)
+                        {
+                            Ok(()) => true,
+                            // Can't hold the pending blocks: emit them unreserved, like
+                            // the flat path emits the whole batch
+                            Err(DataFusionError::ResourcesExhausted(_)) => {
+                                self.reservation.try_resize(hash_table.memory_size())?;
+                                false
+                            }
+                            Err(e) => return Err(e),
+                        };

                         for (i, batch) in
                             materialized_group_states.into_iter().enumerate()
                         {
-                            if i != 0 {
+                            if i != 0 && hold_pending {
                                 self.reservation.try_shrink(batch.memory_size)?;
                             }

The two hash_stream tests also assume a single state batch (first.num_rows() == num_groups, and a shared data_ptr between the first two outputs). Please update them for one batch per block, e.g. assert that the total rows match and that reserved covers the blocks not emitted yet.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But you can't do that, this is "cheating",
you basically say, ok, the memory did not allow us to reserve memory for the batches we are holding between emits so we just don't count them.

this is just ignoring the problem, not a fix.

Main could do what it did since it only had 1 batch, so it just emitted it without holding on it, but now we have to hold it since we can't emit multiple batches at once

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed that the held blocks must be counted, and my earlier suggestion (emit without reserving) didn't do that. But the error doesn't count them either: at this point they were never reserved. Logged in the failing tests: aggregate_grouping_sets_with_yielding_with_spill has 0 B reserved while the blocks hold 872 B (pool 500 B), and ..._does_not_fit has 3.5 KB reserved for 4 × 12.5 KB blocks, so ? frees nothing and fails a query main finishes. MemoryReservation::grow is infallible, so you can count them honestly, over the limit, until they drain:

-                        self.reservation.try_resize(
-                            hash_table.memory_size()
-                              // The batches that need to be held while emitting each one
-                              + pending_memory,
-                        )?;
+                        let held = hash_table.memory_size() + pending_memory;
+                        match self.reservation.try_resize(held) {
+                            Ok(()) => {}
+                            // The blocks already exist, so failing would not free
+                            // them: count them anyway, over the pool limit, until
+                            // they are emitted
+                            Err(DataFusionError::ResourcesExhausted(_)) => {
+                                self.reservation.grow(held - self.reservation.size())
+                            }
+                            Err(e) => return Err(e),
+                        }

With this both aggregate_grouping_sets_*_with_spill pass. The two test_partial_hash_stream_* failures assume one flat state batch (as their own comment anticipates): a Utf8 key + sum in partial_stream_under_memory_limit keeps the slicing coverage, plus an Int32 + count variant asserting no error at 32 KB.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Only this issue is left!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

grow() should really not be used anywhere, you are suggesting adding a panic here, which is much more severe than returning an error for the query - which is also bad in my opinion, but preferable to exceeding the expected memory, which can lead to SIGKILLs that are far harder to debug.
Erroring out if the memory is umavailable is the way forward in my opinion.
The next step is to ensure this situation never happens if we are able to hold 2 batches in memory

// Add one to each group's counter for each non null, non filtered value
// SAFETY: group_index is guaranteed to be in bounds and less than total_num_groups
unsafe {
self.counts.update_unchecked(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

update_batch calls update_unchecked with only a SAFETY comment about the caller's contract. That means a safe public path (create_blocked_groups_accumulator → update_batch) can write out of bounds. Passing BlocksIndex::new(0, 1) with total_num_groups = 1 aborts in debug with unsafe precondition(s) violated: slice::get_unchecked_mut at blocked_vec.rs:282. merge_batch already goes through the checked update_with, and all_in_bounds is one vectorized pass, so the checked path should cost little here too. The flat count on main has the same pattern, so this is fine to handle in a follow-up if you'd rather.

-        self.counts.grow_to(total_num_groups, 0);
-
-        // Add one to each group's counter for each non null, non filtered value
-        // SAFETY: group_index is guaranteed to be in bounds and less than total_num_groups
-        unsafe {
-            self.counts.update_unchecked(
-                group_indices,
-                nulls.as_ref(),
-                opt_filter,
-                |count| *count += 1,
-            );
-        }
+        // Add one to each group's counter for each non null, non filtered value
+        self.counts.update(
+            total_num_groups,
+            0,
+            group_indices,
+            nulls.as_ref(),
+            opt_filter,
+            |count| *count += 1,
+        );

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I had unchecked here to keep the same optimizations used as before so this can be as a follow up pr but not in this

self.null_group = None;
blocks
}
BlockedEmitTo::NextBlock => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

NextBlock runs map.retain over every entry to renumber it, so draining n groups one block at a time costs O(n²/block_size). With 10M groups and 8192-row blocks that's about 1.2k passes over the full map. Incremental emit is what NextBlock is for, so this will show up once ordered or external callers use it. Nothing calls it yet, so a follow-up is fine.

A suggestion: store absolute block numbers in the map plus a first_block offset. intern subtracts the offset when it returns an index and adds it when it inserts one. NextBlock then removes only the emitted block's keys by lookup, at O(block_size) per block:

BlockedEmitTo::NextBlock => {
    let Some(block) = self.values.take_next_block() else {
        return Ok(vec![]);
    };
    let first_block = self.first_block;
    let state = &self.random_state;
    for &key in &block {
        // The null group's slot holds `Default`, so match on the block too
        if let Ok(entry) = self.map.find_entry(key.hash(state), |&(g, k)| {
            g.block_index() == first_block && k.is_eq(key)
        }) {
            entry.remove();
        }
    }
    self.first_block += 1;
    let len = block.len();
    let array = self.build_block(block, 0);
    self.shift_null_group(len);
    vec![vec![array]]
}

All and clear_shrink reset first_block to 0. First(n) would also need to account for the offset.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the shifting every entry exists in the non blocked implementation (which is why I did that), so to avoid a lot of changes I will keep that for now and this can be in a later PR

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But you are right, just not in this PR

# Conflicts:
#	datafusion/physical-plan/src/aggregates/single_stream.rs
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

functions Changes to functions implementation logical-expr Logical plan and expressions physical-expr Changes to the physical-expr crates physical-plan Changes to the physical-plan crate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants