Skip to content

perf: unroll PrimitiveArray hashing for the no-null path - #26191

Draft
Rich-T-kid wants to merge 1 commit into
apache:mainfrom
Rich-T-kid:rich-T-kid/primitive-hash-no-null-unroll
Draft

Rich-T-kid wants to merge 1 commit into
apache:mainfrom
Rich-T-kid:rich-T-kid/primitive-hash-no-null-unroll

Conversation

@Rich-T-kid

Copy link
Copy Markdown
Contributor

Splits the single scalar loop in hash_array_primitive's no-null branch into two const-generic-free helpers that batch independent hashes so the pipeline isn't serialized on store-to-load deps:

  • hash_prim_fresh_dense: overwrite mode, 16 rows per iteration.
  • hash_prim_rehash_dense: fold-with-prev mode, 8 rows per iteration (narrower because each lane holds both prev and value live).

Both use get_unchecked / get_unchecked_mut with proven invariants (hashes.len() == values.len()) and a scalar tail for the remainder.

Benchmarks (int64, 8192 rows, no nulls) on Apple M4 Max:
single, no nulls : neutral to small regression (<+7%)
multiple, no nulls : ~-14%

The nullable path is unchanged -- see PR #26143 for the sibling change.

Which issue does this PR close?

  • Closes #.

Rationale for this change

What changes are included in this PR?

What is the testing strategy for this PR?

Are there any user-facing changes?

@github-actions github-actions Bot added the common Related to common crate label Oct 10, 2026
Splits the single scalar loop in `hash_array_primitive`'s no-null branch
into two const-generic-free helpers that batch independent hashes so the
pipeline isn't serialized on store-to-load deps:

- `hash_prim_fresh_dense`: overwrite mode, 16 rows per iteration.
- `hash_prim_rehash_dense`: fold-with-prev mode, 8 rows per iteration
  (narrower because each lane holds both `prev` and `value` live).

Both use `get_unchecked` / `get_unchecked_mut` with proven invariants
(hashes.len() == values.len()) and a scalar tail for the remainder.

Benchmarks (int64, 8192 rows, no nulls) on Apple M4 Max:
    single, no nulls   : neutral to small regression (<+7%)
    multiple, no nulls : ~-14%

The nullable path is unchanged -- see PR apache#26143 for the sibling change.
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark with_hashes

@Rich-T-kid
Rich-T-kid force-pushed the rich-T-kid/primitive-hash-no-null-unroll branch from 5027475 to 4d6ff20 Compare October 10, 2026 19:03
@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c6101105285-3271-psgvz 6.12.94+ #1 SMP Fri Aug 21 08:00:16 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing rich-T-kid/primitive-hash-no-null-unroll (4d6ff20) to 1c49b7f (merge-base) diff

Run configuration
run benchmark with_hashes

Results will be posted here when complete


File an issue against this benchmark runner

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark with_hashes

@codecov-commenter

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 82.79%. Comparing base (1c49b7f) to head (4d6ff20).
⚠️ Report is 3 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main   #26191      +/-   ##
==========================================
+ Coverage   82.78%   82.79%   +0.01%     
==========================================
  Files        1148     1149       +1     
  Lines      451049   451695     +646     
  Branches   451049   451695     +646     
==========================================
+ Hits       373381   373983     +602     
- Misses      54928    54934       +6     
- Partials    22740    22778      +38     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing rich-T-kid/primitive-hash-no-null-unroll (4d6ff20) to 1c49b7f (merge-base) diff

Run configuration
run benchmark with_hashes
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                         HEAD                                   rich-T-kid_primitive-hash-no-null-unroll
-----                                         ----                                   ----------------------------------------
dense_union: multiple, no nulls               1.06     44.0±0.06µs        ? ?/sec    1.00     41.6±0.07µs        ? ?/sec
dense_union: single, no nulls                 1.06     15.0±0.02µs        ? ?/sec    1.00     14.2±0.03µs        ? ?/sec
dense_union_sliced: 1/10 of 81920 rows        1.04     22.7±0.02µs        ? ?/sec    1.00     21.8±0.04µs        ? ?/sec
dense_union_sliced: 1/2 of 16384 rows         1.04     22.7±0.03µs        ? ?/sec    1.00     21.7±0.03µs        ? ?/sec
dense_union_sliced: 1/5 of 40960 rows         1.04     22.7±0.03µs        ? ?/sec    1.00     21.9±0.03µs        ? ?/sec
dictionary_utf8_int32: multiple, no nulls     1.05     14.8±0.01µs        ? ?/sec    1.00     14.1±0.12µs        ? ?/sec
dictionary_utf8_int32: multiple, nulls        1.00     29.1±0.06µs        ? ?/sec    1.00     29.2±0.02µs        ? ?/sec
dictionary_utf8_int32: single, no nulls       1.18      4.7±0.06µs        ? ?/sec    1.00      4.0±0.17µs        ? ?/sec
dictionary_utf8_int32: single, nulls          1.00      9.5±0.02µs        ? ?/sec    1.02      9.7±0.01µs        ? ?/sec
fixed_size_list_array: multiple, no nulls     1.04    111.6±0.56µs        ? ?/sec    1.00    107.0±0.82µs        ? ?/sec
fixed_size_list_array: multiple, nulls        1.11    123.7±6.04µs        ? ?/sec    1.00    111.9±0.27µs        ? ?/sec
fixed_size_list_array: single, no nulls       1.12     37.3±0.96µs        ? ?/sec    1.00     33.4±0.16µs        ? ?/sec
fixed_size_list_array: single, nulls          1.08     40.3±0.07µs        ? ?/sec    1.00     37.3±0.09µs        ? ?/sec
int64: multiple, no nulls                     1.44     17.0±0.01µs        ? ?/sec    1.00     11.8±0.01µs        ? ?/sec
int64: multiple, nulls                        1.00     26.7±0.02µs        ? ?/sec    1.00     26.8±0.02µs        ? ?/sec
int64: single, no nulls                       1.21      4.6±0.01µs        ? ?/sec    1.00      3.8±0.01µs        ? ?/sec
int64: single, nulls                          1.00      8.8±0.02µs        ? ?/sec    1.00      8.7±0.01µs        ? ?/sec
large_utf8: multiple, no nulls                1.00    119.6±2.36µs        ? ?/sec    1.01    120.9±0.98µs        ? ?/sec
large_utf8: multiple, nulls                   1.00    145.3±0.94µs        ? ?/sec    1.00    144.9±0.77µs        ? ?/sec
large_utf8: single, no nulls                  1.02     38.9±0.54µs        ? ?/sec    1.00     38.2±0.33µs        ? ?/sec
large_utf8: single, nulls                     1.00     34.8±0.25µs        ? ?/sec    1.00     34.7±0.24µs        ? ?/sec
list_array: multiple, no nulls                1.02    140.6±0.64µs        ? ?/sec    1.00    137.9±0.41µs        ? ?/sec
list_array: multiple, nulls                   1.06    169.2±0.70µs        ? ?/sec    1.00    159.0±1.16µs        ? ?/sec
list_array: single, no nulls                  1.08     47.8±0.32µs        ? ?/sec    1.00     44.1±0.10µs        ? ?/sec
list_array: single, nulls                     1.07     57.0±0.20µs        ? ?/sec    1.00     53.2±0.12µs        ? ?/sec
list_array_sliced: 1/10 of 81920 rows         1.11     48.7±1.26µs        ? ?/sec    1.00     43.8±0.20µs        ? ?/sec
list_array_sliced: 1/2 of 16384 rows          1.13     49.3±1.04µs        ? ?/sec    1.00     43.5±0.14µs        ? ?/sec
list_array_sliced: 1/5 of 40960 rows          1.08     47.4±0.08µs        ? ?/sec    1.00     44.0±0.17µs        ? ?/sec
list_view_array: multiple, no nulls           1.03    137.6±4.05µs        ? ?/sec    1.00    133.7±1.18µs        ? ?/sec
list_view_array: multiple, nulls              1.14    152.3±1.53µs        ? ?/sec    1.00    133.1±0.40µs        ? ?/sec
list_view_array: single, no nulls             1.08     45.6±0.21µs        ? ?/sec    1.00     42.2±0.44µs        ? ?/sec
list_view_array: single, nulls                1.10     48.8±0.11µs        ? ?/sec    1.00     44.5±0.17µs        ? ?/sec
map_array: multiple, no nulls                 1.24    244.7±1.84µs        ? ?/sec    1.00    196.9±1.17µs        ? ?/sec
map_array: multiple, nulls                    1.23    251.6±0.77µs        ? ?/sec    1.00    205.1±1.00µs        ? ?/sec
map_array: single, no nulls                   1.23     81.0±0.41µs        ? ?/sec    1.00     66.1±0.25µs        ? ?/sec
map_array: single, nulls                      1.21     83.8±0.28µs        ? ?/sec    1.00     69.0±0.21µs        ? ?/sec
map_array_sliced: 1/10 of 81920 rows          1.20     81.2±0.37µs        ? ?/sec    1.00     67.6±0.24µs        ? ?/sec
map_array_sliced: 1/2 of 16384 rows           1.21     81.5±0.46µs        ? ?/sec    1.00     67.4±0.22µs        ? ?/sec
map_array_sliced: 1/5 of 40960 rows           1.22     81.6±0.34µs        ? ?/sec    1.00     66.8±0.25µs        ? ?/sec
mixed: 3 columns (int64, utf8, utf8_view)     1.01     82.5±0.32µs        ? ?/sec    1.00     81.7±0.28µs        ? ?/sec
null: multiple, no nulls                      1.00      2.5±0.00µs        ? ?/sec    1.00      2.5±0.00µs        ? ?/sec
null: single, no nulls                        1.00   1108.2±2.14ns        ? ?/sec    1.00   1107.4±2.76ns        ? ?/sec
run_array_int32: multiple, no nulls           1.03      5.1±0.01µs        ? ?/sec    1.00      4.9±0.01µs        ? ?/sec
run_array_int32: multiple, nulls              1.00      6.8±0.01µs        ? ?/sec    1.01      6.9±0.01µs        ? ?/sec
run_array_int32: single, no nulls             1.05   1899.0±3.22ns        ? ?/sec    1.00   1809.7±3.02ns        ? ?/sec
run_array_int32: single, nulls                1.00      2.4±0.01µs        ? ?/sec    1.04      2.5±0.00µs        ? ?/sec
sparse_union (2 types): multiple, no nulls    1.00    102.4±0.36µs        ? ?/sec    1.16    119.2±1.26µs        ? ?/sec
sparse_union (2 types): single, no nulls      1.03     35.0±0.82µs        ? ?/sec    1.00     33.9±0.16µs        ? ?/sec
sparse_union (utf8): multiple, no nulls       1.01    241.1±2.24µs        ? ?/sec    1.00    239.6±1.15µs        ? ?/sec
sparse_union (utf8): single, no nulls         1.00     79.8±0.46µs        ? ?/sec    1.01     80.9±0.47µs        ? ?/sec
sparse_union: multiple, no nulls              1.00     98.6±0.77µs        ? ?/sec    1.02    100.7±4.71µs        ? ?/sec
sparse_union: single, no nulls                1.01     33.4±0.29µs        ? ?/sec    1.00     33.0±0.16µs        ? ?/sec
sparse_union_sliced: 1/10 of 81920 rows       1.03     33.9±1.36µs        ? ?/sec    1.00     33.0±0.20µs        ? ?/sec
sparse_union_sliced: 1/2 of 16384 rows        1.03     33.6±0.18µs        ? ?/sec    1.00     32.6±0.20µs        ? ?/sec
sparse_union_sliced: 1/5 of 40960 rows        1.02     33.5±0.82µs        ? ?/sec    1.00     32.7±0.34µs        ? ?/sec
struct_array: multiple, no nulls              1.06    161.8±0.53µs        ? ?/sec    1.00    152.4±0.92µs        ? ?/sec
struct_array: multiple, nulls                 1.08    181.4±0.62µs        ? ?/sec    1.00    167.4±0.82µs        ? ?/sec
struct_array: single, no nulls                1.09     55.3±0.51µs        ? ?/sec    1.00     50.8±0.27µs        ? ?/sec
struct_array: single, nulls                   1.08     60.3±0.20µs        ? ?/sec    1.00     56.0±0.27µs        ? ?/sec
utf8: multiple, no nulls                      1.00    115.1±0.41µs        ? ?/sec    1.00    115.4±0.62µs        ? ?/sec
utf8: multiple, nulls                         1.00    161.5±0.91µs        ? ?/sec    1.00    161.3±0.90µs        ? ?/sec
utf8: single, no nulls                        1.01     30.0±0.28µs        ? ?/sec    1.00     29.6±0.26µs        ? ?/sec
utf8: single, nulls                           1.01     45.4±1.08µs        ? ?/sec    1.00     45.1±0.89µs        ? ?/sec
utf8_view (small): multiple, no nulls         1.00     19.1±0.04µs        ? ?/sec    1.00     19.1±0.02µs        ? ?/sec
utf8_view (small): multiple, nulls            1.00     29.2±0.05µs        ? ?/sec    1.00     29.3±0.02µs        ? ?/sec
utf8_view (small): single, no nulls           1.00      5.1±0.01µs        ? ?/sec    1.00      5.0±0.00µs        ? ?/sec
utf8_view (small): single, nulls              1.00      9.7±0.03µs        ? ?/sec    1.01      9.8±0.01µs        ? ?/sec
utf8_view: multiple, no nulls                 1.04    102.5±0.41µs        ? ?/sec    1.00     98.4±0.69µs        ? ?/sec
utf8_view: multiple, nulls                    1.15    124.0±0.48µs        ? ?/sec    1.00    107.8±0.26µs        ? ?/sec
utf8_view: single, no nulls                   1.16     35.5±0.65µs        ? ?/sec    1.00     30.6±0.25µs        ? ?/sec
utf8_view: single, nulls                      1.12     36.7±0.05µs        ? ?/sec    1.00     32.8±0.14µs        ? ?/sec

Resource Usage

with_hashes — base (merge-base)

Metric Value
Wall time 755.2s
Peak memory 43.3 MiB
Avg memory 36.9 MiB
CPU user 885.5s
CPU sys 0.8s
Peak spill 0 B

with_hashes — branch

Metric Value
Wall time 770.2s
Peak memory 42.1 MiB
Avg memory 34.0 MiB
CPU user 882.2s
CPU sys 0.8s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c6101232844-3273-n95rb 6.12.94+ #1 SMP Fri Aug 21 08:00:16 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing rich-T-kid/primitive-hash-no-null-unroll (4d6ff20) to 1c49b7f (merge-base) diff

Run configuration
run benchmark with_hashes

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing rich-T-kid/primitive-hash-no-null-unroll (4d6ff20) to 1c49b7f (merge-base) diff

Run configuration
run benchmark with_hashes
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                         HEAD                                   rich-T-kid_primitive-hash-no-null-unroll
-----                                         ----                                   ----------------------------------------
dense_union: multiple, no nulls               1.05     43.9±0.09µs        ? ?/sec    1.00     42.0±0.13µs        ? ?/sec
dense_union: single, no nulls                 1.05     15.0±0.04µs        ? ?/sec    1.00     14.3±0.07µs        ? ?/sec
dense_union_sliced: 1/10 of 81920 rows        1.04     22.7±0.03µs        ? ?/sec    1.00     21.8±0.03µs        ? ?/sec
dense_union_sliced: 1/2 of 16384 rows         1.03     22.5±0.02µs        ? ?/sec    1.00     21.8±0.03µs        ? ?/sec
dense_union_sliced: 1/5 of 40960 rows         1.04     22.6±0.06µs        ? ?/sec    1.00     21.8±0.03µs        ? ?/sec
dictionary_utf8_int32: multiple, no nulls     1.06     14.8±0.01µs        ? ?/sec    1.00     14.0±0.03µs        ? ?/sec
dictionary_utf8_int32: multiple, nulls        1.00     28.9±0.04µs        ? ?/sec    1.01     29.2±0.05µs        ? ?/sec
dictionary_utf8_int32: single, no nulls       1.19      4.6±0.01µs        ? ?/sec    1.00      3.9±0.01µs        ? ?/sec
dictionary_utf8_int32: single, nulls          1.00      9.5±0.01µs        ? ?/sec    1.02      9.6±0.01µs        ? ?/sec
fixed_size_list_array: multiple, no nulls     1.06    112.4±0.45µs        ? ?/sec    1.00    105.8±1.68µs        ? ?/sec
fixed_size_list_array: multiple, nulls        1.09    120.4±0.36µs        ? ?/sec    1.00    110.6±0.26µs        ? ?/sec
fixed_size_list_array: single, no nulls       1.09     36.5±0.11µs        ? ?/sec    1.00     33.6±0.07µs        ? ?/sec
fixed_size_list_array: single, nulls          1.08     40.2±0.10µs        ? ?/sec    1.00     37.2±0.07µs        ? ?/sec
int64: multiple, no nulls                     1.43     17.0±0.01µs        ? ?/sec    1.00     11.8±0.02µs        ? ?/sec
int64: multiple, nulls                        1.00     26.8±0.03µs        ? ?/sec    1.00     26.8±0.02µs        ? ?/sec
int64: single, no nulls                       1.20      4.6±0.01µs        ? ?/sec    1.00      3.8±0.01µs        ? ?/sec
int64: single, nulls                          1.01      8.8±0.03µs        ? ?/sec    1.00      8.8±0.01µs        ? ?/sec
large_utf8: multiple, no nulls                1.00    120.0±2.73µs        ? ?/sec    1.00    120.2±1.52µs        ? ?/sec
large_utf8: multiple, nulls                   1.00    144.5±0.67µs        ? ?/sec    1.01    145.3±0.67µs        ? ?/sec
large_utf8: single, no nulls                  1.00     37.9±0.18µs        ? ?/sec    1.02     38.8±0.60µs        ? ?/sec
large_utf8: single, nulls                     1.00     34.8±0.22µs        ? ?/sec    1.00     34.8±0.30µs        ? ?/sec
list_array: multiple, no nulls                1.04    141.2±0.71µs        ? ?/sec    1.00    136.0±0.68µs        ? ?/sec
list_array: multiple, nulls                   1.07    169.5±0.64µs        ? ?/sec    1.00    158.6±0.32µs        ? ?/sec
list_array: single, no nulls                  1.02     47.3±0.22µs        ? ?/sec    1.00     46.5±0.15µs        ? ?/sec
list_array: single, nulls                     1.07     57.0±0.15µs        ? ?/sec    1.00     53.2±0.14µs        ? ?/sec
list_array_sliced: 1/10 of 81920 rows         1.09     47.8±0.19µs        ? ?/sec    1.00     43.8±0.18µs        ? ?/sec
list_array_sliced: 1/2 of 16384 rows          1.09     47.8±0.22µs        ? ?/sec    1.00     44.0±0.11µs        ? ?/sec
list_array_sliced: 1/5 of 40960 rows          1.10     48.1±0.15µs        ? ?/sec    1.00     43.6±0.10µs        ? ?/sec
list_view_array: multiple, no nulls           1.10    136.7±0.40µs        ? ?/sec    1.00    123.9±1.43µs        ? ?/sec
list_view_array: multiple, nulls              1.05    147.4±0.38µs        ? ?/sec    1.00    139.8±0.46µs        ? ?/sec
list_view_array: single, no nulls             1.09     45.8±0.13µs        ? ?/sec    1.00     42.0±0.12µs        ? ?/sec
list_view_array: single, nulls                1.07     48.4±0.13µs        ? ?/sec    1.00     45.3±0.92µs        ? ?/sec
map_array: multiple, no nulls                 1.23    243.7±0.72µs        ? ?/sec    1.00    198.4±0.77µs        ? ?/sec
map_array: multiple, nulls                    1.22    249.7±1.39µs        ? ?/sec    1.00    205.1±0.51µs        ? ?/sec
map_array: single, no nulls                   1.23     81.3±0.32µs        ? ?/sec    1.00     66.2±0.26µs        ? ?/sec
map_array: single, nulls                      1.21     84.1±0.37µs        ? ?/sec    1.00     69.6±0.18µs        ? ?/sec
map_array_sliced: 1/10 of 81920 rows          1.21     81.9±0.27µs        ? ?/sec    1.00     67.6±0.23µs        ? ?/sec
map_array_sliced: 1/2 of 16384 rows           1.23     81.6±0.34µs        ? ?/sec    1.00     66.4±0.23µs        ? ?/sec
map_array_sliced: 1/5 of 40960 rows           1.23     81.6±0.48µs        ? ?/sec    1.00     66.4±0.40µs        ? ?/sec
mixed: 3 columns (int64, utf8, utf8_view)     1.05     82.8±0.57µs        ? ?/sec    1.00     79.2±1.00µs        ? ?/sec
null: multiple, no nulls                      1.01      2.5±0.01µs        ? ?/sec    1.00      2.5±0.00µs        ? ?/sec
null: single, no nulls                        1.00   1109.5±2.15ns        ? ?/sec    1.00   1109.6±2.50ns        ? ?/sec
run_array_int32: multiple, no nulls           1.03      5.1±0.01µs        ? ?/sec    1.00      4.9±0.01µs        ? ?/sec
run_array_int32: multiple, nulls              1.00      6.8±0.01µs        ? ?/sec    1.01      6.9±0.01µs        ? ?/sec
run_array_int32: single, no nulls             1.05   1899.8±2.21ns        ? ?/sec    1.00   1809.9±3.12ns        ? ?/sec
run_array_int32: single, nulls                1.00      2.4±0.00µs        ? ?/sec    1.04      2.5±0.00µs        ? ?/sec
sparse_union (2 types): multiple, no nulls    1.11    111.7±8.52µs        ? ?/sec    1.00    100.6±0.56µs        ? ?/sec
sparse_union (2 types): single, no nulls      1.02     34.6±0.14µs        ? ?/sec    1.00     33.9±0.15µs        ? ?/sec
sparse_union (utf8): multiple, no nulls       1.00    238.3±1.33µs        ? ?/sec    1.00    239.2±1.52µs        ? ?/sec
sparse_union (utf8): single, no nulls         1.00     79.6±0.36µs        ? ?/sec    1.01     80.5±0.64µs        ? ?/sec
sparse_union: multiple, no nulls              1.00     99.1±0.64µs        ? ?/sec    1.04    102.8±5.97µs        ? ?/sec
sparse_union: single, no nulls                1.02     33.6±0.57µs        ? ?/sec    1.00     32.8±0.18µs        ? ?/sec
sparse_union_sliced: 1/10 of 81920 rows       1.14     37.2±0.18µs        ? ?/sec    1.00     32.7±0.20µs        ? ?/sec
sparse_union_sliced: 1/2 of 16384 rows        1.00     33.2±0.19µs        ? ?/sec    1.01     33.5±0.82µs        ? ?/sec
sparse_union_sliced: 1/5 of 40960 rows        1.01     33.2±0.29µs        ? ?/sec    1.00     32.8±0.26µs        ? ?/sec
struct_array: multiple, no nulls              1.09    164.1±0.46µs        ? ?/sec    1.00    150.8±0.63µs        ? ?/sec
struct_array: multiple, nulls                 1.08    179.4±0.52µs        ? ?/sec    1.00    166.5±1.05µs        ? ?/sec
struct_array: single, no nulls                1.07     55.1±0.11µs        ? ?/sec    1.00     51.4±0.23µs        ? ?/sec
struct_array: single, nulls                   1.08     60.5±0.22µs        ? ?/sec    1.00     56.0±0.29µs        ? ?/sec
utf8: multiple, no nulls                      1.00    115.0±0.35µs        ? ?/sec    1.00    115.5±0.49µs        ? ?/sec
utf8: multiple, nulls                         1.00    161.6±1.23µs        ? ?/sec    1.00    161.8±1.18µs        ? ?/sec
utf8: single, no nulls                        1.00     29.8±0.21µs        ? ?/sec    1.01     30.1±0.38µs        ? ?/sec
utf8: single, nulls                           1.00     45.3±0.92µs        ? ?/sec    1.00     45.2±0.76µs        ? ?/sec
utf8_view (small): multiple, no nulls         1.00     19.1±0.03µs        ? ?/sec    1.02     19.4±0.01µs        ? ?/sec
utf8_view (small): multiple, nulls            1.00     28.9±0.04µs        ? ?/sec    1.01     29.4±0.07µs        ? ?/sec
utf8_view (small): single, no nulls           1.00      5.1±0.01µs        ? ?/sec    1.00      5.1±0.01µs        ? ?/sec
utf8_view (small): single, nulls              1.00      9.7±0.01µs        ? ?/sec    1.00      9.8±0.01µs        ? ?/sec
utf8_view: multiple, no nulls                 1.06    103.0±0.91µs        ? ?/sec    1.00     97.6±0.15µs        ? ?/sec
utf8_view: multiple, nulls                    1.15    124.4±0.38µs        ? ?/sec    1.00    108.3±0.23µs        ? ?/sec
utf8_view: single, no nulls                   1.15     35.1±0.14µs        ? ?/sec    1.00     30.5±0.07µs        ? ?/sec
utf8_view: single, nulls                      1.12     36.8±0.15µs        ? ?/sec    1.00     32.7±0.11µs        ? ?/sec

Resource Usage

with_hashes — base (merge-base)

Metric Value
Wall time 745.2s
Peak memory 43.0 MiB
Avg memory 36.7 MiB
CPU user 880.3s
CPU sys 0.8s
Peak spill 0 B

with_hashes — branch

Metric Value
Wall time 735.2s
Peak memory 42.3 MiB
Avg memory 36.2 MiB
CPU user 879.3s
CPU sys 0.8s
Peak spill 0 B

File an issue against this benchmark runner

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

common Related to common crate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants