Repository navigation
perf: unroll PrimitiveArray hashing for the no-null path - #26191
Rich-T-kid wants to merge 1 commit into
Conversation
Splits the single scalar loop in `hash_array_primitive`'s no-null branch
into two const-generic-free helpers that batch independent hashes so the
pipeline isn't serialized on store-to-load deps:
- `hash_prim_fresh_dense`: overwrite mode, 16 rows per iteration.
- `hash_prim_rehash_dense`: fold-with-prev mode, 8 rows per iteration
(narrower because each lane holds both `prev` and `value` live).
Both use `get_unchecked` / `get_unchecked_mut` with proven invariants
(hashes.len() == values.len()) and a scalar tail for the remainder.
Benchmarks (int64, 8192 rows, no nulls) on Apple M4 Max:
single, no nulls : neutral to small regression (<+7%)
multiple, no nulls : ~-14%
The nullable path is unchanged -- see PR apache#26143 for the sibling change.
|
run benchmark with_hashes |
5027475 to
4d6ff20
Compare
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing rich-T-kid/primitive-hash-no-null-unroll (4d6ff20) to 1c49b7f (merge-base) diff Run configurationrun benchmark with_hashesResults will be posted here when complete File an issue against this benchmark runner |
|
run benchmark with_hashes |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #26191 +/- ##
==========================================
+ Coverage 82.78% 82.79% +0.01%
==========================================
Files 1148 1149 +1
Lines 451049 451695 +646
Branches 451049 451695 +646
==========================================
+ Hits 373381 373983 +602
- Misses 54928 54934 +6
- Partials 22740 22778 +38 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing rich-T-kid/primitive-hash-no-null-unroll (4d6ff20) to 1c49b7f (merge-base) diff Run configurationrun benchmark with_hashesCPU Details (lscpu)Details
Resource Usagewith_hashes — base (merge-base)
with_hashes — branch
File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing rich-T-kid/primitive-hash-no-null-unroll (4d6ff20) to 1c49b7f (merge-base) diff Run configurationrun benchmark with_hashesResults will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing rich-T-kid/primitive-hash-no-null-unroll (4d6ff20) to 1c49b7f (merge-base) diff Run configurationrun benchmark with_hashesCPU Details (lscpu)Details
Resource Usagewith_hashes — base (merge-base)
with_hashes — branch
File an issue against this benchmark runner |
Splits the single scalar loop in
hash_array_primitive's no-null branch into two const-generic-free helpers that batch independent hashes so the pipeline isn't serialized on store-to-load deps:hash_prim_fresh_dense: overwrite mode, 16 rows per iteration.hash_prim_rehash_dense: fold-with-prev mode, 8 rows per iteration (narrower because each lane holds bothprevandvaluelive).Both use
get_unchecked/get_unchecked_mutwith proven invariants (hashes.len() == values.len()) and a scalar tail for the remainder.Benchmarks (int64, 8192 rows, no nulls) on Apple M4 Max:
single, no nulls : neutral to small regression (<+7%)
multiple, no nulls : ~-14%
The nullable path is unchanged -- see PR #26143 for the sibling change.
Which issue does this PR close?
Rationale for this change
What changes are included in this PR?
What is the testing strategy for this PR?
Are there any user-facing changes?