Conversation
c-blosc 1.x only provides SSE2/AVX2 shuffle kernels. On aarch64 its dispatcher always selects blosc_internal_(un)shuffle_generic, a scalar byte-by-byte loop that dominates Blosc runtime on ARM. On aarch64, compile a drop-in replacement for c-blosc's shuffle-generic.c that implements those two symbols with NEON kernels (copied from c-blosc2) for type sizes 2, 4, 8 and 16, falling back to the scalar loop otherwise. The c-blosc submodule is left untouched and the output is byte-identical. Controlled by a new `neon` meson option (default: auto). Assisted-by: ClaudeCode:claude-opus-5-5
d-v-b
force-pushed
the
perf/blosc-neon-shuffle
branch
from
September 27, 2026 11:33
2e138e2 to
5c34a8d
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #864 +/- ##
=========================================
Coverage 100.00% 100.00%
=========================================
Files 27 27
Lines 905 905
=========================================
Hits 905 905 🚀 New features to boost your workflow:
|
2 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🤖 AI text below 🤖
Summary
c-blosc 1.x only ships SSE2/AVX2 shuffle kernels. On aarch64 (Apple Silicon, AWS Graviton, …) its dispatcher always falls back to
blosc_internal_(un)shuffle_generic, a scalar byte-by-byte loop. That loop dominates Blosc runtime on ARM: with shuffle disabled, the same build decodes lz4 at ~17 GB/s, but with shuffle enabled it decodes at ~2.5 GB/s.This PR adds
src/numcodecs/_blosc_neon/shuffle-generic-neon.c. On aarch64, meson compiles it instead of c-blosc'sshuffle-generic.c. It defines the same two symbols:blosc/shuffle-neon.c(BSD-3, attributed in the header).The
c-bloscsubmodule is not modified, and c-blosc's dispatcher (shuffle.c) is unchanged. A newneonmeson option (defaultauto) controls this. Bitshuffle is unchanged, since c-blosc2 also found its NEON bitshuffle no faster than scalar.Compatibility
The output is byte-identical to the scalar build. The agent compared SHA-256 hashes of the compressed output from both builds over 1,452 combinations: type sizes 1–32 including odd sizes, 11 buffer lengths, 3 block sizes, shuffle and bitshuffle, lz4 and blosclz. The hashes matched and every case round-tripped. The existing backwards-compatibility fixtures also decode correctly.
Benchmarks
Apple M5, Python 3.14, single thread (
use_threads = False),Blosc(cname, clevel=5, shuffle=SHUFFLE), MB/s of uncompressed data. The numbers were stable to within ~1% across two runs.Testing
test_shuffle_roundtrip_typesizes. It round-trips across type sizes, lengths that aren't multiples of the vector width, and block sizes, to cover the scalar tail._M_ARM64as well as__ARM_NEON.Notes for reviewers
🤖 Generated with Claude Code