Skip to content

perf(blosc): use NEON SIMD for byte-shuffle on aarch64 - #864

Open
d-v-b wants to merge 1 commit into
zarr-developers:mainfrom
d-v-b:perf/blosc-neon-shuffle
Open

d-v-b wants to merge 1 commit into
zarr-developers:mainfrom
d-v-b:perf/blosc-neon-shuffle

Conversation

@d-v-b

@d-v-b d-v-b commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

🤖 AI text below 🤖

Summary

c-blosc 1.x only ships SSE2/AVX2 shuffle kernels. On aarch64 (Apple Silicon, AWS Graviton, …) its dispatcher always falls back to blosc_internal_(un)shuffle_generic, a scalar byte-by-byte loop. That loop dominates Blosc runtime on ARM: with shuffle disabled, the same build decodes lz4 at ~17 GB/s, but with shuffle enabled it decodes at ~2.5 GB/s.

This PR adds src/numcodecs/_blosc_neon/shuffle-generic-neon.c. On aarch64, meson compiles it instead of c-blosc's shuffle-generic.c. It defines the same two symbols:

  • NEON kernels for type sizes 2, 4, 8 and 16, copied verbatim from c-blosc2's blosc/shuffle-neon.c (BSD-3, attributed in the header).
  • c-blosc1's scalar code for other type sizes and for the tail of each block.

The c-blosc submodule is not modified, and c-blosc's dispatcher (shuffle.c) is unchanged. A new neon meson option (default auto) controls this. Bitshuffle is unchanged, since c-blosc2 also found its NEON bitshuffle no faster than scalar.

Compatibility

The output is byte-identical to the scalar build. The agent compared SHA-256 hashes of the compressed output from both builds over 1,452 combinations: type sizes 1–32 including odd sizes, 11 buffer lengths, 3 block sizes, shuffle and bitshuffle, lz4 and blosclz. The hashes matched and every case round-tripped. The existing backwards-compatibility fixtures also decode correctly.

Benchmarks

Apple M5, Python 3.14, single thread (use_threads = False), Blosc(cname, clevel=5, shuffle=SHUFFLE), MB/s of uncompressed data. The numbers were stable to within ~1% across two runs.

data cname encode before → after decode before → after
uint16, 1 MiB lz4 2665 → 5758 2557 → 16024
uint16, 16 MiB lz4 2309 → 4274 2344 → 10497
float32, 1 MiB lz4 2452 → 5021 2498 → 13198
float64, 1 MiB lz4 2812 → 5793 4663 → 14264
uint16, 1 MiB blosclz 1685 → 2462 2210 → 8338
float32, 1 MiB zstd 257 → 273 2033 → 5038

Testing

  • Adds test_shuffle_roundtrip_typesizes. It round-trips across type sizes, lengths that aren't multiples of the vector width, and block sizes, to cover the scalar tail.
  • The full suite passes locally on macOS arm64.
  • Not tested: Linux aarch64 and Windows ARM64. For MSVC, the source guard and the meson check accept _M_ARM64 as well as __ARM_NEON.

Notes for reviewers

  • A longer-term alternative is to upstream these kernels to c-blosc 1.x. The file-substitution approach here avoids carrying a fork of the submodule in the meantime.

🤖 Generated with Claude Code

c-blosc 1.x only provides SSE2/AVX2 shuffle kernels. On aarch64 its
dispatcher always selects blosc_internal_(un)shuffle_generic, a scalar
byte-by-byte loop that dominates Blosc runtime on ARM.

On aarch64, compile a drop-in replacement for c-blosc's shuffle-generic.c
that implements those two symbols with NEON kernels (copied from c-blosc2)
for type sizes 2, 4, 8 and 16, falling back to the scalar loop otherwise.
The c-blosc submodule is left untouched and the output is byte-identical.

Controlled by a new `neon` meson option (default: auto).

Assisted-by: ClaudeCode:claude-opus-5-5
@d-v-b
d-v-b force-pushed the perf/blosc-neon-shuffle branch from 2e138e2 to 5c34a8d Compare September 27, 2026 11:33
@codecov

codecov Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 100.00%. Comparing base (1f83681) to head (5c34a8d).

Additional details and impacted files
@@            Coverage Diff            @@
##              main      #864   +/-   ##
=========================================
  Coverage   100.00%   100.00%           
=========================================
  Files           27        27           
  Lines          905       905           
=========================================
  Hits           905       905           
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant