Skip to content

perf(reverse): use a dedicated kernel for the dims reversal - #118

Open
shreyas-omkar wants to merge 4 commits into
JuliaGPU:mainfrom
shreyas-omkar:sh/reverse-kernel-speedup
Open

perf(reverse): use a dedicated kernel for the dims reversal#118
shreyas-omkar wants to merge 4 commits into
JuliaGPU:mainfrom
shreyas-omkar:sh/reverse-kernel-speedup

Conversation

@shreyas-omkar

Copy link
Copy Markdown
Member

No description provided.

shreyas-omkar and others added 4 commits August 21, 2026 18:04
Add a `dims` keyword to `reverse!`/`reverse`, reaching parity with
`Base.reverse` and the vendor reverse kernels. `dims=:` (the default)
keeps the fast flat path - each thread swaps one mirrored pair - while
`dims=d` (an integer or iterable) reverses only along those dimensions
via a general ND kernel written on `foreachindex`, so it runs on every
backend (CUDA/AMDGPU/oneAPI/Metal/POCL) and the CPU-threaded path from
one implementation. Invalid dims throw `ArgumentError`.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Exercise single-dim, multi-dim ((1,2)/(2,3)/(1,3)/:), size-1 degenerate
dims and 3-D arrays across in-place, out-of-place and allocating forms,
plus ArgumentError on out-of-range dims. Verified on CPU-threaded,
AMDGPU (ROCm) and POCL backends.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants