Skip to content

Support LLVM 23's library-based NVPTX backend - #161

Merged
AntonOresten merged 4 commits into
mainfrom
agent/llvm23-backend
Sep 14, 2026
Merged

AntonOresten merged 4 commits into
mainfrom
agent/llvm23-backend

Conversation

@AntonOresten

@AntonOresten AntonOresten commented Sep 11, 2026 •

Copy link
Copy Markdown
Member

NVPTX_LLVM_Backend_jll 23 replaces the llc executable with libnvptx. Migrate PTX's registry, wrapper adaptations, and conformance tests together so the package works with CUDA.jl 6.4 and its library-based backend (CUDA.jl#3267, GPUCompiler.jl#930).

Require CUDACore/CUDATools 6.4 or later and resolve registered releases. .ci/prepare.jl refreshes cached registries and creates isolated test/documentation environments across supported Julia versions, including Julia 1.10. CUDA and the backend remain weak dependencies of PTX.

  • Regenerate all 2,633 NVVM records from LLVM 23.1.1: 95 additions, 31 removals, and 473 changed existing records. Adapt extraction to LLVM's new overload metadata and vector types; retain the reviewed convergence and memory overlays.
  • Preserve the public scalar result and tuple input conventions for single-register tcgen05 loads/stores despite their new vector intrinsic signatures.
  • Route the newly available mxf4nvf4 4X ue8m0 scaled MMA through NVVM. Its golden changes qualifier ordering and eliminates a redundant zero move. Numerical tests cover unit and non-unit scales plus accumulation on GB10.
  • Replace executable probes with the library API. Add diagnostics/recovery, FMA contraction and option-isolation checks, all 28 new ld.red signatures and selection probes, and tensor-map dimension bounds. Existing ld.red wrappers retain their reviewed assembly route.
  • Use toolkit-free NVPTX targets for LLVM optimizer contract tests: Julia 1.13 rejects NVVM intrinsics during CPU-target lowering. Keep the call-count, convergence, side-effect and raw return-metadata assertions.
  • Strip debug metadata before backend code generation in the six-variant register-wait comparison. Julia 1.12 otherwise produces scheduling differences after the host matrix reflection tests. Retain exact instruction-encoding equality, including spill instructions.

Validation with registered CUDACore/CUDATools/CUPTI/NVML 6.4.0, GPUCompiler 2.8.0, LLVM.jl 9.13.1 and NVPTX_LLVM_Backend_jll 23.1.1+2:

  • Julia 1.10: 26,762/26,762 assertions across host/nvvm, host/conformance, host/nvptx_backend, host/effect_ceiling, host/warp_reduce, host/wrappers, host/tensor_map, host/aqua and ptxas/golden.
  • Julia 1.13: 26,771/26,771 across the same suites plus gpu/sm121a_smoke, executed on GB10 with CUDA compiler 13.4.59.
  • Register-wait regression fix: Julia 1.12 passes 10,179/10,179 assertions with the previously failing order (host/matrix_api_safety followed by ptxas/wait_registers), coverage enabled, and the exact CI backend build 23.1.1+2. Julia 1.10 and 1.13 each pass 301/301 assertions across host/wait_registers and ptxas/wait_registers. All six instruction streams remain byte-identical.
  • Both environments were created by the CI preparation script. All CUDA components resolve from the registry, with no source checkout. The release update required no further golden changes.

Supersedes the compatibility-only updates in #155, #156 and #157.

@AntonOresten
AntonOresten marked this pull request as draft September 11, 2026 11:45
@codecov

codecov Bot commented Sep 11, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

The Julia 1.12 CI worker runs matrix_api_safety before wait_registers.
That sequence perturbs the quarters variant's instruction scheduling and
predicate allocation through debug metadata, although instruction counts
and resource usage match. Running the comparison alone passes.

Strip debug information before backend code generation in both jobs.
Retain the exact encoding comparison across all six attention variants
and its ability to catch added spill instructions.

Validation: the formerly failing Julia 1.12 sequence passes all 10,179
assertions with CUDA 6.4, backend 23.1.1+2, and coverage enabled.
Julia 1.10 and 1.13 each pass all 301 host and offline wait assertions.
@AntonOresten
AntonOresten marked this pull request as ready for review September 14, 2026 14:34
@AntonOresten
AntonOresten merged commit 5d13b07 into main Sep 14, 2026
6 checks passed
@AntonOresten
AntonOresten deleted the agent/llvm23-backend branch September 14, 2026 14:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant