Repository navigation
AutoBatching: batch through a nested loop whose trip count the data gives - #3437
Merged
Merged
Conversation
…ives The parallel-while batcher ran a nested loop over the batched values only where its trip count was a constant, though all it needs is that the count be the same for every iteration of the parallel loop, which its condition check already demands: a CSR transpose padded to its longest row runs its inner loop to a max read from the offsets. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016zErYp7upmqr4NHfhod9UD
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #3437 +/- ##
==========================================
+ Coverage 29.63% 29.96% +0.33%
==========================================
Files 240 240
Lines 48496 48546 +50
==========================================
+ Hits 14371 14548 +177
+ Misses 34125 33998 -127 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Member
Author
|
MFEM GPU unit suite (74 tests) with the runtime built from Enzyme-JAX main + #3437 #3439 #3440 #3441 #3442 and Reactant.jl #3423 + #3425, over objects from main + the open affine-cfg PRs (#3436 #3438 among them): 74/74 (sweep final55, 2026-10-08 17:38). ex1 (Poisson, order 3, PA, PCG to 1e-12), timed solves in one process (JIT excluded), native CUDA for reference:
|
This was referenced Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
ParallelWhileBatcher(#3261) runs a loop nested in anenzymexla.parallelwhile over the batched values — interchanging it with the parallel loop — only when its trip count is a constant. What the interchange needs is that the count be the same for every iteration of the parallel loop, andanalyzeWhilealready refuses a condition reading anything that varies with it; so the gate is now a constant start and step, with any invariant limit.The case: a CSR transpose padded to its longest row (the next affine-cfg PR),
for i: for k < M: j = off[i] + k; if j < off[i+1]: acc += x[idx[j]], whereM = max_i (off[i+1] - off[i])is a reduce over the data. With this the row loop batches intowhile k < Mover gathers of every row at once, where before it stayed a host-driven loop over the rows (MFEM'sElementRestriction::MultTranspose, the dominant cost ofex1's CG iteration on XLA).Test
parallel_while_invariant_trip.mlir(that kernel, full-line golden); checked against NumPy on random CSR structures with empty rows (exact before and after). The batcher's other tests unchanged.🤖 Generated with Claude Code
https://claude.ai/code/session_016zErYp7upmqr4NHfhod9UD