Repository navigation
Conversation
affine.parallel's verifier admits maxs and mins only on an integer explicitly signed, a type arith's maxsi and minsi never produce, so the loop affine-cfg makes of a loop carrying the max of what it reads (the longest row of a CSR structure, say) did not verify, and affine-cfg's own rule for an scf.parallel reduction wanted the same; the raising had no reduce for the integer extrema. The LLVM patch admits a signless integer for the four, affine-cfg likewise, and the raising reduces them: maxs and mins to stablehlo.maximum and minimum from the signed extremes, maxu and minu through arith.maxui and minui, as arith-raise reads them. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016zErYp7upmqr4NHfhod9UD
A loop whose bounds a row of a structure gives, as a CSR transpose's `for (j = I[i]; j < I[i+1]; ++j)`, is not affine, stays an scf.for, and raised runs the rows one at a time. It runs instead to the longest row, M = max over the rows of hi - lo, a parallel reduction before the row loop, with its body under `lo + k < hi`: every row then runs the same k loop, which the raising reads as a lane of it. The bounds may vary with several enclosing loops (the reduction nests them), and their computation is copied: pure ops, and loads of buffers nothing in the row loop writes, which every row runs. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016zErYp7upmqr4NHfhod9UD
wsmoses
force-pushed
the
pb/pad-loop-to-longest-row
branch
from
October 8, 2026 17:40
bee2e63 to
0f035ed
Compare
This was referenced Oct 8, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #3436 (needs the integer max reduction); the diff here is the last commit.
MFEM's
ElementRestriction::MultTranspose(and every CSR-shaped kernel) runs, per dofi,for (j = offsets[i]; j < offsets[i+1]; ++j) acc += x[indices[j]]. The bounds are loads, so the loop is not affine, stays anscf.forinside the parallel row loop, and the raising makes of it a while nested in a (peeled) while over the rows — which at exec time stays a host-driven loop over 185281 rows: the ex1 CG iteration's dominant cost on XLA (Sep 28 census: all 1544 remaining parallel loops of the suite are this CSR family).PadLoopToLongestRow: for anscf.forwith a constant step whose bounds vary with enclosing affine loops (the rows), through pure ops and loads of buffers nothing in the row loop writes (the loads run by every row — a load under a guard refuses; a speculatable op under one is copied), emit before the outermost row loop a reduction nest over the same rows,reduce ("maxs")ofhi - lo(the copied computation), and replace the loop byaffine.for k = 0 to Mwhose body runs underlo + k * step < hi. The rotated form clang makes (if (lo < hi) for (lo+1 .. hi+1)) and a vector dimension between the rows and the entries are handled (@rotated); bounds varying with two loops nest two reductions (@two_rows).Mmust be a symbol, so the outermost row loop sits directly in the affine scope (the GPU kernels do; the host copies underscf.ifguards are left alone).What follows on its own: affine-cfg turns the padded loop's conditional accumulation into
affine.parallel ... reduce ("addf")(#3409), the raising emitsMas astablehlo.reduce maximumand the k loop as a while with the rows as lanes (static rows) or a parallel while over the rows (dynamic rows), and at exec time #3437 batches the latter through the k loop intowhile k < M { gathers over every row }— the ELL form.On
fem/restriction.cppreplayed to its last affine-cfg: 20 CSR loops padded (scf.for130 → 110,affine.parallel68 → 108 with the 20 reductions); the full pipeline raises all 24 kernels.Test
affinecfg_pad_loop_longest_row.mlir(full-line goldens): the plain transpose, the rotated vdim form, two-loop bounds, and two refusals (a bound read under a condition, offsets written in the row loop).affinecfg_copy_carry.mlir's@scf_copy(a loop tonb[j]under a row loop) is now padded too; regolded, the collapse of its copy carry to a select is unchanged. Lit on main + the open affine-cfg PRs: only the known interaction goldens differ.🤖 Generated with Claude Code
https://claude.ai/code/session_016zErYp7upmqr4NHfhod9UD