Skip to content

AffineCFG: a loop over a row's entries runs to the longest row - #3438

Open
wsmoses wants to merge 2 commits into
mainfrom
pb/pad-loop-to-longest-row
Open

wsmoses wants to merge 2 commits into
mainfrom
pb/pad-loop-to-longest-row

Conversation

@wsmoses

@wsmoses wsmoses commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

Stacked on #3436 (needs the integer max reduction); the diff here is the last commit.

MFEM's ElementRestriction::MultTranspose (and every CSR-shaped kernel) runs, per dof i, for (j = offsets[i]; j < offsets[i+1]; ++j) acc += x[indices[j]]. The bounds are loads, so the loop is not affine, stays an scf.for inside the parallel row loop, and the raising makes of it a while nested in a (peeled) while over the rows — which at exec time stays a host-driven loop over 185281 rows: the ex1 CG iteration's dominant cost on XLA (Sep 28 census: all 1544 remaining parallel loops of the suite are this CSR family).

PadLoopToLongestRow: for an scf.for with a constant step whose bounds vary with enclosing affine loops (the rows), through pure ops and loads of buffers nothing in the row loop writes (the loads run by every row — a load under a guard refuses; a speculatable op under one is copied), emit before the outermost row loop a reduction nest over the same rows, reduce ("maxs") of hi - lo (the copied computation), and replace the loop by affine.for k = 0 to M whose body runs under lo + k * step < hi. The rotated form clang makes (if (lo < hi) for (lo+1 .. hi+1)) and a vector dimension between the rows and the entries are handled (@rotated); bounds varying with two loops nest two reductions (@two_rows). M must be a symbol, so the outermost row loop sits directly in the affine scope (the GPU kernels do; the host copies under scf.if guards are left alone).

What follows on its own: affine-cfg turns the padded loop's conditional accumulation into affine.parallel ... reduce ("addf") (#3409), the raising emits M as a stablehlo.reduce maximum and the k loop as a while with the rows as lanes (static rows) or a parallel while over the rows (dynamic rows), and at exec time #3437 batches the latter through the k loop into while k < M { gathers over every row } — the ELL form.

On fem/restriction.cpp replayed to its last affine-cfg: 20 CSR loops padded (scf.for 130 → 110, affine.parallel 68 → 108 with the 20 reductions); the full pipeline raises all 24 kernels.

Test affinecfg_pad_loop_longest_row.mlir (full-line goldens): the plain transpose, the rotated vdim form, two-loop bounds, and two refusals (a bound read under a condition, offsets written in the row loop). affinecfg_copy_carry.mlir's @scf_copy (a loop to nb[j] under a row loop) is now padded too; regolded, the collapse of its copy carry to a select is unchanged. Lit on main + the open affine-cfg PRs: only the known interaction goldens differ.

🤖 Generated with Claude Code

https://claude.ai/code/session_016zErYp7upmqr4NHfhod9UD

wsmoses and others added 2 commits October 8, 2026 12:20
affine.parallel's verifier admits maxs and mins only on an integer
explicitly signed, a type arith's maxsi and minsi never produce, so the
loop affine-cfg makes of a loop carrying the max of what it reads (the
longest row of a CSR structure, say) did not verify, and affine-cfg's own
rule for an scf.parallel reduction wanted the same; the raising had no
reduce for the integer extrema. The LLVM patch admits a signless integer
for the four, affine-cfg likewise, and the raising reduces them: maxs and
mins to stablehlo.maximum and minimum from the signed extremes, maxu and
minu through arith.maxui and minui, as arith-raise reads them.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016zErYp7upmqr4NHfhod9UD
A loop whose bounds a row of a structure gives, as a CSR transpose's
`for (j = I[i]; j < I[i+1]; ++j)`, is not affine, stays an scf.for, and
raised runs the rows one at a time. It runs instead to the longest row,
M = max over the rows of hi - lo, a parallel reduction before the row
loop, with its body under `lo + k < hi`: every row then runs the same k
loop, which the raising reads as a lane of it. The bounds may vary with
several enclosing loops (the reduction nests them), and their computation
is copied: pure ops, and loads of buffers nothing in the row loop writes,
which every row runs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016zErYp7upmqr4NHfhod9UD

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant