Repository navigation
AffineCFG: parallelize an accumulation taken under a condition - #3409
Merged
Merged
Conversation
A value a loop accumulates only where a condition holds, acc = c ? acc + x : acc through an scf.if, affine.if or select, matched no reduction: the loop yields the if or select, not the combination. Such a loop accumulates c ? x : id wherever, with id the combination's identity (-0.0 for addf, so that a -0.0 sum stays one; 1.0 for mulf; 0 for addi/ori/xori; 1 for muli; all ones for andi; a multiply-add adds its product), which is a reduction. isLoopParallel now takes such a value as a reduction of its kind, and AffineParallelizePattern, where it parallelizes the loop, rewrites the if or select to give x or the identity and combines the initial value with the reduction. The if keeps its branches, so nothing in them runs where it did not; a loop that stays serial is left as it is. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016zErYp7upmqr4NHfhod9UD
wsmoses
force-pushed
the
pb/affine-conditional-accumulation
branch
from
October 7, 2026 00:13
4b5b041 to
63a4b4b
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #3409 +/- ##
=======================================
Coverage 29.61% 29.62%
=======================================
Files 240 240
Lines 48460 48474 +14
=======================================
+ Hits 14351 14359 +8
- Misses 34109 34115 +6 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
affine-cfg: parallelize an accumulation taken only under a condition.
A value a loop accumulates only where a condition holds,
acc = c ? acc ⊕ x : acc, through anscf.if,affine.iforselect, matched no reduction: the loop yields the if or select, not the combination. Such a loop accumulatesc ? x : idon every iteration, whereidis the combination's identity, and that is a reduction. The identities used are exact:addf-0.0(adding it leaves any sum, including-0.0, unchanged)mulf1.0addi,ori,xori0muli1andifmuladd-0.0, with the branch yielding the product (as in #3407)This is done in the parallelization itself, as #3407 does for multiply-adds, so a loop that stays serial is left exactly as it is:
isLoopParallel).addConditionalReductionstakes such a carried value as a reduction of its kind, without changing the IR. It applies when the value is read only by the combination and by the unchanged arm or operand.AffineParallelizePattern). Where it parallelizes the loop, the if or select is rewritten to givexon the combining side and the identity on the other. That becomes each iteration's reduced value, and the initial value is combined with the reduction outside the loop (arith::getReductionOp). The if keeps its branches, so nothing in them runs where it did not before.Measured with
affine-cfgon the IR of the 143 raised MFEM translation units just before the lastaffine-cfg, against current main (with #3407): parallel dimensions go from 66,139 to 66,153 (batchitrans+6,bilininteg_mixedcurl_pa+4,bilininteg_curlcurl_pa+4), and no file loses any. Of the 160 carries of this shape, many sit in loops that stay serial for other reasons. Others feed one carried value from another's arm, such asbatchitrans's product-rule recurrences.Test:
affinecfg_conditional_accumulation.mlir(full-line goldens). It fails on main.@scf_if: thescf.ifyields the load or-0.0, and the loop is anaddfreduction.@affine_if: the same under anaffine.ifon the iteration.@select: ani32count,select(c, acc + 1, acc), becomes anaddireduction ofselect(c, 1, 0).@fmuladd: a multiply-add under a condition. Theifyields the product.@last: a value replaced under a condition (the last one taken) is no accumulation and stays carried.@serial: an accumulation under a condition in a loop that is not parallel for other reasons is left exactly as it is.Lit: every RUN line of every test against current main's binary. Nothing fails that passed there.
🤖 Generated with Claude Code
https://claude.ai/code/session_016zErYp7upmqr4NHfhod9UD