Repository navigation
AffineCFG: check lockstep dependences on the parallel check's relations - #3405
Merged
Merged
Conversation
wsmoses
force-pushed
the
pb/affine-lockstep-shared-relations
branch
2 times, most recently
from
October 6, 2026 23:41
49a0f62 to
5119f9d
Compare
isLoopMemoryLockStepExecutable checked dependences with MLIR's checkMemrefAccessDependence, while isLoopMemoryParallel builds its own access relations (InvariantTerms): loop-invariant products of symbols abstracted (#3369), domains that stop at the affine scope and read the bounds defined above it as symbols (#3380), the facts of the nest's in-bounds accesses (#3400) and of the checks on the way to it (#3404). An access whose index multiplies two symbols, or one under an scf.if, failed the lockstep check outright. It now builds the same relations and calls checkAccessDependence; accesses to different memrefs, or two reads, have no dependence, as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016zErYp7upmqr4NHfhod9UD
wsmoses
force-pushed
the
pb/affine-lockstep-shared-relations
branch
from
October 6, 2026 23:42
5119f9d to
4b0f2f4
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #3405 +/- ##
=======================================
Coverage 29.61% 29.61%
=======================================
Files 240 240
Lines 48460 48467 +7
=======================================
+ Hits 14351 14354 +3
- Misses 34109 34113 +4 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Lockstep check: use the same dependence relations as the parallel check.
isLoopMemoryLockStepExecutable, which the raiser asks before running a constant-count loop in lock step, checked dependences with MLIR'scheckMemrefAccessDependence. MeanwhileisLoopMemoryParallelbuilds its own access relations throughInvariantTerms:So an access whose index multiplies two symbols, or one under an
scf.if, failed the lockstep check outright, even where the parallel check could analyze it.The lockstep check now builds the same relations and calls
checkAccessDependence, so all of these facts reach it too.Where the guard facts reach lockstep. The raiser checks lockstep on outlined kernel functions. There, a host-side
MFEM_VERIFYis not in the function being raised, so its facts reach onlyaffine-cfg's parallel check, which runs before outlining. A guard inside the raised function does reach the lockstep check. As before, accesses to different memrefs, or two reads, have no dependence. Its rules for which dependences lockstep execution may break are unchanged.Measured with
affine-cfg,canonicalizeand raising (enable_lockstep_for=true) on all 143 MFEM translation units. The raiser emits 14 fewerstablehlo.whiles (11,077 → 11,063), inbilininteg_dgtrace_ea(−5),bilininteg_mass_ea(−5),bilininteg_convection_ea(−3) andbilininteg_diffusion_ea(−1). No file gains any, and none fail.Tests:
raising/affine_to_stablehlo_lockstep_invariant_terms.mlir(full-line goldens), which fails on main: a 3-trip loop copying at an index offset by a product of two symbols now runs in lock step, as one gather and one scatter, instead of awhile.raising/affine_to_stablehlo_if_masking.mlir:@test_if_else_masking(per element,if (m[i]) a[i] *= 2 else a[i] *= a[i]) now runs in lock step as two whole-tensor selects instead of awhile. That is correct: each iteration reads and writes only its own element.Lit: every
affinecfg*test, plus the raising tests above, against this branch rebased on main (which includes #3404 and #3406).🤖 Generated with Claude Code
https://claude.ai/code/session_016zErYp7upmqr4NHfhod9UD