gdb.rocm: step-schedlock-spurious-waves.cpp: Use inlined asm for breakpoint - #302
Merged
Merged
Conversation
Collaborator
Collaborator
lumachad
reviewed
Aug 24, 2026
lumachad
left a comment
Collaborator
There was a problem hiding this comment.
Some automated comments. Otherwise this seems OK. Have you validated on other gfx architectures?
akondrat-amd
force-pushed
the
users/akondrat/spurious-waves-fix
branch
2 times, most recently
from
August 24, 2026 15:54
546dfad to
5a4cae9
Compare
Collaborator
Thanks for the update @akondrat-amd. How was validation on other gfx arches? |
…kpoint On gfx1250 the original code did not achieve full wavefront occupancy; VGPR pressure was created by the unoptimized function call and for loop. Instead of relying on optnone + an empty end_of_kernel function as a breakpoint site, use __forceinline__ with volatile inline assembly (s_nop 0). This guarantees the instruction is emitted in-line in the kernel and cannot be optimized away, giving the debugger a reliable address to break on. Move optnone to kern to keep the s_sleep calls from being optimized, and replace the counted loop with explicit s_sleep calls to make the stepping sequence clearer.
lumachad
force-pushed
the
users/akondrat/spurious-waves-fix
branch
from
August 25, 2026 12:15
5a4cae9 to
4276b43
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
On gfx1250 the original code did not achieve full wavefront occupancy; VGPR pressure was created by the unoptimized function call and for loop.
Instead of relying on optnone + an empty end_of_kernel function as a breakpoint site, use forceinline with volatile inline assembly (s_nop 0). This guarantees the instruction is emitted in-line in the kernel and cannot be optimized away, giving the debugger a reliable address to break on.
Move optnone to kern to keep the s_sleep calls from being optimized, and replace the counted loop with explicit s_sleep calls to make the stepping sequence clearer.
More details in Jira [AIROCGDB-552]