Skip to content

[INTERNAL][UPSTREAM CANDIDATE] ruby: stop native resume from dropping Ruby frames - #91

Draft
dalehamel wants to merge 1 commit into
dale/ruby41-supportfrom
dale/ruby-native-resume-cfunc-owner
Draft

dalehamel wants to merge 1 commit into
dale/ruby41-supportfrom
dale/ruby-native-resume-cfunc-owner

Conversation

@dalehamel

@dalehamel dalehamel commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

Fork-local scope

Important

This draft targets the Shopify fork branch dale/ruby41-support, not upstream main, and is independent of #89/#90. The regenerated coredump goldens are in stacked draft #92. Until that lands, this PR's coredump test jobs are expected to fail on exactly the 18 Ruby cases #92 updates.

With native resume enabled (the default), samples taken inside a C method lose Ruby frames. I found this while investigating open-telemetry#1934. It's separate from that issue's GC-offset fix (#89).

Cause

Native resume defers a cfunc frame until the next rb_vm_exec, so the cfunc owns the native frames unwound in between. That's only right for the cfunc that called into that VM loop, which sits directly below the loop's base (VM_FRAME_FLAG_FINISH) frame.

When a sample lands inside a C method, that method's native frames were unwound before the Ruby unwinder ran. Deferring it attaches it to the next native segment, shifts every later cfunc by one segment, and drops the outermost Ruby segment. At the top level there's no later rb_vm_exec, so every Ruby frame is lost. That's the Ruby 3.3+ Process.clock_gettime symptom in open-telemetry#1934.

Change

  • Defer a cfunc only when the previous frame walked was a FINISH frame, and push any other cfunc inline. This also covers C methods calling C methods, such as Range#each → <=>.
  • Exception: if at most one native frame was unwound before the first Ruby unwinder entry, a cfunc on top of the VM stack is still deferred. That cfunc isn't running: its block has returned and the inner rb_vm_exec is finishing. A running C method always adds at least two native frames, its own function and the VM call helper reached through an indirect call.
  • Once the last Ruby frame is pushed, mark the Ruby unwinder done, so a further rb_vm_exec frame can't push that frame again. A still-deferred bottom cfunc is pushed instead of dropped.
  • VM_FRAME_FLAG_FINISH is 0x0020 from Ruby 2.6 through master. The skip_native_resume, JIT, and pre-2.6 paths are unchanged.
  • This is a single commit, including the BPF objects rebuilt in the dev image CI uses. unwind_ruby grows from 1,298 to 1,381 instructions on both architectures.

Cherry-picking

ruby_tracer.ebpf.c and types.h merge cleanly onto upstream main and v0.0.202636. Only support/ebpf/tracer.ebpf.{amd64,arm64} conflict. Rebuild them with make amd64 -C support/ebpf && make arm64 -C support/ebpf, ideally in otel/opentelemetry-ebpf-profiler-dev, rather than taking either side. Without the rebuild, the embedded BPF doesn't contain the fix.

Validation

Live profiling on arm64 (kernel 6.8), Ruby 3.3.12 with default native resume, same base before and after:

Workload Before After
Issue workload (hot.rb) 35/75 samples with Ruby frames (47%) 90/96 (94%), all reaching <main>
Reporter's full script 53/117 (45%); 43 reach <main> 128/131 (98%), all reaching <main>
Tight loop (1..n).each { |i| s += i } 64.5% complete; 34.6% with no Ruby frames 98.8% complete and correctly ordered
Block doing C-method work (1..n).each { |i| s += i.to_s.size } 35.3% complete; 57.6% missing frames 92.5% complete and correctly ordered

In the "after" column, every remaining sample is an abort inside a libruby PLT stub. That's an arm64 native-unwinder gap, unrelated to this change and fixed separately in #93. With both changes, every sample in both loops is complete.

  • Coredump replays: only Ruby cases change (9 of 81 on arm64, 9 of 96 on amd64). With the stacked goldens, all 81 arm64 and 96 amd64 cases pass.
  • The committed base BPF objects rebuild byte-for-byte in that image, so the blob check can verify the new ones.

Known limit

In the window between pushing a block's frame and entering its VM loop, the bottom Ruby frames are still dropped, as before. This never occurred in about 9,500 samples of block-heavy loops.

Native resume defers a cfunc frame until the next rb_vm_exec, so the cfunc owns the native frames unwound in between. That is only right for the cfunc that called into that VM loop, which sits directly below the loop's base (VM_FRAME_FLAG_FINISH) frame.

When a sample lands inside a C method, that method's native frames were unwound before the Ruby unwinder ran. Deferring it attached it to the next native segment, shifted every later cfunc by one segment, and dropped the outermost Ruby segment. At the top level no later rb_vm_exec exists, so every Ruby frame was lost.

Push such cfuncs inline. A cfunc on top of the VM stack with at most one native frame above the first rb_vm_exec is still deferred: it is not running, because its block has returned and the inner rb_vm_exec is finishing. A running C method always adds at least two frames, its own function and the VM call helper reached through an indirect call.

Once the last Ruby frame is pushed, mark the Ruby unwinder done so a further rb_vm_exec frame cannot push that frame again, and push a still-deferred bottom cfunc instead of dropping it.

The tracer blobs are rebuilt in otel/opentelemetry-ebpf-profiler-dev@sha256:585409de39191c201f91a8e6ec5294556c8f21c7d2ce0b0f96e838f27e399f82. unwind_ruby grows from 1298 to 1381 instructions.

Found while investigating open-telemetry#1934.

Signed-off-by: Dale Hamel <dale.hamel@shopify.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant