Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .knowledgebase/.cairn
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
{"generator": "cairn", "retired_threshold": 10}
41 changes: 41 additions & 0 deletions .knowledgebase/memories/dead-bitstate-lazy-fastconsume.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
---
id: dead-bitstate-lazy-fastconsume
title: BitState lazy char-class loop fast-consume measured flat on the real corpus — rolled back; lazy-loop consumption is not the lazy families' cost center
kind: dead-end
tags: [bitstate, lazy-loop, perf, measured-negative, rollback]
applies_to: [reggie-runtime/src/main/java/com/datadoghq/reggie/runtime/BitStateMatcher.java]
source: backend-ready/dead-bitstate-lazy-fastconsume
status: active
recorded: 2026-09-28
valid_at: 2026-09-28
---


# BitState lazy char-class loop fast-consume measured flat on the real corpus — rolled back

## What was tried
A lazy mirror of the existing greedy fast-consume path in BitStateMatcher: a lazy
single-char-class loop (`c*?`) whose continuation closure contains no anchor, no accepting
state, and at least one consuming transition gets a tight first pass — consume the whole
`c`-run in a scan, mark the same visited cells, push exit jobs only at positions where the
continuation's first-char set can match input[p] (descending push = ascending pop = Perl
lazy priority). Fully implemented and gate-verified (fuzz seeds, real-corpus parity probes
0 divergences, full suite).

## Why it was rolled back
- Box JMH matched sweep: flat to slightly negative vs the pre-change build; controls flat.
- Per-pattern C2-converged profiling showed the path barely engages on the corpus lazy
families: with an all-optional tail (`\s*((?<function>[^@]*)@)?(?<file>.*?)(:?\d+)?...`)
the lazy loop exits EMPTY almost immediately (the zero-width guard correctly disables the
fast path) — its cost is unanchored-seed DFS overhead + greedy give-back, not loop
consumption. Where it does engage (narrow continuation first-set), loop consumption is a
small share of the pattern's total cost.
- Local single-JVM probe "wins" were noise. Rule reinforced: **local single-JVM probe
deltas are not perf evidence — box JMH + per-pattern C2-converged profiling only.**

## What this rules out / redirects
- Lazy-loop stack round-trips are NOT the lazy families' cost center.
- The real lever: a priority-correct (lazy leftmost-first) AND fast captureless DFA find in
the LazyDFA/RD lane — the chain lane proves leftmost-first lazy is achievable for simple
shapes; the fast-consume design above is the template if revisited (greedy fast-consume
fields + ctor shape detection in BitStateMatcher).
27 changes: 27 additions & 0 deletions .knowledgebase/memories/dead-classwriter-subclass.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
---
id: dead-classwriter-subclass
title: ASM ClassWriter.visit/visitMethod are final (9.10) — method emission cannot be intercepted via subclassing; wrap with a ClassVisitor instead
kind: dead-end
tags: [asm, classwriter, final, refuted-approach]
applies_to: []
source: multi-64k/dead-classwriter-subclass
status: active
recorded: 2026-09-28
valid_at: 2026-09-28
---



# ClassWriter subclassing is impossible — wrap a ClassVisitor

First design: `SplittingClassWriter extends ClassWriter`, override `visitMethod` to return a
buffering MethodVisitor. Died on compilation: ASM 9.10 declares `ClassWriter.visit` and
`visitMethod` `public final`. No interception is possible via subclassing.

The interception point must be a **wrapping `ClassVisitor`** (the standard ASM
instrumentation shape) delegating to the real ClassWriter, which stays the terminal pipeline
stage and `toByteArray()` producer. Chunks are emitted directly on the real writer; the
splitter takes `(node, realWriter, chunk0Visitor)`.

Any future ASM-pipeline work in this repo should start from a ClassVisitor wrapper, not a
ClassWriter subclass.
52 changes: 52 additions & 0 deletions .knowledgebase/memories/dead-env-gotchas.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
---
id: dead-env-gotchas
title: Build/test/harness gotchas for java-reggie and consumer-repo regex audits — stale-jar trap, gradle caching, piped-build chains, JDK-differential tests, JDK matcher statefulness (resolved — don't re-debug)
kind: dead-end
tags: [environment, git, gradle, classpath]
applies_to: []
source: dd_backend_fit/dead-env-gotchas
status: active
recorded: 2026-09-28
valid_at: 2026-09-28
---



# Environment gotchas hit (resolved — don't re-debug)

- Both consumer repos had stale `.git/index.lock` from crashed git processes →
pull failed with "an editor opened by git commit" message; fix: `rm .git/index.lock`.
- Large consumer repos (e.g. logs-backend): remote may carry dead refspecs
(plain fetch dies with "couldn't find remote ref …"; workaround
`git fetch origin <branch> && git merge --ff-only FETCH_HEAD`) and
filesystem-wide greps time out — ALWAYS `git ls-files '*.java' | xargs grep …`.
- reggie-runtime jar does NOT bundle ASM (declared `implementation`): any
standalone harness needs asm/asm-commons/asm-util 9.10.1 on the classpath
(from ~/.gradle/caches), else every compile fails with
NoClassDefFoundError MethodTooLargeException (misleading).
- Harness must emit results incrementally + per-entry timing; buffered output
+ one hung pattern (semver) loses 15 min of work.

Reggie build discipline (don't repeat):
- STALE-JAR TRAP: requesting only the runtime jar after editing reggie-codegen sources can
be an up-to-date NO-OP — the fat runtime jar embeds the codegen classes and did not
repackage. After codegen changes run :reggie-codegen:compileJava explicitly and
sanity-check via a changed generated-class hash.
- `./gradlew jar ... | grep BUILD` in an && chain does NOT stop on BUILD FAILED — grep
exits 0 on match and the following test/verify runs against the STALE jar. Check exit
codes explicitly; never pipe-build-then-run in one chain.
- Gradle test tasks are cached: `./gradlew test` returning BUILD SUCCESSFUL
in <1s means UP-TO-DATE, not re-run. For real verification use
`./gradlew cleanTest test`.
- git commit signing via ~/.ssh/datadog_git_commit_signing can fail
("agent refused operation") until the key passphrase is cached; retry after
unlocking (user unlocked on demand 2026-09-16).
- Writing regression tests: do NOT hand-derive expectations (repeated
test-assertion bugs from misread semantics); write JDK-differential tests
(compute expected values from java.util.regex at runtime).
- JDK Matcher is STATEFUL: matcher.matches() then matcher.find() continues
from the end of the previous match — use a fresh Matcher per operation in
differential tests (ReggieMatcher is stateless per call).
- jstack sampling root-causes compile hangs; run the victim via nohup+disown
in one tool call, sample in the next (the tool kills the process group
when the command ends).
36 changes: 36 additions & 0 deletions .knowledgebase/memories/dead-group-boundary-fusion.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
---
id: dead-group-boundary-fusion
title: Group-boundary state fusion in the capture-tracking DFS is net-negative — group writes sit on loop edges off the hot path, while the extra ops array taxes every push/pop
kind: dead-end
tags: [fusion, group-boundary, dfs, capture, perf, measured-negative, reverted]
applies_to: []
source: backend-ready/dead-group-boundary-fusion
status: active
recorded: 2026-09-28
valid_at: 2026-09-28
---


# Group-boundary state fusion in the capture-tracking DFS is net-negative

## What was tried
Fusing group enter/exit-only states into incoming edges so capture writes ride on edges
instead of dedicated marker frames. Two variants built and fully tested:
- EAGER: ops applied at push — pays ~150 wasted capture clones for never-popped
marker-stop frames on greedy paths.
- DEFERRED: ops row on the frame, applied at pop (sound: push pos == pop pos of the
bypassed state); post-write dedup strictly better.

## Why it was refuted (measured)
- Realistic 313-char accept shape: +1.4–2.4us; reject shape: +6.2–7.8us; only degenerate
5-char inputs win ~0.1us.
- Instrumentation: fusion removes only ~10 pops on the 313-char accept — the pre-release
loop body has NO boundary states (group writes sit on loop EDGES, off the winning path) —
while the extra ops array taxes every push/pop (~0.9ns x ~1700 frames); reject paths pay
on 5–10x more frames.
- State-id bit-packing of ops (~0.3–0.5ns/pop) also estimated net-negative for the same
reason.

## Verdict / redirect
Trades a negligible win on degenerate short inputs for microsecond-class losses on
realistic shapes. Short-input closure cost needs a hybrid engine, not fusion.
29 changes: 29 additions & 0 deletions .knowledgebase/memories/dead-handler-extent-heuristic.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
---
id: dead-handler-extent-heuristic
title: Exception-handler noCut zone must be {region start, region end−1, handler entry} only — a handler-extent heuristic (walk to first terminal) over-conservatively no-cuts the whole method for fall-through handlers
kind: dead-end
tags: [exception-handler, no-cut-zone, refuted-approach]
applies_to: []
source: multi-64k/dead-handler-extent-heuristic
status: active
recorded: 2026-09-28
valid_at: 2026-09-28
---



# noCut = positions separating {region start, end−1, handler entry} — not the handler extent

First attempt at protecting try/handler regions: noCut zone = protected region ∪ handler
*extent*, where the extent was computed by walking forward from the handler label to its first
`athrow`/`return`/`goto`/switch. Two flaws:

1. **Over-conservative**: a fall-through handler (catch sets a flag, then continues into the
method tail) has no terminal → the walk ran to the end of the method → the entire method
became no-cut → `chooseCuts` declined (observed: tryRegion test declined then verbatim
threw).
2. **Wrong model**: the handler body extending across a cut is *fine* — control flows through
the chunk tail chain; only the handler *entry label* must live in the region's chunk.

Correct rule: region + handler entry must share a chunk; bodies may tail-chain. noCut =
positions separating {start, end−1, handler}.
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
---
id: dead-real-generator-differential-test
title: No constructible pattern both overflows a generated method and compiles within the 10s deadline — an end-to-end overflow differential test is infeasible; splitter semantics must be covered at unit level
kind: dead-end
tags: [testing, differential, probe, ci-flake]
applies_to: []
source: multi-64k/dead-real-generator-differential-test
status: active
recorded: 2026-09-28
valid_at: 2026-09-28
---



# End-to-end overflow differential test is infeasible

Test-plan idea: "a pattern whose NFA step method overflows 64KB today compiles,
loads, and matches java.util.regex differentially." No such pattern could be constructed:

- Huge literal capture alternations route to BitStateMatcher (compact) or LiteralAlternation
trie strategies — method size never overflows.
- Pushing counts up (4000 alternatives) hits OOM in the analysis phase, before codegen.
- Compiles that do succeed take 8.5–13.5s against the 10s total-compile deadline — too close
for a stable CI test; the probe test was deleted rather than kept flaky.

Consequence: splitter semantics must be covered at unit level (load real classes so the JVM
verifier validates recomputed frames, check exact values) plus the runtime suite through the
wired-in pipeline. Don't spend more time hunting a triggering pattern within the current
strategy/limit envelope.
30 changes: 30 additions & 0 deletions .knowledgebase/memories/dead-sourceinterpreter-typing.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
---
id: dead-sourceinterpreter-typing
title: SourceInterpreter is unusable as slot-typing basis for large generated methods — copyOperation masks producers (ASTORE) and source sets grow quadratically at loop-carried merges (78k-insn method hung the test worker >120s)
kind: dead-end
tags: [asm, sourceinterpreter, refuted-approach, performance]
applies_to: []
source: multi-64k/dead-sourceinterpreter-typing
status: active
recorded: 2026-09-28
valid_at: 2026-09-28
---


# SourceInterpreter pathologies make it unusable for slot-typing large generated methods

The original typing plan derived live-slot types from `SourceInterpreter` frames (per-source
descriptors + unmask recursion for ASTORE/xLOAD). Two fatal flaws:

1. Quadratic set growth on loop-carried locals at merge points — a 78k-insn realistic test
method hung the test worker (>120s, jstack pinned `Analyzer.merge`, Analyzer.java:642).
Generated NFA-style methods are exactly this shape: each merge unions the previous
iteration's accumulated set with new defs → O(k) set unions per merge → O(n²) total.
2. ASTORE source masking forced an unmask-recursion layer (operand sources at the store, local
sources at loads) — `copyOperation` attributes values to the *copy instruction*, so a
local's type cannot be derived directly from its source insns. Built, working, then deleted
with the interpreter swap.

Consequence: use a descriptor-flowing interpreter instead (see the DescriptorInterpreter
memory). SourceInterpreter remains fine for small methods or one-shot analyses; the pathology
needs loop-carried locals surviving merge points.
35 changes: 35 additions & 0 deletions .knowledgebase/memories/find-api-surface-sufficient.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
---
id: find-api-surface-sufficient
title: ReggieMatcher runtime API covers every JDK Matcher op observed in the dd backends; grok SPI + shadow infra is the insertion seam (top-level reggie.ReggieMatcher facade exposes only matches/find)
kind: finding
tags: [api, compatibility, grok, spi, shadow]
applies_to: [reggie-runtime/src/main/java/com/datadoghq/reggie/runtime/ReggieMatcher.java, reggie-runtime/src/main/java/com/datadoghq/reggie/ReggieMatcher.java]
source: dd_backend_fit/find-api-surface-sufficient
status: active
recorded: 2026-09-28
valid_at: 2026-09-28
---



# API surface no longer a blocker; grok SPI + shadow infra is the insertion seam

Trunk `com.datadoghq.reggie.runtime.ReggieMatcher` covers every JDK Matcher op
observed in both repos: matches/find/findFrom, match/findMatch→MatchResult
(indexed+named groups, spans), matchInto/findMatchInto (alloc-free),
replaceFirst/replaceAll (literal + Function<MatchResult,String> — covers
appendReplacement loops and results() streams), split(input[,limit]), findAll,
cursor() with appendReplacement/appendTail.

Caveat: top-level `com.datadoghq.reggie.ReggieMatcher` facade exposes ONLY
matches/find — consumers must type against runtime.ReggieMatcher.

re2j maps 1:1 (`Pattern.matcher(input).find()/matches()`, groupCount, named
groups). logs-backend grok pipeline: production runs `JdkRegexPatternSupplier`
(re2j supplier is non-production); `GrokPatternBundle` supports a SHADOW
supplier with configurable shadow_ratio + `RegexMatcherShadowActor` — A/B
validation infra already built for a `ReggieRegexPatternSupplier`. The re2j
supplier's JDK→RE2 syntax-conversion hack (`(?P<` rewrite) becomes unnecessary.
Remaining JDK-adjacent needs: Pattern.quote (keep on JDK), asPredicate adapter,
Jackson Pattern deserialization (pb YAMLInsightProvider), Map<FrameField,
Pattern> adapter type.
31 changes: 31 additions & 0 deletions .knowledgebase/memories/find-asm-classwriter-final.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
---
id: find-asm-classwriter-final
title: ASM interception shape — ClassWriter.visit/visitMethod are final in 9.10, so pipeline instrumentation must wrap a ClassVisitor delegating to the terminal ClassWriter
kind: finding
tags: [asm, classwriter, classvisitor, interception]
applies_to: []
source: multi-64k/find-asm-classwriter-final
status: active
recorded: 2026-09-28
valid_at: 2026-09-28
---



# Wrap a ClassVisitor — ClassWriter methods are final

`javap` on asm-9.10.1.jar shows:
`public final void visit(int,int,String,String,String,String[])` and
`public final MethodVisitor visitMethod(int,String,String,String,String[])` in
`org.objectweb.asm.ClassWriter`. A `ClassWriter` subclass therefore cannot intercept method
emission ("overridden method is final" compile errors).

Consequence: the interception point must be a **wrapping `ClassVisitor`** (the standard ASM
instrumentation shape) delegating to the real ClassWriter, which stays the terminal pipeline
stage and `toByteArray()` producer. In the (now dropped) splitter patch, reggie's generators
had parameters typed `ClassWriter` and were mechanically widened to `ClassVisitor`
(~32 files); the only non-visitor use, `RecursiveDescentBytecodeGenerator.generate()`'s
internal `cw.toByteArray()`, was restructured to real-writer + front-visitor. Both pipeline
entry points (`RuntimeCompiler.generateBytecode`, `ReggieMatcherBytecodeGenerator`) construct
`new ClassWriter(COMPUTE_FRAMES|COMPUTE_MAXS)` wrapped by the front visitor — that wiring was
removed when the splitter was dropped and must be re-added on revival (see the L2 memories).
28 changes: 28 additions & 0 deletions .knowledgebase/memories/find-asm-toobytearray-sizecheck.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
---
id: find-asm-toobytearray-sizecheck
title: ASM MethodTooLargeException fires at toByteArray(), NOT at visitMaxs() — any method-size probe must call toByteArray() on its scratch writer; probe results that skip it silently over-report "fits"
kind: finding
tags: [asm, methodtoolargeexception, tobytearray, probe]
applies_to: []
source: multi-64k/find-asm-toobytearray-sizecheck
status: active
recorded: 2026-09-28
valid_at: 2026-09-28
---



# MethodTooLargeException fires at toByteArray(), not at visitMaxs()

Empirically verified with a scratch probe (25,000 × {NOP; ICONST_1; POP} ≈ 75,001 code bytes
into a `ClassWriter(0)`): `visitMaxs` returns normally, `visitEnd` returns normally,
`toByteArray()` throws `MethodTooLargeException` with `getCodeSize()=75001`. The COMPUTE_FRAMES
| COMPUTE_MAXS writers used by Reggie behave the same.

Consequences:
1. **A size probe must call `toByteArray()`** on its scratch writer. A probe that only emits +
visitMaxs silently reports 72–144KB methods as "fits" — the flush then takes the verbatim
path and the entry point's toByteArray throws later.
2. The real writer's guard surfaces at the entry point's `toByteArray()` → inside
`RuntimeCompiler`'s existing try/catch → L3 fallback preserved (a chunk-margin bug cannot
escape to a corrupt class, only to the fallback).
35 changes: 35 additions & 0 deletions .knowledgebase/memories/find-backend-prod-regex-cost.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
---
id: find-backend-prod-regex-cost
title: Backend regex CPU is concentrated in logs-processing grok-on-JDK matching; savings model = cost x share x (1-1/f) with every 1% of logs-processing CPU in JDK regex worth ~$50-60K/yr
kind: finding
tags: [prod-measurement, continuous-profiler, cloud-cost, savings-estimate, logs-processing, grok]
applies_to: []
source: backend-ready/find-backend-prod-regex-cost
status: active
recorded: 2026-09-28
valid_at: 2026-09-28
---


# Backend regex CPU is concentrated in logs-processing grok-on-JDK matching

## Measured distribution (prod, continuous profiles + CCM)
- logs-processing: java.util.regex self = 5.60% of service CPU, essentially ALL match-time
(compile 0.003%); driver ~100% grok (fsmatic GrokModule under Matcher entries); +0.55%
InterruptibleCharSequence.charAt input plumbing. com.google.re2j = 0.00% — a RE2J
migration had NOT landed in prod despite expectations.
- prof-analyzer: java.util.regex = 1.00% (JFR/pprof parsing).
- apm-processing: java.util.regex = 0.38% only — heavy regex there is RUST dd_sds/
regex-automata (~11%), NOT replaceable by a JDK-semantics engine.
- Continuous profiles are not reachable via the Datadog MCP toolset; pulled via the
dodo-cli flamegraph endpoint (env:prod, family java, cpu-time).

## Savings model
cost x share x (1 - 1/f), where f = reggie-vs-JDK speedup on the fleet mix. Initial
f=3–6x (repo benchmark corpus) overstated it; the real-corpus smoke (see
find-rust-engine-crossover) revised the honest central f to ~2x. Rule of thumb: every 1%
of logs-processing CPU in JDK regex = ~$50–60K/yr. Savings materialize only if HPA scales
replicas down (fleet ~50% of requests, diurnal autoscaling visible).

Full write-up: Datadog notebook 15576172; repo doc/investigations/multi-64k/evidence/
ev-reggie-cpu-savings-estimate.md.
Loading
Loading