Skip to content

perf: eliminate slow paths (BitState routing, dotall-sink, SIMD fast-reject, PINNED_BACKREFERENCE) - #94

Merged
jbachorik merged 14 commits into
mainfrom
perf/eliminate-slow-paths
Jul 7, 2026
Merged

jbachorik merged 14 commits into
mainfrom
perf/eliminate-slow-paths

Conversation

@jbachorik

@jbachorik jbachorik commented Jul 6, 2026 •

Copy link
Copy Markdown
Collaborator

What does this PR do?

Fast-path optimizations across the matching engine, most recently adding a new
PINNED_BACKREFERENCE matching strategy: a single forward-scan (no retry/backtracking)
matcher for backreference patterns where the captured group's charset is provably disjoint
from whatever follows it (e.g. \b(\w+)\s+\1\b); plus two correctness bugs found and fixed in
the existing SPECIALIZED_BACKREFERENCE HTML-tag-close path, which led to replacing its
matches() retry loop with an O(1) closed-form check.

Chronological summary of this branch's commits:

  • PINNED_BACKREFERENCE strategy: new MatchingStrategy, PinnedBackreferenceInfo carrier,
    detectPinnedBackreference detector (disjointness proof via existing charset/CharSet.isDisjoint
    toolkit), PinnedBackreferenceBytecodeGenerator, routing/fallback-guard/dispatch wiring, plus
    correctness tests and a JMH benchmark.
  • HTML-tag-close matches() correctness fixes + O(1) rewrite: matches() previously gave up
    on the first candidate closing tag instead of continuing the search when it didn't reach
    end-of-input (wrong on e.g. <a>x</a>y</a> — confirmed diverging from JDK); the rich-API span
    helper (findHtmlTagMatchFrom) had the same class of bug, picking the first valid occurrence
    instead of the greedy-correct last one. Fixing this surfaced a further algorithmic insight:
    matches() requires full input consumption, so the closing tag's position is algebraically
    fixed once the tag-name length is known — no search needed at all. Replaced the
    indexOf-retry loop with a direct O(1) position check.
  • BitState capture-engine routing + PikeVM span tightening — bounded-backtracking DFS
    capture engine (BITSTATE_CAPTURE) with eligibility gating; LazyDFACache.findEnd bounds
    PikeVM's find loop to the DFA-proven match end.
  • BitState capture-engine feasibility study (design doc).
  • Greedy dotall-sink fast-forward — (?s).* tails jump straight to regionEnd instead of
    stepping every remaining char (O(tail) → O(1)) when provably safe.
  • Build capture MatchResult once, not per greedy give-back — cuts throwaway allocation
    on greedy-tail patterns (45x reduction on one measured case).
  • Extend findDfa to handle (?m)^ via split reinject + line-start SIMD fast-reject.
  • SIMD single-char fast-reject in find()/findBoolPosFrom() via String.indexOf
    intrinsic short-circuit.
  • Lazy quantifier PIKEVM_CAPTURE with JDK-compatible \N backreference-digit parsing;
    divergence gate 34 → 28.
  • R1–R3a micro-perf: skip capture-array writes in find(), embed nameToIndex in
    buildResult(), gate VarHandle acquire semantics behind an ARM/x86 flag.
  • Docs slimming (AGENTS.md/CLAUDE.md → lean indexes).

Motivation

Close remaining slow paths in the bytecode-generated matchers so more pattern shapes get
Reggie's native linear-time speed instead of falling back to backtracking-loop or JDK-delegated
paths. See per-commit messages above, doc/2026-07-06-backreference-pinned-boundary-design.md
for the PINNED_BACKREFERENCE design rationale, and
doc/2026-07-06-html-tag-matches-retry-fix-design.md for the HTML-tag-close bug fixes and
rewrite.

Related Issue(s)

N/A

Change Type

  • Bug fix
  • New feature
  • Performance improvement
  • Refactoring (no functional change)
  • Documentation
  • Test improvement
  • Build/CI change

Checklist

  • I have read the CONTRIBUTING.md guidelines
  • All existing tests pass (./gradlew build)
  • I have added tests for my changes
  • I have updated documentation (if applicable)
  • My commits are signed

Performance Impact

PINNED_BACKREFERENCE vs JDK, JMH throughput (ops/ms), 2 forks × 5 warmup + 5 measurement
iterations, adversarial "false start" inputs designed to stress a naive retry loop:

Pattern shape Case Reggie JDK Result
\b(\w+)\s+\1\b (routes to PINNED_BACKREFERENCE) match 93.0 44.6 ~2.1x faster
\b(\w+)\s+\1\b (routes to PINNED_BACKREFERENCE) no-match 89.5 43.7 ~2.1x faster

SPECIALIZED_BACKREFERENCE HTML-tag-close (<(\w+)>.*</\1>), same adversarial-input JMH setup,
before vs. after the correctness fixes + O(1) rewrite (numbers pulled directly from the raw
results.json, not agent-summarized):

Case Reggie (before) Reggie (after) JDK (unchanged control) Result (after)
match 936.2 158,028.3 ± 33,890.8 1,072.0 ± 37.9 ~147x faster (was ~15% slower)
no-match 964.5 430,168.1 ± 26,269.4 152.8 ± 18.7 ~2,816x faster (was ~5.8x faster)

Confidence intervals don't overlap for any row — real effects, not noise. JDK's own numbers are
stable across both runs (its code is untouched), confirming the measurement environment itself
didn't shift. The matches() retry loop was replaced with a direct O(1) computed-position check:
matches() requires full-input consumption, so once the tag-name length is known the closing
tag's position is algebraically fixed (len - group1Len - 3) — no search needed in either
direction. See doc/2026-07-06-html-tag-matches-retry-fix-design.md for the derivation, the two
correctness bugs this uncovered, and why the no-match case improved even more than match (the old
algorithm's worst case was exhaustive-scan-to-failure, which the O(1) check eliminates for both
outcomes equally).

Note on re2j: no three-way reggie/JDK/re2j comparison is possible for these patterns —
re2j has no backreference support at all (com.google.re2j.Pattern.compile("\\b(\\w+)\\s+\\1\\b")
throws PatternSyntaxException: invalid escape sequence: \1, confirmed directly). Prior commits
on this branch (SIMD fast-reject, (?m)^ DFA extension) do include re2j comparisons in their own
benchmark suites for patterns re2j can compile — see individual commit messages above.

Full differential fuzzer gate stays at the pre-existing budget (28 findings, none new/attributable
to this branch, cross-checked against doc/fuzz/2026-06-29.md's catalogue); StrategyCorrectnessMetaTest
4/4 with -Dreggie.metatest.enforce=true; a 5000-input randomized fuzz of <(\w+)>.*</\1> compared
directly against JDK found 0 mismatches.

Additional Notes

  • SPECIALIZED_BACKREFERENCE's hardcoded detectHTMLTagPattern/detectRepeatedWordPattern
    detectors are intentionally left untouched — retiring them in favor of PINNED_BACKREFERENCE
    is a follow-up decision, not part of this PR.
  • The O(1) matches() trick generalizes beyond HTML/XML syntax (any fixed-delimiter,
    single-gap, single-backref shape under full-consumption matching), but as shipped it only
    benefits the literal </>/<//> pattern, since nothing generalizes the detector side —
    a real follow-up if this shape shows up with other delimiters (see
    doc/2026-07-06-html-tag-matches-retry-fix-design.md §6).
  • find()/findFrom()/findMatchFrom()/findMatch() for the HTML-tag shape are unaffected by
    the O(1) rewrite (they don't require full consumption, so the gap length isn't algebraically
    pinned down) but did get the correctness fix (greedy-longest instead of first-found).

jbachorik and others added 12 commits July 3, 2026 10:10
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…er for embedding matchers

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…86 fast path

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- ThompsonBuilder(lazyAware=true): reversed epsilon order for lazy quantifiers
  gives PikeVM shortest-match semantics without separate backtracking
- PatternAnalyzer: route safe lazy patterns to PIKEVM_CAPTURE (before
  requiresRecursiveDescent); set lazyNfa=true on MatchingStrategyResult
- FallbackPatternDetector: B16 nullable-body guard extended to lazy
  repeatable quantifiers; hasCapturingGroupWithNullableBodyInRepeatableQuantifier
- PikeVMMatcher: Option B useBoolFind path (findBoolPosFrom/boolEpsilonClose)
  skips capture tracking in find() for patterns with no assertions or backrefs
- RuntimeCompiler: rebuild NFA with ThompsonBuilder(true) when lazyNfa=true;
  fix missing generateMatchFromMethod call in ONEPASS_NFA case
- RegexParser: switch \N digit-collection to JDK stop-early rule — stop when
  running value exceeds totalGroupCount (fixes ()c|\10? fuzz divergence)
- Divergence gate: 34 → 28; smoke fuzz: 0 findings

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add singleFirstCharAscii field: when exactly one ASCII char can begin a
match (e.g. '(' for LDAP), String.indexOf provides a JVM-intrinsified
SIMD scan that short-circuits before any DFA/NFA work. Eliminates the
per-char acquire-fenced LazyDFACache reads on ARM for no-match inputs.

LDAP no-match: 49K → 683K ops/ms (13.9x), now 17.6x over JDK.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…fast-reject

- Allow START_MULTILINE in findDfaEligible (previously rejected all non-START anchors)
- Split reinject closure: mid-line blocks START_MULTILINE; after-\n crosses it
- findStepClosureMultiline selects reinject based on c == '\n'
- sortedEpsilonClosure gains blockMultilineAnchor parameter
- hasMultilineFirstCharCandidate: SIMD indexOf('\n') + firstByteAscii scan
  short-circuits find() before DFA for no-match inputs (no valid first char
  at any line start); turns 0.63x JDK no-match into 2.68x

Benchmark vs JDK: match +4-10x, no-match +2.68x, RE2J ~20x slower throughout.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
findMatchResultFrom built a MatchResult on every accept. A greedy tail
(e.g. (.*)) makes the top thread accept at every remaining position, so
this allocated O(matchLen) throwaway results. Record the winner into the
preallocated winCaptures and build once after the loop. Cuts COMMAND
capture allocation 5040 -> 112 B/op (45x); 2149 tests pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
When the sole surviving capture thread is parked on a (?s).* tail with no
competing consuming transition, jump straight to regionEnd instead of
stepping every remaining char (O(tail) -> O(1)). Guarded by a precomputed
sink table (computeSinkGroups): the epsilon closure must consume any char
back into itself, reach accept, and contain no anchor/assertion/backref/
atomic state. Flat on the short COMMAND benchmark (prefix-dominated) but
eliminates linear tail-stepping on long dotall tails. 7 new tests incl.
must-not-fire cases (required suffix, non-dotall .).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Add BITSTATE_CAPTURE strategy (bounded-backtracking DFS capture engine)
with eligibility gating in PatternAnalyzer, including exclusion of
quantified bodies that are nullable and contain a nested capturing
group (DFS visited-memo collapses the empty re-entry iteration before
reaching the group's exit write, e.g. wrong span for (?:(.*[_]*))+).

Also adds LazyDFACache.findEnd and wires it into PikeVMMatcher's find
loop to bound the scan to the DFA-proven match end (scanLimit),
independent of regionEnd's role in anchor semantics.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Single forward-scan backreference matching (no retry/backtracking) for
patterns where the captured group's charset is provably disjoint from
what follows it, e.g. \b(\w+)\s+\1\b.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@jbachorik jbachorik added the AI Generated or assisted by AI label Jul 6, 2026
@codecov-commenter

codecov-commenter commented Jul 6, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 95.94395% with 55 lines in your changes missing coverage. Please review.
✅ Project coverage is 83.8%. Comparing base (eedc7c1) to head (f2795c8).

Files with missing lines Patch % Lines
...doghq/reggie/codegen/analysis/PatternAnalyzer.java 87.1% 11 Missing and 14 partials ⚠️
...ava/com/datadoghq/reggie/runtime/LazyDFACache.java 71.7% 4 Missing and 7 partials ⚠️
.../com/datadoghq/reggie/runtime/RuntimeCompiler.java 85.1% 2 Missing and 5 partials ⚠️
...ggie/codegen/analysis/FallbackPatternDetector.java 95.5% 0 Missing and 3 partials ⚠️
...doghq/reggie/codegen/analysis/PatternDebugger.java 0.0% 3 Missing ⚠️
...q/reggie/codegen/codegen/NFABytecodeGenerator.java 98.0% 0 Missing and 2 partials ⚠️
...oghq/reggie/codegen/automaton/ThompsonBuilder.java 98.6% 0 Missing and 1 partial ⚠️
.../datadoghq/reggie/codegen/parsing/RegexParser.java 80.0% 0 Missing and 1 partial ⚠️
...ggie/processor/ReggieMatcherBytecodeGenerator.java 94.4% 0 Missing and 1 partial ⚠️
...ghq/reggie/runtime/LinearTokenSequenceMatcher.java 88.8% 0 Missing and 1 partial ⚠️
Additional details and impacted files
@@            Coverage Diff            @@
##              main     #94     +/-   ##
=========================================
+ Coverage     82.6%   83.8%   +1.1%     
  Complexity       1       1             
=========================================
  Files          130     134      +4     
  Lines        39979   41222   +1243     
  Branches      5225    5522    +297     
=========================================
+ Hits         33042   34560   +1518     
+ Misses        5282    4957    -325     
- Partials      1655    1705     +50     
Files with missing lines Coverage Δ
...ggie/codegen/analysis/PinnedBackreferenceInfo.java 100.0% <100.0%> (ø)
...datadoghq/reggie/codegen/automaton/ProductDFA.java 100.0% <100.0%> (ø)
...odegen/codegen/BackreferenceBytecodeGenerator.java 70.1% <100.0%> (+18.9%) ⬆️
...ggie/codegen/codegen/OnePassBytecodeGenerator.java 99.3% <100.0%> (-0.1%) ⬇️
.../codegen/PinnedBackreferenceBytecodeGenerator.java 100.0% <100.0%> (ø)
...datadoghq/reggie/runtime/BackreferenceHelpers.java 32.8% <100.0%> (+6.4%) ⬆️
.../com/datadoghq/reggie/runtime/BitStateMatcher.java 100.0% <100.0%> (ø)
...va/com/datadoghq/reggie/runtime/PikeVMMatcher.java 87.5% <ø> (-2.0%) ⬇️
...va/com/datadoghq/reggie/runtime/ReggieMatcher.java 82.6% <100.0%> (+0.7%) ⬆️
...oghq/reggie/codegen/automaton/ThompsonBuilder.java 97.2% <98.6%> (+6.5%) ⬆️
... and 9 more

... and 8 files with indirect coverage changes


Continue to review full report in Codecov by Harness.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update eedc7c1...f2795c8. Read the comment docs.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

jbachorik and others added 2 commits July 6, 2026 22:34
matches() gave up on the first candidate closing tag instead of
continuing the search when it didn't reach end-of-input (wrong on
e.g. "<a>x</a>y</a>"). findHtmlTagMatchFrom had the same class of
bug for match() spans, picking the first valid occurrence instead
of the greedy-correct last one.

Fixing the retry logic led to a further insight: matches() requires
full consumption, so the closing tag's position is algebraically
fixed once the tag name length is known - no search needed at all.
Replaced the indexOf-retry loop with a direct O(1) position check.
147x faster than JDK on adversarial match, 2816x on no-match
(previously ~15% slower on match).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
computeSinkGroups checked epsilon-closure membership without crossing the
consuming edge, so the fast-forward for (?s).* tails never fired since
4499480. Rewritten as a single fixed-point closure over epsilon + validated
dotall edges. ~3.1x faster than JDK on long dotall tails once active.

Also adds tests closing Codecov patch-coverage gaps flagged on PR #94.
@jbachorik
jbachorik marked this pull request as ready for review July 7, 2026 10:05
@jbachorik
jbachorik merged commit 6841723 into main Jul 7, 2026
9 checks passed
@jbachorik
jbachorik deleted the perf/eliminate-slow-paths branch July 7, 2026 10:06

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f2795c8cf3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +3574 to +3576
if (groupIndex == -1 || backrefIndex == -1 || groupIndex >= backrefIndex) {
return null;
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reject pinned shapes with ignored prefix/suffix nodes

This eligibility check lets \b(\w+)\s+\1\b route to PINNED_BACKREFERENCE even though the generated matcher only scans the group, optional separator, and backreference; it never evaluates the leading/trailing \b nodes. For example, find("xhello hello") or find("hello helloX") can succeed from the inner hello hello even though JDK rejects it because the word-boundary assertion fails. Either carry those nodes into codegen or reject non-empty prefix/suffix around the group/backref pair.

Useful? React with 👍 / 👎.

return null;
}
QuantifierNode groupQuant = (QuantifierNode) capturingGroup.child;
if (!groupQuant.greedy || groupQuant.min < 1 || groupQuant.max != -1) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Enforce the captured group's minimum length

Allowing groupQuant.min > 1 is unsound because the pinned bytecode later only checks that the scanned group is non-empty, not that it satisfies the quantifier's minimum. A pattern such as ([a-z]{2,})\d+\1 routes here, but the generated matcher accepts a1a with a one-character capture while java.util.regex correctly rejects it.

Useful? React with 👍 / 👎.

Comment on lines +3613 to +3614
separator = sepNodes.get(0);
separatorCharSet = getFirstCharSet(separator);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve separator quantifier bounds

Collapsing any single separator node to just its first character set loses its quantifier bounds, while the generated separator scan accepts one or more characters until the charset changes. This makes patterns like ([a-z]+)\d{2,}\1 accept a1a and patterns like ([a-z]+)\d{1,2}\1 accept a123a, both diverging from JDK because the separator length constraints were ignored.

Useful? React with 👍 / 👎.

Comment on lines +190 to +191
mv.visitVarInsn(ILOAD, 2);
mv.visitVarInsn(ISTORE, startVar);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Clamp negative pinned search offsets

The generated findFrom copies the caller's start offset directly, so a negative start reaches the scan loop and calls charAt(-1). Other matchers and the existing clamping regression tests treat negative starts as 0; Reggie.compile("(\\w+):\\1").findFrom("a:a", -1) should return 0 instead of throwing. The analogous findMatchFrom path needs the same clamp before scanning.

Useful? React with 👍 / 👎.

Comment on lines +435 to +438
className,
"match",
"(Ljava/lang/String;)Lcom/datadoghq/reggie/runtime/MatchResult;",
false);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Shift pinned bounded match result spans

matchBounded delegates to match() on a substring and returns that result unchanged, so successful bounded matches report spans relative to the temporary substring instead of the original input. For a pinned pattern like (\w+):\1, matchBounded("xx aa:aa yy", 3, 8) returns start/end 0..5 even though the ReggieMatcher contract requires 3..8, which also makes group() slice the wrong backing string.

Useful? React with 👍 / 👎.

@jbachorik jbachorik mentioned this pull request Jul 7, 2026
4 of 12 tasks
jbachorik added a commit that referenced this pull request Jul 7, 2026
Require group/backref pair to span the whole pattern, enforce
group/separator length bounds in generated bytecode, and fix
matchBounded to report absolute (not substring-relative) spans.
Patterns with anchors outside the span now correctly fall through
to VARIABLE_CAPTURE_BACKREF.
@jbachorik jbachorik added this to the 0.4.0 milestone Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

AI Generated or assisted by AI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants