Repository navigation
perf: eliminate slow paths (BitState routing, dotall-sink, SIMD fast-reject, PINNED_BACKREFERENCE) - #94
Conversation
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…er for embedding matchers Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…86 fast path Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- ThompsonBuilder(lazyAware=true): reversed epsilon order for lazy quantifiers gives PikeVM shortest-match semantics without separate backtracking - PatternAnalyzer: route safe lazy patterns to PIKEVM_CAPTURE (before requiresRecursiveDescent); set lazyNfa=true on MatchingStrategyResult - FallbackPatternDetector: B16 nullable-body guard extended to lazy repeatable quantifiers; hasCapturingGroupWithNullableBodyInRepeatableQuantifier - PikeVMMatcher: Option B useBoolFind path (findBoolPosFrom/boolEpsilonClose) skips capture tracking in find() for patterns with no assertions or backrefs - RuntimeCompiler: rebuild NFA with ThompsonBuilder(true) when lazyNfa=true; fix missing generateMatchFromMethod call in ONEPASS_NFA case - RegexParser: switch \N digit-collection to JDK stop-early rule — stop when running value exceeds totalGroupCount (fixes ()c|\10? fuzz divergence) - Divergence gate: 34 → 28; smoke fuzz: 0 findings Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add singleFirstCharAscii field: when exactly one ASCII char can begin a
match (e.g. '(' for LDAP), String.indexOf provides a JVM-intrinsified
SIMD scan that short-circuits before any DFA/NFA work. Eliminates the
per-char acquire-fenced LazyDFACache reads on ARM for no-match inputs.
LDAP no-match: 49K → 683K ops/ms (13.9x), now 17.6x over JDK.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…fast-reject
- Allow START_MULTILINE in findDfaEligible (previously rejected all non-START anchors)
- Split reinject closure: mid-line blocks START_MULTILINE; after-\n crosses it
- findStepClosureMultiline selects reinject based on c == '\n'
- sortedEpsilonClosure gains blockMultilineAnchor parameter
- hasMultilineFirstCharCandidate: SIMD indexOf('\n') + firstByteAscii scan
short-circuits find() before DFA for no-match inputs (no valid first char
at any line start); turns 0.63x JDK no-match into 2.68x
Benchmark vs JDK: match +4-10x, no-match +2.68x, RE2J ~20x slower throughout.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
findMatchResultFrom built a MatchResult on every accept. A greedy tail (e.g. (.*)) makes the top thread accept at every remaining position, so this allocated O(matchLen) throwaway results. Record the winner into the preallocated winCaptures and build once after the loop. Cuts COMMAND capture allocation 5040 -> 112 B/op (45x); 2149 tests pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
When the sole surviving capture thread is parked on a (?s).* tail with no competing consuming transition, jump straight to regionEnd instead of stepping every remaining char (O(tail) -> O(1)). Guarded by a precomputed sink table (computeSinkGroups): the epsilon closure must consume any char back into itself, reach accept, and contain no anchor/assertion/backref/ atomic state. Flat on the short COMMAND benchmark (prefix-dominated) but eliminates linear tail-stepping on long dotall tails. 7 new tests incl. must-not-fire cases (required suffix, non-dotall .). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Add BITSTATE_CAPTURE strategy (bounded-backtracking DFS capture engine) with eligibility gating in PatternAnalyzer, including exclusion of quantified bodies that are nullable and contain a nested capturing group (DFS visited-memo collapses the empty re-entry iteration before reaching the group's exit write, e.g. wrong span for (?:(.*[_]*))+). Also adds LazyDFACache.findEnd and wires it into PikeVMMatcher's find loop to bound the scan to the DFA-proven match end (scanLimit), independent of regionEnd's role in anchor semantics. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Single forward-scan backreference matching (no retry/backtracking) for patterns where the captured group's charset is provably disjoint from what follows it, e.g. \b(\w+)\s+\1\b. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
matches() gave up on the first candidate closing tag instead of continuing the search when it didn't reach end-of-input (wrong on e.g. "<a>x</a>y</a>"). findHtmlTagMatchFrom had the same class of bug for match() spans, picking the first valid occurrence instead of the greedy-correct last one. Fixing the retry logic led to a further insight: matches() requires full consumption, so the closing tag's position is algebraically fixed once the tag name length is known - no search needed at all. Replaced the indexOf-retry loop with a direct O(1) position check. 147x faster than JDK on adversarial match, 2816x on no-match (previously ~15% slower on match). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
computeSinkGroups checked epsilon-closure membership without crossing the consuming edge, so the fast-forward for (?s).* tails never fired since 4499480. Rewritten as a single fixed-point closure over epsilon + validated dotall edges. ~3.1x faster than JDK on long dotall tails once active. Also adds tests closing Codecov patch-coverage gaps flagged on PR #94.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f2795c8cf3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if (groupIndex == -1 || backrefIndex == -1 || groupIndex >= backrefIndex) { | ||
| return null; | ||
| } |
There was a problem hiding this comment.
Reject pinned shapes with ignored prefix/suffix nodes
This eligibility check lets \b(\w+)\s+\1\b route to PINNED_BACKREFERENCE even though the generated matcher only scans the group, optional separator, and backreference; it never evaluates the leading/trailing \b nodes. For example, find("xhello hello") or find("hello helloX") can succeed from the inner hello hello even though JDK rejects it because the word-boundary assertion fails. Either carry those nodes into codegen or reject non-empty prefix/suffix around the group/backref pair.
Useful? React with 👍 / 👎.
| return null; | ||
| } | ||
| QuantifierNode groupQuant = (QuantifierNode) capturingGroup.child; | ||
| if (!groupQuant.greedy || groupQuant.min < 1 || groupQuant.max != -1) { |
There was a problem hiding this comment.
Enforce the captured group's minimum length
Allowing groupQuant.min > 1 is unsound because the pinned bytecode later only checks that the scanned group is non-empty, not that it satisfies the quantifier's minimum. A pattern such as ([a-z]{2,})\d+\1 routes here, but the generated matcher accepts a1a with a one-character capture while java.util.regex correctly rejects it.
Useful? React with 👍 / 👎.
| separator = sepNodes.get(0); | ||
| separatorCharSet = getFirstCharSet(separator); |
There was a problem hiding this comment.
Preserve separator quantifier bounds
Collapsing any single separator node to just its first character set loses its quantifier bounds, while the generated separator scan accepts one or more characters until the charset changes. This makes patterns like ([a-z]+)\d{2,}\1 accept a1a and patterns like ([a-z]+)\d{1,2}\1 accept a123a, both diverging from JDK because the separator length constraints were ignored.
Useful? React with 👍 / 👎.
| mv.visitVarInsn(ILOAD, 2); | ||
| mv.visitVarInsn(ISTORE, startVar); |
There was a problem hiding this comment.
Clamp negative pinned search offsets
The generated findFrom copies the caller's start offset directly, so a negative start reaches the scan loop and calls charAt(-1). Other matchers and the existing clamping regression tests treat negative starts as 0; Reggie.compile("(\\w+):\\1").findFrom("a:a", -1) should return 0 instead of throwing. The analogous findMatchFrom path needs the same clamp before scanning.
Useful? React with 👍 / 👎.
| className, | ||
| "match", | ||
| "(Ljava/lang/String;)Lcom/datadoghq/reggie/runtime/MatchResult;", | ||
| false); |
There was a problem hiding this comment.
Shift pinned bounded match result spans
matchBounded delegates to match() on a substring and returns that result unchanged, so successful bounded matches report spans relative to the temporary substring instead of the original input. For a pinned pattern like (\w+):\1, matchBounded("xx aa:aa yy", 3, 8) returns start/end 0..5 even though the ReggieMatcher contract requires 3..8, which also makes group() slice the wrong backing string.
Useful? React with 👍 / 👎.
Require group/backref pair to span the whole pattern, enforce group/separator length bounds in generated bytecode, and fix matchBounded to report absolute (not substring-relative) spans. Patterns with anchors outside the span now correctly fall through to VARIABLE_CAPTURE_BACKREF.
What does this PR do?
Fast-path optimizations across the matching engine, most recently adding a new
PINNED_BACKREFERENCEmatching strategy: a single forward-scan (no retry/backtracking)matcher for backreference patterns where the captured group's charset is provably disjoint
from whatever follows it (e.g.
\b(\w+)\s+\1\b); plus two correctness bugs found and fixed inthe existing
SPECIALIZED_BACKREFERENCEHTML-tag-close path, which led to replacing itsmatches()retry loop with an O(1) closed-form check.Chronological summary of this branch's commits:
PINNED_BACKREFERENCEstrategy: newMatchingStrategy,PinnedBackreferenceInfocarrier,detectPinnedBackreferencedetector (disjointness proof via existing charset/CharSet.isDisjointtoolkit),
PinnedBackreferenceBytecodeGenerator, routing/fallback-guard/dispatch wiring, pluscorrectness tests and a JMH benchmark.
matches()correctness fixes + O(1) rewrite:matches()previously gave upon the first candidate closing tag instead of continuing the search when it didn't reach
end-of-input (wrong on e.g.
<a>x</a>y</a>— confirmed diverging from JDK); the rich-API spanhelper (
findHtmlTagMatchFrom) had the same class of bug, picking the first valid occurrenceinstead of the greedy-correct last one. Fixing this surfaced a further algorithmic insight:
matches()requires full input consumption, so the closing tag's position is algebraicallyfixed once the tag-name length is known — no search needed at all. Replaced the
indexOf-retry loop with a direct O(1) position check.capture engine (
BITSTATE_CAPTURE) with eligibility gating;LazyDFACache.findEndboundsPikeVM's find loop to the DFA-proven match end.
(?s).*tails jump straight toregionEndinstead ofstepping every remaining char (O(tail) → O(1)) when provably safe.
MatchResultonce, not per greedy give-back — cuts throwaway allocationon greedy-tail patterns (45x reduction on one measured case).
findDfato handle(?m)^via split reinject + line-start SIMD fast-reject.find()/findBoolPosFrom()viaString.indexOfintrinsic short-circuit.
PIKEVM_CAPTUREwith JDK-compatible\Nbackreference-digit parsing;divergence gate 34 → 28.
find(), embednameToIndexinbuildResult(), gateVarHandleacquire semantics behind an ARM/x86 flag.AGENTS.md/CLAUDE.md→ lean indexes).Motivation
Close remaining slow paths in the bytecode-generated matchers so more pattern shapes get
Reggie's native linear-time speed instead of falling back to backtracking-loop or JDK-delegated
paths. See per-commit messages above,
doc/2026-07-06-backreference-pinned-boundary-design.mdfor the
PINNED_BACKREFERENCEdesign rationale, anddoc/2026-07-06-html-tag-matches-retry-fix-design.mdfor the HTML-tag-close bug fixes andrewrite.
Related Issue(s)
N/A
Change Type
Checklist
./gradlew build)Performance Impact
PINNED_BACKREFERENCEvs JDK, JMH throughput (ops/ms), 2 forks × 5 warmup + 5 measurementiterations, adversarial "false start" inputs designed to stress a naive retry loop:
\b(\w+)\s+\1\b(routes toPINNED_BACKREFERENCE)\b(\w+)\s+\1\b(routes toPINNED_BACKREFERENCE)SPECIALIZED_BACKREFERENCEHTML-tag-close (<(\w+)>.*</\1>), same adversarial-input JMH setup,before vs. after the correctness fixes + O(1) rewrite (numbers pulled directly from the raw
results.json, not agent-summarized):Confidence intervals don't overlap for any row — real effects, not noise. JDK's own numbers are
stable across both runs (its code is untouched), confirming the measurement environment itself
didn't shift. The
matches()retry loop was replaced with a direct O(1) computed-position check:matches()requires full-input consumption, so once the tag-name length is known the closingtag's position is algebraically fixed (
len - group1Len - 3) — no search needed in eitherdirection. See
doc/2026-07-06-html-tag-matches-retry-fix-design.mdfor the derivation, the twocorrectness bugs this uncovered, and why the no-match case improved even more than match (the old
algorithm's worst case was exhaustive-scan-to-failure, which the O(1) check eliminates for both
outcomes equally).
Note on
re2j: no three-way reggie/JDK/re2j comparison is possible for these patterns —re2jhas no backreference support at all (com.google.re2j.Pattern.compile("\\b(\\w+)\\s+\\1\\b")throws
PatternSyntaxException: invalid escape sequence: \1, confirmed directly). Prior commitson this branch (SIMD fast-reject,
(?m)^DFA extension) do include re2j comparisons in their ownbenchmark suites for patterns re2j can compile — see individual commit messages above.
Full differential fuzzer gate stays at the pre-existing budget (28 findings, none new/attributable
to this branch, cross-checked against
doc/fuzz/2026-06-29.md's catalogue);StrategyCorrectnessMetaTest4/4 with
-Dreggie.metatest.enforce=true; a 5000-input randomized fuzz of<(\w+)>.*</\1>compareddirectly against JDK found 0 mismatches.
Additional Notes
SPECIALIZED_BACKREFERENCE's hardcodeddetectHTMLTagPattern/detectRepeatedWordPatterndetectors are intentionally left untouched — retiring them in favor of
PINNED_BACKREFERENCEis a follow-up decision, not part of this PR.
matches()trick generalizes beyond HTML/XML syntax (any fixed-delimiter,single-gap, single-backref shape under full-consumption matching), but as shipped it only
benefits the literal
</>/<//>pattern, since nothing generalizes the detector side —a real follow-up if this shape shows up with other delimiters (see
doc/2026-07-06-html-tag-matches-retry-fix-design.md§6).find()/findFrom()/findMatchFrom()/findMatch()for the HTML-tag shape are unaffected bythe O(1) rewrite (they don't require full consumption, so the gap length isn't algebraically
pinned down) but did get the correctness fix (greedy-longest instead of first-found).