diff --git a/README.md b/README.md index 90d52672..5fac33b2 100644 --- a/README.md +++ b/README.md @@ -665,6 +665,15 @@ Pattern Analysis Decision Tree: └─ Unsupported features? ─────────────────────► Compile-time error ``` +Beyond this top-level routing, some structural shapes get a dedicated fast path instead of a +general-purpose engine. For example, `BITSTATE_BYTECODE` recognizes patterns of the form +`^(?:leadingWs(kw1|kw2|...)separatorWs)?mandatoryCharSet+trailingWs*(tail)` (an optional +prefix keyword plus a mandatory scan and tail — e.g. shell-command-style patterns) and +compiles them to straight-line, non-backtracking bytecode rather than routing through the +general `BITSTATE_CAPTURE` interpreter. See +[doc/2026-07-08-bitstate-bytecode-generator-design.md](doc/2026-07-08-bitstate-bytecode-generator-design.md) +for the design rationale. + ### Generated Code Examples #### Literal Pattern (`hello`) diff --git a/doc/2026-07-08-bitstate-bytecode-generator-design.md b/doc/2026-07-08-bitstate-bytecode-generator-design.md new file mode 100644 index 00000000..9a312637 --- /dev/null +++ b/doc/2026-07-08-bitstate-bytecode-generator-design.md @@ -0,0 +1,239 @@ +# Design: BitState bytecode generator (revisiting "interpreter, not codegen") + +Date: 2026-07-08 +Status: narrowed and in implementation. See §4a for the scope-narrowing decision made +during implementation (general NFA-splitting deferred; first cut recognizes a specific +structural family directly). + +## 1. Why this doc exists + +`doc/2026-07-03-bitstate-capture-engine-feasibility.md` §3 deliberately chose an interpreter +over a bytecode generator for BitState, to minimize risk for the P1 rollout: + +> "Interpreter, not codegen... This is the lowest-risk integration and avoids touching +> `RuntimeCompiler`'s bytecode paths." + +That shipped as `BitStateMatcher.java` (job-stack DFS + generation-stamped visited bitmap + +undo-log capture restoration, interpreting the shared `NFA` object). Separately, +`doc/2026-07-06-bitstate-span-tightening-design.md` is an in-flight, not-yet-implemented +follow-up that narrows the *span* the interpreter searches over — a real but bounded win +(less redundant work over trailing input), not a change to per-character constant-factor +overhead. + +This doc presents new evidence that the interpreter's per-character constant-factor +overhead — not just span length — is a large, separate cost, and proposes a narrow, +opt-in bytecode path for BitState-eligible patterns whose backtracking is shallow enough to +make generation tractable, rather than reopening the general case. + +## 2. Evidence: hand-written prototype ceiling + +A throwaway, hand-written (non-generated) specialized matcher for the IAST `COMMAND` pattern +(`(?s)(?m)^(?:\s*(?:sudo|doas)\s+)?\b\S+\b\s*(.*)`, currently routed to `BITSTATE_CAPTURE`) +was built to measure the ceiling of eliminating interpreter overhead: direct `char` +comparisons and plain loops instead of `CharSet.contains()` dispatch, job-stack array +indirection, and generic NFA-state lookup. Verified byte-for-byte correct against +`java.util.Pattern` across 11 edge cases (anchors, MULTILINE re-anchoring, empty input, +false-keyword-prefix rejection). JMH (`:reggie-benchmark:jmh`, JDK 21): + +| engine | find (ops/ms) | capture (ops/ms) | +|---|---:|---:| +| hand-written prototype | 92,234 | 92,948 | +| JDK `Pattern` | 20,995 | 20,843 | +| Reggie `BITSTATE_CAPTURE` (interpreter) | 1,809 | 1,668 | +| RE2J | 711 | 345 | + +The prototype is ~51x faster than the current interpreter and ~4.4x faster than JDK on this +pattern. Since the prototype uses the *same* algorithmic shape as `BitStateMatcher` (locate +the mandatory prefix, one optional-group backtrack point, then a dotall tail) with no +asymptotic advantage, the gap is attributable to interpreter overhead: `CharSet` object +dispatch per character, `int[]` job-stack push/pop per state transition, and NFA-state +array indirection per step. + +**Caveat (load-bearing for the scope decision in §4):** this pattern's backtracking is +shallow — exactly one optional-group fallback, no nested or combinatorial choice points. +Patterns with deeper alternation/backtracking structure were not measured and are not +assumed to see the same multiplier; a general bytecode generator would need to reproduce +the interpreter's explicit backtrack stack and undo-log in generated form, which is a +materially harder codegen problem than `OnePassBytecodeGenerator`'s straight-line +switch-per-state (OnePass has no choice points to encode at all). + +## 3. What makes this harder than `OnePassBytecodeGenerator` + +`OnePassBytecodeGenerator` works because OnePass-eligible patterns are unambiguous: at most +one valid transition per state, so codegen is a straight-line `switch`-per-state loop with +no backtracking. BitState exists precisely for ambiguous patterns, so a bytecode version +must additionally encode, per generated method: + +- **Choice points.** Where the interpreter pushes competing jobs onto an explicit stack in + priority order, generated code needs an equivalent — either recursive method calls (using + the JVM call stack as the backtrack stack) or an explicit local/instance array mirroring + `BitStateMatcher.stackA/B/C`. +- **Capture undo-log.** `BitStateMatcher` pushes a `RESTORE` job before overwriting a + capture slot, so a failed branch's writes are undone on backtrack. Recursive-call codegen + gets this for free (save the old value in a local before recursing, restore after the + call returns false) — this is the strongest argument for choosing recursive-descent + codegen over explicit-stack codegen for this generator specifically. +- **Visited-bitmap / ReDoS bound — NOT required at the ≤1-choice-point tier (revised).** + The interpreter's visited bitmap exists to bound work when the *same* `(stateId, pos)` + pair is reachable via multiple different combinations of backtracking decisions — which + only happens when choice points nest or repeat (inside a `*`/`+` loop, or two independent + optionals compounding). Eligibility for this generator (§4) is capped at exactly one + choice point, not nested inside any repeating quantifier, so no `(stateId, pos)` pair can + ever be revisited: at most two straight-line passes are attempted (with the optional + element, then without), bounding total work to `O(2 × pattern length)` at *compile* time, + not run time. Adding a bitmap here would reintroduce the exact dynamic dispatch overhead + this generator exists to eliminate, for no safety benefit — there is nothing unsafe to + bound. +- **Budget fallback — NOT required, for the same reason.** The interpreter's + `exceedsBudget(spanLen)`/PikeVM-delegation exists to catch visited-bitmap blowup. Since + the ≤1-choice-point tier has no such blowup by construction (bounded compile-time work, + independent of input length), there is no budget to exceed and no fallback path needed. + Generated classes for this tier are pure "always-native" classes, the same shape as + `ONEPASS_NFA` classes — this is *only* true because eligibility is capped this tightly; + it would need to be revisited if the choice-point bound is ever widened (§4, "Open + questions"). +- **Mandatory (non-optional) quantifiers still need a narrow backtrack-free proof.** A + quantifier with `min ≥ 1` (e.g. `\S+`) is not a "choice point" under §4's semantic + definition (it has no distinct branches), but it can still in principle need to give back + consumed characters if what follows it would otherwise fail to match. The prototype's + `\S+` never needs to because everything after it (`\s*` then an unconstrained dotall + `.*` capture) can absorb any remaining input regardless of how much `\S+` consumed. This + "the suffix always succeeds" property is not decidable from local AST shape alone in + general; rather than general dataflow analysis, eligibility treats it as a narrow + syntactic special case: a mandatory quantifier is accepted as backtrack-free only when + everything following it is optional whitespace/anchors terminating in an unconstrained + trailing `.*`-style capture (the exact COMMAND shape). Anything else declines eligibility + and falls back to the interpreter. + +None of this makes a bytecode generator infeasible — it makes it a different, larger shape +than `OnePassBytecodeGenerator`, closer in spirit to generating the interpreter's own +control flow than to eliminating it. + +## 4. Proposed scope: narrow, opt-in, falls back to the interpreter + +Given §3, this proposal does **not** replace `BitStateMatcher` or attempt to +bytecode-generate every `BITSTATE_CAPTURE`-eligible pattern. It adds a second, stricter +eligibility tier: + +``` +OnePass (unambiguous) — existing, unchanged + ↓ not one-pass +BitState-bytecode (ambiguous, SHALLOW backtracking) — NEW, this proposal + ↓ not shallow-eligible +BitState-interpreter (ambiguous, bounded budget) — existing, unchanged + ↓ budget exceeded +PikeVM (ambiguous, any size) — existing, unchanged +``` + +"Shallow backtracking" eligibility (a new, stricter predicate layered on top of the +existing `isBitStateEligible`) is deliberately conservative for a first cut: **at most one +semantic choice point** on any path, where a choice point is counted only when the +take-or-skip (or which-alternative) decision is *continuation-dependent* — its correctness +can only be confirmed by whether the rest of the pattern subsequently matches, so it can only +be resolved by attempting a full downstream match and retrying on failure. An alternation +whose branches are distinguishable by local lookahead and lead to a shared continuation +(e.g. `sudo|doas`, disjoint first characters, identical code after either matches) is **not** +a choice point under this definition — it compiles to a plain sequential +try-this-then-that with no retry-the-continuation step, exactly as the hand-written +prototype's `matchKeyword` does. This is intentionally narrower than "everything BitState +currently handles" — widening it is future work, done incrementally as each additional +shape is measured, not assumed. + +Because eligibility is capped at exactly one continuation-dependent choice point, and that +choice point may not sit inside any repeating quantifier, codegen does not need general +recursive backtracking, an undo-log, a visited bitmap, or a dynamic PikeVM-budget fallback +(§3, revised) — the generated shape is two independent straight-line methods (one per +branch of the single choice) selected by a single `if`, matching the hand-written +prototype's structure exactly. This is a stricter and simpler codegen target than the +original recursive-descent sketch; if the choice-point bound is ever widened beyond one +(future work, not this proposal), the recursive-descent-with-undo-log approach and the +visited-bitmap/budget-fallback requirements from §3 would need to be reinstated. + +## 4a. Scope-narrowing decision made during implementation + +While implementing §5 step 2, composing two independently bytecode-generated sub-matchers +generically (split an arbitrary eligible NFA at its one choice point, generate each branch, +glue them with correct Perl-priority semantics for unanchored `find`) turned out to need a +new sub-NFA-splitting/composition layer, not just "two straight-line methods and an `if`" +as §4 assumed — `ReggieMatcher`'s public API has no "attempt an anchored match starting at +exactly this position" hook that two separately-generated `OnePassBytecodeGenerator` +instances could share to interleave branch priority per candidate start position during an +unanchored scan. + +**Decision: defer the general NFA-splitting/composition layer. The first implementation +recognizes one specific structural family directly from the AST** — the exact shape +validated by the prototype and this doc's motivating evidence (§2): an optional, +lookahead-disjoint literal-or-charclass prefix, a mandatory backtrack-free affix (§3's +narrow syntactic special case), and an unconstrained trailing `.*`-style capture. This is +implemented as a dedicated `PatternAnalyzer` detector (`detectXxx(ast)` returning a +descriptor `PatternInfo`, e.g. `PrefixGuardedScanInfo`) in the same style as the existing +`detectStatelessPattern`/`StatelessPatternInfo` and `detectCountingGlushkov`/ +`CountingGlushkovInfo` pairs already in `PatternAnalyzer` — not a generic NFA-consuming +generator like `OnePassBytecodeGenerator`. The detector extracts literals/char-sets from +the AST (data-driven), so it covers any pattern matching this *shape*, not just the literal +COMMAND pattern string — but it declines (falls back to the `BITSTATE_CAPTURE` interpreter) +for any other ≤1-choice-point pattern shape, including ones that would satisfy §4's +semantic choice-point definition but aren't structured as +`prefix? affix (.* capture)`. + +**What this means going forward:** "general ≤1-choice-point NFA bytecode generation" (the +NFA-splitting/composition layer described in §4/§5 step 2 originally) is explicitly +deferred future work, not implemented in this pass. Widening coverage beyond the recognized +structural family should be done by adding new `detectXxx`/`XxxInfo` pairs for additional +*measured* shapes (mirroring this one), or by eventually building the general +splitting/composition layer — whichever a future session decides, informed by which +additional shapes actually show up in real IAST/production patterns. This mirrors how +`PatternAnalyzer` already handles most of its fast-path tiers (`STATELESS_LOOP`, +`COUNTING_GLUSHKOV`, literal alternation, etc.): one detector per recognized shape, not one +generic engine for the whole eligibility class. + +## 5. Concrete plan (as narrowed by §4a) + +1. **`PatternAnalyzer`** (`reggie-codegen/.../analysis/PatternAnalyzer.java`): add + `detectPrefixGuardedScan(ast)` returning a `PrefixGuardedScanInfo implements PatternInfo` + descriptor or `null`, in the same style as `detectStatelessPattern`/ + `StatelessPatternInfo` and `detectCountingGlushkov`/`CountingGlushkovInfo`. Wire into + `routeBitState` as a second substitution: when `result.strategy == BITSTATE_CAPTURE` and + `detectPrefixGuardedScan(ast)` returns non-null, substitute `BITSTATE_BYTECODE` with that + descriptor as `patternInfo`. +2. **`BitStateBytecodeGenerator`** (new, + `reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/codegen/BitStateBytecodeGenerator.java`), + consumes a `PrefixGuardedScanInfo` descriptor (not an `NFA`) and emits the two + straight-line branches (with-prefix, without-prefix) directly, mirroring + `CommandPatternPrototype`'s logic but data-driven from the descriptor's extracted + literals/char-sets. No general NFA-splitting/composition layer (deferred, §4a). Reuses + `BytecodeUtil.pushInt`, `LocalVarAllocator`, and `generateCharSetCheck` from the existing + codegen package. +3. **`RuntimeCompiler`**: add a `case BITSTATE_BYTECODE:` arm in the `generateBytecode` + switch (currently `BITSTATE_CAPTURE` never reaches this switch — see research notes, + §"RuntimeCompiler" below). No fallback `PikeVMMatcher` field is needed (§3, revised) — + generated classes for this tier are pure always-native classes, the same shape as + `ONEPASS_NFA`/`STATELESS_LOOP` classes. +4. **Dual-path rule** (AGENTS.md #2): mirror the same routing/generation change in + `ReggieMatcherBytecodeGenerator.java` (`reggie-processor`). +5. **`StructuralHash`** (AGENTS.md #3): add `PrefixGuardedScanInfo`'s fields if the + structural-hash class-cache path is used for this strategy. +6. **Tests**: correctness parity suite against COMMAND and at least one structurally + similar synthetic pattern (different literals/char-classes) to confirm the detector is + genuinely data-driven and not COMMAND-string-specific, plus negative tests confirming + patterns just outside the recognized shape correctly decline to `BITSTATE_CAPTURE` + (interpreter) rather than misfiring. No ReDoS/budget test is needed for this tier (§3, + revised: work is compile-time-bounded, not input-dependent). +7. **Benchmark**: extend `IastRegexpBenchmark` with `bitstateBytecodeXxx` methods (as the + prototype's `protoCommandFind/Capture` did) to measure the real generator against the + prototype ceiling, JDK, interpreter, and RE2J. + +## 6. Open questions for reviewer sign-off before implementation + +- Is the semantic (continuation-dependent) choice-point definition (§4) correctly and + completely specified for a mechanical AST check, or are there pattern shapes where + "distinguishable by local lookahead with a shared continuation" is ambiguous to decide + syntactically and needs a conservative decline-to-generate default? +- Is the narrow syntactic special-case for backtrack-free mandatory quantifiers (§3: "must + be followed only by optional whitespace/anchors terminating in an unconstrained trailing + `.*`-style capture") too narrow, too broad, or exactly right for a first cut? This is the + main correctness-risk surface since it's a heuristic, not a general proof. +- Should `BITSTATE_BYTECODE` reuse `BITSTATE_NFA_CACHE`'s pattern-keyed caching model, or + move to the structural-hash class-cache model used by other bytecode-backed strategies + (§"RuntimeCompiler" research notes) — the latter is more consistent with how every other + bytecode strategy is cached today. diff --git a/reggie-benchmark/src/main/java/com/datadoghq/reggie/benchmark/CommandPatternPrototype.java b/reggie-benchmark/src/main/java/com/datadoghq/reggie/benchmark/CommandPatternPrototype.java new file mode 100644 index 00000000..0026c7eb --- /dev/null +++ b/reggie-benchmark/src/main/java/com/datadoghq/reggie/benchmark/CommandPatternPrototype.java @@ -0,0 +1,128 @@ +/* + * Copyright 2026-Present Datadog, Inc. + * + * Licensed under the Apache License, Version 2.0 (the "License"); + * you may not use this file except in compliance with the License. + * You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, software + * distributed under the License is distributed on an "AS IS" BASIS, + * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. + * See the License for the specific language governing permissions and + * limitations under the License. + */ +package com.datadoghq.reggie.benchmark; + +/** + * Hand-written specialized matcher for a single pattern: {@code + * (?s)(?m)^(?:\s*(?:sudo|doas)\s+)?\b\S+\b\s*(.*)} (the {@code IastRegexpBenchmark} COMMAND + * pattern). Not a generator output — this is a throwaway prototype to measure the performance + * ceiling of per-pattern generated code (direct char comparisons, no job-stack, no CharSet + * indirection) before committing to building a real bytecode generator for BitState-class patterns. + */ +final class CommandPatternPrototype { + + private CommandPatternPrototype() {} + + /** + * Returns the span of capture group 1 packed as {@code (start << 32) | end}, or {@code -1} if no + * match is found anywhere in the input. + */ + static long findMatch(CharSequence input) { + int n = input.length(); + int lineStart = 0; + while (lineStart <= n) { + long r = tryMatchAt(input, lineStart, n); + if (r != -1L) { + return r; + } + int nl = indexOfNewline(input, lineStart, n); + if (nl < 0) { + break; + } + lineStart = nl + 1; + } + return -1L; + } + + private static int indexOfNewline(CharSequence input, int from, int n) { + for (int i = from; i < n; i++) { + if (input.charAt(i) == '\n') { + return i; + } + } + return -1; + } + + private static long tryMatchAt(CharSequence input, int start, int n) { + // Greedy optional group first: \s*(?:sudo|doas)\s+ + int q = skipSpaces(input, start, n); + int afterKeyword = matchKeyword(input, q, n); + if (afterKeyword >= 0) { + int r = skipSpaces(input, afterKeyword, n); + if (r > afterKeyword) { + long res = matchMandatory(input, r, n); + if (res != -1L) { + return res; + } + } + } + // Backtrack: optional group not taken. + return matchMandatory(input, start, n); + } + + private static int skipSpaces(CharSequence input, int p, int n) { + while (p < n && isSpace(input.charAt(p))) { + p++; + } + return p; + } + + private static boolean isSpace(char c) { + return c == ' ' || c == '\t' || c == '\n' || c == '\u000B' || c == '\f' || c == '\r'; + } + + private static int matchKeyword(CharSequence input, int p, int n) { + if (p + 4 <= n + && input.charAt(p) == 's' + && input.charAt(p + 1) == 'u' + && input.charAt(p + 2) == 'd' + && input.charAt(p + 3) == 'o') { + return p + 4; + } + if (p + 4 <= n + && input.charAt(p) == 'd' + && input.charAt(p + 1) == 'o' + && input.charAt(p + 2) == 'a' + && input.charAt(p + 3) == 's') { + return p + 4; + } + return -1; + } + + // Matches \b\S+\b\s*(.*) starting at p. The trailing \b after \S+ is implied: a non-space + // char followed by a space or end-of-input is always a word boundary, so it needs no check. + private static long matchMandatory(CharSequence input, int p, int n) { + if (p >= n || !isWordBoundary(input, p, n) || isSpace(input.charAt(p))) { + return -1L; + } + int q = p; + while (q < n && !isSpace(input.charAt(q))) { + q++; + } + int r = skipSpaces(input, q, n); + return ((long) r << 32) | (n & 0xFFFFFFFFL); + } + + private static boolean isWordChar(char c) { + return (c >= 'a' && c <= 'z') || (c >= 'A' && c <= 'Z') || (c >= '0' && c <= '9') || c == '_'; + } + + private static boolean isWordBoundary(CharSequence input, int p, int n) { + boolean before = p > 0 && isWordChar(input.charAt(p - 1)); + boolean after = p < n && isWordChar(input.charAt(p)); + return before != after; + } +} diff --git a/reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/analysis/PatternAnalyzer.java b/reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/analysis/PatternAnalyzer.java index a8e40b43..57ff27c8 100644 --- a/reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/analysis/PatternAnalyzer.java +++ b/reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/analysis/PatternAnalyzer.java @@ -324,6 +324,31 @@ private MatchingStrategyResult routeBitState(MatchingStrategyResult result) { replaced.hasAtomicGroups = result.hasAtomicGroups; replaced.lazyNfa = result.lazyNfa; replaced.guardTrace.addAll(result.guardTrace); + + // Second substitution: narrower than BITSTATE_CAPTURE eligibility. Patterns matching the + // "prefix-guarded scan" structural family (see doc/2026-07-08-bitstate-bytecode-generator- + // design.md) can be compiled to pure straight-line, always-native bytecode instead of the + // BitState interpreter. + PrefixGuardedScanInfo prefixGuardedScanInfo = detectPrefixGuardedScan(ast); + if (prefixGuardedScanInfo != null) { + MatchingStrategyResult bytecodeResult = + new MatchingStrategyResult( + MatchingStrategy.BITSTATE_BYTECODE, + replaced.dfa, + prefixGuardedScanInfo, + replaced.useTaggedDFA, + replaced.requiredLiterals, + replaced.lookaheadGreedyInfo, + replaced.usePosixLastMatch); + bytecodeResult.alternationPriorityConflict = replaced.alternationPriorityConflict; + bytecodeResult.captureAmbiguous = replaced.captureAmbiguous; + bytecodeResult.anchorConditionDiluted = replaced.anchorConditionDiluted; + bytecodeResult.hasAtomicGroups = replaced.hasAtomicGroups; + bytecodeResult.lazyNfa = replaced.lazyNfa; + bytecodeResult.guardTrace.addAll(replaced.guardTrace); + return bytecodeResult; + } + return replaced; } @@ -3354,6 +3379,16 @@ public enum MatchingStrategy { */ BITSTATE_CAPTURE, + /** + * Always-native, pure straight-line bytecode for the narrow "prefix-guarded scan" structural + * family recognized by {@link #detectPrefixGuardedScan}: an optional lookahead-disjoint + * literal-keyword prefix, a mandatory backtrack-free scan, and an unconstrained trailing {@code + * .*}-style capture. Selected instead of {@link #BITSTATE_CAPTURE} by {@link #routeBitState} + * when the pattern matches this shape (see + * doc/2026-07-08-bitstate-bytecode-generator-design.md). + */ + BITSTATE_BYTECODE, + /** PikeVM NFA-with-capture: O(n·m) native group extraction, leftmost-greedy, ReDoS-safe. */ PIKEVM_CAPTURE } @@ -7047,6 +7082,103 @@ public int structuralHashCode() { } } + /** + * Information about a "prefix-guarded scan" pattern recognized by {@link + * #detectPrefixGuardedScan}: {@code ^(?:leadingWs(?:keyword1|keyword2|...)separatorWs)? + * \b?mandatoryCharSet+\b?trailingWs*(tailCharSet*)}, e.g. the IAST {@code COMMAND} pattern {@code + * (?s)(?m)^(?:\s*(?:sudo|doas)\s+)?\b\S+\b\s*(.*)}. See + * doc/2026-07-08-bitstate-bytecode-generator-design.md. + */ + public static class PrefixGuardedScanInfo implements PatternInfo { + // Optional prefix group fields (step 2 of the detector); all null/unset when the optional + // group is absent from the pattern. + public final CharSet leadingWsCharSet; // null if optional group absent + public final List keywords; // null if optional group absent + public final CharSet separatorCharSet; // null if optional group absent + public final int separatorMin; // min repetitions of separatorCharSet; unset (0) if absent + + // Mandatory backtrack-free scan (step 4); always present. + public final CharSet mandatoryCharSet; + public final int mandatoryMin; + + // Whether steps 3/5's optional \b anchors (immediately before/after the mandatory scan) were + // present in the source pattern. These are NOT redundant with the mandatory scan's greedy + // stopping point in general (e.g. \S+ can stop on a non-word character, in which case a + // required trailing \b needs the generator to give back characters — see + // BitStateBytecodeGenerator's matchMandatory), so the generator must know whether to enforce + // them at all. + public final boolean leadingWordBoundary; + public final boolean trailingWordBoundary; + + // Optional trailing whitespace (step 6); null if absent. + public final CharSet trailingCharSet; + + // Trailing capturing group (step 7); always present. + public final CharSet tailCharSet; + public final int tailGroupNumber; + + // Whether the leading ^ anchor (step 1) was declared multiline, i.e. (?m)^. + public final boolean multiline; + + public PrefixGuardedScanInfo( + CharSet leadingWsCharSet, + List keywords, + CharSet separatorCharSet, + int separatorMin, + CharSet mandatoryCharSet, + int mandatoryMin, + boolean leadingWordBoundary, + boolean trailingWordBoundary, + CharSet trailingCharSet, + CharSet tailCharSet, + int tailGroupNumber, + boolean multiline) { + this.leadingWsCharSet = leadingWsCharSet; + this.keywords = keywords; + this.separatorCharSet = separatorCharSet; + this.separatorMin = separatorMin; + this.mandatoryCharSet = mandatoryCharSet; + this.mandatoryMin = mandatoryMin; + this.leadingWordBoundary = leadingWordBoundary; + this.trailingWordBoundary = trailingWordBoundary; + this.trailingCharSet = trailingCharSet; + this.tailCharSet = tailCharSet; + this.tailGroupNumber = tailGroupNumber; + this.multiline = multiline; + } + + /** + * True when the pattern has the optional {@code (?:leadingWs(?:kw1|kw2|...)separatorWs)?} + * prefix. + */ + public boolean hasOptionalPrefix() { + return keywords != null; + } + + @Override + public int structuralHashCode() { + int hash = getClass().getName().hashCode(); + hash = 31 * hash + (leadingWsCharSet != null ? leadingWsCharSet.hashCode() : 0); + if (keywords != null) { + hash = 31 * hash + keywords.size(); + for (String keyword : keywords) { + hash = 31 * hash + keyword.hashCode(); + } + } + hash = 31 * hash + (separatorCharSet != null ? separatorCharSet.hashCode() : 0); + hash = 31 * hash + separatorMin; + hash = 31 * hash + mandatoryCharSet.hashCode(); + hash = 31 * hash + mandatoryMin; + hash = 31 * hash + (leadingWordBoundary ? 1 : 0); + hash = 31 * hash + (trailingWordBoundary ? 1 : 0); + hash = 31 * hash + (trailingCharSet != null ? trailingCharSet.hashCode() : 0); + hash = 31 * hash + tailCharSet.hashCode(); + hash = 31 * hash + tailGroupNumber; + hash = 31 * hash + (multiline ? 1 : 0); + return hash; + } + } + /** * Information about pure literal alternation patterns like keyword1|keyword2|...|keywordN. These * patterns can be optimized using trie-based matching instead of DFA state machines. @@ -8585,6 +8717,231 @@ private ConcatQuantifiedGroupsInfo detectConcatQuantifiedGroups(RegexNode ast) { return new ConcatQuantifiedGroupsInfo(groups); } + /** + * Detect the "prefix-guarded scan" structural family: {@code + * ^(?:leadingWs(?:keyword1|keyword2|...)separatorWs)?\b?mandatoryCharSet+\b?trailingWs*(tailCharSet*)}. + * See {@link PrefixGuardedScanInfo} and doc/2026-07-08-bitstate-bytecode-generator-design.md §4a + * for the exact shape and rationale. Returns {@code null} if {@code ast} does not match this + * shape exactly (any deviation declines eligibility rather than guessing). + */ + private PrefixGuardedScanInfo detectPrefixGuardedScan(RegexNode ast) { + if (!(ast instanceof ConcatNode)) { + return null; + } + List children = ((ConcatNode) ast).children; + int idx = 0; + + // Step 1 (required): leading ^ anchor. + if (idx >= children.size() || !(children.get(idx) instanceof AnchorNode)) { + return null; + } + AnchorNode startAnchor = (AnchorNode) children.get(idx); + if (startAnchor.type != AnchorNode.Type.START) { + return null; + } + boolean multiline = startAnchor.multiline; + idx++; + + // Step 2 (optional): (?:leadingWs(?:kw1|kw2|...)separatorWs)? + CharSet leadingWsCharSet = null; + List keywords = null; + CharSet separatorCharSet = null; + int separatorMin = 0; + if (idx < children.size() + && children.get(idx) instanceof QuantifierNode + && ((QuantifierNode) children.get(idx)).greedy) { + QuantifierNode optQuant = (QuantifierNode) children.get(idx); + if (optQuant.min == 0 && optQuant.max == 1 && optQuant.child instanceof GroupNode) { + GroupNode optGroup = (GroupNode) optQuant.child; + if (!optGroup.capturing && optGroup.child instanceof ConcatNode) { + List innerChildren = ((ConcatNode) optGroup.child).children; + if (innerChildren.size() == 3) { + CharSet lws = extractLeadingCharSet(innerChildren.get(0)); + List kws = lws != null ? extractDisjointKeywords(innerChildren.get(1)) : null; + CharSet sep = null; + int sepMin = 0; + if (lws != null && kws != null && innerChildren.get(2) instanceof QuantifierNode) { + QuantifierNode sepQuant = (QuantifierNode) innerChildren.get(2); + // max must be unbounded: the generator only stores separatorMin and emits an + // unbounded greedy scan, so a bounded separator (e.g. \s{1,3}) would let the + // generated matcher over-consume past JDK's upper bound, diverging on inputs like + // "sudo ls" (four spaces) vs. the pattern's \s{1,3}. + if (sepQuant.min >= 1 + && sepQuant.max == -1 + && sepQuant.greedy + && sepQuant.child instanceof CharClassNode) { + sep = effectiveCharSet((CharClassNode) sepQuant.child); + sepMin = sepQuant.min; + } + } + if (lws != null && kws != null && sep != null) { + leadingWsCharSet = lws; + keywords = kws; + separatorCharSet = sep; + separatorMin = sepMin; + idx++; // consume the whole optional group — shape matched + } + } + } + } + // Shape didn't match: leave idx unchanged, the optional element is simply absent. + } + + // Step 3 (optional): \b + int idxBeforeStep3 = idx; + idx = skipOptionalWordBoundary(children, idx); + boolean leadingWordBoundary = idx != idxBeforeStep3; + + // Step 4 (required): mandatory backtrack-free scan charset+ + if (idx >= children.size() || !(children.get(idx) instanceof QuantifierNode)) { + return null; + } + QuantifierNode mandatoryQuant = (QuantifierNode) children.get(idx); + // max must be unbounded (+): the generator's greedy scan (and its bounded trailing-\b + // give-back correction) assumes there is no upper bound to respect while consuming. + if (mandatoryQuant.min < 1 + || mandatoryQuant.max != -1 + || !mandatoryQuant.greedy + || !(mandatoryQuant.child instanceof CharClassNode)) { + return null; + } + CharSet mandatoryCharSet = effectiveCharSet((CharClassNode) mandatoryQuant.child); + int mandatoryMin = mandatoryQuant.min; + idx++; + + // The generator's optional prefix match is a straight-line, non-backtracking attempt: it + // either consumes the whole prefix (leadingWs + keyword + separator) or none of it. If + // leadingWs/separator can consume characters that mandatoryCharSet also accepts, the greedy + // separator scan can over-consume what the mandatory scan needs, with no way to give + // characters back except discarding the entire prefix (including the already-matched + // keyword) — a silent false negative vs. JDK. Decline rather than guess. + if (keywords != null + && (leadingWsCharSet.intersects(mandatoryCharSet) + || separatorCharSet.intersects(mandatoryCharSet))) { + return null; + } + + // Step 5 (optional): \b + int idxBeforeStep5 = idx; + idx = skipOptionalWordBoundary(children, idx); + boolean trailingWordBoundary = idx != idxBeforeStep5; + + // Step 6 (optional): trailing whitespace* + CharSet trailingCharSet = null; + if (idx < children.size() && children.get(idx) instanceof QuantifierNode) { + QuantifierNode trailingQuant = (QuantifierNode) children.get(idx); + if (trailingQuant.min == 0 + && trailingQuant.max == -1 + && trailingQuant.greedy + && trailingQuant.child instanceof CharClassNode) { + trailingCharSet = effectiveCharSet((CharClassNode) trailingQuant.child); + idx++; + } + } + + // Step 7 (required): trailing capturing group must be the LAST child. + if (idx != children.size() - 1) { + return null; + } + RegexNode last = children.get(idx); + if (!(last instanceof GroupNode)) { + return null; + } + GroupNode tailGroup = (GroupNode) last; + if (!tailGroup.capturing || !(tailGroup.child instanceof QuantifierNode)) { + return null; + } + QuantifierNode tailQuant = (QuantifierNode) tailGroup.child; + if (tailQuant.min != 0 + || tailQuant.max != -1 + || !tailQuant.greedy + || !(tailQuant.child instanceof CharClassNode)) { + return null; + } + CharSet tailCharSet = effectiveCharSet((CharClassNode) tailQuant.child); + if (!tailCharSet.equals(CharSet.ANY) && !tailCharSet.equals(CharSet.ANY_EXCEPT_NEWLINE)) { + return null; + } + + return new PrefixGuardedScanInfo( + leadingWsCharSet, + keywords, + separatorCharSet, + separatorMin, + mandatoryCharSet, + mandatoryMin, + leadingWordBoundary, + trailingWordBoundary, + trailingCharSet, + tailCharSet, + tailGroup.groupNumber, + multiline); + } + + /** Resolves a {@link CharClassNode}'s effective runtime {@link CharSet}, applying negation. */ + private CharSet effectiveCharSet(CharClassNode node) { + return node.negated ? node.chars.complement() : node.chars; + } + + /** Step 2 element [0]: {@code QuantifierNode(min=0, CharClassNode)} — leading whitespace-trim. */ + private CharSet extractLeadingCharSet(RegexNode node) { + if (!(node instanceof QuantifierNode)) { + return null; + } + QuantifierNode quant = (QuantifierNode) node; + if (quant.min != 0 + || quant.max != -1 + || !quant.greedy + || !(quant.child instanceof CharClassNode)) { + return null; + } + return effectiveCharSet((CharClassNode) quant.child); + } + + /** + * Step 2 element [1]: {@code GroupNode(non-capturing)} wrapping an {@link AlternationNode} whose + * alternatives are all pure literal keywords, pairwise distinct in their first character. This + * uniqueness lets {@link com.datadoghq.reggie.codegen.codegen.BitStateBytecodeGenerator} tell + * keywords apart from a single input character with no ambiguity, even though the generator + * currently implements the actual per-keyword check as a chain of full-literal comparisons rather + * than a first-character dispatch. + */ + private List extractDisjointKeywords(RegexNode node) { + if (!(node instanceof GroupNode)) { + return null; + } + GroupNode group = (GroupNode) node; + if (group.capturing || !(group.child instanceof AlternationNode)) { + return null; + } + List keywords = new ArrayList<>(); + Set firstChars = new HashSet<>(); + for (RegexNode alternative : ((AlternationNode) group.child).alternatives) { + String keyword = extractLiteralString(alternative); + if (keyword == null || keyword.isEmpty()) { + return null; + } + if (!firstChars.add(keyword.charAt(0))) { + return null; // Two keywords share a first character — not first-char dispatchable. + } + keywords.add(keyword); + } + if (keywords.isEmpty()) { + return null; + } + return keywords; + } + + /** Consumes an optional {@code \b} (word boundary) anchor at {@code idx}, if present. */ + private int skipOptionalWordBoundary(List children, int idx) { + if (idx < children.size() + && children.get(idx) instanceof AnchorNode + && ((AnchorNode) children.get(idx)).type == AnchorNode.Type.WORD_BOUNDARY) { + return idx + 1; + } + return idx; + } + /** * Detect stateless pattern that doesn't require state tracking. Returns StatelessPatternInfo if * pattern matches, null otherwise. diff --git a/reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/codegen/BitStateBytecodeGenerator.java b/reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/codegen/BitStateBytecodeGenerator.java new file mode 100644 index 00000000..96dc78a6 --- /dev/null +++ b/reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/codegen/BitStateBytecodeGenerator.java @@ -0,0 +1,1134 @@ +/* + * Copyright 2026-Present Datadog, Inc. + * + * Licensed under the Apache License, Version 2.0 (the "License"); + * you may not use this file except in compliance with the License. + * You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, software + * distributed under the License is distributed on an "AS IS" BASIS, + * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. + * See the License for the specific language governing permissions and + * limitations under the License. + */ +package com.datadoghq.reggie.codegen.codegen; + +import static com.datadoghq.reggie.codegen.codegen.BytecodeUtil.pushInt; +import static org.objectweb.asm.Opcodes.*; + +import com.datadoghq.reggie.codegen.analysis.PatternAnalyzer.PrefixGuardedScanInfo; +import com.datadoghq.reggie.codegen.automaton.CharSet; +import org.objectweb.asm.ClassWriter; +import org.objectweb.asm.Label; +import org.objectweb.asm.MethodVisitor; + +/** + * Generates pure straight-line, always-native bytecode for the "prefix-guarded scan" structural + * family recognized by {@code PatternAnalyzer.detectPrefixGuardedScan} (see + * doc/2026-07-08-bitstate-bytecode-generator-design.md), industrializing the hand-written reference + * at {@code reggie-benchmark/.../CommandPatternPrototype.java}: {@code + * ^(?:leadingWs(?:keyword1|keyword2|...)separatorWs)?\b?mandatoryCharSet+\b?trailingWs*(tailCharSet*)}. + * + *

The single optional prefix is resolved by two independent straight-line attempts (with the + * prefix, then without) selected by one {@code if} — no job-stack, no visited bitmap, no undo-log. + * + *

The mandatory scan's greedy stopping point is corrected for a required trailing {@code \b} + * with a single bounded backward give-back scan, not general backtracking. {@code + * CommandPatternPrototype}'s comment claiming the trailing {@code \b} after {@code \S+} is always + * implied by "non-space followed by space/EOF" does not hold for every {@code mandatoryCharSet}: + * e.g. input {@code "a!"} against {@code \b\S+\b}, the maximal {@code \S+} match ends on the + * non-word character {@code '!'} followed by end-of-input — not a word-boundary transition — so + * real {@code java.util.regex} backtracks {@code \S+} by one character to land on a boundary. This + * generator reproduces that via the bounded give-back scan in {@code matchMandatory} below, rather + * than the prototype's shortcut (which would silently produce a wrong capture-group span for such + * inputs). + */ +public class BitStateBytecodeGenerator { + + private final PrefixGuardedScanInfo info; + private final int groupCount; + + public BitStateBytecodeGenerator(PrefixGuardedScanInfo info, int groupCount) { + this.info = info; + this.groupCount = groupCount; + } + + private static String cn(String className) { + return className.replace('.', '/'); + } + + /** Registers all helper + public entry-point methods on {@code cw}. */ + public void generateAll(ClassWriter cw, String className) { + generateIsWordCharMethod(cw); + generateIsBoundaryMethod(cw, className); + if (info.hasOptionalPrefix()) { + generateSkipLeadingWsMethod(cw); + generateMatchKeywordMethod(cw); + generateSkipSeparatorMethod(cw); + } + generateMatchMandatoryMethod(cw, className); + generateTryMatchAtMethod(cw, className); + + generateMatchesMethod(cw, className); + generateFindMethod(cw, className); + generateFindFromMethod(cw, className); + generateMatchMethod(cw, className); + generateMatchesBoundedMethod(cw, className); + generateMatchBoundedMethod(cw, className); + generateFindMatchMethod(cw, className); + generateFindMatchFromMethod(cw, className); + generateFindBoundsFromMethod(cw, className); + } + + // --------------------------------------------------------------------------------------- + // Charset check emission (mirrors GreedyCharClassBytecodeGenerator.generateInlineCharSetCheck, + // parameterized by an explicit CharSet argument since this generator handles several distinct + // charsets in one class instead of one fixed instance field). PrefixGuardedScanInfo's charsets + // are already negation-resolved (PatternAnalyzer.effectiveCharSet), so this never needs a + // negated variant. + // --------------------------------------------------------------------------------------- + + /** Emits: {@code if (charVar not in charSet) goto exitLabel;} */ + private static void emitCharSetCheck( + MethodVisitor mv, int charVar, Label exitLabel, CharSet charSet) { + if (charSet.isSingleChar()) { + char c = charSet.getSingleChar(); + mv.visitVarInsn(ILOAD, charVar); + pushInt(mv, c); + mv.visitJumpInsn(IF_ICMPNE, exitLabel); + } else if (charSet.isSimpleRange()) { + CharSet.Range range = charSet.getSimpleRange(); + mv.visitVarInsn(ILOAD, charVar); + pushInt(mv, range.start); + mv.visitJumpInsn(IF_ICMPLT, exitLabel); + mv.visitVarInsn(ILOAD, charVar); + pushInt(mv, range.end); + mv.visitJumpInsn(IF_ICMPGT, exitLabel); + } else { + Label matches = new Label(); + for (CharSet.Range range : charSet.getRanges()) { + Label tryNext = new Label(); + mv.visitVarInsn(ILOAD, charVar); + pushInt(mv, range.start); + mv.visitJumpInsn(IF_ICMPLT, tryNext); + mv.visitVarInsn(ILOAD, charVar); + pushInt(mv, range.end); + mv.visitJumpInsn(IF_ICMPLE, matches); + mv.visitLabel(tryNext); + } + mv.visitJumpInsn(GOTO, exitLabel); + mv.visitLabel(matches); + } + } + + /** + * Emits a greedy "while (p < n && charAt(p) in charSet) p++;" loop, leaving the final p in pVar. + */ + private static void emitGreedyScanLoop( + MethodVisitor mv, int inputVar, int pVar, int nVar, int cVar, CharSet charSet) { + Label loopStart = new Label(); + Label loopEnd = new Label(); + mv.visitLabel(loopStart); + mv.visitVarInsn(ILOAD, pVar); + mv.visitVarInsn(ILOAD, nVar); + mv.visitJumpInsn(IF_ICMPGE, loopEnd); + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, pVar); + mv.visitMethodInsn(INVOKEVIRTUAL, "java/lang/String", "charAt", "(I)C", false); + mv.visitVarInsn(ISTORE, cVar); + emitCharSetCheck(mv, cVar, loopEnd, charSet); + mv.visitIincInsn(pVar, 1); + mv.visitJumpInsn(GOTO, loopStart); + mv.visitLabel(loopEnd); + } + + // --------------------------------------------------------------------------------------- + // Helper methods, generated once per class, shared by tryMatchAt/matchMandatory and the public + // entry points. All are private static: the generated class has no runtime instance state (all + // pattern-specific data — charsets, keyword literals, boundary requirements — is baked directly + // into instructions at generation time). + // --------------------------------------------------------------------------------------- + + /** {@code private static boolean isWordChar(char c)} — the fixed PCRE/Java \b word-char set. */ + private void generateIsWordCharMethod(ClassWriter cw) { + MethodVisitor mv = cw.visitMethod(ACC_PRIVATE | ACC_STATIC, "isWordChar", "(C)Z", null, null); + mv.visitCode(); + int cVar = 0; + Label checkUpper = new Label(); + Label checkDigit = new Label(); + Label checkUnderscore = new Label(); + Label yes = new Label(); + Label no = new Label(); + + mv.visitVarInsn(ILOAD, cVar); + pushInt(mv, 'a'); + mv.visitJumpInsn(IF_ICMPLT, checkUpper); + mv.visitVarInsn(ILOAD, cVar); + pushInt(mv, 'z'); + mv.visitJumpInsn(IF_ICMPLE, yes); + + mv.visitLabel(checkUpper); + mv.visitVarInsn(ILOAD, cVar); + pushInt(mv, 'A'); + mv.visitJumpInsn(IF_ICMPLT, checkDigit); + mv.visitVarInsn(ILOAD, cVar); + pushInt(mv, 'Z'); + mv.visitJumpInsn(IF_ICMPLE, yes); + + mv.visitLabel(checkDigit); + mv.visitVarInsn(ILOAD, cVar); + pushInt(mv, '0'); + mv.visitJumpInsn(IF_ICMPLT, checkUnderscore); + mv.visitVarInsn(ILOAD, cVar); + pushInt(mv, '9'); + mv.visitJumpInsn(IF_ICMPLE, yes); + + mv.visitLabel(checkUnderscore); + mv.visitVarInsn(ILOAD, cVar); + pushInt(mv, '_'); + mv.visitJumpInsn(IF_ICMPEQ, yes); + mv.visitJumpInsn(GOTO, no); + + mv.visitLabel(yes); + mv.visitInsn(ICONST_1); + mv.visitInsn(IRETURN); + mv.visitLabel(no); + mv.visitInsn(ICONST_0); + mv.visitInsn(IRETURN); + + mv.visitMaxs(0, 0); + mv.visitEnd(); + } + + /** + * {@code private static boolean isBoundary(String input, int p, int n)} — true iff exactly one of + * {@code (p>0 && isWordChar(input.charAt(p-1)))} / {@code (p 0 && isWordChar(input.charAt(p - 1)); + mv.visitInsn(ICONST_0); + mv.visitVarInsn(ISTORE, beforeVar); + Label skipBefore = new Label(); + mv.visitVarInsn(ILOAD, pVar); + mv.visitJumpInsn(IFLE, skipBefore); + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, pVar); + pushInt(mv, 1); + mv.visitInsn(ISUB); + mv.visitMethodInsn(INVOKEVIRTUAL, "java/lang/String", "charAt", "(I)C", false); + mv.visitMethodInsn(INVOKESTATIC, cn(className), "isWordChar", "(C)Z", false); + mv.visitVarInsn(ISTORE, beforeVar); + mv.visitLabel(skipBefore); + + // after = p < n && isWordChar(input.charAt(p)); + mv.visitInsn(ICONST_0); + mv.visitVarInsn(ISTORE, afterVar); + Label skipAfter = new Label(); + mv.visitVarInsn(ILOAD, pVar); + mv.visitVarInsn(ILOAD, nVar); + mv.visitJumpInsn(IF_ICMPGE, skipAfter); + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, pVar); + mv.visitMethodInsn(INVOKEVIRTUAL, "java/lang/String", "charAt", "(I)C", false); + mv.visitMethodInsn(INVOKESTATIC, cn(className), "isWordChar", "(C)Z", false); + mv.visitVarInsn(ISTORE, afterVar); + mv.visitLabel(skipAfter); + + // return before != after; + Label returnTrue = new Label(); + mv.visitVarInsn(ILOAD, beforeVar); + mv.visitVarInsn(ILOAD, afterVar); + mv.visitJumpInsn(IF_ICMPNE, returnTrue); + mv.visitInsn(ICONST_0); + mv.visitInsn(IRETURN); + mv.visitLabel(returnTrue); + mv.visitInsn(ICONST_1); + mv.visitInsn(IRETURN); + + mv.visitMaxs(0, 0); + mv.visitEnd(); + } + + /** + * {@code private static int skipLeadingWs(String input, int p, int n)} — greedy {@code + * leadingWsCharSet*} scan, returns the position after the scan. + */ + private void generateSkipLeadingWsMethod(ClassWriter cw) { + MethodVisitor mv = + cw.visitMethod( + ACC_PRIVATE | ACC_STATIC, "skipLeadingWs", "(Ljava/lang/String;II)I", null, null); + mv.visitCode(); + int inputVar = 0; + int pVar = 1; + int nVar = 2; + LocalVarAllocator allocator = new LocalVarAllocator(3); + int cVar = allocator.allocate(); + + emitGreedyScanLoop(mv, inputVar, pVar, nVar, cVar, info.leadingWsCharSet); + mv.visitVarInsn(ILOAD, pVar); + mv.visitInsn(IRETURN); + + mv.visitMaxs(0, 0); + mv.visitEnd(); + } + + /** + * {@code private static int matchKeyword(String input, int p, int n)} — first-char-dispatched + * literal comparisons per keyword (mirrors {@code CommandPatternPrototype.matchKeyword}, but + * data-driven from {@code info.keywords}), returns {@code p + keyword.length()} on a match, or + * {@code -1}. + */ + private void generateMatchKeywordMethod(ClassWriter cw) { + MethodVisitor mv = + cw.visitMethod( + ACC_PRIVATE | ACC_STATIC, "matchKeyword", "(Ljava/lang/String;II)I", null, null); + mv.visitCode(); + int inputVar = 0; + int pVar = 1; + int nVar = 2; + + for (String keyword : info.keywords) { + Label tryNext = new Label(); + int len = keyword.length(); + + // if (p + len > n) goto tryNext; + mv.visitVarInsn(ILOAD, pVar); + pushInt(mv, len); + mv.visitInsn(IADD); + mv.visitVarInsn(ILOAD, nVar); + mv.visitJumpInsn(IF_ICMPGT, tryNext); + + for (int i = 0; i < len; i++) { + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, pVar); + pushInt(mv, i); + mv.visitInsn(IADD); + mv.visitMethodInsn(INVOKEVIRTUAL, "java/lang/String", "charAt", "(I)C", false); + pushInt(mv, keyword.charAt(i)); + mv.visitJumpInsn(IF_ICMPNE, tryNext); + } + + // success: return p + len; + mv.visitVarInsn(ILOAD, pVar); + pushInt(mv, len); + mv.visitInsn(IADD); + mv.visitInsn(IRETURN); + + mv.visitLabel(tryNext); + } + + pushInt(mv, -1); + mv.visitInsn(IRETURN); + + mv.visitMaxs(0, 0); + mv.visitEnd(); + } + + /** + * {@code private static int skipSeparator(String input, int p, int n)} — greedy {@code + * separatorCharSet{separatorMin,}} scan; returns the position after the scan, or {@code -1} if + * fewer than {@code separatorMin} characters matched. + */ + private void generateSkipSeparatorMethod(ClassWriter cw) { + MethodVisitor mv = + cw.visitMethod( + ACC_PRIVATE | ACC_STATIC, "skipSeparator", "(Ljava/lang/String;II)I", null, null); + mv.visitCode(); + int inputVar = 0; + int pVar = 1; + int nVar = 2; + LocalVarAllocator allocator = new LocalVarAllocator(3); + int startVar = allocator.allocate(); + int cVar = allocator.allocate(); + + // int start = p; + mv.visitVarInsn(ILOAD, pVar); + mv.visitVarInsn(ISTORE, startVar); + + emitGreedyScanLoop(mv, inputVar, pVar, nVar, cVar, info.separatorCharSet); + + // if (p - start < separatorMin) return -1; + Label minOk = new Label(); + mv.visitVarInsn(ILOAD, pVar); + mv.visitVarInsn(ILOAD, startVar); + mv.visitInsn(ISUB); + pushInt(mv, info.separatorMin); + mv.visitJumpInsn(IF_ICMPGE, minOk); + pushInt(mv, -1); + mv.visitInsn(IRETURN); + mv.visitLabel(minOk); + + mv.visitVarInsn(ILOAD, pVar); + mv.visitInsn(IRETURN); + + mv.visitMaxs(0, 0); + mv.visitEnd(); + } + + /** + * {@code private static long matchMandatory(String input, int p, int n)} — matches {@code + * \b?mandatoryCharSet+\b?trailingWs*(tailCharSet*)} starting at {@code p}. Returns the tail + * group's span packed as {@code (tailStart << 32) | tailEnd}, or {@code -1L} on failure. + * + *

Unlike {@code CommandPatternPrototype.matchMandatory}, this performs a bounded backward + * give-back scan when {@code trailingWordBoundary} is required and the maximal greedy match does + * not land on a boundary (see class javadoc for why this is necessary for correctness). + */ + private void generateMatchMandatoryMethod(ClassWriter cw, String className) { + MethodVisitor mv = + cw.visitMethod( + ACC_PRIVATE | ACC_STATIC, "matchMandatory", "(Ljava/lang/String;II)J", null, null); + mv.visitCode(); + int inputVar = 0; + int pVar = 1; + int nVar = 2; + LocalVarAllocator allocator = new LocalVarAllocator(3); + int qVar = allocator.allocate(); + int cVar = allocator.allocate(); + int minEndVar = allocator.allocate(); // p + mandatoryMin, the give-back floor + int rVar = allocator.allocate(); + + Label fail = new Label(); + + // if (leadingWordBoundary && !isBoundary(input, p, n)) return -1L; + if (info.leadingWordBoundary) { + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, pVar); + mv.visitVarInsn(ILOAD, nVar); + mv.visitMethodInsn( + INVOKESTATIC, cn(className), "isBoundary", "(Ljava/lang/String;II)Z", false); + mv.visitJumpInsn(IFEQ, fail); + } + + // int q = p; + mv.visitVarInsn(ILOAD, pVar); + mv.visitVarInsn(ISTORE, qVar); + + emitGreedyScanLoop(mv, inputVar, qVar, nVar, cVar, info.mandatoryCharSet); + + // int minEnd = p + mandatoryMin; + mv.visitVarInsn(ILOAD, pVar); + pushInt(mv, info.mandatoryMin); + mv.visitInsn(IADD); + mv.visitVarInsn(ISTORE, minEndVar); + + // if (q < minEnd) return -1L; (fewer than mandatoryMin chars matched) + mv.visitVarInsn(ILOAD, qVar); + mv.visitVarInsn(ILOAD, minEndVar); + mv.visitJumpInsn(IF_ICMPLT, fail); + + if (info.trailingWordBoundary) { + // while (q > minEnd && !isBoundary(input, q, n)) q--; + Label backLoop = new Label(); + Label backEnd = new Label(); + mv.visitLabel(backLoop); + mv.visitVarInsn(ILOAD, qVar); + mv.visitVarInsn(ILOAD, minEndVar); + mv.visitJumpInsn(IF_ICMPLE, backEnd); + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, qVar); + mv.visitVarInsn(ILOAD, nVar); + mv.visitMethodInsn( + INVOKESTATIC, cn(className), "isBoundary", "(Ljava/lang/String;II)Z", false); + mv.visitJumpInsn(IFNE, backEnd); + mv.visitIincInsn(qVar, -1); + mv.visitJumpInsn(GOTO, backLoop); + mv.visitLabel(backEnd); + + // if (!isBoundary(input, q, n)) return -1L; + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, qVar); + mv.visitVarInsn(ILOAD, nVar); + mv.visitMethodInsn( + INVOKESTATIC, cn(className), "isBoundary", "(Ljava/lang/String;II)Z", false); + mv.visitJumpInsn(IFEQ, fail); + } + + // int r = q; + mv.visitVarInsn(ILOAD, qVar); + mv.visitVarInsn(ISTORE, rVar); + + if (info.trailingCharSet != null) { + emitGreedyScanLoop(mv, inputVar, rVar, nVar, cVar, info.trailingCharSet); + } + + // Tail scan: greedily consume tailCharSet from r; tailStart = r (before the scan). + int tailStartVar = allocator.allocate(); + mv.visitVarInsn(ILOAD, rVar); + mv.visitVarInsn(ISTORE, tailStartVar); + emitGreedyScanLoop(mv, inputVar, rVar, nVar, cVar, info.tailCharSet); + + // return ((long) tailStart << 32) | (long) tailEnd; (tailEnd is now in rVar) + mv.visitVarInsn(ILOAD, tailStartVar); + mv.visitInsn(I2L); + pushInt(mv, 32); + mv.visitInsn(LSHL); + mv.visitVarInsn(ILOAD, rVar); + mv.visitInsn(I2L); + mv.visitInsn(LOR); + mv.visitInsn(LRETURN); + + mv.visitLabel(fail); + pushInt(mv, -1); + mv.visitInsn(I2L); + mv.visitInsn(LRETURN); + + mv.visitMaxs(0, 0); + mv.visitEnd(); + } + + /** + * {@code private static long tryMatchAt(String input, int start, int n)} — the single choice + * point: attempt the optional prefix then the mandatory scan; on failure, retry the mandatory + * scan alone. Two independent straight-line paths, no recursion (mirrors {@code + * CommandPatternPrototype.tryMatchAt}). + */ + private void generateTryMatchAtMethod(ClassWriter cw, String className) { + MethodVisitor mv = + cw.visitMethod( + ACC_PRIVATE | ACC_STATIC, "tryMatchAt", "(Ljava/lang/String;II)J", null, null); + mv.visitCode(); + int inputVar = 0; + int startVar = 1; + int nVar = 2; + LocalVarAllocator allocator = new LocalVarAllocator(3); + + if (info.hasOptionalPrefix()) { + int qVar = allocator.allocate(); + int afterKeywordVar = allocator.allocate(); + int rVar = allocator.allocate(); + int resVar = allocator.allocateWide(); + + Label fallback = new Label(); + + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, startVar); + mv.visitVarInsn(ILOAD, nVar); + mv.visitMethodInsn( + INVOKESTATIC, cn(className), "skipLeadingWs", "(Ljava/lang/String;II)I", false); + mv.visitVarInsn(ISTORE, qVar); + + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, qVar); + mv.visitVarInsn(ILOAD, nVar); + mv.visitMethodInsn( + INVOKESTATIC, cn(className), "matchKeyword", "(Ljava/lang/String;II)I", false); + mv.visitVarInsn(ISTORE, afterKeywordVar); + + mv.visitVarInsn(ILOAD, afterKeywordVar); + mv.visitJumpInsn(IFLT, fallback); + + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, afterKeywordVar); + mv.visitVarInsn(ILOAD, nVar); + mv.visitMethodInsn( + INVOKESTATIC, cn(className), "skipSeparator", "(Ljava/lang/String;II)I", false); + mv.visitVarInsn(ISTORE, rVar); + + mv.visitVarInsn(ILOAD, rVar); + mv.visitJumpInsn(IFLT, fallback); + + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, rVar); + mv.visitVarInsn(ILOAD, nVar); + mv.visitMethodInsn( + INVOKESTATIC, cn(className), "matchMandatory", "(Ljava/lang/String;II)J", false); + mv.visitVarInsn(LSTORE, resVar); + + // if (res != -1L) return res; + mv.visitVarInsn(LLOAD, resVar); + pushInt(mv, -1); + mv.visitInsn(I2L); + mv.visitInsn(LCMP); + mv.visitJumpInsn(IFEQ, fallback); + mv.visitVarInsn(LLOAD, resVar); + mv.visitInsn(LRETURN); + + mv.visitLabel(fallback); + } + + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, startVar); + mv.visitVarInsn(ILOAD, nVar); + mv.visitMethodInsn( + INVOKESTATIC, cn(className), "matchMandatory", "(Ljava/lang/String;II)J", false); + mv.visitInsn(LRETURN); + + mv.visitMaxs(0, 0); + mv.visitEnd(); + } + + // --------------------------------------------------------------------------------------- + // Public entry points (mirror GreedyCharClassBytecodeGenerator's method set: matches, find, + // findFrom, match, matchesBounded, matchBounded, findMatch, findMatchFrom, findBoundsFrom). + // --------------------------------------------------------------------------------------- + + /** {@code public boolean matches(String input)} — full-string, position-0-anchored match. */ + public void generateMatchesMethod(ClassWriter cw, String className) { + MethodVisitor mv = cw.visitMethod(ACC_PUBLIC, "matches", "(Ljava/lang/String;)Z", null, null); + mv.visitCode(); + int inputVar = 1; + LocalVarAllocator allocator = new LocalVarAllocator(2); + int nVar = allocator.allocate(); + int resVar = allocator.allocateWide(); + + Label notNull = new Label(); + mv.visitVarInsn(ALOAD, inputVar); + mv.visitJumpInsn(IFNONNULL, notNull); + mv.visitInsn(ICONST_0); + mv.visitInsn(IRETURN); + mv.visitLabel(notNull); + + mv.visitVarInsn(ALOAD, inputVar); + mv.visitMethodInsn(INVOKEVIRTUAL, "java/lang/String", "length", "()I", false); + mv.visitVarInsn(ISTORE, nVar); + + mv.visitVarInsn(ALOAD, inputVar); + mv.visitInsn(ICONST_0); + mv.visitVarInsn(ILOAD, nVar); + mv.visitMethodInsn(INVOKESTATIC, cn(className), "tryMatchAt", "(Ljava/lang/String;II)J", false); + mv.visitVarInsn(LSTORE, resVar); + + // if (res == -1L) return false; + Label ok = new Label(); + mv.visitVarInsn(LLOAD, resVar); + pushInt(mv, -1); + mv.visitInsn(I2L); + mv.visitInsn(LCMP); + mv.visitJumpInsn(IFNE, ok); + mv.visitInsn(ICONST_0); + mv.visitInsn(IRETURN); + mv.visitLabel(ok); + + // return ((int) res) == n; (tailEnd must reach the end of input for matches()) + mv.visitVarInsn(LLOAD, resVar); + mv.visitInsn(L2I); + mv.visitVarInsn(ILOAD, nVar); + Label eq = new Label(); + mv.visitJumpInsn(IF_ICMPEQ, eq); + mv.visitInsn(ICONST_0); + mv.visitInsn(IRETURN); + mv.visitLabel(eq); + mv.visitInsn(ICONST_1); + mv.visitInsn(IRETURN); + + mv.visitMaxs(0, 0); + mv.visitEnd(); + } + + /** {@code public boolean find(String input)} — delegates to {@code findFrom(input, 0) >= 0}. */ + public void generateFindMethod(ClassWriter cw, String className) { + MethodVisitor mv = cw.visitMethod(ACC_PUBLIC, "find", "(Ljava/lang/String;)Z", null, null); + mv.visitCode(); + + mv.visitVarInsn(ALOAD, 0); + mv.visitVarInsn(ALOAD, 1); + mv.visitInsn(ICONST_0); + mv.visitMethodInsn(INVOKEVIRTUAL, cn(className), "findFrom", "(Ljava/lang/String;I)I", false); + + Label returnTrue = new Label(); + Label end = new Label(); + mv.visitJumpInsn(IFGE, returnTrue); + mv.visitInsn(ICONST_0); + mv.visitJumpInsn(GOTO, end); + mv.visitLabel(returnTrue); + mv.visitInsn(ICONST_1); + mv.visitLabel(end); + mv.visitInsn(IRETURN); + + mv.visitMaxs(0, 0); + mv.visitEnd(); + } + + /** + * {@code public int findFrom(String input, int start)} — returns the position of the first anchor + * point (line start under {@code multiline}, else only position 0) at or after {@code start} + * where {@code tryMatchAt} succeeds, or {@code -1}. + */ + public void generateFindFromMethod(ClassWriter cw, String className) { + MethodVisitor mv = cw.visitMethod(ACC_PUBLIC, "findFrom", "(Ljava/lang/String;I)I", null, null); + mv.visitCode(); + int inputVar = 1; + int startVar = 2; + LocalVarAllocator allocator = new LocalVarAllocator(3); + int nVar = allocator.allocate(); + + // if (input == null || start < 0 || start > input.length()) return -1; + Label checksPass = new Label(); + Label returnMinusOne = new Label(); + mv.visitVarInsn(ALOAD, inputVar); + mv.visitJumpInsn(IFNULL, returnMinusOne); + mv.visitVarInsn(ILOAD, startVar); + mv.visitJumpInsn(IFLT, returnMinusOne); + mv.visitVarInsn(ILOAD, startVar); + mv.visitVarInsn(ALOAD, inputVar); + mv.visitMethodInsn(INVOKEVIRTUAL, "java/lang/String", "length", "()I", false); + mv.visitJumpInsn(IF_ICMPLE, checksPass); + mv.visitLabel(returnMinusOne); + mv.visitInsn(ICONST_M1); + mv.visitInsn(IRETURN); + mv.visitLabel(checksPass); + + mv.visitVarInsn(ALOAD, inputVar); + mv.visitMethodInsn(INVOKEVIRTUAL, "java/lang/String", "length", "()I", false); + mv.visitVarInsn(ISTORE, nVar); + + if (!info.multiline) { + // Only position 0 is ever a valid ^ anchor point without MULTILINE. + Label tryZero = new Label(); + mv.visitVarInsn(ILOAD, startVar); + mv.visitJumpInsn(IFLE, tryZero); + mv.visitInsn(ICONST_M1); + mv.visitInsn(IRETURN); + mv.visitLabel(tryZero); + + int resVar = allocator.allocateWide(); + mv.visitVarInsn(ALOAD, inputVar); + mv.visitInsn(ICONST_0); + mv.visitVarInsn(ILOAD, nVar); + mv.visitMethodInsn( + INVOKESTATIC, cn(className), "tryMatchAt", "(Ljava/lang/String;II)J", false); + mv.visitVarInsn(LSTORE, resVar); + + Label fail = new Label(); + mv.visitVarInsn(LLOAD, resVar); + pushInt(mv, -1); + mv.visitInsn(I2L); + mv.visitInsn(LCMP); + mv.visitJumpInsn(IFEQ, fail); + mv.visitInsn(ICONST_0); + mv.visitInsn(IRETURN); + mv.visitLabel(fail); + mv.visitInsn(ICONST_M1); + mv.visitInsn(IRETURN); + } else { + int lineStartVar = allocator.allocate(); + int resVar = allocator.allocateWide(); + int nlVar = allocator.allocate(); + + // int lineStart; + // if (start <= 0) lineStart = 0; + // else if (input.charAt(start - 1) == '\n') lineStart = start; + // else { nl = input.indexOf('\n', start); if (nl < 0) return -1; lineStart = nl + 1; } + Label startLeZero = new Label(); + Label afterNewline = new Label(); + Label searchNewline = new Label(); + Label lineStartSet = new Label(); + + mv.visitVarInsn(ILOAD, startVar); + mv.visitJumpInsn(IFLE, startLeZero); + + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, startVar); + pushInt(mv, 1); + mv.visitInsn(ISUB); + mv.visitMethodInsn(INVOKEVIRTUAL, "java/lang/String", "charAt", "(I)C", false); + pushInt(mv, '\n'); + mv.visitJumpInsn(IF_ICMPNE, searchNewline); + + mv.visitJumpInsn(GOTO, afterNewline); + + mv.visitLabel(searchNewline); + mv.visitVarInsn(ALOAD, inputVar); + pushInt(mv, '\n'); + mv.visitVarInsn(ILOAD, startVar); + mv.visitMethodInsn(INVOKEVIRTUAL, "java/lang/String", "indexOf", "(II)I", false); + mv.visitVarInsn(ISTORE, nlVar); + Label nlFound = new Label(); + mv.visitVarInsn(ILOAD, nlVar); + mv.visitJumpInsn(IFGE, nlFound); + mv.visitInsn(ICONST_M1); + mv.visitInsn(IRETURN); + mv.visitLabel(nlFound); + mv.visitVarInsn(ILOAD, nlVar); + pushInt(mv, 1); + mv.visitInsn(IADD); + mv.visitVarInsn(ISTORE, lineStartVar); + mv.visitJumpInsn(GOTO, lineStartSet); + + mv.visitLabel(afterNewline); + mv.visitVarInsn(ILOAD, startVar); + mv.visitVarInsn(ISTORE, lineStartVar); + mv.visitJumpInsn(GOTO, lineStartSet); + + mv.visitLabel(startLeZero); + mv.visitInsn(ICONST_0); + mv.visitVarInsn(ISTORE, lineStartVar); + + mv.visitLabel(lineStartSet); + + // while (lineStart <= n) { try; advance past next '\n' or break; } + Label loop = new Label(); + Label loopEnd = new Label(); + mv.visitLabel(loop); + mv.visitVarInsn(ILOAD, lineStartVar); + mv.visitVarInsn(ILOAD, nVar); + mv.visitJumpInsn(IF_ICMPGT, loopEnd); + + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, lineStartVar); + mv.visitVarInsn(ILOAD, nVar); + mv.visitMethodInsn( + INVOKESTATIC, cn(className), "tryMatchAt", "(Ljava/lang/String;II)J", false); + mv.visitVarInsn(LSTORE, resVar); + + Label noMatch = new Label(); + mv.visitVarInsn(LLOAD, resVar); + pushInt(mv, -1); + mv.visitInsn(I2L); + mv.visitInsn(LCMP); + mv.visitJumpInsn(IFEQ, noMatch); + mv.visitVarInsn(ILOAD, lineStartVar); + mv.visitInsn(IRETURN); + mv.visitLabel(noMatch); + + mv.visitVarInsn(ALOAD, inputVar); + pushInt(mv, '\n'); + mv.visitVarInsn(ILOAD, lineStartVar); + mv.visitMethodInsn(INVOKEVIRTUAL, "java/lang/String", "indexOf", "(II)I", false); + mv.visitVarInsn(ISTORE, nlVar); + mv.visitVarInsn(ILOAD, nlVar); + mv.visitJumpInsn(IFLT, loopEnd); + mv.visitVarInsn(ILOAD, nlVar); + pushInt(mv, 1); + mv.visitInsn(IADD); + mv.visitVarInsn(ISTORE, lineStartVar); + mv.visitJumpInsn(GOTO, loop); + + mv.visitLabel(loopEnd); + mv.visitInsn(ICONST_M1); + mv.visitInsn(IRETURN); + } + + mv.visitMaxs(0, 0); + mv.visitEnd(); + } + + /** Emits: {@code long res = tryMatchAt(input, matchStart, n);} and unpacks tailStart/tailEnd. */ + private void emitTryMatchAndUnpack( + MethodVisitor mv, + String className, + int inputVar, + int matchStartVar, + int nVar, + int resVar, + int tailStartVar, + int tailEndVar) { + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, matchStartVar); + mv.visitVarInsn(ILOAD, nVar); + mv.visitMethodInsn(INVOKESTATIC, cn(className), "tryMatchAt", "(Ljava/lang/String;II)J", false); + mv.visitVarInsn(LSTORE, resVar); + + // tailStart = (int) (res >>> 32); tailEnd = (int) res; + mv.visitVarInsn(LLOAD, resVar); + pushInt(mv, 32); + mv.visitInsn(LUSHR); + mv.visitInsn(L2I); + mv.visitVarInsn(ISTORE, tailStartVar); + mv.visitVarInsn(LLOAD, resVar); + mv.visitInsn(L2I); + mv.visitVarInsn(ISTORE, tailEndVar); + } + + /** + * Emits construction of a {@code MatchResultImpl} for a 1-capturing-group result and returns it. + */ + private void emitBuildAndReturnMatchResult( + MethodVisitor mv, int inputVar, int startVar, int tailStartVar, int endVar) { + mv.visitTypeInsn(NEW, "com/datadoghq/reggie/runtime/MatchResultImpl"); + mv.visitInsn(DUP); + mv.visitVarInsn(ALOAD, inputVar); + + // starts: {start, ..., tailStart} at index tailGroupNumber + pushInt(mv, groupCount + 1); + mv.visitIntInsn(NEWARRAY, T_INT); + mv.visitInsn(DUP); + mv.visitInsn(ICONST_0); + mv.visitVarInsn(ILOAD, startVar); + mv.visitInsn(IASTORE); + mv.visitInsn(DUP); + pushInt(mv, info.tailGroupNumber); + mv.visitVarInsn(ILOAD, tailStartVar); + mv.visitInsn(IASTORE); + + // ends: {end, ..., end} — group 0 and the tail group share the same end position. + pushInt(mv, groupCount + 1); + mv.visitIntInsn(NEWARRAY, T_INT); + mv.visitInsn(DUP); + mv.visitInsn(ICONST_0); + mv.visitVarInsn(ILOAD, endVar); + mv.visitInsn(IASTORE); + mv.visitInsn(DUP); + pushInt(mv, info.tailGroupNumber); + mv.visitVarInsn(ILOAD, endVar); + mv.visitInsn(IASTORE); + + pushInt(mv, groupCount); + + mv.visitMethodInsn( + INVOKESPECIAL, + "com/datadoghq/reggie/runtime/MatchResultImpl", + "", + "(Ljava/lang/String;[I[II)V", + false); + mv.visitInsn(ARETURN); + } + + /** {@code public MatchResult match(String input)} — full-string, position-0-anchored match. */ + public void generateMatchMethod(ClassWriter cw, String className) { + MethodVisitor mv = + cw.visitMethod( + ACC_PUBLIC, + "match", + "(Ljava/lang/String;)Lcom/datadoghq/reggie/runtime/MatchResult;", + null, + null); + mv.visitCode(); + int inputVar = 1; + LocalVarAllocator allocator = new LocalVarAllocator(2); + int nVar = allocator.allocate(); + int resVar = allocator.allocateWide(); + int tailStartVar = allocator.allocate(); + int tailEndVar = allocator.allocate(); + int startVar = allocator.allocate(); + + Label notNull = new Label(); + mv.visitVarInsn(ALOAD, inputVar); + mv.visitJumpInsn(IFNONNULL, notNull); + mv.visitInsn(ACONST_NULL); + mv.visitInsn(ARETURN); + mv.visitLabel(notNull); + + mv.visitVarInsn(ALOAD, inputVar); + mv.visitMethodInsn(INVOKEVIRTUAL, "java/lang/String", "length", "()I", false); + mv.visitVarInsn(ISTORE, nVar); + + mv.visitInsn(ICONST_0); + mv.visitVarInsn(ISTORE, startVar); + + emitTryMatchAndUnpack( + mv, className, inputVar, startVar, nVar, resVar, tailStartVar, tailEndVar); + + // if (res == -1L || tailEnd != n) return null; + Label ok = new Label(); + mv.visitVarInsn(LLOAD, resVar); + pushInt(mv, -1); + mv.visitInsn(I2L); + mv.visitInsn(LCMP); + mv.visitJumpInsn(IFNE, ok); + mv.visitInsn(ACONST_NULL); + mv.visitInsn(ARETURN); + mv.visitLabel(ok); + + Label full = new Label(); + mv.visitVarInsn(ILOAD, tailEndVar); + mv.visitVarInsn(ILOAD, nVar); + mv.visitJumpInsn(IF_ICMPEQ, full); + mv.visitInsn(ACONST_NULL); + mv.visitInsn(ARETURN); + mv.visitLabel(full); + + emitBuildAndReturnMatchResult(mv, inputVar, startVar, tailStartVar, tailEndVar); + + mv.visitMaxs(0, 0); + mv.visitEnd(); + } + + /** {@code public boolean matchesBounded(CharSequence, int, int)} — delegates to matches(). */ + public void generateMatchesBoundedMethod(ClassWriter cw, String className) { + MethodVisitor mv = + cw.visitMethod(ACC_PUBLIC, "matchesBounded", "(Ljava/lang/CharSequence;II)Z", null, null); + mv.visitCode(); + mv.visitVarInsn(ALOAD, 0); + mv.visitVarInsn(ALOAD, 1); + mv.visitVarInsn(ILOAD, 2); + mv.visitVarInsn(ILOAD, 3); + mv.visitMethodInsn( + INVOKEINTERFACE, + "java/lang/CharSequence", + "subSequence", + "(II)Ljava/lang/CharSequence;", + true); + mv.visitMethodInsn( + INVOKEINTERFACE, "java/lang/CharSequence", "toString", "()Ljava/lang/String;", true); + mv.visitMethodInsn(INVOKEVIRTUAL, cn(className), "matches", "(Ljava/lang/String;)Z", false); + mv.visitInsn(IRETURN); + mv.visitMaxs(0, 0); + mv.visitEnd(); + } + + /** {@code public MatchResult matchBounded(CharSequence, int, int)} — delegates to match(). */ + public void generateMatchBoundedMethod(ClassWriter cw, String className) { + MethodVisitor mv = + cw.visitMethod( + ACC_PUBLIC, + "matchBounded", + "(Ljava/lang/CharSequence;II)Lcom/datadoghq/reggie/runtime/MatchResult;", + null, + null); + mv.visitCode(); + mv.visitVarInsn(ALOAD, 0); + mv.visitVarInsn(ALOAD, 1); + mv.visitVarInsn(ILOAD, 2); + mv.visitVarInsn(ILOAD, 3); + mv.visitMethodInsn( + INVOKEINTERFACE, + "java/lang/CharSequence", + "subSequence", + "(II)Ljava/lang/CharSequence;", + true); + mv.visitMethodInsn( + INVOKEINTERFACE, "java/lang/CharSequence", "toString", "()Ljava/lang/String;", true); + mv.visitMethodInsn( + INVOKEVIRTUAL, + cn(className), + "match", + "(Ljava/lang/String;)Lcom/datadoghq/reggie/runtime/MatchResult;", + false); + mv.visitInsn(ARETURN); + mv.visitMaxs(0, 0); + mv.visitEnd(); + } + + /** {@code public MatchResult findMatch(String input)} — delegates to findMatchFrom(input, 0). */ + public void generateFindMatchMethod(ClassWriter cw, String className) { + MethodVisitor mv = + cw.visitMethod( + ACC_PUBLIC, + "findMatch", + "(Ljava/lang/String;)Lcom/datadoghq/reggie/runtime/MatchResult;", + null, + null); + mv.visitCode(); + mv.visitVarInsn(ALOAD, 0); + mv.visitVarInsn(ALOAD, 1); + mv.visitInsn(ICONST_0); + mv.visitMethodInsn( + INVOKEVIRTUAL, + cn(className), + "findMatchFrom", + "(Ljava/lang/String;I)Lcom/datadoghq/reggie/runtime/MatchResult;", + false); + mv.visitInsn(ARETURN); + mv.visitMaxs(0, 0); + mv.visitEnd(); + } + + /** + * {@code public MatchResult findMatchFrom(String input, int start)} — locates the next line start + * via {@code findFrom} and re-runs {@code tryMatchAt} there to recover the tail span. + */ + public void generateFindMatchFromMethod(ClassWriter cw, String className) { + MethodVisitor mv = + cw.visitMethod( + ACC_PUBLIC, + "findMatchFrom", + "(Ljava/lang/String;I)Lcom/datadoghq/reggie/runtime/MatchResult;", + null, + null); + mv.visitCode(); + int inputVar = 1; + int startVar = 2; + LocalVarAllocator allocator = new LocalVarAllocator(3); + int matchStartVar = allocator.allocate(); + int nVar = allocator.allocate(); + int resVar = allocator.allocateWide(); + int tailStartVar = allocator.allocate(); + int tailEndVar = allocator.allocate(); + + mv.visitVarInsn(ALOAD, 0); + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, startVar); + mv.visitMethodInsn(INVOKEVIRTUAL, cn(className), "findFrom", "(Ljava/lang/String;I)I", false); + mv.visitVarInsn(ISTORE, matchStartVar); + + Label found = new Label(); + mv.visitVarInsn(ILOAD, matchStartVar); + mv.visitJumpInsn(IFGE, found); + mv.visitInsn(ACONST_NULL); + mv.visitInsn(ARETURN); + mv.visitLabel(found); + + mv.visitVarInsn(ALOAD, inputVar); + mv.visitMethodInsn(INVOKEVIRTUAL, "java/lang/String", "length", "()I", false); + mv.visitVarInsn(ISTORE, nVar); + + emitTryMatchAndUnpack( + mv, className, inputVar, matchStartVar, nVar, resVar, tailStartVar, tailEndVar); + + emitBuildAndReturnMatchResult(mv, inputVar, matchStartVar, tailStartVar, tailEndVar); + + mv.visitMaxs(0, 0); + mv.visitEnd(); + } + + /** + * {@code public boolean findBoundsFrom(String input, int start, int[] bounds)} — allocation-free + * alternative to {@code findMatchFrom}; only fills group 0's span, matching {@code + * GreedyCharClassBytecodeGenerator}'s contract. + */ + public void generateFindBoundsFromMethod(ClassWriter cw, String className) { + MethodVisitor mv = + cw.visitMethod(ACC_PUBLIC, "findBoundsFrom", "(Ljava/lang/String;I[I)Z", null, null); + mv.visitCode(); + int inputVar = 1; + int startVar = 2; + int boundsVar = 3; + LocalVarAllocator allocator = new LocalVarAllocator(4); + int matchStartVar = allocator.allocate(); + int nVar = allocator.allocate(); + int resVar = allocator.allocateWide(); + int tailStartVar = allocator.allocate(); + int tailEndVar = allocator.allocate(); + + mv.visitVarInsn(ALOAD, 0); + mv.visitVarInsn(ALOAD, inputVar); + mv.visitVarInsn(ILOAD, startVar); + mv.visitMethodInsn(INVOKEVIRTUAL, cn(className), "findFrom", "(Ljava/lang/String;I)I", false); + mv.visitVarInsn(ISTORE, matchStartVar); + + Label found = new Label(); + mv.visitVarInsn(ILOAD, matchStartVar); + mv.visitJumpInsn(IFGE, found); + mv.visitInsn(ICONST_0); + mv.visitInsn(IRETURN); + mv.visitLabel(found); + + mv.visitVarInsn(ALOAD, inputVar); + mv.visitMethodInsn(INVOKEVIRTUAL, "java/lang/String", "length", "()I", false); + mv.visitVarInsn(ISTORE, nVar); + + emitTryMatchAndUnpack( + mv, className, inputVar, matchStartVar, nVar, resVar, tailStartVar, tailEndVar); + + mv.visitVarInsn(ALOAD, boundsVar); + mv.visitInsn(ICONST_0); + mv.visitVarInsn(ILOAD, matchStartVar); + mv.visitInsn(IASTORE); + mv.visitVarInsn(ALOAD, boundsVar); + mv.visitInsn(ICONST_1); + mv.visitVarInsn(ILOAD, tailEndVar); + mv.visitInsn(IASTORE); + + mv.visitInsn(ICONST_1); + mv.visitInsn(IRETURN); + + mv.visitMaxs(0, 0); + mv.visitEnd(); + } +} diff --git a/reggie-processor/src/main/java/com/datadoghq/reggie/processor/ReggieMatcherBytecodeGenerator.java b/reggie-processor/src/main/java/com/datadoghq/reggie/processor/ReggieMatcherBytecodeGenerator.java index 7cd4e574..b1c0e2e0 100644 --- a/reggie-processor/src/main/java/com/datadoghq/reggie/processor/ReggieMatcherBytecodeGenerator.java +++ b/reggie-processor/src/main/java/com/datadoghq/reggie/processor/ReggieMatcherBytecodeGenerator.java @@ -32,6 +32,7 @@ import com.datadoghq.reggie.codegen.automaton.NFA; import com.datadoghq.reggie.codegen.automaton.ThompsonBuilder; import com.datadoghq.reggie.codegen.codegen.BackreferenceBytecodeGenerator; +import com.datadoghq.reggie.codegen.codegen.BitStateBytecodeGenerator; import com.datadoghq.reggie.codegen.codegen.BoundedQuantifierBytecodeGenerator; import com.datadoghq.reggie.codegen.codegen.DFASwitchBytecodeGenerator; import com.datadoghq.reggie.codegen.codegen.DFATableBytecodeGenerator; @@ -346,6 +347,14 @@ public byte[] generate() throws Exception { multiGroupGen.generateTryMatchBoundsFromPositionMethod(cw, getJavaClassName()); break; + case BITSTATE_BYTECODE: + PatternAnalyzer.PrefixGuardedScanInfo prefixGuardedInfo = + (PatternAnalyzer.PrefixGuardedScanInfo) result.patternInfo; + BitStateBytecodeGenerator bitStateGen = + new BitStateBytecodeGenerator(prefixGuardedInfo, nfa.getGroupCount()); + bitStateGen.generateAll(cw, getJavaClassName()); + break; + case SPECIALIZED_FIXED_SEQUENCE: PatternAnalyzer.FixedSequenceInfo fixedInfo = (PatternAnalyzer.FixedSequenceInfo) result.patternInfo; diff --git a/reggie-processor/src/test/java/com/datadoghq/reggie/processor/ReggieMatcherBytecodeGeneratorTest.java b/reggie-processor/src/test/java/com/datadoghq/reggie/processor/ReggieMatcherBytecodeGeneratorTest.java index 1963e3a9..42f33f3f 100644 --- a/reggie-processor/src/test/java/com/datadoghq/reggie/processor/ReggieMatcherBytecodeGeneratorTest.java +++ b/reggie-processor/src/test/java/com/datadoghq/reggie/processor/ReggieMatcherBytecodeGeneratorTest.java @@ -353,6 +353,21 @@ void testOnePassNfaMatchFromMethod() throws Exception { assertEquals(-1, (Integer) matchFrom.invoke(matcher, (String) null, 0)); } + @Test + void testBitStateBytecodeStrategy() throws Exception { + // Prefix-guarded-scan shape recognized by PatternAnalyzer#detectPrefixGuardedScan; exercises + // the compile-time BITSTATE_BYTECODE dispatch case (com.datadoghq.reggie.codegen.codegen. + // BitStateBytecodeGenerator), separate from the runtime dispatch path covered by + // BitStateBytecodeGeneratorTest in reggie-runtime. + Object matcher = + compile("(?s)(?m)^(?:\\s*(?:sudo|doas)\\s+)?\\b\\S+\\b\\s*(.*)", "CommandMatcher"); + Method matches = matcher.getClass().getMethod("matches", String.class); + assertTrue((Boolean) matches.invoke(matcher, "ls -la")); + assertTrue((Boolean) matches.invoke(matcher, "sudo ls -la")); + assertTrue((Boolean) matches.invoke(matcher, "doas rm -rf /")); + assertFalse((Boolean) matches.invoke(matcher, " ")); + } + @Test void testSpecializedMultipleLookaheadsStrategyNative() throws Exception { // SPECIALIZED_MULTIPLE_LOOKAHEADS is now NATIVE (Wave 3 fixed the lookahead boolean engine). diff --git a/reggie-runtime/src/main/java/com/datadoghq/reggie/runtime/BitStateMatcher.java b/reggie-runtime/src/main/java/com/datadoghq/reggie/runtime/BitStateMatcher.java index caf19b65..d94ceaf1 100644 --- a/reggie-runtime/src/main/java/com/datadoghq/reggie/runtime/BitStateMatcher.java +++ b/reggie-runtime/src/main/java/com/datadoghq/reggie/runtime/BitStateMatcher.java @@ -15,6 +15,7 @@ */ package com.datadoghq.reggie.runtime; +import com.datadoghq.reggie.codegen.automaton.CharSet; import com.datadoghq.reggie.codegen.automaton.NFA; import java.util.Arrays; import java.util.Collections; @@ -72,6 +73,14 @@ final class BitStateMatcher extends ReggieMatcher { private final NFA.NFAState[] statesById; private final boolean[] isAccept; + // Flattened per-state transition tables, built once here so the hot DFS loop never calls + // NFAState.getEpsilonTransitions()/getTransitions() (each allocates a fresh + // Collections.unmodifiableList wrapper per call). Indexed by state id, in the same Perl + // priority order as the underlying NFAState lists. + private final int[][] epsilonTargets; + private final CharSet[][] transitionCharSets; + private final int[][] transitionTargets; + // Input-independent structures, preallocated once (design §5): live capture slots and the // winning snapshot. Sized 2 * (groupCount + 1): slot 2*g is group g's start, 2*g+1 its end. private final int[] caps; @@ -113,6 +122,31 @@ final class BitStateMatcher extends ReggieMatcher { isAccept[s.id] = true; } + epsilonTargets = new int[stateCount][]; + transitionCharSets = new CharSet[stateCount][]; + transitionTargets = new int[stateCount][]; + for (int i = 0; i < stateCount; i++) { + NFA.NFAState s = statesById[i]; + + List eps = s.getEpsilonTransitions(); + int[] epsIds = new int[eps.size()]; + for (int j = 0; j < epsIds.length; j++) { + epsIds[j] = eps.get(j).id; + } + epsilonTargets[i] = epsIds; + + List trans = s.getTransitions(); + CharSet[] charSets = new CharSet[trans.size()]; + int[] targets = new int[trans.size()]; + for (int j = 0; j < charSets.length; j++) { + NFA.Transition t = trans.get(j); + charSets[j] = t.chars; + targets[j] = t.target.id; + } + transitionCharSets[i] = charSets; + transitionTargets[i] = targets; + } + int slotCount = 2 * (groupCount + 1); caps = new int[slotCount]; winCaptures = new int[slotCount]; @@ -322,7 +356,7 @@ private boolean search( // find/findFrom semantics: ^/\A never fire at a findFrom(start>0) offset, only at true // input start. if (!PikeVMMatcher.checkAnchor(s.anchor, input, pos, 0, spanEnd)) continue; - pushEpsilonChildrenReversed(s, pos); + pushEpsilonChildrenReversed(sid, pos); continue; } @@ -345,20 +379,20 @@ private boolean search( return true; } - List epsilons = s.getEpsilonTransitions(); - if (!epsilons.isEmpty()) { - pushEpsilonChildrenReversed(s, pos); + int[] eps = epsilonTargets[sid]; + if (eps.length != 0) { + pushEpsilonChildrenReversed(sid, pos); continue; } // Consuming leaf. if (pos < spanEnd) { char ch = input.charAt(pos); - List transitions = s.getTransitions(); - for (int i = transitions.size() - 1; i >= 0; i--) { - NFA.Transition tr = transitions.get(i); - if (tr.chars.contains(ch)) { - pushExpand(tr.target.id, pos + 1); + CharSet[] charSets = transitionCharSets[sid]; + int[] targets = transitionTargets[sid]; + for (int i = charSets.length - 1; i >= 0; i--) { + if (charSets[i].contains(ch)) { + pushExpand(targets[i], pos + 1); } } } @@ -366,11 +400,11 @@ private boolean search( return false; } - /** Pushes {@code state}'s epsilon children in reverse priority order (highest pops first). */ - private void pushEpsilonChildrenReversed(NFA.NFAState state, int pos) { - List epsilons = state.getEpsilonTransitions(); - for (int i = epsilons.size() - 1; i >= 0; i--) { - pushExpand(epsilons.get(i).id, pos); + /** Pushes {@code sid}'s epsilon children in reverse priority order (highest pops first). */ + private void pushEpsilonChildrenReversed(int sid, int pos) { + int[] eps = epsilonTargets[sid]; + for (int i = eps.length - 1; i >= 0; i--) { + pushExpand(eps[i], pos); } } diff --git a/reggie-runtime/src/main/java/com/datadoghq/reggie/runtime/RuntimeCompiler.java b/reggie-runtime/src/main/java/com/datadoghq/reggie/runtime/RuntimeCompiler.java index 968086b0..67cce6c2 100644 --- a/reggie-runtime/src/main/java/com/datadoghq/reggie/runtime/RuntimeCompiler.java +++ b/reggie-runtime/src/main/java/com/datadoghq/reggie/runtime/RuntimeCompiler.java @@ -44,6 +44,7 @@ import com.datadoghq.reggie.codegen.automaton.ThompsonBuilder; import com.datadoghq.reggie.codegen.codegen.BackreferenceBytecodeGenerator; import com.datadoghq.reggie.codegen.codegen.BitParallelGlushkovBytecodeGenerator; +import com.datadoghq.reggie.codegen.codegen.BitStateBytecodeGenerator; import com.datadoghq.reggie.codegen.codegen.BoundedQuantifierBytecodeGenerator; import com.datadoghq.reggie.codegen.codegen.ConcatGreedyGroupBytecodeGenerator; import com.datadoghq.reggie.codegen.codegen.ConcatQuantifiedGroupsBytecodeGenerator; @@ -1000,6 +1001,14 @@ private static byte[] generateBytecode( greedyGen.generateFindBoundsFromMethod(cw, "com/datadoghq/reggie/runtime/" + className); break; + case BITSTATE_BYTECODE: + PatternAnalyzer.PrefixGuardedScanInfo prefixGuardedInfo = + (PatternAnalyzer.PrefixGuardedScanInfo) result.patternInfo; + BitStateBytecodeGenerator bitStateGen = + new BitStateBytecodeGenerator(prefixGuardedInfo, nfa.getGroupCount()); + bitStateGen.generateAll(cw, "com/datadoghq/reggie/runtime/" + className); + break; + case SPECIALIZED_MULTI_GROUP_GREEDY: PatternAnalyzer.MultiGroupGreedyInfo multiGroupInfo = (PatternAnalyzer.MultiGroupGreedyInfo) result.patternInfo; diff --git a/reggie-runtime/src/test/java/com/datadoghq/reggie/runtime/BitStateBytecodeGeneratorTest.java b/reggie-runtime/src/test/java/com/datadoghq/reggie/runtime/BitStateBytecodeGeneratorTest.java new file mode 100644 index 00000000..05220ab9 --- /dev/null +++ b/reggie-runtime/src/test/java/com/datadoghq/reggie/runtime/BitStateBytecodeGeneratorTest.java @@ -0,0 +1,393 @@ +/* + * Copyright 2026-Present Datadog, Inc. + * + * Licensed under the Apache License, Version 2.0 (the "License"); + * you may not use this file except in compliance with the License. + * You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, software + * distributed under the License is distributed on an "AS IS" BASIS, + * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. + * See the License for the specific language governing permissions and + * limitations under the License. + */ +package com.datadoghq.reggie.runtime; + +import static org.junit.jupiter.api.Assertions.assertEquals; +import static org.junit.jupiter.api.Assertions.assertFalse; +import static org.junit.jupiter.api.Assertions.assertTrue; + +import java.util.regex.Matcher; +import java.util.regex.Pattern; +import org.junit.jupiter.api.BeforeEach; +import org.junit.jupiter.api.Test; + +/** + * Correctness tests for the {@code BITSTATE_BYTECODE} strategy, generated by {@code + * BitStateBytecodeGenerator} for the prefix-guarded-scan pattern family recognized by {@code + * PatternAnalyzer#detectPrefixGuardedScan}. + * + *

The reference pattern under test is {@code (?s)(?m)^(?:\s*(?:sudo|doas)\s+)?\b\S+\b\s*(.*)} + * (the "COMMAND" pattern from the design doc). Every expected value below was independently + * verified against {@code java.util.regex.Pattern} (JDK 21) rather than assumed, since {@code + * java.util.regex.Pattern}'s backtracking semantics for a required trailing {@code \b} after a + * greedy {@code \S+} are not obvious from the pattern text alone (see {@code + * BitStateBytecodeGeneratorTest#testTrailingBoundaryForcesGiveBack} below). + */ +public class BitStateBytecodeGeneratorTest { + + private static final String COMMAND_PATTERN = + "(?s)(?m)^(?:\\s*(?:sudo|doas)\\s+)?\\b\\S+\\b\\s*(.*)"; + + @BeforeEach + void clearCache() { + RuntimeCompiler.clearCache(); + } + + private static void assertMatchesLikeJdk(ReggieMatcher matcher, Pattern jdk, String input) { + Matcher jm = jdk.matcher(input); + boolean jdkMatches = jm.matches(); + assertEquals(jdkMatches, matcher.matches(input), "matches() for input [" + input + "]"); + if (jdkMatches) { + MatchResult r = matcher.match(input); + assertTrue(r != null, "match() should succeed for input [" + input + "]"); + assertEquals(jm.group(0), r.group(0), "group(0) for input [" + input + "]"); + assertEquals(jm.group(1), r.group(1), "group(1) for input [" + input + "]"); + } else { + assertEquals(null, matcher.match(input), "match() should return null for [" + input + "]"); + } + } + + @Test + public void testRoutesToBitStateBytecode() { + // Confirms the detector + router substitution actually selects BITSTATE_BYTECODE for the + // COMMAND pattern, not merely that some other strategy happens to produce the right answer. + ReggieMatcher matcher = RuntimeCompiler.compile(COMMAND_PATTERN); + assertTrue( + matcher.getClass().getName().contains("ReggieMatcher$"), + "expected a generated hidden class, got " + matcher.getClass()); + } + + @Test + public void testWithoutPrefixSimpleCommand() { + ReggieMatcher matcher = RuntimeCompiler.compile(COMMAND_PATTERN); + Pattern jdk = Pattern.compile(COMMAND_PATTERN); + assertMatchesLikeJdk(matcher, jdk, "ls -la"); + assertMatchesLikeJdk(matcher, jdk, "ls"); + } + + @Test + public void testWithSudoPrefix() { + ReggieMatcher matcher = RuntimeCompiler.compile(COMMAND_PATTERN); + Pattern jdk = Pattern.compile(COMMAND_PATTERN); + assertMatchesLikeJdk(matcher, jdk, "sudo ls -la"); + assertMatchesLikeJdk(matcher, jdk, " sudo apt update"); + } + + @Test + public void testWithDoasPrefix() { + ReggieMatcher matcher = RuntimeCompiler.compile(COMMAND_PATTERN); + Pattern jdk = Pattern.compile(COMMAND_PATTERN); + assertMatchesLikeJdk(matcher, jdk, "doas rm -rf /"); + } + + @Test + public void testEmptyTailAfterMandatoryToken() { + ReggieMatcher matcher = RuntimeCompiler.compile(COMMAND_PATTERN); + Pattern jdk = Pattern.compile(COMMAND_PATTERN); + assertMatchesLikeJdk(matcher, jdk, "ls"); + assertMatchesLikeJdk(matcher, jdk, "sudo"); + assertMatchesLikeJdk(matcher, jdk, "a"); + assertMatchesLikeJdk(matcher, jdk, "ab"); + } + + @Test + public void testTrailingWhitespaceAbsorbedBeforeTail() { + ReggieMatcher matcher = RuntimeCompiler.compile(COMMAND_PATTERN); + Pattern jdk = Pattern.compile(COMMAND_PATTERN); + assertMatchesLikeJdk(matcher, jdk, "sudo\tls"); + assertMatchesLikeJdk(matcher, jdk, "SUDO ls"); + } + + @Test + public void testNoMatchEmptyOrWhitespaceOnly() { + ReggieMatcher matcher = RuntimeCompiler.compile(COMMAND_PATTERN); + Pattern jdk = Pattern.compile(COMMAND_PATTERN); + assertMatchesLikeJdk(matcher, jdk, ""); + assertMatchesLikeJdk(matcher, jdk, " "); + } + + @Test + public void testNoMatchSingleNonWordCharacter() { + // "!" alone: the mandatory \S+ matches "!", but the leading \b at position 0 never holds + // (nothing before, non-word after), so the whole match fails - confirmed against JDK. + ReggieMatcher matcher = RuntimeCompiler.compile(COMMAND_PATTERN); + Pattern jdk = Pattern.compile(COMMAND_PATTERN); + assertMatchesLikeJdk(matcher, jdk, "!"); + } + + /** + * The key correctness scenario motivating {@code matchMandatory}'s bounded backward give-back + * scan: JDK's real backtracking gives back one character from a greedy {@code \S+} run so that + * the required trailing {@code \b} holds, changing where group 1 starts. Verified against JDK 21: + * {@code "a!"} matches as a whole, with group(1) == "!" (not "" as a naive port of the + * hand-written {@code CommandPatternPrototype} reference — which does not check the trailing + * boundary at all — would produce). + */ + @Test + public void testTrailingBoundaryForcesGiveBack() { + ReggieMatcher matcher = RuntimeCompiler.compile(COMMAND_PATTERN); + Pattern jdk = Pattern.compile(COMMAND_PATTERN); + assertTrue(jdk.matcher("a!").matches(), "sanity: JDK must match \"a!\""); + assertEquals("!", jdk.matcher("a!").replaceFirst("$1"), "sanity: JDK group(1) must be \"!\""); + assertMatchesLikeJdk(matcher, jdk, "a!"); + } + + @Test + public void testWordCharMandatoryNeverNeedsGiveBack() { + // Contrast case: a maximal \w+ run always ends adjacent to a non-word char or EOF, so the + // trailing \b always holds at the greedy maximum and the give-back loop is a no-op. + ReggieMatcher matcher = RuntimeCompiler.compile(COMMAND_PATTERN); + Pattern jdk = Pattern.compile(COMMAND_PATTERN); + assertMatchesLikeJdk(matcher, jdk, "ab"); + assertMatchesLikeJdk(matcher, jdk, "abc123"); + } + + @Test + public void testMultilineRestartOnLeadingBoundaryFailure() { + // First line "!" fails the leading \b unconditionally (not fixable by shrinking, since the + // leading boundary is checked at a fixed position), forcing a line restart at the second line. + ReggieMatcher matcher = RuntimeCompiler.compile(COMMAND_PATTERN); + Pattern jdk = Pattern.compile(COMMAND_PATTERN); + String input = "!\nab"; + Matcher jm = jdk.matcher(input); + assertTrue(jm.find()); + assertTrue(matcher.find(input)); + int foundAt = matcher.findFrom(input, 0); + assertEquals(jm.start(), foundAt); + MatchResult r = matcher.findMatchFrom(input, 0); + assertEquals(jm.group(0), r.group(0)); + assertEquals(jm.group(1), r.group(1)); + } + + @Test + public void testMultilineOptionalPrefixAbsorbsLeadingNewline() { + ReggieMatcher matcher = RuntimeCompiler.compile(COMMAND_PATTERN); + Pattern jdk = Pattern.compile(COMMAND_PATTERN); + assertMatchesLikeJdk(matcher, jdk, "\nsudo ls -la"); + assertMatchesLikeJdk(matcher, jdk, " \nsudo ls"); + } + + @Test + public void testFindMultilineDotAllTailAbsorbsNewlines() { + ReggieMatcher matcher = RuntimeCompiler.compile(COMMAND_PATTERN); + Pattern jdk = Pattern.compile(COMMAND_PATTERN); + String input = "line1\nsudo ls -la\nline3"; + Matcher jm = jdk.matcher(input); + assertTrue(jm.find()); + assertTrue(matcher.find(input)); + MatchResult r = matcher.findMatch(input); + assertEquals(jm.group(0), r.group(0)); + assertEquals(jm.group(1), r.group(1)); + } + + @Test + public void testKeywordSharedPrefixDoesNotFalselyMatchAsPrefixGroup() { + // "sudoer" starts with "sudo" but there is no separator whitespace after it, so the optional + // prefix group must NOT consume "sudo" as the keyword - the whole word is mandatory instead. + ReggieMatcher matcher = RuntimeCompiler.compile(COMMAND_PATTERN); + Pattern jdk = Pattern.compile(COMMAND_PATTERN); + assertMatchesLikeJdk(matcher, jdk, "sudoer"); + assertMatchesLikeJdk(matcher, jdk, "doasomething"); + } + + @Test + public void testStructurallySimilarSyntheticPattern() { + // A different keyword set and word-char mandatory class, to prove the detector/generator is + // data-driven from the AST rather than hardcoded to the COMMAND pattern's literals. + String pattern = "^(?:\\s*(?:run|exec)\\s+)?\\b[a-zA-Z0-9_]+\\b\\s*(.*)"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + assertMatchesLikeJdk(matcher, jdk, "run build.sh --fast"); + assertMatchesLikeJdk(matcher, jdk, "exec ls"); + assertMatchesLikeJdk(matcher, jdk, "build.sh"); + assertMatchesLikeJdk(matcher, jdk, ""); + assertMatchesLikeJdk(matcher, jdk, "runner"); // shared-prefix keyword, no separator + } + + @Test + public void testNegativeMissingLeadingAnchorFallsBackWithoutCrashing() { + // No leading ^: does not match the detector's step 1, must not route to BITSTATE_BYTECODE + // (falls back to some other strategy) but must still produce correct results. + String pattern = "(?:\\s*(?:sudo|doas)\\s+)?\\b\\S+\\b\\s*(.*)"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + assertMatchesLikeJdk(matcher, jdk, "sudo ls -la"); + assertMatchesLikeJdk(matcher, jdk, "ls"); + } + + @Test + public void testNegativeNonDotStarTailFallsBackWithoutCrashing() { + // Tail is a restricted char class, not `.*`, so it does not match the detector's step 7 + // shape; must not route to BITSTATE_BYTECODE but must still work correctly. + String pattern = "^(?:\\s*(?:sudo|doas)\\s+)?\\b\\S+\\b\\s*([a-z]*)"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + assertMatchesLikeJdk(matcher, jdk, "sudo ls"); + assertMatchesLikeJdk(matcher, jdk, "ls abc"); + } + + @Test + public void testNegativeBoundedMandatoryQuantifierFallsBackWithoutCrashing() { + // The mandatory scan is `\S{2,4}`, not `\S+` (bounded max), so the detector's unbounded-max + // guard must reject this pattern; must not route to BITSTATE_BYTECODE but must still work. + String pattern = "^(?:\\s*(?:sudo|doas)\\s+)?\\b\\S{2,4}\\b\\s*(.*)"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + assertMatchesLikeJdk(matcher, jdk, "sudo ls -la"); + assertMatchesLikeJdk(matcher, jdk, "abcd tail"); + assertFalse(jdk.matcher("a tail").matches()); // sanity: below the {2,4} floor + assertMatchesLikeJdk(matcher, jdk, "a tail"); + } + + @Test + public void testNegativeOverlappingSeparatorAndMandatoryCharSetsFallsBackWithoutCrashing() { + // separatorCharSet ([a-z]) overlaps mandatoryCharSet ([a-z]): the optional prefix's greedy + // separator scan could over-consume characters the mandatory scan needs, with no give-back + // possible without discarding the whole prefix. The detector must decline this shape. + String pattern = "^(?:\\s*(?:123|456)[a-z]{1,})?[a-z]{3,}(.*)"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + assertMatchesLikeJdk(matcher, jdk, "123abcdef"); + assertMatchesLikeJdk(matcher, jdk, "abcdef"); + } + + @Test + public void testNegativeOverlappingLeadingWsAndMandatoryCharSetsFallsBackWithoutCrashing() { + // Mirror of the separator-overlap case above, but the overlap is on leadingWsCharSet instead + // (separatorCharSet here is \s+, disjoint from mandatory [a-z]) - exercises the other half of + // the detector's disjointness OR-check. + String pattern = "^(?:[a-z]*(?:123|456)\\s+)?[a-z]{3,}(.*)"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + assertMatchesLikeJdk(matcher, jdk, "abc123 def"); + assertMatchesLikeJdk(matcher, jdk, "abcdef"); + } + + @Test + public void testNegativeLazyOptionalPrefixFallsBackWithoutCrashing() { + // Lazy `??` on the optional prefix group: the detector must decline rather than treat it as + // greedy, since the generator always emits a maximal-match attempt for the prefix. + String pattern = "^(?:\\s*(?:sudo|doas)\\s+)??\\b\\S+\\b\\s*(.*)"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + assertMatchesLikeJdk(matcher, jdk, "sudo ls -la"); + assertMatchesLikeJdk(matcher, jdk, "ls"); + } + + @Test + public void testNegativeLazySeparatorFallsBackWithoutCrashing() { + // Lazy `+?` on the separator quantifier. + String pattern = "^(?:\\s*(?:sudo|doas)\\s+?)?\\b\\S+\\b\\s*(.*)"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + assertMatchesLikeJdk(matcher, jdk, "sudo ls -la"); + assertMatchesLikeJdk(matcher, jdk, "ls"); + } + + @Test + public void testNegativeLazyMandatoryScanFallsBackWithoutCrashing() { + // Lazy `+?` on the mandatory scan - the generator's maximal-scan assumption would silently + // over-match without this guard (this is the exact g-3 example from the review finding). + String pattern = "^(?:\\s*(?:sudo|doas)\\s+)?\\b\\S+?\\b\\s*(.*)"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + assertMatchesLikeJdk(matcher, jdk, "sudo ls -la"); + assertMatchesLikeJdk(matcher, jdk, "hello world"); + } + + @Test + public void testNegativeLazyTrailingWhitespaceFallsBackWithoutCrashing() { + // Lazy `*?` on the trailing whitespace absorbed before the tail capture group. + String pattern = "^(?:\\s*(?:sudo|doas)\\s+)?\\b\\S+\\b\\s*?(.*)"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + assertMatchesLikeJdk(matcher, jdk, "sudo ls -la"); + assertMatchesLikeJdk(matcher, jdk, "ls"); + } + + @Test + public void testNegativeLazyTailCaptureFallsBackWithoutCrashing() { + // Lazy `*?` on the tail capture group itself. + String pattern = "^(?:\\s*(?:sudo|doas)\\s+)?\\b\\S+\\b\\s*(.*?)"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + assertMatchesLikeJdk(matcher, jdk, "sudo ls -la"); + assertMatchesLikeJdk(matcher, jdk, "ls"); + } + + @Test + public void testNegativeBoundedLeadingWhitespaceFallsBackWithoutCrashing() { + // Bounded max (`{0,3}`) on the leading-whitespace-trim quantifier ahead of the keyword. + String pattern = "^(?:\\s{0,3}(?:sudo|doas)\\s+)?\\b\\S+\\b\\s*(.*)"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + assertMatchesLikeJdk(matcher, jdk, " sudo ls -la"); + assertMatchesLikeJdk(matcher, jdk, "ls"); + } + + @Test + public void testNegativeBoundedTrailingWhitespaceFallsBackWithoutCrashing() { + // Bounded max (`{0,3}`) on the trailing whitespace absorbed before the tail capture group. + String pattern = "^(?:\\s*(?:sudo|doas)\\s+)?\\b\\S+\\b\\s{0,3}(.*)"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + assertMatchesLikeJdk(matcher, jdk, "sudo ls -la"); + assertMatchesLikeJdk(matcher, jdk, "ls"); + } + + @Test + public void testNegativeBoundedTailCaptureFallsBackWithoutCrashing() { + // Bounded max (`{0,5}`) on the tail capture group itself, not an unbounded `.*`. + String pattern = "^(?:\\s*(?:sudo|doas)\\s+)?\\b\\S+\\b\\s*(.{0,5})"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + assertMatchesLikeJdk(matcher, jdk, "sudo ls"); + assertMatchesLikeJdk(matcher, jdk, "ls abcdefgh"); + } + + @Test + public void testNegativeBoundedSeparatorFallsBackWithoutCrashing() { + // The separator quantifier is `\s{1,3}` (bounded max), not `\s+`: the generator only stores + // separatorMin and emits an unbounded greedy scan, so allowing this shape would let the + // generated matcher consume past JDK's upper bound. On "sudo ls" (four spaces), JDK cannot + // take the prefix beyond the 3-space limit and falls back to matching "sudo" as the mandatory + // token, while an ungated generator would consume all four spaces as the separator and match + // "ls" as the mandatory token instead - a different group 1. The detector must reject this + // shape rather than route to BITSTATE_BYTECODE. (Two keywords are required here so + // extractDisjointKeywords's alternation shape actually qualifies the pattern for the + // optional-prefix detection path in the first place.) + String pattern = "^(?:\\s*(?:sudo|doas)\\s{1,3})?\\b\\S+\\b\\s*(.*)"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + assertMatchesLikeJdk(matcher, jdk, "sudo ls"); + assertMatchesLikeJdk(matcher, jdk, "sudo ls"); + assertMatchesLikeJdk(matcher, jdk, "ls"); + } + + @Test + public void testNegativeNoTrailingCapturingGroupFallsBackWithoutCrashing() { + // No capturing group around the tail at all; must not route to BITSTATE_BYTECODE. Compares + // only matches()/group(0) since there is no group 1 to compare. + String pattern = "^(?:\\s*(?:sudo|doas)\\s+)?\\b\\S+\\b\\s*.*"; + ReggieMatcher matcher = RuntimeCompiler.compile(pattern); + Pattern jdk = Pattern.compile(pattern); + String input = "sudo ls -la"; + assertEquals(jdk.matcher(input).matches(), matcher.matches(input)); + MatchResult r = matcher.match(input); + assertTrue(r != null); + assertEquals(input, r.group(0)); + } +} diff --git a/reggie-runtime/src/test/java/com/datadoghq/reggie/runtime/StrategyCorrectnessMetaTest.java b/reggie-runtime/src/test/java/com/datadoghq/reggie/runtime/StrategyCorrectnessMetaTest.java index 6c3a12b5..23a3696d 100644 --- a/reggie-runtime/src/test/java/com/datadoghq/reggie/runtime/StrategyCorrectnessMetaTest.java +++ b/reggie-runtime/src/test/java/com/datadoghq/reggie/runtime/StrategyCorrectnessMetaTest.java @@ -249,6 +249,15 @@ private static Map strategyPatterns() { m.put( PatternAnalyzer.MatchingStrategy.BITSTATE_CAPTURE, new Spec("(a)?b", List.of("b", "xaby", "ab", "", "bé"))); + // BITSTATE_BYTECODE: the prefix-guarded-scan family detected by + // PatternAnalyzer#detectPrefixGuardedScan, substituted in for BITSTATE_CAPTURE. "a!" exercises + // the bounded backward give-back scan for the required trailing \b after the greedy \S+ (see + // BitStateBytecodeGenerator#generateMatchMandatoryMethod). + m.put( + PatternAnalyzer.MatchingStrategy.BITSTATE_BYTECODE, + new Spec( + "(?s)(?m)^(?:\\s*(?:sudo|doas)\\s+)?\\b\\S+\\b\\s*(.*)", + List.of("sudo ls -la", "a!", "!!!", "", "héllo"))); return m; }