diff --git a/doc/fuzz/2026-06-29.md b/doc/fuzz/2026-06-29.md new file mode 100644 index 00000000..7ebb9b66 --- /dev/null +++ b/doc/fuzz/2026-06-29.md @@ -0,0 +1,196 @@ +# Fuzz Sweep — 2026-06-29 + +**Seed:** `0xC0DEFEED_DEADBEEFL` (`BASE_SEED`) +**Range:** patterns 25 001–50 000 (`skip=25_000`, `count=25_000`) +**Depth:** 3 · **Inputs/pattern:** 16 · **Input max length:** 16 +**Raw findings:** 34 · **Unique minimal repros:** 29 +**Baseline (0–25 000):** 0 divergences (established after B3a/B3b/B4/B5/B6 fixes) + +> **Note on skip semantics:** `skip=N` advances both the pattern RNG and the input RNG by +> `N × inputsPerPattern` steps before the main loop. This is conservative relative to a real +> 50k run (which advances the input RNG by fewer steps for compile-rejected patterns), so the +> inputs for this window differ from a full 0–50k run. A one-shot 50k run produces 43 raw +> findings over this range; with skip=25 000 the canonical count is 34. Findings are fully +> reproducible given the same `(seed, skip, count, inputsPerPattern, depth)` tuple. + +--- + +## How to reproduce + +Exact reproduction (hardcoded in `divergenceGate`, `KNOWN_FINDINGS_BUDGET=34`): + +```bash +cd +./gradlew :reggie-integration-tests:test \ + --tests "*.AlgorithmicFuzzTest.divergenceGate" \ + --no-daemon +``` + +Manual equivalent with explicit parameters: + +```bash +./gradlew :reggie-integration-tests:test \ + --tests "*.AlgorithmicFuzzTest.divergenceGate" \ + -Dreggie.fuzz.skip=25000 \ + -Dreggie.fuzz.size=25000 \ + -Dreggie.fuzz.maxFindings=9999 \ + --no-daemon +``` + +Advance to the next window once this budget reaches 0 (patterns 50 001–75 000): + +```bash +./gradlew :reggie-integration-tests:test \ + --tests "*.AlgorithmicFuzzTest.divergenceGate" \ + -Dreggie.fuzz.skip=50000 \ + -Dreggie.fuzz.size=25000 \ + -Dreggie.fuzz.maxFindings=9999 \ + --no-daemon +``` + +Then document new findings in `doc/fuzz/YYYY-MM-DD.md`, update `largeSweepConfig` default skip +to 50 000, and set `KNOWN_FINDINGS_BUDGET` to the new count. + +--- + +## Summary + +29 unique minimal repros across six root-cause classes (E1–E6). All are pre-existing bugs in +native strategies — none are regressions introduced by the B3a/B3b/B4/B5/B6 routing fixes. + +--- + +## E1 — Find-path group span overextension (2 patterns, 3 findings) + +A quantified prefix is followed by a capturing group. On the `findAll()` path the group span +extends beyond the chars it should own. Same root-cause class as A1/A2; those fixes addressed +`matches()` and the first `find()`. These patterns exercise subsequent `findAll()` iterations. + +| Pattern | Input | Symptom | +|---------|-------|---------| +| `[^-]{3,}([b].)` | `0acbb1` | `findAll() match 0 group 1 span differs` | +| `[^-]{3,}([b].)` | `c1cb-` | `findAll() match 0 group 1 span differs` | +| `[^-]{3,}([c].)` | `a0bc-` | `findAll() match 0 group 1 span differs` | + +--- + +## E2 — Anchor at unusual position (4 patterns, 9 findings) + +An anchor appears inside a group body, after a quantified group, or combined with alternation +and a backreference. The routing logic does not cover all these forms. + +### E2a — `\A` + repeated capturing group + +`\A` with `{n,}` on a capturing group. The strategy treats `\A` as handled but repeated-group +span bookkeeping conflicts with the start-anchor assertion. + +| Pattern | Input | Symptom | +|---------|-------|---------| +| `\A(c){1,}` | `ac` | `find() boolean differs` | +| `\A(c){1,}` | `ac` | `findAll() count differs` | +| `\A(c){1,}` | `0c` | `find() boolean differs` | + +### E2b — Quantified charset + `\Z`, no capturing group + +The outer match span (group 0) is wrong — a match-boundary error, not a capture-tracking issue. +`hasStringEndAnchorInAlternation` does not apply (no alternation); a separate guard for `\Z` +after a plain quantifier is needed. + +| Pattern | Input | Symptom | +|---------|-------|---------| +| `[^c]*\Z` | `` | `first-match span differs` | +| `[^c]*\Z` | `a` | `findAll() count differs` | + +### E2c — `^` after a quantified group + +`^` appears as a zero-width assertion immediately after a quantified capturing group. Not routed +to a strategy that handles post-quantifier anchors. + +| Pattern | Input | Symptom | +|---------|-------|---------| +| `(0)+^` | `0` | `find() boolean differs` | +| `(0)+^` | `0` | `findAll() count differs` | + +### E2d — Anchor inside alternation combined with backreference + +The routing logic handles `\A`-in-alternation but not when a backreference is also present. + +| Pattern | Input | Symptom | +|---------|-------|---------| +| `(c\|a?){3}\A\1?` | `c` | `find() boolean differs` | +| `(c\|a?){3}\A\1?` | `c` | `findAll() count differs` | + +--- + +## E3 — Backreference divergence (5 patterns, 9 findings) + +The backreference resolves to a wrong value or the match boolean is wrong. Distinct from E1: +the entire match succeeds or fails where the JDK says otherwise. + +| Pattern | Input | Symptom | +|---------|-------|---------| +| `(.b{0})?\1` | `00` | `find() boolean differs` | +| `b(-)\1{1}` | `--` | `find() boolean differs` | +| `b(-)\1{1}` | `--` | `findAll() count differs` | +| `${1}[^a]` | `` | `find() boolean differs` | +| `${1}[^a]` | `` | `findAll() count differs` | +| `(b{1,}){1}\1+\|(])` | `bb` | `findAll() count differs` | +| `^(.{1}\|.{0}){4}\1{3}` | `1cb0` | `find() boolean differs` | +| `^(.{1}\|.{0}){4}\1{3}` | `1cb0` | `findAll() count differs` | +| `^(.{1}\|.{0}){4}\1{3}` | `00_1` | `find() boolean differs` | + +--- + +## E4 — Repeated group last-iteration span (2 patterns, 3 findings) + +After a group is matched multiple times the final captured span is wrong on `findAll()`. +The per-iteration span-reset logic does not fire correctly for the last iteration before the +quantifier exits. + +| Pattern | Input | Symptom | +|---------|-------|---------| +| `(-+)*` | `1` | `findAll() match 4 group 1 span differs` | +| `(c{2}){1,}` | `c` | `find() boolean differs` | +| `(c{2}){1,}` | `c` | `findAll() count differs` | + +--- + +## E5 — Alternation with mixed-width branches (1 pattern, 2 findings) + +Complex alternation where branches have different widths; `find()` boolean and `findAll()` count +are both wrong. Distinct from B5 (variable-length body inside a capturing group): here the outer +match boolean is wrong, not an inner span. + +| Pattern | Input | Symptom | +|---------|-------|---------| +| `(-{0}[1]{1})([_][^_]\|\1*).{2}` | `1_0` | `find() boolean differs` | +| `(-{0}[1]{1})([_][^_]\|\1*).{2}` | `1_0` | `findAll() count differs` | + +--- + +## E6 — Simple group routing gap (1 pattern, 3 findings) + +A straightforward pattern routes to a strategy that produces wrong `find()` results. Likely a +missing routing guard; use `debugPattern` to confirm the assigned strategy. + +| Pattern | Input | Symptom | +|---------|-------|---------| +| `[1]([^b]{2})` | `110` | `find() boolean differs` | +| `[1]([^b]{2})` | `110` | `findAll() count differs` | +| `[1]([^b]{2})` | `10c` | `find() boolean differs` | + +--- + +## Next steps (priority order) + +1. **E2b** — add a `\Z`-after-quantifier routing guard (no alternation required); companion to + `hasStringEndAnchorInAlternation`. +2. **E6** — run `./gradlew :reggie-runtime:debugPattern -Ppattern='[1]([^b]{2})'` to identify + the misrouted strategy; likely a one-line guard fix. +3. **E1** — find-path group span: remaining cases after A1/A2; same root-cause, new trigger + shapes (`[^-]{3,}([b].)` — quantified non-capturing prefix followed by a capturing group). +4. **E2a/E2c** — `\A`/`^` + repeated group: new anchor-in-repetition routing class. +5. **E3** — backreference edge cases: optional-group backref, quantified backref in alternation. +6. **E4** — repeated group final-iteration span: per-iteration span-reset in TDFA. +7. **E5** — mixed-width alternation outer match: `find()` boolean wrong on alternation with + backreference branch. diff --git a/docs/sphinx/specs/2026-06-30-preserve-fallback-for-pikevm-anchor-quantifi.md b/docs/sphinx/specs/2026-06-30-preserve-fallback-for-pikevm-anchor-quantifi.md new file mode 100644 index 00000000..79e26509 --- /dev/null +++ b/docs/sphinx/specs/2026-06-30-preserve-fallback-for-pikevm-anchor-quantifi.md @@ -0,0 +1,73 @@ +# Spec: preserve-fallback-for-pikevm-anchor-quantifi + +## Problem + +`ReggieMatcherBytecodeGenerator.resolveRealization()` returns `DELEGATE_PIKEVM` for any +pattern whose `PatternAnalyzer` result is `PIKEVM_CAPTURE` **without first consulting +`FallbackPatternDetector.needsFallback()`**. This means patterns that trigger the B3a route +in `PatternAnalyzer` (anchor-only capturing group, e.g. `($)`) but which also carry a +quantifier on that group (e.g. `($){2}`, `(^)?`) reach the compile-time PIKEVM path even +though `FallbackPatternDetector` marks them unsafe via the B2 +(`hasAnchorInQuantifierInCapturingGroup`) and B3 (`hasAnchorInQuantifier`) KEEP-PERMANENT +guards. + +The same gap exists in `RuntimeCompiler.compilePikeVm()` (called by generated stubs): it +only applies the B16 nullable-group-content guard, missing B2/B3. + +The runtime `Reggie.compile()` path is correct: it calls +`FallbackPatternDetector.needsFallback(ast, PIKEVM_CAPTURE)` at line 494 and falls back to +JDK when non-null. + +## Correct behaviour + +1. When `@RegexPattern` annotation processing routes a pattern to `PIKEVM_CAPTURE`: + - `resolveRealization()` must call `FallbackPatternDetector.needsFallback(ast, PIKEVM_CAPTURE)`. + - If the result is non-null, the pattern requires JDK fallback: + - If `allowJdkFallback` is set → return `DELEGATE_FALLBACK`. + - Otherwise → throw `UnsupportedOperationException` with the fallback reason, consistent + with the existing `needsJdk` error message format. + - If the result is null → return `DELEGATE_PIKEVM` as before. + +2. `RuntimeCompiler.compilePikeVm()` must apply the full `needsFallback()` guard (not just + the ad-hoc `hasNullableGroupContentWithNullableQuantifier` check). If the result is + non-null → throw `UnsupportedPatternException` with the reason string. + +3. Patterns like `($)` (no quantifier on the anchor-only group) remain unaffected: + `hasAnchorInQuantifier` returns false for them → `needsFallback()` returns null → + they continue to reach `DELEGATE_PIKEVM`. + +## Constraints + +- No changes to `FallbackPatternDetector` or `PatternAnalyzer`. +- Error message format in `resolveRealization()` must be consistent with the existing + `needsJdk` path (folding the PIKEVM fallback reason into the `needsJdk` boolean or + reusing the same throw is preferred). +- `FallbackPatternDetector` is already imported in both files — no new dependencies. + +## Scope + +### Primary fixes + +- `reggie-processor/src/main/java/com/datadoghq/reggie/processor/ReggieMatcherBytecodeGenerator.java:103` + — missing fallback guard on PIKEVM_CAPTURE early-exit: add `needsFallback()` check + before returning `DELEGATE_PIKEVM`. + +### Auto-expanded sibling fixes + +- `reggie-runtime/src/main/java/com/datadoghq/reggie/runtime/RuntimeCompiler.java:299` + — partial guard (B16 only) in `compilePikeVm()`: replace `hasNullableGroupContentWithNullableQuantifier` + check with full `FallbackPatternDetector.needsFallback(ast, MatchingStrategy.PIKEVM_CAPTURE)` call. + FLOW: annotation-processor-generated stub calls `compilePikeVm()` directly, bypassing + `RuntimeCompiler.compile()` which has the correct full guard. + PRECONDITION: pattern has PIKEVM_CAPTURE result from PatternAnalyzer AND triggers B2 or B3. + REACHABLE: yes — `($){2}` at an `@RegexPattern` site reaches `compilePikeVm()` with no B2/B3 check. + CONCLUSION: medium-severity defensive guard that catches patterns slipping through Fix 1 if any + future code path creates a PIKEVM stub without going through `resolveRealization()`. + +## Assumptions + +- `FallbackPatternDetector.needsFallback(ast, PIKEVM_CAPTURE)` subsumes + `hasNullableGroupContentWithNullableQuantifier` (confirmed: B16 guard at + `FallbackPatternDetector.java:287` includes `strategy == PIKEVM_CAPTURE`). +- Fix 1 folds the PIKEVM fallback check into the existing `needsJdk` boolean rather than + adding a separate throw, to keep uniform error message formatting. diff --git a/reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/analysis/FallbackPatternDetector.java b/reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/analysis/FallbackPatternDetector.java index 2d303df4..3c4b8d0e 100644 --- a/reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/analysis/FallbackPatternDetector.java +++ b/reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/analysis/FallbackPatternDetector.java @@ -1408,6 +1408,344 @@ private static boolean containsNullableCapturingGroup(RegexNode node) { return false; } + /** + * Returns true if any {@link AlternationNode} in {@code ast} contains an alternative that is + * missing a capturing group number present in another alternative of the same alternation, AND at + * least one of the alternatives in that alternation consumes more than one character (minimum + * match length > 1). Single-character-per-alt patterns (e.g. {@code (a)|b}, {@code (b)|b}) are + * correctly handled by the TDFA 1B alternation-binding fix and do not require PIKEVM routing. + * + *

Examples that fire: {@code [a][1-b]|(.)}, {@code _.|(_)}, {@code (1)c|10} — at least one alt + * is a multi-char sequence. Examples that do NOT fire: {@code (a)|b}, {@code (b)|b}, {@code + * b|(b)}, {@code .|([^c])} — all single-char alts. {@code (a|b)} does NOT fire (alternation + * inside the group; inner branches add no separate group). Patterns with no alternation (e.g. + * {@code (ab)c}) do not fire. + */ + public static boolean hasCapturingGroupAbsentFromSomeAlternative(RegexNode ast) { + if (ast instanceof AlternationNode) { + AlternationNode alt = (AlternationNode) ast; + List alts = alt.alternatives; + @SuppressWarnings("unchecked") + Set[] groups = new Set[alts.size()]; + Set union = new HashSet<>(); + for (int i = 0; i < alts.size(); i++) { + groups[i] = new HashSet<>(); + collectGroupsInSubtree(alts.get(i), groups[i]); + union.addAll(groups[i]); + } + if (!union.isEmpty()) { + boolean anyAbsent = false; + for (Set altGroups : groups) { + if (!altGroups.containsAll(union)) { + anyAbsent = true; + break; + } + } + if (anyAbsent) { + // Only flag when at least one alternative consumes more than one character. Single-char + // alternations are handled correctly by the TDFA 1B fix. + for (RegexNode alternative : alts) { + if (altMinLength(alternative) > 1) return true; + } + } + } + for (RegexNode alternative : alts) { + if (hasCapturingGroupAbsentFromSomeAlternative(alternative)) return true; + } + return false; + } + if (ast instanceof ConcatNode) { + for (RegexNode child : ((ConcatNode) ast).children) { + if (hasCapturingGroupAbsentFromSomeAlternative(child)) return true; + } + return false; + } + if (ast instanceof GroupNode) { + return hasCapturingGroupAbsentFromSomeAlternative(((GroupNode) ast).child); + } + if (ast instanceof QuantifierNode) { + return hasCapturingGroupAbsentFromSomeAlternative(((QuantifierNode) ast).child); + } + return false; + } + + /** + * Returns true if any capturing {@link GroupNode} in {@code ast} has a body whose leading element + * is nullable (i.e. can match the empty string). This identifies patterns where the TDFA fires + * the group-start tag at a state that is epsilon-reachable from before the group, causing the + * recorded start to equal the overall match start rather than the group's actual start. + * + *

Examples: {@code -{1}(a?.*)} — group body starts with {@code a?} (min=0, nullable). {@code + * (0{0}[^_]{1,})-} — group body starts with {@code 0{0}} (max=0, nullable). + */ + public static boolean hasCapturingGroupWithNullableFirstElement(RegexNode ast) { + if (ast instanceof GroupNode) { + GroupNode g = (GroupNode) ast; + if (g.capturing) { + RegexNode body = g.child; + boolean nullableFirst; + if (body instanceof ConcatNode) { + List children = ((ConcatNode) body).children; + nullableFirst = !children.isEmpty() && isNullable(children.get(0)); + } else { + nullableFirst = isNullable(body); + } + if (nullableFirst) return true; + } + return hasCapturingGroupWithNullableFirstElement(g.child); + } + if (ast instanceof AlternationNode) { + for (RegexNode alt : ((AlternationNode) ast).alternatives) { + if (hasCapturingGroupWithNullableFirstElement(alt)) return true; + } + return false; + } + if (ast instanceof ConcatNode) { + for (RegexNode child : ((ConcatNode) ast).children) { + if (hasCapturingGroupWithNullableFirstElement(child)) return true; + } + return false; + } + if (ast instanceof QuantifierNode) { + return hasCapturingGroupWithNullableFirstElement(((QuantifierNode) ast).child); + } + return false; + } + + /** + * Returns the minimum number of characters that {@code node} must consume (ignoring zero-width + * anchors and assertions). Used to distinguish single-character alternation alternatives from + * multi-character ones. + */ + private static int altMinLength(RegexNode node) { + if (node instanceof LiteralNode) { + return ((LiteralNode) node).ch == 0 ? 0 : 1; // epsilon is 0; normal literal is 1 + } + if (node instanceof CharClassNode) return 1; + if (node instanceof AnchorNode || node instanceof AssertionNode) return 0; // zero-width + if (node instanceof GroupNode) return altMinLength(((GroupNode) node).child); + if (node instanceof QuantifierNode) { + QuantifierNode q = (QuantifierNode) node; + if (q.min == 0) return 0; + return q.min * altMinLength(q.child); + } + if (node instanceof ConcatNode) { + int total = 0; + for (RegexNode c : ((ConcatNode) node).children) { + total += altMinLength(c); + } + return total; + } + if (node instanceof AlternationNode) { + int min = Integer.MAX_VALUE; + for (RegexNode a : ((AlternationNode) node).alternatives) { + min = Math.min(min, altMinLength(a)); + } + return min == Integer.MAX_VALUE ? 0 : min; + } + return 0; + } + + /** + * Returns true if any capturing {@link GroupNode} in {@code ast} has a body that consists solely + * of an anchor (e.g. {@code ($)}, {@code (^)}). The OnePass NFA emits a wrong zero-width span for + * such groups; routing to PIKEVM_CAPTURE gives correct results. + */ + public static boolean hasAnchorOnlyCapturingGroup(RegexNode ast) { + if (ast instanceof GroupNode g && g.capturing) { + if (isAnchorOnlyBody(g.child)) return true; + return hasAnchorOnlyCapturingGroup(g.child); + } + if (ast instanceof ConcatNode c) { + for (RegexNode child : c.children) if (hasAnchorOnlyCapturingGroup(child)) return true; + } + if (ast instanceof AlternationNode a) { + for (RegexNode alt : a.alternatives) if (hasAnchorOnlyCapturingGroup(alt)) return true; + } + if (ast instanceof QuantifierNode q) return hasAnchorOnlyCapturingGroup(q.child); + return false; + } + + private static boolean isAnchorOnlyBody(RegexNode node) { + if (node instanceof AnchorNode) return true; + if (node instanceof ConcatNode c) + return c.children.size() == 1 && c.children.get(0) instanceof AnchorNode; + return false; + } + + /** + * Peels any number of non-capturing {@link GroupNode} wrappers from {@code node} and returns the + * innermost node that is not a non-capturing group. Capturing groups and all other node types are + * returned as-is. + */ + private static RegexNode unwrapNonCapturing(RegexNode node) { + while (node instanceof GroupNode g && !g.capturing) { + node = g.child; + } + return node; + } + + /** + * Returns true if the top-level concat contains a capturing group whose body is a greedy + * quantifier with min≥1, infinite max, and a broad charset (.+ or .+[DOTALL]), followed by at + * least one more node (the suffix). The TDFA extends the group-end tag into the suffix for these + * patterns, producing wrong capture spans. + * + *

Non-capturing wrappers around the capturing group (e.g. {@code (?:(.+))_}) or around the + * quantifier body inside the capturing group (e.g. {@code ((?:.+))_}) are transparently unwrapped + * before the check, so both forms are correctly detected. + */ + public static boolean hasGreedyDotPlusGroupWithSuffix(RegexNode ast) { + if (!(ast instanceof ConcatNode concat)) return false; + List children = concat.children; + for (int i = 0; i < children.size() - 1; i++) { + RegexNode node = unwrapNonCapturing(children.get(i)); + if (!(node instanceof GroupNode g) || !g.capturing) continue; + RegexNode body = unwrapNonCapturing(g.child); + if (!(body instanceof QuantifierNode q)) continue; + if (q.min < 1 || (q.max != Integer.MAX_VALUE && q.max != -1)) continue; + if (!(q.child instanceof CharClassNode cc)) continue; + CharSet cs = cc.chars; + if (cs.equals(CharSet.ANY) || cs.equals(CharSet.ANY_EXCEPT_NEWLINE)) return true; + } + return false; + } + + /** + * Returns the minimum number of characters that {@code node} must consume. AnchorNodes are + * zero-width and contribute 0. Returns 0 for unknown node types (conservative). + */ + private static int minLength(RegexNode node) { + if (node instanceof LiteralNode) return 1; + if (node instanceof CharClassNode) return 1; + if (node instanceof AnchorNode) return 0; + if (node instanceof QuantifierNode q) return q.min * minLength(q.child); + if (node instanceof ConcatNode c) { + int total = 0; + for (RegexNode child : c.children) total += minLength(child); + return total; + } + if (node instanceof AlternationNode a) { + int min = Integer.MAX_VALUE; + for (RegexNode alt : a.alternatives) min = Math.min(min, minLength(alt)); + return min == Integer.MAX_VALUE ? 0 : min; + } + if (node instanceof GroupNode g) return minLength(g.child); + return 0; + } + + /** + * Returns the maximum number of characters that {@code node} can consume, or {@link + * Integer#MAX_VALUE} for unbounded. Returns {@link Integer#MAX_VALUE} for unknown node types + * (conservative). + */ + private static int maxLength(RegexNode node) { + if (node instanceof LiteralNode) return 1; + if (node instanceof CharClassNode) return 1; + if (node instanceof AnchorNode) return 0; + if (node instanceof QuantifierNode q) { + if (q.max == Integer.MAX_VALUE || q.max == -1) return Integer.MAX_VALUE; + int childMax = maxLength(q.child); + if (childMax == Integer.MAX_VALUE) return Integer.MAX_VALUE; + return q.max * childMax; + } + if (node instanceof ConcatNode c) { + int total = 0; + for (RegexNode child : c.children) { + int cm = maxLength(child); + if (cm == Integer.MAX_VALUE) return Integer.MAX_VALUE; + total += cm; + if (total < 0) return Integer.MAX_VALUE; // overflow guard + } + return total; + } + if (node instanceof AlternationNode a) { + int max = 0; + for (RegexNode alt : a.alternatives) { + int am = maxLength(alt); + if (am == Integer.MAX_VALUE) return Integer.MAX_VALUE; + if (am > max) max = am; + } + return max; + } + if (node instanceof GroupNode g) return maxLength(g.child); + return Integer.MAX_VALUE; + } + + /** + * Returns true if {@code node} or any of its descendants is a {@link CharClassNode} whose charset + * matches more than one character (e.g. {@code .}, {@code [a-z]}, or any negated class). + * Single-character classes like {@code [1]} or {@code [a]} are not considered broad. + */ + private static boolean containsBroadCharClass(RegexNode node) { + if (node instanceof CharClassNode cc) return cc.negated || !cc.chars.isSingleChar(); + if (node instanceof LiteralNode || node instanceof AnchorNode) return false; + if (node instanceof GroupNode g) return containsBroadCharClass(g.child); + if (node instanceof ConcatNode c) { + for (RegexNode child : c.children) if (containsBroadCharClass(child)) return true; + return false; + } + if (node instanceof AlternationNode a) { + for (RegexNode alt : a.alternatives) if (containsBroadCharClass(alt)) return true; + return false; + } + if (node instanceof QuantifierNode q) return containsBroadCharClass(q.child); + return false; + } + + /** + * Returns true if any capturing {@link GroupNode} in {@code ast} has a body that, after + * unwrapping transparent non-capturing wrappers, is an {@link AlternationNode} whose branches + * differ in minimum or maximum length AND at least one branch contains a broad charset (a {@link + * CharClassNode} matching more than one character). Pure-literal variable-length alternations + * like {@code (aa|a)} are handled correctly by the TDFA priority-cut and do not require PIKEVM + * routing. The broad-charset condition ensures only patterns where the TDFA cannot + * deterministically assign the group-end tag are routed to PIKEVM_CAPTURE. PikeVM gives correct + * spans for all cases. + * + *

Example: {@code ([1]|1.)} — branch {@code [1]} has length 1, branch {@code 1.} has length 2, + * and {@code 1.} contains {@code .} (broad charset) → variable-length with broad charset → route + * to PIKEVM_CAPTURE. + * + *

Non-capturing wrappers around the alternation body (e.g. {@code ((?:[1]|1.))}) are + * transparently unwrapped before the alternation check so that wrapped forms are also detected. + */ + public static boolean hasGroupWithVariableLengthAlternationBody(RegexNode ast) { + if (ast instanceof GroupNode g && g.groupNumber > 0) { + if (unwrapNonCapturing(g.child) instanceof AlternationNode a) { + List alts = a.alternatives; + if (alts.size() >= 2) { + int firstMin = minLength(alts.get(0)); + int firstMax = maxLength(alts.get(0)); + boolean hasVariableLength = false; + for (int i = 1; i < alts.size(); i++) { + int altMin = minLength(alts.get(i)); + int altMax = maxLength(alts.get(i)); + if (altMin != firstMin || altMax != firstMax) { + hasVariableLength = true; + break; + } + } + if (hasVariableLength + && alts.stream().anyMatch(FallbackPatternDetector::containsBroadCharClass)) + return true; + } + } + return hasGroupWithVariableLengthAlternationBody(g.child); + } + if (ast instanceof AlternationNode a) { + for (RegexNode alt : a.alternatives) + if (hasGroupWithVariableLengthAlternationBody(alt)) return true; + } + if (ast instanceof ConcatNode c) { + for (RegexNode child : c.children) + if (hasGroupWithVariableLengthAlternationBody(child)) return true; + } + if (ast instanceof QuantifierNode q) return hasGroupWithVariableLengthAlternationBody(q.child); + if (ast instanceof GroupNode g) return hasGroupWithVariableLengthAlternationBody(g.child); + return false; + } + /** * Returns true if any capturing GroupNode is directly wrapped by a QuantifierNode with min=0 AND * the group's content is itself nullable (can match the empty string). Example: {@code diff --git a/reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/analysis/PatternAnalyzer.java b/reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/analysis/PatternAnalyzer.java index b8c8a6e2..5f7ad578 100644 --- a/reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/analysis/PatternAnalyzer.java +++ b/reggie-codegen/src/main/java/com/datadoghq/reggie/codegen/analysis/PatternAnalyzer.java @@ -41,11 +41,22 @@ public class PatternAnalyzer { private final RegexNode ast; private final NFA nfa; + /** Accumulated guard trace entries for the most recent {@link #analyzeAndRecommend} call. */ + private final List guardTrace = new ArrayList<>(); + public PatternAnalyzer(RegexNode ast, NFA nfa) { this.ast = ast; this.nfa = nfa; } + /** + * Records one guard evaluation entry. A checkmark prefix marks guards that fired (caused a + * routing decision); a space prefix marks guards that were evaluated but did not fire. + */ + private void addTrace(String guardName, boolean fired) { + guardTrace.add((fired ? "✓ " : " ") + guardName); + } + /** * Check if a pattern requires recursive descent parsing (context-free features). This can be * called before building NFA to avoid unnecessary work. @@ -279,6 +290,13 @@ public MatchingStrategyResult analyzeAndRecommend() { * RuntimeCompiler for hybrid DFA+NFA approach) */ public MatchingStrategyResult analyzeAndRecommend(boolean ignoreGroupCount) { + guardTrace.clear(); + MatchingStrategyResult result = doAnalyze(ignoreGroupCount); + result.guardTrace.addAll(guardTrace); + return result; + } + + private MatchingStrategyResult doAnalyze(boolean ignoreGroupCount) { // Extract required literals for indexOf optimization (applicable to all strategies) java.util.Set requiredLiterals = extractRequiredLiterals(ast); @@ -312,6 +330,17 @@ public MatchingStrategyResult analyzeAndRecommend(boolean ignoreGroupCount) { if (hasBackrefs && hasQuantifiedBackrefs) { FixedRepetitionBackrefInfo fixedRepBackrefInfo = detectFixedRepetitionBackref(ast); if (fixedRepBackrefInfo != null) { + // B6: decline when a non-empty suffix exists — the bytecode generator places the + // group-end tag after the suffix is consumed, not after the initial group match. + // Fall through to OPTIMIZED_NFA_WITH_BACKREFS, which handles group spans correctly. + if (!fixedRepBackrefInfo.suffix.isEmpty()) { + return new MatchingStrategyResult( + MatchingStrategy.OPTIMIZED_NFA_WITH_BACKREFS, + null, + null, + false, + java.util.Collections.emptySet()); + } return new MatchingStrategyResult( MatchingStrategy.FIXED_REPETITION_BACKREF, null, @@ -321,11 +350,14 @@ public MatchingStrategyResult analyzeAndRecommend(boolean ignoreGroupCount) { } } - if (hasSubroutines(ast) - || hasConditionals(ast) - || hasBranchReset(ast) - || (hasNonGreedyQuantifiers(ast) && !hasBackrefs) - || hasQuantifiedBackrefs) { + boolean requiresRecursiveDescentFlag = + hasSubroutines(ast) + || hasConditionals(ast) + || hasBranchReset(ast) + || (hasNonGreedyQuantifiers(ast) && !hasBackrefs) + || hasQuantifiedBackrefs; + addTrace("requiresRecursiveDescent", requiresRecursiveDescentFlag); + if (requiresRecursiveDescentFlag) { return new MatchingStrategyResult( MatchingStrategy.RECURSIVE_DESCENT, null, // no DFA @@ -415,7 +447,9 @@ public MatchingStrategyResult analyzeAndRecommend(boolean ignoreGroupCount) { // Check for lookaround assertions // Try DFA first for simple literal assertions (e.g., (?<=ab)c, a(?=bc)) // Fall back to NFA for complex assertions (e.g., (?=.*[A-Z])) - if (hasLookaround(ast)) { + boolean hasLookaroundFlag = hasLookaround(ast); + addTrace("hasLookaround", hasLookaroundFlag); + if (hasLookaroundFlag) { // CRITICAL: Check if backrefs reference groups inside lookaheads // DFA-based strategies can't track capturing groups, so we must use NFA if (hasBackrefToLookaheadCapture(ast)) { @@ -665,7 +699,7 @@ public MatchingStrategyResult analyzeAndRecommend(boolean ignoreGroupCount) { // Try to detect fixed-repetition backreference patterns: (a)\1{8,}, (abc)\1{3} // These don't require backtracking - just verification loop FixedRepetitionBackrefInfo fixedRepBackrefInfo = detectFixedRepetitionBackref(ast); - if (fixedRepBackrefInfo != null) { + if (fixedRepBackrefInfo != null && fixedRepBackrefInfo.suffix.isEmpty()) { return new MatchingStrategyResult( MatchingStrategy.FIXED_REPETITION_BACKREF, null, @@ -673,6 +707,7 @@ public MatchingStrategyResult analyzeAndRecommend(boolean ignoreGroupCount) { false, requiredLiterals); } + // B6: if suffix is non-empty, fall through to OPTIMIZED_NFA_WITH_BACKREFS below. // Try to detect variable-capture backreference patterns: (.*)\d+\1, (.+)=\1 // These require backtracking from longest to shortest capture @@ -717,7 +752,20 @@ public MatchingStrategyResult analyzeAndRecommend(boolean ignoreGroupCount) { } // Check for OnePass eligibility (highest priority for patterns with groups) - if (!ignoreGroupCount && nfa.getGroupCount() > 0 && isOnePassEligible()) { + // B3a: skip OnePass for anchor-only capturing group bodies — the OnePass NFA emits a wrong + // zero-width span for such groups; they are handled below via PIKEVM_CAPTURE. + boolean hasAnchorOnlyCapturingGroupFlag = + nfa != null + && nfa.getGroupCount() > 0 + && FallbackPatternDetector.hasAnchorOnlyCapturingGroup(ast); + addTrace("B3a: hasAnchorOnlyCapturingGroup", hasAnchorOnlyCapturingGroupFlag); + boolean isOnePassEligibleFlag = + !ignoreGroupCount && nfa.getGroupCount() > 0 && isOnePassEligible(); + addTrace("isOnePassEligible", isOnePassEligibleFlag && !hasAnchorOnlyCapturingGroupFlag); + if (!ignoreGroupCount + && nfa.getGroupCount() > 0 + && isOnePassEligibleFlag + && !hasAnchorOnlyCapturingGroupFlag) { return new MatchingStrategyResult( MatchingStrategy.ONEPASS_NFA, null, null, false, requiredLiterals); } @@ -728,6 +776,18 @@ public MatchingStrategyResult analyzeAndRecommend(boolean ignoreGroupCount) { // Check if pattern has groups inside repeating quantifiers (needs POSIX semantics) boolean needsPosixSemantics = hasGroupsInRepeatingQuantifiers(ast); + // B3a: anchor-only group body — OnePass NFA emits wrong zero-width span. + if (hasAnchorOnlyCapturingGroupFlag) { + return new MatchingStrategyResult( + MatchingStrategy.PIKEVM_CAPTURE, + null, + null, + false, + requiredLiterals, + null, + needsPosixSemantics); + } + // Try specialized strategies for patterns with quantified groups // These provide correct POSIX last-match semantics for group capture if (needsPosixSemantics) { @@ -765,6 +825,25 @@ public MatchingStrategyResult analyzeAndRecommend(boolean ignoreGroupCount) { // POSIX last-match semantics (groups may contain first match instead of last) } + // B4: greedy .+ group followed by a non-group suffix — the GREEDY_BACKTRACK indexOf scan + // overshoots on inputs ending with '\n' because '.' (ANY_EXCEPT_NEWLINE) cannot consume '\n' + // but the literal-suffix scan stops there. Route to PIKEVM_CAPTURE which handles this + // correctly. Patterns with a capturing-group suffix are excluded (different scan path). + boolean b4Flag = + FallbackPatternDetector.hasGreedyDotPlusGroupWithSuffix(ast) + && !FallbackPatternDetector.hasNullableGroupContentWithNullableQuantifier(ast); + addTrace("B4: hasGreedyDotPlusGroupWithSuffix", b4Flag); + if (b4Flag) { + return new MatchingStrategyResult( + MatchingStrategy.PIKEVM_CAPTURE, + null, + null, + false, + requiredLiterals, + null, + needsPosixSemantics); + } + // Check if pattern requires backtracking for correct group capture // Pattern a([bc]*)(c+d) needs backtracking: ([bc]*) must give back chars to allow (c+d) to // match @@ -797,13 +876,15 @@ public MatchingStrategyResult analyzeAndRecommend(boolean ignoreGroupCount) { null, needsPosixSemantics); } - if (hasStringEndAnchorInAlternation(ast) && !dfaHasAcceptingStateWithTransitions(dfa)) { - // \Z or $ in alternation with capturing groups: OPTIMIZED_NFA handles anchors as - // zero-width NFA assertions. The nfa.getGroupCount() == 0 branch that previously - // appeared here was unreachable (this block is guarded by nfa.getGroupCount() > 0). - // Zero-group patterns with \Z in alternation are handled outside this block. + boolean b3bCapFlag = + (hasStringEndAnchorInAlternation(ast) || hasEndAnchorLeadingInAlternationBranch(ast)) + && !dfaHasAcceptingStateWithTransitions(dfa); + addTrace("B3b: hasStringEndAnchorInAlternation", b3bCapFlag); + if (b3bCapFlag) { + // \Z in alternation with capturing groups: PIKEVM_CAPTURE handles anchors correctly. + // OPTIMIZED_NFA would be rejected by needsFallback for this combination. return new MatchingStrategyResult( - MatchingStrategy.OPTIMIZED_NFA, + MatchingStrategy.PIKEVM_CAPTURE, null, null, false, @@ -817,6 +898,7 @@ public MatchingStrategyResult analyzeAndRecommend(boolean ignoreGroupCount) { // the DFA to lose the anchor guard. PikeVMMatcher.checkAnchor evaluates all anchor types // correctly against the actual search position, so PIKEVM is safe for all diluted shapes — // not just alternation patterns. The alternation+accepting-transitions guard is removed. + addTrace("isAnchorConditionDiluted", dfa.isAnchorConditionDiluted()); if (dfa.isAnchorConditionDiluted()) { // Anchor condition diluted in DFA: capture-ambiguous patterns are safe for PikeVM // because PikeVM evaluates anchors natively per position (via checkAnchor) and tracks @@ -909,6 +991,7 @@ && containsAnyQuantifier(ast) return r; } + addTrace("isCaptureAmbiguous → DFA_UNROLLED_WITH_GROUPS", dfa.isCaptureAmbiguous()); if (dfa.isCaptureAmbiguous()) { // For pure-regular, anchor-free patterns the C2 priority-ordered TDFA gives correct // spans and can use an inline DFA strategy when the state count is small enough. @@ -977,8 +1060,75 @@ && containsAnyQuantifier(ast) null, needsPosixSemantics); } + // A1: group body starts with a nullable first element — TDFA fires the group-start + // tag at an epsilon-reachable state, recording match-start instead of group-start. + boolean a1Flag = + FallbackPatternDetector.hasCapturingGroupWithNullableFirstElement(ast) + && !FallbackPatternDetector.hasNullableGroupContentWithNullableQuantifier(ast); + addTrace("A1: hasCapturingGroupWithNullableFirstElement", a1Flag); + if (a1Flag) { + return new MatchingStrategyResult( + MatchingStrategy.PIKEVM_CAPTURE, + null, + null, + false, + requiredLiterals, + null, + needsPosixSemantics); + } + // A2: capturing group absent from some alternation branch — TDFA binds the absent + // group to a wrong span when the branch that lacks the group wins. + boolean a2Flag = + FallbackPatternDetector.hasCapturingGroupAbsentFromSomeAlternative(ast) + && !FallbackPatternDetector.hasNullableGroupContentWithNullableQuantifier(ast); + addTrace("A2: hasCapturingGroupAbsentFromSomeAlternative", a2Flag); + if (a2Flag) { + return new MatchingStrategyResult( + MatchingStrategy.PIKEVM_CAPTURE, + null, + null, + false, + requiredLiterals, + null, + needsPosixSemantics); + } + // B4: greedy .+ group followed by a suffix — TDFA extends group-end into the suffix. + boolean b4CaFlag = + FallbackPatternDetector.hasGreedyDotPlusGroupWithSuffix(ast) + && !FallbackPatternDetector.hasNullableGroupContentWithNullableQuantifier(ast); + addTrace("B4: hasGreedyDotPlusGroupWithSuffix (captureAmbiguous)", b4CaFlag); + if (b4CaFlag) { + return new MatchingStrategyResult( + MatchingStrategy.PIKEVM_CAPTURE, + null, + null, + false, + requiredLiterals, + null, + needsPosixSemantics); + } + // B5: group body is an alternation whose branches have different min- or max-lengths. + boolean b5CaFlag = + FallbackPatternDetector.hasGroupWithVariableLengthAlternationBody(ast) + && !FallbackPatternDetector.hasNullableGroupContentWithNullableQuantifier(ast); + addTrace("B5: hasGroupWithVariableLengthAlternationBody (captureAmbiguous)", b5CaFlag); + if (b5CaFlag) { + return new MatchingStrategyResult( + MatchingStrategy.PIKEVM_CAPTURE, + null, + null, + false, + requiredLiterals, + null, + needsPosixSemantics); + } // Pure-regular, anchor-free: C2 priority-ordered TDFA gives correct spans. int stateCount = dfa.getStateCount(); + addTrace( + "DFA state count → DFA_UNROLLED / DFA_SWITCH / OPTIMIZED_NFA (stateCount=" + + stateCount + + ")", + true); if (stateCount < DFA_UNROLLED_STATE_LIMIT) { return new MatchingStrategyResult( MatchingStrategy.DFA_UNROLLED_WITH_GROUPS, @@ -1046,6 +1196,46 @@ && containsAnyQuantifier(ast) null, needsPosixSemantics); } + // A1: group body starts with a nullable first element — TDFA fires the group-start + // tag at an epsilon-reachable state, recording match-start instead of group-start. + if (FallbackPatternDetector.hasCapturingGroupWithNullableFirstElement(ast) + && !FallbackPatternDetector.hasNullableGroupContentWithNullableQuantifier(ast)) { + return new MatchingStrategyResult( + MatchingStrategy.PIKEVM_CAPTURE, + null, + null, + false, + requiredLiterals, + null, + needsPosixSemantics); + } + // A2: capturing group absent from some alternation branch — TDFA binds the absent + // group to a wrong span when the branch that lacks the group wins. + if (FallbackPatternDetector.hasCapturingGroupAbsentFromSomeAlternative(ast) + && !FallbackPatternDetector.hasNullableGroupContentWithNullableQuantifier(ast)) { + return new MatchingStrategyResult( + MatchingStrategy.PIKEVM_CAPTURE, + null, + null, + false, + requiredLiterals, + null, + needsPosixSemantics); + } + // B5: group body is an alternation whose branches have different min- or max-lengths + // and at least one branch contains a broad charset — TDFA cannot deterministically assign + // the group-end tag when a broad-charset alternative competes with the suffix. + if (FallbackPatternDetector.hasGroupWithVariableLengthAlternationBody(ast) + && !FallbackPatternDetector.hasNullableGroupContentWithNullableQuantifier(ast)) { + return new MatchingStrategyResult( + MatchingStrategy.PIKEVM_CAPTURE, + null, + null, + false, + requiredLiterals, + null, + needsPosixSemantics); + } // Class E: two interacting variable-length capturing alternations (e.g. (a|ab)(c|bcd)). The // first alternation's branches share a prefix, so its capture span is ambiguous until the // second alternation resolves it — which the single-register TDFA cannot track @@ -1129,7 +1319,11 @@ && containsAnyQuantifier(ast) return new MatchingStrategyResult( MatchingStrategy.OPTIMIZED_NFA, null, null, false, requiredLiterals); } - if (hasStringEndAnchorInAlternation(ast) && !dfaHasAcceptingStateWithTransitions(dfa)) { + boolean b3bFlag = + (hasStringEndAnchorInAlternation(ast) || hasEndAnchorLeadingInAlternationBranch(ast)) + && !dfaHasAcceptingStateWithTransitions(dfa); + addTrace("B3b: hasStringEndAnchorInAlternation", b3bFlag); + if (b3bFlag) { // \Z or $ in alternation: OPTIMIZED_NFA mishandles find() anchor semantics; // route to PIKEVM_CAPTURE which handles \Z/$ correctly. return new MatchingStrategyResult( @@ -1146,7 +1340,11 @@ && containsAnyQuantifier(ast) // This block runs BEFORE the isAnchorConditionDiluted guard below: a diluted-anchor // pattern (e.g. ^c|[^1][b]) is handled correctly by PIKEVM, whereas OPTIMIZED_NFA // (the dilution fallback target) shares the old find() anchor bug. - if (containsAlternation(ast) && dfaHasAcceptingStateWithTransitions(dfa)) { + boolean altWithAcceptingTransFlag = + containsAlternation(ast) && dfaHasAcceptingStateWithTransitions(dfa); + addTrace( + "containsAlternation && dfaHasAcceptingStateWithTransitions", altWithAcceptingTransFlag); + if (altWithAcceptingTransFlag) { return new MatchingStrategyResult( MatchingStrategy.PIKEVM_CAPTURE, null, null, false, requiredLiterals); } @@ -1154,6 +1352,7 @@ && containsAnyQuantifier(ast) // correctly at each search position, whereas OPTIMIZED_NFA mishandles diluted conditions. // anchorConditionDiluted=true on the result signals RuntimeCompiler's hybrid pre-check to // skip the hybrid DFA path (a diluted DFA is not safe for the fast-matching pass). + addTrace("isAnchorConditionDiluted (no-group path)", dfa.isAnchorConditionDiluted()); if (dfa.isAnchorConditionDiluted()) { MatchingStrategyResult r = new MatchingStrategyResult( @@ -1169,6 +1368,11 @@ && containsAnyQuantifier(ast) } // Choose DFA strategy based on state count int stateCount = dfa.getStateCount(); + addTrace( + "DFA state count → DFA_UNROLLED / DFA_SWITCH / OPTIMIZED_NFA (stateCount=" + + stateCount + + ")", + true); if (stateCount < DFA_UNROLLED_STATE_LIMIT) { return new MatchingStrategyResult( MatchingStrategy.DFA_UNROLLED, dfa, null, false, requiredLiterals); @@ -2863,6 +3067,9 @@ public static class MatchingStrategyResult { */ public boolean anchorConditionDiluted; + /** Per-guard routing trace populated by {@link PatternAnalyzer#analyzeAndRecommend}. */ + public final List guardTrace = new ArrayList<>(); + public MatchingStrategyResult(MatchingStrategy strategy, DFA dfa) { this(strategy, dfa, null, false, java.util.Collections.emptySet(), null, false); } @@ -7051,6 +7258,22 @@ private GreedyBacktrackInfo detectGreedyBacktrackPattern(RegexNode ast) { return null; // No suffix } + // Decline .+ (min>=1) with broad charset when the suffix is a literal (not a capturing group). + // The GREEDY_BACKTRACK indexOf scan cannot enforce the min=1 lower bound while surrendering + // chars to a non-group suffix: on inputs with a trailing newline the scan overshoots, matching + // the newline via the suffix delimiter check instead of stopping at the non-newline boundary of + // '.' (ANY_EXCEPT_NEWLINE). Capturing-group suffixes use a different scan path unaffected by + // this. + boolean suffixStartsWithCapturingGroup = + allNodes.get(greedyGroupIndex + 1) instanceof GroupNode sg && sg.capturing; + if (!suffixStartsWithCapturingGroup + && greedyMinCount > 0 + && greedyCharSet != null + && (greedyCharSet.equals(CharSet.ANY) + || greedyCharSet.equals(CharSet.ANY_EXCEPT_NEWLINE))) { + return null; + } + // Extract prefix (everything before the greedy group) List prefix = new ArrayList<>(); for (int i = 0; i < greedyGroupIndex; i++) { diff --git a/reggie-integration-tests/src/main/java/com/datadoghq/reggie/integration/fuzz/FuzzRunner.java b/reggie-integration-tests/src/main/java/com/datadoghq/reggie/integration/fuzz/FuzzRunner.java index d891a29f..7ee5b2a6 100644 --- a/reggie-integration-tests/src/main/java/com/datadoghq/reggie/integration/fuzz/FuzzRunner.java +++ b/reggie-integration-tests/src/main/java/com/datadoghq/reggie/integration/fuzz/FuzzRunner.java @@ -56,6 +56,15 @@ public static final class Config { public int patternDepth = 3; public int inputMaxLength = 12; + /** + * Number of (pattern, input) batches to skip at the start of the sequence. Both the pattern RNG + * and input RNG are advanced by {@code patternSkip * inputsPerPattern} steps so the remaining + * run covers fresh territory not exercised by a sweep of the same seed with a lower pattern + * count. Reproducible: given identical (seed, patternSkip, patternCount) the findings are + * always the same. + */ + public int patternSkip = 0; + /** Cap the number of findings retained per pattern to avoid quadratic-style log explosions. */ public int findingsPerPatternCap = 3; } @@ -68,6 +77,17 @@ public Report run(Config cfg) { RandomInputGenerator inputGen = new RandomInputGenerator(inputRng, cfg.inputMaxLength); RegexFuzzOracle oracle = new RegexFuzzOracle(); + // Advance both RNGs past the skip window so the active range starts at a fresh position. + // inputsPerPattern steps per skipped pattern is conservative (ignores compile-time rejects + // that would consume fewer inputs in a real run) but keeps the skip deterministic without + // running the oracle. + for (int p = 0; p < cfg.patternSkip; p++) { + regexGen.generate(); + for (int i = 0; i < cfg.inputsPerPattern; i++) { + inputGen.generate(); + } + } + int skipped = 0; int inputs = 0; List findings = new ArrayList<>(); diff --git a/reggie-integration-tests/src/test/java/com/datadoghq/reggie/integration/AlgorithmicFuzzTest.java b/reggie-integration-tests/src/test/java/com/datadoghq/reggie/integration/AlgorithmicFuzzTest.java index b023428a..affcd92b 100644 --- a/reggie-integration-tests/src/test/java/com/datadoghq/reggie/integration/AlgorithmicFuzzTest.java +++ b/reggie-integration-tests/src/test/java/com/datadoghq/reggie/integration/AlgorithmicFuzzTest.java @@ -47,21 +47,19 @@ public class AlgorithmicFuzzTest { private static final long BASE_SEED = 0xC0DEFEED_DEADBEEFL; /** - * Known pre-existing divergence budget for {@link #BASE_SEED} at the default sweep dimensions - * (25k patterns × 16 inputs × max-length 16). Every finding here is a known, tracked bug in a - * native strategy — not a regression. When this count changes, update the budget and document the - * new/fixed finding in {@code doc/temp/prod-readiness/fuzz-inventory.md}. Override via {@code - * -Dreggie.fuzz.maxFindings=N} for stricter local runs. + * Known pre-existing divergence budget for the active fuzz window: {@link #BASE_SEED}, patterns + * 25 001–50 000 (skip=25 000, count=25 000, depth=3, 16 inputs × max-length 16). The window + * starts after the range already cleared to zero by the B3a/B3b/B4/B5/B6 fixes. Every finding is + * a tracked bug — not a regression. Clustered inventory: {@code doc/fuzz/2026-06-29.md}. * - *

Raised 18→78 when {@link RegexFuzzOracle} gained a {@code findAll()} differential that - * checks per-match group spans (≥1) on the FIND path — the first oracle to do so. It surfaced - * pre-existing find-path group-capture bugs in the codegen TDFA / PikeVM (untaken-branch group - * not reset to −1; empty-iteration binding; greedy give-back inner-span). These are tracked as - * the capture-correctness effort and ratchet this budget back toward 0 as each root-cause class - * is fixed. Ratcheted 78→69→65: Class A (nullable capturing group in an alternation branch, e.g. - * {@code 1|()b}) now routes to PIKEVM_CAPTURE for correct spans. + *

When this reaches 0: advance the window — run {@code -Dreggie.fuzz.skip=50000 + * -Dreggie.fuzz.size=25000 -Dreggie.fuzz.maxFindings=9999}, document new findings in {@code + * doc/fuzz/YYYY-MM-DD.md}, then update skip and this budget to match. + * + *

History: 18→78 (findAll group-span oracle) → 69→65→13→0 (B3a/B3b/B4/B5/B6, window 0–25k) → + * window advanced to 25k–50k → 34 (E1–E6 found; see {@code doc/fuzz/2026-06-29.md}). */ - private static final int KNOWN_FINDINGS_BUDGET = 65; + private static final int KNOWN_FINDINGS_BUDGET = 34; @Test @Timeout(value = 300, unit = TimeUnit.SECONDS) @@ -221,6 +219,8 @@ private void runDivergenceGate(FuzzRunner.Config cfg, String tag, int maxFinding * *