Repository navigation
feat: add BITSTATE_BYTECODE strategy for prefix-guarded scan patterns - #99
Conversation
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #99 +/- ##
=========================================
+ Coverage 83.1% 83.4% +0.2%
Complexity 1 1
=========================================
Files 134 135 +1
Lines 41223 42147 +924
Branches 5548 5605 +57
=========================================
+ Hits 34288 35166 +878
- Misses 5217 5227 +10
- Partials 1718 1754 +36
Continue to review full report in Codecov by Harness.
🚀 New features to boost your workflow:
|
Detects the ^(?:leadingWs(kw1|kw2|...)separatorWs)?mandatoryCharSet+(tail) AST shape and compiles it to straight-line bytecode instead of routing through the general BitState interpreter. Review fixes: decline when leadingWs/separator charsets overlap the mandatory charset (no give-back possible), require unbounded+greedy quantifiers throughout, correct extractDisjointKeywords javadoc, and guard against an empty keyword list.
515c482 to
7db5913
Compare
…d branches Adds a compile-time ReggieMatcherBytecodeGenerator test for the case BITSTATE_BYTECODE dispatch arm (previously 0% patch coverage — only the runtime dispatch path was tested), plus PatternAnalyzer detector negative tests for each quantifier's greedy/max guard and the leadingWs-overlap branch, per the codecov coverage regression on PR #99.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0c69a83e02
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
skipSeparator only stores separatorMin and emits an unbounded greedy
scan, so a bounded separator quantifier (e.g. \s{1,3}) let the
generated matcher over-consume past JDK's upper bound, diverging on
inputs where the separator run exceeds the bound.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
What does this PR do?
Adds a new
BITSTATE_BYTECODEmatching strategy:PatternAnalyzer.detectPrefixGuardedScanrecognizes the
^(?:leadingWs(kw1|kw2|...)separatorWs)?mandatoryCharSet+trailingWs*(tail)AST shape and
BitStateBytecodeGeneratorcompiles it to straight-line, non-backtracking JVMbytecode instead of routing through the general
BitStateinterpreter. Wired into both theruntime (
RuntimeCompiler) and compile-time (ReggieMatcherBytecodeGenerator) dispatch paths.Motivation
This pattern shape (an optional prefix keyword + separator, then a mandatory greedy scan and a
.*tail) is common enough in practice to warrant a dedicated fast path over the generalBitState interpreter.
Related Issue(s)
Change Type
Checklist
./gradlew build)Performance Impact
Not benchmarked in this PR.
Additional Notes
Detector-side hardening applied during review: decline when the leading-whitespace or separator
charset overlaps the mandatory charset (no give-back is possible without discarding the whole
prefix), require unbounded max and greedy quantifiers throughout the recognized shape, corrected
extractDisjointKeywords' javadoc to match the generator's actual sequential-comparisonimplementation, and guard against an empty keyword list.