Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,11 @@ All notable changes to this repository are documented in this file.

### Changed

- Updated review guidance to use the public `--committed` and `--uncommitted`
selectors, allow `--dir` paths inside a Git working tree, and rely on the
review command's built-in authentication flow. Documented untracked-file scope,
NDJSON completion and severities, saved prompts, and current configuration and
account-command contracts.
- Aligned the shared code-review subagent metadata with Gemini CLI's schema.
- Removed alternate detailed-output guidance so review agents use `--agent`
exclusively.
Expand Down
20 changes: 8 additions & 12 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,20 +13,15 @@ CodeRabbit detects bugs, security issues, and quality risks before you merge.
## Quickstart

Install the CodeRabbit CLI via the [CLI docs](https://docs.coderabbit.ai/cli),
then authenticate:

```bash
coderabbit auth login
```

Then tell your agent: **“Review my code.”**
then tell your agent: **“Review my code.”** The review command starts the
authentication flow when needed.

## Installation

### 1. Install the CodeRabbit CLI

Use the [CLI docs](https://docs.coderabbit.ai/cli) for the primary install path.
They cover Homebrew, the install script, authentication, and CLI usage.
They cover Homebrew, the install script, and CLI usage.

### 2. Install the agent integration

Expand Down Expand Up @@ -142,13 +137,14 @@ Review the directory at ../my-service

The agent will automatically:

1. Check if CodeRabbit CLI is installed and authenticated
1. Check if CodeRabbit CLI is installed
2. Run the review on your changes
3. Present findings grouped by severity
4. Optionally fix issues and re-review

When you ask for a specific review directory, the agent can pass CodeRabbit CLI
`--dir <path>` after confirming that path is an initialized Git repository.
`--dir <path>` after confirming that path is inside an initialized Git working
tree.

## Supported Agents

Expand Down Expand Up @@ -212,9 +208,9 @@ AI-powered code review that finds bugs, security issues, and suggests improvemen
**Capabilities:**

- Analyzes code changes for bugs, security issues, and anti-patterns
- Groups findings by severity (critical, warning, info)
- Preserves finding severities (critical, major, minor, trivial, info, none)
- Supports autonomous fix-review cycles
- Works with staged, committed, or all changes
- Reviews tracked changes by default, with committed and uncommitted scopes
- Supports directory-scoped reviews through CodeRabbit CLI `--dir <path>`

### [autofix](skills/autofix/SKILL.md)
Expand Down
36 changes: 6 additions & 30 deletions agents/code-reviewer.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,14 +39,16 @@ Prefer a package manager or a verified binary over piping a remote script to a s

1. **Gather Context**
- Identify changed files and their scope
- Identify any requested review directory and confirm it contains an initialized Git repository
- Identify any requested review directory and confirm it is inside an initialized Git working tree
- Understand the type of changes (feature, bugfix, refactor)
- Check for related configuration files

2. **Run CodeRabbit Review**
- Execute `coderabbit review --agent` to get structured review output
- Add `--dir <path>` when the user requests a specific review directory
- Parse and categorize findings by severity and type
- Review starts browser authentication if needed; honor no-login restrictions and use the host authentication path when a sandbox hides credentials
- Raw untracked files require `--include-untracked`, which conflicts with `--committed`; staged new files are included by default
- Parse NDJSON and preserve critical, major, minor, trivial, info, or none severity

3. **Analyze Findings**
- Prioritize critical security issues
Expand All @@ -63,32 +65,6 @@ Prefer a package manager or a verified binary over piping a remote script to a s
- Explain complex issues in detail
- Help implement suggested changes

## Review Categories
## Completion and scope

### Critical (Must Fix)

- Security vulnerabilities
- Data exposure risks
- Authentication/authorization flaws
- Injection vulnerabilities

### High Priority

- Bug-prone code patterns
- Missing error handling
- Resource leaks
- Race conditions

### Medium Priority

- Code duplication
- Complex/hard-to-maintain code
- Missing tests
- Documentation gaps

### Low Priority (Suggestions)

- Style improvements
- Minor optimizations
- Naming conventions
- Code organization
Wait for a successful completion; heartbeats only indicate liveness. `status: review_skipped` with zero findings means no review ran, not that code is clean. Preserve requested scope on retries and report incomplete reviews. Prioritize the returned severity rather than inventing a separate category system.
39 changes: 17 additions & 22 deletions commands/coderabbit-review.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,7 @@ Review code based on: **$ARGUMENTS**
Otherwise, run:

```bash
coderabbit --version 2>/dev/null && coderabbit auth status 2>&1 | head -3
coderabbit --version 2>/dev/null
```

**If CLI not found**, tell user:
Expand All @@ -36,46 +36,41 @@ coderabbit --version 2>/dev/null && coderabbit auth status 2>&1 | head -3
>
> Prefer a package manager or a verified binary, then restart your shell and try again.

**If "Not logged in"**, tell user:
> You need to authenticate. Run in your terminal:
>
> ```bash
> coderabbit auth login
> ```
>
> Then try again.

### Run Review

Once prerequisites are met:
Once prerequisites are met, run review directly; it starts browser authentication when needed. Honor no-login restrictions and use the host flow if a sandbox hides credentials; never read credential files or request pasted tokens.

```bash
# type defaults to "all"; add --base and --dir only when specified
args=(review --agent -t "${type:-all}")
# type defaults to "all"; use a public scope option only when requested
args=(review --agent)
case "${type:-all}" in
committed) args+=(--committed) ;;
uncommitted) args+=(--uncommitted) ;;
all) ;;
*) printf 'Unsupported review type: %s\n' "$type" >&2; exit 2 ;;
esac
[ -n "${base:-}" ] && args+=(--base "$base")
[ -n "${dir:-}" ] && args+=(--dir "$dir")
coderabbit "${args[@]}"
```

Where `type`, `base`, and `dir` come from `$ARGUMENTS`:

- `all` (default) - All changes
- `all` (default) - All tracked changes
- `committed` - Committed changes only
- `uncommitted` - Uncommitted only
- `uncommitted` - Staged changes and unstaged edits to tracked files

Add `--base <branch>` only when a base branch is specified.
Add `--dir <path>` only when a review directory is specified. The directory must contain an initialized Git repository; verify it first:
Raw untracked files are excluded by default; staged new files are included. Add `--include-untracked` only when requested; it conflicts with `--committed` but can combine with `--uncommitted`. Never combine committed and uncommitted selectors or silently shrink the requested scope.

Append any requested `--include-untracked`, `--light`, or `--base-commit <commit>` option to the argument array; do not discard these when translating `$ARGUMENTS`. Add `--base <branch>` only when a base branch is specified.
Add `--dir <path>` only when a review directory is specified. The directory must be inside an initialized Git working tree; verify it first:

```bash
git -C "$dir" rev-parse --is-inside-work-tree
```

### Present Results

Group findings by severity:

1. **Critical** - Security vulnerabilities, data loss risks, crashes
2. **Warning** - Bugs, performance issues, anti-patterns
3. **Info** - Style issues, suggestions, minor improvements
Parse `--agent` as NDJSON and preserve the returned severity (`critical`, `major`, `minor`, `trivial`, `info`, or `none`). Heartbeats are liveness only. A `complete` event with `status: review_skipped` is not a clean review.

Offer to apply fixes from the `--agent` findings when the output includes actionable remediation details.
1 change: 1 addition & 0 deletions evals/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
results/
21 changes: 21 additions & 0 deletions evals/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# CLI behavior evaluations

Six offline cases exercise public scope flags, default and uncommitted untracked-file inclusion, local
versus PR prompt retrieval, EU browser authentication, and incomplete/skipped
review output. No shell, writes, network tools, or production reviews are granted.

With Claude Code 2.1.269+ and an authenticated account, run from the plugin root:

```sh
claude plugin eval . --tag cli-parity --runs 1 --ablation with-without --no-publish --max-cost-usd 10 --keep-temp
```

Pin `--model` for comparisons. Positive skill activation is diagnostic and does
not contribute to the outcome score. Deterministic graders check specific command
contracts; advisory LLM graders are with-only and excluded from the ablation
score. Inspect the actual answers and retained transcripts: regex checks and LLM
judges do not establish complete semantic correctness. One run per arm is a smoke
evaluation, not a reliable effect-size estimate. Results stay under ignored
`evals/results/`; do not commit account metadata or private source provenance.

See the [official evaluator documentation](https://code.claude.com/docs/en/plugin-evals).
41 changes: 41 additions & 0 deletions evals/review-default-untracked/case.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
{
"schema_version": "1.1",
"name": "review-default-untracked",
"description": "Offline default-scope contract; no shell, network, or file mutations.",
"tags": [
"cli-parity",
"code-review"
],
"execution": {
"prompt": "Give a CodeRabbit CLI command for the current repo that preserves its default review scope and also includes brand-new untracked files. Use output suitable for a coding agent. Assume a current official CLI and valid login. Do not execute anything; give the command and a short scope explanation.",
"max_turns": 15,
"timeout_seconds": 240,
"allowed_tools": [
"Read",
"Glob",
"Grep",
"Skill"
]
},
"graders": [
{
"name": "skill-activation",
"type": "tool_used",
"tool": "Skill",
"input_match": "\"skill\"\\s*:\\s*\"(?:[\\w-]+:)?code-review\""
},
{
"name": "default-plus-untracked-command",
"type": "regex",
"pattern": "^(?=[^\\n]*\\b(?:coderabbit|cr)\\s+review\\b)(?=[^\\n]*--agent\\b)(?=[^\\n]*--include-untracked\\b)(?![^\\n]*--(?:uncommitted|committed|type)\\b)(?![^\\n]*\\s-t\\s)[^\\n]+$",
"flags": "mi",
"match": "contains"
},
{
"name": "no-untracked-selector-requirement",
"type": "llm",
"arm": "with-only",
"criteria": "PASS if the proposed command uses --agent and --include-untracked without selecting committed-only or uncommitted-only, and the explanation preserves committed plus staged/tracked-unstaged changes while adding untracked files. FAIL if it says --include-untracked requires --uncommitted or silently narrows the default scope."
}
]
}
42 changes: 42 additions & 0 deletions evals/review-eu-auth/case.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
{
"schema_version": "1.1",
"name": "review-eu-auth",
"description": "Offline CLI contract evaluation; no shell, network, or file mutations.",
"tags": [
"cli-parity",
"code-review"
],
"execution": {
"prompt": "Prepare a CodeRabbit CLI runbook without executing it. We need a new EU SaaS browser login from an agent with a working local browser callback, then review committed changes. We do not use API keys or self-hosting. Give the login command and review command, keeping both suitable for the agent.",
"max_turns": 15,
"timeout_seconds": 240,
"allowed_tools": [
"Read",
"Glob",
"Grep",
"Skill"
]
},
"graders": [
{
"name": "skill-activation",
"type": "tool_used",
"tool": "Skill",
"input_match": "\"skill\"\\s*:\\s*\"(?:[\\w-]+:)?code-review\""
},
{
"name": "eu-agent-login",
"type": "regex",
"pattern": "^(?=[^\\n]*\\b(?:coderabbit|cr)\\s+auth\\s+login\\b)(?=[^\\n]*--agent\\b)(?=[^\\n]*--region\\s+eu\\b)[^\\n]+$",
"match": "contains",
"flags": "mi"
},
{
"name": "review-saved-region",
"type": "regex",
"pattern": "^(?=[^\\n]*\\b(?:coderabbit|cr)\\s+review\\b)(?=[^\\n]*--committed\\b)(?=[^\\n]*--agent\\b)(?![^\\n]*--region\\b)[^\\n]+$",
"match": "contains",
"flags": "mi"
}
]
}
49 changes: 49 additions & 0 deletions evals/review-saved-prompts/case.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,49 @@
{
"schema_version": "1.1",
"name": "review-saved-prompts",
"description": "Offline CLI contract evaluation; no shell, network, or file mutations.",
"tags": [
"cli-parity",
"code-review"
],
"execution": {
"prompt": "Using CodeRabbit CLI, I need two runbook commands, not a new review: (1) get saved local fix prompts for packages/api, (2) get the consolidated prompt for https://github.com/example/widgets/pull/123 as machine-readable output while outside any checkout. Assume a current official CLI and valid existing CodeRabbit login. Do not execute anything; give one command per operation and the output format.",
"max_turns": 15,
"timeout_seconds": 240,
"allowed_tools": [
"Read",
"Glob",
"Grep",
"Skill"
]
},
"graders": [
{
"name": "skill-activation",
"type": "tool_used",
"tool": "Skill",
"input_match": "\"skill\"\\s*:\\s*\"(?:[\\w-]+:)?code-review\""
},
{
"name": "local-saved-command",
"type": "regex",
"pattern": "^(?=[^\\n]*\\b(?:coderabbit|cr)\\s+review\\s)(?=[^\\n]*--show-prompts\\b)(?=[^\\n]*--dir\\s+[\"\\x27]?packages/api\\b)(?![^\\n]*--agent\\b)[^\\n]+$",
"match": "contains",
"flags": "mi"
},
{
"name": "pr-prompt-command",
"type": "regex",
"pattern": "^(?=[^\\n]*\\b(?:coderabbit|cr)\\s+pullrequest\\s+[\"\\x27]?https://github.com/example/widgets/pull/123)(?=[^\\n]*--show-prompts\\b)(?=[^\\n]*--agent\\b)[^\\n]+$",
Comment thread
coderabbitai[bot] marked this conversation as resolved.
"match": "contains",
"flags": "mi"
},
{
"name": "pr-prompt-output-format",
"type": "regex",
"pattern": "(?=[\\s\\S]*\\b(?:NDJSON|newline[- ]delimited JSON)\\b)(?=[\\s\\S]*\\btype[\"\\x27`]*\\s*[:=]\\s*[\"\\x27`]*prompt\\b)",
"flags": "i",
"match": "contains"
}
]
}
Loading
Loading