Skip to content

Load the CodeRabbit review skill whenever code needs checking - #50

Draft
nehal-a2z wants to merge 2 commits into
nehal/lightsage-skills-evalfrom
nehal/skills-sticky-review
Draft

nehal-a2z wants to merge 2 commits into
nehal/lightsage-skills-evalfrom
nehal/skills-sticky-review

Conversation

@nehal-a2z

@nehal-a2z nehal-a2z commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Stacked on #44. The goal is to get CodeRabbit invoked whenever the user or the agent wants code reviewed, checked or verified, and to keep it invoked for the rest of the session.

The skill was losing to Claude Code's built-in code-review. In claude plugin eval traces, "check my changes before I push" called Skill {"skill": "code-review"} in 4/4 runs. That bare name resolves to the bundled reviewer, not coderabbit:code-review. The suite's activation graders accepted the bare name, so they had been counting the built-in as this skill.

  • Rename: skills/code-review becomes skills/coderabbit-review. The skill now provides /coderabbit:coderabbit-review itself, so commands/coderabbit-review.md is removed (same slash command, no duplicate). The Gemini /coderabbit:review command is unchanged.
  • Description: it now says to review any code change, verify a fix, or confirm the agent's own changes before a task is reported done, even when CodeRabbit isn't named. It says to prefer this skill over the built-in code-review and verify skills when CodeRabbit is installed. It still leads with answering any CodeRabbit CLI question. It is 667 characters, well under the 1,024-character limit.
  • Standing rule: loaded skill content stays in context, so after changing code the agent reviews it with CodeRabbit before saying it's done.
    • once per task, scoped to what changed;
    • skipped when the user said not to, only docs changed, or there's no diff;
    • never uses --use-credits without approval;
    • never presents a manual reading of the diff as CodeRabbit's review.
  • Subagent: the code-reviewer subagent is marked "use proactively after writing or modifying code".
  • Evals:
    • routing-* cases: four checks that don't name CodeRabbit, the agent's own edit, and three controls, with scaffolded Git repos.
    • shell-* cases: a stand-in coderabbit in the run's ~/.local/bin, passing only if CodeRabbit actually runs.
    • Activation graders now require coderabbit:coderabbit-review.
    • prepare_comparison.py copies both old and new skill directories.

The agent's own edits aren't reached by routing (0/6 across both descriptions). That's handled by the Stop-hook reminder in #46, which gets 3/3.

Sharing the name gained nothing. Same prompts, each run with a shell, a stand-in CLI and 3 repetitions:

  • Old name: "Review my code", "Do a code review of my changes", "Can you review this diff?" and "Check my changes before I push" invoked the built-in in 12/12 runs, and CodeRabbit ran 0/12.
  • Renamed skill: the same prompts reached this skill 12/12 times, and CodeRabbit ran 11/12.
  • Typed /code-review runs the built-in under either name.
  • Only prompts naming CodeRabbit reached this skill before the rename.

The built-in exists from v2.1.147, and Claude could start it on its own everywhere from v2.1.246. See evals/RESULTS.md.

Affected surfaces

Public references

Validation

  • claude plugin validate . passes, both skills' frontmatter parses, and git diff --check is clean.
  • claude plugin eval setup: Claude Code 2.1.282, Sonnet 4.6, Haiku 4.5 judge, --ablation none --scaffold --allow-tools Edit Bash, 2 runs per case, strict graders, Align skills with CLI workflows and add reproducible comparisons #44 head vs this branch:
    • Whole suite: 28 → 32 of 50 cases fully passed; mean score 0.71 → 0.82.
    • Loaded this skill: pre-push check 0/2 → 2/2; "verify my fix" 0/2 → 2/2.
    • Ran coderabbit review with a shell: 0/2 → 2/2.
    • Controls: 6/6 stayed quiet in both.
  • Final wording: it puts "answer any CodeRabbit CLI question" first, which keeps readiness-auth-denied at 4/6 (the old description pooled 9/14). The shell pre-push case ran CodeRabbit 2/2.
  • Still weak: "sanity-check my diff" (1/4) and "ready for a PR?" (0/4) mostly get a manual review with no skill.
  • Full numbers are in evals/RESULTS.md. These are small samples.

Risks

  • Existing CLI-installed skills: coderabbit skills installs by directory name, so existing installs keep a stale code-review folder next to coderabbit-review until the installer removes it.
  • Muscle memory: anyone who typed /coderabbit:code-review now uses /coderabbit:coderabbit-review, which was already the documented command.
  • Quota: reviews run more often, which uses the user's review quota.

Checklist

  • SKILL.md stays focused on activation, routing, domain context, and workflow framing.
  • Detailed material uses focused references and progressive disclosure.
  • Repeatable deterministic operations use scripts or tools when practical.
  • Every referenced file, script, tool, command, and option exists.
  • Native commands, agents, manifests, docs, and distribution records remain aligned.
  • The change follows the Agent Skills specification, AGENTS.md open format, and current public guidance for every declared host.
  • The pull request contains no credentials, private links, private configuration, or private operational details.

🤖 Generated with Claude Code

…ecking

In Claude Code the skill named `code-review` lost review requests to the
built-in `code-review`: traces showed the bare name resolving to the
bundled reviewer in 4/4 "check my changes" runs, and activation graders
counted that as this skill.

- Rename the skill to `coderabbit-review`. It now also provides
  `/coderabbit:coderabbit-review`, replacing the duplicate command.
- Description: review or check any code change, verify a fix, or confirm
  the agent's own changes before reporting done, even when CodeRabbit
  isn't named; prefer it over the built-in code-review and verify skills;
  still answer any CodeRabbit CLI question.
- Skill body: after changing code, review it once per task before saying
  it's done; never present a manual reading of the diff as CodeRabbit's
  review. The code-reviewer subagent is marked for proactive use.
- Evals: routing and shell cases with scaffolded repositories and a
  stand-in CLI; activation graders require this plugin's skill name.

Over the suite (2 runs per case) cases fully passed rose from 28 to 32 of
50; the pre-push check loaded this skill 0/2 -> 2/2 and ran CodeRabbit
0/2 -> 2/2; controls stayed clean. Details in evals/RESULTS.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository: coderabbitai/skills/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: 70af8eb2-1d88-4d7a-99db-acc037ca3c5c

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

With the old name, generic review requests invoked Claude Code's built-in
code-review 12/12 times and CodeRabbit ran 0/12; after the rename this
skill ran 11/12. A typed /code-review runs the built-in either way.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant