From 793dd8e5663dd2655361bef16de2ad2058ad3db0 Mon Sep 17 00:00:00 2001 From: Alan Sikora Date: Fri, 4 Sep 2026 20:35:23 -0300 Subject: [PATCH 1/3] feat(eval): golden-prompt harness over frozen review inputs MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Prompt edits are invisible in code review. Their effect surfaces later as a change in the findings, by which point the cause is hard to attribute — and instruction overhead is paid on every review forever, so growth that nobody measured compounds silently. This adds the layer that makes prompt construction observable: - internal/evalcorpus — the on-disk fixture format. A fixture is everything a review needs to build its prompt (PR metadata, diff, the file contents the reviewer read, project docs, config), so rendering one touches neither the network nor a working tree. It declares its own structs rather than reusing review's types, which keeps it dependency-free (no import cycle for review's own tests) and pins the serialised format against refactors inside review. - cmd/evalsnap — captures a fixture by reusing review's real fetch path, so a fixture is what a review would have seen rather than an approximation. Not part of the shipped binary. - TestPromptGolden — renders every fixture and diffs against a checked-in golden, reporting the first differing line and the size delta rather than dumping two multi-kilobyte prompts. SIZES.txt records each prompt's rendered size so growth lands in the diff. Validated against the open prompt chain (#176-#180): the harness reports +4,599 characters, identical across all three fixtures, and names the section responsible. This measures prompt construction, not review quality — no model runs. Judging whether a prompt change helps needs labelled findings, repeated runs to establish variance, and real model calls. That layer does not exist yet, and the variance measurement has to come first: with ~5 findings per PR, run-to-run noise can swallow the effect being looked for. Corpus fixtures embed full file contents, so a fixture from a private repository must never be committed to this public one. evalsnap checks the source repo's visibility with GitHub and refuses to write into any git-tracked directory, before it fetches or reads anything. --- CLAUDE.md | 19 + cmd/evalsnap/main.go | 312 ++ cmd/evalsnap/main_test.go | 81 + internal/evalcorpus/corpus.go | 168 + internal/review/prompt_golden_test.go | 183 ++ internal/review/testdata/corpus/README.md | 73 + internal/review/testdata/corpus/SIZES.txt | 5 + .../corpus/alansikora-codecanary-pr165.json | 27 + .../alansikora-codecanary-pr165.prompt.golden | 1684 ++++++++++ .../corpus/alansikora-codecanary-pr173.json | 35 + .../alansikora-codecanary-pr173.prompt.golden | 2727 +++++++++++++++++ .../corpus/alansikora-codecanary-pr175.json | 29 + .../alansikora-codecanary-pr175.prompt.golden | 1860 +++++++++++ 13 files changed, 7203 insertions(+) create mode 100644 cmd/evalsnap/main.go create mode 100644 cmd/evalsnap/main_test.go create mode 100644 internal/evalcorpus/corpus.go create mode 100644 internal/review/prompt_golden_test.go create mode 100644 internal/review/testdata/corpus/README.md create mode 100644 internal/review/testdata/corpus/SIZES.txt create mode 100644 internal/review/testdata/corpus/alansikora-codecanary-pr165.json create mode 100644 internal/review/testdata/corpus/alansikora-codecanary-pr165.prompt.golden create mode 100644 internal/review/testdata/corpus/alansikora-codecanary-pr173.json create mode 100644 internal/review/testdata/corpus/alansikora-codecanary-pr173.prompt.golden create mode 100644 internal/review/testdata/corpus/alansikora-codecanary-pr175.json create mode 100644 internal/review/testdata/corpus/alansikora-codecanary-pr175.prompt.golden diff --git a/CLAUDE.md b/CLAUDE.md index c952683..9db65c5 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -16,6 +16,8 @@ cmd/ install_skill.go # codecanary install-skill — write embedded Claude skill to disk setup.go # codecanary setup [local|github] auth.go # codecanary auth [status|delete] + evalsnap/ # Dev tool — freeze a PR's review inputs into a corpus fixture + main.go # NOT part of the codecanary binary internal/ review/ runner.go # Core review pipeline — single Run() entry point @@ -43,8 +45,12 @@ internal/ local.go # Local diff & git operations state.go # Local state persistence docs.go # Project doc discovery + prompt_golden_test.go # Renders corpus fixtures into prompts, diffs vs goldens + testdata/corpus/ # Frozen review inputs + prompt goldens + SIZES.txt (public PRs only) credentials/ # Credential storage (keychain with file fallback) keyring.go # Store/Retrieve/Delete — keychain first, ~/.codecanary/credentials.json fallback + evalcorpus/ # Frozen review inputs — fixture format + load/save (dev tooling) + corpus.go # Fixture type, Dir/Load/Save/List skills/ # Claude Code skills embedded in the binary via //go:embed skills.go # Exports CodecanaryFix() returning the skill body codecanary-fix/SKILL.md # Canonical skill source (duplicated at .claude/skills/codecanary-fix/SKILL.md; parity enforced by skills_test.go) @@ -71,6 +77,19 @@ install.sh # Downloads and installs codecanary binary permanently ## Binary - **`codecanary`** — single binary for reviews, setup, and credential management. Installed locally via `install.sh`, also used by the GitHub Action. +- **`evalsnap`** (`cmd/evalsnap`) — developer tool, not shipped. `go build ./cmd/review` does not pull it in. + +## Prompt evaluation harness + +`internal/review/testdata/corpus/` holds frozen review inputs captured from real PRs. `TestPromptGolden` renders each into a prompt and diffs it against a checked-in golden, and `SIZES.txt` records every prompt's rendered size. + +This exists because prompt edits are otherwise invisible in code review: their effect surfaces later as a change in the findings, when the cause is hard to attribute. Instruction overhead is paid on every review forever, so growth should be a number in the diff, not a surprise. + +It measures prompt *construction*, not review *quality* — no model runs. Judging whether a prompt change helps needs labelled findings, repeated runs to establish variance, and real model calls; that is a separate layer that does not exist yet. + +Regenerate goldens with `go test ./internal/review/ -run Golden -update`, and state the size delta in the PR description. + +**Fixtures committed to this repo must come from public repositories only.** A fixture embeds the PR diff and the full contents of every file the reviewer read; committing one from a private repo publishes that source permanently. `evalsnap` refuses to write a private repo's fixture into any git-tracked directory. Keep those outside the repo and point `$CODECANARY_EVAL_CORPUS` at them. See `internal/review/testdata/corpus/README.md`. ## Build diff --git a/cmd/evalsnap/main.go b/cmd/evalsnap/main.go new file mode 100644 index 0000000..fa26b17 --- /dev/null +++ b/cmd/evalsnap/main.go @@ -0,0 +1,312 @@ +// Command evalsnap freezes a pull request's review inputs into a corpus +// fixture, so a prompt can later be rendered from it without touching the +// network or a working tree. +// +// It is developer tooling and is deliberately not part of the codecanary +// binary: `go build ./cmd/review` does not pull it in. +// +// Fidelity comes from reusing the review package's own fetch path rather than +// reimplementing it — the same FetchPR, the same file-content reader with the +// same size and ignore filtering, the same project-doc discovery. A fixture is +// therefore what a real review would have seen, not an approximation of it. +// That reuse is also why the tool must run from inside a checkout of the +// target repository at the PR's head commit: file contents and project docs +// are read from the working tree, and reading them at the wrong commit would +// bake a mismatched snapshot into the corpus. The tool verifies HEAD before +// capturing rather than trusting the caller. +// +// Usage: +// +// cd /path/to/target-repo +// git fetch origin pull/1234/head && git checkout FETCH_HEAD +// evalsnap --repo owner/name --pr 1234 --out "$CODECANARY_EVAL_CORPUS" +// +// --out has no default: a fixture embeds the PR's diff and the full contents +// of every file the reviewer read, so where it lands is a decision to make +// deliberately rather than one to fall into. Fixtures from a private +// repository must go somewhere git does not track — guardPrivateSource +// enforces that rather than trusting it. +package main + +import ( + "flag" + "fmt" + "os" + "os/exec" + "path/filepath" + "strings" + "time" + + "github.com/alansikora/codecanary/internal/evalcorpus" + "github.com/alansikora/codecanary/internal/review" +) + +func main() { + if err := run(); err != nil { + fmt.Fprintf(os.Stderr, "evalsnap: %v\n", err) + os.Exit(1) + } +} + +func run() error { + var ( + repo = flag.String("repo", "", "GitHub repo as owner/name (required)") + pr = flag.Int("pr", 0, "pull request number (required)") + out = flag.String("out", "", "corpus directory to write into (required)") + name = flag.String("name", "", "fixture name (default: --pr)") + configPath = flag.String("config", "", "review config path (auto-detected when empty)") + force = flag.Bool("force", false, "capture even if HEAD is not the PR's head commit") + ) + flag.Parse() + + switch { + case *repo == "": + return fmt.Errorf("--repo is required") + case *pr == 0: + return fmt.Errorf("--pr is required") + case *out == "": + return fmt.Errorf("--out is required (the corpus directory to write into)") + } + + // Safety before work: this is the check that prevents publishing private + // source, so it runs before any fetch, any read, and any write. + if err := guardPrivateSource(*repo, *out); err != nil { + return err + } + + prData, err := review.FetchPR(*repo, *pr) + if err != nil { + return fmt.Errorf("fetching PR: %w", err) + } + + headSHA, err := review.HeadSHA() + if err != nil { + return fmt.Errorf("reading HEAD (run this from inside a checkout of %s): %w", *repo, err) + } + prHead, err := prHeadSHA(*repo, *pr) + if err != nil { + return err + } + if headSHA != prHead && !*force { + fmt.Fprintf(os.Stderr, ` +File contents and project docs are read from the working tree, so capturing +now would freeze a snapshot that does not match the diff. Check out the PR +head first: + + git fetch origin pull/%d/head && git checkout FETCH_HEAD + +Pass --force to capture anyway. + +`, *pr) + return fmt.Errorf("HEAD is %s but PR #%d is at %s", short(headSHA), *pr, short(prHead)) + } + + cfg, err := loadConfig(*configPath) + if err != nil { + return fmt.Errorf("loading review config: %w", err) + } + + fileContents, skipped := review.FetchFileContents( + prData.Files, cfg.Ignore, cfg.EffectiveMaxFileSize(), cfg.EffectiveMaxTotalSize()) + if len(skipped) > 0 { + fmt.Fprintf(os.Stderr, "skipped %d large/ignored file(s): %s\n", + len(skipped), strings.Join(skipped, ", ")) + } + + fixture := &evalcorpus.Fixture{ + Name: fixtureName(*name, *repo, *pr), + Repo: *repo, + PRNumber: *pr, + HeadSHA: headSHA, + CapturedAt: time.Now().UTC().Format(time.RFC3339), + PR: toPRInput(prData, fileContents), + Config: toConfigInput(cfg), + ProjectDocs: review.ReadProjectDocs(prData.Files), + } + + path, err := evalcorpus.Save(*out, fixture) + if err != nil { + return err + } + + info, err := os.Stat(path) + if err != nil { + return err + } + fmt.Printf("wrote %s (%s, %d files, %d project doc(s))\n", + path, humanBytes(info.Size()), len(fixture.PR.Files), len(fixture.ProjectDocs)) + return nil +} + +// loadConfig resolves the review config the same way a real review does, +// falling back to an empty config when the target repo has none — a repo +// without a .codecanary config is a perfectly valid thing to capture, and a +// nil config is exactly what the prompt builder would receive there. +func loadConfig(path string) (*review.ReviewConfig, error) { + if path == "" { + found, err := review.FindConfig() + if err != nil { + fmt.Fprintf(os.Stderr, "no review config found; capturing with an empty config\n") + return &review.ReviewConfig{}, nil + } + path = found + } + return review.LoadConfig(path) +} + +// guardPrivateSource refuses to write a fixture captured from a private +// repository into a directory that git tracks. +// +// A fixture embeds the PR's diff and the full contents of every file the +// reviewer read. Committing one captured from a private repository publishes +// that source, and git history makes the mistake permanent — so this is +// checked before any file content is read, not after. +// +// The test is "does git ignore the destination", not "is the destination +// inside this repository": a corpus kept anywhere git would track it carries +// the same risk, and an ignored path is the arrangement that is actually safe. +func guardPrivateSource(repo, out string) error { + private, err := isPrivateRepo(repo) + if err != nil { + return fmt.Errorf("could not determine whether %s is private (refusing to guess): %w", repo, err) + } + if !private { + return nil + } + tracked, err := gitWouldTrack(out) + if err != nil { + return err + } + if !tracked { + return nil + } + fmt.Fprintf(os.Stderr, ` +A fixture embeds the PR's diff and the full contents of every file the +reviewer read. Committing it would publish that source permanently. + +Write the corpus somewhere git does not track — a private repository, or a +gitignored directory — and point the harness at it: + + export %s=/path/to/private/corpus + evalsnap --repo %s --pr N --out "$%s" + +`, evalcorpus.EnvCorpusDir, repo, evalcorpus.EnvCorpusDir) + return fmt.Errorf("refusing to capture: %s is private but %s is tracked by git", repo, out) +} + +// isPrivateRepo asks GitHub rather than inferring from the name, so a rename +// or a transfer cannot quietly turn the guard off. +func isPrivateRepo(repo string) (bool, error) { + cmd := exec.Command("gh", "repo", "view", repo, "--json", "isPrivate", "--jq", ".isPrivate") + var stderr strings.Builder + cmd.Stderr = &stderr + out, err := cmd.Output() + if err != nil { + return false, fmt.Errorf("%w: %s", err, strings.TrimSpace(stderr.String())) + } + return strings.TrimSpace(string(out)) == "true", nil +} + +// gitWouldTrack reports whether git would track files written to dir: it is +// inside a work tree and not ignored. A path outside any repository, or one +// git ignores, is safe to write a private corpus into. +func gitWouldTrack(dir string) (bool, error) { + abs, err := filepath.Abs(dir) + if err != nil { + return false, err + } + // check-ignore needs an existing ancestor to resolve against. + probe := abs + for { + if _, err := os.Stat(probe); err == nil { + break + } + parent := filepath.Dir(probe) + if parent == probe { + return false, nil // nothing on this path exists yet + } + probe = parent + } + + inTree := exec.Command("git", "-C", probe, "rev-parse", "--is-inside-work-tree") + if out, err := inTree.Output(); err != nil || strings.TrimSpace(string(out)) != "true" { + return false, nil // not a git work tree at all + } + + // check-ignore exits 0 when the path IS ignored, 1 when it is not. + ignored := exec.Command("git", "-C", probe, "check-ignore", "-q", abs) + if err := ignored.Run(); err == nil { + return false, nil // ignored, therefore not tracked + } + return true, nil +} + +// prHeadSHA asks GitHub for the PR's head commit so the working tree can be +// checked against it before anything is read from disk. +func prHeadSHA(repo string, pr int) (string, error) { + cmd := exec.Command("gh", "api", + fmt.Sprintf("repos/%s/pulls/%d", repo, pr), "--jq", ".head.sha") + var stderr strings.Builder + cmd.Stderr = &stderr + out, err := cmd.Output() + if err != nil { + return "", fmt.Errorf("resolving PR head sha: %w: %s", err, strings.TrimSpace(stderr.String())) + } + return strings.TrimSpace(string(out)), nil +} + +func fixtureName(explicit, repo string, pr int) string { + if explicit != "" { + return explicit + } + return fmt.Sprintf("%s-pr%d", strings.ReplaceAll(repo, "/", "-"), pr) +} + +func toPRInput(pr *review.PRData, contents map[string]string) evalcorpus.PRInput { + return evalcorpus.PRInput{ + Number: pr.Number, + Title: pr.Title, + Body: pr.Body, + Author: pr.Author, + BaseBranch: pr.BaseBranch, + HeadBranch: pr.HeadBranch, + Diff: pr.Diff, + Files: pr.Files, + FileContents: contents, + } +} + +func toConfigInput(cfg *review.ReviewConfig) *evalcorpus.ConfigInput { + if cfg == nil { + return nil + } + rules := make([]evalcorpus.RuleInput, 0, len(cfg.Rules)) + for _, r := range cfg.Rules { + rules = append(rules, evalcorpus.RuleInput{ + ID: r.ID, + Description: r.Description, + Severity: r.Severity, + Paths: r.Paths, + ExcludePaths: r.ExcludePaths, + }) + } + return &evalcorpus.ConfigInput{Rules: rules, Context: cfg.Context, Ignore: cfg.Ignore} +} + +func short(sha string) string { + if len(sha) > 7 { + return sha[:7] + } + return sha +} + +func humanBytes(n int64) string { + switch { + case n >= 1<<20: + return fmt.Sprintf("%.1f MB", float64(n)/(1<<20)) + case n >= 1<<10: + return fmt.Sprintf("%.1f KB", float64(n)/(1<<10)) + default: + return fmt.Sprintf("%d B", n) + } +} diff --git a/cmd/evalsnap/main_test.go b/cmd/evalsnap/main_test.go new file mode 100644 index 0000000..c5987ba --- /dev/null +++ b/cmd/evalsnap/main_test.go @@ -0,0 +1,81 @@ +package main + +import ( + "os" + "os/exec" + "path/filepath" + "testing" +) + +// gitWouldTrack is what stands between a private repository's source and a +// public git history, so its edges are worth pinning down: a destination that +// does not exist yet, one git ignores, one git would track, and one outside +// any repository at all. +func TestGitWouldTrack(t *testing.T) { + repo := t.TempDir() + runGit(t, repo, "init", "-q") + if err := os.WriteFile(filepath.Join(repo, ".gitignore"), []byte("ignored/\n"), 0o644); err != nil { + t.Fatalf("writing .gitignore: %v", err) + } + for _, d := range []string{"tracked", "ignored"} { + if err := os.MkdirAll(filepath.Join(repo, d), 0o755); err != nil { + t.Fatalf("creating %s: %v", d, err) + } + } + + cases := []struct { + name string + dir string + want bool + }{ + {"tracked directory inside a repo", filepath.Join(repo, "tracked"), true}, + {"gitignored directory", filepath.Join(repo, "ignored"), false}, + {"gitignored directory that does not exist yet", filepath.Join(repo, "ignored", "corpus"), false}, + {"path outside any git work tree", t.TempDir(), false}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + got, err := gitWouldTrack(tc.dir) + if err != nil { + t.Fatalf("gitWouldTrack: %v", err) + } + if got != tc.want { + t.Errorf("gitWouldTrack(%s) = %v, want %v", tc.dir, got, tc.want) + } + }) + } +} + +// A destination that does not exist yet inside a tracked directory must still +// read as tracked: creating it is exactly what evalsnap is about to do, and +// answering "false" because it is absent would open the hole the guard exists +// to close. +func TestGitWouldTrack_UncreatedPathUnderTrackedDir(t *testing.T) { + repo := t.TempDir() + runGit(t, repo, "init", "-q") + + got, err := gitWouldTrack(filepath.Join(repo, "corpus", "nested")) + if err != nil { + t.Fatalf("gitWouldTrack: %v", err) + } + if !got { + t.Error("an uncreated path under a tracked repo should read as tracked") + } +} + +func TestFixtureName(t *testing.T) { + if got := fixtureName("", "owner/name", 42); got != "owner-name-pr42" { + t.Errorf("fixtureName = %q, want owner-name-pr42", got) + } + if got := fixtureName("custom", "owner/name", 42); got != "custom" { + t.Errorf("explicit name should win, got %q", got) + } +} + +func runGit(t *testing.T, dir string, args ...string) { + t.Helper() + cmd := exec.Command("git", append([]string{"-C", dir}, args...)...) + if out, err := cmd.CombinedOutput(); err != nil { + t.Fatalf("git %v: %v\n%s", args, err, out) + } +} diff --git a/internal/evalcorpus/corpus.go b/internal/evalcorpus/corpus.go new file mode 100644 index 0000000..99ccbde --- /dev/null +++ b/internal/evalcorpus/corpus.go @@ -0,0 +1,168 @@ +// Package evalcorpus defines the on-disk format for frozen review inputs and +// the helpers to read and write them. +// +// A fixture is everything a review needs in order to build its prompt: the PR +// metadata and diff, the file contents the reviewer would have read, the +// project docs it would have picked up, and the review config in force. Once +// captured, rendering a prompt from a fixture touches neither the network nor +// a working tree, so the same input produces the same prompt on any machine. +// +// The package deliberately declares its own plain structs rather than reusing +// the review package's types. Keeping it dependency-free means the review +// package's own tests can import it without an import cycle, and it pins the +// serialised format so a refactor inside review can't silently invalidate a +// corpus captured months earlier. +// +// # Corpus location +// +// A small corpus captured from this repository's own public pull requests is +// committed under internal/review/testdata/corpus, so the harness runs in CI +// and anyone reading the tests can see the shape of a fixture without having +// to capture one first. +// +// Richer corpora are usually captured from private repositories, and those +// must not be committed here. $CODECANARY_EVAL_CORPUS points the harness at +// one kept outside this repository; see Dir. +package evalcorpus + +import ( + "encoding/json" + "fmt" + "os" + "path/filepath" + "sort" + "strings" +) + +// FormatVersion is bumped whenever Fixture changes shape in a way that makes +// previously captured fixtures unreadable. Load refuses a fixture from a newer +// format rather than silently misreading it. +const FormatVersion = 1 + +// EnvCorpusDir is the environment variable naming the directory that holds the +// corpus. See Dir. +const EnvCorpusDir = "CODECANARY_EVAL_CORPUS" + +// Fixture is one frozen review input. +type Fixture struct { + FormatVersion int `json:"format_version"` + Name string `json:"name"` + Repo string `json:"repo"` + PRNumber int `json:"pr_number"` + HeadSHA string `json:"head_sha"` + CapturedAt string `json:"captured_at"` + + PR PRInput `json:"pr"` + Config *ConfigInput `json:"config,omitempty"` + ProjectDocs map[string]string `json:"project_docs,omitempty"` +} + +// PRInput mirrors the fields of review.PRData that feed prompt construction. +type PRInput struct { + Number int `json:"number"` + Title string `json:"title"` + Body string `json:"body"` + Author string `json:"author"` + BaseBranch string `json:"base_branch"` + HeadBranch string `json:"head_branch"` + Diff string `json:"diff"` + Files []string `json:"files"` + FileContents map[string]string `json:"file_contents,omitempty"` +} + +// ConfigInput mirrors the review config fields that reach the prompt. Model, +// provider and budget settings are deliberately absent: they steer how the +// review runs, not what the prompt says, and capturing them would invite +// fixtures that drift with unrelated config changes. +type ConfigInput struct { + Rules []RuleInput `json:"rules,omitempty"` + Context string `json:"context,omitempty"` + Ignore []string `json:"ignore,omitempty"` +} + +// RuleInput mirrors review.Rule. +type RuleInput struct { + ID string `json:"id"` + Description string `json:"description"` + Severity string `json:"severity"` + Paths []string `json:"paths,omitempty"` + ExcludePaths []string `json:"exclude_paths,omitempty"` +} + +// Dir resolves the corpus directory, preferring $CODECANARY_EVAL_CORPUS over +// the committed corpus at fallback. +// +// The override exists so a larger corpus captured from a private repository +// can drive the same harness from outside this repository. When the variable +// is set but does not name a directory, Dir reports it as not-ok rather than +// quietly falling back: someone who pointed the harness somewhere specific +// should hear that the path is wrong, not silently get different results. +// +// ok=false means no corpus is readable. For the committed fallback that +// should not happen in a normal checkout, so callers are better off failing +// than skipping — a corpus that vanished is a broken checkout, not a +// legitimate state. +func Dir(fallback string) (dir string, ok bool) { + if env := strings.TrimSpace(os.Getenv(EnvCorpusDir)); env != "" { + info, err := os.Stat(env) + return env, err == nil && info.IsDir() + } + info, err := os.Stat(fallback) + return fallback, err == nil && info.IsDir() +} + +// Save writes a fixture as indented JSON under dir, named after f.Name. +func Save(dir string, f *Fixture) (string, error) { + if f.Name == "" { + return "", fmt.Errorf("fixture has no name") + } + f.FormatVersion = FormatVersion + if err := os.MkdirAll(dir, 0o755); err != nil { + return "", fmt.Errorf("creating corpus dir: %w", err) + } + data, err := json.MarshalIndent(f, "", " ") + if err != nil { + return "", fmt.Errorf("encoding fixture: %w", err) + } + path := filepath.Join(dir, f.Name+".json") + if err := os.WriteFile(path, append(data, '\n'), 0o644); err != nil { + return "", fmt.Errorf("writing fixture: %w", err) + } + return path, nil +} + +// Load reads one fixture from disk. +func Load(path string) (*Fixture, error) { + data, err := os.ReadFile(path) + if err != nil { + return nil, err + } + var f Fixture + if err := json.Unmarshal(data, &f); err != nil { + return nil, fmt.Errorf("decoding %s: %w", filepath.Base(path), err) + } + if f.FormatVersion > FormatVersion { + return nil, fmt.Errorf("%s: fixture format v%d is newer than this build understands (v%d); update codecanary or re-capture the corpus", + filepath.Base(path), f.FormatVersion, FormatVersion) + } + return &f, nil +} + +// List loads every fixture in dir, sorted by name so callers iterate in a +// stable order. +func List(dir string) ([]*Fixture, error) { + matches, err := filepath.Glob(filepath.Join(dir, "*.json")) + if err != nil { + return nil, err + } + sort.Strings(matches) + out := make([]*Fixture, 0, len(matches)) + for _, m := range matches { + f, err := Load(m) + if err != nil { + return nil, err + } + out = append(out, f) + } + return out, nil +} diff --git a/internal/review/prompt_golden_test.go b/internal/review/prompt_golden_test.go new file mode 100644 index 0000000..e306304 --- /dev/null +++ b/internal/review/prompt_golden_test.go @@ -0,0 +1,183 @@ +package review + +import ( + "flag" + "fmt" + "os" + "path/filepath" + "sort" + "strings" + "testing" + + "github.com/alansikora/codecanary/internal/evalcorpus" +) + +// -update rewrites the golden files instead of comparing against them. +// +// go test ./internal/review/ -run Golden -update +var updateGolden = flag.Bool("update", false, "rewrite prompt golden files") + +// defaultCorpusDir holds fixtures captured from this repository's own public +// pull requests. It is committed so the harness runs in CI and so the shape of +// a fixture is visible to anyone reading these tests. Fixtures from private +// repositories belong in $CODECANARY_EVAL_CORPUS instead — see the package +// doc on internal/evalcorpus and testdata/corpus/README.md. +const defaultCorpusDir = "testdata/corpus" + +// TestPromptGolden renders the review prompt for every frozen fixture and +// compares it against a checked-in golden file. +// +// This does not measure review quality — no model runs, nothing is judged. It +// answers a narrower question that nothing else in the suite answers: did a +// change to prompt construction alter the prompt in a way its author did not +// intend, and by how much? Prompt edits are otherwise invisible in review; +// their effect only shows up later as a change in the findings, by which point +// the cause is hard to attribute. +// +// The size report the test writes alongside the goldens is the point as much +// as the diff is. Instruction overhead is paid on every review forever, so a +// PR that adds 4KB of prompt should have to say so out loud. +func TestPromptGolden(t *testing.T) { + dir, ok := evalcorpus.Dir(defaultCorpusDir) + if !ok { + if custom := os.Getenv(evalcorpus.EnvCorpusDir); custom != "" { + t.Fatalf("%s points at %q, which is not a directory", evalcorpus.EnvCorpusDir, custom) + } + t.Fatalf("corpus missing at %s — this directory is committed; the checkout looks incomplete", dir) + } + + fixtures, err := evalcorpus.List(dir) + if err != nil { + t.Fatalf("loading corpus from %s: %v", dir, err) + } + if len(fixtures) == 0 { + t.Fatalf("no fixtures in %s; capture one with `go run ./cmd/evalsnap --help`", dir) + } + + sizes := make([]string, 0, len(fixtures)) + for _, f := range fixtures { + t.Run(f.Name, func(t *testing.T) { + got := BuildPrompt(toPRData(f), toReviewConfig(f.Config), 0, f.ProjectDocs) + sizes = append(sizes, fmt.Sprintf("%-40s %7d", f.Name, len(got))) + compareGolden(t, filepath.Join(dir, f.Name+".prompt.golden"), got) + }) + } + + writeSizeReport(t, dir, sizes) +} + +// compareGolden diffs rendered output against its golden file, reporting the +// first differing line rather than dumping two multi-kilobyte prompts. +func compareGolden(t *testing.T, path, got string) { + t.Helper() + + if *updateGolden { + if err := os.WriteFile(path, []byte(got), 0o644); err != nil { + t.Fatalf("writing golden: %v", err) + } + return + } + + wantBytes, err := os.ReadFile(path) + if err != nil { + if os.IsNotExist(err) { + t.Fatalf("no golden file at %s; run `go test ./internal/review/ -run Golden -update`", path) + } + t.Fatalf("reading golden: %v", err) + } + want := string(wantBytes) + if got == want { + return + } + + gotLines, wantLines := strings.Split(got, "\n"), strings.Split(want, "\n") + delta := len(got) - len(want) + for i := 0; i < len(gotLines) || i < len(wantLines); i++ { + g, w := lineAt(gotLines, i), lineAt(wantLines, i) + if g == w { + continue + } + t.Errorf(`prompt changed (%+d chars, %d -> %d) + +first difference at line %d: + want: %s + got: %s + +If this change is intended, re-run with -update and make sure the size delta +is called out in the PR description.`, delta, len(want), len(got), i+1, truncate(w), truncate(g)) + return + } + t.Errorf("prompt changed (%+d chars) with no differing line; check trailing whitespace", delta) +} + +// writeSizeReport records the rendered size of every prompt so growth shows up +// as a reviewable diff rather than as something nobody measured. +func writeSizeReport(t *testing.T, dir string, sizes []string) { + t.Helper() + if len(sizes) == 0 { + return + } + sort.Strings(sizes) + body := "# Rendered prompt sizes, in characters.\n" + + "# Regenerate: go test ./internal/review/ -run Golden -update\n" + + strings.Join(sizes, "\n") + "\n" + + path := filepath.Join(dir, "SIZES.txt") + if *updateGolden { + if err := os.WriteFile(path, []byte(body), 0o644); err != nil { + t.Fatalf("writing size report: %v", err) + } + return + } + prev, err := os.ReadFile(path) + if err != nil || string(prev) == body { + return + } + t.Errorf("prompt sizes changed:\n\n%s\nwant:\n\n%s\nre-run with -update", body, prev) +} + +func lineAt(lines []string, i int) string { + if i < len(lines) { + return lines[i] + } + return "" +} + +func truncate(s string) string { + const max = 120 + if len(s) <= max { + return fmt.Sprintf("%q", s) + } + return fmt.Sprintf("%q…", s[:max]) +} + +func toPRData(f *evalcorpus.Fixture) *PRData { + return &PRData{ + Number: f.PR.Number, + Title: f.PR.Title, + Body: f.PR.Body, + Author: f.PR.Author, + BaseBranch: f.PR.BaseBranch, + HeadBranch: f.PR.HeadBranch, + Diff: f.PR.Diff, + Files: f.PR.Files, + FileContents: f.PR.FileContents, + } +} + +func toReviewConfig(c *evalcorpus.ConfigInput) *ReviewConfig { + if c == nil { + return nil + } + rules := make([]Rule, 0, len(c.Rules)) + for _, r := range c.Rules { + rules = append(rules, Rule{ + ID: r.ID, + Description: r.Description, + Severity: r.Severity, + Paths: r.Paths, + ExcludePaths: r.ExcludePaths, + }) + } + return &ReviewConfig{Rules: rules, Context: c.Context, Ignore: c.Ignore} +} diff --git a/internal/review/testdata/corpus/README.md b/internal/review/testdata/corpus/README.md new file mode 100644 index 0000000..93ff7fd --- /dev/null +++ b/internal/review/testdata/corpus/README.md @@ -0,0 +1,73 @@ +# Prompt evaluation corpus + +Frozen review inputs. Each `*.json` here is everything a review needs to build +its prompt — PR metadata, diff, the file contents the reviewer would have read, +project docs, and the review config in force — captured from a real pull +request. Rendering a prompt from one touches neither the network nor a working +tree, so the same fixture produces the same prompt on any machine. + +`TestPromptGolden` in `internal/review/prompt_golden_test.go` renders each +fixture and compares the result against its `*.prompt.golden` file. + +## What this catches, and what it does not + +It catches **unintended changes to prompt construction**, and it puts a number +on intended ones. `SIZES.txt` records the rendered size of every prompt, so a +change that adds 4KB of instruction has to show that in its diff. + +That number is worth watching. Instruction overhead is paid on every review +forever, and prompt edits are otherwise invisible in code review — their effect +surfaces later as a change in the findings, by which point the cause is hard to +attribute. + +It does **not** measure review quality. No model runs and nothing is judged. +Whether a prompt change makes reviews better or worse is a separate question +that needs labelled findings, repeated runs to establish variance, and real +model calls. + +## Updating the goldens + +When a prompt change is intentional: + +```bash +go test ./internal/review/ -run Golden -update +``` + +Commit the regenerated goldens together with the change, and say what the size +delta is in the PR description. + +## Capturing a fixture + +`evalsnap` reuses the review package's own fetch path, so a fixture is what a +real review would have seen rather than an approximation. Because file contents +and project docs are read from the working tree, it must run from a checkout of +the target repository at the PR's head commit — it verifies this rather than +trusting you: + +```bash +cd /path/to/target-repo +git fetch origin pull/1234/head && git checkout FETCH_HEAD +go run github.com/alansikora/codecanary/cmd/evalsnap \ + --repo owner/name --pr 1234 --out /path/to/corpus +``` + +## Private repositories + +**Fixtures committed here must come from public repositories only.** A fixture +embeds the PR's diff and the full contents of every file the reviewer read; +committing one captured from a private repository publishes that source to this +public repo, and git history makes it permanent. + +`evalsnap` refuses to write a fixture from a private repository into any +directory git tracks. Keep those outside this repository and point the harness +at them: + +```bash +export CODECANARY_EVAL_CORPUS=/path/to/private/corpus +go test ./internal/review/ -run Golden +``` + +When that variable is set it replaces this directory for the run. The fixtures +committed here are deliberately small and few — enough to exercise the harness +in CI and to show what a fixture looks like, not enough to be a serious +evaluation set. diff --git a/internal/review/testdata/corpus/SIZES.txt b/internal/review/testdata/corpus/SIZES.txt new file mode 100644 index 0000000..38bee3b --- /dev/null +++ b/internal/review/testdata/corpus/SIZES.txt @@ -0,0 +1,5 @@ +# Rendered prompt sizes, in characters. +# Regenerate: go test ./internal/review/ -run Golden -update +alansikora-codecanary-pr165 70300 +alansikora-codecanary-pr173 130438 +alansikora-codecanary-pr175 86970 diff --git a/internal/review/testdata/corpus/alansikora-codecanary-pr165.json b/internal/review/testdata/corpus/alansikora-codecanary-pr165.json new file mode 100644 index 0000000..d731044 --- /dev/null +++ b/internal/review/testdata/corpus/alansikora-codecanary-pr165.json @@ -0,0 +1,27 @@ +{ + "format_version": 1, + "name": "alansikora-codecanary-pr165", + "repo": "alansikora/codecanary", + "pr_number": 165, + "head_sha": "49b69470772e6de5a01ec684f9f5a74c5ca310ce", + "captured_at": "2026-09-04T23:32:01Z", + "pr": { + "number": 165, + "title": "fix: surface gh stderr when FetchPR fails", + "body": "## Summary\n- `gh pr view` / `gh pr diff` errors in `FetchPR` were wrapped with only the exit status, so Actions logs showed `gh pr diff: exit status 1` with no detail (stderr from `.Output()` was dropped).\n- Added a small `runGH` helper that pipes stderr into the wrapped error, so the real gh message (rate limit, 5xx, token scope, etc.) shows up in logs.\n- For `gh pr diff` specifically, append a hint pointing at GitHub's pull request diff API (`GET /repos/{owner}/{repo}/pulls/{n}` with `Accept: application/vnd.github.v3.diff`), since transient failures there are the usual cause and a job retry typically resolves them.\n\nMotivated by [this failing run](https://github.com/thetechfx/bedrock/actions/runs/24791879406/job/72551238868?pr=1581): the PR was small and mergeable, neighboring codecanary runs succeeded the same minute, and no stderr was visible to confirm it was just a transient API blip.\n\n## Test plan\n- [ ] `go build ./...` and `go vet ./...` pass (done locally).\n- [ ] On the next real `gh pr diff` failure in Actions, verify the error line now includes gh's stderr.\n\n🤖 Generated with [Claude Code](https://claude.com/claude-code)", + "author": "alansikora", + "base_branch": "main", + "head_branch": "fix/surface-gh-stderr-in-fetchpr", + "diff": "diff --git a/internal/review/github.go b/internal/review/github.go\nindex cf0dee8..00081a6 100644\n--- a/internal/review/github.go\n+++ b/internal/review/github.go\n@@ -14,6 +14,24 @@ import (\n \t\"github.com/bmatcuk/doublestar/v4\"\n )\n \n+// runGH runs a gh command and returns stdout; on failure, the returned error\n+// includes gh's stderr so upstream API messages surface in logs instead of just\n+// \"exit status 1\".\n+func runGH(label string, args ...string) ([]byte, error) {\n+\tcmd := exec.Command(\"gh\", args...)\n+\tvar stderr bytes.Buffer\n+\tcmd.Stderr = \u0026stderr\n+\tout, err := cmd.Output()\n+\tif err != nil {\n+\t\tmsg := strings.TrimSpace(stderr.String())\n+\t\tif msg == \"\" {\n+\t\t\treturn nil, fmt.Errorf(\"%s: %w\", label, err)\n+\t\t}\n+\t\treturn nil, fmt.Errorf(\"%s: %s: %w\", label, msg, err)\n+\t}\n+\treturn out, nil\n+}\n+\n // parseRepoSlug splits a \"owner/name\" repository slug into its two parts.\n func parseRepoSlug(repo string) (owner, name string, err error) {\n \tparts := strings.SplitN(repo, \"/\", 2)\n@@ -84,12 +102,12 @@ func FetchPR(repo string, number int) (*PRData, error) {\n \tnumStr := fmt.Sprintf(\"%d\", number)\n \n \t// Fetch PR metadata as JSON.\n-\tviewOut, err := exec.Command(\"gh\", \"pr\", \"view\", numStr,\n+\tviewOut, err := runGH(\"gh pr view\", \"pr\", \"view\", numStr,\n \t\t\"--repo\", repo,\n \t\t\"--json\", \"title,body,author,baseRefName,headRefName,files\",\n-\t).Output()\n+\t)\n \tif err != nil {\n-\t\treturn nil, fmt.Errorf(\"gh pr view: %w\", err)\n+\t\treturn nil, err\n \t}\n \n \tvar view ghPRView\n@@ -97,12 +115,11 @@ func FetchPR(repo string, number int) (*PRData, error) {\n \t\treturn nil, fmt.Errorf(\"parsing gh pr view output: %w\", err)\n \t}\n \n-\t// Fetch the diff.\n-\tdiffOut, err := exec.Command(\"gh\", \"pr\", \"diff\", numStr,\n-\t\t\"--repo\", repo,\n-\t).Output()\n+\t// Fetch the diff. This hits GitHub's pull request diff API\n+\t// (GET /repos/{owner}/{repo}/pulls/{n} with Accept: application/vnd.github.v3.diff).\n+\tdiffOut, err := runGH(\"gh pr diff\", \"pr\", \"diff\", numStr, \"--repo\", repo)\n \tif err != nil {\n-\t\treturn nil, fmt.Errorf(\"gh pr diff: %w\", err)\n+\t\treturn nil, fmt.Errorf(\"%w (GitHub pull request diff API; if this looks like a transient 5xx/timeout, retrying the job may help)\", err)\n \t}\n \n \tfiles := make([]string, len(view.Files))\n", + "files": [ + "internal/review/github.go" + ], + "file_contents": { + "internal/review/github.go": "package review\n\nimport (\n\t\"bytes\"\n\t\"encoding/json\"\n\t\"fmt\"\n\t\"os\"\n\t\"os/exec\"\n\t\"path/filepath\"\n\t\"regexp\"\n\t\"strconv\"\n\t\"strings\"\n\n\t\"github.com/bmatcuk/doublestar/v4\"\n)\n\n// runGH runs a gh command and returns stdout; on failure, the returned error\n// includes gh's stderr so upstream API messages surface in logs instead of just\n// \"exit status 1\".\nfunc runGH(label string, args ...string) ([]byte, error) {\n\tcmd := exec.Command(\"gh\", args...)\n\tvar stderr bytes.Buffer\n\tcmd.Stderr = \u0026stderr\n\tout, err := cmd.Output()\n\tif err != nil {\n\t\tmsg := strings.TrimSpace(stderr.String())\n\t\tif msg == \"\" {\n\t\t\treturn nil, fmt.Errorf(\"%s: %w\", label, err)\n\t\t}\n\t\treturn nil, fmt.Errorf(\"%s: %s: %w\", label, msg, err)\n\t}\n\treturn out, nil\n}\n\n// parseRepoSlug splits a \"owner/name\" repository slug into its two parts.\nfunc parseRepoSlug(repo string) (owner, name string, err error) {\n\tparts := strings.SplitN(repo, \"/\", 2)\n\tif len(parts) != 2 {\n\t\treturn \"\", \"\", fmt.Errorf(\"invalid repo format %q, expected owner/name\", repo)\n\t}\n\treturn parts[0], parts[1], nil\n}\n\n// MaxFindingProximity is the maximum number of lines a finding may be from the\n// nearest changed line in the PR diff. Findings beyond this distance are dropped\n// (runner.go) or demoted from inline to body (PostReview). This enforces review\n// scope — keeping findings anchored to the PR's actual changes — and catches\n// hallucinated line numbers. A single constant ensures both checks stay in sync.\nconst MaxFindingProximity = 20\n\n// HTML comment markers for embedding and detecting review data.\n// Dual prefixes support both current (codecanary) and legacy (clanopy) markers.\nvar reviewMarkerPrefixes = []string{\"\u003c!-- codecanary:review \", \"\u003c!-- clanopy:review \"}\n\nconst (\n\treviewMarkerSuffix = \" --\u003e\"\n\tfindingMarkerPrefix = \"\u003c!-- codecanary:finding \"\n\tackMarkerPrefix = \"\u003c!-- codecanary:ack:\"\n\tlegacyAckPrefix = \"\u003c!-- clanopy:ack:\"\n)\n\n// PRData holds PR metadata and diff.\ntype PRData struct {\n\tNumber int\n\tTitle string\n\tBody string\n\tAuthor string\n\tBaseBranch string\n\tHeadBranch string\n\tDiff string\n\tFullDiff string // unfiltered diff for finding validation (set by prepareReview)\n\tFiles []string\n\tFileContents map[string]string // path -\u003e full file content\n}\n\n// ValidationDiff returns the unfiltered diff for finding validation. When\n// files were skipped during prepareReview, FullDiff holds the original diff\n// while Diff is filtered for the LLM prompt.\nfunc (pr *PRData) ValidationDiff() string {\n\tif pr.FullDiff != \"\" {\n\t\treturn pr.FullDiff\n\t}\n\treturn pr.Diff\n}\n\n// ghPRView is the JSON shape returned by gh pr view.\ntype ghPRView struct {\n\tTitle string `json:\"title\"`\n\tBody string `json:\"body\"`\n\tAuthor struct {\n\t\tLogin string `json:\"login\"`\n\t} `json:\"author\"`\n\tBaseRefName string `json:\"baseRefName\"`\n\tHeadRefName string `json:\"headRefName\"`\n\tFiles []struct {\n\t\tPath string `json:\"path\"`\n\t} `json:\"files\"`\n}\n\n// FetchPR gets PR metadata and diff using the gh CLI.\nfunc FetchPR(repo string, number int) (*PRData, error) {\n\tnumStr := fmt.Sprintf(\"%d\", number)\n\n\t// Fetch PR metadata as JSON.\n\tviewOut, err := runGH(\"gh pr view\", \"pr\", \"view\", numStr,\n\t\t\"--repo\", repo,\n\t\t\"--json\", \"title,body,author,baseRefName,headRefName,files\",\n\t)\n\tif err != nil {\n\t\treturn nil, err\n\t}\n\n\tvar view ghPRView\n\tif err := json.Unmarshal(viewOut, \u0026view); err != nil {\n\t\treturn nil, fmt.Errorf(\"parsing gh pr view output: %w\", err)\n\t}\n\n\t// Fetch the diff. This hits GitHub's pull request diff API\n\t// (GET /repos/{owner}/{repo}/pulls/{n} with Accept: application/vnd.github.v3.diff).\n\tdiffOut, err := runGH(\"gh pr diff\", \"pr\", \"diff\", numStr, \"--repo\", repo)\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"%w (GitHub pull request diff API; if this looks like a transient 5xx/timeout, retrying the job may help)\", err)\n\t}\n\n\tfiles := make([]string, len(view.Files))\n\tfor i, f := range view.Files {\n\t\tfiles[i] = f.Path\n\t}\n\n\treturn \u0026PRData{\n\t\tNumber: number,\n\t\tTitle: view.Title,\n\t\tBody: view.Body,\n\t\tAuthor: view.Author.Login,\n\t\tBaseBranch: view.BaseRefName,\n\t\tHeadBranch: view.HeadRefName,\n\t\tDiff: string(diffOut),\n\t\tFiles: files,\n\t}, nil\n}\n\n// PostComment posts a review comment on a PR.\nfunc PostComment(repo string, number int, body string) error {\n\tnumStr := fmt.Sprintf(\"%d\", number)\n\tcmd := exec.Command(\"gh\", \"pr\", \"comment\", numStr,\n\t\t\"--repo\", repo,\n\t\t\"--body\", body,\n\t)\n\tif out, err := cmd.CombinedOutput(); err != nil {\n\t\treturn fmt.Errorf(\"gh pr comment: %w\\n%s\", err, string(out))\n\t}\n\treturn nil\n}\n\n// reviewPayload is the JSON structure for the GitHub PR review API.\ntype reviewPayload struct {\n\tEvent string `json:\"event\"`\n\tBody string `json:\"body\"`\n\tComments []reviewComment `json:\"comments\"`\n\tCommitID string `json:\"commit_id,omitempty\"`\n}\n\n// reviewComment is a single inline comment in a PR review.\ntype reviewComment struct {\n\tPath string `json:\"path\"`\n\tLine int `json:\"line\"`\n\tBody string `json:\"body\"`\n}\n\n// countDiffLines counts the number of added and removed lines in a unified diff.\nfunc countDiffLines(diff string) (added, removed int) {\n\tfor _, line := range strings.Split(diff, \"\\n\") {\n\t\tif strings.HasPrefix(line, \"+++\") || strings.HasPrefix(line, \"---\") {\n\t\t\tcontinue\n\t\t}\n\t\tif strings.HasPrefix(line, \"+\") {\n\t\t\tadded++\n\t\t} else if strings.HasPrefix(line, \"-\") {\n\t\t\tremoved++\n\t\t}\n\t}\n\treturn\n}\n\n// diffLineMap maps each file to its sorted list of valid line numbers from the diff.\ntype diffLineMap map[string][]int\n\n// parseDiffLines extracts valid line numbers per file from a unified diff.\nfunc parseDiffLines(diff string) diffLineMap {\n\tvalid := make(diffLineMap)\n\tvar currentFile string\n\tvar lineNum int\n\n\tfor _, line := range strings.Split(diff, \"\\n\") {\n\t\tif strings.HasPrefix(line, \"+++ b/\") {\n\t\t\tcurrentFile = line[6:]\n\t\t\tcontinue\n\t\t}\n\t\tif strings.HasPrefix(line, \"@@ \") {\n\t\t\t// Parse hunk header: @@ -old,count +new,count @@\n\t\t\tif idx := strings.Index(line, \"+\"); idx \u003e= 0 {\n\t\t\t\trest := line[idx+1:]\n\t\t\t\tif comma := strings.IndexAny(rest, \", \"); comma \u003e= 0 {\n\t\t\t\t\trest = rest[:comma]\n\t\t\t\t}\n\t\t\t\t_, _ = fmt.Sscanf(rest, \"%d\", \u0026lineNum)\n\t\t\t}\n\t\t\tcontinue\n\t\t}\n\t\tif currentFile == \"\" || lineNum == 0 {\n\t\t\tcontinue\n\t\t}\n\t\tif strings.HasPrefix(line, \"-\") {\n\t\t\tcontinue\n\t\t}\n\t\tif strings.HasPrefix(line, \"+\") {\n\t\t\tvalid[currentFile] = append(valid[currentFile], lineNum)\n\t\t\tlineNum++\n\t\t\tcontinue\n\t\t}\n\t\tif strings.HasPrefix(line, \" \") {\n\t\t\tlineNum++\n\t\t\tcontinue\n\t\t}\n\t}\n\treturn valid\n}\n\n// nearestLine returns the closest valid diff line for a file:line pair.\n// Returns the line itself if valid, the nearest valid line in that file,\n// or 0 if the file is not in the diff at all.\nfunc (d diffLineMap) nearestLine(file string, line int) int {\n\tlines, ok := d[file]\n\tif !ok || len(lines) == 0 {\n\t\treturn 0\n\t}\n\tbest := lines[0]\n\tbestDist := abs(line - best)\n\tfor _, l := range lines[1:] {\n\t\tdist := abs(line - l)\n\t\tif dist \u003c bestDist {\n\t\t\tbest = l\n\t\t\tbestDist = dist\n\t\t}\n\t}\n\treturn best\n}\n\nfunc abs(x int) int {\n\tif x \u003c 0 {\n\t\treturn -x\n\t}\n\treturn x\n}\n\n// PostReview posts a PR review with inline comments using the GitHub API.\n// Findings with file and line information become inline comments; others are\n// included in the review body. The summary block is appended to the body so\n// the status dashboard appears on every CodeCanary top-level review.\nfunc PostReview(repo string, prNumber int, result *ReviewResult, diff string, commitSHA string, summary ReviewSummary) error {\n\t// Sort findings by severity before formatting.\n\tsortFindings(result.Findings)\n\n\t// Parse the diff to find valid line positions for inline comments.\n\tvalidLines := parseDiffLines(diff)\n\n\t// A finding can be inlined if its file is in the diff and the nearest\n\t// valid line is within a reasonable distance. Without a bound, findings\n\t// about code far from the diff get silently snapped to unrelated lines.\n\tcanInline := func(f Finding) bool {\n\t\tif f.File == \"\" || f.Line \u003c= 0 {\n\t\t\treturn false\n\t\t}\n\t\tnearest := validLines.nearestLine(f.File, f.Line)\n\t\treturn nearest \u003e 0 \u0026\u0026 abs(f.Line-nearest) \u003c= MaxFindingProximity\n\t}\n\n\tcomments := make([]reviewComment, 0)\n\tfor _, f := range result.Findings {\n\t\tif canInline(f) {\n\t\t\tcomments = append(comments, reviewComment{\n\t\t\t\tPath: f.File,\n\t\t\t\tLine: validLines.nearestLine(f.File, f.Line),\n\t\t\t\tBody: FormatFindingComment(\u0026f),\n\t\t\t})\n\t\t}\n\t}\n\n\tbody := withSummary(FormatReviewBody(result, canInline), summary)\n\n\tpayload := reviewPayload{\n\t\tEvent: \"COMMENT\",\n\t\tBody: body,\n\t\tComments: comments,\n\t\tCommitID: commitSHA,\n\t}\n\n\tpayloadJSON, err := json.Marshal(payload)\n\tif err != nil {\n\t\treturn fmt.Errorf(\"marshaling review payload: %w\", err)\n\t}\n\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn err\n\t}\n\n\t// No fallback — if this fails, the apiError carries stderr and response\n\t// body so the caller surfaces full diagnostics for debugging.\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, name, prNumber)\n\t_, err = ghAPIPOST(apiPath, payloadJSON)\n\treturn err\n}\n\n// FetchReviewFromPR extracts cached review data from a PR review's hidden HTML tag.\nfunc FetchReviewFromPR(repo string, prNumber int) (*ReviewResult, error) {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn nil, err\n\t}\n\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, name, prNumber)\n\tout, err := exec.Command(\"gh\", \"api\", apiPath).Output()\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"fetching PR reviews: %w\", err)\n\t}\n\n\tvar reviews []ghReview\n\tif err := json.Unmarshal(out, \u0026reviews); err != nil {\n\t\treturn nil, fmt.Errorf(\"parsing PR reviews: %w\", err)\n\t}\n\n\tprefixes := reviewMarkerPrefixes\n\tconst suffix = \" --\u003e\"\n\n\t// Search from most recent to oldest.\n\tfor i := len(reviews) - 1; i \u003e= 0; i-- {\n\t\tbody := reviews[i].Body\n\t\tfor _, prefix := range prefixes {\n\t\t\tidx := strings.Index(body, prefix)\n\t\t\tif idx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tstart := idx + len(prefix)\n\t\t\tendIdx := strings.Index(body[start:], suffix)\n\t\t\tif endIdx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tjsonData := body[start : start+endIdx]\n\t\t\tvar result ReviewResult\n\t\t\tif err := json.Unmarshal([]byte(jsonData), \u0026result); err != nil {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\treturn \u0026result, nil\n\t\t}\n\t}\n\n\treturn nil, fmt.Errorf(\"no review data found in PR #%d reviews\", prNumber)\n}\n\n// FetchFindingFromPR searches all reviews on a PR for a specific fix_ref.\n// Unlike FetchReviewFromPR (which returns the latest review), this searches every\n// review so that fix_ref links from older review rounds still resolve correctly.\nfunc FetchFindingFromPR(repo string, prNumber int, fixRef string) (*Finding, error) {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn nil, err\n\t}\n\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, name, prNumber)\n\tout, err := exec.Command(\"gh\", \"api\", apiPath).Output()\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"fetching PR reviews: %w\", err)\n\t}\n\n\tvar reviews []ghReview\n\tif err := json.Unmarshal(out, \u0026reviews); err != nil {\n\t\treturn nil, fmt.Errorf(\"parsing PR reviews: %w\", err)\n\t}\n\n\tprefixes := reviewMarkerPrefixes\n\tconst suffix = \" --\u003e\"\n\n\tfor _, rev := range reviews {\n\t\tfor _, prefix := range prefixes {\n\t\t\tidx := strings.Index(rev.Body, prefix)\n\t\t\tif idx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tstart := idx + len(prefix)\n\t\t\tendIdx := strings.Index(rev.Body[start:], suffix)\n\t\t\tif endIdx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tvar result ReviewResult\n\t\t\tif err := json.Unmarshal([]byte(rev.Body[start:start+endIdx]), \u0026result); err != nil {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tfor i := range result.Findings {\n\t\t\t\tif result.Findings[i].FixRef == fixRef {\n\t\t\t\t\treturn \u0026result.Findings[i], nil\n\t\t\t\t}\n\t\t\t}\n\t\t}\n\t}\n\n\treturn nil, fmt.Errorf(\"fix_ref %q not found in any review on PR #%d\", fixRef, prNumber)\n}\n\n// DetectRepo gets owner/name from the current git remote.\nfunc DetectRepo() (string, error) {\n\tout, err := exec.Command(\"gh\", \"repo\", \"view\",\n\t\t\"--json\", \"nameWithOwner\",\n\t\t\"--jq\", \".nameWithOwner\",\n\t).Output()\n\tif err != nil {\n\t\treturn \"\", fmt.Errorf(\"gh repo view: %w\", err)\n\t}\n\treturn strings.TrimSpace(string(out)), nil\n}\n\n// DetectPRNumber detects the PR number for the current branch using gh.\n// If repo is non-empty, it is passed as --repo to scope the lookup.\nfunc DetectPRNumber(repo string) (int, error) {\n\targs := []string{\"pr\", \"view\", \"--json\", \"number\", \"--jq\", \".number\"}\n\tif repo != \"\" {\n\t\targs = append(args, \"--repo\", repo)\n\t}\n\tout, err := exec.Command(\"gh\", args...).Output()\n\tif err != nil {\n\t\treturn 0, fmt.Errorf(\"no open pull request found for the current branch\")\n\t}\n\tnum, err := strconv.Atoi(strings.TrimSpace(string(out)))\n\tif err != nil {\n\t\treturn 0, fmt.Errorf(\"unexpected PR number from gh: %w\", err)\n\t}\n\treturn num, nil\n}\n\n// ThreadReply represents a reply to a review thread (i.e. any comment after the first).\ntype ThreadReply struct {\n\tAuthor string\n\tBody string\n}\n\n// ReviewThread represents a review thread from a PR.\ntype ReviewThread struct {\n\tID string\n\tPath string\n\tLine int\n\tBody string\n\tAuthor string // login of the first comment author (the bot for review threads)\n\tOutdated bool // true if GitHub marked the comment position as outdated (code changed)\n\tResolved bool\n\tReplies []ThreadReply\n}\n\n// graphQLThreadsResponse is the JSON shape returned by the review threads query.\ntype graphQLThreadsResponse struct {\n\tData struct {\n\t\tRepository struct {\n\t\t\tPullRequest struct {\n\t\t\t\tReviewThreads struct {\n\t\t\t\t\tNodes []struct {\n\t\t\t\t\t\tID string `json:\"id\"`\n\t\t\t\t\t\tIsResolved bool `json:\"isResolved\"`\n\t\t\t\t\t\tComments struct {\n\t\t\t\t\t\t\tNodes []struct {\n\t\t\t\t\t\t\t\tBody string `json:\"body\"`\n\t\t\t\t\t\t\t\tPath string `json:\"path\"`\n\t\t\t\t\t\t\t\tLine int `json:\"line\"`\n\t\t\t\t\t\t\t\tOriginalLine int `json:\"originalLine\"`\n\t\t\t\t\t\t\t\tOutdated bool `json:\"outdated\"`\n\t\t\t\t\t\t\t\tAuthor struct {\n\t\t\t\t\t\t\t\t\tLogin string `json:\"login\"`\n\t\t\t\t\t\t\t\t} `json:\"author\"`\n\t\t\t\t\t\t\t} `json:\"nodes\"`\n\t\t\t\t\t\t} `json:\"comments\"`\n\t\t\t\t\t} `json:\"nodes\"`\n\t\t\t\t} `json:\"reviewThreads\"`\n\t\t\t} `json:\"pullRequest\"`\n\t\t} `json:\"repository\"`\n\t} `json:\"data\"`\n}\n\n// FetchReviewThreads gets all review threads from a PR via GraphQL.\nfunc FetchReviewThreads(repo string, prNumber int) ([]ReviewThread, error) {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn nil, err\n\t}\n\n\tquery := `query($owner:String!,$name:String!,$pr:Int!){\n repository(owner:$owner,name:$name){\n pullRequest(number:$pr){\n reviewThreads(first:100){\n nodes{\n id\n isResolved\n comments(first:100){\n nodes{body path line originalLine outdated author{login}}\n }\n }\n }\n }\n }\n}`\n\n\tcmd := exec.Command(\"gh\", \"api\", \"graphql\",\n\t\t\"-f\", \"query=\"+query,\n\t\t\"-f\", fmt.Sprintf(\"owner=%s\", owner),\n\t\t\"-f\", fmt.Sprintf(\"name=%s\", name),\n\t\t\"-F\", fmt.Sprintf(\"pr=%d\", prNumber),\n\t)\n\tout, err := cmd.Output()\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"gh api graphql: %w\", err)\n\t}\n\n\tvar resp graphQLThreadsResponse\n\tif err := json.Unmarshal(out, \u0026resp); err != nil {\n\t\treturn nil, fmt.Errorf(\"parsing graphql response: %w\", err)\n\t}\n\n\tvar threads []ReviewThread\n\tfor _, node := range resp.Data.Repository.PullRequest.ReviewThreads.Nodes {\n\t\tif len(node.Comments.Nodes) == 0 {\n\t\t\tcontinue\n\t\t}\n\t\tcomment := node.Comments.Nodes[0]\n\n\t\t// Filter to review threads only (new marker + legacy markers for backward compat).\n\t\tif !strings.Contains(comment.Body, findingMarkerPrefix) \u0026\u0026\n\t\t\t!strings.Contains(comment.Body, \"codecanary fix\") \u0026\u0026\n\t\t\t!strings.Contains(comment.Body, \"clanopy fix\") {\n\t\t\tcontinue\n\t\t}\n\n\t\tvar replies []ThreadReply\n\t\tfor _, c := range node.Comments.Nodes[1:] {\n\t\t\treplies = append(replies, ThreadReply{\n\t\t\t\tAuthor: c.Author.Login,\n\t\t\t\tBody: c.Body,\n\t\t\t})\n\t\t}\n\n\t\tline := comment.Line\n\t\tif comment.Outdated \u0026\u0026 line == 0 \u0026\u0026 comment.OriginalLine \u003e 0 {\n\t\t\tline = comment.OriginalLine\n\t\t}\n\n\t\tthreads = append(threads, ReviewThread{\n\t\t\tID: node.ID,\n\t\t\tPath: comment.Path,\n\t\t\tLine: line,\n\t\t\tBody: comment.Body,\n\t\t\tAuthor: comment.Author.Login,\n\t\t\tOutdated: comment.Outdated,\n\t\t\tResolved: node.IsResolved,\n\t\t\tReplies: replies,\n\t\t})\n\t}\n\n\treturn threads, nil\n}\n\n// ResolveThread resolves a review thread via GraphQL mutation.\nfunc ResolveThread(threadID string) error {\n\tcmd := exec.Command(\"gh\", \"api\", \"graphql\",\n\t\t\"-f\", \"query=mutation($threadId:ID!){resolveReviewThread(input:{threadId:$threadId}){thread{isResolved}}}\",\n\t\t\"-f\", fmt.Sprintf(\"threadId=%s\", threadID),\n\t)\n\tif out, err := cmd.CombinedOutput(); err != nil {\n\t\treturn fmt.Errorf(\"gh api graphql resolve: %w\\n%s\", err, string(out))\n\t}\n\treturn nil\n}\n\n// ReplyToThread posts a reply on a review thread via GraphQL.\nfunc ReplyToThread(threadID, body string) error {\n\tcmd := exec.Command(\"gh\", \"api\", \"graphql\",\n\t\t\"-f\", \"query=mutation($threadId:ID!,$body:String!){addPullRequestReviewThreadReply(input:{pullRequestReviewThreadId:$threadId,body:$body}){comment{id}}}\",\n\t\t\"-f\", fmt.Sprintf(\"threadId=%s\", threadID),\n\t\t\"-f\", fmt.Sprintf(\"body=%s\", body),\n\t)\n\tif out, err := cmd.CombinedOutput(); err != nil {\n\t\treturn fmt.Errorf(\"gh api graphql reply: %w\\n%s\", err, string(out))\n\t}\n\treturn nil\n}\n\n// ghReview is the JSON shape for a PR review from the REST API.\ntype ghReview struct {\n\tID int64 `json:\"id\"`\n\tNodeID string `json:\"node_id\"`\n\tBody string `json:\"body\"`\n}\n\n// LatestCodecanaryReview captures the identity and body of the most recent\n// CodeCanary top-level review, together with the commit SHA embedded in its\n// hidden marker. Returned by FetchLatestCodecanaryReview for the\n// \"edit if same SHA, otherwise post new\" publish decision.\ntype LatestCodecanaryReview struct {\n\tID int64\n\tSHA string\n\tBody string\n}\n\n// FetchLatestCodecanaryReview returns the most recent top-level review on\n// the PR that carries a CodeCanary marker, along with the commit SHA\n// embedded in that marker. Returns (nil, nil) when no matching review\n// exists. A non-nil error is returned only for transport/parsing failures.\nfunc FetchLatestCodecanaryReview(repo string, prNumber int) (*LatestCodecanaryReview, error) {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn nil, err\n\t}\n\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, name, prNumber)\n\tout, err := exec.Command(\"gh\", \"api\", apiPath).Output()\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"fetching PR reviews: %w\", err)\n\t}\n\n\tvar reviews []ghReview\n\tif err := json.Unmarshal(out, \u0026reviews); err != nil {\n\t\treturn nil, fmt.Errorf(\"parsing PR reviews: %w\", err)\n\t}\n\n\t// Walk newest → oldest so we return the latest CodeCanary review.\n\tfor i := len(reviews) - 1; i \u003e= 0; i-- {\n\t\trev := reviews[i]\n\t\tfor _, prefix := range reviewMarkerPrefixes {\n\t\t\tidx := strings.Index(rev.Body, prefix)\n\t\t\tif idx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tstart := idx + len(prefix)\n\t\t\tendIdx := strings.Index(rev.Body[start:], reviewMarkerSuffix)\n\t\t\tif endIdx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tjsonData := rev.Body[start : start+endIdx]\n\t\t\tvar payload struct {\n\t\t\t\tSHA string `json:\"sha\"`\n\t\t\t}\n\t\t\t// Ignore unmarshal errors so old reviews that embed a richer\n\t\t\t// ReviewResult (also containing a \"sha\" field) still match.\n\t\t\t_ = json.Unmarshal([]byte(jsonData), \u0026payload)\n\t\t\treturn \u0026LatestCodecanaryReview{\n\t\t\t\tID: rev.ID,\n\t\t\t\tSHA: payload.SHA,\n\t\t\t\tBody: rev.Body,\n\t\t\t}, nil\n\t\t}\n\t}\n\n\treturn nil, nil\n}\n\n// UpdateReviewBody replaces the body of an existing PR review via the\n// GitHub REST API. Used when a reply-only run (or a duplicate synchronize\n// webhook on the same HEAD) needs to refresh the latest CodeCanary review's\n// counts instead of posting a new top-level comment.\nfunc UpdateReviewBody(repo string, prNumber int, reviewID int64, body string) error {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn err\n\t}\n\n\tpayload, err := json.Marshal(struct {\n\t\tBody string `json:\"body\"`\n\t}{Body: body})\n\tif err != nil {\n\t\treturn fmt.Errorf(\"marshaling update payload: %w\", err)\n\t}\n\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews/%d\", owner, name, prNumber, reviewID)\n\t_, err = ghAPIRequest(\"PUT\", apiPath, payload)\n\treturn err\n}\n\n// ghAPIRequest is a generic gh api wrapper that preserves the temp-file\n// pattern used by ghAPIPOST (avoids stdin pipe truncation on large bodies).\nfunc ghAPIRequest(method, apiPath string, payloadJSON []byte) ([]byte, error) {\n\tdir, err := os.MkdirTemp(\"\", \"codecanary-*\")\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"creating temp dir: %w\", err)\n\t}\n\tdefer func() { _ = os.RemoveAll(dir) }()\n\n\ttmpFile, err := os.CreateTemp(dir, \"payload.json\")\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"creating temp file: %w\", err)\n\t}\n\tif _, err := tmpFile.Write(payloadJSON); err != nil {\n\t\t_ = tmpFile.Close()\n\t\treturn nil, fmt.Errorf(\"writing payload to temp file: %w\", err)\n\t}\n\t_ = tmpFile.Close()\n\n\tcmd := exec.Command(\"gh\", \"api\", apiPath, \"--method\", method, \"--input\", tmpFile.Name())\n\tvar stdout, stderr bytes.Buffer\n\tcmd.Stdout = \u0026stdout\n\tcmd.Stderr = \u0026stderr\n\tif err := cmd.Run(); err != nil {\n\t\treturn stdout.Bytes(), \u0026apiError{Err: err, Stderr: stderr.String(), Response: stdout.String()}\n\t}\n\treturn stdout.Bytes(), nil\n}\n\n// FetchPreviousReviewSHA gets the SHA from the last review's hidden data.\nfunc FetchPreviousReviewSHA(repo string, prNumber int) string {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn \"\"\n\t}\n\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, name, prNumber)\n\tout, err := exec.Command(\"gh\", \"api\", apiPath).Output()\n\tif err != nil {\n\t\treturn \"\"\n\t}\n\n\tvar reviews []ghReview\n\tif err := json.Unmarshal(out, \u0026reviews); err != nil {\n\t\treturn \"\"\n\t}\n\n\tprefixes := reviewMarkerPrefixes\n\tconst suffix = \" --\u003e\"\n\n\t// Search from most recent to oldest.\n\tfor i := len(reviews) - 1; i \u003e= 0; i-- {\n\t\tbody := reviews[i].Body\n\t\tfor _, prefix := range prefixes {\n\t\t\tidx := strings.Index(body, prefix)\n\t\t\tif idx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tstart := idx + len(prefix)\n\t\t\tendIdx := strings.Index(body[start:], suffix)\n\t\t\tif endIdx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tjsonData := body[start : start+endIdx]\n\t\t\tvar result ReviewResult\n\t\t\tif err := json.Unmarshal([]byte(jsonData), \u0026result); err != nil {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tif result.SHA != \"\" {\n\t\t\t\treturn result.SHA\n\t\t\t}\n\t\t}\n\t}\n\n\treturn \"\"\n}\n\n// FilesFromDiff extracts the list of file paths touched in a unified diff.\nfunc FilesFromDiff(diff string) []string {\n\tvar files []string\n\tseen := make(map[string]bool)\n\tfor _, line := range strings.Split(diff, \"\\n\") {\n\t\tif strings.HasPrefix(line, \"+++ b/\") {\n\t\t\tpath := strings.TrimRight(line[6:], \"\\r\")\n\t\t\tif !seen[path] {\n\t\t\t\tseen[path] = true\n\t\t\t\tfiles = append(files, path)\n\t\t\t}\n\t\t}\n\t}\n\treturn files\n}\n\n// ScopeDiffToFiles filters a unified diff to only include hunks for files in\n// the allowed set. This prevents rebase noise (main-branch changes) from\n// leaking into incremental reviews.\nfunc ScopeDiffToFiles(diff string, allowedFiles map[string]bool) string {\n\tif len(allowedFiles) == 0 {\n\t\treturn diff\n\t}\n\n\tlines := strings.Split(diff, \"\\n\")\n\tvar result []string\n\tblockStart := -1\n\tblockAllowed := false\n\n\tfor i := 0; i \u003c len(lines); i++ {\n\t\tif strings.HasPrefix(lines[i], \"diff --git\") {\n\t\t\t// Flush previous block if allowed.\n\t\t\tif blockStart \u003e= 0 \u0026\u0026 blockAllowed {\n\t\t\t\tresult = append(result, lines[blockStart:i]...)\n\t\t\t}\n\t\t\tblockStart = i\n\t\t\tblockAllowed = false\n\n\t\t\t// Look ahead for +++ b/\u003cpath\u003e to determine if this block is allowed.\n\t\t\tfor j := i + 1; j \u003c len(lines) \u0026\u0026 !strings.HasPrefix(lines[j], \"diff --git\"); j++ {\n\t\t\t\tif strings.HasPrefix(lines[j], \"+++ b/\") {\n\t\t\t\t\tpath := strings.TrimRight(lines[j][6:], \"\\r\")\n\t\t\t\t\tif allowedFiles[path] {\n\t\t\t\t\t\tblockAllowed = true\n\t\t\t\t\t}\n\t\t\t\t\tbreak\n\t\t\t\t}\n\t\t\t}\n\t\t\tcontinue\n\t\t}\n\t}\n\n\t// Flush last block.\n\tif blockStart \u003e= 0 \u0026\u0026 blockAllowed {\n\t\tresult = append(result, lines[blockStart:]...)\n\t}\n\n\tif len(result) == 0 {\n\t\treturn \"\"\n\t}\n\treturn strings.Join(result, \"\\n\")\n}\n\n// validSHA matches a full-length lowercase hex Git SHA.\nvar validSHA = regexp.MustCompile(`^[0-9a-f]{40}$`)\n\n// GetIncrementalDiff gets the diff since a given SHA.\nfunc GetIncrementalDiff(baseSHA string) (string, error) {\n\tif !validSHA.MatchString(baseSHA) {\n\t\treturn \"\", fmt.Errorf(\"invalid SHA format: %q\", baseSHA)\n\t}\n\tout, err := exec.Command(\"git\", \"diff\", baseSHA+\"..HEAD\").Output()\n\tif err != nil {\n\t\treturn \"\", fmt.Errorf(\"git diff: %w\", err)\n\t}\n\treturn string(out), nil\n}\n\n// PostCleanReview posts a review when the first review finds no issues. The\n// commitSHA is embedded in a hidden marker so future runs treat it as the\n// baseline for incremental reviews, avoiding a redundant full re-review on the\n// next push.\nfunc PostCleanReview(repo string, prNumber int, commitSHA string, summary ReviewSummary) error {\n\treturn postSimpleReview(repo, prNumber, buildCleanReviewBody(commitSHA, summary))\n}\n\n// PostAllClearReview posts a review when all previous findings have been\n// resolved. If minimizeFailed is true, a note is appended warning about\n// visible old reviews. The commitSHA is embedded in a hidden marker so future\n// runs treat it as the baseline for incremental reviews; without it, the next\n// push would fall back to reviewing the entire PR again.\nfunc PostAllClearReview(repo string, prNumber int, commitSHA string, minimizeFailed bool, summary ReviewSummary) error {\n\treturn postSimpleReview(repo, prNumber, buildAllClearReviewBody(commitSHA, minimizeFailed, summary))\n}\n\n// PostActivityReview posts a review when no new findings were raised but\n// there is cycle activity worth surfacing (dismissals, acknowledgments,\n// rebuttals, still-open threads). This keeps every commit push producing a\n// visible top-level status comment instead of silently logging.\nfunc PostActivityReview(repo string, prNumber int, commitSHA string, summary ReviewSummary) error {\n\treturn postSimpleReview(repo, prNumber, buildActivityReviewBody(commitSHA, summary))\n}\n\n// buildCleanReviewBody renders the full Markdown body posted by\n// PostCleanReview. Split out from the poster so tests can assert the exact\n// string that lands on GitHub without having to mock gh.\nfunc buildCleanReviewBody(commitSHA string, summary ReviewSummary) string {\n\treturn withSummary(\"CodeCanary reviewed this PR \\u2014 no issues found.\", summary) + embedBaselineMarker(commitSHA)\n}\n\n// buildAllClearReviewBody renders the full Markdown body posted by\n// PostAllClearReview. Split out for the same reason as buildCleanReviewBody.\nfunc buildAllClearReviewBody(commitSHA string, minimizeFailed bool, summary ReviewSummary) string {\n\tbody := \"## \\U0001F425 CodeCanary\\n\\n\\u2705 All previous findings have been addressed. No new issues found. \\u2728\"\n\tif minimizeFailed {\n\t\tbody += \"\\n\\n\u003e \\u26A0\\uFE0F Some previous review comments could not be minimized and may still be visible.\"\n\t}\n\treturn withSummary(body, summary) + embedBaselineMarker(commitSHA)\n}\n\n// buildActivityReviewBody renders the body for a commit push that raised no\n// new findings but has cycle activity (dismissals/acknowledgments/rebuttals\n// or still-open threads carried forward).\nfunc buildActivityReviewBody(commitSHA string, summary ReviewSummary) string {\n\tbody := \"## \\U0001F425 CodeCanary\\n\\nReviewed this push \\u2014 no new issues found.\"\n\treturn withSummary(body, summary) + embedBaselineMarker(commitSHA)\n}\n\n// withSummary appends the status summary block to a review body. The block\n// is skipped when the summary has no non-zero counts, so existing clean/\n// all-clear bodies render identically when nothing happened.\nfunc withSummary(body string, summary ReviewSummary) string {\n\tblock := renderSummaryBlock(summary)\n\tif block == \"\" {\n\t\treturn body\n\t}\n\treturn body + block\n}\n\n// embedBaselineMarker returns a hidden HTML comment containing the commitSHA\n// so FetchPreviousReviewSHA can use this review as the incremental baseline.\n// Returns an empty string if commitSHA is empty (local mode, dry run). The\n// marker carries only the SHA — FetchPreviousReviewSHA is the sole reader and\n// it only needs that field.\nfunc embedBaselineMarker(commitSHA string) string {\n\tif commitSHA == \"\" {\n\t\treturn \"\"\n\t}\n\tdata, err := json.Marshal(struct {\n\t\tSHA string `json:\"sha\"`\n\t}{SHA: commitSHA})\n\tif err != nil {\n\t\treturn \"\"\n\t}\n\treturn fmt.Sprintf(\"\\n%s%s%s\\n\", reviewMarkerPrefixes[0], string(data), reviewMarkerSuffix)\n}\n\nfunc postSimpleReview(repo string, prNumber int, body string) error {\n\towner, repoName, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn err\n\t}\n\n\tpayload := reviewPayload{\n\t\tEvent: \"COMMENT\",\n\t\tBody: body,\n\t\tComments: make([]reviewComment, 0),\n\t}\n\n\tpayloadJSON, err := json.Marshal(payload)\n\tif err != nil {\n\t\treturn fmt.Errorf(\"marshaling payload: %w\", err)\n\t}\n\n\t_, err = ghAPIPOST(fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, repoName, prNumber), payloadJSON)\n\treturn err\n}\n\n// apiError is returned by ghAPIPOST so callers can inspect the stderr output\n// from gh (which contains the HTTP status line) separately from the response body.\ntype apiError struct {\n\tErr error\n\tStderr string\n\tResponse string\n}\n\nfunc (e *apiError) Error() string {\n\treturn fmt.Sprintf(\"gh api: %v\\nstderr: %s\\nresponse: %s\", e.Err, e.Stderr, e.Response)\n}\n\nfunc (e *apiError) Unwrap() error { return e.Err }\n\n// ghAPIPOST sends a JSON payload to the GitHub API via gh, using a temp file\n// to avoid stdin pipe issues that can cause \"unexpected end of JSON input\"\n// errors on large payloads.\nfunc ghAPIPOST(apiPath string, payloadJSON []byte) ([]byte, error) {\n\tdir, err := os.MkdirTemp(\"\", \"codecanary-*\")\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"creating temp dir: %w\", err)\n\t}\n\tdefer func() { _ = os.RemoveAll(dir) }()\n\n\ttmpFile, err := os.CreateTemp(dir, \"payload.json\")\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"creating temp file: %w\", err)\n\t}\n\n\tif _, err := tmpFile.Write(payloadJSON); err != nil {\n\t\t_ = tmpFile.Close()\n\t\treturn nil, fmt.Errorf(\"writing payload to temp file: %w\", err)\n\t}\n\t_ = tmpFile.Close()\n\n\tcmd := exec.Command(\"gh\", \"api\", apiPath, \"--method\", \"POST\", \"--input\", tmpFile.Name())\n\tvar stdout, stderr bytes.Buffer\n\tcmd.Stdout = \u0026stdout\n\tcmd.Stderr = \u0026stderr\n\tif err := cmd.Run(); err != nil {\n\t\treturn stdout.Bytes(), \u0026apiError{Err: err, Stderr: stderr.String(), Response: stdout.String()}\n\t}\n\treturn stdout.Bytes(), nil\n}\n\n// FindReviewNodeIDs returns the node_ids of all reviews on a PR.\nfunc FindReviewNodeIDs(repo string, prNumber int) ([]string, error) {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn nil, err\n\t}\n\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, name, prNumber)\n\tout, err := exec.Command(\"gh\", \"api\", apiPath).Output()\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"fetching PR reviews: %w\", err)\n\t}\n\n\tvar reviews []ghReview\n\tif err := json.Unmarshal(out, \u0026reviews); err != nil {\n\t\treturn nil, fmt.Errorf(\"parsing PR reviews: %w\", err)\n\t}\n\n\tprefixes := reviewMarkerPrefixes\n\n\tvar nodeIDs []string\n\tfor _, rev := range reviews {\n\t\tfor _, prefix := range prefixes {\n\t\t\tif strings.Contains(rev.Body, prefix) {\n\t\t\t\tnodeIDs = append(nodeIDs, rev.NodeID)\n\t\t\t\tbreak\n\t\t\t}\n\t\t}\n\t}\n\n\treturn nodeIDs, nil\n}\n\n// threadHeaderLine returns the first non-empty, non-HTML-comment line from\n// a thread body. This is the header line containing severity and finding ID.\nfunc threadHeaderLine(body string) string {\n\tfor _, line := range strings.SplitN(body, \"\\n\", 5) {\n\t\tline = strings.TrimSpace(line)\n\t\tif line == \"\" || strings.HasPrefix(line, \"\u003c!--\") {\n\t\t\tcontinue\n\t\t}\n\t\treturn line\n\t}\n\treturn \"\"\n}\n\n// FindingIDFromThread extracts the finding ID from a thread body.\n// Thread bodies follow the format: 🟠 **bug** — `finding-id`\nfunc FindingIDFromThread(body string) string {\n\tfirstLine := threadHeaderLine(body)\n\t// The separator is \" — `\" (space, em-dash U+2014, space, backtick).\n\tmarker := \" \\u2014 `\"\n\tstart := strings.Index(firstLine, marker)\n\tif start \u003c 0 {\n\t\treturn \"\"\n\t}\n\tstart += len(marker)\n\tend := strings.Index(firstLine[start:], \"`\")\n\tif end \u003c 0 {\n\t\treturn \"\"\n\t}\n\treturn firstLine[start : start+end]\n}\n\n// severityFromThreadBody extracts the severity string from a thread body.\n// Thread bodies follow the format: {icon} **severity** — `id`\nvar threadSeverityRe = regexp.MustCompile(`\\*\\*(\\w+)\\*\\*`)\n\nfunc severityFromThreadBody(body string) string {\n\tfirstLine := threadHeaderLine(body)\n\tif m := threadSeverityRe.FindStringSubmatch(firstLine); len(m) \u003e 1 {\n\t\treturn strings.ToLower(m[1])\n\t}\n\treturn \"warning\"\n}\n\n// findingFromEmbeddedJSON tries to parse a Finding from the JSON embedded in the\n// codecanary:finding HTML comment marker. Returns the finding and true if successful.\nfunc findingFromEmbeddedJSON(body string) (Finding, bool) {\n\tprefix := findingMarkerPrefix\n\tsuffix := reviewMarkerSuffix\n\tstart := strings.Index(body, prefix)\n\tif start \u003c 0 {\n\t\treturn Finding{}, false\n\t}\n\tstart += len(prefix)\n\tend := strings.Index(body[start:], suffix)\n\tif end \u003c 0 {\n\t\treturn Finding{}, false\n\t}\n\traw := body[start : start+end]\n\tif len(raw) == 0 || raw[0] != '{' {\n\t\treturn Finding{}, false\n\t}\n\tvar f Finding\n\tif err := json.Unmarshal([]byte(raw), \u0026f); err != nil {\n\t\treturn Finding{}, false\n\t}\n\treturn f, true\n}\n\n// parseThreadBody extracts the description and suggestion from a thread comment\n// body. This is the fallback parser for older comments that don't embed JSON.\nfunc parseThreadBody(body string) (description, suggestion string) {\n\tlines := strings.Split(body, \"\\n\")\n\n\t// Skip leading HTML markers and the header line (icon **sev** — `id`).\n\tcontentStart := 0\n\tpastHeader := false\n\tfor i, line := range lines {\n\t\ttrimmed := strings.TrimSpace(line)\n\t\tif trimmed == \"\" {\n\t\t\tcontinue\n\t\t}\n\t\tif strings.HasPrefix(trimmed, \"\u003c!--\") {\n\t\t\tcontinue\n\t\t}\n\t\tif !pastHeader {\n\t\t\t// First non-empty non-marker line is the header — skip it.\n\t\t\tpastHeader = true\n\t\t\tcontentStart = i + 1\n\t\t\tcontinue\n\t\t}\n\t\tcontentStart = i\n\t\tbreak\n\t}\n\n\tif contentStart \u003e= len(lines) {\n\t\treturn \"\", \"\"\n\t}\n\n\t// Split content into description and suggestion.\n\tvar descLines []string\n\tvar suggLines []string\n\tinSuggestion := false\n\tsuggestionPrefix := \"\u003e **Suggestion**: \"\n\n\tfor _, line := range lines[contentStart:] {\n\t\ttrimmed := strings.TrimSpace(line)\n\t\tif strings.HasPrefix(trimmed, suggestionPrefix) {\n\t\t\tinSuggestion = true\n\t\t\tsuggLines = append(suggLines, strings.TrimPrefix(trimmed, suggestionPrefix))\n\t\t\tcontinue\n\t\t}\n\t\tif inSuggestion {\n\t\t\t// Continuation of suggestion (blockquote lines).\n\t\t\tif strings.HasPrefix(trimmed, \"\u003e \") {\n\t\t\t\tsuggLines = append(suggLines, strings.TrimPrefix(trimmed, \"\u003e \"))\n\t\t\t} else if trimmed == \"\u003e\" {\n\t\t\t\tsuggLines = append(suggLines, \"\")\n\t\t\t} else {\n\t\t\t\tsuggLines = append(suggLines, line)\n\t\t\t}\n\t\t\tcontinue\n\t\t}\n\t\tdescLines = append(descLines, line)\n\t}\n\n\tdescription = strings.TrimSpace(strings.Join(descLines, \"\\n\"))\n\tsuggestion = strings.TrimSpace(strings.Join(suggLines, \"\\n\"))\n\treturn description, suggestion\n}\n\n// FindingFromThread extracts a Finding from a ReviewThread. It first tries to\n// parse the embedded JSON (new format), then falls back to body parsing.\nfunc FindingFromThread(t ReviewThread) Finding {\n\t// Try embedded JSON first (lossless roundtrip).\n\tif f, ok := findingFromEmbeddedJSON(t.Body); ok {\n\t\tf.File = t.Path\n\t\tf.Line = t.Line\n\t\tf.Status = \"still open\"\n\t\treturn f\n\t}\n\n\t// Fallback: parse the markdown body.\n\tdesc, suggestion := parseThreadBody(t.Body)\n\treturn Finding{\n\t\tID: FindingIDFromThread(t.Body),\n\t\tFile: t.Path,\n\t\tLine: t.Line,\n\t\tSeverity: severityFromThreadBody(t.Body),\n\t\tTitle: firstSentence(desc),\n\t\tDescription: desc,\n\t\tSuggestion: suggestion,\n\t\tStatus: \"still open\",\n\t}\n}\n\n// firstSentence returns the first sentence (or first line) of text as a title.\nfunc firstSentence(text string) string {\n\tif text == \"\" {\n\t\treturn \"\"\n\t}\n\t// Use first line.\n\tline := text\n\tif idx := strings.IndexByte(text, '\\n'); idx \u003e= 0 {\n\t\tline = text[:idx]\n\t}\n\tline = strings.TrimSpace(line)\n\t// Truncate to 120 runes if very long.\n\trunes := []rune(line)\n\tif len(runes) \u003e 120 {\n\t\treturn string(runes[:117]) + \"...\"\n\t}\n\treturn line\n}\n\n// ReviewInfo represents a review with its node ID and finding IDs.\ntype ReviewInfo struct {\n\tNodeID string\n\tFindingIDs []string\n}\n\n// FindReviews returns reviews with their parsed finding IDs.\nfunc FindReviews(repo string, prNumber int) ([]ReviewInfo, error) {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn nil, err\n\t}\n\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, name, prNumber)\n\tout, err := exec.Command(\"gh\", \"api\", apiPath).Output()\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"fetching PR reviews: %w\", err)\n\t}\n\n\tvar reviews []ghReview\n\tif err := json.Unmarshal(out, \u0026reviews); err != nil {\n\t\treturn nil, fmt.Errorf(\"parsing PR reviews: %w\", err)\n\t}\n\n\tprefixes := reviewMarkerPrefixes\n\tconst suffix = \" --\u003e\"\n\n\tvar result []ReviewInfo\n\tfor _, rev := range reviews {\n\t\tfor _, prefix := range prefixes {\n\t\t\tidx := strings.Index(rev.Body, prefix)\n\t\t\tif idx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tstart := idx + len(prefix)\n\t\t\tendIdx := strings.Index(rev.Body[start:], suffix)\n\t\t\tif endIdx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tjsonData := rev.Body[start : start+endIdx]\n\t\t\tvar rr ReviewResult\n\t\t\tif err := json.Unmarshal([]byte(jsonData), \u0026rr); err != nil {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tvar ids []string\n\t\t\tfor _, f := range rr.Findings {\n\t\t\t\tids = append(ids, f.ID)\n\t\t\t}\n\t\t\tresult = append(result, ReviewInfo{\n\t\t\t\tNodeID: rev.NodeID,\n\t\t\t\tFindingIDs: ids,\n\t\t\t})\n\t\t\tbreak\n\t\t}\n\t}\n\n\treturn result, nil\n}\n\n// MinimizeComment hides a comment on GitHub using the minimizeComment GraphQL mutation.\nfunc MinimizeComment(nodeID string) error {\n\tcmd := exec.Command(\"gh\", \"api\", \"graphql\",\n\t\t\"-f\", \"query=mutation($id:ID!){minimizeComment(input:{subjectId:$id,classifier:RESOLVED}){minimizedComment{isMinimized}}}\",\n\t\t\"-F\", \"id=\"+nodeID,\n\t)\n\tif out, err := cmd.CombinedOutput(); err != nil {\n\t\treturn fmt.Errorf(\"gh api graphql minimize: %w\\n%s\", err, string(out))\n\t}\n\treturn nil\n}\n\n// FetchFileContents reads the full contents of changed files from disk.\n// It skips files that are too large, binary, deleted, or match ignore patterns.\n// Returns a map of path-\u003econtent and a list of skipped file paths.\nfunc FetchFileContents(files []string, ignorePatterns []string, maxPerFile, maxTotal int) (map[string]string, []string) {\n\tcontents := make(map[string]string)\n\tvar skipped []string\n\ttotalSize := 0\n\n\tfor _, path := range files {\n\t\t// Check ignore patterns.\n\t\tif matchesIgnore(path, ignorePatterns) {\n\t\t\tskipped = append(skipped, path)\n\t\t\tcontinue\n\t\t}\n\n\t\tdata, err := os.ReadFile(path)\n\t\tif err != nil {\n\t\t\t// File may have been deleted in this PR — skip gracefully.\n\t\t\tcontinue\n\t\t}\n\n\t\t// Skip binary files (null bytes in first 512 bytes).\n\t\tpeek := data\n\t\tif len(peek) \u003e 512 {\n\t\t\tpeek = peek[:512]\n\t\t}\n\t\tif bytes.ContainsRune(peek, 0) {\n\t\t\tskipped = append(skipped, path)\n\t\t\tcontinue\n\t\t}\n\n\t\tsize := len(data)\n\n\t\t// Skip files exceeding per-file limit.\n\t\tif size \u003e maxPerFile {\n\t\t\tskipped = append(skipped, path)\n\t\t\tcontinue\n\t\t}\n\n\t\t// Stop if total budget would be exceeded.\n\t\tif totalSize+size \u003e maxTotal {\n\t\t\tskipped = append(skipped, path)\n\t\t\tcontinue\n\t\t}\n\n\t\tcontents[path] = string(data)\n\t\ttotalSize += size\n\t}\n\n\treturn contents, skipped\n}\n\n// isSetupPR detects whether this is the initial setup PR.\n// Returns true only when a new workflow file referencing codecanary is added AND\n// the PR contains no other files beyond expected setup artifacts (workflow +\n// config), so that PRs bundling real code changes are never silently skipped.\nfunc isSetupPR(diff string, files []string) bool {\n\t// All files must be known setup paths.\n\tfor _, f := range files {\n\t\tif !isSetupFile(f) {\n\t\t\treturn false\n\t\t}\n\t}\n\n\t// At least one newly added workflow file must reference codecanary.\n\tlines := strings.Split(diff, \"\\n\")\n\tfor i := 0; i \u003c len(lines)-1; i++ {\n\t\tif lines[i] != \"--- /dev/null\" {\n\t\t\tcontinue\n\t\t}\n\t\tplusLine := lines[i+1]\n\t\tif !strings.HasPrefix(plusLine, \"+++ b/.github/workflows/\") {\n\t\t\tcontinue\n\t\t}\n\t\tfor j := i + 2; j \u003c len(lines); j++ {\n\t\t\tif strings.HasPrefix(lines[j], \"--- \") || strings.HasPrefix(lines[j], \"diff --git\") {\n\t\t\t\tbreak\n\t\t\t}\n\t\t\tif strings.HasPrefix(lines[j], \"+\") \u0026\u0026 (strings.Contains(lines[j], \"codecanary\") || strings.Contains(lines[j], \"clanopy\")) {\n\t\t\t\treturn true\n\t\t\t}\n\t\t}\n\t}\n\treturn false\n}\n\n// isSetupFile returns true if the file path is a known setup artifact.\nfunc isSetupFile(path string) bool {\n\treturn strings.HasPrefix(path, \".github/workflows/\") ||\n\t\tstrings.HasPrefix(path, \".codecanary/\") || path == \".codecanary.yml\" ||\n\t\tstrings.HasPrefix(path, \".clanopy/\")\n}\n\n// matchesIgnore checks if a path matches any of the ignore glob patterns.\n// Uses doublestar to support ** recursive globs (e.g. \"dist/**\", \"src/**/*.test.*\").\nfunc matchesIgnore(path string, patterns []string) bool {\n\tfor _, pat := range patterns {\n\t\tif matched, _ := doublestar.Match(pat, path); matched {\n\t\t\treturn true\n\t\t}\n\t\t// Also try matching against just the filename.\n\t\tif matched, _ := doublestar.Match(pat, filepath.Base(path)); matched {\n\t\t\treturn true\n\t\t}\n\t}\n\treturn false\n}\n" + } + }, + "config": {}, + "project_docs": { + "CLAUDE.md": "# CodeCanary\n\nAI-powered code review for GitHub pull requests.\n\n## Project structure\n\n```\ncmd/\n review/ # Main binary — review CLI + setup wizard\n main.go # Entry point\n cli/ # Cobra commands\n root.go # Root \"codecanary\" command\n review.go # codecanary review \u003cpr\u003e\n findings.go # codecanary findings \u003cpr\u003e — fetch bot findings for the review skill\n reply.go # codecanary reply --url \u003ccomment-url\u003e --body \u003ctext\u003e — post a reply on a review thread\n install_skill.go # codecanary install-skill — write embedded Claude skill to disk\n setup.go # codecanary setup [local|github]\n auth.go # codecanary auth [status|delete]\ninternal/\n review/\n runner.go # Core review pipeline — single Run() entry point\n config.go # Config loading, validation, defaults\n # Provider layer (LLM abstraction)\n provider.go # ModelProvider interface + factory registry\n provider_anthropic.go\n provider_openai.go\n provider_openrouter.go\n provider_claude.go # Claude CLI wrapper\n provider_compat.go # Shared types for OpenAI-compatible APIs\n pricing.go # Token-based cost estimation\n # Platform layer (environment abstraction)\n platform.go # ReviewPlatform interface\n platform_github.go # GitHub Actions implementation\n platform_local.go # Local CLI implementation\n # Supporting modules\n prompt.go # Prompt building (review, incremental, per-thread)\n findings.go # Finding parsing, filtering, result structures\n triage.go # Thread classification + parallel LLM evaluation\n formatter.go # JSON/Markdown/Terminal output formatting\n usage.go # Token tracking, budget checking\n github.go # GitHub API calls (fetch threads, post reviews)\n comments.go # PR review comment fetch + finding marker parser + review-check watcher\n local.go # Local diff \u0026 git operations\n state.go # Local state persistence\n docs.go # Project doc discovery\n credentials/ # Credential storage (keychain with file fallback)\n keyring.go # Store/Retrieve/Delete — keychain first, ~/.codecanary/credentials.json fallback\n skills/ # Claude Code skills embedded in the binary via //go:embed\n skills.go # Exports CodecanaryFix() returning the skill body\n codecanary-fix/SKILL.md # Canonical skill source (duplicated at .claude/skills/codecanary-fix/SKILL.md; parity enforced by skills_test.go)\n setup/ # Setup wizard logic (huh forms)\n forms.go # Shared huh form components\n validate.go # API key validation via test calls\n guidance.go # Token/permissions guidance text\n workflow.go # GitHub Actions workflow template\n local.go # RunLocal() — local setup flow\n github.go # RunGitHub() — GitHub Actions setup flow\n auth/ # OAuth PKCE flow, GitHub App installation\ntelemetry/ # Telemetry domain (anonymous usage analytics)\n worker/ # Cloudflare Worker — telemetry ingestion (TypeScript)\n dashboard/ # Cloudflare Pages — internal analytics dashboard (vanilla JS + Chart.js)\noidc/ # OIDC domain\n worker/ # Cloudflare Worker — OIDC token exchange proxy (TypeScript)\naction.yml # GitHub Action definition (composite action)\ninstall.sh # Downloads and installs codecanary binary permanently\n.claude/\n skills/\n codecanary-fix/ # Claude Code skill — drives review→fix→push loop using `codecanary findings` + `codecanary reply`\n```\n\n## Binary\n\n- **`codecanary`** — single binary for reviews, setup, and credential management. Installed locally via `install.sh`, also used by the GitHub Action.\n\n## Build\n\n```sh\ngo build ./cmd/review # builds codecanary\n```\n\nVersion is set via ldflags: `-X main.version=v{version}`\n\n## Lint\n\n```sh\ngolangci-lint run ./...\n```\n\nAll code must pass `golangci-lint` with default linters (errcheck, staticcheck, etc.). Run this before committing.\n\n## Key dependencies\n\n- `spf13/cobra` — CLI framework\n- `charmbracelet/huh` — terminal form builder (setup wizard)\n- `zalando/go-keyring` — OS keychain (with file-based fallback for systems without one)\n- `bmatcuk/doublestar` — glob pattern matching for ignore rules\n- `gopkg.in/yaml.v3` — config parsing\n- `golang.org/x/term` — terminal detection\n\n## Architecture\n\n### Core principle: adapters keep the engine agnostic\n\nThe review engine (`runner.go`) is provider- and platform-agnostic. It depends only on two interfaces — never on concrete GitHub APIs, LLM SDKs, or environment-specific logic. All environment and provider specifics live behind adapters.\n\n### Provider layer — `ModelProvider` interface (`provider.go`)\n\nAbstracts LLM invocations. The core engine calls `provider.Run(ctx, prompt, opts)` and gets back text + usage metadata. It never knows which LLM backend is being used.\n\n**Implementations**: `anthropic`, `openai`, `openrouter`, `claude` (CLI).\n**Selection**: factory registry in `provider.go` — `NewProviderForRole(mc, env)` returns the right implementation based on `mc.Provider`.\n\nAdding a new LLM provider means: create `provider_\u003cname\u003e.go` and register a `ProviderFactory` (constructor, validation, pricing, default models) via `init()`.\n\n### Platform layer — `ReviewPlatform` interface (`platform.go`)\n\nAbstracts environment-specific operations: loading previous findings, publishing results, saving state, resolving threads, reporting usage.\n\n**Implementations**: `GithubPlatform` (posts to PRs, reads threads via API), `LocalPlatform` (prints to terminal, persists state to `.codecanary/state/`).\n\nAdding a new platform (e.g., GitLab) means: implement `ReviewPlatform`, wire it in the CLI.\n\n### Unified review pipeline (`runner.go`)\n\nThere is a **single `Run()` function** — not separate paths for GitHub vs. local. The pipeline is:\n\n1. Fetch PR data (or local diff)\n2. Load config, project docs, file contents\n3. Create providers via `NewProviderForRole()` (factory, provider-agnostic)\n4. Load previous findings via `platform.LoadPreviousFindings()`\n5. If incremental: triage threads, evaluate via provider, handle resolutions\n6. Build and execute main review prompt\n7. Parse findings, filter non-actionable\n8. `platform.Publish()` → `platform.SaveState()` → `platform.ReportUsage()`\n\n### Other architecture notes\n\n- **Config** is split across two files in `.codecanary/`: `config.yml` (provider, models, budgets, timeouts) and `review.yml` (rules, context, ignore patterns). `review.yml` is optional — if present, its fields override rules/context/ignore in `config.yml`. A personal `review.local.yml` can add rules, context, and ignore patterns on top of `review.yml` (append semantics, not replacement). Legacy `.codecanary.yml` at repo root is still supported with a deprecation warning.\n- **Incremental reviews**: on re-push, triage existing threads (Go-driven classifier in `triage.go`), evaluate changed threads via provider (triage model), then review only new code\n- **Dual marker detection**: reads both `codecanary:review` and legacy `clanopy:review` HTML markers for backward compatibility\n- **Anti-hallucination**: explicit file allowlist, line validation against diff, max finding distance threshold\n- **OIDC worker** (`oidc/worker/`): OIDC token exchange proxy at `oidc.codecanary.sh` — verifies GitHub Actions OIDC token, returns GitHub App installation token\n- **Telemetry worker** (`telemetry/worker/`): anonymous usage ingest at `telemetry.codecanary.sh` — writes to a Cloudflare Analytics Engine dataset\n- **Telemetry dashboard** (`telemetry/dashboard/`): internal analytics view at `dashboard.codecanary.sh` — Cloudflare Pages + Pages Functions, gated by Cloudflare Access. Reads the AE dataset via the SQL HTTP API with a read-only API token (no AE binding, so writes are platform-impossible)\n- **Setup** is a subcommand (`codecanary setup`) using `charmbracelet/huh` forms, with `local` and `github` sub-flows\n- **Credentials** use a single env var `CODECANARY_PROVIDER_SECRET` for all providers. Stored via `go-keyring` (OS keychain) with a file-based fallback (`~/.codecanary/credentials.json`, mode `0600`). `resolveEnv()` in `runner.go` injects the stored credential into the filtered env when not already set.\n\n## Rules\n\n- **Keep the core engine agnostic.** `runner.go`, `triage.go`, `prompt.go`, `findings.go` must never import or reference a specific LLM provider or platform. All provider/platform specifics go behind the `ModelProvider` or `ReviewPlatform` interfaces. No `if provider == \"openai\"` in core logic.\n- **Use the adapter/provider pattern for new integrations.** New LLM backends → create `provider_\u003cname\u003e.go` with a `ProviderFactory` registration in `init()`. New deployment targets → implement `ReviewPlatform` + wire in CLI. Never fork the pipeline.\n- **One pipeline, not two.** There must be a single `Run()` path. GitHub and local modes differ only in which `ReviewPlatform` implementation is injected — the orchestration logic is shared.\n- **Shared types for similar providers.** OpenAI-compatible APIs share request/response types via `provider_compat.go`. Don't duplicate HTTP client logic across providers.\n- **Don't repeat yourself.** Before writing new code, search the codebase for existing functions, mappings, or logic that already does what you need — then call it instead of reimplementing it. This applies to everything: switch statements, helper functions, validation logic, data mappings, HTTP calls. One source of truth, callers import it. Don't merge scaffolding or unused exports — if it's not called yet, it doesn't ship yet.\n- **File names are ownership boundaries.** A function defined in `local.go` implies it belongs to the local flow; one in `github.go` implies it belongs to GitHub. If a function is called by multiple files in the same package, it belongs in a shared file (e.g., `forms.go` for setup helpers, `platform.go` for platform-shared logic). Never define shared infrastructure in a flow-specific file — move it to the file that matches its actual scope.\n- **Canonical provider registration points.** Provider names live in the factory map in `provider.go`. All providers use `CODECANARY_PROVIDER_SECRET` for credentials (defined in `internal/credentials/keyring.go`). When adding a new provider, register a `ProviderFactory` in `provider.go` and add config validation in `config.go`.\n- **Minimize shell code.** `install.sh` and the GitHub Action (`action.yml`) should be kept as thin as possible. All logic must live in Go.\n- **Workflow template is embedded.** `internal/setup/codecanary.yml` is the single source of truth for the GitHub Actions workflow, embedded via `//go:embed`. `.github/workflows/codecanary.yml` must be identical — `go test ./internal/setup/` enforces this. When changing the workflow, edit either file and copy to the other.\n- **Claude skills are embedded.** `internal/skills/codecanary-fix/SKILL.md` is the single source of truth for the codecanary-fix skill, embedded via `//go:embed` and materialized by `codecanary install-skill`. `.claude/skills/codecanary-fix/SKILL.md` must be identical so Claude Code's project-mode discovery finds it when working in this repo — `go test ./internal/skills/` enforces this. When changing the skill, edit either file and copy to the other.\n- **Keep `docs/review-flow.md` in sync.** This document describes the full review pipeline — every step, the triage flow, platform differences, and key design decisions. When changing `runner.go`, `triage.go`, `prompt.go`, `findings.go`, `github.go`, `local.go`, `platform.go`, or the `ReviewPlatform` implementations, update the doc to reflect the new behavior.\n- Tests exist for config, findings, formatting, and triage. Be careful with refactors — run `go test ./...` and `go vet ./...`.\n" + } +} diff --git a/internal/review/testdata/corpus/alansikora-codecanary-pr165.prompt.golden b/internal/review/testdata/corpus/alansikora-codecanary-pr165.prompt.golden new file mode 100644 index 0000000..4ed3bb1 --- /dev/null +++ b/internal/review/testdata/corpus/alansikora-codecanary-pr165.prompt.golden @@ -0,0 +1,1684 @@ +You are a code reviewer. Review the following pull request and report findings. +You will be given the full contents of changed files for context, along with the diff. Only report issues that are directly related to the changes in the diff — do not flag pre-existing issues in unchanged code. Do not report a finding if your analysis concludes that the code is correct and no action is needed — only report findings that require the author to make a change or consider a specific alternative. +Also consider whether the changes could cause side effects in other files that depend on or interact with the modified code (e.g. callers, importers, shared state). If you identify a potential side effect, anchor your finding to the relevant line in the diff and describe the affected downstream code in the description. + +## Pull Request #165 +fix: surface gh stderr when FetchPR fails +**Author:** alansikora + +## Summary +- `gh pr view` / `gh pr diff` errors in `FetchPR` were wrapped with only the exit status, so Actions logs showed `gh pr diff: exit status 1` with no detail (stderr from `.Output()` was dropped). +- Added a small `runGH` helper that pipes stderr into the wrapped error, so the real gh message (rate limit, 5xx, token scope, etc.) shows up in logs. +- For `gh pr diff` specifically, append a hint pointing at GitHub's pull request diff API (`GET /repos/{owner}/{repo}/pulls/{n}` with `Accept: application/vnd.github.v3.diff`), since transient failures there are the usual cause and a job retry typically resolves them. + +Motivated by [this failing run](https://github.com/thetechfx/bedrock/actions/runs/24791879406/job/72551238868?pr=1581): the PR was small and mergeable, neighboring codecanary runs succeeded the same minute, and no stderr was visible to confirm it was just a transient API blip. + +## Test plan +- [ ] `go build ./...` and `go vet ./...` pass (done locally). +- [ ] On the next real `gh pr diff` failure in Actions, verify the error line now includes gh's stderr. + +🤖 Generated with [Claude Code](https://claude.com/claude-code) + + +## Project Documentation +The following project documentation describes conventions and standards for this codebase. Use these to inform your review — flag violations of these conventions when relevant. + + +# CodeCanary + +AI-powered code review for GitHub pull requests. + +## Project structure + +``` +cmd/ + review/ # Main binary — review CLI + setup wizard + main.go # Entry point + cli/ # Cobra commands + root.go # Root "codecanary" command + review.go # codecanary review + findings.go # codecanary findings — fetch bot findings for the review skill + reply.go # codecanary reply --url --body — post a reply on a review thread + install_skill.go # codecanary install-skill — write embedded Claude skill to disk + setup.go # codecanary setup [local|github] + auth.go # codecanary auth [status|delete] +internal/ + review/ + runner.go # Core review pipeline — single Run() entry point + config.go # Config loading, validation, defaults + # Provider layer (LLM abstraction) + provider.go # ModelProvider interface + factory registry + provider_anthropic.go + provider_openai.go + provider_openrouter.go + provider_claude.go # Claude CLI wrapper + provider_compat.go # Shared types for OpenAI-compatible APIs + pricing.go # Token-based cost estimation + # Platform layer (environment abstraction) + platform.go # ReviewPlatform interface + platform_github.go # GitHub Actions implementation + platform_local.go # Local CLI implementation + # Supporting modules + prompt.go # Prompt building (review, incremental, per-thread) + findings.go # Finding parsing, filtering, result structures + triage.go # Thread classification + parallel LLM evaluation + formatter.go # JSON/Markdown/Terminal output formatting + usage.go # Token tracking, budget checking + github.go # GitHub API calls (fetch threads, post reviews) + comments.go # PR review comment fetch + finding marker parser + review-check watcher + local.go # Local diff & git operations + state.go # Local state persistence + docs.go # Project doc discovery + credentials/ # Credential storage (keychain with file fallback) + keyring.go # Store/Retrieve/Delete — keychain first, ~/.codecanary/credentials.json fallback + skills/ # Claude Code skills embedded in the binary via //go:embed + skills.go # Exports CodecanaryFix() returning the skill body + codecanary-fix/SKILL.md # Canonical skill source (duplicated at .claude/skills/codecanary-fix/SKILL.md; parity enforced by skills_test.go) + setup/ # Setup wizard logic (huh forms) + forms.go # Shared huh form components + validate.go # API key validation via test calls + guidance.go # Token/permissions guidance text + workflow.go # GitHub Actions workflow template + local.go # RunLocal() — local setup flow + github.go # RunGitHub() — GitHub Actions setup flow + auth/ # OAuth PKCE flow, GitHub App installation +telemetry/ # Telemetry domain (anonymous usage analytics) + worker/ # Cloudflare Worker — telemetry ingestion (TypeScript) + dashboard/ # Cloudflare Pages — internal analytics dashboard (vanilla JS + Chart.js) +oidc/ # OIDC domain + worker/ # Cloudflare Worker — OIDC token exchange proxy (TypeScript) +action.yml # GitHub Action definition (composite action) +install.sh # Downloads and installs codecanary binary permanently +.claude/ + skills/ + codecanary-fix/ # Claude Code skill — drives review→fix→push loop using `codecanary findings` + `codecanary reply` +``` + +## Binary + +- **`codecanary`** — single binary for reviews, setup, and credential management. Installed locally via `install.sh`, also used by the GitHub Action. + +## Build + +```sh +go build ./cmd/review # builds codecanary +``` + +Version is set via ldflags: `-X main.version=v{version}` + +## Lint + +```sh +golangci-lint run ./... +``` + +All code must pass `golangci-lint` with default linters (errcheck, staticcheck, etc.). Run this before committing. + +## Key dependencies + +- `spf13/cobra` — CLI framework +- `charmbracelet/huh` — terminal form builder (setup wizard) +- `zalando/go-keyring` — OS keychain (with file-based fallback for systems without one) +- `bmatcuk/doublestar` — glob pattern matching for ignore rules +- `gopkg.in/yaml.v3` — config parsing +- `golang.org/x/term` — terminal detection + +## Architecture + +### Core principle: adapters keep the engine agnostic + +The review engine (`runner.go`) is provider- and platform-agnostic. It depends only on two interfaces — never on concrete GitHub APIs, LLM SDKs, or environment-specific logic. All environment and provider specifics live behind adapters. + +### Provider layer — `ModelProvider` interface (`provider.go`) + +Abstracts LLM invocations. The core engine calls `provider.Run(ctx, prompt, opts)` and gets back text + usage metadata. It never knows which LLM backend is being used. + +**Implementations**: `anthropic`, `openai`, `openrouter`, `claude` (CLI). +**Selection**: factory registry in `provider.go` — `NewProviderForRole(mc, env)` returns the right implementation based on `mc.Provider`. + +Adding a new LLM provider means: create `provider_.go` and register a `ProviderFactory` (constructor, validation, pricing, default models) via `init()`. + +### Platform layer — `ReviewPlatform` interface (`platform.go`) + +Abstracts environment-specific operations: loading previous findings, publishing results, saving state, resolving threads, reporting usage. + +**Implementations**: `GithubPlatform` (posts to PRs, reads threads via API), `LocalPlatform` (prints to terminal, persists state to `.codecanary/state/`). + +Adding a new platform (e.g., GitLab) means: implement `ReviewPlatform`, wire it in the CLI. + +### Unified review pipeline (`runner.go`) + +There is a **single `Run()` function** — not separate paths for GitHub vs. local. The pipeline is: + +1. Fetch PR data (or local diff) +2. Load config, project docs, file contents +3. Create providers via `NewProviderForRole()` (factory, provider-agnostic) +4. Load previous findings via `platform.LoadPreviousFindings()` +5. If incremental: triage threads, evaluate via provider, handle resolutions +6. Build and execute main review prompt +7. Parse findings, filter non-actionable +8. `platform.Publish()` → `platform.SaveState()` → `platform.ReportUsage()` + +### Other architecture notes + +- **Config** is split across two files in `.codecanary/`: `config.yml` (provider, models, budgets, timeouts) and `review.yml` (rules, context, ignore patterns). `review.yml` is optional — if present, its fields override rules/context/ignore in `config.yml`. A personal `review.local.yml` can add rules, context, and ignore patterns on top of `review.yml` (append semantics, not replacement). Legacy `.codecanary.yml` at repo root is still supported with a deprecation warning. +- **Incremental reviews**: on re-push, triage existing threads (Go-driven classifier in `triage.go`), evaluate changed threads via provider (triage model), then review only new code +- **Dual marker detection**: reads both `codecanary:review` and legacy `clanopy:review` HTML markers for backward compatibility +- **Anti-hallucination**: explicit file allowlist, line validation against diff, max finding distance threshold +- **OIDC worker** (`oidc/worker/`): OIDC token exchange proxy at `oidc.codecanary.sh` — verifies GitHub Actions OIDC token, returns GitHub App installation token +- **Telemetry worker** (`telemetry/worker/`): anonymous usage ingest at `telemetry.codecanary.sh` — writes to a Cloudflare Analytics Engine dataset +- **Telemetry dashboard** (`telemetry/dashboard/`): internal analytics view at `dashboard.codecanary.sh` — Cloudflare Pages + Pages Functions, gated by Cloudflare Access. Reads the AE dataset via the SQL HTTP API with a read-only API token (no AE binding, so writes are platform-impossible) +- **Setup** is a subcommand (`codecanary setup`) using `charmbracelet/huh` forms, with `local` and `github` sub-flows +- **Credentials** use a single env var `CODECANARY_PROVIDER_SECRET` for all providers. Stored via `go-keyring` (OS keychain) with a file-based fallback (`~/.codecanary/credentials.json`, mode `0600`). `resolveEnv()` in `runner.go` injects the stored credential into the filtered env when not already set. + +## Rules + +- **Keep the core engine agnostic.** `runner.go`, `triage.go`, `prompt.go`, `findings.go` must never import or reference a specific LLM provider or platform. All provider/platform specifics go behind the `ModelProvider` or `ReviewPlatform` interfaces. No `if provider == "openai"` in core logic. +- **Use the adapter/provider pattern for new integrations.** New LLM backends → create `provider_.go` with a `ProviderFactory` registration in `init()`. New deployment targets → implement `ReviewPlatform` + wire in CLI. Never fork the pipeline. +- **One pipeline, not two.** There must be a single `Run()` path. GitHub and local modes differ only in which `ReviewPlatform` implementation is injected — the orchestration logic is shared. +- **Shared types for similar providers.** OpenAI-compatible APIs share request/response types via `provider_compat.go`. Don't duplicate HTTP client logic across providers. +- **Don't repeat yourself.** Before writing new code, search the codebase for existing functions, mappings, or logic that already does what you need — then call it instead of reimplementing it. This applies to everything: switch statements, helper functions, validation logic, data mappings, HTTP calls. One source of truth, callers import it. Don't merge scaffolding or unused exports — if it's not called yet, it doesn't ship yet. +- **File names are ownership boundaries.** A function defined in `local.go` implies it belongs to the local flow; one in `github.go` implies it belongs to GitHub. If a function is called by multiple files in the same package, it belongs in a shared file (e.g., `forms.go` for setup helpers, `platform.go` for platform-shared logic). Never define shared infrastructure in a flow-specific file — move it to the file that matches its actual scope. +- **Canonical provider registration points.** Provider names live in the factory map in `provider.go`. All providers use `CODECANARY_PROVIDER_SECRET` for credentials (defined in `internal/credentials/keyring.go`). When adding a new provider, register a `ProviderFactory` in `provider.go` and add config validation in `config.go`. +- **Minimize shell code.** `install.sh` and the GitHub Action (`action.yml`) should be kept as thin as possible. All logic must live in Go. +- **Workflow template is embedded.** `internal/setup/codecanary.yml` is the single source of truth for the GitHub Actions workflow, embedded via `//go:embed`. `.github/workflows/codecanary.yml` must be identical — `go test ./internal/setup/` enforces this. When changing the workflow, edit either file and copy to the other. +- **Claude skills are embedded.** `internal/skills/codecanary-fix/SKILL.md` is the single source of truth for the codecanary-fix skill, embedded via `//go:embed` and materialized by `codecanary install-skill`. `.claude/skills/codecanary-fix/SKILL.md` must be identical so Claude Code's project-mode discovery finds it when working in this repo — `go test ./internal/skills/` enforces this. When changing the skill, edit either file and copy to the other. +- **Keep `docs/review-flow.md` in sync.** This document describes the full review pipeline — every step, the triage flow, platform differences, and key design decisions. When changing `runner.go`, `triage.go`, `prompt.go`, `findings.go`, `github.go`, `local.go`, `platform.go`, or the `ReviewPlatform` implementations, update the doc to reflect the new behavior. +- Tests exist for config, findings, formatting, and triage. Be careful with refactors — run `go test ./...` and `go vet ./...`. + + + +## Review Rules +No specific rules are defined. Perform a general code review covering correctness, security, performance, and maintainability. + +## Files in This Diff +The following files — and ONLY these files — are part of this diff. Every finding you report MUST reference one of these exact paths. Do NOT reference any file that is not in this list. + +- `internal/review/github.go` + +## Changed File Contents +Below are the full contents of changed files. Use these to understand surrounding code, types, imports, and control flow. Do NOT report findings on unchanged code — only flag issues directly related to changes in the diff. + +### `internal/review/github.go` +``` +1: package review +2: +3: import ( +4: "bytes" +5: "encoding/json" +6: "fmt" +7: "os" +8: "os/exec" +9: "path/filepath" +10: "regexp" +11: "strconv" +12: "strings" +13: +14: "github.com/bmatcuk/doublestar/v4" +15: ) +16: +17: // runGH runs a gh command and returns stdout; on failure, the returned error +18: // includes gh's stderr so upstream API messages surface in logs instead of just +19: // "exit status 1". +20: func runGH(label string, args ...string) ([]byte, error) { +21: cmd := exec.Command("gh", args...) +22: var stderr bytes.Buffer +23: cmd.Stderr = &stderr +24: out, err := cmd.Output() +25: if err != nil { +26: msg := strings.TrimSpace(stderr.String()) +27: if msg == "" { +28: return nil, fmt.Errorf("%s: %w", label, err) +29: } +30: return nil, fmt.Errorf("%s: %s: %w", label, msg, err) +31: } +32: return out, nil +33: } +34: +35: // parseRepoSlug splits a "owner/name" repository slug into its two parts. +36: func parseRepoSlug(repo string) (owner, name string, err error) { +37: parts := strings.SplitN(repo, "/", 2) +38: if len(parts) != 2 { +39: return "", "", fmt.Errorf("invalid repo format %q, expected owner/name", repo) +40: } +41: return parts[0], parts[1], nil +42: } +43: +44: // MaxFindingProximity is the maximum number of lines a finding may be from the +45: // nearest changed line in the PR diff. Findings beyond this distance are dropped +46: // (runner.go) or demoted from inline to body (PostReview). This enforces review +47: // scope — keeping findings anchored to the PR's actual changes — and catches +48: // hallucinated line numbers. A single constant ensures both checks stay in sync. +49: const MaxFindingProximity = 20 +50: +51: // HTML comment markers for embedding and detecting review data. +52: // Dual prefixes support both current (codecanary) and legacy (clanopy) markers. +53: var reviewMarkerPrefixes = []string{"" +57: findingMarkerPrefix = "" +335: +336: // Search from most recent to oldest. +337: for i := len(reviews) - 1; i >= 0; i-- { +338: body := reviews[i].Body +339: for _, prefix := range prefixes { +340: idx := strings.Index(body, prefix) +341: if idx < 0 { +342: continue +343: } +344: start := idx + len(prefix) +345: endIdx := strings.Index(body[start:], suffix) +346: if endIdx < 0 { +347: continue +348: } +349: jsonData := body[start : start+endIdx] +350: var result ReviewResult +351: if err := json.Unmarshal([]byte(jsonData), &result); err != nil { +352: continue +353: } +354: return &result, nil +355: } +356: } +357: +358: return nil, fmt.Errorf("no review data found in PR #%d reviews", prNumber) +359: } +360: +361: // FetchFindingFromPR searches all reviews on a PR for a specific fix_ref. +362: // Unlike FetchReviewFromPR (which returns the latest review), this searches every +363: // review so that fix_ref links from older review rounds still resolve correctly. +364: func FetchFindingFromPR(repo string, prNumber int, fixRef string) (*Finding, error) { +365: owner, name, err := parseRepoSlug(repo) +366: if err != nil { +367: return nil, err +368: } +369: +370: apiPath := fmt.Sprintf("repos/%s/%s/pulls/%d/reviews", owner, name, prNumber) +371: out, err := exec.Command("gh", "api", apiPath).Output() +372: if err != nil { +373: return nil, fmt.Errorf("fetching PR reviews: %w", err) +374: } +375: +376: var reviews []ghReview +377: if err := json.Unmarshal(out, &reviews); err != nil { +378: return nil, fmt.Errorf("parsing PR reviews: %w", err) +379: } +380: +381: prefixes := reviewMarkerPrefixes +382: const suffix = " -->" +383: +384: for _, rev := range reviews { +385: for _, prefix := range prefixes { +386: idx := strings.Index(rev.Body, prefix) +387: if idx < 0 { +388: continue +389: } +390: start := idx + len(prefix) +391: endIdx := strings.Index(rev.Body[start:], suffix) +392: if endIdx < 0 { +393: continue +394: } +395: var result ReviewResult +396: if err := json.Unmarshal([]byte(rev.Body[start:start+endIdx]), &result); err != nil { +397: continue +398: } +399: for i := range result.Findings { +400: if result.Findings[i].FixRef == fixRef { +401: return &result.Findings[i], nil +402: } +403: } +404: } +405: } +406: +407: return nil, fmt.Errorf("fix_ref %q not found in any review on PR #%d", fixRef, prNumber) +408: } +409: +410: // DetectRepo gets owner/name from the current git remote. +411: func DetectRepo() (string, error) { +412: out, err := exec.Command("gh", "repo", "view", +413: "--json", "nameWithOwner", +414: "--jq", ".nameWithOwner", +415: ).Output() +416: if err != nil { +417: return "", fmt.Errorf("gh repo view: %w", err) +418: } +419: return strings.TrimSpace(string(out)), nil +420: } +421: +422: // DetectPRNumber detects the PR number for the current branch using gh. +423: // If repo is non-empty, it is passed as --repo to scope the lookup. +424: func DetectPRNumber(repo string) (int, error) { +425: args := []string{"pr", "view", "--json", "number", "--jq", ".number"} +426: if repo != "" { +427: args = append(args, "--repo", repo) +428: } +429: out, err := exec.Command("gh", args...).Output() +430: if err != nil { +431: return 0, fmt.Errorf("no open pull request found for the current branch") +432: } +433: num, err := strconv.Atoi(strings.TrimSpace(string(out))) +434: if err != nil { +435: return 0, fmt.Errorf("unexpected PR number from gh: %w", err) +436: } +437: return num, nil +438: } +439: +440: // ThreadReply represents a reply to a review thread (i.e. any comment after the first). +441: type ThreadReply struct { +442: Author string +443: Body string +444: } +445: +446: // ReviewThread represents a review thread from a PR. +447: type ReviewThread struct { +448: ID string +449: Path string +450: Line int +451: Body string +452: Author string // login of the first comment author (the bot for review threads) +453: Outdated bool // true if GitHub marked the comment position as outdated (code changed) +454: Resolved bool +455: Replies []ThreadReply +456: } +457: +458: // graphQLThreadsResponse is the JSON shape returned by the review threads query. +459: type graphQLThreadsResponse struct { +460: Data struct { +461: Repository struct { +462: PullRequest struct { +463: ReviewThreads struct { +464: Nodes []struct { +465: ID string `json:"id"` +466: IsResolved bool `json:"isResolved"` +467: Comments struct { +468: Nodes []struct { +469: Body string `json:"body"` +470: Path string `json:"path"` +471: Line int `json:"line"` +472: OriginalLine int `json:"originalLine"` +473: Outdated bool `json:"outdated"` +474: Author struct { +475: Login string `json:"login"` +476: } `json:"author"` +477: } `json:"nodes"` +478: } `json:"comments"` +479: } `json:"nodes"` +480: } `json:"reviewThreads"` +481: } `json:"pullRequest"` +482: } `json:"repository"` +483: } `json:"data"` +484: } +485: +486: // FetchReviewThreads gets all review threads from a PR via GraphQL. +487: func FetchReviewThreads(repo string, prNumber int) ([]ReviewThread, error) { +488: owner, name, err := parseRepoSlug(repo) +489: if err != nil { +490: return nil, err +491: } +492: +493: query := `query($owner:String!,$name:String!,$pr:Int!){ +494: repository(owner:$owner,name:$name){ +495: pullRequest(number:$pr){ +496: reviewThreads(first:100){ +497: nodes{ +498: id +499: isResolved +500: comments(first:100){ +501: nodes{body path line originalLine outdated author{login}} +502: } +503: } +504: } +505: } +506: } +507: }` +508: +509: cmd := exec.Command("gh", "api", "graphql", +510: "-f", "query="+query, +511: "-f", fmt.Sprintf("owner=%s", owner), +512: "-f", fmt.Sprintf("name=%s", name), +513: "-F", fmt.Sprintf("pr=%d", prNumber), +514: ) +515: out, err := cmd.Output() +516: if err != nil { +517: return nil, fmt.Errorf("gh api graphql: %w", err) +518: } +519: +520: var resp graphQLThreadsResponse +521: if err := json.Unmarshal(out, &resp); err != nil { +522: return nil, fmt.Errorf("parsing graphql response: %w", err) +523: } +524: +525: var threads []ReviewThread +526: for _, node := range resp.Data.Repository.PullRequest.ReviewThreads.Nodes { +527: if len(node.Comments.Nodes) == 0 { +528: continue +529: } +530: comment := node.Comments.Nodes[0] +531: +532: // Filter to review threads only (new marker + legacy markers for backward compat). +533: if !strings.Contains(comment.Body, findingMarkerPrefix) && +534: !strings.Contains(comment.Body, "codecanary fix") && +535: !strings.Contains(comment.Body, "clanopy fix") { +536: continue +537: } +538: +539: var replies []ThreadReply +540: for _, c := range node.Comments.Nodes[1:] { +541: replies = append(replies, ThreadReply{ +542: Author: c.Author.Login, +543: Body: c.Body, +544: }) +545: } +546: +547: line := comment.Line +548: if comment.Outdated && line == 0 && comment.OriginalLine > 0 { +549: line = comment.OriginalLine +550: } +551: +552: threads = append(threads, ReviewThread{ +553: ID: node.ID, +554: Path: comment.Path, +555: Line: line, +556: Body: comment.Body, +557: Author: comment.Author.Login, +558: Outdated: comment.Outdated, +559: Resolved: node.IsResolved, +560: Replies: replies, +561: }) +562: } +563: +564: return threads, nil +565: } +566: +567: // ResolveThread resolves a review thread via GraphQL mutation. +568: func ResolveThread(threadID string) error { +569: cmd := exec.Command("gh", "api", "graphql", +570: "-f", "query=mutation($threadId:ID!){resolveReviewThread(input:{threadId:$threadId}){thread{isResolved}}}", +571: "-f", fmt.Sprintf("threadId=%s", threadID), +572: ) +573: if out, err := cmd.CombinedOutput(); err != nil { +574: return fmt.Errorf("gh api graphql resolve: %w\n%s", err, string(out)) +575: } +576: return nil +577: } +578: +579: // ReplyToThread posts a reply on a review thread via GraphQL. +580: func ReplyToThread(threadID, body string) error { +581: cmd := exec.Command("gh", "api", "graphql", +582: "-f", "query=mutation($threadId:ID!,$body:String!){addPullRequestReviewThreadReply(input:{pullRequestReviewThreadId:$threadId,body:$body}){comment{id}}}", +583: "-f", fmt.Sprintf("threadId=%s", threadID), +584: "-f", fmt.Sprintf("body=%s", body), +585: ) +586: if out, err := cmd.CombinedOutput(); err != nil { +587: return fmt.Errorf("gh api graphql reply: %w\n%s", err, string(out)) +588: } +589: return nil +590: } +591: +592: // ghReview is the JSON shape for a PR review from the REST API. +593: type ghReview struct { +594: ID int64 `json:"id"` +595: NodeID string `json:"node_id"` +596: Body string `json:"body"` +597: } +598: +599: // LatestCodecanaryReview captures the identity and body of the most recent +600: // CodeCanary top-level review, together with the commit SHA embedded in its +601: // hidden marker. Returned by FetchLatestCodecanaryReview for the +602: // "edit if same SHA, otherwise post new" publish decision. +603: type LatestCodecanaryReview struct { +604: ID int64 +605: SHA string +606: Body string +607: } +608: +609: // FetchLatestCodecanaryReview returns the most recent top-level review on +610: // the PR that carries a CodeCanary marker, along with the commit SHA +611: // embedded in that marker. Returns (nil, nil) when no matching review +612: // exists. A non-nil error is returned only for transport/parsing failures. +613: func FetchLatestCodecanaryReview(repo string, prNumber int) (*LatestCodecanaryReview, error) { +614: owner, name, err := parseRepoSlug(repo) +615: if err != nil { +616: return nil, err +617: } +618: +619: apiPath := fmt.Sprintf("repos/%s/%s/pulls/%d/reviews", owner, name, prNumber) +620: out, err := exec.Command("gh", "api", apiPath).Output() +621: if err != nil { +622: return nil, fmt.Errorf("fetching PR reviews: %w", err) +623: } +624: +625: var reviews []ghReview +626: if err := json.Unmarshal(out, &reviews); err != nil { +627: return nil, fmt.Errorf("parsing PR reviews: %w", err) +628: } +629: +630: // Walk newest → oldest so we return the latest CodeCanary review. +631: for i := len(reviews) - 1; i >= 0; i-- { +632: rev := reviews[i] +633: for _, prefix := range reviewMarkerPrefixes { +634: idx := strings.Index(rev.Body, prefix) +635: if idx < 0 { +636: continue +637: } +638: start := idx + len(prefix) +639: endIdx := strings.Index(rev.Body[start:], reviewMarkerSuffix) +640: if endIdx < 0 { +641: continue +642: } +643: jsonData := rev.Body[start : start+endIdx] +644: var payload struct { +645: SHA string `json:"sha"` +646: } +647: // Ignore unmarshal errors so old reviews that embed a richer +648: // ReviewResult (also containing a "sha" field) still match. +649: _ = json.Unmarshal([]byte(jsonData), &payload) +650: return &LatestCodecanaryReview{ +651: ID: rev.ID, +652: SHA: payload.SHA, +653: Body: rev.Body, +654: }, nil +655: } +656: } +657: +658: return nil, nil +659: } +660: +661: // UpdateReviewBody replaces the body of an existing PR review via the +662: // GitHub REST API. Used when a reply-only run (or a duplicate synchronize +663: // webhook on the same HEAD) needs to refresh the latest CodeCanary review's +664: // counts instead of posting a new top-level comment. +665: func UpdateReviewBody(repo string, prNumber int, reviewID int64, body string) error { +666: owner, name, err := parseRepoSlug(repo) +667: if err != nil { +668: return err +669: } +670: +671: payload, err := json.Marshal(struct { +672: Body string `json:"body"` +673: }{Body: body}) +674: if err != nil { +675: return fmt.Errorf("marshaling update payload: %w", err) +676: } +677: +678: apiPath := fmt.Sprintf("repos/%s/%s/pulls/%d/reviews/%d", owner, name, prNumber, reviewID) +679: _, err = ghAPIRequest("PUT", apiPath, payload) +680: return err +681: } +682: +683: // ghAPIRequest is a generic gh api wrapper that preserves the temp-file +684: // pattern used by ghAPIPOST (avoids stdin pipe truncation on large bodies). +685: func ghAPIRequest(method, apiPath string, payloadJSON []byte) ([]byte, error) { +686: dir, err := os.MkdirTemp("", "codecanary-*") +687: if err != nil { +688: return nil, fmt.Errorf("creating temp dir: %w", err) +689: } +690: defer func() { _ = os.RemoveAll(dir) }() +691: +692: tmpFile, err := os.CreateTemp(dir, "payload.json") +693: if err != nil { +694: return nil, fmt.Errorf("creating temp file: %w", err) +695: } +696: if _, err := tmpFile.Write(payloadJSON); err != nil { +697: _ = tmpFile.Close() +698: return nil, fmt.Errorf("writing payload to temp file: %w", err) +699: } +700: _ = tmpFile.Close() +701: +702: cmd := exec.Command("gh", "api", apiPath, "--method", method, "--input", tmpFile.Name()) +703: var stdout, stderr bytes.Buffer +704: cmd.Stdout = &stdout +705: cmd.Stderr = &stderr +706: if err := cmd.Run(); err != nil { +707: return stdout.Bytes(), &apiError{Err: err, Stderr: stderr.String(), Response: stdout.String()} +708: } +709: return stdout.Bytes(), nil +710: } +711: +712: // FetchPreviousReviewSHA gets the SHA from the last review's hidden data. +713: func FetchPreviousReviewSHA(repo string, prNumber int) string { +714: owner, name, err := parseRepoSlug(repo) +715: if err != nil { +716: return "" +717: } +718: +719: apiPath := fmt.Sprintf("repos/%s/%s/pulls/%d/reviews", owner, name, prNumber) +720: out, err := exec.Command("gh", "api", apiPath).Output() +721: if err != nil { +722: return "" +723: } +724: +725: var reviews []ghReview +726: if err := json.Unmarshal(out, &reviews); err != nil { +727: return "" +728: } +729: +730: prefixes := reviewMarkerPrefixes +731: const suffix = " -->" +732: +733: // Search from most recent to oldest. +734: for i := len(reviews) - 1; i >= 0; i-- { +735: body := reviews[i].Body +736: for _, prefix := range prefixes { +737: idx := strings.Index(body, prefix) +738: if idx < 0 { +739: continue +740: } +741: start := idx + len(prefix) +742: endIdx := strings.Index(body[start:], suffix) +743: if endIdx < 0 { +744: continue +745: } +746: jsonData := body[start : start+endIdx] +747: var result ReviewResult +748: if err := json.Unmarshal([]byte(jsonData), &result); err != nil { +749: continue +750: } +751: if result.SHA != "" { +752: return result.SHA +753: } +754: } +755: } +756: +757: return "" +758: } +759: +760: // FilesFromDiff extracts the list of file paths touched in a unified diff. +761: func FilesFromDiff(diff string) []string { +762: var files []string +763: seen := make(map[string]bool) +764: for _, line := range strings.Split(diff, "\n") { +765: if strings.HasPrefix(line, "+++ b/") { +766: path := strings.TrimRight(line[6:], "\r") +767: if !seen[path] { +768: seen[path] = true +769: files = append(files, path) +770: } +771: } +772: } +773: return files +774: } +775: +776: // ScopeDiffToFiles filters a unified diff to only include hunks for files in +777: // the allowed set. This prevents rebase noise (main-branch changes) from +778: // leaking into incremental reviews. +779: func ScopeDiffToFiles(diff string, allowedFiles map[string]bool) string { +780: if len(allowedFiles) == 0 { +781: return diff +782: } +783: +784: lines := strings.Split(diff, "\n") +785: var result []string +786: blockStart := -1 +787: blockAllowed := false +788: +789: for i := 0; i < len(lines); i++ { +790: if strings.HasPrefix(lines[i], "diff --git") { +791: // Flush previous block if allowed. +792: if blockStart >= 0 && blockAllowed { +793: result = append(result, lines[blockStart:i]...) +794: } +795: blockStart = i +796: blockAllowed = false +797: +798: // Look ahead for +++ b/ to determine if this block is allowed. +799: for j := i + 1; j < len(lines) && !strings.HasPrefix(lines[j], "diff --git"); j++ { +800: if strings.HasPrefix(lines[j], "+++ b/") { +801: path := strings.TrimRight(lines[j][6:], "\r") +802: if allowedFiles[path] { +803: blockAllowed = true +804: } +805: break +806: } +807: } +808: continue +809: } +810: } +811: +812: // Flush last block. +813: if blockStart >= 0 && blockAllowed { +814: result = append(result, lines[blockStart:]...) +815: } +816: +817: if len(result) == 0 { +818: return "" +819: } +820: return strings.Join(result, "\n") +821: } +822: +823: // validSHA matches a full-length lowercase hex Git SHA. +824: var validSHA = regexp.MustCompile(`^[0-9a-f]{40}$`) +825: +826: // GetIncrementalDiff gets the diff since a given SHA. +827: func GetIncrementalDiff(baseSHA string) (string, error) { +828: if !validSHA.MatchString(baseSHA) { +829: return "", fmt.Errorf("invalid SHA format: %q", baseSHA) +830: } +831: out, err := exec.Command("git", "diff", baseSHA+"..HEAD").Output() +832: if err != nil { +833: return "", fmt.Errorf("git diff: %w", err) +834: } +835: return string(out), nil +836: } +837: +838: // PostCleanReview posts a review when the first review finds no issues. The +839: // commitSHA is embedded in a hidden marker so future runs treat it as the +840: // baseline for incremental reviews, avoiding a redundant full re-review on the +841: // next push. +842: func PostCleanReview(repo string, prNumber int, commitSHA string, summary ReviewSummary) error { +843: return postSimpleReview(repo, prNumber, buildCleanReviewBody(commitSHA, summary)) +844: } +845: +846: // PostAllClearReview posts a review when all previous findings have been +847: // resolved. If minimizeFailed is true, a note is appended warning about +848: // visible old reviews. The commitSHA is embedded in a hidden marker so future +849: // runs treat it as the baseline for incremental reviews; without it, the next +850: // push would fall back to reviewing the entire PR again. +851: func PostAllClearReview(repo string, prNumber int, commitSHA string, minimizeFailed bool, summary ReviewSummary) error { +852: return postSimpleReview(repo, prNumber, buildAllClearReviewBody(commitSHA, minimizeFailed, summary)) +853: } +854: +855: // PostActivityReview posts a review when no new findings were raised but +856: // there is cycle activity worth surfacing (dismissals, acknowledgments, +857: // rebuttals, still-open threads). This keeps every commit push producing a +858: // visible top-level status comment instead of silently logging. +859: func PostActivityReview(repo string, prNumber int, commitSHA string, summary ReviewSummary) error { +860: return postSimpleReview(repo, prNumber, buildActivityReviewBody(commitSHA, summary)) +861: } +862: +863: // buildCleanReviewBody renders the full Markdown body posted by +864: // PostCleanReview. Split out from the poster so tests can assert the exact +865: // string that lands on GitHub without having to mock gh. +866: func buildCleanReviewBody(commitSHA string, summary ReviewSummary) string { +867: return withSummary("CodeCanary reviewed this PR \u2014 no issues found.", summary) + embedBaselineMarker(commitSHA) +868: } +869: +870: // buildAllClearReviewBody renders the full Markdown body posted by +871: // PostAllClearReview. Split out for the same reason as buildCleanReviewBody. +872: func buildAllClearReviewBody(commitSHA string, minimizeFailed bool, summary ReviewSummary) string { +873: body := "## \U0001F425 CodeCanary\n\n\u2705 All previous findings have been addressed. No new issues found. \u2728" +874: if minimizeFailed { +875: body += "\n\n> \u26A0\uFE0F Some previous review comments could not be minimized and may still be visible." +876: } +877: return withSummary(body, summary) + embedBaselineMarker(commitSHA) +878: } +879: +880: // buildActivityReviewBody renders the body for a commit push that raised no +881: // new findings but has cycle activity (dismissals/acknowledgments/rebuttals +882: // or still-open threads carried forward). +883: func buildActivityReviewBody(commitSHA string, summary ReviewSummary) string { +884: body := "## \U0001F425 CodeCanary\n\nReviewed this push \u2014 no new issues found." +885: return withSummary(body, summary) + embedBaselineMarker(commitSHA) +886: } +887: +888: // withSummary appends the status summary block to a review body. The block +889: // is skipped when the summary has no non-zero counts, so existing clean/ +890: // all-clear bodies render identically when nothing happened. +891: func withSummary(body string, summary ReviewSummary) string { +892: block := renderSummaryBlock(summary) +893: if block == "" { +894: return body +895: } +896: return body + block +897: } +898: +899: // embedBaselineMarker returns a hidden HTML comment containing the commitSHA +900: // so FetchPreviousReviewSHA can use this review as the incremental baseline. +901: // Returns an empty string if commitSHA is empty (local mode, dry run). The +902: // marker carries only the SHA — FetchPreviousReviewSHA is the sole reader and +903: // it only needs that field. +904: func embedBaselineMarker(commitSHA string) string { +905: if commitSHA == "" { +906: return "" +907: } +908: data, err := json.Marshal(struct { +909: SHA string `json:"sha"` +910: }{SHA: commitSHA}) +911: if err != nil { +912: return "" +913: } +914: return fmt.Sprintf("\n%s%s%s\n", reviewMarkerPrefixes[0], string(data), reviewMarkerSuffix) +915: } +916: +917: func postSimpleReview(repo string, prNumber int, body string) error { +918: owner, repoName, err := parseRepoSlug(repo) +919: if err != nil { +920: return err +921: } +922: +923: payload := reviewPayload{ +924: Event: "COMMENT", +925: Body: body, +926: Comments: make([]reviewComment, 0), +927: } +928: +929: payloadJSON, err := json.Marshal(payload) +930: if err != nil { +931: return fmt.Errorf("marshaling payload: %w", err) +932: } +933: +934: _, err = ghAPIPOST(fmt.Sprintf("repos/%s/%s/pulls/%d/reviews", owner, repoName, prNumber), payloadJSON) +935: return err +936: } +937: +938: // apiError is returned by ghAPIPOST so callers can inspect the stderr output +939: // from gh (which contains the HTTP status line) separately from the response body. +940: type apiError struct { +941: Err error +942: Stderr string +943: Response string +944: } +945: +946: func (e *apiError) Error() string { +947: return fmt.Sprintf("gh api: %v\nstderr: %s\nresponse: %s", e.Err, e.Stderr, e.Response) +948: } +949: +950: func (e *apiError) Unwrap() error { return e.Err } +951: +952: // ghAPIPOST sends a JSON payload to the GitHub API via gh, using a temp file +953: // to avoid stdin pipe issues that can cause "unexpected end of JSON input" +954: // errors on large payloads. +955: func ghAPIPOST(apiPath string, payloadJSON []byte) ([]byte, error) { +956: dir, err := os.MkdirTemp("", "codecanary-*") +957: if err != nil { +958: return nil, fmt.Errorf("creating temp dir: %w", err) +959: } +960: defer func() { _ = os.RemoveAll(dir) }() +961: +962: tmpFile, err := os.CreateTemp(dir, "payload.json") +963: if err != nil { +964: return nil, fmt.Errorf("creating temp file: %w", err) +965: } +966: +967: if _, err := tmpFile.Write(payloadJSON); err != nil { +968: _ = tmpFile.Close() +969: return nil, fmt.Errorf("writing payload to temp file: %w", err) +970: } +971: _ = tmpFile.Close() +972: +973: cmd := exec.Command("gh", "api", apiPath, "--method", "POST", "--input", tmpFile.Name()) +974: var stdout, stderr bytes.Buffer +975: cmd.Stdout = &stdout +976: cmd.Stderr = &stderr +977: if err := cmd.Run(); err != nil { +978: return stdout.Bytes(), &apiError{Err: err, Stderr: stderr.String(), Response: stdout.String()} +979: } +980: return stdout.Bytes(), nil +981: } +982: +983: // FindReviewNodeIDs returns the node_ids of all reviews on a PR. +984: func FindReviewNodeIDs(repo string, prNumber int) ([]string, error) { +985: owner, name, err := parseRepoSlug(repo) +986: if err != nil { +987: return nil, err +988: } +989: +990: apiPath := fmt.Sprintf("repos/%s/%s/pulls/%d/reviews", owner, name, prNumber) +991: out, err := exec.Command("gh", "api", apiPath).Output() +992: if err != nil { +993: return nil, fmt.Errorf("fetching PR reviews: %w", err) +994: } +995: +996: var reviews []ghReview +997: if err := json.Unmarshal(out, &reviews); err != nil { +998: return nil, fmt.Errorf("parsing PR reviews: %w", err) +999: } +1000: +1001: prefixes := reviewMarkerPrefixes +1002: +1003: var nodeIDs []string +1004: for _, rev := range reviews { +1005: for _, prefix := range prefixes { +1006: if strings.Contains(rev.Body, prefix) { +1007: nodeIDs = append(nodeIDs, rev.NodeID) +1008: break +1009: } +1010: } +1011: } +1012: +1013: return nodeIDs, nil +1014: } +1015: +1016: // threadHeaderLine returns the first non-empty, non-HTML-comment line from +1017: // a thread body. This is the header line containing severity and finding ID. +1018: func threadHeaderLine(body string) string { +1019: for _, line := range strings.SplitN(body, "\n", 5) { +1020: line = strings.TrimSpace(line) +1021: if line == "" || strings.HasPrefix(line, "" +1216: +1217: var result []ReviewInfo +1218: for _, rev := range reviews { +1219: for _, prefix := range prefixes { +1220: idx := strings.Index(rev.Body, prefix) +1221: if idx < 0 { +1222: continue +1223: } +1224: start := idx + len(prefix) +1225: endIdx := strings.Index(rev.Body[start:], suffix) +1226: if endIdx < 0 { +1227: continue +1228: } +1229: jsonData := rev.Body[start : start+endIdx] +1230: var rr ReviewResult +1231: if err := json.Unmarshal([]byte(jsonData), &rr); err != nil { +1232: continue +1233: } +1234: var ids []string +1235: for _, f := range rr.Findings { +1236: ids = append(ids, f.ID) +1237: } +1238: result = append(result, ReviewInfo{ +1239: NodeID: rev.NodeID, +1240: FindingIDs: ids, +1241: }) +1242: break +1243: } +1244: } +1245: +1246: return result, nil +1247: } +1248: +1249: // MinimizeComment hides a comment on GitHub using the minimizeComment GraphQL mutation. +1250: func MinimizeComment(nodeID string) error { +1251: cmd := exec.Command("gh", "api", "graphql", +1252: "-f", "query=mutation($id:ID!){minimizeComment(input:{subjectId:$id,classifier:RESOLVED}){minimizedComment{isMinimized}}}", +1253: "-F", "id="+nodeID, +1254: ) +1255: if out, err := cmd.CombinedOutput(); err != nil { +1256: return fmt.Errorf("gh api graphql minimize: %w\n%s", err, string(out)) +1257: } +1258: return nil +1259: } +1260: +1261: // FetchFileContents reads the full contents of changed files from disk. +1262: // It skips files that are too large, binary, deleted, or match ignore patterns. +1263: // Returns a map of path->content and a list of skipped file paths. +1264: func FetchFileContents(files []string, ignorePatterns []string, maxPerFile, maxTotal int) (map[string]string, []string) { +1265: contents := make(map[string]string) +1266: var skipped []string +1267: totalSize := 0 +1268: +1269: for _, path := range files { +1270: // Check ignore patterns. +1271: if matchesIgnore(path, ignorePatterns) { +1272: skipped = append(skipped, path) +1273: continue +1274: } +1275: +1276: data, err := os.ReadFile(path) +1277: if err != nil { +1278: // File may have been deleted in this PR — skip gracefully. +1279: continue +1280: } +1281: +1282: // Skip binary files (null bytes in first 512 bytes). +1283: peek := data +1284: if len(peek) > 512 { +1285: peek = peek[:512] +1286: } +1287: if bytes.ContainsRune(peek, 0) { +1288: skipped = append(skipped, path) +1289: continue +1290: } +1291: +1292: size := len(data) +1293: +1294: // Skip files exceeding per-file limit. +1295: if size > maxPerFile { +1296: skipped = append(skipped, path) +1297: continue +1298: } +1299: +1300: // Stop if total budget would be exceeded. +1301: if totalSize+size > maxTotal { +1302: skipped = append(skipped, path) +1303: continue +1304: } +1305: +1306: contents[path] = string(data) +1307: totalSize += size +1308: } +1309: +1310: return contents, skipped +1311: } +1312: +1313: // isSetupPR detects whether this is the initial setup PR. +1314: // Returns true only when a new workflow file referencing codecanary is added AND +1315: // the PR contains no other files beyond expected setup artifacts (workflow + +1316: // config), so that PRs bundling real code changes are never silently skipped. +1317: func isSetupPR(diff string, files []string) bool { +1318: // All files must be known setup paths. +1319: for _, f := range files { +1320: if !isSetupFile(f) { +1321: return false +1322: } +1323: } +1324: +1325: // At least one newly added workflow file must reference codecanary. +1326: lines := strings.Split(diff, "\n") +1327: for i := 0; i < len(lines)-1; i++ { +1328: if lines[i] != "--- /dev/null" { +1329: continue +1330: } +1331: plusLine := lines[i+1] +1332: if !strings.HasPrefix(plusLine, "+++ b/.github/workflows/") { +1333: continue +1334: } +1335: for j := i + 2; j < len(lines); j++ { +1336: if strings.HasPrefix(lines[j], "--- ") || strings.HasPrefix(lines[j], "diff --git") { +1337: break +1338: } +1339: if strings.HasPrefix(lines[j], "+") && (strings.Contains(lines[j], "codecanary") || strings.Contains(lines[j], "clanopy")) { +1340: return true +1341: } +1342: } +1343: } +1344: return false +1345: } +1346: +1347: // isSetupFile returns true if the file path is a known setup artifact. +1348: func isSetupFile(path string) bool { +1349: return strings.HasPrefix(path, ".github/workflows/") || +1350: strings.HasPrefix(path, ".codecanary/") || path == ".codecanary.yml" || +1351: strings.HasPrefix(path, ".clanopy/") +1352: } +1353: +1354: // matchesIgnore checks if a path matches any of the ignore glob patterns. +1355: // Uses doublestar to support ** recursive globs (e.g. "dist/**", "src/**/*.test.*"). +1356: func matchesIgnore(path string, patterns []string) bool { +1357: for _, pat := range patterns { +1358: if matched, _ := doublestar.Match(pat, path); matched { +1359: return true +1360: } +1361: // Also try matching against just the filename. +1362: if matched, _ := doublestar.Match(pat, filepath.Base(path)); matched { +1363: return true +1364: } +1365: } +1366: return false +1367: } +``` + +## Diff +```diff +diff --git a/internal/review/github.go b/internal/review/github.go +index cf0dee8..00081a6 100644 +--- a/internal/review/github.go ++++ b/internal/review/github.go +@@ -14,6 +14,24 @@ import ( + "github.com/bmatcuk/doublestar/v4" + ) + ++// runGH runs a gh command and returns stdout; on failure, the returned error ++// includes gh's stderr so upstream API messages surface in logs instead of just ++// "exit status 1". ++func runGH(label string, args ...string) ([]byte, error) { ++ cmd := exec.Command("gh", args...) ++ var stderr bytes.Buffer ++ cmd.Stderr = &stderr ++ out, err := cmd.Output() ++ if err != nil { ++ msg := strings.TrimSpace(stderr.String()) ++ if msg == "" { ++ return nil, fmt.Errorf("%s: %w", label, err) ++ } ++ return nil, fmt.Errorf("%s: %s: %w", label, msg, err) ++ } ++ return out, nil ++} ++ + // parseRepoSlug splits a "owner/name" repository slug into its two parts. + func parseRepoSlug(repo string) (owner, name string, err error) { + parts := strings.SplitN(repo, "/", 2) +@@ -84,12 +102,12 @@ func FetchPR(repo string, number int) (*PRData, error) { + numStr := fmt.Sprintf("%d", number) + + // Fetch PR metadata as JSON. +- viewOut, err := exec.Command("gh", "pr", "view", numStr, ++ viewOut, err := runGH("gh pr view", "pr", "view", numStr, + "--repo", repo, + "--json", "title,body,author,baseRefName,headRefName,files", +- ).Output() ++ ) + if err != nil { +- return nil, fmt.Errorf("gh pr view: %w", err) ++ return nil, err + } + + var view ghPRView +@@ -97,12 +115,11 @@ func FetchPR(repo string, number int) (*PRData, error) { + return nil, fmt.Errorf("parsing gh pr view output: %w", err) + } + +- // Fetch the diff. +- diffOut, err := exec.Command("gh", "pr", "diff", numStr, +- "--repo", repo, +- ).Output() ++ // Fetch the diff. This hits GitHub's pull request diff API ++ // (GET /repos/{owner}/{repo}/pulls/{n} with Accept: application/vnd.github.v3.diff). ++ diffOut, err := runGH("gh pr diff", "pr", "diff", numStr, "--repo", repo) + if err != nil { +- return nil, fmt.Errorf("gh pr diff: %w", err) ++ return nil, fmt.Errorf("%w (GitHub pull request diff API; if this looks like a transient 5xx/timeout, retrying the job may help)", err) + } + + files := make([]string, len(view.Files)) +``` + +## Output Format +Return your findings as a JSON array inside a ```json code fence. Each finding must have these fields: + +- `id` (string): The rule ID that was violated, or a short kebab-case identifier for general findings. +- `file` (string): The file path where the issue was found. **Must be one of the exact paths listed in "Files in This Diff" above.** If a file path does not appear in that list, do NOT reference it. If your finding relates to a file not in the diff (e.g. a downstream consequence), set `file` and `line` to the diff location that triggers the issue and mention the affected file in `description`. +- `line` (int): The line number in the file. **Must be a line that was added or modified in the diff** (a `+` line in the diff hunk). If your finding is about a side effect on a distant line, set `line` to the diff line that *causes* the issue and describe the affected location in `description`. +- `severity` (string): One of "critical", "bug", "warning", "suggestion", or "nitpick". + - "critical": Security vulnerabilities, data loss, crashes. + - "bug": A logic error that causes incorrect runtime behavior for real inputs. Missing test coverage, unused parameters, typos in identifiers that happen to compile, or "what if a future caller…" concerns do NOT qualify — use "suggestion" or "nitpick" for those. If you cannot name the concrete input and the concrete wrong output, it is not a bug. + - "warning": Potential issues, performance problems, code smells. + - "suggestion": Better patterns, readability improvements. + - "nitpick": Minor style, naming, formatting. +- `title` (string): A short title for the finding. +- `description` (string): A concise explanation of the issue — 2-3 sentences max. State what is wrong and why it matters. Do not repeat the code or walk through the logic step by step. +- `suggestion` (string, optional): A concise suggested fix — 1-2 sentences of prose, then a code block if helpful. Do not explain what the code block does. For suggestions about broader patterns or improvements beyond the current PR scope, recommend opening a separate PR — do not imply they should fix it here. +- `fix_ref` (string): A reference ID in the format `165-` where index starts at 1 (e.g. `165-1`, `165-2`). +- `actionable` (boolean): Set to `false` if your analysis concludes the code is correct and no change is needed. Set to `true` if the finding requires the author to act. **Prefer returning an empty array over emitting findings with `actionable: false`.** + +**IMPORTANT — JSON escaping:** When your description or suggestion references code containing backslash sequences (e.g. `\n`, `\t`, `\"`), you MUST double-escape the backslash in the JSON string value. For example, to mention `fmt.Print("\n")` in a JSON string, write `fmt.Print("\\n")`. A single `\n` in JSON is a newline character, not the literal text `\n`. + +**Do not include findings where your conclusion is that the code is correct or no action is needed.** If you evaluate something and determine it is fine, omit it entirely rather than reporting it. Specifically: if you begin analyzing a potential issue but then realize the code handles it correctly, do NOT emit a finding that walks through the concern and then concludes "this is actually fine" or "no bug here" — simply drop it. Every finding you emit must represent a real, actionable problem. + +**Check against project documentation before emitting.** The "Project Documentation" section above defines conventions for this codebase (e.g. "don't add error handling for scenarios that can't happen", "keep the core engine agnostic"). Before emitting a finding, verify it does not contradict those conventions. If your suggested fix would violate a project-doc rule, drop the finding — the author has already made that tradeoff deliberately. + +**Label uncertainty from external behavior.** If your finding's validity depends on the behavior of a third-party API, webhook payload shape, framework internal, or other system you cannot verify from the diff, file contents, and project docs above, you MUST (a) cap severity at "suggestion" and (b) state the assumption in `description` (e.g. "Assumes `github.event.pull_request.number` is unset on `pull_request_review_comment` events — verify against GitHub's webhook docs before acting."). A finding that asserts external behavior as fact without this label is a false-positive risk. + +**CRITICAL: Do NOT invent or hallucinate file paths, function names, or code that does not appear in the diff or the provided file contents. If a file or function is not shown above, do not reference it.** + +If there are no findings, return an empty array: `[]`. + +Example: +```json +[ + { + "id": "rule-id", + "file": "src/main.go", + "line": 42, + "severity": "warning", + "title": "Short title", + "description": "The value is used after the error check, so a non-nil error silently proceeds with stale data.", + "suggestion": "Return early on error.\n\n```go\nif err != nil {\n return err\n}\n```", + "fix_ref": "165-1", + "actionable": true + } +] +``` diff --git a/internal/review/testdata/corpus/alansikora-codecanary-pr173.json b/internal/review/testdata/corpus/alansikora-codecanary-pr173.json new file mode 100644 index 0000000..9aed202 --- /dev/null +++ b/internal/review/testdata/corpus/alansikora-codecanary-pr173.json @@ -0,0 +1,35 @@ +{ + "format_version": 1, + "name": "alansikora-codecanary-pr173", + "repo": "alansikora/codecanary", + "pr_number": 173, + "head_sha": "628ba9440af6e9b5b07ac05c1c4f18a1a9c44fbd", + "captured_at": "2026-09-04T23:31:55Z", + "pr": { + "number": 173, + "title": "feat: prettier commit status context (CodeCanary / review)", + "body": "## Summary\n\nRenames the commit status context posted by `GithubPlatform.Publish` (and `codecanary signoff`) from `codecanary/review` to `CodeCanary / review`. Capital C, spaces around the slash — matches how GitHub renders workflow check runs, so the status reads as cleanly as the existing CI entries in the PR checks list.\n\nNo behavior change. Same success/failure logic, same call sites, same pure-helper test coverage.\n\n## Heads-up: clash with the Action's check run\n\nThe workflow job `review` in workflow `CodeCanary` also renders as `CodeCanary / review` in the PR checks UI. They're technically separate entities (Check Run vs Commit Status), and the branch-protection status-check picker distinguishes them by source, but they now share an identical display label. Discussed in chat; accepted for now on the theory that the two entries both represent \"CodeCanary did its thing for this commit\" and a single unified label is fine. Easy to split later by renaming the workflow job id or the status context if it becomes confusing.\n\n## Migration\n\nExisting installs that wired `codecanary/review` into branch-protection rules will need to update the required-check name to `CodeCanary / review` after this merges. Documented in the README and `docs/review-flow.md` alongside the rename.\n\n## Test plan\n\n- [x] `go build ./...`\n- [x] `go vet ./...`\n- [x] `go test ./...`\n- [ ] End-to-end: merge + push, confirm `CodeCanary / review = success` / `failure` lands on HEAD\n- [ ] Confirm branch-protection picker lists the new context", + "author": "alansikora", + "base_branch": "main", + "head_branch": "pretty-status-context", + "diff": "diff --git a/README.md b/README.md\nindex b0d4ab1..d15aad7 100644\n--- a/README.md\n+++ b/README.md\n@@ -81,12 +81,12 @@ Once merged, CodeCanary reviews every PR on open and push. Draft PRs are skipped\n \n ### Gating merges on clean reviews\n \n-CodeCanary can block merges until a review comes back clean. After every review, the bot (and the local `codecanary signoff` command) posts a GitHub commit status under the context `codecanary/review`:\n+CodeCanary can block merges until a review comes back clean. After every review, the bot (and the local `codecanary signoff` command) posts a GitHub commit status under the context `CodeCanary / review`:\n \n - `success` — no unresolved findings (everything is either unraised, fixed by code, or handled by the author)\n - `failure` — one or more findings remain unresolved, with a description like `\"3 unresolved findings\"`\n \n-To turn this into a required check, add `codecanary/review` to your repo's required status checks via whichever branch protection mechanism you use (rulesets, classic branch protection rules, etc.). GitHub accepts any context name; if a review has already run, it will also show up in autocomplete.\n+To turn this into a required check, add `CodeCanary / review` to your repo's required status checks via whichever branch protection mechanism you use (rulesets, classic branch protection rules, etc.). GitHub accepts any context name; if a review has already run, it will also show up in autocomplete.\n \n Statuses are keyed by commit SHA, so staling is automatic: a new push has no status until the next review run posts one, which re-blocks the merge button.\n \n@@ -95,7 +95,7 @@ Statuses are keyed by commit SHA, so staling is automatic: a new push has no sta\n ```sh\n codecanary review # reviews locally, stores findings in ~/.codecanary/...\n # fix anything that came up, commit\n-codecanary signoff # posts codecanary/review = success on HEAD\n+codecanary signoff # posts CodeCanary / review = success on HEAD\n ```\n \n The command refuses to sign off unless HEAD matches the reviewed SHA and the tree is clean — this prevents attesting a review of code that isn't actually in the commit. It needs `gh` authenticated with `repo:status` scope (default `gh auth login` covers it).\n@@ -109,7 +109,7 @@ Same required-check config works for both paths: the bot satisfies the check on\n | `codecanary review [pr-number]` | Review a PR or local diff |\n | `codecanary findings [pr-number]` | Fetch bot findings for a PR (markdown or JSON) |\n | `codecanary reply --url \u003cURL\u003e --body \u003ctext\u003e` | Post a reply on a review-comment thread (used by the skill when skipping) |\n-| `codecanary signoff` | Post a `codecanary/review` commit status from the last local review (see [gating merges](#gating-merges-on-clean-reviews)) |\n+| `codecanary signoff` | Post a `CodeCanary / review` commit status from the last local review (see [gating merges](#gating-merges-on-clean-reviews)) |\n | `codecanary install-skill` | Install the `codecanary-fix` Claude Code skill |\n | `codecanary setup [local\\|github]` | Interactive setup wizard |\n | `codecanary auth status` | Show stored credential info |\ndiff --git a/cmd/review/cli/signoff.go b/cmd/review/cli/signoff.go\nindex 82a1c57..3621db9 100644\n--- a/cmd/review/cli/signoff.go\n+++ b/cmd/review/cli/signoff.go\n@@ -25,7 +25,7 @@ Requires:\n - the working tree is clean and HEAD matches the reviewed SHA\n - 'gh' is installed and authenticated with repo:status scope\n \n-Combine with a required 'codecanary/review' check in branch protection to\n+Combine with a required 'CodeCanary / review' check in branch protection to\n block merges until a clean local review exists for the tip commit.`,\n \tRunE: func(cmd *cobra.Command, args []string) error {\n \t\tforce, _ := cmd.Flags().GetBool(\"force\")\ndiff --git a/docs/review-flow.md b/docs/review-flow.md\nindex 3f398d3..3973019 100644\n--- a/docs/review-flow.md\n+++ b/docs/review-flow.md\n@@ -165,7 +165,7 @@ Per-thread ack replies for dismissed/acknowledged/rebutted resolutions are poste\n \n `codecanary findings` applies the same marker to filter deferrals out of its default output: threads with any `codecanary:ack:*` reply are treated as handled and omitted alongside GitHub-resolved threads. Pass `--include-resolved` to see them. This keeps the codecanary-fix skill from re-prompting on findings the operator already deferred.\n \n-After the review is posted (or updated in place), `GithubPlatform.Publish` also POSTs a `codecanary/review` commit status on the reviewed SHA via `PostReviewCommitStatus`. State is `success` when `NewFindings + StillOpen == 0` (everything is either new-and-green, fixed by code, or explicitly handled by the author), and `failure` otherwise. Description is the unresolved count or \"all findings resolved\" / \"no findings\". Teams that add `codecanary/review` as a required status check in branch protection get auto-gating: merges are blocked until a review run posts a green status on HEAD. Status posting failures are logged as warnings and do not abort Publish — the review itself has already landed. The local `codecanary signoff` command posts a status under the same context, so a team can rely on a single required check that either the bot (pr-loop) or a local reviewer (local-loop) satisfies.\n+After the review is posted (or updated in place), `GithubPlatform.Publish` also POSTs a `CodeCanary / review` commit status on the reviewed SHA via `PostReviewCommitStatus`. State is `success` when `NewFindings + StillOpen == 0` (everything is either new-and-green, fixed by code, or explicitly handled by the author), and `failure` otherwise. Description is the unresolved count or \"all findings resolved\" / \"no findings\". Teams that add `CodeCanary / review` as a required status check in branch protection get auto-gating: merges are blocked until a review run posts a green status on HEAD. Status posting failures are logged as warnings and do not abort Publish — the review itself has already landed. The local `codecanary signoff` command posts a status under the same context, so a team can rely on a single required check that either the bot (pr-loop) or a local reviewer (local-loop) satisfies.\n \n **Local**: Prints the formatted result to stdout. Format depends on context: terminal (colored, human-readable), markdown, or JSON.\n \ndiff --git a/internal/review/github.go b/internal/review/github.go\nindex b626a82..80b0015 100644\n--- a/internal/review/github.go\n+++ b/internal/review/github.go\n@@ -314,11 +314,11 @@ func PostReview(repo string, prNumber int, result *ReviewResult, diff string, co\n \n // ReviewCommitStatusContext is the commit-status context the review bot\n // writes when a review run completes. Using a stable string lets teams\n-// require it in branch protection — \"codecanary/review must be green\n+// require it in branch protection — \"CodeCanary / review must be green\n // before merge.\" Matches the context used by `codecanary signoff`, so\n // both the bot (on PRs with the workflow) and a local signoff can\n // satisfy the same required check.\n-const ReviewCommitStatusContext = \"codecanary/review\"\n+const ReviewCommitStatusContext = \"CodeCanary / review\"\n \n // PostReviewCommitStatus POSTs a commit status on the given SHA summarising\n // the review outcome. state is \"success\" when all findings are resolved or\ndiff --git a/internal/review/platform_github.go b/internal/review/platform_github.go\nindex 01843e2..37c4f7f 100644\n--- a/internal/review/platform_github.go\n+++ b/internal/review/platform_github.go\n@@ -289,7 +289,7 @@ func (g *GithubPlatform) ReportUsage(tracker *UsageTracker) {\n \t}\n }\n \n-// postReviewCommitStatus POSTs a `codecanary/review` commit status on the\n+// postReviewCommitStatus POSTs a `CodeCanary / review` commit status on the\n // reviewed SHA. state=success when no unresolved findings remain for the\n // PR (new findings this cycle + threads still open with no classification\n // both at zero); state=failure otherwise. Teams can require this check in\n", + "files": [ + "README.md", + "cmd/review/cli/signoff.go", + "docs/review-flow.md", + "internal/review/github.go", + "internal/review/platform_github.go" + ], + "file_contents": { + "README.md": "# \u003cimg width=\"75\" alt=\"codecanary\" src=\"https://github.com/user-attachments/assets/bb494aa1-9bb2-486c-a253-ba8a9a2939e4\" /\u003e CodeCanary\n\nAI-powered code review for GitHub pull requests. Catch bugs, security issues, and quality problems before they land in main.\n\n## Quick Start\n\n```sh\ncurl -fsSL https://codecanary.sh/install | sh\ncodecanary setup local\ncodecanary review\n```\n\nThat's it. CodeCanary diffs your branch against main and reviews the changes locally.\n\n## Why CodeCanary?\n\n- **Fully automated** — runs as a GitHub Action on every push, or locally from the terminal.\n- **Multi-provider** — bring your own LLM: Anthropic, OpenAI, OpenRouter, Grok (xAI), or Claude CLI. No vendor lock-in.\n- **Incremental reviews** — on re-push, Go-driven triage classifies existing threads at zero LLM cost. Only changed code gets re-evaluated.\n- **Conversational** — when authors reply to a finding, CodeCanary re-evaluates in context. It distinguishes code fixes, dismissals, acknowledgments, and rebuttals.\n- **Native PR integration** — posts inline comments on exact diff lines, auto-resolves threads when code is fixed, and minimizes stale reviews.\n- **Anti-hallucination** — explicit file allowlists, line validation against the diff, and distance thresholds prevent fabricated findings.\n- **Cost-efficient** — uses a fast triage model for thread re-evaluation and a full model for review. Tracks per-invocation usage so you see what you spend.\n- **Configuration-as-code** — project-specific rules, severity levels, ignore patterns, and context in `.codecanary/config.yml`.\n- **Agentic loop** — pairs with Claude Code via the bundled `codecanary-fix` skill to review, triage, fix, and push until the PR is clean.\n\n## Installation\n\n```sh\ncurl -fsSL https://codecanary.sh/install | sh\n```\n\nInstalls the `codecanary` binary to `/usr/local/bin` (or `~/.local/bin`). Supports Linux and macOS (amd64/arm64).\n\nTo self-update later:\n\n```sh\ncodecanary upgrade\n```\n\n### Canary builds\n\n```sh\ncurl -fsSL https://codecanary.sh/install | sh -s -- --canary\ncodecanary upgrade --canary\n```\n\n## Setup\n\n### Local reviews\n\n```sh\ncodecanary setup local\n```\n\nThe setup wizard walks you through choosing a provider, entering your API key (stored in your system keychain), and selecting models. It creates a `.codecanary/config.yml` in your repo.\n\nThen review your changes:\n\n```sh\ncodecanary review # diff against main, print to terminal\ncodecanary review --post # same, but also post findings to the PR on GitHub\ncodecanary review --output json # machine-readable output\n```\n\nWithout `--post`, `codecanary review` is always local: it diffs your branch (with uncommitted changes) against the default branch and keeps state in `~/.codecanary/state/\u003cbranch\u003e.json` for incremental re-runs. With `--post`, it fetches the PR from GitHub — pass a number or let it auto-detect from the current branch — and posts findings as review comments.\n\n### GitHub Actions\n\n```sh\ncodecanary setup github\n```\n\nThis runs the same provider and key selection, then:\n1. Installs the CodeCanary Review GitHub App\n2. Sets your API key as a GitHub repo secret\n3. Creates the workflow (`.github/workflows/codecanary.yml`)\n4. Opens a PR with everything ready to merge\n\nOnce merged, CodeCanary reviews every PR on open and push. Draft PRs are skipped by default.\n\n### Gating merges on clean reviews\n\nCodeCanary can block merges until a review comes back clean. After every review, the bot (and the local `codecanary signoff` command) posts a GitHub commit status under the context `CodeCanary / review`:\n\n- `success` — no unresolved findings (everything is either unraised, fixed by code, or handled by the author)\n- `failure` — one or more findings remain unresolved, with a description like `\"3 unresolved findings\"`\n\nTo turn this into a required check, add `CodeCanary / review` to your repo's required status checks via whichever branch protection mechanism you use (rulesets, classic branch protection rules, etc.). GitHub accepts any context name; if a review has already run, it will also show up in autocomplete.\n\nStatuses are keyed by commit SHA, so staling is automatic: a new push has no status until the next review run posts one, which re-blocks the merge button.\n\n**Without a workflow.** If your repo doesn't use the GitHub Action, `codecanary signoff` posts the same status from your local machine after a clean `codecanary review`:\n\n```sh\ncodecanary review # reviews locally, stores findings in ~/.codecanary/...\n# fix anything that came up, commit\ncodecanary signoff # posts CodeCanary / review = success on HEAD\n```\n\nThe command refuses to sign off unless HEAD matches the reviewed SHA and the tree is clean — this prevents attesting a review of code that isn't actually in the commit. It needs `gh` authenticated with `repo:status` scope (default `gh auth login` covers it).\n\nSame required-check config works for both paths: the bot satisfies the check on PRs that run the workflow; `codecanary signoff` satisfies it for repos without the workflow, branches where the workflow doesn't run, or PRs from forks where secrets are unavailable.\n\n## CLI Reference\n\n| Command | Description |\n|---------|-------------|\n| `codecanary review [pr-number]` | Review a PR or local diff |\n| `codecanary findings [pr-number]` | Fetch bot findings for a PR (markdown or JSON) |\n| `codecanary reply --url \u003cURL\u003e --body \u003ctext\u003e` | Post a reply on a review-comment thread (used by the skill when skipping) |\n| `codecanary signoff` | Post a `CodeCanary / review` commit status from the last local review (see [gating merges](#gating-merges-on-clean-reviews)) |\n| `codecanary install-skill` | Install the `codecanary-fix` Claude Code skill |\n| `codecanary setup [local\\|github]` | Interactive setup wizard |\n| `codecanary auth status` | Show stored credential info |\n| `codecanary auth delete` | Remove a stored API key |\n| `codecanary auth refresh` | Validate and update stored credentials |\n| `codecanary upgrade` | Update to the latest release |\n\n### Review flags\n\n| Flag | Description |\n|------|-------------|\n| `--repo, -r` | GitHub repo (owner/name) |\n| `--output, -o` | Output format: `terminal`, `markdown`, or `json` (auto-detects TTY) |\n| `--post` | Post findings as a PR review comment |\n| `--config, -c` | Path to config file (auto-detected if empty) |\n| `--reply-only` | Re-evaluate thread replies only, skip new findings |\n| `--dry-run` | Show the prompt without calling the LLM |\n\n## Configuration\n\nCodeCanary uses `.codecanary/config.yml` in your repo. The `provider` field is required.\n\n### Minimal config\n\n```yaml\nversion: 1\nprovider: anthropic\nreview_model: claude-sonnet-4-6\ntriage_model: claude-haiku-4-5-20251001\n```\n\n### Config with rules and context\n\n```yaml\nversion: 1\nprovider: anthropic\nreview_model: claude-sonnet-4-6\ntriage_model: claude-haiku-4-5-20251001\n\ncontext: |\n Go REST API using chi router. Tests use testify.\n\nrules:\n - id: error-handling\n description: \"Errors must be wrapped with context using fmt.Errorf\"\n severity: warning\n paths: [\"**/*.go\"]\n\n - id: sql-injection\n description: \"Database queries must use parameterized statements\"\n severity: critical\n\nignore:\n - \"dist/**\"\n - \"*.lock\"\n - \"vendor/**\"\n```\n\nRules support `paths` and `exclude_paths` globs, and five severity levels: `critical`, `bug`, `warning`, `suggestion`, `nitpick`.\n\n### Provider examples\n\n**OpenAI** (also works with Azure, Ollama, or any OpenAI-compatible endpoint via `api_base`):\n```yaml\nversion: 1\nprovider: openai\nreview_model: gpt-5.4\ntriage_model: gpt-5.4-mini\n# api_base: https://your-endpoint.com/v1\n```\n\n**OpenRouter**:\n```yaml\nversion: 1\nprovider: openrouter\nreview_model: anthropic/claude-sonnet-4-6\ntriage_model: anthropic/claude-haiku-4-5-20251001\n```\n\n**Grok (xAI)**:\n```yaml\nversion: 1\nprovider: grok\nreview_model: grok-4.20-0309-non-reasoning\ntriage_model: grok-4-1-fast-non-reasoning\n```\n\n**Claude CLI** (uses your logged-in `claude` session, no API key needed):\n```yaml\nversion: 1\nprovider: claude\nreview_model: claude-sonnet-4-6\ntriage_model: haiku\n```\n\nYou can also create a `.codecanary/review.local.yml` for personal overrides (gitignored) — its rules, context, and ignore patterns are appended to the shared `review.yml`.\n\nFor the full config reference including budget controls, size limits, timeouts, evaluation context, and the `review.yml` override file, see [docs/configuration.md](docs/configuration.md).\n\n## Credential Management\n\n```sh\ncodecanary auth status # show which API keys are stored\ncodecanary auth delete # remove a stored API key\ncodecanary auth refresh # validate and update credentials\n```\n\nKeys are stored in your system keychain (macOS Keychain, GNOME Keyring, KDE Wallet) with a fallback to `~/.codecanary/credentials.json`. Environment variables always override stored credentials.\n\n| Provider | What you need | Where to get it |\n|----------|--------------|-----------------|\n| Anthropic | API key | [console.anthropic.com](https://console.anthropic.com) |\n| OpenAI | API key | [platform.openai.com](https://platform.openai.com) |\n| OpenRouter | API key | [openrouter.ai](https://openrouter.ai) |\n| Grok (xAI) | API key | [console.x.ai](https://console.x.ai) |\n| Claude CLI | Logged-in `claude` binary | Run `claude` and complete the login flow |\n\n## How It Works\n\n### First review\n\n1. Fetches PR metadata and diff (via `gh` CLI or local git)\n2. Reads file contents for context (respecting ignore patterns and size limits)\n3. Auto-discovers project docs (CLAUDE.md files) for additional context\n4. Calls your configured LLM to analyze the changes\n5. Posts findings as inline PR review comments (or prints to terminal)\n\n### Incremental reviews (on re-push)\n\n1. **Go-driven triage** classifies existing threads — no LLM calls for unchanged code\n2. **Parallel evaluation** re-checks threads where code changed or the author replied (using the triage model)\n3. **New code review** covers only the incremental diff, excluding known issues\n4. **Auto-resolution** marks threads as resolved when the code fix addresses the finding\n\n### Thread lifecycle\n\n| Event | Result |\n|-------|--------|\n| Code fix detected | Thread auto-resolved |\n| Author dismisses | Acknowledged, kept open for re-check |\n| Author acknowledges | Noted, kept open |\n| Author rebuts | Evaluated for technical merit, kept open |\n| No changes | Skipped (zero LLM cost) |\n\n### Safety\n\n- **Anti-hallucination**: explicit file allowlist, line number validation against diff, max finding distance threshold\n- **Anti-ping-pong**: resolved findings injected as context to prevent re-raising\n- **Prompt injection protection**: repository content escaped before inclusion in prompts\n\n## Agentic review loop\n\nCodeCanary ships with a [Claude Code](https://docs.claude.com/en/docs/claude-code) skill, `codecanary-fix`, that drives a review → triage → fix → push cycle until the PR is clean. You stay in the loop — every fix is confirmed before it's applied — but the polling, fetching, and CI watching is handled by the CLI.\n\nInstall the skill once:\n\n```sh\ncodecanary install-skill\n```\n\nThis writes the embedded skill to `~/.claude/skills/codecanary-fix/SKILL.md`, where Claude Code discovers it in every session. Re-run the command after `codecanary upgrade` to pick up new versions.\n\nThen in Claude Code, ask it to `handle codecanary` on your PR (or invoke `/codecanary-fix` directly) — the skill is auto-discovered and matched to your request via its frontmatter description. Two modes:\n\n- **PR mode** (default) — watches the GitHub Actions review check via `codecanary findings --watch`, renders a triage table, asks you to confirm which fixes to apply, commits and pushes, then loops on the next review. Every finding you defer gets a reply posted on its review thread explaining why, via `codecanary reply`.\n- **Local mode** — triggered automatically when no PR is detected for the current branch. Single pass against your dirty working tree. Applies approved fixes without committing or pushing.\n\nThe full skill contract lives at [internal/skills/codecanary-fix/SKILL.md](internal/skills/codecanary-fix/SKILL.md).\n\n## Contributing\n\nSee [CONTRIBUTING.md](CONTRIBUTING.md) for architecture details and how to add new LLM providers or platforms.\n\n## License\n\nMIT\n", + "cmd/review/cli/signoff.go": "package cli\n\nimport (\n\t\"errors\"\n\t\"fmt\"\n\n\t\"github.com/alansikora/codecanary/internal/review\"\n\t\"github.com/spf13/cobra\"\n)\n\n// signoffCmd posts a GitHub commit status on the current HEAD based on the\n// most recent local review for the current branch. Inspired by Rails 8.1's\n// `gh signoff` (Basecamp): combined with a required status check in branch\n// protection, this lets a team gate merges on \"a human actually ran\n// codecanary locally and it came back clean\", without paying for a cloud\n// review on every push.\nvar signoffCmd = \u0026cobra.Command{\n\tUse: \"signoff\",\n\tShort: \"Post a GitHub commit status from the last local review\",\n\tLong: `Post a GitHub commit status on HEAD reflecting the most recent local\ncodecanary review for the current branch.\n\nRequires:\n - you ran 'codecanary review' on this branch\n - the working tree is clean and HEAD matches the reviewed SHA\n - 'gh' is installed and authenticated with repo:status scope\n\nCombine with a required 'CodeCanary / review' check in branch protection to\nblock merges until a clean local review exists for the tip commit.`,\n\tRunE: func(cmd *cobra.Command, args []string) error {\n\t\tforce, _ := cmd.Flags().GetBool(\"force\")\n\n\t\tslug, err := review.RepoSlug()\n\t\tif err != nil {\n\t\t\treturn err\n\t\t}\n\t\tbranch, err := review.CurrentBranch()\n\t\tif err != nil {\n\t\t\treturn err\n\t\t}\n\t\tsha, err := review.HeadSHA()\n\t\tif err != nil {\n\t\t\treturn err\n\t\t}\n\n\t\tstate, err := review.LoadLocalState(branch)\n\t\tif err != nil {\n\t\t\treturn fmt.Errorf(\"loading local state: %w\", err)\n\t\t}\n\t\tif state == nil {\n\t\t\treturn fmt.Errorf(\"no local review found for branch %q — run 'codecanary review' first\", branch)\n\t\t}\n\n\t\tif !force {\n\t\t\tif state.SHA != sha {\n\t\t\t\treturn fmt.Errorf(\n\t\t\t\t\t\"local review was for %s but HEAD is %s — re-run 'codecanary review' (or pass --force)\",\n\t\t\t\t\tshortSHA(state.SHA), shortSHA(sha))\n\t\t\t}\n\t\t\tdirty, err := review.WorkingTreeDirty()\n\t\t\tif err != nil {\n\t\t\t\treturn fmt.Errorf(\"checking working tree: %w\", err)\n\t\t\t}\n\t\t\tif dirty {\n\t\t\t\treturn errors.New(\n\t\t\t\t\t\"working tree has uncommitted changes — commit them, \" +\n\t\t\t\t\t\t\"run 'git stash --include-untracked' (plain 'git stash' leaves untracked files behind), \" +\n\t\t\t\t\t\t\"or pass --force if you know the dirty files are unrelated to the PR\")\n\t\t\t}\n\t\t}\n\n\t\tunresolved := 0\n\t\tfor _, f := range state.Findings {\n\t\t\tif f.Actionable != nil \u0026\u0026 !*f.Actionable {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tunresolved++\n\t\t}\n\n\t\tstatusState, desc := \"success\", \"0 findings\"\n\t\tif unresolved \u003e 0 {\n\t\t\tstatusState = \"failure\"\n\t\t\tsuffix := \"s\"\n\t\t\tif unresolved == 1 {\n\t\t\t\tsuffix = \"\"\n\t\t\t}\n\t\t\tdesc = fmt.Sprintf(\"%d unresolved finding%s\", unresolved, suffix)\n\t\t}\n\n\t\tif err := review.PostReviewCommitStatus(slug, sha, statusState, desc); err != nil {\n\t\t\treturn fmt.Errorf(\"posting commit status: %w\", err)\n\t\t}\n\n\t\t_, _ = fmt.Fprintf(cmd.OutOrStdout(),\n\t\t\t\"✓ %s = %s on %s@%s (%s)\\n\",\n\t\t\treview.ReviewCommitStatusContext, statusState, slug, shortSHA(sha), desc)\n\t\treturn nil\n\t},\n}\n\nfunc init() {\n\tsignoffCmd.Flags().Bool(\"force\", false,\n\t\t\"Sign off even if HEAD doesn't match the reviewed SHA or the tree is dirty\")\n\trootCmd.AddCommand(signoffCmd)\n}\n", + "docs/review-flow.md": "# Review Flow\n\nHow CodeCanary reviews a pull request, step by step.\n\n## Overview\n\nThe review pipeline has two modes of operation:\n\n- **First review**: Reviews the full PR diff against the base branch.\n- **Incremental review**: Re-evaluates previous findings and reviews only new changes since the last review.\n\nBoth modes run through the same `Run()` function in `runner.go`. The pipeline is platform-agnostic -- GitHub and local modes differ only in which `ReviewPlatform` adapter is injected.\n\n## Platforms\n\nTwo platforms, routed strictly by `--post`:\n\n| Context | Platform | How it runs | State storage | Output |\n|---------|----------|-------------|---------------|--------|\n| **GitHub PR** | `GithubPlatform` | `codecanary review --post` (locally or in CI) | PR review threads via API | Posts review comments on the PR |\n| **Local** | `LocalPlatform` | `codecanary review` (with or without a PR for the branch) | `~/.codecanary/repos/\u003cowner\u003e/\u003crepo\u003e/state/\u003cbranch\u003e.json` | Prints to terminal |\n\n`codecanary review` without `--post` is always local — even if the branch has an open PR. \"Local is local\": the branch diff (including uncommitted changes) is reviewed against the default base, and previous findings come from the repo-scoped state file. There is no hybrid mode that reads GitHub but writes local state. Two consecutive local runs go incremental off the saved state (locked in by `TestLocalPlatformIncrementalHandoff` in `state_test.go`).\n\nState files are keyed by `owner/repo/branch` so the same branch name across different repos (e.g. `main` in repo A vs. repo B) no longer collides. When the git remote can't be resolved (detached worktree, no remote), state falls back to the legacy `~/.codecanary/state/\u003cbranch\u003e.json` path. On first save after the upgrade, any existing legacy file is read once, migrated to the repo-scoped path, and the legacy copy is removed — transparent to the operator.\n\n## Pipeline Steps\n\n### 1. Fetch PR data\n\n**GitHub PR** (`--post`): Fetches PR metadata (title, body, author, branches) and diff via `gh pr view` and `gh pr diff`.\n\n**Local**: Detects the default branch (`main`, falling back to `master`, or the explicit `--base`) and computes diff from merge-base to HEAD via `git diff $(git merge-base HEAD \u003cdefault-branch\u003e)..HEAD`. Uncommitted working-tree changes scoped to the branch files are appended to the incremental diff on subsequent runs. Uses current branch name as the title and `git config user.name` as the author.\n\nIf the PR is a setup PR (only adds workflow files with no real code changes), the review is skipped with an informational comment.\n\n### 2. Prepare review context\n\n`prepareReview()` loads everything the review needs:\n\n- **Config**: Reads `config.yml` (provider, models, budgets, timeouts). If a `review.yml` exists alongside it, its rules/context/ignore fields override the config. If a `review.local.yml` also exists, its fields are appended (not replaced) on top of `review.yml`.\n- **Project docs**: Discovers CLAUDE.md files at the repo root and in every ancestor directory of a changed PR file. Skips `vendor/`, `node_modules/`, hidden dirs, and other build artifacts. Up to 10 files, 16 KB each, 48 KB total. Monorepos commonly keep per-app conventions (e.g. `apps/exchange-api/CLAUDE.md`) — those load automatically when a PR touches files under that directory, so the reviewer sees the conventions specific to the code being changed rather than only the repo-root overview.\n- **File contents**: Reads changed files from disk with size limits (default 100KB per file, 500KB total). Skips binary files, ignored patterns, and files exceeding limits. When files are skipped, the diff is also filtered to remove their hunks (via `ScopeDiffToFiles`) and they are removed from the file list. The original unfiltered diff is preserved in `FullDiff` for finding validation.\n- **Environment**: Builds a filtered env for LLM subprocesses (only allowed prefixes like `CODECANARY_`, `GITHUB_`, plus essential vars like `PATH`). Injects keychain credentials if not already set.\n\n### 3. Create providers\n\nTwo `ModelProvider` instances are created from config:\n\n- **Review provider**: The main model that reviews code (configured via `review_model` in config). When `advisor_model` is set (anthropic or claude provider only), the review provider also enables Anthropic's server-side advisor tool so a stronger advisor model can weigh in mid-generation — the triage provider never uses advisor, since its classifier turns are too short to benefit.\n- **Triage provider**: A cheaper model for re-evaluating previous findings (configured via `triage_model` in config).\n\nEach provider is constructed via the factory registry in `provider.go`. The provider name determines which adapter handles the API call (Anthropic, OpenAI, OpenRouter, or Claude CLI).\n\n### 4. Load previous findings\n\nThe platform adapter loads unresolved findings from the last review:\n\n**GitHub PR** (`--post`): Fetches review threads via GraphQL. Filters to CodeCanary findings only (detected by HTML marker comments). Extracts the previous review's HEAD SHA from the most recent review body — clean and all-clear reviews embed this marker too, so the baseline advances even when a push produced no findings. Returns unresolved threads, the SHA, and a count for fix_ref numbering.\n\n**Local**: Reads `~/.codecanary/repos/\u003cowner\u003e/\u003crepo\u003e/state/\u003cbranch\u003e.json`, which stores the SHA, branch name, and findings array from the previous review. Falls back to `~/.codecanary/state/\u003cbranch\u003e.json` when the repo slug can't be resolved or a pre-migration state file is present. Converts saved findings into `ReviewThread` shape for the triage pipeline.\n\nIf no previous findings exist, this is a first review.\n\n### 5. Triage and build prompt\n\nThis step diverges based on whether a previous review SHA exists. A previous SHA alone is enough to enter the incremental path — if previous findings were all resolved (no open threads), the incremental diff still scopes the review to commits since the last baseline, avoiding a redundant full re-review.\n\n#### First review path\n\nCalls `BuildPrompt()` to assemble the full review prompt. The prompt includes (in order):\n\n1. System instructions (reviewer role, diff-only rules, side-effect awareness)\n2. PR metadata (number, title, author, description)\n3. Additional context from config\n4. Project documentation (CLAUDE.md files in `\u003cproject-doc\u003e` tags)\n5. Review rules (from config) — filtered to rules whose `paths:` / `exclude_paths:` globs match at least one PR file. Rules scoped to file types not in the diff (e.g. CSS rules on a Ruby-only change) are omitted to keep LLM attention focused. Falls back to a general review instruction when no rules apply.\n6. Ignore patterns\n7. Explicit file allowlist (anti-hallucination)\n8. Full contents of changed files with line numbers\n9. The unified diff\n10. Output format instructions (JSON schema, examples, escaping rules)\n\nAfter building, `fitPromptForModel()` checks whether the prompt fits the review model's context window (context window minus max output tokens). If it exceeds the budget, it progressively drops the largest file contents first, then truncates the diff as a last resort.\n\n#### Incremental review path (triage)\n\n`runTriage()` handles the incremental case in two phases.\n\n**Phase 1 -- Classify and evaluate previous findings**\n\nFirst, an incremental diff is computed (`git diff \u003cpreviousSHA\u003e..HEAD`). Two diffs serve different purposes:\n\n- **Activity diff** (incremental): Determines whether there's new activity to evaluate. If empty, threads with no replies are skipped (no LLM cost).\n- **Context diff** (full PR diff): Used for classification and evaluation context. Ensures fixes from earlier pushes are visible even if they predate the incremental window.\n\n`ClassifyThreads()` assigns each unresolved thread one of six classifications:\n\n| Classification | Condition | Evaluation |\n|---|---|---|\n| `TriageSkip` | No activity diff, not outdated, no replies | Skipped (no LLM) |\n| `TriageCodeChanged` | GitHub outdated flag, or file in PR diff | LLM evaluates with file-scoped diff + file snippet |\n| `TriageHasReply` | Human replied (no code changes) | LLM evaluates reply intent |\n| `TriageCodeChangedReply` | Both code changed and human replied | LLM evaluates both |\n| `TriageCrossFileChange` | Changes in other files only | LLM evaluates with full PR diff |\n| `TriageFileRemovedFromPR` | File no longer in PR | Auto-resolved by Go code (no LLM) -- thread resolved on GitHub |\n\nThreads classified as `TriageFileRemovedFromPR` are auto-resolved without an LLM call. The Go code sets reason `file_removed` and resolves the thread directly.\n\nFor remaining threads, `EvaluateThreadsParallel()` runs up to 3 concurrent LLM calls using the triage model. Evaluation uses a **two-level approach** to balance precision and coverage:\n\n- **Level 1 (file-scoped)**: The LLM receives the finding, the current file content (presented first), and a file-scoped diff. This catches same-file fixes with minimal noise. Most evaluations resolve here.\n- **Level 2 (widened scope)**: Only when level 1 says \"not resolved\" and the thread is `TriageCodeChanged` (file is in the PR diff). The LLM receives the full PR diff with a prompt primed to look for cross-file fixes. This catches the edge case where the finding's file has unrelated changes but the actual fix is in a different file.\n\nFor `TriageCrossFileChange`, only the full PR diff is used (no level 1 — there's no file-scoped diff to show).\n\nThe LLM returns JSON: `{\"resolved\": true, \"reason\": \"code_change\"}` or `{\"resolved\": false}`.\n\nLLM resolution reasons and their effects:\n\n| Reason | Effect | Thread stays open? |\n|---|---|---|\n| `code_change` | Thread resolved on GitHub | No |\n| `dismissed` | Ack reply posted | Yes (re-triaged on next push) |\n| `acknowledged` | Ack reply posted | Yes |\n| `rebutted` | Ack reply posted | Yes |\n\n**Phase 2 -- Build prompt for new findings**\n\nAfter triage, the pipeline builds an incremental review prompt using `BuildIncrementalPrompt()`. This is similar to `BuildPrompt()` but:\n\n- Uses the incremental diff (or falls back to full PR diff if the incremental diff failed)\n- Includes a \"Known Issues\" section listing unresolved threads (prevents duplicating them)\n- Includes a \"Recently Resolved Issues\" section with findings fixed by code changes (prevents re-raising similar issues -- anti-ping-pong)\n- Only includes file contents for files touched in the incremental diff\n\nThe prompt is then fitted to the context window, same as the first review path.\n\n### 6. LLM call\n\nIf not a dry run and budget permits, the review prompt is sent to the review provider. The provider handles API communication (Anthropic Messages API, OpenAI Chat Completions, OpenRouter, or Claude CLI).\n\nIf the response is truncated (hit max output tokens), a warning is logged. The pipeline attempts to salvage complete findings from the truncated JSON by scanning backward for valid objects.\n\n### 7. Process findings\n\n`processFindings()` parses and validates the LLM's output:\n\n1. **Parse JSON**: Extracts the findings array from the ```json fence. Falls back to bracket-matching if embedded code blocks break the regex.\n2. **File validation**: Drops findings referencing files not in the PR.\n3. **Line validation**: Drops findings whose line number is more than 20 lines from any changed line in the PR diff. This catches hallucinated line numbers and scope creep.\n4. **Actionable filter**: Removes findings where `actionable: false`.\n5. **Status tagging**: Tags all findings as `\"new\"` if this is an incremental review.\n\n### 8. Publish results\n\n**GitHub PR** (`--post`): Every cycle emits exactly one top-level CodeCanary review, decided by an edit-vs-post rule. `FetchLatestCodecanaryReview` reads the commit SHA from the most recent CodeCanary review's hidden marker:\n\n- **Same SHA** (reply-only run, or a duplicate `synchronize` webhook on the same HEAD): the existing body is updated in place with `UpdateReviewBody`. Only the status block between the `\u003c!-- codecanary:status --\u003e` markers is swapped — inline comments and prior findings text are untouched.\n- **Different or no SHA** (new commits, or first review on the PR): a fresh review is posted. The body variant depends on the cycle outcome — findings review, all-clear, activity summary (no new findings but cycle activity to surface), or clean review. All variants carry the same status block and baseline SHA marker. Older CodeCanary reviews are minimized (collapsed) before posting.\n\nThe status block lists non-zero counts for: new findings, resolved by code, file removed, dismissed by author, acknowledged by author, rebutted by author, still unresolved. The block renders nothing when all counts are zero, so clean reviews remain copy-exact.\n\nPer-thread ack replies for dismissed/acknowledged/rebutted resolutions are posted earlier in the pipeline (`HandleResolutions`). Dedup is reason-agnostic: if the thread already carries *any* `\u003c!-- codecanary:ack:... --\u003e` marker, no further ack reply is posted. Reasons can shift across triage runs (LLM non-determinism), and all three convey the same outcome (\"keeping open\"), so one ack per thread is enough. Reply-only runs skip `SaveState` so the empty findings slice doesn't overwrite persisted state.\n\n`codecanary findings` applies the same marker to filter deferrals out of its default output: threads with any `codecanary:ack:*` reply are treated as handled and omitted alongside GitHub-resolved threads. Pass `--include-resolved` to see them. This keeps the codecanary-fix skill from re-prompting on findings the operator already deferred.\n\nAfter the review is posted (or updated in place), `GithubPlatform.Publish` also POSTs a `CodeCanary / review` commit status on the reviewed SHA via `PostReviewCommitStatus`. State is `success` when `NewFindings + StillOpen == 0` (everything is either new-and-green, fixed by code, or explicitly handled by the author), and `failure` otherwise. Description is the unresolved count or \"all findings resolved\" / \"no findings\". Teams that add `CodeCanary / review` as a required status check in branch protection get auto-gating: merges are blocked until a review run posts a green status on HEAD. Status posting failures are logged as warnings and do not abort Publish — the review itself has already landed. The local `codecanary signoff` command posts a status under the same context, so a team can rely on a single required check that either the bot (pr-loop) or a local reviewer (local-loop) satisfies.\n\n**Local**: Prints the formatted result to stdout. Format depends on context: terminal (colored, human-readable), markdown, or JSON.\n\n### 9. Save state\n\n**GitHub PR** (`--post`): No-op. State is stored in the review threads themselves (the embedded JSON marker contains the SHA and findings).\n\n**Local**: Writes `~/.codecanary/repos/\u003cowner\u003e/\u003crepo\u003e/state/\u003cbranch\u003e.json` with the current HEAD SHA, branch name, and combined findings (still-open + new). If a legacy `~/.codecanary/state/\u003cbranch\u003e.json` exists from a pre-migration run, it is removed after the new file lands. This enables incremental reviews on the next run.\n\n### 10. Report usage\n\n**GitHub PR** (`--post`): Writes token counts and cost to `GITHUB_ENV` for downstream workflow steps.\n\n**Local**: Prints a usage summary table to stderr (model, tokens, cost, duration) if running in a terminal.\n\n### 11. Telemetry\n\nIf telemetry is enabled (opt-in), fires an anonymous event with aggregate stats: provider, platform, finding counts by severity, token counts, cost, and duration. No code content is sent.\n\n## Key Design Decisions\n\n**Single pipeline, two platforms.** `Run()` never branches on \"am I on GitHub?\" The `ReviewPlatform` interface absorbs all environment differences. Adding a new platform (e.g. GitLab) means implementing the interface, not forking the pipeline.\n\n**Two diffs for triage.** The incremental diff (changes since last review) decides whether to skip evaluation. The full PR diff (all changes) provides context for evaluation. This prevents the \"triage horizon\" bug where fixes committed before the triage baseline become invisible.\n\n**Two-level triage evaluation.** Same-file evaluations (`TriageCodeChanged`) start with a file-scoped diff (level 1) to reduce noise — the full PR diff can drown out the relevant fix with changes from unrelated files. If level 1 finds no fix, a widened-scope fallback (level 2) sends the full PR diff to catch cross-file fixes. Cross-file evaluations (`TriageCrossFileChange`) go straight to the full diff. The file snippet (current code state) is presented first in all evaluation prompts, so the LLM checks whether the issue still exists before analyzing the diff.\n\n**Per-thread evaluation.** Each unresolved thread gets its own LLM call with tailored context, rather than one bulk prompt. This allows fine-grained classification, parallel execution, and per-thread budget control.\n\n**Anti-ping-pong.** The incremental prompt includes recently resolved findings so the LLM doesn't re-raise similar issues. Non-code resolutions (dismissed, acknowledged, rebutted) keep threads open for re-triage on future pushes, but post ack replies to avoid duplicate acknowledgments.\n\n**Context window fitting.** After building the prompt, the pipeline estimates token count and progressively trims file contents (largest first) then diff to fit the model's context window. This prevents API failures on large PRs.\n\n**Finding validation.** All findings are validated against the PR diff regardless of what diff the LLM prompt contained. Line proximity checks (within 20 lines of a changed line) catch hallucinated line numbers and prevent scope creep from rebase noise.\n\n## The codecanary-fix loop\n\nThe `codecanary-fix` Claude skill wraps the review pipeline in a confirm-and-apply loop. The skill calls `codecanary mode --output json` once at startup; the CLI returns one of three modes based on whether an open PR exists for the current branch and whether a CodeCanary workflow file is detected under `.github/workflows/`:\n\n| Mode | PR | Workflow | Findings source | Cycle finalization |\n|---|---|---|---|---|\n| `pr-loop` | yes | yes | `codecanary findings --watch` (bot posts on push) | commit + push; bot re-runs |\n| `local-loop-git` | yes | no | `codecanary review` (local engine) | commit on PR branch, **no push**; operator is asked at session end whether to push accumulated commits |\n| `local-loop-nogit` | no | — | `codecanary review` (local engine) | no commits, no pushes; fixes applied in place |\n\nWorkflow detection is a textual scan for a non-commented `uses: alansikora/codecanary...` step in any workflow file on the current branch. All three modes share the same triage UX (Markdown table, `AskUserQuestion` confirmation). The bot's ack layer (`\u003c!-- codecanary:ack:* --\u003e` markers) handles deferral persistence in `pr-loop`; local modes use an in-memory `DEFERRED_FIX_REFS` set in the skill so operator-skipped findings don't re-surface within a session. `pr-loop` failures (GHA broken, `conclusion: failure`) never silently fall back to a local mode — the operator is asked to investigate.\n", + "internal/review/github.go": "package review\n\nimport (\n\t\"bytes\"\n\t\"encoding/json\"\n\t\"fmt\"\n\t\"os\"\n\t\"os/exec\"\n\t\"path/filepath\"\n\t\"regexp\"\n\t\"strconv\"\n\t\"strings\"\n\n\t\"github.com/bmatcuk/doublestar/v4\"\n)\n\n// runGH runs a gh command and returns stdout; on failure, the returned error\n// includes gh's stderr so upstream API messages surface in logs instead of just\n// \"exit status 1\".\nfunc runGH(label string, args ...string) ([]byte, error) {\n\tcmd := exec.Command(\"gh\", args...)\n\tvar stderr bytes.Buffer\n\tcmd.Stderr = \u0026stderr\n\tout, err := cmd.Output()\n\tif err != nil {\n\t\tmsg := strings.TrimSpace(stderr.String())\n\t\tif msg == \"\" {\n\t\t\treturn nil, fmt.Errorf(\"%s: %w\", label, err)\n\t\t}\n\t\treturn nil, fmt.Errorf(\"%s: %s: %w\", label, msg, err)\n\t}\n\treturn out, nil\n}\n\n// parseRepoSlug splits a \"owner/name\" repository slug into its two parts.\nfunc parseRepoSlug(repo string) (owner, name string, err error) {\n\tparts := strings.SplitN(repo, \"/\", 2)\n\tif len(parts) != 2 {\n\t\treturn \"\", \"\", fmt.Errorf(\"invalid repo format %q, expected owner/name\", repo)\n\t}\n\treturn parts[0], parts[1], nil\n}\n\n// MaxFindingProximity is the maximum number of lines a finding may be from the\n// nearest changed line in the PR diff. Findings beyond this distance are dropped\n// (runner.go) or demoted from inline to body (PostReview). This enforces review\n// scope — keeping findings anchored to the PR's actual changes — and catches\n// hallucinated line numbers. A single constant ensures both checks stay in sync.\nconst MaxFindingProximity = 20\n\n// HTML comment markers for embedding and detecting review data.\n// Dual prefixes support both current (codecanary) and legacy (clanopy) markers.\nvar reviewMarkerPrefixes = []string{\"\u003c!-- codecanary:review \", \"\u003c!-- clanopy:review \"}\n\nconst (\n\treviewMarkerSuffix = \" --\u003e\"\n\tfindingMarkerPrefix = \"\u003c!-- codecanary:finding \"\n\tackMarkerPrefix = \"\u003c!-- codecanary:ack:\"\n\tlegacyAckPrefix = \"\u003c!-- clanopy:ack:\"\n)\n\n// PRData holds PR metadata and diff.\ntype PRData struct {\n\tNumber int\n\tTitle string\n\tBody string\n\tAuthor string\n\tBaseBranch string\n\tHeadBranch string\n\tDiff string\n\tFullDiff string // unfiltered diff for finding validation (set by prepareReview)\n\tFiles []string\n\tFileContents map[string]string // path -\u003e full file content\n}\n\n// ValidationDiff returns the unfiltered diff for finding validation. When\n// files were skipped during prepareReview, FullDiff holds the original diff\n// while Diff is filtered for the LLM prompt.\nfunc (pr *PRData) ValidationDiff() string {\n\tif pr.FullDiff != \"\" {\n\t\treturn pr.FullDiff\n\t}\n\treturn pr.Diff\n}\n\n// ghPRView is the JSON shape returned by gh pr view.\ntype ghPRView struct {\n\tTitle string `json:\"title\"`\n\tBody string `json:\"body\"`\n\tAuthor struct {\n\t\tLogin string `json:\"login\"`\n\t} `json:\"author\"`\n\tBaseRefName string `json:\"baseRefName\"`\n\tHeadRefName string `json:\"headRefName\"`\n\tFiles []struct {\n\t\tPath string `json:\"path\"`\n\t} `json:\"files\"`\n}\n\n// FetchPR gets PR metadata and diff using the gh CLI.\nfunc FetchPR(repo string, number int) (*PRData, error) {\n\tnumStr := fmt.Sprintf(\"%d\", number)\n\n\t// Fetch PR metadata as JSON.\n\tviewOut, err := runGH(\"gh pr view\", \"pr\", \"view\", numStr,\n\t\t\"--repo\", repo,\n\t\t\"--json\", \"title,body,author,baseRefName,headRefName,files\",\n\t)\n\tif err != nil {\n\t\treturn nil, err\n\t}\n\n\tvar view ghPRView\n\tif err := json.Unmarshal(viewOut, \u0026view); err != nil {\n\t\treturn nil, fmt.Errorf(\"parsing gh pr view output: %w\", err)\n\t}\n\n\t// Fetch the diff. This hits GitHub's pull request diff API\n\t// (GET /repos/{owner}/{repo}/pulls/{n} with Accept: application/vnd.github.v3.diff).\n\tdiffOut, err := runGH(\"gh pr diff\", \"pr\", \"diff\", numStr, \"--repo\", repo)\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"%w (GitHub pull request diff API; if this looks like a transient 5xx/timeout, retrying the job may help)\", err)\n\t}\n\n\tfiles := make([]string, len(view.Files))\n\tfor i, f := range view.Files {\n\t\tfiles[i] = f.Path\n\t}\n\n\treturn \u0026PRData{\n\t\tNumber: number,\n\t\tTitle: view.Title,\n\t\tBody: view.Body,\n\t\tAuthor: view.Author.Login,\n\t\tBaseBranch: view.BaseRefName,\n\t\tHeadBranch: view.HeadRefName,\n\t\tDiff: string(diffOut),\n\t\tFiles: files,\n\t}, nil\n}\n\n// PostComment posts a review comment on a PR.\nfunc PostComment(repo string, number int, body string) error {\n\tnumStr := fmt.Sprintf(\"%d\", number)\n\tcmd := exec.Command(\"gh\", \"pr\", \"comment\", numStr,\n\t\t\"--repo\", repo,\n\t\t\"--body\", body,\n\t)\n\tif out, err := cmd.CombinedOutput(); err != nil {\n\t\treturn fmt.Errorf(\"gh pr comment: %w\\n%s\", err, string(out))\n\t}\n\treturn nil\n}\n\n// reviewPayload is the JSON structure for the GitHub PR review API.\ntype reviewPayload struct {\n\tEvent string `json:\"event\"`\n\tBody string `json:\"body\"`\n\tComments []reviewComment `json:\"comments\"`\n\tCommitID string `json:\"commit_id,omitempty\"`\n}\n\n// reviewComment is a single inline comment in a PR review.\ntype reviewComment struct {\n\tPath string `json:\"path\"`\n\tLine int `json:\"line\"`\n\tBody string `json:\"body\"`\n}\n\n// countDiffLines counts the number of added and removed lines in a unified diff.\nfunc countDiffLines(diff string) (added, removed int) {\n\tfor _, line := range strings.Split(diff, \"\\n\") {\n\t\tif strings.HasPrefix(line, \"+++\") || strings.HasPrefix(line, \"---\") {\n\t\t\tcontinue\n\t\t}\n\t\tif strings.HasPrefix(line, \"+\") {\n\t\t\tadded++\n\t\t} else if strings.HasPrefix(line, \"-\") {\n\t\t\tremoved++\n\t\t}\n\t}\n\treturn\n}\n\n// diffLineMap maps each file to its sorted list of valid line numbers from the diff.\ntype diffLineMap map[string][]int\n\n// parseDiffLines extracts valid line numbers per file from a unified diff.\nfunc parseDiffLines(diff string) diffLineMap {\n\tvalid := make(diffLineMap)\n\tvar currentFile string\n\tvar lineNum int\n\n\tfor _, line := range strings.Split(diff, \"\\n\") {\n\t\tif strings.HasPrefix(line, \"+++ b/\") {\n\t\t\tcurrentFile = line[6:]\n\t\t\tcontinue\n\t\t}\n\t\tif strings.HasPrefix(line, \"@@ \") {\n\t\t\t// Parse hunk header: @@ -old,count +new,count @@\n\t\t\tif idx := strings.Index(line, \"+\"); idx \u003e= 0 {\n\t\t\t\trest := line[idx+1:]\n\t\t\t\tif comma := strings.IndexAny(rest, \", \"); comma \u003e= 0 {\n\t\t\t\t\trest = rest[:comma]\n\t\t\t\t}\n\t\t\t\t_, _ = fmt.Sscanf(rest, \"%d\", \u0026lineNum)\n\t\t\t}\n\t\t\tcontinue\n\t\t}\n\t\tif currentFile == \"\" || lineNum == 0 {\n\t\t\tcontinue\n\t\t}\n\t\tif strings.HasPrefix(line, \"-\") {\n\t\t\tcontinue\n\t\t}\n\t\tif strings.HasPrefix(line, \"+\") {\n\t\t\tvalid[currentFile] = append(valid[currentFile], lineNum)\n\t\t\tlineNum++\n\t\t\tcontinue\n\t\t}\n\t\tif strings.HasPrefix(line, \" \") {\n\t\t\tlineNum++\n\t\t\tcontinue\n\t\t}\n\t}\n\treturn valid\n}\n\n// nearestLine returns the closest valid diff line for a file:line pair.\n// Returns the line itself if valid, the nearest valid line in that file,\n// or 0 if the file is not in the diff at all.\nfunc (d diffLineMap) nearestLine(file string, line int) int {\n\tlines, ok := d[file]\n\tif !ok || len(lines) == 0 {\n\t\treturn 0\n\t}\n\tbest := lines[0]\n\tbestDist := abs(line - best)\n\tfor _, l := range lines[1:] {\n\t\tdist := abs(line - l)\n\t\tif dist \u003c bestDist {\n\t\t\tbest = l\n\t\t\tbestDist = dist\n\t\t}\n\t}\n\treturn best\n}\n\nfunc abs(x int) int {\n\tif x \u003c 0 {\n\t\treturn -x\n\t}\n\treturn x\n}\n\n// PostReview posts a PR review with inline comments using the GitHub API.\n// Findings with file and line information become inline comments; others are\n// included in the review body. The summary block is appended to the body so\n// the status dashboard appears on every CodeCanary top-level review.\nfunc PostReview(repo string, prNumber int, result *ReviewResult, diff string, commitSHA string, summary ReviewSummary) error {\n\t// Sort findings by severity before formatting.\n\tsortFindings(result.Findings)\n\n\t// Parse the diff to find valid line positions for inline comments.\n\tvalidLines := parseDiffLines(diff)\n\n\t// A finding can be inlined if its file is in the diff and the nearest\n\t// valid line is within a reasonable distance. Without a bound, findings\n\t// about code far from the diff get silently snapped to unrelated lines.\n\tcanInline := func(f Finding) bool {\n\t\tif f.File == \"\" || f.Line \u003c= 0 {\n\t\t\treturn false\n\t\t}\n\t\tnearest := validLines.nearestLine(f.File, f.Line)\n\t\treturn nearest \u003e 0 \u0026\u0026 abs(f.Line-nearest) \u003c= MaxFindingProximity\n\t}\n\n\tcomments := make([]reviewComment, 0)\n\tfor _, f := range result.Findings {\n\t\tif canInline(f) {\n\t\t\tcomments = append(comments, reviewComment{\n\t\t\t\tPath: f.File,\n\t\t\t\tLine: validLines.nearestLine(f.File, f.Line),\n\t\t\t\tBody: FormatFindingComment(\u0026f),\n\t\t\t})\n\t\t}\n\t}\n\n\tbody := withSummary(FormatReviewBody(result, canInline), summary)\n\n\tpayload := reviewPayload{\n\t\tEvent: \"COMMENT\",\n\t\tBody: body,\n\t\tComments: comments,\n\t\tCommitID: commitSHA,\n\t}\n\n\tpayloadJSON, err := json.Marshal(payload)\n\tif err != nil {\n\t\treturn fmt.Errorf(\"marshaling review payload: %w\", err)\n\t}\n\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn err\n\t}\n\n\t// No fallback — if this fails, the apiError carries stderr and response\n\t// body so the caller surfaces full diagnostics for debugging.\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, name, prNumber)\n\t_, err = ghAPIPOST(apiPath, payloadJSON)\n\treturn err\n}\n\n// ReviewCommitStatusContext is the commit-status context the review bot\n// writes when a review run completes. Using a stable string lets teams\n// require it in branch protection — \"CodeCanary / review must be green\n// before merge.\" Matches the context used by `codecanary signoff`, so\n// both the bot (on PRs with the workflow) and a local signoff can\n// satisfy the same required check.\nconst ReviewCommitStatusContext = \"CodeCanary / review\"\n\n// PostReviewCommitStatus POSTs a commit status on the given SHA summarising\n// the review outcome. state is \"success\" when all findings are resolved or\n// handled by the author, \"failure\" otherwise. Description is shown on the\n// PR checks list, so keep it short and human-readable. Errors are returned\n// so the caller can log them — a failed status post should not abort the\n// larger Publish flow, since the review itself has already been posted.\nfunc PostReviewCommitStatus(repo, sha, state, description string) error {\n\tif sha == \"\" {\n\t\treturn fmt.Errorf(\"commit status requires a SHA\")\n\t}\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn err\n\t}\n\tpayload := map[string]string{\n\t\t\"state\": state,\n\t\t\"context\": ReviewCommitStatusContext,\n\t\t\"description\": description,\n\t}\n\tpayloadJSON, err := json.Marshal(payload)\n\tif err != nil {\n\t\treturn fmt.Errorf(\"marshaling commit status payload: %w\", err)\n\t}\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/statuses/%s\", owner, name, sha)\n\t_, err = ghAPIPOST(apiPath, payloadJSON)\n\treturn err\n}\n\n// FetchReviewFromPR extracts cached review data from a PR review's hidden HTML tag.\nfunc FetchReviewFromPR(repo string, prNumber int) (*ReviewResult, error) {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn nil, err\n\t}\n\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, name, prNumber)\n\tout, err := exec.Command(\"gh\", \"api\", apiPath).Output()\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"fetching PR reviews: %w\", err)\n\t}\n\n\tvar reviews []ghReview\n\tif err := json.Unmarshal(out, \u0026reviews); err != nil {\n\t\treturn nil, fmt.Errorf(\"parsing PR reviews: %w\", err)\n\t}\n\n\tprefixes := reviewMarkerPrefixes\n\tconst suffix = \" --\u003e\"\n\n\t// Search from most recent to oldest.\n\tfor i := len(reviews) - 1; i \u003e= 0; i-- {\n\t\tbody := reviews[i].Body\n\t\tfor _, prefix := range prefixes {\n\t\t\tidx := strings.Index(body, prefix)\n\t\t\tif idx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tstart := idx + len(prefix)\n\t\t\tendIdx := strings.Index(body[start:], suffix)\n\t\t\tif endIdx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tjsonData := body[start : start+endIdx]\n\t\t\tvar result ReviewResult\n\t\t\tif err := json.Unmarshal([]byte(jsonData), \u0026result); err != nil {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\treturn \u0026result, nil\n\t\t}\n\t}\n\n\treturn nil, fmt.Errorf(\"no review data found in PR #%d reviews\", prNumber)\n}\n\n// FetchFindingFromPR searches all reviews on a PR for a specific fix_ref.\n// Unlike FetchReviewFromPR (which returns the latest review), this searches every\n// review so that fix_ref links from older review rounds still resolve correctly.\nfunc FetchFindingFromPR(repo string, prNumber int, fixRef string) (*Finding, error) {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn nil, err\n\t}\n\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, name, prNumber)\n\tout, err := exec.Command(\"gh\", \"api\", apiPath).Output()\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"fetching PR reviews: %w\", err)\n\t}\n\n\tvar reviews []ghReview\n\tif err := json.Unmarshal(out, \u0026reviews); err != nil {\n\t\treturn nil, fmt.Errorf(\"parsing PR reviews: %w\", err)\n\t}\n\n\tprefixes := reviewMarkerPrefixes\n\tconst suffix = \" --\u003e\"\n\n\tfor _, rev := range reviews {\n\t\tfor _, prefix := range prefixes {\n\t\t\tidx := strings.Index(rev.Body, prefix)\n\t\t\tif idx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tstart := idx + len(prefix)\n\t\t\tendIdx := strings.Index(rev.Body[start:], suffix)\n\t\t\tif endIdx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tvar result ReviewResult\n\t\t\tif err := json.Unmarshal([]byte(rev.Body[start:start+endIdx]), \u0026result); err != nil {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tfor i := range result.Findings {\n\t\t\t\tif result.Findings[i].FixRef == fixRef {\n\t\t\t\t\treturn \u0026result.Findings[i], nil\n\t\t\t\t}\n\t\t\t}\n\t\t}\n\t}\n\n\treturn nil, fmt.Errorf(\"fix_ref %q not found in any review on PR #%d\", fixRef, prNumber)\n}\n\n// DetectRepo gets owner/name from the current git remote.\nfunc DetectRepo() (string, error) {\n\tout, err := exec.Command(\"gh\", \"repo\", \"view\",\n\t\t\"--json\", \"nameWithOwner\",\n\t\t\"--jq\", \".nameWithOwner\",\n\t).Output()\n\tif err != nil {\n\t\treturn \"\", fmt.Errorf(\"gh repo view: %w\", err)\n\t}\n\treturn strings.TrimSpace(string(out)), nil\n}\n\n// DetectPRNumber detects the PR number for the current branch using gh.\n// If repo is non-empty, it is passed as --repo to scope the lookup.\nfunc DetectPRNumber(repo string) (int, error) {\n\targs := []string{\"pr\", \"view\", \"--json\", \"number\", \"--jq\", \".number\"}\n\tif repo != \"\" {\n\t\targs = append(args, \"--repo\", repo)\n\t}\n\tout, err := exec.Command(\"gh\", args...).Output()\n\tif err != nil {\n\t\treturn 0, fmt.Errorf(\"no open pull request found for the current branch\")\n\t}\n\tnum, err := strconv.Atoi(strings.TrimSpace(string(out)))\n\tif err != nil {\n\t\treturn 0, fmt.Errorf(\"unexpected PR number from gh: %w\", err)\n\t}\n\treturn num, nil\n}\n\n// ThreadReply represents a reply to a review thread (i.e. any comment after the first).\ntype ThreadReply struct {\n\tAuthor string\n\tBody string\n}\n\n// ReviewThread represents a review thread from a PR.\ntype ReviewThread struct {\n\tID string\n\tPath string\n\tLine int\n\tBody string\n\tAuthor string // login of the first comment author (the bot for review threads)\n\tOutdated bool // true if GitHub marked the comment position as outdated (code changed)\n\tResolved bool\n\tReplies []ThreadReply\n}\n\n// graphQLThreadsResponse is the JSON shape returned by the review threads query.\ntype graphQLThreadsResponse struct {\n\tData struct {\n\t\tRepository struct {\n\t\t\tPullRequest struct {\n\t\t\t\tReviewThreads struct {\n\t\t\t\t\tNodes []struct {\n\t\t\t\t\t\tID string `json:\"id\"`\n\t\t\t\t\t\tIsResolved bool `json:\"isResolved\"`\n\t\t\t\t\t\tComments struct {\n\t\t\t\t\t\t\tNodes []struct {\n\t\t\t\t\t\t\t\tBody string `json:\"body\"`\n\t\t\t\t\t\t\t\tPath string `json:\"path\"`\n\t\t\t\t\t\t\t\tLine int `json:\"line\"`\n\t\t\t\t\t\t\t\tOriginalLine int `json:\"originalLine\"`\n\t\t\t\t\t\t\t\tOutdated bool `json:\"outdated\"`\n\t\t\t\t\t\t\t\tAuthor struct {\n\t\t\t\t\t\t\t\t\tLogin string `json:\"login\"`\n\t\t\t\t\t\t\t\t} `json:\"author\"`\n\t\t\t\t\t\t\t} `json:\"nodes\"`\n\t\t\t\t\t\t} `json:\"comments\"`\n\t\t\t\t\t} `json:\"nodes\"`\n\t\t\t\t} `json:\"reviewThreads\"`\n\t\t\t} `json:\"pullRequest\"`\n\t\t} `json:\"repository\"`\n\t} `json:\"data\"`\n}\n\n// FetchReviewThreads gets all review threads from a PR via GraphQL.\nfunc FetchReviewThreads(repo string, prNumber int) ([]ReviewThread, error) {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn nil, err\n\t}\n\n\tquery := `query($owner:String!,$name:String!,$pr:Int!){\n repository(owner:$owner,name:$name){\n pullRequest(number:$pr){\n reviewThreads(first:100){\n nodes{\n id\n isResolved\n comments(first:100){\n nodes{body path line originalLine outdated author{login}}\n }\n }\n }\n }\n }\n}`\n\n\tcmd := exec.Command(\"gh\", \"api\", \"graphql\",\n\t\t\"-f\", \"query=\"+query,\n\t\t\"-f\", fmt.Sprintf(\"owner=%s\", owner),\n\t\t\"-f\", fmt.Sprintf(\"name=%s\", name),\n\t\t\"-F\", fmt.Sprintf(\"pr=%d\", prNumber),\n\t)\n\tout, err := cmd.Output()\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"gh api graphql: %w\", err)\n\t}\n\n\tvar resp graphQLThreadsResponse\n\tif err := json.Unmarshal(out, \u0026resp); err != nil {\n\t\treturn nil, fmt.Errorf(\"parsing graphql response: %w\", err)\n\t}\n\n\tvar threads []ReviewThread\n\tfor _, node := range resp.Data.Repository.PullRequest.ReviewThreads.Nodes {\n\t\tif len(node.Comments.Nodes) == 0 {\n\t\t\tcontinue\n\t\t}\n\t\tcomment := node.Comments.Nodes[0]\n\n\t\t// Filter to review threads only (new marker + legacy markers for backward compat).\n\t\tif !strings.Contains(comment.Body, findingMarkerPrefix) \u0026\u0026\n\t\t\t!strings.Contains(comment.Body, \"codecanary fix\") \u0026\u0026\n\t\t\t!strings.Contains(comment.Body, \"clanopy fix\") {\n\t\t\tcontinue\n\t\t}\n\n\t\tvar replies []ThreadReply\n\t\tfor _, c := range node.Comments.Nodes[1:] {\n\t\t\treplies = append(replies, ThreadReply{\n\t\t\t\tAuthor: c.Author.Login,\n\t\t\t\tBody: c.Body,\n\t\t\t})\n\t\t}\n\n\t\tline := comment.Line\n\t\tif comment.Outdated \u0026\u0026 line == 0 \u0026\u0026 comment.OriginalLine \u003e 0 {\n\t\t\tline = comment.OriginalLine\n\t\t}\n\n\t\tthreads = append(threads, ReviewThread{\n\t\t\tID: node.ID,\n\t\t\tPath: comment.Path,\n\t\t\tLine: line,\n\t\t\tBody: comment.Body,\n\t\t\tAuthor: comment.Author.Login,\n\t\t\tOutdated: comment.Outdated,\n\t\t\tResolved: node.IsResolved,\n\t\t\tReplies: replies,\n\t\t})\n\t}\n\n\treturn threads, nil\n}\n\n// ResolveThread resolves a review thread via GraphQL mutation.\nfunc ResolveThread(threadID string) error {\n\tcmd := exec.Command(\"gh\", \"api\", \"graphql\",\n\t\t\"-f\", \"query=mutation($threadId:ID!){resolveReviewThread(input:{threadId:$threadId}){thread{isResolved}}}\",\n\t\t\"-f\", fmt.Sprintf(\"threadId=%s\", threadID),\n\t)\n\tif out, err := cmd.CombinedOutput(); err != nil {\n\t\treturn fmt.Errorf(\"gh api graphql resolve: %w\\n%s\", err, string(out))\n\t}\n\treturn nil\n}\n\n// ReplyToThread posts a reply on a review thread via GraphQL.\nfunc ReplyToThread(threadID, body string) error {\n\tcmd := exec.Command(\"gh\", \"api\", \"graphql\",\n\t\t\"-f\", \"query=mutation($threadId:ID!,$body:String!){addPullRequestReviewThreadReply(input:{pullRequestReviewThreadId:$threadId,body:$body}){comment{id}}}\",\n\t\t\"-f\", fmt.Sprintf(\"threadId=%s\", threadID),\n\t\t\"-f\", fmt.Sprintf(\"body=%s\", body),\n\t)\n\tif out, err := cmd.CombinedOutput(); err != nil {\n\t\treturn fmt.Errorf(\"gh api graphql reply: %w\\n%s\", err, string(out))\n\t}\n\treturn nil\n}\n\n// ghReview is the JSON shape for a PR review from the REST API.\ntype ghReview struct {\n\tID int64 `json:\"id\"`\n\tNodeID string `json:\"node_id\"`\n\tBody string `json:\"body\"`\n}\n\n// LatestCodecanaryReview captures the identity and body of the most recent\n// CodeCanary top-level review, together with the commit SHA embedded in its\n// hidden marker. Returned by FetchLatestCodecanaryReview for the\n// \"edit if same SHA, otherwise post new\" publish decision.\ntype LatestCodecanaryReview struct {\n\tID int64\n\tSHA string\n\tBody string\n}\n\n// FetchLatestCodecanaryReview returns the most recent top-level review on\n// the PR that carries a CodeCanary marker, along with the commit SHA\n// embedded in that marker. Returns (nil, nil) when no matching review\n// exists. A non-nil error is returned only for transport/parsing failures.\nfunc FetchLatestCodecanaryReview(repo string, prNumber int) (*LatestCodecanaryReview, error) {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn nil, err\n\t}\n\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, name, prNumber)\n\tout, err := exec.Command(\"gh\", \"api\", apiPath).Output()\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"fetching PR reviews: %w\", err)\n\t}\n\n\tvar reviews []ghReview\n\tif err := json.Unmarshal(out, \u0026reviews); err != nil {\n\t\treturn nil, fmt.Errorf(\"parsing PR reviews: %w\", err)\n\t}\n\n\t// Walk newest → oldest so we return the latest CodeCanary review.\n\tfor i := len(reviews) - 1; i \u003e= 0; i-- {\n\t\trev := reviews[i]\n\t\tfor _, prefix := range reviewMarkerPrefixes {\n\t\t\tidx := strings.Index(rev.Body, prefix)\n\t\t\tif idx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tstart := idx + len(prefix)\n\t\t\tendIdx := strings.Index(rev.Body[start:], reviewMarkerSuffix)\n\t\t\tif endIdx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tjsonData := rev.Body[start : start+endIdx]\n\t\t\tvar payload struct {\n\t\t\t\tSHA string `json:\"sha\"`\n\t\t\t}\n\t\t\t// Ignore unmarshal errors so old reviews that embed a richer\n\t\t\t// ReviewResult (also containing a \"sha\" field) still match.\n\t\t\t_ = json.Unmarshal([]byte(jsonData), \u0026payload)\n\t\t\treturn \u0026LatestCodecanaryReview{\n\t\t\t\tID: rev.ID,\n\t\t\t\tSHA: payload.SHA,\n\t\t\t\tBody: rev.Body,\n\t\t\t}, nil\n\t\t}\n\t}\n\n\treturn nil, nil\n}\n\n// UpdateReviewBody replaces the body of an existing PR review via the\n// GitHub REST API. Used when a reply-only run (or a duplicate synchronize\n// webhook on the same HEAD) needs to refresh the latest CodeCanary review's\n// counts instead of posting a new top-level comment.\nfunc UpdateReviewBody(repo string, prNumber int, reviewID int64, body string) error {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn err\n\t}\n\n\tpayload, err := json.Marshal(struct {\n\t\tBody string `json:\"body\"`\n\t}{Body: body})\n\tif err != nil {\n\t\treturn fmt.Errorf(\"marshaling update payload: %w\", err)\n\t}\n\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews/%d\", owner, name, prNumber, reviewID)\n\t_, err = ghAPIRequest(\"PUT\", apiPath, payload)\n\treturn err\n}\n\n// ghAPIRequest is a generic gh api wrapper that preserves the temp-file\n// pattern used by ghAPIPOST (avoids stdin pipe truncation on large bodies).\nfunc ghAPIRequest(method, apiPath string, payloadJSON []byte) ([]byte, error) {\n\tdir, err := os.MkdirTemp(\"\", \"codecanary-*\")\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"creating temp dir: %w\", err)\n\t}\n\tdefer func() { _ = os.RemoveAll(dir) }()\n\n\ttmpFile, err := os.CreateTemp(dir, \"payload.json\")\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"creating temp file: %w\", err)\n\t}\n\tif _, err := tmpFile.Write(payloadJSON); err != nil {\n\t\t_ = tmpFile.Close()\n\t\treturn nil, fmt.Errorf(\"writing payload to temp file: %w\", err)\n\t}\n\t_ = tmpFile.Close()\n\n\tcmd := exec.Command(\"gh\", \"api\", apiPath, \"--method\", method, \"--input\", tmpFile.Name())\n\tvar stdout, stderr bytes.Buffer\n\tcmd.Stdout = \u0026stdout\n\tcmd.Stderr = \u0026stderr\n\tif err := cmd.Run(); err != nil {\n\t\treturn stdout.Bytes(), \u0026apiError{Err: err, Stderr: stderr.String(), Response: stdout.String()}\n\t}\n\treturn stdout.Bytes(), nil\n}\n\n// FetchPreviousReviewSHA gets the SHA from the last review's hidden data.\nfunc FetchPreviousReviewSHA(repo string, prNumber int) string {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn \"\"\n\t}\n\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, name, prNumber)\n\tout, err := exec.Command(\"gh\", \"api\", apiPath).Output()\n\tif err != nil {\n\t\treturn \"\"\n\t}\n\n\tvar reviews []ghReview\n\tif err := json.Unmarshal(out, \u0026reviews); err != nil {\n\t\treturn \"\"\n\t}\n\n\tprefixes := reviewMarkerPrefixes\n\tconst suffix = \" --\u003e\"\n\n\t// Search from most recent to oldest.\n\tfor i := len(reviews) - 1; i \u003e= 0; i-- {\n\t\tbody := reviews[i].Body\n\t\tfor _, prefix := range prefixes {\n\t\t\tidx := strings.Index(body, prefix)\n\t\t\tif idx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tstart := idx + len(prefix)\n\t\t\tendIdx := strings.Index(body[start:], suffix)\n\t\t\tif endIdx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tjsonData := body[start : start+endIdx]\n\t\t\tvar result ReviewResult\n\t\t\tif err := json.Unmarshal([]byte(jsonData), \u0026result); err != nil {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tif result.SHA != \"\" {\n\t\t\t\treturn result.SHA\n\t\t\t}\n\t\t}\n\t}\n\n\treturn \"\"\n}\n\n// FilesFromDiff extracts the list of file paths touched in a unified diff.\nfunc FilesFromDiff(diff string) []string {\n\tvar files []string\n\tseen := make(map[string]bool)\n\tfor _, line := range strings.Split(diff, \"\\n\") {\n\t\tif strings.HasPrefix(line, \"+++ b/\") {\n\t\t\tpath := strings.TrimRight(line[6:], \"\\r\")\n\t\t\tif !seen[path] {\n\t\t\t\tseen[path] = true\n\t\t\t\tfiles = append(files, path)\n\t\t\t}\n\t\t}\n\t}\n\treturn files\n}\n\n// ScopeDiffToFiles filters a unified diff to only include hunks for files in\n// the allowed set. This prevents rebase noise (main-branch changes) from\n// leaking into incremental reviews.\nfunc ScopeDiffToFiles(diff string, allowedFiles map[string]bool) string {\n\tif len(allowedFiles) == 0 {\n\t\treturn diff\n\t}\n\n\tlines := strings.Split(diff, \"\\n\")\n\tvar result []string\n\tblockStart := -1\n\tblockAllowed := false\n\n\tfor i := 0; i \u003c len(lines); i++ {\n\t\tif strings.HasPrefix(lines[i], \"diff --git\") {\n\t\t\t// Flush previous block if allowed.\n\t\t\tif blockStart \u003e= 0 \u0026\u0026 blockAllowed {\n\t\t\t\tresult = append(result, lines[blockStart:i]...)\n\t\t\t}\n\t\t\tblockStart = i\n\t\t\tblockAllowed = false\n\n\t\t\t// Look ahead for +++ b/\u003cpath\u003e to determine if this block is allowed.\n\t\t\tfor j := i + 1; j \u003c len(lines) \u0026\u0026 !strings.HasPrefix(lines[j], \"diff --git\"); j++ {\n\t\t\t\tif strings.HasPrefix(lines[j], \"+++ b/\") {\n\t\t\t\t\tpath := strings.TrimRight(lines[j][6:], \"\\r\")\n\t\t\t\t\tif allowedFiles[path] {\n\t\t\t\t\t\tblockAllowed = true\n\t\t\t\t\t}\n\t\t\t\t\tbreak\n\t\t\t\t}\n\t\t\t}\n\t\t\tcontinue\n\t\t}\n\t}\n\n\t// Flush last block.\n\tif blockStart \u003e= 0 \u0026\u0026 blockAllowed {\n\t\tresult = append(result, lines[blockStart:]...)\n\t}\n\n\tif len(result) == 0 {\n\t\treturn \"\"\n\t}\n\treturn strings.Join(result, \"\\n\")\n}\n\n// validSHA matches a full-length lowercase hex Git SHA.\nvar validSHA = regexp.MustCompile(`^[0-9a-f]{40}$`)\n\n// GetIncrementalDiff gets the diff since a given SHA.\nfunc GetIncrementalDiff(baseSHA string) (string, error) {\n\tif !validSHA.MatchString(baseSHA) {\n\t\treturn \"\", fmt.Errorf(\"invalid SHA format: %q\", baseSHA)\n\t}\n\tout, err := exec.Command(\"git\", \"diff\", baseSHA+\"..HEAD\").Output()\n\tif err != nil {\n\t\treturn \"\", fmt.Errorf(\"git diff: %w\", err)\n\t}\n\treturn string(out), nil\n}\n\n// PostCleanReview posts a review when the first review finds no issues. The\n// commitSHA is embedded in a hidden marker so future runs treat it as the\n// baseline for incremental reviews, avoiding a redundant full re-review on the\n// next push.\nfunc PostCleanReview(repo string, prNumber int, commitSHA string, summary ReviewSummary) error {\n\treturn postSimpleReview(repo, prNumber, buildCleanReviewBody(commitSHA, summary))\n}\n\n// PostAllClearReview posts a review when all previous findings have been\n// resolved. If minimizeFailed is true, a note is appended warning about\n// visible old reviews. The commitSHA is embedded in a hidden marker so future\n// runs treat it as the baseline for incremental reviews; without it, the next\n// push would fall back to reviewing the entire PR again.\nfunc PostAllClearReview(repo string, prNumber int, commitSHA string, minimizeFailed bool, summary ReviewSummary) error {\n\treturn postSimpleReview(repo, prNumber, buildAllClearReviewBody(commitSHA, minimizeFailed, summary))\n}\n\n// PostActivityReview posts a review when no new findings were raised but\n// there is cycle activity worth surfacing (dismissals, acknowledgments,\n// rebuttals, still-open threads). This keeps every commit push producing a\n// visible top-level status comment instead of silently logging.\nfunc PostActivityReview(repo string, prNumber int, commitSHA string, summary ReviewSummary) error {\n\treturn postSimpleReview(repo, prNumber, buildActivityReviewBody(commitSHA, summary))\n}\n\n// buildCleanReviewBody renders the full Markdown body posted by\n// PostCleanReview. Split out from the poster so tests can assert the exact\n// string that lands on GitHub without having to mock gh.\nfunc buildCleanReviewBody(commitSHA string, summary ReviewSummary) string {\n\treturn withSummary(\"CodeCanary reviewed this PR \\u2014 no issues found.\", summary) + embedBaselineMarker(commitSHA)\n}\n\n// buildAllClearReviewBody renders the full Markdown body posted by\n// PostAllClearReview. Split out for the same reason as buildCleanReviewBody.\nfunc buildAllClearReviewBody(commitSHA string, minimizeFailed bool, summary ReviewSummary) string {\n\tbody := \"## \\U0001F425 CodeCanary\\n\\n\\u2705 All previous findings have been addressed. No new issues found. \\u2728\"\n\tif minimizeFailed {\n\t\tbody += \"\\n\\n\u003e \\u26A0\\uFE0F Some previous review comments could not be minimized and may still be visible.\"\n\t}\n\treturn withSummary(body, summary) + embedBaselineMarker(commitSHA)\n}\n\n// buildActivityReviewBody renders the body for a commit push that raised no\n// new findings but has cycle activity (dismissals/acknowledgments/rebuttals\n// or still-open threads carried forward).\nfunc buildActivityReviewBody(commitSHA string, summary ReviewSummary) string {\n\tbody := \"## \\U0001F425 CodeCanary\\n\\nReviewed this push \\u2014 no new issues found.\"\n\treturn withSummary(body, summary) + embedBaselineMarker(commitSHA)\n}\n\n// withSummary appends the status summary block to a review body. The block\n// is skipped when the summary has no non-zero counts, so existing clean/\n// all-clear bodies render identically when nothing happened.\nfunc withSummary(body string, summary ReviewSummary) string {\n\tblock := renderSummaryBlock(summary)\n\tif block == \"\" {\n\t\treturn body\n\t}\n\treturn body + block\n}\n\n// embedBaselineMarker returns a hidden HTML comment containing the commitSHA\n// so FetchPreviousReviewSHA can use this review as the incremental baseline.\n// Returns an empty string if commitSHA is empty (local mode, dry run). The\n// marker carries only the SHA — FetchPreviousReviewSHA is the sole reader and\n// it only needs that field.\nfunc embedBaselineMarker(commitSHA string) string {\n\tif commitSHA == \"\" {\n\t\treturn \"\"\n\t}\n\tdata, err := json.Marshal(struct {\n\t\tSHA string `json:\"sha\"`\n\t}{SHA: commitSHA})\n\tif err != nil {\n\t\treturn \"\"\n\t}\n\treturn fmt.Sprintf(\"\\n%s%s%s\\n\", reviewMarkerPrefixes[0], string(data), reviewMarkerSuffix)\n}\n\nfunc postSimpleReview(repo string, prNumber int, body string) error {\n\towner, repoName, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn err\n\t}\n\n\tpayload := reviewPayload{\n\t\tEvent: \"COMMENT\",\n\t\tBody: body,\n\t\tComments: make([]reviewComment, 0),\n\t}\n\n\tpayloadJSON, err := json.Marshal(payload)\n\tif err != nil {\n\t\treturn fmt.Errorf(\"marshaling payload: %w\", err)\n\t}\n\n\t_, err = ghAPIPOST(fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, repoName, prNumber), payloadJSON)\n\treturn err\n}\n\n// apiError is returned by ghAPIPOST so callers can inspect the stderr output\n// from gh (which contains the HTTP status line) separately from the response body.\ntype apiError struct {\n\tErr error\n\tStderr string\n\tResponse string\n}\n\nfunc (e *apiError) Error() string {\n\treturn fmt.Sprintf(\"gh api: %v\\nstderr: %s\\nresponse: %s\", e.Err, e.Stderr, e.Response)\n}\n\nfunc (e *apiError) Unwrap() error { return e.Err }\n\n// ghAPIPOST sends a JSON payload to the GitHub API via gh, using a temp file\n// to avoid stdin pipe issues that can cause \"unexpected end of JSON input\"\n// errors on large payloads.\nfunc ghAPIPOST(apiPath string, payloadJSON []byte) ([]byte, error) {\n\tdir, err := os.MkdirTemp(\"\", \"codecanary-*\")\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"creating temp dir: %w\", err)\n\t}\n\tdefer func() { _ = os.RemoveAll(dir) }()\n\n\ttmpFile, err := os.CreateTemp(dir, \"payload.json\")\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"creating temp file: %w\", err)\n\t}\n\n\tif _, err := tmpFile.Write(payloadJSON); err != nil {\n\t\t_ = tmpFile.Close()\n\t\treturn nil, fmt.Errorf(\"writing payload to temp file: %w\", err)\n\t}\n\t_ = tmpFile.Close()\n\n\tcmd := exec.Command(\"gh\", \"api\", apiPath, \"--method\", \"POST\", \"--input\", tmpFile.Name())\n\tvar stdout, stderr bytes.Buffer\n\tcmd.Stdout = \u0026stdout\n\tcmd.Stderr = \u0026stderr\n\tif err := cmd.Run(); err != nil {\n\t\treturn stdout.Bytes(), \u0026apiError{Err: err, Stderr: stderr.String(), Response: stdout.String()}\n\t}\n\treturn stdout.Bytes(), nil\n}\n\n// FindReviewNodeIDs returns the node_ids of all reviews on a PR.\nfunc FindReviewNodeIDs(repo string, prNumber int) ([]string, error) {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn nil, err\n\t}\n\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, name, prNumber)\n\tout, err := exec.Command(\"gh\", \"api\", apiPath).Output()\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"fetching PR reviews: %w\", err)\n\t}\n\n\tvar reviews []ghReview\n\tif err := json.Unmarshal(out, \u0026reviews); err != nil {\n\t\treturn nil, fmt.Errorf(\"parsing PR reviews: %w\", err)\n\t}\n\n\tprefixes := reviewMarkerPrefixes\n\n\tvar nodeIDs []string\n\tfor _, rev := range reviews {\n\t\tfor _, prefix := range prefixes {\n\t\t\tif strings.Contains(rev.Body, prefix) {\n\t\t\t\tnodeIDs = append(nodeIDs, rev.NodeID)\n\t\t\t\tbreak\n\t\t\t}\n\t\t}\n\t}\n\n\treturn nodeIDs, nil\n}\n\n// threadHeaderLine returns the first non-empty, non-HTML-comment line from\n// a thread body. This is the header line containing severity and finding ID.\nfunc threadHeaderLine(body string) string {\n\tfor _, line := range strings.SplitN(body, \"\\n\", 5) {\n\t\tline = strings.TrimSpace(line)\n\t\tif line == \"\" || strings.HasPrefix(line, \"\u003c!--\") {\n\t\t\tcontinue\n\t\t}\n\t\treturn line\n\t}\n\treturn \"\"\n}\n\n// FindingIDFromThread extracts the finding ID from a thread body.\n// Thread bodies follow the format: 🟠 **bug** — `finding-id`\nfunc FindingIDFromThread(body string) string {\n\tfirstLine := threadHeaderLine(body)\n\t// The separator is \" — `\" (space, em-dash U+2014, space, backtick).\n\tmarker := \" \\u2014 `\"\n\tstart := strings.Index(firstLine, marker)\n\tif start \u003c 0 {\n\t\treturn \"\"\n\t}\n\tstart += len(marker)\n\tend := strings.Index(firstLine[start:], \"`\")\n\tif end \u003c 0 {\n\t\treturn \"\"\n\t}\n\treturn firstLine[start : start+end]\n}\n\n// severityFromThreadBody extracts the severity string from a thread body.\n// Thread bodies follow the format: {icon} **severity** — `id`\nvar threadSeverityRe = regexp.MustCompile(`\\*\\*(\\w+)\\*\\*`)\n\nfunc severityFromThreadBody(body string) string {\n\tfirstLine := threadHeaderLine(body)\n\tif m := threadSeverityRe.FindStringSubmatch(firstLine); len(m) \u003e 1 {\n\t\treturn strings.ToLower(m[1])\n\t}\n\treturn \"warning\"\n}\n\n// findingFromEmbeddedJSON tries to parse a Finding from the JSON embedded in the\n// codecanary:finding HTML comment marker. Returns the finding and true if successful.\nfunc findingFromEmbeddedJSON(body string) (Finding, bool) {\n\tprefix := findingMarkerPrefix\n\tsuffix := reviewMarkerSuffix\n\tstart := strings.Index(body, prefix)\n\tif start \u003c 0 {\n\t\treturn Finding{}, false\n\t}\n\tstart += len(prefix)\n\tend := strings.Index(body[start:], suffix)\n\tif end \u003c 0 {\n\t\treturn Finding{}, false\n\t}\n\traw := body[start : start+end]\n\tif len(raw) == 0 || raw[0] != '{' {\n\t\treturn Finding{}, false\n\t}\n\tvar f Finding\n\tif err := json.Unmarshal([]byte(raw), \u0026f); err != nil {\n\t\treturn Finding{}, false\n\t}\n\treturn f, true\n}\n\n// parseThreadBody extracts the description and suggestion from a thread comment\n// body. This is the fallback parser for older comments that don't embed JSON.\nfunc parseThreadBody(body string) (description, suggestion string) {\n\tlines := strings.Split(body, \"\\n\")\n\n\t// Skip leading HTML markers and the header line (icon **sev** — `id`).\n\tcontentStart := 0\n\tpastHeader := false\n\tfor i, line := range lines {\n\t\ttrimmed := strings.TrimSpace(line)\n\t\tif trimmed == \"\" {\n\t\t\tcontinue\n\t\t}\n\t\tif strings.HasPrefix(trimmed, \"\u003c!--\") {\n\t\t\tcontinue\n\t\t}\n\t\tif !pastHeader {\n\t\t\t// First non-empty non-marker line is the header — skip it.\n\t\t\tpastHeader = true\n\t\t\tcontentStart = i + 1\n\t\t\tcontinue\n\t\t}\n\t\tcontentStart = i\n\t\tbreak\n\t}\n\n\tif contentStart \u003e= len(lines) {\n\t\treturn \"\", \"\"\n\t}\n\n\t// Split content into description and suggestion.\n\tvar descLines []string\n\tvar suggLines []string\n\tinSuggestion := false\n\tsuggestionPrefix := \"\u003e **Suggestion**: \"\n\n\tfor _, line := range lines[contentStart:] {\n\t\ttrimmed := strings.TrimSpace(line)\n\t\tif strings.HasPrefix(trimmed, suggestionPrefix) {\n\t\t\tinSuggestion = true\n\t\t\tsuggLines = append(suggLines, strings.TrimPrefix(trimmed, suggestionPrefix))\n\t\t\tcontinue\n\t\t}\n\t\tif inSuggestion {\n\t\t\t// Continuation of suggestion (blockquote lines).\n\t\t\tif strings.HasPrefix(trimmed, \"\u003e \") {\n\t\t\t\tsuggLines = append(suggLines, strings.TrimPrefix(trimmed, \"\u003e \"))\n\t\t\t} else if trimmed == \"\u003e\" {\n\t\t\t\tsuggLines = append(suggLines, \"\")\n\t\t\t} else {\n\t\t\t\tsuggLines = append(suggLines, line)\n\t\t\t}\n\t\t\tcontinue\n\t\t}\n\t\tdescLines = append(descLines, line)\n\t}\n\n\tdescription = strings.TrimSpace(strings.Join(descLines, \"\\n\"))\n\tsuggestion = strings.TrimSpace(strings.Join(suggLines, \"\\n\"))\n\treturn description, suggestion\n}\n\n// FindingFromThread extracts a Finding from a ReviewThread. It first tries to\n// parse the embedded JSON (new format), then falls back to body parsing.\nfunc FindingFromThread(t ReviewThread) Finding {\n\t// Try embedded JSON first (lossless roundtrip).\n\tif f, ok := findingFromEmbeddedJSON(t.Body); ok {\n\t\tf.File = t.Path\n\t\tf.Line = t.Line\n\t\tf.Status = \"still open\"\n\t\treturn f\n\t}\n\n\t// Fallback: parse the markdown body.\n\tdesc, suggestion := parseThreadBody(t.Body)\n\treturn Finding{\n\t\tID: FindingIDFromThread(t.Body),\n\t\tFile: t.Path,\n\t\tLine: t.Line,\n\t\tSeverity: severityFromThreadBody(t.Body),\n\t\tTitle: firstSentence(desc),\n\t\tDescription: desc,\n\t\tSuggestion: suggestion,\n\t\tStatus: \"still open\",\n\t}\n}\n\n// firstSentence returns the first sentence (or first line) of text as a title.\nfunc firstSentence(text string) string {\n\tif text == \"\" {\n\t\treturn \"\"\n\t}\n\t// Use first line.\n\tline := text\n\tif idx := strings.IndexByte(text, '\\n'); idx \u003e= 0 {\n\t\tline = text[:idx]\n\t}\n\tline = strings.TrimSpace(line)\n\t// Truncate to 120 runes if very long.\n\trunes := []rune(line)\n\tif len(runes) \u003e 120 {\n\t\treturn string(runes[:117]) + \"...\"\n\t}\n\treturn line\n}\n\n// ReviewInfo represents a review with its node ID and finding IDs.\ntype ReviewInfo struct {\n\tNodeID string\n\tFindingIDs []string\n}\n\n// FindReviews returns reviews with their parsed finding IDs.\nfunc FindReviews(repo string, prNumber int) ([]ReviewInfo, error) {\n\towner, name, err := parseRepoSlug(repo)\n\tif err != nil {\n\t\treturn nil, err\n\t}\n\n\tapiPath := fmt.Sprintf(\"repos/%s/%s/pulls/%d/reviews\", owner, name, prNumber)\n\tout, err := exec.Command(\"gh\", \"api\", apiPath).Output()\n\tif err != nil {\n\t\treturn nil, fmt.Errorf(\"fetching PR reviews: %w\", err)\n\t}\n\n\tvar reviews []ghReview\n\tif err := json.Unmarshal(out, \u0026reviews); err != nil {\n\t\treturn nil, fmt.Errorf(\"parsing PR reviews: %w\", err)\n\t}\n\n\tprefixes := reviewMarkerPrefixes\n\tconst suffix = \" --\u003e\"\n\n\tvar result []ReviewInfo\n\tfor _, rev := range reviews {\n\t\tfor _, prefix := range prefixes {\n\t\t\tidx := strings.Index(rev.Body, prefix)\n\t\t\tif idx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tstart := idx + len(prefix)\n\t\t\tendIdx := strings.Index(rev.Body[start:], suffix)\n\t\t\tif endIdx \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tjsonData := rev.Body[start : start+endIdx]\n\t\t\tvar rr ReviewResult\n\t\t\tif err := json.Unmarshal([]byte(jsonData), \u0026rr); err != nil {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\tvar ids []string\n\t\t\tfor _, f := range rr.Findings {\n\t\t\t\tids = append(ids, f.ID)\n\t\t\t}\n\t\t\tresult = append(result, ReviewInfo{\n\t\t\t\tNodeID: rev.NodeID,\n\t\t\t\tFindingIDs: ids,\n\t\t\t})\n\t\t\tbreak\n\t\t}\n\t}\n\n\treturn result, nil\n}\n\n// MinimizeComment hides a comment on GitHub using the minimizeComment GraphQL mutation.\nfunc MinimizeComment(nodeID string) error {\n\tcmd := exec.Command(\"gh\", \"api\", \"graphql\",\n\t\t\"-f\", \"query=mutation($id:ID!){minimizeComment(input:{subjectId:$id,classifier:RESOLVED}){minimizedComment{isMinimized}}}\",\n\t\t\"-F\", \"id=\"+nodeID,\n\t)\n\tif out, err := cmd.CombinedOutput(); err != nil {\n\t\treturn fmt.Errorf(\"gh api graphql minimize: %w\\n%s\", err, string(out))\n\t}\n\treturn nil\n}\n\n// FetchFileContents reads the full contents of changed files from disk.\n// It skips files that are too large, binary, deleted, or match ignore patterns.\n// Returns a map of path-\u003econtent and a list of skipped file paths.\nfunc FetchFileContents(files []string, ignorePatterns []string, maxPerFile, maxTotal int) (map[string]string, []string) {\n\tcontents := make(map[string]string)\n\tvar skipped []string\n\ttotalSize := 0\n\n\tfor _, path := range files {\n\t\t// Check ignore patterns.\n\t\tif matchesIgnore(path, ignorePatterns) {\n\t\t\tskipped = append(skipped, path)\n\t\t\tcontinue\n\t\t}\n\n\t\tdata, err := os.ReadFile(path)\n\t\tif err != nil {\n\t\t\t// File may have been deleted in this PR — skip gracefully.\n\t\t\tcontinue\n\t\t}\n\n\t\t// Skip binary files (null bytes in first 512 bytes).\n\t\tpeek := data\n\t\tif len(peek) \u003e 512 {\n\t\t\tpeek = peek[:512]\n\t\t}\n\t\tif bytes.ContainsRune(peek, 0) {\n\t\t\tskipped = append(skipped, path)\n\t\t\tcontinue\n\t\t}\n\n\t\tsize := len(data)\n\n\t\t// Skip files exceeding per-file limit.\n\t\tif size \u003e maxPerFile {\n\t\t\tskipped = append(skipped, path)\n\t\t\tcontinue\n\t\t}\n\n\t\t// Stop if total budget would be exceeded.\n\t\tif totalSize+size \u003e maxTotal {\n\t\t\tskipped = append(skipped, path)\n\t\t\tcontinue\n\t\t}\n\n\t\tcontents[path] = string(data)\n\t\ttotalSize += size\n\t}\n\n\treturn contents, skipped\n}\n\n// isSetupPR detects whether this is the initial setup PR.\n// Returns true only when a new workflow file referencing codecanary is added AND\n// the PR contains no other files beyond expected setup artifacts (workflow +\n// config), so that PRs bundling real code changes are never silently skipped.\nfunc isSetupPR(diff string, files []string) bool {\n\t// All files must be known setup paths.\n\tfor _, f := range files {\n\t\tif !isSetupFile(f) {\n\t\t\treturn false\n\t\t}\n\t}\n\n\t// At least one newly added workflow file must reference codecanary.\n\tlines := strings.Split(diff, \"\\n\")\n\tfor i := 0; i \u003c len(lines)-1; i++ {\n\t\tif lines[i] != \"--- /dev/null\" {\n\t\t\tcontinue\n\t\t}\n\t\tplusLine := lines[i+1]\n\t\tif !strings.HasPrefix(plusLine, \"+++ b/.github/workflows/\") {\n\t\t\tcontinue\n\t\t}\n\t\tfor j := i + 2; j \u003c len(lines); j++ {\n\t\t\tif strings.HasPrefix(lines[j], \"--- \") || strings.HasPrefix(lines[j], \"diff --git\") {\n\t\t\t\tbreak\n\t\t\t}\n\t\t\tif strings.HasPrefix(lines[j], \"+\") \u0026\u0026 (strings.Contains(lines[j], \"codecanary\") || strings.Contains(lines[j], \"clanopy\")) {\n\t\t\t\treturn true\n\t\t\t}\n\t\t}\n\t}\n\treturn false\n}\n\n// isSetupFile returns true if the file path is a known setup artifact.\nfunc isSetupFile(path string) bool {\n\treturn strings.HasPrefix(path, \".github/workflows/\") ||\n\t\tstrings.HasPrefix(path, \".codecanary/\") || path == \".codecanary.yml\" ||\n\t\tstrings.HasPrefix(path, \".clanopy/\")\n}\n\n// matchesIgnore checks if a path matches any of the ignore glob patterns.\n// Uses doublestar to support ** recursive globs (e.g. \"dist/**\", \"src/**/*.test.*\").\nfunc matchesIgnore(path string, patterns []string) bool {\n\tfor _, pat := range patterns {\n\t\tif matched, _ := doublestar.Match(pat, path); matched {\n\t\t\treturn true\n\t\t}\n\t\t// Also try matching against just the filename.\n\t\tif matched, _ := doublestar.Match(pat, filepath.Base(path)); matched {\n\t\t\treturn true\n\t\t}\n\t}\n\treturn false\n}\n", + "internal/review/platform_github.go": "package review\n\nimport (\n\t\"fmt\"\n\t\"os\"\n\t\"strings\"\n)\n\n// allResolved checks if all review threads have been resolved by code changes.\n// Threads resolved by other reasons (dismissed, acknowledged, rebutted) are kept\n// open and do not count as resolved.\nfunc allResolved(threads []ReviewThread, fixed []fixedThread) bool {\n\tfixedSet := make(map[int]bool, len(fixed))\n\tfor _, f := range fixed {\n\t\tif isTrueResolution(f.Reason) {\n\t\t\tfixedSet[f.Index] = true\n\t\t}\n\t}\n\tfor i := range threads {\n\t\tif !fixedSet[i] {\n\t\t\treturn false\n\t\t}\n\t}\n\treturn true\n}\n\n// resolvedFindingIDs builds the set of finding IDs that are resolved.\n// It combines threads already resolved on GitHub with threads just fixed by code changes.\n// Threads resolved by other reasons (dismissed, acknowledged, rebutted) are not included\n// since they are kept open for re-triage.\nfunc resolvedFindingIDs(allThreads, unresolved []ReviewThread, fixed []fixedThread) map[string]bool {\n\tresolved := make(map[string]bool)\n\tfor _, t := range allThreads {\n\t\tif t.Resolved {\n\t\t\tif id := FindingIDFromThread(t.Body); id != \"\" {\n\t\t\t\tresolved[id] = true\n\t\t\t}\n\t\t}\n\t}\n\tfixedSet := make(map[int]bool, len(fixed))\n\tfor _, f := range fixed {\n\t\tif isTrueResolution(f.Reason) {\n\t\t\tfixedSet[f.Index] = true\n\t\t}\n\t}\n\tfor i, t := range unresolved {\n\t\tif fixedSet[i] {\n\t\t\tif id := FindingIDFromThread(t.Body); id != \"\" {\n\t\t\t\tresolved[id] = true\n\t\t\t}\n\t\t}\n\t}\n\treturn resolved\n}\n\n// minimizeFullyResolvedReviews minimizes reviews whose findings are all resolved.\nfunc minimizeFullyResolvedReviews(repo string, prNumber int, resolvedIDs map[string]bool) {\n\treviews, err := FindReviews(repo, prNumber)\n\tif err != nil {\n\t\tfmt.Fprintf(os.Stderr, \"Warning: could not fetch reviews for minimization: %v\\n\", err)\n\t\treturn\n\t}\n\tminimized := 0\n\tfor _, rev := range reviews {\n\t\tallResolved := len(rev.FindingIDs) \u003e 0\n\t\tfor _, fid := range rev.FindingIDs {\n\t\t\tif !resolvedIDs[fid] {\n\t\t\t\tallResolved = false\n\t\t\t\tbreak\n\t\t\t}\n\t\t}\n\t\tif !allResolved {\n\t\t\tcontinue\n\t\t}\n\t\tif err := MinimizeComment(rev.NodeID); err != nil {\n\t\t\tfmt.Fprintf(os.Stderr, \"Warning: could not minimize review: %v\\n\", err)\n\t\t} else {\n\t\t\tminimized++\n\t\t}\n\t}\n\tif minimized \u003e 0 {\n\t\tfmt.Fprintf(os.Stderr, \"Minimized %d previous review(s)\\n\", minimized)\n\t}\n}\n\n// acknowledgmentMessage returns a reply body for non-code-change resolutions.\n// Each message includes a hidden HTML marker for dedup detection.\nfunc acknowledgmentMessage(reason string) string {\n\tmarker := fmt.Sprintf(\"%s%s --\u003e\", ackMarkerPrefix, reason)\n\tswitch reason {\n\tcase \"dismissed\":\n\t\treturn marker + \"\\nAuthor dismissed this finding. Keeping open \\u2014 will re-check if related code changes.\"\n\tcase \"acknowledged\":\n\t\treturn marker + \"\\nAuthor acknowledged this finding. Keeping open \\u2014 will re-check on future pushes.\"\n\tcase \"rebutted\":\n\t\treturn marker + \"\\nAuthor provided a technical rebuttal. Keeping open \\u2014 will re-check if related code changes.\"\n\tdefault:\n\t\treturn fmt.Sprintf(\"%sunknown --\u003e\", ackMarkerPrefix) + \"\\nFinding acknowledged. Keeping open \\u2014 will re-check on future pushes.\"\n\t}\n}\n\n// hasAcknowledgmentReply reports whether the thread already carries any\n// codecanary ack reply. Dedup is intentionally reason-agnostic: all three\n// ack reasons (dismissed/rebutted/acknowledged) keep the thread open and\n// convey the same outcome to the author, and triage classification across\n// runs is not deterministic — so checking only the same reason let a\n// rebutted ack get followed by a dismissed ack on the next run, stacking\n// two replies on one skip. One ack per thread is enough.\nfunc hasAcknowledgmentReply(t ReviewThread) bool {\n\tfor _, r := range t.Replies {\n\t\tif strings.Contains(r.Body, ackMarkerPrefix) || strings.Contains(r.Body, legacyAckPrefix) {\n\t\t\treturn true\n\t\t}\n\t}\n\treturn false\n}\n\n// GithubPlatform implements ReviewPlatform for GitHub PR mode. It always\n// posts to the PR; \"local preview\" against a GitHub PR is no longer a thing —\n// `codecanary review` without --post uses LocalPlatform.\ntype GithubPlatform struct {\n\tRepo string\n\tPRNumber int\n\tDryRun bool\n}\n\nfunc (g *GithubPlatform) LoadPreviousFindings() ([]ReviewThread, string, int) {\n\tallThreads, err := FetchReviewThreads(g.Repo, g.PRNumber)\n\tif err != nil {\n\t\tfmt.Fprintf(os.Stderr, \"Warning: could not fetch review threads: %v\\n\", err)\n\t\treturn nil, \"\", 0\n\t}\n\tstartIndex := len(allThreads)\n\tvar unresolved []ReviewThread\n\tfor _, t := range allThreads {\n\t\tif !t.Resolved {\n\t\t\tunresolved = append(unresolved, t)\n\t\t}\n\t}\n\tpreviousSHA := FetchPreviousReviewSHA(g.Repo, g.PRNumber)\n\treturn unresolved, previousSHA, startIndex\n}\n\nfunc (g *GithubPlatform) ExcludedAuthor(threads []ReviewThread) string {\n\tif len(threads) \u003e 0 {\n\t\tif login := threads[0].Author; login != \"\" {\n\t\t\treturn login\n\t\t}\n\t\tfmt.Fprintf(os.Stderr, \"Warning: could not determine bot login from thread author\\n\")\n\t}\n\treturn \"\"\n}\n\nfunc (g *GithubPlatform) HandleResolutions(threads []ReviewThread, fixed []fixedThread) {\n\tfor _, f := range fixed {\n\t\tif f.Index \u003c 0 || f.Index \u003e= len(threads) {\n\t\t\tcontinue\n\t\t}\n\t\tt := threads[f.Index]\n\t\tlabel := threadLabel(t)\n\t\tif isTrueResolution(f.Reason) {\n\t\t\t// Post an explanatory comment for file_removed before resolving.\n\t\t\tif f.Reason == \"file_removed\" {\n\t\t\t\tmsg := \"File removed from PR — resolving.\"\n\t\t\t\tif err := ReplyToThread(t.ID, msg); err != nil {\n\t\t\t\t\tfmt.Fprintf(os.Stderr, \" ! %s — failed to post file-removed comment: %v\\n\", label, err)\n\t\t\t\t}\n\t\t\t}\n\t\t\tif err := ResolveThread(t.ID); err != nil {\n\t\t\t\tif strings.Contains(err.Error(), \"Resource not accessible\") {\n\t\t\t\t\tfmt.Fprintf(os.Stderr, \" ~ %s (auto-resolve unavailable: token lacks permission)\\n\", label)\n\t\t\t\t} else {\n\t\t\t\t\tfmt.Fprintf(os.Stderr, \" ! %s — resolved, but failed to update thread: %v\\n\", label, err)\n\t\t\t\t}\n\t\t\t}\n\t\t} else {\n\t\t\tif !hasAcknowledgmentReply(t) {\n\t\t\t\tmsg := acknowledgmentMessage(f.Reason)\n\t\t\t\tif err := ReplyToThread(t.ID, msg); err != nil {\n\t\t\t\t\tfmt.Fprintf(os.Stderr, \" ! %s — failed to post acknowledgment: %v\\n\", label, err)\n\t\t\t\t}\n\t\t\t}\n\t\t}\n\t}\n}\n\nfunc (g *GithubPlatform) Publish(result *ReviewResult, pr *PRData, threads []ReviewThread, fixed []fixedThread) error {\n\tsummary := computeReviewSummary(threads, fixed, result.Findings)\n\n\t// Decide edit-vs-post: if the latest CodeCanary review on the PR carries\n\t// the current HEAD SHA in its marker, this is either a reply-only run or\n\t// a duplicate synchronize webhook — refresh that review in place rather\n\t// than stacking another top-level comment.\n\tlatest, latestErr := FetchLatestCodecanaryReview(g.Repo, g.PRNumber)\n\tif latestErr != nil {\n\t\tfmt.Fprintf(os.Stderr, \"Warning: could not fetch latest review for dedup: %v\\n\", latestErr)\n\t}\n\tif latest != nil \u0026\u0026 result.SHA != \"\" \u0026\u0026 latest.SHA == result.SHA {\n\t\tupdated := replaceSummaryBlock(latest.Body, summary)\n\t\tif updated == latest.Body {\n\t\t\tStderrf(ansiGreen, \"Latest review already current for %s — no update needed\\n\", shortSHA(result.SHA))\n\t\t\treturn nil\n\t\t}\n\t\tif err := UpdateReviewBody(g.Repo, g.PRNumber, latest.ID, updated); err != nil {\n\t\t\treturn fmt.Errorf(\"updating review body: %w\", err)\n\t\t}\n\t\tStderrf(ansiGreen, \"Updated latest review on PR #%d\\n\", g.PRNumber)\n\t\tg.postReviewCommitStatus(result.SHA, summary)\n\t\treturn nil\n\t}\n\n\t// Minimize previous reviews before posting a fresh one.\n\tminimizeFailed := false\n\tif len(threads) \u003e 0 {\n\t\tif nodeIDs, err := FindReviewNodeIDs(g.Repo, g.PRNumber); err == nil {\n\t\t\tif allResolved(threads, fixed) {\n\t\t\t\tfor _, nodeID := range nodeIDs {\n\t\t\t\t\tif err := MinimizeComment(nodeID); err != nil {\n\t\t\t\t\t\tfmt.Fprintf(os.Stderr, \"Warning: could not minimize review: %v\\n\", err)\n\t\t\t\t\t\tminimizeFailed = true\n\t\t\t\t\t}\n\t\t\t\t}\n\t\t\t\tif len(nodeIDs) \u003e 0 \u0026\u0026 !minimizeFailed {\n\t\t\t\t\tfmt.Fprintf(os.Stderr, \"Minimized %d previous review(s)\\n\", len(nodeIDs))\n\t\t\t\t}\n\t\t\t} else {\n\t\t\t\tallThreads, err := FetchReviewThreads(g.Repo, g.PRNumber)\n\t\t\t\tif err != nil {\n\t\t\t\t\tfmt.Fprintf(os.Stderr, \"Warning: could not fetch review threads for minimization: %v\\n\", err)\n\t\t\t\t} else {\n\t\t\t\t\tresolvedIDs := resolvedFindingIDs(allThreads, threads, fixed)\n\t\t\t\t\tminimizeFullyResolvedReviews(g.Repo, g.PRNumber, resolvedIDs)\n\t\t\t\t}\n\t\t\t}\n\t\t} else {\n\t\t\tfmt.Fprintf(os.Stderr, \"Warning: could not fetch reviews for minimization: %v\\n\", err)\n\t\t\tminimizeFailed = true\n\t\t}\n\t}\n\n\t// POST path — pick the body shape that fits the cycle outcome. Every\n\t// branch emits a top-level review so each push lands a visible status\n\t// comment on the PR.\n\tswitch {\n\tcase len(result.Findings) \u003e 0:\n\t\tif err := PostReview(g.Repo, g.PRNumber, result, pr.ValidationDiff(), result.SHA, summary); err != nil {\n\t\t\treturn fmt.Errorf(\"posting review: %w\", err)\n\t\t}\n\t\tStderrf(ansiGreen, \"Review posted to PR #%d\\n\", g.PRNumber)\n\tcase len(threads) \u003e 0 \u0026\u0026 allResolved(threads, fixed):\n\t\tif err := PostAllClearReview(g.Repo, g.PRNumber, result.SHA, minimizeFailed, summary); err != nil {\n\t\t\treturn fmt.Errorf(\"posting all-clear review: %w\", err)\n\t\t}\n\t\tStderrf(ansiGreen, \"All clear! No issues remaining.\\n\")\n\tcase len(threads) \u003e 0:\n\t\tif err := PostActivityReview(g.Repo, g.PRNumber, result.SHA, summary); err != nil {\n\t\t\treturn fmt.Errorf(\"posting activity review: %w\", err)\n\t\t}\n\t\tStderrf(ansiGreen, \"Posted activity summary to PR #%d\\n\", g.PRNumber)\n\tdefault:\n\t\tif err := PostCleanReview(g.Repo, g.PRNumber, result.SHA, summary); err != nil {\n\t\t\treturn fmt.Errorf(\"posting review: %w\", err)\n\t\t}\n\t\tStderrf(ansiGreen, \"Review posted to PR #%d\\n\", g.PRNumber)\n\t}\n\n\tg.postReviewCommitStatus(result.SHA, summary)\n\treturn nil\n}\n\nfunc (g *GithubPlatform) SaveState(_ *ReviewResult, _ []Finding, _ bool) error {\n\t// No-op: GitHub mode stores state in PR review threads (embedded JSON\n\t// markers carry the SHA and findings). Local state files are owned by\n\t// LocalPlatform — this keeps the two adapters from fighting over the\n\t// same ~/.codecanary/state/\u003cbranch\u003e.json file.\n\treturn nil\n}\n\nfunc (g *GithubPlatform) GetIncrementalDiff(baseSHA string, _ []string) (string, error) {\n\treturn GetIncrementalDiff(baseSHA)\n}\n\nfunc (g *GithubPlatform) ReportUsage(tracker *UsageTracker) {\n\treport := tracker.Report(g.Repo, g.PRNumber)\n\tif len(report.Calls) \u003e 0 {\n\t\tif err := WriteUsageEnv(report); err != nil {\n\t\t\tfmt.Fprintf(os.Stderr, \"Warning: could not write usage env: %v\\n\", err)\n\t\t}\n\t}\n}\n\n// postReviewCommitStatus POSTs a `CodeCanary / review` commit status on the\n// reviewed SHA. state=success when no unresolved findings remain for the\n// PR (new findings this cycle + threads still open with no classification\n// both at zero); state=failure otherwise. Teams can require this check in\n// branch protection to gate merges on a clean review.\n//\n// Skipped silently when the SHA is empty (non-pr-loop contexts that\n// accidentally share the adapter). Failures are logged as warnings — the\n// review itself has already been published, so a flaky status post should\n// not abort the whole Publish flow.\nfunc (g *GithubPlatform) postReviewCommitStatus(sha string, summary ReviewSummary) {\n\tif sha == \"\" {\n\t\treturn\n\t}\n\tstate, desc := commitStatusFromSummary(summary)\n\tif err := PostReviewCommitStatus(g.Repo, sha, state, desc); err != nil {\n\t\tfmt.Fprintf(os.Stderr, \"Warning: could not post %s commit status: %v\\n\",\n\t\t\tReviewCommitStatusContext, err)\n\t\treturn\n\t}\n\tStderrf(ansiGreen, \"Posted %s = %s on %s (%s)\\n\",\n\t\tReviewCommitStatusContext, state, shortSHA(sha), desc)\n}\n\n// commitStatusFromSummary maps a ReviewSummary to the (state, description)\n// pair sent to the commit status API. Pulled out so the mapping is\n// unit-testable without network access.\n//\n// An \"unresolved\" count combines new findings this cycle with threads that\n// were already open and remain unclassified — either kind should fail the\n// required check. Everything classified by triage (resolved by code, file\n// removed, dismissed, acknowledged, rebutted) counts as handled.\nfunc commitStatusFromSummary(summary ReviewSummary) (state, desc string) {\n\tunresolved := summary.NewFindings + summary.StillOpen\n\tif unresolved \u003e 0 {\n\t\tsuffix := \"s\"\n\t\tif unresolved == 1 {\n\t\t\tsuffix = \"\"\n\t\t}\n\t\treturn \"failure\", fmt.Sprintf(\"%d unresolved finding%s\", unresolved, suffix)\n\t}\n\tif summary.ResolvedByCode+summary.FileRemoved+summary.Dismissed+summary.Acknowledged+summary.Rebutted \u003e 0 {\n\t\treturn \"success\", \"all findings resolved\"\n\t}\n\treturn \"success\", \"no findings\"\n}\n" + } + }, + "config": {}, + "project_docs": { + "CLAUDE.md": "# CodeCanary\n\nAI-powered code review for GitHub pull requests.\n\n## Project structure\n\n```\ncmd/\n review/ # Main binary — review CLI + setup wizard\n main.go # Entry point\n cli/ # Cobra commands\n root.go # Root \"codecanary\" command\n review.go # codecanary review \u003cpr\u003e\n findings.go # codecanary findings \u003cpr\u003e — fetch bot findings for the review skill\n reply.go # codecanary reply --url \u003ccomment-url\u003e --body \u003ctext\u003e — post a reply on a review thread\n install_skill.go # codecanary install-skill — write embedded Claude skill to disk\n setup.go # codecanary setup [local|github]\n auth.go # codecanary auth [status|delete]\ninternal/\n review/\n runner.go # Core review pipeline — single Run() entry point\n config.go # Config loading, validation, defaults\n # Provider layer (LLM abstraction)\n provider.go # ModelProvider interface + factory registry\n provider_anthropic.go\n provider_openai.go\n provider_openrouter.go\n provider_claude.go # Claude CLI wrapper\n provider_compat.go # Shared types for OpenAI-compatible APIs\n pricing.go # Token-based cost estimation\n # Platform layer (environment abstraction)\n platform.go # ReviewPlatform interface\n platform_github.go # GitHub Actions implementation\n platform_local.go # Local CLI implementation\n # Supporting modules\n prompt.go # Prompt building (review, incremental, per-thread)\n findings.go # Finding parsing, filtering, result structures\n triage.go # Thread classification + parallel LLM evaluation\n formatter.go # JSON/Markdown/Terminal output formatting\n usage.go # Token tracking, budget checking\n github.go # GitHub API calls (fetch threads, post reviews)\n comments.go # PR review comment fetch + finding marker parser + review-check watcher\n local.go # Local diff \u0026 git operations\n state.go # Local state persistence\n docs.go # Project doc discovery\n credentials/ # Credential storage (keychain with file fallback)\n keyring.go # Store/Retrieve/Delete — keychain first, ~/.codecanary/credentials.json fallback\n skills/ # Claude Code skills embedded in the binary via //go:embed\n skills.go # Exports CodecanaryFix() returning the skill body\n codecanary-fix/SKILL.md # Canonical skill source (duplicated at .claude/skills/codecanary-fix/SKILL.md; parity enforced by skills_test.go)\n setup/ # Setup wizard logic (huh forms)\n forms.go # Shared huh form components\n validate.go # API key validation via test calls\n guidance.go # Token/permissions guidance text\n workflow.go # GitHub Actions workflow template\n local.go # RunLocal() — local setup flow\n github.go # RunGitHub() — GitHub Actions setup flow\n auth/ # OAuth PKCE flow, GitHub App installation\ntelemetry/ # Telemetry domain (anonymous usage analytics)\n worker/ # Cloudflare Worker — telemetry ingestion (TypeScript)\n dashboard/ # Cloudflare Pages — internal analytics dashboard (vanilla JS + Chart.js)\noidc/ # OIDC domain\n worker/ # Cloudflare Worker — OIDC token exchange proxy (TypeScript)\naction.yml # GitHub Action definition (composite action)\ninstall.sh # Downloads and installs codecanary binary permanently\n.claude/\n skills/\n codecanary-fix/ # Claude Code skill — drives review→fix→push loop using `codecanary findings` + `codecanary reply`\n```\n\n## Binary\n\n- **`codecanary`** — single binary for reviews, setup, and credential management. Installed locally via `install.sh`, also used by the GitHub Action.\n\n## Build\n\n```sh\ngo build ./cmd/review # builds codecanary\n```\n\nVersion is set via ldflags: `-X main.version=v{version}`\n\n## Lint\n\n```sh\ngolangci-lint run ./...\n```\n\nAll code must pass `golangci-lint` with default linters (errcheck, staticcheck, etc.). Run this before committing.\n\n## Key dependencies\n\n- `spf13/cobra` — CLI framework\n- `charmbracelet/huh` — terminal form builder (setup wizard)\n- `zalando/go-keyring` — OS keychain (with file-based fallback for systems without one)\n- `bmatcuk/doublestar` — glob pattern matching for ignore rules\n- `gopkg.in/yaml.v3` — config parsing\n- `golang.org/x/term` — terminal detection\n\n## Architecture\n\n### Core principle: adapters keep the engine agnostic\n\nThe review engine (`runner.go`) is provider- and platform-agnostic. It depends only on two interfaces — never on concrete GitHub APIs, LLM SDKs, or environment-specific logic. All environment and provider specifics live behind adapters.\n\n### Provider layer — `ModelProvider` interface (`provider.go`)\n\nAbstracts LLM invocations. The core engine calls `provider.Run(ctx, prompt, opts)` and gets back text + usage metadata. It never knows which LLM backend is being used.\n\n**Implementations**: `anthropic`, `openai`, `openrouter`, `claude` (CLI).\n**Selection**: factory registry in `provider.go` — `NewProviderForRole(mc, env)` returns the right implementation based on `mc.Provider`.\n\nAdding a new LLM provider means: create `provider_\u003cname\u003e.go` and register a `ProviderFactory` (constructor, validation, pricing, default models) via `init()`.\n\n### Platform layer — `ReviewPlatform` interface (`platform.go`)\n\nAbstracts environment-specific operations: loading previous findings, publishing results, saving state, resolving threads, reporting usage.\n\n**Implementations**: `GithubPlatform` (posts to PRs, reads threads via API), `LocalPlatform` (prints to terminal, persists state to `~/.codecanary/repos/\u003cowner\u003e/\u003crepo\u003e/state/\u003cbranch\u003e.json` — per-repo scoping keeps branch names like `main` from colliding across repos; falls back to `~/.codecanary/state/\u003cbranch\u003e.json` when no git remote is resolvable).\n\nRouting is strict: `codecanary review --post` → `GithubPlatform`; `codecanary review` (no `--post`) → `LocalPlatform`, even when the branch has an open PR. Local is local — the branch diff (with uncommitted changes) is reviewed against the default base, previous findings come from local state, nothing is fetched from or posted to GitHub. The old `GithubPlatform`-with-`Post=false` hybrid is gone; it was the source of the \"state written locally, read from GitHub\" asymmetry that kept breaking incremental local reviews.\n\nAdding a new platform (e.g., GitLab) means: implement `ReviewPlatform`, wire it in the CLI.\n\n### Unified review pipeline (`runner.go`)\n\nThere is a **single `Run()` function** — not separate paths for GitHub vs. local. The pipeline is:\n\n1. Fetch PR data (or local diff)\n2. Load config, project docs, file contents\n3. Create providers via `NewProviderForRole()` (factory, provider-agnostic)\n4. Load previous findings via `platform.LoadPreviousFindings()`\n5. If incremental: triage threads, evaluate via provider, handle resolutions\n6. Build and execute main review prompt\n7. Parse findings, filter non-actionable\n8. `platform.Publish()` → `platform.SaveState()` → `platform.ReportUsage()`\n\n### Other architecture notes\n\n- **Config** is split across two files in `.codecanary/`: `config.yml` (provider, models, budgets, timeouts) and `review.yml` (rules, context, ignore patterns). `review.yml` is optional — if present, its fields override rules/context/ignore in `config.yml`. A personal `review.local.yml` can add rules, context, and ignore patterns on top of `review.yml` (append semantics, not replacement). Legacy `.codecanary.yml` at repo root is still supported with a deprecation warning.\n- **Incremental reviews**: on re-push, triage existing threads (Go-driven classifier in `triage.go`), evaluate changed threads via provider (triage model), then review only new code\n- **Dual marker detection**: reads both `codecanary:review` and legacy `clanopy:review` HTML markers for backward compatibility\n- **Anti-hallucination**: explicit file allowlist, line validation against diff, max finding distance threshold\n- **OIDC worker** (`oidc/worker/`): OIDC token exchange proxy at `oidc.codecanary.sh` — verifies GitHub Actions OIDC token, returns GitHub App installation token\n- **Telemetry worker** (`telemetry/worker/`): anonymous usage ingest at `telemetry.codecanary.sh` — writes to a Cloudflare Analytics Engine dataset\n- **Telemetry dashboard** (`telemetry/dashboard/`): internal analytics view at `dashboard.codecanary.sh` — Cloudflare Pages + Pages Functions, gated by Cloudflare Access. Reads the AE dataset via the SQL HTTP API with a read-only API token (no AE binding, so writes are platform-impossible)\n- **Setup** is a subcommand (`codecanary setup`) using `charmbracelet/huh` forms, with `local` and `github` sub-flows\n- **Credentials** use a single env var `CODECANARY_PROVIDER_SECRET` for all providers. Stored via `go-keyring` (OS keychain) with a file-based fallback (`~/.codecanary/credentials.json`, mode `0600`). `resolveEnv()` in `runner.go` injects the stored credential into the filtered env when not already set.\n\n## Rules\n\n- **Keep the core engine agnostic.** `runner.go`, `triage.go`, `prompt.go`, `findings.go` must never import or reference a specific LLM provider or platform. All provider/platform specifics go behind the `ModelProvider` or `ReviewPlatform` interfaces. No `if provider == \"openai\"` in core logic.\n- **Use the adapter/provider pattern for new integrations.** New LLM backends → create `provider_\u003cname\u003e.go` with a `ProviderFactory` registration in `init()`. New deployment targets → implement `ReviewPlatform` + wire in CLI. Never fork the pipeline.\n- **One pipeline, not two.** There must be a single `Run()` path. GitHub and local modes differ only in which `ReviewPlatform` implementation is injected — the orchestration logic is shared.\n- **Shared types for similar providers.** OpenAI-compatible APIs share request/response types via `provider_compat.go`. Don't duplicate HTTP client logic across providers.\n- **Don't repeat yourself.** Before writing new code, search the codebase for existing functions, mappings, or logic that already does what you need — then call it instead of reimplementing it. This applies to everything: switch statements, helper functions, validation logic, data mappings, HTTP calls. One source of truth, callers import it. Don't merge scaffolding or unused exports — if it's not called yet, it doesn't ship yet.\n- **File names are ownership boundaries.** A function defined in `local.go` implies it belongs to the local flow; one in `github.go` implies it belongs to GitHub. If a function is called by multiple files in the same package, it belongs in a shared file (e.g., `forms.go` for setup helpers, `platform.go` for platform-shared logic). Never define shared infrastructure in a flow-specific file — move it to the file that matches its actual scope.\n- **Canonical provider registration points.** Provider names live in the factory map in `provider.go`. All providers use `CODECANARY_PROVIDER_SECRET` for credentials (defined in `internal/credentials/keyring.go`). When adding a new provider, register a `ProviderFactory` in `provider.go` and add config validation in `config.go`.\n- **Minimize shell code.** `install.sh` and the GitHub Action (`action.yml`) should be kept as thin as possible. All logic must live in Go.\n- **Workflow template is embedded.** `internal/setup/codecanary.yml` is the single source of truth for the GitHub Actions workflow, embedded via `//go:embed`. `.github/workflows/codecanary.yml` must be identical — `go test ./internal/setup/` enforces this. When changing the workflow, edit either file and copy to the other.\n- **Claude skills are embedded.** `internal/skills/codecanary-fix/SKILL.md` is the single source of truth for the codecanary-fix skill, embedded via `//go:embed` and materialized by `codecanary install-skill`. `.claude/skills/codecanary-fix/SKILL.md` must be identical so Claude Code's project-mode discovery finds it when working in this repo — `go test ./internal/skills/` enforces this. When changing the skill, edit either file and copy to the other.\n- **Keep `docs/review-flow.md` in sync.** This document describes the full review pipeline — every step, the triage flow, platform differences, and key design decisions. When changing `runner.go`, `triage.go`, `prompt.go`, `findings.go`, `github.go`, `local.go`, `platform.go`, or the `ReviewPlatform` implementations, update the doc to reflect the new behavior.\n- Tests exist for config, findings, formatting, and triage. Be careful with refactors — run `go test ./...` and `go vet ./...`.\n" + } +} diff --git a/internal/review/testdata/corpus/alansikora-codecanary-pr173.prompt.golden b/internal/review/testdata/corpus/alansikora-codecanary-pr173.prompt.golden new file mode 100644 index 0000000..9ceaa0f --- /dev/null +++ b/internal/review/testdata/corpus/alansikora-codecanary-pr173.prompt.golden @@ -0,0 +1,2727 @@ +You are a code reviewer. Review the following pull request and report findings. +You will be given the full contents of changed files for context, along with the diff. Only report issues that are directly related to the changes in the diff — do not flag pre-existing issues in unchanged code. Do not report a finding if your analysis concludes that the code is correct and no action is needed — only report findings that require the author to make a change or consider a specific alternative. +Also consider whether the changes could cause side effects in other files that depend on or interact with the modified code (e.g. callers, importers, shared state). If you identify a potential side effect, anchor your finding to the relevant line in the diff and describe the affected downstream code in the description. + +## Pull Request #173 +feat: prettier commit status context (CodeCanary / review) +**Author:** alansikora + +## Summary + +Renames the commit status context posted by `GithubPlatform.Publish` (and `codecanary signoff`) from `codecanary/review` to `CodeCanary / review`. Capital C, spaces around the slash — matches how GitHub renders workflow check runs, so the status reads as cleanly as the existing CI entries in the PR checks list. + +No behavior change. Same success/failure logic, same call sites, same pure-helper test coverage. + +## Heads-up: clash with the Action's check run + +The workflow job `review` in workflow `CodeCanary` also renders as `CodeCanary / review` in the PR checks UI. They're technically separate entities (Check Run vs Commit Status), and the branch-protection status-check picker distinguishes them by source, but they now share an identical display label. Discussed in chat; accepted for now on the theory that the two entries both represent "CodeCanary did its thing for this commit" and a single unified label is fine. Easy to split later by renaming the workflow job id or the status context if it becomes confusing. + +## Migration + +Existing installs that wired `codecanary/review` into branch-protection rules will need to update the required-check name to `CodeCanary / review` after this merges. Documented in the README and `docs/review-flow.md` alongside the rename. + +## Test plan + +- [x] `go build ./...` +- [x] `go vet ./...` +- [x] `go test ./...` +- [ ] End-to-end: merge + push, confirm `CodeCanary / review = success` / `failure` lands on HEAD +- [ ] Confirm branch-protection picker lists the new context + + +## Project Documentation +The following project documentation describes conventions and standards for this codebase. Use these to inform your review — flag violations of these conventions when relevant. + + +# CodeCanary + +AI-powered code review for GitHub pull requests. + +## Project structure + +``` +cmd/ + review/ # Main binary — review CLI + setup wizard + main.go # Entry point + cli/ # Cobra commands + root.go # Root "codecanary" command + review.go # codecanary review + findings.go # codecanary findings — fetch bot findings for the review skill + reply.go # codecanary reply --url --body — post a reply on a review thread + install_skill.go # codecanary install-skill — write embedded Claude skill to disk + setup.go # codecanary setup [local|github] + auth.go # codecanary auth [status|delete] +internal/ + review/ + runner.go # Core review pipeline — single Run() entry point + config.go # Config loading, validation, defaults + # Provider layer (LLM abstraction) + provider.go # ModelProvider interface + factory registry + provider_anthropic.go + provider_openai.go + provider_openrouter.go + provider_claude.go # Claude CLI wrapper + provider_compat.go # Shared types for OpenAI-compatible APIs + pricing.go # Token-based cost estimation + # Platform layer (environment abstraction) + platform.go # ReviewPlatform interface + platform_github.go # GitHub Actions implementation + platform_local.go # Local CLI implementation + # Supporting modules + prompt.go # Prompt building (review, incremental, per-thread) + findings.go # Finding parsing, filtering, result structures + triage.go # Thread classification + parallel LLM evaluation + formatter.go # JSON/Markdown/Terminal output formatting + usage.go # Token tracking, budget checking + github.go # GitHub API calls (fetch threads, post reviews) + comments.go # PR review comment fetch + finding marker parser + review-check watcher + local.go # Local diff & git operations + state.go # Local state persistence + docs.go # Project doc discovery + credentials/ # Credential storage (keychain with file fallback) + keyring.go # Store/Retrieve/Delete — keychain first, ~/.codecanary/credentials.json fallback + skills/ # Claude Code skills embedded in the binary via //go:embed + skills.go # Exports CodecanaryFix() returning the skill body + codecanary-fix/SKILL.md # Canonical skill source (duplicated at .claude/skills/codecanary-fix/SKILL.md; parity enforced by skills_test.go) + setup/ # Setup wizard logic (huh forms) + forms.go # Shared huh form components + validate.go # API key validation via test calls + guidance.go # Token/permissions guidance text + workflow.go # GitHub Actions workflow template + local.go # RunLocal() — local setup flow + github.go # RunGitHub() — GitHub Actions setup flow + auth/ # OAuth PKCE flow, GitHub App installation +telemetry/ # Telemetry domain (anonymous usage analytics) + worker/ # Cloudflare Worker — telemetry ingestion (TypeScript) + dashboard/ # Cloudflare Pages — internal analytics dashboard (vanilla JS + Chart.js) +oidc/ # OIDC domain + worker/ # Cloudflare Worker — OIDC token exchange proxy (TypeScript) +action.yml # GitHub Action definition (composite action) +install.sh # Downloads and installs codecanary binary permanently +.claude/ + skills/ + codecanary-fix/ # Claude Code skill — drives review→fix→push loop using `codecanary findings` + `codecanary reply` +``` + +## Binary + +- **`codecanary`** — single binary for reviews, setup, and credential management. Installed locally via `install.sh`, also used by the GitHub Action. + +## Build + +```sh +go build ./cmd/review # builds codecanary +``` + +Version is set via ldflags: `-X main.version=v{version}` + +## Lint + +```sh +golangci-lint run ./... +``` + +All code must pass `golangci-lint` with default linters (errcheck, staticcheck, etc.). Run this before committing. + +## Key dependencies + +- `spf13/cobra` — CLI framework +- `charmbracelet/huh` — terminal form builder (setup wizard) +- `zalando/go-keyring` — OS keychain (with file-based fallback for systems without one) +- `bmatcuk/doublestar` — glob pattern matching for ignore rules +- `gopkg.in/yaml.v3` — config parsing +- `golang.org/x/term` — terminal detection + +## Architecture + +### Core principle: adapters keep the engine agnostic + +The review engine (`runner.go`) is provider- and platform-agnostic. It depends only on two interfaces — never on concrete GitHub APIs, LLM SDKs, or environment-specific logic. All environment and provider specifics live behind adapters. + +### Provider layer — `ModelProvider` interface (`provider.go`) + +Abstracts LLM invocations. The core engine calls `provider.Run(ctx, prompt, opts)` and gets back text + usage metadata. It never knows which LLM backend is being used. + +**Implementations**: `anthropic`, `openai`, `openrouter`, `claude` (CLI). +**Selection**: factory registry in `provider.go` — `NewProviderForRole(mc, env)` returns the right implementation based on `mc.Provider`. + +Adding a new LLM provider means: create `provider_.go` and register a `ProviderFactory` (constructor, validation, pricing, default models) via `init()`. + +### Platform layer — `ReviewPlatform` interface (`platform.go`) + +Abstracts environment-specific operations: loading previous findings, publishing results, saving state, resolving threads, reporting usage. + +**Implementations**: `GithubPlatform` (posts to PRs, reads threads via API), `LocalPlatform` (prints to terminal, persists state to `~/.codecanary/repos///state/.json` — per-repo scoping keeps branch names like `main` from colliding across repos; falls back to `~/.codecanary/state/.json` when no git remote is resolvable). + +Routing is strict: `codecanary review --post` → `GithubPlatform`; `codecanary review` (no `--post`) → `LocalPlatform`, even when the branch has an open PR. Local is local — the branch diff (with uncommitted changes) is reviewed against the default base, previous findings come from local state, nothing is fetched from or posted to GitHub. The old `GithubPlatform`-with-`Post=false` hybrid is gone; it was the source of the "state written locally, read from GitHub" asymmetry that kept breaking incremental local reviews. + +Adding a new platform (e.g., GitLab) means: implement `ReviewPlatform`, wire it in the CLI. + +### Unified review pipeline (`runner.go`) + +There is a **single `Run()` function** — not separate paths for GitHub vs. local. The pipeline is: + +1. Fetch PR data (or local diff) +2. Load config, project docs, file contents +3. Create providers via `NewProviderForRole()` (factory, provider-agnostic) +4. Load previous findings via `platform.LoadPreviousFindings()` +5. If incremental: triage threads, evaluate via provider, handle resolutions +6. Build and execute main review prompt +7. Parse findings, filter non-actionable +8. `platform.Publish()` → `platform.SaveState()` → `platform.ReportUsage()` + +### Other architecture notes + +- **Config** is split across two files in `.codecanary/`: `config.yml` (provider, models, budgets, timeouts) and `review.yml` (rules, context, ignore patterns). `review.yml` is optional — if present, its fields override rules/context/ignore in `config.yml`. A personal `review.local.yml` can add rules, context, and ignore patterns on top of `review.yml` (append semantics, not replacement). Legacy `.codecanary.yml` at repo root is still supported with a deprecation warning. +- **Incremental reviews**: on re-push, triage existing threads (Go-driven classifier in `triage.go`), evaluate changed threads via provider (triage model), then review only new code +- **Dual marker detection**: reads both `codecanary:review` and legacy `clanopy:review` HTML markers for backward compatibility +- **Anti-hallucination**: explicit file allowlist, line validation against diff, max finding distance threshold +- **OIDC worker** (`oidc/worker/`): OIDC token exchange proxy at `oidc.codecanary.sh` — verifies GitHub Actions OIDC token, returns GitHub App installation token +- **Telemetry worker** (`telemetry/worker/`): anonymous usage ingest at `telemetry.codecanary.sh` — writes to a Cloudflare Analytics Engine dataset +- **Telemetry dashboard** (`telemetry/dashboard/`): internal analytics view at `dashboard.codecanary.sh` — Cloudflare Pages + Pages Functions, gated by Cloudflare Access. Reads the AE dataset via the SQL HTTP API with a read-only API token (no AE binding, so writes are platform-impossible) +- **Setup** is a subcommand (`codecanary setup`) using `charmbracelet/huh` forms, with `local` and `github` sub-flows +- **Credentials** use a single env var `CODECANARY_PROVIDER_SECRET` for all providers. Stored via `go-keyring` (OS keychain) with a file-based fallback (`~/.codecanary/credentials.json`, mode `0600`). `resolveEnv()` in `runner.go` injects the stored credential into the filtered env when not already set. + +## Rules + +- **Keep the core engine agnostic.** `runner.go`, `triage.go`, `prompt.go`, `findings.go` must never import or reference a specific LLM provider or platform. All provider/platform specifics go behind the `ModelProvider` or `ReviewPlatform` interfaces. No `if provider == "openai"` in core logic. +- **Use the adapter/provider pattern for new integrations.** New LLM backends → create `provider_.go` with a `ProviderFactory` registration in `init()`. New deployment targets → implement `ReviewPlatform` + wire in CLI. Never fork the pipeline. +- **One pipeline, not two.** There must be a single `Run()` path. GitHub and local modes differ only in which `ReviewPlatform` implementation is injected — the orchestration logic is shared. +- **Shared types for similar providers.** OpenAI-compatible APIs share request/response types via `provider_compat.go`. Don't duplicate HTTP client logic across providers. +- **Don't repeat yourself.** Before writing new code, search the codebase for existing functions, mappings, or logic that already does what you need — then call it instead of reimplementing it. This applies to everything: switch statements, helper functions, validation logic, data mappings, HTTP calls. One source of truth, callers import it. Don't merge scaffolding or unused exports — if it's not called yet, it doesn't ship yet. +- **File names are ownership boundaries.** A function defined in `local.go` implies it belongs to the local flow; one in `github.go` implies it belongs to GitHub. If a function is called by multiple files in the same package, it belongs in a shared file (e.g., `forms.go` for setup helpers, `platform.go` for platform-shared logic). Never define shared infrastructure in a flow-specific file — move it to the file that matches its actual scope. +- **Canonical provider registration points.** Provider names live in the factory map in `provider.go`. All providers use `CODECANARY_PROVIDER_SECRET` for credentials (defined in `internal/credentials/keyring.go`). When adding a new provider, register a `ProviderFactory` in `provider.go` and add config validation in `config.go`. +- **Minimize shell code.** `install.sh` and the GitHub Action (`action.yml`) should be kept as thin as possible. All logic must live in Go. +- **Workflow template is embedded.** `internal/setup/codecanary.yml` is the single source of truth for the GitHub Actions workflow, embedded via `//go:embed`. `.github/workflows/codecanary.yml` must be identical — `go test ./internal/setup/` enforces this. When changing the workflow, edit either file and copy to the other. +- **Claude skills are embedded.** `internal/skills/codecanary-fix/SKILL.md` is the single source of truth for the codecanary-fix skill, embedded via `//go:embed` and materialized by `codecanary install-skill`. `.claude/skills/codecanary-fix/SKILL.md` must be identical so Claude Code's project-mode discovery finds it when working in this repo — `go test ./internal/skills/` enforces this. When changing the skill, edit either file and copy to the other. +- **Keep `docs/review-flow.md` in sync.** This document describes the full review pipeline — every step, the triage flow, platform differences, and key design decisions. When changing `runner.go`, `triage.go`, `prompt.go`, `findings.go`, `github.go`, `local.go`, `platform.go`, or the `ReviewPlatform` implementations, update the doc to reflect the new behavior. +- Tests exist for config, findings, formatting, and triage. Be careful with refactors — run `go test ./...` and `go vet ./...`. + + + +## Review Rules +No specific rules are defined. Perform a general code review covering correctness, security, performance, and maintainability. + +## Files in This Diff +The following files — and ONLY these files — are part of this diff. Every finding you report MUST reference one of these exact paths. Do NOT reference any file that is not in this list. + +- `README.md` +- `cmd/review/cli/signoff.go` +- `docs/review-flow.md` +- `internal/review/github.go` +- `internal/review/platform_github.go` + +## Changed File Contents +Below are the full contents of changed files. Use these to understand surrounding code, types, imports, and control flow. Do NOT report findings on unchanged code — only flag issues directly related to changes in the diff. + +### `README.md` +```` +1: # codecanary CodeCanary +2: +3: AI-powered code review for GitHub pull requests. Catch bugs, security issues, and quality problems before they land in main. +4: +5: ## Quick Start +6: +7: ```sh +8: curl -fsSL https://codecanary.sh/install | sh +9: codecanary setup local +10: codecanary review +11: ``` +12: +13: That's it. CodeCanary diffs your branch against main and reviews the changes locally. +14: +15: ## Why CodeCanary? +16: +17: - **Fully automated** — runs as a GitHub Action on every push, or locally from the terminal. +18: - **Multi-provider** — bring your own LLM: Anthropic, OpenAI, OpenRouter, Grok (xAI), or Claude CLI. No vendor lock-in. +19: - **Incremental reviews** — on re-push, Go-driven triage classifies existing threads at zero LLM cost. Only changed code gets re-evaluated. +20: - **Conversational** — when authors reply to a finding, CodeCanary re-evaluates in context. It distinguishes code fixes, dismissals, acknowledgments, and rebuttals. +21: - **Native PR integration** — posts inline comments on exact diff lines, auto-resolves threads when code is fixed, and minimizes stale reviews. +22: - **Anti-hallucination** — explicit file allowlists, line validation against the diff, and distance thresholds prevent fabricated findings. +23: - **Cost-efficient** — uses a fast triage model for thread re-evaluation and a full model for review. Tracks per-invocation usage so you see what you spend. +24: - **Configuration-as-code** — project-specific rules, severity levels, ignore patterns, and context in `.codecanary/config.yml`. +25: - **Agentic loop** — pairs with Claude Code via the bundled `codecanary-fix` skill to review, triage, fix, and push until the PR is clean. +26: +27: ## Installation +28: +29: ```sh +30: curl -fsSL https://codecanary.sh/install | sh +31: ``` +32: +33: Installs the `codecanary` binary to `/usr/local/bin` (or `~/.local/bin`). Supports Linux and macOS (amd64/arm64). +34: +35: To self-update later: +36: +37: ```sh +38: codecanary upgrade +39: ``` +40: +41: ### Canary builds +42: +43: ```sh +44: curl -fsSL https://codecanary.sh/install | sh -s -- --canary +45: codecanary upgrade --canary +46: ``` +47: +48: ## Setup +49: +50: ### Local reviews +51: +52: ```sh +53: codecanary setup local +54: ``` +55: +56: The setup wizard walks you through choosing a provider, entering your API key (stored in your system keychain), and selecting models. It creates a `.codecanary/config.yml` in your repo. +57: +58: Then review your changes: +59: +60: ```sh +61: codecanary review # diff against main, print to terminal +62: codecanary review --post # same, but also post findings to the PR on GitHub +63: codecanary review --output json # machine-readable output +64: ``` +65: +66: Without `--post`, `codecanary review` is always local: it diffs your branch (with uncommitted changes) against the default branch and keeps state in `~/.codecanary/state/.json` for incremental re-runs. With `--post`, it fetches the PR from GitHub — pass a number or let it auto-detect from the current branch — and posts findings as review comments. +67: +68: ### GitHub Actions +69: +70: ```sh +71: codecanary setup github +72: ``` +73: +74: This runs the same provider and key selection, then: +75: 1. Installs the CodeCanary Review GitHub App +76: 2. Sets your API key as a GitHub repo secret +77: 3. Creates the workflow (`.github/workflows/codecanary.yml`) +78: 4. Opens a PR with everything ready to merge +79: +80: Once merged, CodeCanary reviews every PR on open and push. Draft PRs are skipped by default. +81: +82: ### Gating merges on clean reviews +83: +84: CodeCanary can block merges until a review comes back clean. After every review, the bot (and the local `codecanary signoff` command) posts a GitHub commit status under the context `CodeCanary / review`: +85: +86: - `success` — no unresolved findings (everything is either unraised, fixed by code, or handled by the author) +87: - `failure` — one or more findings remain unresolved, with a description like `"3 unresolved findings"` +88: +89: To turn this into a required check, add `CodeCanary / review` to your repo's required status checks via whichever branch protection mechanism you use (rulesets, classic branch protection rules, etc.). GitHub accepts any context name; if a review has already run, it will also show up in autocomplete. +90: +91: Statuses are keyed by commit SHA, so staling is automatic: a new push has no status until the next review run posts one, which re-blocks the merge button. +92: +93: **Without a workflow.** If your repo doesn't use the GitHub Action, `codecanary signoff` posts the same status from your local machine after a clean `codecanary review`: +94: +95: ```sh +96: codecanary review # reviews locally, stores findings in ~/.codecanary/... +97: # fix anything that came up, commit +98: codecanary signoff # posts CodeCanary / review = success on HEAD +99: ``` +100: +101: The command refuses to sign off unless HEAD matches the reviewed SHA and the tree is clean — this prevents attesting a review of code that isn't actually in the commit. It needs `gh` authenticated with `repo:status` scope (default `gh auth login` covers it). +102: +103: Same required-check config works for both paths: the bot satisfies the check on PRs that run the workflow; `codecanary signoff` satisfies it for repos without the workflow, branches where the workflow doesn't run, or PRs from forks where secrets are unavailable. +104: +105: ## CLI Reference +106: +107: | Command | Description | +108: |---------|-------------| +109: | `codecanary review [pr-number]` | Review a PR or local diff | +110: | `codecanary findings [pr-number]` | Fetch bot findings for a PR (markdown or JSON) | +111: | `codecanary reply --url --body ` | Post a reply on a review-comment thread (used by the skill when skipping) | +112: | `codecanary signoff` | Post a `CodeCanary / review` commit status from the last local review (see [gating merges](#gating-merges-on-clean-reviews)) | +113: | `codecanary install-skill` | Install the `codecanary-fix` Claude Code skill | +114: | `codecanary setup [local\|github]` | Interactive setup wizard | +115: | `codecanary auth status` | Show stored credential info | +116: | `codecanary auth delete` | Remove a stored API key | +117: | `codecanary auth refresh` | Validate and update stored credentials | +118: | `codecanary upgrade` | Update to the latest release | +119: +120: ### Review flags +121: +122: | Flag | Description | +123: |------|-------------| +124: | `--repo, -r` | GitHub repo (owner/name) | +125: | `--output, -o` | Output format: `terminal`, `markdown`, or `json` (auto-detects TTY) | +126: | `--post` | Post findings as a PR review comment | +127: | `--config, -c` | Path to config file (auto-detected if empty) | +128: | `--reply-only` | Re-evaluate thread replies only, skip new findings | +129: | `--dry-run` | Show the prompt without calling the LLM | +130: +131: ## Configuration +132: +133: CodeCanary uses `.codecanary/config.yml` in your repo. The `provider` field is required. +134: +135: ### Minimal config +136: +137: ```yaml +138: version: 1 +139: provider: anthropic +140: review_model: claude-sonnet-4-6 +141: triage_model: claude-haiku-4-5-20251001 +142: ``` +143: +144: ### Config with rules and context +145: +146: ```yaml +147: version: 1 +148: provider: anthropic +149: review_model: claude-sonnet-4-6 +150: triage_model: claude-haiku-4-5-20251001 +151: +152: context: | +153: Go REST API using chi router. Tests use testify. +154: +155: rules: +156: - id: error-handling +157: description: "Errors must be wrapped with context using fmt.Errorf" +158: severity: warning +159: paths: ["**/*.go"] +160: +161: - id: sql-injection +162: description: "Database queries must use parameterized statements" +163: severity: critical +164: +165: ignore: +166: - "dist/**" +167: - "*.lock" +168: - "vendor/**" +169: ``` +170: +171: Rules support `paths` and `exclude_paths` globs, and five severity levels: `critical`, `bug`, `warning`, `suggestion`, `nitpick`. +172: +173: ### Provider examples +174: +175: **OpenAI** (also works with Azure, Ollama, or any OpenAI-compatible endpoint via `api_base`): +176: ```yaml +177: version: 1 +178: provider: openai +179: review_model: gpt-5.4 +180: triage_model: gpt-5.4-mini +181: # api_base: https://your-endpoint.com/v1 +182: ``` +183: +184: **OpenRouter**: +185: ```yaml +186: version: 1 +187: provider: openrouter +188: review_model: anthropic/claude-sonnet-4-6 +189: triage_model: anthropic/claude-haiku-4-5-20251001 +190: ``` +191: +192: **Grok (xAI)**: +193: ```yaml +194: version: 1 +195: provider: grok +196: review_model: grok-4.20-0309-non-reasoning +197: triage_model: grok-4-1-fast-non-reasoning +198: ``` +199: +200: **Claude CLI** (uses your logged-in `claude` session, no API key needed): +201: ```yaml +202: version: 1 +203: provider: claude +204: review_model: claude-sonnet-4-6 +205: triage_model: haiku +206: ``` +207: +208: You can also create a `.codecanary/review.local.yml` for personal overrides (gitignored) — its rules, context, and ignore patterns are appended to the shared `review.yml`. +209: +210: For the full config reference including budget controls, size limits, timeouts, evaluation context, and the `review.yml` override file, see [docs/configuration.md](docs/configuration.md). +211: +212: ## Credential Management +213: +214: ```sh +215: codecanary auth status # show which API keys are stored +216: codecanary auth delete # remove a stored API key +217: codecanary auth refresh # validate and update credentials +218: ``` +219: +220: Keys are stored in your system keychain (macOS Keychain, GNOME Keyring, KDE Wallet) with a fallback to `~/.codecanary/credentials.json`. Environment variables always override stored credentials. +221: +222: | Provider | What you need | Where to get it | +223: |----------|--------------|-----------------| +224: | Anthropic | API key | [console.anthropic.com](https://console.anthropic.com) | +225: | OpenAI | API key | [platform.openai.com](https://platform.openai.com) | +226: | OpenRouter | API key | [openrouter.ai](https://openrouter.ai) | +227: | Grok (xAI) | API key | [console.x.ai](https://console.x.ai) | +228: | Claude CLI | Logged-in `claude` binary | Run `claude` and complete the login flow | +229: +230: ## How It Works +231: +232: ### First review +233: +234: 1. Fetches PR metadata and diff (via `gh` CLI or local git) +235: 2. Reads file contents for context (respecting ignore patterns and size limits) +236: 3. Auto-discovers project docs (CLAUDE.md files) for additional context +237: 4. Calls your configured LLM to analyze the changes +238: 5. Posts findings as inline PR review comments (or prints to terminal) +239: +240: ### Incremental reviews (on re-push) +241: +242: 1. **Go-driven triage** classifies existing threads — no LLM calls for unchanged code +243: 2. **Parallel evaluation** re-checks threads where code changed or the author replied (using the triage model) +244: 3. **New code review** covers only the incremental diff, excluding known issues +245: 4. **Auto-resolution** marks threads as resolved when the code fix addresses the finding +246: +247: ### Thread lifecycle +248: +249: | Event | Result | +250: |-------|--------| +251: | Code fix detected | Thread auto-resolved | +252: | Author dismisses | Acknowledged, kept open for re-check | +253: | Author acknowledges | Noted, kept open | +254: | Author rebuts | Evaluated for technical merit, kept open | +255: | No changes | Skipped (zero LLM cost) | +256: +257: ### Safety +258: +259: - **Anti-hallucination**: explicit file allowlist, line number validation against diff, max finding distance threshold +260: - **Anti-ping-pong**: resolved findings injected as context to prevent re-raising +261: - **Prompt injection protection**: repository content escaped before inclusion in prompts +262: +263: ## Agentic review loop +264: +265: CodeCanary ships with a [Claude Code](https://docs.claude.com/en/docs/claude-code) skill, `codecanary-fix`, that drives a review → triage → fix → push cycle until the PR is clean. You stay in the loop — every fix is confirmed before it's applied — but the polling, fetching, and CI watching is handled by the CLI. +266: +267: Install the skill once: +268: +269: ```sh +270: codecanary install-skill +271: ``` +272: +273: This writes the embedded skill to `~/.claude/skills/codecanary-fix/SKILL.md`, where Claude Code discovers it in every session. Re-run the command after `codecanary upgrade` to pick up new versions. +274: +275: Then in Claude Code, ask it to `handle codecanary` on your PR (or invoke `/codecanary-fix` directly) — the skill is auto-discovered and matched to your request via its frontmatter description. Two modes: +276: +277: - **PR mode** (default) — watches the GitHub Actions review check via `codecanary findings --watch`, renders a triage table, asks you to confirm which fixes to apply, commits and pushes, then loops on the next review. Every finding you defer gets a reply posted on its review thread explaining why, via `codecanary reply`. +278: - **Local mode** — triggered automatically when no PR is detected for the current branch. Single pass against your dirty working tree. Applies approved fixes without committing or pushing. +279: +280: The full skill contract lives at [internal/skills/codecanary-fix/SKILL.md](internal/skills/codecanary-fix/SKILL.md). +281: +282: ## Contributing +283: +284: See [CONTRIBUTING.md](CONTRIBUTING.md) for architecture details and how to add new LLM providers or platforms. +285: +286: ## License +287: +288: MIT +```` + +### `cmd/review/cli/signoff.go` +``` +1: package cli +2: +3: import ( +4: "errors" +5: "fmt" +6: +7: "github.com/alansikora/codecanary/internal/review" +8: "github.com/spf13/cobra" +9: ) +10: +11: // signoffCmd posts a GitHub commit status on the current HEAD based on the +12: // most recent local review for the current branch. Inspired by Rails 8.1's +13: // `gh signoff` (Basecamp): combined with a required status check in branch +14: // protection, this lets a team gate merges on "a human actually ran +15: // codecanary locally and it came back clean", without paying for a cloud +16: // review on every push. +17: var signoffCmd = &cobra.Command{ +18: Use: "signoff", +19: Short: "Post a GitHub commit status from the last local review", +20: Long: `Post a GitHub commit status on HEAD reflecting the most recent local +21: codecanary review for the current branch. +22: +23: Requires: +24: - you ran 'codecanary review' on this branch +25: - the working tree is clean and HEAD matches the reviewed SHA +26: - 'gh' is installed and authenticated with repo:status scope +27: +28: Combine with a required 'CodeCanary / review' check in branch protection to +29: block merges until a clean local review exists for the tip commit.`, +30: RunE: func(cmd *cobra.Command, args []string) error { +31: force, _ := cmd.Flags().GetBool("force") +32: +33: slug, err := review.RepoSlug() +34: if err != nil { +35: return err +36: } +37: branch, err := review.CurrentBranch() +38: if err != nil { +39: return err +40: } +41: sha, err := review.HeadSHA() +42: if err != nil { +43: return err +44: } +45: +46: state, err := review.LoadLocalState(branch) +47: if err != nil { +48: return fmt.Errorf("loading local state: %w", err) +49: } +50: if state == nil { +51: return fmt.Errorf("no local review found for branch %q — run 'codecanary review' first", branch) +52: } +53: +54: if !force { +55: if state.SHA != sha { +56: return fmt.Errorf( +57: "local review was for %s but HEAD is %s — re-run 'codecanary review' (or pass --force)", +58: shortSHA(state.SHA), shortSHA(sha)) +59: } +60: dirty, err := review.WorkingTreeDirty() +61: if err != nil { +62: return fmt.Errorf("checking working tree: %w", err) +63: } +64: if dirty { +65: return errors.New( +66: "working tree has uncommitted changes — commit them, " + +67: "run 'git stash --include-untracked' (plain 'git stash' leaves untracked files behind), " + +68: "or pass --force if you know the dirty files are unrelated to the PR") +69: } +70: } +71: +72: unresolved := 0 +73: for _, f := range state.Findings { +74: if f.Actionable != nil && !*f.Actionable { +75: continue +76: } +77: unresolved++ +78: } +79: +80: statusState, desc := "success", "0 findings" +81: if unresolved > 0 { +82: statusState = "failure" +83: suffix := "s" +84: if unresolved == 1 { +85: suffix = "" +86: } +87: desc = fmt.Sprintf("%d unresolved finding%s", unresolved, suffix) +88: } +89: +90: if err := review.PostReviewCommitStatus(slug, sha, statusState, desc); err != nil { +91: return fmt.Errorf("posting commit status: %w", err) +92: } +93: +94: _, _ = fmt.Fprintf(cmd.OutOrStdout(), +95: "✓ %s = %s on %s@%s (%s)\n", +96: review.ReviewCommitStatusContext, statusState, slug, shortSHA(sha), desc) +97: return nil +98: }, +99: } +100: +101: func init() { +102: signoffCmd.Flags().Bool("force", false, +103: "Sign off even if HEAD doesn't match the reviewed SHA or the tree is dirty") +104: rootCmd.AddCommand(signoffCmd) +105: } +``` + +### `docs/review-flow.md` +```` +1: # Review Flow +2: +3: How CodeCanary reviews a pull request, step by step. +4: +5: ## Overview +6: +7: The review pipeline has two modes of operation: +8: +9: - **First review**: Reviews the full PR diff against the base branch. +10: - **Incremental review**: Re-evaluates previous findings and reviews only new changes since the last review. +11: +12: Both modes run through the same `Run()` function in `runner.go`. The pipeline is platform-agnostic -- GitHub and local modes differ only in which `ReviewPlatform` adapter is injected. +13: +14: ## Platforms +15: +16: Two platforms, routed strictly by `--post`: +17: +18: | Context | Platform | How it runs | State storage | Output | +19: |---------|----------|-------------|---------------|--------| +20: | **GitHub PR** | `GithubPlatform` | `codecanary review --post` (locally or in CI) | PR review threads via API | Posts review comments on the PR | +21: | **Local** | `LocalPlatform` | `codecanary review` (with or without a PR for the branch) | `~/.codecanary/repos///state/.json` | Prints to terminal | +22: +23: `codecanary review` without `--post` is always local — even if the branch has an open PR. "Local is local": the branch diff (including uncommitted changes) is reviewed against the default base, and previous findings come from the repo-scoped state file. There is no hybrid mode that reads GitHub but writes local state. Two consecutive local runs go incremental off the saved state (locked in by `TestLocalPlatformIncrementalHandoff` in `state_test.go`). +24: +25: State files are keyed by `owner/repo/branch` so the same branch name across different repos (e.g. `main` in repo A vs. repo B) no longer collides. When the git remote can't be resolved (detached worktree, no remote), state falls back to the legacy `~/.codecanary/state/.json` path. On first save after the upgrade, any existing legacy file is read once, migrated to the repo-scoped path, and the legacy copy is removed — transparent to the operator. +26: +27: ## Pipeline Steps +28: +29: ### 1. Fetch PR data +30: +31: **GitHub PR** (`--post`): Fetches PR metadata (title, body, author, branches) and diff via `gh pr view` and `gh pr diff`. +32: +33: **Local**: Detects the default branch (`main`, falling back to `master`, or the explicit `--base`) and computes diff from merge-base to HEAD via `git diff $(git merge-base HEAD )..HEAD`. Uncommitted working-tree changes scoped to the branch files are appended to the incremental diff on subsequent runs. Uses current branch name as the title and `git config user.name` as the author. +34: +35: If the PR is a setup PR (only adds workflow files with no real code changes), the review is skipped with an informational comment. +36: +37: ### 2. Prepare review context +38: +39: `prepareReview()` loads everything the review needs: +40: +41: - **Config**: Reads `config.yml` (provider, models, budgets, timeouts). If a `review.yml` exists alongside it, its rules/context/ignore fields override the config. If a `review.local.yml` also exists, its fields are appended (not replaced) on top of `review.yml`. +42: - **Project docs**: Discovers CLAUDE.md files at the repo root and in every ancestor directory of a changed PR file. Skips `vendor/`, `node_modules/`, hidden dirs, and other build artifacts. Up to 10 files, 16 KB each, 48 KB total. Monorepos commonly keep per-app conventions (e.g. `apps/exchange-api/CLAUDE.md`) — those load automatically when a PR touches files under that directory, so the reviewer sees the conventions specific to the code being changed rather than only the repo-root overview. +43: - **File contents**: Reads changed files from disk with size limits (default 100KB per file, 500KB total). Skips binary files, ignored patterns, and files exceeding limits. When files are skipped, the diff is also filtered to remove their hunks (via `ScopeDiffToFiles`) and they are removed from the file list. The original unfiltered diff is preserved in `FullDiff` for finding validation. +44: - **Environment**: Builds a filtered env for LLM subprocesses (only allowed prefixes like `CODECANARY_`, `GITHUB_`, plus essential vars like `PATH`). Injects keychain credentials if not already set. +45: +46: ### 3. Create providers +47: +48: Two `ModelProvider` instances are created from config: +49: +50: - **Review provider**: The main model that reviews code (configured via `review_model` in config). When `advisor_model` is set (anthropic or claude provider only), the review provider also enables Anthropic's server-side advisor tool so a stronger advisor model can weigh in mid-generation — the triage provider never uses advisor, since its classifier turns are too short to benefit. +51: - **Triage provider**: A cheaper model for re-evaluating previous findings (configured via `triage_model` in config). +52: +53: Each provider is constructed via the factory registry in `provider.go`. The provider name determines which adapter handles the API call (Anthropic, OpenAI, OpenRouter, or Claude CLI). +54: +55: ### 4. Load previous findings +56: +57: The platform adapter loads unresolved findings from the last review: +58: +59: **GitHub PR** (`--post`): Fetches review threads via GraphQL. Filters to CodeCanary findings only (detected by HTML marker comments). Extracts the previous review's HEAD SHA from the most recent review body — clean and all-clear reviews embed this marker too, so the baseline advances even when a push produced no findings. Returns unresolved threads, the SHA, and a count for fix_ref numbering. +60: +61: **Local**: Reads `~/.codecanary/repos///state/.json`, which stores the SHA, branch name, and findings array from the previous review. Falls back to `~/.codecanary/state/.json` when the repo slug can't be resolved or a pre-migration state file is present. Converts saved findings into `ReviewThread` shape for the triage pipeline. +62: +63: If no previous findings exist, this is a first review. +64: +65: ### 5. Triage and build prompt +66: +67: This step diverges based on whether a previous review SHA exists. A previous SHA alone is enough to enter the incremental path — if previous findings were all resolved (no open threads), the incremental diff still scopes the review to commits since the last baseline, avoiding a redundant full re-review. +68: +69: #### First review path +70: +71: Calls `BuildPrompt()` to assemble the full review prompt. The prompt includes (in order): +72: +73: 1. System instructions (reviewer role, diff-only rules, side-effect awareness) +74: 2. PR metadata (number, title, author, description) +75: 3. Additional context from config +76: 4. Project documentation (CLAUDE.md files in `` tags) +77: 5. Review rules (from config) — filtered to rules whose `paths:` / `exclude_paths:` globs match at least one PR file. Rules scoped to file types not in the diff (e.g. CSS rules on a Ruby-only change) are omitted to keep LLM attention focused. Falls back to a general review instruction when no rules apply. +78: 6. Ignore patterns +79: 7. Explicit file allowlist (anti-hallucination) +80: 8. Full contents of changed files with line numbers +81: 9. The unified diff +82: 10. Output format instructions (JSON schema, examples, escaping rules) +83: +84: After building, `fitPromptForModel()` checks whether the prompt fits the review model's context window (context window minus max output tokens). If it exceeds the budget, it progressively drops the largest file contents first, then truncates the diff as a last resort. +85: +86: #### Incremental review path (triage) +87: +88: `runTriage()` handles the incremental case in two phases. +89: +90: **Phase 1 -- Classify and evaluate previous findings** +91: +92: First, an incremental diff is computed (`git diff ..HEAD`). Two diffs serve different purposes: +93: +94: - **Activity diff** (incremental): Determines whether there's new activity to evaluate. If empty, threads with no replies are skipped (no LLM cost). +95: - **Context diff** (full PR diff): Used for classification and evaluation context. Ensures fixes from earlier pushes are visible even if they predate the incremental window. +96: +97: `ClassifyThreads()` assigns each unresolved thread one of six classifications: +98: +99: | Classification | Condition | Evaluation | +100: |---|---|---| +101: | `TriageSkip` | No activity diff, not outdated, no replies | Skipped (no LLM) | +102: | `TriageCodeChanged` | GitHub outdated flag, or file in PR diff | LLM evaluates with file-scoped diff + file snippet | +103: | `TriageHasReply` | Human replied (no code changes) | LLM evaluates reply intent | +104: | `TriageCodeChangedReply` | Both code changed and human replied | LLM evaluates both | +105: | `TriageCrossFileChange` | Changes in other files only | LLM evaluates with full PR diff | +106: | `TriageFileRemovedFromPR` | File no longer in PR | Auto-resolved by Go code (no LLM) -- thread resolved on GitHub | +107: +108: Threads classified as `TriageFileRemovedFromPR` are auto-resolved without an LLM call. The Go code sets reason `file_removed` and resolves the thread directly. +109: +110: For remaining threads, `EvaluateThreadsParallel()` runs up to 3 concurrent LLM calls using the triage model. Evaluation uses a **two-level approach** to balance precision and coverage: +111: +112: - **Level 1 (file-scoped)**: The LLM receives the finding, the current file content (presented first), and a file-scoped diff. This catches same-file fixes with minimal noise. Most evaluations resolve here. +113: - **Level 2 (widened scope)**: Only when level 1 says "not resolved" and the thread is `TriageCodeChanged` (file is in the PR diff). The LLM receives the full PR diff with a prompt primed to look for cross-file fixes. This catches the edge case where the finding's file has unrelated changes but the actual fix is in a different file. +114: +115: For `TriageCrossFileChange`, only the full PR diff is used (no level 1 — there's no file-scoped diff to show). +116: +117: The LLM returns JSON: `{"resolved": true, "reason": "code_change"}` or `{"resolved": false}`. +118: +119: LLM resolution reasons and their effects: +120: +121: | Reason | Effect | Thread stays open? | +122: |---|---|---| +123: | `code_change` | Thread resolved on GitHub | No | +124: | `dismissed` | Ack reply posted | Yes (re-triaged on next push) | +125: | `acknowledged` | Ack reply posted | Yes | +126: | `rebutted` | Ack reply posted | Yes | +127: +128: **Phase 2 -- Build prompt for new findings** +129: +130: After triage, the pipeline builds an incremental review prompt using `BuildIncrementalPrompt()`. This is similar to `BuildPrompt()` but: +131: +132: - Uses the incremental diff (or falls back to full PR diff if the incremental diff failed) +133: - Includes a "Known Issues" section listing unresolved threads (prevents duplicating them) +134: - Includes a "Recently Resolved Issues" section with findings fixed by code changes (prevents re-raising similar issues -- anti-ping-pong) +135: - Only includes file contents for files touched in the incremental diff +136: +137: The prompt is then fitted to the context window, same as the first review path. +138: +139: ### 6. LLM call +140: +141: If not a dry run and budget permits, the review prompt is sent to the review provider. The provider handles API communication (Anthropic Messages API, OpenAI Chat Completions, OpenRouter, or Claude CLI). +142: +143: If the response is truncated (hit max output tokens), a warning is logged. The pipeline attempts to salvage complete findings from the truncated JSON by scanning backward for valid objects. +144: +145: ### 7. Process findings +146: +147: `processFindings()` parses and validates the LLM's output: +148: +149: 1. **Parse JSON**: Extracts the findings array from the ```json fence. Falls back to bracket-matching if embedded code blocks break the regex. +150: 2. **File validation**: Drops findings referencing files not in the PR. +151: 3. **Line validation**: Drops findings whose line number is more than 20 lines from any changed line in the PR diff. This catches hallucinated line numbers and scope creep. +152: 4. **Actionable filter**: Removes findings where `actionable: false`. +153: 5. **Status tagging**: Tags all findings as `"new"` if this is an incremental review. +154: +155: ### 8. Publish results +156: +157: **GitHub PR** (`--post`): Every cycle emits exactly one top-level CodeCanary review, decided by an edit-vs-post rule. `FetchLatestCodecanaryReview` reads the commit SHA from the most recent CodeCanary review's hidden marker: +158: +159: - **Same SHA** (reply-only run, or a duplicate `synchronize` webhook on the same HEAD): the existing body is updated in place with `UpdateReviewBody`. Only the status block between the `` markers is swapped — inline comments and prior findings text are untouched. +160: - **Different or no SHA** (new commits, or first review on the PR): a fresh review is posted. The body variant depends on the cycle outcome — findings review, all-clear, activity summary (no new findings but cycle activity to surface), or clean review. All variants carry the same status block and baseline SHA marker. Older CodeCanary reviews are minimized (collapsed) before posting. +161: +162: The status block lists non-zero counts for: new findings, resolved by code, file removed, dismissed by author, acknowledged by author, rebutted by author, still unresolved. The block renders nothing when all counts are zero, so clean reviews remain copy-exact. +163: +164: Per-thread ack replies for dismissed/acknowledged/rebutted resolutions are posted earlier in the pipeline (`HandleResolutions`). Dedup is reason-agnostic: if the thread already carries *any* `` marker, no further ack reply is posted. Reasons can shift across triage runs (LLM non-determinism), and all three convey the same outcome ("keeping open"), so one ack per thread is enough. Reply-only runs skip `SaveState` so the empty findings slice doesn't overwrite persisted state. +165: +166: `codecanary findings` applies the same marker to filter deferrals out of its default output: threads with any `codecanary:ack:*` reply are treated as handled and omitted alongside GitHub-resolved threads. Pass `--include-resolved` to see them. This keeps the codecanary-fix skill from re-prompting on findings the operator already deferred. +167: +168: After the review is posted (or updated in place), `GithubPlatform.Publish` also POSTs a `CodeCanary / review` commit status on the reviewed SHA via `PostReviewCommitStatus`. State is `success` when `NewFindings + StillOpen == 0` (everything is either new-and-green, fixed by code, or explicitly handled by the author), and `failure` otherwise. Description is the unresolved count or "all findings resolved" / "no findings". Teams that add `CodeCanary / review` as a required status check in branch protection get auto-gating: merges are blocked until a review run posts a green status on HEAD. Status posting failures are logged as warnings and do not abort Publish — the review itself has already landed. The local `codecanary signoff` command posts a status under the same context, so a team can rely on a single required check that either the bot (pr-loop) or a local reviewer (local-loop) satisfies. +169: +170: **Local**: Prints the formatted result to stdout. Format depends on context: terminal (colored, human-readable), markdown, or JSON. +171: +172: ### 9. Save state +173: +174: **GitHub PR** (`--post`): No-op. State is stored in the review threads themselves (the embedded JSON marker contains the SHA and findings). +175: +176: **Local**: Writes `~/.codecanary/repos///state/.json` with the current HEAD SHA, branch name, and combined findings (still-open + new). If a legacy `~/.codecanary/state/.json` exists from a pre-migration run, it is removed after the new file lands. This enables incremental reviews on the next run. +177: +178: ### 10. Report usage +179: +180: **GitHub PR** (`--post`): Writes token counts and cost to `GITHUB_ENV` for downstream workflow steps. +181: +182: **Local**: Prints a usage summary table to stderr (model, tokens, cost, duration) if running in a terminal. +183: +184: ### 11. Telemetry +185: +186: If telemetry is enabled (opt-in), fires an anonymous event with aggregate stats: provider, platform, finding counts by severity, token counts, cost, and duration. No code content is sent. +187: +188: ## Key Design Decisions +189: +190: **Single pipeline, two platforms.** `Run()` never branches on "am I on GitHub?" The `ReviewPlatform` interface absorbs all environment differences. Adding a new platform (e.g. GitLab) means implementing the interface, not forking the pipeline. +191: +192: **Two diffs for triage.** The incremental diff (changes since last review) decides whether to skip evaluation. The full PR diff (all changes) provides context for evaluation. This prevents the "triage horizon" bug where fixes committed before the triage baseline become invisible. +193: +194: **Two-level triage evaluation.** Same-file evaluations (`TriageCodeChanged`) start with a file-scoped diff (level 1) to reduce noise — the full PR diff can drown out the relevant fix with changes from unrelated files. If level 1 finds no fix, a widened-scope fallback (level 2) sends the full PR diff to catch cross-file fixes. Cross-file evaluations (`TriageCrossFileChange`) go straight to the full diff. The file snippet (current code state) is presented first in all evaluation prompts, so the LLM checks whether the issue still exists before analyzing the diff. +195: +196: **Per-thread evaluation.** Each unresolved thread gets its own LLM call with tailored context, rather than one bulk prompt. This allows fine-grained classification, parallel execution, and per-thread budget control. +197: +198: **Anti-ping-pong.** The incremental prompt includes recently resolved findings so the LLM doesn't re-raise similar issues. Non-code resolutions (dismissed, acknowledged, rebutted) keep threads open for re-triage on future pushes, but post ack replies to avoid duplicate acknowledgments. +199: +200: **Context window fitting.** After building the prompt, the pipeline estimates token count and progressively trims file contents (largest first) then diff to fit the model's context window. This prevents API failures on large PRs. +201: +202: **Finding validation.** All findings are validated against the PR diff regardless of what diff the LLM prompt contained. Line proximity checks (within 20 lines of a changed line) catch hallucinated line numbers and prevent scope creep from rebase noise. +203: +204: ## The codecanary-fix loop +205: +206: The `codecanary-fix` Claude skill wraps the review pipeline in a confirm-and-apply loop. The skill calls `codecanary mode --output json` once at startup; the CLI returns one of three modes based on whether an open PR exists for the current branch and whether a CodeCanary workflow file is detected under `.github/workflows/`: +207: +208: | Mode | PR | Workflow | Findings source | Cycle finalization | +209: |---|---|---|---|---| +210: | `pr-loop` | yes | yes | `codecanary findings --watch` (bot posts on push) | commit + push; bot re-runs | +211: | `local-loop-git` | yes | no | `codecanary review` (local engine) | commit on PR branch, **no push**; operator is asked at session end whether to push accumulated commits | +212: | `local-loop-nogit` | no | — | `codecanary review` (local engine) | no commits, no pushes; fixes applied in place | +213: +214: Workflow detection is a textual scan for a non-commented `uses: alansikora/codecanary...` step in any workflow file on the current branch. All three modes share the same triage UX (Markdown table, `AskUserQuestion` confirmation). The bot's ack layer (`` markers) handles deferral persistence in `pr-loop`; local modes use an in-memory `DEFERRED_FIX_REFS` set in the skill so operator-skipped findings don't re-surface within a session. `pr-loop` failures (GHA broken, `conclusion: failure`) never silently fall back to a local mode — the operator is asked to investigate. +```` + +### `internal/review/github.go` +``` +1: package review +2: +3: import ( +4: "bytes" +5: "encoding/json" +6: "fmt" +7: "os" +8: "os/exec" +9: "path/filepath" +10: "regexp" +11: "strconv" +12: "strings" +13: +14: "github.com/bmatcuk/doublestar/v4" +15: ) +16: +17: // runGH runs a gh command and returns stdout; on failure, the returned error +18: // includes gh's stderr so upstream API messages surface in logs instead of just +19: // "exit status 1". +20: func runGH(label string, args ...string) ([]byte, error) { +21: cmd := exec.Command("gh", args...) +22: var stderr bytes.Buffer +23: cmd.Stderr = &stderr +24: out, err := cmd.Output() +25: if err != nil { +26: msg := strings.TrimSpace(stderr.String()) +27: if msg == "" { +28: return nil, fmt.Errorf("%s: %w", label, err) +29: } +30: return nil, fmt.Errorf("%s: %s: %w", label, msg, err) +31: } +32: return out, nil +33: } +34: +35: // parseRepoSlug splits a "owner/name" repository slug into its two parts. +36: func parseRepoSlug(repo string) (owner, name string, err error) { +37: parts := strings.SplitN(repo, "/", 2) +38: if len(parts) != 2 { +39: return "", "", fmt.Errorf("invalid repo format %q, expected owner/name", repo) +40: } +41: return parts[0], parts[1], nil +42: } +43: +44: // MaxFindingProximity is the maximum number of lines a finding may be from the +45: // nearest changed line in the PR diff. Findings beyond this distance are dropped +46: // (runner.go) or demoted from inline to body (PostReview). This enforces review +47: // scope — keeping findings anchored to the PR's actual changes — and catches +48: // hallucinated line numbers. A single constant ensures both checks stay in sync. +49: const MaxFindingProximity = 20 +50: +51: // HTML comment markers for embedding and detecting review data. +52: // Dual prefixes support both current (codecanary) and legacy (clanopy) markers. +53: var reviewMarkerPrefixes = []string{"" +57: findingMarkerPrefix = "" +371: +372: // Search from most recent to oldest. +373: for i := len(reviews) - 1; i >= 0; i-- { +374: body := reviews[i].Body +375: for _, prefix := range prefixes { +376: idx := strings.Index(body, prefix) +377: if idx < 0 { +378: continue +379: } +380: start := idx + len(prefix) +381: endIdx := strings.Index(body[start:], suffix) +382: if endIdx < 0 { +383: continue +384: } +385: jsonData := body[start : start+endIdx] +386: var result ReviewResult +387: if err := json.Unmarshal([]byte(jsonData), &result); err != nil { +388: continue +389: } +390: return &result, nil +391: } +392: } +393: +394: return nil, fmt.Errorf("no review data found in PR #%d reviews", prNumber) +395: } +396: +397: // FetchFindingFromPR searches all reviews on a PR for a specific fix_ref. +398: // Unlike FetchReviewFromPR (which returns the latest review), this searches every +399: // review so that fix_ref links from older review rounds still resolve correctly. +400: func FetchFindingFromPR(repo string, prNumber int, fixRef string) (*Finding, error) { +401: owner, name, err := parseRepoSlug(repo) +402: if err != nil { +403: return nil, err +404: } +405: +406: apiPath := fmt.Sprintf("repos/%s/%s/pulls/%d/reviews", owner, name, prNumber) +407: out, err := exec.Command("gh", "api", apiPath).Output() +408: if err != nil { +409: return nil, fmt.Errorf("fetching PR reviews: %w", err) +410: } +411: +412: var reviews []ghReview +413: if err := json.Unmarshal(out, &reviews); err != nil { +414: return nil, fmt.Errorf("parsing PR reviews: %w", err) +415: } +416: +417: prefixes := reviewMarkerPrefixes +418: const suffix = " -->" +419: +420: for _, rev := range reviews { +421: for _, prefix := range prefixes { +422: idx := strings.Index(rev.Body, prefix) +423: if idx < 0 { +424: continue +425: } +426: start := idx + len(prefix) +427: endIdx := strings.Index(rev.Body[start:], suffix) +428: if endIdx < 0 { +429: continue +430: } +431: var result ReviewResult +432: if err := json.Unmarshal([]byte(rev.Body[start:start+endIdx]), &result); err != nil { +433: continue +434: } +435: for i := range result.Findings { +436: if result.Findings[i].FixRef == fixRef { +437: return &result.Findings[i], nil +438: } +439: } +440: } +441: } +442: +443: return nil, fmt.Errorf("fix_ref %q not found in any review on PR #%d", fixRef, prNumber) +444: } +445: +446: // DetectRepo gets owner/name from the current git remote. +447: func DetectRepo() (string, error) { +448: out, err := exec.Command("gh", "repo", "view", +449: "--json", "nameWithOwner", +450: "--jq", ".nameWithOwner", +451: ).Output() +452: if err != nil { +453: return "", fmt.Errorf("gh repo view: %w", err) +454: } +455: return strings.TrimSpace(string(out)), nil +456: } +457: +458: // DetectPRNumber detects the PR number for the current branch using gh. +459: // If repo is non-empty, it is passed as --repo to scope the lookup. +460: func DetectPRNumber(repo string) (int, error) { +461: args := []string{"pr", "view", "--json", "number", "--jq", ".number"} +462: if repo != "" { +463: args = append(args, "--repo", repo) +464: } +465: out, err := exec.Command("gh", args...).Output() +466: if err != nil { +467: return 0, fmt.Errorf("no open pull request found for the current branch") +468: } +469: num, err := strconv.Atoi(strings.TrimSpace(string(out))) +470: if err != nil { +471: return 0, fmt.Errorf("unexpected PR number from gh: %w", err) +472: } +473: return num, nil +474: } +475: +476: // ThreadReply represents a reply to a review thread (i.e. any comment after the first). +477: type ThreadReply struct { +478: Author string +479: Body string +480: } +481: +482: // ReviewThread represents a review thread from a PR. +483: type ReviewThread struct { +484: ID string +485: Path string +486: Line int +487: Body string +488: Author string // login of the first comment author (the bot for review threads) +489: Outdated bool // true if GitHub marked the comment position as outdated (code changed) +490: Resolved bool +491: Replies []ThreadReply +492: } +493: +494: // graphQLThreadsResponse is the JSON shape returned by the review threads query. +495: type graphQLThreadsResponse struct { +496: Data struct { +497: Repository struct { +498: PullRequest struct { +499: ReviewThreads struct { +500: Nodes []struct { +501: ID string `json:"id"` +502: IsResolved bool `json:"isResolved"` +503: Comments struct { +504: Nodes []struct { +505: Body string `json:"body"` +506: Path string `json:"path"` +507: Line int `json:"line"` +508: OriginalLine int `json:"originalLine"` +509: Outdated bool `json:"outdated"` +510: Author struct { +511: Login string `json:"login"` +512: } `json:"author"` +513: } `json:"nodes"` +514: } `json:"comments"` +515: } `json:"nodes"` +516: } `json:"reviewThreads"` +517: } `json:"pullRequest"` +518: } `json:"repository"` +519: } `json:"data"` +520: } +521: +522: // FetchReviewThreads gets all review threads from a PR via GraphQL. +523: func FetchReviewThreads(repo string, prNumber int) ([]ReviewThread, error) { +524: owner, name, err := parseRepoSlug(repo) +525: if err != nil { +526: return nil, err +527: } +528: +529: query := `query($owner:String!,$name:String!,$pr:Int!){ +530: repository(owner:$owner,name:$name){ +531: pullRequest(number:$pr){ +532: reviewThreads(first:100){ +533: nodes{ +534: id +535: isResolved +536: comments(first:100){ +537: nodes{body path line originalLine outdated author{login}} +538: } +539: } +540: } +541: } +542: } +543: }` +544: +545: cmd := exec.Command("gh", "api", "graphql", +546: "-f", "query="+query, +547: "-f", fmt.Sprintf("owner=%s", owner), +548: "-f", fmt.Sprintf("name=%s", name), +549: "-F", fmt.Sprintf("pr=%d", prNumber), +550: ) +551: out, err := cmd.Output() +552: if err != nil { +553: return nil, fmt.Errorf("gh api graphql: %w", err) +554: } +555: +556: var resp graphQLThreadsResponse +557: if err := json.Unmarshal(out, &resp); err != nil { +558: return nil, fmt.Errorf("parsing graphql response: %w", err) +559: } +560: +561: var threads []ReviewThread +562: for _, node := range resp.Data.Repository.PullRequest.ReviewThreads.Nodes { +563: if len(node.Comments.Nodes) == 0 { +564: continue +565: } +566: comment := node.Comments.Nodes[0] +567: +568: // Filter to review threads only (new marker + legacy markers for backward compat). +569: if !strings.Contains(comment.Body, findingMarkerPrefix) && +570: !strings.Contains(comment.Body, "codecanary fix") && +571: !strings.Contains(comment.Body, "clanopy fix") { +572: continue +573: } +574: +575: var replies []ThreadReply +576: for _, c := range node.Comments.Nodes[1:] { +577: replies = append(replies, ThreadReply{ +578: Author: c.Author.Login, +579: Body: c.Body, +580: }) +581: } +582: +583: line := comment.Line +584: if comment.Outdated && line == 0 && comment.OriginalLine > 0 { +585: line = comment.OriginalLine +586: } +587: +588: threads = append(threads, ReviewThread{ +589: ID: node.ID, +590: Path: comment.Path, +591: Line: line, +592: Body: comment.Body, +593: Author: comment.Author.Login, +594: Outdated: comment.Outdated, +595: Resolved: node.IsResolved, +596: Replies: replies, +597: }) +598: } +599: +600: return threads, nil +601: } +602: +603: // ResolveThread resolves a review thread via GraphQL mutation. +604: func ResolveThread(threadID string) error { +605: cmd := exec.Command("gh", "api", "graphql", +606: "-f", "query=mutation($threadId:ID!){resolveReviewThread(input:{threadId:$threadId}){thread{isResolved}}}", +607: "-f", fmt.Sprintf("threadId=%s", threadID), +608: ) +609: if out, err := cmd.CombinedOutput(); err != nil { +610: return fmt.Errorf("gh api graphql resolve: %w\n%s", err, string(out)) +611: } +612: return nil +613: } +614: +615: // ReplyToThread posts a reply on a review thread via GraphQL. +616: func ReplyToThread(threadID, body string) error { +617: cmd := exec.Command("gh", "api", "graphql", +618: "-f", "query=mutation($threadId:ID!,$body:String!){addPullRequestReviewThreadReply(input:{pullRequestReviewThreadId:$threadId,body:$body}){comment{id}}}", +619: "-f", fmt.Sprintf("threadId=%s", threadID), +620: "-f", fmt.Sprintf("body=%s", body), +621: ) +622: if out, err := cmd.CombinedOutput(); err != nil { +623: return fmt.Errorf("gh api graphql reply: %w\n%s", err, string(out)) +624: } +625: return nil +626: } +627: +628: // ghReview is the JSON shape for a PR review from the REST API. +629: type ghReview struct { +630: ID int64 `json:"id"` +631: NodeID string `json:"node_id"` +632: Body string `json:"body"` +633: } +634: +635: // LatestCodecanaryReview captures the identity and body of the most recent +636: // CodeCanary top-level review, together with the commit SHA embedded in its +637: // hidden marker. Returned by FetchLatestCodecanaryReview for the +638: // "edit if same SHA, otherwise post new" publish decision. +639: type LatestCodecanaryReview struct { +640: ID int64 +641: SHA string +642: Body string +643: } +644: +645: // FetchLatestCodecanaryReview returns the most recent top-level review on +646: // the PR that carries a CodeCanary marker, along with the commit SHA +647: // embedded in that marker. Returns (nil, nil) when no matching review +648: // exists. A non-nil error is returned only for transport/parsing failures. +649: func FetchLatestCodecanaryReview(repo string, prNumber int) (*LatestCodecanaryReview, error) { +650: owner, name, err := parseRepoSlug(repo) +651: if err != nil { +652: return nil, err +653: } +654: +655: apiPath := fmt.Sprintf("repos/%s/%s/pulls/%d/reviews", owner, name, prNumber) +656: out, err := exec.Command("gh", "api", apiPath).Output() +657: if err != nil { +658: return nil, fmt.Errorf("fetching PR reviews: %w", err) +659: } +660: +661: var reviews []ghReview +662: if err := json.Unmarshal(out, &reviews); err != nil { +663: return nil, fmt.Errorf("parsing PR reviews: %w", err) +664: } +665: +666: // Walk newest → oldest so we return the latest CodeCanary review. +667: for i := len(reviews) - 1; i >= 0; i-- { +668: rev := reviews[i] +669: for _, prefix := range reviewMarkerPrefixes { +670: idx := strings.Index(rev.Body, prefix) +671: if idx < 0 { +672: continue +673: } +674: start := idx + len(prefix) +675: endIdx := strings.Index(rev.Body[start:], reviewMarkerSuffix) +676: if endIdx < 0 { +677: continue +678: } +679: jsonData := rev.Body[start : start+endIdx] +680: var payload struct { +681: SHA string `json:"sha"` +682: } +683: // Ignore unmarshal errors so old reviews that embed a richer +684: // ReviewResult (also containing a "sha" field) still match. +685: _ = json.Unmarshal([]byte(jsonData), &payload) +686: return &LatestCodecanaryReview{ +687: ID: rev.ID, +688: SHA: payload.SHA, +689: Body: rev.Body, +690: }, nil +691: } +692: } +693: +694: return nil, nil +695: } +696: +697: // UpdateReviewBody replaces the body of an existing PR review via the +698: // GitHub REST API. Used when a reply-only run (or a duplicate synchronize +699: // webhook on the same HEAD) needs to refresh the latest CodeCanary review's +700: // counts instead of posting a new top-level comment. +701: func UpdateReviewBody(repo string, prNumber int, reviewID int64, body string) error { +702: owner, name, err := parseRepoSlug(repo) +703: if err != nil { +704: return err +705: } +706: +707: payload, err := json.Marshal(struct { +708: Body string `json:"body"` +709: }{Body: body}) +710: if err != nil { +711: return fmt.Errorf("marshaling update payload: %w", err) +712: } +713: +714: apiPath := fmt.Sprintf("repos/%s/%s/pulls/%d/reviews/%d", owner, name, prNumber, reviewID) +715: _, err = ghAPIRequest("PUT", apiPath, payload) +716: return err +717: } +718: +719: // ghAPIRequest is a generic gh api wrapper that preserves the temp-file +720: // pattern used by ghAPIPOST (avoids stdin pipe truncation on large bodies). +721: func ghAPIRequest(method, apiPath string, payloadJSON []byte) ([]byte, error) { +722: dir, err := os.MkdirTemp("", "codecanary-*") +723: if err != nil { +724: return nil, fmt.Errorf("creating temp dir: %w", err) +725: } +726: defer func() { _ = os.RemoveAll(dir) }() +727: +728: tmpFile, err := os.CreateTemp(dir, "payload.json") +729: if err != nil { +730: return nil, fmt.Errorf("creating temp file: %w", err) +731: } +732: if _, err := tmpFile.Write(payloadJSON); err != nil { +733: _ = tmpFile.Close() +734: return nil, fmt.Errorf("writing payload to temp file: %w", err) +735: } +736: _ = tmpFile.Close() +737: +738: cmd := exec.Command("gh", "api", apiPath, "--method", method, "--input", tmpFile.Name()) +739: var stdout, stderr bytes.Buffer +740: cmd.Stdout = &stdout +741: cmd.Stderr = &stderr +742: if err := cmd.Run(); err != nil { +743: return stdout.Bytes(), &apiError{Err: err, Stderr: stderr.String(), Response: stdout.String()} +744: } +745: return stdout.Bytes(), nil +746: } +747: +748: // FetchPreviousReviewSHA gets the SHA from the last review's hidden data. +749: func FetchPreviousReviewSHA(repo string, prNumber int) string { +750: owner, name, err := parseRepoSlug(repo) +751: if err != nil { +752: return "" +753: } +754: +755: apiPath := fmt.Sprintf("repos/%s/%s/pulls/%d/reviews", owner, name, prNumber) +756: out, err := exec.Command("gh", "api", apiPath).Output() +757: if err != nil { +758: return "" +759: } +760: +761: var reviews []ghReview +762: if err := json.Unmarshal(out, &reviews); err != nil { +763: return "" +764: } +765: +766: prefixes := reviewMarkerPrefixes +767: const suffix = " -->" +768: +769: // Search from most recent to oldest. +770: for i := len(reviews) - 1; i >= 0; i-- { +771: body := reviews[i].Body +772: for _, prefix := range prefixes { +773: idx := strings.Index(body, prefix) +774: if idx < 0 { +775: continue +776: } +777: start := idx + len(prefix) +778: endIdx := strings.Index(body[start:], suffix) +779: if endIdx < 0 { +780: continue +781: } +782: jsonData := body[start : start+endIdx] +783: var result ReviewResult +784: if err := json.Unmarshal([]byte(jsonData), &result); err != nil { +785: continue +786: } +787: if result.SHA != "" { +788: return result.SHA +789: } +790: } +791: } +792: +793: return "" +794: } +795: +796: // FilesFromDiff extracts the list of file paths touched in a unified diff. +797: func FilesFromDiff(diff string) []string { +798: var files []string +799: seen := make(map[string]bool) +800: for _, line := range strings.Split(diff, "\n") { +801: if strings.HasPrefix(line, "+++ b/") { +802: path := strings.TrimRight(line[6:], "\r") +803: if !seen[path] { +804: seen[path] = true +805: files = append(files, path) +806: } +807: } +808: } +809: return files +810: } +811: +812: // ScopeDiffToFiles filters a unified diff to only include hunks for files in +813: // the allowed set. This prevents rebase noise (main-branch changes) from +814: // leaking into incremental reviews. +815: func ScopeDiffToFiles(diff string, allowedFiles map[string]bool) string { +816: if len(allowedFiles) == 0 { +817: return diff +818: } +819: +820: lines := strings.Split(diff, "\n") +821: var result []string +822: blockStart := -1 +823: blockAllowed := false +824: +825: for i := 0; i < len(lines); i++ { +826: if strings.HasPrefix(lines[i], "diff --git") { +827: // Flush previous block if allowed. +828: if blockStart >= 0 && blockAllowed { +829: result = append(result, lines[blockStart:i]...) +830: } +831: blockStart = i +832: blockAllowed = false +833: +834: // Look ahead for +++ b/ to determine if this block is allowed. +835: for j := i + 1; j < len(lines) && !strings.HasPrefix(lines[j], "diff --git"); j++ { +836: if strings.HasPrefix(lines[j], "+++ b/") { +837: path := strings.TrimRight(lines[j][6:], "\r") +838: if allowedFiles[path] { +839: blockAllowed = true +840: } +841: break +842: } +843: } +844: continue +845: } +846: } +847: +848: // Flush last block. +849: if blockStart >= 0 && blockAllowed { +850: result = append(result, lines[blockStart:]...) +851: } +852: +853: if len(result) == 0 { +854: return "" +855: } +856: return strings.Join(result, "\n") +857: } +858: +859: // validSHA matches a full-length lowercase hex Git SHA. +860: var validSHA = regexp.MustCompile(`^[0-9a-f]{40}$`) +861: +862: // GetIncrementalDiff gets the diff since a given SHA. +863: func GetIncrementalDiff(baseSHA string) (string, error) { +864: if !validSHA.MatchString(baseSHA) { +865: return "", fmt.Errorf("invalid SHA format: %q", baseSHA) +866: } +867: out, err := exec.Command("git", "diff", baseSHA+"..HEAD").Output() +868: if err != nil { +869: return "", fmt.Errorf("git diff: %w", err) +870: } +871: return string(out), nil +872: } +873: +874: // PostCleanReview posts a review when the first review finds no issues. The +875: // commitSHA is embedded in a hidden marker so future runs treat it as the +876: // baseline for incremental reviews, avoiding a redundant full re-review on the +877: // next push. +878: func PostCleanReview(repo string, prNumber int, commitSHA string, summary ReviewSummary) error { +879: return postSimpleReview(repo, prNumber, buildCleanReviewBody(commitSHA, summary)) +880: } +881: +882: // PostAllClearReview posts a review when all previous findings have been +883: // resolved. If minimizeFailed is true, a note is appended warning about +884: // visible old reviews. The commitSHA is embedded in a hidden marker so future +885: // runs treat it as the baseline for incremental reviews; without it, the next +886: // push would fall back to reviewing the entire PR again. +887: func PostAllClearReview(repo string, prNumber int, commitSHA string, minimizeFailed bool, summary ReviewSummary) error { +888: return postSimpleReview(repo, prNumber, buildAllClearReviewBody(commitSHA, minimizeFailed, summary)) +889: } +890: +891: // PostActivityReview posts a review when no new findings were raised but +892: // there is cycle activity worth surfacing (dismissals, acknowledgments, +893: // rebuttals, still-open threads). This keeps every commit push producing a +894: // visible top-level status comment instead of silently logging. +895: func PostActivityReview(repo string, prNumber int, commitSHA string, summary ReviewSummary) error { +896: return postSimpleReview(repo, prNumber, buildActivityReviewBody(commitSHA, summary)) +897: } +898: +899: // buildCleanReviewBody renders the full Markdown body posted by +900: // PostCleanReview. Split out from the poster so tests can assert the exact +901: // string that lands on GitHub without having to mock gh. +902: func buildCleanReviewBody(commitSHA string, summary ReviewSummary) string { +903: return withSummary("CodeCanary reviewed this PR \u2014 no issues found.", summary) + embedBaselineMarker(commitSHA) +904: } +905: +906: // buildAllClearReviewBody renders the full Markdown body posted by +907: // PostAllClearReview. Split out for the same reason as buildCleanReviewBody. +908: func buildAllClearReviewBody(commitSHA string, minimizeFailed bool, summary ReviewSummary) string { +909: body := "## \U0001F425 CodeCanary\n\n\u2705 All previous findings have been addressed. No new issues found. \u2728" +910: if minimizeFailed { +911: body += "\n\n> \u26A0\uFE0F Some previous review comments could not be minimized and may still be visible." +912: } +913: return withSummary(body, summary) + embedBaselineMarker(commitSHA) +914: } +915: +916: // buildActivityReviewBody renders the body for a commit push that raised no +917: // new findings but has cycle activity (dismissals/acknowledgments/rebuttals +918: // or still-open threads carried forward). +919: func buildActivityReviewBody(commitSHA string, summary ReviewSummary) string { +920: body := "## \U0001F425 CodeCanary\n\nReviewed this push \u2014 no new issues found." +921: return withSummary(body, summary) + embedBaselineMarker(commitSHA) +922: } +923: +924: // withSummary appends the status summary block to a review body. The block +925: // is skipped when the summary has no non-zero counts, so existing clean/ +926: // all-clear bodies render identically when nothing happened. +927: func withSummary(body string, summary ReviewSummary) string { +928: block := renderSummaryBlock(summary) +929: if block == "" { +930: return body +931: } +932: return body + block +933: } +934: +935: // embedBaselineMarker returns a hidden HTML comment containing the commitSHA +936: // so FetchPreviousReviewSHA can use this review as the incremental baseline. +937: // Returns an empty string if commitSHA is empty (local mode, dry run). The +938: // marker carries only the SHA — FetchPreviousReviewSHA is the sole reader and +939: // it only needs that field. +940: func embedBaselineMarker(commitSHA string) string { +941: if commitSHA == "" { +942: return "" +943: } +944: data, err := json.Marshal(struct { +945: SHA string `json:"sha"` +946: }{SHA: commitSHA}) +947: if err != nil { +948: return "" +949: } +950: return fmt.Sprintf("\n%s%s%s\n", reviewMarkerPrefixes[0], string(data), reviewMarkerSuffix) +951: } +952: +953: func postSimpleReview(repo string, prNumber int, body string) error { +954: owner, repoName, err := parseRepoSlug(repo) +955: if err != nil { +956: return err +957: } +958: +959: payload := reviewPayload{ +960: Event: "COMMENT", +961: Body: body, +962: Comments: make([]reviewComment, 0), +963: } +964: +965: payloadJSON, err := json.Marshal(payload) +966: if err != nil { +967: return fmt.Errorf("marshaling payload: %w", err) +968: } +969: +970: _, err = ghAPIPOST(fmt.Sprintf("repos/%s/%s/pulls/%d/reviews", owner, repoName, prNumber), payloadJSON) +971: return err +972: } +973: +974: // apiError is returned by ghAPIPOST so callers can inspect the stderr output +975: // from gh (which contains the HTTP status line) separately from the response body. +976: type apiError struct { +977: Err error +978: Stderr string +979: Response string +980: } +981: +982: func (e *apiError) Error() string { +983: return fmt.Sprintf("gh api: %v\nstderr: %s\nresponse: %s", e.Err, e.Stderr, e.Response) +984: } +985: +986: func (e *apiError) Unwrap() error { return e.Err } +987: +988: // ghAPIPOST sends a JSON payload to the GitHub API via gh, using a temp file +989: // to avoid stdin pipe issues that can cause "unexpected end of JSON input" +990: // errors on large payloads. +991: func ghAPIPOST(apiPath string, payloadJSON []byte) ([]byte, error) { +992: dir, err := os.MkdirTemp("", "codecanary-*") +993: if err != nil { +994: return nil, fmt.Errorf("creating temp dir: %w", err) +995: } +996: defer func() { _ = os.RemoveAll(dir) }() +997: +998: tmpFile, err := os.CreateTemp(dir, "payload.json") +999: if err != nil { +1000: return nil, fmt.Errorf("creating temp file: %w", err) +1001: } +1002: +1003: if _, err := tmpFile.Write(payloadJSON); err != nil { +1004: _ = tmpFile.Close() +1005: return nil, fmt.Errorf("writing payload to temp file: %w", err) +1006: } +1007: _ = tmpFile.Close() +1008: +1009: cmd := exec.Command("gh", "api", apiPath, "--method", "POST", "--input", tmpFile.Name()) +1010: var stdout, stderr bytes.Buffer +1011: cmd.Stdout = &stdout +1012: cmd.Stderr = &stderr +1013: if err := cmd.Run(); err != nil { +1014: return stdout.Bytes(), &apiError{Err: err, Stderr: stderr.String(), Response: stdout.String()} +1015: } +1016: return stdout.Bytes(), nil +1017: } +1018: +1019: // FindReviewNodeIDs returns the node_ids of all reviews on a PR. +1020: func FindReviewNodeIDs(repo string, prNumber int) ([]string, error) { +1021: owner, name, err := parseRepoSlug(repo) +1022: if err != nil { +1023: return nil, err +1024: } +1025: +1026: apiPath := fmt.Sprintf("repos/%s/%s/pulls/%d/reviews", owner, name, prNumber) +1027: out, err := exec.Command("gh", "api", apiPath).Output() +1028: if err != nil { +1029: return nil, fmt.Errorf("fetching PR reviews: %w", err) +1030: } +1031: +1032: var reviews []ghReview +1033: if err := json.Unmarshal(out, &reviews); err != nil { +1034: return nil, fmt.Errorf("parsing PR reviews: %w", err) +1035: } +1036: +1037: prefixes := reviewMarkerPrefixes +1038: +1039: var nodeIDs []string +1040: for _, rev := range reviews { +1041: for _, prefix := range prefixes { +1042: if strings.Contains(rev.Body, prefix) { +1043: nodeIDs = append(nodeIDs, rev.NodeID) +1044: break +1045: } +1046: } +1047: } +1048: +1049: return nodeIDs, nil +1050: } +1051: +1052: // threadHeaderLine returns the first non-empty, non-HTML-comment line from +1053: // a thread body. This is the header line containing severity and finding ID. +1054: func threadHeaderLine(body string) string { +1055: for _, line := range strings.SplitN(body, "\n", 5) { +1056: line = strings.TrimSpace(line) +1057: if line == "" || strings.HasPrefix(line, "" +1252: +1253: var result []ReviewInfo +1254: for _, rev := range reviews { +1255: for _, prefix := range prefixes { +1256: idx := strings.Index(rev.Body, prefix) +1257: if idx < 0 { +1258: continue +1259: } +1260: start := idx + len(prefix) +1261: endIdx := strings.Index(rev.Body[start:], suffix) +1262: if endIdx < 0 { +1263: continue +1264: } +1265: jsonData := rev.Body[start : start+endIdx] +1266: var rr ReviewResult +1267: if err := json.Unmarshal([]byte(jsonData), &rr); err != nil { +1268: continue +1269: } +1270: var ids []string +1271: for _, f := range rr.Findings { +1272: ids = append(ids, f.ID) +1273: } +1274: result = append(result, ReviewInfo{ +1275: NodeID: rev.NodeID, +1276: FindingIDs: ids, +1277: }) +1278: break +1279: } +1280: } +1281: +1282: return result, nil +1283: } +1284: +1285: // MinimizeComment hides a comment on GitHub using the minimizeComment GraphQL mutation. +1286: func MinimizeComment(nodeID string) error { +1287: cmd := exec.Command("gh", "api", "graphql", +1288: "-f", "query=mutation($id:ID!){minimizeComment(input:{subjectId:$id,classifier:RESOLVED}){minimizedComment{isMinimized}}}", +1289: "-F", "id="+nodeID, +1290: ) +1291: if out, err := cmd.CombinedOutput(); err != nil { +1292: return fmt.Errorf("gh api graphql minimize: %w\n%s", err, string(out)) +1293: } +1294: return nil +1295: } +1296: +1297: // FetchFileContents reads the full contents of changed files from disk. +1298: // It skips files that are too large, binary, deleted, or match ignore patterns. +1299: // Returns a map of path->content and a list of skipped file paths. +1300: func FetchFileContents(files []string, ignorePatterns []string, maxPerFile, maxTotal int) (map[string]string, []string) { +1301: contents := make(map[string]string) +1302: var skipped []string +1303: totalSize := 0 +1304: +1305: for _, path := range files { +1306: // Check ignore patterns. +1307: if matchesIgnore(path, ignorePatterns) { +1308: skipped = append(skipped, path) +1309: continue +1310: } +1311: +1312: data, err := os.ReadFile(path) +1313: if err != nil { +1314: // File may have been deleted in this PR — skip gracefully. +1315: continue +1316: } +1317: +1318: // Skip binary files (null bytes in first 512 bytes). +1319: peek := data +1320: if len(peek) > 512 { +1321: peek = peek[:512] +1322: } +1323: if bytes.ContainsRune(peek, 0) { +1324: skipped = append(skipped, path) +1325: continue +1326: } +1327: +1328: size := len(data) +1329: +1330: // Skip files exceeding per-file limit. +1331: if size > maxPerFile { +1332: skipped = append(skipped, path) +1333: continue +1334: } +1335: +1336: // Stop if total budget would be exceeded. +1337: if totalSize+size > maxTotal { +1338: skipped = append(skipped, path) +1339: continue +1340: } +1341: +1342: contents[path] = string(data) +1343: totalSize += size +1344: } +1345: +1346: return contents, skipped +1347: } +1348: +1349: // isSetupPR detects whether this is the initial setup PR. +1350: // Returns true only when a new workflow file referencing codecanary is added AND +1351: // the PR contains no other files beyond expected setup artifacts (workflow + +1352: // config), so that PRs bundling real code changes are never silently skipped. +1353: func isSetupPR(diff string, files []string) bool { +1354: // All files must be known setup paths. +1355: for _, f := range files { +1356: if !isSetupFile(f) { +1357: return false +1358: } +1359: } +1360: +1361: // At least one newly added workflow file must reference codecanary. +1362: lines := strings.Split(diff, "\n") +1363: for i := 0; i < len(lines)-1; i++ { +1364: if lines[i] != "--- /dev/null" { +1365: continue +1366: } +1367: plusLine := lines[i+1] +1368: if !strings.HasPrefix(plusLine, "+++ b/.github/workflows/") { +1369: continue +1370: } +1371: for j := i + 2; j < len(lines); j++ { +1372: if strings.HasPrefix(lines[j], "--- ") || strings.HasPrefix(lines[j], "diff --git") { +1373: break +1374: } +1375: if strings.HasPrefix(lines[j], "+") && (strings.Contains(lines[j], "codecanary") || strings.Contains(lines[j], "clanopy")) { +1376: return true +1377: } +1378: } +1379: } +1380: return false +1381: } +1382: +1383: // isSetupFile returns true if the file path is a known setup artifact. +1384: func isSetupFile(path string) bool { +1385: return strings.HasPrefix(path, ".github/workflows/") || +1386: strings.HasPrefix(path, ".codecanary/") || path == ".codecanary.yml" || +1387: strings.HasPrefix(path, ".clanopy/") +1388: } +1389: +1390: // matchesIgnore checks if a path matches any of the ignore glob patterns. +1391: // Uses doublestar to support ** recursive globs (e.g. "dist/**", "src/**/*.test.*"). +1392: func matchesIgnore(path string, patterns []string) bool { +1393: for _, pat := range patterns { +1394: if matched, _ := doublestar.Match(pat, path); matched { +1395: return true +1396: } +1397: // Also try matching against just the filename. +1398: if matched, _ := doublestar.Match(pat, filepath.Base(path)); matched { +1399: return true +1400: } +1401: } +1402: return false +1403: } +``` + +### `internal/review/platform_github.go` +``` +1: package review +2: +3: import ( +4: "fmt" +5: "os" +6: "strings" +7: ) +8: +9: // allResolved checks if all review threads have been resolved by code changes. +10: // Threads resolved by other reasons (dismissed, acknowledged, rebutted) are kept +11: // open and do not count as resolved. +12: func allResolved(threads []ReviewThread, fixed []fixedThread) bool { +13: fixedSet := make(map[int]bool, len(fixed)) +14: for _, f := range fixed { +15: if isTrueResolution(f.Reason) { +16: fixedSet[f.Index] = true +17: } +18: } +19: for i := range threads { +20: if !fixedSet[i] { +21: return false +22: } +23: } +24: return true +25: } +26: +27: // resolvedFindingIDs builds the set of finding IDs that are resolved. +28: // It combines threads already resolved on GitHub with threads just fixed by code changes. +29: // Threads resolved by other reasons (dismissed, acknowledged, rebutted) are not included +30: // since they are kept open for re-triage. +31: func resolvedFindingIDs(allThreads, unresolved []ReviewThread, fixed []fixedThread) map[string]bool { +32: resolved := make(map[string]bool) +33: for _, t := range allThreads { +34: if t.Resolved { +35: if id := FindingIDFromThread(t.Body); id != "" { +36: resolved[id] = true +37: } +38: } +39: } +40: fixedSet := make(map[int]bool, len(fixed)) +41: for _, f := range fixed { +42: if isTrueResolution(f.Reason) { +43: fixedSet[f.Index] = true +44: } +45: } +46: for i, t := range unresolved { +47: if fixedSet[i] { +48: if id := FindingIDFromThread(t.Body); id != "" { +49: resolved[id] = true +50: } +51: } +52: } +53: return resolved +54: } +55: +56: // minimizeFullyResolvedReviews minimizes reviews whose findings are all resolved. +57: func minimizeFullyResolvedReviews(repo string, prNumber int, resolvedIDs map[string]bool) { +58: reviews, err := FindReviews(repo, prNumber) +59: if err != nil { +60: fmt.Fprintf(os.Stderr, "Warning: could not fetch reviews for minimization: %v\n", err) +61: return +62: } +63: minimized := 0 +64: for _, rev := range reviews { +65: allResolved := len(rev.FindingIDs) > 0 +66: for _, fid := range rev.FindingIDs { +67: if !resolvedIDs[fid] { +68: allResolved = false +69: break +70: } +71: } +72: if !allResolved { +73: continue +74: } +75: if err := MinimizeComment(rev.NodeID); err != nil { +76: fmt.Fprintf(os.Stderr, "Warning: could not minimize review: %v\n", err) +77: } else { +78: minimized++ +79: } +80: } +81: if minimized > 0 { +82: fmt.Fprintf(os.Stderr, "Minimized %d previous review(s)\n", minimized) +83: } +84: } +85: +86: // acknowledgmentMessage returns a reply body for non-code-change resolutions. +87: // Each message includes a hidden HTML marker for dedup detection. +88: func acknowledgmentMessage(reason string) string { +89: marker := fmt.Sprintf("%s%s -->", ackMarkerPrefix, reason) +90: switch reason { +91: case "dismissed": +92: return marker + "\nAuthor dismissed this finding. Keeping open \u2014 will re-check if related code changes." +93: case "acknowledged": +94: return marker + "\nAuthor acknowledged this finding. Keeping open \u2014 will re-check on future pushes." +95: case "rebutted": +96: return marker + "\nAuthor provided a technical rebuttal. Keeping open \u2014 will re-check if related code changes." +97: default: +98: return fmt.Sprintf("%sunknown -->", ackMarkerPrefix) + "\nFinding acknowledged. Keeping open \u2014 will re-check on future pushes." +99: } +100: } +101: +102: // hasAcknowledgmentReply reports whether the thread already carries any +103: // codecanary ack reply. Dedup is intentionally reason-agnostic: all three +104: // ack reasons (dismissed/rebutted/acknowledged) keep the thread open and +105: // convey the same outcome to the author, and triage classification across +106: // runs is not deterministic — so checking only the same reason let a +107: // rebutted ack get followed by a dismissed ack on the next run, stacking +108: // two replies on one skip. One ack per thread is enough. +109: func hasAcknowledgmentReply(t ReviewThread) bool { +110: for _, r := range t.Replies { +111: if strings.Contains(r.Body, ackMarkerPrefix) || strings.Contains(r.Body, legacyAckPrefix) { +112: return true +113: } +114: } +115: return false +116: } +117: +118: // GithubPlatform implements ReviewPlatform for GitHub PR mode. It always +119: // posts to the PR; "local preview" against a GitHub PR is no longer a thing — +120: // `codecanary review` without --post uses LocalPlatform. +121: type GithubPlatform struct { +122: Repo string +123: PRNumber int +124: DryRun bool +125: } +126: +127: func (g *GithubPlatform) LoadPreviousFindings() ([]ReviewThread, string, int) { +128: allThreads, err := FetchReviewThreads(g.Repo, g.PRNumber) +129: if err != nil { +130: fmt.Fprintf(os.Stderr, "Warning: could not fetch review threads: %v\n", err) +131: return nil, "", 0 +132: } +133: startIndex := len(allThreads) +134: var unresolved []ReviewThread +135: for _, t := range allThreads { +136: if !t.Resolved { +137: unresolved = append(unresolved, t) +138: } +139: } +140: previousSHA := FetchPreviousReviewSHA(g.Repo, g.PRNumber) +141: return unresolved, previousSHA, startIndex +142: } +143: +144: func (g *GithubPlatform) ExcludedAuthor(threads []ReviewThread) string { +145: if len(threads) > 0 { +146: if login := threads[0].Author; login != "" { +147: return login +148: } +149: fmt.Fprintf(os.Stderr, "Warning: could not determine bot login from thread author\n") +150: } +151: return "" +152: } +153: +154: func (g *GithubPlatform) HandleResolutions(threads []ReviewThread, fixed []fixedThread) { +155: for _, f := range fixed { +156: if f.Index < 0 || f.Index >= len(threads) { +157: continue +158: } +159: t := threads[f.Index] +160: label := threadLabel(t) +161: if isTrueResolution(f.Reason) { +162: // Post an explanatory comment for file_removed before resolving. +163: if f.Reason == "file_removed" { +164: msg := "File removed from PR — resolving." +165: if err := ReplyToThread(t.ID, msg); err != nil { +166: fmt.Fprintf(os.Stderr, " ! %s — failed to post file-removed comment: %v\n", label, err) +167: } +168: } +169: if err := ResolveThread(t.ID); err != nil { +170: if strings.Contains(err.Error(), "Resource not accessible") { +171: fmt.Fprintf(os.Stderr, " ~ %s (auto-resolve unavailable: token lacks permission)\n", label) +172: } else { +173: fmt.Fprintf(os.Stderr, " ! %s — resolved, but failed to update thread: %v\n", label, err) +174: } +175: } +176: } else { +177: if !hasAcknowledgmentReply(t) { +178: msg := acknowledgmentMessage(f.Reason) +179: if err := ReplyToThread(t.ID, msg); err != nil { +180: fmt.Fprintf(os.Stderr, " ! %s — failed to post acknowledgment: %v\n", label, err) +181: } +182: } +183: } +184: } +185: } +186: +187: func (g *GithubPlatform) Publish(result *ReviewResult, pr *PRData, threads []ReviewThread, fixed []fixedThread) error { +188: summary := computeReviewSummary(threads, fixed, result.Findings) +189: +190: // Decide edit-vs-post: if the latest CodeCanary review on the PR carries +191: // the current HEAD SHA in its marker, this is either a reply-only run or +192: // a duplicate synchronize webhook — refresh that review in place rather +193: // than stacking another top-level comment. +194: latest, latestErr := FetchLatestCodecanaryReview(g.Repo, g.PRNumber) +195: if latestErr != nil { +196: fmt.Fprintf(os.Stderr, "Warning: could not fetch latest review for dedup: %v\n", latestErr) +197: } +198: if latest != nil && result.SHA != "" && latest.SHA == result.SHA { +199: updated := replaceSummaryBlock(latest.Body, summary) +200: if updated == latest.Body { +201: Stderrf(ansiGreen, "Latest review already current for %s — no update needed\n", shortSHA(result.SHA)) +202: return nil +203: } +204: if err := UpdateReviewBody(g.Repo, g.PRNumber, latest.ID, updated); err != nil { +205: return fmt.Errorf("updating review body: %w", err) +206: } +207: Stderrf(ansiGreen, "Updated latest review on PR #%d\n", g.PRNumber) +208: g.postReviewCommitStatus(result.SHA, summary) +209: return nil +210: } +211: +212: // Minimize previous reviews before posting a fresh one. +213: minimizeFailed := false +214: if len(threads) > 0 { +215: if nodeIDs, err := FindReviewNodeIDs(g.Repo, g.PRNumber); err == nil { +216: if allResolved(threads, fixed) { +217: for _, nodeID := range nodeIDs { +218: if err := MinimizeComment(nodeID); err != nil { +219: fmt.Fprintf(os.Stderr, "Warning: could not minimize review: %v\n", err) +220: minimizeFailed = true +221: } +222: } +223: if len(nodeIDs) > 0 && !minimizeFailed { +224: fmt.Fprintf(os.Stderr, "Minimized %d previous review(s)\n", len(nodeIDs)) +225: } +226: } else { +227: allThreads, err := FetchReviewThreads(g.Repo, g.PRNumber) +228: if err != nil { +229: fmt.Fprintf(os.Stderr, "Warning: could not fetch review threads for minimization: %v\n", err) +230: } else { +231: resolvedIDs := resolvedFindingIDs(allThreads, threads, fixed) +232: minimizeFullyResolvedReviews(g.Repo, g.PRNumber, resolvedIDs) +233: } +234: } +235: } else { +236: fmt.Fprintf(os.Stderr, "Warning: could not fetch reviews for minimization: %v\n", err) +237: minimizeFailed = true +238: } +239: } +240: +241: // POST path — pick the body shape that fits the cycle outcome. Every +242: // branch emits a top-level review so each push lands a visible status +243: // comment on the PR. +244: switch { +245: case len(result.Findings) > 0: +246: if err := PostReview(g.Repo, g.PRNumber, result, pr.ValidationDiff(), result.SHA, summary); err != nil { +247: return fmt.Errorf("posting review: %w", err) +248: } +249: Stderrf(ansiGreen, "Review posted to PR #%d\n", g.PRNumber) +250: case len(threads) > 0 && allResolved(threads, fixed): +251: if err := PostAllClearReview(g.Repo, g.PRNumber, result.SHA, minimizeFailed, summary); err != nil { +252: return fmt.Errorf("posting all-clear review: %w", err) +253: } +254: Stderrf(ansiGreen, "All clear! No issues remaining.\n") +255: case len(threads) > 0: +256: if err := PostActivityReview(g.Repo, g.PRNumber, result.SHA, summary); err != nil { +257: return fmt.Errorf("posting activity review: %w", err) +258: } +259: Stderrf(ansiGreen, "Posted activity summary to PR #%d\n", g.PRNumber) +260: default: +261: if err := PostCleanReview(g.Repo, g.PRNumber, result.SHA, summary); err != nil { +262: return fmt.Errorf("posting review: %w", err) +263: } +264: Stderrf(ansiGreen, "Review posted to PR #%d\n", g.PRNumber) +265: } +266: +267: g.postReviewCommitStatus(result.SHA, summary) +268: return nil +269: } +270: +271: func (g *GithubPlatform) SaveState(_ *ReviewResult, _ []Finding, _ bool) error { +272: // No-op: GitHub mode stores state in PR review threads (embedded JSON +273: // markers carry the SHA and findings). Local state files are owned by +274: // LocalPlatform — this keeps the two adapters from fighting over the +275: // same ~/.codecanary/state/.json file. +276: return nil +277: } +278: +279: func (g *GithubPlatform) GetIncrementalDiff(baseSHA string, _ []string) (string, error) { +280: return GetIncrementalDiff(baseSHA) +281: } +282: +283: func (g *GithubPlatform) ReportUsage(tracker *UsageTracker) { +284: report := tracker.Report(g.Repo, g.PRNumber) +285: if len(report.Calls) > 0 { +286: if err := WriteUsageEnv(report); err != nil { +287: fmt.Fprintf(os.Stderr, "Warning: could not write usage env: %v\n", err) +288: } +289: } +290: } +291: +292: // postReviewCommitStatus POSTs a `CodeCanary / review` commit status on the +293: // reviewed SHA. state=success when no unresolved findings remain for the +294: // PR (new findings this cycle + threads still open with no classification +295: // both at zero); state=failure otherwise. Teams can require this check in +296: // branch protection to gate merges on a clean review. +297: // +298: // Skipped silently when the SHA is empty (non-pr-loop contexts that +299: // accidentally share the adapter). Failures are logged as warnings — the +300: // review itself has already been published, so a flaky status post should +301: // not abort the whole Publish flow. +302: func (g *GithubPlatform) postReviewCommitStatus(sha string, summary ReviewSummary) { +303: if sha == "" { +304: return +305: } +306: state, desc := commitStatusFromSummary(summary) +307: if err := PostReviewCommitStatus(g.Repo, sha, state, desc); err != nil { +308: fmt.Fprintf(os.Stderr, "Warning: could not post %s commit status: %v\n", +309: ReviewCommitStatusContext, err) +310: return +311: } +312: Stderrf(ansiGreen, "Posted %s = %s on %s (%s)\n", +313: ReviewCommitStatusContext, state, shortSHA(sha), desc) +314: } +315: +316: // commitStatusFromSummary maps a ReviewSummary to the (state, description) +317: // pair sent to the commit status API. Pulled out so the mapping is +318: // unit-testable without network access. +319: // +320: // An "unresolved" count combines new findings this cycle with threads that +321: // were already open and remain unclassified — either kind should fail the +322: // required check. Everything classified by triage (resolved by code, file +323: // removed, dismissed, acknowledged, rebutted) counts as handled. +324: func commitStatusFromSummary(summary ReviewSummary) (state, desc string) { +325: unresolved := summary.NewFindings + summary.StillOpen +326: if unresolved > 0 { +327: suffix := "s" +328: if unresolved == 1 { +329: suffix = "" +330: } +331: return "failure", fmt.Sprintf("%d unresolved finding%s", unresolved, suffix) +332: } +333: if summary.ResolvedByCode+summary.FileRemoved+summary.Dismissed+summary.Acknowledged+summary.Rebutted > 0 { +334: return "success", "all findings resolved" +335: } +336: return "success", "no findings" +337: } +``` + +## Diff +````diff +diff --git a/README.md b/README.md +index b0d4ab1..d15aad7 100644 +--- a/README.md ++++ b/README.md +@@ -81,12 +81,12 @@ Once merged, CodeCanary reviews every PR on open and push. Draft PRs are skipped + + ### Gating merges on clean reviews + +-CodeCanary can block merges until a review comes back clean. After every review, the bot (and the local `codecanary signoff` command) posts a GitHub commit status under the context `codecanary/review`: ++CodeCanary can block merges until a review comes back clean. After every review, the bot (and the local `codecanary signoff` command) posts a GitHub commit status under the context `CodeCanary / review`: + + - `success` — no unresolved findings (everything is either unraised, fixed by code, or handled by the author) + - `failure` — one or more findings remain unresolved, with a description like `"3 unresolved findings"` + +-To turn this into a required check, add `codecanary/review` to your repo's required status checks via whichever branch protection mechanism you use (rulesets, classic branch protection rules, etc.). GitHub accepts any context name; if a review has already run, it will also show up in autocomplete. ++To turn this into a required check, add `CodeCanary / review` to your repo's required status checks via whichever branch protection mechanism you use (rulesets, classic branch protection rules, etc.). GitHub accepts any context name; if a review has already run, it will also show up in autocomplete. + + Statuses are keyed by commit SHA, so staling is automatic: a new push has no status until the next review run posts one, which re-blocks the merge button. + +@@ -95,7 +95,7 @@ Statuses are keyed by commit SHA, so staling is automatic: a new push has no sta + ```sh + codecanary review # reviews locally, stores findings in ~/.codecanary/... + # fix anything that came up, commit +-codecanary signoff # posts codecanary/review = success on HEAD ++codecanary signoff # posts CodeCanary / review = success on HEAD + ``` + + The command refuses to sign off unless HEAD matches the reviewed SHA and the tree is clean — this prevents attesting a review of code that isn't actually in the commit. It needs `gh` authenticated with `repo:status` scope (default `gh auth login` covers it). +@@ -109,7 +109,7 @@ Same required-check config works for both paths: the bot satisfies the check on + | `codecanary review [pr-number]` | Review a PR or local diff | + | `codecanary findings [pr-number]` | Fetch bot findings for a PR (markdown or JSON) | + | `codecanary reply --url --body ` | Post a reply on a review-comment thread (used by the skill when skipping) | +-| `codecanary signoff` | Post a `codecanary/review` commit status from the last local review (see [gating merges](#gating-merges-on-clean-reviews)) | ++| `codecanary signoff` | Post a `CodeCanary / review` commit status from the last local review (see [gating merges](#gating-merges-on-clean-reviews)) | + | `codecanary install-skill` | Install the `codecanary-fix` Claude Code skill | + | `codecanary setup [local\|github]` | Interactive setup wizard | + | `codecanary auth status` | Show stored credential info | +diff --git a/cmd/review/cli/signoff.go b/cmd/review/cli/signoff.go +index 82a1c57..3621db9 100644 +--- a/cmd/review/cli/signoff.go ++++ b/cmd/review/cli/signoff.go +@@ -25,7 +25,7 @@ Requires: + - the working tree is clean and HEAD matches the reviewed SHA + - 'gh' is installed and authenticated with repo:status scope + +-Combine with a required 'codecanary/review' check in branch protection to ++Combine with a required 'CodeCanary / review' check in branch protection to + block merges until a clean local review exists for the tip commit.`, + RunE: func(cmd *cobra.Command, args []string) error { + force, _ := cmd.Flags().GetBool("force") +diff --git a/docs/review-flow.md b/docs/review-flow.md +index 3f398d3..3973019 100644 +--- a/docs/review-flow.md ++++ b/docs/review-flow.md +@@ -165,7 +165,7 @@ Per-thread ack replies for dismissed/acknowledged/rebutted resolutions are poste + + `codecanary findings` applies the same marker to filter deferrals out of its default output: threads with any `codecanary:ack:*` reply are treated as handled and omitted alongside GitHub-resolved threads. Pass `--include-resolved` to see them. This keeps the codecanary-fix skill from re-prompting on findings the operator already deferred. + +-After the review is posted (or updated in place), `GithubPlatform.Publish` also POSTs a `codecanary/review` commit status on the reviewed SHA via `PostReviewCommitStatus`. State is `success` when `NewFindings + StillOpen == 0` (everything is either new-and-green, fixed by code, or explicitly handled by the author), and `failure` otherwise. Description is the unresolved count or "all findings resolved" / "no findings". Teams that add `codecanary/review` as a required status check in branch protection get auto-gating: merges are blocked until a review run posts a green status on HEAD. Status posting failures are logged as warnings and do not abort Publish — the review itself has already landed. The local `codecanary signoff` command posts a status under the same context, so a team can rely on a single required check that either the bot (pr-loop) or a local reviewer (local-loop) satisfies. ++After the review is posted (or updated in place), `GithubPlatform.Publish` also POSTs a `CodeCanary / review` commit status on the reviewed SHA via `PostReviewCommitStatus`. State is `success` when `NewFindings + StillOpen == 0` (everything is either new-and-green, fixed by code, or explicitly handled by the author), and `failure` otherwise. Description is the unresolved count or "all findings resolved" / "no findings". Teams that add `CodeCanary / review` as a required status check in branch protection get auto-gating: merges are blocked until a review run posts a green status on HEAD. Status posting failures are logged as warnings and do not abort Publish — the review itself has already landed. The local `codecanary signoff` command posts a status under the same context, so a team can rely on a single required check that either the bot (pr-loop) or a local reviewer (local-loop) satisfies. + + **Local**: Prints the formatted result to stdout. Format depends on context: terminal (colored, human-readable), markdown, or JSON. + +diff --git a/internal/review/github.go b/internal/review/github.go +index b626a82..80b0015 100644 +--- a/internal/review/github.go ++++ b/internal/review/github.go +@@ -314,11 +314,11 @@ func PostReview(repo string, prNumber int, result *ReviewResult, diff string, co + + // ReviewCommitStatusContext is the commit-status context the review bot + // writes when a review run completes. Using a stable string lets teams +-// require it in branch protection — "codecanary/review must be green ++// require it in branch protection — "CodeCanary / review must be green + // before merge." Matches the context used by `codecanary signoff`, so + // both the bot (on PRs with the workflow) and a local signoff can + // satisfy the same required check. +-const ReviewCommitStatusContext = "codecanary/review" ++const ReviewCommitStatusContext = "CodeCanary / review" + + // PostReviewCommitStatus POSTs a commit status on the given SHA summarising + // the review outcome. state is "success" when all findings are resolved or +diff --git a/internal/review/platform_github.go b/internal/review/platform_github.go +index 01843e2..37c4f7f 100644 +--- a/internal/review/platform_github.go ++++ b/internal/review/platform_github.go +@@ -289,7 +289,7 @@ func (g *GithubPlatform) ReportUsage(tracker *UsageTracker) { + } + } + +-// postReviewCommitStatus POSTs a `codecanary/review` commit status on the ++// postReviewCommitStatus POSTs a `CodeCanary / review` commit status on the + // reviewed SHA. state=success when no unresolved findings remain for the + // PR (new findings this cycle + threads still open with no classification + // both at zero); state=failure otherwise. Teams can require this check in +```` + +## Output Format +Return your findings as a JSON array inside a ```json code fence. Each finding must have these fields: + +- `id` (string): The rule ID that was violated, or a short kebab-case identifier for general findings. +- `file` (string): The file path where the issue was found. **Must be one of the exact paths listed in "Files in This Diff" above.** If a file path does not appear in that list, do NOT reference it. If your finding relates to a file not in the diff (e.g. a downstream consequence), set `file` and `line` to the diff location that triggers the issue and mention the affected file in `description`. +- `line` (int): The line number in the file. **Must be a line that was added or modified in the diff** (a `+` line in the diff hunk). If your finding is about a side effect on a distant line, set `line` to the diff line that *causes* the issue and describe the affected location in `description`. +- `severity` (string): One of "critical", "bug", "warning", "suggestion", or "nitpick". + - "critical": Security vulnerabilities, data loss, crashes. + - "bug": A logic error that causes incorrect runtime behavior for real inputs. Missing test coverage, unused parameters, typos in identifiers that happen to compile, or "what if a future caller…" concerns do NOT qualify — use "suggestion" or "nitpick" for those. If you cannot name the concrete input and the concrete wrong output, it is not a bug. + - "warning": Potential issues, performance problems, code smells. + - "suggestion": Better patterns, readability improvements. + - "nitpick": Minor style, naming, formatting. +- `title` (string): A short title for the finding. +- `description` (string): A concise explanation of the issue — 2-3 sentences max. State what is wrong and why it matters. Do not repeat the code or walk through the logic step by step. +- `suggestion` (string, optional): A concise suggested fix — 1-2 sentences of prose, then a code block if helpful. Do not explain what the code block does. For suggestions about broader patterns or improvements beyond the current PR scope, recommend opening a separate PR — do not imply they should fix it here. +- `fix_ref` (string): A reference ID in the format `173-` where index starts at 1 (e.g. `173-1`, `173-2`). +- `actionable` (boolean): Set to `false` if your analysis concludes the code is correct and no change is needed. Set to `true` if the finding requires the author to act. **Prefer returning an empty array over emitting findings with `actionable: false`.** + +**IMPORTANT — JSON escaping:** When your description or suggestion references code containing backslash sequences (e.g. `\n`, `\t`, `\"`), you MUST double-escape the backslash in the JSON string value. For example, to mention `fmt.Print("\n")` in a JSON string, write `fmt.Print("\\n")`. A single `\n` in JSON is a newline character, not the literal text `\n`. + +**Do not include findings where your conclusion is that the code is correct or no action is needed.** If you evaluate something and determine it is fine, omit it entirely rather than reporting it. Specifically: if you begin analyzing a potential issue but then realize the code handles it correctly, do NOT emit a finding that walks through the concern and then concludes "this is actually fine" or "no bug here" — simply drop it. Every finding you emit must represent a real, actionable problem. + +**Check against project documentation before emitting.** The "Project Documentation" section above defines conventions for this codebase (e.g. "don't add error handling for scenarios that can't happen", "keep the core engine agnostic"). Before emitting a finding, verify it does not contradict those conventions. If your suggested fix would violate a project-doc rule, drop the finding — the author has already made that tradeoff deliberately. + +**Label uncertainty from external behavior.** If your finding's validity depends on the behavior of a third-party API, webhook payload shape, framework internal, or other system you cannot verify from the diff, file contents, and project docs above, you MUST (a) cap severity at "suggestion" and (b) state the assumption in `description` (e.g. "Assumes `github.event.pull_request.number` is unset on `pull_request_review_comment` events — verify against GitHub's webhook docs before acting."). A finding that asserts external behavior as fact without this label is a false-positive risk. + +**CRITICAL: Do NOT invent or hallucinate file paths, function names, or code that does not appear in the diff or the provided file contents. If a file or function is not shown above, do not reference it.** + +If there are no findings, return an empty array: `[]`. + +Example: +```json +[ + { + "id": "rule-id", + "file": "src/main.go", + "line": 42, + "severity": "warning", + "title": "Short title", + "description": "The value is used after the error check, so a non-nil error silently proceeds with stale data.", + "suggestion": "Return early on error.\n\n```go\nif err != nil {\n return err\n}\n```", + "fix_ref": "173-1", + "actionable": true + } +] +``` diff --git a/internal/review/testdata/corpus/alansikora-codecanary-pr175.json b/internal/review/testdata/corpus/alansikora-codecanary-pr175.json new file mode 100644 index 0000000..60e7009 --- /dev/null +++ b/internal/review/testdata/corpus/alansikora-codecanary-pr175.json @@ -0,0 +1,29 @@ +{ + "format_version": 1, + "name": "alansikora-codecanary-pr175", + "repo": "alansikora/codecanary", + "pr_number": 175, + "head_sha": "921c880eb435710b24113c883b1133c2299c2321", + "captured_at": "2026-09-04T23:31:50Z", + "pr": { + "number": 175, + "title": "fix(triage): include PreviouslyAcked in countNonSkipped", + "body": "Follow-up to #174.\n\n## Summary\n\nThe first sticky-ack fix worked at the triage layer (logs showed `[sticky] ... carried forward`) but the published summary still said **Still unresolved: 2** — confirmed in production on thetechfx/bedrock#1671 right after v0.6.23 shipped.\n\nRoot cause: `countNonSkipped` excluded `TriagePreviouslyAcked`. When *every* remaining thread is sticky-ack, `needsEval` came back as zero, the runner skipped `EvaluateThreadsParallel` entirely, the fast path I added there never fired, no `fixedThread` was emitted, and `computeReviewSummary` lumped the threads into `StillOpen`.\n\nIncluding PreviouslyAcked in `countNonSkipped` restores the round trip: eval runs, the fast path returns the prior reason without an LLM call, `toFixedThreads` collects the resolutions, and the summary lands them in Acknowledged/Rebutted/Dismissed.\n\n## Test plan\n\n- [x] `go test ./...`\n- [x] New regression test `TestCountNonSkipped_IncludesPreviouslyAcked`\n- [ ] Re-trigger a review on bedrock#1671 with v0.6.24 to confirm the summary flips to \"Acknowledged by author: 2\"", + "author": "alansikora", + "base_branch": "main", + "head_branch": "worktree-structured-spinning-wreath", + "diff": "diff --git a/internal/review/triage.go b/internal/review/triage.go\nindex dc20e2a..485db27 100644\n--- a/internal/review/triage.go\n+++ b/internal/review/triage.go\n@@ -912,12 +912,17 @@ func LogResolutions(triaged []TriagedThread, resolutions []ThreadResolution) {\n \t}\n }\n \n-// countNonSkipped returns the number of triaged threads that need LLM evaluation.\n+// countNonSkipped returns the number of triaged threads that EvaluateThreadsParallel\n+// must process. TriagePreviouslyAcked is included even though it short-circuits\n+// without an LLM call: the function still has to run to emit the carried-forward\n+// fixedThread, which is what feeds the summary's Acknowledged/Rebutted/Dismissed\n+// buckets. Excluding it here was the bug behind the \"Still unresolved: 2\" stuck\n+// status on already-acked threads.\n func countNonSkipped(triaged []TriagedThread) int {\n \tn := 0\n \tfor _, t := range triaged {\n \t\tswitch t.Class {\n-\t\tcase TriageSkip, TriageFileRemovedFromPR, TriagePreviouslyAcked:\n+\t\tcase TriageSkip, TriageFileRemovedFromPR:\n \t\t\tcontinue\n \t\tdefault:\n \t\t\tn++\ndiff --git a/internal/review/triage_test.go b/internal/review/triage_test.go\nindex be37fcb..79a5950 100644\n--- a/internal/review/triage_test.go\n+++ b/internal/review/triage_test.go\n@@ -540,6 +540,23 @@ func TestClassifyThreads_NewHumanReplyBreaksStickiness(t *testing.T) {\n \t}\n }\n \n+func TestCountNonSkipped_IncludesPreviouslyAcked(t *testing.T) {\n+\t// Regression: sticky-ack threads must count toward needsEval so\n+\t// EvaluateThreadsParallel runs and emits the carried-forward fixedThread.\n+\t// Excluding them caused the runner to skip eval entirely, leaving the\n+\t// summary stuck on \"Still unresolved\" for threads the bot had already\n+\t// acked.\n+\ttriaged := []TriagedThread{\n+\t\t{Class: TriageSkip},\n+\t\t{Class: TriageFileRemovedFromPR},\n+\t\t{Class: TriagePreviouslyAcked, PriorAckReason: \"acknowledged\"},\n+\t\t{Class: TriagePreviouslyAcked, PriorAckReason: \"rebutted\"},\n+\t}\n+\tif got := countNonSkipped(triaged); got != 2 {\n+\t\tt.Errorf(\"countNonSkipped = %d, want 2 (the two PreviouslyAcked threads)\", got)\n+\t}\n+}\n+\n func TestEvaluateThreadsParallel_StickyAckShortCircuits(t *testing.T) {\n \t// Provider that panics if called — sticky-ack must skip the LLM.\n \tprovider := \u0026stickyAckPanicProvider{t: t}\n", + "files": [ + "internal/review/triage.go", + "internal/review/triage_test.go" + ], + "file_contents": { + "internal/review/triage.go": "package review\n\nimport (\n\t\"context\"\n\t\"encoding/json\"\n\t\"fmt\"\n\t\"os\"\n\t\"sort\"\n\t\"strconv\"\n\t\"strings\"\n\t\"sync\"\n)\n\n// ThreadClassification is the triage result for a single unresolved thread.\ntype ThreadClassification int\n\nconst (\n\tTriageSkip ThreadClassification = iota // no code changes at all\n\tTriageCodeChanged // diff touches finding location (outdated)\n\tTriageHasReply // thread has human replies\n\tTriageCodeChangedReply // both code changed AND has replies\n\tTriageCrossFileChange // diff has changes but NOT in this thread's file\n\tTriageFileRemovedFromPR // file no longer in the PR\n\tTriagePreviouslyAcked // bot already ack'd a deferral; no new human reply since\n)\n\n// TriagedThread pairs a ReviewThread with its classification and context.\ntype TriagedThread struct {\n\tThread ReviewThread\n\tIndex int // original index in the unresolved slice\n\tClass ThreadClassification\n\tFileDiff string // file-scoped diff (level 1) or full diff (cross-file)\n\tFullDiff string // full PR diff for widened-scope fallback; empty when not applicable\n\tFileSnippet string // windowed file content around finding + diff hunks\n\tBotLogin string // login of the review bot, for filtering replies\n\tPriorAckReason string // for TriagePreviouslyAcked: the previously-recorded reason (acknowledged/rebutted/dismissed)\n}\n\n// ThreadResolution is the result of a per-thread Claude evaluation.\ntype ThreadResolution struct {\n\tIndex int\n\tResolved bool\n\tReason string // \"code_change\", \"acknowledged\", \"rebutted\", \"dismissed\"\n\tRationale string // one-sentence narrative from the evaluator, surfaced in the incremental review to prevent ping-ponging\n\tError error\n}\n\n// ExtractFileDiff extracts all diff hunks for a specific file from a unified diff.\nfunc ExtractFileDiff(fullDiff, filePath string) string {\n\tlines := strings.Split(fullDiff, \"\\n\")\n\tvar result []string\n\tcapturing := false\n\n\tfor i := 0; i \u003c len(lines); i++ {\n\t\t// Detect start of a new file in the diff.\n\t\tif strings.HasPrefix(lines[i], \"diff --git\") {\n\t\t\tif capturing {\n\t\t\t\t// We were capturing — new file starts, stop.\n\t\t\t\tbreak\n\t\t\t}\n\t\t\t// Search ahead for the \"+++ b/\u003cpath\u003e\" header line.\n\t\t\t// The offset varies (index line, mode lines, etc.) so scan\n\t\t\t// forward instead of using a fixed offset.\n\t\t\tfor j := i + 1; j \u003c len(lines) \u0026\u0026 !strings.HasPrefix(lines[j], \"diff --git\"); j++ {\n\t\t\t\tif strings.HasPrefix(lines[j], \"+++ b/\"+filePath) {\n\t\t\t\t\tcapturing = true\n\t\t\t\t\tresult = append(result, lines[i])\n\t\t\t\t\tbreak\n\t\t\t\t}\n\t\t\t}\n\t\t\tcontinue\n\t\t}\n\t\tif capturing {\n\t\t\tresult = append(result, lines[i])\n\t\t}\n\t}\n\n\treturn strings.Join(result, \"\\n\")\n}\n\n// lineRange represents an inclusive range of 1-based line numbers.\ntype lineRange struct{ start, end int }\n\n// parseHunkNewRanges extracts the new-file line ranges from unified diff hunk headers.\n// Each @@ -X,Y +N,M @@ header yields a range [N, N+M-1].\nfunc parseHunkNewRanges(diffText string) []lineRange {\n\tvar ranges []lineRange\n\tfor _, line := range strings.Split(diffText, \"\\n\") {\n\t\tif !strings.HasPrefix(line, \"@@ \") {\n\t\t\tcontinue\n\t\t}\n\t\t// Find +N or +N,M in the hunk header.\n\t\tidx := strings.Index(line, \"+\")\n\t\tif idx \u003c 0 {\n\t\t\tcontinue\n\t\t}\n\t\trest := line[idx+1:]\n\t\t// Trim everything after the space/comma/@@ that ends the range spec.\n\t\tif sp := strings.IndexAny(rest, \" @\"); sp \u003e= 0 {\n\t\t\trest = rest[:sp]\n\t\t}\n\t\tparts := strings.SplitN(rest, \",\", 2)\n\t\tstart, err := strconv.Atoi(parts[0])\n\t\tif err != nil {\n\t\t\tcontinue\n\t\t}\n\t\tcount := 1\n\t\tif len(parts) == 2 {\n\t\t\tif c, err := strconv.Atoi(parts[1]); err == nil {\n\t\t\t\tif c == 0 {\n\t\t\t\t\tcontinue // pure deletion hunk — no new lines\n\t\t\t\t}\n\t\t\t\tif c \u003e 0 {\n\t\t\t\t\tcount = c\n\t\t\t\t}\n\t\t\t}\n\t\t}\n\t\tranges = append(ranges, lineRange{start: start, end: start + count - 1})\n\t}\n\treturn ranges\n}\n\n// mergeRanges merges overlapping or adjacent line ranges, sorted by start.\nfunc mergeRanges(ranges []lineRange) []lineRange {\n\tif len(ranges) == 0 {\n\t\treturn nil\n\t}\n\tsort.Slice(ranges, func(i, j int) bool { return ranges[i].start \u003c ranges[j].start })\n\tmerged := []lineRange{ranges[0]}\n\tfor _, r := range ranges[1:] {\n\t\tlast := \u0026merged[len(merged)-1]\n\t\tif r.start \u003c= last.end+1 {\n\t\t\tif r.end \u003e last.end {\n\t\t\t\tlast.end = r.end\n\t\t\t}\n\t\t} else {\n\t\t\tmerged = append(merged, r)\n\t\t}\n\t}\n\treturn merged\n}\n\n// ExtractFileSnippet extracts a windowed snippet from file content centered around\n// the finding line and expanded to cover diff hunk ranges. Returns an empty string\n// if content is empty. findingLine is 1-based. diffText is the file-scoped diff\n// (used to parse hunk ranges; may be empty for cross-file cases). maxLines caps the\n// total snippet length.\nfunc ExtractFileSnippet(content string, findingLine int, diffText string, maxLines int) string {\n\tif content == \"\" {\n\t\treturn \"\"\n\t}\n\tlines := strings.Split(content, \"\\n\")\n\ttotalLines := len(lines)\n\n\t// Build interesting ranges: finding line ± 50, each hunk range ± 30.\n\tconst findingPad = 50\n\tconst hunkPad = 30\n\n\tvar ranges []lineRange\n\tif findingLine \u003e 0 {\n\t\tranges = append(ranges, lineRange{\n\t\t\tstart: max(1, findingLine-findingPad),\n\t\t\tend: min(totalLines, findingLine+findingPad),\n\t\t})\n\t}\n\tfor _, hr := range parseHunkNewRanges(diffText) {\n\t\tranges = append(ranges, lineRange{\n\t\t\tstart: max(1, hr.start-hunkPad),\n\t\t\tend: min(totalLines, hr.end+hunkPad),\n\t\t})\n\t}\n\tif len(ranges) == 0 {\n\t\t// Fallback: center on line 1 if nothing else.\n\t\tranges = append(ranges, lineRange{start: 1, end: min(totalLines, maxLines)})\n\t}\n\n\tmerged := mergeRanges(ranges)\n\n\t// If total exceeds maxLines, prioritize the range containing the finding line,\n\t// then include other ranges in order until the budget is exhausted.\n\ttotal := 0\n\tfor _, r := range merged {\n\t\ttotal += r.end - r.start + 1\n\t}\n\tif total \u003e maxLines {\n\t\t// Find which merged range contains the finding line.\n\t\t// When findingLine is invalid (\u003c= 0), default to the first range.\n\t\tfindingIdx := 0\n\t\tif findingLine \u003e 0 {\n\t\t\tfor i, r := range merged {\n\t\t\t\tif findingLine \u003e= r.start \u0026\u0026 findingLine \u003c= r.end {\n\t\t\t\t\tfindingIdx = i\n\t\t\t\t\tbreak\n\t\t\t\t}\n\t\t\t}\n\t\t}\n\t\t// Start with the finding range, then add others.\n\t\tbudget := maxLines\n\t\tkept := make([]bool, len(merged))\n\t\tkept[findingIdx] = true\n\t\tsize := merged[findingIdx].end - merged[findingIdx].start + 1\n\t\tif size \u003e budget \u0026\u0026 findingLine \u003e 0 {\n\t\t\t// Truncate the finding range around the finding line.\n\t\t\thalf := budget / 2\n\t\t\tmerged[findingIdx] = lineRange{\n\t\t\t\tstart: max(1, findingLine-half),\n\t\t\t\tend: min(totalLines, findingLine+half),\n\t\t\t}\n\t\t\tsize = merged[findingIdx].end - merged[findingIdx].start + 1\n\t\t} else if size \u003e budget {\n\t\t\t// No valid finding line — just take the first maxLines of the range.\n\t\t\tmerged[findingIdx] = lineRange{\n\t\t\t\tstart: merged[findingIdx].start,\n\t\t\t\tend: min(totalLines, merged[findingIdx].start+budget-1),\n\t\t\t}\n\t\t\tsize = merged[findingIdx].end - merged[findingIdx].start + 1\n\t\t}\n\t\tbudget -= size\n\t\tfor i, r := range merged {\n\t\t\tif kept[i] {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t\trSize := r.end - r.start + 1\n\t\t\tif rSize \u003c= budget {\n\t\t\t\tkept[i] = true\n\t\t\t\tbudget -= rSize\n\t\t\t}\n\t\t}\n\t\tvar trimmed []lineRange\n\t\tfor i, r := range merged {\n\t\t\tif kept[i] {\n\t\t\t\ttrimmed = append(trimmed, r)\n\t\t\t}\n\t\t}\n\t\tmerged = trimmed\n\t}\n\n\t// Format with line numbers, inserting omission markers between gaps.\n\tvar b strings.Builder\n\tfor i, r := range merged {\n\t\tif i \u003e 0 {\n\t\t\tgap := r.start - merged[i-1].end - 1\n\t\t\tfmt.Fprintf(\u0026b, \"... (%d lines omitted) ...\\n\", gap)\n\t\t}\n\t\tfor ln := r.start; ln \u003c= r.end \u0026\u0026 ln \u003c= totalLines; ln++ {\n\t\t\tfmt.Fprintf(\u0026b, \"%d: %s\\n\", ln, lines[ln-1])\n\t\t}\n\t}\n\treturn b.String()\n}\n\n// ClassifyThreads triages unresolved threads using GitHub's outdated flag and reply presence.\n//\n// Two diffs serve different purposes:\n// - activityDiff: the incremental diff (changes since last review). Used to decide whether\n// to skip evaluation — when empty, non-outdated threads with no replies are TriageSkip.\n// - contextDiff: the full PR diff (all changes). Used to determine the correct classification\n// (TriageCodeChanged vs TriageCrossFileChange) and to extract FileDiff/FileSnippet for the\n// evaluation prompt. This ensures fixes from earlier pushes are visible to the evaluator.\n//\n// prFiles is the current set of files in the PR; threads on files no longer in the PR\n// are classified as TriageFileRemovedFromPR and auto-resolved without an LLM call.\n// fileContents provides current file contents for building context snippets in triage prompts.\nfunc ClassifyThreads(threads []ReviewThread, activityDiff, contextDiff, botLogin string, prFiles []string, fileContents map[string]string) []TriagedThread {\n\tprFileSet := make(map[string]bool, len(prFiles))\n\tfor _, f := range prFiles {\n\t\tprFileSet[f] = true\n\t}\n\n\tresult := make([]TriagedThread, len(threads))\n\n\tfor i, t := range threads {\n\t\t// File no longer in the PR — auto-resolve without LLM.\n\t\t// When prFiles is empty (e.g. upstream fetch returned no file list),\n\t\t// we skip this check to avoid incorrectly resolving all threads.\n\t\tif len(prFileSet) \u003e 0 \u0026\u0026 !prFileSet[t.Path] {\n\t\t\tresult[i] = TriagedThread{\n\t\t\t\tThread: t,\n\t\t\t\tIndex: i,\n\t\t\t\tClass: TriageFileRemovedFromPR,\n\t\t\t}\n\t\t\tcontinue\n\t\t}\n\n\t\thasReply := hasNewHumanReply(t, botLogin)\n\t\toutdated := t.Outdated\n\t\tdeleted := fileDeletedInDiff(contextDiff, t.Path)\n\n\t\t// Sticky-ack: when the bot already recorded a deferral and the\n\t\t// author has not added a fresh reply since, preserve the prior\n\t\t// classification instead of rerunning triage. This prevents the\n\t\t// \"Acknowledged by author: 2\" → \"Still unresolved: 2\" regression\n\t\t// that happened on subsequent pushes when no new signal was present.\n\t\tvar priorAck string\n\t\tif !hasReply {\n\t\t\tpriorAck = parsePriorAckReason(t, botLogin)\n\t\t}\n\n\t\tvar class ThreadClassification\n\t\tswitch {\n\t\tcase priorAck != \"\":\n\t\t\tclass = TriagePreviouslyAcked\n\t\tcase deleted:\n\t\t\t// File was deleted — evaluate with full diff so Claude can check\n\t\t\t// whether the code moved to a replacement file with the fix applied.\n\t\t\tclass = TriageCrossFileChange\n\t\tcase outdated \u0026\u0026 hasReply:\n\t\t\tclass = TriageCodeChangedReply\n\t\tcase outdated:\n\t\t\tclass = TriageCodeChanged\n\t\tcase hasReply:\n\t\t\tclass = TriageHasReply\n\t\tdefault:\n\t\t\t// No GitHub outdated flag, no replies. Use the incremental diff to\n\t\t\t// decide whether there is new activity worth evaluating.\n\t\t\tif activityDiff == \"\" {\n\t\t\t\t// No changes since last review — nothing new to evaluate.\n\t\t\t\tclass = TriageSkip\n\t\t\t} else if fileInDiff(contextDiff, t.Path) {\n\t\t\t\t// File was changed in the PR — classify as code-changed so the\n\t\t\t\t// evaluator sees the file-scoped diff (which may include fixes\n\t\t\t\t// from earlier pushes that the incremental diff missed).\n\t\t\t\tclass = TriageCodeChanged\n\t\t\t} else {\n\t\t\t\t// Code changed in other files — evaluate in case the fix is cross-file.\n\t\t\t\tclass = TriageCrossFileChange\n\t\t\t}\n\t\t}\n\n\t\t// Use file-scoped diff for same-file evaluations to reduce noise.\n\t\t// The full PR diff can drown out the relevant fix with changes from\n\t\t// unrelated files, causing false \"not resolved\" verdicts. Cross-file\n\t\t// evaluations keep the full diff since the fix is in a different file.\n\t\tvar fileDiff, fullDiff string\n\t\tswitch class {\n\t\tcase TriageCodeChanged, TriageCodeChangedReply:\n\t\t\tfileDiff = ExtractFileDiff(contextDiff, t.Path)\n\t\t\tfullDiff = contextDiff // retained for widened-scope fallback\n\t\tcase TriageCrossFileChange:\n\t\t\tfileDiff = contextDiff\n\t\t}\n\n\t\t// Build a windowed file snippet for code-change evaluations.\n\t\t// Reuse fileDiff (already file-scoped) for snippet range calculation\n\t\t// so hunk line numbers from other files don't pull in wrong sections.\n\t\tvar fileSnippet string\n\t\tif content, ok := fileContents[t.Path]; ok {\n\t\t\tswitch class {\n\t\t\tcase TriageCodeChanged, TriageCodeChangedReply:\n\t\t\t\tfileSnippet = ExtractFileSnippet(content, t.Line, fileDiff, 300)\n\t\t\tcase TriageCrossFileChange:\n\t\t\t\t// Show finding's file context even though the diff is in other files.\n\t\t\t\tfileSnippet = ExtractFileSnippet(content, t.Line, \"\", 200)\n\t\t\t}\n\t\t}\n\n\t\tresult[i] = TriagedThread{\n\t\t\tThread: t,\n\t\t\tIndex: i,\n\t\t\tClass: class,\n\t\t\tFileDiff: fileDiff,\n\t\t\tFullDiff: fullDiff,\n\t\t\tFileSnippet: fileSnippet,\n\t\t\tBotLogin: botLogin,\n\t\t\tPriorAckReason: priorAck,\n\t\t}\n\t}\n\n\treturn result\n}\n\n// parsePriorAckReason returns the reason recorded in the most recent\n// codecanary ack reply on the thread (acknowledged/rebutted/dismissed),\n// or \"\" if no ack reply is present. Handles both the current marker\n// (\u003c!-- codecanary:ack:\u003creason\u003e --\u003e) and the legacy clanopy form.\nfunc parsePriorAckReason(t ReviewThread, botLogin string) string {\n\tvar latest string\n\tfor _, r := range t.Replies {\n\t\tif r.Author != botLogin {\n\t\t\tcontinue\n\t\t}\n\t\tif reason := extractAckReason(r.Body); reason != \"\" {\n\t\t\tlatest = reason\n\t\t}\n\t}\n\treturn latest\n}\n\n// extractAckReason pulls the reason value out of a codecanary ack marker.\n// Returns \"\" when the body has no marker. Unknown reasons are returned\n// as-is so callers can decide how to handle them; in practice the bot\n// only writes \"acknowledged\", \"rebutted\", \"dismissed\", or \"unknown\".\nfunc extractAckReason(body string) string {\n\tfor _, prefix := range []string{ackMarkerPrefix, legacyAckPrefix} {\n\t\tidx := strings.Index(body, prefix)\n\t\tif idx \u003c 0 {\n\t\t\tcontinue\n\t\t}\n\t\trest := body[idx+len(prefix):]\n\t\tend := strings.Index(rest, \" --\u003e\")\n\t\tif end \u003c 0 {\n\t\t\tend = strings.Index(rest, \"--\u003e\")\n\t\t\tif end \u003c 0 {\n\t\t\t\tcontinue\n\t\t\t}\n\t\t}\n\t\treturn strings.TrimSpace(rest[:end])\n\t}\n\treturn \"\"\n}\n\n// fileInDiff checks if the diff contains changes to the given file path.\nfunc fileInDiff(diff, path string) bool {\n\ttarget := \"+++ b/\" + path\n\treturn strings.Contains(diff, target+\"\\n\") || strings.Contains(diff, target+\"\\t\") || strings.HasSuffix(diff, target)\n}\n\n// fileDeletedInDiff checks if the diff shows the given file was deleted.\n// A deleted file has \"--- a/\u003cpath\u003e\" followed by \"+++ /dev/null\".\nfunc fileDeletedInDiff(diff, path string) bool {\n\tmarker := \"--- a/\" + path\n\tidx := strings.Index(diff, marker)\n\tif idx \u003c 0 {\n\t\treturn false\n\t}\n\t// Ensure full path match (not a prefix of a longer filename).\n\trest := diff[idx+len(marker):]\n\tif len(rest) \u003e 0 \u0026\u0026 rest[0] != '\\n' \u0026\u0026 rest[0] != '\\r' {\n\t\treturn false\n\t}\n\tnl := strings.Index(rest, \"\\n\")\n\tif nl \u003c 0 {\n\t\treturn false\n\t}\n\tnextLine := \"\"\n\trest = rest[nl+1:]\n\tif eol := strings.Index(rest, \"\\n\"); eol \u003e= 0 {\n\t\tnextLine = rest[:eol]\n\t} else {\n\t\tnextLine = rest\n\t}\n\treturn nextLine == \"+++ /dev/null\"\n}\n\n// hasHumanReply checks if a thread has at least one reply from a non-bot author.\nfunc hasHumanReply(t ReviewThread, botLogin string) bool {\n\tfor _, r := range t.Replies {\n\t\tif r.Author != botLogin {\n\t\t\treturn true\n\t\t}\n\t}\n\treturn false\n}\n\n// isAckReply checks if a reply body contains an acknowledgment marker.\nfunc isAckReply(body string) bool {\n\treturn strings.Contains(body, ackMarkerPrefix) || strings.Contains(body, legacyAckPrefix)\n}\n\n// hasNewHumanReply checks if a thread has a human reply AFTER the last\n// ack reply. If no ack reply exists, it falls back to hasHumanReply behavior.\n// Replies are in chronological order.\nfunc hasNewHumanReply(t ReviewThread, botLogin string) bool {\n\tlastAckIdx := -1\n\tfor i, r := range t.Replies {\n\t\tif r.Author == botLogin \u0026\u0026 isAckReply(r.Body) {\n\t\t\tlastAckIdx = i\n\t\t}\n\t}\n\tif lastAckIdx == -1 {\n\t\t// No ack reply exists — fall back to standard check.\n\t\treturn hasHumanReply(t, botLogin)\n\t}\n\t// Check for human replies after the last ack.\n\tfor _, r := range t.Replies[lastAckIdx+1:] {\n\t\tif r.Author != botLogin {\n\t\t\treturn true\n\t\t}\n\t}\n\treturn false\n}\n\n// BuildPerThreadPrompt dispatches to the appropriate prompt builder based on classification.\nfunc BuildPerThreadPrompt(t TriagedThread, cfg *ReviewConfig) string {\n\tswitch t.Class {\n\tcase TriageCodeChanged:\n\t\treturn buildCodeChangePrompt(t, cfg)\n\tcase TriageHasReply:\n\t\treturn buildReplyPrompt(t, cfg)\n\tcase TriageCodeChangedReply:\n\t\treturn buildCodeChangeReplyPrompt(t, cfg)\n\tcase TriageCrossFileChange:\n\t\treturn buildCrossFilePrompt(t, cfg)\n\tdefault:\n\t\treturn \"\" // TriageSkip — should not be called\n\t}\n}\n\nfunc buildCodeChangePrompt(t TriagedThread, cfg *ReviewConfig) string {\n\tvar b strings.Builder\n\n\tb.WriteString(\"You are a code reviewer. You previously raised a finding on a pull request. The author pushed new code.\\n\\n\")\n\n\twriteFinding(\u0026b, t.Thread)\n\n\twriteFileSnippet(\u0026b, t.FileSnippet)\n\n\tb.WriteString(\"## Code Changes\\n```diff\\n\")\n\tb.WriteString(t.FileDiff)\n\tb.WriteString(\"\\n```\\n\\n\")\n\n\tif ctx := evalContext(cfg, \"code_change\"); ctx != \"\" {\n\t\tfmt.Fprintf(\u0026b, \"## Additional Context\\n%s\\n\\n\", ctx)\n\t}\n\n\tb.WriteString(\"## Task\\n\")\n\tb.WriteString(\"Determine whether the issue you raised has been resolved.\\n\\n\")\n\tb.WriteString(\"**Start with the current file content** (if provided). Read the code around the finding location and determine whether the issue still exists. If the problematic code has been fixed, removed, or restructured so the finding no longer applies, the answer is YES — regardless of which specific diff line produced the fix.\\n\\n\")\n\tb.WriteString(\"If file content is not available, examine the diff for evidence that the issue was addressed.\\n\")\n\tb.WriteString(\"- A change that fixes the root cause, removes the problematic code, or meaningfully changes the code so the finding no longer applies → YES.\\n\")\n\tb.WriteString(\"- A change to nearby or adjacent code counts IF it effectively resolves the concern (e.g. fixing the logic, adding the missing check, refactoring the problematic pattern).\\n\")\n\tb.WriteString(\"- A structural change also counts — for example, if code was moved before a guard condition, control flow was reordered, or the code was refactored so the finding no longer applies.\\n\")\n\tb.WriteString(\"- Answer NO only if the concern is still present in the current code (when file content is provided) or if the diff does not address the finding.\\n\\n\")\n\twriteCodeChangeResolutionFormat(\u0026b)\n\n\treturn b.String()\n}\n\nfunc buildReplyPrompt(t TriagedThread, cfg *ReviewConfig) string {\n\tvar b strings.Builder\n\n\tb.WriteString(\"You are a code reviewer. You previously raised a finding on a pull request. The author replied.\\n\\n\")\n\n\twriteFinding(\u0026b, t.Thread)\n\twriteReplies(\u0026b, t.Thread, t.BotLogin)\n\n\tif ctx := evalContext(cfg, \"reply\"); ctx != \"\" {\n\t\tfmt.Fprintf(\u0026b, \"## Additional Context\\n%s\\n\\n\", ctx)\n\t}\n\n\tb.WriteString(\"## Task\\n\")\n\tb.WriteString(\"Does the author's reply resolve the finding?\\n\")\n\tb.WriteString(\"- **Dismissed**: Reply explicitly asks the reviewer to dismiss, ignore, or skip the finding (e.g. \\\"dismiss this\\\", \\\"you can safely dismiss\\\", \\\"please ignore\\\", \\\"skip this one\\\"). The author is exercising their authority to close the thread without further justification.\\n\")\n\tb.WriteString(\"- **Acknowledged**: Reply indicates the finding is intentional, accepted, or tracked elsewhere (e.g. \\\"intentional\\\", \\\"will fix in a future PR\\\", \\\"tracked in issue #N\\\").\\n\")\n\tb.WriteString(\"- **Rebutted**: Reply provides concrete technical reasoning showing the finding is not applicable. Vague disagreement (\\\"I don't think so\\\") does NOT qualify — the reply must cite specific technical details, framework behavior, or project constraints.\\n\")\n\tb.WriteString(\"- **Not resolved**: Reply is a question, vague disagreement, or does not address the finding.\\n\\n\")\n\twriteResolutionFormat(\u0026b)\n\n\treturn b.String()\n}\n\nfunc buildCodeChangeReplyPrompt(t TriagedThread, cfg *ReviewConfig) string {\n\tvar b strings.Builder\n\n\tb.WriteString(\"You are a code reviewer. You previously raised a finding on a pull request. The author pushed new code AND replied.\\n\\n\")\n\n\twriteFinding(\u0026b, t.Thread)\n\twriteReplies(\u0026b, t.Thread, t.BotLogin)\n\n\twriteFileSnippet(\u0026b, t.FileSnippet)\n\n\tb.WriteString(\"## Code Changes\\n```diff\\n\")\n\tb.WriteString(t.FileDiff)\n\tb.WriteString(\"\\n```\\n\\n\")\n\n\tif ctx := evalContext(cfg, \"code_change\"); ctx != \"\" {\n\t\tfmt.Fprintf(\u0026b, \"## Additional Context (Code Changes)\\n%s\\n\\n\", ctx)\n\t}\n\tif ctx := evalContext(cfg, \"reply\"); ctx != \"\" {\n\t\tfmt.Fprintf(\u0026b, \"## Additional Context (Replies)\\n%s\\n\\n\", ctx)\n\t}\n\n\tb.WriteString(\"## Task\\n\")\n\tb.WriteString(\"Is the finding resolved? It may be resolved by the code change, the reply, or both.\\n\\n\")\n\tb.WriteString(\"**Start with the current file content** (if provided). If the issue no longer exists in the current code — the root cause is fixed, the code was removed, or the code was restructured — it is resolved by code change regardless of which diff line produced the fix.\\n\\n\")\n\tb.WriteString(\"If the code still shows the issue, evaluate the reply:\\n\")\n\tb.WriteString(\"- **Dismissed**: Reply explicitly asks the reviewer to dismiss, ignore, or skip the finding (e.g. \\\"dismiss this\\\", \\\"you can safely dismiss\\\", \\\"please ignore\\\", \\\"skip this one\\\"). The author is exercising their authority to close the thread without further justification.\\n\")\n\tb.WriteString(\"- **Acknowledged**: Reply indicates the finding is intentional, accepted, or tracked elsewhere (e.g. \\\"intentional\\\", \\\"will fix in a future PR\\\", \\\"tracked in issue #N\\\").\\n\")\n\tb.WriteString(\"- **Rebutted**: Reply provides concrete technical reasoning showing the finding is not applicable. Vague disagreement (\\\"I don't think so\\\") does NOT qualify — the reply must cite specific technical details, framework behavior, or project constraints.\\n\")\n\tb.WriteString(\"- **Not resolved**: The issue is still in the code, and the reply is a question, vague disagreement, or does not address the finding.\\n\\n\")\n\twriteResolutionFormat(\u0026b)\n\n\treturn b.String()\n}\n\nfunc buildCrossFilePrompt(t TriagedThread, cfg *ReviewConfig) string {\n\tvar b strings.Builder\n\n\tb.WriteString(\"You are a code reviewer. You previously raised a finding on a pull request. The author pushed new code, but the changes are in DIFFERENT files from where you left your finding.\\n\\n\")\n\n\twriteFinding(\u0026b, t.Thread)\n\n\twriteFileSnippet(\u0026b, t.FileSnippet)\n\n\tb.WriteString(\"## All Code Changes\\n```diff\\n\")\n\tb.WriteString(t.FileDiff)\n\tb.WriteString(\"\\n```\\n\\n\")\n\n\tif ctx := evalContext(cfg, \"code_change\"); ctx != \"\" {\n\t\tfmt.Fprintf(\u0026b, \"## Additional Context\\n%s\\n\\n\", ctx)\n\t}\n\n\tb.WriteString(\"## Task\\n\")\n\tb.WriteString(\"Determine whether the issue you raised has been resolved, even though the changes are in different files from where you left your finding.\\n\")\n\tb.WriteString(\"**Start with the current file content** (if provided). If the issue no longer exists in the current code, the answer is YES.\\n\\n\")\n\tb.WriteString(\"Otherwise, examine the code changes for evidence of a cross-file fix.\\n\")\n\twriteCrossFileCriteria(\u0026b)\n\tb.WriteString(\"- Answer NO if none of the changes in this diff are related to the finding, or if file context is provided and the concern is still present in the current code.\\n\\n\")\n\twriteCodeChangeResolutionFormat(\u0026b)\n\n\treturn b.String()\n}\n\n// buildWidenedScopePrompt is the level-2 fallback prompt used when the file-scoped\n// evaluation (level 1) found no fix. It sends the full PR diff so the evaluator can\n// detect cross-file fixes that the narrower scope missed.\nfunc buildWidenedScopePrompt(t TriagedThread, cfg *ReviewConfig) string {\n\tvar b strings.Builder\n\n\tb.WriteString(\"You are a code reviewer. You previously raised a finding on a pull request. A focused check of the finding's file did not find a fix. Now examine ALL code changes across the PR — the fix may be in a different file.\\n\\n\")\n\n\twriteFinding(\u0026b, t.Thread)\n\n\twriteFileSnippet(\u0026b, t.FileSnippet)\n\n\tb.WriteString(\"## All Code Changes (full PR)\\n```diff\\n\")\n\tb.WriteString(t.FullDiff)\n\tb.WriteString(\"\\n```\\n\\n\")\n\n\tif ctx := evalContext(cfg, \"code_change\"); ctx != \"\" {\n\t\tfmt.Fprintf(\u0026b, \"## Additional Context\\n%s\\n\\n\", ctx)\n\t}\n\n\tb.WriteString(\"## Task\\n\")\n\tb.WriteString(\"The finding's file was already checked and the issue appears unresolved there. Examine the full PR diff to determine if a change in another file resolves the concern.\\n\\n\")\n\twriteCrossFileCriteria(\u0026b)\n\tb.WriteString(\"- Answer NO if none of the changes in the diff are related to the finding.\\n\\n\")\n\twriteCodeChangeResolutionFormat(\u0026b)\n\n\treturn b.String()\n}\n\n// writeCrossFileCriteria writes the shared YES-criteria bullets for cross-file evaluations.\n// Used by both buildCrossFilePrompt and buildWidenedScopePrompt.\nfunc writeCrossFileCriteria(b *strings.Builder) {\n\tb.WriteString(\"- Answer YES if a change in another file effectively resolves the concern (e.g. fixing the caller instead of the callee, adding validation in a different layer, removing the code path that triggers the issue).\\n\")\n\tb.WriteString(\"- A structural change also counts — for example, if code was moved, control flow was reordered, or the code was refactored so the finding no longer applies.\\n\")\n}\n\n// writeFileSnippet adds the current file content section to the prompt when available.\nfunc writeFileSnippet(b *strings.Builder, snippet string) {\n\tif snippet == \"\" {\n\t\treturn\n\t}\n\tb.WriteString(\"## Current File Content (around finding)\\n\")\n\tb.WriteString(\"This shows the file as it exists NOW (after the changes). Use it to understand the final code structure and control flow.\\n\\n~~~\\n\")\n\tb.WriteString(snippet)\n\tb.WriteString(\"~~~\\n\\n\")\n}\n\n// writeFinding writes the finding section to the prompt.\nfunc writeFinding(b *strings.Builder, t ReviewThread) {\n\tb.WriteString(\"## Finding\\n\")\n\tfmt.Fprintf(b, \"File: `%s:%d`\\n\", t.Path, t.Line)\n\tb.WriteString(t.Body)\n\tb.WriteString(\"\\n\\n\")\n}\n\n// writeReplies writes the author replies section to the prompt.\n// Bot replies are filtered out so the bot's own acknowledgment messages\n// don't leak into the Claude prompt and bias evaluation.\nfunc writeReplies(b *strings.Builder, t ReviewThread, botLogin string) {\n\tif len(t.Replies) == 0 {\n\t\treturn\n\t}\n\t// Filter out bot replies using explicit botLogin, consistent with\n\t// hasHumanReply and hasNewHumanReply.\n\tvar humanReplies []ThreadReply\n\tfor _, r := range t.Replies {\n\t\tif r.Author != botLogin {\n\t\t\thumanReplies = append(humanReplies, r)\n\t\t}\n\t}\n\tif len(humanReplies) == 0 {\n\t\treturn\n\t}\n\tb.WriteString(\"## Author Replies\\n\")\n\tfor _, r := range humanReplies {\n\t\tnormalizedBody := strings.ReplaceAll(r.Body, \"\\n\", \" \")\n\t\tfmt.Fprintf(b, \"\u003e **@%s**: %s\\n\", r.Author, normalizedBody)\n\t}\n\tb.WriteString(\"\\n\")\n}\n\n// writeResolutionFormat writes the expected JSON response format with all reason options.\n// Used for prompts where author replies are present (TriageHasReply, TriageCodeChangedReply).\nfunc writeResolutionFormat(b *strings.Builder) {\n\tb.WriteString(\"Return a JSON object inside a ```json code fence:\\n\")\n\tb.WriteString(\"- If resolved: `{\\\"resolved\\\": true, \\\"reason\\\": \\\"\u003creason\u003e\\\", \\\"rationale\\\": \\\"\u003cone sentence\u003e\\\"}` — `reason` is one of `\\\"code_change\\\"`, `\\\"dismissed\\\"`, `\\\"acknowledged\\\"`, `\\\"rebutted\\\"`. `rationale` must be a single sentence naming the specific change, reply, or reasoning that resolved the finding (e.g. \\\"Removed the three `original_error_*` fields from the service logger\\\").\\n\")\n\tb.WriteString(\"- If NOT resolved: `{\\\"resolved\\\": false}`\\n\")\n}\n\n// writeCodeChangeResolutionFormat writes a restricted JSON response format\n// for code-change-only evaluations (no author reply). Only allows code_change\n// as a resolution reason since there is no author reply to acknowledge/dismiss/rebut.\nfunc writeCodeChangeResolutionFormat(b *strings.Builder) {\n\tb.WriteString(\"Return a JSON object inside a ```json code fence:\\n\")\n\tb.WriteString(\"- If resolved: `{\\\"resolved\\\": true, \\\"reason\\\": \\\"code_change\\\", \\\"rationale\\\": \\\"\u003cone sentence\u003e\\\"}` — `rationale` must be a single sentence naming the specific change that resolved the finding (e.g. \\\"Removed the three `original_error_*` fields from the service logger\\\").\\n\")\n\tb.WriteString(\"- If NOT resolved: `{\\\"resolved\\\": false}`\\n\")\n}\n\n// evalContext returns the evaluation context string for a given type from config.\nfunc evalContext(cfg *ReviewConfig, evalType string) string {\n\tif cfg == nil || cfg.Evaluation == nil {\n\t\treturn \"\"\n\t}\n\tswitch evalType {\n\tcase \"code_change\":\n\t\treturn cfg.Evaluation.CodeChange.Context\n\tcase \"reply\":\n\t\treturn cfg.Evaluation.Reply.Context\n\t}\n\treturn \"\"\n}\n\n// EvaluateThreadsParallel runs the LLM in parallel for threads that need evaluation.\n// When maxBudgetUSD \u003e 0, new goroutines are not launched once the budget is exceeded\n// (already-running goroutines are allowed to finish).\nfunc EvaluateThreadsParallel(triaged []TriagedThread, provider ModelProvider, cfg *ReviewConfig, maxConcurrent int, tracker *UsageTracker, maxBudgetUSD float64) []ThreadResolution {\n\tresults := make([]ThreadResolution, len(triaged))\n\n\tsem := make(chan struct{}, maxConcurrent)\n\tvar wg sync.WaitGroup\n\n\tfor i, t := range triaged {\n\t\tif t.Class == TriageSkip || t.Class == TriageFileRemovedFromPR {\n\t\t\tresults[i] = ThreadResolution{Index: t.Index, Resolved: false}\n\t\t\tcontinue\n\t\t}\n\t\tif t.Class == TriagePreviouslyAcked {\n\t\t\t// Sticky: carry the prior ack reason forward without burning\n\t\t\t// triage tokens. computeReviewSummary will route it to the\n\t\t\t// matching bucket (Dismissed/Acknowledged/Rebutted), so the\n\t\t\t// commit status check stays green across pushes that don't\n\t\t\t// touch the deferred finding.\n\t\t\tresults[i] = ThreadResolution{\n\t\t\t\tIndex: t.Index,\n\t\t\t\tResolved: true,\n\t\t\t\tReason: t.PriorAckReason,\n\t\t\t}\n\t\t\tcontinue\n\t\t}\n\t\t// Soft budget cap: skip remaining evaluations if budget is exceeded.\n\t\tif err := CheckBudget(tracker, maxBudgetUSD); err != nil {\n\t\t\tresults[i] = ThreadResolution{Index: t.Index, Error: err}\n\t\t\tcontinue\n\t\t}\n\t\twg.Add(1)\n\t\tgo func(idx int, tt TriagedThread) {\n\t\t\tdefer wg.Done()\n\t\t\tsem \u003c- struct{}{}\n\t\t\tdefer func() { \u003c-sem }()\n\n\t\t\t// Level 1: file-scoped evaluation.\n\t\t\tprompt := BuildPerThreadPrompt(tt, cfg)\n\t\t\tresult, err := provider.Run(context.Background(), prompt, RunOpts{})\n\t\t\tif err != nil {\n\t\t\t\tresults[idx] = ThreadResolution{Index: tt.Index, Error: err}\n\t\t\t\treturn\n\t\t\t}\n\t\t\ttrackUsage(tracker, result, \"triage\")\n\t\t\tres := parseThreadResolution(result.Text, tt.Index)\n\t\t\tres = validateResolutionReason(res, tt.Class)\n\n\t\t\t// Level 2: widen to full PR diff if file-scoped check was inconclusive.\n\t\t\t// Only for TriageCodeChanged — TriageCodeChangedReply has author replies\n\t\t\t// that buildWidenedScopePrompt doesn't include, so widening would drop\n\t\t\t// the reply context and prevent dismissed/acknowledged/rebutted resolutions.\n\t\t\tif !res.Resolved \u0026\u0026 tt.FullDiff != \"\" \u0026\u0026 tt.Class == TriageCodeChanged {\n\t\t\t\tif err := CheckBudget(tracker, maxBudgetUSD); err == nil {\n\t\t\t\t\tfmt.Fprintf(os.Stderr, \" [widen] %s — checking full PR diff\\n\", threadLabel(tt.Thread))\n\t\t\t\t\tprompt2 := buildWidenedScopePrompt(tt, cfg)\n\t\t\t\t\tresult2, err := provider.Run(context.Background(), prompt2, RunOpts{})\n\t\t\t\t\tif err == nil {\n\t\t\t\t\t\ttrackUsage(tracker, result2, \"triage\")\n\t\t\t\t\t\tres2 := parseThreadResolution(result2.Text, tt.Index)\n\t\t\t\t\t\tres2 = validateResolutionReason(res2, tt.Class)\n\t\t\t\t\t\tif res2.Resolved {\n\t\t\t\t\t\t\tres = res2\n\t\t\t\t\t\t}\n\t\t\t\t\t}\n\t\t\t\t}\n\t\t\t}\n\n\t\t\tresults[idx] = res\n\t\t}(i, t)\n\t}\n\n\twg.Wait()\n\treturn results\n}\n\n// parseThreadResolution parses Claude's JSON response for a single thread evaluation.\nfunc parseThreadResolution(output string, index int) ThreadResolution {\n\tallMatches := jsonFenceRe.FindAllStringSubmatch(output, -1)\n\tfor _, matches := range allMatches {\n\t\traw := matches[1]\n\n\t\tvar resp struct {\n\t\t\tResolved bool `json:\"resolved\"`\n\t\t\tReason string `json:\"reason\"`\n\t\t\tRationale string `json:\"rationale\"`\n\t\t}\n\t\tif err := json.Unmarshal([]byte(raw), \u0026resp); err != nil {\n\t\t\tcontinue\n\t\t}\n\t\treturn ThreadResolution{\n\t\t\tIndex: index,\n\t\t\tResolved: resp.Resolved,\n\t\t\tReason: resp.Reason,\n\t\t\tRationale: strings.TrimSpace(resp.Rationale),\n\t\t}\n\t}\n\n\t// If parsing fails, treat as unresolved (conservative).\n\treturn ThreadResolution{Index: index, Resolved: false}\n}\n\n// validateResolutionReason enforces that code-change-only classifications\n// (no author reply) can only resolve with reason \"code_change\". If Claude\n// returns \"acknowledged\"/\"dismissed\"/\"rebutted\", treat it as unresolved.\nfunc validateResolutionReason(res ThreadResolution, class ThreadClassification) ThreadResolution {\n\tif res.Resolved \u0026\u0026 res.Reason != \"code_change\" \u0026\u0026\n\t\t(class == TriageCodeChanged || class == TriageCrossFileChange) {\n\t\tres.Resolved = false\n\t\tres.Reason = \"\"\n\t}\n\treturn res\n}\n\n// LogTriage prints structured triage results to stderr.\nfunc LogTriage(triaged []TriagedThread) {\n\tfmt.Fprintf(os.Stderr, \"Re-evaluating %d unresolved thread(s)...\\n\\n\", len(triaged))\n\n\tfor _, t := range triaged {\n\t\tlabel := threadLabel(t.Thread)\n\t\tswitch t.Class {\n\t\tcase TriageSkip:\n\t\t\tfmt.Fprintf(os.Stderr, \" [skip] %s — no code changes, no human replies\\n\", label)\n\t\tcase TriageCodeChanged:\n\t\t\tfmt.Fprintf(os.Stderr, \" [evaluate] %s — code changes detected\\n\", label)\n\t\tcase TriageHasReply:\n\t\t\tfmt.Fprintf(os.Stderr, \" [evaluate] %s — human reply detected\\n\", label)\n\t\tcase TriageCodeChangedReply:\n\t\t\tfmt.Fprintf(os.Stderr, \" [evaluate] %s — code changes + human reply detected\\n\", label)\n\t\tcase TriageCrossFileChange:\n\t\t\tfmt.Fprintf(os.Stderr, \" [evaluate] %s — cross-file changes detected\\n\", label)\n\t\tcase TriageFileRemovedFromPR:\n\t\t\tfmt.Fprintf(os.Stderr, \" [resolve] %s — file removed from PR\\n\", label)\n\t\tcase TriagePreviouslyAcked:\n\t\t\tfmt.Fprintf(os.Stderr, \" [sticky] %s — prior ack (%s) carried forward\\n\", label, t.PriorAckReason)\n\t\t}\n\t}\n\n\tskipped := 0\n\tautoResolved := 0\n\tneedsEval := 0\n\tfor _, t := range triaged {\n\t\tswitch t.Class {\n\t\tcase TriageSkip:\n\t\t\tskipped++\n\t\tcase TriageFileRemovedFromPR, TriagePreviouslyAcked:\n\t\t\tautoResolved++\n\t\tdefault:\n\t\t\tneedsEval++\n\t\t}\n\t}\n\tfmt.Fprintf(os.Stderr, \"\\nTriage result: %d skipped, %d auto-resolved, %d need evaluation\\n\", skipped, autoResolved, needsEval)\n}\n\n// LogResolutions prints structured evaluation results to stderr.\nfunc LogResolutions(triaged []TriagedThread, resolutions []ThreadResolution) {\n\tfmt.Fprintf(os.Stderr, \"\\n\")\n\tfor i, r := range resolutions {\n\t\tswitch triaged[i].Class {\n\t\tcase TriageSkip, TriageFileRemovedFromPR, TriagePreviouslyAcked:\n\t\t\tcontinue\n\t\t}\n\t\tlabel := threadLabel(triaged[i].Thread)\n\t\tif r.Error != nil {\n\t\t\tif isBudgetError(r.Error) {\n\t\t\t\tfmt.Fprintf(os.Stderr, \" [skip] %s — %v\\n\", label, r.Error)\n\t\t\t} else {\n\t\t\t\tfmt.Fprintf(os.Stderr, \" [error] %s — evaluation failed: %v\\n\", label, r.Error)\n\t\t\t}\n\t\t} else if r.Resolved {\n\t\t\tswitch r.Reason {\n\t\t\tcase \"code_change\":\n\t\t\t\tfmt.Fprintf(os.Stderr, \" [resolved] %s — fixed by code change\\n\", label)\n\t\t\tcase \"dismissed\":\n\t\t\t\tfmt.Fprintf(os.Stderr, \" [ack] %s — dismissed by author (keeping open)\\n\", label)\n\t\t\tcase \"acknowledged\":\n\t\t\t\tfmt.Fprintf(os.Stderr, \" [ack] %s — acknowledged by author (keeping open)\\n\", label)\n\t\t\tcase \"rebutted\":\n\t\t\t\tfmt.Fprintf(os.Stderr, \" [ack] %s — rebutted by author (keeping open)\\n\", label)\n\t\t\tdefault:\n\t\t\t\tfmt.Fprintf(os.Stderr, \" [resolved] %s — resolved\\n\", label)\n\t\t\t}\n\t\t} else {\n\t\t\tfmt.Fprintf(os.Stderr, \" [open] %s — not resolved\\n\", label)\n\t\t}\n\t}\n}\n\n// countNonSkipped returns the number of triaged threads that EvaluateThreadsParallel\n// must process. TriagePreviouslyAcked is included even though it short-circuits\n// without an LLM call: the function still has to run to emit the carried-forward\n// fixedThread, which is what feeds the summary's Acknowledged/Rebutted/Dismissed\n// buckets. Excluding it here was the bug behind the \"Still unresolved: 2\" stuck\n// status on already-acked threads.\nfunc countNonSkipped(triaged []TriagedThread) int {\n\tn := 0\n\tfor _, t := range triaged {\n\t\tswitch t.Class {\n\t\tcase TriageSkip, TriageFileRemovedFromPR:\n\t\t\tcontinue\n\t\tdefault:\n\t\t\tn++\n\t\t}\n\t}\n\treturn n\n}\n\n// toFixedThreads converts thread resolutions to the fixedThread type used by downstream code.\nfunc toFixedThreads(resolutions []ThreadResolution) []fixedThread {\n\tvar result []fixedThread\n\tfor _, r := range resolutions {\n\t\tif r.Resolved {\n\t\t\tresult = append(result, fixedThread{Index: r.Index, Reason: r.Reason, Rationale: r.Rationale})\n\t\t}\n\t}\n\treturn result\n}\n", + "internal/review/triage_test.go": "package review\n\nimport (\n\t\"context\"\n\t\"fmt\"\n\t\"strings\"\n\t\"testing\"\n)\n\n// --- ExtractFileSnippet tests ---\n\nfunc TestExtractFileSnippet_Basic(t *testing.T) {\n\t// 100-line file, finding at line 50, hunk at lines 45-55.\n\tvar lines []string\n\tfor i := 1; i \u003c= 100; i++ {\n\t\tlines = append(lines, fmt.Sprintf(\"line %d content\", i))\n\t}\n\tcontent := strings.Join(lines, \"\\n\")\n\tdiff := \"@@ -40,10 +45,11 @@ func foo() {\\n+added line\\n\"\n\n\tsnippet := ExtractFileSnippet(content, 50, diff, 300)\n\tif snippet == \"\" {\n\t\tt.Fatal(\"expected non-empty snippet\")\n\t}\n\t// Should contain the finding line.\n\tif !strings.Contains(snippet, \"50: line 50 content\") {\n\t\tt.Error(\"snippet should contain the finding line\")\n\t}\n\t// Should contain hunk area.\n\tif !strings.Contains(snippet, \"45: line 45 content\") {\n\t\tt.Error(\"snippet should contain hunk start area\")\n\t}\n}\n\nfunc TestExtractFileSnippet_MergesOverlappingRanges(t *testing.T) {\n\tvar lines []string\n\tfor i := 1; i \u003c= 200; i++ {\n\t\tlines = append(lines, fmt.Sprintf(\"line %d\", i))\n\t}\n\tcontent := strings.Join(lines, \"\\n\")\n\t// Two hunks close together — should merge into one contiguous range.\n\tdiff := \"@@ -10,5 +10,5 @@\\n+a\\n@@ -20,5 +20,5 @@\\n+b\\n\"\n\n\tsnippet := ExtractFileSnippet(content, 15, diff, 300)\n\t// Should NOT contain omission markers since ranges overlap/merge.\n\tif strings.Contains(snippet, \"lines omitted\") {\n\t\tt.Error(\"close hunks should merge without omission markers\")\n\t}\n}\n\nfunc TestExtractFileSnippet_CapsAtMaxLines(t *testing.T) {\n\tvar lines []string\n\tfor i := 1; i \u003c= 1000; i++ {\n\t\tlines = append(lines, fmt.Sprintf(\"line %d\", i))\n\t}\n\tcontent := strings.Join(lines, \"\\n\")\n\t// Hunks spread across the file.\n\tdiff := \"@@ -10,5 +10,5 @@\\n+a\\n@@ -500,5 +500,5 @@\\n+b\\n@@ -900,5 +900,5 @@\\n+c\\n\"\n\n\tsnippet := ExtractFileSnippet(content, 50, diff, 150)\n\tsnippetLines := strings.Split(strings.TrimRight(snippet, \"\\n\"), \"\\n\")\n\tif len(snippetLines) \u003e 160 { // small buffer for omission markers\n\t\tt.Errorf(\"snippet should respect maxLines cap, got %d lines\", len(snippetLines))\n\t}\n\t// Must contain the finding line.\n\tif !strings.Contains(snippet, \"50: line 50\") {\n\t\tt.Error(\"snippet must prioritize the finding line\")\n\t}\n}\n\nfunc TestExtractFileSnippet_NoDiff(t *testing.T) {\n\tvar lines []string\n\tfor i := 1; i \u003c= 100; i++ {\n\t\tlines = append(lines, fmt.Sprintf(\"line %d\", i))\n\t}\n\tcontent := strings.Join(lines, \"\\n\")\n\n\t// Cross-file case: no diff for this file.\n\tsnippet := ExtractFileSnippet(content, 50, \"\", 300)\n\tif snippet == \"\" {\n\t\tt.Fatal(\"expected non-empty snippet for cross-file case\")\n\t}\n\tif !strings.Contains(snippet, \"50: line 50\") {\n\t\tt.Error(\"snippet should center on finding line\")\n\t}\n}\n\nfunc TestExtractFileSnippet_ZeroCountHunk(t *testing.T) {\n\t// A hunk with count=0 (pure deletion) should not produce a range.\n\tranges := parseHunkNewRanges(\"@@ -5,3 +10,0 @@\\n-deleted line\\n\")\n\tif len(ranges) != 0 {\n\t\tt.Errorf(\"expected 0 ranges for zero-count hunk, got %d\", len(ranges))\n\t}\n}\n\nfunc TestExtractFileSnippet_FindingLineZero(t *testing.T) {\n\tvar lines []string\n\tfor i := 1; i \u003c= 100; i++ {\n\t\tlines = append(lines, fmt.Sprintf(\"line %d\", i))\n\t}\n\tcontent := strings.Join(lines, \"\\n\")\n\tdiff := \"@@ -40,10 +40,10 @@\\n+changed\\n\"\n\n\t// findingLine=0 should not panic and should anchor to hunk area.\n\tsnippet := ExtractFileSnippet(content, 0, diff, 50)\n\tif snippet == \"\" {\n\t\tt.Fatal(\"expected non-empty snippet even with findingLine=0\")\n\t}\n\t// Should contain hunk area, not be anchored to line 0.\n\tif !strings.Contains(snippet, \"40: line 40\") {\n\t\tt.Error(\"snippet should include hunk area when findingLine is 0\")\n\t}\n}\n\nfunc TestExtractFileSnippet_EmptyContent(t *testing.T) {\n\tsnippet := ExtractFileSnippet(\"\", 10, \"@@ -1,5 +1,5 @@\\n\", 300)\n\tif snippet != \"\" {\n\t\tt.Error(\"expected empty snippet for empty content\")\n\t}\n}\n\n// --- Prompt builder tests ---\n\nfunc TestBuildCodeChangePrompt_IncludesFileContext(t *testing.T) {\n\ttt := TriagedThread{\n\t\tThread: ReviewThread{\n\t\t\tPath: \"main.go\",\n\t\t\tLine: 10,\n\t\t\tBody: \"Found a bug\",\n\t\t},\n\t\tFileDiff: \"+ fixed line\",\n\t\tFileSnippet: \"9: before\\n10: the line\\n11: after\\n\",\n\t}\n\tprompt := buildCodeChangePrompt(tt, nil)\n\n\tif !strings.Contains(prompt, \"## Current File Content (around finding)\") {\n\t\tt.Error(\"prompt should include file context section when FileSnippet is set\")\n\t}\n\tif !strings.Contains(prompt, \"10: the line\") {\n\t\tt.Error(\"prompt should include the file snippet content\")\n\t}\n}\n\nfunc TestBuildCodeChangePrompt_NoFileContextWhenEmpty(t *testing.T) {\n\ttt := TriagedThread{\n\t\tThread: ReviewThread{\n\t\t\tPath: \"main.go\",\n\t\t\tLine: 10,\n\t\t\tBody: \"Found a bug\",\n\t\t},\n\t\tFileDiff: \"+ fixed line\",\n\t}\n\tprompt := buildCodeChangePrompt(tt, nil)\n\n\tif strings.Contains(prompt, \"## Current File Content\") {\n\t\tt.Error(\"prompt should NOT include file context section when FileSnippet is empty\")\n\t}\n}\n\nfunc TestBuildCodeChangePrompt_StructuralChangeInstruction(t *testing.T) {\n\ttt := TriagedThread{\n\t\tThread: ReviewThread{\n\t\t\tPath: \"main.go\",\n\t\t\tLine: 10,\n\t\t\tBody: \"Found a bug\",\n\t\t},\n\t\tFileDiff: \"+ fixed line\",\n\t}\n\tprompt := buildCodeChangePrompt(tt, nil)\n\n\tif !strings.Contains(prompt, \"structural change\") {\n\t\tt.Error(\"prompt should include structural change guidance\")\n\t}\n}\n\nfunc TestBuildCrossFilePrompt_IncludesFileContext(t *testing.T) {\n\ttt := TriagedThread{\n\t\tThread: ReviewThread{\n\t\t\tPath: \"main.go\",\n\t\t\tLine: 10,\n\t\t\tBody: \"Found a bug\",\n\t\t},\n\t\tFileDiff: \"+ change in other file\",\n\t\tFileSnippet: \"9: before\\n10: the line\\n11: after\\n\",\n\t}\n\tprompt := buildCrossFilePrompt(tt, nil)\n\n\tif !strings.Contains(prompt, \"## Current File Content (around finding)\") {\n\t\tt.Error(\"cross-file prompt should include file context section\")\n\t}\n\tif !strings.Contains(prompt, \"structural change\") {\n\t\tt.Error(\"cross-file prompt should include structural change guidance\")\n\t}\n}\n\nfunc TestBuildCodeChangePrompt_OnlyAllowsCodeChangeReason(t *testing.T) {\n\ttt := TriagedThread{\n\t\tThread: ReviewThread{\n\t\t\tPath: \"main.go\",\n\t\t\tLine: 10,\n\t\t\tBody: \"Found a bug\",\n\t\t},\n\t\tFileDiff: \"+ fixed line\",\n\t}\n\tprompt := buildCodeChangePrompt(tt, nil)\n\n\tif strings.Contains(prompt, `\"acknowledged\"`) {\n\t\tt.Error(\"buildCodeChangePrompt should not offer 'acknowledged' as a reason\")\n\t}\n\tif strings.Contains(prompt, `\"dismissed\"`) {\n\t\tt.Error(\"buildCodeChangePrompt should not offer 'dismissed' as a reason\")\n\t}\n\tif strings.Contains(prompt, `\"rebutted\"`) {\n\t\tt.Error(\"buildCodeChangePrompt should not offer 'rebutted' as a reason\")\n\t}\n\tif !strings.Contains(prompt, `\"code_change\"`) {\n\t\tt.Error(\"buildCodeChangePrompt must offer 'code_change' as a reason\")\n\t}\n}\n\nfunc TestBuildCrossFilePrompt_OnlyAllowsCodeChangeReason(t *testing.T) {\n\ttt := TriagedThread{\n\t\tThread: ReviewThread{\n\t\t\tPath: \"main.go\",\n\t\t\tLine: 10,\n\t\t\tBody: \"Found a bug\",\n\t\t},\n\t\tFileDiff: \"+ fixed in other file\",\n\t}\n\tprompt := buildCrossFilePrompt(tt, nil)\n\n\tif strings.Contains(prompt, `\"acknowledged\"`) {\n\t\tt.Error(\"buildCrossFilePrompt should not offer 'acknowledged' as a reason\")\n\t}\n\tif strings.Contains(prompt, `\"dismissed\"`) {\n\t\tt.Error(\"buildCrossFilePrompt should not offer 'dismissed' as a reason\")\n\t}\n\tif strings.Contains(prompt, `\"rebutted\"`) {\n\t\tt.Error(\"buildCrossFilePrompt should not offer 'rebutted' as a reason\")\n\t}\n\tif !strings.Contains(prompt, `\"code_change\"`) {\n\t\tt.Error(\"buildCrossFilePrompt must offer 'code_change' as a reason\")\n\t}\n}\n\nfunc TestBuildReplyPrompt_AllowsAllReasons(t *testing.T) {\n\ttt := TriagedThread{\n\t\tThread: ReviewThread{\n\t\t\tPath: \"main.go\",\n\t\t\tLine: 10,\n\t\t\tBody: \"Found a bug\",\n\t\t\tReplies: []ThreadReply{\n\t\t\t\t{Author: \"user1\", Body: \"Will fix later\"},\n\t\t\t},\n\t\t},\n\t\tBotLogin: \"codecanary-bot\",\n\t}\n\tprompt := buildReplyPrompt(tt, nil)\n\n\tfor _, reason := range []string{`\"code_change\"`, `\"acknowledged\"`, `\"dismissed\"`, `\"rebutted\"`} {\n\t\tif !strings.Contains(prompt, reason) {\n\t\t\tt.Errorf(\"buildReplyPrompt must offer %s as a reason\", reason)\n\t\t}\n\t}\n}\n\nfunc TestBuildCodeChangeReplyPrompt_AllowsAllReasons(t *testing.T) {\n\ttt := TriagedThread{\n\t\tThread: ReviewThread{\n\t\t\tPath: \"main.go\",\n\t\t\tLine: 10,\n\t\t\tBody: \"Found a bug\",\n\t\t\tReplies: []ThreadReply{\n\t\t\t\t{Author: \"user1\", Body: \"Fixed it\"},\n\t\t\t},\n\t\t},\n\t\tFileDiff: \"+ fixed line\",\n\t\tBotLogin: \"codecanary-bot\",\n\t}\n\tprompt := buildCodeChangeReplyPrompt(tt, nil)\n\n\tfor _, reason := range []string{`\"code_change\"`, `\"acknowledged\"`, `\"dismissed\"`, `\"rebutted\"`} {\n\t\tif !strings.Contains(prompt, reason) {\n\t\t\tt.Errorf(\"buildCodeChangeReplyPrompt must offer %s as a reason\", reason)\n\t\t}\n\t}\n}\n\n// --- ClassifyThreads diff scoping tests ---\n\nfunc TestClassifyThreads_FileScopedDiffForCodeChanged(t *testing.T) {\n\tthreads := []ReviewThread{\n\t\t{Path: \"a.go\", Line: 10, Body: \"Issue in a.go\", Outdated: true},\n\t}\n\tfullDiff := \"diff --git a/a.go b/a.go\\n--- a/a.go\\n+++ b/a.go\\n@@ -10,3 +10,3 @@\\n-old\\n+new\\n\" +\n\t\t\"diff --git a/b.go b/b.go\\n--- a/b.go\\n+++ b/b.go\\n@@ -5,3 +5,3 @@\\n-old b\\n+new b\\n\"\n\n\ttriaged := ClassifyThreads(threads, fullDiff, fullDiff, \"bot\", []string{\"a.go\", \"b.go\"}, nil)\n\n\tif triaged[0].Class != TriageCodeChanged {\n\t\tt.Fatalf(\"expected TriageCodeChanged, got %d\", triaged[0].Class)\n\t}\n\tif strings.Contains(triaged[0].FileDiff, \"b.go\") {\n\t\tt.Error(\"FileDiff for TriageCodeChanged should be file-scoped, not full PR diff\")\n\t}\n\tif !strings.Contains(triaged[0].FileDiff, \"a.go\") {\n\t\tt.Error(\"FileDiff should contain the finding's file diff\")\n\t}\n\t// FullDiff should contain the entire PR diff for widened-scope fallback.\n\tif !strings.Contains(triaged[0].FullDiff, \"b.go\") {\n\t\tt.Error(\"FullDiff should contain the full PR diff for fallback\")\n\t}\n}\n\nfunc TestClassifyThreads_NoFullDiffForCrossFile(t *testing.T) {\n\tthreads := []ReviewThread{\n\t\t{Path: \"a.go\", Line: 10, Body: \"Issue in a.go\"},\n\t}\n\tdiff := \"diff --git a/b.go b/b.go\\n--- a/b.go\\n+++ b/b.go\\n@@ -5,3 +5,3 @@\\n-old b\\n+new b\\n\"\n\n\ttriaged := ClassifyThreads(threads, diff, diff, \"bot\", []string{\"a.go\", \"b.go\"}, nil)\n\n\tif triaged[0].Class != TriageCrossFileChange {\n\t\tt.Fatalf(\"expected TriageCrossFileChange, got %d\", triaged[0].Class)\n\t}\n\t// TriageCrossFileChange already gets the full diff as FileDiff — no fallback needed.\n\tif triaged[0].FullDiff != \"\" {\n\t\tt.Error(\"FullDiff should be empty for TriageCrossFileChange (no fallback needed)\")\n\t}\n}\n\nfunc TestBuildWidenedScopePrompt(t *testing.T) {\n\ttt := TriagedThread{\n\t\tThread: ReviewThread{\n\t\t\tPath: \"main.go\",\n\t\t\tLine: 10,\n\t\t\tBody: \"Found a bug\",\n\t\t},\n\t\tFullDiff: \"+ fix in other file\",\n\t\tFileSnippet: \"9: before\\n10: the line\\n11: after\\n\",\n\t}\n\tprompt := buildWidenedScopePrompt(tt, nil)\n\n\tif !strings.Contains(prompt, \"full PR diff\") {\n\t\tt.Error(\"widened prompt should mention full PR diff\")\n\t}\n\tif !strings.Contains(prompt, \"another file\") {\n\t\tt.Error(\"widened prompt should guide LLM to check other files\")\n\t}\n\tif !strings.Contains(prompt, \"+ fix in other file\") {\n\t\tt.Error(\"widened prompt should include FullDiff content\")\n\t}\n\tif !strings.Contains(prompt, \"## Current File Content\") {\n\t\tt.Error(\"widened prompt should include file snippet\")\n\t}\n\t// Should only allow code_change reason (no reply-based reasons).\n\tif strings.Contains(prompt, `\"acknowledged\"`) || strings.Contains(prompt, `\"dismissed\"`) {\n\t\tt.Error(\"widened prompt should not offer reply-based resolution reasons\")\n\t}\n}\n\nfunc TestClassifyThreads_FullDiffForCrossFile(t *testing.T) {\n\tthreads := []ReviewThread{\n\t\t{Path: \"a.go\", Line: 10, Body: \"Issue in a.go\"},\n\t}\n\t// Only b.go changed — finding's file (a.go) is not in the diff.\n\tdiff := \"diff --git a/b.go b/b.go\\n--- a/b.go\\n+++ b/b.go\\n@@ -5,3 +5,3 @@\\n-old b\\n+new b\\n\"\n\n\ttriaged := ClassifyThreads(threads, diff, diff, \"bot\", []string{\"a.go\", \"b.go\"}, nil)\n\n\tif triaged[0].Class != TriageCrossFileChange {\n\t\tt.Fatalf(\"expected TriageCrossFileChange, got %d\", triaged[0].Class)\n\t}\n\tif !strings.Contains(triaged[0].FileDiff, \"b.go\") {\n\t\tt.Error(\"FileDiff for TriageCrossFileChange should contain the full PR diff\")\n\t}\n}\n\nfunc TestValidateResolutionReason_RejectsInvalidReasonForCodeChangeOnly(t *testing.T) {\n\t// Simulate Claude returning \"acknowledged\" for a code-change-only thread.\n\toutput := \"```json\\n{\\\"resolved\\\": true, \\\"reason\\\": \\\"acknowledged\\\"}\\n```\"\n\tparsed := parseThreadResolution(output, 0)\n\n\t// parseThreadResolution itself accepts any reason (it's just a parser).\n\tif !parsed.Resolved || parsed.Reason != \"acknowledged\" {\n\t\tt.Fatal(\"parseThreadResolution should parse the raw response as-is\")\n\t}\n\n\t// For code-change-only classifications, invalid reasons should be rejected.\n\tfor _, class := range []ThreadClassification{TriageCodeChanged, TriageCrossFileChange} {\n\t\tres := validateResolutionReason(parsed, class)\n\t\tif res.Resolved {\n\t\t\tt.Errorf(\"class %d: resolution with reason 'acknowledged' should be rejected\", class)\n\t\t}\n\t\tif res.Reason != \"\" {\n\t\t\tt.Errorf(\"class %d: reason should be cleared, got %q\", class, res.Reason)\n\t\t}\n\t}\n\n\t// For reply-based classifications, the same reason should be accepted.\n\tfor _, class := range []ThreadClassification{TriageHasReply, TriageCodeChangedReply} {\n\t\tres := validateResolutionReason(parsed, class)\n\t\tif !res.Resolved {\n\t\t\tt.Errorf(\"class %d: resolution with reason 'acknowledged' should be accepted\", class)\n\t\t}\n\t\tif res.Reason != \"acknowledged\" {\n\t\t\tt.Errorf(\"class %d: reason should be 'acknowledged', got %q\", class, res.Reason)\n\t\t}\n\t}\n}\n\nfunc TestParseThreadResolution_CapturesRationale(t *testing.T) {\n\toutput := \"```json\\n{\\\"resolved\\\": true, \\\"reason\\\": \\\"code_change\\\", \\\"rationale\\\": \\\" Removed the three original_error_* fields \\\"}\\n```\"\n\tparsed := parseThreadResolution(output, 7)\n\n\tif !parsed.Resolved || parsed.Reason != \"code_change\" {\n\t\tt.Fatalf(\"expected resolved code_change, got %+v\", parsed)\n\t}\n\tif parsed.Rationale != \"Removed the three original_error_* fields\" {\n\t\tt.Errorf(\"rationale should be trimmed, got %q\", parsed.Rationale)\n\t}\n\tif parsed.Index != 7 {\n\t\tt.Errorf(\"index should be preserved, got %d\", parsed.Index)\n\t}\n}\n\nfunc TestToFixedThreads_PropagatesRationale(t *testing.T) {\n\tresolutions := []ThreadResolution{\n\t\t{Index: 0, Resolved: true, Reason: \"code_change\", Rationale: \"dropped dead fields\"},\n\t\t{Index: 1, Resolved: false},\n\t\t{Index: 2, Resolved: true, Reason: \"dismissed\"},\n\t}\n\tfixed := toFixedThreads(resolutions)\n\n\tif len(fixed) != 2 {\n\t\tt.Fatalf(\"expected 2 fixed threads, got %d\", len(fixed))\n\t}\n\tif fixed[0].Rationale != \"dropped dead fields\" {\n\t\tt.Errorf(\"first rationale should carry over, got %q\", fixed[0].Rationale)\n\t}\n\tif fixed[1].Rationale != \"\" {\n\t\tt.Errorf(\"empty rationale should stay empty, got %q\", fixed[1].Rationale)\n\t}\n}\n\nfunc TestExtractAckReason(t *testing.T) {\n\tcases := []struct {\n\t\tname string\n\t\tbody string\n\t\twant string\n\t}{\n\t\t{\"current marker\", \"\u003c!-- codecanary:ack:rebutted --\u003e\\nKeeping open.\", \"rebutted\"},\n\t\t{\"acknowledged\", \"\u003c!-- codecanary:ack:acknowledged --\u003e\\nKeeping open.\", \"acknowledged\"},\n\t\t{\"dismissed\", \"\u003c!-- codecanary:ack:dismissed --\u003e\\nKeeping open.\", \"dismissed\"},\n\t\t{\"legacy marker\", \"\u003c!-- clanopy:ack:rebutted --\u003e\\nKeeping open.\", \"rebutted\"},\n\t\t{\"no marker\", \"Some normal reply\", \"\"},\n\t\t{\"no closing tag\", \"\u003c!-- codecanary:ack:acknowledged with no closer\", \"\"},\n\t\t{\"unknown reason\", \"\u003c!-- codecanary:ack:unknown --\u003e\\n\", \"unknown\"},\n\t}\n\tfor _, tc := range cases {\n\t\tt.Run(tc.name, func(t *testing.T) {\n\t\t\tif got := extractAckReason(tc.body); got != tc.want {\n\t\t\t\tt.Errorf(\"extractAckReason(%q) = %q, want %q\", tc.body, got, tc.want)\n\t\t\t}\n\t\t})\n\t}\n}\n\nfunc TestParsePriorAckReason_PicksLatestBotAck(t *testing.T) {\n\tthread := ReviewThread{\n\t\tReplies: []ThreadReply{\n\t\t\t{Author: \"user\", Body: \"Deferring — see rationale.\"},\n\t\t\t{Author: \"bot\", Body: \"\u003c!-- codecanary:ack:acknowledged --\u003e\\nAuthor acknowledged this finding.\"},\n\t\t\t{Author: \"bot\", Body: \"\u003c!-- codecanary:ack:rebutted --\u003e\\nAuthor provided rebuttal.\"},\n\t\t},\n\t}\n\tif got := parsePriorAckReason(thread, \"bot\"); got != \"rebutted\" {\n\t\tt.Errorf(\"expected latest reason 'rebutted', got %q\", got)\n\t}\n}\n\nfunc TestParsePriorAckReason_IgnoresNonBotMarkers(t *testing.T) {\n\tthread := ReviewThread{\n\t\tReplies: []ThreadReply{\n\t\t\t{Author: \"attacker\", Body: \"\u003c!-- codecanary:ack:dismissed --\u003e\\nFake.\"},\n\t\t},\n\t}\n\tif got := parsePriorAckReason(thread, \"bot\"); got != \"\" {\n\t\tt.Errorf(\"non-bot ack should not be sticky, got %q\", got)\n\t}\n}\n\nfunc TestClassifyThreads_StickyAckSurvivesNextPush(t *testing.T) {\n\tthreads := []ReviewThread{\n\t\t{\n\t\t\tPath: \"CLAUDE.md\",\n\t\t\tLine: 724,\n\t\t\tBody: \"Original finding\",\n\t\t\tReplies: []ThreadReply{\n\t\t\t\t{Author: \"user\", Body: \"Deferring — forward recommendation, not yet implemented.\"},\n\t\t\t\t{Author: \"bot\", Body: \"\u003c!-- codecanary:ack:acknowledged --\u003e\\nAuthor acknowledged.\"},\n\t\t\t},\n\t\t},\n\t}\n\t// File changed in the next push — without sticky-ack this would\n\t// classify as TriageCodeChanged and lose the prior reason.\n\tdiff := \"diff --git a/CLAUDE.md b/CLAUDE.md\\n--- a/CLAUDE.md\\n+++ b/CLAUDE.md\\n@@ -700,3 +700,3 @@\\n-old\\n+new\\n\"\n\n\ttriaged := ClassifyThreads(threads, diff, diff, \"bot\", []string{\"CLAUDE.md\"}, nil)\n\n\tif triaged[0].Class != TriagePreviouslyAcked {\n\t\tt.Fatalf(\"expected TriagePreviouslyAcked, got %d\", triaged[0].Class)\n\t}\n\tif triaged[0].PriorAckReason != \"acknowledged\" {\n\t\tt.Errorf(\"expected PriorAckReason 'acknowledged', got %q\", triaged[0].PriorAckReason)\n\t}\n}\n\nfunc TestClassifyThreads_NewHumanReplyBreaksStickiness(t *testing.T) {\n\t// Author replied AGAIN after the bot's ack — that's a fresh signal\n\t// and should reopen the thread for evaluation, not stay sticky.\n\tthreads := []ReviewThread{\n\t\t{\n\t\t\tPath: \"CLAUDE.md\",\n\t\t\tLine: 724,\n\t\t\tReplies: []ThreadReply{\n\t\t\t\t{Author: \"user\", Body: \"Original deferring rationale.\"},\n\t\t\t\t{Author: \"bot\", Body: \"\u003c!-- codecanary:ack:acknowledged --\u003e\\nAck.\"},\n\t\t\t\t{Author: \"user\", Body: \"Wait, actually this one needs another look.\"},\n\t\t\t},\n\t\t},\n\t}\n\ttriaged := ClassifyThreads(threads, \"\", \"\", \"bot\", []string{\"CLAUDE.md\"}, nil)\n\n\tif triaged[0].Class != TriageHasReply {\n\t\tt.Fatalf(\"expected TriageHasReply (new reply after ack), got %d\", triaged[0].Class)\n\t}\n\tif triaged[0].PriorAckReason != \"\" {\n\t\tt.Errorf(\"PriorAckReason should be empty when new reply present, got %q\", triaged[0].PriorAckReason)\n\t}\n}\n\nfunc TestCountNonSkipped_IncludesPreviouslyAcked(t *testing.T) {\n\t// Regression: sticky-ack threads must count toward needsEval so\n\t// EvaluateThreadsParallel runs and emits the carried-forward fixedThread.\n\t// Excluding them caused the runner to skip eval entirely, leaving the\n\t// summary stuck on \"Still unresolved\" for threads the bot had already\n\t// acked.\n\ttriaged := []TriagedThread{\n\t\t{Class: TriageSkip},\n\t\t{Class: TriageFileRemovedFromPR},\n\t\t{Class: TriagePreviouslyAcked, PriorAckReason: \"acknowledged\"},\n\t\t{Class: TriagePreviouslyAcked, PriorAckReason: \"rebutted\"},\n\t}\n\tif got := countNonSkipped(triaged); got != 2 {\n\t\tt.Errorf(\"countNonSkipped = %d, want 2 (the two PreviouslyAcked threads)\", got)\n\t}\n}\n\nfunc TestEvaluateThreadsParallel_StickyAckShortCircuits(t *testing.T) {\n\t// Provider that panics if called — sticky-ack must skip the LLM.\n\tprovider := \u0026stickyAckPanicProvider{t: t}\n\n\ttriaged := []TriagedThread{\n\t\t{Index: 0, Class: TriagePreviouslyAcked, PriorAckReason: \"rebutted\"},\n\t\t{Index: 1, Class: TriagePreviouslyAcked, PriorAckReason: \"dismissed\"},\n\t}\n\tresults := EvaluateThreadsParallel(triaged, provider, nil, 3, \u0026UsageTracker{}, 0)\n\n\tif len(results) != 2 {\n\t\tt.Fatalf(\"expected 2 results, got %d\", len(results))\n\t}\n\tif !results[0].Resolved || results[0].Reason != \"rebutted\" {\n\t\tt.Errorf(\"first result: expected resolved=true reason=rebutted, got %+v\", results[0])\n\t}\n\tif !results[1].Resolved || results[1].Reason != \"dismissed\" {\n\t\tt.Errorf(\"second result: expected resolved=true reason=dismissed, got %+v\", results[1])\n\t}\n}\n\ntype stickyAckPanicProvider struct{ t *testing.T }\n\nfunc (p *stickyAckPanicProvider) Run(_ context.Context, _ string, _ RunOpts) (*providerResult, error) {\n\tp.t.Fatalf(\"provider.Run must not be called for TriagePreviouslyAcked threads\")\n\treturn nil, nil\n}\n\nfunc TestResolutionFormat_RequestsRationale(t *testing.T) {\n\tfor _, fn := range map[string]func(*strings.Builder){\n\t\t\"writeResolutionFormat\": writeResolutionFormat,\n\t\t\"writeCodeChangeResolutionFormat\": writeCodeChangeResolutionFormat,\n\t} {\n\t\tvar b strings.Builder\n\t\tfn(\u0026b)\n\t\tout := b.String()\n\t\tif !strings.Contains(out, `\"rationale\"`) {\n\t\t\tt.Errorf(\"prompt format should request a rationale field, got:\\n%s\", out)\n\t\t}\n\t}\n}\n" + } + }, + "config": {}, + "project_docs": { + "CLAUDE.md": "# CodeCanary\n\nAI-powered code review for GitHub pull requests.\n\n## Project structure\n\n```\ncmd/\n review/ # Main binary — review CLI + setup wizard\n main.go # Entry point\n cli/ # Cobra commands\n root.go # Root \"codecanary\" command\n review.go # codecanary review \u003cpr\u003e\n findings.go # codecanary findings \u003cpr\u003e — fetch bot findings for the review skill\n reply.go # codecanary reply --url \u003ccomment-url\u003e --body \u003ctext\u003e — post a reply on a review thread\n install_skill.go # codecanary install-skill — write embedded Claude skill to disk\n setup.go # codecanary setup [local|github]\n auth.go # codecanary auth [status|delete]\ninternal/\n review/\n runner.go # Core review pipeline — single Run() entry point\n config.go # Config loading, validation, defaults\n # Provider layer (LLM abstraction)\n provider.go # ModelProvider interface + factory registry\n provider_anthropic.go\n provider_openai.go\n provider_openrouter.go\n provider_claude.go # Claude CLI wrapper\n provider_compat.go # Shared types for OpenAI-compatible APIs\n pricing.go # Token-based cost estimation\n # Platform layer (environment abstraction)\n platform.go # ReviewPlatform interface\n platform_github.go # GitHub Actions implementation\n platform_local.go # Local CLI implementation\n # Supporting modules\n prompt.go # Prompt building (review, incremental, per-thread)\n findings.go # Finding parsing, filtering, result structures\n triage.go # Thread classification + parallel LLM evaluation\n formatter.go # JSON/Markdown/Terminal output formatting\n usage.go # Token tracking, budget checking\n github.go # GitHub API calls (fetch threads, post reviews)\n comments.go # PR review comment fetch + finding marker parser + review-check watcher\n local.go # Local diff \u0026 git operations\n state.go # Local state persistence\n docs.go # Project doc discovery\n credentials/ # Credential storage (keychain with file fallback)\n keyring.go # Store/Retrieve/Delete — keychain first, ~/.codecanary/credentials.json fallback\n skills/ # Claude Code skills embedded in the binary via //go:embed\n skills.go # Exports CodecanaryFix() returning the skill body\n codecanary-fix/SKILL.md # Canonical skill source (duplicated at .claude/skills/codecanary-fix/SKILL.md; parity enforced by skills_test.go)\n setup/ # Setup wizard logic (huh forms)\n forms.go # Shared huh form components\n validate.go # API key validation via test calls\n guidance.go # Token/permissions guidance text\n workflow.go # GitHub Actions workflow template\n local.go # RunLocal() — local setup flow\n github.go # RunGitHub() — GitHub Actions setup flow\n auth/ # OAuth PKCE flow, GitHub App installation\ntelemetry/ # Telemetry domain (anonymous usage analytics)\n worker/ # Cloudflare Worker — telemetry ingestion (TypeScript)\n dashboard/ # Cloudflare Pages — internal analytics dashboard (vanilla JS + Chart.js)\noidc/ # OIDC domain\n worker/ # Cloudflare Worker — OIDC token exchange proxy (TypeScript)\naction.yml # GitHub Action definition (composite action)\ninstall.sh # Downloads and installs codecanary binary permanently\n.claude/\n skills/\n codecanary-fix/ # Claude Code skill — drives review→fix→push loop using `codecanary findings` + `codecanary reply`\n```\n\n## Binary\n\n- **`codecanary`** — single binary for reviews, setup, and credential management. Installed locally via `install.sh`, also used by the GitHub Action.\n\n## Build\n\n```sh\ngo build ./cmd/review # builds codecanary\n```\n\nVersion is set via ldflags: `-X main.version=v{version}`\n\n## Lint\n\n```sh\ngolangci-lint run ./...\n```\n\nAll code must pass `golangci-lint` with default linters (errcheck, staticcheck, etc.). Run this before committing.\n\n## Key dependencies\n\n- `spf13/cobra` — CLI framework\n- `charmbracelet/huh` — terminal form builder (setup wizard)\n- `zalando/go-keyring` — OS keychain (with file-based fallback for systems without one)\n- `bmatcuk/doublestar` — glob pattern matching for ignore rules\n- `gopkg.in/yaml.v3` — config parsing\n- `golang.org/x/term` — terminal detection\n\n## Architecture\n\n### Core principle: adapters keep the engine agnostic\n\nThe review engine (`runner.go`) is provider- and platform-agnostic. It depends only on two interfaces — never on concrete GitHub APIs, LLM SDKs, or environment-specific logic. All environment and provider specifics live behind adapters.\n\n### Provider layer — `ModelProvider` interface (`provider.go`)\n\nAbstracts LLM invocations. The core engine calls `provider.Run(ctx, prompt, opts)` and gets back text + usage metadata. It never knows which LLM backend is being used.\n\n**Implementations**: `anthropic`, `openai`, `openrouter`, `claude` (CLI).\n**Selection**: factory registry in `provider.go` — `NewProviderForRole(mc, env)` returns the right implementation based on `mc.Provider`.\n\nAdding a new LLM provider means: create `provider_\u003cname\u003e.go` and register a `ProviderFactory` (constructor, validation, pricing, default models) via `init()`.\n\n### Platform layer — `ReviewPlatform` interface (`platform.go`)\n\nAbstracts environment-specific operations: loading previous findings, publishing results, saving state, resolving threads, reporting usage.\n\n**Implementations**: `GithubPlatform` (posts to PRs, reads threads via API), `LocalPlatform` (prints to terminal, persists state to `~/.codecanary/repos/\u003cowner\u003e/\u003crepo\u003e/state/\u003cbranch\u003e.json` — per-repo scoping keeps branch names like `main` from colliding across repos; falls back to `~/.codecanary/state/\u003cbranch\u003e.json` when no git remote is resolvable).\n\nRouting is strict: `codecanary review --post` → `GithubPlatform`; `codecanary review` (no `--post`) → `LocalPlatform`, even when the branch has an open PR. Local is local — the branch diff (with uncommitted changes) is reviewed against the default base, previous findings come from local state, nothing is fetched from or posted to GitHub. The old `GithubPlatform`-with-`Post=false` hybrid is gone; it was the source of the \"state written locally, read from GitHub\" asymmetry that kept breaking incremental local reviews.\n\nAdding a new platform (e.g., GitLab) means: implement `ReviewPlatform`, wire it in the CLI.\n\n### Unified review pipeline (`runner.go`)\n\nThere is a **single `Run()` function** — not separate paths for GitHub vs. local. The pipeline is:\n\n1. Fetch PR data (or local diff)\n2. Load config, project docs, file contents\n3. Create providers via `NewProviderForRole()` (factory, provider-agnostic)\n4. Load previous findings via `platform.LoadPreviousFindings()`\n5. If incremental: triage threads, evaluate via provider, handle resolutions\n6. Build and execute main review prompt\n7. Parse findings, filter non-actionable\n8. `platform.Publish()` → `platform.SaveState()` → `platform.ReportUsage()`\n\n### Other architecture notes\n\n- **Config** is split across two files in `.codecanary/`: `config.yml` (provider, models, budgets, timeouts) and `review.yml` (rules, context, ignore patterns). `review.yml` is optional — if present, its fields override rules/context/ignore in `config.yml`. A personal `review.local.yml` can add rules, context, and ignore patterns on top of `review.yml` (append semantics, not replacement). Legacy `.codecanary.yml` at repo root is still supported with a deprecation warning.\n- **Incremental reviews**: on re-push, triage existing threads (Go-driven classifier in `triage.go`), evaluate changed threads via provider (triage model), then review only new code\n- **Dual marker detection**: reads both `codecanary:review` and legacy `clanopy:review` HTML markers for backward compatibility\n- **Anti-hallucination**: explicit file allowlist, line validation against diff, max finding distance threshold\n- **OIDC worker** (`oidc/worker/`): OIDC token exchange proxy at `oidc.codecanary.sh` — verifies GitHub Actions OIDC token, returns GitHub App installation token\n- **Telemetry worker** (`telemetry/worker/`): anonymous usage ingest at `telemetry.codecanary.sh` — writes to a Cloudflare Analytics Engine dataset\n- **Telemetry dashboard** (`telemetry/dashboard/`): internal analytics view at `dashboard.codecanary.sh` — Cloudflare Pages + Pages Functions, gated by Cloudflare Access. Reads the AE dataset via the SQL HTTP API with a read-only API token (no AE binding, so writes are platform-impossible)\n- **Setup** is a subcommand (`codecanary setup`) using `charmbracelet/huh` forms, with `local` and `github` sub-flows\n- **Credentials** use a single env var `CODECANARY_PROVIDER_SECRET` for all providers. Stored via `go-keyring` (OS keychain) with a file-based fallback (`~/.codecanary/credentials.json`, mode `0600`). `resolveEnv()` in `runner.go` injects the stored credential into the filtered env when not already set.\n\n## Rules\n\n- **Keep the core engine agnostic.** `runner.go`, `triage.go`, `prompt.go`, `findings.go` must never import or reference a specific LLM provider or platform. All provider/platform specifics go behind the `ModelProvider` or `ReviewPlatform` interfaces. No `if provider == \"openai\"` in core logic.\n- **Use the adapter/provider pattern for new integrations.** New LLM backends → create `provider_\u003cname\u003e.go` with a `ProviderFactory` registration in `init()`. New deployment targets → implement `ReviewPlatform` + wire in CLI. Never fork the pipeline.\n- **One pipeline, not two.** There must be a single `Run()` path. GitHub and local modes differ only in which `ReviewPlatform` implementation is injected — the orchestration logic is shared.\n- **Shared types for similar providers.** OpenAI-compatible APIs share request/response types via `provider_compat.go`. Don't duplicate HTTP client logic across providers.\n- **Don't repeat yourself.** Before writing new code, search the codebase for existing functions, mappings, or logic that already does what you need — then call it instead of reimplementing it. This applies to everything: switch statements, helper functions, validation logic, data mappings, HTTP calls. One source of truth, callers import it. Don't merge scaffolding or unused exports — if it's not called yet, it doesn't ship yet.\n- **File names are ownership boundaries.** A function defined in `local.go` implies it belongs to the local flow; one in `github.go` implies it belongs to GitHub. If a function is called by multiple files in the same package, it belongs in a shared file (e.g., `forms.go` for setup helpers, `platform.go` for platform-shared logic). Never define shared infrastructure in a flow-specific file — move it to the file that matches its actual scope.\n- **Canonical provider registration points.** Provider names live in the factory map in `provider.go`. All providers use `CODECANARY_PROVIDER_SECRET` for credentials (defined in `internal/credentials/keyring.go`). When adding a new provider, register a `ProviderFactory` in `provider.go` and add config validation in `config.go`.\n- **Minimize shell code.** `install.sh` and the GitHub Action (`action.yml`) should be kept as thin as possible. All logic must live in Go.\n- **Workflow template is embedded.** `internal/setup/codecanary.yml` is the single source of truth for the GitHub Actions workflow, embedded via `//go:embed`. `.github/workflows/codecanary.yml` must be identical — `go test ./internal/setup/` enforces this. When changing the workflow, edit either file and copy to the other.\n- **Claude skills are embedded.** `internal/skills/codecanary-fix/SKILL.md` is the single source of truth for the codecanary-fix skill, embedded via `//go:embed` and materialized by `codecanary install-skill`. `.claude/skills/codecanary-fix/SKILL.md` must be identical so Claude Code's project-mode discovery finds it when working in this repo — `go test ./internal/skills/` enforces this. When changing the skill, edit either file and copy to the other.\n- **Keep `docs/review-flow.md` in sync.** This document describes the full review pipeline — every step, the triage flow, platform differences, and key design decisions. When changing `runner.go`, `triage.go`, `prompt.go`, `findings.go`, `github.go`, `local.go`, `platform.go`, or the `ReviewPlatform` implementations, update the doc to reflect the new behavior.\n- Tests exist for config, findings, formatting, and triage. Be careful with refactors — run `go test ./...` and `go vet ./...`.\n" + } +} diff --git a/internal/review/testdata/corpus/alansikora-codecanary-pr175.prompt.golden b/internal/review/testdata/corpus/alansikora-codecanary-pr175.prompt.golden new file mode 100644 index 0000000..2d69ffa --- /dev/null +++ b/internal/review/testdata/corpus/alansikora-codecanary-pr175.prompt.golden @@ -0,0 +1,1860 @@ +You are a code reviewer. Review the following pull request and report findings. +You will be given the full contents of changed files for context, along with the diff. Only report issues that are directly related to the changes in the diff — do not flag pre-existing issues in unchanged code. Do not report a finding if your analysis concludes that the code is correct and no action is needed — only report findings that require the author to make a change or consider a specific alternative. +Also consider whether the changes could cause side effects in other files that depend on or interact with the modified code (e.g. callers, importers, shared state). If you identify a potential side effect, anchor your finding to the relevant line in the diff and describe the affected downstream code in the description. + +## Pull Request #175 +fix(triage): include PreviouslyAcked in countNonSkipped +**Author:** alansikora + +Follow-up to #174. + +## Summary + +The first sticky-ack fix worked at the triage layer (logs showed `[sticky] ... carried forward`) but the published summary still said **Still unresolved: 2** — confirmed in production on thetechfx/bedrock#1671 right after v0.6.23 shipped. + +Root cause: `countNonSkipped` excluded `TriagePreviouslyAcked`. When *every* remaining thread is sticky-ack, `needsEval` came back as zero, the runner skipped `EvaluateThreadsParallel` entirely, the fast path I added there never fired, no `fixedThread` was emitted, and `computeReviewSummary` lumped the threads into `StillOpen`. + +Including PreviouslyAcked in `countNonSkipped` restores the round trip: eval runs, the fast path returns the prior reason without an LLM call, `toFixedThreads` collects the resolutions, and the summary lands them in Acknowledged/Rebutted/Dismissed. + +## Test plan + +- [x] `go test ./...` +- [x] New regression test `TestCountNonSkipped_IncludesPreviouslyAcked` +- [ ] Re-trigger a review on bedrock#1671 with v0.6.24 to confirm the summary flips to "Acknowledged by author: 2" + + +## Project Documentation +The following project documentation describes conventions and standards for this codebase. Use these to inform your review — flag violations of these conventions when relevant. + + +# CodeCanary + +AI-powered code review for GitHub pull requests. + +## Project structure + +``` +cmd/ + review/ # Main binary — review CLI + setup wizard + main.go # Entry point + cli/ # Cobra commands + root.go # Root "codecanary" command + review.go # codecanary review + findings.go # codecanary findings — fetch bot findings for the review skill + reply.go # codecanary reply --url --body — post a reply on a review thread + install_skill.go # codecanary install-skill — write embedded Claude skill to disk + setup.go # codecanary setup [local|github] + auth.go # codecanary auth [status|delete] +internal/ + review/ + runner.go # Core review pipeline — single Run() entry point + config.go # Config loading, validation, defaults + # Provider layer (LLM abstraction) + provider.go # ModelProvider interface + factory registry + provider_anthropic.go + provider_openai.go + provider_openrouter.go + provider_claude.go # Claude CLI wrapper + provider_compat.go # Shared types for OpenAI-compatible APIs + pricing.go # Token-based cost estimation + # Platform layer (environment abstraction) + platform.go # ReviewPlatform interface + platform_github.go # GitHub Actions implementation + platform_local.go # Local CLI implementation + # Supporting modules + prompt.go # Prompt building (review, incremental, per-thread) + findings.go # Finding parsing, filtering, result structures + triage.go # Thread classification + parallel LLM evaluation + formatter.go # JSON/Markdown/Terminal output formatting + usage.go # Token tracking, budget checking + github.go # GitHub API calls (fetch threads, post reviews) + comments.go # PR review comment fetch + finding marker parser + review-check watcher + local.go # Local diff & git operations + state.go # Local state persistence + docs.go # Project doc discovery + credentials/ # Credential storage (keychain with file fallback) + keyring.go # Store/Retrieve/Delete — keychain first, ~/.codecanary/credentials.json fallback + skills/ # Claude Code skills embedded in the binary via //go:embed + skills.go # Exports CodecanaryFix() returning the skill body + codecanary-fix/SKILL.md # Canonical skill source (duplicated at .claude/skills/codecanary-fix/SKILL.md; parity enforced by skills_test.go) + setup/ # Setup wizard logic (huh forms) + forms.go # Shared huh form components + validate.go # API key validation via test calls + guidance.go # Token/permissions guidance text + workflow.go # GitHub Actions workflow template + local.go # RunLocal() — local setup flow + github.go # RunGitHub() — GitHub Actions setup flow + auth/ # OAuth PKCE flow, GitHub App installation +telemetry/ # Telemetry domain (anonymous usage analytics) + worker/ # Cloudflare Worker — telemetry ingestion (TypeScript) + dashboard/ # Cloudflare Pages — internal analytics dashboard (vanilla JS + Chart.js) +oidc/ # OIDC domain + worker/ # Cloudflare Worker — OIDC token exchange proxy (TypeScript) +action.yml # GitHub Action definition (composite action) +install.sh # Downloads and installs codecanary binary permanently +.claude/ + skills/ + codecanary-fix/ # Claude Code skill — drives review→fix→push loop using `codecanary findings` + `codecanary reply` +``` + +## Binary + +- **`codecanary`** — single binary for reviews, setup, and credential management. Installed locally via `install.sh`, also used by the GitHub Action. + +## Build + +```sh +go build ./cmd/review # builds codecanary +``` + +Version is set via ldflags: `-X main.version=v{version}` + +## Lint + +```sh +golangci-lint run ./... +``` + +All code must pass `golangci-lint` with default linters (errcheck, staticcheck, etc.). Run this before committing. + +## Key dependencies + +- `spf13/cobra` — CLI framework +- `charmbracelet/huh` — terminal form builder (setup wizard) +- `zalando/go-keyring` — OS keychain (with file-based fallback for systems without one) +- `bmatcuk/doublestar` — glob pattern matching for ignore rules +- `gopkg.in/yaml.v3` — config parsing +- `golang.org/x/term` — terminal detection + +## Architecture + +### Core principle: adapters keep the engine agnostic + +The review engine (`runner.go`) is provider- and platform-agnostic. It depends only on two interfaces — never on concrete GitHub APIs, LLM SDKs, or environment-specific logic. All environment and provider specifics live behind adapters. + +### Provider layer — `ModelProvider` interface (`provider.go`) + +Abstracts LLM invocations. The core engine calls `provider.Run(ctx, prompt, opts)` and gets back text + usage metadata. It never knows which LLM backend is being used. + +**Implementations**: `anthropic`, `openai`, `openrouter`, `claude` (CLI). +**Selection**: factory registry in `provider.go` — `NewProviderForRole(mc, env)` returns the right implementation based on `mc.Provider`. + +Adding a new LLM provider means: create `provider_.go` and register a `ProviderFactory` (constructor, validation, pricing, default models) via `init()`. + +### Platform layer — `ReviewPlatform` interface (`platform.go`) + +Abstracts environment-specific operations: loading previous findings, publishing results, saving state, resolving threads, reporting usage. + +**Implementations**: `GithubPlatform` (posts to PRs, reads threads via API), `LocalPlatform` (prints to terminal, persists state to `~/.codecanary/repos///state/.json` — per-repo scoping keeps branch names like `main` from colliding across repos; falls back to `~/.codecanary/state/.json` when no git remote is resolvable). + +Routing is strict: `codecanary review --post` → `GithubPlatform`; `codecanary review` (no `--post`) → `LocalPlatform`, even when the branch has an open PR. Local is local — the branch diff (with uncommitted changes) is reviewed against the default base, previous findings come from local state, nothing is fetched from or posted to GitHub. The old `GithubPlatform`-with-`Post=false` hybrid is gone; it was the source of the "state written locally, read from GitHub" asymmetry that kept breaking incremental local reviews. + +Adding a new platform (e.g., GitLab) means: implement `ReviewPlatform`, wire it in the CLI. + +### Unified review pipeline (`runner.go`) + +There is a **single `Run()` function** — not separate paths for GitHub vs. local. The pipeline is: + +1. Fetch PR data (or local diff) +2. Load config, project docs, file contents +3. Create providers via `NewProviderForRole()` (factory, provider-agnostic) +4. Load previous findings via `platform.LoadPreviousFindings()` +5. If incremental: triage threads, evaluate via provider, handle resolutions +6. Build and execute main review prompt +7. Parse findings, filter non-actionable +8. `platform.Publish()` → `platform.SaveState()` → `platform.ReportUsage()` + +### Other architecture notes + +- **Config** is split across two files in `.codecanary/`: `config.yml` (provider, models, budgets, timeouts) and `review.yml` (rules, context, ignore patterns). `review.yml` is optional — if present, its fields override rules/context/ignore in `config.yml`. A personal `review.local.yml` can add rules, context, and ignore patterns on top of `review.yml` (append semantics, not replacement). Legacy `.codecanary.yml` at repo root is still supported with a deprecation warning. +- **Incremental reviews**: on re-push, triage existing threads (Go-driven classifier in `triage.go`), evaluate changed threads via provider (triage model), then review only new code +- **Dual marker detection**: reads both `codecanary:review` and legacy `clanopy:review` HTML markers for backward compatibility +- **Anti-hallucination**: explicit file allowlist, line validation against diff, max finding distance threshold +- **OIDC worker** (`oidc/worker/`): OIDC token exchange proxy at `oidc.codecanary.sh` — verifies GitHub Actions OIDC token, returns GitHub App installation token +- **Telemetry worker** (`telemetry/worker/`): anonymous usage ingest at `telemetry.codecanary.sh` — writes to a Cloudflare Analytics Engine dataset +- **Telemetry dashboard** (`telemetry/dashboard/`): internal analytics view at `dashboard.codecanary.sh` — Cloudflare Pages + Pages Functions, gated by Cloudflare Access. Reads the AE dataset via the SQL HTTP API with a read-only API token (no AE binding, so writes are platform-impossible) +- **Setup** is a subcommand (`codecanary setup`) using `charmbracelet/huh` forms, with `local` and `github` sub-flows +- **Credentials** use a single env var `CODECANARY_PROVIDER_SECRET` for all providers. Stored via `go-keyring` (OS keychain) with a file-based fallback (`~/.codecanary/credentials.json`, mode `0600`). `resolveEnv()` in `runner.go` injects the stored credential into the filtered env when not already set. + +## Rules + +- **Keep the core engine agnostic.** `runner.go`, `triage.go`, `prompt.go`, `findings.go` must never import or reference a specific LLM provider or platform. All provider/platform specifics go behind the `ModelProvider` or `ReviewPlatform` interfaces. No `if provider == "openai"` in core logic. +- **Use the adapter/provider pattern for new integrations.** New LLM backends → create `provider_.go` with a `ProviderFactory` registration in `init()`. New deployment targets → implement `ReviewPlatform` + wire in CLI. Never fork the pipeline. +- **One pipeline, not two.** There must be a single `Run()` path. GitHub and local modes differ only in which `ReviewPlatform` implementation is injected — the orchestration logic is shared. +- **Shared types for similar providers.** OpenAI-compatible APIs share request/response types via `provider_compat.go`. Don't duplicate HTTP client logic across providers. +- **Don't repeat yourself.** Before writing new code, search the codebase for existing functions, mappings, or logic that already does what you need — then call it instead of reimplementing it. This applies to everything: switch statements, helper functions, validation logic, data mappings, HTTP calls. One source of truth, callers import it. Don't merge scaffolding or unused exports — if it's not called yet, it doesn't ship yet. +- **File names are ownership boundaries.** A function defined in `local.go` implies it belongs to the local flow; one in `github.go` implies it belongs to GitHub. If a function is called by multiple files in the same package, it belongs in a shared file (e.g., `forms.go` for setup helpers, `platform.go` for platform-shared logic). Never define shared infrastructure in a flow-specific file — move it to the file that matches its actual scope. +- **Canonical provider registration points.** Provider names live in the factory map in `provider.go`. All providers use `CODECANARY_PROVIDER_SECRET` for credentials (defined in `internal/credentials/keyring.go`). When adding a new provider, register a `ProviderFactory` in `provider.go` and add config validation in `config.go`. +- **Minimize shell code.** `install.sh` and the GitHub Action (`action.yml`) should be kept as thin as possible. All logic must live in Go. +- **Workflow template is embedded.** `internal/setup/codecanary.yml` is the single source of truth for the GitHub Actions workflow, embedded via `//go:embed`. `.github/workflows/codecanary.yml` must be identical — `go test ./internal/setup/` enforces this. When changing the workflow, edit either file and copy to the other. +- **Claude skills are embedded.** `internal/skills/codecanary-fix/SKILL.md` is the single source of truth for the codecanary-fix skill, embedded via `//go:embed` and materialized by `codecanary install-skill`. `.claude/skills/codecanary-fix/SKILL.md` must be identical so Claude Code's project-mode discovery finds it when working in this repo — `go test ./internal/skills/` enforces this. When changing the skill, edit either file and copy to the other. +- **Keep `docs/review-flow.md` in sync.** This document describes the full review pipeline — every step, the triage flow, platform differences, and key design decisions. When changing `runner.go`, `triage.go`, `prompt.go`, `findings.go`, `github.go`, `local.go`, `platform.go`, or the `ReviewPlatform` implementations, update the doc to reflect the new behavior. +- Tests exist for config, findings, formatting, and triage. Be careful with refactors — run `go test ./...` and `go vet ./...`. + + + +## Review Rules +No specific rules are defined. Perform a general code review covering correctness, security, performance, and maintainability. + +## Files in This Diff +The following files — and ONLY these files — are part of this diff. Every finding you report MUST reference one of these exact paths. Do NOT reference any file that is not in this list. + +- `internal/review/triage.go` +- `internal/review/triage_test.go` + +## Changed File Contents +Below are the full contents of changed files. Use these to understand surrounding code, types, imports, and control flow. Do NOT report findings on unchanged code — only flag issues directly related to changes in the diff. + +### `internal/review/triage.go` +```` +1: package review +2: +3: import ( +4: "context" +5: "encoding/json" +6: "fmt" +7: "os" +8: "sort" +9: "strconv" +10: "strings" +11: "sync" +12: ) +13: +14: // ThreadClassification is the triage result for a single unresolved thread. +15: type ThreadClassification int +16: +17: const ( +18: TriageSkip ThreadClassification = iota // no code changes at all +19: TriageCodeChanged // diff touches finding location (outdated) +20: TriageHasReply // thread has human replies +21: TriageCodeChangedReply // both code changed AND has replies +22: TriageCrossFileChange // diff has changes but NOT in this thread's file +23: TriageFileRemovedFromPR // file no longer in the PR +24: TriagePreviouslyAcked // bot already ack'd a deferral; no new human reply since +25: ) +26: +27: // TriagedThread pairs a ReviewThread with its classification and context. +28: type TriagedThread struct { +29: Thread ReviewThread +30: Index int // original index in the unresolved slice +31: Class ThreadClassification +32: FileDiff string // file-scoped diff (level 1) or full diff (cross-file) +33: FullDiff string // full PR diff for widened-scope fallback; empty when not applicable +34: FileSnippet string // windowed file content around finding + diff hunks +35: BotLogin string // login of the review bot, for filtering replies +36: PriorAckReason string // for TriagePreviouslyAcked: the previously-recorded reason (acknowledged/rebutted/dismissed) +37: } +38: +39: // ThreadResolution is the result of a per-thread Claude evaluation. +40: type ThreadResolution struct { +41: Index int +42: Resolved bool +43: Reason string // "code_change", "acknowledged", "rebutted", "dismissed" +44: Rationale string // one-sentence narrative from the evaluator, surfaced in the incremental review to prevent ping-ponging +45: Error error +46: } +47: +48: // ExtractFileDiff extracts all diff hunks for a specific file from a unified diff. +49: func ExtractFileDiff(fullDiff, filePath string) string { +50: lines := strings.Split(fullDiff, "\n") +51: var result []string +52: capturing := false +53: +54: for i := 0; i < len(lines); i++ { +55: // Detect start of a new file in the diff. +56: if strings.HasPrefix(lines[i], "diff --git") { +57: if capturing { +58: // We were capturing — new file starts, stop. +59: break +60: } +61: // Search ahead for the "+++ b/" header line. +62: // The offset varies (index line, mode lines, etc.) so scan +63: // forward instead of using a fixed offset. +64: for j := i + 1; j < len(lines) && !strings.HasPrefix(lines[j], "diff --git"); j++ { +65: if strings.HasPrefix(lines[j], "+++ b/"+filePath) { +66: capturing = true +67: result = append(result, lines[i]) +68: break +69: } +70: } +71: continue +72: } +73: if capturing { +74: result = append(result, lines[i]) +75: } +76: } +77: +78: return strings.Join(result, "\n") +79: } +80: +81: // lineRange represents an inclusive range of 1-based line numbers. +82: type lineRange struct{ start, end int } +83: +84: // parseHunkNewRanges extracts the new-file line ranges from unified diff hunk headers. +85: // Each @@ -X,Y +N,M @@ header yields a range [N, N+M-1]. +86: func parseHunkNewRanges(diffText string) []lineRange { +87: var ranges []lineRange +88: for _, line := range strings.Split(diffText, "\n") { +89: if !strings.HasPrefix(line, "@@ ") { +90: continue +91: } +92: // Find +N or +N,M in the hunk header. +93: idx := strings.Index(line, "+") +94: if idx < 0 { +95: continue +96: } +97: rest := line[idx+1:] +98: // Trim everything after the space/comma/@@ that ends the range spec. +99: if sp := strings.IndexAny(rest, " @"); sp >= 0 { +100: rest = rest[:sp] +101: } +102: parts := strings.SplitN(rest, ",", 2) +103: start, err := strconv.Atoi(parts[0]) +104: if err != nil { +105: continue +106: } +107: count := 1 +108: if len(parts) == 2 { +109: if c, err := strconv.Atoi(parts[1]); err == nil { +110: if c == 0 { +111: continue // pure deletion hunk — no new lines +112: } +113: if c > 0 { +114: count = c +115: } +116: } +117: } +118: ranges = append(ranges, lineRange{start: start, end: start + count - 1}) +119: } +120: return ranges +121: } +122: +123: // mergeRanges merges overlapping or adjacent line ranges, sorted by start. +124: func mergeRanges(ranges []lineRange) []lineRange { +125: if len(ranges) == 0 { +126: return nil +127: } +128: sort.Slice(ranges, func(i, j int) bool { return ranges[i].start < ranges[j].start }) +129: merged := []lineRange{ranges[0]} +130: for _, r := range ranges[1:] { +131: last := &merged[len(merged)-1] +132: if r.start <= last.end+1 { +133: if r.end > last.end { +134: last.end = r.end +135: } +136: } else { +137: merged = append(merged, r) +138: } +139: } +140: return merged +141: } +142: +143: // ExtractFileSnippet extracts a windowed snippet from file content centered around +144: // the finding line and expanded to cover diff hunk ranges. Returns an empty string +145: // if content is empty. findingLine is 1-based. diffText is the file-scoped diff +146: // (used to parse hunk ranges; may be empty for cross-file cases). maxLines caps the +147: // total snippet length. +148: func ExtractFileSnippet(content string, findingLine int, diffText string, maxLines int) string { +149: if content == "" { +150: return "" +151: } +152: lines := strings.Split(content, "\n") +153: totalLines := len(lines) +154: +155: // Build interesting ranges: finding line ± 50, each hunk range ± 30. +156: const findingPad = 50 +157: const hunkPad = 30 +158: +159: var ranges []lineRange +160: if findingLine > 0 { +161: ranges = append(ranges, lineRange{ +162: start: max(1, findingLine-findingPad), +163: end: min(totalLines, findingLine+findingPad), +164: }) +165: } +166: for _, hr := range parseHunkNewRanges(diffText) { +167: ranges = append(ranges, lineRange{ +168: start: max(1, hr.start-hunkPad), +169: end: min(totalLines, hr.end+hunkPad), +170: }) +171: } +172: if len(ranges) == 0 { +173: // Fallback: center on line 1 if nothing else. +174: ranges = append(ranges, lineRange{start: 1, end: min(totalLines, maxLines)}) +175: } +176: +177: merged := mergeRanges(ranges) +178: +179: // If total exceeds maxLines, prioritize the range containing the finding line, +180: // then include other ranges in order until the budget is exhausted. +181: total := 0 +182: for _, r := range merged { +183: total += r.end - r.start + 1 +184: } +185: if total > maxLines { +186: // Find which merged range contains the finding line. +187: // When findingLine is invalid (<= 0), default to the first range. +188: findingIdx := 0 +189: if findingLine > 0 { +190: for i, r := range merged { +191: if findingLine >= r.start && findingLine <= r.end { +192: findingIdx = i +193: break +194: } +195: } +196: } +197: // Start with the finding range, then add others. +198: budget := maxLines +199: kept := make([]bool, len(merged)) +200: kept[findingIdx] = true +201: size := merged[findingIdx].end - merged[findingIdx].start + 1 +202: if size > budget && findingLine > 0 { +203: // Truncate the finding range around the finding line. +204: half := budget / 2 +205: merged[findingIdx] = lineRange{ +206: start: max(1, findingLine-half), +207: end: min(totalLines, findingLine+half), +208: } +209: size = merged[findingIdx].end - merged[findingIdx].start + 1 +210: } else if size > budget { +211: // No valid finding line — just take the first maxLines of the range. +212: merged[findingIdx] = lineRange{ +213: start: merged[findingIdx].start, +214: end: min(totalLines, merged[findingIdx].start+budget-1), +215: } +216: size = merged[findingIdx].end - merged[findingIdx].start + 1 +217: } +218: budget -= size +219: for i, r := range merged { +220: if kept[i] { +221: continue +222: } +223: rSize := r.end - r.start + 1 +224: if rSize <= budget { +225: kept[i] = true +226: budget -= rSize +227: } +228: } +229: var trimmed []lineRange +230: for i, r := range merged { +231: if kept[i] { +232: trimmed = append(trimmed, r) +233: } +234: } +235: merged = trimmed +236: } +237: +238: // Format with line numbers, inserting omission markers between gaps. +239: var b strings.Builder +240: for i, r := range merged { +241: if i > 0 { +242: gap := r.start - merged[i-1].end - 1 +243: fmt.Fprintf(&b, "... (%d lines omitted) ...\n", gap) +244: } +245: for ln := r.start; ln <= r.end && ln <= totalLines; ln++ { +246: fmt.Fprintf(&b, "%d: %s\n", ln, lines[ln-1]) +247: } +248: } +249: return b.String() +250: } +251: +252: // ClassifyThreads triages unresolved threads using GitHub's outdated flag and reply presence. +253: // +254: // Two diffs serve different purposes: +255: // - activityDiff: the incremental diff (changes since last review). Used to decide whether +256: // to skip evaluation — when empty, non-outdated threads with no replies are TriageSkip. +257: // - contextDiff: the full PR diff (all changes). Used to determine the correct classification +258: // (TriageCodeChanged vs TriageCrossFileChange) and to extract FileDiff/FileSnippet for the +259: // evaluation prompt. This ensures fixes from earlier pushes are visible to the evaluator. +260: // +261: // prFiles is the current set of files in the PR; threads on files no longer in the PR +262: // are classified as TriageFileRemovedFromPR and auto-resolved without an LLM call. +263: // fileContents provides current file contents for building context snippets in triage prompts. +264: func ClassifyThreads(threads []ReviewThread, activityDiff, contextDiff, botLogin string, prFiles []string, fileContents map[string]string) []TriagedThread { +265: prFileSet := make(map[string]bool, len(prFiles)) +266: for _, f := range prFiles { +267: prFileSet[f] = true +268: } +269: +270: result := make([]TriagedThread, len(threads)) +271: +272: for i, t := range threads { +273: // File no longer in the PR — auto-resolve without LLM. +274: // When prFiles is empty (e.g. upstream fetch returned no file list), +275: // we skip this check to avoid incorrectly resolving all threads. +276: if len(prFileSet) > 0 && !prFileSet[t.Path] { +277: result[i] = TriagedThread{ +278: Thread: t, +279: Index: i, +280: Class: TriageFileRemovedFromPR, +281: } +282: continue +283: } +284: +285: hasReply := hasNewHumanReply(t, botLogin) +286: outdated := t.Outdated +287: deleted := fileDeletedInDiff(contextDiff, t.Path) +288: +289: // Sticky-ack: when the bot already recorded a deferral and the +290: // author has not added a fresh reply since, preserve the prior +291: // classification instead of rerunning triage. This prevents the +292: // "Acknowledged by author: 2" → "Still unresolved: 2" regression +293: // that happened on subsequent pushes when no new signal was present. +294: var priorAck string +295: if !hasReply { +296: priorAck = parsePriorAckReason(t, botLogin) +297: } +298: +299: var class ThreadClassification +300: switch { +301: case priorAck != "": +302: class = TriagePreviouslyAcked +303: case deleted: +304: // File was deleted — evaluate with full diff so Claude can check +305: // whether the code moved to a replacement file with the fix applied. +306: class = TriageCrossFileChange +307: case outdated && hasReply: +308: class = TriageCodeChangedReply +309: case outdated: +310: class = TriageCodeChanged +311: case hasReply: +312: class = TriageHasReply +313: default: +314: // No GitHub outdated flag, no replies. Use the incremental diff to +315: // decide whether there is new activity worth evaluating. +316: if activityDiff == "" { +317: // No changes since last review — nothing new to evaluate. +318: class = TriageSkip +319: } else if fileInDiff(contextDiff, t.Path) { +320: // File was changed in the PR — classify as code-changed so the +321: // evaluator sees the file-scoped diff (which may include fixes +322: // from earlier pushes that the incremental diff missed). +323: class = TriageCodeChanged +324: } else { +325: // Code changed in other files — evaluate in case the fix is cross-file. +326: class = TriageCrossFileChange +327: } +328: } +329: +330: // Use file-scoped diff for same-file evaluations to reduce noise. +331: // The full PR diff can drown out the relevant fix with changes from +332: // unrelated files, causing false "not resolved" verdicts. Cross-file +333: // evaluations keep the full diff since the fix is in a different file. +334: var fileDiff, fullDiff string +335: switch class { +336: case TriageCodeChanged, TriageCodeChangedReply: +337: fileDiff = ExtractFileDiff(contextDiff, t.Path) +338: fullDiff = contextDiff // retained for widened-scope fallback +339: case TriageCrossFileChange: +340: fileDiff = contextDiff +341: } +342: +343: // Build a windowed file snippet for code-change evaluations. +344: // Reuse fileDiff (already file-scoped) for snippet range calculation +345: // so hunk line numbers from other files don't pull in wrong sections. +346: var fileSnippet string +347: if content, ok := fileContents[t.Path]; ok { +348: switch class { +349: case TriageCodeChanged, TriageCodeChangedReply: +350: fileSnippet = ExtractFileSnippet(content, t.Line, fileDiff, 300) +351: case TriageCrossFileChange: +352: // Show finding's file context even though the diff is in other files. +353: fileSnippet = ExtractFileSnippet(content, t.Line, "", 200) +354: } +355: } +356: +357: result[i] = TriagedThread{ +358: Thread: t, +359: Index: i, +360: Class: class, +361: FileDiff: fileDiff, +362: FullDiff: fullDiff, +363: FileSnippet: fileSnippet, +364: BotLogin: botLogin, +365: PriorAckReason: priorAck, +366: } +367: } +368: +369: return result +370: } +371: +372: // parsePriorAckReason returns the reason recorded in the most recent +373: // codecanary ack reply on the thread (acknowledged/rebutted/dismissed), +374: // or "" if no ack reply is present. Handles both the current marker +375: // () and the legacy clanopy form. +376: func parsePriorAckReason(t ReviewThread, botLogin string) string { +377: var latest string +378: for _, r := range t.Replies { +379: if r.Author != botLogin { +380: continue +381: } +382: if reason := extractAckReason(r.Body); reason != "" { +383: latest = reason +384: } +385: } +386: return latest +387: } +388: +389: // extractAckReason pulls the reason value out of a codecanary ack marker. +390: // Returns "" when the body has no marker. Unknown reasons are returned +391: // as-is so callers can decide how to handle them; in practice the bot +392: // only writes "acknowledged", "rebutted", "dismissed", or "unknown". +393: func extractAckReason(body string) string { +394: for _, prefix := range []string{ackMarkerPrefix, legacyAckPrefix} { +395: idx := strings.Index(body, prefix) +396: if idx < 0 { +397: continue +398: } +399: rest := body[idx+len(prefix):] +400: end := strings.Index(rest, " -->") +401: if end < 0 { +402: end = strings.Index(rest, "-->") +403: if end < 0 { +404: continue +405: } +406: } +407: return strings.TrimSpace(rest[:end]) +408: } +409: return "" +410: } +411: +412: // fileInDiff checks if the diff contains changes to the given file path. +413: func fileInDiff(diff, path string) bool { +414: target := "+++ b/" + path +415: return strings.Contains(diff, target+"\n") || strings.Contains(diff, target+"\t") || strings.HasSuffix(diff, target) +416: } +417: +418: // fileDeletedInDiff checks if the diff shows the given file was deleted. +419: // A deleted file has "--- a/" followed by "+++ /dev/null". +420: func fileDeletedInDiff(diff, path string) bool { +421: marker := "--- a/" + path +422: idx := strings.Index(diff, marker) +423: if idx < 0 { +424: return false +425: } +426: // Ensure full path match (not a prefix of a longer filename). +427: rest := diff[idx+len(marker):] +428: if len(rest) > 0 && rest[0] != '\n' && rest[0] != '\r' { +429: return false +430: } +431: nl := strings.Index(rest, "\n") +432: if nl < 0 { +433: return false +434: } +435: nextLine := "" +436: rest = rest[nl+1:] +437: if eol := strings.Index(rest, "\n"); eol >= 0 { +438: nextLine = rest[:eol] +439: } else { +440: nextLine = rest +441: } +442: return nextLine == "+++ /dev/null" +443: } +444: +445: // hasHumanReply checks if a thread has at least one reply from a non-bot author. +446: func hasHumanReply(t ReviewThread, botLogin string) bool { +447: for _, r := range t.Replies { +448: if r.Author != botLogin { +449: return true +450: } +451: } +452: return false +453: } +454: +455: // isAckReply checks if a reply body contains an acknowledgment marker. +456: func isAckReply(body string) bool { +457: return strings.Contains(body, ackMarkerPrefix) || strings.Contains(body, legacyAckPrefix) +458: } +459: +460: // hasNewHumanReply checks if a thread has a human reply AFTER the last +461: // ack reply. If no ack reply exists, it falls back to hasHumanReply behavior. +462: // Replies are in chronological order. +463: func hasNewHumanReply(t ReviewThread, botLogin string) bool { +464: lastAckIdx := -1 +465: for i, r := range t.Replies { +466: if r.Author == botLogin && isAckReply(r.Body) { +467: lastAckIdx = i +468: } +469: } +470: if lastAckIdx == -1 { +471: // No ack reply exists — fall back to standard check. +472: return hasHumanReply(t, botLogin) +473: } +474: // Check for human replies after the last ack. +475: for _, r := range t.Replies[lastAckIdx+1:] { +476: if r.Author != botLogin { +477: return true +478: } +479: } +480: return false +481: } +482: +483: // BuildPerThreadPrompt dispatches to the appropriate prompt builder based on classification. +484: func BuildPerThreadPrompt(t TriagedThread, cfg *ReviewConfig) string { +485: switch t.Class { +486: case TriageCodeChanged: +487: return buildCodeChangePrompt(t, cfg) +488: case TriageHasReply: +489: return buildReplyPrompt(t, cfg) +490: case TriageCodeChangedReply: +491: return buildCodeChangeReplyPrompt(t, cfg) +492: case TriageCrossFileChange: +493: return buildCrossFilePrompt(t, cfg) +494: default: +495: return "" // TriageSkip — should not be called +496: } +497: } +498: +499: func buildCodeChangePrompt(t TriagedThread, cfg *ReviewConfig) string { +500: var b strings.Builder +501: +502: b.WriteString("You are a code reviewer. You previously raised a finding on a pull request. The author pushed new code.\n\n") +503: +504: writeFinding(&b, t.Thread) +505: +506: writeFileSnippet(&b, t.FileSnippet) +507: +508: b.WriteString("## Code Changes\n```diff\n") +509: b.WriteString(t.FileDiff) +510: b.WriteString("\n```\n\n") +511: +512: if ctx := evalContext(cfg, "code_change"); ctx != "" { +513: fmt.Fprintf(&b, "## Additional Context\n%s\n\n", ctx) +514: } +515: +516: b.WriteString("## Task\n") +517: b.WriteString("Determine whether the issue you raised has been resolved.\n\n") +518: b.WriteString("**Start with the current file content** (if provided). Read the code around the finding location and determine whether the issue still exists. If the problematic code has been fixed, removed, or restructured so the finding no longer applies, the answer is YES — regardless of which specific diff line produced the fix.\n\n") +519: b.WriteString("If file content is not available, examine the diff for evidence that the issue was addressed.\n") +520: b.WriteString("- A change that fixes the root cause, removes the problematic code, or meaningfully changes the code so the finding no longer applies → YES.\n") +521: b.WriteString("- A change to nearby or adjacent code counts IF it effectively resolves the concern (e.g. fixing the logic, adding the missing check, refactoring the problematic pattern).\n") +522: b.WriteString("- A structural change also counts — for example, if code was moved before a guard condition, control flow was reordered, or the code was refactored so the finding no longer applies.\n") +523: b.WriteString("- Answer NO only if the concern is still present in the current code (when file content is provided) or if the diff does not address the finding.\n\n") +524: writeCodeChangeResolutionFormat(&b) +525: +526: return b.String() +527: } +528: +529: func buildReplyPrompt(t TriagedThread, cfg *ReviewConfig) string { +530: var b strings.Builder +531: +532: b.WriteString("You are a code reviewer. You previously raised a finding on a pull request. The author replied.\n\n") +533: +534: writeFinding(&b, t.Thread) +535: writeReplies(&b, t.Thread, t.BotLogin) +536: +537: if ctx := evalContext(cfg, "reply"); ctx != "" { +538: fmt.Fprintf(&b, "## Additional Context\n%s\n\n", ctx) +539: } +540: +541: b.WriteString("## Task\n") +542: b.WriteString("Does the author's reply resolve the finding?\n") +543: b.WriteString("- **Dismissed**: Reply explicitly asks the reviewer to dismiss, ignore, or skip the finding (e.g. \"dismiss this\", \"you can safely dismiss\", \"please ignore\", \"skip this one\"). The author is exercising their authority to close the thread without further justification.\n") +544: b.WriteString("- **Acknowledged**: Reply indicates the finding is intentional, accepted, or tracked elsewhere (e.g. \"intentional\", \"will fix in a future PR\", \"tracked in issue #N\").\n") +545: b.WriteString("- **Rebutted**: Reply provides concrete technical reasoning showing the finding is not applicable. Vague disagreement (\"I don't think so\") does NOT qualify — the reply must cite specific technical details, framework behavior, or project constraints.\n") +546: b.WriteString("- **Not resolved**: Reply is a question, vague disagreement, or does not address the finding.\n\n") +547: writeResolutionFormat(&b) +548: +549: return b.String() +550: } +551: +552: func buildCodeChangeReplyPrompt(t TriagedThread, cfg *ReviewConfig) string { +553: var b strings.Builder +554: +555: b.WriteString("You are a code reviewer. You previously raised a finding on a pull request. The author pushed new code AND replied.\n\n") +556: +557: writeFinding(&b, t.Thread) +558: writeReplies(&b, t.Thread, t.BotLogin) +559: +560: writeFileSnippet(&b, t.FileSnippet) +561: +562: b.WriteString("## Code Changes\n```diff\n") +563: b.WriteString(t.FileDiff) +564: b.WriteString("\n```\n\n") +565: +566: if ctx := evalContext(cfg, "code_change"); ctx != "" { +567: fmt.Fprintf(&b, "## Additional Context (Code Changes)\n%s\n\n", ctx) +568: } +569: if ctx := evalContext(cfg, "reply"); ctx != "" { +570: fmt.Fprintf(&b, "## Additional Context (Replies)\n%s\n\n", ctx) +571: } +572: +573: b.WriteString("## Task\n") +574: b.WriteString("Is the finding resolved? It may be resolved by the code change, the reply, or both.\n\n") +575: b.WriteString("**Start with the current file content** (if provided). If the issue no longer exists in the current code — the root cause is fixed, the code was removed, or the code was restructured — it is resolved by code change regardless of which diff line produced the fix.\n\n") +576: b.WriteString("If the code still shows the issue, evaluate the reply:\n") +577: b.WriteString("- **Dismissed**: Reply explicitly asks the reviewer to dismiss, ignore, or skip the finding (e.g. \"dismiss this\", \"you can safely dismiss\", \"please ignore\", \"skip this one\"). The author is exercising their authority to close the thread without further justification.\n") +578: b.WriteString("- **Acknowledged**: Reply indicates the finding is intentional, accepted, or tracked elsewhere (e.g. \"intentional\", \"will fix in a future PR\", \"tracked in issue #N\").\n") +579: b.WriteString("- **Rebutted**: Reply provides concrete technical reasoning showing the finding is not applicable. Vague disagreement (\"I don't think so\") does NOT qualify — the reply must cite specific technical details, framework behavior, or project constraints.\n") +580: b.WriteString("- **Not resolved**: The issue is still in the code, and the reply is a question, vague disagreement, or does not address the finding.\n\n") +581: writeResolutionFormat(&b) +582: +583: return b.String() +584: } +585: +586: func buildCrossFilePrompt(t TriagedThread, cfg *ReviewConfig) string { +587: var b strings.Builder +588: +589: b.WriteString("You are a code reviewer. You previously raised a finding on a pull request. The author pushed new code, but the changes are in DIFFERENT files from where you left your finding.\n\n") +590: +591: writeFinding(&b, t.Thread) +592: +593: writeFileSnippet(&b, t.FileSnippet) +594: +595: b.WriteString("## All Code Changes\n```diff\n") +596: b.WriteString(t.FileDiff) +597: b.WriteString("\n```\n\n") +598: +599: if ctx := evalContext(cfg, "code_change"); ctx != "" { +600: fmt.Fprintf(&b, "## Additional Context\n%s\n\n", ctx) +601: } +602: +603: b.WriteString("## Task\n") +604: b.WriteString("Determine whether the issue you raised has been resolved, even though the changes are in different files from where you left your finding.\n") +605: b.WriteString("**Start with the current file content** (if provided). If the issue no longer exists in the current code, the answer is YES.\n\n") +606: b.WriteString("Otherwise, examine the code changes for evidence of a cross-file fix.\n") +607: writeCrossFileCriteria(&b) +608: b.WriteString("- Answer NO if none of the changes in this diff are related to the finding, or if file context is provided and the concern is still present in the current code.\n\n") +609: writeCodeChangeResolutionFormat(&b) +610: +611: return b.String() +612: } +613: +614: // buildWidenedScopePrompt is the level-2 fallback prompt used when the file-scoped +615: // evaluation (level 1) found no fix. It sends the full PR diff so the evaluator can +616: // detect cross-file fixes that the narrower scope missed. +617: func buildWidenedScopePrompt(t TriagedThread, cfg *ReviewConfig) string { +618: var b strings.Builder +619: +620: b.WriteString("You are a code reviewer. You previously raised a finding on a pull request. A focused check of the finding's file did not find a fix. Now examine ALL code changes across the PR — the fix may be in a different file.\n\n") +621: +622: writeFinding(&b, t.Thread) +623: +624: writeFileSnippet(&b, t.FileSnippet) +625: +626: b.WriteString("## All Code Changes (full PR)\n```diff\n") +627: b.WriteString(t.FullDiff) +628: b.WriteString("\n```\n\n") +629: +630: if ctx := evalContext(cfg, "code_change"); ctx != "" { +631: fmt.Fprintf(&b, "## Additional Context\n%s\n\n", ctx) +632: } +633: +634: b.WriteString("## Task\n") +635: b.WriteString("The finding's file was already checked and the issue appears unresolved there. Examine the full PR diff to determine if a change in another file resolves the concern.\n\n") +636: writeCrossFileCriteria(&b) +637: b.WriteString("- Answer NO if none of the changes in the diff are related to the finding.\n\n") +638: writeCodeChangeResolutionFormat(&b) +639: +640: return b.String() +641: } +642: +643: // writeCrossFileCriteria writes the shared YES-criteria bullets for cross-file evaluations. +644: // Used by both buildCrossFilePrompt and buildWidenedScopePrompt. +645: func writeCrossFileCriteria(b *strings.Builder) { +646: b.WriteString("- Answer YES if a change in another file effectively resolves the concern (e.g. fixing the caller instead of the callee, adding validation in a different layer, removing the code path that triggers the issue).\n") +647: b.WriteString("- A structural change also counts — for example, if code was moved, control flow was reordered, or the code was refactored so the finding no longer applies.\n") +648: } +649: +650: // writeFileSnippet adds the current file content section to the prompt when available. +651: func writeFileSnippet(b *strings.Builder, snippet string) { +652: if snippet == "" { +653: return +654: } +655: b.WriteString("## Current File Content (around finding)\n") +656: b.WriteString("This shows the file as it exists NOW (after the changes). Use it to understand the final code structure and control flow.\n\n~~~\n") +657: b.WriteString(snippet) +658: b.WriteString("~~~\n\n") +659: } +660: +661: // writeFinding writes the finding section to the prompt. +662: func writeFinding(b *strings.Builder, t ReviewThread) { +663: b.WriteString("## Finding\n") +664: fmt.Fprintf(b, "File: `%s:%d`\n", t.Path, t.Line) +665: b.WriteString(t.Body) +666: b.WriteString("\n\n") +667: } +668: +669: // writeReplies writes the author replies section to the prompt. +670: // Bot replies are filtered out so the bot's own acknowledgment messages +671: // don't leak into the Claude prompt and bias evaluation. +672: func writeReplies(b *strings.Builder, t ReviewThread, botLogin string) { +673: if len(t.Replies) == 0 { +674: return +675: } +676: // Filter out bot replies using explicit botLogin, consistent with +677: // hasHumanReply and hasNewHumanReply. +678: var humanReplies []ThreadReply +679: for _, r := range t.Replies { +680: if r.Author != botLogin { +681: humanReplies = append(humanReplies, r) +682: } +683: } +684: if len(humanReplies) == 0 { +685: return +686: } +687: b.WriteString("## Author Replies\n") +688: for _, r := range humanReplies { +689: normalizedBody := strings.ReplaceAll(r.Body, "\n", " ") +690: fmt.Fprintf(b, "> **@%s**: %s\n", r.Author, normalizedBody) +691: } +692: b.WriteString("\n") +693: } +694: +695: // writeResolutionFormat writes the expected JSON response format with all reason options. +696: // Used for prompts where author replies are present (TriageHasReply, TriageCodeChangedReply). +697: func writeResolutionFormat(b *strings.Builder) { +698: b.WriteString("Return a JSON object inside a ```json code fence:\n") +699: b.WriteString("- If resolved: `{\"resolved\": true, \"reason\": \"\", \"rationale\": \"\"}` — `reason` is one of `\"code_change\"`, `\"dismissed\"`, `\"acknowledged\"`, `\"rebutted\"`. `rationale` must be a single sentence naming the specific change, reply, or reasoning that resolved the finding (e.g. \"Removed the three `original_error_*` fields from the service logger\").\n") +700: b.WriteString("- If NOT resolved: `{\"resolved\": false}`\n") +701: } +702: +703: // writeCodeChangeResolutionFormat writes a restricted JSON response format +704: // for code-change-only evaluations (no author reply). Only allows code_change +705: // as a resolution reason since there is no author reply to acknowledge/dismiss/rebut. +706: func writeCodeChangeResolutionFormat(b *strings.Builder) { +707: b.WriteString("Return a JSON object inside a ```json code fence:\n") +708: b.WriteString("- If resolved: `{\"resolved\": true, \"reason\": \"code_change\", \"rationale\": \"\"}` — `rationale` must be a single sentence naming the specific change that resolved the finding (e.g. \"Removed the three `original_error_*` fields from the service logger\").\n") +709: b.WriteString("- If NOT resolved: `{\"resolved\": false}`\n") +710: } +711: +712: // evalContext returns the evaluation context string for a given type from config. +713: func evalContext(cfg *ReviewConfig, evalType string) string { +714: if cfg == nil || cfg.Evaluation == nil { +715: return "" +716: } +717: switch evalType { +718: case "code_change": +719: return cfg.Evaluation.CodeChange.Context +720: case "reply": +721: return cfg.Evaluation.Reply.Context +722: } +723: return "" +724: } +725: +726: // EvaluateThreadsParallel runs the LLM in parallel for threads that need evaluation. +727: // When maxBudgetUSD > 0, new goroutines are not launched once the budget is exceeded +728: // (already-running goroutines are allowed to finish). +729: func EvaluateThreadsParallel(triaged []TriagedThread, provider ModelProvider, cfg *ReviewConfig, maxConcurrent int, tracker *UsageTracker, maxBudgetUSD float64) []ThreadResolution { +730: results := make([]ThreadResolution, len(triaged)) +731: +732: sem := make(chan struct{}, maxConcurrent) +733: var wg sync.WaitGroup +734: +735: for i, t := range triaged { +736: if t.Class == TriageSkip || t.Class == TriageFileRemovedFromPR { +737: results[i] = ThreadResolution{Index: t.Index, Resolved: false} +738: continue +739: } +740: if t.Class == TriagePreviouslyAcked { +741: // Sticky: carry the prior ack reason forward without burning +742: // triage tokens. computeReviewSummary will route it to the +743: // matching bucket (Dismissed/Acknowledged/Rebutted), so the +744: // commit status check stays green across pushes that don't +745: // touch the deferred finding. +746: results[i] = ThreadResolution{ +747: Index: t.Index, +748: Resolved: true, +749: Reason: t.PriorAckReason, +750: } +751: continue +752: } +753: // Soft budget cap: skip remaining evaluations if budget is exceeded. +754: if err := CheckBudget(tracker, maxBudgetUSD); err != nil { +755: results[i] = ThreadResolution{Index: t.Index, Error: err} +756: continue +757: } +758: wg.Add(1) +759: go func(idx int, tt TriagedThread) { +760: defer wg.Done() +761: sem <- struct{}{} +762: defer func() { <-sem }() +763: +764: // Level 1: file-scoped evaluation. +765: prompt := BuildPerThreadPrompt(tt, cfg) +766: result, err := provider.Run(context.Background(), prompt, RunOpts{}) +767: if err != nil { +768: results[idx] = ThreadResolution{Index: tt.Index, Error: err} +769: return +770: } +771: trackUsage(tracker, result, "triage") +772: res := parseThreadResolution(result.Text, tt.Index) +773: res = validateResolutionReason(res, tt.Class) +774: +775: // Level 2: widen to full PR diff if file-scoped check was inconclusive. +776: // Only for TriageCodeChanged — TriageCodeChangedReply has author replies +777: // that buildWidenedScopePrompt doesn't include, so widening would drop +778: // the reply context and prevent dismissed/acknowledged/rebutted resolutions. +779: if !res.Resolved && tt.FullDiff != "" && tt.Class == TriageCodeChanged { +780: if err := CheckBudget(tracker, maxBudgetUSD); err == nil { +781: fmt.Fprintf(os.Stderr, " [widen] %s — checking full PR diff\n", threadLabel(tt.Thread)) +782: prompt2 := buildWidenedScopePrompt(tt, cfg) +783: result2, err := provider.Run(context.Background(), prompt2, RunOpts{}) +784: if err == nil { +785: trackUsage(tracker, result2, "triage") +786: res2 := parseThreadResolution(result2.Text, tt.Index) +787: res2 = validateResolutionReason(res2, tt.Class) +788: if res2.Resolved { +789: res = res2 +790: } +791: } +792: } +793: } +794: +795: results[idx] = res +796: }(i, t) +797: } +798: +799: wg.Wait() +800: return results +801: } +802: +803: // parseThreadResolution parses Claude's JSON response for a single thread evaluation. +804: func parseThreadResolution(output string, index int) ThreadResolution { +805: allMatches := jsonFenceRe.FindAllStringSubmatch(output, -1) +806: for _, matches := range allMatches { +807: raw := matches[1] +808: +809: var resp struct { +810: Resolved bool `json:"resolved"` +811: Reason string `json:"reason"` +812: Rationale string `json:"rationale"` +813: } +814: if err := json.Unmarshal([]byte(raw), &resp); err != nil { +815: continue +816: } +817: return ThreadResolution{ +818: Index: index, +819: Resolved: resp.Resolved, +820: Reason: resp.Reason, +821: Rationale: strings.TrimSpace(resp.Rationale), +822: } +823: } +824: +825: // If parsing fails, treat as unresolved (conservative). +826: return ThreadResolution{Index: index, Resolved: false} +827: } +828: +829: // validateResolutionReason enforces that code-change-only classifications +830: // (no author reply) can only resolve with reason "code_change". If Claude +831: // returns "acknowledged"/"dismissed"/"rebutted", treat it as unresolved. +832: func validateResolutionReason(res ThreadResolution, class ThreadClassification) ThreadResolution { +833: if res.Resolved && res.Reason != "code_change" && +834: (class == TriageCodeChanged || class == TriageCrossFileChange) { +835: res.Resolved = false +836: res.Reason = "" +837: } +838: return res +839: } +840: +841: // LogTriage prints structured triage results to stderr. +842: func LogTriage(triaged []TriagedThread) { +843: fmt.Fprintf(os.Stderr, "Re-evaluating %d unresolved thread(s)...\n\n", len(triaged)) +844: +845: for _, t := range triaged { +846: label := threadLabel(t.Thread) +847: switch t.Class { +848: case TriageSkip: +849: fmt.Fprintf(os.Stderr, " [skip] %s — no code changes, no human replies\n", label) +850: case TriageCodeChanged: +851: fmt.Fprintf(os.Stderr, " [evaluate] %s — code changes detected\n", label) +852: case TriageHasReply: +853: fmt.Fprintf(os.Stderr, " [evaluate] %s — human reply detected\n", label) +854: case TriageCodeChangedReply: +855: fmt.Fprintf(os.Stderr, " [evaluate] %s — code changes + human reply detected\n", label) +856: case TriageCrossFileChange: +857: fmt.Fprintf(os.Stderr, " [evaluate] %s — cross-file changes detected\n", label) +858: case TriageFileRemovedFromPR: +859: fmt.Fprintf(os.Stderr, " [resolve] %s — file removed from PR\n", label) +860: case TriagePreviouslyAcked: +861: fmt.Fprintf(os.Stderr, " [sticky] %s — prior ack (%s) carried forward\n", label, t.PriorAckReason) +862: } +863: } +864: +865: skipped := 0 +866: autoResolved := 0 +867: needsEval := 0 +868: for _, t := range triaged { +869: switch t.Class { +870: case TriageSkip: +871: skipped++ +872: case TriageFileRemovedFromPR, TriagePreviouslyAcked: +873: autoResolved++ +874: default: +875: needsEval++ +876: } +877: } +878: fmt.Fprintf(os.Stderr, "\nTriage result: %d skipped, %d auto-resolved, %d need evaluation\n", skipped, autoResolved, needsEval) +879: } +880: +881: // LogResolutions prints structured evaluation results to stderr. +882: func LogResolutions(triaged []TriagedThread, resolutions []ThreadResolution) { +883: fmt.Fprintf(os.Stderr, "\n") +884: for i, r := range resolutions { +885: switch triaged[i].Class { +886: case TriageSkip, TriageFileRemovedFromPR, TriagePreviouslyAcked: +887: continue +888: } +889: label := threadLabel(triaged[i].Thread) +890: if r.Error != nil { +891: if isBudgetError(r.Error) { +892: fmt.Fprintf(os.Stderr, " [skip] %s — %v\n", label, r.Error) +893: } else { +894: fmt.Fprintf(os.Stderr, " [error] %s — evaluation failed: %v\n", label, r.Error) +895: } +896: } else if r.Resolved { +897: switch r.Reason { +898: case "code_change": +899: fmt.Fprintf(os.Stderr, " [resolved] %s — fixed by code change\n", label) +900: case "dismissed": +901: fmt.Fprintf(os.Stderr, " [ack] %s — dismissed by author (keeping open)\n", label) +902: case "acknowledged": +903: fmt.Fprintf(os.Stderr, " [ack] %s — acknowledged by author (keeping open)\n", label) +904: case "rebutted": +905: fmt.Fprintf(os.Stderr, " [ack] %s — rebutted by author (keeping open)\n", label) +906: default: +907: fmt.Fprintf(os.Stderr, " [resolved] %s — resolved\n", label) +908: } +909: } else { +910: fmt.Fprintf(os.Stderr, " [open] %s — not resolved\n", label) +911: } +912: } +913: } +914: +915: // countNonSkipped returns the number of triaged threads that EvaluateThreadsParallel +916: // must process. TriagePreviouslyAcked is included even though it short-circuits +917: // without an LLM call: the function still has to run to emit the carried-forward +918: // fixedThread, which is what feeds the summary's Acknowledged/Rebutted/Dismissed +919: // buckets. Excluding it here was the bug behind the "Still unresolved: 2" stuck +920: // status on already-acked threads. +921: func countNonSkipped(triaged []TriagedThread) int { +922: n := 0 +923: for _, t := range triaged { +924: switch t.Class { +925: case TriageSkip, TriageFileRemovedFromPR: +926: continue +927: default: +928: n++ +929: } +930: } +931: return n +932: } +933: +934: // toFixedThreads converts thread resolutions to the fixedThread type used by downstream code. +935: func toFixedThreads(resolutions []ThreadResolution) []fixedThread { +936: var result []fixedThread +937: for _, r := range resolutions { +938: if r.Resolved { +939: result = append(result, fixedThread{Index: r.Index, Reason: r.Reason, Rationale: r.Rationale}) +940: } +941: } +942: return result +943: } +```` + +### `internal/review/triage_test.go` +```` +1: package review +2: +3: import ( +4: "context" +5: "fmt" +6: "strings" +7: "testing" +8: ) +9: +10: // --- ExtractFileSnippet tests --- +11: +12: func TestExtractFileSnippet_Basic(t *testing.T) { +13: // 100-line file, finding at line 50, hunk at lines 45-55. +14: var lines []string +15: for i := 1; i <= 100; i++ { +16: lines = append(lines, fmt.Sprintf("line %d content", i)) +17: } +18: content := strings.Join(lines, "\n") +19: diff := "@@ -40,10 +45,11 @@ func foo() {\n+added line\n" +20: +21: snippet := ExtractFileSnippet(content, 50, diff, 300) +22: if snippet == "" { +23: t.Fatal("expected non-empty snippet") +24: } +25: // Should contain the finding line. +26: if !strings.Contains(snippet, "50: line 50 content") { +27: t.Error("snippet should contain the finding line") +28: } +29: // Should contain hunk area. +30: if !strings.Contains(snippet, "45: line 45 content") { +31: t.Error("snippet should contain hunk start area") +32: } +33: } +34: +35: func TestExtractFileSnippet_MergesOverlappingRanges(t *testing.T) { +36: var lines []string +37: for i := 1; i <= 200; i++ { +38: lines = append(lines, fmt.Sprintf("line %d", i)) +39: } +40: content := strings.Join(lines, "\n") +41: // Two hunks close together — should merge into one contiguous range. +42: diff := "@@ -10,5 +10,5 @@\n+a\n@@ -20,5 +20,5 @@\n+b\n" +43: +44: snippet := ExtractFileSnippet(content, 15, diff, 300) +45: // Should NOT contain omission markers since ranges overlap/merge. +46: if strings.Contains(snippet, "lines omitted") { +47: t.Error("close hunks should merge without omission markers") +48: } +49: } +50: +51: func TestExtractFileSnippet_CapsAtMaxLines(t *testing.T) { +52: var lines []string +53: for i := 1; i <= 1000; i++ { +54: lines = append(lines, fmt.Sprintf("line %d", i)) +55: } +56: content := strings.Join(lines, "\n") +57: // Hunks spread across the file. +58: diff := "@@ -10,5 +10,5 @@\n+a\n@@ -500,5 +500,5 @@\n+b\n@@ -900,5 +900,5 @@\n+c\n" +59: +60: snippet := ExtractFileSnippet(content, 50, diff, 150) +61: snippetLines := strings.Split(strings.TrimRight(snippet, "\n"), "\n") +62: if len(snippetLines) > 160 { // small buffer for omission markers +63: t.Errorf("snippet should respect maxLines cap, got %d lines", len(snippetLines)) +64: } +65: // Must contain the finding line. +66: if !strings.Contains(snippet, "50: line 50") { +67: t.Error("snippet must prioritize the finding line") +68: } +69: } +70: +71: func TestExtractFileSnippet_NoDiff(t *testing.T) { +72: var lines []string +73: for i := 1; i <= 100; i++ { +74: lines = append(lines, fmt.Sprintf("line %d", i)) +75: } +76: content := strings.Join(lines, "\n") +77: +78: // Cross-file case: no diff for this file. +79: snippet := ExtractFileSnippet(content, 50, "", 300) +80: if snippet == "" { +81: t.Fatal("expected non-empty snippet for cross-file case") +82: } +83: if !strings.Contains(snippet, "50: line 50") { +84: t.Error("snippet should center on finding line") +85: } +86: } +87: +88: func TestExtractFileSnippet_ZeroCountHunk(t *testing.T) { +89: // A hunk with count=0 (pure deletion) should not produce a range. +90: ranges := parseHunkNewRanges("@@ -5,3 +10,0 @@\n-deleted line\n") +91: if len(ranges) != 0 { +92: t.Errorf("expected 0 ranges for zero-count hunk, got %d", len(ranges)) +93: } +94: } +95: +96: func TestExtractFileSnippet_FindingLineZero(t *testing.T) { +97: var lines []string +98: for i := 1; i <= 100; i++ { +99: lines = append(lines, fmt.Sprintf("line %d", i)) +100: } +101: content := strings.Join(lines, "\n") +102: diff := "@@ -40,10 +40,10 @@\n+changed\n" +103: +104: // findingLine=0 should not panic and should anchor to hunk area. +105: snippet := ExtractFileSnippet(content, 0, diff, 50) +106: if snippet == "" { +107: t.Fatal("expected non-empty snippet even with findingLine=0") +108: } +109: // Should contain hunk area, not be anchored to line 0. +110: if !strings.Contains(snippet, "40: line 40") { +111: t.Error("snippet should include hunk area when findingLine is 0") +112: } +113: } +114: +115: func TestExtractFileSnippet_EmptyContent(t *testing.T) { +116: snippet := ExtractFileSnippet("", 10, "@@ -1,5 +1,5 @@\n", 300) +117: if snippet != "" { +118: t.Error("expected empty snippet for empty content") +119: } +120: } +121: +122: // --- Prompt builder tests --- +123: +124: func TestBuildCodeChangePrompt_IncludesFileContext(t *testing.T) { +125: tt := TriagedThread{ +126: Thread: ReviewThread{ +127: Path: "main.go", +128: Line: 10, +129: Body: "Found a bug", +130: }, +131: FileDiff: "+ fixed line", +132: FileSnippet: "9: before\n10: the line\n11: after\n", +133: } +134: prompt := buildCodeChangePrompt(tt, nil) +135: +136: if !strings.Contains(prompt, "## Current File Content (around finding)") { +137: t.Error("prompt should include file context section when FileSnippet is set") +138: } +139: if !strings.Contains(prompt, "10: the line") { +140: t.Error("prompt should include the file snippet content") +141: } +142: } +143: +144: func TestBuildCodeChangePrompt_NoFileContextWhenEmpty(t *testing.T) { +145: tt := TriagedThread{ +146: Thread: ReviewThread{ +147: Path: "main.go", +148: Line: 10, +149: Body: "Found a bug", +150: }, +151: FileDiff: "+ fixed line", +152: } +153: prompt := buildCodeChangePrompt(tt, nil) +154: +155: if strings.Contains(prompt, "## Current File Content") { +156: t.Error("prompt should NOT include file context section when FileSnippet is empty") +157: } +158: } +159: +160: func TestBuildCodeChangePrompt_StructuralChangeInstruction(t *testing.T) { +161: tt := TriagedThread{ +162: Thread: ReviewThread{ +163: Path: "main.go", +164: Line: 10, +165: Body: "Found a bug", +166: }, +167: FileDiff: "+ fixed line", +168: } +169: prompt := buildCodeChangePrompt(tt, nil) +170: +171: if !strings.Contains(prompt, "structural change") { +172: t.Error("prompt should include structural change guidance") +173: } +174: } +175: +176: func TestBuildCrossFilePrompt_IncludesFileContext(t *testing.T) { +177: tt := TriagedThread{ +178: Thread: ReviewThread{ +179: Path: "main.go", +180: Line: 10, +181: Body: "Found a bug", +182: }, +183: FileDiff: "+ change in other file", +184: FileSnippet: "9: before\n10: the line\n11: after\n", +185: } +186: prompt := buildCrossFilePrompt(tt, nil) +187: +188: if !strings.Contains(prompt, "## Current File Content (around finding)") { +189: t.Error("cross-file prompt should include file context section") +190: } +191: if !strings.Contains(prompt, "structural change") { +192: t.Error("cross-file prompt should include structural change guidance") +193: } +194: } +195: +196: func TestBuildCodeChangePrompt_OnlyAllowsCodeChangeReason(t *testing.T) { +197: tt := TriagedThread{ +198: Thread: ReviewThread{ +199: Path: "main.go", +200: Line: 10, +201: Body: "Found a bug", +202: }, +203: FileDiff: "+ fixed line", +204: } +205: prompt := buildCodeChangePrompt(tt, nil) +206: +207: if strings.Contains(prompt, `"acknowledged"`) { +208: t.Error("buildCodeChangePrompt should not offer 'acknowledged' as a reason") +209: } +210: if strings.Contains(prompt, `"dismissed"`) { +211: t.Error("buildCodeChangePrompt should not offer 'dismissed' as a reason") +212: } +213: if strings.Contains(prompt, `"rebutted"`) { +214: t.Error("buildCodeChangePrompt should not offer 'rebutted' as a reason") +215: } +216: if !strings.Contains(prompt, `"code_change"`) { +217: t.Error("buildCodeChangePrompt must offer 'code_change' as a reason") +218: } +219: } +220: +221: func TestBuildCrossFilePrompt_OnlyAllowsCodeChangeReason(t *testing.T) { +222: tt := TriagedThread{ +223: Thread: ReviewThread{ +224: Path: "main.go", +225: Line: 10, +226: Body: "Found a bug", +227: }, +228: FileDiff: "+ fixed in other file", +229: } +230: prompt := buildCrossFilePrompt(tt, nil) +231: +232: if strings.Contains(prompt, `"acknowledged"`) { +233: t.Error("buildCrossFilePrompt should not offer 'acknowledged' as a reason") +234: } +235: if strings.Contains(prompt, `"dismissed"`) { +236: t.Error("buildCrossFilePrompt should not offer 'dismissed' as a reason") +237: } +238: if strings.Contains(prompt, `"rebutted"`) { +239: t.Error("buildCrossFilePrompt should not offer 'rebutted' as a reason") +240: } +241: if !strings.Contains(prompt, `"code_change"`) { +242: t.Error("buildCrossFilePrompt must offer 'code_change' as a reason") +243: } +244: } +245: +246: func TestBuildReplyPrompt_AllowsAllReasons(t *testing.T) { +247: tt := TriagedThread{ +248: Thread: ReviewThread{ +249: Path: "main.go", +250: Line: 10, +251: Body: "Found a bug", +252: Replies: []ThreadReply{ +253: {Author: "user1", Body: "Will fix later"}, +254: }, +255: }, +256: BotLogin: "codecanary-bot", +257: } +258: prompt := buildReplyPrompt(tt, nil) +259: +260: for _, reason := range []string{`"code_change"`, `"acknowledged"`, `"dismissed"`, `"rebutted"`} { +261: if !strings.Contains(prompt, reason) { +262: t.Errorf("buildReplyPrompt must offer %s as a reason", reason) +263: } +264: } +265: } +266: +267: func TestBuildCodeChangeReplyPrompt_AllowsAllReasons(t *testing.T) { +268: tt := TriagedThread{ +269: Thread: ReviewThread{ +270: Path: "main.go", +271: Line: 10, +272: Body: "Found a bug", +273: Replies: []ThreadReply{ +274: {Author: "user1", Body: "Fixed it"}, +275: }, +276: }, +277: FileDiff: "+ fixed line", +278: BotLogin: "codecanary-bot", +279: } +280: prompt := buildCodeChangeReplyPrompt(tt, nil) +281: +282: for _, reason := range []string{`"code_change"`, `"acknowledged"`, `"dismissed"`, `"rebutted"`} { +283: if !strings.Contains(prompt, reason) { +284: t.Errorf("buildCodeChangeReplyPrompt must offer %s as a reason", reason) +285: } +286: } +287: } +288: +289: // --- ClassifyThreads diff scoping tests --- +290: +291: func TestClassifyThreads_FileScopedDiffForCodeChanged(t *testing.T) { +292: threads := []ReviewThread{ +293: {Path: "a.go", Line: 10, Body: "Issue in a.go", Outdated: true}, +294: } +295: fullDiff := "diff --git a/a.go b/a.go\n--- a/a.go\n+++ b/a.go\n@@ -10,3 +10,3 @@\n-old\n+new\n" + +296: "diff --git a/b.go b/b.go\n--- a/b.go\n+++ b/b.go\n@@ -5,3 +5,3 @@\n-old b\n+new b\n" +297: +298: triaged := ClassifyThreads(threads, fullDiff, fullDiff, "bot", []string{"a.go", "b.go"}, nil) +299: +300: if triaged[0].Class != TriageCodeChanged { +301: t.Fatalf("expected TriageCodeChanged, got %d", triaged[0].Class) +302: } +303: if strings.Contains(triaged[0].FileDiff, "b.go") { +304: t.Error("FileDiff for TriageCodeChanged should be file-scoped, not full PR diff") +305: } +306: if !strings.Contains(triaged[0].FileDiff, "a.go") { +307: t.Error("FileDiff should contain the finding's file diff") +308: } +309: // FullDiff should contain the entire PR diff for widened-scope fallback. +310: if !strings.Contains(triaged[0].FullDiff, "b.go") { +311: t.Error("FullDiff should contain the full PR diff for fallback") +312: } +313: } +314: +315: func TestClassifyThreads_NoFullDiffForCrossFile(t *testing.T) { +316: threads := []ReviewThread{ +317: {Path: "a.go", Line: 10, Body: "Issue in a.go"}, +318: } +319: diff := "diff --git a/b.go b/b.go\n--- a/b.go\n+++ b/b.go\n@@ -5,3 +5,3 @@\n-old b\n+new b\n" +320: +321: triaged := ClassifyThreads(threads, diff, diff, "bot", []string{"a.go", "b.go"}, nil) +322: +323: if triaged[0].Class != TriageCrossFileChange { +324: t.Fatalf("expected TriageCrossFileChange, got %d", triaged[0].Class) +325: } +326: // TriageCrossFileChange already gets the full diff as FileDiff — no fallback needed. +327: if triaged[0].FullDiff != "" { +328: t.Error("FullDiff should be empty for TriageCrossFileChange (no fallback needed)") +329: } +330: } +331: +332: func TestBuildWidenedScopePrompt(t *testing.T) { +333: tt := TriagedThread{ +334: Thread: ReviewThread{ +335: Path: "main.go", +336: Line: 10, +337: Body: "Found a bug", +338: }, +339: FullDiff: "+ fix in other file", +340: FileSnippet: "9: before\n10: the line\n11: after\n", +341: } +342: prompt := buildWidenedScopePrompt(tt, nil) +343: +344: if !strings.Contains(prompt, "full PR diff") { +345: t.Error("widened prompt should mention full PR diff") +346: } +347: if !strings.Contains(prompt, "another file") { +348: t.Error("widened prompt should guide LLM to check other files") +349: } +350: if !strings.Contains(prompt, "+ fix in other file") { +351: t.Error("widened prompt should include FullDiff content") +352: } +353: if !strings.Contains(prompt, "## Current File Content") { +354: t.Error("widened prompt should include file snippet") +355: } +356: // Should only allow code_change reason (no reply-based reasons). +357: if strings.Contains(prompt, `"acknowledged"`) || strings.Contains(prompt, `"dismissed"`) { +358: t.Error("widened prompt should not offer reply-based resolution reasons") +359: } +360: } +361: +362: func TestClassifyThreads_FullDiffForCrossFile(t *testing.T) { +363: threads := []ReviewThread{ +364: {Path: "a.go", Line: 10, Body: "Issue in a.go"}, +365: } +366: // Only b.go changed — finding's file (a.go) is not in the diff. +367: diff := "diff --git a/b.go b/b.go\n--- a/b.go\n+++ b/b.go\n@@ -5,3 +5,3 @@\n-old b\n+new b\n" +368: +369: triaged := ClassifyThreads(threads, diff, diff, "bot", []string{"a.go", "b.go"}, nil) +370: +371: if triaged[0].Class != TriageCrossFileChange { +372: t.Fatalf("expected TriageCrossFileChange, got %d", triaged[0].Class) +373: } +374: if !strings.Contains(triaged[0].FileDiff, "b.go") { +375: t.Error("FileDiff for TriageCrossFileChange should contain the full PR diff") +376: } +377: } +378: +379: func TestValidateResolutionReason_RejectsInvalidReasonForCodeChangeOnly(t *testing.T) { +380: // Simulate Claude returning "acknowledged" for a code-change-only thread. +381: output := "```json\n{\"resolved\": true, \"reason\": \"acknowledged\"}\n```" +382: parsed := parseThreadResolution(output, 0) +383: +384: // parseThreadResolution itself accepts any reason (it's just a parser). +385: if !parsed.Resolved || parsed.Reason != "acknowledged" { +386: t.Fatal("parseThreadResolution should parse the raw response as-is") +387: } +388: +389: // For code-change-only classifications, invalid reasons should be rejected. +390: for _, class := range []ThreadClassification{TriageCodeChanged, TriageCrossFileChange} { +391: res := validateResolutionReason(parsed, class) +392: if res.Resolved { +393: t.Errorf("class %d: resolution with reason 'acknowledged' should be rejected", class) +394: } +395: if res.Reason != "" { +396: t.Errorf("class %d: reason should be cleared, got %q", class, res.Reason) +397: } +398: } +399: +400: // For reply-based classifications, the same reason should be accepted. +401: for _, class := range []ThreadClassification{TriageHasReply, TriageCodeChangedReply} { +402: res := validateResolutionReason(parsed, class) +403: if !res.Resolved { +404: t.Errorf("class %d: resolution with reason 'acknowledged' should be accepted", class) +405: } +406: if res.Reason != "acknowledged" { +407: t.Errorf("class %d: reason should be 'acknowledged', got %q", class, res.Reason) +408: } +409: } +410: } +411: +412: func TestParseThreadResolution_CapturesRationale(t *testing.T) { +413: output := "```json\n{\"resolved\": true, \"reason\": \"code_change\", \"rationale\": \" Removed the three original_error_* fields \"}\n```" +414: parsed := parseThreadResolution(output, 7) +415: +416: if !parsed.Resolved || parsed.Reason != "code_change" { +417: t.Fatalf("expected resolved code_change, got %+v", parsed) +418: } +419: if parsed.Rationale != "Removed the three original_error_* fields" { +420: t.Errorf("rationale should be trimmed, got %q", parsed.Rationale) +421: } +422: if parsed.Index != 7 { +423: t.Errorf("index should be preserved, got %d", parsed.Index) +424: } +425: } +426: +427: func TestToFixedThreads_PropagatesRationale(t *testing.T) { +428: resolutions := []ThreadResolution{ +429: {Index: 0, Resolved: true, Reason: "code_change", Rationale: "dropped dead fields"}, +430: {Index: 1, Resolved: false}, +431: {Index: 2, Resolved: true, Reason: "dismissed"}, +432: } +433: fixed := toFixedThreads(resolutions) +434: +435: if len(fixed) != 2 { +436: t.Fatalf("expected 2 fixed threads, got %d", len(fixed)) +437: } +438: if fixed[0].Rationale != "dropped dead fields" { +439: t.Errorf("first rationale should carry over, got %q", fixed[0].Rationale) +440: } +441: if fixed[1].Rationale != "" { +442: t.Errorf("empty rationale should stay empty, got %q", fixed[1].Rationale) +443: } +444: } +445: +446: func TestExtractAckReason(t *testing.T) { +447: cases := []struct { +448: name string +449: body string +450: want string +451: }{ +452: {"current marker", "\nKeeping open.", "rebutted"}, +453: {"acknowledged", "\nKeeping open.", "acknowledged"}, +454: {"dismissed", "\nKeeping open.", "dismissed"}, +455: {"legacy marker", "\nKeeping open.", "rebutted"}, +456: {"no marker", "Some normal reply", ""}, +457: {"no closing tag", "\n", "unknown"}, +459: } +460: for _, tc := range cases { +461: t.Run(tc.name, func(t *testing.T) { +462: if got := extractAckReason(tc.body); got != tc.want { +463: t.Errorf("extractAckReason(%q) = %q, want %q", tc.body, got, tc.want) +464: } +465: }) +466: } +467: } +468: +469: func TestParsePriorAckReason_PicksLatestBotAck(t *testing.T) { +470: thread := ReviewThread{ +471: Replies: []ThreadReply{ +472: {Author: "user", Body: "Deferring — see rationale."}, +473: {Author: "bot", Body: "\nAuthor acknowledged this finding."}, +474: {Author: "bot", Body: "\nAuthor provided rebuttal."}, +475: }, +476: } +477: if got := parsePriorAckReason(thread, "bot"); got != "rebutted" { +478: t.Errorf("expected latest reason 'rebutted', got %q", got) +479: } +480: } +481: +482: func TestParsePriorAckReason_IgnoresNonBotMarkers(t *testing.T) { +483: thread := ReviewThread{ +484: Replies: []ThreadReply{ +485: {Author: "attacker", Body: "\nFake."}, +486: }, +487: } +488: if got := parsePriorAckReason(thread, "bot"); got != "" { +489: t.Errorf("non-bot ack should not be sticky, got %q", got) +490: } +491: } +492: +493: func TestClassifyThreads_StickyAckSurvivesNextPush(t *testing.T) { +494: threads := []ReviewThread{ +495: { +496: Path: "CLAUDE.md", +497: Line: 724, +498: Body: "Original finding", +499: Replies: []ThreadReply{ +500: {Author: "user", Body: "Deferring — forward recommendation, not yet implemented."}, +501: {Author: "bot", Body: "\nAuthor acknowledged."}, +502: }, +503: }, +504: } +505: // File changed in the next push — without sticky-ack this would +506: // classify as TriageCodeChanged and lose the prior reason. +507: diff := "diff --git a/CLAUDE.md b/CLAUDE.md\n--- a/CLAUDE.md\n+++ b/CLAUDE.md\n@@ -700,3 +700,3 @@\n-old\n+new\n" +508: +509: triaged := ClassifyThreads(threads, diff, diff, "bot", []string{"CLAUDE.md"}, nil) +510: +511: if triaged[0].Class != TriagePreviouslyAcked { +512: t.Fatalf("expected TriagePreviouslyAcked, got %d", triaged[0].Class) +513: } +514: if triaged[0].PriorAckReason != "acknowledged" { +515: t.Errorf("expected PriorAckReason 'acknowledged', got %q", triaged[0].PriorAckReason) +516: } +517: } +518: +519: func TestClassifyThreads_NewHumanReplyBreaksStickiness(t *testing.T) { +520: // Author replied AGAIN after the bot's ack — that's a fresh signal +521: // and should reopen the thread for evaluation, not stay sticky. +522: threads := []ReviewThread{ +523: { +524: Path: "CLAUDE.md", +525: Line: 724, +526: Replies: []ThreadReply{ +527: {Author: "user", Body: "Original deferring rationale."}, +528: {Author: "bot", Body: "\nAck."}, +529: {Author: "user", Body: "Wait, actually this one needs another look."}, +530: }, +531: }, +532: } +533: triaged := ClassifyThreads(threads, "", "", "bot", []string{"CLAUDE.md"}, nil) +534: +535: if triaged[0].Class != TriageHasReply { +536: t.Fatalf("expected TriageHasReply (new reply after ack), got %d", triaged[0].Class) +537: } +538: if triaged[0].PriorAckReason != "" { +539: t.Errorf("PriorAckReason should be empty when new reply present, got %q", triaged[0].PriorAckReason) +540: } +541: } +542: +543: func TestCountNonSkipped_IncludesPreviouslyAcked(t *testing.T) { +544: // Regression: sticky-ack threads must count toward needsEval so +545: // EvaluateThreadsParallel runs and emits the carried-forward fixedThread. +546: // Excluding them caused the runner to skip eval entirely, leaving the +547: // summary stuck on "Still unresolved" for threads the bot had already +548: // acked. +549: triaged := []TriagedThread{ +550: {Class: TriageSkip}, +551: {Class: TriageFileRemovedFromPR}, +552: {Class: TriagePreviouslyAcked, PriorAckReason: "acknowledged"}, +553: {Class: TriagePreviouslyAcked, PriorAckReason: "rebutted"}, +554: } +555: if got := countNonSkipped(triaged); got != 2 { +556: t.Errorf("countNonSkipped = %d, want 2 (the two PreviouslyAcked threads)", got) +557: } +558: } +559: +560: func TestEvaluateThreadsParallel_StickyAckShortCircuits(t *testing.T) { +561: // Provider that panics if called — sticky-ack must skip the LLM. +562: provider := &stickyAckPanicProvider{t: t} +563: +564: triaged := []TriagedThread{ +565: {Index: 0, Class: TriagePreviouslyAcked, PriorAckReason: "rebutted"}, +566: {Index: 1, Class: TriagePreviouslyAcked, PriorAckReason: "dismissed"}, +567: } +568: results := EvaluateThreadsParallel(triaged, provider, nil, 3, &UsageTracker{}, 0) +569: +570: if len(results) != 2 { +571: t.Fatalf("expected 2 results, got %d", len(results)) +572: } +573: if !results[0].Resolved || results[0].Reason != "rebutted" { +574: t.Errorf("first result: expected resolved=true reason=rebutted, got %+v", results[0]) +575: } +576: if !results[1].Resolved || results[1].Reason != "dismissed" { +577: t.Errorf("second result: expected resolved=true reason=dismissed, got %+v", results[1]) +578: } +579: } +580: +581: type stickyAckPanicProvider struct{ t *testing.T } +582: +583: func (p *stickyAckPanicProvider) Run(_ context.Context, _ string, _ RunOpts) (*providerResult, error) { +584: p.t.Fatalf("provider.Run must not be called for TriagePreviouslyAcked threads") +585: return nil, nil +586: } +587: +588: func TestResolutionFormat_RequestsRationale(t *testing.T) { +589: for _, fn := range map[string]func(*strings.Builder){ +590: "writeResolutionFormat": writeResolutionFormat, +591: "writeCodeChangeResolutionFormat": writeCodeChangeResolutionFormat, +592: } { +593: var b strings.Builder +594: fn(&b) +595: out := b.String() +596: if !strings.Contains(out, `"rationale"`) { +597: t.Errorf("prompt format should request a rationale field, got:\n%s", out) +598: } +599: } +600: } +```` + +## Diff +```diff +diff --git a/internal/review/triage.go b/internal/review/triage.go +index dc20e2a..485db27 100644 +--- a/internal/review/triage.go ++++ b/internal/review/triage.go +@@ -912,12 +912,17 @@ func LogResolutions(triaged []TriagedThread, resolutions []ThreadResolution) { + } + } + +-// countNonSkipped returns the number of triaged threads that need LLM evaluation. ++// countNonSkipped returns the number of triaged threads that EvaluateThreadsParallel ++// must process. TriagePreviouslyAcked is included even though it short-circuits ++// without an LLM call: the function still has to run to emit the carried-forward ++// fixedThread, which is what feeds the summary's Acknowledged/Rebutted/Dismissed ++// buckets. Excluding it here was the bug behind the "Still unresolved: 2" stuck ++// status on already-acked threads. + func countNonSkipped(triaged []TriagedThread) int { + n := 0 + for _, t := range triaged { + switch t.Class { +- case TriageSkip, TriageFileRemovedFromPR, TriagePreviouslyAcked: ++ case TriageSkip, TriageFileRemovedFromPR: + continue + default: + n++ +diff --git a/internal/review/triage_test.go b/internal/review/triage_test.go +index be37fcb..79a5950 100644 +--- a/internal/review/triage_test.go ++++ b/internal/review/triage_test.go +@@ -540,6 +540,23 @@ func TestClassifyThreads_NewHumanReplyBreaksStickiness(t *testing.T) { + } + } + ++func TestCountNonSkipped_IncludesPreviouslyAcked(t *testing.T) { ++ // Regression: sticky-ack threads must count toward needsEval so ++ // EvaluateThreadsParallel runs and emits the carried-forward fixedThread. ++ // Excluding them caused the runner to skip eval entirely, leaving the ++ // summary stuck on "Still unresolved" for threads the bot had already ++ // acked. ++ triaged := []TriagedThread{ ++ {Class: TriageSkip}, ++ {Class: TriageFileRemovedFromPR}, ++ {Class: TriagePreviouslyAcked, PriorAckReason: "acknowledged"}, ++ {Class: TriagePreviouslyAcked, PriorAckReason: "rebutted"}, ++ } ++ if got := countNonSkipped(triaged); got != 2 { ++ t.Errorf("countNonSkipped = %d, want 2 (the two PreviouslyAcked threads)", got) ++ } ++} ++ + func TestEvaluateThreadsParallel_StickyAckShortCircuits(t *testing.T) { + // Provider that panics if called — sticky-ack must skip the LLM. + provider := &stickyAckPanicProvider{t: t} +``` + +## Output Format +Return your findings as a JSON array inside a ```json code fence. Each finding must have these fields: + +- `id` (string): The rule ID that was violated, or a short kebab-case identifier for general findings. +- `file` (string): The file path where the issue was found. **Must be one of the exact paths listed in "Files in This Diff" above.** If a file path does not appear in that list, do NOT reference it. If your finding relates to a file not in the diff (e.g. a downstream consequence), set `file` and `line` to the diff location that triggers the issue and mention the affected file in `description`. +- `line` (int): The line number in the file. **Must be a line that was added or modified in the diff** (a `+` line in the diff hunk). If your finding is about a side effect on a distant line, set `line` to the diff line that *causes* the issue and describe the affected location in `description`. +- `severity` (string): One of "critical", "bug", "warning", "suggestion", or "nitpick". + - "critical": Security vulnerabilities, data loss, crashes. + - "bug": A logic error that causes incorrect runtime behavior for real inputs. Missing test coverage, unused parameters, typos in identifiers that happen to compile, or "what if a future caller…" concerns do NOT qualify — use "suggestion" or "nitpick" for those. If you cannot name the concrete input and the concrete wrong output, it is not a bug. + - "warning": Potential issues, performance problems, code smells. + - "suggestion": Better patterns, readability improvements. + - "nitpick": Minor style, naming, formatting. +- `title` (string): A short title for the finding. +- `description` (string): A concise explanation of the issue — 2-3 sentences max. State what is wrong and why it matters. Do not repeat the code or walk through the logic step by step. +- `suggestion` (string, optional): A concise suggested fix — 1-2 sentences of prose, then a code block if helpful. Do not explain what the code block does. For suggestions about broader patterns or improvements beyond the current PR scope, recommend opening a separate PR — do not imply they should fix it here. +- `fix_ref` (string): A reference ID in the format `175-` where index starts at 1 (e.g. `175-1`, `175-2`). +- `actionable` (boolean): Set to `false` if your analysis concludes the code is correct and no change is needed. Set to `true` if the finding requires the author to act. **Prefer returning an empty array over emitting findings with `actionable: false`.** + +**IMPORTANT — JSON escaping:** When your description or suggestion references code containing backslash sequences (e.g. `\n`, `\t`, `\"`), you MUST double-escape the backslash in the JSON string value. For example, to mention `fmt.Print("\n")` in a JSON string, write `fmt.Print("\\n")`. A single `\n` in JSON is a newline character, not the literal text `\n`. + +**Do not include findings where your conclusion is that the code is correct or no action is needed.** If you evaluate something and determine it is fine, omit it entirely rather than reporting it. Specifically: if you begin analyzing a potential issue but then realize the code handles it correctly, do NOT emit a finding that walks through the concern and then concludes "this is actually fine" or "no bug here" — simply drop it. Every finding you emit must represent a real, actionable problem. + +**Check against project documentation before emitting.** The "Project Documentation" section above defines conventions for this codebase (e.g. "don't add error handling for scenarios that can't happen", "keep the core engine agnostic"). Before emitting a finding, verify it does not contradict those conventions. If your suggested fix would violate a project-doc rule, drop the finding — the author has already made that tradeoff deliberately. + +**Label uncertainty from external behavior.** If your finding's validity depends on the behavior of a third-party API, webhook payload shape, framework internal, or other system you cannot verify from the diff, file contents, and project docs above, you MUST (a) cap severity at "suggestion" and (b) state the assumption in `description` (e.g. "Assumes `github.event.pull_request.number` is unset on `pull_request_review_comment` events — verify against GitHub's webhook docs before acting."). A finding that asserts external behavior as fact without this label is a false-positive risk. + +**CRITICAL: Do NOT invent or hallucinate file paths, function names, or code that does not appear in the diff or the provided file contents. If a file or function is not shown above, do not reference it.** + +If there are no findings, return an empty array: `[]`. + +Example: +```json +[ + { + "id": "rule-id", + "file": "src/main.go", + "line": 42, + "severity": "warning", + "title": "Short title", + "description": "The value is used after the error check, so a non-nil error silently proceeds with stale data.", + "suggestion": "Return early on error.\n\n```go\nif err != nil {\n return err\n}\n```", + "fix_ref": "175-1", + "actionable": true + } +] +``` From e193347c4a333839c72e4566e43f4414d0939663 Mon Sep 17 00:00:00 2001 From: Alan Sikora Date: Fri, 4 Sep 2026 20:37:38 -0300 Subject: [PATCH 2/3] fix(eval): fail when the size report is missing MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A deleted or never-committed SIZES.txt made writeSizeReport return without reporting anything, so the check passed silently — defeating the one thing the file is for. Distinguish a missing report from a matching one and fail on the former. --- internal/review/prompt_golden_test.go | 12 +++++++++--- 1 file changed, 9 insertions(+), 3 deletions(-) diff --git a/internal/review/prompt_golden_test.go b/internal/review/prompt_golden_test.go index e306304..05d1868 100644 --- a/internal/review/prompt_golden_test.go +++ b/internal/review/prompt_golden_test.go @@ -130,10 +130,16 @@ func writeSizeReport(t *testing.T, dir string, sizes []string) { return } prev, err := os.ReadFile(path) - if err != nil || string(prev) == body { - return + switch { + case os.IsNotExist(err): + // The point of this file is that prompt growth cannot pass unnoticed, + // so a missing one is a failure rather than a skip. + t.Errorf("no size report at %s; run `go test ./internal/review/ -run Golden -update`", path) + case err != nil: + t.Errorf("reading size report: %v", err) + case string(prev) != body: + t.Errorf("prompt sizes changed:\n\n%s\nwant:\n\n%s\nre-run with -update", body, prev) } - t.Errorf("prompt sizes changed:\n\n%s\nwant:\n\n%s\nre-run with -update", body, prev) } func lineAt(lines []string, i int) string { From a27dcf0d53838a735a5fb2f0489991cf9ababe26 Mon Sep 17 00:00:00 2001 From: Alan Sikora Date: Mon, 5 Oct 2026 15:55:00 -0300 Subject: [PATCH 3/3] fix(eval): adapt harness to FileContentsResult and production prompt scoping - evalsnap: use FetchFileContents' FileContentsResult; record diff_only, excluded and max_diff_size in the fixture. - TestPromptGolden: scope the PR with scopePRForPrompt before BuildPrompt, as prepareReview does, so goldens match what a real review sends. - Regenerate goldens: +761 chars per fixture from the needs_verification field and the broadened 'Label uncertainty' instruction. - README: note the incremental prompt is not covered (no thread data). Co-Authored-By: Claude --- cmd/evalsnap/main.go | 31 ++++++++++++++----- internal/evalcorpus/corpus.go | 11 +++++++ internal/review/prompt_golden_test.go | 31 ++++++++++++------- internal/review/testdata/corpus/README.md | 11 +++++++ internal/review/testdata/corpus/SIZES.txt | 6 ++-- .../alansikora-codecanary-pr165.prompt.golden | 3 +- .../alansikora-codecanary-pr173.prompt.golden | 3 +- .../alansikora-codecanary-pr175.prompt.golden | 3 +- 8 files changed, 74 insertions(+), 25 deletions(-) diff --git a/cmd/evalsnap/main.go b/cmd/evalsnap/main.go index fa26b17..604eec6 100644 --- a/cmd/evalsnap/main.go +++ b/cmd/evalsnap/main.go @@ -106,11 +106,19 @@ Pass --force to capture anyway. return fmt.Errorf("loading review config: %w", err) } - fileContents, skipped := review.FetchFileContents( + // Freeze the raw PR plus what the content reader left out, not an + // already-scoped PR: the harness applies the review's own scoping when it + // renders, so a change to that scoping shows up in the goldens instead of + // being baked into the fixture. + fc := review.FetchFileContents( prData.Files, cfg.Ignore, cfg.EffectiveMaxFileSize(), cfg.EffectiveMaxTotalSize()) - if len(skipped) > 0 { - fmt.Fprintf(os.Stderr, "skipped %d large/ignored file(s): %s\n", - len(skipped), strings.Join(skipped, ", ")) + if len(fc.Excluded) > 0 { + fmt.Fprintf(os.Stderr, "excluded %d ignored/binary file(s): %s\n", + len(fc.Excluded), strings.Join(fc.Excluded, ", ")) + } + if len(fc.DiffOnly) > 0 { + fmt.Fprintf(os.Stderr, "%d file(s) over the size limits, diff only: %s\n", + len(fc.DiffOnly), strings.Join(fc.DiffOnly, ", ")) } fixture := &evalcorpus.Fixture{ @@ -119,7 +127,7 @@ Pass --force to capture anyway. PRNumber: *pr, HeadSHA: headSHA, CapturedAt: time.Now().UTC().Format(time.RFC3339), - PR: toPRInput(prData, fileContents), + PR: toPRInput(prData, fc), Config: toConfigInput(cfg), ProjectDocs: review.ReadProjectDocs(prData.Files), } @@ -262,7 +270,7 @@ func fixtureName(explicit, repo string, pr int) string { return fmt.Sprintf("%s-pr%d", strings.ReplaceAll(repo, "/", "-"), pr) } -func toPRInput(pr *review.PRData, contents map[string]string) evalcorpus.PRInput { +func toPRInput(pr *review.PRData, fc review.FileContentsResult) evalcorpus.PRInput { return evalcorpus.PRInput{ Number: pr.Number, Title: pr.Title, @@ -272,7 +280,9 @@ func toPRInput(pr *review.PRData, contents map[string]string) evalcorpus.PRInput HeadBranch: pr.HeadBranch, Diff: pr.Diff, Files: pr.Files, - FileContents: contents, + FileContents: fc.Contents, + DiffOnly: fc.DiffOnly, + Excluded: fc.Excluded, } } @@ -290,7 +300,12 @@ func toConfigInput(cfg *review.ReviewConfig) *evalcorpus.ConfigInput { ExcludePaths: r.ExcludePaths, }) } - return &evalcorpus.ConfigInput{Rules: rules, Context: cfg.Context, Ignore: cfg.Ignore} + return &evalcorpus.ConfigInput{ + Rules: rules, + Context: cfg.Context, + Ignore: cfg.Ignore, + MaxDiffSize: cfg.MaxDiffSize, + } } func short(sha string) string { diff --git a/internal/evalcorpus/corpus.go b/internal/evalcorpus/corpus.go index 99ccbde..d362e81 100644 --- a/internal/evalcorpus/corpus.go +++ b/internal/evalcorpus/corpus.go @@ -68,6 +68,14 @@ type PRInput struct { Diff string `json:"diff"` Files []string `json:"files"` FileContents map[string]string `json:"file_contents,omitempty"` + + // DiffOnly and Excluded record what the content reader left out, as + // review.FetchFileContents reported it at capture time. They cannot be + // recomputed from the fixture (binary detection and the size budget need + // the working tree), and the review scopes its prompt from them: Excluded + // files are dropped from the file list and the diff. + DiffOnly []string `json:"diff_only,omitempty"` + Excluded []string `json:"excluded,omitempty"` } // ConfigInput mirrors the review config fields that reach the prompt. Model, @@ -78,6 +86,9 @@ type ConfigInput struct { Rules []RuleInput `json:"rules,omitempty"` Context string `json:"context,omitempty"` Ignore []string `json:"ignore,omitempty"` + // MaxDiffSize is the configured max_diff_size; zero means the review's + // default. It trims the prompt's diff, so it shapes what the prompt says. + MaxDiffSize int `json:"max_diff_size,omitempty"` } // RuleInput mirrors review.Rule. diff --git a/internal/review/prompt_golden_test.go b/internal/review/prompt_golden_test.go index 05d1868..d27aa2f 100644 --- a/internal/review/prompt_golden_test.go +++ b/internal/review/prompt_golden_test.go @@ -57,7 +57,16 @@ func TestPromptGolden(t *testing.T) { sizes := make([]string, 0, len(fixtures)) for _, f := range fixtures { t.Run(f.Name, func(t *testing.T) { - got := BuildPrompt(toPRData(f), toReviewConfig(f.Config), 0, f.ProjectDocs) + pr, cfg := toPRData(f), toReviewConfig(f.Config) + // Scope the PR the way prepareReview does before BuildPrompt, so + // the golden is the prompt a real review sends: ignored and + // binary files dropped, the diff trimmed to max_diff_size. + scopePRForPrompt(pr, FileContentsResult{ + Contents: f.PR.FileContents, + DiffOnly: f.PR.DiffOnly, + Excluded: f.PR.Excluded, + }, cfg.EffectiveMaxDiffSize()) + got := BuildPrompt(pr, cfg, 0, f.ProjectDocs) sizes = append(sizes, fmt.Sprintf("%-40s %7d", f.Name, len(got))) compareGolden(t, filepath.Join(dir, f.Name+".prompt.golden"), got) }) @@ -159,15 +168,15 @@ func truncate(s string) string { func toPRData(f *evalcorpus.Fixture) *PRData { return &PRData{ - Number: f.PR.Number, - Title: f.PR.Title, - Body: f.PR.Body, - Author: f.PR.Author, - BaseBranch: f.PR.BaseBranch, - HeadBranch: f.PR.HeadBranch, - Diff: f.PR.Diff, - Files: f.PR.Files, - FileContents: f.PR.FileContents, + Number: f.PR.Number, + Title: f.PR.Title, + Body: f.PR.Body, + Author: f.PR.Author, + BaseBranch: f.PR.BaseBranch, + HeadBranch: f.PR.HeadBranch, + Diff: f.PR.Diff, + Files: f.PR.Files, + // FileContents is set by scopePRForPrompt, as in a real review. } } @@ -185,5 +194,5 @@ func toReviewConfig(c *evalcorpus.ConfigInput) *ReviewConfig { ExcludePaths: r.ExcludePaths, }) } - return &ReviewConfig{Rules: rules, Context: c.Context, Ignore: c.Ignore} + return &ReviewConfig{Rules: rules, Context: c.Context, Ignore: c.Ignore, MaxDiffSize: c.MaxDiffSize} } diff --git a/internal/review/testdata/corpus/README.md b/internal/review/testdata/corpus/README.md index 93ff7fd..3b5835f 100644 --- a/internal/review/testdata/corpus/README.md +++ b/internal/review/testdata/corpus/README.md @@ -25,6 +25,11 @@ Whether a prompt change makes reviews better or worse is a separate question that needs labelled findings, repeated runs to establish variance, and real model calls. +It also covers only the **first-review prompt** (`BuildPrompt`). Re-pushes use +`BuildIncrementalPrompt`, which adds known issues (with author replies) and +recently resolved findings; fixtures do not capture review threads, so that +prompt is not rendered here. + ## Updating the goldens When a prompt change is intentional: @@ -51,6 +56,12 @@ go run github.com/alansikora/codecanary/cmd/evalsnap \ --repo owner/name --pr 1234 --out /path/to/corpus ``` +A fixture holds the raw PR diff and file list plus the files the content reader +left out (`excluded`: ignored or binary; `diff_only`: over the size limits). +The harness scopes the PR with the review's own `scopePRForPrompt` before +rendering, exactly as a real review does, so changes to that scoping (dropping +excluded files, trimming to `max_diff_size`) show up in the goldens. + ## Private repositories **Fixtures committed here must come from public repositories only.** A fixture diff --git a/internal/review/testdata/corpus/SIZES.txt b/internal/review/testdata/corpus/SIZES.txt index 38bee3b..f830ced 100644 --- a/internal/review/testdata/corpus/SIZES.txt +++ b/internal/review/testdata/corpus/SIZES.txt @@ -1,5 +1,5 @@ # Rendered prompt sizes, in characters. # Regenerate: go test ./internal/review/ -run Golden -update -alansikora-codecanary-pr165 70300 -alansikora-codecanary-pr173 130438 -alansikora-codecanary-pr175 86970 +alansikora-codecanary-pr165 71061 +alansikora-codecanary-pr173 131199 +alansikora-codecanary-pr175 87731 diff --git a/internal/review/testdata/corpus/alansikora-codecanary-pr165.prompt.golden b/internal/review/testdata/corpus/alansikora-codecanary-pr165.prompt.golden index 4ed3bb1..e4d743f 100644 --- a/internal/review/testdata/corpus/alansikora-codecanary-pr165.prompt.golden +++ b/internal/review/testdata/corpus/alansikora-codecanary-pr165.prompt.golden @@ -1653,6 +1653,7 @@ Return your findings as a JSON array inside a ```json code fence. Each finding m - `suggestion` (string, optional): A concise suggested fix — 1-2 sentences of prose, then a code block if helpful. Do not explain what the code block does. For suggestions about broader patterns or improvements beyond the current PR scope, recommend opening a separate PR — do not imply they should fix it here. - `fix_ref` (string): A reference ID in the format `165-` where index starts at 1 (e.g. `165-1`, `165-2`). - `actionable` (boolean): Set to `false` if your analysis concludes the code is correct and no change is needed. Set to `true` if the finding requires the author to act. **Prefer returning an empty array over emitting findings with `actionable: false`.** +- `needs_verification` (boolean, optional): Set to `true` when the finding is an open question you could not settle — its validity depends on code, callers, configuration, or behavior that is not in the diff, file contents, or project docs above (e.g. "is anything else still calling the removed helper?", "does the dispatcher pass this argument through?"). These are not posted as review threads and do not block the PR; they are listed as open questions in the review summary. Omit the field when the diff and files above are enough to confirm the issue. Never title a finding "Verify …" without setting this flag. **IMPORTANT — JSON escaping:** When your description or suggestion references code containing backslash sequences (e.g. `\n`, `\t`, `\"`), you MUST double-escape the backslash in the JSON string value. For example, to mention `fmt.Print("\n")` in a JSON string, write `fmt.Print("\\n")`. A single `\n` in JSON is a newline character, not the literal text `\n`. @@ -1660,7 +1661,7 @@ Return your findings as a JSON array inside a ```json code fence. Each finding m **Check against project documentation before emitting.** The "Project Documentation" section above defines conventions for this codebase (e.g. "don't add error handling for scenarios that can't happen", "keep the core engine agnostic"). Before emitting a finding, verify it does not contradict those conventions. If your suggested fix would violate a project-doc rule, drop the finding — the author has already made that tradeoff deliberately. -**Label uncertainty from external behavior.** If your finding's validity depends on the behavior of a third-party API, webhook payload shape, framework internal, or other system you cannot verify from the diff, file contents, and project docs above, you MUST (a) cap severity at "suggestion" and (b) state the assumption in `description` (e.g. "Assumes `github.event.pull_request.number` is unset on `pull_request_review_comment` events — verify against GitHub's webhook docs before acting."). A finding that asserts external behavior as fact without this label is a false-positive risk. +**Label uncertainty.** If your finding's validity depends on something you cannot confirm from the diff, file contents, and project docs above — a third-party API, webhook payload shape, framework internal, or code and callers not shown — you MUST (a) cap severity at "suggestion", (b) state the assumption in `description` (e.g. "Assumes `github.event.pull_request.number` is unset on `pull_request_review_comment` events — verify against GitHub's webhook docs before acting."), and (c) set `needs_verification: true`. A finding that asserts unverified behavior as fact is a false-positive risk. If the diff and files above already show the problem, it is not uncertain — report it as a regular finding without the flag. **CRITICAL: Do NOT invent or hallucinate file paths, function names, or code that does not appear in the diff or the provided file contents. If a file or function is not shown above, do not reference it.** diff --git a/internal/review/testdata/corpus/alansikora-codecanary-pr173.prompt.golden b/internal/review/testdata/corpus/alansikora-codecanary-pr173.prompt.golden index 9ceaa0f..9effb79 100644 --- a/internal/review/testdata/corpus/alansikora-codecanary-pr173.prompt.golden +++ b/internal/review/testdata/corpus/alansikora-codecanary-pr173.prompt.golden @@ -2696,6 +2696,7 @@ Return your findings as a JSON array inside a ```json code fence. Each finding m - `suggestion` (string, optional): A concise suggested fix — 1-2 sentences of prose, then a code block if helpful. Do not explain what the code block does. For suggestions about broader patterns or improvements beyond the current PR scope, recommend opening a separate PR — do not imply they should fix it here. - `fix_ref` (string): A reference ID in the format `173-` where index starts at 1 (e.g. `173-1`, `173-2`). - `actionable` (boolean): Set to `false` if your analysis concludes the code is correct and no change is needed. Set to `true` if the finding requires the author to act. **Prefer returning an empty array over emitting findings with `actionable: false`.** +- `needs_verification` (boolean, optional): Set to `true` when the finding is an open question you could not settle — its validity depends on code, callers, configuration, or behavior that is not in the diff, file contents, or project docs above (e.g. "is anything else still calling the removed helper?", "does the dispatcher pass this argument through?"). These are not posted as review threads and do not block the PR; they are listed as open questions in the review summary. Omit the field when the diff and files above are enough to confirm the issue. Never title a finding "Verify …" without setting this flag. **IMPORTANT — JSON escaping:** When your description or suggestion references code containing backslash sequences (e.g. `\n`, `\t`, `\"`), you MUST double-escape the backslash in the JSON string value. For example, to mention `fmt.Print("\n")` in a JSON string, write `fmt.Print("\\n")`. A single `\n` in JSON is a newline character, not the literal text `\n`. @@ -2703,7 +2704,7 @@ Return your findings as a JSON array inside a ```json code fence. Each finding m **Check against project documentation before emitting.** The "Project Documentation" section above defines conventions for this codebase (e.g. "don't add error handling for scenarios that can't happen", "keep the core engine agnostic"). Before emitting a finding, verify it does not contradict those conventions. If your suggested fix would violate a project-doc rule, drop the finding — the author has already made that tradeoff deliberately. -**Label uncertainty from external behavior.** If your finding's validity depends on the behavior of a third-party API, webhook payload shape, framework internal, or other system you cannot verify from the diff, file contents, and project docs above, you MUST (a) cap severity at "suggestion" and (b) state the assumption in `description` (e.g. "Assumes `github.event.pull_request.number` is unset on `pull_request_review_comment` events — verify against GitHub's webhook docs before acting."). A finding that asserts external behavior as fact without this label is a false-positive risk. +**Label uncertainty.** If your finding's validity depends on something you cannot confirm from the diff, file contents, and project docs above — a third-party API, webhook payload shape, framework internal, or code and callers not shown — you MUST (a) cap severity at "suggestion", (b) state the assumption in `description` (e.g. "Assumes `github.event.pull_request.number` is unset on `pull_request_review_comment` events — verify against GitHub's webhook docs before acting."), and (c) set `needs_verification: true`. A finding that asserts unverified behavior as fact is a false-positive risk. If the diff and files above already show the problem, it is not uncertain — report it as a regular finding without the flag. **CRITICAL: Do NOT invent or hallucinate file paths, function names, or code that does not appear in the diff or the provided file contents. If a file or function is not shown above, do not reference it.** diff --git a/internal/review/testdata/corpus/alansikora-codecanary-pr175.prompt.golden b/internal/review/testdata/corpus/alansikora-codecanary-pr175.prompt.golden index 2d69ffa..9b9ab5f 100644 --- a/internal/review/testdata/corpus/alansikora-codecanary-pr175.prompt.golden +++ b/internal/review/testdata/corpus/alansikora-codecanary-pr175.prompt.golden @@ -1829,6 +1829,7 @@ Return your findings as a JSON array inside a ```json code fence. Each finding m - `suggestion` (string, optional): A concise suggested fix — 1-2 sentences of prose, then a code block if helpful. Do not explain what the code block does. For suggestions about broader patterns or improvements beyond the current PR scope, recommend opening a separate PR — do not imply they should fix it here. - `fix_ref` (string): A reference ID in the format `175-` where index starts at 1 (e.g. `175-1`, `175-2`). - `actionable` (boolean): Set to `false` if your analysis concludes the code is correct and no change is needed. Set to `true` if the finding requires the author to act. **Prefer returning an empty array over emitting findings with `actionable: false`.** +- `needs_verification` (boolean, optional): Set to `true` when the finding is an open question you could not settle — its validity depends on code, callers, configuration, or behavior that is not in the diff, file contents, or project docs above (e.g. "is anything else still calling the removed helper?", "does the dispatcher pass this argument through?"). These are not posted as review threads and do not block the PR; they are listed as open questions in the review summary. Omit the field when the diff and files above are enough to confirm the issue. Never title a finding "Verify …" without setting this flag. **IMPORTANT — JSON escaping:** When your description or suggestion references code containing backslash sequences (e.g. `\n`, `\t`, `\"`), you MUST double-escape the backslash in the JSON string value. For example, to mention `fmt.Print("\n")` in a JSON string, write `fmt.Print("\\n")`. A single `\n` in JSON is a newline character, not the literal text `\n`. @@ -1836,7 +1837,7 @@ Return your findings as a JSON array inside a ```json code fence. Each finding m **Check against project documentation before emitting.** The "Project Documentation" section above defines conventions for this codebase (e.g. "don't add error handling for scenarios that can't happen", "keep the core engine agnostic"). Before emitting a finding, verify it does not contradict those conventions. If your suggested fix would violate a project-doc rule, drop the finding — the author has already made that tradeoff deliberately. -**Label uncertainty from external behavior.** If your finding's validity depends on the behavior of a third-party API, webhook payload shape, framework internal, or other system you cannot verify from the diff, file contents, and project docs above, you MUST (a) cap severity at "suggestion" and (b) state the assumption in `description` (e.g. "Assumes `github.event.pull_request.number` is unset on `pull_request_review_comment` events — verify against GitHub's webhook docs before acting."). A finding that asserts external behavior as fact without this label is a false-positive risk. +**Label uncertainty.** If your finding's validity depends on something you cannot confirm from the diff, file contents, and project docs above — a third-party API, webhook payload shape, framework internal, or code and callers not shown — you MUST (a) cap severity at "suggestion", (b) state the assumption in `description` (e.g. "Assumes `github.event.pull_request.number` is unset on `pull_request_review_comment` events — verify against GitHub's webhook docs before acting."), and (c) set `needs_verification: true`. A finding that asserts unverified behavior as fact is a false-positive risk. If the diff and files above already show the problem, it is not uncertain — report it as a regular finding without the flag. **CRITICAL: Do NOT invent or hallucinate file paths, function names, or code that does not appear in the diff or the provided file contents. If a file or function is not shown above, do not reference it.**