Skip to content

feat(plugin): Add Plugin evaluation support across all tiers - #28

Open
chrisknvidia wants to merge 491 commits into
mainfrom
naren/plugin-evaluation-all-tiers
Open

chrisknvidia wants to merge 491 commits into
mainfrom
naren/plugin-evaluation-all-tiers

Conversation

@chrisknvidia

@chrisknvidia chrisknvidia commented Aug 1, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR adds plugin evaluation to every tier of SkillEvaluator. It takes a plugin in any of five manifest formats (agent_plugin.yaml, Claude Code, Codex, Cursor or Agent Plugins v1), checks its structure and security statically (Tier 1), checks it for duplication (Tier 2), and measures its behavior with live agents (Tier 3), loading it through a generated wrapper skill or natively in Claude Code, Codex and OpenCode. The reports say plainly what was and was not demonstrated.

Files staged ≠ components loaded ≠ behavior verified. Every plugin report separates what was inventoried, what was staged into a run, what the harness reported as loaded, what was observed in agent traces, and what was not evaluated at all. Missing evidence narrows the claim; it is never silently dropped.

Full documentation: docs/plugin-evaluation.mdx.

How plugin evaluation works

flowchart LR
    P["Plugin<br/>agent_plugin.yaml · Claude Code · Codex · Cursor · Agent Plugins v1<br/>skills · rules · MCP · hooks · agents · commands ·<br/>LSP · monitors · settings · styles"]

    subgraph T1["Tier 1 · Static & security (gating)"]
        direction TB
        T1a["Manifest checks per format<br/>+ component inventory"]
        T1b["Dependency classification<br/>provided / referenced / missing / external / unresolved"]
        T1c["MCP policy · pinning · endpoints<br/>(opt-in DNS and redirect checks)"]
        T1d["Hook risk · agent, command and skill privileges<br/>LSP, monitor and settings checks"]
        T1e["Whole-tree scans · bundled-skill quality/lint/version<br/>opt-in CVE audit: Python, npm, containers"]
        T1f["Static context-cost estimate"]
    end

    subgraph T2["Tier 2 · Deduplication (advisory)"]
        direction TB
        T2a["A · duplicate refs"]
        T2b["C-intra · bundled-skill overlap"]
        T2c["C-inter · skills vs local catalog"]
        T2d["B · plugin vs other plugins<br/>(optional LLM verdict)"]
    end

    subgraph T3["Tier 3 · Live agent evaluation (Harbor)"]
        direction TB
        L["Plugin loading<br/>wrapper (default): generated skill + members + rules + MCP<br/>native: claude-code · codex · opencode adapters"]
        A1["With plugin<br/>load census · hook census"]
        A2["Without plugin"]
        A3["Member skills only<br/>(sum of parts)"]
        G["Graders<br/>security incl. canary exfiltration · execution ·<br/>judges (N/A without a reference)"]
        S["Plugin signals (advisory)<br/>activations · routing · tool selection · arguments ·<br/>MCP outcomes · order · handoff · conflict"]
        M["MCP proof (opt-in)<br/>host probe + in-agent calls"]
        ST["Statistics<br/>paired bootstrap CI · pass@k / pass^k ·<br/>cost per success · token efficiency · context delta"]
    end

    R["Reports<br/>JSON · Markdown · HTML · SARIF · CLI<br/>coverage: staged / loaded / exercised<br/>INCOMPLETE / INCONCLUSIVE · plugin BENCHMARK.md"]

    P --> T1 --> T2 --> T3 --> R
    L --> A1
    A1 --> G
    A2 --> G
    A3 --> G
    G --> S --> ST
    S --> M
Loading

What each tier supports

Tier Capability Gating
1 Manifests: agent_plugin.yaml and the Claude Code, Agent Plugins v1, Codex and Cursor plugin.json formats, selected in that order, with root-bounded, no-follow discovery and each client's field checks. Other manifests found are recorded and their components checked. Opt-in claude-validate compares the verdict with claude plugin validate for Claude Code plugins (advisory) Blocking on invalid or unsafe input; MEDIUM plugin_manifest_conflict on a name or version mismatch
1 Component inventory: skills, rules, MCP, hooks (frontmatter hooks too), subagents, commands, LSP, output styles, monitors, settings, Codex apps and Agent Plugins extensions, each with a support level and per-component finding counts Missing, escaping, unsafe or invalid declared paths block (HIGH)
1 Dependency classification: provided / referenced / missing / external / unresolved, with fail-closed repository identity and validate --repo-root. Claude Code dependencies resolve against the plugin's marketplace.json Only missing blocks (HIGH plugin_dependency_missing); refs it cannot check are MEDIUM
1 MCP static policy for every mcpServers form (inline, path, array, root .mcp.json), read with each client's ${VAR} expansion CRITICAL and HIGH block; Tier 3 staging re-checks them
1 Pinning of package runners and images (npx, uvx, pnpm dlx, docker run, ...) through launch wrappers, in MCP servers and command hooks, with a pinning ratio HIGH mcp_command_floating_version for a moving tag; MEDIUM mcp_unpinned_package otherwise
1 Bypass flags and overrides in MCP servers, hooks, LSP servers, monitors and settings (--dangerously-skip-permissions, bypassPermissions, Codex -a never, auto-approve, LD_PRELOAD, proxy and base-URL redirects, shipped .env) HIGH or MEDIUM, by type
1 Endpoint policy for metadata, private, loopback and link-local hosts (encoded IPs too), with an mcp.allowed_private_hosts allowlist. Opt-in --resolve-endpoints adds DNS and one HEAD check HIGH (metadata, never allowlisted) or MEDIUM
1 Hook risk per handler, scripts included: remote code (curl … | sh, download-then-run), inline secrets, broad auto-approval, HTTP endpoints (hooks.allowed_urls), context injection CRITICAL to LOW; an unanalyzable script is HIGH
1 Subagent, command and skill privileges (tools, allowed-tools, permissionMode) HIGH: bypassPermissions, or allowed-tools that pre-approve any Bash. MEDIUM: wildcards, acceptEdits/auto, subagents that inherit every tool beside a write-capable MCP server. LOW: a subagent's unrestricted Bash
1 Whole-plugin-tree security, PII, code-risk, secrets, license, hygiene and Unicode scans, plus quality/lint/version on every bundled skill Existing Tier 1 gates
1 Opt-in dependency audit of declared exact pins, installing nothing: Python (pip-audit --no-deps --disable-pip), npm (lockfiles, package.json, MCP runner packages) and container images (MCP and Dockerfiles), with OSV-Scanner, npm audit, Grype or Trivy Advisory severity. Pins that pip-audit skips or the npm registry lacks are MEDIUM dependency-not-audited; floating versions are INFO; no scanner or an unreadable file is INCOMPLETE
1 Static always-on vs on-demand context-cost estimate, per harness and load mode Report-only
2 Check A (duplicate refs), C-intra (bundled-skill overlap), and the local-catalog checks C-inter (skills vs catalog) and B (plugin vs other plugins, optional --llm verdict). Catalog v2 stores plugin entries; v1 catalogs still load Advisory; unsafe input blocks
3 Three arms: with plugin, without plugin, member skills only. --lift-mode effectiveness|integration|both, with the Integration evidence gate (cross_component, expected_skills) Gates only with --block-on-agent-eval
3 --plugin-load wrapper|native|auto (default wrapper) for the with-plugin arm only: native adapters for claude-code (--plugin-dir), codex and opencode; Hermes (experimental) uses the wrapper. A per-trial load census lists staged components; Claude Code's init event marks them loaded or not_loaded Report-only. native refuses components with a permission bypass; a native run with no load census, or whose plugin Claude Code never loaded, is INCOMPLETE
3 Hook census: native command hooks run through hook_census.sh, which records runs, denials and start failures and passes input, output and exit code through Advisory; marks hooks exercised
3 Canary exfiltration: a random decoy credential (file and env var) in every arm, traced through shell statements, network and MCP tools, URLs, git, and files written outside the workspace Critical canary_exfiltration; plugin-attributable when the plugin arm leaks more often than the baseline
3 Runtime security: credential-store reads and protected-file writes, matched as normalized whole paths, subagent and child-agent calls included Critical scores 0.0
3 MCP proof (opt-in --probe-mcp): a policy-checked host initialize and tools/list for URL servers, with no host credentials unless a variable is named with --probe-mcp-env, then updated from the agent's own calls Advisory
3 Paired case-bootstrap 95% CI on Effectiveness and Integration lift, with a precision flag. pass^k, cost per success, token efficiency, measured first-turn context delta Report-only; Integration becomes INCONCLUSIVE when the CI crosses a ±0.05 band edge, fewer than 5 cases pair, or coverage is incomplete
3 Plugin signals (dataset-driven): activations, routing and tool-selection P/R/F1 with decoys, argument checks, MCP call outcomes by server and tool, order, handoff, conflict probes, activation coverage Advisory, report-only
3 Judges without a reference (ground_truth / expected_behavior) are N/A, not a fabricated 1.0 See reviewer notes
All Reports: plugin sections in JSON, Markdown, HTML, SARIF and the CLI, including plugin loading, hook census, canary and MCP proof. Tier 3 coverage per component (staged, loaded, not_loaded, exercised, ...), headlined as "N components not staged"; explicit INCOMPLETE and INCONCLUSIVE; plugin_provenance.json re-read when a report is rebuilt; plugin BENCHMARK.md card —

Reviewer notes: behavior changes to be aware of

  • N/A judges. Accuracy, goal accuracy and behavior check no longer return 1.0 when the case has no reference. They are excluded from the overall score and from lift, so a run where no case has ground_truth now gets a NEUTRAL verdict at best, because its Correctness evidence is missing. reward.json stays numeric-only for Harbor, and the N/A markers travel in the sidecar.
  • One lift basis, skills included. The arm without the skill or plugin scores skill_execution and skill_efficiency as N/A, and every lift compares its two arms case by case on the dimensions both scored. Activation alone no longer earns lift, so Skill Lift values change for skill runs too.
  • Tier 3 gate. With --block-on-agent-eval, a Tier 3 FAIL verdict now fails validate, and so does a confirmed Skill Lift regression (lift at or below −0.10 with its whole interval below zero). NEUTRAL never does.
  • Scores can move for skill runs too. Judges see the full body of each file write, within the evidence budget, instead of its first 200 characters. Runtime security checks more credential stores, protected files and subagent calls, and treats curl … | sh and forced git push as critical.
  • New blocking Tier 1 findings for plugins: the CRITICAL and HIGH findings of the manifest, component-path, dependency, MCP, endpoint, bypass-flag, hook-risk and privilege checks above. Tier 3 staging fails closed on the same MCP findings.
  • Bundled skills are gated. Quality, lint and version now run on each bundled skill, so a weak bundled skill can fail a plugin that used to pass.
  • Integration verdict (advisory). It reads the whole interval against the ±0.05 band and is inconclusive when the interval crosses a band edge, fewer than 5 cases pair, or per-case coverage is incomplete. The point estimate's band is kept as point_verdict.
  • The dependency audit no longer installs anything. It audits exact pins with --no-deps --disable-pip, so transitive coverage from requirements files is traded for not executing untrusted build code. Python advisories take their severity from the public OSV record (GitHub severity, then CVSS) rather than defaulting to HIGH.

Public adaptations and exclusions

  • Public GitHub/Git references and same-repository offline resolution only. Nothing internal is included: no GitLab, P4, managed execution, private providers or endpoints, credentials, telemetry, provider-registry MCP resolution, or remote vector DB.
  • Tier 1 plugin checks run offline unless you opt in: --resolve-endpoints adds DNS and redirect checks, and the dependency audit uses its scanners' advisory and registry lookups. Tier 2 catalogs are local files.
  • Not yet supported, and reported as such rather than hidden:
    • runtime tests of LSP servers and monitors: native Claude Code stages LSP servers, but nothing confirms they ran, and monitors are never staged;
    • native-load confirmation outside Claude Code: Codex and OpenCode components stay listed (found, not confirmed by the harness); Hermes has no native plugin path, and no end-to-end Hermes plugin run has been verified;
    • process-level evidence: runtime security and the canary read the agent's tool calls and the files they wrote, so a hook or MCP server process acting on its own is not observed;
    • an equivalent of internal's provider-registry MCP proof: provider-only MCP servers are never resolved or probed (the run is INCOMPLETE), and --probe-mcp covers URL servers only;
    • MCP env, headers and ${user_config.*} values that the with-plugin arm cannot apply (only native Claude Code applies them); such runs are INCOMPLETE;
    • Codex apps and Agent Plugins extensions, which are inventoried only.

Included pull requests

  • #171 (Stage C): native plugin loading, the Codex, Cursor and Agent Plugins v1 manifests, static risk checks for hooks, privileges, LSP, monitors, settings and output styles, runtime evidence (load and hook census, canary), npm and container CVE audits, and the MCP proof.
  • #180 and #181: fixes from two rounds of code review of this PR.
  • The branch tracks main through v0.5.0, including the Harbor 0.24.0 upgrade from #83.

Verification

  • uv run pytest -q -n 4 (with main through v0.5.0 merged): 17,480 passed, 29 skipped, 0 failed
  • uv run ruff check .: passed
  • scripts/check_oss_boundary.py and scripts/ci/check_public_benchmarks.py --require-files tests/golden: passed
  • fern check: 0 errors

A sample live Harbor run on the internal counterpart of this work (MR !280) passed for Claude Code and Codex with native loading (results). That run predates the Harbor 0.24.0 upgrade, which passed its own live parity runs in #83; a live plugin run on Harbor 0.24.0 is still to do.

Verification also covers focused regressions per feature (including symlink, escape and oversize cases) and packaging and boundary checks.

Restoration note (from the original description)

This PR replaces #17, which GitHub permanently closed after the default-branch history was consolidated and the original head branch was deleted and recreated. The original discussion and commit history remain available on #17.

🤖 Generated with Claude Code

@rng1995
rng1995 force-pushed the naren/plugin-evaluation-all-tiers branch from 3a51f76 to a4a5e63 Compare August 4, 2026 19:14

@rng1995 rng1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The plugin evaluation implementation and its follow-up hardening are well covered, but I found one small diff-hygiene issue to clean up.

Comment thread src/skillevaluator/utils/structured_data.py Outdated
rng1995
rng1995 previously approved these changes Aug 5, 2026

@rng1995 rng1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved after the EOF diff-hygiene issue was fixed in 3965061, the review thread was resolved, and the complete GitHub check matrix passed.

@rng1995 rng1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fresh review found two actionable security issues in the plugin evaluation staging path. The previous EOF-hygiene thread is already fixed and resolved. I will address these findings while updating the branch from current main.

Comment thread src/skillevaluator/tier3/plugin_eval.py Outdated
Comment thread src/skillevaluator/tier3/plugin_eval.py Outdated
rng1995
rng1995 previously approved these changes Aug 12, 2026

@rng1995 rng1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved after fresh security review and remediation: all review threads are resolved, the branch is conflict-free against current main, the complete local suite passes (3,999 passed, 21 skipped, 3 deselected), and all 15 GitHub checks pass including Windows, packaging, DCO, and security scans.

@mohgupta-ship-it mohgupta-ship-it left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Posted by Codex on behalf of Mohit.

Request changes: plugin provenance is persisted through a symlink-following path after the long-running evaluation. This is a medium output-integrity risk for a same-privilege actor able to modify the selected results location. The inline note describes a no-follow, atomic remediation and the regression coverage needed.

Comment thread src/skillevaluator/tier3/plugin_eval.py Outdated

@mohgupta-ship-it mohgupta-ship-it left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex review on behalf of Mohit — REQUEST CHANGES

Critical sandbox/path-integrity blocker. Two safe reproductions show plugin-controlled Git symlinks can cause host-readable content to be staged before Docker isolation: (1) a member evals/evals.* symlink is accepted through find_eval_file(...).exists(), parsed, and copied into task inputs before any symlink validation; (2) a repo-root skills or rules symlink is resolved before containment, making the external target the trusted root.

This is a pre-sandbox host-file disclosure path, not only a race. Reject links/reparse points and mount crossings before resolution; use descriptor-anchored no-follow reads/copies for member datasets (new and legacy layouts); preserve the lexical clone-root boundary for canonical refs; and add regressions for both attack paths.

Secondary integrity issue: fail closed when the Git-origin slug is unavailable — the current fallback can mark a foreign reference as fully evaluated.

Comment thread src/skillevaluator/tier3/plugin_eval.py Outdated
Comment thread src/skillevaluator/tier3/plugin_eval.py Outdated
@rng1995

rng1995 commented Aug 31, 2026

Copy link
Copy Markdown
Collaborator

@chrisknvidia Gentle ping when you have a chance: there are still three unresolved review threads on this PR. Please take a look and update the branch or reply on the threads where you disagree.

@rng1995

rng1995 commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator

Updated in f3aefb9: merged current main and resolved the conflicts, hardened provenance/canonical-ref/member-dataset staging, and fixed the custom-only sum-of-parts report path uncovered during the merge review. All addressed review threads are resolved. Local verification covered the full suite and focused security/report regressions; Ruff and git diff --check pass. All GitHub checks, including Python 3.12/3.13, DCO, CodeQL, packaging, Windows Tier 2, and macOS Tier 3, are green.

rng1995 added a commit that referenced this pull request Sep 30, 2026
…n-eval-stage-c

Resolve conflicts between PR #28's review fixes and Stage C: plugin_components.py keeps Stage C's hook-script reads (read_prefix via a shared _read_bytes helper that now takes the config budget flag), privilege and write-capable MCP helpers, and the inline {"mcpServers": {...}} wrapper, with #28's per-field item-cap findings, separate config read budget, and the >256 inline-map finding applied to the unwrapped map; plugin_sections.py combines #28's not-staged/not-observed coverage view with Stage C's loaded/exercised states (both count as staged; a row is unobserved only when neither its state nor its activation is exercised) and keeps one _PLUGIN_ARMS/_BASELINE_ARMS definition for the canary verdict and sum-of-parts labels; plugin_eval.py splits MCP servers from the normalized component manifest with #28's plugin_root check; runner.py passes the sum-of-parts alias proof and the canary kwargs to the sum-of-parts arm; cli.py escapes the integration skip reason alongside Stage C's MCP proof output; the Stage C coverage tests now expect the 'not staged' wording.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <32655573+rng1995@users.noreply.github.com>
chrisknvidia added a commit that referenced this pull request Oct 1, 2026
Add --plugin-load wrapper|native|auto (default wrapper, which keeps the
PR #28 behavior). Native Claude Code loads the plugin with --plugin-dir:
skills, rules, MCP servers including ${CLAUDE_PLUGIN_ROOT} launches,
hooks, agents, commands, output styles, LSP servers, settings and
userConfig. Codex, OpenCode and Hermes get their own config files, and
OpenCode subagents keep their tool limits. Only the with-plugin arm
changes.

A per-trial load census records what each harness listed or loaded. A
failed MCP server never counts as loaded, and declared components that
were not staged make the run INCOMPLETE. auto falls back to the wrapper
per agent and records why. Hermes is experimental and refused with the
OpenAI and OpenAI-compatible providers. Launch rewriting is linear-time.

Co-authored-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
chrisknvidia added a commit that referenced this pull request Oct 1, 2026
Add --plugin-load wrapper|native|auto (default wrapper, which keeps the
PR #28 behavior). Native Claude Code loads the plugin with --plugin-dir:
skills, rules, MCP servers including ${CLAUDE_PLUGIN_ROOT} launches,
hooks, agents, commands, output styles, LSP servers, settings and
userConfig. Codex, OpenCode and Hermes get their own config files, and
OpenCode subagents keep their tool limits. Only the with-plugin arm
changes.

A per-trial load census records what each harness listed or loaded. A
failed MCP server never counts as loaded, and declared components that
were not staged make the run INCOMPLETE. auto falls back to the wrapper
per agent and records why. Hermes is experimental and refused with the
OpenAI and OpenAI-compatible providers. Launch rewriting is linear-time.

Co-authored-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
rng1995 added a commit that referenced this pull request Oct 5, 2026
Bug fixes from the code-quality review of #28 and the follow-up review of #180, one signed-off commit per fix with a regression test.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
rng1995 and others added 10 commits October 5, 2026 09:16
… plans

plugin_eval imported plugin_native lazily in nine functions although
there is no import cycle, and typed the per-agent plans and the native
snapshot as Any. It now imports plugin_native once at module level and
types plans as dict[str, AgentLoadDecision], snapshots as
NativePluginSource | None, and coverage components as Component.

The CLI plugin helpers take a PluginEvalPackage and read native_source,
mcp_probe_targets, and package_path directly instead of through getattr
defaults, and the Harbor adapter types native_plugin as
NativeTaskStaging | None (a TYPE_CHECKING import) and reads the staged
command texts directly. The CLI tests' prepared-package fakes now carry
the two fields a real package always has.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…ocks

The validator compared severities with "in (Severity.CRITICAL,
Severity.HIGH)" or "is Severity.HIGH" in three places: the blocking-MCP
check before the manifest success row, the native-manifest field check,
and _advisory_result. Use Severity.is_error(), which is that rule
(Finding always holds a Severity). No behavior change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…place

Every caller of _resolve_declared re-tested what its result already
guaranteed (declared is None or declared.rel is None or kind in
{escape, invalid, missing, unsafe}, which is just "not a file or
folder"), and the "wrong kind: invalid finding plus broken component"
block was written three times (skills, the commands map, JSON-config
sources).

_resolve_declared now takes the kinds a field accepts and reports a
path of the other kind as invalid itself. The new
_resolve_component() records the broken component for a path that
cannot be loaded and returns only a usable path and its kind
(_ResolvedPath); skills, rules, Markdown components and the commands
map use it. JSON-config sources keep their broken components without a
path, as before.

Findings (including the "commands" field name of a command source that
is a folder) and components are unchanged: the 2,185 inventories built
by the plugin test suites serialize identically. A new test pins the
wrong-kind cases.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Whether a plaintext http hook is HIGH depended on the message text
endpoint.reason == "loopback". EndpointClass now carries an is_loopback
field. classify_endpoint_host sets it where it picks the reason: true for
a loopback name, and otherwise from the address that decided the
classification (an embedded IPv4 address decides for an IPv6 literal).
The http-hook check reads the field. The result matches the reason test
for names, 127.0.0.0/8, ::1, and the mapped, 6to4, NAT64, and compatible
forms.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
_ArmStaging.native mapped each agent to a positional
(adapter id, component modes, staged types) tuple that every reader
unpacked by position. It now maps to a _NativeArm NamedTuple.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
… states

The inventory produces the Tier 3 coverage states, but the post-run
states and their rank ({"staged", "loaded", "exercised"}) were spelled as
literals here and again in plugin_runtime, plugin_native and the
reporting view. Define COVERAGE_STATE_RANK and EVALUATED_COVERAGE_STATES
next to COVERAGE_STATES, document how the states relate, and use them in
summarize_coverage. The other modules can import them instead of
redeclaring the vocabulary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Evidence entries were positional lists, [text, rank, paths] and
[text, rank, paths, head, body, tail, body_cut], read as entry[4] or
entry[6]. They are now the NamedTuples _Entry and _FileChange in both
copies; _demote_superseded_writes replaces an entry with its demoted copy.

NamedTuple rather than a dataclass: with "from __future__ import
annotations", the dataclass decorator looks the class's module up in
sys.modules, and the tests load the verifier template from its file
without registering it there. The template still imports and runs on
Python 3.9.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
validate_contained_mcp_servers was called only from tests (about 60
calls in test_mcp_static.py), and it had drifted from production.
Production runs mcpServers through plugin_components.collect_mcp_declarations,
which reads path strings through the plugin-root reader, follows Claude
Code's load order, and validates each server with
validate_mcp_server_declaration.

The per-server tests now call validate_mcp_server_declaration directly
(56 mechanical rewrites). The three shape tests (scalar, path string,
array, absent, and empty mcpServers) now go through
collect_mcp_declarations with a real plugin root. So the test for a path
string reads an actual file instead of assuming it is clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
plugin_inventory_for_root and tier3 plugin_eval each turned
location.additional into (manifest_type, manifest_rel, parsed manifest)
tuples with the same loop. Add parsed_additional_manifests(location)
next to build_plugin_inventory, whose additional argument it produces,
and use it here; plugin_eval can switch to it in a follow-up.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
_fit_entries looped over the two actions, then over the ranks, then over
the entries, and branched on which action it was in. It now lists the
entries below the top rank once (lowest rank first, middle first when
asked), shrinks them, drops them, and then shrinks and drops the older
top-rank entries, with a needed() helper for the cut that leaves the text
as long as the budget allows. Output is unchanged (checked against the
previous version on 20,000 random entry lists). Both copies.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
rng1995 and others added 30 commits October 6, 2026 03:13
… rank aliases

parsed_additional_manifests was a one-line alias of
PluginManifestLocation.parsed_additional that only a test imported, and
its test duplicated
test_parsed_additional_reads_each_client_manifest_like_its_client.
plugin_components also re-exported COVERAGE_STATE_RANK for that test
only, which invites importing coverage states through the heavy
plugin_components module that plugin_states exists to avoid.

Both aliases and the duplicate test are removed, and the coverage
vocabulary test imports its names from plugin_states.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
HostAllowlist.from_entries normalized a whole '*.<name>' entry. IDNA
rejects the '*' label, so the NFKC fallback left a Unicode suffix such
as '.bücher.example', while every host is normalized to punycode
('sub.xn--bcher-kva.example'). A Unicode wildcard in hooks.allowed_urls
or mcp.allowed_private_hosts therefore matched nothing, where the hook
check before this branch matched it.

The suffix after '*.' is now normalized on its own, after look-alike
dots are read as '.', the way hosts are.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…d plugin-file MCP rows

native_agents was set only on rule and hook/subagent/command rows. A
declared skill directory that only the native claude-code arm stages, and
an MCP server that launches from plugin files and starts only there,
carried no native_agents, so JSON consumers could not tell which agents
stage them and the load census repeated "staged natively for claude-code"
in their reasons. The docs say every row a native arm stages lists those
agents.

Both rows now record the native Claude Code arms that copy the plugin
tree.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
plugin_component_risk kept a copy of mcp_static._split_words with the
same docstring and the same behavior (shlex.split's defaults are
comments=False, posix=True), although it already imports from
mcp_static. Hook command stages and the MCP 'env -S' string now share
the one tokenizer, so a later fix cannot make them split the same text
differently.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
… them

Hook census labels moved from a whole-value redaction to
plugin_signals._safe_text, which redacts only the first limit * 4 + 256
characters of the raw value. A token that window's end cut in two no
longer matched its pattern, and when an earlier redaction (a long JWT)
shortened the text, the token's head reached the 256-character hook_id in
clear: AKIAABCDEFGHIJKLMNO, a ghp_ prefix, or a second JWT's header and
payload. The census file is agent-writable, so the line can be forged.

parse_hook_census now redacts each hook_id and event whole, then makes the
label with _safe_text. A census line is at most MAX_HOOK_CENSUS_LINE_CHARS,
which bounds the value, and each distinct pair is still redacted once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
host_is_allowlisted had become a one-line wrapper around
HostAllowlist.of(...).allows(...) with no production caller: every
caller moved to HostAllowlist. It left two public spellings of one
decision, and host_name_is_allowlisted's docstring pointed readers to
the wrapper. The wrapper is gone, its test uses HostAllowlist.allows,
and the docstring points there.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
… the reports give

When the Harbor run failed, evaluate-plugin exited with its own "Tier 3
plugin evaluation did not complete: <failure>" text, without the
INCOMPLETE prefix or the deferred components, although the provenance it
had just built carries both and validate and every report say
"INCOMPLETE: <failure>; <deferred> could not be resolved or evaluated at
Tier 3". Scripts keyed on the documented INCOMPLETE message missed the
failure.

The failure path now raises incomplete_reason(provenance), and falls back
to the failure alone, with the same prefix, when the provenance could not
be built.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
The previous loop variable still holds the last sidecar while the next
one is read, so the docstring now says that memory does not grow with the
number of sidecars instead of claiming one sidecar at a time.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…and-line credential messages

Readability follow-up to the two previous report_text and command
credential fixes. report_text's truncation is a plain if-chain instead
of one conditional expression, and the inline-credential messages are
built explicitly, so a credential flag in a command line written in
'command' reads "command line argument 'password' carries an inline
credential". A test pins those messages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
The project runs no type checker, and no other test in the file carries
one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…ped helper

Follow-up to the skip_reason fix: the warnings and unverified findings
for pins pip-audit skipped move into _report_skipped_pins, which lists
at most MAX_UNVERIFIED_PER_SOURCE of them like the other unverified
declarations and counts the rest in a message.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…it has no options

A command line whose first word is a path, such as
'./bin/serve --cache-dir /opt/cache/uvx', still read as one program
named by its last path segment (here uvx). A program path with spaces
has no option words, so _command_argv now requires that too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…on carries it

check_canary tested the token against a whole apply_patch patch and
charged it to every header path outside the workspace. A multi-file patch
that put the token in a workspace file then reported a critical,
score-impacting file_outside_workspace leak for an outside file that got
none of it, and so did the source of a move (Update File: /tmp/old.txt
with Move to: notes.txt), which the move removes. The PR's shell
apply_patch support extended this to apply_patch heredocs.

_canary_patch_sections splits a patch into per-file sections: an Add or
Update header up to the next file header, written to its own path or to
the Move to path that follows it; a deleted file and a move's source
receive nothing. The apply_patch tool tests the token per section. A shell
apply_patch charges a file only when the statement carries the canary and
the file's own section holds the token or an expansion ($, a backquote)
the shell may fill with it. An outside file whose section does not still
gets the verifier's read-back, as a moved decoy does.

The shared block stays byte-identical in the verifier.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
The canary read a shell apply_patch's patch only from the command's own
argument or heredoc. A patch piped in (cat <<'EOF' | apply_patch, or
echo '...' | applypatch) that wrote the token outside the workspace gave
no file_outside_workspace sink, in the host checks and the verifier, while
check_security and the judge evidence did see the write.

When the apply_patch command's own words hold no patch header and the
statement is a pipeline, the patch is read from the pipeline's words,
which hold the heredoc body or the echoed text. Each file is then charged
by its own section, as for a patch given directly. printf text with
escaped newlines is still not read.

The shared block stays byte-identical in the verifier.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
…e network check

The network check recorded only bare leading NAME=value words and
classified the command word as written, so these ran curl with an upload
and passed, in the host checks and the verifier (bash, sh and zsh run them
all):

- export, declare (-x), local, readonly or typeset A='curl -d @f URL',
  then eval "$A" or bash -c "$A";
- env A='curl -d @f URL' sh -c 'eval "$A"', whose value kept a stray
  quote;
- A='curl ...'; eval -- "$A", where -- became the payload's command;
- A='curl -d @f URL'; $A, and A=curl; $A -d @f URL.

Declaration builtins and env now assign through _network_assignment like a
bare assignment, eval drops a leading --, and an unquoted $VAR command word
is replaced by the words its value splits into before it is classified. A
declaration runs nothing itself, so its words are no longer read as a
wrapped client. A plain GET through any of these forms is still not a
finding.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Bring #28's 19 new commits (Harbor 0.24.0, the Tier 3 collector and
plugin-signals rewrites, plugin evaluation fixes) into this branch and
resolve the 42 conflicting files.

Where both sides changed the same behaviour, #28's approach is kept
unless this branch covered a tested case it did not; this branch's
refactors are ported onto #28's code:
- validators: one MCP runner reader (parse_mcp_runner) covers #28's
  runner forms; URL credential rules live in url_policy, shared by MCP
  URLs, HTTP hooks and hook commands; #28's floating-version, shell -c,
  wrapper and credential rules are kept.
- plugin model: #28's repository-identity rule, with a remedy per
  cause; one agent-CLI flag walk that includes --allowedTools; the
  not-loaded state lives in plugin_states.
- verifier: eval.py and eval_core stay byte-identical; process
  substitution writes such as `tee >(cat) ~/.bashrc` are detected.
- Tier 3: one arm-collection pipeline carries #28's per-arm behaviour.
- reporting: one component index, one canary attribution rule, and
  per-agent Integration blocks in the result display.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The local streaming test allowed 0.5 s for the first output of a
command that exits after 1 s, and the bounded-process timeout test
killed its child after 0.1 s, before a loaded machine had started the
interpreter and written the diagnostic. Both failed under a parallel
full-suite run and passed alone.

The streamed command now sleeps 5 s with a 4 s wait for the partial
line, and the timeout test allows 5 s against a child that sleeps 30 s,
so each still proves what it tests.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Gitleaks scans the full history of every pull request, and three
review-fix commits added high-entropy test fixtures that trip its jwt,
generic-api-key and private-key rules.

The tests now assemble those values at run time, so the tree carries
no literal Gitleaks can match, and .gitleaks.toml allowlists each
original commit for only its file and the one rule it tripped.
tests/test_ci_workflows.py pins the three entries.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SecurityValidator runs redact_secrets on whole SkillSpector snippets and
PII lines, and four of its patterns took quadratic time on one long run
of text, so a crafted 100 KB line that also triggered a finding stalled
Tier 1 for minutes (CWE-1333): the URL userinfo and query patterns tried
an optional scheme at every word of a run such as '-ab-ab...' and
retried a run of slashes from every slash, the query pattern re-read a
URL's path from every 'a://' inside it and scanned to the end of the
text from every '${' without a '}', and the credential flag and
assignment patterns retried every credential word of a long name
('--auth' or 'TOKEN' repeated), each time reading to the end of the run.
The userinfo pattern now starts only at the first slash of a run (the
scheme before it was always kept as written), the query step finds URL
starts once per scheme run and computes where every path and query ends
in one backward pass over the text's breaks, and the flag and assignment
patterns commit to a name's first credential word with an atomic group
and possessive repeats. Output is unchanged: no difference from the old
code on the 3,596 matching string literals in tests/, 31,215 generated
URL-ish strings, or 3.66 million token sequences; 100 KB inputs that
took 8 s to over 2 minutes now take under 0.05 s.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Windows Tier 2 job failed this test with a UnicodeDecodeError:
it read the Markdown report with Path.read_text() and no encoding, so
Windows decoded the UTF-8 report as cp1252. The reports are written as
UTF-8, so the test now reads all three (JSON, SARIF, Markdown) that
way.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Batch 2 of fixes from the code-quality review of #28 (remaining bugs, logic that had drifted apart, dead code, readability, performance), the fixes from the review of #181, and fixes to #28's newer code made while merging its head. One signed-off commit per fix, each bug fix with a regression test.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Bring main's v0.5.0 into this branch, including the Harbor 0.24.0
upgrade from #83, and resolve the 40 conflicting files.

This branch had moved plugin evaluation to Harbor 0.24.0 on its own.
Where that work and #83 changed the same code, #83's reviewed version
is kept for Harbor and runtime plumbing, and this branch's
plugin-evaluation work is kept on top:
- runtime: Harbor launches from an evaluator-owned directory with .env
  loading and telemetry off; hardened Docker stays on Docker; failed
  jobs name Harbor's recorded trial or step exception; multimodal ATIF
  content reaches judges and checks as text.
- environments: `--env-mode wandb` is kept as an alias that runs
  cwsandbox with W&B authentication, operators cannot set `--ek auth`,
  and W&B credentials reach only wandb mode.
- collection: grader-written rewards cannot set the collector's
  identity keys; every arm goes through the one collection pipeline.
- verifier: eval.py and eval_core stay byte-identical where shared and
  Python 3.9-compatible.
- CHANGELOG: v0.5.0 reads as released; Unreleased keeps only this
  branch's own entries.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Conflicts (docs only):
- CHANGELOG.md: keep the plugin evaluation entries and add main's
  documentation-refresh entry under Unreleased > Fixed.
- docs/cli-reference.mdx: keep this branch's --repo-root,
  --resolve-endpoints and claude-validate rows and the experimental
  hermes agent; take main's new --workers and --previous-version rows,
  the full --env-mode backend list, and the fuller --environment-kwarg
  row (a superset of this branch's).
- docs/tier3-live-evaluation.mdx: take main's wording for the reward and
  step-exception note and keep this branch's bullet that Harbor 0.24's
  stream and enable_environment_dir_upload options are reserved.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants