Skip to content

feat(plugin): Stage C — native loading, more manifests, static risk, runtime evidence - #171

Merged
chrisknvidia merged 5 commits into
naren/plugin-evaluation-all-tiersfrom
naren/plugin-eval-stage-c
Oct 1, 2026
Merged

chrisknvidia merged 5 commits into
naren/plugin-evaluation-all-tiersfrom
naren/plugin-eval-stage-c

Conversation

@rng1995

@rng1995 rng1995 commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Stage C and open-gap work for plugin evaluation, stacked on #28. The base is naren/plugin-evaluation-all-tiers; it retargets to main once #28 merges. It closes the remaining gaps from the plugin-evaluation review: more manifest formats, deeper static risk checks, runtime evidence, and native plugin loading for each harness.

Draft for review. Do not merge. Nothing here has been run against a live Harbor run yet; live validation follows review.

What's included

Area Capability Gating
Manifests Codex .codex-plugin/plugin.json, Cursor .cursor-plugin/plugin.json, and Agent Plugins v1 (plugin.json with an agent-plugins.org $schema), in every tier. Precedence when several are present: agent_plugin.yaml > .claude-plugin > Agent Plugins > .codex-plugin > .cursor-plugin. The other manifests are recorded under manifest_declarations, and name/version conflicts are flagged. Components flow into the existing inventory, MCP policy and Tier 3 staging. Codex http_headers are checked together with headers. A root plugin.json that is not UTF-8 or exceeds 1 MiB no longer fails discovery; an additional manifest like that is a MEDIUM plugin_manifest_additional_invalid. An Agent Plugins plugin.json with a JSON syntax error is still auto-detected as a plugin, so the error is reported. Per-format required-field checks (HIGH/MEDIUM)
Subagent and command privileges Frontmatter tools / allowed-tools / permissionMode. Flags unrestricted Bash, wildcard tools, bypassPermissions / acceptEdits, and subagents that inherit all tools while write-capable MCP servers are present. MCP denials are matched against Claude Code's mcp__plugin_<plugin>_<server>__* tool names, --read-only=false is not treated as read-only, and tool lists are checked in full, so padding cannot hide unrestricted Bash. HIGH / MEDIUM, overridable by policy
Hook risk model Parses events, matchers and handler types (command/http/prompt/agent). Flags broad-matcher auto-approve, context injection, remote code (curl … | sh, including path-prefixed shells, env/sudo/xargs wrappers, intermediate pipe stages, and download-then-run), http hook endpoint policy with a hooks.allowed_urls allowlist compared as parsed URLs (scheme, host, port, and path on a segment boundary), and inline secrets, including credentials embedded in URLs. Matchers are classified without a backtracking regex engine and fail closed when too long or complex. Flat hooks in Cursor and Agent Plugins extensions get the same checks. A hook script that can't be analyzed is a HIGH plugin_hook_script_unanalyzed, and hitting a scan limit is a HIGH plugin_hook_scan_truncated. URL credentials are stripped from finding messages. CRITICAL / HIGH / MEDIUM / LOW
npm and container CVEs The opt-in dependency check audits npm exact pins and container image refs with OSV-Scanner, falling back to npm audit --package-lock-only --ignore-scripts, then Grype or Trivy for images. Images from every manifest the plugin ships are covered. Plugin files never reach a scanner directly (synthesized lockfile). The plugin audit never passes silently: no scanner, a failed scanner run (pip-audit included), or a manifest, lockfile or Dockerfile that can't be read within bounds makes the check INCOMPLETE. Scanner severity
Endpoint DNS/redirect policy Opt-in validate --resolve-endpoints: resolves MCP (including servers that only an additional manifest declares) and http-hook hosts, flags private or metadata answers, and sends one HEAD request (TLS 1.2 minimum, no redirect following, no credentials). Non-public hosts are never contacted, and hosts with an unexpanded ${VAR} are skipped. HIGH / MEDIUM
Validator parity Opt-in --checks claude-validate runs claude plugin validate --strict --json when the CLI is present and reports disagreements. MEDIUM / LOW / INFO
Hook execution census hook_census.sh wraps natively loaded hooks and records runs, failures and duration without changing stdin, stdout or the exit code. Natively staged hooks that ran become exercised in coverage. Hooks that weren't staged or loaded are never promoted from census lines. Advisory
Canary exfiltration Each task plants a random canary (file plus env var) in both arms. The finding fires when the canary reaches network commands or tools, MCP args, URLs, git, or files outside the workspace, and reports whether the difference is attributable to the plugin. Shell commands are checked per statement, including sh -c/bash -lc/eval payloads, heredocs and argv-list commands, with taint tracked across statements. Whole-environment dumps count; reading one variable by name does not. URLs count only for network, MCP and command tools. Critical canary_exfiltration
Public MCP proof Opt-in --probe-mcp: probes the author's URL MCP servers from the host (initialize + tools/list, endpoint policy first), then upgrades status from in-agent evidence: declared → reachable-host → reachable-in-agent → used-successfully. The probe sends only literal declared headers; host environment variables are never expanded unless named with --probe-mcp-env NAME, which can be repeated. Advisory
Coverage New states loaded and exercised (precedence exercised > loaded > staged). Subagents and commands observed at runtime count as exercised. The headline counts every evaluated state (staged, loaded, exercised). Report-only
Native loading + Hermes --plugin-load {wrapper,native,auto} (default wrapper). The with-plugin arm loads the plugin the way each harness does: claude-code via --plugin-dir (skills, commands, agents, hooks wrapped by the census, .mcp.json, output styles; rules in $CLAUDE_CONFIG_DIR/rules); codex (skills, $CODEX_HOME/AGENTS.md, config.toml mcp_servers); opencode (OPENCODE_CONFIG mcp and instructions, agents, commands); hermes (skills and MCP; rules stay in the wrapper). Codex, Cursor and Agent Plugins manifests are staged through their normalized Claude view. A load census inside the container marks components loaded. Member-skills and no-plugin arms are unchanged, so lift stays comparable. Hermes is added as a Tier 3 agent (containers only). No bypass flags are ever added, and configs carrying them fail closed. Report-only

All new data is rendered in the JSON, Markdown, HTML and CLI plugin sections (plugin-controlled text is escaped in the CLI report), and documented in docs/plugin-evaluation.mdx, including updated limitations.

Review follow-ups

All 30 comments from the review and the 3 CodeQL alerts are addressed and resolved. The fixes are folded into the table above. The most important:

  • --probe-mcp credential exfiltration (High): headers declared as ${VAR} were filled from the host environment and sent to the plugin's own URL. Now only variables named with --probe-mcp-env are expanded.
  • hooks.allowed_urls bypass (High): matching was a string prefix, so https://hooks.example.com.evil.net passed. Now parsed URLs are compared.
  • Fail-open limits and ReDoS: scan and size limits now produce HIGH or INCOMPLETE results instead of silently passing. Plugin-supplied regexes and the remote-code patterns are bounded; one reviewer case dropped from 18 s to 0.01 s.
  • CodeQL: TLS 1.2 minimum for the endpoint HEAD check, and exact-hostname assertions in tests. No open alerts.

Needs live validation (after review)

  • The hook census inside real harness hooks: a writable /logs/agent and passthrough of hook decisions.
  • Canary placement on every Dockerfile path and in local mode, plus the false-positive rate on real trajectories.
  • --probe-mcp against real streamable-HTTP/SSE servers and behind proxies.
  • claude plugin validate --json output on current Claude Code builds.
  • Native loading on each harness: whether each harness parses the staged config (--plugin-dir in -p mode, $CLAUDE_CONFIG_DIR/rules, the Codex config.toml/AGENTS.md appends, the OpenCode config merge), hooks firing under Harbor's --permission-mode=bypassPermissions, the Hermes install and provider routing, and /skilleval and /logs/agent permissions for non-root agent users. The census proves files are present, not that they were loaded.

Verification

  • After the review fixes (63f9f93):
    • ruff check .: passes.
    • Full suite (pytest -n auto): 10,729 passed, 26 skipped, 0 failed.
    • Golden CLI surface: regenerated for --probe-mcp-env.
    • CodeQL: 0 open alerts.
  • At submission: the OSS boundary scan and the public benchmark scan passed.

Review fixes

A review of the Stage C code (same design as this PR) turned up correctness and security defects. This push ports the 14 fixes that apply here, and merges the updated naren/plugin-evaluation-all-tiers (#28 with its own ported review fixes) to clear this PR's conflict with its base. Each fix was confirmed against this branch before porting, and each ported regression test fails on the previous commit and passes after.

Area Ported Already fixed or not applicable here
Static risk 5: npm audit reports its exit code and stderr when it prints no JSON; npm/image audits are INCOMPLETE past their discovery caps; hooks.allowed_urls entries with userinfo, a query or a fragment match nothing; hook scripts are read under each format's root placeholder (${PLUGIN_ROOT}, ${CURSOR_PLUGIN_ROOT}, relative Cursor scripts, Agent Plugins namespaces); every resolved address must be allowlisted, whatever the DNS answer order ExternalTool.run env layering, MCP from additional manifests, the additional agent_plugin.yml, and the unreadable root plugin.json were already handled
Runtime 5: --probe-mcp connections are pinned to the policy-checked addresses (answers the DNS-rebinding review comment); grading data (results dirs, generated output, evals source) is kept out of the native Claude Code plugin copy; only the selected manifest's components are staged natively, and hook sources with no staged handlers are reported not_loaded; only the Hermes chat launch gets the native setup prefix; docs canary sink scoping and the header-variable allowlist were already handled (--probe-mcp-env)
Reports 4: the canary verdict comes from the per-arm rows (no green verdict when the plugin arm leaked); hook-census qualifiers are kept in Markdown and CLI; endpoint rows take their redirect target's classification; docs say where the --probe-mcp private-host allowlist applies the hook-matcher markup escape and the coverage count were already fixed

🤖 Generated with Claude Code

Comment thread src/skillevaluator/validators/endpoint_resolution.py Fixed
Comment thread tests/validators/test_plugin_component_risk.py Fixed
Comment thread tests/validators/test_plugin_manifest_formats.py Fixed

@rng1995 rng1995 left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review of Stage C at 16ea9ba: correctness and security findings. Each inline comment was reproduced against a snapshot of this SHA or confirmed by reading the code, and every one has a severity, a failure scenario and a suggested fix.

Highest priority:

  • --probe-mcp expands host environment variables into plugin-declared headers and sends them to the plugin's own URL.
  • The hooks.allowed_urls prefix match can be bypassed with a look-alike host.

Beyond those, most findings are detection bypasses in plugin_component_risk.py (caps and truncation that fail open, ReDoS), CVE-audit coverage gaps, and canary false negatives and false positives.

e1a0006..28e804a (native loading) was pushed after this SHA. I re-checked every finding against 28e804a, and all still apply. I dropped one docs finding (the hook census never being deployed) because native staging now wires hook_census.sh. The existing CodeQL alerts are not repeated here.

Comment thread src/skillevaluator/tier3/mcp_proof.py
Comment thread src/skillevaluator/plugin_component_risk.py Outdated
Comment thread src/skillevaluator/plugin_component_risk.py Outdated
Comment thread src/skillevaluator/plugin_component_risk.py Outdated
Comment thread src/skillevaluator/plugin_component_risk.py Outdated
Comment thread src/skillevaluator/cli_core.py Outdated
Comment thread src/skillevaluator/reporting/plugin_sections.py Outdated
Comment thread docs/cli-reference.mdx
Comment thread docs/plugin-evaluation.mdx Outdated
Comment thread docs/tier1-validation.mdx
@rng1995
rng1995 marked this pull request as ready for review September 30, 2026 07:07
@rng1995
rng1995 marked this pull request as draft September 30, 2026 07:21
@rng1995
rng1995 marked this pull request as ready for review September 30, 2026 16:51
@rng1995

rng1995 commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Live Harbor sample (non-blocking)

A small Harbor sample ran on the equivalent internal code with Claude Code (native and wrapper loading) and Codex (native loading). It confirmed that native loading works (Claude Code reports the plugin, the skill and the MCP server as connected), that hooks run under bypassPermissions, that the canary raised no false positives, and that --probe-mcp reached a public server. It also found follow-ups that apply to this design:

  • Markdown reporting crashes when a baseline evaluator score is None.
  • The claude plugin validate child process still receives NVIDIA_INFERENCE_KEY, AWS_ACCESS_KEY_ID and similar variables.
  • Cursor hook events are staged into Claude's hooks.json without translation.
  • Natively loaded skills (<plugin>:<skill>) aren't matched by the component classifier.

🤖 Generated with Claude Code

Comment thread src/skillevaluator/tier3/harbor/native_agents.py Fixed
chrisknvidia and others added 5 commits October 1, 2026 13:04
…atic risk checks

Detect and validate .codex-plugin, .cursor-plugin and the Agent Plugins v1
root plugin.json next to the Claude manifest, with per-format component
profiles and flagged name/version conflicts. Oversize or non-UTF-8 client
manifests no longer hide their hooks or MCP servers.

Add Tier 1 static risk checks: hook events, handlers, auto-approve,
fetch-and-execute, context injection and hooks.allowed_urls; subagent and
command privileges; skill frontmatter hooks and allowed-tools; monitor and
LSP commands; permission-bypass flags in every config node. Unreadable or
oversize component files fail closed, and whole-tree scans cover declared
skill folders.

Co-authored-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Add npm and container image CVE audits for plugins. An audit that runs
offline, hits its discovery cap or cannot find its scanner reports
INCOMPLETE instead of passing.

Add --resolve-endpoints: DNS, redirect and TLS checks against the endpoint
policy for every resolved address, with WHATWG redirect parsing and time
budgets that report unchecked endpoints. Add opt-in parity with
claude plugin validate --strict; its child process gets a scrubbed
environment and a throwaway HOME.

Co-authored-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Add --plugin-load wrapper|native|auto (default wrapper, which keeps the
PR #28 behavior). Native Claude Code loads the plugin with --plugin-dir:
skills, rules, MCP servers including ${CLAUDE_PLUGIN_ROOT} launches,
hooks, agents, commands, output styles, LSP servers, settings and
userConfig. Codex, OpenCode and Hermes get their own config files, and
OpenCode subagents keep their tool limits. Only the with-plugin arm
changes.

A per-trial load census records what each harness listed or loaded. A
failed MCP server never counts as loaded, and declared components that
were not staged make the run INCOMPLETE. auto falls back to the wrapper
per agent and records why. Hermes is experimental and refused with the
OpenAI and OpenAI-compatible providers. Launch rewriting is linear-time.

Co-authored-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Add per-arm canary exfiltration checks (incl. Hermes, OpenCode and
subagent tool calls), a hook execution census, and --probe-mcp proof that
connects only to policy-checked addresses and never sends host secrets.
Native skill and MCP tool names are matched for every harness, so native
runs credit the components they used.

Render load, census, canary, signals and coverage in JSON, Markdown, HTML,
SARIF, the CLI summary and the BENCHMARK card, with untrusted text
escaped. Reports no longer crash on a missing baseline score, and a failed
arm no longer drops the with-plugin evidence.

Co-authored-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
The accuracy, goal accuracy and behavior judges now see write-tool
content under every harness's argument names (for example OpenCode
patchText), instead of a silent 200-character cut. Long text keeps its
start and end around a visible truncation marker, budgets are filled by
priority so the final answer, the skill call, test results and the latest
write to each path survive, and judge text is redacted. The configurable
evidence budgets from #160 drive this logic. Host and verifier copies stay
identical, and the verifier runs on Python 3.9 and 3.11 (CI smoke added).

Also update the docs, the changelog and the Gitleaks allowlist for
synthetic test fixtures.

Co-authored-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
@chrisknvidia
chrisknvidia force-pushed the naren/plugin-eval-stage-c branch from 6f0e1a7 to 429aa26 Compare October 1, 2026 20:07
@chrisknvidia
chrisknvidia merged commit 38a4f64 into naren/plugin-evaluation-all-tiers Oct 1, 2026
17 checks passed
@chrisknvidia
chrisknvidia deleted the naren/plugin-eval-stage-c branch October 1, 2026 20:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants