Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -72,6 +72,27 @@ jobs:
- name: Lint
if: ${{ needs.classify-changes.outputs.docs_only != 'true' }}
run: uv run ruff check .
# The Harbor verifier runs these files with the task image's python3, which
# can be older than the Python SkillEvaluator itself needs. Compiling catches
# syntax; the smoke script also runs code that needs newer runtime features.
- name: Compile Harbor verifier files on Python 3.9 and 3.11
if: ${{ needs.classify-changes.outputs.docs_only != 'true' }}
shell: bash
run: |
set -euo pipefail
files=(
src/skillevaluator/tier3/harbor/templates/*.py
src/skillevaluator/tier3/eval_core/log_converters.py
src/skillevaluator/tier3/eval_core/codex_tool_call_normalizer.py
src/skillevaluator/evidence.py
)
for version in 3.9 3.11; do
uv run --no-project --python "$version" -- python -I -c \
"import sys; [compile(open(p, encoding='utf-8').read(), p, 'exec') for p in sys.argv[1:]]" \
"${files[@]}"
uv run --no-project --python "$version" --with "idna>=3.10,<4" -- python -I \
scripts/ci/smoke_harbor_verifier.py
done
- name: Run tests with coverage
if: ${{ needs.classify-changes.outputs.docs_only != 'true' }}
run: >-
Expand Down
14 changes: 14 additions & 0 deletions .gitleaks.toml
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,20 @@ targetRules = ["generic-api-key"]
regexes = ['''^sk-AbCdEf1234567890$''']
paths = ['''^tests/test_tier3_(progress|result_display)\.py$''']

[[allowlists]]
description = "Synthetic API key used by plugin hook inline-secret tests"
condition = "AND"
targetRules = ["generic-api-key"]
regexes = ['''^abcd1234secret$''']
paths = ['''^tests/validators/test_plugin_component_risk\.py$''']

[[allowlists]]
description = "Synthetic Docker registry login (base64 of a fake user:password) used by the scanner-credential tests"
condition = "AND"
targetRules = ["generic-api-key"]
regexes = ['''^dXNlcjpodW50ZXIy$''']
paths = ['''^tests/validators/test_audits_endpoints_review_fixes\.py$''']

[[allowlists]]
description = "Synthetic basic-auth, cloud-key and URL-token fixtures in the first revisions of the security-evidence and MCP-hardening tests; both files now assemble them from split literals"
condition = "AND"
Expand Down
200 changes: 200 additions & 0 deletions CHANGELOG.md

Large diffs are not rendered by default.

24 changes: 17 additions & 7 deletions docs/cli-reference.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -202,7 +202,7 @@ Applies to the whole run: target typing, policy profile, reports, tier selection
| `--policy FILE` | none | Custom policy YAML overlaid on top of `--profile`. |
| `--repo-root DIRECTORY` | git top-level of the plugin | Plugin only: repository root used to resolve same-repository skill and rule references in Tier 1 and Tier 3. The missing-dependency gate also requires this root to have a git `origin` remote; otherwise references stay `unresolved` (advisory). |

Auto-detection recognizes: a `SKILL.md` for skills, `.mdc` files for rules, `workflow-rules.mdc` for workflows, and — for plugins — a bundle-reference `agent_plugin.yaml`/`.yml` manifest or a contained `.claude-plugin/plugin.json` manifest. Plugins are validated against their public contract; quality, lint, and version checks run on each skill bundled under the plugin's `skills/` directory. Plugins run Tiers 1 and 2 by default; request plugin Tier 3 with `--tier3`. See [Plugin Evaluation](plugin-evaluation.mdx).
Auto-detection recognizes: a `SKILL.md` for skills, `.mdc` files for rules, `workflow-rules.mdc` for workflows, and — for plugins — a bundle-reference `agent_plugin.yaml`/`.yml` manifest, or a contained `.claude-plugin/`, `.codex-plugin/`, or `.cursor-plugin/` `plugin.json` or Agent Plugins root `plugin.json` manifest ([precedence](plugin-evaluation.mdx#manifest-detection)). Plugins are validated against their public contract; quality, lint, and version checks run on each skill bundled under the plugin's `skills/` directory. Plugins run Tiers 1 and 2 by default; request plugin Tier 3 with `--tier3`. See [Plugin Evaluation](plugin-evaluation.mdx).

`validate` also takes the standard `-r/--report` and `-o/--output-dir` options described under [Global conventions](#global-conventions).

Expand All @@ -212,7 +212,8 @@ Static checks; LLM-free by default. Tier 1 gates the exit code and always runs.

| Flag | Default | Effect |
| --- | --- | --- |
| `--checks, --tier1-checks TEXT` | all applicable | Comma-separated subset of Tier 1 checks. Default choices: `schema`, `version`, `security`, `pii`, `license`, `code-integrity`, `unicode`, `quality`, `lint`; opt-in (not run by default): `dependency`. `quality`/`lint`/`version` are skill-only and skipped for rules and workflows. |
| `--checks, --tier1-checks TEXT` | all applicable | Comma-separated subset of Tier 1 checks. Default choices: `schema`, `version`, `security`, `pii`, `license`, `code-integrity`, `unicode`, `quality`, `lint`; opt-in (not run by default): `dependency`, and, for plugins, `claude-validate` (parity with `claude plugin validate --strict`). `quality`/`lint`/`version` are skill-only and skipped for rules and workflows. |
| `--resolve-endpoints` | off | Plugin only, opt-in network check: resolve MCP and HTTP hook URL hosts and send one credential-free `HEAD` (no redirects followed) to flag names or redirects that reach private, link-local, or cloud-metadata addresses. Also enabled by `endpoints.resolve: true` in the policy. See [Plugin Evaluation](plugin-evaluation.mdx#opt-in-endpoint-dns-and-redirect-checks). |
Comment thread
rng1995 marked this conversation as resolved.
| `--fail-fast` | off | Stop on the first failing check instead of collecting all issues. |
| `-c, --continue-on-failure` | off | Run the full pipeline without stopping early; record all issues in the reports. Overrides `--fail-fast`, and for folder validation keeps scanning every skill past a CRITICAL finding. |
| `--llm, --tier1-llm / --no-llm, --no-tier1-llm` | `no-llm` | Enable LLM-backed security analysis (requires a configured public provider — see [Providers & Credentials](configuration.mdx)). |
Expand Down Expand Up @@ -242,6 +243,9 @@ The following flags are forwarded to the live-eval engine when Tier 3 is enabled
| `-a, --agents TEXT` | provider-native | Comma-separated Harbor agents to evaluate. Defaults: NVIDIA Build=`opencode`, OpenAI=`codex`, Anthropic=`claude-code`. |
| `--env-mode` | `docker` | Harbor environment backend (full list under [tier3 evaluate](#tier3-evaluate)). |
| `--lift-mode [effectiveness\|integration\|both]` | `effectiveness` | Plugin only: compare the coordinated plugin with no plugin, its member skills staged individually, or both. Integration requires explicit cross-component dataset evidence. |
| `--plugin-load [wrapper\|native\|auto]` | `wrapper` | Plugin only: how the with-plugin arm loads the plugin: the generated wrapper skill, the harness's own plugin layout, or native where supported. See [Native loading](plugin-evaluation.mdx#native-loading). |
| `--probe-mcp` | off | Plugin only: before Tier 3, probe each author-supplied URL MCP server from the host (`initialize` + `tools/list`, bounded, no redirects) under the endpoint policy and `mcp.allowed_private_hosts` from `--policy`. Advisory; recorded as `mcp_proof`. See [Public MCP proof](plugin-evaluation.mdx#public-mcp-proof). |
| `--probe-mcp-env NAME` | none | Plugin only, with `--probe-mcp`: a host environment variable the probe may expand into a server's declared `${NAME}` headers (repeatable). Without it, the probe sends only literal header values; a header that references any other variable is not sent. |
| `--skip-baseline` | off | Skip the without-skill baseline (no lift analysis, faster). |
| `--n-concurrent INTEGER` | unset | Concurrent eval cases per agent. |
| `--max-agents INTEGER` | unset | Maximum agents to run in parallel. |
Expand Down Expand Up @@ -394,7 +398,7 @@ Use `skillevaluator tier3 PATH` for the complete workflow with automatic missing

| Flag | Default | Effect |
| --- | --- | --- |
| `-a, --agents TEXT` | provider-native | Comma-separated Harbor agents. Defaults: NVIDIA Build=`opencode`, OpenAI=`codex`, Anthropic=`claude-code`. Supported: `claude-code`, `codex`, `opencode`; the alias `claude` is accepted for `claude-code`. See [Agents & Sandboxes](agents-and-sandboxes.mdx). |
| `-a, --agents TEXT` | provider-native | Comma-separated Harbor agents. Defaults: NVIDIA Build=`opencode`, OpenAI=`codex`, Anthropic=`claude-code`. Supported: `claude-code`, `codex`, `opencode`; experimental: `hermes` (container environments only, Anthropic provider); the alias `claude` is accepted for `claude-code`. See [Agents & Sandboxes](agents-and-sandboxes.mdx). |
| `--env-mode` | `docker` | Where trials run. All 16 values: `docker`, `daytona`, `e2b`, `modal`, `runloop`, `langsmith`, `gke`, `novita`, `apple-container`, `singularity`, `islo`, `tensorlake`, `cwsandbox`, `wandb`, `use-computer`, `local`. Cloud modes are provider-managed Harbor backends enabled by the matching Harbor extra; `local` runs on your host. See [Agents & Sandboxes](agents-and-sandboxes.mdx). |
| `--autopilot` | off | Create one eval case when no dataset/task source exists, then evaluate. The case is LLM-generated with the configured provider, with a deterministic keyless template fallback; an existing source is never overwritten. |
| `--skip-baseline` | off | Skip the without-skill baseline (no lift analysis, faster). |
Expand Down Expand Up @@ -433,10 +437,13 @@ Flags marked "unset" fall back to their matching keys in `evals/config.yml` wher
## tier3 evaluate-plugin

Run Tier 3 against a public plugin without fetching remote components. The
command accepts a bundle-reference `agent_plugin.yaml`/`.yml` target or a
contained `.claude-plugin/plugin.json` target, stages the locally evaluable
skills, rules, and MCP declarations into a temporary wrapper, and records any
unresolved remote references in plugin provenance.
command accepts a plugin root or its manifest in any supported format: a
bundle-reference `agent_plugin.yaml`/`.yml`, a contained `.claude-plugin/`,
`.codex-plugin/`, or `.cursor-plugin/` `plugin.json`, or an Agent Plugins root
`plugin.json` ([precedence](plugin-evaluation.mdx#manifest-detection)). It
stages the locally evaluable skills, rules, and MCP declarations into a
temporary wrapper (or the harness's own layout with `--plugin-load`), and
records any unresolved remote references in plugin provenance.

```bash title="Evaluate plugin effectiveness and Integration"
skillevaluator tier3 evaluate-plugin ./my-plugin --lift-mode both \
Expand All @@ -450,6 +457,9 @@ skillevaluator tier3 evaluate-plugin ./my-plugin --lift-mode both \
| `--env-mode` | `docker` | Harbor environment backend; accepts the same values as [tier3 evaluate](#tier3-evaluate). |
| `--skip-baseline` | off | Skip the no-plugin baseline. Invalid with `integration` or `both`, which require a baseline. |
| `--lift-mode [effectiveness\|integration\|both]` | `effectiveness` | Compare the coordinated plugin with no plugin, its member skills staged individually, or both. `both` falls back to effectiveness when composition evidence is missing. |
| `--plugin-load [wrapper\|native\|auto]` | `wrapper` | How the with-plugin arm loads the plugin. `native` fails for agents or environments without a native adapter; `auto` falls back to the wrapper. Baseline arms are unchanged. See [Native loading](plugin-evaluation.mdx#native-loading). |
| `--probe-mcp` | off | Before the run, probe each author-supplied URL MCP server from the host (`initialize` + `tools/list`, bounded, no redirects) under the endpoint policy. This command takes no policy file and uses a bundled profile (`SKILLEVALUATOR_PROFILE` or the default), none of which allowlists private hosts, so private hosts are not probed; a custom `mcp.allowed_private_hosts` allowlist applies to `--probe-mcp` only through `validate --tier3 --policy <file>`. Advisory; recorded as `mcp_proof` in plugin provenance. See [Public MCP proof](plugin-evaluation.mdx#public-mcp-proof). |
| `--probe-mcp-env NAME` | none | With `--probe-mcp`: a host environment variable the probe may expand into a server's declared `${NAME}` headers (repeatable). Without it, the probe sends only literal header values; a header that references any other variable is not sent. |
| `--n-attempts INTEGER` | unset | Attempts per eval case (pass@k). |
| `--pass-threshold FLOAT` | unset | Score threshold (0.0–1.0) for a case to count as passed. |
| `--stop-on-pass / --no-stop-on-pass` | unset | Stop a case's remaining attempts once one passes. |
Expand Down
1 change: 1 addition & 0 deletions docs/eval-datasets.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -283,6 +283,7 @@ grading:
| `harbor.max_agents` | integer ≥ 1 | Cap on agents evaluated in one run. |
| `harbor.timeout_multiplier` | number > 0 | Scales task timeouts. |
| `harbor.agent_runtime_preflight` | boolean | Opt-in extra agent-only execution of the first staged task before the measured evaluation. Default: `false`; the CLI flag overrides this value. |
| `harbor.plugin_canary` | boolean | Plant the canary decoy credential in every arm of a plugin run. Default: `true`. Set `false` to run a plugin evaluation without it. |
| `harbor.agent_workdir` | string | Working directory for the agent inside the container. |
| `harbor.resources` | `cpus`, `memory_mb`, `storage_mb` | Per-container resource requests. |
| `harbor.runtime_env` | list or mapping | Non-credential task values passed into the container. Prefer a list of plain names — each expands to `${NAME}` from your shell; a mapping sets explicit templates. Entries that name or reference operator-owned credentials (`OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `NVIDIA_API_KEY`, base-URL variables, AWS credential variables) fail with a hard error. Alias: `passthrough_env`. |
Expand Down
Loading
Loading