Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 8 additions & 4 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,10 @@ All notable changes to SkillEvaluator are documented in this file.
- Tier 3 log converters now rebuild ATIF trajectories from OpenCode JSON streams
(`opencode.txt`) and structured Codex tee logs (`codex.txt`) when
`trajectory.json` is missing or empty.
- `--env-mode runta` and `--env-mode mosaic` run Tier 3 trials on Harbor's Runta

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] The CHANGELOG contradicts itself

The same Unreleased section still says "Exposed 23 Harbor 0.24 backends" and lists runta and mosaic as disabled.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in d245d41: the entry now says 25 backends and no longer lists runta or mosaic as disabled.

and Mosaic backends. Runta's `mode` and Mosaic's `volume`, `persist`,
`enable_ssh`, `build_args`, and `build_target` kwargs are reserved because they
bypass the task image or let state and access outlive one isolated trial.

### Changed

Expand All @@ -55,10 +59,10 @@ All notable changes to SkillEvaluator are documented in this file.
Native tasks' own test scripts must write numeric `reward.json` values, which
Harbor 0.24 enforces; a failed job now names the trial or step exception that
Harbor recorded.
- Exposed 23 Harbor 0.24 backends alongside local mode. `cua-cloud`,
`opensandbox`, `hf-sandbox`, `podman`, `kata`, `runta`, `prime`, `mosaic`,
and `smol` remain disabled until generated tasks can be projected through a
trusted image or backend-native provisioning path. Non-secret backend
- Exposed 25 Harbor 0.24 backends alongside local mode. `cua-cloud`,
`opensandbox`, `hf-sandbox`, `podman`, `kata`, `prime`, and `smol` remain
disabled until generated tasks can be projected through a trusted image or
backend-native provisioning path. Non-secret backend
constructor options can be supplied with repeatable, operator-only
`--environment-kwarg` / `--ek` flags; skill-owned configuration,
credentials, and sandbox-policy overrides remain outside that surface.
Expand Down
51 changes: 44 additions & 7 deletions docs/agents-and-sandboxes.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -127,7 +127,7 @@ Anthropic evaluator with `opencode` also does not support local mode.

## Where trials run

SkillEvaluator exposes 24 environment modes: 23 Harbor-native backends,
SkillEvaluator exposes 26 environment modes: 25 Harbor-native backends,
including Docker, plus SkillEvaluator's own `local` host-execution mode.
Docker is the default. The `tier3` extra installs base `harbor==0.24.0`; install
the environment extra shown below when a managed backend needs one.
Expand Down Expand Up @@ -157,6 +157,8 @@ the environment extra shown below when a managed backend needs one.
| `skypilot` | Harbor-native | `harbor[skypilot]==0.24.0` |
| `hyperbrowser` | Harbor-native | `harbor[hyperbrowser]==0.24.0` |
| `vercel` | Harbor-native | `harbor[vercel]==0.24.0` |
| `runta` | Harbor-native | `harbor[runta]==0.24.0` |
| `mosaic` | Harbor-native | `harbor[mosaic]==0.24.0` |
| `local` | SkillEvaluator | Your machine, under an OS sandbox policy — see [Local mode](#local-mode). |

<Tip>
Expand All @@ -173,12 +175,16 @@ authentication (`auth=wandb`), using `WANDB_API_KEY` or the netrc file that
is rejected for both `cwsandbox` and `wandb`.

Harbor 0.24 also defines `cua-cloud`, `opensandbox`, `hf-sandbox`, `podman`,
`kata`, `runta`, `prime`, `mosaic`, and `smol`, but those backends are not
exposed until SkillEvaluator can project the complete task bundle — skill,
verifier, dependencies, and inputs — through a trusted image or
backend-native provisioning path, with credentials and constructor options
reviewed per backend. Selecting one of those names is rejected before a run
starts.
`kata`, `prime`, and `smol`, but those backends are not exposed until
SkillEvaluator can project the complete task bundle — skill, verifier,
dependencies, and inputs — through a trusted image or backend-native
provisioning path, with credentials and constructor options reviewed per
backend. Selecting one of those names is rejected before a run starts. Prime
cannot build a Dockerfile-only task, which is how SkillEvaluator stages most
tasks, and Smol runs only a published `[environment].docker_image`, which
SkillEvaluator rejects because it bypasses the evaluator-staged image. Exposing
either one needs Compose-only staging or a trusted image-publishing step,
validated live.

### Native backend constructor options

Expand Down Expand Up @@ -254,6 +260,37 @@ Harbor 0.24 Modal preflight requires `~/.modal.toml` or `MODAL_TOKEN_ID` plus
custom config path by itself; configure one of Harbor's supported credential
forms before launching trials.

Runta authenticates with `RUNTA_TOKEN` or the configuration file named by
`RUNTA_CONFIG` (default `~/.config/runta/config.toml`), and Harbor's preflight
requires one of the two variables. `RUNTA_ENDPOINT` selects a non-default API
endpoint. A relative `RUNTA_CONFIG` path is resolved against the directory you
run SkillEvaluator from, and a configured file must exist. Runta's `mode`
kwarg is reserved: `mode=direct` runs commands in Runta's base runtime and
ignores the task image, while Harbor's automatic selection builds the task's
Dockerfile or Compose definition inside the runtime.

Mosaic authenticates with `MOSAIC_API_TOKEN` or the token `mos login` saves in
`~/.config/mosaic-sandbox/config.json`; `MOSAIC_CONFIG` names a different file,
which must exist and is resolved against the directory you run SkillEvaluator
from when relative.
`MOSAIC_API_URL` and `MOSAIC_RETRIES` tune the API client, the SDK's older
`MAR_API_TOKEN`, `MAR_ENDPOINT`, and `MAR_CONFIG` names are forwarded too, and
`MOSAIC_REGISTRY_USERNAME` and `MOSAIC_REGISTRY_PASSWORD` cover private base
images. Mosaic's `volume`, `persist`, and `enable_ssh` kwargs are
reserved because they let state or access outlive one isolated trial, and
`build_args` and `build_target` are reserved because they can skip or alter
evaluator-staged Dockerfile layers. Mosaic's `secrets` kwarg is rejected because
it injects provider secrets into the agent's sandbox, and a Mosaic token stored
only in `E2B_API_KEY` is rejected because that variable is never forwarded to
Mosaic. Mosaic runs one VM per trial, so tasks with an
`environment/docker-compose.yaml` are rejected before the run starts.

Harbor builds each Mosaic task image into a snapshot named from the task and its
build context, reuses an existing snapshot with that name in later runs, and
never deletes snapshots. Run evaluations in a Mosaic account that only trusted
operators can write to, and delete old `harbor-*` snapshots when you no longer
need the staged skills and inputs they contain.

Other eligible, backend-specific constructor options are forwarded after
validation. A key that the selected backend's Harbor 0.24.0 constructor does not
accept is rejected before the run starts; for example, Harbor 0.24 removed
Expand Down
6 changes: 3 additions & 3 deletions docs/cli-reference.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -386,7 +386,7 @@ Use `skillevaluator tier3 PATH` for the complete workflow with automatic missing
| Flag | Default | Effect |
| --- | --- | --- |
| `-a, --agents TEXT` | provider-native | Comma-separated Harbor agents. Defaults: NVIDIA Build=`opencode`, OpenAI=`codex`, Anthropic=`claude-code`. Supported: `claude-code`, `codex`, `opencode`; the alias `claude` is accepted for `claude-code`. See [Agents & Sandboxes](agents-and-sandboxes.mdx). |
| `--env-mode` | `docker` | Where trials run. SkillEvaluator exposes 24 environment modes: 23 Harbor-native modes — `docker`, `daytona`, `e2b`, `modal`, `runloop`, `langsmith`, `ec2`, `gke`, `ack`, `openshift`, `novita`, `apple-container`, `singularity`, `islo`, `tensorlake`, `cwsandbox`, `wandb`, `use-computer`, `blaxel`, `beam`, `skypilot`, `hyperbrowser`, `vercel` — plus SkillEvaluator `local` mode. `wandb` runs Harbor's `cwsandbox` backend with W&B authentication, so it needs the `harbor[cwsandbox]==0.24.0` extra; Harbor 0.24 has no `wandb` extra. Harbor's `cua-cloud`, `opensandbox`, `hf-sandbox`, `podman`, `kata`, `runta`, `prime`, `mosaic`, and `smol` backends are not selectable. See the complete extras and prerequisites matrix in [Agents & Sandboxes](agents-and-sandboxes.mdx#where-trials-run). |
| `--env-mode` | `docker` | Where trials run. SkillEvaluator exposes 26 environment modes: 25 Harbor-native modes — `docker`, `daytona`, `e2b`, `modal`, `runloop`, `langsmith`, `ec2`, `gke`, `ack`, `openshift`, `novita`, `apple-container`, `singularity`, `islo`, `tensorlake`, `cwsandbox`, `wandb`, `use-computer`, `blaxel`, `beam`, `skypilot`, `hyperbrowser`, `vercel`, `runta`, `mosaic` — plus SkillEvaluator `local` mode. `wandb` runs Harbor's `cwsandbox` backend with W&B authentication, so it needs the `harbor[cwsandbox]==0.24.0` extra; Harbor 0.24 has no `wandb` extra. Harbor's `cua-cloud`, `opensandbox`, `hf-sandbox`, `podman`, `kata`, `prime`, and `smol` backends are not selectable. See the complete extras and prerequisites matrix in [Agents & Sandboxes](agents-and-sandboxes.mdx#where-trials-run). |
| `--environment-kwarg`, `--ek` | none | Operator-only `KEY=VALUE` constructor option for Harbor-native non-Docker backends (repeatable). Values accept JSON-compatible types; skill-owned `evals/config.yml` cannot set them. Never pass secrets. Docker and local reject all kwargs. Harbor runtime-policy fields for overrides, mounts, networks, extra Compose, lifecycle, and pod security are reserved, as are `stream`, `enable_environment_dir_upload`, and `auth` for `cwsandbox`/`wandb`. Names the selected backend does not accept in Harbor 0.24.0 are rejected. See [Native backend constructor options](agents-and-sandboxes.mdx#native-backend-constructor-options). |
| `--autopilot` | off | Create one eval case when no dataset/task source exists, then evaluate. The case is LLM-generated with the configured provider, with a deterministic keyless template fallback; an existing source is never overwritten. |
| `--skip-baseline` | off | Skip the without-skill baseline (no lift analysis, faster). |
Expand Down Expand Up @@ -519,7 +519,7 @@ skillevaluator doctor --env-mode docker
| Flag | Default | Effect |
| --- | --- | --- |
| `-a, --agents TEXT` | provider-native | Comma-separated agents to check. Defaults: NVIDIA Build=`opencode`, OpenAI=`codex`, Anthropic=`claude-code`. |
| `--env-mode` | `docker` | Backend to check (same 24 modes as [tier3 evaluate](#tier3-evaluate)). |
| `--env-mode` | `docker` | Backend to check (same 26 modes as [tier3 evaluate](#tier3-evaluate)). |
| `--environment-kwarg`, `--ek` | none | `KEY=VALUE` non-secret constructor option for a Harbor-native non-Docker backend (repeatable). Pass the same values intended for evaluation. Docker and local reject all kwargs; policy-owned fields are rejected for every eligible backend. |
| `--agent-model TEXT` | unset | Per-agent model override, `AGENT=MODEL` (repeatable) — check readiness with the model each agent will actually run. |
| `--verify-models` | off | Live per-agent catalog probe. Prints pass for verified access, warn for an inconclusive or non-authoritative result (success or failure), and fail only for a definitive credential, model, or configuration rejection. |
Expand All @@ -537,7 +537,7 @@ skillevaluator health-check
| Flag | Default | Effect |
| --- | --- | --- |
| `-a, --agents TEXT` | provider-native | Comma-separated agents to check. Defaults: NVIDIA Build=`opencode`, OpenAI=`codex`, Anthropic=`claude-code`. |
| `--env-mode` | `docker` | Backend to check (same 24 modes as [tier3 evaluate](#tier3-evaluate)). |
| `--env-mode` | `docker` | Backend to check (same 26 modes as [tier3 evaluate](#tier3-evaluate)). |
| `--environment-kwarg`, `--ek` | none | `KEY=VALUE` non-secret constructor option for a Harbor-native non-Docker backend (repeatable). Pass the same values intended for evaluation. Docker and local reject all kwargs; policy-owned fields are rejected for every eligible backend. |

## models
Expand Down
6 changes: 3 additions & 3 deletions docs/tier3-live-evaluation.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -390,7 +390,7 @@ explicit models — see [Agents & Sandboxes](agents-and-sandboxes.mdx).

## Where agents run

`--env-mode` selects one of 24 environment modes: 23 Harbor-native backends,
`--env-mode` selects one of 26 environment modes: 25 Harbor-native backends,
including the default Docker backend, plus SkillEvaluator's **local mode**,
which runs the agent CLI directly on the host under an OS-sandbox policy.
Managed backends such as `daytona`, `e2b`, and `modal` require their mapped
Expand All @@ -402,8 +402,8 @@ validate any managed backend's extra, credentials, infrastructure, and
constructor settings with `doctor` and its Harbor preflight before relying on
it.

Harbor's `cua-cloud`, `opensandbox`, `hf-sandbox`, `podman`, `kata`, `runta`,
`prime`, `mosaic`, and `smol` backends are not selectable because they cannot
Harbor's `cua-cloud`, `opensandbox`, `hf-sandbox`, `podman`, `kata`, `prime`,
and `smol` backends are not selectable because they cannot
yet receive SkillEvaluator's complete generated task bundle through a trusted
image or backend-native provisioning path.

Expand Down
89 changes: 87 additions & 2 deletions src/skillevaluator/tier3/harbor/runner.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@
import threading
import time
import tomllib
from collections.abc import Iterator, Mapping
from collections.abc import Iterable, Iterator, Mapping
from concurrent.futures import FIRST_COMPLETED, ThreadPoolExecutor, as_completed, wait
from contextlib import ExitStack, contextmanager, suppress
from dataclasses import dataclass, field
Expand Down Expand Up @@ -889,6 +889,21 @@ def _reserve_run_dir(results_root: Path, timestamp: str) -> Path:
),
"hyperbrowser": _DOCKER_HOST_ENV_VARS | frozenset({"HYPERBROWSER_API_KEY", "HYPERBROWSER_BASE_URL"}),
"vercel": frozenset({"VERCEL_OIDC_TOKEN", "VERCEL_PROJECT_ID", "VERCEL_TEAM_ID", "VERCEL_TOKEN"}),
"runta": frozenset({"RUNTA_CONFIG", "RUNTA_ENDPOINT", "RUNTA_TOKEN"}),

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] A relative RUNTA_CONFIG is forwarded but not anchored

The Runta SDK opens Path(RUNTA_CONFIG).expanduser() relative to Harbor's empty, evaluator-owned working directory. A relative or missing RUNTA_CONFIG still passes doctor, and then every trial fails with a missing token. Worse, if RUNTA_TOKEN is also set, the file's endpoint is silently ignored and the token goes to the default API.

Fix: anchor RUNTA_CONFIG (and MOSAIC_CONFIG/MAR_CONFIG) like the other host-path variables, and check that a configured file exists, as for MODAL_CONFIG_PATH.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in d245d41: RUNTA_CONFIG, MOSAIC_CONFIG and MAR_CONFIG are anchored like the other host-path variables, and the prerequisite check requires a configured file to exist. Tests: test_runta_and_mosaic_config_paths_survive_the_evaluator_owned_launch_directory and test_runta_and_mosaic_config_files_must_exist. The docs drop the absolute-path caveat.

# The Mosaic SDK still reads the MAR_* names its MOSAIC_* names replaced.
"mosaic": frozenset(
{
"MAR_API_TOKEN",
"MAR_CONFIG",
"MAR_ENDPOINT",
"MOSAIC_API_TOKEN",

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] A skill's runtime_env can redirect the operator's Mosaic token

This set also decides which names a skill's harbor.runtime_env may not set, through _is_operator_owned_runtime_name. It leaves out MOSAIC_CONFIG and the SDK's legacy names MAR_ENDPOINT and MAR_CONFIG (LEGACY_ENV_NAMES in mosaic_sandbox/credentials.py).

Scenario:

  1. A skill sets runtime_env: {MAR_ENDPOINT: https://attacker.example}. It is accepted, while the same value under MOSAIC_API_URL is rejected.
  2. _harbor_subprocess_environment places it in Harbor's environment.
  3. With MOSAIC_API_URL unset, the SDK resolves the attacker URL.
  4. The first snapshot call sends authorization: Bearer <MOSAIC_API_TOKEN or the mos login token> there.

A skill-controlled MOSAIC_CONFIG can point at a config file with an attacker endpoint to the same effect.

Fix: add MAR_API_TOKEN, MAR_ENDPOINT, MAR_CONFIG and MOSAIC_CONFIG to this set, so they are both forwarded for operators and blocked for skills. Add a _resolve_runtime_env regression test for each name.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in d245d41: MAR_API_TOKEN, MAR_ENDPOINT, MAR_CONFIG and MOSAIC_CONFIG are now in the Mosaic set. Operators' values are forwarded, and a skill's runtime_env can no longer set them. test_skill_runtime_env_cannot_redirect_runta_or_mosaic_credentials covers each name. test_mosaic_allowlist_covers_the_sdk_credential_names checks the set against the SDK's LEGACY_ENV_NAMES when mosaic-sandbox is installed; it passes against 0.14.7.

"MOSAIC_API_URL",
"MOSAIC_CONFIG",
"MOSAIC_REGISTRY_PASSWORD",
"MOSAIC_REGISTRY_USERNAME",
"MOSAIC_RETRIES",
}
),
}
_BEDROCK_HOST_ENV_VARS = _AWS_HOST_ENV_VARS | {
"AWS_BEARER_TOKEN_BEDROCK",
Expand Down Expand Up @@ -961,9 +976,12 @@ def _harbor_bin() -> str:
"DOCKER_CONFIG",
"GOOGLE_APPLICATION_CREDENTIALS",
"LANGSMITH_CONFIG_FILE",
"MAR_CONFIG",
"MODAL_CONFIG_PATH",
"MOSAIC_CONFIG",
"NETRC",
"REQUESTS_CA_BUNDLE",
"RUNTA_CONFIG",
"SINGULARITY_AUTHFILE",
"SINGULARITY_CONFIGDIR",
"SKYPILOT_GLOBAL_CONFIG",
Expand Down Expand Up @@ -1124,7 +1142,14 @@ def _nvidia_build_key_handoff(
"ec2": frozenset({"iam_instance_profile", "strict_host_key_checking"}),
"gke": frozenset({"memory_limit_multiplier"}),
"modal": frozenset({"volumes"}),
# Persistent volumes, persisted sandboxes, SSH access, and alternate build
# inputs would let state or access outlive one isolated trial or skip
# evaluator-staged Dockerfile layers.
"mosaic": frozenset({"build_args", "build_target", "enable_ssh", "persist", "volume"}),

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Mosaic snapshots are reused account-wide and never deleted

Harbor reuses any existing snapshot named harbor-<task name>-<12 hex of the environment id> and never deletes Mosaic snapshots, so the staged skill, inputs and repository context stay in the account. The with-skill and baseline arms never share a snapshot, because their build contexts differ. But anyone with write access to the account who creates that name first controls the image later evaluations run in.

Fix: either pass --force-build for mosaic, which gives each job its own names but retains a snapshot per job, or document that evaluations need a dedicated Mosaic account and that snapshots are reused and retained.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in d245d41 with documentation rather than --force-build. Forced builds give each job its own snapshot names, but Harbor still never deletes them, so they would multiply the retained skills and inputs and rebuild for every job. Snapshot names derive from the task and its build context, so the two arms never share one, and claiming a name requires write access to the operator's Mosaic account. The Mosaic section now says to run evaluations in an account only trusted operators can write to, and to delete old harbor-* snapshots.

"openshift": frozenset({"service_account_name"}),
# Direct mode runs commands in Runta's base runtime and ignores the task
# image; Harbor's automatic selection honors the task's Dockerfile or Compose.
"runta": frozenset({"mode"}),
"singularity": frozenset({"singularity_no_mount"}),
"use-computer": frozenset({"resources"}),
"vercel": frozenset({"ports"}),
Expand Down Expand Up @@ -1566,6 +1591,62 @@ def invalid_strings(*names: str) -> list[str]:
return []


# Backend SDK configuration files that operators may name in the host environment.
_BACKEND_CONFIG_FILE_ENV_VARS: Mapping[str, tuple[str, ...]] = MappingProxyType(
{"mosaic": ("MOSAIC_CONFIG", "MAR_CONFIG"), "runta": ("RUNTA_CONFIG",)}
)


def _backend_config_prerequisite_errors(env_mode: str) -> list[str]:
"""Check Runta and Mosaic credential inputs against what Harbor's child will see."""
errors: list[str] = []
for name in _BACKEND_CONFIG_FILE_ENV_VARS.get(env_mode, ()):
raw = os.environ.get(name)
if raw is None:
continue
try:
is_file = bool(raw.strip()) and Path(raw).expanduser().is_file()
except (OSError, RuntimeError):
is_file = False
if not is_file:
errors.append(f"Harbor environment '{env_mode}' requires {name} to name an existing regular file.")
if env_mode == "mosaic" and not errors:
try:
from mosaic_sandbox.credentials import resolve_token
except ImportError:
return errors # Harbor's preflight reports the missing extra.
_token, source = resolve_token()
# The SDK accepts a Mosaic token in E2B_API_KEY for E2B migrations, but
# that variable belongs to the e2b backend and never reaches Mosaic.
if source == "e2b_environment":
errors.append(
"Harbor environment 'mosaic' does not forward E2B_API_KEY; set MOSAIC_API_TOKEN or run `mos login`."
)
return errors


def _staged_task_environment_error(env_mode: str, task_dirs: Iterable[Path]) -> str | None:
"""Reject staged task environments the selected backend cannot run."""
if env_mode != "mosaic":
return None
for task_dir in task_dirs:
if (task_dir / "environment" / "docker-compose.yaml").exists():
return (
f"Harbor environment 'mosaic' runs one VM per trial and cannot run Docker Compose task "
f"'{task_dir.name}'; use a Dockerfile-only environment or another backend."
)
return None


def _harbor_missing_dependency_summary(exc: ImportError) -> str:
"""Keep Harbor's diagnosis, without its unpinned install commands or cloud-bundle hint."""
lines = str(exc).strip().splitlines()
summary = lines[0] if lines else type(exc).__name__
for suffix in (" Install it with:", " Install them with:"):
summary = summary.removesuffix(suffix)
return summary.strip().rstrip(".")


def _cwsandbox_prerequisite_errors(env_mode: str) -> list[str]:
"""Check what Harbor's cwsandbox backend needs; Harbor 0.24 no longer checks it."""
if harbor_environment_type(env_mode) != "cwsandbox":
Expand Down Expand Up @@ -1838,6 +1919,8 @@ def _check_prerequisites(

if cwsandbox_errors := _cwsandbox_prerequisite_errors(env_mode):
return cwsandbox_errors
if backend_config_errors := _backend_config_prerequisite_errors(env_mode):
return backend_config_errors
EnvironmentFactory.run_preflight(EnvironmentType(harbor_environment_type(env_mode)))
if env_mode == "ack":
ack_subprocess_env = (
Expand All @@ -1854,7 +1937,7 @@ def _check_prerequisites(
)
except ImportError as exc:
detail = redact_progress_detail(
exc,
_harbor_missing_dependency_summary(exc),
secret_values=secret_values_from_environment(os.environ),
)
return [
Expand Down Expand Up @@ -4167,6 +4250,8 @@ def _persist_pre_execution_failure(errors: list[str]) -> dict[str, Any]:
evaluator_skill_path=evaluator_skill_path,
arm_suffix=with_arm_suffix,
)
if staged_environment_error := _staged_task_environment_error(env_mode, task_paths):
raise ValueError(staged_environment_error)
task_selectors = validate_case_ids(task.name for task in task_paths)
logical_case_ids = validate_case_ids(_native_entry_id(task) for task in task_paths)
case_id_by_task_selector = dict(zip(task_selectors, logical_case_ids, strict=True))
Expand Down
Loading
Loading