Repository navigation
🤖 tests: run make bug-bash in the bug-bash sandbox, with the mock app AI - #6076
Conversation
|
@codex review |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
🛡️ Codex Security Review · Automatically triggeredSecurity review completed. No security issues were found in this pull request. Reviewed commit: Only the user who started this review can view the report in Codex. ℹ️ About Codex security reviews in GitHubThis is an experimental Codex feature. Security reviews are triggered when:
Once complete, Codex will leave suggestions, or a comment if no findings are found. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d07865edf3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Paid bug bash (functional acceptance of B2)One
Generated with |
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2366c986c9
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
🛡️ Codex Security Review · Automatically triggeredSecurity review completed. No security issues were found in this pull request. Reviewed commit: Only the user who started this review can view the report in Codex. ℹ️ About Codex security reviews in GitHubThis is an experimental Codex feature. Security reviews are triggered when:
Once complete, Codex will leave suggestions, or a comment if no findings are found. |
- run.ts no longer spawns `e2e explore` on the host. Each charter is one in-process launchJob() (sandbox/launch.ts) with its own stop, its own proxy and the run's one Ledger (BUGBASH_BUDGET_USD). A proxy fault starts no new charter and stops the running ones. Exit precedence: unknown cleanup (3), then a stop, then a proxy fault (5). findings.md gains a cost line. - launch.ts: the `explore` job kind with its argument allowlist (inputs.ts exploreRefusal); launchJob() returns a JobOutcome and exitFor() keeps the CLI exit codes. Options for a shared ledger, a log sink and a stderr fd (runner.ts). - hostPause.ts allows `e2e explore` only in a model-driven sandbox job, and e2e.config.ts gives its agents the proxy model only there (sandbox/explorerModel.ts). BUGBASH_MODEL_DRIVEN reaches the app (S3). - Only the mock app AI runs: run.ts refuses another BUGBASH_AI before anything. --- _Generated with [`xum`](https://github.com/coder/xum) • Model: `anthropic:claude-opus-5-5` • Thinking: `high`_
…onfig up front - run.ts dropped every charter exit 130 from the final code, so a charter that was interrupted on its own (not by the run's stop or a proxy fault) could end in a successful run. Those two cases already return earlier, so the final max keeps every code now. - run.ts accepted any existing --config, but the launcher runs only the two bug-bash configs: the run made its folders and then every job refused. It now refuses another config before any folder (exit 2), and the header says so. --- _Generated with [`xum`](https://github.com/coder/xum) • Model: `anthropic:claude-opus-5-5` • Thinking: `high`_
2366c98 to
0927e1b
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0927e1bfc6
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…riven runs in the sandbox - proxy.ts reports a bound miss at once (onFault). The launcher stops that job and tells its caller (onProxyFault); a failed record write takes the same path. run.ts then halts every running job instead of waiting for the faulty charter to end (Codex P2). - run.ts gives each run folder a random suffix: two runs in one millisecond collided on it (a flake in run.test.ts). - docs/AGENTS.md: the host-pause and Agent E2E lines now say that model-driven runs run only in the bug-bash sandbox, refuse on the host, and need the budget and the provider settings (Codex P1). builtInSkillContent regenerated. --- _Generated with [`xum`](https://github.com/coder/xum) • Model: `anthropic:claude-opus-5-5` • Thinking: `high`_
|
Preview deployment for your docs. Learn more about Mintlify Previews.
💡 Tip: Enable Automations to automatically generate PRs for you. |
Summary
Runs
make bug-bash(e2e explorecharters) in the bug-bash sandbox: each charter is one sandboxed job with its own provider proxy, and all jobs spend from one list-price budget. Nothing model-driven runs on the host any more, and the app AI is the mock. This is PR B2 of the #5714 plan, stacked on B1 (#6073).Refs #5714
Background
B1 turned the sandbox on for the MCP Apps suite.
make bug-bashstill spawnede2e exploreon the host, where the host pause refused it. B2 moves every charter into the sandbox. The real app AI through the proxy is the next step (B3), so this PR runs the mock app AI only.Implementation
run.tsno longer spawns e2e. Each charter is one in-processlaunchJob()(sandbox/launch.ts) with its own stop signal and its own proxy, and every job reserves from oneLedgerfunded byBUGBASH_BUDGET_USD. A proxy fault (a call over its bound, a call record that was not written, a call that outlivedclose()) starts no new charter and stops the running ones.findings.mdgains a cost line: calls, budget refusals, list-price spend and tokens per model.run.tsrefuses anyBUGBASH_AIother than unset ormock, and checks every model, the budget and the provider settings, before it creates a folder or starts a job.launch.tsgains theexplorejob kind, with an argument allowlist (inputs.tsexploreRefusal: one goal, the two bug-bash configs,--target,--agent,--output,--max-steps1-12,--video=on,--reporter).launchJob()returns aJobOutcomewith the job code, cleanup state, stop and proxy fault kept apart.exitFor()keeps the command-line exit codes. New options: a shared ledger, a log sink, and a file descriptor for the container's stderr (the charter log).run.ts: unknown cleanup (3), then a stop (130 or 143), then a proxy fault (5), then the charters' codes.hostPause.tsallowse2e exploreonly in a model-driven sandbox job.e2e.config.tsgives its agents the proxy model only there (sandbox/explorerModel.ts, shared with the MCP Apps config), soe2e runwith this config never gets a model.BUGBASH_MODEL_DRIVENnow reaches the app for explore jobs.Threat-model claims
e2eCommandRefusalallowsexploreonly whenmodelDrivenSandbox()holds. Tests: a fake sandbox root passes, the host and a container without the marker refuseexplorerModel()returns a model only forexplorein the sandbox, and never fore2e run.run.tsreaches the model only throughlaunchJob()(a test asserts no host spawn)BUGBASH_MODEL_DRIVENis in e2e.config.ts's app env. The transport check read it in every container (below)Validation
tests/bugbash/. These tests fail when their fix is reverted: the exit order, the halt after a proxy fault, andexploreallowed only in the sandbox.--parallel 8. 8 containers ran at once, 352 calls settled through 8 proxies into one ledger, every job exited 0, and the run took 15 s with 0 containers left. A read-onlydocker execfoundXUM_DISABLE_AGENT_TOOLS=1,XUM_DISABLE_PROJECT_AUTOMATION=1andXUM_DISABLE_TERMINALS=1in the app environ of all 8 containers. Peak memory was 506 MiB per container and 2,327 MiB for the 8, so the default--parallel 8stays.Risks
Test-only code. The new runtime surface is up to
--parallelconcurrent sandbox jobs, each with a proxy socket on the host, only during amake bug-bashrun with an explicit budget and key.Known gaps (tracked in the plan for B3 and PR D):
run.tskeeps them apart.bug-bashandmcp-apps-e2eare paused. PR D updates the docs.Generated with
xum• Model:anthropic:claude-opus-5-5• Thinking:high• Cost:$97.07