Skip to content

🤖 tests: run make bug-bash in the bug-bash sandbox, with the mock app AI - #6076

Merged
ThomasK33 merged 3 commits into
mainfrom
tests/5714-b2-bug-bash
Oct 10, 2026
Merged

ThomasK33 merged 3 commits into
mainfrom
tests/5714-b2-bug-bash

Conversation

@ThomasK33

Copy link
Copy Markdown
Member

Summary

Runs make bug-bash (e2e explore charters) in the bug-bash sandbox: each charter is one sandboxed job with its own provider proxy, and all jobs spend from one list-price budget. Nothing model-driven runs on the host any more, and the app AI is the mock. This is PR B2 of the #5714 plan, stacked on B1 (#6073).

Refs #5714

Background

B1 turned the sandbox on for the MCP Apps suite. make bug-bash still spawned e2e explore on the host, where the host pause refused it. B2 moves every charter into the sandbox. The real app AI through the proxy is the next step (B3), so this PR runs the mock app AI only.

Implementation

  1. run.ts no longer spawns e2e. Each charter is one in-process launchJob() (sandbox/launch.ts) with its own stop signal and its own proxy, and every job reserves from one Ledger funded by BUGBASH_BUDGET_USD. A proxy fault (a call over its bound, a call record that was not written, a call that outlived close()) starts no new charter and stops the running ones. findings.md gains a cost line: calls, budget refusals, list-price spend and tokens per model.
  2. run.ts refuses any BUGBASH_AI other than unset or mock, and checks every model, the budget and the provider settings, before it creates a folder or starts a job.
  3. launch.ts gains the explore job kind, with an argument allowlist (inputs.ts exploreRefusal: one goal, the two bug-bash configs, --target, --agent, --output, --max-steps 1-12, --video=on, --reporter). launchJob() returns a JobOutcome with the job code, cleanup state, stop and proxy fault kept apart. exitFor() keeps the command-line exit codes. New options: a shared ledger, a log sink, and a file descriptor for the container's stderr (the charter log).
  4. Exit precedence, the same in the launcher and in run.ts: unknown cleanup (3), then a stop (130 or 143), then a proxy fault (5), then the charters' codes.
  5. hostPause.ts allows e2e explore only in a model-driven sandbox job. e2e.config.ts gives its agents the proxy model only there (sandbox/explorerModel.ts, shared with the MCP Apps config), so e2e run with this config never gets a model. BUGBASH_MODEL_DRIVEN now reaches the app for explore jobs.

Threat-model claims

ID Claim Where B2 wires or tests it
S1 Model-driven commands run only in a sandbox container with the proxy socket e2eCommandRefusal allows explore only when modelDrivenSandbox() holds. Tests: a fake sandbox root passes, the host and a container without the marker refuse
S2 Model-driven jobs start only with the proxy, and repro agents stay model-free explorerModel() returns a model only for explore in the sandbox, and never for e2e run. run.ts reaches the model only through launchJob() (a test asserts no host spawn)
S3 Agent tools, terminals and project automation off in every model-driven job BUGBASH_MODEL_DRIVEN is in e2e.config.ts's app env. The transport check read it in every container (below)

Validation

  1. Unit tests: 219 pass in tests/bugbash/. These tests fail when their fix is reverted: the exit order, the halt after a proxy fault, and explore allowed only in the sandbox.
  2. Transport check with the fake upstream, real Docker, at this head: 4 charters × 2 models with --parallel 8. 8 containers ran at once, 352 calls settled through 8 proxies into one ledger, every job exited 0, and the run took 15 s with 0 containers left. A read-only docker exec found XUM_DISABLE_AGENT_TOOLS=1, XUM_DISABLE_PROJECT_AUTOMATION=1 and XUM_DISABLE_TERMINALS=1 in the app environ of all 8 containers. Peak memory was 506 MiB per container and 2,327 MiB for the 8, so the default --parallel 8 stays.
  3. Functional acceptance (one paid bug bash, mock app AI) follows as a PR comment.

Risks

Test-only code. The new runtime surface is up to --parallel concurrent sandbox jobs, each with a proxy socket on the host, only during a make bug-bash run with an explicit budget and key.

Known gaps (tracked in the plan for B3 and PR D):

  1. The real app AI is refused here. It comes in B3.
  2. e2e's own exit 3 still shares the launcher's exit 3 on the command line. run.ts keeps them apart.
  3. After a SIGKILL, the launcher's temp folders stay behind.
  4. AGENTS.md still says bug-bash and mcp-apps-e2e are paused. PR D updates the docs.

Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high • Cost: $97.07

@ThomasK33
ThomasK33 added this pull request to stack #6070 October 10, 2026 18:11
@ThomasK33

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-10T18:43:57.126202Z c976412 New commits
🔒 Security Review ✅ Completed 2026-10-10T18:44:51.306676Z c976412 New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: d07865edf3

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d07865edf3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tests/bugbash/run.ts Outdated
Comment thread tests/bugbash/sandbox/inputs.ts Outdated
@ThomasK33

Copy link
Copy Markdown
Member Author

Paid bug bash (functional acceptance of B2)

One make bug-bash at d07865edf3: the 6 charters × anthropic:claude-sonnet-5-5, --max-steps 6, BUGBASH_BUDGET_USD=10, mock app AI. make exited 0 after 290 s. All 6 containers ran at once, and 0 were left afterwards.

Charter Job exit Steps Proxy calls Cleanup
artifacts-keyboard 1 (1 issue) 5 of 6, goal covered 40 settled removed
composer 0 3 of 6, goal covered 29 settled removed
workspace-lifecycle 0 4 of 6, goal covered 48 settled removed
settings 0 6 of 6, step limit 51 settled removed
mock-errors 0 2 of 6, goal covered 18 settled removed
phone-layout 0 3 of 6, goal covered 24 settled removed
  1. Cost: findings.md reports 210 proxied calls, 0 refused by the budget, $2.3614 of $10.0000 at list price. In the proxy records, all 210 calls are HTTP 200 and settled, 0 are kept in full, 0 are refused, and boundExceeded = 0. The highest reservation was $0.6436.
  2. S3: a read-only docker exec read the app environ of all 6 job containers. Each had XUM_DISABLE_AGENT_TOOLS=1, XUM_DISABLE_PROJECT_AUTOMATION=1 and XUM_DISABLE_TERMINALS=1.
  3. Functional proof: every charter took steps against the app, for example "Open Settings overview", "Create workspace" and "Rename back, archive, restore", "Empty and whitespace messages", "Artifacts view on phone". There was one finding: "Ctrl+Shift+K does nothing once Artifacts is open" (an explorer claim, not yet triaged).

Generated with xum • Model: anthropic:claude-opus-5-5 • Thinking: high

@ThomasK33

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 2366c986c9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tests/bugbash/run.ts Outdated
@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: 2366c986c9

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

Base automatically changed from tests/5714-b1-mcp-apps to main October 10, 2026 18:31
- run.ts no longer spawns `e2e explore` on the host. Each charter is one
  in-process launchJob() (sandbox/launch.ts) with its own stop, its own proxy
  and the run's one Ledger (BUGBASH_BUDGET_USD). A proxy fault starts no new
  charter and stops the running ones. Exit precedence: unknown cleanup (3),
  then a stop, then a proxy fault (5). findings.md gains a cost line.
- launch.ts: the `explore` job kind with its argument allowlist
  (inputs.ts exploreRefusal); launchJob() returns a JobOutcome and exitFor()
  keeps the CLI exit codes. Options for a shared ledger, a log sink and a
  stderr fd (runner.ts).
- hostPause.ts allows `e2e explore` only in a model-driven sandbox job, and
  e2e.config.ts gives its agents the proxy model only there
  (sandbox/explorerModel.ts). BUGBASH_MODEL_DRIVEN reaches the app (S3).
- Only the mock app AI runs: run.ts refuses another BUGBASH_AI before anything.

---
_Generated with [`xum`](https://github.com/coder/xum) • Model: `anthropic:claude-opus-5-5` • Thinking: `high`_
…onfig up front

- run.ts dropped every charter exit 130 from the final code, so a charter that
  was interrupted on its own (not by the run's stop or a proxy fault) could end
  in a successful run. Those two cases already return earlier, so the final
  max keeps every code now.
- run.ts accepted any existing --config, but the launcher runs only the two
  bug-bash configs: the run made its folders and then every job refused. It
  now refuses another config before any folder (exit 2), and the header says so.

---
_Generated with [`xum`](https://github.com/coder/xum) • Model: `anthropic:claude-opus-5-5` • Thinking: `high`_
@ThomasK33
ThomasK33 force-pushed the tests/5714-b2-bug-bash branch from 2366c98 to 0927e1b Compare October 10, 2026 18:31

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0927e1bfc6

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tests/bugbash/hostPause.ts
…riven runs in the sandbox

- proxy.ts reports a bound miss at once (onFault). The launcher stops that job
  and tells its caller (onProxyFault); a failed record write takes the same
  path. run.ts then halts every running job instead of waiting for the faulty
  charter to end (Codex P2).
- run.ts gives each run folder a random suffix: two runs in one millisecond
  collided on it (a flake in run.test.ts).
- docs/AGENTS.md: the host-pause and Agent E2E lines now say that model-driven
  runs run only in the bug-bash sandbox, refuse on the host, and need the
  budget and the provider settings (Codex P1). builtInSkillContent regenerated.

---
_Generated with [`xum`](https://github.com/coder/xum) • Model: `anthropic:claude-opus-5-5` • Thinking: `high`_
@mintlify

mintlify Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Preview deployment for your docs. Learn more about Mintlify Previews.

Project Status Preview Updated
Mux 🟢 Ready View Preview Oct 10, 2026, 6:40 PM

💡 Tip: Enable Automations to automatically generate PRs for you.

@ThomasK33
ThomasK33 added this pull request to the merge queue Oct 10, 2026
Merged via the queue into main with commit b34713f Oct 10, 2026
33 checks passed
@ThomasK33
ThomasK33 deleted the tests/5714-b2-bug-bash branch October 10, 2026 19:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant