Skip to content

test(e2e): add deterministic end-to-end suite and required CI check - #41

Merged
cocofhu merged 1 commit into
Tencent:mainfrom
Yanami-2K:test/e2e-suite
Oct 11, 2026
Merged

cocofhu merged 1 commit into
Tencent:mainfrom
Yanami-2K:test/e2e-suite

Conversation

@Yanami-2K

Copy link
Copy Markdown
Contributor

What changed

Closes the P0 part of #31: merges currently rely on static checks, install smoke tests and manual testing. No test installs the package the way users do and then runs the workflow end to end. This PR adds a deterministic E2E suite, bash scripts/e2e.sh, and a CI workflow so the suite can be required before a merge and a human reviews after it passes.

How it stays deterministic (no API keys, no real Agent CLI)

  • tests/e2e/fake_agent.py stands in for the host Agent. It loads the installed Skill and follows the same resume-driven protocol a real host follows: init, prepare --emit-prompt, assign, start, finish, approve --user-confirmed, decide-knowledge and rotate-team. The only difference is that it writes stage outputs recorded in fixtures/scenario-medium.json instead of calling an LLM.
  • The developer stage really patches a sample project whose tests start out failing, and the test stage really runs those tests.
  • Fault-injection flags (--crash-in, --stop-after, --no-user-confirmation, --review-reject, --leave-placeholders, --fail-stage) drive the failure paths.
  • The CLI under test comes from npm pack, installed into a temporary prefix, which is the same path users take with npx loopforge-cli.

Scenarios (13 tests)

Area Coverage
Install (test_e2e_install.py) All 9 edition × host combinations: entrypoint files, status --json, doctor, idempotent update, clean uninstall that keeps user files. Rejects Classic+pi and a second edition without --force.
Portable workflow (test_e2e_portable_workflow.py) A medium task through clarify → user confirmation → design → implement → independent review → test → knowledge (skipped) → summary on all 5 hosts. Executor type and dispatch tool match the installed adapter, and each agent stage gets its own executor. Resume after a crash mid-stage (claude: continue-executor; codebuddy: --fresh-team → rotate-team). Resume after an interruption between actions, with no stage run twice. The requirement gate blocks DESIGN until the user confirms. A rejected review goes back to IMPLEMENT. Unfilled artifacts are rejected. Repeated failures end in blocked.
Classic runtime and dashboard (test_e2e_classic_runtime.py) CodeBuddy auto-dispatch routes TASK-01…TASK-05 to the next role and records the decisions. It falls back to main on a target mismatch, a failed event, workflow end and manual gates. A recorded session goes through the 5 observability hooks, then build_dashboard_data.py, then the dashboard served over HTTP. The test checks sessions, tool calls, tokens, the devflow run and the auto-dispatch stats, with CLS env vars removed so the run stays offline.

CI

  • New .github/workflows/e2e.yml: runs on Python 3.8 and 3.12 with Node 22 and a 15-minute timeout. An e2e-required gate job gives branch protection one stable check name.
  • release.yml runs the suite before publishing.
  • AGENTS.md, CONTRIBUTING (EN/ZH) and the PR template list bash scripts/e2e.sh.

Maintainer action needed: a PR can't change branch protection. Please mark e2e-required as a required status check on main, ideally together with validate and secrets.

Scope

  • Classic
  • Portable
  • Shared behavior contract
  • Installer / CLI
  • Documentation only

Tests and CI only. Runtime files are unchanged, and tests/e2e is not part of the npm package (checked with npm pack --dry-run).

Compatibility

No runtime behavior change. The suite uses only the Python standard library (3.8+), Node.js 20+ and npm.

Validation

Run locally on Linux, with the CLI installed from npm pack:

bash scripts/e2e.sh                         # 13 tests OK on Python 3.13 (~64s) and 3.8.20 (~65s)
bash scripts/validate.sh                    # passed on Python 3.13 and 3.8
bash scripts/smoke-install.sh               # passed
bash scripts/scan-secrets.sh                # no leaks found

Mutation check: with required_approval_gates set to [], test_requirement_gate_waits_for_explicit_user_confirmation fails (expected exit 3, got 0).

Not run: macOS runners and real Agent CLIs.

Follow-ups (separate PR)

  • small/SOLO and overflow, large with KNOWLEDGE=run, design-artifact tasks, and manual-mode gates.
  • Render the dashboard in Playwright.
  • Node 20 and macOS runners.
  • Classic link-closure checks.
  • Bring the 129 agent-observability unit tests into CI. 6 fail locally today: 5 in test_cls_sink, 1 in test_collector.
  • Those unit tests write logs/cls-push-debug.ndjson into the source tree. .npmignore overrides .gitignore, so the log can end up in a published package. The npm package also ships the observability tests/ folders.
  • Optional nightly smoke with a real Agent, non-blocking.

Safety

  • I did not include credentials, internal endpoints, production data, or organization-specific infrastructure.
  • I preserved user files and unrelated changes.
  • I updated generated Classic host bundles when applicable. (N/A: no Classic changes)
  • I documented third-party content and licensing changes when applicable. (N/A)

Install the packed npm CLI into throw-away projects for every host and
edition, drive the installed Portable workflow with a fake host Agent that
follows the resume protocol using recorded stage outputs, and exercise the
Classic CodeBuddy auto-dispatch and observability hooks through the local
dashboard. Covers resume after crash/interruption, the requirement
confirmation gate, review rejection, unfilled artifacts and repeated
failures. No network access or API keys are needed.

Adds .github/workflows/e2e.yml with a stable 'e2e-required' gate job and
runs the suite before releases.

Refs Tencent#31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants