Skip to content

feat: add direct BrowserGym MiniWoB evaluation - #1530

Merged
Yunnglin merged 8 commits into
mainfrom
feat/openenv-miniwob
Jul 31, 2026
Merged

Yunnglin merged 8 commits into
mainfrom
feat/openenv-miniwob

Conversation

@Yunnglin

Copy link
Copy Markdown
Collaborator

Summary

  • add a native MiniWoB benchmark backed directly by BrowserGym, with pinned task metadata and MiniWoB assets verified by SHA-256
  • add a reusable BrowserGym adapter that keeps Playwright lifecycle operations on a thread-affine worker while allowing model requests to remain concurrent
  • record reward-based success/error metrics, AgentTrace events, accessibility-tree observations, and per-step screenshots
  • remove the intermediate OpenEnv/runtime abstraction introduced earlier in this branch and consolidate URL/media helpers into uri_utils
  • update the report UI to render environment observations separately from user messages, preserve their metadata/tool-call linkage, show screenshots in the correct trace step, and format success-rate metrics correctly
  • update MiniWoB benchmark metadata, installation guidance, and English/Chinese documentation

Why

MiniWoB is already implemented by BrowserGym, so routing it through a second OpenEnv service added lifecycle, Docker, configuration, and compatibility complexity without improving the evaluation contract. Direct BrowserGym integration is smaller and preserves BrowserGym's task validation and rewards.

The dashboard also previously lost list-valued tool_call_id and message metadata during serialization. As a result, screenshot observations emitted after tool calls appeared as blank user messages. The serializer and trace grouping now retain that relationship and present these messages as environment observations.

User impact

Users can run MiniWoB through the normal EvalScope benchmark flow with image-capable function-calling models. Reports include BrowserGym rewards, errors, accessibility trees, screenshots, and a visualizable agent trace. Existing public APIs on main are not removed; the discarded OpenEnv APIs existed only in earlier commits of this feature branch.

Validation

  • python -m pre_commit run --all-files
  • pytest tests/benchmark/test_miniwob.py tests/benchmark/test_agent.py::TestAgentBenchmark::test_miniwob tests/agent/test_agent_loop.py tests/agent/test_t2_environment.py tests/api/test_text2speech_model.py tests/test_download_utils.py -q — 94 passed
  • npm test -- --run src/api/schemas/reports.schema.test.ts src/components/single/ChatView.test.tsx src/domain/metric/registry.test.ts — 22 passed
  • npm run build
  • manual MiniWoB E2E run through evalscope service, including trace rendering, two environment screenshot observations, and 100.0% score display

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@Yunnglin
Yunnglin marked this pull request as ready for review July 30, 2026 12:57
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@Yunnglin Yunnglin added the qoder-review Add to a PR to trigger Qoder code review label Jul 30, 2026

@qoderai qoderai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👋 Review Summary

This PR adds a native BrowserGym-backed MiniWoB benchmark, centralises URL/data/file/media helpers, and extends the agent loop and report schemas to capture environment reset events, rich tool outputs, and environment observations. Overall the design is thoughtful, with strong integrity checks and good test coverage for MiniWoB, download utilities, and trace/UI behaviour.

🛡️ Key Risks & Issues

  • ToolExecutionOutput now allows tools to attach arbitrary ContentImage entries whose image field is later rendered via the reports/media endpoint. In the MiniWoB flow those paths are under /tmp/miniwob/..., but if a future tool (or benchmark) sets image to a host path outside the run’s artifact directory, the reporting API could unintentionally serve arbitrary host files. It would be safer if the backend strictly whitelisted media paths under an outputs root, rejected absolute/parent-traversing paths, and documented ToolExecutionOutput/ContentImage usage as “artifact paths only”.
  • The new uri_utils/file_as_data helper transparently handles data URIs, HTTP(S) URLs, and local paths, and is reused across audio/vision metrics and model utilities. If any caller ever passes attacker-controlled URLs or file paths (for example from dataset metadata or tool outputs), this function will happily perform HTTP GETs or read arbitrary files. Although this pattern existed in url_utils before, centralising it widens its usage; it’s worth treating this as a privileged API, documenting the risk, and considering domain/path restrictions or size limits.
  • TaskConfig.update now has special logic for agent_config: when merging two dicts with different mode values, it silently resets the current agent_config before deep-merge. This is a reasonable attempt to separate native vs external configs, but it can surprise callers who expect incremental overrides and may inadvertently drop pre-existing agent settings. Explicit tests and documentation for this behaviour would help avoid subtle configuration regressions.
  • The download_url helper is now used by multiple benchmarks and performs HEAD/GET requests against hardcoded external URLs when sha256 is not provided. While AA-LCR, ClawEval, and MiniWoB are legitimate benchmarks, a central framework helper that performs network downloads should ideally have clear documentation around trust and possibly a central allowlist of domains, so that new benchmarks cannot accidentally introduce risky download endpoints without review.

🧪 Verification Advice

  • For environment-observation flows (MiniWoB and future benchmarks), manually verify that only files under the run’s outputs/artifacts directory can be fetched through the reports/media endpoint, and that absolute or parent-traversing paths are rejected. Try a misconfigured tool that sets an attachment path outside the run root.
  • Exercise MiniWoB with both successful and intentionally failing BrowserGym episodes (e.g., mock reset/step exceptions) to confirm success_rate/error_rate behaviour, AgentTrace error events, and UI rendering of failures. This will help ensure the scoring and trace semantics are correct under error conditions.
  • For agent_config merging, run a small suite where you start from a native agent TaskConfig and update it with an external agent_config dict, then inspect the resulting config to confirm that only the intended fields remain and that mode switches don’t leave stale options behind.
  • In environments where EvalScope is deployed on shared or sensitive machines, review the set of benchmarks that call download_url and file_as_data, and validate that their URLs/paths point only to trusted sources and bounded assets.

💡 Thoughts & Suggestions

  • The MiniWoB/BrowserGym integration, deterministic repeat scheduling, and action validation look solid and well-tested, and the rich ToolExecutionOutput/trace wiring is a nice step toward more transparent agent benchmarking. Adding a bit more documentation around the security assumptions (local vs sandboxed environments, trusted asset domains, and allowed media paths) would make it easier for users to understand where they might need to tighten controls in their own deployments.

🤖 Generated by Qoder • View workflow run

expect(screen.queryByText('User')).not.toBeInTheDocument()
const images = container.querySelectorAll('img')
expect(images).toHaveLength(2)
expect(images[0]).toHaveAttribute(

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ToolExecutionOutput and the agent loop correctly propagate attachments and metadata, but the reports/media pipeline now trusts whatever paths tools put into ContentImage.image.

In this MiniWoB flow the images are under /tmp/miniwob/..., which is fine for a controlled benchmark, but if a future tool sets image to a path outside the run's artifact directory (e.g., /etc/passwd or a user home file), the /api/v1/reports/media/file?path=... endpoint could end up serving arbitrary host files unless the backend route enforces strict whitelisting.

Given ToolExecutionOutput is now a general contract for tools, it would be safer if the media endpoint only served files under a known outputs root and rejected absolute or parent-traversing paths, and if ToolExecutionOutput/ContentImage usage were documented as "artifact paths only" to avoid leaking host filesystem contents through the UI.


🤖 Generated by Qoder

@Yunnglin
Yunnglin merged commit e66e415 into main Jul 31, 2026
5 checks passed
@Yunnglin
Yunnglin deleted the feat/openenv-miniwob branch September 10, 2026 07:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

qoder-review Add to a PR to trigger Qoder code review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant