Skip to content

Add JobBench benchmark - #1509

Merged
Yunnglin merged 2 commits into
mainfrom
agent/add-job-bench
Jul 23, 2026
Merged

Yunnglin merged 2 commits into
mainfrom
agent/add-job-bench

Conversation

@Yunnglin

Copy link
Copy Markdown
Collaborator

Summary

  • add the native job_bench benchmark backed by the ModelScope evalscope/job-bench mirror
  • run tasks through EvalScope's agent loop with local or Docker workspaces and persistent deliverable artifacts
  • port JobBench rubric judging, document extraction, visual attachments, and weighted score aggregation
  • add generated benchmark documentation and a real API Docker E2E test

User impact

Users can evaluate agents on JobBench with EvalScope's existing runner and judge configuration. Reference files are
resolved from the cached ModelScope snapshot, while task deliverables remain available under the EvalScope output
directory after the sandbox is removed.

Validation

  • pre-commit run --files ... passed
  • make docs-pipeline BENCHMARK=job_bench FORCE=1 passed
  • real API E2E passed with main split, limit=5, eval_batch_size=5, Docker, and max_steps=80
  • E2E produced 5 predictions, 5 reviews, 5 artifact directories, and 11 deliverable files

@Yunnglin Yunnglin added the qoder-review Add to a PR to trigger Qoder code review label Jul 23, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@qoderai qoderai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👋 Review Summary

This PR adds the JobBench agentic benchmark, including dataset meta/docs, a new adapter that integrates with the agent loop and LLM judging, artifact handling utilities, and an end-to-end Docker-based test. Overall the design aligns well with existing agent benchmarks and provides a solid foundation for rubric-based deliverable evaluation.

🛡️ Key Risks & Issues

  • JobBench’s scoring relies heavily on the rubric parsing and evaluation pipeline in evalscope/benchmarks/job_bench/utils.py. While the happy path looks good, error handling for SQLite databases in _sqlite_to_text currently leaves connections open when an exception is thrown, which can lead to accumulated file descriptors and locked DB files in longer or parallel runs. This is mainly a robustness/performance concern rather than a correctness bug, but it is worth tightening.
  • The adapter’s match_score returns a placeholder zeroed value map and marks official_score_computed=False, while llm_match_score computes the true rubric metrics. When combined with LLMJudge’s generic LLM_RECALL merging strategy, there’s some risk of configuration where the main score comes from the LLM but the value dict still reflects the placeholder zeros. For JobBench we should treat LLM judging as the authoritative source and ensure score merging doesn’t silently report inconsistent metric values.

🧪 Verification Advice

  • Continue running the new test_job_bench end-to-end test in environments where DASHSCOPE_API_KEY and Docker are available; it’s a valuable integration check for dataset loading, agent loop execution, artifact creation, and LLM rubric scoring.
  • Add focused unit tests around parse_rubrics, evaluate_job_bench_output, and judge_rubric using a fake judge to validate rubric parsing variants, error branches (no rubrics, no output files, unreadable outputs), and scorecard aggregation. This will help catch subtle logic issues without relying on external APIs.
  • Consider small adapter-level tests that exercise both the local and Docker environment paths for JobBench to ensure reference file mounting, artifact directory setup, and jobbench_output wiring behave correctly across configuration variants.

💡 Thoughts & Suggestions

  • To harden resource handling, update _sqlite_to_text to always close its connection using a context manager (with sqlite.connect(...) as con:) or a try/finally block; this will keep JobBench evaluations resilient when encountering malformed or large SQLite files.
  • For JobBench specifically, it may be safer to either avoid judge_strategy='llm_recall' or adjust LLMJudge’s merging behavior so that when a benchmark provides its own LLM-based metric values, those values become the authoritative Score.value rather than placeholder zeros. Documenting the recommended judge strategy for JobBench in the benchmark docs would also help users avoid misconfiguration.
  • The artifact capture via JobBenchArtifactEnvironment is a nice touch; as you evolve this benchmark, keeping output file formats and directory structure stable will make downstream analysis and debugging much easier.

🤖 Generated by Qoder • View workflow run

con.close()
return '\n'.join(parts)
except Exception as exc:
return f'[ERROR: Failed to read SQLite {path.name}: {exc}]'

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The _sqlite_to_text helper opens a SQLite connection but only calls con.close() on the success path. If any exception is raised after sqlite.connect (for example while reading schema, counting rows, or sampling table data), execution jumps to the except block, returns the error string, and the connection is never explicitly closed.

Because this utility can be used repeatedly on problematic databases in longer JobBench runs, leaked connections and file descriptors may accumulate and keep OS-level locks on the DB files, which can degrade performance or interfere with later reads.

To make this more robust, it would be safer to ensure the connection is always closed via a context manager or a finally block. For example, using with sqlite.connect(str(path)) as con: or moving con.close() into finally so both success and failure paths reliably release the database handle.


🤖 Generated by Qoder • Fix in Qoder

@Yunnglin
Yunnglin marked this pull request as ready for review July 23, 2026 09:36
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@Yunnglin
Yunnglin merged commit c897ec7 into main Jul 23, 2026
3 checks passed
@Yunnglin
Yunnglin deleted the agent/add-job-bench branch September 10, 2026 07:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

qoder-review Add to a PR to trigger Qoder code review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant