From 7b561631730f31675248862595b3334f350c58c6 Mon Sep 17 00:00:00 2001 From: genitrix Date: Wed, 29 Jul 2026 23:44:27 +0800 Subject: [PATCH] Add ClawBench to benchmark integrity resources --- README.md | 1 + 1 file changed, 1 insertion(+) diff --git a/README.md b/README.md index 4ea2ad2..ee58e3c 100644 --- a/README.md +++ b/README.md @@ -268,6 +268,7 @@ Most "awesome" lists are link dumps. This one is **annotated and verified**: eve - **[Search-Time Contamination in Deep Research Agents](https://arxiv.org/abs/2606.05241)** โ€” Wang, Zhang, Yao, Zeng, Song, Lin, Shen โ€” ยท *paper* โ€” Defines three contamination types (Benchmark Metadata Leakage, Question-Context Leakage, Explicit Answer Leakage) for agents that web-search during evaluation; applies detection algorithms across six benchmarks and finds performance inflation up to 4%. "Such agents may retrieve public benchmark metadata, question context, or even ground-truth answers via web search. This gives rise to Search-Time Contamination (STC), where external retrieval bypasses intended reasoning and inflates measured performance." Distinct from training contamination; advocates isolated sandboxes and transparent search trajectories. ๐Ÿ†• - **[RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents](https://arxiv.org/abs/2603.11337)** โ€” Yonas Atinafu, Robin Cohen โ€” ยท *paper* โ€” Treats evaluation integrity as a first-class benchmarked outcome instead of an assumption: two compromise vectors (evaluator tampering, train/test leakage) instrumented via patch tracking and runtime file-access logging; in natural agent runs evaluator-tampering attempts appear in ~50% of episodes, and evaluator locking eliminates them at a 25โ€“31% median runtime overhead. ๐Ÿ†• - **[Prediction: A Frontier Open Source LLM Will Be Released On 3rd December 2026](https://blog.doubleword.ai/frontier-os-llm)** โ€” Jamie Dborin (Doubleword) โ€” ยท *blog* โ€” Extrapolates the Artificial Analysis Intelligence Index across 18 constituent benchmarks to forecast the open-vs-closed capability gap; a naive fit of the headline index points to convergence by 3 Dec 2026, but the average per-benchmark lag holds steady at ~5 months โ€” a worked cautionary case in reading trends off an aggregate leaderboard index rather than its components. ๐Ÿ†• +- **[ClawBench: Can AI Agents Complete Everyday Online Tasks?](https://arxiv.org/abs/2604.08523)** โ€” Zhang et al. (TIGER-AI-Lab) โ€” ยท *benchmark* โ€” Evaluates browser agents on 281 everyday tasks (V1 152 + V2 129) across 163 live production websites and 15 categories. A submission-interception layer prevents real-world side effects while preserving realistic workflows, and the harness captures browser actions, HTTP traffic, screenshots, recordings, and agent messages for reproducible, fine-grained analysis. ๐Ÿ†• **Must-reads:** Press ยท Kapoor et al. ยท OpenAI (SWE-bench Verified) ยท Leaderboard Illusion