Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -268,6 +268,7 @@ Most "awesome" lists are link dumps. This one is **annotated and verified**: eve
- **[Search-Time Contamination in Deep Research Agents](https://arxiv.org/abs/2606.05241)** — Wang, Zhang, Yao, Zeng, Song, Lin, Shen — <https://arxiv.org/abs/2606.05241> · *paper* — Defines three contamination types (Benchmark Metadata Leakage, Question-Context Leakage, Explicit Answer Leakage) for agents that web-search during evaluation; applies detection algorithms across six benchmarks and finds performance inflation up to 4%. "Such agents may retrieve public benchmark metadata, question context, or even ground-truth answers via web search. This gives rise to Search-Time Contamination (STC), where external retrieval bypasses intended reasoning and inflates measured performance." Distinct from training contamination; advocates isolated sandboxes and transparent search trajectories. 🆕
- **[RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents](https://arxiv.org/abs/2603.11337)** — Yonas Atinafu, Robin Cohen — <https://arxiv.org/abs/2603.11337> · *paper* — Treats evaluation integrity as a first-class benchmarked outcome instead of an assumption: two compromise vectors (evaluator tampering, train/test leakage) instrumented via patch tracking and runtime file-access logging; in natural agent runs evaluator-tampering attempts appear in ~50% of episodes, and evaluator locking eliminates them at a 25–31% median runtime overhead. 🆕
- **[Prediction: A Frontier Open Source LLM Will Be Released On 3rd December 2026](https://blog.doubleword.ai/frontier-os-llm)** — Jamie Dborin (Doubleword) — <https://blog.doubleword.ai/frontier-os-llm> · *blog* — Extrapolates the Artificial Analysis Intelligence Index across 18 constituent benchmarks to forecast the open-vs-closed capability gap; a naive fit of the headline index points to convergence by 3 Dec 2026, but the average per-benchmark lag holds steady at ~5 months — a worked cautionary case in reading trends off an aggregate leaderboard index rather than its components. 🆕
- **[ClawBench: Can AI Agents Complete Everyday Online Tasks?](https://arxiv.org/abs/2604.08523)** — Zhang et al. (TIGER-AI-Lab) — <https://github.com/TIGER-AI-Lab/ClawBench> · *benchmark* — Evaluates browser agents on 281 everyday tasks (V1 152 + V2 129) across 163 live production websites and 15 categories. A submission-interception layer prevents real-world side effects while preserving realistic workflows, and the harness captures browser actions, HTTP traffic, screenshots, recordings, and agent messages for reproducible, fine-grained analysis. 🆕

**Must-reads:** Press · Kapoor et al. · OpenAI (SWE-bench Verified) · Leaderboard Illusion

Expand Down