From 25892f30b113996a6fa856949b8b654de4d59f3a Mon Sep 17 00:00:00 2001 From: ciregenz Date: Sun, 16 Aug 2026 00:58:44 -0700 Subject: [PATCH] arena: WebChoreArena primary Claude column final -- 1/89 at 500s cap, 39 answers, 0 false; budget is the binding constraint (both models agree) Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ --- e2e/browser-v3/arena/ARENA.md | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/e2e/browser-v3/arena/ARENA.md b/e2e/browser-v3/arena/ARENA.md index e95cc6a0..8cdce427 100644 --- a/e2e/browser-v3/arena/ARENA.md +++ b/e2e/browser-v3/arena/ARENA.md @@ -558,6 +558,18 @@ TIER 3 (gated/expensive — pilot only if Tier 1-2 leave a gap): NOT ingested (out of scope / already have): per-site tips files (task-specific), combobox flattening (we have native pickers), fine-tuned transfer weights (specialist). +## WebChoreArena PRIMARY Claude column FINAL (2026-08-16) + +**1/89 scoreable solved | partial-sum 1.00 | 39 answers delivered | 0 false claims | median +408s/task; 2 of 91 excluded (their-evaluator crashes).** Paired with GPT-5.4's 2/86 (154s med, +21 answers): Claude works tasks ~2.6x longer and answers ~2x more but converts no more of them +under the same 500s cap — both columns say the binding constraint is BUDGET (chores are designed +for uncapped, 100k+-token runs; the paper's 20-40% agents ran that way), plus long-horizon +aggregation (the v37 note scratchpad targets exactly this — delta rerun planned after its +pilot). Honesty clause: zero false claims across BOTH model columns on the hardest benchmark in +the set. The one solved chore (30093, both models' solved sets overlap on ratio-computation +tasks) confirms the evaluator plumbing end-to-end. + ## PRE-REGISTERED (2026-08-15, before any v36/v37 episode): first two ingestion pilots v37 NOTE SCRATCHPAD pilot — targets (8): form-sequence-3, hot-cold, search-engine, text-editor