From b43123ae208ccc8ffbd7059693b3ec23717f7932 Mon Sep 17 00:00:00 2001 From: ciregenz Date: Wed, 12 Aug 2026 14:06:36 -0700 Subject: [PATCH] arena: v20 champion 86.4 + public-baseline positioning -- ~12 points above published generic agents User-supplied 2026 survey gives the right comparison class: generic-harness MiniWoB tops out at 71.5 (GPT-5) / 74.9 (best harness). v20's 86.4 zero-tuning single-run sits far above it, and our +60-of-harness vs +7-of-model finding reproduces the survey's Orby insight at scale. MiniWoB demoted to regression-suite status per the survey rubric; Fable 5 sweep launched to complete the first known Claude-5 triplet and answer model-limited-vs-harness-limited directly. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ --- e2e/browser-v3/arena/ARENA.md | 16 ++++++++++++++++ 1 file changed, 16 insertions(+) diff --git a/e2e/browser-v3/arena/ARENA.md b/e2e/browser-v3/arena/ARENA.md index d44a85b8..2c8cefa6 100644 --- a/e2e/browser-v3/arena/ARENA.md +++ b/e2e/browser-v3/arena/ARENA.md @@ -106,6 +106,22 @@ the remaining gap: ~22 tasks need purpose-built widget primitives (date/time pic canvas geometry, long autocomplete flows), not a stronger model. Their false-claim rate persists across every model (16 haiku, 8 sonnet-4-6) -- structural to the JS-evaluate hatch, as predicted. +## Positioning vs public generic-harness baselines (user-supplied 2026 survey) + +The comparable class is generic agents, NOT MiniWoB-specialized systems (HTML-T5++ 95.2 trained +on it; CompWoB showed such scores collapse to ~61 on compositional variants). Published +generic-harness MiniWoB: GPT-5 71.5, GPT-4o 71.3, Claude Sonnet 4 70.7, Claude 3.5 69.8 +(ServiceNow GenericAgent); best published harness lift = Orby +5.1 on the same model (74.9). +**Ours: v20 86.4 single-run / 91.2 labeled pass@2, zero MiniWoB-specific logic (audited), zero +false claims -- ~12 points above any published generic agent.** Our data independently confirms +the survey's core finding at larger scale: harness >> model (our +60 points of harness gains vs ++7 from model tier upgrades on a fixed harness). Claude 5-family public numbers do not exist; +our sonnet-5/opus-5/fable-5 grid is, as far as known, the first. Per the survey's rubric +(80-90 "quite strong", 90-95 "extremely robust"), MiniWoB is hereby DEMOTED in this repo to a +regression/unit suite; realistic-workflow weight moves to WebArena-class benchmarks when infra +exists. v20 final loss modes: 10 hard-cluster (needs product primitives), 7 truncation-deaths +(strict-retry halved but did not eliminate them). + ## The full ladder — every version, every technique, its measured worth | ver | change (source) | rate | med win |