arena: v20 champion 86.4 + public-baseline positioning -- ~12 points above published generic agents

User-supplied 2026 survey gives the right comparison class: generic-harness MiniWoB
tops out at 71.5 (GPT-5) / 74.9 (best harness). v20's 86.4 zero-tuning single-run sits
far above it, and our +60-of-harness vs +7-of-model finding reproduces the survey's
Orby insight at scale. MiniWoB demoted to regression-suite status per the survey
rubric; Fable 5 sweep launched to complete the first known Claude-5 triplet and answer
model-limited-vs-harness-limited directly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
This commit is contained in:
ciregenz
2026-08-12 14:06:36 -07:00
co-authored by Claude Fable 5
parent 94fd5d5185
commit b43123ae20
+16
View File
@@ -106,6 +106,22 @@ the remaining gap: ~22 tasks need purpose-built widget primitives (date/time pic
canvas geometry, long autocomplete flows), not a stronger model. Their false-claim rate persists
across every model (16 haiku, 8 sonnet-4-6) -- structural to the JS-evaluate hatch, as predicted.
## Positioning vs public generic-harness baselines (user-supplied 2026 survey)
The comparable class is generic agents, NOT MiniWoB-specialized systems (HTML-T5++ 95.2 trained
on it; CompWoB showed such scores collapse to ~61 on compositional variants). Published
generic-harness MiniWoB: GPT-5 71.5, GPT-4o 71.3, Claude Sonnet 4 70.7, Claude 3.5 69.8
(ServiceNow GenericAgent); best published harness lift = Orby +5.1 on the same model (74.9).
**Ours: v20 86.4 single-run / 91.2 labeled pass@2, zero MiniWoB-specific logic (audited), zero
false claims -- ~12 points above any published generic agent.** Our data independently confirms
the survey's core finding at larger scale: harness >> model (our +60 points of harness gains vs
+7 from model tier upgrades on a fixed harness). Claude 5-family public numbers do not exist;
our sonnet-5/opus-5/fable-5 grid is, as far as known, the first. Per the survey's rubric
(80-90 "quite strong", 90-95 "extremely robust"), MiniWoB is hereby DEMOTED in this repo to a
regression/unit suite; realistic-workflow weight moves to WebArena-class benchmarks when infra
exists. v20 final loss modes: 10 hard-cluster (needs product primitives), 7 truncation-deaths
(strict-retry halved but did not eliminate them).
## The full ladder — every version, every technique, its measured worth
| ver | change (source) | rate | med win |