arena: v22 champion at 89.6 -- the structural fix lands one task shy of 90

Action-first ordering (+3.2 over v20) proved the design-out-the-class principle:
truncation deaths ended, book-flight fell for the first time in twenty sweeps, forms
19/22, drag and email perfect, zero false claims at 8.5s median. Seed-43 running as
the cross-seed decider; the residue is the named product-primitive cluster plus an
empty-reply subclass whose structural close is reply prefill.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
This commit is contained in:
ciregenz
2026-08-12 16:56:04 -07:00
co-authored by Claude Fable 5
parent 8cc2c99c32
commit 832c4a0362
+9
View File
@@ -106,6 +106,15 @@ the remaining gap: ~22 tasks need purpose-built widget primitives (date/time pic
canvas geometry, long autocomplete flows), not a stronger model. Their false-claim rate persists
across every model (16 haiku, 8 sonnet-4-6) -- structural to the JS-evaluate hatch, as predicted.
## Champion: v22 -- 89.6% single-run, the structural-fix payoff
Action-first reply order (truncation structurally impossible) + every prior rung: **112/125 =
89.6% @ 8.5s median, 0 false claims** -- +3.2 over v20, one task from the 90 line. book-flight
fell for the FIRST time in 20+ sweeps (forms 19/22); drag 13/13 and email 10/10 remain perfect.
The 13 residual losses are the irreducible product-primitive cluster (games, console, precision
geometry, two long forms) plus a residual empty-reply subclass (model prose with no action at
all -- prefill-forcing is the structural close). Seed-43 confirmation in flight.
## Positioning vs public generic-harness baselines (user-supplied 2026 survey)
The comparable class is generic agents, NOT MiniWoB-specialized systems (HTML-T5++ 95.2 trained