mirror of
https://github.com/openswarm-ai/openswarm.git
synced 2026-09-08 02:37:45 +02:00
arena: v22 champion at 89.6 -- the structural fix lands one task shy of 90
Action-first ordering (+3.2 over v20) proved the design-out-the-class principle: truncation deaths ended, book-flight fell for the first time in twenty sweeps, forms 19/22, drag and email perfect, zero false claims at 8.5s median. Seed-43 running as the cross-seed decider; the residue is the named product-primitive cluster plus an empty-reply subclass whose structural close is reply prefill. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
This commit is contained in:
co-authored by
Claude Fable 5
parent
8cc2c99c32
commit
832c4a0362
@@ -106,6 +106,15 @@ the remaining gap: ~22 tasks need purpose-built widget primitives (date/time pic
|
||||
canvas geometry, long autocomplete flows), not a stronger model. Their false-claim rate persists
|
||||
across every model (16 haiku, 8 sonnet-4-6) -- structural to the JS-evaluate hatch, as predicted.
|
||||
|
||||
## Champion: v22 -- 89.6% single-run, the structural-fix payoff
|
||||
|
||||
Action-first reply order (truncation structurally impossible) + every prior rung: **112/125 =
|
||||
89.6% @ 8.5s median, 0 false claims** -- +3.2 over v20, one task from the 90 line. book-flight
|
||||
fell for the FIRST time in 20+ sweeps (forms 19/22); drag 13/13 and email 10/10 remain perfect.
|
||||
The 13 residual losses are the irreducible product-primitive cluster (games, console, precision
|
||||
geometry, two long forms) plus a residual empty-reply subclass (model prose with no action at
|
||||
all -- prefill-forcing is the structural close). Seed-43 confirmation in flight.
|
||||
|
||||
## Positioning vs public generic-harness baselines (user-supplied 2026 survey)
|
||||
|
||||
The comparable class is generic agents, NOT MiniWoB-specialized systems (HTML-T5++ 95.2 trained
|
||||
|
||||
Reference in New Issue
Block a user