arena: v19 confirms the ~85 plateau -- targeted losses flip, variance pays it back

Off-screen rows + group ordinals solved the social-media class outright (first time in
any single run) and nudged forms; the headline held at 84.8 because decode variance
returned equivalent tasks elsewhere. v17/v18/v19 = 85.2/85.5/84.8: the
prompt-and-perception ceiling is ~85 pass@1, 91.2 labeled pass@2. What remains is
engineering, not tuning -- the four mechanical primitives, which belong in the product.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
This commit is contained in:
ciregenz
2026-08-12 12:02:36 -07:00
co-authored by Claude Fable 5
parent 3a34795ba8
commit 84af515859
+9
View File
@@ -54,6 +54,15 @@ text_entry 15/17**, every one a lead over browser-use's best cell. Forms stayed
need the missing *primitive* (field-by-field controller with readback), not a better prompt.
Category ledger vs bu-real-opus5 (66.4%): LEAD 7, tie 1, BEHIND 1.
**v19 (off-screen rows + group ordinals) FINAL: 84.8%, 0 false.** The targeted flips landed --
social-media-all and social-media-some both solved for the first time in any single run (the
@ashlea class: the goal's target was below the fold and previously absent from the menu), forms
ticked 15->16 -- but variance gave back equivalent tasks elsewhere. **Three consecutive versions
now sit at 84.8-85.5: the prompt-and-perception plateau is ~85 single-run (91.2 labeled pass@2),
and the residual is decode variance plus the four engineering clusters.** Further headline gains
require the mechanical primitives (form-flow controller, game-state loop, console rung,
pixel-feedback geometry) -- product-level rungs, not agent tuning.
**v17 (isolation + mechanical fill-verify), two seeds, every episode clean, 0 false claims:**
| protocol | result |