mirror of
https://github.com/openswarm-ai/openswarm.git
synced 2026-08-26 14:32:22 +02:00
arena: clean bu-real-opus5 cell lands -- 66.4% at 28.3s; temperature lever dead on Claude 5
The untainted re-run replaces the asterisked 48%: their best-model cell is 16 points behind ours at 4.5x the wall. Claude-5 lanes reject the temperature param outright -- v17's 400-loop root-caused, field omitted; decode variance now attacked at the action layer (fill-verify, verify-terminal) instead. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
This commit is contained in:
co-authored by
Claude Fable 5
parent
bc6f88d694
commit
9db052cf56
@@ -44,7 +44,7 @@ it — the constraint is flow competence, not steps.
|
||||
| haiku-4-5 | 71.2% @ 4.8s, 0 false (v10: 75.2% @ 5.2s) | 69.6% @ 44.5s, 16 false | ours leads all axes |
|
||||
| sonnet-4-6 | **77.6% @ 6.5s, 0 false** | 74.4% @ 37.4s, 8 false | ours leads all axes |
|
||||
| sonnet-5 | 76.0% @ 5.5s, 0 false | 63% running @ 42s | their loop DEGRADES on the newest model |
|
||||
| **opus-5** | **82.4% @ 6.3s, 0 false** (v14); v15 81.6; v16 82.1 @ 9.5s | 48.0%* tainted; clean re-run in flight | ours scales with the model |
|
||||
| **opus-5** | **82.4% @ 6.3s, 0 false** (v14); v15 81.6; v16 82.1 @ 9.5s | 66.4% @ 28.3s, 2 false (clean re-run; replaces tainted 48.0%*) | ours leads +16 pts at 4.5x speed |
|
||||
|
||||
v16 (verify-terminal, look-act-look, rapid-fire, sub-step confirm) FINAL: 82.1% on clean episodes
|
||||
-- statistically tied with v14, but the fixes hit their targets: email 10/10 (their best 6/10),
|
||||
|
||||
Reference in New Issue
Block a user