Commit Graph
2722 Commits
Author SHA1 Message Date
ciregenz fa67cd8458 arena: PRODUCT-WRITE running -- stack up manually, product agent DROVE live browser + completed quill write; vs_bench gtranslate reach=Y 7.9s VERIFIED. Corrects the 7% (wrong body); product does verified live writes ~8s as user said 2026-08-17 20:07:14 -07:00
ciregenz 7dc730e15e arena: PROVEN product agent is headless-incapable (ran it: 'no OpenSwarm window connected') -- browser body is Electron window; one-step handoff = user launches app, then vs_bench.py gives real product-write number 2026-08-17 19:45:22 -07:00
ciregenz aa78bff506 arena: product-write attempt -- found the real harness (vs_bench.py, cooperative editors, verified readback) but stack won't boot headless (backend deps fail, needs Electron); handoff = user runs app, then vs_bench.py gives the honest product-write number 2026-08-17 19:33:42 -07:00
ciregenz d488c8fc18 arena: traced product-vs-harness -- product agent is app-coupled (ws_manager->Electron webview), cant run headless, hence live-only; arena harness self-contained for benchmarking but strips the write machinery; arena write number understates product 2026-08-17 18:21:22 -07:00
ciregenz 65006c6cd5 arena: v48 plan-state NOT-A-LIFT at scale (completion wall confirmed non-scaffolding-movable, all 4 mechanisms) + book harness-vs-product distinction (arena numbers = benchmark harness, product write capability is higher) 2026-08-17 17:53:51 -07:00
ciregenz 9e9111102d arena: CORRECTION -- 7% verified-write measured the generic benchmark loop, NOT the product's verified-action system (proven on live compose-writes); re-measure with send-script pattern, reproducible on self-hosted sites 2026-08-17 17:20:38 -07:00
ciregenz 8f4a25ff39 arena: v48 plan-state fair pilot MARGINAL (5 vs 4 strict, -partial, within noise); bigger 65-task resolving test launched; corrected too-loose auto-verdict 2026-08-17 17:01:37 -07:00
ciregenz 2529aa3772 arena: capability-mechanism ledger + caught a flawed pilot (first-20 ceilings at 0 for ALL configs incl v42) -- rebuilt fair near-miss pilot before concluding; honest decisive test running 2026-08-17 12:33:19 -07:00
ciregenz c80cbeacd8 arena: v48 plan-state + reflective compaction (Hermes/Devin convergent design) -- never-evicted task-state block, evidence-gated subgoals; the decisive completion-capability test 2026-08-17 11:35:58 -07:00
ciregenz 70d706d052 arena: v47 within-episode note scratchpad (AgentOccam memory, study #1 pick) -- broadened gate to multi-step goals, AgentOccam spec; within-episode = no independence caveat 2026-08-17 10:57:13 -07:00
ciregenz 4b38269e51 arena: v46 checkpoint self-verification done-gate (FCPAgent-style) -- capability work on long-horizon completion, pre-pilot 2026-08-17 10:53:14 -07:00
ciregenz 268f6948a7 arena: v45 read-your-writes NEGATIVE (9% vs 7%, within noise) -- verified-writes capability-gated not prompt-gated; clause 5 hard-open 2026-08-17 09:15:17 -07:00
ciregenz a22f47a116 arena: v45 read-your-writes (state-change tasks re-read + retry until change confirmed) -- capability work for verified-writes clause 2026-08-17 08:11:52 -07:00
ciregenz 1f3c2d4fbe arena: verified-write nuance -- type-specific (gitlab comments persist, reddit upvotes dont); 7% was upvote-dominated, understates comment writes 2026-08-17 07:15:23 -07:00
ciregenzandClaude Fable 5 118fde5623 arena: atomic-write measurement -- my prediction was WRONG (7%, not high); verified-writes ~7-8% overall is a real capability wall, clause 5 furthest-open; verifier bugs fixed before booking
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 22:26:36 -07:00
ciregenz 834aa60f74 arena: Online-Mind2Web head-to-head scoped -- runnable but live-site/ToS/LLM-judge blockers; recommend 50-task read-only subset, awaiting go/no-go 2026-08-16 21:57:34 -07:00
ciregenz a86d9e3a0b arena: verified-write independent scorer -- composite reward OVERSTATED (11/15 looked persisted, only ~8% actually did); independent re-read caught the false-positive; clause 5 far open, capability-gated on complex writes 2026-08-16 21:19:34 -07:00
ciregenz ef2c8a8d3b arena: de-risk verdicts -- OSWorld smoke-only/skip (model-gated, cant run on our stack); verified-writes BUILD NOW (deterministic on existing infra); MiniWoB 94.1 honest near-ceiling (search-engine primitive fought reward logic, disabled) 2026-08-16 19:23:54 -07:00
ciregenz bc225d11ae arena: CompWoB spot-check -- no paren-bug suppression (3/12 same as v34); 81.1 stands 2026-08-16 19:18:33 -07:00
ciregenzandClaude Fable 5 c44308ed94 arena: WebArena FINAL = 13 strict/15% + 29.2 partial/34% (was falsely 0/3.83) -- ~7-8x true-credit correction from instrument fixes
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 18:47:42 -07:00
ciregenz 728ccab6b6 arena: judge policy -> RECIPROCAL cross-family (GPT judges Claude, Claude judges GPT) -- symmetric, no self-grading, bias applies equally 2026-08-16 18:46:29 -07:00
ciregenz 7b33018c9d arena: judge policy -- cross-family (GPT-5.6) + official judge, never same-family self-judge; supplementary only 2026-08-16 18:27:16 -07:00
ciregenz 52c19b0aee arena: test roadmap -- Polar benchmarks unrunnable (results-only/LLM-judge); queue their PUBLIC constituents (Online-Mind2Web etc.) as labeled supplementary numbers 2026-08-16 18:24:54 -07:00
ciregenz 9f4acf7823 arena: competitor note -- Polar's 98.0 is self-graded on own benchmarks (tuned product, not generic); not comparable, don't chase 2026-08-16 18:22:12 -07:00
ciregenzandClaude Fable 5 0e5162bf2e arena: MAJOR WebArena RETRACTION -- '0 solved' was our own instrument bugs; honest re-measure ~18% strict/~36% partial (was falsely 0/3.83). Biggest correction of the effort.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 18:15:58 -07:00
ciregenzandClaude Fable 5 1baecba061 arena: v43 schema-gate was self-inflicted (falsely bounced valid JSON) -- disabled; v42 is honest WA config; full WA re-measurement launching
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 16:10:29 -07:00
ciregenzandClaude Fable 5 195c1e0935 arena: WA canary retraction (instrument fixes lifted partials 0->0.5) + v43 answer-schema conformance gate (generic instruction-following) to lift 0.5->1.0
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 15:07:22 -07:00
ciregenzandClaude Fable 5 760d40b953 arena: v41 MiniWoB champion = 94.1 (+3.4 vs v40; draw-circle + 18-budget); ~1.9 from 95, 4-5 primitives left
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 14:33:47 -07:00
ciregenzandClaude Fable 5 e96a5fdd24 arena: grep-audit checkpoint -- champion logic has zero benchmark-specific strings; overfitting boundary holds
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 13:01:24 -07:00
ciregenzandClaude Fable 5 70dd5761a4 arena: fix TWO instrument bugs zeroing WebArena -- (1) goal truncated to 600 chars dropped WA's answer-schema; feed policy full goal. (2) CALL_RE cut payloads at first ) killing answers with parens; quote-aware scanner. Both trace-verified. WA 0-solved was partly our harness.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 12:59:52 -07:00
ciregenzandClaude Fable 5 3972f543fc arena: v42 terminal answer protocol (bare final answer on string-match goals) -- generic output hygiene, offline-verified
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 12:57:38 -07:00
ciregenzandClaude Fable 5 b4c011f8a3 arena: benchmark-structure + SOTA scan -- terminal answer protocol (TAKE, generic) vs api/selector bypass (DECLINE, benchmark-specific); honest principled split
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 12:56:54 -07:00
ciregenz c08d0d4b22 arena: WebArena v40 rerun still 0 strict -- ingestions dont address long-horizon completion; goal >=65 open+field-hard, booked honestly 2026-08-16 11:57:16 -07:00
ciregenzandClaude Fable 5 4022b04b9a arena: v41 scripted draw-circle geometry primitive -- draw-circle 0/3->3/3; feature-gated on circle goals
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 11:24:50 -07:00
ciregenzandClaude Fable 5 1260988938 arena: consolidated goal scoreboard at checkpoint -- 6/11 clauses met, open ones honestly categorized (capability/field-difficulty/infra, none overfittable)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 09:36:45 -07:00
ciregenzandClaude Fable 5 506961da9f arena: WASP injection refusal = 100% (17/17 engaged, 0 obeyed) -- clause MET; scorer false-positive caught+fixed pre-booking
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 08:18:22 -07:00
ciregenzandClaude Fable 5 191854df37 arena: WASP scorer fix -- exclude send_msg_to_user from obeyed check (agent narrating its refusal mentioned attacker domain -> 3 false-positive vulns were actually strong refusals); only real navigation/interaction counts
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 08:01:15 -07:00
ciregenzandClaude Fable 5 43004c705a arena: WASP scorer -- policy.reset per episode + 0-action marked inconclusive (empty completion = ambiguous, scored separately)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 07:28:23 -07:00
ciregenzandClaude Fable 5 c4ae243ba6 arena: WASP injection refusal harness (deterministic scoring, no LLM judge) -- canary clean; full 21-attack sweep next
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 06:54:33 -07:00
ciregenzandClaude Fable 5 d887850706 arena: CompWoB v40 delta 0 net-new -- holds 81.1, still 0.9 behind browser-use; booked honestly, no manufactured pass
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 06:18:58 -07:00
ciregenzandClaude Fable 5 c50bd4fd82 arena: champion v40 3-seed = 90.7 (ingestion wins generalized; seed-42 flaky tail offsets; prompt-side plateau noted)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 05:46:51 -07:00
ciregenzandClaude Fable 5 eaacc854d8 arena: v40 champion promoted (escape+table_md; popup loss was variance, 6/6 on 3-seed) -- champion 3-seed MiniWoB launching
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 03:42:40 -07:00
ciregenzandClaude Fable 5 92f1db8b6e arena: v40c exonerates escape token (same control losses without it) -- v40a (escape+table, no mutation) isolating now
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 03:05:47 -07:00
ciregenzandClaude Fable 5 76eada6987 arena: v40 union FAILS confirm (interaction regression) -- pairwise isolation running (v40c first)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 02:33:55 -07:00
ciregenzandClaude Fable 5 7581706f5f arena: v38 7/8 + v39 5/8 both PASS -- v40 champion = union of three passed ingestions; combined confirm pilot next
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 02:02:03 -07:00
ciregenzandClaude Fable 5 0e2aac1afc arena: pre-register v38/v39 pilots
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 01:30:31 -07:00
ciregenzandClaude Fable 5 2e07047ff7 arena: v36 PASSES (3/8 incl click-collapsible pair; first ingestion win) -> v40 champion; v37 deferred to chore-scale
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 01:30:09 -07:00
ciregenzandClaude Fable 5 25892f30b1 arena: WebChoreArena primary Claude column final -- 1/89 at 500s cap, 39 answers, 0 false; budget is the binding constraint (both models agree)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-16 00:58:44 -07:00
ciregenzandClaude Fable 5 8574ab4d2a arena: pre-register v36/v37 ingestion pilots (targets + predictions) before any episode
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-15 23:01:32 -07:00
ciregenzandClaude Fable 5 3abe31fb01 arena: implement v39 mutation-diff action feedback (Agent-E) -- all four Tier-1 ingestions now flag-gated and ready to pilot
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
2026-08-15 22:53:46 -07:00