Numbers from this box, with the unverifiable ones marked as such.
The result that changes priorities: a deterministic accessibility-tree snapshot
(agent-browser, Rust CLI, no LLM) perceives in ~50ms median against our
BrowserFindComposer's 13,991ms median / 30,003ms p95 over 119 calls, and fills+verifies
onlinegdb in 326ms -- a site our agent cannot reach at all. Also fills w3schools (810ms)
and regex101 (31ms). That is a 40-450x perception gap with no model in the loop.
Also recorded: all nine claimed repos verified to exist with matching stars and
AGPL-compatible licences; browser-use benchmarked as an agent over 43 trials (we are
faster on every site we reach, 9.8s vs 35.7s median, but our arm ran with its fallback
disabled by dry-run so reach is not comparable); direct CDP evidence that deepl serves
our browser a Cloudflare challenge while browser-use's fresh-per-run profile is never
flagged; OpenCLI's API-first-with-browser-fallback design as the route to a learned path
that actually replays; and why Polar was not benchmarked (their ToS forbids using the
product to build a competing one, and their 98.0 concatenates two benchmarks and appears
on neither leaderboard).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The wait was a bare `timeout=300.0` at the call site: invisible, unchangeable, and
longer than WALL_BUDGET_S (180s) itself, so a single prompt could outlive the whole
run's budget.
Measured 2026-08-08. deepl bot-detected the browser profile, the agent correctly
refused to solve the challenge ("handing to the user, not solving it") and asked for
help via RequestHumanIntervention. Headless, nobody answered, so it burned the full
306s and then denied -- the identical verdict it can reach instantly. That single
wait consumed the entire 420s task budget and was the whole of what looked like a
"249s spawn stall" while profiling. Cron runs, scheduled agents, CI and benchmarks
all sit in exactly this position.
Two changes, neither of which weakens the gate:
- ws_manager.has_listener(session_id) reports whether ANY socket would receive the
session's events, reading the same two lists send_to_session broadcasts to so it
cannot drift from where messages actually go.
- p_request_browser_approval checks it BEFORE building a request, and declines with
an honest reason when no UI is attached. The decision is unchanged (deny); only
the five minutes of waiting for it are gone.
The timeout is now P_APPROVAL_TIMEOUT_S, overridable via OSW_APPROVAL_TIMEOUT_S for
automation contexts that want a different budget.
A human at the keyboard sees no change: with a socket attached the request is sent
and awaited exactly as before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A trial whose card was reaped before the readback could not be verified either way,
and excluding those rows was too charitable: it flipped openswarm 60.9% -> 100% and
McNemar p=0.02 -> 1.0 purely on my read timing. The backend's own dryrun-report
(composer/textboxes/filled) is the same bar the card readback applies, so it stands
in when the card is gone.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Accumulating and resumable: a 95% CI narrow enough to separate two systems needs tens
of trials per site per arm, which is hours, so rows append to results/vs_raw.jsonl and
the report is computed from everything on disk. `vs_bench.py report` re-reads without
running anything.
Paired by construction: both arms run the same site back-to-back in one iteration, so
a site that changes mid-day (deepl started redirecting to a locale path during this
work) moves BOTH arms rather than one. Discordant pairs are tested with McNemar exact,
which is the right test for paired binary outcomes at small discordant counts, and
proportions use Wilson rather than the normal approximation -- at 20/20 the naive
interval is [100%, 100%], a certainty no 20-trial sample carries.
Three confounds found while building it, each of which had produced a wrong number:
1. The stop directive cannot be shared. "Do NOT submit" trips OpenSwarm's is_readonly()
and the send-script declines outright ("read-only directive in user request"), so
our arm measured 0/2 while browser-use scored 2/2 on the identical prompt. That is
the prompt disabling one side, not a capability gap. OpenSwarm now gets the bare task
and is held to reach-only by the backend's own OSW_SENDSCRIPT_DRYRUN=1; browser-use
has no dry-run mode, so its constraint stays in the prompt. Verified at the
classifier: is_readonly(BU)=True, is_readonly(OSW)=False.
2. Verification must be scoped to the cards a trial CREATED. A dashboard accumulates one
card per run and they persist in the saved layout; reading "all cards" checked seven
stale pages, and a regex101 run graded 0/1 because its own card had been reaped while
leftovers answered instead.
3. browser-use's readback goes through ITS session (get_tabs + get_or_create_cdp_session)
rather than a second websocket. Every earlier verifier bug -- hardcoded port, ws:// vs
http://, page-vs-webview target filtering -- lived in hand-rolled target discovery.
The browser is deliberately NOT held constant. browser-use cannot screenshot an Electron
webview and falls back to the app's page target on 100% of steps (49/49 measured), so a
shared browser blinds it; each side runs in the browser it was built for and that is
recorded as a caveat rather than hidden.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Added to compare browser-use against us with the BROWSER held constant, and the
result is that it cannot be held constant, which is worth writing down.
What works: with the port exposed, browser-use attaches, sees the browser cards
(they are CDP targets of type 'webview'), navigates one, and completes a task that
verifies independently. Cards persist across restarts because they live in the saved
dashboard layout, so there is something to attach to.
What does not: browser-use's screenshot path cannot capture a webview and falls back
to the app's own 'page' target -- 49 fallbacks in 49 agent steps, i.e. it drove the
web page while LOOKING at the OpenSwarm UI for every single step. A comparison run
that way handicaps the rival for an infrastructural reason and measures neither
agent, so the shared-browser numbers are not publishable as a fair head-to-head.
Two hazards found the hard way, both worth knowing before anyone repeats this:
- browser-use filters CDP targets to page/tab, so 'webview' is invisible to it, and
Electron refuses Target.createTarget ("Not supported") so it cannot open its own.
Pointed at Electron unrestricted it therefore grabs the app's page target and
navigates the OpenSwarm UI away, killing the renderer: the next sweep failed
infra_no_renderer 1/1 because of exactly that.
- An external agent must never be allowed to kill the shared session. Doing so after
trial 1 made trials 2-15 fail ConnectError, which a naive summary published as
"browser-use REACH 4/15" -- ten rows of my own teardown scored against the agent.
Off by default: an open debugging port is a local attack surface, and this only
exists for measurement.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>