The wait was a bare `timeout=300.0` at the call site: invisible, unchangeable, and
longer than WALL_BUDGET_S (180s) itself, so a single prompt could outlive the whole
run's budget.
Measured 2026-08-08. deepl bot-detected the browser profile, the agent correctly
refused to solve the challenge ("handing to the user, not solving it") and asked for
help via RequestHumanIntervention. Headless, nobody answered, so it burned the full
306s and then denied -- the identical verdict it can reach instantly. That single
wait consumed the entire 420s task budget and was the whole of what looked like a
"249s spawn stall" while profiling. Cron runs, scheduled agents, CI and benchmarks
all sit in exactly this position.
Two changes, neither of which weakens the gate:
- ws_manager.has_listener(session_id) reports whether ANY socket would receive the
session's events, reading the same two lists send_to_session broadcasts to so it
cannot drift from where messages actually go.
- p_request_browser_approval checks it BEFORE building a request, and declines with
an honest reason when no UI is attached. The decision is unchanged (deny); only
the five minutes of waiting for it are gone.
The timeout is now P_APPROVAL_TIMEOUT_S, overridable via OSW_APPROVAL_TIMEOUT_S for
automation contexts that want a different budget.
A human at the keyboard sees no change: with a socket attached the request is sent
and awaited exactly as before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Six of the eleven defects here were in the MEASUREMENT, not the product, and they
were wrong in both directions.
Harness, all of which silently produced wrong numbers:
- coverage.py preflight refused every sweep on a box holding exactly one backend:
stack.sh's supervisor is a `bash -c` quoting the whole uvicorn line, so it carries
both "-m uvicorn backend.main" AND the venv python path. Discriminate on POSITION.
- stack.sh status reported 2 backends over 1 and 0 webpack over a live dev server
(webpack retitles its process). A status check whose job is preventing a second
stack, failing in the direction that lets one land.
- c7_run.sh/c8_run.sh slice r6_be.log while stack.sh names logs by TAG: a stack under
any other tag hands every trial an empty slice and the sweep reports a confident
0/108. Now refuses loudly; it caught this exact mistake on first use.
- "Browser command timed out" was bucketed infra. It is ONE command blowing its own
budget, not a dead webview: all 4 such rows were BrowserFindComposer at exactly its
30s cap, every run completed after, zero card-gone markers in the whole log. Filed
as infra it read as 11.8% flake AND lifted holdout reach 70% -> 84%.
- api_retry / rate_limit_error now grade as infra. A provider 429 storm turned clean
15-21s exclusions into 188s product_no_composer rows.
- bench.py prints reach BOTH ways when a row is UNVERIFIED. An exclusion resting on
the agent's own word quietly flatters the score, and coverage.py's own instruction
to confirm it by hand goes unread (I quoted a 100% that excluded onlinegdb).
Timing was measuring 0.2% of the run: prestage completes BEFORE metrics_started_at,
so other_ms was 25ms of a 12700ms median while prestage (4146ms, ~61%) sat in no
bucket at all. prestage_ms/task_ms are now recorded; total_ms is deliberately NOT
redefined, which would invalidate every before/after already taken against it.
Product:
- find_composer rejected ACE/CodeMirror-5/Monaco composers. Their input is a ~1x1
offscreen textarea that paints into a sibling div, so it can never pass a size
gate. Accept it when a VISIBLE ancestor is composer-sized; honeypots stay out
because the input itself must not be display:none/visibility:hidden/opacity:0.
anon reach 80% -> 100%, holdout 89% -> 90%, p95 38.6s -> 9.7s.
- the composer poll slept a blind 0+1.2+1.4 = 2.6s whenever prestage staged nothing,
which is nearly every run, and it was the whole of other_ms's suspicious constancy
(2610-2613ms regardless of tools_ms). Stop when two reads are identical, the rule
the opener poll 40 lines below already applies. other_ms -53.7%, tools_ms flat.
- prestage no longer navigates to the page it is already on, nor sleeps 0.35s before
its first settle probe.
- is_replay_boundary reasoned from the NAME alone, so x.com's composer textbox named
"Post text" was ruled an irreversible send and truncated its replay to a bare
navigate. Excluded by ROLE; first_unsafe_step now passes role through at all.
Measured on this box: reach 100% (83% if onlinegdb's unverified exclusion is bogus),
0 false successes in ~155 runs, prestage tier-0/1 2702ms, other_ms -53.7%, infra
flake 0/158, holdout 18/20. Criteria 2/4/9 need live writes and are untouched.
Full evidence, including what did NOT work, in e2e/browser-v3/RESULTS_2026-08-06.md.
Also drops the tracked electron/node_modules symlink pointing at another machine's
Downloads folder; it is dangling on every other checkout and re-breaks the install on
any stash or checkout.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>