11 KiB
Handoff: browser v3, state as of 2026-08-06
Branch eric/browser-merged, 45 commits past origin/eric/dev, pushed. Browser suite 707 passing, tsc 0,
linter no new violations, tree clean. The goal and the harness are described in README.md.
Six of nine criteria pass. Three are open, and none of them is open because the fix is unknown.
Scorecard
| # | criterion | target | before | now | |
|---|---|---|---|---|---|
| 1 | composer reach | >=90% | 57% | 93% (28/30) | PASS |
| 2 | verified writes | >=95% | no honest data | LinkedIn proven end to end; N too small | OPEN |
| 3 | false success | 0 | unknown | 0 in every live round | PASS |
| 4 | median write | <=12s | ~21s | 1.99s (p95 12.2s) | PASS |
| 5 | cold start tier-0/1 | <=3s | ~16s | 0.13s | PASS |
| 6 | other_ms | -50% | none | -70% (417 -> 126ms) | PASS |
| 7 | infra flake | <=1% | 60% | 34 runs, 0 failures; sample too small | OPEN |
| 8 | holdout | >=80%, <=10pt | untested | 87%, 6pt gap | PASS |
| 9 | learned path | remove or >=50% | 0/55 | recording 4/4 = 100%; replays 0 of 2 | OPEN |
Per-site reach at N=5 (criterion 1): x 5/5, linkedin 5/5, reddit 5/5, instagram 5/5 (was 1/10), youtube 4/5, twitch 4/5 (was 0/3). gmail, tiktok and substack are excluded, see Exclusions.
Timing at n=28 successful runs (criterion 4/6): wall median 1988ms / p95 12180ms; browser tools
median 1621ms; other_ms median 126ms, down from 417ms. Wall fell too, so nothing moved from one
bucket into another.
The finding that matters most
Several "product failures" were the measuring instrument. This is the single most useful thing to carry forward, because it recurred six times in one session and each instance cost hours chasing a bug that did not exist.
The worst case: criterion 2 failed on LinkedIn for days. LinkedIn was never broken. The canary
proved a write landed by grepping the backend log for the marker string, and the backend does not log
page text. A grep for every canary marker ever generated, across every backend log on the machine,
returns zero lines. The audit could only ever answer "could not look", and that was being read as a
product failure. probe_evidence.py settled it: the session API carries the model's answer and
never the tool results, so auditing it is asking the same model whose claim is under audit.
Every dead grep found, and what each did:
| where | needle | occurrences in source |
|---|---|---|
canary delivered |
DELIVERY CONFIRMED |
0 |
canary saw_page |
two [browser-action] X strings |
0 |
| canary receipt | only the fast-lane string; 2 of 3 producers missed | 1 of 3 |
skillstats replay_full |
replay(ed|ing) N steps |
0 |
bench infra_browser |
card is unavailable |
0 |
stack.sh status |
pgrep -fc (no such flag on macOS) |
printed 0 over a live stack |
verify_markers.py now checks all 32 harness literals against the source that prints them: 32/32
present. Run it before trusting any number.
What is open, why, and exactly what to do
Criterion 2, verified writes
The instrument was rebuilt (950755af) and the first result was linkedin PASS: posted,
receipt-verified, deleted, verified gone. That is the first clean end-to-end LinkedIn round this
project has recorded. One site is not >=95% across sites, so the criterion is not met.
Blocked on nothing technical. Two things to do first:
- Verify reddit's Title field name.
af68dcd0prefers the textbox whose accessible name contains the field word the task used, and requires exactly one match. The unit tests use an assumed listing ([21]<textbox "Title" />) that has never been confirmed against the live page; the backend only logstextboxes=3. If reddit's title input is named something else, the fix silently does nothing. reddit is 1 of the 3 sites in this denominator. - Then
ROUNDS=7 c2_rounds.shandc2_tally.py.
Criterion 7, infrastructure flake
Last clean sample: 34 of 108 runs, 0 infra failures, 0 backend restarts. Encouraging but partial. An earlier 5.4% reading was contaminated: a second OpenSwarm checkout was up on :8324 and the shared 9router on :20128 was being evicted.
Needs a box with nothing on :8324, then N=12 c7_run.sh for the full 108.
Criterion 9, learned fast path
The starting baseline, over 427 gate decisions on dry sweeps: 95 eligible, 0 recorded, 0
replays, with one refusal reason (host empty or no robust steps). Read on its own that number is
misleading, see "measure this on a LIVE run" below; but the refusal reason was a real bug.
The gate was never the problem. Every eligible run died inside record_skill: distill_steps reads
clicked_name, and browser_send_script.py wrote the element name into result_summary prose only.
An unnameable click truncates the distillation, truncation drops the typing steps, and the
navigation-only remainder is correctly refused. Fixed in 438a96eb, proven directly:
OLD shape (no clicked_name) -> []
NEW shape -> ['BrowserNavigate', 'BrowserClickByName', 'BrowserClickByName']
Safety held rather than added: a BrowserClickIndex distills to a name-only BrowserClickByName, so
the payload never enters the skill, and the send click's tool name matches no distill branch, so a
replay cannot re-fire a send.
Recording is PROVEN live, at 100%. On the live canary rounds (r4): 14 runs reached the gate,
4 were eligible, 4 recorded, 0 refused. Two skills persist on disk with exactly the shape the
unit test predicts, both timestamped after the fix:
x.com BrowserNavigate(https://x.com/compose/post)
BrowserClickByName(textbox "Post text")
www.linkedin.com BrowserNavigate(https://www.linkedin.com/feed/?shareActive=true)
BrowserClickByName(textbox "Text editor for creating content")
Measure this on a LIVE run, never a dry one. Recording is gated on delivery_verified, which a
dry run can never produce, so a dry sweep correctly records nothing. Reading a dry log as a verdict
is how this was first misreported as "0/95, the fix did not fire". skillstats.py now prints a note
when it sees that shape. (It also had two counting bugs of its own, fixed: it counted the
slots unfillable FAILURE line under a label that read "skill matched", and it never surfaced
quarantines.)
What still fails is REPLAY: 0 of 2.
- x.com truncates to a bare navigate.
is_replay_boundaryflags the composer click as irreversible because the accessible namePost textcontains "post". It is a textbox, and focusing a textbox is reversible, but the guard matches on the name only. The prefix therefore replays a single navigate, which prestage already does in 0.13s. - linkedin quarantines.
replay step failed (BrowserClickByName: No element matching role="textbox" name="Text editor for creating content"). LinkedIn's composer is lazily mounted and does not exist at navigate time; the recorded skill never learned the opener click that reveals it. Quarantine is correct behaviour here, not a bug.
So criterion 9 clears the >=50% recording bar and fails the "successful replays and measured
benefit" half. Before investing in either replay fix, answer the question that decides it: prestage
already reaches a tier-0/1 composer in 0.13s, so what is a replay actually worth? If the answer is
"nothing", removal is the honest path and the criterion explicitly allows it. The two
browser-memory endpoints have no frontend caller (grep of frontend/src and electron/ returns
nothing), so removal is not user-visible.
Product bugs found by the fixed instrument
- reddit: the task named a field and nobody read it (
af68dcd0). "create a text post whose title is exactly X" filled the body, Title stayed empty, and reddit's submit stayed DISABLED on every attempt. A field word from the user's own sentence now beats the compose-shape guess. 6 tests, including that no-hint behaviour is byte-identical. - disqus: a stalled top document denied its own child frames (
7ba99ba4). The still-loading retry rethrew on a second failure, killingfind_composerbefore the child-frame search ran. An ad-heavy page keeps the top document loading while the composer, in an embedded iframe with its own load state, is perfectly readable.
Known false negatives, filed not fixed
Both under-claim rather than over-claim, so neither violates criterion 3, but both read as drift on every live round.
- A torn-down browser card wedges, and
navigatelies about it. Reproduced 4x: the reply is{"text":"Navigated to https://x.com/home","url":"https://x.com/home"}in ~80ms while the webview never leaves reddit. Anything trusting that reply reads the wrong page and answers confidently about it. The audit now re-readslocation.hrefand requires a host match. - Delete reports failure on deletions that provably worked. X, 2 for 2: "the deletion was never confirmed", and an independent read shows the post gone both times. LinkedIn showed the same shape.
Exclusions (predefined, never quietly dropped)
gmail, substack and tiktok are signed out; tiktok is additionally captcha-walled. That is 14 of 45 rows on the known suite. Each exclusion is judged on the page's own evidence, never on the product's claim. Solving a bot-detection challenge is off-limits, so a captcha-walled page is one this system is choosing not to reach, and scoring it against reach would charge us for a rule we intend to keep.
Account hygiene
Every marker written to a real account during this work was cleaned up and independently verified gone: X profile scan shows 0 markers with the newest genuine post predating the tests, and the LinkedIn post permalink returns nav chrome with no content.
Personal handles are not committed. Both live sites read their account from the environment
(OSW_CANARY_X_HANDLE, OSW_CANARY_REDDIT_HANDLE), and an unset handle makes the round refuse
loudly rather than audit a malformed URL.
Environment notes
- The stack is backend :8326, webpack :3026, Electron on its own
--user-data-dir. It never touches :8324 or :3000, which belong to whatever else is on the machine. - The backend was SIGTERMed three times mid-measurement, always cleanly, always with nothing in its
log. Cause unknown.
stack.shnow supervises it and writes[stack] BACKEND RESTARTEDinto the log the harness slices, sobench.py'sinfra_backendbucket can exclude any trial spanning a restart instead of scoring it as a product failure. stack.sh downonce matchedpkill -f "uvicorn backend.main", which is exactly what the other checkout runs, and it executed while that checkout was live. It is now scoped to this stack's own ports and profile and cannot match a process it did not start.
Next executable steps, in order
verify_markers.py(32/32 expected).- Confirm reddit's Title field name against the live page.
- Decide criterion 9: is a replay worth anything against a 0.13s prestage? If not, remove.
N=12 c7_run.shon a box with nothing on :8324. Gives criteria 1, 4, 5, 6, 7.skillstats.pyon a LIVE round's log (not the dry sweep) for criterion 9.ROUNDS=7 c2_rounds.shthenc2_tally.py. Gives criterion 2.N=2 c8_run.shto re-check criterion 8 after the disqus fix.