Files
openswarm/e2e
ciregenzandClaude Opus 5 9f7035da6b [eric] harness: paired OpenSwarm-vs-browser-use benchmark with Wilson CIs and McNemar
Accumulating and resumable: a 95% CI narrow enough to separate two systems needs tens
of trials per site per arm, which is hours, so rows append to results/vs_raw.jsonl and
the report is computed from everything on disk. `vs_bench.py report` re-reads without
running anything.

Paired by construction: both arms run the same site back-to-back in one iteration, so
a site that changes mid-day (deepl started redirecting to a locale path during this
work) moves BOTH arms rather than one. Discordant pairs are tested with McNemar exact,
which is the right test for paired binary outcomes at small discordant counts, and
proportions use Wilson rather than the normal approximation -- at 20/20 the naive
interval is [100%, 100%], a certainty no 20-trial sample carries.

Three confounds found while building it, each of which had produced a wrong number:

1. The stop directive cannot be shared. "Do NOT submit" trips OpenSwarm's is_readonly()
   and the send-script declines outright ("read-only directive in user request"), so
   our arm measured 0/2 while browser-use scored 2/2 on the identical prompt. That is
   the prompt disabling one side, not a capability gap. OpenSwarm now gets the bare task
   and is held to reach-only by the backend's own OSW_SENDSCRIPT_DRYRUN=1; browser-use
   has no dry-run mode, so its constraint stays in the prompt. Verified at the
   classifier: is_readonly(BU)=True, is_readonly(OSW)=False.
2. Verification must be scoped to the cards a trial CREATED. A dashboard accumulates one
   card per run and they persist in the saved layout; reading "all cards" checked seven
   stale pages, and a regex101 run graded 0/1 because its own card had been reaped while
   leftovers answered instead.
3. browser-use's readback goes through ITS session (get_tabs + get_or_create_cdp_session)
   rather than a second websocket. Every earlier verifier bug -- hardcoded port, ws:// vs
   http://, page-vs-webview target filtering -- lived in hand-rolled target discovery.

The browser is deliberately NOT held constant. browser-use cannot screenshot an Electron
webview and falls back to the app's page target on 100% of steps (49/49 measured), so a
shared browser blinds it; each side runs in the browser it was built for and that is
recorded as a caveat rather than hidden.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 00:21:28 -07:00
..

End-to-end tests (packaged app, macOS + Windows)

Playwright tests that launch the packaged OpenSwarm desktop app (the real built binary, asar + bundled python-env + real paths) and drive it the way a user would. The same specs run unchanged on macOS and Windows; CI builds the artifact per-OS, then runs these. No provider API key is needed (no agent turn), so the suite is hermetic and deterministic on a clean machine.

What it checks (per OS)

  • Main window paints the React shell (first meaningful paint).
  • The preload bridge (window.openswarm) is exposed.
  • The real backend the app spawned reaches HTTP-ready (/api/health/check -> 200).
  • Provenance: the running app's getBuildInfo() sha matches electron/build-info.json.
  • App version is reported.

Run locally

  1. Build the app first (produces electron/dist/...):
    • Windows: pwsh scripts/build-app-win.ps1
    • macOS: bash scripts/build-app.sh
  2. Then:
    cd e2e
    PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD=1 npm ci   # Electron ships its own Chromium
    npm test
    

Override the binary location with E2E_APP_PATH=/path/to/app if your build output lives elsewhere. Auto-detection covers win-unpacked/OpenSwarm.exe and the mac OpenSwarm.app variants.

CI

.github/workflows/e2e.yml runs this on a windows-latest + macos-latest matrix: it builds the unsigned app, then runs the suite. Tag-driven signed releases are covered separately by release-windows.yml / release-macos.yml.