mirror of
https://github.com/openswarm-ai/openswarm.git
synced 2026-09-09 03:07:45 +02:00
Six of the eleven defects here were in the MEASUREMENT, not the product, and they were wrong in both directions. Harness, all of which silently produced wrong numbers: - coverage.py preflight refused every sweep on a box holding exactly one backend: stack.sh's supervisor is a `bash -c` quoting the whole uvicorn line, so it carries both "-m uvicorn backend.main" AND the venv python path. Discriminate on POSITION. - stack.sh status reported 2 backends over 1 and 0 webpack over a live dev server (webpack retitles its process). A status check whose job is preventing a second stack, failing in the direction that lets one land. - c7_run.sh/c8_run.sh slice r6_be.log while stack.sh names logs by TAG: a stack under any other tag hands every trial an empty slice and the sweep reports a confident 0/108. Now refuses loudly; it caught this exact mistake on first use. - "Browser command timed out" was bucketed infra. It is ONE command blowing its own budget, not a dead webview: all 4 such rows were BrowserFindComposer at exactly its 30s cap, every run completed after, zero card-gone markers in the whole log. Filed as infra it read as 11.8% flake AND lifted holdout reach 70% -> 84%. - api_retry / rate_limit_error now grade as infra. A provider 429 storm turned clean 15-21s exclusions into 188s product_no_composer rows. - bench.py prints reach BOTH ways when a row is UNVERIFIED. An exclusion resting on the agent's own word quietly flatters the score, and coverage.py's own instruction to confirm it by hand goes unread (I quoted a 100% that excluded onlinegdb). Timing was measuring 0.2% of the run: prestage completes BEFORE metrics_started_at, so other_ms was 25ms of a 12700ms median while prestage (4146ms, ~61%) sat in no bucket at all. prestage_ms/task_ms are now recorded; total_ms is deliberately NOT redefined, which would invalidate every before/after already taken against it. Product: - find_composer rejected ACE/CodeMirror-5/Monaco composers. Their input is a ~1x1 offscreen textarea that paints into a sibling div, so it can never pass a size gate. Accept it when a VISIBLE ancestor is composer-sized; honeypots stay out because the input itself must not be display:none/visibility:hidden/opacity:0. anon reach 80% -> 100%, holdout 89% -> 90%, p95 38.6s -> 9.7s. - the composer poll slept a blind 0+1.2+1.4 = 2.6s whenever prestage staged nothing, which is nearly every run, and it was the whole of other_ms's suspicious constancy (2610-2613ms regardless of tools_ms). Stop when two reads are identical, the rule the opener poll 40 lines below already applies. other_ms -53.7%, tools_ms flat. - prestage no longer navigates to the page it is already on, nor sleeps 0.35s before its first settle probe. - is_replay_boundary reasoned from the NAME alone, so x.com's composer textbox named "Post text" was ruled an irreversible send and truncated its replay to a bare navigate. Excluded by ROLE; first_unsafe_step now passes role through at all. Measured on this box: reach 100% (83% if onlinegdb's unverified exclusion is bogus), 0 false successes in ~155 runs, prestage tier-0/1 2702ms, other_ms -53.7%, infra flake 0/158, holdout 18/20. Criteria 2/4/9 need live writes and are untouched. Full evidence, including what did NOT work, in e2e/browser-v3/RESULTS_2026-08-06.md. Also drops the tracked electron/node_modules symlink pointing at another machine's Downloads folder; it is dangling on every other checkout and re-breaks the install on any stash or checkout. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
89 lines
4.2 KiB
Python
89 lines
4.2 KiB
Python
"""Every literal this harness greps for must exist in the code that is supposed to print it.
|
|
|
|
Written after the fourth instance of the same bug. `DELIVERY CONFIRMED` appeared in the canary and
|
|
nowhere in the backend. `saw_page` listed two `[browser-action] X` strings that are never emitted.
|
|
`replay_full` matched a phrase no logger uses. Each one silently turned a measurement into a
|
|
constant, and each cost hours of chasing a product bug that did not exist.
|
|
|
|
A grep whose needle is absent from the source cannot fail loudly, so make it fail here instead.
|
|
"""
|
|
|
|
import os
|
|
import re
|
|
import subprocess
|
|
import sys
|
|
|
|
TREE = os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
|
|
HERE = os.path.dirname(os.path.abspath(__file__))
|
|
SP = os.path.dirname(HERE)
|
|
|
|
# (label, literal, where it must appear). A literal is a fixed substring of a real log line, with
|
|
# the f-string holes cut out, so it can be checked with a plain fixed-string grep.
|
|
MARKERS = [
|
|
# canary receipt detection
|
|
("receipt/sendscript", "done sent_receipt=", "backend/apps"),
|
|
("receipt/agent-loop", "two-sided receipt passed", "backend/apps"),
|
|
("receipt/autosend", "code-send delivered (receipt verified)", "backend/apps"),
|
|
# coverage.py grading
|
|
("dryrun report", "DRYRUN: WOULD send (fill committed", "backend/apps"),
|
|
("decline", "[browser-sendscript] decline: ", "backend/apps"),
|
|
("disabled submit", "is present but DISABLED", "backend/apps"),
|
|
("fill target", "[browser-sendscript] fill target ", "backend/apps"),
|
|
("fill errored", "fill errored (", "backend/apps"),
|
|
("fill tier", "[browser-sendscript] fill ok via ", "backend/apps"),
|
|
("prestage cost", "[browser-prestage] cost ", "backend/apps"),
|
|
("login wall", "decline: login/auth wall", "backend/apps"),
|
|
("signed out", "decline: signed OUT", "backend/apps"),
|
|
("recovery", "one recovery dispatch", "backend/apps"),
|
|
# coverage.py UNHEALTHY (infrastructure)
|
|
("router watchdog", "9Router watchdog", "backend/apps"),
|
|
("router died", "9Router process died", "backend/apps"),
|
|
("no provider", "No AI provider connected", "backend/apps"),
|
|
("no dashboard", "dispatch refused: no dashboard", "backend/apps"),
|
|
# bench.py infra buckets
|
|
("browser timeout", "Browser command timed out", "backend/apps"),
|
|
("card gone/webview", "not an electron webview", "backend/apps"),
|
|
("card gone/dashboard", "no dashboard is connected", "backend/apps"),
|
|
("card gone/unresponsive", "page unresponsive", "backend/apps"),
|
|
("busy to read", "too busy to read", "frontend/src"),
|
|
# skillstats.py
|
|
("skill record gate", "record gate: honest=", "backend/apps"),
|
|
("skill recorded", "-step skill for ", "backend/apps"),
|
|
("skill re-derived", "re-derived identical ", "backend/apps"),
|
|
("skill not recorded", "NOT recorded (", "backend/apps"),
|
|
("skill matched", "skill matched on ", "backend/apps"),
|
|
("no skill", "no skill for host=", "backend/apps"),
|
|
("prefix replay", "PREFIX replay: ", "backend/apps"),
|
|
("replay attempt", "REPLAY attempt: ", "backend/apps"),
|
|
("replay succeeded", "REPLAY SUCCEEDED in ", "backend/apps"),
|
|
("replay step failed", "replay step failed (", "backend/apps"),
|
|
("not replayed", " not replayed: ", "backend/apps"),
|
|
# the honesty marker the packaged smokes also grep for
|
|
("send-not-verified", "[send clicked, NOT verified]", "backend/apps"),
|
|
]
|
|
|
|
|
|
def main() -> int:
|
|
bad = []
|
|
for label, needle, where in MARKERS:
|
|
r = subprocess.run(["grep", "-rF", "--", needle, os.path.join(TREE, where)],
|
|
capture_output=True, text=True)
|
|
n = len([ln for ln in r.stdout.splitlines() if ln.strip()])
|
|
status = "ok " if n else "DEAD"
|
|
if not n:
|
|
bad.append((label, needle))
|
|
print(f" {status} {label:<22} x{n:<3} {needle!r}")
|
|
print()
|
|
if bad:
|
|
print(f"{len(bad)} DEAD marker(s): a grep for these can never match, so whatever they")
|
|
print("measure is a constant. Fix the needle or delete the check.")
|
|
for label, needle in bad:
|
|
print(f" - {label}: {needle!r}")
|
|
return 1
|
|
print(f"all {len(MARKERS)} markers exist in the source that prints them")
|
|
return 0
|
|
|
|
|
|
if __name__ == "__main__":
|
|
sys.exit(main())
|