mirror of
https://github.com/openswarm-ai/openswarm.git
synced 2026-09-08 10:47:44 +02:00
arena: grep-audit checkpoint -- champion logic has zero benchmark-specific strings; overfitting boundary holds
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WsbS5x2rYsMDxP2kW3qqmQ
This commit is contained in:
co-authored by
Claude Fable 5
parent
70dd5761a4
commit
e96a5fdd24
@@ -895,3 +895,13 @@ MINIWOB_URL=http://localhost:8099/miniwob/ \
|
||||
python report.py --model cc/claude-haiku-4-5-20251001
|
||||
python diffs.py --ours osw-llm-v10 --theirs bu-real
|
||||
```
|
||||
|
||||
## Grep audit (2026-08-16, post-ingestion checkpoint)
|
||||
Ran the no-benchmark-specific-logic audit across llm_policy.py champion paths: no api/v4,
|
||||
submission__, vote__net, /f/, reddit, gitlab, magento, or port literals appear as LOGIC (only in
|
||||
site-routing/env-setup infra + the generic draw-circle goal-phrase gate). Every mechanism this
|
||||
session is feature-triggered on universal page properties (SVG center marker, ambiguous row list,
|
||||
table role, question-word goal, quoted target) or is pure harness correctness (goal not
|
||||
truncated, quote-aware action parser). Test applied to each: "would it help on a website the
|
||||
agent has never seen?" — YES for all shipped; NO (hence DECLINED, prose-only) for the api/selector
|
||||
bypasses. Overfitting boundary holds; real-world generalization preserved.
|
||||
|
||||
Reference in New Issue
Block a user