mirror of
https://github.com/affaan-m/ECC.git
synced 2026-09-29 21:15:16 +02:00
fix(tests): reconcile launcher coverage with current main
This commit is contained in:
@@ -20,7 +20,7 @@ jobs:
|
||||
test:
|
||||
name: Test (${{ matrix.os }}, Node ${{ matrix.node }}, ${{ matrix.pm }})
|
||||
runs-on: ${{ matrix.os }}
|
||||
timeout-minutes: 20
|
||||
timeout-minutes: 30
|
||||
|
||||
strategy:
|
||||
fail-fast: false
|
||||
|
||||
@@ -48,7 +48,7 @@ jobs:
|
||||
name: Stale Issues/PRs
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/stale@1e223db275d687790206a7acac4d1a11bd6fe629 # v10.4.0
|
||||
- uses: actions/stale@4391f3da665fdf50b6810c1a66712fb9ba21aa93 # v11.0.0
|
||||
with:
|
||||
stale-issue-message: 'This issue is stale due to inactivity.'
|
||||
stale-pr-message: 'This PR is stale due to inactivity.'
|
||||
|
||||
@@ -0,0 +1,19 @@
|
||||
ARG NODE_IMAGE=node:22-bookworm-slim
|
||||
FROM ${NODE_IMAGE}
|
||||
ARG CODEX_VERSION=0.154.0
|
||||
WORKDIR /consumer
|
||||
COPY package.tgz /tmp/ecc-context-package.tgz
|
||||
RUN npm install --ignore-scripts --omit=dev --no-audit --no-fund --fetch-timeout=30000 --fetch-retries=1 /tmp/ecc-context-package.tgz \
|
||||
&& task_arch=$(node -p process.arch) \
|
||||
&& npm install --global --ignore-scripts --no-audit --no-fund --fetch-timeout=30000 --fetch-retries=1 \
|
||||
@openai/codex@${CODEX_VERSION} "@openai/codex-linux-${task_arch}@npm:@openai/codex@${CODEX_VERSION}-linux-${task_arch}" \
|
||||
&& codex --version
|
||||
COPY native-probe.js /consumer/node_modules/ecc-universal/docker/context-profiles/native-probe.js
|
||||
COPY native-switch-probe.js /consumer/node_modules/ecc-universal/docker/context-profiles/native-switch-probe.js
|
||||
COPY packed-smoke.js /consumer/node_modules/ecc-universal/docker/context-profiles/packed-smoke.js
|
||||
COPY context-carrier-fixture.js /consumer/node_modules/ecc-universal/tests/lib/helpers/context-carrier-fixture.js
|
||||
COPY expected-carriers.json /tmp/ecc-expected-carriers.json
|
||||
ENV ECC_EXPECTED_CARRIERS=/tmp/ecc-expected-carriers.json
|
||||
ENV PATH="/consumer/node_modules/.bin:${PATH}"
|
||||
USER node
|
||||
CMD ["node", "/consumer/node_modules/ecc-universal/docker/context-profiles/packed-smoke.js"]
|
||||
@@ -0,0 +1,69 @@
|
||||
# Context profile native and fresh install checks
|
||||
|
||||
These opt-in probes exercise real native discovery without creating a model
|
||||
thread or copying credentials. They are separate from the default unit suite.
|
||||
|
||||
```sh
|
||||
node docker/context-profiles/native-probe.js
|
||||
node docker/context-profiles/native-probe.js --claude
|
||||
node docker/context-profiles/native-switch-probe.js
|
||||
node docker/context-profiles/run-podman.js
|
||||
```
|
||||
|
||||
The first command uses the locally installed Codex executable, a new private
|
||||
temporary home for each case, a local marketplace, and the native plugin cache.
|
||||
It starts a new app-server process and calls only `initialize` and `skills/list`.
|
||||
Lean, Lean with Angular's bundled resources, and Full excluding Python patterns
|
||||
must expose exactly their selected plugin skill names. Provider-owned system
|
||||
skills are reported separately. Every installed resource is checked against its
|
||||
source digest after removing the local marketplace's carrier source.
|
||||
|
||||
The Claude command uses the locally installed Claude executable, a private
|
||||
temporary home, empty setting sources, `plugin validate`, and `plugin details`
|
||||
with an inline plugin directory. It checks exact Lean/Full-with-exclusion skill
|
||||
inventories and zero agent, hook, MCP, and LSP components. Reported token costs
|
||||
are the provider's projections, not measured usage. Manifest attribution and
|
||||
version warnings remain visible.
|
||||
|
||||
The switch probe uses the product's managed store and isolated native adapter for
|
||||
Full, Lean, and rollback to Full. Preparation creates a separate provider home
|
||||
and registers the selected carrier, then opens a fresh app-server to verify
|
||||
discovery. Rollback first restores managed authority, then re-verifies the prior
|
||||
native home and selects it. The Full Python exclusion and unrelated bytes in the
|
||||
prior home must survive every transition. Each native pointer binds its managed
|
||||
store revision, carrier digest, exact provider version, and native executable
|
||||
SHA-256. Read-only status rechecks receipts, native configuration, cached resource
|
||||
bytes, and the pinned executable. Existing sessions and host registration remain
|
||||
unchanged.
|
||||
|
||||
The Podman runner runs the normal `npm pack` lifecycle, reports its archive
|
||||
SHA-256, and builds an isolated consumer from that archive. It installs runtime
|
||||
dependencies and pinned Codex 0.154.0 during the image build. The final container
|
||||
runs as the image's unprivileged `node` user, with networking disabled, all Linux
|
||||
capabilities dropped, no added host mounts, and no copied credentials. It checks
|
||||
all ten target/profile combinations through the packed public CLI and independent
|
||||
structural oracle, including exact carrier equality with the source checkout.
|
||||
It also checks the packed CLI's Full/Lean/rollback lifecycle, idempotency, stale
|
||||
revision rejection, Auto context loading, Suggest/Manual/dry-run boundaries,
|
||||
pinned receipt reuse, and no-workflow reset. It then repeats native Codex discovery
|
||||
and product native preparation/rollback. The packed CLI also prepares a native
|
||||
generation and verifies an isolated launch dry-run with no provider on PATH.
|
||||
Test helpers are
|
||||
copied separately into the image; they are not part of the published package.
|
||||
|
||||
An existing compatible Node image can be selected with
|
||||
`ECC_CONTEXT_NODE_IMAGE=<image-id>`. The default is `node:22-bookworm-slim`.
|
||||
The task image and private temporary build directory are removed afterward.
|
||||
Dependency download layers can remain in Podman's ordinary build cache. The
|
||||
runner never changes host harness configuration or mounts a host home.
|
||||
|
||||
The outcome evaluator (`ai-eval.js`) measures graded task success and provider
|
||||
usage across install arms; see `ai-corpus.json` for the 30-task repair corpus
|
||||
and `complex-eval/DESIGN.md` for the preregistered three-task complex-task
|
||||
benchmark (feature build, incident triage, security hardening) with scored
|
||||
hidden graders, reference solutions, and reproduction instructions.
|
||||
|
||||
These checks certify the observed discovery paths for the reported exact provider
|
||||
versions. They do not certify model invocation, skill workflow outcomes,
|
||||
implicit provider invocation of Auto, host activation, crash recovery, permission consent, or actual token
|
||||
savings. CLI-provided system skills still contribute to whole-session context.
|
||||
@@ -0,0 +1,415 @@
|
||||
{
|
||||
"schemaVersion": "ecc.context-eval-corpus.v2",
|
||||
"id": "coding-tasks@1",
|
||||
"sampling": "Purposive coding-task corpus fixed before any provider call: 22 small JavaScript repairs paired with one plausibly helpful ECC skill, 8 trivial no-workflow fixes (some with misleading workflow vocabulary), and selection probes for exact names, paraphrases, no-workflow queries and policy blocks; equal weight per distinct task and no population-representativeness claim.",
|
||||
"minimumDistinctTasks": 30,
|
||||
"nonInferiorityMargin": 0.05,
|
||||
"selection": [
|
||||
{ "id":"exact-python", "category":"exact", "query":"Use python-patterns to review typed Python functions.", "expectedIds":["skill:python-patterns"] },
|
||||
{ "id":"exact-api", "category":"exact", "query":"Use api-design for REST pagination.", "expectedIds":["skill:api-design"] },
|
||||
{ "id":"paraphrase-tests", "category":"paraphrase", "query":"Write pytest fixtures and parametrized regression tests for a Python package.", "expectedIds":["skill:python-testing"] },
|
||||
{ "id":"paraphrase-api", "category":"paraphrase", "query":"Design REST endpoints with pagination and status codes.", "expectedIds":["skill:api-design"] },
|
||||
{ "id":"plain-arithmetic", "category":"no-workflow", "query":"What is 17 times 24?", "expectedIds":[] },
|
||||
{ "id":"ambiguous-vocabulary", "category":"no-workflow", "query":"Count words in this literal text: database testing security review. Do not perform any of those activities.", "expectedIds":[] },
|
||||
{ "id":"negative-skill-mention", "category":"no-workflow", "query":"Do not use python-patterns. Return only OK.", "expectedIds":[] },
|
||||
{ "id":"explicit-no-workflow", "category":"no-workflow", "query":"Use python-patterns as plain text only. Add 7 and 5.", "noWorkflow":true,"expectedIds":[] },
|
||||
{ "id":"excluded-name", "category":"policy", "query":"Use python-patterns to simplify Python code.", "exclude":["skill:python-patterns"],"expectedIds":[] },
|
||||
{ "id":"excluded-explicit", "category":"policy", "query":"Use python-patterns.", "exclude":["skill:python-patterns"],"explicitIds":["skill:python-patterns"],"expectedBlock":"excluded" },
|
||||
{ "id":"authority-boundary", "category":"policy", "query":"Use inherit-legacy-style to preserve conventions.", "explicitIds":["skill:inherit-legacy-style"],"expectedBlock":"native-authority" },
|
||||
{ "id":"opt-out-conflict", "category":"policy", "query":"Use python-patterns.", "noWorkflow":true,"explicitIds":["skill:python-patterns"],"expectedBlock":"opt-out-conflict" },
|
||||
{ "id":"unknown-explicit", "category":"policy", "query":"Use an unavailable workflow.", "explicitIds":["skill:ecc-eval-nonexistent"],"expectedBlock":"unknown-id" },
|
||||
{ "id":"exact-security-review", "category":"exact", "query":"Use security-review to check this login handler for SQL injection and leaked secrets.", "expectedIds":["skill:security-review"] },
|
||||
{ "id":"exact-error-handling", "category":"exact", "query":"Use error-handling to add typed error classes to the config loader.", "expectedIds":["skill:error-handling"] },
|
||||
{ "id":"exact-database-migrations", "category":"exact", "query":"Use database-migrations to add a NOT NULL column to a large Postgres table.", "expectedIds":["skill:database-migrations"] },
|
||||
{ "id":"exact-regex-structured-text", "category":"exact", "query":"Use regex-vs-llm-structured-text to decide how to parse vendor invoice lines.", "expectedIds":["skill:regex-vs-llm-structured-text"] },
|
||||
{ "id":"exact-content-hash-cache", "category":"exact", "query":"Use content-hash-cache-pattern to cache PDF text extraction results.", "expectedIds":["skill:content-hash-cache-pattern"] },
|
||||
{ "id":"exact-hexagonal", "category":"exact", "query":"Use hexagonal-architecture to separate the signup use case from its database and email adapters.", "expectedIds":["skill:hexagonal-architecture"] },
|
||||
{ "id":"paraphrase-sql-injection", "category":"paraphrase", "query":"User input is concatenated into SQL strings in our login endpoint; audit the handler for injection and hardcoded credentials before release.", "expectedIds":["skill:security-review"] },
|
||||
{ "id":"paraphrase-retry", "category":"paraphrase", "query":"Wrap a flaky payment provider call with exponential backoff retries and typed error classes so callers get useful failure messages.", "expectedIds":["skill:error-handling"] },
|
||||
{ "id":"paraphrase-zero-downtime-rename", "category":"paraphrase", "query":"Rename a column on a busy PostgreSQL table without downtime, with reversible up and down schema changes.", "expectedIds":["skill:database-migrations"] },
|
||||
{ "id":"paraphrase-redis-cache", "category":"paraphrase", "query":"Add a Redis cache-aside layer with key expiry and a distributed lock for our profile reads.", "expectedIds":["skill:redis-patterns"] },
|
||||
{ "id":"paraphrase-token-decimals", "category":"paraphrase", "query":"Our dashboard shows USDC balances wrong on some EVM chains because token decimals differ; normalize amounts across chains safely.", "expectedIds":["skill:evm-token-decimals"] },
|
||||
{ "id":"paraphrase-keccak", "category":"paraphrase", "query":"Compute Ethereum function selectors in Node without confusing NIST SHA3-256 with Keccak-256.", "expectedIds":["skill:nodejs-keccak256"] },
|
||||
{ "id":"paraphrase-content-hash", "category":"paraphrase", "query":"Cache slow document parsing so results are keyed by the SHA-256 of file content instead of the file path.", "expectedIds":["skill:content-hash-cache-pattern"] },
|
||||
{ "id":"paraphrase-ports-adapters", "category":"paraphrase", "query":"Refactor toward ports and adapters so the domain use case no longer imports the database driver directly.", "expectedIds":["skill:hexagonal-architecture"] },
|
||||
{ "id":"paraphrase-structured-text", "category":"paraphrase", "query":"Should I parse these semi-structured quiz and invoice text lines with regular expressions or an LLM? Start with the cheapest reliable option.", "expectedIds":["skill:regex-vs-llm-structured-text"] },
|
||||
{ "id":"rename-variable", "category":"no-workflow", "query":"Rename the local variable tmp to total in this three-line function.", "expectedIds":[] },
|
||||
{ "id":"misleading-security-typo", "category":"no-workflow", "query":"Fix the spelling of \"recieve\" in the footer text of the security settings page. Nothing else.", "expectedIds":[] },
|
||||
{ "id":"misleading-tests-heading", "category":"no-workflow", "query":"Change the README heading \"Running tests\" to \"Running checks\". Do not write or run any tests.", "expectedIds":[] },
|
||||
{ "id":"explicit-no-workflow-migration", "category":"no-workflow", "query":"Treat database-migrations as plain words. Reverse the string abc.", "noWorkflow":true,"expectedIds":[] },
|
||||
{ "id":"excluded-api-explicit", "category":"policy", "query":"Use api-design.", "exclude":["skill:api-design"],"explicitIds":["skill:api-design"],"expectedBlock":"excluded" },
|
||||
{ "id":"authority-latency", "category":"policy", "query":"Use latency-critical-systems to tune the quote cache.", "explicitIds":["skill:latency-critical-systems"],"expectedBlock":"native-authority" },
|
||||
{ "id":"authority-rust-testing", "category":"policy", "query":"Use rust-testing for property tests.", "explicitIds":["skill:rust-testing"],"expectedBlock":"native-authority" },
|
||||
{ "id":"opt-out-conflict-security", "category":"policy", "query":"Use security-review.", "noWorkflow":true,"explicitIds":["skill:security-review"],"expectedBlock":"opt-out-conflict" },
|
||||
{ "id":"unknown-typo-id", "category":"policy", "query":"Use security-reveiw.", "explicitIds":["skill:security-reveiw"],"expectedBlock":"unknown-id" },
|
||||
{ "id":"explicit-allowed", "category":"policy", "query":"Use error-handling for the retry wrapper.", "explicitIds":["skill:error-handling"],"expectedIds":["skill:error-handling"] }
|
||||
],
|
||||
"tasks": [
|
||||
{
|
||||
"id": "sql-injection-query",
|
||||
"category": "security",
|
||||
"manualIds": [
|
||||
"skill:security-review"
|
||||
],
|
||||
"query": "src/users.js builds SQL for a node-postgres style driver: each builder returns { text, values } where text uses $1, $2 placeholders. Both buildFindUserQuery(email) and buildSearchUsersQuery(nameFragment, limit) interpolate caller input into the SQL text. Fix them so no caller-supplied string is ever placed in the SQL text; pass it through values instead. The search must still match names containing the fragment case-insensitively. limit must be an integer from 1 to 100; throw a RangeError for anything else (including numeric strings). Keep both exports and the selected columns. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/users.js": "'use strict';\n\n// Query builders used by the /users routes. The db layer calls\n// pool.query(q.text, q.values) with node-postgres.\nfunction buildFindUserQuery(email) {\n return {\n text: \"SELECT id, email, name FROM users WHERE email = '\" + email + \"'\",\n values: [],\n };\n}\n\nfunction buildSearchUsersQuery(nameFragment, limit) {\n return {\n text: \"SELECT id, email, name FROM users WHERE name ILIKE '%\" + nameFragment + \"%' ORDER BY name LIMIT \" + limit,\n values: [],\n };\n}\n\nmodule.exports = { buildFindUserQuery, buildSearchUsersQuery };\n",
|
||||
"README.md": "# users-api\n\nSmall query-builder layer for the users service. Builders return\n{ text, values } objects that are passed directly to node-postgres.\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst m = require(path.join(process.cwd(), 'src/users.js'));\nconst evil = \"a' OR '1'='1\";\nconst q = m.buildFindUserQuery(evil);\nassert.ok(!q.text.includes(\"'1'='1\") && !q.text.includes(evil), 'email leaked into SQL text');\nassert.match(q.text, /\\$1/);\nassert.match(q.text, /FROM\\s+users/i);\nassert.ok(q.values.includes(evil));\nconst frag = \"bob'; DROP TABLE users;--\";\nconst s = m.buildSearchUsersQuery(frag, 10);\nassert.ok(!/DROP\\s+TABLE/i.test(s.text), 'fragment leaked into SQL text');\nassert.match(s.text, /ILIKE|lower\\s*\\(/i);\nassert.match(s.text, /LIMIT\\s+(\\$\\d+|10)\\b/i);\nassert.ok(s.values.some(v => typeof v === 'string' && v.includes(frag)));\nfor (const bad of [0, 101, 2.5, '10', '10; DROP TABLE users', NaN, undefined]) {\n assert.throws(() => m.buildSearchUsersQuery('x', bad), RangeError);\n}\nconst max = Math.max(0, ...[...s.text.matchAll(/\\$(\\d+)/g)].map(x => Number(x[1])));\nassert.equal(max, s.values.length, 'placeholders and values disagree');\n"
|
||||
},
|
||||
{
|
||||
"id": "path-traversal-guard",
|
||||
"category": "security",
|
||||
"manualIds": [
|
||||
"skill:security-review"
|
||||
],
|
||||
"query": "src/static.js exports resolvePublicPath(requestPath, root) used by our static file server. requestPath is the raw URL path (for example \"/css/site.css\", possibly percent-encoded). It currently joins it onto root, which allows escaping the public directory. Make it return the absolute file path when the decoded path stays inside root (root itself counts as inside), and return null (never throw) when the path escapes root, contains a NUL byte, or cannot be percent-decoded. Watch out for sibling directories that share root as a string prefix. Keep the export name and signature. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/static.js": "'use strict';\nconst path = require('path');\n\nconst PUBLIC_ROOT = path.resolve(__dirname, '..', 'public');\n\n// Maps a request path such as \"/css/site.css\" to a file on disk.\nfunction resolvePublicPath(requestPath, root = PUBLIC_ROOT) {\n return path.join(root, decodeURIComponent(requestPath));\n}\n\nmodule.exports = { resolvePublicPath, PUBLIC_ROOT };\n",
|
||||
"public/index.html": "<!doctype html><title>home</title>\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { resolvePublicPath } = require(path.join(process.cwd(), 'src/static.js'));\nconst root = path.resolve(path.sep + 'srv', 'app', 'public');\nassert.equal(resolvePublicPath('/css/site.css', root), path.join(root, 'css', 'site.css'));\nassert.equal(resolvePublicPath('/css/../index.html', root), path.join(root, 'index.html'));\nassert.equal(resolvePublicPath('/a%20b.txt', root), path.join(root, 'a b.txt'));\nfor (const bad of ['/../secret.env', '/%2e%2e/%2e%2e/etc/passwd', '/css/../../x', '/../public-evil/x',\n '/a%00.txt', '/%E0%A4%A', '..%2f..%2fetc%2fpasswd']) {\n let out;\n assert.doesNotThrow(() => { out = resolvePublicPath(bad, root); }, bad);\n assert.equal(out, null, bad);\n}\n"
|
||||
},
|
||||
{
|
||||
"id": "escape-comment-html",
|
||||
"category": "security",
|
||||
"manualIds": [
|
||||
"skill:security-review"
|
||||
],
|
||||
"query": "src/render.js exports renderComment({ author, body, website }) which returns an HTML string for a user comment. All three fields are untrusted user input and are currently inserted raw. Fix it so author and body are HTML-escaped (at least & < > \" and '), and website is only used as the link href when it is an absolute http: or https: URL; otherwise the href must be \"#\". The href value must also be escaped. Keep the existing markup structure (li.comment containing an a element and a p element). Do not add dependencies.",
|
||||
"files": {
|
||||
"src/render.js": "'use strict';\n\nfunction renderComment({ author, body, website }) {\n return '<li class=\"comment\"><a href=\"' + website + '\">' + author + '</a><p>' + body + '</p></li>';\n}\n\nmodule.exports = { renderComment };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { renderComment } = require(path.join(process.cwd(), 'src/render.js'));\nconst a = renderComment({ author: '<script>alert(1)</script>', body: 'Tom & \"Jerry\" \\'s', website: 'https://ex.com/' });\nassert.ok(a.startsWith('<li class=\"comment\">'));\nassert.ok(!a.includes('<script'));\nassert.ok(a.includes('<script>'));\nassert.ok(a.includes('&') && a.includes('"'));\nassert.ok(/�*39;|�*27;|'/i.test(a));\nassert.ok(a.includes('href=\"https://ex.com/\"'));\nfor (const w of ['javascript:alert(1)', ' JavaScript:alert(1)', 'data:text/html,x', 'vbscript:x', '//evil.com', '']) {\n const out = renderComment({ author: 'a', body: 'b', website: w });\n assert.ok(!/javascript:|data:|vbscript:/i.test(out), w);\n assert.ok(out.includes('href=\"#\"'), w);\n}\nconst q = renderComment({ author: 'a', body: 'b', website: 'https://ex.com/?a=1&b=\"x\"' });\nassert.ok(!q.includes('\"x\"'));\nassert.ok(/href=\"https:\\/\\/ex\\.com\\/\\?a=1&b=/.test(q));\nassert.ok(/<a [^>]*>a<\\/a>/.test(q) && /<p>b<\\/p>/.test(q));\n"
|
||||
},
|
||||
{
|
||||
"id": "list-pagination",
|
||||
"category": "api",
|
||||
"manualIds": [
|
||||
"skill:api-design"
|
||||
],
|
||||
"query": "src/listProducts.js exports listProducts(query, store) for GET /products. query holds raw query-string values (strings or undefined); store.all() returns the full array. Implement offset pagination: limit defaults to 20 and must be an integer 1..100, offset defaults to 0 and must be an integer >= 0. Success returns { status: 200, body: { data, meta: { total, limit, offset, hasMore } } }. Invalid values return { status: 400, body: { error: { code: \"VALIDATION_ERROR\", message, details: [{ field, message }] } } } with one details entry per invalid field (\"limit\" or \"offset\"). Do not mutate the store array. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/listProducts.js": "'use strict';\n\n// GET /products?limit=&offset=\nfunction listProducts(query, store) {\n const items = store.all();\n const page = items.slice(query.offset, query.offset + query.limit);\n return { status: 200, body: page };\n}\n\nmodule.exports = { listProducts };\n",
|
||||
"src/store.js": "'use strict';\n\nfunction createStore(items) {\n return { all: () => items };\n}\n\nmodule.exports = { createStore };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { listProducts } = require(path.join(process.cwd(), 'src/listProducts.js'));\nconst items = Array.from({ length: 45 }, (_, i) => ({ id: i + 1 }));\nconst copy = JSON.stringify(items);\nconst store = { all: () => items };\nlet r = listProducts({}, store);\nassert.equal(r.status, 200);\nassert.equal(r.body.data.length, 20);\nassert.deepEqual(r.body.meta, { total: 45, limit: 20, offset: 0, hasMore: true });\nr = listProducts({ limit: '10', offset: '40' }, store);\nassert.deepEqual(r.body.data.map(x => x.id), [41, 42, 43, 44, 45]);\nassert.deepEqual(r.body.meta, { total: 45, limit: 10, offset: 40, hasMore: false });\nr = listProducts({ limit: '5', offset: '35' }, store);\nassert.equal(r.body.meta.hasMore, true);\nr = listProducts({ limit: '100', offset: '100' }, store);\nassert.equal(r.status, 200);\nassert.deepEqual(r.body.data, []);\nassert.equal(r.body.meta.hasMore, false);\nfor (const [q, fields] of [[{ limit: '0' }, ['limit']], [{ limit: '101' }, ['limit']], [{ limit: 'abc' }, ['limit']],\n [{ limit: '2.5' }, ['limit']], [{ offset: '-1' }, ['offset']], [{ limit: '-3', offset: 'x' }, ['limit', 'offset']]]) {\n const bad = listProducts(q, store);\n assert.equal(bad.status, 400, JSON.stringify(q));\n assert.equal(bad.body.error.code, 'VALIDATION_ERROR');\n assert.equal(typeof bad.body.error.message, 'string');\n assert.deepEqual(bad.body.error.details.map(d => d.field).sort(), fields);\n assert.ok(bad.body.error.details.every(d => typeof d.message === 'string'));\n}\nassert.equal(JSON.stringify(items), copy);\n"
|
||||
},
|
||||
{
|
||||
"id": "create-user-status-codes",
|
||||
"category": "api",
|
||||
"manualIds": [
|
||||
"skill:api-design"
|
||||
],
|
||||
"query": "src/usersRoute.js exports async createUser(req, repo) for POST /users and async getUser(req, repo) for GET /users/:id. Both return { status, headers?, body }. They currently return 200 for everything and 500 on duplicates. Fix them to use proper REST semantics. createUser: body { email, name }; email must be a string containing \"@\" and name a non-empty trimmed string, otherwise 400 with body { error: { code: \"VALIDATION_ERROR\", message, details: [{ field, message }] } } listing each bad field; if repo.findByEmail(email) returns a user, 409 with error code \"CONFLICT\"; otherwise call repo.create({ email, name }) and return 201 with headers { Location: \"/users/<id>\" } and body { data: user }. getUser: req.params.id; missing user gives 404 with error code \"NOT_FOUND\", found user gives 200 { data: user }. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/usersRoute.js": "'use strict';\n\nasync function createUser(req, repo) {\n try {\n const { email, name } = req.body || {};\n const existing = await repo.findByEmail(email);\n if (existing) throw new Error('duplicate');\n const user = await repo.create({ email, name });\n return { status: 200, body: user };\n } catch (err) {\n return { status: 500, body: { message: err.message } };\n }\n}\n\nasync function getUser(req, repo) {\n const user = await repo.findById(req.params.id);\n return { status: 200, body: user };\n}\n\nmodule.exports = { createUser, getUser };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { createUser, getUser } = require(path.join(process.cwd(), 'src/usersRoute.js'));\nfunction repo() {\n const users = [{ id: 1, email: 'ada@example.com', name: 'Ada' }];\n return { created: 0, async findByEmail(e) { return users.find(u => u.email === e) || null; },\n async findById(id) { return users.find(u => String(u.id) === String(id)) || null; },\n async create(u) { this.created++; const user = { id: users.length + 1, ...u }; users.push(user); return user; } };\n}\n(async () => {\n const r = repo();\n let res = await createUser({ body: { email: 'lin@example.com', name: 'Lin' } }, r);\n assert.equal(res.status, 201);\n assert.equal(res.headers.Location, '/users/2');\n assert.deepEqual(res.body.data, { id: 2, email: 'lin@example.com', name: 'Lin' });\n res = await createUser({ body: { email: 'ada@example.com', name: 'Ada2' } }, r);\n assert.equal(res.status, 409);\n assert.equal(res.body.error.code, 'CONFLICT');\n res = await createUser({ body: { email: 'nope', name: ' ' } }, r);\n assert.equal(res.status, 400);\n assert.equal(res.body.error.code, 'VALIDATION_ERROR');\n assert.deepEqual(res.body.error.details.map(d => d.field).sort(), ['email', 'name']);\n res = await createUser({ body: { email: 'x@y.z' } }, r);\n assert.equal(res.status, 400);\n assert.deepEqual(res.body.error.details.map(d => d.field), ['name']);\n assert.equal(r.created, 1);\n res = await getUser({ params: { id: '99' } }, r);\n assert.equal(res.status, 404);\n assert.equal(res.body.error.code, 'NOT_FOUND');\n res = await getUser({ params: { id: '1' } }, r);\n assert.equal(res.status, 200);\n assert.equal(res.body.data.email, 'ada@example.com');\n})().catch(err => { console.error(err); process.exitCode = 1; });\n"
|
||||
},
|
||||
{
|
||||
"id": "retry-with-backoff",
|
||||
"category": "errors",
|
||||
"manualIds": [
|
||||
"skill:error-handling"
|
||||
],
|
||||
"query": "src/retry.js exports async withRetry(fn, options) used around calls to a flaky payments API. It currently retries every error immediately and throws a generic Error(\"failed\"), losing the cause. Rewrite it: options are { retries = 3, baseDelayMs = 100, maxDelayMs = 2000, sleep } where sleep(ms) returns a promise (default: a real setTimeout sleep). Call fn(attempt) with attempt starting at 1, for at most retries + 1 attempts. Only retry when the error is retryable: err.retryable === true, or err.status is 429 or >= 500. Non-retryable errors must be rethrown immediately (the same error object). Before retry n (n = 1, 2, ...) await sleep(d) where d is between half and all of min(baseDelayMs * 2^(n-1), maxDelayMs) (jitter optional). When retries are exhausted, rethrow the last error object. Return fn's resolved value on success. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/retry.js": "'use strict';\n\nasync function withRetry(fn, options = {}) {\n const retries = options.retries || 3;\n for (let i = 0; i < retries; i++) {\n try {\n return await fn(i);\n } catch (err) {\n // try again\n }\n }\n throw new Error('failed');\n}\n\nmodule.exports = { withRetry };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { withRetry } = require(path.join(process.cwd(), 'src/retry.js'));\nconst mk = (status, extra = {}) => Object.assign(new Error('e' + status), { status }, extra);\n(async () => {\n let delays = [];\n const sleep = ms => { delays.push(ms); return Promise.resolve(); };\n let calls = [];\n const out = await withRetry(async a => { calls.push(a); if (a < 3) throw mk(503); return 'ok'; }, { sleep });\n assert.equal(out, 'ok');\n assert.deepEqual(calls, [1, 2, 3]);\n assert.equal(delays.length, 2);\n assert.ok(delays[0] >= 50 && delays[0] <= 100 && delays[1] >= 100 && delays[1] <= 200, String(delays));\n delays = []; calls = [];\n const last = mk(500);\n let n = 0;\n await assert.rejects(withRetry(async a => { calls.push(a); n++; throw n === 5 ? last : mk(502); },\n { retries: 4, baseDelayMs: 1000, maxDelayMs: 3000, sleep }), e => e === last);\n assert.deepEqual(calls, [1, 2, 3, 4, 5]);\n const caps = [1000, 2000, 3000, 3000];\n assert.equal(delays.length, 4);\n delays.forEach((d, i) => assert.ok(d >= caps[i] / 2 && d <= caps[i], 'delay ' + i + '=' + d));\n delays = []; calls = [];\n const bad = mk(400);\n await assert.rejects(withRetry(async a => { calls.push(a); throw bad; }, { sleep }), e => e === bad);\n assert.deepEqual(calls, [1]);\n assert.equal(delays.length, 0);\n calls = [];\n const plain = new Error('boom');\n await assert.rejects(withRetry(async a => { calls.push(a); throw plain; }, { sleep }), e => e === plain);\n assert.equal(calls.length, 1);\n calls = [];\n await withRetry(async a => { calls.push(a); if (a === 1) throw mk(429); if (a === 2) throw Object.assign(new Error('r'), { retryable: true }); return 1; }, { sleep });\n assert.deepEqual(calls, [1, 2, 3]);\n calls = [];\n await assert.rejects(withRetry(async a => { calls.push(a); throw mk(503); }, { retries: 0, sleep }));\n assert.deepEqual(calls, [1]);\n})().catch(err => { console.error(err); process.exitCode = 1; });\n"
|
||||
},
|
||||
{
|
||||
"id": "typed-config-errors",
|
||||
"category": "errors",
|
||||
"manualIds": [
|
||||
"skill:error-handling"
|
||||
],
|
||||
"query": "src/config.js exports loadConfig(text), which parses a JSON config string. Today it silently returns {} on bad JSON and accepts missing fields. Add and export a ConfigError class (extends Error, name \"ConfigError\") with a code property, and make loadConfig throw it: code \"CONFIG_PARSE\" for invalid JSON (with the original SyntaxError as error.cause); code \"CONFIG_MISSING\" with error.field set when a required field is missing (required: apiUrl, then timeoutMs, checked in that order); code \"CONFIG_INVALID\" with error.field = \"timeoutMs\" when timeoutMs is not a positive integer. On success return { apiUrl, timeoutMs, retries } where retries defaults to 2. Messages should be human readable. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/config.js": "'use strict';\n\nfunction loadConfig(text) {\n let raw;\n try {\n raw = JSON.parse(text);\n } catch (e) {\n return {};\n }\n return { apiUrl: raw.apiUrl, timeoutMs: raw.timeoutMs, retries: raw.retries };\n}\n\nmodule.exports = { loadConfig };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { loadConfig, ConfigError } = require(path.join(process.cwd(), 'src/config.js'));\nassert.equal(typeof ConfigError, 'function');\nassert.deepEqual(loadConfig('{\"apiUrl\":\"https://x\",\"timeoutMs\":500}'), { apiUrl: 'https://x', timeoutMs: 500, retries: 2 });\nassert.deepEqual(loadConfig('{\"apiUrl\":\"https://x\",\"timeoutMs\":5,\"retries\":0}'), { apiUrl: 'https://x', timeoutMs: 5, retries: 0 });\nfunction thrown(text) { try { loadConfig(text); } catch (e) { return e; } assert.fail('expected throw for ' + text); }\nlet e = thrown('{bad json');\nassert.ok(e instanceof ConfigError && e instanceof Error);\nassert.equal(e.name, 'ConfigError');\nassert.equal(e.code, 'CONFIG_PARSE');\nassert.ok(e.cause instanceof SyntaxError);\nassert.ok(e.message.length > 0);\ne = thrown('{\"timeoutMs\":1}');\nassert.equal(e.code, 'CONFIG_MISSING');\nassert.equal(e.field, 'apiUrl');\ne = thrown('{\"apiUrl\":\"u\"}');\nassert.equal(e.code, 'CONFIG_MISSING');\nassert.equal(e.field, 'timeoutMs');\nfor (const t of ['0', '-5', '1.5', '\"100\"']) {\n e = thrown('{\"apiUrl\":\"u\",\"timeoutMs\":' + t + '}');\n assert.ok(e instanceof ConfigError);\n assert.equal(e.code, 'CONFIG_INVALID');\n assert.equal(e.field, 'timeoutMs');\n}\n"
|
||||
},
|
||||
{
|
||||
"id": "batch-partial-failures",
|
||||
"category": "errors",
|
||||
"manualIds": [
|
||||
"skill:error-handling"
|
||||
],
|
||||
"query": "src/batch.js exports async processAll(items, worker). items are objects with an id; worker(item) returns a promise. The current version swallows errors inside an empty catch and returns only a count, so failed webhook deliveries vanish. Change it to process every item (a failure must not stop the others) and resolve to { succeeded: [{ id, result }], failed: [{ id, error }] }, both in input order, where error is the thrown error's message (or String(value) if a non-Error was thrown). It must never reject because of a worker failure, and a worker that throws synchronously must be treated like a rejection. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/batch.js": "'use strict';\n\nasync function processAll(items, worker) {\n let done = 0;\n for (const item of items) {\n try {\n await worker(item);\n done++;\n } catch (e) {}\n }\n return done;\n}\n\nmodule.exports = { processAll };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { processAll } = require(path.join(process.cwd(), 'src/batch.js'));\n(async () => {\n const seen = [];\n const items = [1, 2, 3, 4, 5].map(id => ({ id }));\n const out = await processAll(items, item => {\n seen.push(item.id);\n if (item.id === 2) throw new Error('sync boom');\n if (item.id === 4) return Promise.reject('plain string');\n if (item.id === 5) return Promise.reject(new TypeError('bad payload'));\n return Promise.resolve(item.id * 10);\n });\n assert.deepEqual(seen.slice().sort(), [1, 2, 3, 4, 5]);\n assert.deepEqual(out.succeeded, [{ id: 1, result: 10 }, { id: 3, result: 30 }]);\n assert.deepEqual(out.failed, [{ id: 2, error: 'sync boom' }, { id: 4, error: 'plain string' }, { id: 5, error: 'bad payload' }]);\n assert.deepEqual(await processAll([], () => 1), { succeeded: [], failed: [] });\n})().catch(err => { console.error(err); process.exitCode = 1; });\n"
|
||||
},
|
||||
{
|
||||
"id": "access-log-parser",
|
||||
"category": "parsing",
|
||||
"manualIds": [
|
||||
"skill:regex-vs-llm-structured-text"
|
||||
],
|
||||
"query": "src/parseLog.js parses web server access logs in Common Log Format, optionally extended to Combined Log Format with a quoted referrer and a quoted user agent. The current parseLine(line) splits on spaces and breaks on user agents and timestamps that contain spaces. Rewrite parseLine(line) to return { ip, user, time, method, path, protocol, status, bytes, referrer, userAgent } or null for any line that does not match the format. user, referrer and userAgent are null when the field is \"-\" or absent; time is the text inside the square brackets; status is a number (three digits); bytes is a number and \"-\" means 0. Also export parseLog(text) returning { entries, invalid } where blank lines (LF or CRLF endings) are skipped and invalid counts non-matching lines. See README.md for examples. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/parseLog.js": "'use strict';\n\nfunction parseLine(line) {\n const parts = line.split(' ');\n return {\n ip: parts[0],\n user: parts[2],\n time: parts[3],\n method: parts[5],\n path: parts[6],\n protocol: parts[7],\n status: Number(parts[8]),\n bytes: Number(parts[9]),\n };\n}\n\nmodule.exports = { parseLine };\n",
|
||||
"README.md": "# log-stats\n\nAccess log examples we must support:\n\n 127.0.0.1 - frank [10/Oct/2000:13:55:36 -0700] \"GET /apache_pb.gif HTTP/1.0\" 200 2326 \"http://www.example.com/start.html\" \"Mozilla/4.08 [en] (Win98; I ;Nav)\"\n 10.0.0.2 - - [11/Oct/2000:08:00:01 +0000] \"POST /api/login HTTP/1.1\" 401 -\n\nThe first is Combined Log Format, the second plain Common Log Format.\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { parseLine, parseLog } = require(path.join(process.cwd(), 'src/parseLog.js'));\nconst a = '127.0.0.1 - frank [10/Oct/2000:13:55:36 -0700] \"GET /apache_pb.gif HTTP/1.0\" 200 2326 \"http://www.example.com/start.html\" \"Mozilla/4.08 [en] (Win98; I ;Nav)\"';\nassert.deepEqual(parseLine(a), { ip: '127.0.0.1', user: 'frank', time: '10/Oct/2000:13:55:36 -0700', method: 'GET',\n path: '/apache_pb.gif', protocol: 'HTTP/1.0', status: 200, bytes: 2326,\n referrer: 'http://www.example.com/start.html', userAgent: 'Mozilla/4.08 [en] (Win98; I ;Nav)' });\nconst b = '10.0.0.2 - - [11/Oct/2000:08:00:01 +0000] \"POST /api/login HTTP/1.1\" 401 -';\nassert.deepEqual(parseLine(b), { ip: '10.0.0.2', user: null, time: '11/Oct/2000:08:00:01 +0000', method: 'POST',\n path: '/api/login', protocol: 'HTTP/1.1', status: 401, bytes: 0, referrer: null, userAgent: null });\nconst c = '::1 - - [01/Jan/2024:00:00:00 +0000] \"DELETE /items/9?force=1 HTTP/2.0\" 204 0 \"-\" \"curl/8.4.0\"';\nconst pc = parseLine(c);\nassert.equal(pc.ip, '::1');\nassert.equal(pc.path, '/items/9?force=1');\nassert.equal(pc.referrer, null);\nassert.equal(pc.userAgent, 'curl/8.4.0');\nassert.equal(pc.status, 204);\nfor (const bad of ['garbage line', '', '10.0.0.2 - - 11/Oct/2000:08:00:01 +0000 \"GET / HTTP/1.1\" 200 5',\n '10.0.0.2 - - [11/Oct/2000:08:00:01 +0000] \"GET / HTTP/1.1\" 2000 5', '10.0.0.2 - - [x] \"GET / HTTP/1.1\" 200 abc',\n '\"GET / HTTP/1.1\" 200 12']) {\n assert.equal(parseLine(bad), null, bad);\n}\nconst log = [a, '', 'nonsense', b + '\\r', ' ', c, ''].join('\\n');\nconst out = parseLog(log);\nassert.equal(out.entries.length, 3);\nassert.equal(out.invalid, 1);\nassert.equal(out.entries[1].bytes, 0);\n"
|
||||
},
|
||||
{
|
||||
"id": "invoice-field-extraction",
|
||||
"category": "parsing",
|
||||
"manualIds": [
|
||||
"skill:regex-vs-llm-structured-text"
|
||||
],
|
||||
"query": "src/extract.js exports extractInvoice(text), which pulls fields out of plain-text invoices from several vendors. It only handles one vendor today. Make it return { invoiceNumber, date, total, currency } for all layouts documented in FORMATS.md: invoiceNumber is the identifier string; date is normalized to YYYY-MM-DD; total is a number (thousands separators removed) taken from the grand total line, never from Subtotal or Tax lines; currency is a three-letter code (\"$\" means USD). Any field that cannot be found is null. Labels are case-insensitive. Keep it deterministic and offline. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/extract.js": "'use strict';\n\nfunction extractInvoice(text) {\n const num = /Invoice #: (\\S+)/.exec(text);\n const date = /Date: (\\d{4}-\\d{2}-\\d{2})/.exec(text);\n const total = /Total: \\$([\\d.]+)/.exec(text);\n return {\n invoiceNumber: num ? num[1] : null,\n date: date ? date[1] : null,\n total: total ? Number(total[1]) : null,\n currency: total ? 'USD' : null,\n };\n}\n\nmodule.exports = { extractInvoice };\n",
|
||||
"FORMATS.md": "# Invoice layouts\n\nInvoice number labels: \"Invoice #:\", \"Invoice No.\", \"Invoice Number:\".\nIdentifiers use letters, digits and hyphens, for example INV-2024-0042, INV-7, A-19.\n\nDate labels: \"Date:\", \"Invoice Date:\", \"Issued:\". Values appear as\n2024-03-05 (ISO), 05/03/2024 (DD/MM/YYYY, day first) or 7 November 2023\n(day, full English month name, year).\n\nGrand total labels: \"Total:\", \"Total due:\", \"Amount due:\". Amounts look like\n$1,234.50 or EUR 99.00 (code before) or 1,000.00 GBP (code after).\nInvoices may also contain \"Subtotal:\" and \"Tax:\" lines, which are not totals.\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { extractInvoice } = require(path.join(process.cwd(), 'src/extract.js'));\nassert.deepEqual(extractInvoice(['ACME Corp', 'Invoice #: INV-2024-0042', 'Date: 2024-03-05', 'Subtotal: $1,100.00',\n 'Tax: $134.50', 'Total: $1,234.50'].join('\\n')), { invoiceNumber: 'INV-2024-0042', date: '2024-03-05', total: 1234.5, currency: 'USD' });\nassert.deepEqual(extractInvoice(['Globex GmbH', 'invoice no. INV-7', 'Invoice Date: 05/03/2024', 'Subtotal: EUR 90.00',\n 'TOTAL DUE: EUR 99.00'].join('\\r\\n')), { invoiceNumber: 'INV-7', date: '2024-03-05', total: 99, currency: 'EUR' });\nassert.deepEqual(extractInvoice(['Initech Ltd', 'Invoice Number: A-19', 'Issued: 7 November 2023', 'Tax: 0.00 GBP',\n 'Amount due: 1,000.00 GBP'].join('\\n')), { invoiceNumber: 'A-19', date: '2023-11-07', total: 1000, currency: 'GBP' });\nassert.deepEqual(extractInvoice('Thanks for your business!'), { invoiceNumber: null, date: null, total: null, currency: null });\nconst partial = extractInvoice('Invoice #: Z-1\\nSubtotal: $5.00');\nassert.equal(partial.invoiceNumber, 'Z-1');\nassert.equal(partial.total, null);\nassert.equal(partial.date, null);\n"
|
||||
},
|
||||
{
|
||||
"id": "add-column-migration",
|
||||
"category": "database",
|
||||
"manualIds": [
|
||||
"skill:database-migrations"
|
||||
],
|
||||
"query": "This repo keeps PostgreSQL migrations in migrations/ as NNN_name.up.sql plus NNN_name.down.sql (see README.md). Add migration 002 (one .up.sql and one .down.sql with the same NNN_name stem) that adds users.email_verified as a boolean that is NOT NULL with default false, and a unique index named users_email_lower_key on lower(email). The users table is large and takes writes constantly, so the index must be built without blocking writes, and the runner does not wrap files in a transaction. The down migration must fully reverse 002 and nothing else. Do not modify migration 001. Do not add dependencies.",
|
||||
"files": {
|
||||
"migrations/001_create_users.up.sql": "CREATE TABLE users (\n id bigserial PRIMARY KEY,\n email text NOT NULL,\n name text NOT NULL,\n created_at timestamptz NOT NULL DEFAULT now()\n);\n",
|
||||
"migrations/001_create_users.down.sql": "DROP TABLE users;\n",
|
||||
"README.md": "# accounts-db\n\nPostgreSQL 15. Migrations live in migrations/ and are applied in filename order.\nEach migration is a pair: NNN_name.up.sql and NNN_name.down.sql.\nThe runner sends each file as-is (no implicit BEGIN/COMMIT).\nProduction: users has about 40 million rows and receives writes all day.\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst fs = require('node:fs');\nconst path = require('node:path');\nconst dir = path.join(process.cwd(), 'migrations');\nconst names = fs.readdirSync(dir);\nconst ups = names.filter(n => /^002_[A-Za-z0-9_-]+\\.up\\.sql$/.test(n));\nassert.equal(ups.length, 1, 'expected one 002 up migration');\nconst stem = ups[0].slice(0, -'.up.sql'.length);\nassert.ok(names.includes(stem + '.down.sql'), 'matching down migration missing');\nconst strip = s => s.replace(/--[^\\n]*/g, '').replace(/\\/\\*[\\s\\S]*?\\*\\//g, '');\nconst up = strip(fs.readFileSync(path.join(dir, ups[0]), 'utf8'));\nconst down = strip(fs.readFileSync(path.join(dir, stem + '.down.sql'), 'utf8'));\nconst add = /ALTER\\s+TABLE\\s+(?:IF\\s+EXISTS\\s+)?(?:ONLY\\s+)?\"?users\"?\\s+ADD\\s+(?:COLUMN\\s+)?(?:IF\\s+NOT\\s+EXISTS\\s+)?\"?email_verified\"?\\s+(?:boolean|bool)\\b([^;]*)/i.exec(up);\nassert.ok(add, 'ADD COLUMN email_verified boolean missing');\nconst col = '\"?email_verified\"?';\nassert.ok(/NOT\\s+NULL/i.test(add[1]) || new RegExp('ALTER\\\\s+COLUMN\\\\s+' + col + '\\\\s+SET\\\\s+NOT\\\\s+NULL', 'i').test(up), 'NOT NULL missing');\nassert.ok(/DEFAULT\\s+(?:false|'f'|'false')/i.test(add[1]) || new RegExp('ALTER\\\\s+COLUMN\\\\s+' + col + '\\\\s+SET\\\\s+DEFAULT\\\\s+false', 'i').test(up), 'DEFAULT false missing');\nassert.match(up, /CREATE\\s+UNIQUE\\s+INDEX\\s+CONCURRENTLY\\s+(?:IF\\s+NOT\\s+EXISTS\\s+)?\"?users_email_lower_key\"?\\s+ON\\s+(?:ONLY\\s+)?\"?users\"?\\s*(?:USING\\s+btree\\s*)?\\(\\s*lower\\s*\\(\\s*\"?email\"?\\s*\\)\\s*\\)/i);\nconst idx = up.search(/CREATE\\s+UNIQUE\\s+INDEX\\s+CONCURRENTLY/i);\nconst opened = [...up.slice(0, idx).matchAll(/\\b(BEGIN|START\\s+TRANSACTION|COMMIT|END|ROLLBACK)\\b\\s*;/gi)].map(x => x[1].toUpperCase());\nassert.ok(!opened.length || !/^(BEGIN|START)/.test(opened[opened.length - 1]), 'concurrent index inside a transaction');\nassert.doesNotMatch(up, /DROP\\s+(?:COLUMN|TABLE|INDEX)/i);\nassert.match(down, /DROP\\s+INDEX\\s+(?:CONCURRENTLY\\s+)?(?:IF\\s+EXISTS\\s+)?\"?users_email_lower_key\"?/i);\nassert.match(down, /ALTER\\s+TABLE\\s+(?:IF\\s+EXISTS\\s+)?\"?users\"?\\s+DROP\\s+(?:COLUMN\\s+)?(?:IF\\s+EXISTS\\s+)?\"?email_verified\"?/i);\nassert.doesNotMatch(down, /DROP\\s+TABLE/i);\nconst original = \"CREATE TABLE users (\\n id bigserial PRIMARY KEY,\\n email text NOT NULL,\\n name text NOT NULL,\\n created_at timestamptz NOT NULL DEFAULT now()\\n);\\n\";\nassert.equal(fs.readFileSync(path.join(dir, '001_create_users.up.sql'), 'utf8'), original);\nassert.equal(fs.readFileSync(path.join(dir, '001_create_users.down.sql'), 'utf8'), 'DROP TABLE users;\\n');\n"
|
||||
},
|
||||
{
|
||||
"id": "rename-column-expand",
|
||||
"category": "database",
|
||||
"manualIds": [
|
||||
"skill:database-migrations"
|
||||
],
|
||||
"query": "We want PostgreSQL column customers.full_name renamed to display_name, but old app instances keep reading and writing full_name for hours during the rolling deploy (see README.md). Do only the zero-downtime expand step. 1) Add migrations/002_<name>.up.sql and matching .down.sql: the up adds a nullable display_name text column and backfills it from full_name; it must not rename or drop full_name. The down removes display_name only. 2) Update src/customerRepo.js: buildInsert(customer) and buildUpdateName(id, name) must write the name to both full_name and display_name (still parameterized { text, values } with $n placeholders), and mapRow(row) must return name from display_name, falling back to full_name when display_name is null. Keep all exports. Do not add dependencies.",
|
||||
"files": {
|
||||
"migrations/001_create_customers.up.sql": "CREATE TABLE customers (\n id bigserial PRIMARY KEY,\n email text NOT NULL,\n full_name text NOT NULL\n);\n",
|
||||
"migrations/001_create_customers.down.sql": "DROP TABLE customers;\n",
|
||||
"src/customerRepo.js": "'use strict';\n\nfunction buildInsert(customer) {\n return { text: 'INSERT INTO customers (email, full_name) VALUES ($1, $2) RETURNING id', values: [customer.email, customer.name] };\n}\n\nfunction buildUpdateName(id, name) {\n return { text: 'UPDATE customers SET full_name = $1 WHERE id = $2', values: [name, id] };\n}\n\nfunction mapRow(row) {\n return { id: row.id, email: row.email, name: row.full_name };\n}\n\nmodule.exports = { buildInsert, buildUpdateName, mapRow };\n",
|
||||
"README.md": "# customers-service\n\nPostgreSQL 15. Migrations: migrations/NNN_name.up.sql and NNN_name.down.sql.\nDeploys are rolling: the previous app version keeps serving traffic (reading\nand writing full_name) until every instance is replaced.\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst fs = require('node:fs');\nconst path = require('node:path');\nconst dir = path.join(process.cwd(), 'migrations');\nconst names = fs.readdirSync(dir);\nconst ups = names.filter(n => /^002_[A-Za-z0-9_-]+\\.up\\.sql$/.test(n));\nassert.equal(ups.length, 1);\nconst stem = ups[0].slice(0, -'.up.sql'.length);\nconst strip = s => s.replace(/--[^\\n]*/g, '').replace(/\\/\\*[\\s\\S]*?\\*\\//g, '');\nconst up = strip(fs.readFileSync(path.join(dir, ups[0]), 'utf8'));\nconst down = strip(fs.readFileSync(path.join(dir, stem + '.down.sql'), 'utf8'));\nconst add = /ALTER\\s+TABLE\\s+(?:IF\\s+EXISTS\\s+)?\"?customers\"?\\s+ADD\\s+(?:COLUMN\\s+)?(?:IF\\s+NOT\\s+EXISTS\\s+)?\"?display_name\"?\\s+(?:text|varchar|character\\s+varying)\\b([^;]*)/i.exec(up);\nassert.ok(add, 'ADD COLUMN display_name missing');\nassert.doesNotMatch(add[1], /NOT\\s+NULL/i);\nassert.match(up, /UPDATE\\s+\"?customers\"?\\s+SET\\s+\"?display_name\"?\\s*=\\s*\"?full_name\"?/i);\nassert.doesNotMatch(up, /RENAME\\s+(?:COLUMN\\s+)?\"?full_name/i);\nassert.doesNotMatch(up, /DROP\\s+(?:COLUMN|TABLE)|DROP\\s+\"?full_name/i);\nassert.match(down, /DROP\\s+(?:COLUMN\\s+)?(?:IF\\s+EXISTS\\s+)?\"?display_name\"?/i);\nassert.doesNotMatch(down, /full_name|DROP\\s+TABLE/i);\nconst repo = require(path.join(process.cwd(), 'src/customerRepo.js'));\nconst maxParam = t => Math.max(0, ...[...t.matchAll(/\\$(\\d+)/g)].map(x => Number(x[1])));\nconst ins = repo.buildInsert({ email: 'a@x.io', name: \"O'Hara\" });\nassert.match(ins.text, /INSERT\\s+INTO\\s+\"?customers\"?/i);\nassert.match(ins.text, /full_name/);\nassert.match(ins.text, /display_name/);\nassert.ok(!ins.text.includes(\"O'Hara\"));\nassert.ok(ins.values.includes(\"O'Hara\") && ins.values.includes('a@x.io'));\nassert.equal(maxParam(ins.text), ins.values.length);\nconst upd = repo.buildUpdateName(7, 'Bo');\nassert.match(upd.text, /UPDATE\\s+\"?customers\"?\\s+SET/i);\nassert.match(upd.text, /full_name\\s*=\\s*\\$\\d+/);\nassert.match(upd.text, /display_name\\s*=\\s*\\$\\d+/);\nassert.match(upd.text, /WHERE\\s+\"?id\"?\\s*=\\s*\\$\\d+/i);\nassert.ok(upd.values.includes('Bo') && upd.values.includes(7));\nassert.equal(maxParam(upd.text), upd.values.length);\nassert.equal(repo.mapRow({ id: 1, email: 'e', full_name: 'Old', display_name: null }).name, 'Old');\nassert.equal(repo.mapRow({ id: 1, email: 'e', full_name: 'Old' }).name, 'Old');\nassert.equal(repo.mapRow({ id: 1, email: 'e', full_name: 'Old', display_name: 'New' }).name, 'New');\nassert.equal(repo.mapRow({ id: 2, email: 'e', full_name: 'Old', display_name: 'New' }).id, 2);\n"
|
||||
},
|
||||
{
|
||||
"id": "keyset-feed-query",
|
||||
"category": "database",
|
||||
"manualIds": [
|
||||
"skill:postgres-patterns"
|
||||
],
|
||||
"query": "src/feedQuery.js builds the PostgreSQL query for a user's post feed using OFFSET, which gets slow and skips rows on deep pages. Switch to keyset (cursor) pagination ordered by created_at DESC, id DESC. Export encodeCursor(row) (row has created_at as an ISO string and id) returning an opaque string, and buildFeedQuery({ userId, limit, cursor }) returning { text, values } for node-postgres ($n placeholders; no caller value inlined into text). cursor is undefined for the first page; otherwise it comes from encodeCursor and the query must return only rows strictly after that row in the sort order. Throw an Error for a malformed cursor and a RangeError unless limit is an integer 1..50. Also add migrations/002_<name>.sql creating a composite index on posts that supports this query (single-file migrations, see 001). Do not add dependencies.",
|
||||
"files": {
|
||||
"src/feedQuery.js": "'use strict';\n\n// page is 0-based\nfunction buildFeedQuery({ userId, limit, page = 0 }) {\n return {\n text: 'SELECT id, user_id, body, created_at FROM posts WHERE user_id = $1 ORDER BY created_at DESC LIMIT $2 OFFSET $3',\n values: [userId, limit, page * limit],\n };\n}\n\nmodule.exports = { buildFeedQuery };\n",
|
||||
"migrations/001_create_posts.sql": "CREATE TABLE posts (\n id bigserial PRIMARY KEY,\n user_id bigint NOT NULL,\n body text NOT NULL,\n created_at timestamptz NOT NULL DEFAULT now()\n);\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst fs = require('node:fs');\nconst path = require('node:path');\nconst { buildFeedQuery, encodeCursor } = require(path.join(process.cwd(), 'src/feedQuery.js'));\nconst maxParam = t => Math.max(0, ...[...t.matchAll(/\\$(\\d+)/g)].map(x => Number(x[1])));\nconst ORDER = /ORDER\\s+BY\\s+\"?created_at\"?\\s+DESC\\s*,\\s*\"?id\"?\\s+DESC/i;\nconst first = buildFeedQuery({ userId: 7, limit: 20 });\nassert.doesNotMatch(first.text, /OFFSET/i);\nassert.match(first.text, ORDER);\nassert.match(first.text, /user_id\\s*=\\s*\\$\\d+/i);\nassert.ok(first.values.includes(7));\nassert.match(first.text, /LIMIT\\s+(\\$\\d+|20)\\b/i);\nassert.equal(maxParam(first.text), first.values.length);\nconst cur = encodeCursor({ id: 42, user_id: 7, body: 'hi', created_at: '2024-05-01T10:00:00.000Z' });\nassert.equal(typeof cur, 'string');\nconst next = buildFeedQuery({ userId: 7, limit: 20, cursor: cur });\nassert.doesNotMatch(next.text, /OFFSET/i);\nassert.match(next.text, ORDER);\nassert.ok(!next.text.includes('2024-05-01') && !/\\b42\\b/.test(next.text));\nconst row = /\\(\\s*\"?created_at\"?\\s*,\\s*\"?id\"?\\s*\\)\\s*<\\s*\\(\\s*\\$(\\d+)(?:::\\w+)?\\s*,\\s*\\$(\\d+)(?:::\\w+)?\\s*\\)/i.exec(next.text);\nconst expanded = /\"?created_at\"?\\s*<\\s*\\$(\\d+)[\\s\\S]*\"?created_at\"?\\s*=\\s*\\$(\\d+)[\\s\\S]*\"?id\"?\\s*<\\s*\\$(\\d+)/i.exec(next.text);\nassert.ok(row || expanded, 'keyset predicate missing: ' + next.text);\nconst vals = next.values.map(v => (v instanceof Date ? v.toISOString() : String(v)));\nassert.ok(vals.includes('2024-05-01T10:00:00.000Z'));\nassert.ok(vals.includes('42'));\nassert.ok(next.values.includes(7));\nassert.equal(maxParam(next.text), next.values.length);\nassert.throws(() => buildFeedQuery({ userId: 7, limit: 20, cursor: 'not-a-cursor' }));\nfor (const bad of [0, 51, '20', 1.5]) assert.throws(() => buildFeedQuery({ userId: 7, limit: bad }), RangeError);\nconst dir = path.join(process.cwd(), 'migrations');\nconst mig = fs.readdirSync(dir).filter(n => /^002_[A-Za-z0-9_-]+\\.sql$/.test(n));\nassert.equal(mig.length, 1);\nconst sql = fs.readFileSync(path.join(dir, mig[0]), 'utf8').replace(/--[^\\n]*/g, '');\nassert.match(sql, /CREATE\\s+(?:UNIQUE\\s+)?INDEX\\s+[\\s\\S]*?ON\\s+(?:ONLY\\s+)?\"?posts\"?\\s*(?:USING\\s+btree\\s*)?\\(\\s*\"?user_id\"?\\s*,\\s*\"?created_at\"?(?:\\s+DESC)?\\s*,\\s*\"?id\"?(?:\\s+DESC)?\\s*\\)/i);\n"
|
||||
},
|
||||
{
|
||||
"id": "upsert-inventory-sql",
|
||||
"category": "database",
|
||||
"manualIds": [
|
||||
"skill:postgres-patterns"
|
||||
],
|
||||
"query": "src/inventory.js exports async syncStock(db, items), where items are { sku, quantity } and db.query(text, values) runs a parameterized PostgreSQL statement (node-postgres style, $n placeholders). It currently does a SELECT and then an UPDATE or INSERT per item, which is slow and races with concurrent syncs. Replace it with a single INSERT INTO inventory (sku, quantity, updated_at) ... ON CONFLICT (sku) DO UPDATE statement for the whole batch that sets quantity from the incoming row and updated_at to now(). Exactly one db.query call per non-empty batch and none for an empty batch. If the same sku appears more than once in items, the last occurrence wins (PostgreSQL rejects affecting a row twice in one statement). No caller value may be inlined into the SQL text. Resolve to the number of distinct skus written. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/inventory.js": "'use strict';\n\nasync function syncStock(db, items) {\n let count = 0;\n for (const item of items) {\n const found = await db.query('SELECT sku FROM inventory WHERE sku = $1', [item.sku]);\n if (found.rows.length) {\n await db.query('UPDATE inventory SET quantity = $1, updated_at = now() WHERE sku = $2', [item.quantity, item.sku]);\n } else {\n await db.query('INSERT INTO inventory (sku, quantity, updated_at) VALUES ($1, $2, now())', [item.sku, item.quantity]);\n }\n count++;\n }\n return count;\n}\n\nmodule.exports = { syncStock };\n",
|
||||
"schema.sql": "CREATE TABLE inventory (\n sku text PRIMARY KEY,\n quantity integer NOT NULL,\n updated_at timestamptz NOT NULL\n);\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { syncStock } = require(path.join(process.cwd(), 'src/inventory.js'));\nfunction fakeDb() {\n const calls = [];\n return { calls, async query(text, values) { calls.push({ text, values }); return { rows: [], rowCount: 0 }; } };\n}\n(async () => {\n let db = fakeDb();\n assert.equal(await syncStock(db, []), 0);\n assert.equal(db.calls.length, 0);\n db = fakeDb();\n const n = await syncStock(db, [{ sku: 'SKU-A', quantity: 11 }, { sku: \"SKU-'B\", quantity: 55 }, { sku: 'SKU-A', quantity: 7 }]);\n assert.equal(n, 2);\n assert.equal(db.calls.length, 1);\n const { text, values } = db.calls[0];\n assert.match(text, /INSERT\\s+INTO\\s+\"?inventory\"?/i);\n assert.match(text, /ON\\s+CONFLICT\\s*\\(\\s*\"?sku\"?\\s*\\)\\s*DO\\s+UPDATE\\s+SET/i);\n assert.match(text, /\"?quantity\"?\\s*=\\s*EXCLUDED\\.\"?quantity\"?/i);\n assert.match(text, /\"?updated_at\"?\\s*=\\s*(?:now\\(\\)|CURRENT_TIMESTAMP|EXCLUDED\\.\"?updated_at\"?)/i);\n assert.ok(!text.includes('SKU-'), 'sku inlined into SQL');\n const flat = values.flat(Infinity).map(v => (typeof v === 'string' && /^\\d+$/.test(v) ? Number(v) : v));\n assert.equal(flat.filter(v => v === 'SKU-A').length, 1);\n assert.equal(flat.filter(v => v === \"SKU-'B\").length, 1);\n assert.ok(flat.includes(7) && flat.includes(55));\n assert.ok(!flat.includes(11), 'stale duplicate quantity sent');\n const maxParam = Math.max(0, ...[...text.matchAll(/\\$(\\d+)/g)].map(x => Number(x[1])));\n assert.equal(maxParam, values.length);\n db = fakeDb();\n assert.equal(await syncStock(db, [{ sku: 'X', quantity: 1 }]), 1);\n assert.equal(db.calls.length, 1);\n})().catch(err => { console.error(err); process.exitCode = 1; });\n"
|
||||
},
|
||||
{
|
||||
"id": "slugify-regression-tests",
|
||||
"category": "testing",
|
||||
"manualIds": [
|
||||
"skill:tdd-workflow"
|
||||
],
|
||||
"query": "Bug report in BUGS.md: src/slugify.js produces leading and trailing hyphens and mangles accented letters. Work test-first: add test/slugify.test.js using the built-in node:test runner and node:assert, requiring ../src/slugify, with at least three separate test cases that reproduce the reported bugs and cover edge cases (empty input, repeated separators), then fix slugify(input) so they pass. Expected behavior: lowercase ASCII output; accented Latin letters lose their accents (e with grave becomes e); every run of non-alphanumeric characters becomes a single hyphen; no leading or trailing hyphens; empty or separator-only input returns an empty string. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/slugify.js": "'use strict';\n\nfunction slugify(input) {\n return String(input).toLowerCase().replace(/[^a-z0-9]+/g, '-');\n}\n\nmodule.exports = { slugify };\n",
|
||||
"BUGS.md": "# Open bugs\n\n1. slugify(' Hello, World! ') returns '-hello-world-' (expected 'hello-world').\n2. slugify('Cr\\u00e8me Br\\u00fbl\\u00e9e') (accented) returns 'cr-me-br-l-e' (expected 'creme-brulee').\n",
|
||||
"package.json": "{\n \"name\": \"slugs\",\n \"version\": \"1.0.0\",\n \"private\": true,\n \"scripts\": { \"test\": \"node --test test/\" }\n}\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst fs = require('node:fs');\nconst path = require('node:path');\nconst { slugify } = require(path.join(process.cwd(), 'src/slugify.js'));\nassert.equal(slugify(' Hello, World! '), 'hello-world');\nassert.equal(slugify('Cr\\u00e8me Br\\u00fbl\\u00e9e'), 'creme-brulee');\nassert.equal(slugify('D\\u00e9j\\u00e0 Vu 2024'), 'deja-vu-2024');\nassert.equal(slugify('a--b__c'), 'a-b-c');\nassert.equal(slugify(''), '');\nassert.equal(slugify(' -- !! '), '');\nassert.equal(slugify('already-slugged'), 'already-slugged');\nconst testFile = path.join(process.cwd(), 'test', 'slugify.test.js');\nassert.ok(fs.existsSync(testFile), 'test/slugify.test.js missing');\nconst src = fs.readFileSync(testFile, 'utf8');\nassert.match(src, /node:test/);\nassert.match(src, /require\\(\\s*['\"]\\.\\.\\/src\\/slugify(?:\\.js)?['\"]\\s*\\)/);\nassert.ok((src.match(/\\b(?:test|it)\\s*\\(/g) || []).length >= 3, 'expected at least three test cases');\n"
|
||||
},
|
||||
{
|
||||
"id": "content-hash-cache",
|
||||
"category": "performance",
|
||||
"manualIds": [
|
||||
"skill:content-hash-cache-pattern"
|
||||
],
|
||||
"query": "src/extractor.js exports createExtractor({ readFile, parse }). readFile(filePath) returns a Buffer and parse(text) is an expensive document parser. The cache is keyed by file path, so edited files return stale results and renamed or copied files are parsed again. Re-key the cache by the SHA-256 hex digest of the file bytes (use node:crypto) so identical content at any path is parsed once and changed content is re-parsed. Also export cacheKeyFor(buffer) returning that hex digest. extract(filePath) must still return the parse result, and stats() must return { hits, misses } counting cache hits and parses. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/extractor.js": "'use strict';\n\nfunction createExtractor({ readFile, parse }) {\n const cache = new Map();\n let hits = 0;\n let misses = 0;\n return {\n extract(filePath) {\n if (cache.has(filePath)) {\n hits++;\n return cache.get(filePath);\n }\n misses++;\n const result = parse(readFile(filePath).toString('utf8'));\n cache.set(filePath, result);\n return result;\n },\n stats: () => ({ hits, misses }),\n };\n}\n\nmodule.exports = { createExtractor };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { createExtractor, cacheKeyFor } = require(path.join(process.cwd(), 'src/extractor.js'));\nassert.equal(cacheKeyFor(Buffer.from('hello')), '2cf24dba5fb0a30e26e83b2ac5b9e29e1b161e5c1fa7425e73043362938b9824');\nassert.notEqual(cacheKeyFor(Buffer.from('a')), cacheKeyFor(Buffer.from('b')));\nconst disk = { 'a.txt': Buffer.from('report one'), 'b.txt': Buffer.from('report one') };\nlet parses = 0;\nconst ex = createExtractor({ readFile: p => Buffer.from(disk[p]), parse: t => { parses++; return { words: t.split(' ').length, text: t }; } });\nassert.deepEqual(ex.extract('a.txt'), { words: 2, text: 'report one' });\nassert.deepEqual(ex.extract('b.txt'), { words: 2, text: 'report one' });\nassert.equal(parses, 1);\ndisk['a.txt'] = Buffer.from('report one edited');\nassert.deepEqual(ex.extract('a.txt'), { words: 3, text: 'report one edited' });\nassert.equal(parses, 2);\nex.extract('a.txt');\nex.extract('b.txt');\nassert.equal(parses, 2);\nassert.deepEqual(ex.stats(), { hits: 3, misses: 2 });\n"
|
||||
},
|
||||
{
|
||||
"id": "batch-customer-lookup",
|
||||
"category": "performance",
|
||||
"manualIds": [
|
||||
"skill:backend-patterns"
|
||||
],
|
||||
"query": "src/orders.js exports async getOrdersWithCustomers(repo) for the orders dashboard endpoint. It calls repo.findCustomerById once per order, which is an N+1 query pattern and times out for large accounts. The repo (see src/repo.js for the interface) also offers findCustomersByIds(ids), which resolves to the matching customers in any order and omits unknown ids. Rewrite the function to load all customers with a single findCustomersByIds call using the distinct customer ids (and no call at all when there are no orders), never calling findCustomerById. Return the orders in their original order, each as a new object with a customer property (null when the customer does not exist). Do not add dependencies.",
|
||||
"files": {
|
||||
"src/orders.js": "'use strict';\n\nasync function getOrdersWithCustomers(repo) {\n const orders = await repo.listOrders();\n const result = [];\n for (const order of orders) {\n const customer = await repo.findCustomerById(order.customerId);\n result.push({ ...order, customer });\n }\n return result;\n}\n\nmodule.exports = { getOrdersWithCustomers };\n",
|
||||
"src/repo.js": "'use strict';\n\n// Interface implemented by the SQL repository in production.\n// listOrders(): Promise<Array<{ id, customerId, total }>>\n// findCustomerById(id): Promise<{ id, name } | null> -- one query per call\n// findCustomersByIds(ids): Promise<Array<{ id, name }>> -- one query, WHERE id = ANY($1)\nmodule.exports = {};\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { getOrdersWithCustomers } = require(path.join(process.cwd(), 'src/orders.js'));\nfunction repo(orders) {\n const customers = [{ id: 'c1', name: 'Ada' }, { id: 'c2', name: 'Lin' }, { id: 'c3', name: 'Bo' }];\n const r = { single: 0, batch: [], async listOrders() { return orders; },\n async findCustomerById(id) { r.single++; return customers.find(c => c.id === id) || null; },\n async findCustomersByIds(ids) { r.batch.push([...ids]); return customers.filter(c => ids.includes(c.id)).reverse(); } };\n return r;\n}\n(async () => {\n const orders = [{ id: 1, customerId: 'c2', total: 5 }, { id: 2, customerId: 'c1', total: 7 },\n { id: 3, customerId: 'c2', total: 1 }, { id: 4, customerId: 'gone', total: 2 }];\n const snapshot = JSON.stringify(orders);\n const r = repo(orders);\n const out = await getOrdersWithCustomers(r);\n assert.equal(r.single, 0);\n assert.equal(r.batch.length, 1);\n assert.deepEqual(r.batch[0].slice().sort(), ['c1', 'c2', 'gone']);\n assert.deepEqual(out.map(o => o.id), [1, 2, 3, 4]);\n assert.deepEqual(out.map(o => o.customer && o.customer.name), ['Lin', 'Ada', 'Lin', null]);\n assert.equal(out[0].total, 5);\n assert.equal(JSON.stringify(orders), snapshot);\n const empty = repo([]);\n assert.deepEqual(await getOrdersWithCustomers(empty), []);\n assert.equal(empty.batch.length + empty.single, 0);\n})().catch(err => { console.error(err); process.exitCode = 1; });\n"
|
||||
},
|
||||
{
|
||||
"id": "rbac-middleware",
|
||||
"category": "auth",
|
||||
"manualIds": [
|
||||
"skill:backend-patterns"
|
||||
],
|
||||
"query": "src/auth.js exports requirePermission(permission), an Express-style middleware factory, and ROLE_PERMISSIONS. It only checks that req.user exists and never checks the role. Implement role-based access control: calling requirePermission with a permission that no role grants must throw immediately. The returned middleware (req, res, next) must respond res.status(401).json({ error: { code: \"UNAUTHENTICATED\", message } }) when req.user is missing; res.status(403).json({ error: { code: \"FORBIDDEN\", message } }) when req.user.role is unknown or lacks the permission (role names must be looked up safely, so values such as \"constructor\" or \"__proto__\" are simply unknown roles); otherwise call next() exactly once without responding. Do not change ROLE_PERMISSIONS. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/auth.js": "'use strict';\n\nconst ROLE_PERMISSIONS = {\n admin: ['read', 'write', 'delete'],\n editor: ['read', 'write'],\n viewer: ['read'],\n};\n\nfunction requirePermission(permission) {\n return (req, res, next) => {\n if (!req.user) return res.status(401).json({ error: 'unauthorized' });\n return next();\n };\n}\n\nmodule.exports = { requirePermission, ROLE_PERMISSIONS };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { requirePermission } = require(path.join(process.cwd(), 'src/auth.js'));\nfunction run(permission, user) {\n const res = { code: null, body: null, status(c) { this.code = c; return this; }, json(b) { this.body = b; return this; } };\n let nexts = 0;\n requirePermission(permission)(user === undefined ? {} : { user }, res, () => { nexts++; });\n return { res, nexts };\n}\nlet r = run('read');\nassert.equal(r.res.code, 401);\nassert.equal(r.res.body.error.code, 'UNAUTHENTICATED');\nassert.equal(typeof r.res.body.error.message, 'string');\nassert.equal(r.nexts, 0);\nr = run('write', { id: 1, role: 'viewer' });\nassert.equal(r.res.code, 403);\nassert.equal(r.res.body.error.code, 'FORBIDDEN');\nassert.equal(r.nexts, 0);\nfor (const role of ['root', 'constructor', '__proto__', 'toString', undefined, 'hasOwnProperty']) {\n let out;\n assert.doesNotThrow(() => { out = run('read', { id: 2, role }); }, String(role));\n assert.equal(out.res.code, 403, String(role));\n assert.equal(out.nexts, 0);\n}\nr = run('write', { id: 3, role: 'editor' });\nassert.equal(r.nexts, 1);\nassert.equal(r.res.code, null);\nr = run('delete', { id: 4, role: 'admin' });\nassert.equal(r.nexts, 1);\nr = run('delete', { id: 5, role: 'editor' });\nassert.equal(r.res.code, 403);\nassert.throws(() => requirePermission('fly'));\nassert.throws(() => requirePermission('constructor'));\n"
|
||||
},
|
||||
{
|
||||
"id": "immutable-cart-update",
|
||||
"category": "refactor",
|
||||
"manualIds": [
|
||||
"skill:coding-standards"
|
||||
],
|
||||
"query": "src/cart.js exports addItem(cart, item), removeItem(cart, sku), applyDiscount(cart, pct) and total(cart). A cart is { items: [{ sku, price, quantity }], discountPct }. The update functions mutate their arguments, which causes stale UI state bugs. Refactor them to be pure: never mutate the cart, its items array, any item object, or the item argument; always return a new cart object. Keep the behavior: addItem adds the item, or increases quantity when the sku already exists; removeItem drops the sku; applyDiscount sets discountPct and must throw a RangeError unless pct is a number from 0 to 100; total returns the discounted sum rounded to 2 decimal places. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/cart.js": "'use strict';\n\nfunction addItem(cart, item) {\n const existing = cart.items.find(i => i.sku === item.sku);\n if (existing) existing.quantity += item.quantity;\n else cart.items.push(item);\n return cart;\n}\n\nfunction removeItem(cart, sku) {\n cart.items = cart.items.filter(i => i.sku !== sku);\n return cart;\n}\n\nfunction applyDiscount(cart, pct) {\n cart.discountPct = pct;\n return cart;\n}\n\nfunction total(cart) {\n const sum = cart.items.reduce((acc, i) => acc + i.price * i.quantity, 0);\n return Math.round(sum * (1 - (cart.discountPct || 0) / 100) * 100) / 100;\n}\n\nmodule.exports = { addItem, removeItem, applyDiscount, total };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst cart = require(path.join(process.cwd(), 'src/cart.js'));\nconst deepFreeze = o => { Object.values(o).forEach(v => { if (v && typeof v === 'object') deepFreeze(v); }); return Object.freeze(o); };\nconst base = deepFreeze({ items: [{ sku: 'a', price: 10, quantity: 1 }, { sku: 'b', price: 2.5, quantity: 2 }], discountPct: 0 });\nconst snap = JSON.stringify(base);\nconst item = deepFreeze({ sku: 'a', price: 10, quantity: 2 });\nconst c1 = cart.addItem(base, item);\nassert.notEqual(c1, base);\nassert.deepEqual(c1.items.find(i => i.sku === 'a').quantity, 3);\nassert.equal(c1.items.length, 2);\nconst newItem = deepFreeze({ sku: 'c', price: 1, quantity: 1 });\nconst c2 = cart.addItem(c1, newItem);\nassert.equal(c2.items.length, 3);\nassert.equal(c1.items.length, 2);\nconst c3 = cart.removeItem(c2, 'b');\nassert.deepEqual(c3.items.map(i => i.sku), ['a', 'c']);\nassert.equal(c2.items.length, 3);\nconst c4 = cart.applyDiscount(c3, 10);\nassert.equal(c4.discountPct, 10);\nassert.equal(c3.discountPct, 0);\nassert.equal(cart.total(c4), 27.9);\nassert.equal(cart.total(base), 15);\nfor (const bad of [-1, 101, '10', NaN]) assert.throws(() => cart.applyDiscount(base, bad), RangeError);\nassert.equal(JSON.stringify(base), snap);\nconst m = { items: [{ sku: 'z', price: 1, quantity: 1 }], discountPct: 0 };\nconst m2 = cart.addItem(m, { sku: 'z', price: 1, quantity: 4 });\nassert.equal(m.items[0].quantity, 1);\nassert.equal(m2.items[0].quantity, 5);\nconst added = { sku: 'y', price: 3, quantity: 1 };\nconst m3 = cart.addItem(m, added);\ncart.addItem(m3, { sku: 'y', price: 3, quantity: 5 });\nassert.equal(added.quantity, 1);\n"
|
||||
},
|
||||
{
|
||||
"id": "inject-signup-deps",
|
||||
"category": "refactor",
|
||||
"manualIds": [
|
||||
"skill:hexagonal-architecture"
|
||||
],
|
||||
"query": "src/signup.js hard-requires the Postgres and SMTP adapters in src/adapters/, which fail at import time without infrastructure, so the sign-up use case cannot be unit tested. Refactor to ports and adapters. src/signup.js must export createSignupService({ userRepository, mailer, clock }) returning { signUp({ email, name }) } and must not import anything from src/adapters or read environment variables. Ports: userRepository.findByEmail(email) and userRepository.save(user) (resolves to the stored user including id), mailer.sendWelcome({ to, name }), clock.now() returning a Date. signUp trims and lowercases the email; rejects with an error whose code is \"INVALID_EMAIL\" if it lacks \"@\", or \"EMAIL_TAKEN\" if findByEmail finds a user (without saving or mailing); otherwise saves { email, name, createdAt: clock.now().toISOString() }, sends the welcome email to the saved user, and resolves to the saved user. Add src/main.js as the composition root that wires the real adapters. Keep the adapters as they are. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/signup.js": "'use strict';\nconst store = require('./adapters/pgUserStore');\nconst mailer = require('./adapters/smtpMailer');\n\nasync function signUp({ email, name }) {\n const normalized = email.trim().toLowerCase();\n if (await store.findByEmail(normalized)) throw new Error('taken');\n const user = await store.insert({ email: normalized, name, createdAt: new Date().toISOString() });\n await mailer.sendWelcome(user.email, user.name);\n return user;\n}\n\nmodule.exports = { signUp };\n",
|
||||
"src/adapters/pgUserStore.js": "'use strict';\n// Connects at import time, like our real pool module.\nif (!process.env.DATABASE_URL) throw new Error('DATABASE_URL is not configured');\n\nmodule.exports = {\n async findByEmail(email) { throw new Error('not implemented in this repo snapshot: ' + email); },\n async insert(user) { throw new Error('not implemented in this repo snapshot: ' + user.email); },\n};\n",
|
||||
"src/adapters/smtpMailer.js": "'use strict';\nif (!process.env.SMTP_URL) throw new Error('SMTP_URL is not configured');\n\nmodule.exports = {\n async sendWelcome(to, name) { throw new Error('not implemented in this repo snapshot: ' + to + name); },\n};\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst fs = require('node:fs');\nconst path = require('node:path');\ndelete process.env.DATABASE_URL;\ndelete process.env.SMTP_URL;\nconst file = path.join(process.cwd(), 'src/signup.js');\nconst source = fs.readFileSync(file, 'utf8');\nassert.doesNotMatch(source, /require\\([^)]*adapters|from\\s+['\"][^'\"]*adapters/, 'domain imports an adapter');\nassert.doesNotMatch(source, /process\\.env/, 'domain reads the environment');\nassert.ok(fs.existsSync(path.join(process.cwd(), 'src/main.js')), 'composition root missing');\nconst { createSignupService } = require(file);\nfunction setup(existing = []) {\n const users = [...existing];\n const log = { saved: [], mails: [] };\n const svc = createSignupService({\n userRepository: { async findByEmail(e) { return users.find(u => u.email === e) || null; },\n async save(u) { const s = { id: 'u' + (users.length + 1), ...u }; users.push(s); log.saved.push(u); return s; } },\n mailer: { async sendWelcome(msg) { log.mails.push(msg); } },\n clock: { now: () => new Date(Date.UTC(2024, 0, 2, 3, 4, 5)) },\n });\n return { svc, log };\n}\n(async () => {\n let { svc, log } = setup();\n const user = await svc.signUp({ email: ' Ada@Example.COM ', name: 'Ada' });\n assert.deepEqual(user, { id: 'u1', email: 'ada@example.com', name: 'Ada', createdAt: '2024-01-02T03:04:05.000Z' });\n assert.deepEqual(log.saved, [{ email: 'ada@example.com', name: 'Ada', createdAt: '2024-01-02T03:04:05.000Z' }]);\n assert.deepEqual(log.mails, [{ to: 'ada@example.com', name: 'Ada' }]);\n ({ svc, log } = setup([{ id: 'x', email: 'lin@example.com', name: 'Lin' }]));\n await assert.rejects(svc.signUp({ email: 'LIN@example.com', name: 'Lin 2' }), e => e.code === 'EMAIL_TAKEN');\n await assert.rejects(svc.signUp({ email: 'nope', name: 'N' }), e => e.code === 'INVALID_EMAIL');\n assert.equal(log.saved.length, 0);\n assert.equal(log.mails.length, 0);\n})().catch(err => { console.error(err); process.exitCode = 1; });\n"
|
||||
},
|
||||
{
|
||||
"id": "cache-aside-user",
|
||||
"category": "caching",
|
||||
"manualIds": [
|
||||
"skill:redis-patterns"
|
||||
],
|
||||
"query": "src/userCache.js exports createUserCache({ redis, db, ttlSeconds = 300 }). redis is a node-redis v4 style client (async get(key), set(key, value, { EX }), del(key)) and db has async findUser(id) and updateUser(id, patch). Profile reads are hammering the database. Implement cache-aside: getUser(id) uses key \"user:\" + id, returns the parsed cached JSON on a hit without touching db, and on a miss loads from db and caches JSON with an expiry of ttlSeconds (do not cache a missing user; return null). updateUser(id, patch) writes to db first, then deletes the cache key, and resolves to the updated user. Redis is an optimization, not a dependency: if any redis call rejects, getUser and updateUser must still return the correct db result. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/userCache.js": "'use strict';\n\nfunction createUserCache({ redis, db, ttlSeconds = 300 }) {\n return {\n async getUser(id) {\n return db.findUser(id);\n },\n async updateUser(id, patch) {\n return db.updateUser(id, patch);\n },\n };\n}\n\nmodule.exports = { createUserCache };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { createUserCache } = require(path.join(process.cwd(), 'src/userCache.js'));\nfunction fakes(broken = false) {\n const store = new Map();\n const log = [];\n const redis = {\n async get(k) { log.push(['get', k]); if (broken) throw new Error('ECONNREFUSED'); return store.has(k) ? store.get(k) : null; },\n async set(k, v, opts) { log.push(['set', k, opts]); if (broken) throw new Error('ECONNREFUSED'); store.set(k, v); return 'OK'; },\n async del(k) { log.push(['del', k]); if (broken) throw new Error('ECONNREFUSED'); return store.delete(k) ? 1 : 0; },\n };\n const rows = { 1: { id: 1, name: 'Ada' } };\n const db = { reads: 0, async findUser(id) { db.reads++; return rows[id] ? { ...rows[id] } : null; },\n async updateUser(id, patch) { log.push(['db-update', id]); rows[id] = { ...rows[id], ...patch }; return { ...rows[id] }; } };\n return { store, log, redis, db };\n}\n(async () => {\n let f = fakes();\n const cache = createUserCache({ redis: f.redis, db: f.db, ttlSeconds: 60 });\n assert.deepEqual(await cache.getUser(1), { id: 1, name: 'Ada' });\n assert.equal(f.db.reads, 1);\n const set = f.log.find(e => e[0] === 'set');\n assert.equal(set[1], 'user:1');\n assert.deepEqual(set[2], { EX: 60 });\n assert.deepEqual(JSON.parse(f.store.get('user:1')), { id: 1, name: 'Ada' });\n assert.deepEqual(await cache.getUser(1), { id: 1, name: 'Ada' });\n assert.equal(f.db.reads, 1);\n assert.equal(await cache.getUser(2), null);\n assert.ok(!f.store.has('user:2'));\n const updated = await cache.updateUser(1, { name: 'Ada L' });\n assert.deepEqual(updated, { id: 1, name: 'Ada L' });\n const iUpd = f.log.findIndex(e => e[0] === 'db-update');\n const iDel = f.log.findIndex(e => e[0] === 'del' && e[1] === 'user:1');\n assert.ok(iUpd >= 0 && iDel > iUpd, 'must invalidate after the db write');\n assert.deepEqual(await cache.getUser(1), { id: 1, name: 'Ada L' });\n f = fakes();\n const dflt = createUserCache({ redis: f.redis, db: f.db });\n await dflt.getUser(1);\n assert.deepEqual(f.log.find(e => e[0] === 'set')[2], { EX: 300 });\n f = fakes(true);\n const broken = createUserCache({ redis: f.redis, db: f.db });\n assert.deepEqual(await broken.getUser(1), { id: 1, name: 'Ada' });\n assert.deepEqual(await broken.updateUser(1, { name: 'X' }), { id: 1, name: 'X' });\n})().catch(err => { console.error(err); process.exitCode = 1; });\n"
|
||||
},
|
||||
{
|
||||
"id": "token-units-bigint",
|
||||
"category": "data",
|
||||
"manualIds": [
|
||||
"skill:evm-token-decimals"
|
||||
],
|
||||
"query": "src/units.js converts ERC-20 token amounts for our portfolio dashboard, but it uses floating point, so 18-decimal balances lose precision. Rewrite it with exact BigInt math. formatUnits(raw, decimals): raw is a bigint or an integer string in base units; return a decimal string with no trailing fractional zeros and no trailing \".\", keeping a leading \"-\" for negatives. parseUnits(value, decimals): value is a decimal string such as \"1.5\" or \"-0.25\"; return a bigint in base units; throw a RangeError if it has more fractional digits than decimals, and throw an Error for anything that is not a plain decimal number (e.g. \"\", \"abc\", \"1e5\", \"1.2.3\"). Also export normalizeAmount(raw, fromDecimals, toDecimals) returning a bigint rescaled between token precisions, truncating toward zero when precision is reduced. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/units.js": "'use strict';\n\nfunction formatUnits(raw, decimals) {\n return String(Number(raw) / 10 ** decimals);\n}\n\nfunction parseUnits(value, decimals) {\n return BigInt(Math.round(parseFloat(value) * 10 ** decimals));\n}\n\nmodule.exports = { formatUnits, parseUnits };\n",
|
||||
"README.md": "# portfolio-units\n\nToken decimals differ per token and per chain: USDC uses 6 on Ethereum mainnet,\nWETH uses 18, and some bridged tokens differ from their native versions.\nAlways pass the decimals value read from the token contract.\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { formatUnits, parseUnits, normalizeAmount } = require(path.join(process.cwd(), 'src/units.js'));\nassert.equal(formatUnits(123456789012345678901234567n, 18), '123456789.012345678901234567');\nassert.equal(formatUnits('1000000', 6), '1');\nassert.equal(formatUnits(1500000n, 6), '1.5');\nassert.equal(formatUnits(0n, 18), '0');\nassert.equal(formatUnits(-1n, 18), '-0.000000000000000001');\nassert.equal(formatUnits(-1500000n, 6), '-1.5');\nassert.equal(formatUnits(5n, 0), '5');\nassert.equal(parseUnits('1.5', 6), 1500000n);\nassert.equal(parseUnits('0.000000000000000001', 18), 1n);\nassert.equal(parseUnits('123456789.012345678901234567', 18), 123456789012345678901234567n);\nassert.equal(parseUnits('-0.25', 6), -250000n);\nassert.equal(parseUnits('100', 0), 100n);\nassert.throws(() => parseUnits('1.1234567', 6), RangeError);\nfor (const bad of ['', 'abc', '1e5', '1.2.3', '0x10', ' 1']) assert.throws(() => parseUnits(bad, 6), Error, bad);\nassert.equal(normalizeAmount(1234567n, 6, 18), 1234567000000000000n);\nassert.equal(normalizeAmount(1234567890123456789n, 18, 6), 1234567n);\nassert.equal(normalizeAmount(-1234567890123456789n, 18, 6), -1234567n);\nassert.equal(normalizeAmount(42n, 8, 8), 42n);\nassert.equal(typeof normalizeAmount(1n, 6, 6), 'bigint');\n"
|
||||
},
|
||||
{
|
||||
"id": "inclusive-range",
|
||||
"category": "no-workflow",
|
||||
"manualIds": [],
|
||||
"query": "range(start, end) in src/range.js is documented as inclusive of end, but it stops one short. Fix it so range(1, 5) returns [1, 2, 3, 4, 5]; when start > end it must return an empty array. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/range.js": "'use strict';\n\n/** Returns the integers from start to end, inclusive. */\nfunction range(start, end) {\n const out = [];\n for (let i = start; i < end; i++) out.push(i);\n return out;\n}\n\nmodule.exports = { range };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { range } = require(path.join(process.cwd(), 'src/range.js'));\nassert.deepEqual(range(1, 5), [1, 2, 3, 4, 5]);\nassert.deepEqual(range(3, 3), [3]);\nassert.deepEqual(range(-2, 0), [-2, -1, 0]);\nassert.deepEqual(range(5, 1), []);\n"
|
||||
},
|
||||
{
|
||||
"id": "export-name-typo",
|
||||
"category": "no-workflow",
|
||||
"manualIds": [],
|
||||
"query": "src/report.js crashes with \"formatDate is not a function\" because src/dates.js exports its formatter under a misspelled name. Export it as formatDate, and keep the misspelled export as an alias of the same function so older callers keep working. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/dates.js": "'use strict';\n\nfunction formatDate(date) {\n const pad = n => String(n).padStart(2, '0');\n return date.getUTCFullYear() + '-' + pad(date.getUTCMonth() + 1) + '-' + pad(date.getUTCDate());\n}\n\nmodule.exports = { fromatDate: formatDate };\n",
|
||||
"src/report.js": "'use strict';\nconst { formatDate } = require('./dates');\n\nfunction reportHeader(title, date) {\n return title + ' (' + formatDate(date) + ')';\n}\n\nmodule.exports = { reportHeader };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst dates = require(path.join(process.cwd(), 'src/dates.js'));\nconst { reportHeader } = require(path.join(process.cwd(), 'src/report.js'));\nconst d = new Date(Date.UTC(2024, 0, 5, 12));\nassert.equal(dates.formatDate(d), '2024-01-05');\nassert.equal(dates.fromatDate, dates.formatDate);\nassert.equal(reportHeader('Weekly', d), 'Weekly (2024-01-05)');\n"
|
||||
},
|
||||
{
|
||||
"id": "default-greeting",
|
||||
"category": "no-workflow",
|
||||
"manualIds": [],
|
||||
"noWorkflow": true,
|
||||
"query": "Small fix, no workflow needed. greet(name) in src/greet.js returns \"Hello, undefined!\" when called without a name. Make it trim the name and fall back to \"world\" when the name is missing, null, empty or only whitespace, so greet() returns \"Hello, world!\" and greet(\" Ada \") returns \"Hello, Ada!\". Do not add dependencies.",
|
||||
"files": {
|
||||
"src/greet.js": "'use strict';\n\nfunction greet(name) {\n return 'Hello, ' + name + '!';\n}\n\nmodule.exports = { greet };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { greet } = require(path.join(process.cwd(), 'src/greet.js'));\nassert.equal(greet(), 'Hello, world!');\nassert.equal(greet(null), 'Hello, world!');\nassert.equal(greet(''), 'Hello, world!');\nassert.equal(greet(' '), 'Hello, world!');\nassert.equal(greet(' Ada '), 'Hello, Ada!');\nassert.equal(greet('Lin'), 'Hello, Lin!');\n"
|
||||
},
|
||||
{
|
||||
"id": "sum-form-values",
|
||||
"category": "no-workflow",
|
||||
"manualIds": [],
|
||||
"query": "total(values) in src/total.js sums amounts typed into a form, but the inputs arrive as strings so it returns \"0123.5\" for [\"1\", \"2\", \"3.5\"]. Make it return the numeric sum (6.5 in that example). Empty strings count as 0, plain numbers must still work, and an empty array returns 0. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/total.js": "'use strict';\n\nfunction total(values) {\n return values.reduce((sum, v) => sum + v, 0);\n}\n\nmodule.exports = { total };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { total } = require(path.join(process.cwd(), 'src/total.js'));\nassert.equal(total(['1', '2', '3.5']), 6.5);\nassert.equal(total([]), 0);\nassert.equal(total(['', '4']), 4);\nassert.equal(total([2, '3']), 5);\n"
|
||||
},
|
||||
{
|
||||
"id": "changelog-capitalize",
|
||||
"category": "no-workflow",
|
||||
"manualIds": [],
|
||||
"query": "The security team's release-notes script imports src/changelog.js, and it crashes when a changelog entry has an empty title because capitalize(\"\") throws. Fix capitalize so an empty string returns \"\", while other strings still get only their first character uppercased with the rest unchanged. formatEntry must keep its current output format. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/changelog.js": "'use strict';\n\nfunction capitalize(text) {\n return text[0].toUpperCase() + text.slice(1);\n}\n\nfunction formatEntry(entry) {\n return '- ' + capitalize(entry.title) + ' (' + entry.type + ')';\n}\n\nmodule.exports = { capitalize, formatEntry };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { capitalize, formatEntry } = require(path.join(process.cwd(), 'src/changelog.js'));\nassert.equal(capitalize(''), '');\nassert.equal(capitalize('x'), 'X');\nassert.equal(capitalize('hello World'), 'Hello World');\nassert.equal(formatEntry({ title: 'fix xss in footer', type: 'security' }), '- Fix xss in footer (security)');\nassert.equal(formatEntry({ title: '', type: 'chore' }), '- (chore)');\n"
|
||||
},
|
||||
{
|
||||
"id": "test-summary-plural",
|
||||
"category": "no-workflow",
|
||||
"manualIds": [],
|
||||
"query": "Our test runner prints \"1 tests passed, 1 tests failed\". In src/summary.js, fix formatSummary(passed, failed) to use \"test\" when a count is exactly 1 and \"tests\" otherwise, e.g. \"1 test passed, 0 tests failed\". Keep the rest of the wording identical. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/summary.js": "'use strict';\n\nfunction formatSummary(passed, failed) {\n return passed + ' tests passed, ' + failed + ' tests failed';\n}\n\nmodule.exports = { formatSummary };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { formatSummary } = require(path.join(process.cwd(), 'src/summary.js'));\nassert.equal(formatSummary(1, 0), '1 test passed, 0 tests failed');\nassert.equal(formatSummary(2, 1), '2 tests passed, 1 test failed');\nassert.equal(formatSummary(0, 0), '0 tests passed, 0 tests failed');\nassert.equal(formatSummary(12, 3), '12 tests passed, 3 tests failed');\n"
|
||||
},
|
||||
{
|
||||
"id": "database-label-typo",
|
||||
"category": "no-workflow",
|
||||
"manualIds": [],
|
||||
"noWorkflow": true,
|
||||
"query": "No workflow needed. In src/options.js the settings dropdown shows \"Databse\" for the database option; correct the label to \"Database\". Also make labelFor(value) return the value itself when no option matches, instead of throwing. Do not change the option values or their order. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/options.js": "'use strict';\n\nconst OPTIONS = [\n { value: 'database', label: 'Databse' },\n { value: 'api', label: 'API' },\n { value: 'cache', label: 'Cache' },\n];\n\nfunction labelFor(value) {\n return OPTIONS.find(o => o.value === value).label;\n}\n\nmodule.exports = { OPTIONS, labelFor };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { OPTIONS, labelFor } = require(path.join(process.cwd(), 'src/options.js'));\nassert.deepEqual(OPTIONS, [{ value: 'database', label: 'Database' }, { value: 'api', label: 'API' }, { value: 'cache', label: 'Cache' }]);\nassert.equal(labelFor('database'), 'Database');\nassert.equal(labelFor('api'), 'API');\nassert.equal(labelFor('queue'), 'queue');\n"
|
||||
},
|
||||
{
|
||||
"id": "port-from-env",
|
||||
"category": "no-workflow",
|
||||
"manualIds": [],
|
||||
"noWorkflow": true,
|
||||
"query": "Do not select a workflow for this one-line style fix. getPort(env) in src/server-config.js returns env.PORT as a string or 3000. Make it return a number: the integer value of env.PORT when it consists only of decimal digits and is between 1 and 65535, otherwise 3000. Do not add dependencies.",
|
||||
"files": {
|
||||
"src/server-config.js": "'use strict';\n\nfunction getPort(env = process.env) {\n return env.PORT || 3000;\n}\n\nmodule.exports = { getPort };\n"
|
||||
},
|
||||
"check": "'use strict';\nconst assert = require('node:assert/strict');\nconst path = require('node:path');\nconst { getPort } = require(path.join(process.cwd(), 'src/server-config.js'));\nassert.equal(getPort({ PORT: '8080' }), 8080);\nassert.equal(getPort({}), 3000);\nassert.equal(getPort({ PORT: '' }), 3000);\nassert.equal(getPort({ PORT: 'abc' }), 3000);\nassert.equal(getPort({ PORT: '70000' }), 3000);\nassert.equal(getPort({ PORT: '0' }), 3000);\nassert.equal(getPort({ PORT: '80.5' }), 3000);\nassert.equal(getPort({ PORT: '65535' }), 65535);\n"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,847 @@
|
||||
'use strict';
|
||||
|
||||
// Development-only evaluator. It lives under docker/ so the npm package never ships it.
|
||||
const fs = require('node:fs');
|
||||
const os = require('node:os');
|
||||
const path = require('node:path');
|
||||
const { spawnSync } = require('node:child_process');
|
||||
const { isDeepStrictEqual } = require('node:util');
|
||||
const LIB = path.join(__dirname, '../../scripts/lib');
|
||||
const { loadContextRegistry } = require(path.join(LIB, 'context-pack-registry'));
|
||||
const { compileContextProfile } = require(path.join(LIB, 'context-profiles'));
|
||||
const { resolveTaskContext, resolveDeclinedFallback } = require(path.join(LIB, 'context-selection'));
|
||||
const { proposeTaskContext } = require(path.join(LIB, 'context-profile-proposal'));
|
||||
const { resolveExecutable, fingerprintExecutable } = require(path.join(LIB, 'context-profile-native-executable'));
|
||||
const { launchTaskContext } = require(path.join(LIB, 'context-profile-launch'));
|
||||
const { applyStore } = require(path.join(LIB, 'context-profile-store'));
|
||||
const { prepareNativeProfile, getNativeProfileStatus } = require(path.join(LIB, 'context-profile-native'));
|
||||
const { DEFAULT_REPO_ROOT, digestObject, createSourceReader } = require(path.join(LIB, 'context-profile-support'));
|
||||
const io = require(path.join(LIB, 'context-profile-store-fs'));
|
||||
|
||||
const ARMS = Object.freeze(['full', 'manual-lean', 'auto-lean', 'ecc-legacy', 'baseline']);
|
||||
const CORPUS_PATH = path.join(__dirname, 'ai-corpus.json');
|
||||
const LEGACY_PIN_PATH = path.join(__dirname, 'legacy-source.json');
|
||||
const CHECK_FILE = '.ecc-eval-check.cjs';
|
||||
const IMPLEMENTATION = ['docker/context-profiles/ai-eval-lib.js', 'docker/context-profiles/ai-eval.js',
|
||||
'docker/context-profiles/legacy-source.json',
|
||||
'manifests/context-packs/skill-triggers@1.json',
|
||||
'scripts/lib/context-profile-launch.js', 'scripts/lib/context-selection.js',
|
||||
'scripts/lib/context-retrieval.js',
|
||||
'scripts/lib/context-profile-proposal.js', 'scripts/lib/context-profiles.js',
|
||||
'scripts/lib/context-profile-support.js', 'scripts/lib/context-pack-registry.js',
|
||||
'scripts/lib/context-profile-native-executable.js', 'scripts/lib/context-profile-native.js',
|
||||
'scripts/lib/context-profile-store.js', 'scripts/lib/context-profile-store-fs.js'];
|
||||
const BLOCKS = Object.freeze({ excluded: /Context ID is excluded:/,
|
||||
'native-authority': /requires native authority or dynamic-content review/,
|
||||
'manual-only': /Context ID is manual-only:/, 'opt-out-conflict': /noWorkflow conflicts/, 'unknown-id': /Unknown context ID:/ });
|
||||
const ENV_KEYS = ['PATH', 'HOME', 'USERPROFILE', 'CODEX_HOME', 'TMPDIR', 'LANG', 'SystemRoot'];
|
||||
const CLAUDE_ENV_KEYS = ['PATH', 'HOME', 'USERPROFILE', 'CLAUDE_CONFIG_DIR', 'TMPDIR', 'LANG', 'SystemRoot'];
|
||||
const bounded = (value, min, max) => Number.isSafeInteger(value) && value >= min && value <= max;
|
||||
const exists = file => Boolean(fs.lstatSync(file, { throwIfNoEntry: false }));
|
||||
|
||||
function loadCorpus(file = CORPUS_PATH) { return JSON.parse(fs.readFileSync(file, 'utf8')); }
|
||||
|
||||
function safeRelative(file) {
|
||||
return typeof file === 'string' && file.length > 0 && file.length <= 200 && !path.isAbsolute(file)
|
||||
&& !file.startsWith('.') && !file.includes('\\') && file.split('/').every(part => part && part !== '..' && part !== '.');
|
||||
}
|
||||
|
||||
function validateCorpus(corpus) {
|
||||
if (corpus?.schemaVersion === 'ecc.context-eval-complex-corpus.v1') return validateComplexCorpus(corpus);
|
||||
if (corpus?.schemaVersion !== 'ecc.context-eval-corpus.v2'
|
||||
|| !Array.isArray(corpus.selection) || !Array.isArray(corpus.tasks)
|
||||
|| !bounded(corpus.selection.length, 1, 200) || !bounded(corpus.tasks.length, 1, 200)
|
||||
|| corpus.minimumDistinctTasks !== 30 || corpus.nonInferiorityMargin !== 0.05) {
|
||||
throw new Error('Invalid preregistered corpus');
|
||||
}
|
||||
for (const cases of [corpus.selection, corpus.tasks]) validateCorpusIds(cases);
|
||||
for (const task of corpus.tasks) {
|
||||
const files = Object.entries(task.files || {});
|
||||
if (!Array.isArray(task.manualIds) || task.manualIds.length > 1 || !bounded(files.length, 1, 8)
|
||||
|| files.some(([file, content]) => !safeRelative(file) || typeof content !== 'string' || Buffer.byteLength(content) > 16384)
|
||||
|| typeof task.check !== 'string' || !bounded(Buffer.byteLength(task.check), 1, 16384)) {
|
||||
throw new Error('Invalid corpus task');
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
function validateCorpusIds(cases) {
|
||||
if (new Set(cases.map(c => c.id)).size !== cases.length) throw new Error('Duplicate corpus ID');
|
||||
for (const item of cases) {
|
||||
if (!/^[a-z][a-z0-9-]{0,63}$/.test(item.id) || typeof item.query !== 'string'
|
||||
|| !bounded(Buffer.byteLength(item.query), 1, 8192)) throw new Error('Invalid corpus case');
|
||||
}
|
||||
}
|
||||
|
||||
// Complex corpora hold a few realistic multi-file tasks with scored hidden graders. Sample gates
|
||||
// are descriptive at this size, so the distinct-task minimum relaxes to the corpus itself.
|
||||
function validateComplexCorpus(corpus) {
|
||||
if (!Array.isArray(corpus.selection) || !Array.isArray(corpus.tasks)
|
||||
|| !bounded(corpus.selection.length, 0, 50) || !bounded(corpus.tasks.length, 1, 10)
|
||||
|| corpus.minimumDistinctTasks !== corpus.tasks.length || corpus.nonInferiorityMargin !== 0.05) {
|
||||
throw new Error('Invalid preregistered corpus');
|
||||
}
|
||||
validateCorpusIds(corpus.selection);
|
||||
if (new Set(corpus.tasks.map(c => c.id)).size !== corpus.tasks.length) throw new Error('Duplicate corpus ID');
|
||||
for (const task of corpus.tasks) {
|
||||
if (!/^[a-z][a-z0-9-]{0,63}$/.test(task.id)) throw new Error('Invalid corpus case');
|
||||
if (task.steps === undefined
|
||||
&& (typeof task.query !== 'string' || !bounded(Buffer.byteLength(task.query), 1, 8192))) throw new Error('Invalid corpus case');
|
||||
const files = Object.entries(task.files || {});
|
||||
if (!Array.isArray(task.manualIds) || task.manualIds.length > 3 || !bounded(files.length, 1, 24)
|
||||
|| files.some(([file, content]) => !safeRelative(file) || typeof content !== 'string' || Buffer.byteLength(content) > 65536)) {
|
||||
throw new Error('Invalid corpus task');
|
||||
}
|
||||
if (task.steps !== undefined) {
|
||||
// Stepped (chained) task: sequential tickets graded in one accumulating workspace.
|
||||
if (!Array.isArray(task.steps) || !bounded(task.steps.length, 2, 8)
|
||||
|| task.steps.some(step => typeof step.query !== 'string' || !bounded(Buffer.byteLength(step.query), 1, 8192)
|
||||
|| typeof step.check !== 'string' || !bounded(Buffer.byteLength(step.check), 1, 65536)
|
||||
|| (step.checkTimeoutMs !== undefined && !bounded(step.checkTimeoutMs, 1, 120000))
|
||||
|| (step.manualIds !== undefined && (!Array.isArray(step.manualIds) || step.manualIds.length > 3)))) {
|
||||
throw new Error('Invalid corpus task');
|
||||
}
|
||||
} else if (typeof task.check !== 'string' || !bounded(Buffer.byteLength(task.check), 1, 65536)
|
||||
|| (task.checkTimeoutMs !== undefined && !bounded(task.checkTimeoutMs, 1, 120000))) {
|
||||
throw new Error('Invalid corpus task');
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
function sourceSnapshot(repoRoot) {
|
||||
const registry = loadContextRegistry({ repoRoot });
|
||||
const profiles = ['full@1', 'lean@1'].map(profileId => compileContextProfile({ repoRoot, profileId }));
|
||||
// Implementation modules are loaded from this evaluator's checkout; repoRoot may be a fixture registry.
|
||||
const reader = createSourceReader(DEFAULT_REPO_ROOT);
|
||||
const implementation = IMPLEMENTATION.map(file => ({ path: file, digest: reader.read(file).digest }));
|
||||
const packageJson = JSON.parse(reader.read('package.json').content.toString('utf8'));
|
||||
const runtime = { node: process.versions.node, dependencies: {
|
||||
ajv: packageJson.dependencies.ajv, 'js-yaml': packageJson.dependencies['js-yaml'] } };
|
||||
return { registry, profiles, sourceDigest: digestObject({ registryDigest: registry.registryDigest,
|
||||
planDigests: profiles.map(p => p.planDigest), implementation, runtime }), runtime };
|
||||
}
|
||||
|
||||
const EFFORTS = ['low', 'medium', 'high', 'xhigh', 'max', 'ultra'];
|
||||
|
||||
function providerFamily(executable) {
|
||||
const base = path.basename(String(executable || '')).toLowerCase();
|
||||
if (base.includes('claude')) return 'claude';
|
||||
if (base.includes('codex')) return 'codex';
|
||||
throw new Error('Provider executable must name a Claude or Codex CLI');
|
||||
}
|
||||
|
||||
function resolveFamily(provider, executable) {
|
||||
if (provider !== undefined && provider !== null) {
|
||||
if (!['claude', 'codex'].includes(provider)) throw new Error('Provider must be claude or codex');
|
||||
return provider;
|
||||
}
|
||||
if (executable) return providerFamily(executable);
|
||||
return 'codex';
|
||||
}
|
||||
|
||||
function providerPin(model, executable, effort) {
|
||||
if (model === undefined && executable === undefined && effort === undefined) return null;
|
||||
if (typeof model !== 'string' || !/^[a-zA-Z0-9][a-zA-Z0-9._:-]{0,99}$/.test(model)
|
||||
|| !path.isAbsolute(executable || '')) throw new Error('Provider pin requires model and absolute executable');
|
||||
if (effort !== undefined && !EFFORTS.includes(effort)) throw new Error('Invalid reasoning effort');
|
||||
return { modelDigest: digestObject(model), executableDigest: resolveExecutable(executable).digest,
|
||||
...(effort === undefined ? {} : { effort }) };
|
||||
}
|
||||
|
||||
function preregister({ repoRoot = DEFAULT_REPO_ROOT, corpus = loadCorpus(), repeats = 1, model, executable, effort, arms } = {}) {
|
||||
validateCorpus(corpus);
|
||||
if (!bounded(repeats, 1, 20)) throw new Error('Invalid repeat count');
|
||||
const armList = arms === undefined ? [...ARMS] : arms;
|
||||
if (!Array.isArray(armList) || !armList.length || new Set(armList).size !== armList.length
|
||||
|| armList.some(arm => !ARMS.includes(arm))) throw new Error('Invalid arm subset');
|
||||
const source = sourceSnapshot(repoRoot);
|
||||
const value = { schemaVersion: 'ecc.context-eval-registration.v2', corpusDigest: digestObject(corpus),
|
||||
sourceDigest: source.sourceDigest, registryDigest: source.registry.registryDigest,
|
||||
providerPin: providerPin(model, executable, effort), runtime: source.runtime,
|
||||
arms: armList, repeats, minimumDistinctTasks: corpus.minimumDistinctTasks, nonInferiorityMargin: 0.05,
|
||||
confidence: 0.95, sampling: 'fixed-purposive-pilot',
|
||||
design: corpus.schemaVersion === 'ecc.context-eval-complex-corpus.v1'
|
||||
? 'paired-native-installs-hidden-scored-complex-tasks'
|
||||
: 'paired-native-installs-hidden-graded-coding-tasks',
|
||||
order: corpus.tasks.flatMap((task, index) => Array.from({ length: repeats }, (_, repeat) => ({
|
||||
id: task.id, repeat, arms: armList.map((_, offset) => armList[(index + repeat + offset) % armList.length]),
|
||||
}))), selectionIds: corpus.selection.map(c => c.id) };
|
||||
return { ...value, registrationDigest: digestObject(value) };
|
||||
}
|
||||
|
||||
// Parse in memory only. No event objects, paths, provider messages or error text enter reports.
|
||||
function parseCodexJsonl(stdout) {
|
||||
const invalid = { valid: false, text: '', usage: null };
|
||||
if (typeof stdout !== 'string' || Buffer.byteLength(stdout) > 1024 * 1024) return invalid;
|
||||
let text = '';
|
||||
let completions = 0;
|
||||
let usage = { inputTokens: 0, cachedInputTokens: 0, outputTokens: 0 };
|
||||
try {
|
||||
for (const line of stdout.split('\n').filter(line => line.trim())) {
|
||||
const event = JSON.parse(line);
|
||||
if (!event || typeof event !== 'object' || ['error', 'turn.failed'].includes(event.type)) return invalid;
|
||||
if (event.type === 'item.completed' && event.item?.type === 'agent_message') {
|
||||
if (typeof event.item.text !== 'string') return invalid;
|
||||
text = event.item.text;
|
||||
}
|
||||
if (event.type !== 'turn.completed') continue;
|
||||
const u = event.usage;
|
||||
if (!u || ![u.input_tokens, u.cached_input_tokens, u.output_tokens].every(v => bounded(v, 0, 1e9))
|
||||
|| u.cached_input_tokens > u.input_tokens) return invalid;
|
||||
completions++;
|
||||
usage = { inputTokens: usage.inputTokens + u.input_tokens,
|
||||
cachedInputTokens: usage.cachedInputTokens + u.cached_input_tokens,
|
||||
outputTokens: usage.outputTokens + u.output_tokens };
|
||||
}
|
||||
} catch { return invalid; }
|
||||
return completions === 1 ? { valid: true, text, usage } : invalid;
|
||||
}
|
||||
|
||||
// Claude print-mode emits exactly one result JSON object. Fresh input folds cache creations;
|
||||
// cache reads are reported separately. is_error results are provider failures, not parse failures.
|
||||
function parseClaudeJson(stdout) {
|
||||
const invalid = { valid: false, text: '', usage: null };
|
||||
if (typeof stdout !== 'string' || Buffer.byteLength(stdout) > 1024 * 1024) return invalid;
|
||||
let result = null;
|
||||
let results = 0;
|
||||
try {
|
||||
for (const line of stdout.split('\n').filter(line => line.trim())) {
|
||||
const event = JSON.parse(line);
|
||||
if (!event || typeof event !== 'object' || Array.isArray(event)) return invalid;
|
||||
if (event.type !== 'result') continue;
|
||||
results++;
|
||||
result = event;
|
||||
}
|
||||
} catch { return invalid; }
|
||||
if (results !== 1) return invalid;
|
||||
if (result.is_error !== false || typeof result.result !== 'string') return { ...invalid, error: true };
|
||||
const u = result.usage;
|
||||
if (!u || ![u.input_tokens, u.cache_creation_input_tokens, u.cache_read_input_tokens, u.output_tokens]
|
||||
.every(value => bounded(value, 0, 1e9))) return { ...invalid, error: true };
|
||||
return { valid: true, text: result.result,
|
||||
usage: { inputTokens: u.input_tokens + u.cache_creation_input_tokens,
|
||||
cachedInputTokens: u.cache_read_input_tokens, outputTokens: u.output_tokens } };
|
||||
}
|
||||
|
||||
function privateEntry(file, directory) {
|
||||
const stat = fs.lstatSync(file, { throwIfNoEntry: false });
|
||||
return Boolean(stat) && !stat.isSymbolicLink() && (directory ? stat.isDirectory() : stat.isFile())
|
||||
&& (process.platform === 'win32' || ((stat.mode & 0o077) === 0 && (!process.getuid || stat.uid === process.getuid())));
|
||||
}
|
||||
|
||||
/**
|
||||
* Subscription credentials stay in a dedicated evaluator login home. Each call leases auth.json into the
|
||||
* isolated CODEX_HOME, returns refreshed tokens afterwards and always removes the leased copy.
|
||||
*/
|
||||
function createAuthLease(authHome) {
|
||||
if (typeof authHome !== 'string' || !path.isAbsolute(authHome)) throw new Error('Auth home must be an absolute path');
|
||||
const real = fs.realpathSync(authHome);
|
||||
const forbidden = [path.join(os.homedir(), '.codex'), process.env.CODEX_HOME].filter(Boolean)
|
||||
.map(file => (exists(file) ? fs.realpathSync(file) : path.resolve(file)));
|
||||
if (forbidden.includes(real)) throw new Error('Auth home must be a dedicated evaluator login home, not your Codex home');
|
||||
const source = path.join(real, 'auth.json');
|
||||
if (!privateEntry(real, true) || !privateEntry(source, false)) {
|
||||
throw new Error('Auth home must be a private directory containing a private auth.json; see the evaluation guide');
|
||||
}
|
||||
return {
|
||||
mode: 'subscription-lease',
|
||||
run(codexHome, work) {
|
||||
const leased = path.join(codexHome, 'auth.json');
|
||||
const original = fs.readFileSync(source);
|
||||
fs.writeFileSync(leased, original, { flag: 'wx', mode: 0o600 });
|
||||
try { return work(); } finally {
|
||||
try {
|
||||
const after = fs.readFileSync(leased);
|
||||
if (!after.equals(original)) {
|
||||
JSON.parse(after.toString('utf8'));
|
||||
const temp = `${source}.${process.pid}.tmp`;
|
||||
try {
|
||||
fs.writeFileSync(temp, after, { flag: 'wx', mode: 0o600 });
|
||||
fs.renameSync(temp, source);
|
||||
} finally { fs.rmSync(temp, { force: true }); }
|
||||
}
|
||||
} catch { /* An unreadable refresh keeps the previous login; the next call reports any auth failure. */ }
|
||||
fs.rmSync(leased, { force: true });
|
||||
}
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* Claude subscription logins live in the macOS Keychain as a JSON wrapper. The lease reads the
|
||||
* current access token per call into the child environment only; it is never persisted or reported.
|
||||
*/
|
||||
function readClaudeKeychainToken() {
|
||||
if (process.platform !== 'darwin') throw new Error('Claude Keychain login requires macOS; provide CLAUDE_CODE_OAUTH_TOKEN or ANTHROPIC_API_KEY');
|
||||
const result = spawnSync('security', ['find-generic-password', '-s', 'Claude Code-credentials', '-w'],
|
||||
{ encoding: 'utf8', shell: false, timeout: 15000, killSignal: 'SIGKILL', maxBuffer: 65536 });
|
||||
if (result.status !== 0 || result.error) throw new Error('Claude Keychain login is unavailable; provide CLAUDE_CODE_OAUTH_TOKEN or ANTHROPIC_API_KEY');
|
||||
let parsed;
|
||||
try { parsed = JSON.parse(result.stdout); }
|
||||
catch { throw new Error('Claude Keychain login is unreadable; provide CLAUDE_CODE_OAUTH_TOKEN or ANTHROPIC_API_KEY'); }
|
||||
const token = parsed?.claudeAiOauth?.accessToken;
|
||||
if (typeof token !== 'string' || !token) throw new Error('Claude Keychain login is unrecognized; provide CLAUDE_CODE_OAUTH_TOKEN or ANTHROPIC_API_KEY');
|
||||
return token;
|
||||
}
|
||||
|
||||
function createClaudeProvider({ allowRealProvider = false, allowCredentialedTools = false, executable, model,
|
||||
apiKey = process.env.ANTHROPIC_API_KEY, oauthToken = process.env.CLAUDE_CODE_OAUTH_TOKEN,
|
||||
tokenSource = readClaudeKeychainToken, persistSessions = false, execute = spawnSync } = {}) {
|
||||
if (allowRealProvider !== true) throw new Error('Real provider requires explicit opt-in');
|
||||
if (!model || !executable) throw new Error('Real provider requires a model and absolute executable');
|
||||
let lease = null;
|
||||
let authentication;
|
||||
if (oauthToken) authentication = 'oauth-env';
|
||||
else if (apiKey) authentication = 'api-key';
|
||||
else if (typeof tokenSource === 'function') {
|
||||
lease = { mode: 'subscription-keychain-lease',
|
||||
run(env, work) { env.CLAUDE_CODE_OAUTH_TOKEN = tokenSource(); return work(); } };
|
||||
authentication = lease.mode;
|
||||
} else throw new Error('Real provider requires CLAUDE_CODE_OAUTH_TOKEN, ANTHROPIC_API_KEY, or the Claude Keychain login');
|
||||
const pin = providerPin(model, executable, undefined);
|
||||
const binary = resolveExecutable(executable);
|
||||
const provider = request => {
|
||||
if (fingerprintExecutable(binary.path).digest !== pin.executableDigest) fail('source-drift');
|
||||
const selection = request.phase === 'selection';
|
||||
if (!selection && !allowCredentialedTools) {
|
||||
throw new Error('Claude task tools can read provider credentials; explicit credentialed-tool opt-in is required');
|
||||
}
|
||||
// Selection is tool-free and read-only; task execution may edit and run commands in the workspace.
|
||||
// Claude has no cwd-write sandbox flag, so containment relies on the isolated home and temp workspace.
|
||||
const args = ['--print', '--output-format', 'json',
|
||||
...(persistSessions ? [] : ['--no-session-persistence']),
|
||||
...(selection ? ['--tools', ''] : ['--permission-mode', 'bypassPermissions']),
|
||||
'--model', model];
|
||||
const env = Object.fromEntries(CLAUDE_ENV_KEYS.filter(key => typeof request.env?.[key] === 'string')
|
||||
.map(key => [key, request.env[key]]));
|
||||
env.DISABLE_NON_ESSENTIAL_MODEL_CALLS = '1';
|
||||
if (authentication === 'oauth-env') env.CLAUDE_CODE_OAUTH_TOKEN = oauthToken;
|
||||
if (authentication === 'api-key') env.ANTHROPIC_API_KEY = apiKey;
|
||||
const call = () => execute(binary.path, args, { input: request.input, cwd: request.cwd, env,
|
||||
encoding: 'utf8', shell: false, timeout: request.timeoutMs, killSignal: 'SIGKILL',
|
||||
maxBuffer: request.maxBuffer });
|
||||
return lease ? lease.run(env, call) : call();
|
||||
};
|
||||
provider.authentication = authentication;
|
||||
return provider;
|
||||
}
|
||||
|
||||
function createCodexProvider({ allowRealProvider = false, executable, model, effort, authHome,
|
||||
apiKey = process.env.CODEX_API_KEY, execute = spawnSync } = {}) {
|
||||
if (allowRealProvider !== true) throw new Error('Real provider requires explicit opt-in');
|
||||
if (!model || !executable) throw new Error('Real provider requires a model and absolute executable');
|
||||
if (!authHome && !apiKey) throw new Error('Real provider requires --auth-home (subscription login) or CODEX_API_KEY');
|
||||
const lease = authHome ? createAuthLease(authHome) : null;
|
||||
const pin = providerPin(model, executable, effort);
|
||||
const binary = resolveExecutable(executable);
|
||||
const provider = request => {
|
||||
if (fingerprintExecutable(binary.path).digest !== pin.executableDigest) fail('source-drift');
|
||||
const args = ['exec', '--json', '--ephemeral', '--skip-git-repo-check',
|
||||
'--sandbox', request.phase === 'selection' ? 'read-only' : 'workspace-write',
|
||||
// Connected ChatGPT apps and account plugin installs stay out of every arm.
|
||||
'--disable', 'apps', '--disable', 'remote_plugin',
|
||||
'-c', 'approval_policy="never"', ...(effort ? ['-c', `model_reasoning_effort="${effort}"`] : []),
|
||||
'--model', model, '-'];
|
||||
const env = Object.fromEntries(ENV_KEYS.filter(key => typeof request.env?.[key] === 'string')
|
||||
.map(key => [key, request.env[key]]));
|
||||
if (!lease) env.CODEX_API_KEY = apiKey;
|
||||
const call = () => execute(binary.path, args, { input: request.input, cwd: request.cwd, env,
|
||||
encoding: 'utf8', shell: false, timeout: request.timeoutMs, killSignal: 'SIGKILL',
|
||||
maxBuffer: request.maxBuffer });
|
||||
return lease ? lease.run(env.CODEX_HOME, call) : call();
|
||||
};
|
||||
provider.authentication = lease ? lease.mode : 'api-key';
|
||||
return provider;
|
||||
}
|
||||
|
||||
/** Real Lean and Full installs, prepared through the same isolated native adapter users get. */
|
||||
function prepareEnvironments({ repoRoot, executable, root }) {
|
||||
const binary = resolveExecutable(executable);
|
||||
const environments = {};
|
||||
for (const [name, profileId, selectionMode] of [['full', 'full@1', 'manual'], ['lean', 'lean@1', 'auto']]) {
|
||||
const options = { stateRoot: path.join(root, name, 'managed'), nativeRoot: path.join(root, name, 'native') };
|
||||
fs.mkdirSync(path.join(root, name), { mode: 0o700 });
|
||||
applyStore({ repoRoot, stateRoot: options.stateRoot, target: 'codex', selectionMode, profileId });
|
||||
const status = prepareNativeProfile({ ...options, codexPath: executable });
|
||||
if (!status.ready) throw new Error(`Native ${name} install is not ready`);
|
||||
// A signed-in Codex records task-directory trust in config.toml and downloads account-provided
|
||||
// plugins into plugins/. Restoring the prepared state after every call keeps trials identical;
|
||||
// any other change still fails verification as drift.
|
||||
const config = path.join(status.codexHome, 'config.toml');
|
||||
const prepared = fs.readFileSync(config);
|
||||
const plugins = path.join(status.codexHome, 'plugins');
|
||||
const listing = directory => (exists(directory) ? fs.readdirSync(directory) : []);
|
||||
const preparedPlugins = new Set(listing(plugins));
|
||||
const preparedCache = new Set(listing(path.join(plugins, 'cache')));
|
||||
environments[name] = { profileId, skills: status.selectedIds.length,
|
||||
launch: { home: status.home, codexHome: status.codexHome, codexPath: status.codexPath,
|
||||
executableDigest: status.executableDigest },
|
||||
restore() {
|
||||
fs.writeFileSync(config, prepared);
|
||||
for (const entry of listing(plugins)) if (!preparedPlugins.has(entry)) fs.rmSync(path.join(plugins, entry), { recursive: true, force: true });
|
||||
for (const entry of listing(path.join(plugins, 'cache'))) {
|
||||
if (!preparedCache.has(entry)) fs.rmSync(path.join(plugins, 'cache', entry), { recursive: true, force: true });
|
||||
}
|
||||
},
|
||||
verify() {
|
||||
let ready = false;
|
||||
try { ready = getNativeProfileStatus(options).ready; } catch { ready = false; }
|
||||
if (!ready) fail('environment-drift');
|
||||
} };
|
||||
}
|
||||
// Baseline arm: an empty native home with no ECC install, for provider-overhead subtraction.
|
||||
const home = path.join(root, 'baseline', 'home');
|
||||
fs.mkdirSync(path.join(home, '.codex'), { recursive: true, mode: 0o700 });
|
||||
environments.baseline = { profileId: null, skills: 0, restore() {},
|
||||
launch: { home, codexHome: path.join(home, '.codex'), codexPath: binary.path, executableDigest: binary.digest },
|
||||
verify() { if (fingerprintExecutable(binary.path).digest !== binary.digest) fail('environment-drift'); } };
|
||||
return environments;
|
||||
}
|
||||
|
||||
function installClaudeSkills({ payload, home }) {
|
||||
const config = path.join(home, '.claude');
|
||||
const installed = path.join(config, 'skills');
|
||||
fs.mkdirSync(installed, { recursive: true, mode: 0o700 });
|
||||
for (const entry of fs.readdirSync(payload)) {
|
||||
fs.cpSync(path.join(payload, entry), path.join(installed, entry), { recursive: true, errorOnExist: true, force: false });
|
||||
}
|
||||
return { config, installed };
|
||||
}
|
||||
|
||||
function claudeEnvironment({ name, binary, home, config, installed, profileId, skills, sourceSha = null }) {
|
||||
const managed = () => digestObject(io.inventory(installed));
|
||||
const prepared = managed();
|
||||
return [name, { profileId, skills, sourceSha,
|
||||
launch: { home, claudeConfigDir: config, claudePath: binary.path, executableDigest: binary.digest },
|
||||
restore() {},
|
||||
verify() {
|
||||
if (fingerprintExecutable(binary.path).digest !== binary.digest) fail('environment-drift');
|
||||
let observed = null;
|
||||
try { observed = managed(); } catch { observed = null; }
|
||||
if (observed !== prepared) fail('environment-drift');
|
||||
} }];
|
||||
}
|
||||
|
||||
/** The pre-scoping ECC source, pinned by commit so the ecc-legacy arm is reproducible. */
|
||||
function exportLegacySource({ repoRoot = DEFAULT_REPO_ROOT, destination,
|
||||
pin = JSON.parse(fs.readFileSync(LEGACY_PIN_PATH, 'utf8')) } = {}) {
|
||||
if (!/^[a-f0-9]{40}$/.test(pin?.sha || '')) throw new Error('Invalid legacy source pin');
|
||||
if (!path.isAbsolute(destination || '')) throw new Error('Legacy destination must be absolute');
|
||||
const resolved = spawnSync('git', ['-C', repoRoot, 'rev-parse', '--verify', `${pin.sha}^{commit}`],
|
||||
{ encoding: 'utf8', shell: false, timeout: 30000, killSignal: 'SIGKILL' });
|
||||
if (resolved.status !== 0 || resolved.error || resolved.stdout.trim() !== pin.sha) {
|
||||
throw new Error('Legacy source pin is unavailable in this repository');
|
||||
}
|
||||
fs.mkdirSync(destination, { recursive: true, mode: 0o700 });
|
||||
const tar = path.join(destination, 'legacy.tar');
|
||||
const archive = spawnSync('git', ['-C', repoRoot, 'archive', '--format=tar', '-o', tar, pin.sha, 'skills'],
|
||||
{ encoding: 'utf8', shell: false, timeout: 60000, killSignal: 'SIGKILL' });
|
||||
const extract = archive.status === 0 && !archive.error
|
||||
? spawnSync('tar', ['-xf', tar, '-C', destination], { encoding: 'utf8', shell: false, timeout: 60000, killSignal: 'SIGKILL' })
|
||||
: archive;
|
||||
fs.rmSync(tar, { force: true });
|
||||
const payload = path.join(destination, 'skills');
|
||||
if (extract.status !== 0 || extract.error || !exists(payload) || !fs.readdirSync(payload).length) {
|
||||
throw new Error('Legacy source export failed');
|
||||
}
|
||||
return { root: destination, sha: pin.sha };
|
||||
}
|
||||
|
||||
/** Real Claude installs in isolated config homes. Managed-skill drift aborts; there is no
|
||||
* provider bookkeeping to restore because isolated Claude runs do not mutate the managed tree. */
|
||||
function prepareClaudeEnvironments({ repoRoot, executable, root, legacySource = null }) {
|
||||
const binary = resolveExecutable(executable);
|
||||
const environments = {};
|
||||
for (const [name, profileId, selectionMode] of [['full', 'full@1', 'manual'], ['lean', 'lean@1', 'auto']]) {
|
||||
const stateRoot = path.join(root, name, 'managed');
|
||||
fs.mkdirSync(path.join(root, name), { mode: 0o700 });
|
||||
const status = applyStore({ repoRoot, stateRoot, target: 'claude', selectionMode, profileId });
|
||||
const home = path.join(root, name, 'home');
|
||||
const { config, installed } = installClaudeSkills({ payload: path.join(status.generationRoot, 'skills'), home });
|
||||
const [key, env] = claudeEnvironment({ name, binary, home, config, installed, profileId, skills: status.selectedIds.length });
|
||||
environments[key] = env;
|
||||
}
|
||||
if (legacySource) {
|
||||
// ecc-legacy: the typical pre-scoping install — the full skill library from the pinned
|
||||
// pre-ECC-029 commit, launched bare with no ECC context block.
|
||||
const home = path.join(root, 'ecc-legacy', 'home');
|
||||
const { config, installed } = installClaudeSkills({ payload: path.join(legacySource.root, 'skills'), home });
|
||||
const [key, env] = claudeEnvironment({ name: 'ecc-legacy', binary, home, config, installed,
|
||||
profileId: null, skills: fs.readdirSync(installed).length, sourceSha: legacySource.sha });
|
||||
environments[key] = env;
|
||||
}
|
||||
// Baseline arm: an empty config home with no ECC install, for provider-overhead subtraction.
|
||||
const baselineHome = path.join(root, 'baseline', 'home');
|
||||
const baselineConfig = path.join(baselineHome, '.claude');
|
||||
fs.mkdirSync(baselineConfig, { recursive: true, mode: 0o700 });
|
||||
environments.baseline = { profileId: null, skills: 0, sourceSha: null, restore() {},
|
||||
launch: { home: baselineHome, claudeConfigDir: baselineConfig, claudePath: binary.path, executableDigest: binary.digest },
|
||||
verify() { if (fingerprintExecutable(binary.path).digest !== binary.digest) fail('environment-drift'); } };
|
||||
return environments;
|
||||
}
|
||||
|
||||
function syntheticEnvironments(root) {
|
||||
const executable = resolveExecutable(process.execPath);
|
||||
return Object.fromEntries(['full', 'lean', 'ecc-legacy', 'baseline'].map(name => {
|
||||
const home = path.join(root, name, 'home');
|
||||
fs.mkdirSync(path.join(home, '.codex'), { recursive: true, mode: 0o700 });
|
||||
return [name, { profileId: ['baseline', 'ecc-legacy'].includes(name) ? null : `${name}@1`, skills: null, sourceSha: null,
|
||||
verify() {}, restore() {},
|
||||
launch: { home, codexHome: path.join(home, '.codex'), codexPath: executable.path, executableDigest: executable.digest } }];
|
||||
}));
|
||||
}
|
||||
|
||||
function checkArguments(cwd, file = CHECK_FILE, writable = false) {
|
||||
const major = Number(process.versions.node.split('.')[0]);
|
||||
const flag = major >= 22 ? '--permission' : major >= 20 ? '--experimental-permission' : null;
|
||||
// A directory grant covers its children. Node 20.20.2 can abort in its native
|
||||
// permission radix tree when the same directory is also granted as "cwd/*".
|
||||
return flag ? [flag, `--allow-fs-read=${cwd}`,
|
||||
// Stepped graders exercise stateful apps (persistence); single-step graders stay read-only.
|
||||
...(writable ? [`--allow-fs-write=${cwd}`] : []), file] : [file];
|
||||
}
|
||||
|
||||
// The hidden grader enters the workspace only after the agent exits, and runs read-only where Node supports it.
|
||||
// A grader may print one `ECC_EVAL_SCORE {"score":0..1}` line for partial credit; without it the exit
|
||||
// status alone decides (exit 0 scores 1). Outcome success still requires a full score. Stepped tasks
|
||||
// grade each step with a distinct grader file so earlier graders stay readable in the workspace.
|
||||
const SCORE_LINE = /^\s*ECC_EVAL_SCORE\s+(\{[^\n]*\})\s*$/m;
|
||||
function runScoredCheck(cwd, source, timeoutMs = 10000, step = null) {
|
||||
const name = step === null ? CHECK_FILE : `.ecc-eval-check-${step}.cjs`;
|
||||
const file = path.join(cwd, name);
|
||||
if (exists(file)) return { passed: false, score: 0 };
|
||||
fs.writeFileSync(file, source, { flag: 'wx' });
|
||||
const result = spawnSync(process.execPath, checkArguments(fs.realpathSync(cwd), name, step !== null), { cwd, encoding: 'utf8',
|
||||
env: { LANG: 'C.UTF-8' }, shell: false, timeout: timeoutMs, killSignal: 'SIGKILL', maxBuffer: 65536 });
|
||||
// Grader files never linger: in stepped tasks the workspace accumulates, and a later ticket's
|
||||
// agent could read or replay an earlier grader. The planted-grader guard above still applies.
|
||||
fs.rmSync(file, { force: true });
|
||||
const passed = result.status === 0 && !result.error;
|
||||
let score = passed ? 1 : 0;
|
||||
const match = SCORE_LINE.exec(result.stdout || '');
|
||||
// A grader that advertises ECC_EVAL_SCORE but never printed it died mid-run (e.g. the graded
|
||||
// server crashed the process): that is a zero, never a silent pass. A printed but malformed
|
||||
// line keeps the exit-status score.
|
||||
const graderDied = passed && !match && source.includes('ECC_EVAL_SCORE')
|
||||
&& !(result.stdout || '').includes('ECC_EVAL_SCORE');
|
||||
if (passed && match) {
|
||||
try {
|
||||
const parsed = JSON.parse(match[1]);
|
||||
if (typeof parsed?.score === 'number' && parsed.score >= 0 && parsed.score <= 1) score = parsed.score;
|
||||
} catch { /* A malformed score line keeps the exit-status score. */ }
|
||||
}
|
||||
if (graderDied) score = 0;
|
||||
return { passed, score };
|
||||
}
|
||||
|
||||
function runCheck(cwd, source) { return runScoredCheck(cwd, source).passed; }
|
||||
|
||||
function writeWorkspace(cwd, files) {
|
||||
for (const [relative, content] of Object.entries(files)) {
|
||||
fs.mkdirSync(path.dirname(path.join(cwd, relative)), { recursive: true });
|
||||
fs.writeFileSync(path.join(cwd, relative), content, { flag: 'wx' });
|
||||
}
|
||||
}
|
||||
|
||||
function wilson(successes, n) {
|
||||
if (!n) return [0, 1];
|
||||
const z = 1.959963984540054;
|
||||
const p = successes / n;
|
||||
const denominator = 1 + z * z / n;
|
||||
const center = (p + z * z / (2 * n)) / denominator;
|
||||
const radius = z * Math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / denominator;
|
||||
return [Math.max(0, center - radius), Math.min(1, center + radius)];
|
||||
}
|
||||
|
||||
function summarize(outcomes, arms = ARMS) {
|
||||
const ids = [...new Set(outcomes.map(row => row.id))];
|
||||
// Reference arm: full when present (all-arms runs), otherwise the last registered arm (baseline in subset runs).
|
||||
const reference = arms.includes('full') ? 'full' : arms[arms.length - 1];
|
||||
const rates = arms.map(arm => {
|
||||
const rows = outcomes.filter(row => row.arm === arm);
|
||||
return { arm, attempts: rows.length, successes: rows.filter(row => row.passed).length,
|
||||
rate: rows.length ? rows.filter(row => row.passed).length / rows.length : null,
|
||||
meanScore: rows.length ? rows.reduce((sum, row) => sum + (typeof row.score === 'number' ? row.score : Number(row.passed)), 0) / rows.length : null };
|
||||
});
|
||||
const pairs = arms.filter(arm => arm !== reference).map(arm => {
|
||||
const differences = ids.map(id => {
|
||||
const rows = outcomes.filter(row => row.id === id);
|
||||
const baseline = rows.filter(row => row.arm === reference);
|
||||
const delta = baseline.map(row => Number(rows.find(r => r.arm === arm && r.repeat === row.repeat)?.passed === true)
|
||||
- Number(row.passed === true));
|
||||
return delta.length ? delta.reduce((a, b) => a + b, 0) / delta.length : null;
|
||||
}).filter(value => value !== null);
|
||||
const n = differences.length;
|
||||
const delta = n ? differences.reduce((a, b) => a + b, 0) / n : null;
|
||||
// Paired task-cluster means in [-1,1]. Hoeffding with Bonferroni for the arm comparisons.
|
||||
const radius = n ? Math.sqrt(2 * Math.log(80) / n) : 2;
|
||||
return { arm, reference, n, delta, interval: [Math.max(-1, (delta || 0) - radius), Math.min(1, (delta || 0) + radius)],
|
||||
method: 'paired-task-cluster-hoeffding-familywise-95' };
|
||||
});
|
||||
return { distinctTasks: ids.length, rates, pairs };
|
||||
}
|
||||
|
||||
function selectionTask(item) {
|
||||
return { sessionId: 'ecc-eval', taskId: item.id, revision: 1, phase: 'evaluate', query: item.query,
|
||||
...(item.noWorkflow === undefined ? {} : { noWorkflow: item.noWorkflow }),
|
||||
...(item.explicitIds ? { explicitIds: item.explicitIds } : {}) };
|
||||
}
|
||||
|
||||
function failureCode(error) {
|
||||
if (['call-budget', 'deadline', 'source-drift', 'environment-drift', 'provider-failed', 'invalid-jsonl'].includes(error?.code)) return error.code;
|
||||
for (const [code, pattern] of Object.entries(BLOCKS)) if (pattern.test(error?.message || '')) return code;
|
||||
return 'evaluation-failed';
|
||||
}
|
||||
function fail(code) { const error = new Error(code); error.code = code; throw error; }
|
||||
|
||||
function launchEnvironment(launch) {
|
||||
return { PATH: process.env.PATH, HOME: launch.home,
|
||||
...(launch.codexHome ? { CODEX_HOME: launch.codexHome } : {}),
|
||||
...(launch.claudeConfigDir ? { CLAUDE_CONFIG_DIR: launch.claudeConfigDir } : {}),
|
||||
TMPDIR: launch.home, LANG: 'C.UTF-8' };
|
||||
}
|
||||
|
||||
function executeAdapter(state, cwd, environment) {
|
||||
return (_command, args, options) => {
|
||||
if (state.calls >= state.maxCalls) fail('call-budget');
|
||||
state.assertCurrent();
|
||||
environment.verify();
|
||||
const remaining = state.deadline - Date.now();
|
||||
if (remaining <= 0) fail('deadline');
|
||||
const phase = options.phase || (args.includes('read-only') ? 'selection' : 'task');
|
||||
state.calls++;
|
||||
const started = Date.now();
|
||||
let raw;
|
||||
// Coding tasks outgrow the launcher's interactive default, so the evaluator's own call bound governs them.
|
||||
const timeoutMs = Math.min(phase === 'task' ? state.callTimeoutMs : options.timeout, state.callTimeoutMs, remaining);
|
||||
const env = options.env || launchEnvironment(environment.launch);
|
||||
try {
|
||||
raw = state.provider({ phase, input: options.input, cwd, env, timeoutMs, maxBuffer: 1024 * 1024 });
|
||||
} catch (error) {
|
||||
state.metrics.push({ phase, elapsedMs: Date.now() - started, usage: null });
|
||||
if (error?.code === 'source-drift') throw error;
|
||||
fail('provider-failed');
|
||||
} finally { environment.restore(); }
|
||||
const elapsedMs = Date.now() - started;
|
||||
const parsed = state.family === 'claude' ? parseClaudeJson(raw?.stdout) : parseCodexJsonl(raw?.stdout);
|
||||
state.metrics.push({ phase, elapsedMs, usage: parsed.valid && raw?.status === 0 && !raw?.error ? parsed.usage : null });
|
||||
if (Date.now() >= state.deadline || elapsedMs > timeoutMs) fail('deadline');
|
||||
state.assertCurrent();
|
||||
if (raw?.status !== 0 || raw?.error) fail('provider-failed');
|
||||
if (!parsed.valid) fail(parsed.error ? 'provider-failed' : 'invalid-jsonl');
|
||||
return { status: 0, stdout: parsed.text };
|
||||
};
|
||||
}
|
||||
|
||||
function selectionProbe(item, repoRoot, execute, environment, target) {
|
||||
const options = { repoRoot, task: selectionTask(item), exclude: item.exclude || [], load: true };
|
||||
try {
|
||||
let selection = resolveTaskContext(options);
|
||||
if (selection.reason === 'agent-selection-required') {
|
||||
const proposedIds = proposeTaskContext({ target, query: item.query, candidates: selection.candidates, execute,
|
||||
executable: environment.launch.codexPath || environment.launch.claudePath });
|
||||
// An empty proposal is an explicit decline: honor it (inject nothing).
|
||||
// The tier-2 fallback only applies when a non-empty proposal admitted
|
||||
// nothing — never to override a decline.
|
||||
const declined = proposedIds.length === 0;
|
||||
const next = resolveTaskContext({ ...options, task: { ...options.task, proposedIds, noWorkflow: declined } });
|
||||
if (next.selectedIds.length) selection = next;
|
||||
else if (declined) selection = { ...next, reason: 'agent-declined-selection' };
|
||||
else selection = resolveDeclinedFallback(options, selection);
|
||||
}
|
||||
return { id: item.id, category: item.category, passed: !item.expectedBlock
|
||||
&& isDeepStrictEqual(selection.selectedIds, item.expectedIds), selectedIds: selection.selectedIds, failure: null };
|
||||
} catch (error) {
|
||||
const failure = failureCode(error);
|
||||
return { id: item.id, category: item.category, passed: Boolean(item.expectedBlock && failure === item.expectedBlock),
|
||||
selectedIds: [], failure };
|
||||
}
|
||||
}
|
||||
|
||||
// Full relies on native discovery of the whole install; the Lean arms receive ECC-selected skill bodies;
|
||||
// ecc-legacy runs bare against the pinned pre-scoping skill library; Baseline runs the bare task query.
|
||||
// Stepped tasks run each ticket in the same accumulating workspace, grading after every step.
|
||||
function outcomeTrial(item, arm, repeat, repoRoot, execute, cwd, environment, target, harvest, metrics = null) {
|
||||
const launchStep = (query, manualIds) => {
|
||||
const task = { sessionId: 'ecc-eval', taskId: item.id, revision: 1, phase: 'evaluate', query };
|
||||
return launchTaskContext({ repoRoot, execute, nativeEnvironment: environment.launch, target,
|
||||
bare: arm === 'baseline' || arm === 'ecc-legacy',
|
||||
task: { ...task, ...(arm === 'manual-lean' && manualIds?.length ? { explicitIds: manualIds } : {}) },
|
||||
profileId: arm === 'full' ? 'full@1' : 'lean@1', selectionMode: arm === 'auto-lean' ? 'auto' : 'manual' });
|
||||
};
|
||||
try {
|
||||
if (!item.steps) {
|
||||
const result = launchStep(item.query, item.manualIds);
|
||||
if (harvest) harvest(arm, item.id, repeat, environment);
|
||||
const verdict = runScoredCheck(cwd, item.check, item.checkTimeoutMs);
|
||||
const passed = result.status === 'completed' && verdict.passed && verdict.score >= 0.999;
|
||||
return { id: item.id, arm, repeat, passed, score: result.status === 'completed' ? verdict.score : 0,
|
||||
selectedIds: result.selection.selectedIds, failure: passed ? null : 'hidden-check' };
|
||||
}
|
||||
const steps = [];
|
||||
const selectedIds = [];
|
||||
for (let index = 0; index < item.steps.length; index++) {
|
||||
const step = item.steps[index];
|
||||
const start = metrics ? metrics.length : 0;
|
||||
const result = launchStep(step.query, step.manualIds || item.manualIds);
|
||||
if (harvest) harvest(arm, `${item.id}--step${index + 1}`, repeat, environment);
|
||||
if (result.status !== 'completed') {
|
||||
// A failed ticket ends the chain; remaining tickets are unscored.
|
||||
steps.push({ score: 0, ...(metrics ? metricsSince(metrics, start) : {}) });
|
||||
for (let rest = index + 1; rest < item.steps.length; rest++) {
|
||||
steps.push({ score: 0, ...(metrics ? metricsSince(metrics, metrics.length) : {}) });
|
||||
}
|
||||
break;
|
||||
}
|
||||
selectedIds.push(...result.selection.selectedIds);
|
||||
const verdict = runScoredCheck(cwd, step.check, step.checkTimeoutMs, index + 1);
|
||||
steps.push({ score: verdict.passed ? verdict.score : 0, ...(metrics ? metricsSince(metrics, start) : {}) });
|
||||
}
|
||||
const score = steps.reduce((sum, step) => sum + step.score, 0) / item.steps.length;
|
||||
const passed = steps.length === item.steps.length && steps.every(step => step.score >= 0.999);
|
||||
return { id: item.id, arm, repeat, passed, score, selectedIds: [...new Set(selectedIds)], steps,
|
||||
failure: passed ? null : 'hidden-check' };
|
||||
} catch (error) {
|
||||
if (harvest) harvest(arm, item.id, repeat, environment);
|
||||
return { id: item.id, arm, repeat, passed: false, score: 0, selectedIds: [], failure: failureCode(error) };
|
||||
}
|
||||
}
|
||||
|
||||
function metricsSince(metrics, start) {
|
||||
const calls = metrics.slice(start);
|
||||
const complete = calls.length > 0 && calls.every(call => call.usage !== null);
|
||||
return { calls: calls.length, elapsedMs: calls.reduce((sum, c) => sum + c.elapsedMs, 0),
|
||||
usage: complete ? calls.reduce((sum, c) => ({ inputTokens: sum.inputTokens + c.usage.inputTokens,
|
||||
cachedInputTokens: sum.cachedInputTokens + c.usage.cachedInputTokens,
|
||||
outputTokens: sum.outputTokens + c.usage.outputTokens }), { inputTokens: 0, cachedInputTokens: 0, outputTokens: 0 }) : null };
|
||||
}
|
||||
|
||||
// Transcript retention is opt-in (--artifact-dir) and file-only: reports never embed session content or paths.
|
||||
function createHarvester(artifactDir, envs) {
|
||||
if (typeof artifactDir !== 'string' || !path.isAbsolute(artifactDir)) throw new Error('Artifact directory must be absolute');
|
||||
fs.mkdirSync(artifactDir, { recursive: true });
|
||||
const sessionsOf = env => {
|
||||
const config = env.launch.claudeConfigDir;
|
||||
const projects = config ? path.join(config, 'projects') : null;
|
||||
if (!projects || !exists(projects)) return new Set();
|
||||
const found = new Set();
|
||||
const walk = directory => {
|
||||
for (const entry of fs.readdirSync(directory, { withFileTypes: true })) {
|
||||
const item = path.join(directory, entry.name);
|
||||
if (entry.isDirectory()) walk(item);
|
||||
else if (entry.name.endsWith('.jsonl')) found.add(item);
|
||||
}
|
||||
};
|
||||
walk(projects);
|
||||
return found;
|
||||
};
|
||||
const seen = new Map(Object.entries(envs).map(([name, env]) => [name, sessionsOf(env)]));
|
||||
const index = [];
|
||||
return {
|
||||
record(arm, id, repeat, env) {
|
||||
const before = seen.get(arm) || new Set();
|
||||
const now = sessionsOf(env);
|
||||
seen.set(arm, now);
|
||||
const fresh = [...now].filter(file => !before.has(file));
|
||||
if (!fresh.length) return;
|
||||
const directory = path.join(artifactDir, `${id}--${arm}--${repeat}`);
|
||||
fs.mkdirSync(directory, { recursive: true });
|
||||
for (const file of fresh) fs.copyFileSync(file, path.join(directory, path.basename(file)));
|
||||
index.push({ id, arm, repeat, files: fresh.map(file => path.basename(file)) });
|
||||
},
|
||||
writeIndex() { fs.writeFileSync(path.join(artifactDir, 'artifact-index.json'), `${JSON.stringify(index, null, 1)}\n`); },
|
||||
};
|
||||
}
|
||||
|
||||
function runEvaluation({ repoRoot = DEFAULT_REPO_ROOT, corpus = loadCorpus(), registration,
|
||||
repeats = 1, provider, family, allowRealProvider = false, allowCredentialedTools = false,
|
||||
executable, model, effort, authHome, environments,
|
||||
arms = undefined, artifactDir = null, maxCalls = 300, deadlineMs = 3600000, callTimeoutMs = 300000 } = {}) {
|
||||
if (!provider && !allowRealProvider) throw new Error('Evaluation requires an injected provider or explicit opt-in');
|
||||
if (!bounded(maxCalls, 1, 2000) || !bounded(deadlineMs, 1, 8 * 3600000)
|
||||
|| !bounded(callTimeoutMs, 1, 600000)) throw new Error('Invalid call or deadline bound');
|
||||
if (!provider && !registration) throw new Error('Real evaluation requires prior registration');
|
||||
const resolvedFamily = provider ? (family || 'codex') : resolveFamily(family, executable);
|
||||
if (resolvedFamily === 'claude' && effort !== undefined) throw new Error('Reasoning effort applies only to the Codex provider');
|
||||
if (!provider && resolvedFamily === 'claude' && !allowCredentialedTools) {
|
||||
throw new Error('Claude task tools can read provider credentials; explicit credentialed-tool opt-in is required');
|
||||
}
|
||||
const pin = preregister({ repoRoot, corpus, repeats, model, executable, effort, arms });
|
||||
if (!provider && resolvedFamily === 'codex' && pin.arms.includes('ecc-legacy')) {
|
||||
throw new Error('Codex real evaluation requires --arms without ecc-legacy; the pinned legacy skills arm is Claude-only');
|
||||
}
|
||||
if (registration && !isDeepStrictEqual(registration, pin)) throw new Error('Registration pin mismatch');
|
||||
const injected = Boolean(provider);
|
||||
const liveProvider = provider || (resolvedFamily === 'claude'
|
||||
? createClaudeProvider({ allowRealProvider, allowCredentialedTools, executable, model,
|
||||
persistSessions: Boolean(artifactDir) })
|
||||
: createCodexProvider({ allowRealProvider, executable, model, effort, authHome }));
|
||||
const state = { calls: 0, metrics: [], maxCalls, callTimeoutMs, family: resolvedFamily,
|
||||
deadline: Date.now() + deadlineMs, provider: liveProvider,
|
||||
assertCurrent() {
|
||||
if (digestObject(corpus) !== pin.corpusDigest || sourceSnapshot(repoRoot).sourceDigest !== pin.sourceDigest) fail('source-drift');
|
||||
} };
|
||||
const temp = fs.realpathSync(fs.mkdtempSync(path.join(os.tmpdir(), 'ecc-ai-eval-')));
|
||||
const selection = [];
|
||||
const outcomes = [];
|
||||
let installs = null;
|
||||
let harvester = null;
|
||||
try {
|
||||
const installRoot = path.join(temp, 'installs');
|
||||
fs.mkdirSync(installRoot, { mode: 0o700 });
|
||||
const envs = environments || (injected ? syntheticEnvironments(installRoot)
|
||||
: resolvedFamily === 'claude'
|
||||
? prepareClaudeEnvironments({ repoRoot, executable, root: installRoot,
|
||||
...(pin.arms.includes('ecc-legacy')
|
||||
? { legacySource: exportLegacySource({ repoRoot, destination: path.join(installRoot, 'legacy-source') }) }
|
||||
: {}) })
|
||||
: prepareEnvironments({ repoRoot, executable, root: installRoot }));
|
||||
installs = Object.fromEntries(Object.entries(envs).map(([name, env]) => [name,
|
||||
{ profileId: env.profileId, skills: env.skills, ...(env.sourceSha ? { sourceSha: env.sourceSha } : {}) }]));
|
||||
harvester = artifactDir && resolvedFamily === 'claude' && !injected ? createHarvester(artifactDir, envs) : null;
|
||||
const harvest = harvester ? (arm, id, repeat, env) => harvester.record(arm, id, repeat, env) : null;
|
||||
for (const item of corpus.selection) {
|
||||
const cwd = path.join(temp, `${item.id}--selection`);
|
||||
fs.mkdirSync(cwd);
|
||||
const start = state.metrics.length;
|
||||
selection.push({ ...selectionProbe(item, repoRoot, executeAdapter(state, cwd, envs.lean), envs.lean, resolvedFamily),
|
||||
...metricsSince(state.metrics, start) });
|
||||
}
|
||||
for (const scheduled of pin.order) {
|
||||
const item = corpus.tasks.find(c => c.id === scheduled.id);
|
||||
for (const arm of scheduled.arms) {
|
||||
const cwd = path.join(temp, `${item.id}--${arm}--${scheduled.repeat}`);
|
||||
const environment = envs[['full', 'baseline', 'ecc-legacy'].includes(arm) ? arm : 'lean'];
|
||||
fs.mkdirSync(cwd);
|
||||
writeWorkspace(cwd, item.files);
|
||||
const start = state.metrics.length;
|
||||
outcomes.push({ ...outcomeTrial(item, arm, scheduled.repeat, repoRoot,
|
||||
executeAdapter(state, cwd, environment), cwd, environment, resolvedFamily, harvest, state.metrics),
|
||||
...metricsSince(state.metrics, start) });
|
||||
fs.rmSync(cwd, { recursive: true, force: true });
|
||||
}
|
||||
}
|
||||
if (harvester) harvester.writeIndex();
|
||||
} finally { if (harvester) harvester.writeIndex(); fs.rmSync(temp, { recursive: true, force: true }); }
|
||||
const summary = summarize(outcomes, pin.arms);
|
||||
const insufficient = summary.distinctTasks < pin.minimumDistinctTasks || selection.length < pin.minimumDistinctTasks;
|
||||
const selectionSuccesses = selection.filter(row => row.passed).length;
|
||||
return { schemaVersion: 'ecc.context-eval.v2', registration: pin,
|
||||
evidence: injected ? 'injected-provider' : resolvedFamily === 'claude' ? 'claude-json' : 'codex-jsonl', installs,
|
||||
authentication: injected ? 'injected' : liveProvider.authentication, credentialsRetained: false,
|
||||
calls: state.calls, bounds: { maxCalls, deadlineMs, callTimeoutMs }, selection, outcomes, summary,
|
||||
selectionSummary: { n: selection.length, successes: selectionSuccesses,
|
||||
categories: [...new Set(selection.map(row => row.category))].map(category => ({ category,
|
||||
n: selection.filter(row => row.category === category).length,
|
||||
successes: selection.filter(row => row.category === category && row.passed).length })),
|
||||
interval: wilson(selectionSuccesses, selection.length), method: 'wilson-95-descriptive-purposive-sample' },
|
||||
gate: { status: insufficient ? 'insufficient-sample' : injected ? 'synthetic-only' : 'review-required',
|
||||
nonInferioritySupported: !insufficient && !injected && summary.pairs.every(p => p.interval[0] >= -pin.nonInferiorityMargin),
|
||||
releaseApproved: false }, nativeInvocation: 'unobserved',
|
||||
measurementScope: 'native-install-hidden-graded-coding-tasks',
|
||||
artifactRetention: harvester ? 'session-jsonl-per-task-trial' : 'none', ...metricsSince(state.metrics, 0) };
|
||||
}
|
||||
|
||||
module.exports = { loadCorpus, preregister, runEvaluation, parseCodexJsonl, parseClaudeJson, summarize, wilson,
|
||||
runCheck, runScoredCheck, createAuthLease, createCodexProvider, createClaudeProvider, prepareEnvironments,
|
||||
prepareClaudeEnvironments, exportLegacySource, providerFamily, resolveFamily, readClaudeKeychainToken };
|
||||
@@ -0,0 +1,52 @@
|
||||
#!/usr/bin/env node
|
||||
'use strict';
|
||||
const fs = require('node:fs');
|
||||
const { preregister, runEvaluation, loadCorpus } = require('./ai-eval-lib');
|
||||
|
||||
function main(argv = process.argv.slice(2), injected = {}) {
|
||||
const flags = new Map();
|
||||
const switches = new Set(['--plan', '--allow-real-provider', '--allow-credentialed-tools', '--help']);
|
||||
const values = new Set(['--registration', '--model', '--executable', '--provider', '--auth-home', '--effort', '--repeats', '--max-calls', '--deadline-ms', '--artifact-dir', '--corpus', '--call-timeout-ms', '--arms']);
|
||||
for (let i = 0; i < argv.length; i++) {
|
||||
const flag = argv[i];
|
||||
if (flags.has(flag) || (!switches.has(flag) && !values.has(flag))) throw new Error('Invalid evaluation arguments');
|
||||
if (values.has(flag) && (!argv[i + 1] || argv[i + 1].startsWith('--'))) throw new Error('Missing evaluation argument');
|
||||
flags.set(flag, switches.has(flag) ? true : argv[++i]);
|
||||
}
|
||||
if (flags.has('--help')) {
|
||||
return { usage: 'ai-eval.js --plan [--corpus FILE] [--arms a,b] [--repeats N] [--model MODEL --executable ABSOLUTE_PATH [--provider claude|codex] [--effort LEVEL]] | --allow-real-provider --registration FILE --model MODEL --executable ABSOLUTE_PATH [--provider claude|codex] [--allow-credentialed-tools (Claude only)] [--effort LEVEL (Codex only)] [--auth-home ABSOLUTE_DIR (Codex only)] [--corpus FILE] [--arms a,b] [--repeats N] [--max-calls N] [--deadline-ms N] [--call-timeout-ms N]. Claude auth: CLAUDE_CODE_OAUTH_TOKEN, ANTHROPIC_API_KEY, or the macOS Keychain login.' };
|
||||
}
|
||||
if (flags.get('--provider') !== undefined && !['claude', 'codex'].includes(flags.get('--provider'))) throw new Error('Provider must be claude or codex');
|
||||
if (flags.get('--provider') === 'claude' && flags.has('--effort')) throw new Error('Reasoning effort applies only to the Codex provider');
|
||||
if (flags.has('--allow-credentialed-tools') && (!flags.has('--allow-real-provider') || flags.get('--provider') !== 'claude')) {
|
||||
throw new Error('Credentialed-tool opt-in requires a real Claude evaluation');
|
||||
}
|
||||
const repeats = flags.has('--repeats') ? Number(flags.get('--repeats')) : 1;
|
||||
const corpus = flags.has('--corpus') ? loadCorpus(flags.get('--corpus')) : undefined;
|
||||
const arms = flags.has('--arms') ? flags.get('--arms').split(',').map(a => a.trim()).filter(Boolean) : undefined;
|
||||
if (flags.has('--plan')) {
|
||||
if (flags.has('--allow-real-provider')) throw new Error('Plan and provider execution are separate actions');
|
||||
return preregister({ repeats, model: flags.get('--model'), executable: flags.get('--executable'), effort: flags.get('--effort'),
|
||||
...(corpus ? { corpus } : {}), ...(arms ? { arms } : {}) });
|
||||
}
|
||||
if (!flags.has('--allow-real-provider') && !injected.provider) throw new Error('Real evaluation requires explicit opt-in');
|
||||
if (!flags.has('--registration')) throw new Error('Evaluation requires a preregistration file');
|
||||
const registration = JSON.parse(fs.readFileSync(flags.get('--registration'), 'utf8'));
|
||||
return runEvaluation({ ...injected, registration, repeats, allowRealProvider: flags.has('--allow-real-provider'),
|
||||
allowCredentialedTools: flags.has('--allow-credentialed-tools'),
|
||||
executable: flags.get('--executable'), model: flags.get('--model'), family: flags.get('--provider'), effort: flags.get('--effort'), authHome: flags.get('--auth-home'),
|
||||
artifactDir: flags.get('--artifact-dir'), ...(corpus ? { corpus } : {}), ...(arms ? { arms } : {}),
|
||||
...(flags.has('--max-calls') ? { maxCalls: Number(flags.get('--max-calls')) } : {}),
|
||||
...(flags.has('--deadline-ms') ? { deadlineMs: Number(flags.get('--deadline-ms')) } : {}),
|
||||
...(flags.has('--call-timeout-ms') ? { callTimeoutMs: Number(flags.get('--call-timeout-ms')) } : {}) });
|
||||
}
|
||||
if (require.main === module) {
|
||||
try { process.stdout.write(`${JSON.stringify(main())}\n`); }
|
||||
catch (error) {
|
||||
// Only fixed messages from this evaluator are shown; provider output and paths never reach stderr.
|
||||
const known = /^(Invalid|Missing|Real|Evaluation|Plan|Registration|Provider|Auth home|Native Codex version|Reasoning effort|Claude Keychain login|Claude)[^/\\]*$/.test(error?.message || '');
|
||||
process.stderr.write(`Evaluation stopped: ${known ? error.message : 'invalid arguments, registration, source, or provider configuration'}. Use --help.\n`);
|
||||
process.exitCode = 1;
|
||||
}
|
||||
}
|
||||
module.exports = { main };
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1,293 @@
|
||||
# ECC Complex-Task Evaluation (complex-tasks@1)
|
||||
|
||||
A reproducible, public benchmark of what ECC's context scoping does for **realistic
|
||||
agent work** — as opposed to the 30-task repair corpus (`ai-corpus.json`), which
|
||||
measures small, single-file fixes. This document is the preregistered methodology:
|
||||
it was written before the first provider call against this corpus, and it is the
|
||||
reference for anyone who wants to audit or rerun the evaluation.
|
||||
|
||||
## Research question
|
||||
|
||||
Does ECC's context engineering — the full skill library, manually picked skills
|
||||
(manual-lean), automatic skill matching (auto-lean), and the ECC-029 changes
|
||||
themselves — change what a frontier coding agent delivers on multi-step
|
||||
engineering tasks, and at what cost in tokens, time, and dollars?
|
||||
|
||||
## Arms
|
||||
|
||||
Five conditions, all launched through the same evaluator with real installs in
|
||||
isolated config homes, paired per task and repeat:
|
||||
|
||||
| Arm | What the agent gets | What it represents |
|
||||
|---|---|---|
|
||||
| `full` | Branch skill library installed + ECC context block (catalog/resources) | ECC with scoping machinery present but everything loaded |
|
||||
| `manual-lean` | lean profile + the maintainer-chosen canonical skill(s) injected | A user who knows exactly which ECC skill applies |
|
||||
| `auto-lean` | lean profile; ECC's trigger/proposal machinery picks and injects skills | The "auto" experience: no ECC knowledge required |
|
||||
| `ecc-legacy` | The full skill library **from the pinned pre-ECC-029 commit** (`legacy-source.json`, currently `e482e579` = `origin/main`), bare prompt, no context block | The typical current ECC user experience before the scoping work |
|
||||
| `baseline` | No ECC install, bare prompt | The provider with no ECC at all (overhead subtraction) |
|
||||
|
||||
`ecc-legacy` doubles as a replication control: where its install content matches
|
||||
`full`, score differences between them isolate the ECC-029 deltas (rewritten
|
||||
skill descriptions, scoping layer) rather than provider noise.
|
||||
|
||||
## The three tasks
|
||||
|
||||
Chosen to be the kind of work ECC exists for — multi-step, judgment-heavy,
|
||||
checkpointable — while deliberately **not** shaped around ECC's current skill
|
||||
list. Queries are written as a real user would phrase them, with no ECC
|
||||
vocabulary, no hints about which skill applies, and no instruction to use any
|
||||
particular methodology. Each task has one clear correct outcome and a
|
||||
deterministic, dependency-free grader.
|
||||
|
||||
1. **`webhook-relay`** (feature build). Finish an asynchronous webhook delivery
|
||||
worker: retries with exponential backoff, dead-lettering after 5 attempts,
|
||||
status reporting, under load. Graded by 9 in-process behavioral probes
|
||||
(delivery after failures, exact attempt counts, backoff timing window,
|
||||
dead-lettering, error capture, API preservation, concurrency).
|
||||
*Why it belongs here:* everyday backend feature work where test discipline
|
||||
and backend patterns genuinely change outcomes; canonical skill:
|
||||
`tdd-workflow` (a second skill would exceed the 32 KB selection budget —
|
||||
itself a measured constraint of the scoping layer).
|
||||
|
||||
2. **`incident-triage`** (debugging / root cause). Finance reports one-cent
|
||||
total errors since yesterday's deploy. The repo contains three changelog
|
||||
entries (two red herrings), an incident log with concrete amounts, and a
|
||||
regression: a "readability" refactor that switched integer-cent math to
|
||||
decimal-factor floats, which under-rounds exact half-cent boundaries.
|
||||
Graded by 5 boundary-value totals the float path provably gets wrong, one
|
||||
regression probe, and 2 deterministic checks on the required `INCIDENT.md`
|
||||
(names the right changelog entry, explains the rounding mechanism).
|
||||
*Why it belongs here:* evidence-driven diagnosis under uncertainty is the
|
||||
highest-leverage agent workflow; guessing is penalized because red herrings
|
||||
are plausible; canonical skill: `orch-fix-defect`.
|
||||
|
||||
3. **`sentinel-api`** (security review + hardening). A paste service whose
|
||||
README documents the secure contract while the code violates it five ways:
|
||||
hardcoded admin token, path traversal, reflected XSS, predictable delete
|
||||
tokens, no body-size limit. Graded by 10 exploit probes (each vulnerability
|
||||
must actually be closed) plus functional regression probes (the documented
|
||||
API must still work), including one encoded-traversal variant so partial
|
||||
fixes score partially.
|
||||
*Why it belongs here:* security review is a canonical agent task with
|
||||
objectively checkable outcomes; canonical skill: `security-review`.
|
||||
|
||||
### Why these tests are effective
|
||||
|
||||
- **Realism over benchmark gaming.** Each task is a small production-shaped
|
||||
repo with docs, tests, logs, and changelogs — the inputs a real engineer (or
|
||||
a real user of an agent harness) actually has. Nothing references ECC.
|
||||
- **Correctness is decidable.** Every grader assertion is deterministic:
|
||||
behavioral probes against the agent's own running service, exact numeric
|
||||
answers on boundary cases, static source checks, exploit probes. No LLM
|
||||
judges, no rubrics, no human scoring.
|
||||
- **Partial credit.** Graders emit `ECC_EVAL_SCORE {"score": 0..1}`, so "found
|
||||
4 of 5 vulnerabilities" registers as 0.9-of-task progress instead of a binary
|
||||
failure. Pass/fail (score = 1.0) is reported alongside the mean score.
|
||||
- **Hard to luck into.** Red herrings (incident-triage), timing windows
|
||||
(webhook-relay), and exploit-verified fixes (sentinel-api) mean superficial
|
||||
plausible work scores low.
|
||||
- **Fair across arms.** Hidden graders run only after the agent exits, from a
|
||||
read-only sandbox; the agent never sees the grader. The same grader scores
|
||||
every arm identically. Reference solutions score 1.0 and as-shipped fixtures
|
||||
score ≤ 0.3 (`verify-checks.js` proves both before any provider call).
|
||||
|
||||
## Measured variables
|
||||
|
||||
Per trial (one task × arm × repeat), from the provider's own usage events:
|
||||
|
||||
- **Fresh input tokens** (input + cache-creation), **cache-read tokens**,
|
||||
**output tokens** — the context-cost story.
|
||||
- **Provider calls** per trial (1, or 2 when auto-lean needs a routing proposal).
|
||||
- **Wall-clock time** per provider call and per trial (ms) — time to completion.
|
||||
- **Score** (0..1) and **pass** (score = 1.0) from the hidden grader.
|
||||
- **API-equivalent cost**, derived at analysis time at Anthropic Opus list
|
||||
prices ($15 / $1.50 / $75 per million fresh-input / cache-read / output
|
||||
tokens). This is an accounting convention for comparison, not a billing
|
||||
claim; subscription pricing differs.
|
||||
- **Skill routing** (auto-lean): which skills the trigger/proposal machinery
|
||||
selected vs the maintainer-chosen canonical set, reported as the selection
|
||||
probe accuracy — the direct measure of "automatic skill matching".
|
||||
|
||||
Comparisons are **within-run only**: same provider, model, executable digest,
|
||||
corpus digest, and source digest, paired by task and repeat. Cross-run and
|
||||
cross-provider comparisons are invalid by design. This is a descriptive pilot
|
||||
(3 tasks × 5 arms × 4 repeats = 60 trials): it estimates direction and
|
||||
magnitude, not population statistics, and the report says so in its gate block.
|
||||
|
||||
## Reproducing or auditing
|
||||
|
||||
Everything below is committed; there are no hidden inputs.
|
||||
|
||||
```bash
|
||||
# 1. Inspect the tasks: fixtures, queries, graders, and reference solutions.
|
||||
ls docker/context-profiles/complex-eval/cases/
|
||||
ls docker/context-profiles/complex-eval/reference/
|
||||
|
||||
# 2. Prove the graders: reference solutions must score 1.0, fixtures below 1.0.
|
||||
node docker/context-profiles/complex-eval/verify-checks.js
|
||||
|
||||
# 3. Rebuild the corpus after any fixture edit (digest-pinned at registration).
|
||||
node docker/context-profiles/complex-eval/build-corpus.js
|
||||
|
||||
# 4. Preregister (pins corpus, source, model, executable digests; no provider).
|
||||
node docker/context-profiles/ai-eval.js --plan \
|
||||
--corpus docker/context-profiles/complex-corpus.json --repeats 4 \
|
||||
--provider claude --model <model> --executable /absolute/path/to/claude \
|
||||
> registration.json
|
||||
|
||||
# 5. Run (requires your own Claude subscription login or API key).
|
||||
node docker/context-profiles/ai-eval.js --allow-real-provider --allow-credentialed-tools \
|
||||
--registration registration.json \
|
||||
--corpus docker/context-profiles/complex-corpus.json \
|
||||
--provider claude --model <model> --executable /absolute/path/to/claude \
|
||||
--repeats 4 --max-calls 400 --deadline-ms 25200000 --call-timeout-ms 600000 \
|
||||
--artifact-dir /absolute/path/for/transcripts > report.json
|
||||
```
|
||||
|
||||
Claude task tools inherit the provider credential through the CLI process and can read it. Use
|
||||
`--allow-credentialed-tools` only with a trusted local corpus and credential. Without that
|
||||
explicit flag, real Claude task evaluation stops before a provider call; selection-only calls
|
||||
remain tool-free. This development evaluator does not provide a credential isolation boundary.
|
||||
|
||||
The registration digest binds the exact corpus, evaluator source, model, and
|
||||
executable; the run refuses to start if any of them drift, and aborts if the
|
||||
tree changes mid-run. `--artifact-dir` retains per-trial session transcripts
|
||||
for independent inspection (they never enter the report). The `ecc-legacy` arm
|
||||
is pinned by commit in `legacy-source.json` and exported from git objects at
|
||||
run time. The Codex provider is unsupported for this corpus (the legacy arm has
|
||||
no Codex install path); `--provider claude` is required.
|
||||
|
||||
## Known limits
|
||||
|
||||
- Three tasks is a probe, not a census: treat intervals as descriptive.
|
||||
- Tasks are Node.js/stdlib by construction (graders must be hermetic); results
|
||||
say nothing about other ecosystems directly.
|
||||
- `webhook-relay` uses wall-clock backoff windows; bounds are wide (250–5000ms)
|
||||
but loaded machines could in principle flake a timing probe. The grader
|
||||
reports each probe individually so flakes are visible.
|
||||
- Provider behavior varies week to week; the pinned model/executable digests
|
||||
make a rerun comparable only within the same pin.
|
||||
- Fixture wart observed in the 2026-09-25 run: on Node 24, `node --test test/`
|
||||
no longer scans the directory the way Node 22 did, so `npm test` fails as
|
||||
shipped. This is identical for every arm (the task says to make `npm test`
|
||||
pass, and agents fix the script), so fairness holds, but it adds unplanned
|
||||
work per trial. A future corpus revision should ship a portable test script.
|
||||
|
||||
## complex-tasks@2 (discriminative revision)
|
||||
|
||||
The @1 run saturated: every arm scored 1.000 on every task, so only economics
|
||||
and routing differed. @2 (`cases2/`, built to `complex-corpus-v2.json`) is
|
||||
designed to discriminate on the axes users actually pay for — correctness on
|
||||
traps, solution efficiency, spec thoroughness — with wide partial-credit
|
||||
spreads. The @1 corpus and its report stay untouched for comparability.
|
||||
|
||||
1. **`keccak-selector`** (domain-knowledge trap). Implement Ethereum function
|
||||
selectors from scratch, stdlib only. The trap: Node's crypto offers
|
||||
SHA3-256, which shares the Keccak-f[1600] permutation but differs in
|
||||
padding — the naive one-liner is wrong for every vector (verified: the
|
||||
naive control scores 0.25, format checks only). Graded by 9 selector
|
||||
vectors including a padding edge case, all cross-validated against Node's
|
||||
SHA3-256 on shared-permutation inputs. Canonical skill: `nodejs-keccak256`.
|
||||
*Hypothesis:* the skill body carries exactly this knowledge; bare agents
|
||||
must rediscover it.
|
||||
|
||||
2. **`event-stats-api`** (correctness edges + measured efficiency). A shipped
|
||||
implementation that is both wrong on the documented edge semantics
|
||||
(interpolated instead of nearest-rank percentiles, zeros instead of nulls,
|
||||
unrounded averages, missing 400s) and algorithmically naive (full-log scan
|
||||
and sort per query). Graded by 10 independently computed correctness probes
|
||||
plus a measured 2,000-query performance budget (threshold 6s; shipped naive
|
||||
~7.7s, reference ~1.5s — calibrated on the grading machine in
|
||||
`calibrate-stats.js`). Canonical skill: `backend-patterns`. *Hypothesis:*
|
||||
solution *efficiency* separates arms even when correctness doesn't.
|
||||
|
||||
3. **`forge-cli`** (spec thoroughness + robustness). Twelve contractual
|
||||
behaviors with exact messages, exit codes, sorting, and a never-throw
|
||||
guarantee, graded by 26 checks including junk-input fuzzing and static
|
||||
hygiene (no leftover TODO/FIXME, no new dependencies). Canonical skill:
|
||||
`tdd-workflow`. *Hypothesis:* checklist discipline shows up as breadth of
|
||||
completion, and partial credit spreads the distribution.
|
||||
|
||||
First @2 run uses `claude-opus-4-8` (cost discipline); the corpus is
|
||||
provider- and model-pinned per run, so a later Opus 5.5 rerun on the same
|
||||
digest measures the model difference directly. repeats=2 (30 trials): simple
|
||||
experimentation, expand later.
|
||||
|
||||
## complex-tasks@3 (vagueness and horizon; arms: auto-lean vs baseline)
|
||||
|
||||
@2 still saturated on outcomes (30/30) — enumerated specs are within the
|
||||
model's cold competence. @3 (`cases3/`, built to `complex-corpus-v3.json`)
|
||||
moves grading to what users actually complain about (see the complaint
|
||||
taxonomy in this file's discussion: happy-path-only work, unverified
|
||||
completion, skipped implied work, convention drift, concurrency blindness).
|
||||
Everything graded is discoverable from repo docs visible to every arm — the
|
||||
question is whether agents reliably *do* all of it under vague instruction.
|
||||
|
||||
1. **`chained-tickets`** (long horizon). Four sequential tickets in one
|
||||
accumulating workspace — build a link shortener core, then vague tickets:
|
||||
"links need to survive a restart", "we're seeing abuse, deal with it",
|
||||
"track redirect hits, consistent with the existing API". 33 hidden probes
|
||||
across the four steps grade function, convention compliance (error
|
||||
envelope, layering — pinned in a visible CONTRIBUTING.md), and implied
|
||||
work (changelog entries, growing tests, accurate README). Stepped trials
|
||||
grade each ticket after its call; a failed ticket ends the chain.
|
||||
2. **`production-ready`** (vague prompt, heavy implication). "This goes to
|
||||
production Monday — get it ready." A documented production bar
|
||||
(validation envelopes, body limits, /health, structured request logs, env
|
||||
config, graceful SIGTERM, nosniff, error-path tests, changelog) graded by
|
||||
16 probes against a naive prototype. Fixture scores 0.063.
|
||||
3. **`idempotent-webhooks`** (the "almost right" trap). A payment receiver
|
||||
whose shipped code has a textbook check-then-act race (INC-104). Hidden
|
||||
grader fires 50 concurrent identical deliveries plus replay, already-paid,
|
||||
mixed-storm, and contract probes. The naive fixture double-applies and
|
||||
crashes on unknown orders (0.25). Exactly-once requires claiming events
|
||||
synchronously — the discipline skills like `error-handling` encode.
|
||||
|
||||
Grader robustness (hard-won, now fixed and unit-tested): a graded server runs
|
||||
in-process, so a crashing server kills the grader. Graders install
|
||||
uncaughtException/unhandledRejection handlers, emit their score line via
|
||||
`process.stdout.write` (immune to the log-capture patching used in probes),
|
||||
pre-declare their check totals (unreached checks score zero), and the
|
||||
evaluator itself treats a score-advertising grader that printed nothing as a
|
||||
zero (`graderDied` guard in `runScoredCheck`). Stepped graders may write to
|
||||
the workspace (persistence probes); single-step graders stay read-only.
|
||||
|
||||
First @3 run: arms `auto-lean` and `baseline` only, repeats=1,
|
||||
`claude-opus-4-8` — the direct test of "ECC auto-routing vs no harness" on
|
||||
quality, time, and tokens. Full-arm and Opus 5.5 replications follow if the
|
||||
spread shows up.
|
||||
|
||||
## complex-tasks@4 (learning loops; adds recurring-incident)
|
||||
|
||||
@4 (`cases4/`, built to `complex-corpus-v4.json`) keeps the three @3 cases
|
||||
unchanged and adds a fourth targeting a different ECC value prop: converting
|
||||
a fix into durable, reusable prevention — and *reusing your own artifacts*
|
||||
later in the session. Baseline agents can hold this in context; ECC's claim
|
||||
is that skills/workflows make it systematic.
|
||||
|
||||
4. **`recurring-incident`** (learning loop / institutional memory). Three
|
||||
chained steps against a dependency-free payments service whose gateway
|
||||
records side effects in an append-only JSONL ledger. Step 1: keyless
|
||||
refund retries double-refund (INC-201/214/227 "third time this quarter"
|
||||
trail in `docs/incidents.md`); the vague ask is "make sure this stops
|
||||
being a recurring incident." Probes: functional correctness across a
|
||||
module reload (kills in-memory-only fixes) [0.40], regression test wired
|
||||
into the suite + mutation probe [0.30], a durable prevention runbook
|
||||
[0.20], and the mechanism living in one shared helper module [0.10].
|
||||
Step 2: payout retries, "same family of problem" — graded on REUSE of
|
||||
the step-1 helper (static import check + no divergent inline
|
||||
reimplementation) [0.30] alongside function [0.40], test+mutation [0.20],
|
||||
doc update [0.10]. Step 3: "write the handoff note" — graded on
|
||||
existence [0.20], every referenced path actually existing on disk [0.30],
|
||||
naming the helper + prevention procedure [0.30], and covering both
|
||||
incidents [0.20]. Manual skills: `error-handling`, `continuous-learning`.
|
||||
*Hypothesis:* learning-loop behavior (abstract once, reuse, document,
|
||||
hand off) separates harnessed arms from baseline even when raw bug-fix
|
||||
competence doesn't.
|
||||
|
||||
Verification: reference 1.000 on all steps of all four cases; naive
|
||||
recurring-incident scores 0.20 / 0.00 / 0.20 per step; fixtures 0.00–0.25.
|
||||
|
||||
First @4 run: arm `auto-lean` only, repeats=1, `claude-opus-5-5` — the
|
||||
model-difference probe against the @3 opus-4-8 numbers on the shared cases,
|
||||
plus first signal on the learning-loop case.
|
||||
@@ -0,0 +1,67 @@
|
||||
'use strict';
|
||||
// Development tool: assembles a complex corpus JSON from a reviewed fixture
|
||||
// tree. Usage: node build-corpus.js [casesDir=cases] [outFile=complex-corpus.json] [corpusId=complex-tasks@1]
|
||||
// Run after editing any fixture, query, or grader; commit the tree and the
|
||||
// regenerated corpus together.
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const root = __dirname;
|
||||
const casesDir = path.join(root, process.argv[2] || 'cases');
|
||||
const OUT = path.join(root, '..', process.argv[3] || 'complex-corpus.json');
|
||||
const corpusId = process.argv[4] || 'complex-tasks@1';
|
||||
|
||||
function collect(directory, prefix = '') {
|
||||
const files = {};
|
||||
for (const entry of fs.readdirSync(directory, { withFileTypes: true }).sort((a, b) => a.name.localeCompare(b.name))) {
|
||||
const relative = prefix ? `${prefix}/${entry.name}` : entry.name;
|
||||
if (entry.isDirectory()) Object.assign(files, collect(path.join(directory, entry.name), relative));
|
||||
else if (entry.isFile()) files[relative] = fs.readFileSync(path.join(directory, entry.name), 'utf8');
|
||||
}
|
||||
return files;
|
||||
}
|
||||
|
||||
const tasks = [];
|
||||
const selection = [];
|
||||
for (const id of fs.readdirSync(casesDir).sort()) {
|
||||
const directory = path.join(casesDir, id);
|
||||
const meta = JSON.parse(fs.readFileSync(path.join(directory, 'meta.json'), 'utf8'));
|
||||
if (meta.id !== id || !/^[a-z][a-z0-9-]{0,63}$/.test(id)) throw new Error(`Invalid task metadata in ${id}`);
|
||||
const files = collect(path.join(directory, 'files'));
|
||||
const stepsDir = path.join(directory, 'steps');
|
||||
let task;
|
||||
if (fs.existsSync(stepsDir)) {
|
||||
const steps = fs.readdirSync(stepsDir).sort().map((name, index) => ({
|
||||
query: fs.readFileSync(path.join(stepsDir, name, 'query.md'), 'utf8').trim(),
|
||||
check: fs.readFileSync(path.join(stepsDir, name, 'check.cjs'), 'utf8'),
|
||||
...(meta.steps?.[index]?.manualIds ? { manualIds: meta.steps[index].manualIds } : {}),
|
||||
...((meta.steps?.[index]?.checkTimeoutMs || meta.checkTimeoutMs)
|
||||
? { checkTimeoutMs: meta.steps?.[index]?.checkTimeoutMs || meta.checkTimeoutMs } : {}),
|
||||
}));
|
||||
task = { id, category: meta.category, manualIds: meta.manualIds || [], files, steps };
|
||||
} else {
|
||||
const query = fs.readFileSync(path.join(directory, 'query.md'), 'utf8').trim();
|
||||
task = { id, category: meta.category, manualIds: meta.manualIds,
|
||||
...(meta.checkTimeoutMs ? { checkTimeoutMs: meta.checkTimeoutMs } : {}),
|
||||
query, files, check: fs.readFileSync(path.join(directory, 'check.cjs'), 'utf8') };
|
||||
}
|
||||
tasks.push(task);
|
||||
selection.push({ id: meta.selection.id, category: meta.selection.category,
|
||||
query: meta.selection.query || task.query || task.steps.map(step => step.query).join(' '),
|
||||
expectedIds: meta.selection.expectedIds });
|
||||
}
|
||||
|
||||
const corpus = {
|
||||
schemaVersion: 'ecc.context-eval-complex-corpus.v1',
|
||||
id: corpusId,
|
||||
sampling: 'Realistic multi-file engineering tasks, fixed before any provider call, with deterministic '
|
||||
+ 'hidden graders scoring partial credit (ECC_EVAL_SCORE). Descriptive pilot: no '
|
||||
+ 'population-representativeness claim. See complex-eval/DESIGN.md for the preregistered methodology.',
|
||||
minimumDistinctTasks: tasks.length,
|
||||
nonInferiorityMargin: 0.05,
|
||||
selection,
|
||||
tasks,
|
||||
};
|
||||
fs.writeFileSync(OUT, `${JSON.stringify(corpus, null, 1)}\n`);
|
||||
console.log(`wrote ${path.basename(OUT)} (${corpusId}): ${tasks.length} tasks, ${selection.length} selection probes, `
|
||||
+ `${tasks.reduce((sum, task) => sum + Object.keys(task.files).length, 0)} fixture files`);
|
||||
@@ -0,0 +1,73 @@
|
||||
'use strict';
|
||||
// Calibration harness (not shipped in the corpus): measures the 2,000-query
|
||||
// workload wall time for the shipped naive app and the reference app, each
|
||||
// staged as a standalone copy (fixture; fixture + reference overlay).
|
||||
const fs = require('node:fs');
|
||||
const os = require('node:os');
|
||||
const path = require('node:path');
|
||||
|
||||
const root = __dirname;
|
||||
const fixture = path.join(root, 'cases2', 'event-stats-api', 'files');
|
||||
const overlay = path.join(root, 'reference2', 'event-stats-api');
|
||||
|
||||
function stage(withOverlay) {
|
||||
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'ecc-calib-'));
|
||||
const copy = (from, to) => {
|
||||
for (const entry of fs.readdirSync(from, { withFileTypes: true })) {
|
||||
const target = path.join(to, entry.name);
|
||||
if (entry.isDirectory()) { fs.mkdirSync(target, { recursive: true }); copy(path.join(from, entry.name), target); }
|
||||
else fs.copyFileSync(path.join(from, entry.name), target);
|
||||
}
|
||||
};
|
||||
copy(fixture, dir);
|
||||
if (withOverlay) copy(overlay, dir);
|
||||
return dir;
|
||||
}
|
||||
|
||||
function lcg(seed) {
|
||||
let state = seed >>> 0;
|
||||
return () => {
|
||||
state = (Math.imul(state, 1664525) + 1013904223) >>> 0;
|
||||
return state / 2 ** 32;
|
||||
};
|
||||
}
|
||||
|
||||
function workload(types, epoch, span) {
|
||||
const rand = lcg(777);
|
||||
const queries = [];
|
||||
for (let i = 0; i < 2000; i++) {
|
||||
const type = types[Math.floor(rand() * types.length)];
|
||||
const start = epoch + Math.floor(rand() * span * 0.7);
|
||||
queries.push({ type, from: start, to: start + Math.floor(rand() * span * 0.5) });
|
||||
}
|
||||
return queries;
|
||||
}
|
||||
|
||||
async function measure(label, dir) {
|
||||
const { createApp } = require(path.join(dir, 'src', 'app.js'));
|
||||
const { TYPES, EPOCH_MS, SPAN_MS } = require(path.join(dir, 'src', 'data.js'));
|
||||
const app = createApp();
|
||||
await new Promise(resolve => app.listen(0, '127.0.0.1', resolve));
|
||||
const port = app.address().port;
|
||||
const queries = workload(TYPES, EPOCH_MS, SPAN_MS);
|
||||
const started = Date.now();
|
||||
for (let i = 0; i < queries.length; i += 20) {
|
||||
await Promise.all(queries.slice(i, i + 20).map(q =>
|
||||
fetch(`http://127.0.0.1:${port}/stats?type=${q.type}&from=${q.from}&to=${q.to}`).then(r => r.json())));
|
||||
}
|
||||
const elapsed = Date.now() - started;
|
||||
app.close();
|
||||
console.log(`${label}: ${elapsed}ms for 2000 queries`);
|
||||
return elapsed;
|
||||
}
|
||||
|
||||
(async () => {
|
||||
const naiveDir = stage(false);
|
||||
const refDir = stage(true);
|
||||
await measure('naive 1 ', naiveDir);
|
||||
await measure('naive 2 ', naiveDir);
|
||||
await measure('reference 1 ', refDir);
|
||||
await measure('reference 2 ', refDir);
|
||||
fs.rmSync(naiveDir, { recursive: true, force: true });
|
||||
fs.rmSync(refDir, { recursive: true, force: true });
|
||||
})();
|
||||
@@ -0,0 +1,45 @@
|
||||
'use strict';
|
||||
// Hidden grader for incident-triage: checks exact totals on boundary orders and
|
||||
// the root-cause report. Prints ECC_EVAL_SCORE and always exits 0.
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const checks = [];
|
||||
const record = (name, ok) => checks.push({ name, ok: Boolean(ok) });
|
||||
|
||||
let computeOrderTotal;
|
||||
try { ({ computeOrderTotal } = require(path.join(process.cwd(), 'src', 'totals.js'))); } catch { /* scored below */ }
|
||||
|
||||
// Boundary orders where decimal-factor float math under-rounds by a cent;
|
||||
// expected values follow the README pricing rules (integer cents, half-up per line).
|
||||
const boundary = [
|
||||
{ lines: [{ priceCents: 165, quantity: 1 }], discountPercent: 30, expected: 116 },
|
||||
{ lines: [{ priceCents: 250, quantity: 1 }], discountPercent: 7, expected: 233 },
|
||||
{ lines: [{ priceCents: 325, quantity: 1 }], discountPercent: 30, expected: 228 },
|
||||
{ lines: [{ priceCents: 345, quantity: 1 }], discountPercent: 30, expected: 242 },
|
||||
{ lines: [{ priceCents: 165, quantity: 1 }, { priceCents: 325, quantity: 1 }], discountPercent: 30, expected: 344 },
|
||||
];
|
||||
|
||||
if (typeof computeOrderTotal === 'function') {
|
||||
boundary.forEach((order, index) => {
|
||||
let actual = NaN;
|
||||
try { actual = computeOrderTotal({ lines: order.lines, discountPercent: order.discountPercent }); } catch { /* wrong */ }
|
||||
record(`boundary-total-${index + 1}`, actual === order.expected);
|
||||
});
|
||||
let plain = NaN;
|
||||
try { plain = computeOrderTotal({ lines: [{ priceCents: 1000, quantity: 2 }], discountPercent: 0 }); } catch { /* wrong */ }
|
||||
record('undiscounted-total-unchanged', plain === 2000);
|
||||
} else {
|
||||
for (let index = 0; index < boundary.length; index++) record(`boundary-total-${index + 1}`, false);
|
||||
record('undiscounted-total-unchanged', false);
|
||||
}
|
||||
|
||||
let incident = '';
|
||||
try { incident = fs.readFileSync(path.join(process.cwd(), 'INCIDENT.md'), 'utf8'); } catch { /* missing */ }
|
||||
record('incident-identifies-C-2', /C-2/.test(incident));
|
||||
record('incident-explains-rounding', /round|float|decimal|cent/i.test(incident));
|
||||
|
||||
const ok = checks.filter(c => c.ok).length;
|
||||
for (const c of checks) console.log(`${c.ok ? 'ok' : 'not ok'} - ${c.name}`);
|
||||
console.log(`ECC_EVAL_SCORE ${JSON.stringify({ score: ok / checks.length, passed: ok, total: checks.length })}`);
|
||||
process.exit(0);
|
||||
@@ -0,0 +1,11 @@
|
||||
# Changelog
|
||||
|
||||
## 2026-09-23 deploy
|
||||
|
||||
- **C-1**: request logging switched to JSON lines (`src/request-log.js`).
|
||||
Log volume and format only; no request-handling behavior changed.
|
||||
- **C-2**: totals computation refactored for readability (`src/totals.js`).
|
||||
The old cents-as-integers helper was replaced with a direct decimal
|
||||
expression that reviewers found easier to follow. No behavior change intended.
|
||||
- **C-3**: inventory client timeout raised from 2s to 5s (`src/inventory-client.js`).
|
||||
Reduces spurious failures when the inventory service is slow.
|
||||
@@ -0,0 +1,21 @@
|
||||
# order-service
|
||||
|
||||
Computes order totals for the checkout service.
|
||||
|
||||
## Pricing rules
|
||||
|
||||
An order is `{ "lines": [{ "priceCents": number, "quantity": number }], "discountPercent": number }`.
|
||||
|
||||
- All prices are integer cents. There is no such thing as a fraction of a cent
|
||||
in an order total.
|
||||
- The discount applies per line: `lineCents = priceCents * quantity * (100 - discountPercent) / 100`,
|
||||
rounded **half-up** to the nearest cent (0.5 rounds up).
|
||||
- The order total is the sum of the rounded line totals, in integer cents.
|
||||
|
||||
`src/totals.js` is CommonJS and exports `computeOrderTotal(order)` returning the
|
||||
total in integer cents. Run the tests with `npm test`.
|
||||
|
||||
## Operations
|
||||
|
||||
- `CHANGELOG.md` records what shipped in each deploy.
|
||||
- `evidence/incident.txt` holds the finance team's findings for the current incident.
|
||||
@@ -0,0 +1,5 @@
|
||||
2026-09-24T08:57:11Z finance-review order=ORD-2204 note="charged_total_cents=115 expected_total_cents=116 lines=[{priceCents:165,quantity:1}] discountPercent=30"
|
||||
2026-09-24T09:14:02Z finance-review order=ORD-2291 note="charged_total_cents=232 expected_total_cents=233 lines=[{priceCents:250,quantity:1}] discountPercent=7"
|
||||
2026-09-24T09:41:37Z finance-review order=ORD-2310 note="charged_total_cents=227 expected_total_cents=228 lines=[{priceCents:325,quantity:1}] discountPercent=30"
|
||||
2026-09-24T10:05:19Z support-ticket customer="ORDER-2310 looks like it undercharged me by a cent vs the invoice email"
|
||||
2026-09-24T10:22:48Z finance-review summary="12 of 4,813 orders since the 2026-09-23 deploy are off by exactly one cent, always in the store's favor; all pre-deploy orders reconcile"
|
||||
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"name": "order-service",
|
||||
"private": true,
|
||||
"type": "commonjs",
|
||||
"scripts": { "test": "node --test test/" }
|
||||
}
|
||||
+11
@@ -0,0 +1,11 @@
|
||||
'use strict';
|
||||
|
||||
// Changed 2026-09-23 (C-3): the inventory service has been slow this week;
|
||||
// give it 5s instead of 2s before declaring a failure.
|
||||
const INVENTORY_TIMEOUT_MS = 5000;
|
||||
|
||||
function inventoryClientOptions() {
|
||||
return { timeoutMs: INVENTORY_TIMEOUT_MS, retries: 2 };
|
||||
}
|
||||
|
||||
module.exports = { inventoryClientOptions };
|
||||
@@ -0,0 +1,13 @@
|
||||
'use strict';
|
||||
|
||||
// Changed 2026-09-23 (C-1): emit request logs as JSON lines so the log
|
||||
// pipeline can parse them without regexes.
|
||||
function logRequest(req) {
|
||||
console.log(JSON.stringify({
|
||||
method: req.method,
|
||||
url: req.url,
|
||||
at: new Date().toISOString(),
|
||||
}));
|
||||
}
|
||||
|
||||
module.exports = { logRequest };
|
||||
@@ -0,0 +1,14 @@
|
||||
'use strict';
|
||||
|
||||
// Refactored 2026-09-23 (C-2): express the discount math directly with a
|
||||
// decimal factor instead of the old integer-cents helper, which reviewers
|
||||
// found hard to follow.
|
||||
function computeOrderTotal(order) {
|
||||
let total = 0;
|
||||
for (const line of order.lines) {
|
||||
total += Math.round(line.priceCents * line.quantity * (1 - order.discountPercent / 100));
|
||||
}
|
||||
return total;
|
||||
}
|
||||
|
||||
module.exports = { computeOrderTotal };
|
||||
@@ -0,0 +1,16 @@
|
||||
'use strict';
|
||||
const test = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const { computeOrderTotal } = require('../src/totals');
|
||||
|
||||
test('sums lines without a discount', () => {
|
||||
assert.equal(computeOrderTotal({ lines: [{ priceCents: 1000, quantity: 2 }], discountPercent: 0 }), 2000);
|
||||
});
|
||||
|
||||
test('applies a clean quarter discount', () => {
|
||||
assert.equal(computeOrderTotal({ lines: [{ priceCents: 2000, quantity: 1 }], discountPercent: 25 }), 1500);
|
||||
});
|
||||
|
||||
test('multiplies quantity before discounting', () => {
|
||||
assert.equal(computeOrderTotal({ lines: [{ priceCents: 400, quantity: 3 }], discountPercent: 50 }), 600);
|
||||
});
|
||||
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"id": "incident-triage",
|
||||
"category": "debugging-incident",
|
||||
"manualIds": ["skill:orch-fix-defect"],
|
||||
"checkTimeoutMs": 30000,
|
||||
"selection": {
|
||||
"id": "complex-incident-triage",
|
||||
"category": "complex-debugging-incident",
|
||||
"expectedIds": ["skill:orch-fix-defect"]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
Finance flagged that some order totals have been off by a cent since yesterday's deploy — details are in evidence/incident.txt. Three changes shipped yesterday (CHANGELOG.md, entries C-1 to C-3). Find the root cause, fix it so totals are computed exactly per the pricing rules in the README, keep `npm test` green, and write INCIDENT.md at the repo root identifying which changelog entry introduced the regression, with a short explanation of why it produces wrong totals.
|
||||
@@ -0,0 +1,123 @@
|
||||
'use strict';
|
||||
// Hidden grader for sentinel-api: runs exploit probes and functional regression
|
||||
// probes against the agent's service, in-process, plus static source checks.
|
||||
// Prints ECC_EVAL_SCORE and always exits 0.
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const checks = [];
|
||||
const record = (name, ok) => checks.push({ name, ok: Boolean(ok) });
|
||||
let finished = false;
|
||||
|
||||
function finish() {
|
||||
if (finished) return;
|
||||
finished = true;
|
||||
const ok = checks.filter(c => c.ok).length;
|
||||
for (const c of checks) console.log(`${c.ok ? 'ok' : 'not ok'} - ${c.name}`);
|
||||
console.log(`ECC_EVAL_SCORE ${JSON.stringify({ score: checks.length ? ok / checks.length : 0, passed: ok, total: checks.length })}`);
|
||||
process.exit(0);
|
||||
}
|
||||
setTimeout(finish, 45000).unref();
|
||||
|
||||
const HARDCODED_TOKEN = 'fixture-token';
|
||||
const GRADER_TOKEN = 'ecc-grader-admin-token-31f7';
|
||||
|
||||
function sourceFiles(directory, found = []) {
|
||||
for (const entry of fs.readdirSync(directory, { withFileTypes: true })) {
|
||||
const item = path.join(directory, entry.name);
|
||||
if (entry.isDirectory()) sourceFiles(item, found);
|
||||
else if (entry.name.endsWith('.js')) found.push(item);
|
||||
}
|
||||
return found;
|
||||
}
|
||||
|
||||
function listen(server) {
|
||||
return new Promise((resolve, reject) => {
|
||||
server.once('error', reject);
|
||||
server.listen(0, '127.0.0.1', () => resolve(server.address().port));
|
||||
});
|
||||
}
|
||||
|
||||
async function createPaste(port, content) {
|
||||
const response = await fetch(`http://127.0.0.1:${port}/pastes`, {
|
||||
method: 'POST', headers: { 'content-type': 'application/json' }, body: JSON.stringify({ content }) });
|
||||
return { status: response.status, body: await response.json().catch(() => null) };
|
||||
}
|
||||
|
||||
(async () => {
|
||||
const root = process.cwd();
|
||||
const sources = sourceFiles(path.join(root, 'src'))
|
||||
.map(file => fs.readFileSync(file, 'utf8')).join('\n');
|
||||
|
||||
// Static: the committed admin token must be gone.
|
||||
record('hardcoded-admin-token-removed', !sources.includes(HARDCODED_TOKEN));
|
||||
|
||||
// Static: delete tokens must come from a cryptographic source.
|
||||
record('delete-token-from-crypto', /crypto\.(randomBytes|randomUUID|randomInt)/.test(sources));
|
||||
|
||||
// The service reads ADMIN_TOKEN from the environment at startup.
|
||||
process.env.ADMIN_TOKEN = GRADER_TOKEN;
|
||||
let createApp;
|
||||
try { ({ createApp } = require(path.join(root, 'src', 'app.js'))); } catch { finish(); return; }
|
||||
if (typeof createApp !== 'function') { finish(); return; }
|
||||
|
||||
try {
|
||||
const app = createApp();
|
||||
const port = await listen(app);
|
||||
|
||||
// Admin route: rejected without the token, allowed with the environment token.
|
||||
const denied = await fetch(`http://127.0.0.1:${port}/admin/stats`);
|
||||
const wrong = await fetch(`http://127.0.0.1:${port}/admin/stats`, { headers: { 'x-admin-token': HARDCODED_TOKEN } });
|
||||
const allowed = await fetch(`http://127.0.0.1:${port}/admin/stats`, { headers: { 'x-admin-token': GRADER_TOKEN } });
|
||||
record('admin-requires-env-token', denied.status === 401 && wrong.status === 401 && allowed.status === 200);
|
||||
|
||||
// Path traversal: the canary outside data/files must never be served.
|
||||
const traversal = await fetch(`http://127.0.0.1:${port}/files?name=../secret.txt`);
|
||||
const traversalBody = await traversal.text();
|
||||
record('path-traversal-blocked', [400, 401, 403, 404].includes(traversal.status) && !traversalBody.includes('CANARY-7f3c9d'));
|
||||
const nested = await fetch(`http://127.0.0.1:${port}/files?name=${encodeURIComponent('..%2f..%2fsecret.txt')}`);
|
||||
record('encoded-traversal-blocked', !(await nested.text()).includes('CANARY-7f3c9d'));
|
||||
|
||||
// Legitimate files still work.
|
||||
const legit = await fetch(`http://127.0.0.1:${port}/files?name=welcome.txt`);
|
||||
record('public-files-still-served', legit.status === 200 && (await legit.text()).includes('hello from sentinel'));
|
||||
|
||||
// Reflected content is escaped in the HTML view.
|
||||
const xss = `<script>alert('xss')</script>`;
|
||||
const made = await createPaste(port, xss);
|
||||
const view = made.body && made.body.id
|
||||
? await fetch(`http://127.0.0.1:${port}/p/${made.body.id}`) : null;
|
||||
const html = view ? await view.text() : '';
|
||||
record('html-view-escapes-content', Boolean(view) && !html.includes(xss) && html.includes('<script>'));
|
||||
|
||||
// Delete tokens: unique, long, and not derived from the paste id.
|
||||
const first = await createPaste(port, 'one');
|
||||
const second = await createPaste(port, 'two');
|
||||
const t1 = first.body && first.body.deleteToken;
|
||||
const t2 = second.body && second.body.deleteToken;
|
||||
record('delete-token-unpredictable', Boolean(t1 && t2) && t1 !== t2 && t1.length >= 16
|
||||
&& !t1.includes(String(first.body.id)) && t1 !== `tok_${first.body.id}`);
|
||||
|
||||
// Oversized bodies are rejected per the contract.
|
||||
const big = await fetch(`http://127.0.0.1:${port}/pastes`, {
|
||||
method: 'POST', headers: { 'content-type': 'application/json' },
|
||||
body: JSON.stringify({ content: 'x'.repeat(200 * 1024) }) });
|
||||
record('oversized-body-rejected-413', big.status === 413);
|
||||
|
||||
// Functional regression: the documented flow still works end to end.
|
||||
const flow = await createPaste(port, 'roundtrip content');
|
||||
const readBack = flow.body && flow.body.id
|
||||
? await fetch(`http://127.0.0.1:${port}/pastes/${flow.body.id}`) : null;
|
||||
const readJson = readBack ? await readBack.json().catch(() => null) : null;
|
||||
const deleted = flow.body && flow.body.id
|
||||
? await fetch(`http://127.0.0.1:${port}/pastes/${flow.body.id}`, {
|
||||
method: 'DELETE', headers: { 'x-delete-token': flow.body.deleteToken } }) : null;
|
||||
const afterDelete = flow.body && flow.body.id
|
||||
? await fetch(`http://127.0.0.1:${port}/pastes/${flow.body.id}`) : null;
|
||||
record('documented-api-still-works', Boolean(readJson) && readJson.content === 'roundtrip content'
|
||||
&& Boolean(deleted) && deleted.status === 204 && Boolean(afterDelete) && afterDelete.status === 404);
|
||||
|
||||
app.close();
|
||||
} catch { /* grader-side failure leaves remaining checks unscored */ }
|
||||
finish();
|
||||
})();
|
||||
@@ -0,0 +1,28 @@
|
||||
# sentinel-paste
|
||||
|
||||
Internal paste-sharing service. In-memory store, Node.js standard library only.
|
||||
|
||||
## API contract
|
||||
|
||||
- `POST /pastes` — JSON body `{ "content": string }`. `201` with
|
||||
`{ "id", "deleteToken" }`. Request bodies larger than 64 KB are rejected with
|
||||
`413 Payload Too Large`.
|
||||
- `GET /pastes/:id` — `200` JSON `{ "id", "content" }`, or `404`.
|
||||
- `DELETE /pastes/:id` — requires the `X-Delete-Token` header to match the
|
||||
token returned at creation. `204` on success, `403` otherwise. Delete tokens
|
||||
are shown once and must be unguessable.
|
||||
- `GET /p/:id` — HTML view of a paste. The paste content is rendered as text:
|
||||
it must never be interpreted as markup.
|
||||
- `GET /files?name=<name>` — serves a public file from `data/files/`. Only
|
||||
files inside that directory may ever be served.
|
||||
- `GET /admin/stats` — operational stats. Requires the `X-Admin-Token` header
|
||||
to match the `ADMIN_TOKEN` environment variable; the service reads it at
|
||||
startup and refuses all admin access (`401`) when it is missing or wrong.
|
||||
The token must never be committed to the repository.
|
||||
|
||||
## Module contract
|
||||
|
||||
- `src/app.js` is CommonJS and exports `createApp()`, which returns an
|
||||
`http.Server` that is not yet listening.
|
||||
- `node src/index.js <port>` starts the service.
|
||||
- Run the tests with `npm test`.
|
||||
@@ -0,0 +1 @@
|
||||
hello from sentinel
|
||||
@@ -0,0 +1 @@
|
||||
CANARY-7f3c9d-internal-only
|
||||
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"name": "sentinel-paste",
|
||||
"private": true,
|
||||
"type": "commonjs",
|
||||
"scripts": { "test": "node --test test/" }
|
||||
}
|
||||
@@ -0,0 +1,99 @@
|
||||
'use strict';
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
const http = require('node:http');
|
||||
const config = require('./config');
|
||||
const store = require('./store');
|
||||
|
||||
function readBody(req, callback) {
|
||||
let body = '';
|
||||
req.on('data', chunk => { body += chunk; });
|
||||
req.on('end', () => callback(body));
|
||||
}
|
||||
|
||||
function sendJson(res, status, value) {
|
||||
res.writeHead(status, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify(value));
|
||||
}
|
||||
|
||||
function page(paste) {
|
||||
return `<!doctype html><html><head><title>paste ${paste.id}</title></head>`
|
||||
+ `<body><main><pre class="paste">${paste.content}</pre></main></body></html>`;
|
||||
}
|
||||
|
||||
function createApp() {
|
||||
return http.createServer((req, res) => {
|
||||
const url = new URL(req.url, 'http://localhost');
|
||||
|
||||
if (req.method === 'POST' && url.pathname === '/pastes') {
|
||||
readBody(req, body => {
|
||||
let parsed;
|
||||
try { parsed = JSON.parse(body); } catch {
|
||||
sendJson(res, 400, { error: 'invalid JSON body' });
|
||||
return;
|
||||
}
|
||||
if (typeof parsed.content !== 'string') {
|
||||
sendJson(res, 400, { error: 'content must be a string' });
|
||||
return;
|
||||
}
|
||||
const paste = store.create(parsed.content);
|
||||
sendJson(res, 201, { id: paste.id, deleteToken: paste.deleteToken });
|
||||
});
|
||||
return;
|
||||
}
|
||||
|
||||
const pasteMatch = /^\/pastes\/([\w-]+)$/.exec(url.pathname);
|
||||
if (pasteMatch && req.method === 'GET') {
|
||||
const paste = store.get(pasteMatch[1]);
|
||||
if (!paste) { sendJson(res, 404, { error: 'not found' }); return; }
|
||||
sendJson(res, 200, { id: paste.id, content: paste.content });
|
||||
return;
|
||||
}
|
||||
if (pasteMatch && req.method === 'DELETE') {
|
||||
const paste = store.get(pasteMatch[1]);
|
||||
if (!paste) { sendJson(res, 404, { error: 'not found' }); return; }
|
||||
if (req.headers['x-delete-token'] !== paste.deleteToken) {
|
||||
sendJson(res, 403, { error: 'bad delete token' });
|
||||
return;
|
||||
}
|
||||
store.remove(paste.id);
|
||||
res.writeHead(204);
|
||||
res.end();
|
||||
return;
|
||||
}
|
||||
|
||||
const pageMatch = /^\/p\/([\w-]+)$/.exec(url.pathname);
|
||||
if (pageMatch && req.method === 'GET') {
|
||||
const paste = store.get(pageMatch[1]);
|
||||
if (!paste) { sendJson(res, 404, { error: 'not found' }); return; }
|
||||
res.writeHead(200, { 'content-type': 'text/html' });
|
||||
res.end(page(paste));
|
||||
return;
|
||||
}
|
||||
|
||||
if (req.method === 'GET' && url.pathname === '/files') {
|
||||
const name = url.searchParams.get('name') || '';
|
||||
try {
|
||||
const content = fs.readFileSync(path.join(config.FILES_DIR, name));
|
||||
res.writeHead(200, { 'content-type': 'text/plain' });
|
||||
res.end(content);
|
||||
} catch {
|
||||
sendJson(res, 404, { error: 'not found' });
|
||||
}
|
||||
return;
|
||||
}
|
||||
|
||||
if (req.method === 'GET' && url.pathname === '/admin/stats') {
|
||||
if (req.headers['x-admin-token'] !== config.ADMIN_TOKEN) {
|
||||
sendJson(res, 401, { error: 'unauthorized' });
|
||||
return;
|
||||
}
|
||||
sendJson(res, 200, store.stats());
|
||||
return;
|
||||
}
|
||||
|
||||
sendJson(res, 404, { error: 'not found' });
|
||||
});
|
||||
}
|
||||
|
||||
module.exports = { createApp };
|
||||
@@ -0,0 +1,9 @@
|
||||
'use strict';
|
||||
const path = require('node:path');
|
||||
|
||||
module.exports = {
|
||||
// TODO: move this out of the repository before the next audit.
|
||||
ADMIN_TOKEN: 'fixture-token',
|
||||
MAX_BODY_BYTES: 64 * 1024,
|
||||
FILES_DIR: path.join(__dirname, '..', 'data', 'files'),
|
||||
};
|
||||
@@ -0,0 +1,7 @@
|
||||
'use strict';
|
||||
const { createApp } = require('./app');
|
||||
|
||||
const port = Number(process.argv[2] || 8080);
|
||||
createApp().listen(port, () => {
|
||||
console.log(`sentinel-paste listening on ${port}`);
|
||||
});
|
||||
@@ -0,0 +1,26 @@
|
||||
'use strict';
|
||||
|
||||
// In-memory paste store.
|
||||
const pastes = new Map();
|
||||
let nextId = 1;
|
||||
|
||||
function create(content) {
|
||||
const id = `p_${nextId++}`;
|
||||
const paste = { id, content, deleteToken: `tok_${id}` };
|
||||
pastes.set(id, paste);
|
||||
return paste;
|
||||
}
|
||||
|
||||
function get(id) {
|
||||
return pastes.get(id) || null;
|
||||
}
|
||||
|
||||
function remove(id) {
|
||||
return pastes.delete(id);
|
||||
}
|
||||
|
||||
function stats() {
|
||||
return { pastes: pastes.size, created: nextId - 1 };
|
||||
}
|
||||
|
||||
module.exports = { create, get, remove, stats };
|
||||
@@ -0,0 +1,28 @@
|
||||
'use strict';
|
||||
const test = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const { createApp } = require('../src/app');
|
||||
|
||||
function listen(server) {
|
||||
return new Promise((resolve, reject) => {
|
||||
server.once('error', reject);
|
||||
server.listen(0, '127.0.0.1', () => resolve(server.address().port));
|
||||
});
|
||||
}
|
||||
|
||||
test('create and read back a paste', async () => {
|
||||
const server = createApp();
|
||||
const port = await listen(server);
|
||||
try {
|
||||
const created = await fetch(`http://127.0.0.1:${port}/pastes`, {
|
||||
method: 'POST', headers: { 'content-type': 'application/json' },
|
||||
body: JSON.stringify({ content: 'hello paste' }) });
|
||||
assert.equal(created.status, 201);
|
||||
const { id } = await created.json();
|
||||
const read = await fetch(`http://127.0.0.1:${port}/pastes/${id}`);
|
||||
assert.equal(read.status, 200);
|
||||
assert.equal((await read.json()).content, 'hello paste');
|
||||
} finally {
|
||||
server.close();
|
||||
}
|
||||
});
|
||||
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"id": "sentinel-api",
|
||||
"category": "security-hardening",
|
||||
"manualIds": ["skill:security-review"],
|
||||
"checkTimeoutMs": 60000,
|
||||
"selection": {
|
||||
"id": "complex-sentinel-api",
|
||||
"category": "complex-security-hardening",
|
||||
"expectedIds": ["skill:security-review"]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
This internal paste-sharing service failed a security review, but the auditors didn't itemize the findings. Review the implementation against the API contract in the README, find every place the code violates the documented security behavior or is otherwise exploitable, and fix all of them without breaking the documented API. `npm test` must stay green.
|
||||
@@ -0,0 +1,125 @@
|
||||
'use strict';
|
||||
// Hidden grader for webhook-relay: drives the agent's relay in-process against
|
||||
// local target servers and prints ECC_EVAL_SCORE. Always exits 0; the score line
|
||||
// carries the result. Runs under Node's read-only permission model, so it only
|
||||
// reads the workspace and talks to 127.0.0.1.
|
||||
const http = require('node:http');
|
||||
const path = require('node:path');
|
||||
|
||||
const checks = [];
|
||||
const record = (name, ok) => checks.push({ name, ok: Boolean(ok) });
|
||||
const sleep = ms => new Promise(resolve => setTimeout(resolve, ms));
|
||||
let finished = false;
|
||||
|
||||
function finish() {
|
||||
if (finished) return;
|
||||
finished = true;
|
||||
const ok = checks.filter(c => c.ok).length;
|
||||
for (const c of checks) console.log(`${c.ok ? 'ok' : 'not ok'} - ${c.name}`);
|
||||
console.log(`ECC_EVAL_SCORE ${JSON.stringify({ score: checks.length ? ok / checks.length : 0, passed: ok, total: checks.length })}`);
|
||||
process.exit(0);
|
||||
}
|
||||
setTimeout(finish, 45000).unref();
|
||||
|
||||
function listen(server) {
|
||||
return new Promise((resolve, reject) => {
|
||||
server.once('error', reject);
|
||||
server.listen(0, '127.0.0.1', () => resolve(server.address().port));
|
||||
});
|
||||
}
|
||||
|
||||
function postJson(port, urlPath, body) {
|
||||
return fetch(`http://127.0.0.1:${port}${urlPath}`, {
|
||||
method: 'POST', headers: { 'content-type': 'application/json' }, body: JSON.stringify(body) })
|
||||
.then(async response => ({ status: response.status, body: await response.json().catch(() => null) }));
|
||||
}
|
||||
|
||||
async function waitForStatus(port, id, wanted, timeoutMs) {
|
||||
const started = Date.now();
|
||||
let last = null;
|
||||
while (Date.now() - started < timeoutMs) {
|
||||
try {
|
||||
const response = await fetch(`http://127.0.0.1:${port}/deliveries/${id}`);
|
||||
if (response.status === 200) {
|
||||
last = await response.json();
|
||||
if (last.status === wanted || last.status === 'dead') return { record: last, elapsedMs: Date.now() - started };
|
||||
}
|
||||
} catch { /* relay not ready yet */ }
|
||||
await sleep(25);
|
||||
}
|
||||
return { record: last, elapsedMs: Date.now() - started };
|
||||
}
|
||||
|
||||
(async () => {
|
||||
let createRelay;
|
||||
try { ({ createRelay } = require(path.join(process.cwd(), 'src', 'app.js'))); } catch { finish(); return; }
|
||||
if (typeof createRelay !== 'function') { finish(); return; }
|
||||
|
||||
// Probe group 1: a target that fails 3 times then succeeds.
|
||||
let calls = 0;
|
||||
const flaky = http.createServer((req, res) => {
|
||||
calls++;
|
||||
req.resume();
|
||||
req.on('end', () => { res.writeHead(calls <= 3 ? 500 : 200); res.end('{}'); });
|
||||
});
|
||||
const relay = createRelay();
|
||||
try {
|
||||
const flakyPort = await listen(flaky);
|
||||
const relayPort = await listen(relay);
|
||||
const started = Date.now();
|
||||
const created = await postJson(relayPort, '/deliveries', { url: `http://127.0.0.1:${flakyPort}/hook`, payload: { hello: 'world' } });
|
||||
record('accepts-delivery-202', created.status === 202 && created.body && typeof created.body.id === 'string');
|
||||
if (created.body && created.body.id) {
|
||||
const { record: rec, elapsedMs } = await waitForStatus(relayPort, created.body.id, 'delivered', 8000);
|
||||
record('delivered-after-retries', rec && rec.status === 'delivered' && calls >= 4);
|
||||
record('attempts-counted', rec && rec.attempts === 4);
|
||||
record('backoff-window-respected', rec && rec.status === 'delivered' && elapsedMs >= 250 && elapsedMs <= 5000 && Date.now() - started >= 250);
|
||||
} else {
|
||||
record('delivered-after-retries', false);
|
||||
record('attempts-counted', false);
|
||||
record('backoff-window-respected', false);
|
||||
}
|
||||
|
||||
// Probe group 2: a target that always fails -> dead after exactly 5 attempts.
|
||||
let deadCalls = 0;
|
||||
const deadEnd = http.createServer((req, res) => {
|
||||
deadCalls++;
|
||||
req.resume();
|
||||
req.on('end', () => { res.writeHead(500); res.end('{}'); });
|
||||
});
|
||||
const deadPort = await listen(deadEnd);
|
||||
const doomed = await postJson(relayPort, '/deliveries', { url: `http://127.0.0.1:${deadPort}/hook`, payload: { x: 1 } });
|
||||
if (doomed.body && doomed.body.id) {
|
||||
const { record: rec } = await waitForStatus(relayPort, doomed.body.id, 'dead', 15000);
|
||||
record('dead-after-retries-exhausted', rec && rec.status === 'dead');
|
||||
record('exactly-five-attempts', rec && rec.status === 'dead' && rec.attempts === 5 && deadCalls === 5);
|
||||
record('last-error-recorded', rec && rec.status === 'dead' && typeof rec.lastError === 'string' && rec.lastError.length > 0);
|
||||
} else {
|
||||
record('dead-after-retries-exhausted', false);
|
||||
record('exactly-five-attempts', false);
|
||||
record('last-error-recorded', false);
|
||||
}
|
||||
deadEnd.close();
|
||||
|
||||
// Probe 3: pre-existing API behavior is preserved.
|
||||
const missing = await fetch(`http://127.0.0.1:${relayPort}/deliveries/00000000-0000-0000-0000-000000000000`);
|
||||
record('unknown-id-still-404', missing.status === 404);
|
||||
|
||||
// Probe 4: concurrent deliveries all complete.
|
||||
let goodCalls = 0;
|
||||
const good = http.createServer((req, res) => {
|
||||
goodCalls++;
|
||||
req.resume();
|
||||
req.on('end', () => { res.writeHead(200); res.end('{}'); });
|
||||
});
|
||||
const goodPort = await listen(good);
|
||||
const batch = await Promise.all(Array.from({ length: 10 }, (_, i) =>
|
||||
postJson(relayPort, '/deliveries', { url: `http://127.0.0.1:${goodPort}/hook`, payload: { i } })));
|
||||
const settled = await Promise.all(batch.map(item => item.body && item.body.id
|
||||
? waitForStatus(relayPort, item.body.id, 'delivered', 10000).then(r => r.record && r.record.status === 'delivered')
|
||||
: false));
|
||||
record('concurrent-deliveries-complete', settled.every(Boolean) && goodCalls === 10);
|
||||
good.close();
|
||||
} catch { /* any grader-side failure leaves the missing checks unscored */ }
|
||||
finish();
|
||||
})();
|
||||
@@ -0,0 +1,33 @@
|
||||
# webhook-relay
|
||||
|
||||
In-memory webhook relay. Accepts delivery requests over HTTP and POSTs each
|
||||
payload to its destination URL, retrying failures with exponential backoff.
|
||||
|
||||
## HTTP API
|
||||
|
||||
- `POST /deliveries` — body `{ "url": string, "payload": any }`. Responds
|
||||
`202` with `{ "id" }` and delivers asynchronously. `400` for invalid JSON.
|
||||
- `GET /deliveries/:id` — `200` with
|
||||
`{ "id", "url", "status", "attempts", "lastError" }`, or `404`.
|
||||
`status` is `pending`, `delivered`, or `dead`.
|
||||
|
||||
## Delivery contract
|
||||
|
||||
- The payload is POSTed to `url` with `content-type: application/json`.
|
||||
- Any 2xx response means success: `status` becomes `delivered`.
|
||||
- Any other outcome (non-2xx, connection error, timeout) is a failure and is
|
||||
retried with exponential backoff: the first retry happens after about
|
||||
100ms and the delay doubles each retry. Up to 20% jitter in either
|
||||
direction is fine.
|
||||
- At most 5 attempts are made in total (the initial try plus 4 retries).
|
||||
- After the final failure the delivery becomes `dead` and `lastError`
|
||||
records a short description of the last failure.
|
||||
- `attempts` always reflects how many delivery attempts were made.
|
||||
|
||||
## Module contract
|
||||
|
||||
- `src/app.js` is CommonJS and exports `createRelay()`, which returns an
|
||||
`http.Server` that is not yet listening.
|
||||
- `node src/index.js <port>` starts the service.
|
||||
- No external dependencies; Node.js standard library only.
|
||||
- Run the tests with `npm test`.
|
||||
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"name": "webhook-relay",
|
||||
"private": true,
|
||||
"type": "commonjs",
|
||||
"scripts": { "test": "node --test test/" }
|
||||
}
|
||||
@@ -0,0 +1,51 @@
|
||||
'use strict';
|
||||
const http = require('node:http');
|
||||
const crypto = require('node:crypto');
|
||||
|
||||
// In-memory webhook relay. See README.md for the delivery contract.
|
||||
//
|
||||
// TODO: deliveries are accepted and stored, but the delivery worker was never
|
||||
// finished — nothing ever POSTs to the destination URL, retries never happen,
|
||||
// and records stay "pending" forever.
|
||||
|
||||
function createRelay() {
|
||||
const deliveries = new Map();
|
||||
|
||||
const server = http.createServer((req, res) => {
|
||||
if (req.method === 'POST' && req.url === '/deliveries') {
|
||||
let body = '';
|
||||
req.on('data', chunk => { body += chunk; });
|
||||
req.on('end', () => {
|
||||
let parsed;
|
||||
try { parsed = JSON.parse(body); } catch {
|
||||
res.writeHead(400, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify({ error: 'invalid JSON body' }));
|
||||
return;
|
||||
}
|
||||
const id = crypto.randomUUID();
|
||||
deliveries.set(id, { id, url: parsed.url, payload: parsed.payload,
|
||||
status: 'pending', attempts: 0, lastError: null });
|
||||
res.writeHead(202, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify({ id }));
|
||||
});
|
||||
return;
|
||||
}
|
||||
const match = /^\/deliveries\/([0-9a-f-]+)$/.exec(req.url || '');
|
||||
if (req.method === 'GET' && match) {
|
||||
const record = deliveries.get(match[1]);
|
||||
if (!record) {
|
||||
res.writeHead(404, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify({ error: 'not found' }));
|
||||
return;
|
||||
}
|
||||
res.writeHead(200, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify(record));
|
||||
return;
|
||||
}
|
||||
res.writeHead(404, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify({ error: 'not found' }));
|
||||
});
|
||||
return server;
|
||||
}
|
||||
|
||||
module.exports = { createRelay };
|
||||
@@ -0,0 +1,7 @@
|
||||
'use strict';
|
||||
const { createRelay } = require('./app');
|
||||
|
||||
const port = Number(process.argv[2] || 8080);
|
||||
createRelay().listen(port, () => {
|
||||
console.log(`webhook-relay listening on ${port}`);
|
||||
});
|
||||
@@ -0,0 +1,41 @@
|
||||
'use strict';
|
||||
const test = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const { createRelay } = require('../src/app');
|
||||
|
||||
function listen(server) {
|
||||
return new Promise((resolve, reject) => {
|
||||
server.once('error', reject);
|
||||
server.listen(0, '127.0.0.1', () => resolve(server.address().port));
|
||||
});
|
||||
}
|
||||
|
||||
test('accepts a delivery and reports it as pending', async () => {
|
||||
const server = createRelay();
|
||||
const port = await listen(server);
|
||||
try {
|
||||
const created = await fetch(`http://127.0.0.1:${port}/deliveries`, {
|
||||
method: 'POST', headers: { 'content-type': 'application/json' },
|
||||
body: JSON.stringify({ url: 'http://127.0.0.1:1/hook', payload: { a: 1 } }) });
|
||||
assert.equal(created.status, 202);
|
||||
const { id } = await created.json();
|
||||
const status = await fetch(`http://127.0.0.1:${port}/deliveries/${id}`);
|
||||
assert.equal(status.status, 200);
|
||||
const record = await status.json();
|
||||
assert.equal(record.status, 'pending');
|
||||
assert.equal(record.attempts, 0);
|
||||
} finally {
|
||||
server.close();
|
||||
}
|
||||
});
|
||||
|
||||
test('unknown delivery id returns 404', async () => {
|
||||
const server = createRelay();
|
||||
const port = await listen(server);
|
||||
try {
|
||||
const response = await fetch(`http://127.0.0.1:${port}/deliveries/00000000-0000-0000-0000-000000000000`);
|
||||
assert.equal(response.status, 404);
|
||||
} finally {
|
||||
server.close();
|
||||
}
|
||||
});
|
||||
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"id": "webhook-relay",
|
||||
"category": "feature-build",
|
||||
"manualIds": ["skill:tdd-workflow"],
|
||||
"checkTimeoutMs": 60000,
|
||||
"selection": {
|
||||
"id": "complex-webhook-relay",
|
||||
"category": "complex-feature-build",
|
||||
"expectedIds": ["skill:tdd-workflow"]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
The webhook relay in this repo accepts delivery requests but never actually sends them — the delivery worker was never finished, and customers are losing notifications. Implement asynchronous delivery per the README: POST each delivery's JSON payload to its URL, retry failures with exponential backoff starting around 100ms and doubling each time, give up after 5 total attempts and mark the delivery dead. Keep the documented module contract, make `npm test` pass, and extend the test suite to cover the retry and dead-letter behavior.
|
||||
@@ -0,0 +1,149 @@
|
||||
'use strict';
|
||||
// Hidden grader for event-stats-api: independent spec-conformant aggregation
|
||||
// over the deterministic event log, plus a measured 2,000-query performance
|
||||
// probe (threshold calibrated on the grading machine: shipped naive ~7.7s,
|
||||
// reference ~1.5s). Prints ECC_EVAL_SCORE and always exits 0.
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const checks = [];
|
||||
const record = (name, ok) => checks.push({ name, ok: Boolean(ok) });
|
||||
let finished = false;
|
||||
function finish() {
|
||||
if (finished) return;
|
||||
finished = true;
|
||||
const ok = checks.filter(c => c.ok).length;
|
||||
for (const c of checks) console.log(`${c.ok ? 'ok' : 'not ok'} - ${c.name}`);
|
||||
console.log(`ECC_EVAL_SCORE ${JSON.stringify({ score: ok / checks.length, passed: ok, total: checks.length })}`);
|
||||
process.exit(0);
|
||||
}
|
||||
setTimeout(finish, 110000).unref();
|
||||
|
||||
const PERF_THRESHOLD_MS = 6000;
|
||||
const PERF_QUERIES = 2000;
|
||||
|
||||
function lcg(seed) {
|
||||
let state = seed >>> 0;
|
||||
return () => {
|
||||
state = (Math.imul(state, 1664525) + 1013904223) >>> 0;
|
||||
return state / 2 ** 32;
|
||||
};
|
||||
}
|
||||
|
||||
const root = process.cwd();
|
||||
const { events, TYPES, EPOCH_MS, SPAN_MS } = require(path.join(root, 'src', 'data.js'));
|
||||
|
||||
// Independent reference semantics per the README: inclusive bounds,
|
||||
// nearest-rank percentiles, half-up two-decimal average via exact integer math.
|
||||
function expected(type, from, to) {
|
||||
const rows = events
|
||||
.filter(e => e.type === type && (from === null || e.ts >= from) && (to === null || e.ts <= to))
|
||||
.map(e => e.value)
|
||||
.sort((a, b) => a - b);
|
||||
const count = rows.length;
|
||||
if (!count) return { count: 0, sum: 0, avg: null, p50: null, p95: null, p99: null, min: null, max: null };
|
||||
const sum = rows.reduce((a, b) => a + b, 0);
|
||||
const rank = p => rows[Math.ceil((p / 100) * count) - 1];
|
||||
const avgCents = Math.floor((sum * 200 + count) / (count * 2));
|
||||
return { count, sum, avg: avgCents / 100,
|
||||
p50: rank(50), p95: rank(95), p99: rank(99), min: rows[0], max: rows[count - 1] };
|
||||
}
|
||||
|
||||
const same = (a, b) => JSON.stringify(a) === JSON.stringify(b);
|
||||
|
||||
async function query(port, params) {
|
||||
const qs = Object.entries(params).map(([k, v]) => `${k}=${v}`).join('&');
|
||||
const response = await fetch(`http://127.0.0.1:${port}/stats?${qs}`);
|
||||
return { status: response.status, body: await response.json().catch(() => null) };
|
||||
}
|
||||
|
||||
(async () => {
|
||||
let createApp;
|
||||
try { ({ createApp } = require(path.join(root, 'src', 'app.js'))); } catch { finish(); return; }
|
||||
if (typeof createApp !== 'function') { finish(); return; }
|
||||
|
||||
try {
|
||||
const app = createApp();
|
||||
await new Promise(resolve => app.listen(0, '127.0.0.1', resolve));
|
||||
const port = app.address().port;
|
||||
|
||||
// 1-2: broad and full-range queries with independently computed expectations.
|
||||
const broadFrom = EPOCH_MS;
|
||||
const broadTo = EPOCH_MS + 30 * 86400000;
|
||||
const broad = await query(port, { type: 'click', from: broadFrom, to: broadTo });
|
||||
record('broad-window-exact', broad.status === 200
|
||||
&& same(broad.body, { type: 'click', from: broadFrom, to: broadTo, ...expected('click', broadFrom, broadTo) }));
|
||||
const full = await query(port, { type: 'purchase' });
|
||||
record('full-range-exact', full.status === 200
|
||||
&& same(full.body, { type: 'purchase', from: null, to: null, ...expected('purchase', null, null) }));
|
||||
|
||||
// 3: nearest-rank vs interpolation is distinguishable on a tiny window.
|
||||
const exportEvents = events.filter(e => e.type === 'export').map(e => e.ts).sort((a, b) => a - b);
|
||||
const pivot = exportEvents[Math.floor(exportEvents.length / 2)];
|
||||
const narrowFrom = pivot - 1;
|
||||
const narrowTo = pivot + 1;
|
||||
const narrow = await query(port, { type: 'export', from: narrowFrom, to: narrowTo });
|
||||
record('narrow-window-nearest-rank', narrow.status === 200
|
||||
&& same(narrow.body, { type: 'export', from: narrowFrom, to: narrowTo, ...expected('export', narrowFrom, narrowTo) }));
|
||||
|
||||
// 4-5: empty range and unknown type return nulls, not zeros or errors.
|
||||
const beyond = await query(port, { type: 'click', from: EPOCH_MS + 200 * 86400000, to: EPOCH_MS + 201 * 86400000 });
|
||||
record('empty-range-nulls', beyond.status === 200 && same(beyond.body,
|
||||
{ type: 'click', from: EPOCH_MS + 200 * 86400000, to: EPOCH_MS + 201 * 86400000, ...expected('click', EPOCH_MS + 200 * 86400000, EPOCH_MS + 201 * 86400000) }));
|
||||
const unknown = await query(port, { type: 'nope' });
|
||||
record('unknown-type-nulls', unknown.status === 200
|
||||
&& same(unknown.body, { type: 'nope', from: null, to: null, ...expected('nope', null, null) }));
|
||||
|
||||
// 6: inclusive bounds — a zero-width window on a real timestamp includes it.
|
||||
const likeTs = events.filter(e => e.type === 'like').map(e => e.ts).sort((a, b) => a - b)[100];
|
||||
const inclusive = await query(port, { type: 'like', from: likeTs, to: likeTs });
|
||||
record('bounds-inclusive', inclusive.status === 200 && inclusive.body.count === expected('like', likeTs, likeTs).count && inclusive.body.count >= 1);
|
||||
|
||||
// 7: average rounding follows half-up two decimals exactly.
|
||||
const rounding = expected('view', EPOCH_MS, EPOCH_MS + 86400000);
|
||||
const rounded = await query(port, { type: 'view', from: EPOCH_MS, to: EPOCH_MS + 86400000 });
|
||||
record('avg-half-up-2dp', rounded.status === 200 && rounded.body.avg === rounding.avg);
|
||||
|
||||
// 8-9: invalid parameters are 400.
|
||||
const inverted = await query(port, { type: 'click', from: 10, to: 5 });
|
||||
record('inverted-bounds-400', inverted.status === 400);
|
||||
const garbage = await query(port, { type: 'click', from: 'abc' });
|
||||
record('non-numeric-bounds-400', garbage.status === 400);
|
||||
|
||||
// 10: performance budget.
|
||||
const rand = lcg(777);
|
||||
const queries = [];
|
||||
for (let i = 0; i < PERF_QUERIES; i++) {
|
||||
const type = TYPES[Math.floor(rand() * TYPES.length)];
|
||||
const start = EPOCH_MS + Math.floor(rand() * SPAN_MS * 0.7);
|
||||
queries.push({ type, from: start, to: start + Math.floor(rand() * SPAN_MS * 0.5) });
|
||||
}
|
||||
const started = Date.now();
|
||||
for (let i = 0; i < queries.length; i += 20) {
|
||||
await Promise.all(queries.slice(i, i + 20).map(q => query(port, q)));
|
||||
}
|
||||
const elapsed = Date.now() - started;
|
||||
console.log(`perf: ${elapsed}ms for ${PERF_QUERIES} queries (threshold ${PERF_THRESHOLD_MS}ms)`);
|
||||
record('performance-budget', elapsed < PERF_THRESHOLD_MS);
|
||||
|
||||
app.close();
|
||||
} catch { /* grader-side failure leaves remaining checks unscored */ }
|
||||
|
||||
// 11: no external dependencies.
|
||||
try {
|
||||
const pkg = JSON.parse(fs.readFileSync(path.join(root, 'package.json'), 'utf8'));
|
||||
const sources = [];
|
||||
const walk = directory => {
|
||||
for (const entry of fs.readdirSync(directory, { withFileTypes: true })) {
|
||||
const item = path.join(directory, entry.name);
|
||||
if (entry.isDirectory()) walk(item);
|
||||
else if (entry.name.endsWith('.js')) sources.push(fs.readFileSync(item, 'utf8'));
|
||||
}
|
||||
};
|
||||
walk(path.join(root, 'src'));
|
||||
const bareImport = sources.some(source => /require\(\s*['"](?!node:)[a-z@][^'./]*['"]\s*\)/.test(source));
|
||||
record('no-external-dependencies', !bareImport && !pkg.dependencies && !pkg.devDependencies);
|
||||
} catch { record('no-external-dependencies', false); }
|
||||
|
||||
finish();
|
||||
})();
|
||||
@@ -0,0 +1,41 @@
|
||||
# event-stats
|
||||
|
||||
Analytics endpoint over an in-memory event log (300,000 events, generated
|
||||
deterministically by `src/data.js`).
|
||||
|
||||
## API
|
||||
|
||||
`GET /stats?type=<t>&from=<ms>&to=<ms>` returns JSON:
|
||||
|
||||
```json
|
||||
{ "type": "click", "from": 1754000000000, "to": 1756592000000,
|
||||
"count": 1234, "sum": 56789, "avg": 46.02,
|
||||
"p50": 123, "p95": 456, "p99": 789, "min": 1, "max": 50000 }
|
||||
```
|
||||
|
||||
Semantics (all pinned; follow them exactly):
|
||||
|
||||
- `from`/`to` are millisecond timestamps, **inclusive**, and optional
|
||||
(absent means unbounded). Non-numeric bounds, or `from > to`, are `400`.
|
||||
- Only events of the given `type` within `[from, to]` are included.
|
||||
- `sum` is the exact integer sum of `value`s.
|
||||
- `avg` is `sum / count` rounded **half-up to two decimals**.
|
||||
- Percentiles use the **nearest-rank** method: sort values ascending, take the
|
||||
value at 1-based rank `ceil(p / 100 * count)`. No interpolation.
|
||||
- If no events match (including an unknown `type`), return `200` with
|
||||
`count: 0, sum: 0` and `avg`, `p50`, `p95`, `p99`, `min`, `max` all `null`.
|
||||
- The response echoes the effective `from`/`to` (`null` when unbounded).
|
||||
|
||||
## Performance requirement
|
||||
|
||||
The endpoint must stay fast at this data size: **2,000 mixed queries complete
|
||||
in under 6 seconds** on this machine (the reference does it in ~1.5s).
|
||||
Precompute whatever you need at startup; per-query work must not scan the
|
||||
whole log.
|
||||
|
||||
## Module contract
|
||||
|
||||
- `src/app.js` is CommonJS and exports `createApp()` returning an
|
||||
`http.Server` that is not yet listening.
|
||||
- `node src/index.js <port>` starts the service.
|
||||
- No external dependencies. Run the tests with `npm test`.
|
||||
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"name": "event-stats",
|
||||
"private": true,
|
||||
"type": "commonjs",
|
||||
"scripts": { "test": "node --test test/" }
|
||||
}
|
||||
@@ -0,0 +1,43 @@
|
||||
'use strict';
|
||||
const http = require('node:http');
|
||||
const { events } = require('./data');
|
||||
|
||||
// Current implementation: scan and sort per query. Known slow, and the
|
||||
// analytics team says edge cases don't match the README semantics.
|
||||
function summarize(type, from, to) {
|
||||
const rows = events
|
||||
.filter(e => e.type === type && (from === null || e.ts >= from) && (to === null || e.ts <= to))
|
||||
.map(e => e.value)
|
||||
.sort((a, b) => a - b);
|
||||
const count = rows.length;
|
||||
const sum = rows.reduce((a, b) => a + b, 0);
|
||||
const interpolate = p => {
|
||||
if (!count) return 0;
|
||||
const rank = (p / 100) * (count - 1);
|
||||
const low = Math.floor(rank);
|
||||
const high = Math.ceil(rank);
|
||||
return rows[low] + (rows[high] - rows[low]) * (rank - low);
|
||||
};
|
||||
return { count, sum, avg: count ? sum / count : 0,
|
||||
p50: interpolate(50), p95: interpolate(95), p99: interpolate(99),
|
||||
min: count ? rows[0] : 0, max: count ? rows[count - 1] : 0 };
|
||||
}
|
||||
|
||||
function createApp() {
|
||||
return http.createServer((req, res) => {
|
||||
const url = new URL(req.url, 'http://localhost');
|
||||
if (req.method === 'GET' && url.pathname === '/stats') {
|
||||
const type = url.searchParams.get('type');
|
||||
const from = url.searchParams.has('from') ? Number(url.searchParams.get('from')) : null;
|
||||
const to = url.searchParams.has('to') ? Number(url.searchParams.get('to')) : null;
|
||||
const body = summarize(type, from, to);
|
||||
res.writeHead(200, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify({ type, from, to, ...body }));
|
||||
return;
|
||||
}
|
||||
res.writeHead(404, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify({ error: 'not found' }));
|
||||
});
|
||||
}
|
||||
|
||||
module.exports = { createApp };
|
||||
@@ -0,0 +1,28 @@
|
||||
'use strict';
|
||||
// Deterministic event log: 300,000 events from a seeded LCG so every run,
|
||||
// grader, and reference sees identical data. Do not change the generator.
|
||||
const TYPES = ['click', 'view', 'signup', 'purchase', 'refund', 'login',
|
||||
'logout', 'share', 'comment', 'like', 'search', 'export'];
|
||||
const DAY_MS = 86400000;
|
||||
const EPOCH_MS = 1754000000000;
|
||||
const SPAN_MS = 90 * DAY_MS;
|
||||
|
||||
function lcg(seed) {
|
||||
let state = seed >>> 0;
|
||||
return () => {
|
||||
state = (Math.imul(state, 1664525) + 1013904223) >>> 0;
|
||||
return state / 2 ** 32;
|
||||
};
|
||||
}
|
||||
|
||||
const rand = lcg(20260925);
|
||||
const events = new Array(300000);
|
||||
for (let i = 0; i < events.length; i++) {
|
||||
events[i] = {
|
||||
type: TYPES[Math.floor(rand() * TYPES.length)],
|
||||
ts: EPOCH_MS + Math.floor(rand() * SPAN_MS),
|
||||
value: Math.floor(rand() * 50000) + 1,
|
||||
};
|
||||
}
|
||||
|
||||
module.exports = { events, TYPES, EPOCH_MS, SPAN_MS };
|
||||
@@ -0,0 +1,7 @@
|
||||
'use strict';
|
||||
const { createApp } = require('./app');
|
||||
|
||||
const port = Number(process.argv[2] || 8080);
|
||||
createApp().listen(port, () => {
|
||||
console.log(`event-stats listening on ${port}`);
|
||||
});
|
||||
@@ -0,0 +1,20 @@
|
||||
'use strict';
|
||||
const test = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const { createApp } = require('../src/app');
|
||||
const { EPOCH_MS } = require('../src/data');
|
||||
|
||||
test('stats endpoint answers a broad query', async () => {
|
||||
const server = createApp();
|
||||
await new Promise(resolve => server.listen(0, '127.0.0.1', resolve));
|
||||
try {
|
||||
const port = server.address().port;
|
||||
const response = await fetch(`http://127.0.0.1:${port}/stats?type=click&from=${EPOCH_MS}&to=${EPOCH_MS + 30 * 86400000}`);
|
||||
assert.equal(response.status, 200);
|
||||
const body = await response.json();
|
||||
assert.equal(body.type, 'click');
|
||||
assert.ok(body.count > 0);
|
||||
} finally {
|
||||
server.close();
|
||||
}
|
||||
});
|
||||
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"id": "event-stats-api",
|
||||
"category": "correctness-and-performance",
|
||||
"manualIds": ["skill:backend-patterns"],
|
||||
"checkTimeoutMs": 120000,
|
||||
"selection": {
|
||||
"id": "complex-event-stats-api",
|
||||
"category": "complex-correctness-performance",
|
||||
"expectedIds": ["skill:backend-patterns"]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
The /stats endpoint in this repo is wrong on edge cases and too slow — customers on big dashboards are timing out. It currently rescans and resorts the whole 300k-event log on every request, and the analytics team says the numbers don't match the documented semantics (nearest-rank percentiles, half-up two-decimal averages, null fields when nothing matches, proper 400s). Make it correct per the README and fast enough to meet the documented performance budget, without changing the API shape. `npm test` must stay green.
|
||||
@@ -0,0 +1,132 @@
|
||||
'use strict';
|
||||
// Hidden grader for forge-cli: drives run(argv, state) through the twelve
|
||||
// contractual behaviors plus never-throw fuzzing and static hygiene. Prints
|
||||
// ECC_EVAL_SCORE and always exits 0.
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const checks = [];
|
||||
const record = (name, ok) => checks.push({ name, ok: Boolean(ok) });
|
||||
|
||||
const root = process.cwd();
|
||||
let run;
|
||||
try { ({ run } = require(path.join(root, 'src', 'cli.js'))); } catch { /* scored below */ }
|
||||
|
||||
const USAGE = 'usage: snippet <add|get|list|remove|search|export|import>\n';
|
||||
const ADD_USAGE = 'usage: add <name> [--tags t1,t2] <text...>\n';
|
||||
|
||||
if (typeof run !== 'function') {
|
||||
for (let i = 0; i < 26; i++) record(`check-${i + 1}`, false);
|
||||
} else {
|
||||
const call = (argv, state) => {
|
||||
try {
|
||||
const result = run(argv, state);
|
||||
if (!result || typeof result.code !== 'number'
|
||||
|| typeof result.stdout !== 'string' || typeof result.stderr !== 'string') return null;
|
||||
return result;
|
||||
} catch { return null; }
|
||||
};
|
||||
|
||||
// Basic lifecycle.
|
||||
let s = {};
|
||||
let r = call(['add', 'hello', 'hello', 'world'], s);
|
||||
record('add-happy', r && r.code === 0 && r.stdout === 'created hello\n' && r.stderr === '');
|
||||
r = call(['add', 'hello', 'different', 'text'], s);
|
||||
const afterDup = call(['get', 'hello'], s);
|
||||
record('add-duplicate-rejected', r && r.code === 1 && r.stderr === "error: snippet 'hello' already exists\n"
|
||||
&& afterDup && afterDup.stdout === 'hello world\n');
|
||||
const m1 = call(['add'], s);
|
||||
const m2 = call(['add', 'justname'], s);
|
||||
record('add-missing-args-usage', m1 && m1.code === 2 && m1.stderr === ADD_USAGE
|
||||
&& m2 && m2.code === 2 && m2.stderr === ADD_USAGE);
|
||||
r = call(['add', 'Bad_Name', 'text'], s);
|
||||
record('invalid-name-rejected', r && r.code === 2 && r.stderr === "error: invalid snippet name 'Bad_Name'\n");
|
||||
r = call(['get', 'hello'], s);
|
||||
record('get-happy', r && r.code === 0 && r.stdout === 'hello world\n');
|
||||
r = call(['get', 'ghost'], s);
|
||||
record('get-unknown', r && r.code === 2 && r.stderr === "error: no snippet named 'ghost'\n");
|
||||
|
||||
// Listing and tags.
|
||||
s = {};
|
||||
call(['add', 'bravo', 'second'], s);
|
||||
call(['add', 'alpha', '--tags', 'x,y', 'first'], s);
|
||||
call(['add', 'charlie', '--tags', 'y', 'third'], s);
|
||||
r = call(['list'], s);
|
||||
record('list-sorted', r && r.code === 0 && r.stdout === 'alpha\nbravo\ncharlie\n');
|
||||
r = call(['list'], {});
|
||||
record('list-empty', r && r.code === 0 && r.stdout === 'no snippets\n');
|
||||
r = call(['list', '--tag', 'y'], s);
|
||||
record('list-tag-filter', r && r.code === 0 && r.stdout === 'alpha\ncharlie\n');
|
||||
|
||||
// Removal.
|
||||
r = call(['remove', 'bravo'], s);
|
||||
const gone = call(['get', 'bravo'], s);
|
||||
record('remove-happy', r && r.code === 0 && r.stdout === 'removed bravo\n' && gone && gone.code === 2);
|
||||
r = call(['remove', 'bravo'], s);
|
||||
record('remove-unknown', r && r.code === 2 && r.stderr === "error: no snippet named 'bravo'\n");
|
||||
|
||||
// Search over name and text, case-insensitive, sorted.
|
||||
r = call(['search', 'FIRST'], s);
|
||||
record('search-text-case-insensitive', r && r.code === 0 && r.stdout === 'alpha\n');
|
||||
r = call(['search', 'char'], s);
|
||||
record('search-name-match', r && r.code === 0 && r.stdout === 'charlie\n');
|
||||
r = call(['search', 'zzz'], s);
|
||||
record('search-no-matches', r && r.code === 0 && r.stdout === 'no matches\n');
|
||||
|
||||
// Export/import round-trip with stable ordering.
|
||||
r = call(['export'], s);
|
||||
let doc = null;
|
||||
try { doc = r && JSON.parse(r.stdout); } catch { /* wrong */ }
|
||||
record('export-json-sorted', doc && r.code === 0 && sameDoc(doc, {
|
||||
snippets: { alpha: { text: 'first', tags: ['x', 'y'] }, charlie: { text: 'third', tags: ['y'] } } })
|
||||
&& r.stdout.indexOf('alpha') < r.stdout.indexOf('charlie'));
|
||||
const importedState = { snippets: { alpha: { text: 'preexisting', tags: [] } } };
|
||||
r = call(['import', JSON.stringify({ snippets: {
|
||||
alpha: { text: 'first', tags: ['x', 'y'] }, delta: { text: 'fourth', tags: ['z'] } } })], importedState);
|
||||
const delta = call(['get', 'delta'], importedState);
|
||||
const alpha = call(['get', 'alpha'], importedState);
|
||||
record('import-merge-skip-existing', r && r.code === 0 && r.stdout === 'imported 1, skipped 1\n'
|
||||
&& delta && delta.stdout === 'fourth\n' && alpha && alpha.stdout === 'preexisting\n');
|
||||
const beforeExport = call(['export'], s);
|
||||
r = call(['import', '{not json'], s);
|
||||
const afterExport = call(['export'], s);
|
||||
record('import-malformed-atomic', r && r.code === 1 && r.stderr === 'error: invalid JSON\n'
|
||||
&& beforeExport && afterExport && beforeExport.stdout === afterExport.stdout);
|
||||
|
||||
// Usage fallbacks.
|
||||
r = call(['bogus'], {});
|
||||
record('unknown-command-usage', r && r.code === 2 && r.stderr === USAGE);
|
||||
r = call([], {});
|
||||
record('no-command-usage', r && r.code === 2 && r.stderr === USAGE);
|
||||
|
||||
// Never-throw fuzzing on junk input.
|
||||
const fuzz = [['--help', 'x'], ['get'], ['add', 'x', 'y', '--tags'], ['import']];
|
||||
fuzz.forEach((argv, index) => {
|
||||
record(`fuzz-never-throws-${index + 1}`, call(argv, {}) !== null);
|
||||
});
|
||||
}
|
||||
|
||||
function sameDoc(a, b) { return JSON.stringify(a) === JSON.stringify(b); }
|
||||
|
||||
// Static hygiene.
|
||||
try {
|
||||
const pkg = JSON.parse(fs.readFileSync(path.join(root, 'package.json'), 'utf8'));
|
||||
record('no-external-dependencies', !pkg.dependencies && !pkg.devDependencies);
|
||||
} catch { record('no-external-dependencies', false); }
|
||||
try {
|
||||
const sources = [];
|
||||
const walk = directory => {
|
||||
for (const entry of fs.readdirSync(directory, { withFileTypes: true })) {
|
||||
const item = path.join(directory, entry.name);
|
||||
if (entry.isDirectory()) walk(item);
|
||||
else if (entry.name.endsWith('.js')) sources.push(fs.readFileSync(item, 'utf8'));
|
||||
}
|
||||
};
|
||||
walk(path.join(root, 'src'));
|
||||
record('no-leftover-todos', sources.every(source => !/TODO|FIXME/.test(source)));
|
||||
} catch { record('no-leftover-todos', false); }
|
||||
|
||||
const okCount = checks.filter(c => c.ok).length;
|
||||
for (const c of checks) console.log(`${c.ok ? 'ok' : 'not ok'} - ${c.name}`);
|
||||
console.log(`ECC_EVAL_SCORE ${JSON.stringify({ score: okCount / checks.length, passed: okCount, total: checks.length })}`);
|
||||
process.exit(0);
|
||||
@@ -0,0 +1,46 @@
|
||||
# snippet-cli
|
||||
|
||||
A small in-process snippet manager. No external dependencies; Node.js standard
|
||||
library only.
|
||||
|
||||
## Contract
|
||||
|
||||
`src/cli.js` is CommonJS and exports `run(argv, state)`:
|
||||
|
||||
- `argv`: array of command-line words (already split, no program name).
|
||||
- `state`: any plain object, created by the caller as `{}`. The CLI keeps its
|
||||
data in it and mutates it in place; it survives across calls.
|
||||
- Returns synchronously: `{ code, stdout, stderr }` — a number and two strings
|
||||
(empty string when there is nothing to print). `run` must **never throw**,
|
||||
on any input.
|
||||
- All printed lines end with `\n`.
|
||||
|
||||
## Commands (all behavior below is contractual)
|
||||
|
||||
1. `add <name> [--tags a,b] <text...>` — creates a snippet from the remaining
|
||||
words joined by single spaces. Prints `created <name>`, code 0.
|
||||
2. Adding an existing name: code 1, stderr `error: snippet '<name>' already exists`,
|
||||
state unchanged.
|
||||
3. `add` with a missing name or missing text: code 2, stderr
|
||||
`usage: add <name> [--tags t1,t2] <text...>`.
|
||||
4. Names must match `^[a-z0-9][a-z0-9-]*$`; otherwise code 2, stderr
|
||||
`error: invalid snippet name '<name>'`.
|
||||
5. `get <name>` — prints the exact text, code 0. Unknown name: code 2, stderr
|
||||
`error: no snippet named '<name>'`.
|
||||
6. `remove <name>` — prints `removed <name>`, code 0. Unknown name: same as `get`.
|
||||
7. `list` — every snippet name, sorted ascending, one per line. With no
|
||||
snippets: prints `no snippets`. Always code 0.
|
||||
8. `list --tag <t>` — only snippets whose tags include `t`.
|
||||
9. `search <term>` — case-insensitive substring match over name **and** text;
|
||||
prints matching names sorted, one per line; prints `no matches` when empty.
|
||||
Code 0.
|
||||
10. `export` — prints `JSON.stringify` of `{ snippets: { <name>: { text, tags } } }`
|
||||
with names sorted and each `tags` array sorted. Code 0.
|
||||
11. `import <json>` — merges an exported document: names not already present
|
||||
are added, existing names are skipped. Prints `imported <N>, skipped <M>`,
|
||||
code 0. Malformed JSON: code 1, stderr `error: invalid JSON`, state
|
||||
unchanged.
|
||||
12. No command or an unknown command: code 2, stderr
|
||||
`usage: snippet <add|get|list|remove|search|export|import>`.
|
||||
|
||||
Run the tests with `npm test`.
|
||||
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"name": "snippet-cli",
|
||||
"private": true,
|
||||
"type": "commonjs",
|
||||
"scripts": { "test": "node --test test/" }
|
||||
}
|
||||
@@ -0,0 +1,8 @@
|
||||
'use strict';
|
||||
|
||||
// TODO: implement per README. The contract is run(argv, state) -> { code, stdout, stderr }.
|
||||
function run(_argv, _state) {
|
||||
throw new Error('not implemented');
|
||||
}
|
||||
|
||||
module.exports = { run };
|
||||
@@ -0,0 +1,20 @@
|
||||
'use strict';
|
||||
const test = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const { run } = require('../src/cli');
|
||||
|
||||
test('add then get round-trips a snippet', () => {
|
||||
const state = {};
|
||||
const added = run(['add', 'hello', 'hello', 'world'], state);
|
||||
assert.equal(added.code, 0);
|
||||
assert.equal(added.stdout, 'created hello\n');
|
||||
const got = run(['get', 'hello'], state);
|
||||
assert.equal(got.code, 0);
|
||||
assert.equal(got.stdout, 'hello world\n');
|
||||
});
|
||||
|
||||
test('list on empty state', () => {
|
||||
const result = run(['list'], {});
|
||||
assert.equal(result.code, 0);
|
||||
assert.equal(result.stdout, 'no snippets\n');
|
||||
});
|
||||
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"id": "forge-cli",
|
||||
"category": "spec-thoroughness",
|
||||
"manualIds": ["skill:tdd-workflow"],
|
||||
"checkTimeoutMs": 30000,
|
||||
"selection": {
|
||||
"id": "complex-forge-cli",
|
||||
"category": "complex-spec-thoroughness",
|
||||
"expectedIds": ["skill:tdd-workflow"]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
Build the snippet manager CLI per the README — all twelve numbered behaviors are contractual, including exact messages, exit codes, sorting, and the never-throw guarantee. `npm test` must pass, and add tests for the tricky edges (duplicates, invalid names, bad imports) so we don't regress them.
|
||||
@@ -0,0 +1,63 @@
|
||||
'use strict';
|
||||
// Hidden grader for keccak-selector. Every vector is independently cross-checked:
|
||||
// the implementation is validated against Node's SHA3-256 (same Keccak-f[1600]
|
||||
// permutation, different padding suffix) including multi-block and q=1 padding
|
||||
// edge inputs. Prints ECC_EVAL_SCORE and always exits 0.
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const checks = [];
|
||||
const record = (name, ok) => checks.push({ name, ok: Boolean(ok) });
|
||||
|
||||
const VECTORS = [
|
||||
['name()', '0x06fdde03'],
|
||||
['symbol()', '0x95d89b41'],
|
||||
['decimals()', '0x313ce567'],
|
||||
['totalSupply()', '0x18160ddd'],
|
||||
['balanceOf(address)', '0x70a08231'],
|
||||
['transfer(address,uint256)', '0xa9059cbb'],
|
||||
['approve(address,uint256)', '0x095ea7b3'],
|
||||
['transferFrom(address,address,uint256)', '0x23b872dd'],
|
||||
// 135-byte signature: padding lands on the q=1 edge case.
|
||||
['someVeryLongFunctionNameForTestingMultiBlockHashingBehavior(address,uint256,string,bytes32,bool,uint8[],int128,(address,uint256),bytes)', '0x2add16ac'],
|
||||
];
|
||||
|
||||
let functionSelector;
|
||||
try { ({ functionSelector } = require(path.join(process.cwd(), 'src', 'selector.js'))); } catch { /* scored below */ }
|
||||
|
||||
if (typeof functionSelector === 'function') {
|
||||
VECTORS.forEach(([signature, expected], index) => {
|
||||
let actual = null;
|
||||
try { actual = functionSelector(signature); } catch { /* wrong */ }
|
||||
record(`selector-vector-${index + 1}`, actual === expected);
|
||||
});
|
||||
try { record('output-format', /^0x[0-9a-f]{8}$/.test(functionSelector('name()'))); }
|
||||
catch { record('output-format', false); }
|
||||
let threw = false;
|
||||
try { functionSelector(42); } catch (error) { threw = error instanceof TypeError; }
|
||||
record('typeerror-on-non-string', threw);
|
||||
} else {
|
||||
for (const [,] of VECTORS) checks.push({ name: `selector-vector-${checks.length + 1}`, ok: false });
|
||||
record('output-format', false);
|
||||
record('typeerror-on-non-string', false);
|
||||
}
|
||||
|
||||
// No external code: every import under src/ must be relative or node:-prefixed.
|
||||
const sources = [];
|
||||
const walk = directory => {
|
||||
for (const entry of fs.readdirSync(directory, { withFileTypes: true })) {
|
||||
const item = path.join(directory, entry.name);
|
||||
if (entry.isDirectory()) walk(item);
|
||||
else if (entry.name.endsWith('.js')) sources.push(fs.readFileSync(item, 'utf8'));
|
||||
}
|
||||
};
|
||||
try { walk(path.join(process.cwd(), 'src')); } catch { /* none */ }
|
||||
const bareImport = sources.some(source => /require\(\s*['"](?!node:)[a-z@][^'./]*['"]\s*\)/.test(source)
|
||||
|| /^\s*import\s/m.test(source) && /from\s*['"](?!node:|\.)[^'"]+['"]/.test(source));
|
||||
const pkg = JSON.parse(fs.readFileSync(path.join(process.cwd(), 'package.json'), 'utf8'));
|
||||
record('no-external-dependencies', !bareImport && !pkg.dependencies && !pkg.devDependencies);
|
||||
|
||||
const ok = checks.filter(c => c.ok).length;
|
||||
for (const c of checks) console.log(`${c.ok ? 'ok' : 'not ok'} - ${c.name}`);
|
||||
console.log(`ECC_EVAL_SCORE ${JSON.stringify({ score: ok / checks.length, passed: ok, total: checks.length })}`);
|
||||
process.exit(0);
|
||||
@@ -0,0 +1,21 @@
|
||||
# abi-selectors
|
||||
|
||||
Contract ABI tooling: compute Ethereum function selectors.
|
||||
|
||||
## Contract
|
||||
|
||||
`src/selector.js` is CommonJS and exports `functionSelector(signature)`:
|
||||
|
||||
- `signature` is the canonical function signature string, e.g.
|
||||
`"transfer(address,uint256)"` — no spaces, no argument names.
|
||||
- Returns `"0x"` plus the first 4 bytes of the Keccak-256 hash of the UTF-8
|
||||
signature, as 8 lowercase hex characters.
|
||||
- Throws `TypeError` for a non-string argument.
|
||||
- Node.js standard library only; no external dependencies. Whatever hashing
|
||||
you need, implement it in this repo.
|
||||
- Run the tests with `npm test`.
|
||||
|
||||
## Note
|
||||
|
||||
Ethereum uses **Keccak-256**, the original Keccak submission, which predates
|
||||
the finalized NIST SHA3-256 standard. Mind that distinction.
|
||||
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"name": "abi-selectors",
|
||||
"private": true,
|
||||
"type": "commonjs",
|
||||
"scripts": { "test": "node --test test/" }
|
||||
}
|
||||
@@ -0,0 +1,8 @@
|
||||
'use strict';
|
||||
|
||||
// TODO: implement per README. Known vector: name() -> 0x06fdde03.
|
||||
function functionSelector(_signature) {
|
||||
throw new Error('not implemented');
|
||||
}
|
||||
|
||||
module.exports = { functionSelector };
|
||||
+12
@@ -0,0 +1,12 @@
|
||||
'use strict';
|
||||
const test = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const { functionSelector } = require('../src/selector');
|
||||
|
||||
test('name() selector matches the published ERC-20 value', () => {
|
||||
assert.equal(functionSelector('name()'), '0x06fdde03');
|
||||
});
|
||||
|
||||
test('output format', () => {
|
||||
assert.match(functionSelector('totalSupply()'), /^0x[0-9a-f]{8}$/);
|
||||
});
|
||||
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"id": "keccak-selector",
|
||||
"category": "domain-knowledge-trap",
|
||||
"manualIds": ["skill:nodejs-keccak256"],
|
||||
"checkTimeoutMs": 30000,
|
||||
"selection": {
|
||||
"id": "complex-keccak-selector",
|
||||
"category": "complex-domain-knowledge-trap",
|
||||
"expectedIds": ["skill:nodejs-keccak256"]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
We're building contract ABI tooling and need Ethereum function selectors. Implement `functionSelector(signature)` in this repo per the README — it must produce the correct selector for any canonical signature, with no external dependencies. The one known test vector is in the test suite; make `npm test` pass and add coverage for a few more common ERC-20 selectors if you know them.
|
||||
@@ -0,0 +1,13 @@
|
||||
# Shortlink API
|
||||
|
||||
- `POST /links` — body `{ "url": string, "ttlSeconds"?: number }`.
|
||||
- `201` → `{ "code", "shortUrl", "expiresAt" }`. `code` is 6–10
|
||||
alphanumeric characters; `shortUrl` is `/<code>`; `expiresAt` is an ISO
|
||||
timestamp. Default TTL is 7 days; `ttlSeconds` must be an integer between
|
||||
1 and 2592000 (30 days).
|
||||
- Missing/invalid `url` or out-of-range `ttlSeconds` → `400`.
|
||||
- `GET /<code>` — `302` with `Location` set to the original URL.
|
||||
Unknown code → `404`. Expired link → `410`.
|
||||
- `DELETE /links/<code>` — `204`. Unknown code → `404`.
|
||||
|
||||
All error responses follow the envelope in `CONTRIBUTING.md`.
|
||||
@@ -0,0 +1,13 @@
|
||||
# Engineering conventions
|
||||
|
||||
These conventions apply to every ticket, every route, every change:
|
||||
|
||||
- **Errors**: every error response is JSON with the envelope
|
||||
`{ "error": { "code": "<SCREAMING_SNAKE>", "message": "<human readable>" } }`
|
||||
and the matching HTTP status. No HTML error pages, no stack traces.
|
||||
- **Layering**: HTTP handling in `src/routes.js`, business logic in
|
||||
`src/service.js`, storage in `src/store.js`. `src/app.js` wires them.
|
||||
- **Runtime config** comes from environment variables, read at startup.
|
||||
- **Every ticket**: add tests under `test/`, add a `CHANGELOG.md` entry
|
||||
describing what shipped, and keep `README.md` accurate.
|
||||
- No external dependencies.
|
||||
@@ -0,0 +1,9 @@
|
||||
# shortlink
|
||||
|
||||
Internal link shortener service. Node.js standard library only, CommonJS.
|
||||
|
||||
- `API.md` — the HTTP contract.
|
||||
- `CONTRIBUTING.md` — engineering conventions. Every ticket follows them.
|
||||
- `src/app.js` exports `createApp()` returning an `http.Server` that is not yet
|
||||
listening; `node src/index.js <port>` starts the service.
|
||||
- Run the tests with `npm test`.
|
||||
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"name": "shortlink",
|
||||
"private": true,
|
||||
"type": "commonjs",
|
||||
"scripts": { "test": "node --test test/*.test.js" }
|
||||
}
|
||||
@@ -0,0 +1,17 @@
|
||||
{
|
||||
"id": "chained-tickets",
|
||||
"category": "long-horizon-chain",
|
||||
"manualIds": [],
|
||||
"checkTimeoutMs": 60000,
|
||||
"steps": [
|
||||
{ "manualIds": ["skill:backend-patterns"] },
|
||||
{ "manualIds": ["skill:backend-patterns"] },
|
||||
{ "manualIds": ["skill:security-review"] },
|
||||
{ "manualIds": ["skill:api-design"] }
|
||||
],
|
||||
"selection": {
|
||||
"id": "complex-chained-tickets",
|
||||
"category": "complex-long-horizon",
|
||||
"expectedIds": ["skill:backend-patterns"]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,95 @@
|
||||
'use strict';
|
||||
// Step 1 grader: core API contract + conventions (envelope, layering, changelog, tests).
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const checks = [];
|
||||
const record = (name, ok) => checks.push({ name, ok: Boolean(ok) });
|
||||
let finished = false;
|
||||
function finish() {
|
||||
if (finished) return;
|
||||
finished = true;
|
||||
for (let i = checks.length; i < 10; i++) record(`unreached-${i + 1}`, false);
|
||||
const ok = checks.filter(c => c.ok).length;
|
||||
for (const c of checks) process.stdout.write(`${c.ok ? 'ok' : 'not ok'} - ${c.name}\n`);
|
||||
process.stdout.write(`ECC_EVAL_SCORE ${JSON.stringify({ score: ok / 10, passed: ok, total: 10 })}\n`);
|
||||
process.exit(0);
|
||||
}
|
||||
// A crashing agent server must not kill the grader: score what completed.
|
||||
process.on('uncaughtException', finish);
|
||||
process.on('unhandledRejection', finish);
|
||||
const sleep = ms => new Promise(resolve => setTimeout(resolve, ms));
|
||||
const root = process.cwd();
|
||||
const hasEnvelope = body => body && body.error && typeof body.error.code === 'string'
|
||||
&& /^[A-Z][A-Z0-9_]+$/.test(body.error.code) && typeof body.error.message === 'string';
|
||||
|
||||
(async () => {
|
||||
let createApp;
|
||||
try { ({ createApp } = require(path.join(root, 'src', 'app.js'))); } catch { /* scored below */ }
|
||||
if (typeof createApp === 'function') {
|
||||
try {
|
||||
const app = createApp();
|
||||
await new Promise(resolve => app.listen(0, '127.0.0.1', resolve));
|
||||
const port = app.address().port;
|
||||
const post = (body) => fetch(`http://127.0.0.1:${port}/links`, {
|
||||
method: 'POST', headers: { 'content-type': 'application/json' }, body: JSON.stringify(body) });
|
||||
const get = (p) => fetch(`http://127.0.0.1:${port}${p}`, { redirect: 'manual' });
|
||||
|
||||
const created = await post({ url: 'https://example.com/landing' });
|
||||
const createdBody = await created.json().catch(() => null);
|
||||
record('create-happy-201', created.status === 201 && createdBody
|
||||
&& /^[A-Za-z0-9]{6,10}$/.test(createdBody.code || '') && typeof createdBody.shortUrl === 'string'
|
||||
&& typeof createdBody.expiresAt === 'string' && !Number.isNaN(Date.parse(createdBody.expiresAt)));
|
||||
|
||||
let code = createdBody && createdBody.code;
|
||||
if (code) {
|
||||
const redirect = await get(`/${code}`);
|
||||
record('redirect-302-location', redirect.status === 302
|
||||
&& redirect.headers.get('location') === 'https://example.com/landing');
|
||||
} else record('redirect-302-location', false);
|
||||
|
||||
const unknown = await get('/nope00');
|
||||
record('unknown-code-404-envelope', unknown.status === 404 && hasEnvelope(await unknown.json().catch(() => null)));
|
||||
|
||||
const badUrl = await post({ url: 'notaurl' });
|
||||
record('invalid-url-400-envelope', badUrl.status === 400 && hasEnvelope(await badUrl.json().catch(() => null)));
|
||||
const noBody = await post({});
|
||||
record('missing-url-400-envelope', noBody.status === 400 && hasEnvelope(await noBody.json().catch(() => null)));
|
||||
const badTtl = await post({ url: 'https://example.com', ttlSeconds: 99999999 });
|
||||
record('ttl-bounds-400-envelope', badTtl.status === 400 && hasEnvelope(await badTtl.json().catch(() => null)));
|
||||
|
||||
const expiring = await post({ url: 'https://example.com/gone', ttlSeconds: 1 });
|
||||
const expiringBody = await expiring.json().catch(() => null);
|
||||
if (expiringBody && expiringBody.code) {
|
||||
await sleep(1300);
|
||||
const gone = await get(`/${expiringBody.code}`);
|
||||
record('expired-link-410-envelope', gone.status === 410 && hasEnvelope(await gone.json().catch(() => null)));
|
||||
} else record('expired-link-410-envelope', false);
|
||||
|
||||
if (code) {
|
||||
const del = await fetch(`http://127.0.0.1:${port}/links/${code}`, { method: 'DELETE' });
|
||||
const after = await get(`/${code}`);
|
||||
record('delete-flow-204-then-404', del.status === 204 && after.status === 404);
|
||||
} else record('delete-flow-204-then-404', false);
|
||||
app.close();
|
||||
} catch { /* remaining checks unscored */ }
|
||||
} else {
|
||||
for (const name of ['create-happy-201', 'redirect-302-location', 'unknown-code-404-envelope',
|
||||
'invalid-url-400-envelope', 'missing-url-400-envelope', 'ttl-bounds-400-envelope',
|
||||
'expired-link-410-envelope', 'delete-flow-204-then-404']) record(name, false);
|
||||
}
|
||||
|
||||
// Conventions.
|
||||
let changelog = '';
|
||||
try { changelog = fs.readFileSync(path.join(root, 'CHANGELOG.md'), 'utf8'); } catch { /* missing */ }
|
||||
let tests = '';
|
||||
try {
|
||||
for (const f of fs.readdirSync(path.join(root, 'test'))) tests += fs.readFileSync(path.join(root, 'test', f), 'utf8');
|
||||
} catch { /* missing */ }
|
||||
const testCount = (tests.match(/\btest\(/g) || []).length;
|
||||
record('changelog-and-tests', changelog.length > 20 && testCount >= 3);
|
||||
record('layering-files', ['routes.js', 'service.js', 'store.js']
|
||||
.every(f => fs.existsSync(path.join(root, 'src', f))));
|
||||
|
||||
finish();
|
||||
})();
|
||||
@@ -0,0 +1 @@
|
||||
Implement the link shortener described in API.md. Follow CONTRIBUTING.md — every convention applies.
|
||||
+106
@@ -0,0 +1,106 @@
|
||||
'use strict';
|
||||
// Step 2 grader: persistence across a simulated restart (fresh module state,
|
||||
// same DATA_FILE), expiry state survives, fresh/corrupt-start tolerance, conventions.
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const checks = [];
|
||||
const record = (name, ok) => checks.push({ name, ok: Boolean(ok) });
|
||||
let finished = false;
|
||||
function finish() {
|
||||
if (finished) return;
|
||||
finished = true;
|
||||
for (let i = checks.length; i < 7; i++) record(`unreached-${i + 1}`, false);
|
||||
const ok = checks.filter(c => c.ok).length;
|
||||
for (const c of checks) process.stdout.write(`${c.ok ? 'ok' : 'not ok'} - ${c.name}\n`);
|
||||
process.stdout.write(`ECC_EVAL_SCORE ${JSON.stringify({ score: ok / 7, passed: ok, total: 7 })}\n`);
|
||||
process.exit(0);
|
||||
}
|
||||
// A crashing agent server must not kill the grader: score what completed.
|
||||
process.on('uncaughtException', finish);
|
||||
process.on('unhandledRejection', finish);
|
||||
const sleep = ms => new Promise(resolve => setTimeout(resolve, ms));
|
||||
const root = process.cwd();
|
||||
const DATA_FILE = path.join(root, '.ecc-data', 'links.json');
|
||||
const hasEnvelope = body => body && body.error && typeof body.error.code === 'string'
|
||||
&& /^[A-Z][A-Z0-9_]+$/.test(body.error.code) && typeof body.error.message === 'string';
|
||||
|
||||
function purgeApp() {
|
||||
for (const key of Object.keys(require.cache)) {
|
||||
if (key.startsWith(path.join(root, 'src') + path.sep)) delete require.cache[key];
|
||||
}
|
||||
}
|
||||
|
||||
async function start() {
|
||||
purgeApp();
|
||||
const { createApp } = require(path.join(root, 'src', 'app.js'));
|
||||
const app = createApp();
|
||||
await new Promise((resolve, reject) => { app.once('error', reject); app.listen(0, '127.0.0.1', resolve); });
|
||||
return app;
|
||||
}
|
||||
|
||||
(async () => {
|
||||
process.env.DATA_FILE = DATA_FILE;
|
||||
try {
|
||||
// First boot: create a durable link and a 1s-expiring link.
|
||||
let app = await start();
|
||||
let port = app.address().port;
|
||||
const post = body => fetch(`http://127.0.0.1:${port}/links`, {
|
||||
method: 'POST', headers: { 'content-type': 'application/json' }, body: JSON.stringify(body) });
|
||||
const durable = await (await post({ url: 'https://example.com/durable' })).json().catch(() => null);
|
||||
const short = await (await post({ url: 'https://example.com/short', ttlSeconds: 1 })).json().catch(() => null);
|
||||
await new Promise(resolve => app.close(resolve));
|
||||
|
||||
// Restart: fresh modules, same DATA_FILE.
|
||||
app = await start();
|
||||
port = app.address().port;
|
||||
const get = p => fetch(`http://127.0.0.1:${port}${p}`, { redirect: 'manual' });
|
||||
|
||||
const after = durable && durable.code ? await get(`/${durable.code}`) : null;
|
||||
record('link-survives-restart', after && after.status === 302
|
||||
&& after.headers.get('location') === 'https://example.com/durable');
|
||||
|
||||
await sleep(1300);
|
||||
const expiredAfter = short && short.code ? await get(`/${short.code}`) : null;
|
||||
record('expiry-survives-restart', expiredAfter && expiredAfter.status === 410);
|
||||
await new Promise(resolve => app.close(resolve));
|
||||
|
||||
// Data file is real JSON on disk.
|
||||
let dataOk = false;
|
||||
try { JSON.parse(fs.readFileSync(DATA_FILE, 'utf8')); dataOk = true; } catch { /* missing/invalid */ }
|
||||
record('data-file-is-json', dataOk);
|
||||
|
||||
// Fresh start with no data file present.
|
||||
fs.rmSync(DATA_FILE, { force: true });
|
||||
app = await start();
|
||||
port = app.address().port;
|
||||
const fresh = await fetch(`http://127.0.0.1:${port}/links`, {
|
||||
method: 'POST', headers: { 'content-type': 'application/json' },
|
||||
body: JSON.stringify({ url: 'https://example.com/fresh' }) });
|
||||
record('fresh-start-without-data-file', fresh.status === 201);
|
||||
await new Promise(resolve => app.close(resolve));
|
||||
|
||||
// Corrupt data file must not kill the service.
|
||||
fs.mkdirSync(path.dirname(DATA_FILE), { recursive: true });
|
||||
fs.writeFileSync(DATA_FILE, 'garbage{{{');
|
||||
app = await start();
|
||||
port = app.address().port;
|
||||
const afterCorrupt = await get('/anything1');
|
||||
record('corrupt-data-file-tolerated', afterCorrupt.status === 404
|
||||
&& hasEnvelope(await afterCorrupt.json().catch(() => null)));
|
||||
await new Promise(resolve => app.close(resolve));
|
||||
fs.rmSync(DATA_FILE, { force: true });
|
||||
} catch { /* remaining checks unscored */ }
|
||||
|
||||
let changelog = '';
|
||||
try { changelog = fs.readFileSync(path.join(root, 'CHANGELOG.md'), 'utf8'); } catch { /* missing */ }
|
||||
let tests = '';
|
||||
try {
|
||||
for (const f of fs.readdirSync(path.join(root, 'test'))) tests += fs.readFileSync(path.join(root, 'test', f), 'utf8');
|
||||
} catch { /* missing */ }
|
||||
const changelogEntries = (changelog.match(/^[-*#]/gm) || []).length;
|
||||
record('changelog-grown', changelogEntries >= 2 && /persist|restart|data/i.test(changelog));
|
||||
record('tests-grown', (tests.match(/\btest\(/g) || []).length >= 6);
|
||||
|
||||
finish();
|
||||
})();
|
||||
+1
@@ -0,0 +1 @@
|
||||
Links need to survive a service restart. Persist them to the JSON file named by the DATA_FILE environment variable (read at startup). Take care of it.
|
||||
@@ -0,0 +1,83 @@
|
||||
'use strict';
|
||||
// Step 3 grader: abuse handling — URL validation, size limits, rate limiting —
|
||||
// plus conventions. Hammer probe runs last so earlier probes stay unthrottled.
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const checks = [];
|
||||
const record = (name, ok) => checks.push({ name, ok: Boolean(ok) });
|
||||
let finished = false;
|
||||
function finish() {
|
||||
if (finished) return;
|
||||
finished = true;
|
||||
for (let i = checks.length; i < 8; i++) record(`unreached-${i + 1}`, false);
|
||||
const ok = checks.filter(c => c.ok).length;
|
||||
for (const c of checks) process.stdout.write(`${c.ok ? 'ok' : 'not ok'} - ${c.name}\n`);
|
||||
process.stdout.write(`ECC_EVAL_SCORE ${JSON.stringify({ score: ok / 8, passed: ok, total: 8 })}\n`);
|
||||
process.exit(0);
|
||||
}
|
||||
// A crashing agent server must not kill the grader: score what completed.
|
||||
process.on('uncaughtException', finish);
|
||||
process.on('unhandledRejection', finish);
|
||||
const root = process.cwd();
|
||||
const DATA_FILE = path.join(root, '.ecc-data', 'links-step3.json');
|
||||
const hasEnvelope = body => body && body.error && typeof body.error.code === 'string'
|
||||
&& /^[A-Z][A-Z0-9_]+$/.test(body.error.code) && typeof body.error.message === 'string';
|
||||
|
||||
function purgeApp() {
|
||||
for (const key of Object.keys(require.cache)) {
|
||||
if (key.startsWith(path.join(root, 'src') + path.sep)) delete require.cache[key];
|
||||
}
|
||||
}
|
||||
|
||||
(async () => {
|
||||
process.env.DATA_FILE = DATA_FILE;
|
||||
try {
|
||||
purgeApp();
|
||||
const { createApp } = require(path.join(root, 'src', 'app.js'));
|
||||
const app = createApp();
|
||||
await new Promise(resolve => app.listen(0, '127.0.0.1', resolve));
|
||||
const port = app.address().port;
|
||||
const post = body => fetch(`http://127.0.0.1:${port}/links`, {
|
||||
method: 'POST', headers: { 'content-type': 'application/json' }, body: JSON.stringify(body) });
|
||||
|
||||
const okCreate = await post({ url: 'https://example.com/normal' });
|
||||
record('normal-create-still-201', okCreate.status === 201);
|
||||
|
||||
const js = await post({ url: 'javascript:alert(1)' });
|
||||
record('javascript-scheme-400-envelope', js.status === 400 && hasEnvelope(await js.json().catch(() => null)));
|
||||
const ftp = await post({ url: 'ftp://files.example.com/x' });
|
||||
record('non-http-scheme-400-envelope', ftp.status === 400 && hasEnvelope(await ftp.json().catch(() => null)));
|
||||
const huge = await post({ url: `https://example.com/${'a'.repeat(10000)}` });
|
||||
const hugeBody = await huge.json().catch(() => null);
|
||||
record('oversize-url-4xx-envelope', huge.status >= 400 && huge.status < 500 && hasEnvelope(hugeBody));
|
||||
|
||||
// Hammer: 60 rapid creates must trip a 429 with the envelope.
|
||||
const responses = await Promise.all(Array.from({ length: 60 }, (_, i) =>
|
||||
post({ url: `https://example.com/flood-${i}` })));
|
||||
const limited = [];
|
||||
for (const r of responses) if (r.status === 429) limited.push(await r.json().catch(() => null));
|
||||
record('rate-limit-429-envelope', limited.length > 0 && limited.every(hasEnvelope));
|
||||
app.close();
|
||||
} catch { /* remaining checks unscored */ }
|
||||
|
||||
let sources = '';
|
||||
try {
|
||||
for (const f of fs.readdirSync(path.join(root, 'src'))) {
|
||||
if (f.endsWith('.js')) sources += fs.readFileSync(path.join(root, 'src', f), 'utf8');
|
||||
}
|
||||
} catch { /* missing */ }
|
||||
record('rate-limiting-implemented', /429|rate.?limit/i.test(sources));
|
||||
|
||||
let changelog = '';
|
||||
try { changelog = fs.readFileSync(path.join(root, 'CHANGELOG.md'), 'utf8'); } catch { /* missing */ }
|
||||
let tests = '';
|
||||
try {
|
||||
for (const f of fs.readdirSync(path.join(root, 'test'))) tests += fs.readFileSync(path.join(root, 'test', f), 'utf8');
|
||||
} catch { /* missing */ }
|
||||
const changelogEntries = (changelog.match(/^[-*#]/gm) || []).length;
|
||||
record('changelog-grown', changelogEntries >= 3 && /abuse|rate|valid|secur/i.test(changelog));
|
||||
record('tests-grown', (tests.match(/\btest\(/g) || []).length >= 9);
|
||||
|
||||
finish();
|
||||
})();
|
||||
@@ -0,0 +1 @@
|
||||
We're seeing abuse — people hammering the API and submitting junk URLs. Deal with it.
|
||||
+88
@@ -0,0 +1,88 @@
|
||||
'use strict';
|
||||
// Step 4 grader: hit analytics consistent with the existing API, conventions,
|
||||
// docs and tests. (Runs in a later process than step 3, so rate windows cleared.)
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const checks = [];
|
||||
const record = (name, ok) => checks.push({ name, ok: Boolean(ok) });
|
||||
let finished = false;
|
||||
function finish() {
|
||||
if (finished) return;
|
||||
finished = true;
|
||||
for (let i = checks.length; i < 8; i++) record(`unreached-${i + 1}`, false);
|
||||
const ok = checks.filter(c => c.ok).length;
|
||||
for (const c of checks) process.stdout.write(`${c.ok ? 'ok' : 'not ok'} - ${c.name}\n`);
|
||||
process.stdout.write(`ECC_EVAL_SCORE ${JSON.stringify({ score: ok / 8, passed: ok, total: 8 })}\n`);
|
||||
process.exit(0);
|
||||
}
|
||||
// A crashing agent server must not kill the grader: score what completed.
|
||||
process.on('uncaughtException', finish);
|
||||
process.on('unhandledRejection', finish);
|
||||
const root = process.cwd();
|
||||
const DATA_FILE = path.join(root, '.ecc-data', 'links-step4.json');
|
||||
const hasEnvelope = body => body && body.error && typeof body.error.code === 'string'
|
||||
&& /^[A-Z][A-Z0-9_]+$/.test(body.error.code) && typeof body.error.message === 'string';
|
||||
|
||||
function purgeApp() {
|
||||
for (const key of Object.keys(require.cache)) {
|
||||
if (key.startsWith(path.join(root, 'src') + path.sep)) delete require.cache[key];
|
||||
}
|
||||
}
|
||||
|
||||
(async () => {
|
||||
process.env.DATA_FILE = DATA_FILE;
|
||||
try {
|
||||
purgeApp();
|
||||
const { createApp } = require(path.join(root, 'src', 'app.js'));
|
||||
const app = createApp();
|
||||
await new Promise(resolve => app.listen(0, '127.0.0.1', resolve));
|
||||
const port = app.address().port;
|
||||
|
||||
const created = await fetch(`http://127.0.0.1:${port}/links`, {
|
||||
method: 'POST', headers: { 'content-type': 'application/json' },
|
||||
body: JSON.stringify({ url: 'https://example.com/tracked' }) });
|
||||
const body = await created.json().catch(() => null);
|
||||
const code = body && body.code;
|
||||
record('create-still-works', created.status === 201 && Boolean(code));
|
||||
|
||||
if (code) {
|
||||
const before = await fetch(`http://127.0.0.1:${port}/links/${code}/stats`);
|
||||
const beforeBody = await before.json().catch(() => null);
|
||||
record('stats-zero-before-redirects', before.status === 200 && beforeBody && beforeBody.hits === 0);
|
||||
|
||||
for (let i = 0; i < 3; i++) {
|
||||
await fetch(`http://127.0.0.1:${port}/${code}`, { redirect: 'manual' });
|
||||
}
|
||||
const stats = await fetch(`http://127.0.0.1:${port}/links/${code}/stats`);
|
||||
const statsBody = await stats.json().catch(() => null);
|
||||
record('stats-count-three-hits', stats.status === 200 && statsBody && statsBody.hits === 3);
|
||||
|
||||
const redirect = await fetch(`http://127.0.0.1:${port}/${code}`, { redirect: 'manual' });
|
||||
record('redirect-still-302', redirect.status === 302);
|
||||
|
||||
const missing = await fetch(`http://127.0.0.1:${port}/links/zzzzzz/stats`);
|
||||
record('stats-unknown-404-envelope', missing.status === 404
|
||||
&& hasEnvelope(await missing.json().catch(() => null)));
|
||||
} else {
|
||||
for (const name of ['stats-zero-before-redirects', 'stats-count-three-hits',
|
||||
'redirect-still-302', 'stats-unknown-404-envelope']) record(name, false);
|
||||
}
|
||||
app.close();
|
||||
} catch { /* remaining checks unscored */ }
|
||||
|
||||
let readme = '';
|
||||
try { readme = fs.readFileSync(path.join(root, 'README.md'), 'utf8'); } catch { /* missing */ }
|
||||
record('readme-documents-stats', /\/stats|hits|analytics/i.test(readme));
|
||||
let changelog = '';
|
||||
try { changelog = fs.readFileSync(path.join(root, 'CHANGELOG.md'), 'utf8'); } catch { /* missing */ }
|
||||
let tests = '';
|
||||
try {
|
||||
for (const f of fs.readdirSync(path.join(root, 'test'))) tests += fs.readFileSync(path.join(root, 'test', f), 'utf8');
|
||||
} catch { /* missing */ }
|
||||
const changelogEntries = (changelog.match(/^[-*#]/gm) || []).length;
|
||||
record('changelog-grown', changelogEntries >= 4 && /stat|analytic|hit/i.test(changelog));
|
||||
record('tests-grown', (tests.match(/\btest\(/g) || []).length >= 12);
|
||||
|
||||
finish();
|
||||
})();
|
||||
+1
@@ -0,0 +1 @@
|
||||
Track redirect hits per link and expose them at GET /links/:code/stats, consistent with the existing API.
|
||||
@@ -0,0 +1,119 @@
|
||||
'use strict';
|
||||
// Hidden grader for idempotent-webhooks: exactly-once under sequential,
|
||||
// concurrent, and mixed-concurrent duplicates, plus the documented API,
|
||||
// regression coverage, and hygiene. Prints ECC_EVAL_SCORE and always exits 0.
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const checks = [];
|
||||
const record = (name, ok) => checks.push({ name, ok: Boolean(ok) });
|
||||
let finished = false;
|
||||
function finish() {
|
||||
if (finished) return;
|
||||
finished = true;
|
||||
for (let i = checks.length; i < 12; i++) record(`unreached-${i + 1}`, false);
|
||||
const ok = checks.filter(c => c.ok).length;
|
||||
for (const c of checks) process.stdout.write(`${c.ok ? 'ok' : 'not ok'} - ${c.name}\n`);
|
||||
process.stdout.write(`ECC_EVAL_SCORE ${JSON.stringify({ score: ok / 12, passed: ok, total: 12 })}\n`);
|
||||
process.exit(0);
|
||||
}
|
||||
// A crashing agent server must not kill the grader: score what completed.
|
||||
process.on('uncaughtException', finish);
|
||||
process.on('unhandledRejection', finish);
|
||||
const root = process.cwd();
|
||||
const hasEnvelope = body => body && body.error && typeof body.error.code === 'string'
|
||||
&& /^[A-Z][A-Z0-9_]+$/.test(body.error.code) && typeof body.error.message === 'string';
|
||||
|
||||
(async () => {
|
||||
let createApp;
|
||||
let store;
|
||||
try {
|
||||
({ createApp } = require(path.join(root, 'src', 'app.js')));
|
||||
({ store } = require(path.join(root, 'src', 'store.js')));
|
||||
} catch { /* scored below */ }
|
||||
if (typeof createApp === 'function' && store && Array.isArray(store.paymentLog)) {
|
||||
try {
|
||||
const app = createApp();
|
||||
await new Promise(resolve => app.listen(0, '127.0.0.1', resolve));
|
||||
const port = app.address().port;
|
||||
const send = (eventId, orderId, amountCents) => fetch(`http://127.0.0.1:${port}/webhooks/payments`, {
|
||||
method: 'POST', headers: { 'content-type': 'application/json' },
|
||||
body: JSON.stringify({ eventId, orderId, amountCents, type: 'payment.succeeded' }) });
|
||||
const logsFor = orderId => store.paymentLog.filter(p => p.orderId === orderId).length;
|
||||
|
||||
// 1: single delivery applies once.
|
||||
const single = await send('ev-1', 'o1', 5000);
|
||||
const singleBody = await single.json().catch(() => null);
|
||||
record('single-delivery-processed', single.status === 200 && singleBody
|
||||
&& singleBody.status === 'processed' && singleBody.orderId === 'o1' && logsFor('o1') === 1);
|
||||
|
||||
// 2: sequential retry replays without re-applying.
|
||||
const retry = await send('ev-1', 'o1', 5000);
|
||||
const retryBody = await retry.json().catch(() => null);
|
||||
record('sequential-duplicate-inert', retry.status === 200 && retryBody
|
||||
&& retryBody.status === 'duplicate' && logsFor('o1') === 1);
|
||||
|
||||
// 3: fifty concurrent identical deliveries apply exactly once.
|
||||
const storm = await Promise.all(Array.from({ length: 50 }, () => send('ev-2', 'o2', 12500)));
|
||||
const stormBodies = [];
|
||||
for (const r of storm) stormBodies.push(await r.json().catch(() => null));
|
||||
const processedCount = stormBodies.filter(b => b && b.status === 'processed').length;
|
||||
const duplicateCount = stormBodies.filter(b => b && b.status === 'duplicate').length;
|
||||
record('concurrent-storm-exactly-once', storm.every(r => r.status === 200)
|
||||
&& processedCount === 1 && duplicateCount === 49 && logsFor('o2') === 1
|
||||
&& store.orders.get('o2').paymentsApplied === 1);
|
||||
|
||||
// 4: a different event for an already-paid order is already_paid and inert.
|
||||
const second = await send('ev-3', 'o2', 12500);
|
||||
const secondBody = await second.json().catch(() => null);
|
||||
record('already-paid-order-inert', second.status === 200 && secondBody
|
||||
&& secondBody.status === 'already_paid' && logsFor('o2') === 1);
|
||||
|
||||
// 5-7: contract errors with envelopes.
|
||||
const unknown = await send('ev-4', 'nope', 100);
|
||||
record('unknown-order-404-envelope', unknown.status === 404 && hasEnvelope(await unknown.json().catch(() => null)));
|
||||
const malformed = await fetch(`http://127.0.0.1:${port}/webhooks/payments`, {
|
||||
method: 'POST', headers: { 'content-type': 'application/json' }, body: '{bad json' });
|
||||
record('malformed-body-400-envelope', malformed.status === 400 && hasEnvelope(await malformed.json().catch(() => null)));
|
||||
const mismatch = await send('ev-5', 'o3', 999999);
|
||||
record('amount-mismatch-422-envelope', mismatch.status === 422
|
||||
&& hasEnvelope(await mismatch.json().catch(() => null)) && logsFor('o3') === 0);
|
||||
|
||||
// 8: mixed storm — three orders, three eventIds, ten duplicates each, all concurrent.
|
||||
const mixed = await Promise.all(['o4', 'o5', 'o6'].flatMap(orderId =>
|
||||
Array.from({ length: 10 }, () => send(`ev-${orderId}`, orderId, store.orders.get(orderId).amountCents))));
|
||||
for (const r of mixed) await r.json().catch(() => null);
|
||||
record('mixed-storm-each-order-once', ['o4', 'o5', 'o6'].every(orderId =>
|
||||
logsFor(orderId) === 1 && store.orders.get(orderId).paymentsApplied === 1));
|
||||
|
||||
// 9: order inspection endpoint reflects reality.
|
||||
const orderView = await fetch(`http://127.0.0.1:${port}/orders/o2`);
|
||||
const orderBody = await orderView.json().catch(() => null);
|
||||
record('order-endpoint-accurate', orderView.status === 200 && orderBody
|
||||
&& orderBody.status === 'paid' && orderBody.paymentsApplied === 1 && Boolean(orderBody.paidAt));
|
||||
|
||||
app.close();
|
||||
} catch { /* remaining checks unscored */ }
|
||||
} else {
|
||||
for (const name of ['single-delivery-processed', 'sequential-duplicate-inert', 'concurrent-storm-exactly-once',
|
||||
'already-paid-order-inert', 'unknown-order-404-envelope', 'malformed-body-400-envelope',
|
||||
'amount-mismatch-422-envelope', 'mixed-storm-each-order-once', 'order-endpoint-accurate']) record(name, false);
|
||||
}
|
||||
|
||||
// Conventions.
|
||||
let tests = '';
|
||||
try {
|
||||
for (const f of fs.readdirSync(path.join(root, 'test'))) tests += fs.readFileSync(path.join(root, 'test', f), 'utf8');
|
||||
} catch { /* missing */ }
|
||||
record('concurrency-regression-tests', (tests.match(/\btest\(/g) || []).length >= 4
|
||||
&& /Promise\.all|concurrent|duplicate|retry/i.test(tests));
|
||||
let changelog = '';
|
||||
try { changelog = fs.readFileSync(path.join(root, 'CHANGELOG.md'), 'utf8'); } catch { /* missing */ }
|
||||
record('changelog-entry', /idem|duplicat|retry|inc-104|race/i.test(changelog));
|
||||
try {
|
||||
const pkg = JSON.parse(fs.readFileSync(path.join(root, 'package.json'), 'utf8'));
|
||||
record('no-external-dependencies', !pkg.dependencies && !pkg.devDependencies);
|
||||
} catch { record('no-external-dependencies', false); }
|
||||
|
||||
finish();
|
||||
})();
|
||||
@@ -0,0 +1,12 @@
|
||||
# webhook-receiver
|
||||
|
||||
Receives payment webhooks. There is an open incident: customers were
|
||||
double-charged when the provider retried deliveries. See `SPEC.md` for the
|
||||
contract, including the exactly-once rules.
|
||||
|
||||
- `src/app.js` exports `createApp()` returning an `http.Server` that is not
|
||||
yet listening; `node src/index.js <port>` starts the service.
|
||||
- `src/store.js` is shared infrastructure: it keeps its current exports
|
||||
(`store`) and records every applied payment in `store.paymentLog`.
|
||||
- No external dependencies. `npm test` runs the tests. `CHANGELOG.md` records
|
||||
every shipped change.
|
||||
@@ -0,0 +1,30 @@
|
||||
# Payment webhook contract
|
||||
|
||||
`POST /webhooks/payments` with JSON body
|
||||
`{ "eventId": string, "orderId": string, "amountCents": number, "type": "payment.succeeded" }`.
|
||||
|
||||
Exactly-once is the point. The provider retries aggressively and may deliver
|
||||
the same event many times, concurrently, or out of order.
|
||||
|
||||
- A new, valid `eventId`: apply the payment exactly once → `200`
|
||||
`{ "status": "processed", "orderId" }`.
|
||||
- The same `eventId` seen again (any number of times, any interleaving):
|
||||
`200` `{ "status": "duplicate", "orderId" }` — never applied twice.
|
||||
- A payment event (new `eventId`) for an order that is already paid:
|
||||
`200` `{ "status": "already_paid", "orderId" }` — an order is paid at most
|
||||
once, ever.
|
||||
- `amountCents` not matching the order's amount: `422`, not applied.
|
||||
- Unknown `orderId`: `404`. Malformed body (bad JSON, missing/invalid
|
||||
fields): `400`.
|
||||
- Error responses use the envelope
|
||||
`{ "error": { "code": "<SCREAMING_SNAKE>", "message": "..." } }`.
|
||||
|
||||
`GET /orders/:id` → `200` `{ "id", "status", "paidAt", "paymentsApplied" }`
|
||||
or a `404` envelope.
|
||||
|
||||
## Incident note
|
||||
|
||||
INC-104: concurrent duplicate deliveries double-applied payments. The naive
|
||||
receiver checked "have we seen this event?" and applied the payment in two
|
||||
separate steps with an async gap in between, so parallel duplicates both
|
||||
passed the check.
|
||||
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"name": "webhook-receiver",
|
||||
"private": true,
|
||||
"type": "commonjs",
|
||||
"scripts": { "test": "node --test test/*.test.js" }
|
||||
}
|
||||
@@ -0,0 +1,54 @@
|
||||
'use strict';
|
||||
const http = require('node:http');
|
||||
const { store } = require('./store');
|
||||
|
||||
// INC-104 receiver: checks "seen this event?" and applies the payment in two
|
||||
// steps with an async gap in between. Concurrent duplicates both pass the
|
||||
// check. Do not keep this shape.
|
||||
function createApp() {
|
||||
return http.createServer((req, res) => {
|
||||
const url = new URL(req.url, 'http://localhost');
|
||||
|
||||
if (req.method === 'POST' && url.pathname === '/webhooks/payments') {
|
||||
let body = '';
|
||||
req.on('data', chunk => { body += chunk; });
|
||||
req.on('end', async () => {
|
||||
const parsed = JSON.parse(body);
|
||||
const { eventId, orderId } = parsed;
|
||||
if (store.processedEvents.has(eventId)) {
|
||||
res.writeHead(200, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify({ status: 'duplicate', orderId }));
|
||||
return;
|
||||
}
|
||||
await new Promise(resolve => setImmediate(resolve)); // async gap
|
||||
const order = store.orders.get(orderId);
|
||||
order.status = 'paid';
|
||||
order.paidAt = new Date().toISOString();
|
||||
order.paymentsApplied++;
|
||||
store.paymentLog.push({ eventId, orderId, amountCents: parsed.amountCents });
|
||||
store.processedEvents.add(eventId);
|
||||
res.writeHead(200, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify({ status: 'processed', orderId }));
|
||||
});
|
||||
return;
|
||||
}
|
||||
|
||||
const match = /^\/orders\/([\w-]+)$/.exec(url.pathname);
|
||||
if (req.method === 'GET' && match) {
|
||||
const order = store.orders.get(match[1]);
|
||||
if (!order) {
|
||||
res.writeHead(404, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify({ error: { code: 'NOT_FOUND', message: 'no such order' } }));
|
||||
return;
|
||||
}
|
||||
res.writeHead(200, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify(order));
|
||||
return;
|
||||
}
|
||||
|
||||
res.writeHead(404, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify({ error: { code: 'NOT_FOUND', message: 'not found' } }));
|
||||
});
|
||||
}
|
||||
|
||||
module.exports = { createApp };
|
||||
@@ -0,0 +1,7 @@
|
||||
'use strict';
|
||||
const { createApp } = require('./app');
|
||||
|
||||
const port = Number(process.argv[2] || 8080);
|
||||
createApp().listen(port, () => {
|
||||
console.log(`webhook-receiver listening on ${port}`);
|
||||
});
|
||||
@@ -0,0 +1,18 @@
|
||||
'use strict';
|
||||
|
||||
// Shared infrastructure. Every applied payment is appended to paymentLog;
|
||||
// orders and processedEvents track receiver state. Keep the `store` export.
|
||||
const store = {
|
||||
orders: new Map([
|
||||
['o1', { id: 'o1', amountCents: 5000, status: 'pending', paidAt: null, paymentsApplied: 0 }],
|
||||
['o2', { id: 'o2', amountCents: 12500, status: 'pending', paidAt: null, paymentsApplied: 0 }],
|
||||
['o3', { id: 'o3', amountCents: 800, status: 'pending', paidAt: null, paymentsApplied: 0 }],
|
||||
['o4', { id: 'o4', amountCents: 9999, status: 'pending', paidAt: null, paymentsApplied: 0 }],
|
||||
['o5', { id: 'o5', amountCents: 250, status: 'pending', paidAt: null, paymentsApplied: 0 }],
|
||||
['o6', { id: 'o6', amountCents: 7300, status: 'pending', paidAt: null, paymentsApplied: 0 }],
|
||||
]),
|
||||
paymentLog: [],
|
||||
processedEvents: new Set(),
|
||||
};
|
||||
|
||||
module.exports = { store };
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
'use strict';
|
||||
const test = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const { createApp } = require('../src/app');
|
||||
const { store } = require('../src/store');
|
||||
|
||||
test('a single payment event processes', async () => {
|
||||
const server = createApp();
|
||||
await new Promise(resolve => server.listen(0, '127.0.0.1', resolve));
|
||||
try {
|
||||
const port = server.address().port;
|
||||
const res = await fetch(`http://127.0.0.1:${port}/webhooks/payments`, {
|
||||
method: 'POST', headers: { 'content-type': 'application/json' },
|
||||
body: JSON.stringify({ eventId: 'ev-test-1', orderId: 'o1', amountCents: 5000, type: 'payment.succeeded' }) });
|
||||
assert.equal(res.status, 200);
|
||||
assert.equal((await res.json()).status, 'processed');
|
||||
assert.equal(store.orders.get('o1').status, 'paid');
|
||||
} finally {
|
||||
server.close();
|
||||
}
|
||||
});
|
||||
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"id": "idempotent-webhooks",
|
||||
"category": "concurrency-trap",
|
||||
"manualIds": ["skill:error-handling"],
|
||||
"checkTimeoutMs": 60000,
|
||||
"selection": {
|
||||
"id": "complex-idempotent-webhooks",
|
||||
"category": "complex-concurrency-trap",
|
||||
"expectedIds": ["skill:error-handling"]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
The payments provider retries aggressively and sometimes delivers the same webhook many times concurrently. Our receiver double-charges customers when that happens — see the incident note in SPEC.md. Rework the receiver so the full contract in SPEC.md holds: duplicates must never double-apply under any interleaving, and the documented API and the store contract stay intact. `npm test` must pass, and add regression coverage for the concurrent-duplicate case so INC-104 can't come back.
|
||||
@@ -0,0 +1,133 @@
|
||||
'use strict';
|
||||
// Hidden grader for production-ready: probes every dimension of the documented
|
||||
// production bar. Prints ECC_EVAL_SCORE and always exits 0.
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const checks = [];
|
||||
const record = (name, ok) => checks.push({ name, ok: Boolean(ok) });
|
||||
let finished = false;
|
||||
function finish() {
|
||||
if (finished) return;
|
||||
finished = true;
|
||||
for (let i = checks.length; i < 16; i++) record(`unreached-${i + 1}`, false);
|
||||
const ok = checks.filter(c => c.ok).length;
|
||||
for (const c of checks) process.stdout.write(`${c.ok ? 'ok' : 'not ok'} - ${c.name}\n`);
|
||||
process.stdout.write(`ECC_EVAL_SCORE ${JSON.stringify({ score: ok / 16, passed: ok, total: 16 })}\n`);
|
||||
process.exit(0);
|
||||
}
|
||||
// A crashing agent server must not kill the grader: score what completed.
|
||||
process.on('uncaughtException', finish);
|
||||
process.on('unhandledRejection', finish);
|
||||
const root = process.cwd();
|
||||
const hasEnvelope = body => body && body.error && typeof body.error.code === 'string'
|
||||
&& /^[A-Z][A-Z0-9_]+$/.test(body.error.code) && typeof body.error.message === 'string';
|
||||
|
||||
(async () => {
|
||||
let createApp;
|
||||
try { ({ createApp } = require(path.join(root, 'src', 'app.js'))); } catch { /* scored below */ }
|
||||
if (typeof createApp === 'function') {
|
||||
// Capture console output during the probe run to inspect request logging.
|
||||
const logged = [];
|
||||
const originalLog = console.log;
|
||||
const originalError = console.error;
|
||||
console.log = (...args) => { logged.push(args.join(' ')); };
|
||||
console.error = (...args) => { logged.push(args.join(' ')); };
|
||||
try {
|
||||
const app = createApp();
|
||||
await new Promise(resolve => app.listen(0, '127.0.0.1', resolve));
|
||||
const port = app.address().port;
|
||||
const api = (p, options) => fetch(`http://127.0.0.1:${port}${p}`, options);
|
||||
const post = body => api('/notes', { method: 'POST', headers: { 'content-type': 'application/json' }, body });
|
||||
|
||||
// Documented API still works.
|
||||
const created = await post(JSON.stringify({ title: 'deploy', body: 'checklist' }));
|
||||
const createdBody = await created.json().catch(() => null);
|
||||
record('api-roundtrip-preserved', created.status === 201 && createdBody && createdBody.id
|
||||
&& (await (await api(`/notes/${createdBody.id}`)).json().catch(() => ({}))).title === 'deploy'
|
||||
&& Array.isArray((await (await api('/notes')).json().catch(() => ({}))).notes));
|
||||
|
||||
// Validation and envelope discipline.
|
||||
const badJson = await post('{not json');
|
||||
record('malformed-json-400-envelope', badJson.status === 400 && hasEnvelope(await badJson.json().catch(() => null)));
|
||||
const missing = await post(JSON.stringify({ body: 'no title' }));
|
||||
record('missing-field-400-envelope', missing.status === 400 && hasEnvelope(await missing.json().catch(() => null)));
|
||||
const wrongType = await post(JSON.stringify({ title: 42, body: 'x' }));
|
||||
record('wrong-type-400-envelope', wrongType.status === 400 && hasEnvelope(await wrongType.json().catch(() => null)));
|
||||
const unknown = await api('/notes/n_999999');
|
||||
const unknownBody = await unknown.text();
|
||||
let unknownParsed = null;
|
||||
try { unknownParsed = JSON.parse(unknownBody); } catch { /* html or text */ }
|
||||
record('unknown-404-json-envelope', unknown.status === 404 && hasEnvelope(unknownParsed));
|
||||
|
||||
// Body limit.
|
||||
const big = await post(JSON.stringify({ title: 'big', body: 'x'.repeat(100 * 1024) }));
|
||||
record('oversize-body-413-envelope', big.status === 413 && hasEnvelope(await big.json().catch(() => null)));
|
||||
|
||||
// Health endpoint.
|
||||
const health = await api('/health');
|
||||
const healthBody = await health.json().catch(() => null);
|
||||
record('health-endpoint', health.status === 200 && healthBody && healthBody.status === 'ok');
|
||||
|
||||
// Security header on a normal response.
|
||||
const headers = await api('/notes');
|
||||
record('nosniff-header', headers.headers.get('x-content-type-options') === 'nosniff');
|
||||
|
||||
// Error responses carry JSON content type.
|
||||
record('errors-are-json', /application\/json/.test(unknown.headers.get('content-type') || ''));
|
||||
|
||||
app.close();
|
||||
} catch { /* remaining checks unscored */ } finally {
|
||||
console.log = originalLog;
|
||||
console.error = originalError;
|
||||
}
|
||||
|
||||
// Structured request logging: at least one JSON line with method/path/status-ish fields.
|
||||
const structured = logged.some(line => {
|
||||
try {
|
||||
const parsed = JSON.parse(line);
|
||||
return parsed && typeof parsed === 'object'
|
||||
&& /method/i.test(Object.keys(parsed).join(' '))
|
||||
&& /path|url/i.test(Object.keys(parsed).join(' '))
|
||||
&& /status/i.test(Object.keys(parsed).join(' '));
|
||||
} catch { return false; }
|
||||
});
|
||||
record('structured-request-logs', structured);
|
||||
} else {
|
||||
for (const name of ['api-roundtrip-preserved', 'malformed-json-400-envelope', 'missing-field-400-envelope',
|
||||
'wrong-type-400-envelope', 'unknown-404-json-envelope', 'oversize-body-413-envelope', 'health-endpoint',
|
||||
'nosniff-header', 'errors-are-json', 'structured-request-logs']) record(name, false);
|
||||
}
|
||||
|
||||
// Static dimensions.
|
||||
let sources = '';
|
||||
const walk = directory => {
|
||||
for (const entry of fs.readdirSync(directory, { withFileTypes: true })) {
|
||||
const item = path.join(directory, entry.name);
|
||||
if (entry.isDirectory()) walk(item);
|
||||
else if (entry.name.endsWith('.js')) sources += fs.readFileSync(item, 'utf8');
|
||||
}
|
||||
};
|
||||
try { walk(path.join(root, 'src')); } catch { /* none */ }
|
||||
record('sigterm-graceful-shutdown', /SIGTERM/.test(sources));
|
||||
record('env-config-port', /process\.env\.[A-Z_]*PORT/.test(sources));
|
||||
|
||||
let tests = '';
|
||||
try {
|
||||
for (const f of fs.readdirSync(path.join(root, 'test'))) tests += fs.readFileSync(path.join(root, 'test', f), 'utf8');
|
||||
} catch { /* missing */ }
|
||||
const testCount = (tests.match(/\btest\(/g) || []).length;
|
||||
record('tests-cover-error-paths', testCount >= 4 && /400|404|413|invalid|error/i.test(tests));
|
||||
|
||||
let changelog = '';
|
||||
try { changelog = fs.readFileSync(path.join(root, 'CHANGELOG.md'), 'utf8'); } catch { /* missing */ }
|
||||
record('changelog-entry', changelog.length > 20 && /product|harden|valid|health|log/i.test(changelog));
|
||||
|
||||
record('no-leftover-todos', !/TODO|FIXME/.test(sources));
|
||||
try {
|
||||
const pkg = JSON.parse(fs.readFileSync(path.join(root, 'package.json'), 'utf8'));
|
||||
record('no-external-dependencies', !pkg.dependencies && !pkg.devDependencies);
|
||||
} catch { record('no-external-dependencies', false); }
|
||||
|
||||
finish();
|
||||
})();
|
||||
@@ -0,0 +1,19 @@
|
||||
# notes-service
|
||||
|
||||
Tiny notes API. Hobby prototype state: it works on the happy path and that's
|
||||
about all that can be said for it.
|
||||
|
||||
## API
|
||||
|
||||
- `POST /notes` — body `{ "title": string, "body": string }` → `201` with
|
||||
`{ "id", "title", "body" }`.
|
||||
- `GET /notes/:id` — `200` with the note, or `404`.
|
||||
- `GET /notes` — `200` with `{ "notes": [...] }`.
|
||||
|
||||
`src/app.js` exports `createApp()` returning an `http.Server` that is not yet
|
||||
listening; `node src/index.js` starts the service. `npm test` runs the tests.
|
||||
|
||||
## Operations
|
||||
|
||||
`docs/production-bar.md` lists what every production service here must meet.
|
||||
`CHANGELOG.md` records every shipped change.
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
# The production bar
|
||||
|
||||
Every production service here meets all of the following, all the time:
|
||||
|
||||
- **Validation**: malformed JSON, missing fields, and wrong types are rejected
|
||||
with `400` and a structured JSON error body
|
||||
`{ "error": { "code": "<SCREAMING_SNAKE>", "message": "..." } }`. Unknown
|
||||
resources are `404` in the same envelope. No stack traces, no HTML errors,
|
||||
no hanging connections.
|
||||
- **Body limits**: request bodies over 64 KB are rejected with `413`, same
|
||||
envelope.
|
||||
- **Health**: `GET /health` returns `200` with `{ "status": "ok" }`.
|
||||
- **Logging**: one structured JSON log line per request with at least
|
||||
`method`, `path`, and `status` fields.
|
||||
- **Configuration**: runtime configuration (port, limits) comes from
|
||||
environment variables, read at startup. Nothing secret is hardcoded.
|
||||
- **Shutdown**: the service closes cleanly on `SIGTERM` (stops accepting,
|
||||
drains, exits).
|
||||
- **Headers**: responses carry `X-Content-Type-Options: nosniff`.
|
||||
- **Tests**: the suite covers error paths, not just the happy path.
|
||||
- **Changelog**: every shipped change has a `CHANGELOG.md` entry.
|
||||
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"name": "notes-service",
|
||||
"private": true,
|
||||
"type": "commonjs",
|
||||
"scripts": { "test": "node --test test/*.test.js" }
|
||||
}
|
||||
@@ -0,0 +1,50 @@
|
||||
'use strict';
|
||||
const http = require('node:http');
|
||||
|
||||
// Prototype state: happy path only.
|
||||
const notes = new Map();
|
||||
let nextId = 1;
|
||||
|
||||
function createApp() {
|
||||
return http.createServer((req, res) => {
|
||||
console.log('got a request');
|
||||
const url = new URL(req.url, 'http://localhost');
|
||||
|
||||
if (req.method === 'POST' && url.pathname === '/notes') {
|
||||
let body = '';
|
||||
req.on('data', chunk => { body += chunk; });
|
||||
req.on('end', () => {
|
||||
const parsed = JSON.parse(body);
|
||||
const id = `n_${nextId++}`;
|
||||
notes.set(id, { id, title: parsed.title, body: parsed.body });
|
||||
res.writeHead(201, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify(notes.get(id)));
|
||||
});
|
||||
return;
|
||||
}
|
||||
|
||||
const match = /^\/notes\/([\w-]+)$/.exec(url.pathname);
|
||||
if (req.method === 'GET' && match) {
|
||||
const note = notes.get(match[1]);
|
||||
if (!note) {
|
||||
res.writeHead(404);
|
||||
res.end('<html><body>not found</body></html>');
|
||||
return;
|
||||
}
|
||||
res.writeHead(200, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify(note));
|
||||
return;
|
||||
}
|
||||
|
||||
if (req.method === 'GET' && url.pathname === '/notes') {
|
||||
res.writeHead(200, { 'content-type': 'application/json' });
|
||||
res.end(JSON.stringify({ notes: [...notes.values()] }));
|
||||
return;
|
||||
}
|
||||
|
||||
res.writeHead(404);
|
||||
res.end('<html><body>not found</body></html>');
|
||||
});
|
||||
}
|
||||
|
||||
module.exports = { createApp };
|
||||
@@ -0,0 +1,6 @@
|
||||
'use strict';
|
||||
const { createApp } = require('./app');
|
||||
|
||||
createApp().listen(8080, () => {
|
||||
console.log('notes listening on 8080');
|
||||
});
|
||||
@@ -0,0 +1,21 @@
|
||||
'use strict';
|
||||
const test = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const { createApp } = require('../src/app');
|
||||
|
||||
test('create and read a note', async () => {
|
||||
const server = createApp();
|
||||
await new Promise(resolve => server.listen(0, '127.0.0.1', resolve));
|
||||
try {
|
||||
const port = server.address().port;
|
||||
const created = await fetch(`http://127.0.0.1:${port}/notes`, {
|
||||
method: 'POST', headers: { 'content-type': 'application/json' },
|
||||
body: JSON.stringify({ title: 'first', body: 'hello' }) });
|
||||
assert.equal(created.status, 201);
|
||||
const { id } = await created.json();
|
||||
const read = await fetch(`http://127.0.0.1:${port}/notes/${id}`);
|
||||
assert.equal((await read.json()).title, 'first');
|
||||
} finally {
|
||||
server.close();
|
||||
}
|
||||
});
|
||||
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"id": "production-ready",
|
||||
"category": "vague-hardening",
|
||||
"manualIds": ["skill:backend-patterns"],
|
||||
"checkTimeoutMs": 60000,
|
||||
"selection": {
|
||||
"id": "complex-production-ready",
|
||||
"category": "complex-vague-hardening",
|
||||
"expectedIds": ["skill:backend-patterns"]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
This goes to production Monday morning. Get it ready. The bar we hold production services to is in docs/production-bar.md — meet all of it, keep the documented API working, and leave the repo in a state you'd be comfortable being on-call for.
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user