From d6b91c5081157e4a2f1411e28f3ab2aff544c15f Mon Sep 17 00:00:00 2001 From: Hussein Mohamed Date: Sat, 5 Sep 2026 15:01:10 +0300 Subject: [PATCH 1/9] test(api-diagnostics): read a Signature B capture from the streams that carry it Four real Signature B occurrences were captured this session, the first ever with a parent-side exit status recorded. Every one of them exits 3221226505 (0xC0000409, the Windows __fastfail status) with the child log holding only `start` and `preload-installed`. Reading them exposed two blind spots that would have made the next capture say the wrong thing, and both are fixed here. The fatal markers were read from the parent runner's stderr, which structurally cannot contain them: Node's runner attaches a readline interface to each child's stderr and re-emits every line as a `test:stderr` reporter event, so a child's `FATAL ERROR:` banner reaches the report and never this process. Measured on a real bounded heap exhaustion as zero bytes on the parent's stderr against the full banner in the report. They now come from the report, attributed per failure rather than per run, and counting only the lines Node emitted as diagnostics. `readChildLogs` keyed on the test file path alone and concatenated event kinds across every log in the directory. Given a stale clean log and a fresh terminated one for the same file it merged them and announced that an external termination and a native fault were ruled out -- a false negative on the one hypothesis still standing, reachable through the preload's own documented default log directory. Events are now bucketed per process, keyed by log file and pid because a pid is not an identity on Windows, and a file with several processes on record reports that it cannot attribute rather than picking one. `--report-on-fatalerror` is added because it was measured to discriminate: a V8 heap exhaustion writes exactly one report named after the dying child's pid, while `process.abort()` and both external kills write none -- and all of them can leave exit code 134. `--package` lets one pass target `packages/persistence`, where Signature B has been observed since Increment 46 and where two of these four captures happened. No file moved and no existing invocation changed meaning. Signature B stays UNRESOLVED. No fix is proposed: the mechanism family is narrowed, the terminating party is not identified. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01CUPqvu66J5ZVyiv4nJ797r --- .../test/diagnostics/run-signature-b-pass.mjs | 92 ++--- .../diagnostics/signature-b-correlate.cjs | 152 +++++++- .../diagnostics/signature-b-correlate.test.ts | 346 ++++++++++++++++++ 3 files changed, 536 insertions(+), 54 deletions(-) diff --git a/packages/api/test/diagnostics/run-signature-b-pass.mjs b/packages/api/test/diagnostics/run-signature-b-pass.mjs index 013d1fca..7be37939 100644 --- a/packages/api/test/diagnostics/run-signature-b-pass.mjs +++ b/packages/api/test/diagnostics/run-signature-b-pass.mjs @@ -13,7 +13,9 @@ * throws, `spec` discards them (its `formatError` replaces the error with `error.cause`, the bare * string `'test failed'`), and `tap` serializes them into its YAML block. Running both recovers * the exit status without a custom reporter, without patching Node internals, and without changing - * what a human sees on stdout. + * what a human sees on stdout. It is also the only place a dying child's own fatal banner + * survives: the runner re-emits child stderr as `test:stderr` reporter events, so a heap + * exhaustion's `FATAL ERROR:` lands in the report and never on this process's stderr. * * The pass is **bounded by construction** and stops at the first capture. The wall-clock ceiling is * an absolute deadline enforced *during* a run, not merely consulted before starting another, and @@ -46,10 +48,11 @@ import { fileURLToPath } from 'node:url'; const HERE = path.dirname(fileURLToPath(import.meta.url)); const API_DIR = path.resolve(HERE, '../..'); -const { correlate, parseTapFailures, readChildLogs } = createRequire(import.meta.url)('./signature-b-correlate.cjs'); +const { correlate, fatalMarkersIn, parseTapFailures, readChildLogs } = + createRequire(import.meta.url)('./signature-b-correlate.cjs'); /** Every option this script accepts, so one cannot be consumed as another's value. */ -const KNOWN_FLAGS = new Set(['runs', 'max-minutes', 'out', 'target']); +const KNOWN_FLAGS = new Set(['runs', 'max-minutes', 'out', 'package', 'target']); /** Refuse before anything is created or spawned; a usage error must not leave artifacts behind. */ function usageError(message) { @@ -78,35 +81,6 @@ const flag = (name, fallback) => { return value; }; -/** - * The only stderr lines a dying child produces that name its own cause. Node prints these itself, - * and they are what separates a V8 heap exhaustion from a JS `process.abort()` when both leave the - * same exit code 134. - */ -const FATAL_MARKERS = [ - 'FATAL ERROR:', - 'JavaScript heap out of memory', - '<--- Last few GCs --->', - '----- Native stack trace -----', - '----- JavaScript stack trace -----', -]; - -/** - * Record which fatal markers a run's stderr contained, and never the stderr itself. - * - * A whitelist rather than a redaction pass: the suite's own output can contain anything a test - * chose to print, so matching known markers is the only way to be sure a token, request body or - * connection string cannot reach the capture file. Presence is all the evidence is worth anyway — - * the marker names the mechanism, the surrounding text does not. - * - * @param {string | undefined} stderr - * @returns {string[]} markers present, in the order listed above - */ -function fatalMarkersIn(stderr) { - const text = String(stderr ?? ''); - return FATAL_MARKERS.filter((marker) => text.includes(marker)); -} - /** * A ceiling that does not describe a real experiment must stop the run, not shrink it. * @@ -144,6 +118,12 @@ const outDir = flag('out', null) ?? fs.mkdtempSync(path.join(tmpdir(), 'sigb-pas // file instead of launching a second copy of the suite against the same database. const target = flag('target', 'dist-test/test/**/*.test.js'); +// Which package a run executes in. Signature B has been observed in `packages/persistence` as well +// as `packages/api`, and a runner that can only ever run one of them cannot capture it in the +// other. Resolved against `packages/api` so `--package ../persistence` reads the way it is written, +// and defaulting to this package so every existing invocation means exactly what it meant before. +const packageDir = path.resolve(API_DIR, String(flag('package', '.'))); + /** * Make the artifact directory owner-only, or refuse to write into it. * @@ -215,6 +195,7 @@ function killTree(child) { * * @param {string} tapPath * @param {string} logDir + * @param {string} reportDir * @param {number} budgetMs * @returns {Promise<{ status: number|null, signal: string|null, timedOut: boolean, error: Error|null, stderr: string }>} */ @@ -238,18 +219,25 @@ for (const signal of ['SIGINT', 'SIGTERM']) { }); } -function runOnce(tapPath, logDir, budgetMs) { +function runOnce(tapPath, logDir, reportDir, budgetMs) { return new Promise((resolve) => { const child = spawn( process.execPath, [ - '--require', './test/diagnostics/signature-b-preload.cjs', + '--require', path.join(HERE, 'signature-b-preload.cjs'), + // Node writes one of these, named after the pid that died, when it dies of its OWN fatal + // error — a V8 heap exhaustion, an internal assertion. Measured here: a heap exhaustion + // leaves a report, while `process.abort()` and an external kill leave none, and all three + // can leave exit code 134. The report is therefore what separates a fault Node suffered + // from a status something else chose, and unlike the printed banner it is attributable to + // an exact process rather than to a position in the report stream. + '--report-on-fatalerror', `--report-directory=${reportDir}`, '--test-reporter=spec', '--test-reporter-destination=stdout', '--test-reporter=tap', `--test-reporter-destination=${tapPath}`, '--test', '--test-concurrency=1', target, ], { - cwd: API_DIR, + cwd: packageDir, env: { ...process.env, SIGB_LOG_DIR: logDir }, // `spec` goes straight to this terminal, which is the whole point of running it alongside // `tap`: the human-facing output must stay exactly what it would be without the diagnostic. @@ -298,11 +286,13 @@ for (let run = 1; run <= maxRuns; run++) { const tapPath = path.join(outDir, `run${run}.tap`); const logDir = path.join(outDir, `child-logs-run${run}`); + const reportDir = path.join(outDir, `reports-run${run}`); fs.mkdirSync(logDir, { recursive: true, mode: 0o700 }); + fs.mkdirSync(reportDir, { recursive: true, mode: 0o700 }); const freeBefore = freemem(); const startedRun = Date.now(); - const result = await runOnce(tapPath, logDir, Math.max(1, maxMs - elapsed)); + const result = await runOnce(tapPath, logDir, reportDir, Math.max(1, maxMs - elapsed)); const durationMs = Date.now() - startedRun; // A run whose report cannot be read has not answered the question, and must not be filed next to @@ -313,8 +303,9 @@ for (let run = 1; run <= maxRuns; run++) { if (collectionError === null) { try { const isWin = process.platform === 'win32'; + const tapText = fs.readFileSync(tapPath, 'utf8'); records = correlate( - parseTapFailures(fs.readFileSync(tapPath, 'utf8'), { isWindows: isWin }), + parseTapFailures(tapText, { isWindows: isWin }), readChildLogs(logDir), { isWindows: isWin }, ); @@ -323,6 +314,15 @@ for (let run = 1; run <= maxRuns; run++) { } } + // Names only. A report body carries the whole command line and every environment variable, so + // the file stays on disk for a reader who wants it and never reaches an artifact this prints. + let reportFiles = []; + try { + reportFiles = fs.readdirSync(reportDir); + } catch { + /* nothing wrote one, which is itself the observation */ + } + const summary = { run, durationMs, @@ -335,6 +335,7 @@ for (let run = 1; run <= maxRuns; run++) { freeMemAfterMb: Math.round(freemem() / 1048576), totalMemMb: Math.round(totalmem() / 1048576), captures: records.length, + reportFiles: reportFiles.length, collectionError, }; runs.push(summary); @@ -352,11 +353,19 @@ for (let run = 1; run <= maxRuns; run++) { } if (records.length > 0) { - capture = { run, records, fatalMarkers: fatalMarkersIn(result.stderr) }; + // A child's fatal banner reaches the REPORT, never this process's stderr: the runner reads each + // child's stderr line by line and re-emits it as a `test:stderr` reporter event. Measured on a + // real bounded heap exhaustion, this process's stderr held zero bytes while the report held the + // whole banner. So the markers come from each record's own region of the report, and this + // process's stderr is recorded separately, as what it actually is: the runner's own. + const fatalMarkers = [...new Set(records.flatMap((record) => record.fatalMarkers))]; + capture = { run, records, fatalMarkers, reportFiles, runnerStderrMarkers: fatalMarkersIn(result.stderr) }; fs.writeFileSync(path.join(outDir, 'capture.json'), `${JSON.stringify(capture, null, 2)}\n`, { mode: 0o600 }); - // The raw TAP has served its purpose the moment `capture.json` exists. It is a full transcript - // of whatever the suite printed, so keeping it past the normalised record retains arbitrary - // test output for no diagnostic gain; the child logs stay, being enumerated lifecycle events. + // The raw TAP has served its purpose the moment `capture.json` exists — which now includes the + // fatal markers read out of it, the evidence the first real capture of this investigation lost + // by looking for them on the parent's stderr instead. It is a full transcript of whatever the + // suite printed, so keeping it past the normalised record retains arbitrary test output for no + // diagnostic gain; the child logs stay, being enumerated lifecycle events. fs.rmSync(tapPath, { force: true }); console.log(`\nCAPTURED on run ${run}:\n${records.map((r) => JSON.stringify(r, null, 2)).join('\n')}`); console.log(`\nRaw TAP discarded; the normalised record is ${path.join(outDir, 'capture.json')}.`); @@ -365,6 +374,7 @@ for (let run = 1; run <= maxRuns; run++) { // A clean run's artifacts answer nothing and would otherwise grow without bound across the pass. fs.rmSync(logDir, { recursive: true, force: true }); + fs.rmSync(reportDir, { recursive: true, force: true }); fs.rmSync(tapPath, { force: true }); } diff --git a/packages/api/test/diagnostics/signature-b-correlate.cjs b/packages/api/test/diagnostics/signature-b-correlate.cjs index 7183662d..f0e12af9 100644 --- a/packages/api/test/diagnostics/signature-b-correlate.cjs +++ b/packages/api/test/diagnostics/signature-b-correlate.cjs @@ -173,6 +173,62 @@ function classifyTermination({ exitCode, signal }) { return { id: 'unclassified', meaning: `exit code ${String(exitCode)} is not in the measured table`, specific: false }; } +/** + * The only lines a dying child prints that name its own cause. + * + * Node prints these itself, and they are what separates a V8 heap exhaustion from a JS + * `process.abort()` or an external kill when all three can leave exit code 134. + */ +const FATAL_MARKERS = [ + 'FATAL ERROR:', + 'JavaScript heap out of memory', + '<--- Last few GCs --->', + '----- Native stack trace -----', + '----- JavaScript stack trace -----', +]; + +/** + * The lines of a TAP region that are Node's own diagnostics, and nothing else. + * + * A child's stderr reaches the report as `# ` comment lines, which is what makes a banner readable + * at all. Ordinary suite output reaches it too — as YAML bodies, as quoted error messages — and a + * substring scan over the raw region cannot tell a test that PRINTED `FATAL ERROR:` from a child + * that died of one. Keeping only the diagnostic lines is not a complete defence (a test that writes + * the banner to its own stderr is indistinguishable by construction) but it removes every case + * where the text was never presented as a diagnostic in the first place. + * + * @param {string} region + * @returns {string} the region's `# ` lines, joined + */ +function childDiagnostics(region) { + return String(region) + .split('\n') + .filter((line) => /^#\s/.test(line)) + .join('\n'); +} + +/** + * Which fatal markers a run's reporter output contained, and never the output itself. + * + * Read the REPORT, not the parent's stderr. Node's runner attaches a readline interface to each + * child's stderr and re-emits every line as a `test:stderr` reporter event, so a child's fatal + * banner reaches the TAP file as `# ` diagnostics and never reaches the parent process's own + * stderr at all — measured here on a real bounded heap exhaustion as zero bytes on the parent's + * stderr against the full banner in the TAP. + * + * A whitelist rather than a redaction pass: a suite's output can contain anything a test chose to + * print, so matching known markers is the only way to be sure a token, request body or connection + * string cannot reach a capture file. Presence is all the evidence is worth anyway — the marker + * names the mechanism, the surrounding text does not. + * + * @param {string | undefined} text Reporter output, typically the TAP report. + * @returns {string[]} markers present, in the order listed above + */ +function fatalMarkersIn(text) { + const haystack = String(text ?? ''); + return FATAL_MARKERS.filter((marker) => haystack.includes(marker)); +} + /** * Events that name *why* the process ended: a call the preload intercepted, or a fault it observed. */ @@ -257,13 +313,35 @@ function narrowFromChildEvidence(classification, childKinds) { /** * @param {string} tapText * @param {{ isWindows?: boolean }} [options] - * @returns {Array<{ file: string, exitCode: number | null, signal: string | null, failureType: string | null, durationMs: number | null, isWindows?: boolean }>} + * @returns {Array<{ file: string, exitCode: number | null, signal: string | null, failureType: string | null, durationMs: number | null, fatalMarkers: string[], isWindows?: boolean }>} */ function parseTapFailures(tapText, options = {}) { - const blocks = String(tapText).matchAll(/^not ok \d+ - (.+?)$\r?\n\s*---\r?\n([\s\S]*?)^\s*\.\.\.$/gm); + const text = String(tapText); + const blocks = text.matchAll(/^not ok \d+ - (.+?)$\r?\n\s*---\r?\n([\s\S]*?)^\s*\.\.\.$/gm); + // Where each reported test's own output ends, so a run's diagnostic lines can be split between + // the files that produced them. Node emits a child's stderr as top-level `# ` diagnostics as it + // arrives, ahead of the point line for the file that produced it, so everything between the end of + // the previous report and this point line belongs to this file. + // + // The boundary is the end of the previous test's YAML block, not the end of its point line: the + // block comes AFTER the point line, so stopping at the line would leave the previous failure's + // whole body — its error message, its stack — inside the next file's region, and a suite that + // fails while quoting one of these banners would hand it to whichever file failed next. + // + // Attribution needs one child reporting at a time. The two packages this runner targets both pass + // `--test-concurrency=1` (`packages/api` and `packages/persistence`; no other workspace package + // does). A report produced at Node's default concurrency interleaves several children's output and + // is not attributable this way — reading one here would be reading a different experiment. + const regionEnds = []; + for (const point of text.matchAll(/^(?:not )?ok \d+ - .*$/gm)) { + const afterPoint = point.index + point[0].length; + const yaml = /^\r?\n\s*---\r?\n[\s\S]*?\r?\n\s*\.\.\.[^\r\n]*/.exec(text.slice(afterPoint)); + regionEnds.push(yaml === null ? afterPoint : afterPoint + yaml[0].length); + } const out = []; const isWin = typeof options?.isWindows === 'boolean' ? options.isWindows : undefined; - for (const [, rawName, body] of blocks) { + for (const match of blocks) { + const [, rawName, body] = match; if (!/^\s*error:\s*['"]?test failed['"]?\s*$/m.test(body)) continue; if (!/^\s*failureType:\s*['"]?testCodeFailure['"]?\s*$/m.test(body)) continue; if (out.length >= MAX_RECORDS) break; @@ -277,7 +355,13 @@ function parseTapFailures(tapText, options = {}) { const dur = unquote(scalar('duration_ms')); const parsedExit = exit === null || exit === '~' || exit === '' ? null : Number(exit); const parsedDur = dur === null || dur === '~' || dur === '' ? null : Number(dur); + let sliceStart = 0; + for (const end of regionEnds) { + if (end >= match.index) break; + sliceStart = end; + } out.push({ + fatalMarkers: fatalMarkersIn(childDiagnostics(text.slice(sliceStart, match.index))), file: rawName.trim().replace(/\\\\/g, '\\'), exitCode: parsedExit !== null && Number.isFinite(parsedExit) ? parsedExit : null, signal: sig === null || sig === '~' ? null : sig, @@ -298,7 +382,9 @@ function parseTapFailures(tapText, options = {}) { * final line, and losing that line must not lose the rest of the record. * * @param {string} logDir - * @returns {Map} keyed by the child's whole normalised path + * @returns {Map} keyed by the + * child's whole normalised path. Each entry counts the processes that ran that file, and reports + * a pid and its events only when there was exactly one. */ function readChildLogs(logDir) { const byFile = new Map(); @@ -337,23 +423,60 @@ function readChildLogs(logDir) { if (isWin) { const lower = normalized.toLowerCase(); for (const existingKey of byFile.keys()) { - const entry = byFile.get(existingKey); - if (entry?.isWindows && existingKey.toLowerCase() === lower) { + const candidate = byFile.get(existingKey); + if (candidate?.isWindows && existingKey.toLowerCase() === lower) { key = existingKey; break; } } } - const existing = byFile.get(key) ?? { pid: null, kinds: [], isWindows: false }; - existing.pid = typeof record.pid === 'number' ? record.pid : existing.pid; - if (typeof record.kind === 'string') existing.kinds.push(record.kind); + // Events are bucketed by the process that emitted them, never appended to one list. One + // directory can hold logs from more than one run of the same file — the preload defaults to a + // shared temp directory nothing cleans — and concatenating those would put an earlier run's + // `exit` into a later run's evidence, which is exactly what makes `narrowFromChildEvidence` + // announce that an external termination and a native fault are ruled out. + // + // The bucket is keyed by the log FILE as well as the pid, because a pid is not an identity: + // Windows reuses them freely, and two runs of one file in one directory can carry the same + // number. Each process writes its own `run--.jsonl`, so the file supplies the part + // the pid cannot. + const pid = typeof record.pid === 'number' ? record.pid : null; + const existing = byFile.get(key) ?? { processes: new Map(), isWindows: false }; + const bucketKey = `${entry}\u0000${pid === null ? 'unknown' : pid}`; + const bucket = existing.processes.get(bucketKey) ?? { pid, kinds: [] }; + if (typeof record.kind === 'string') bucket.kinds.push(record.kind); + existing.processes.set(bucketKey, bucket); if (isWin) existing.isWindows = true; byFile.set(key, existing); } } + // A file with one process on record reports that process. A file with several is not evidence + // about either of them, and says so through `pidCount` rather than by picking one. + for (const record of byFile.values()) { + const buckets = [...record.processes.values()]; + record.pidCount = buckets.length; + record.pid = buckets.length === 1 ? buckets[0].pid : null; + record.kinds = buckets.length === 1 ? buckets[0].kinds : []; + } return byFile; } +/** + * One path matched one entry — but an entry holding several processes is still not evidence. + * + * Two runs of the same file in one log directory match that file equally well, and there is no + * sound way to choose between them: the newest is not necessarily the failing one, and attaching + * either one's lifecycle events to the other's failure would make the diagnostic assert something + * false. Reporting that it does not know is the only honest answer. + * + * @param {{ pidCount?: number, pid: number | null, kinds: string[] }} child + */ +function single(child) { + const count = child?.pidCount ?? 1; + if (count > 1) return { child: null, ambiguous: true, candidates: count }; + return { child, ambiguous: false, candidates: 1 }; +} + /** * Find the child log belonging to one parent failure, or say that it cannot be told apart. * @@ -383,7 +506,7 @@ function matchChild(childLogs, parentFile, parentIsWindows) { const keyNorm = isWin ? key.toLowerCase() : key; return keyNorm === wantedNorm || keyNorm.endsWith(`/${wantedNorm}`); }); - if (bySuffix.length === 1) return { child: bySuffix[0][1], ambiguous: false, candidates: 1 }; + if (bySuffix.length === 1) return single(bySuffix[0][1]); if (bySuffix.length > 1) return { child: null, ambiguous: true, candidates: bySuffix.length }; // Basename fallback is ONLY for a parent path too short to disambiguate (i.e. bare filename with no slashes). @@ -399,7 +522,7 @@ function matchChild(childLogs, parentFile, parentIsWindows) { } return kBase === base; }); - if (byBase.length === 1) return { child: byBase[0][1], ambiguous: false, candidates: 1 }; + if (byBase.length === 1) return single(byBase[0][1]); if (byBase.length > 1) return { child: null, ambiguous: true, candidates: byBase.length }; } @@ -424,7 +547,7 @@ function correlate(failures, childLogs, options = {}) { const defaultIsWin = typeof options?.isWindows === 'boolean' ? options.isWindows : undefined; return failures.map((rawFailure) => { const failure = typeof rawFailure === 'string' - ? { file: rawFailure, exitCode: null, signal: null, failureType: null, durationMs: null } + ? { file: rawFailure, exitCode: null, signal: null, failureType: null, durationMs: null, fatalMarkers: [] } : rawFailure; const parentIsWin = typeof failure?.isWindows === 'boolean' ? failure.isWindows @@ -448,12 +571,13 @@ function correlate(failures, childLogs, options = {}) { pid: match.child?.pid ?? null, kinds, }, + fatalMarkers: failure.fatalMarkers ?? [], classification: classification.id, meaning: classification.meaning, specific: classification.specific, narrowed: match.ambiguous ? false : narrowing.narrowed, statement: match.ambiguous - ? `${match.candidates} child logs match this file's name and none matches its path, so no lifecycle evidence can be attributed to it without guessing` + ? `${match.candidates} child processes are on record for this file — a second directory of the same name, or a second run left in the same log directory — so no lifecycle evidence can be attributed to it without guessing` : narrowing.statement, }; }); @@ -461,9 +585,11 @@ function correlate(failures, childLogs, options = {}) { module.exports = { EXIT_CODE_TABLE, + FATAL_MARKERS, MAX_RECORDS, classifyTermination, correlate, + fatalMarkersIn, isWindowsPath, matchChild, narrowFromChildEvidence, diff --git a/packages/api/test/diagnostics/signature-b-correlate.test.ts b/packages/api/test/diagnostics/signature-b-correlate.test.ts index ac31b24e..7e6b624d 100644 --- a/packages/api/test/diagnostics/signature-b-correlate.test.ts +++ b/packages/api/test/diagnostics/signature-b-correlate.test.ts @@ -31,6 +31,7 @@ interface CorrelatedRecord { pid: number | null; kinds: string[]; }; + fatalMarkers: string[]; classification: string; meaning: string; specific: boolean; @@ -48,6 +49,7 @@ const correlator = require(CORRELATE_PATH) as { specific: boolean; }; correlate(failures: unknown[], childLogs: Map, options?: { isWindows?: boolean }): CorrelatedRecord[]; + fatalMarkersIn(text: string | undefined): string[]; narrowFromChildEvidence( c: { id: string; specific: boolean }, kinds: readonly string[], @@ -58,6 +60,7 @@ const correlator = require(CORRELATE_PATH) as { signal: string | null; failureType: string | null; durationMs: number | null; + fatalMarkers: string[]; isWindows?: boolean; }>; readChildLogs(dir: string): Map; @@ -1179,3 +1182,346 @@ test('correlator: CLI preserves POSIX semantics when log directory contains mixe fs.rmSync(tmpDir, { recursive: true, force: true }); } }); + +test('correlator: a stale log from an earlier run is never merged into a later run\'s evidence', () => { + const logDir = fs.mkdtempSync(path.join(os.tmpdir(), 'sigb-stale-')); + try { + // Two runs of the SAME file, in one directory. That is not a contrivance: the preload defaults + // SIGB_LOG_DIR to os.tmpdir()/sigb-diag, which is its documented manual usage and which nothing + // ever cleans, and the correlator CLI takes --child-logs pointed at whatever a reader chose. + const clean = ['start', 'preload-installed', 'beforeExit', 'exit'].map((kind) => ({ + kind, pid: 1111, testFile: 'C:\\repo\\dist-test\\test\\studies-api.test.js', + })); + const terminated = ['start', 'preload-installed'].map((kind) => ({ + kind, pid: 2222, testFile: 'C:\\repo\\dist-test\\test\\studies-api.test.js', + })); + fs.writeFileSync(path.join(logDir, 'run-1111-1000.jsonl'), `${clean.map((l) => JSON.stringify(l)).join('\n')}\n`); + fs.writeFileSync(path.join(logDir, 'run-2222-2000.jsonl'), `${terminated.map((l) => JSON.stringify(l)).join('\n')}\n`); + + const records = correlator.correlate( + correlator.parseTapFailures(tapFailure('dist-test/test/studies-api.test.js', '1')), + correlator.readChildLogs(logDir), + { isWindows: true }, + ); + + assert.equal(records.length, 1); + // The failure this guards against is not cosmetic. Concatenating the two processes' events puts + // `exit` in the list, and `narrowFromChildEvidence` then states that an external termination and + // a native fault are ruled out — a false negative on the one hypothesis still standing. + assert.equal(records[0]?.child.ambiguous, true, 'two processes ran this file; neither one is the evidence'); + assert.equal(records[0]?.child.candidates, 2); + assert.deepEqual(records[0]?.child.kinds, [], 'events from two different processes must never be concatenated'); + assert.equal(records[0]?.narrowed, false); + assert.doesNotMatch( + String(records[0]?.statement), + /rules out an external termination/, + 'a stale log must never be allowed to rule out the mechanism under investigation', + ); + } finally { + fs.rmSync(logDir, { recursive: true, force: true }); + } +}); + +test('correlator: reads a child fatal banner from the report Node routes it to, not the parent stderr', () => { + // Measured this session on a real bounded V8 heap exhaustion: the parent runner's stderr was zero + // bytes, while the TAP report carried the whole banner as `# ` diagnostic lines. Node's runner + // attaches a readline Interface to each child's stderr and re-emits every line as a `test:stderr` + // report event, so the parent's own stderr structurally cannot hold a child's fatal output. + const tapText = [ + 'TAP version 13', + '# <--- Last few GCs --->', + '# FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory', + '# ----- Native stack trace -----', + tapFailure('dist-test/test/studies-api.test.js', '134'), + '1..1', + ].join('\n'); + + assert.deepEqual( + correlator.fatalMarkersIn(tapText), + ['FATAL ERROR:', 'JavaScript heap out of memory', '<--- Last few GCs --->', '----- Native stack trace -----'], + 'the markers that name a V8 fatal must be recoverable from the reporter output', + ); + assert.deepEqual(correlator.fatalMarkersIn(''), [], 'an empty stream names nothing'); + assert.deepEqual(correlator.fatalMarkersIn(undefined), [], 'a missing stream names nothing'); +}); + +test('pass runner: a capture keeps the fatal markers that name the mechanism', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'sigb-markers-')); + try { + // `fs.writeSync(2, ...)` rather than `process.stderr.write`: stderr is a pipe here, its writes + // are asynchronous, and an exit on the next line would race them away — which would make this + // test flake for a reason that has nothing to do with what it is checking. + const target = path.join(dir, 'fatal.test.cjs'); + fs.writeFileSync( + target, + "require('node:fs').writeSync(2, 'FATAL ERROR: Reached heap limit Allocation failed - " + + "JavaScript heap out of memory\\n');\nprocess.exit(134);\n", + ); + const out = path.join(dir, 'out'); + const result = runPass(['--runs', '1', '--max-minutes', '1', '--out', out, '--target', target], dir); + + const capturePath = path.join(out, 'capture.json'); + assert.ok(fs.existsSync(capturePath), `the pass must capture this file-level failure:\n${result.stdout}${result.stderr}`); + const capture = JSON.parse(fs.readFileSync(capturePath, 'utf8')) as { fatalMarkers: string[] }; + assert.deepEqual( + capture.fatalMarkers, + ['FATAL ERROR:', 'JavaScript heap out of memory'], + 'the capture must record the banner the child actually printed', + ); + } finally { + fs.rmSync(dir, { recursive: true, force: true }); + } +}); + +test('pass runner: --package points a pass at another workspace package', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'sigb-package-')); + try { + // Signature B has been observed in packages/persistence as well as packages/api, and a runner + // that can only ever run one package cannot capture it in the other. The target is resolved + // inside the named package, so a pass that finds this file at all proves it ran there. + const pkg = path.join(dir, 'other-package'); + fs.mkdirSync(pkg, { recursive: true }); + trivialTarget(pkg); + const out = path.join(dir, 'out'); + const result = runPass(['--runs', '1', '--max-minutes', '1', '--out', out, '--package', pkg, '--target', 'noop.test.cjs'], dir); + + assert.equal(result.status, 0, `${result.stdout}${result.stderr}`); + assert.match(String(result.stdout), /1 run\(s\)/); + assert.match(String(result.stdout), /Signature B was not observed/); + // A pass whose runner failed to start still prints that summary: zero captures is what a run + // that never ran also produces. The runner's own exit status is what separates them, and it is + // also what catches a preload this file could no longer resolve from the other package's + // directory — `node --require ` exits non-zero before any test runs. + const summary = JSON.parse(fs.readFileSync(path.join(out, 'summary.json'), 'utf8')) as { + runs: Array<{ runnerExitCode: number | null; collectionError: string | null }>; + }; + assert.equal(summary.runs.length, 1); + assert.equal(summary.runs[0]?.runnerExitCode, 0, 'the run must have executed, preload and all'); + assert.equal(summary.runs[0]?.collectionError, null); + } finally { + fs.rmSync(dir, { recursive: true, force: true }); + } +}); + +test('correlator: a fatal banner from an earlier file is not attributed to a later file\'s failure', () => { + // A whole-run scan cannot survive this suite's own diagnostics: the pass runner lets a child's + // `spec` output through to the terminal, so a test that deliberately prints a fatal banner puts + // that banner into the outer run's report. Measured while hunting a real occurrence — every full + // `packages/api` run reported two fatal markers that this file had printed itself. Under + // `--test-concurrency=1` the child that printed them is the one whose result comes next, so + // attribution is exact rather than a guess. + const tapText = [ + 'TAP version 13', + '# FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory', + '# <--- Last few GCs --->', + '# Subtest: dist-test/test/noisy.test.js', + 'ok 1 - dist-test/test/noisy.test.js', + ' ---', + ' duration_ms: 10', + ' ...', + tapFailure('dist-test/test/studies-api.test.js', '134').replace('not ok 1 -', 'not ok 2 -'), + '1..2', + ].join('\n'); + + const failures = correlator.parseTapFailures(tapText, { isWindows: false }); + + assert.equal(failures.length, 1, 'only the bare file-level failure is a candidate'); + assert.deepEqual( + failures[0]?.fatalMarkers, + [], + 'the banner belongs to the file that printed it, which passed, not to the file that failed later', + ); +}); + +test('correlator: a fatal banner printed by the file that then failed is attributed to it', () => { + const tapText = [ + 'TAP version 13', + '# Subtest: dist-test/test/quiet.test.js', + 'ok 1 - dist-test/test/quiet.test.js', + ' ---', + ' duration_ms: 10', + ' ...', + '# <--- Last few GCs --->', + '# FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory', + tapFailure('dist-test/test/studies-api.test.js', '134').replace('not ok 1 -', 'not ok 2 -'), + '1..2', + ].join('\n'); + + const failures = correlator.parseTapFailures(tapText, { isWindows: false }); + + assert.equal(failures.length, 1); + assert.deepEqual( + failures[0]?.fatalMarkers, + ['FATAL ERROR:', 'JavaScript heap out of memory', '<--- Last few GCs --->'], + 'a banner printed after the previous file finished belongs to the file that failed next', + ); +}); + +test('pass runner: a diagnostic report separates a real V8 fatal from a status that merely looks like one', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'sigb-report-')); + try { + // Exit code 134 is produced by a V8 heap exhaustion, by `process.abort()`, and by any process + // that chooses to exit 134 — measured all three this session. Node writes a diagnostic report + // for the first and not for the others, and names the file after the pid that died, so the + // report is what tells them apart even when nothing was printed. + const reportDir = path.join(dir, 'reports'); + fs.mkdirSync(reportDir); + const fatal = path.join(dir, 'oom.cjs'); + // Bounded by construction: a 64 MB ceiling and an allocation loop that stops at 400 MB of + // intent, so this can never reach for the host's real memory. + fs.writeFileSync(fatal, 'const held = [];\nfor (let i = 0; i < 400; i++) held.push(new Array(131072).fill(i));\n'); + const crashed = spawnSync( + process.execPath, + ['--max-old-space-size=64', '--report-on-fatalerror', `--report-directory=${reportDir}`, fatal], + { encoding: 'utf8', timeout: 60_000 }, + ); + + assert.equal(crashed.status, 134, 'a heap exhaustion leaves 134, the same status three mechanisms leave'); + const reports = fs.readdirSync(reportDir); + assert.equal(reports.length, 1, 'a V8 fatal writes exactly one report'); + assert.ok( + reports[0]?.includes(`.${crashed.pid}.`), + `the report must name the process that died (pid ${crashed.pid}, got ${reports[0]})`, + ); + + // The same status, chosen deliberately rather than reached by a fault, writes none. + const quietDir = path.join(dir, 'quiet-reports'); + fs.mkdirSync(quietDir); + const chosen = path.join(dir, 'chosen.cjs'); + fs.writeFileSync(chosen, 'process.exit(134);\n'); + const exited = spawnSync( + process.execPath, + ['--report-on-fatalerror', `--report-directory=${quietDir}`, chosen], + { encoding: 'utf8', timeout: 60_000 }, + ); + + assert.equal(exited.status, 134); + assert.deepEqual(fs.readdirSync(quietDir), [], 'choosing the status of a fault is not a fault'); + } finally { + fs.rmSync(dir, { recursive: true, force: true }); + } +}); + +test('pass runner: exit 134 alone names nothing, and the capture carries what separates the two', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'sigb-capture-reports-')); + try { + // Two files that die with the SAME status by different means. A capture that could not tell them + // apart would be the whole defect this instrumentation exists to fix, so the test runs both + // through the real pass runner and asserts the pair of channels that separates them. + // + // First: a genuine V8 heap exhaustion. The ceiling reaches the child through NODE_OPTIONS, and + // is bounded by construction — 80 MB, with the allocation loop stopping at 400 MB of intent — + // so this can never reach for the host's real memory. + const oom = path.join(dir, 'oom.test.cjs'); + fs.writeFileSync(oom, 'const held = [];\nfor (let i = 0; i < 400; i++) held.push(new Array(131072).fill(i));\n'); + const oomOut = path.join(dir, 'oom-out'); + const oomResult = spawnSync( + process.execPath, + [path.join(DIAGNOSTICS_DIR, 'run-signature-b-pass.mjs'), '--runs', '1', '--max-minutes', '2', '--out', oomOut, '--target', oom], + { cwd: dir, encoding: 'utf8', env: { ...process.env, NODE_TEST_CONTEXT: undefined, NODE_OPTIONS: '--max-old-space-size=80' } }, + ); + const oomCapturePath = path.join(oomOut, 'capture.json'); + assert.ok(fs.existsSync(oomCapturePath), `a heap exhaustion must be captured:\n${oomResult.stdout}${oomResult.stderr}`); + const fatal = JSON.parse(fs.readFileSync(oomCapturePath, 'utf8')) as { + reportFiles: string[]; + fatalMarkers: string[]; + records: Array<{ parent: { exitCode: number }; child: { kinds: string[] } }>; + }; + + assert.equal(fatal.records[0]?.parent.exitCode, 134); + assert.deepEqual(fatal.records[0]?.child.kinds, ['start', 'preload-installed'], + 'a V8 fatal runs no JS exit path, which is why the status alone cannot name it'); + assert.ok(fatal.fatalMarkers.includes('JavaScript heap out of memory'), + `the banner must be recovered from the report Node routed it to, got ${JSON.stringify(fatal.fatalMarkers)}`); + assert.equal(fatal.reportFiles.length, 1, 'a V8 fatal writes exactly one diagnostic report'); + + // Second: the same status, chosen rather than suffered. Neither channel fires. + const chosen = path.join(dir, 'chosen.test.cjs'); + fs.writeFileSync(chosen, 'process.exit(134);\n'); + const chosenOut = path.join(dir, 'chosen-out'); + const chosenResult = runPass(['--runs', '1', '--max-minutes', '1', '--out', chosenOut, '--target', chosen], dir); + const chosenCapturePath = path.join(chosenOut, 'capture.json'); + assert.ok(fs.existsSync(chosenCapturePath), `${chosenResult.stdout}${chosenResult.stderr}`); + const quiet = JSON.parse(fs.readFileSync(chosenCapturePath, 'utf8')) as { + reportFiles: string[]; + fatalMarkers: string[]; + records: Array<{ parent: { exitCode: number } }>; + }; + + assert.equal(quiet.records[0]?.parent.exitCode, 134, 'the same status as the fault above'); + assert.deepEqual(quiet.fatalMarkers, [], 'nothing was printed'); + assert.deepEqual(quiet.reportFiles, [], 'and no report was written, which is what tells them apart'); + } finally { + fs.rmSync(dir, { recursive: true, force: true }); + } +}); + +test('correlator: a marker inside an earlier failure\'s own report is not attributed to the next file', () => { + // A TAP point line is followed by that test's YAML block, so anchoring the region to the point + // line alone leaves the previous failure's whole body inside the next failure's region. A suite + // that asserts on a formatter, or that fails with a message quoting one of these banners, would + // then hand that banner to whichever file failed next. The region has to start after the block + // ends, and only Node's own `# ` diagnostic lines count as a child's output. + const tapText = [ + 'TAP version 13', + '# Subtest: dist-test/test/formatter.test.js', + 'not ok 1 - dist-test/test/formatter.test.js', + ' ---', + ' duration_ms: 5', + " failureType: 'subtestsFailed'", + " error: 'expected FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory'", + ' ...', + 'a bare line mentioning <--- Last few GCs ---> that Node never emitted as a diagnostic', + tapFailure('dist-test/test/studies-api.test.js', '134').replace('not ok 1 -', 'not ok 2 -'), + '1..2', + ].join('\n'); + + const failures = correlator.parseTapFailures(tapText, { isWindows: false }); + + assert.equal(failures.length, 1, 'the first block failed a subtest, so only the second is a candidate'); + assert.deepEqual( + failures[0]?.fatalMarkers, + [], + 'neither an earlier failure\'s report body nor an undiagnosed bare line is this file\'s banner', + ); +}); + +test('correlator: two runs that reused one pid are not merged into a single process', () => { + const logDir = fs.mkdtempSync(path.join(os.tmpdir(), 'sigb-pidreuse-')); + try { + // Bucketing by pid alone is not enough: Windows reuses pids freely, and the preload's default + // log directory is never cleaned, so the same pid can name two different processes that both + // ran this file. Merging them puts the earlier run's `exit` into the later run's evidence, which + // is the false negative the per-process bucketing exists to prevent. Each process writes its own + // file — `run--.jsonl` — so the file is the identity the pid alone cannot supply. + const line = (kind: string) => JSON.stringify({ + kind, pid: 1111, testFile: 'C:\\repo\\dist-test\\test\\studies-api.test.js', + }); + fs.writeFileSync( + path.join(logDir, 'run-1111-1000.jsonl'), + `${['start', 'preload-installed', 'beforeExit', 'exit'].map(line).join('\n')}\n`, + ); + fs.writeFileSync( + path.join(logDir, 'run-1111-2000.jsonl'), + `${['start', 'preload-installed'].map(line).join('\n')}\n`, + ); + + const records = correlator.correlate( + correlator.parseTapFailures(tapFailure('dist-test/test/studies-api.test.js', '1')), + correlator.readChildLogs(logDir), + { isWindows: true }, + ); + + assert.equal(records.length, 1); + assert.equal(records[0]?.child.ambiguous, true, 'one pid, two processes: neither is the evidence'); + assert.equal(records[0]?.child.candidates, 2); + assert.deepEqual(records[0]?.child.kinds, []); + assert.doesNotMatch( + String(records[0]?.statement), + /rules out an external termination/, + 'a reused pid must not be allowed to rule out the mechanism under investigation', + ); + } finally { + fs.rmSync(logDir, { recursive: true, force: true }); + } +}); From e650c41e047561c05b19ad0d2ab8665642677ad4 Mon Sep 17 00:00:00 2001 From: Hussein Mohamed Date: Sat, 5 Sep 2026 15:11:52 +0300 Subject: [PATCH 2/9] fix(api-diagnostics): make the capture channels correct off Windows and off this cwd CI failed the first push on all three test jobs for one reason: two new tests asserted exit code 134 for a V8 heap exhaustion. That is how Windows reports it. POSIX raises SIGABRT and leaves the exit code null, so both tests failed on Linux while passing on the machine they were written on. They now assert the fault rather than one platform's encoding of it, through a named helper that says why the two are the same event. Three review findings, all real: A relative `--out` meant two different directories. This process creates and reads the artifact paths, while the spawned runner resolves the same strings against the package `--package` selected, so `--out artifacts --package ../persistence` had the child write its report into the other package and this process look for it here -- a pass that cannot read its own run. The path is resolved once, up front. A diagnostic report was recorded as one run-level list. A run can hold several failures, and a flat list cannot say which child suffered the fault, which is the only thing the report was collected to say. Node names each report after the pid that wrote it, so the join needs nothing the capture did not already hold; each record now carries its own. A report body is a credential file: it contains the whole command line and every environment variable, and this suite runs with DATABASE_URL and its password in the environment. The names are the evidence and are kept; the bodies are deleted with the raw transcript, for the same reason and at the same moment. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01CUPqvu66J5ZVyiv4nJ797r --- .../test/diagnostics/run-signature-b-pass.mjs | 26 ++++- .../diagnostics/signature-b-correlate.test.ts | 105 +++++++++++++++++- 2 files changed, 122 insertions(+), 9 deletions(-) diff --git a/packages/api/test/diagnostics/run-signature-b-pass.mjs b/packages/api/test/diagnostics/run-signature-b-pass.mjs index 7be37939..d1b176f6 100644 --- a/packages/api/test/diagnostics/run-signature-b-pass.mjs +++ b/packages/api/test/diagnostics/run-signature-b-pass.mjs @@ -111,7 +111,13 @@ const maxRuns = positiveNumber('runs', flag('runs', 20), 20, { integer: true }); const maxMs = positiveNumber('max-minutes', flag('max-minutes', 45), 45) * 60_000; // The fallback is computed only when --out is absent: an argument is evaluated before the call, // so an eager mkdtemp would create an unused directory on every run that supplies its own. -const outDir = flag('out', null) ?? fs.mkdtempSync(path.join(tmpdir(), 'sigb-pass-')); +// +// Resolved against THIS process's working directory, because the paths derived from it are used by +// two processes with different ones: this process creates and reads them, while the spawned runner +// resolves the same strings against the package it was pointed at. Left relative, `--out artifacts` +// with `--package ../persistence` would have the child write its report into the other package and +// this process look for it here — a pass that cannot read its own run. +const outDir = path.resolve(flag('out', null) ?? fs.mkdtempSync(path.join(tmpdir(), 'sigb-pass-'))); // What each run executes. The default is the whole compiled suite, which is what reproducing // Signature B requires; `--target` exists so this script's own tests can point a run at one trivial @@ -314,8 +320,10 @@ for (let run = 1; run <= maxRuns; run++) { } } - // Names only. A report body carries the whole command line and every environment variable, so - // the file stays on disk for a reader who wants it and never reaches an artifact this prints. + // Names only. A report BODY carries the whole command line and every environment variable — this + // suite runs with DATABASE_URL and its password in the environment — so the body is never retained + // and never printed. The name is the whole of the evidence: Node writes one only for a fault it + // suffered itself, and names it after the pid that died. let reportFiles = []; try { reportFiles = fs.readdirSync(reportDir); @@ -323,6 +331,14 @@ for (let run = 1; run <= maxRuns; run++) { /* nothing wrote one, which is itself the observation */ } + // Attributed to the failure whose child wrote it. A run can hold several failures, and a single + // flat list cannot say which child suffered the fault — which is the only thing the report was + // collected to say. + for (const record of records) { + const pid = record.child?.pid ?? null; + record.reportFiles = pid === null ? [] : reportFiles.filter((name) => name.includes(`.${pid}.`)); + } + const summary = { run, durationMs, @@ -367,6 +383,10 @@ for (let run = 1; run <= maxRuns; run++) { // suite printed, so keeping it past the normalised record retains arbitrary test output for no // diagnostic gain; the child logs stay, being enumerated lifecycle events. fs.rmSync(tapPath, { force: true }); + // Same treatment, same reason: every report name is recorded above, and a report body is a file + // of credentials. Windows cannot be given the POSIX mode this directory was created with, so + // leaving one behind would rest on ACLs this script does not control. + fs.rmSync(reportDir, { recursive: true, force: true }); console.log(`\nCAPTURED on run ${run}:\n${records.map((r) => JSON.stringify(r, null, 2)).join('\n')}`); console.log(`\nRaw TAP discarded; the normalised record is ${path.join(outDir, 'capture.json')}.`); break; diff --git a/packages/api/test/diagnostics/signature-b-correlate.test.ts b/packages/api/test/diagnostics/signature-b-correlate.test.ts index 7e6b624d..8c229553 100644 --- a/packages/api/test/diagnostics/signature-b-correlate.test.ts +++ b/packages/api/test/diagnostics/signature-b-correlate.test.ts @@ -69,6 +69,18 @@ const correlator = require(CORRELATE_PATH) as { isWindowsPath(p: string): boolean; }; +/** + * Whether a child died of a fault it suffered, rather than by choosing its own exit status. + * + * The two platforms report the same event differently and neither encoding is the contract. Windows + * has no signals, so a V8 fatal error reaches the parent as exit code 134 with `signal: null`; POSIX + * raises SIGABRT, and the exit code is then null. Asserting 134 alone passes on the machine this was + * written on and fails everywhere CI runs it — which is exactly how it did fail. + */ +function diedOfAFault(exitCode: number | null | undefined, signal: string | null | undefined): boolean { + return exitCode === 134 || signal === 'SIGABRT'; +} + /** A TAP block in the exact shape Node's built-in reporter emits for a file-level failure. */ function tapFailure(file: string, exitCode: string, signal = '~', failureType = 'testCodeFailure'): string { return [ @@ -1376,7 +1388,10 @@ test('pass runner: a diagnostic report separates a real V8 fatal from a status t { encoding: 'utf8', timeout: 60_000 }, ); - assert.equal(crashed.status, 134, 'a heap exhaustion leaves 134, the same status three mechanisms leave'); + assert.ok( + diedOfAFault(crashed.status, crashed.signal), + `a heap exhaustion must abort, however this platform reports it (status ${crashed.status}, signal ${crashed.signal})`, + ); const reports = fs.readdirSync(reportDir); assert.equal(reports.length, 1, 'a V8 fatal writes exactly one report'); assert.ok( @@ -1395,7 +1410,8 @@ test('pass runner: a diagnostic report separates a real V8 fatal from a status t { encoding: 'utf8', timeout: 60_000 }, ); - assert.equal(exited.status, 134); + assert.equal(exited.status, 134, 'a chosen exit status is the same number on every platform'); + assert.equal(exited.signal, null, 'and it is an exit, not an abort'); assert.deepEqual(fs.readdirSync(quietDir), [], 'choosing the status of a fault is not a fault'); } finally { fs.rmSync(dir, { recursive: true, force: true }); @@ -1425,10 +1441,13 @@ test('pass runner: exit 134 alone names nothing, and the capture carries what se const fatal = JSON.parse(fs.readFileSync(oomCapturePath, 'utf8')) as { reportFiles: string[]; fatalMarkers: string[]; - records: Array<{ parent: { exitCode: number }; child: { kinds: string[] } }>; + records: Array<{ parent: { exitCode: number | null; signal: string | null }; child: { kinds: string[] } }>; }; - assert.equal(fatal.records[0]?.parent.exitCode, 134); + assert.ok( + diedOfAFault(fatal.records[0]?.parent.exitCode, fatal.records[0]?.parent.signal), + `the child must have aborted (exitCode ${fatal.records[0]?.parent.exitCode}, signal ${fatal.records[0]?.parent.signal})`, + ); assert.deepEqual(fatal.records[0]?.child.kinds, ['start', 'preload-installed'], 'a V8 fatal runs no JS exit path, which is why the status alone cannot name it'); assert.ok(fatal.fatalMarkers.includes('JavaScript heap out of memory'), @@ -1445,10 +1464,11 @@ test('pass runner: exit 134 alone names nothing, and the capture carries what se const quiet = JSON.parse(fs.readFileSync(chosenCapturePath, 'utf8')) as { reportFiles: string[]; fatalMarkers: string[]; - records: Array<{ parent: { exitCode: number } }>; + records: Array<{ parent: { exitCode: number | null; signal: string | null } }>; }; - assert.equal(quiet.records[0]?.parent.exitCode, 134, 'the same status as the fault above'); + assert.equal(quiet.records[0]?.parent.exitCode, 134, 'a chosen status, identical on every platform'); + assert.equal(quiet.records[0]?.parent.signal, null, 'and reached by exiting, not by aborting'); assert.deepEqual(quiet.fatalMarkers, [], 'nothing was printed'); assert.deepEqual(quiet.reportFiles, [], 'and no report was written, which is what tells them apart'); } finally { @@ -1525,3 +1545,76 @@ test('correlator: two runs that reused one pid are not merged into a single proc fs.rmSync(logDir, { recursive: true, force: true }); } }); + +test('pass runner: a relative --out belongs to the caller, not to the package being run', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'sigb-relout-')); + try { + // The parent creates and reads the artifact paths itself, while the child resolves the same + // strings against the package it was pointed at. Left relative, those are two different + // directories the moment `--package` names anything but this one: the parent would look for a + // report the child wrote somewhere else, and the pass would report that it could not read its + // own run. + const pkg = path.join(dir, 'other-package'); + fs.mkdirSync(pkg, { recursive: true }); + trivialTarget(pkg); + const result = runPass( + ['--runs', '1', '--max-minutes', '1', '--out', 'artifacts', '--package', pkg, '--target', 'noop.test.cjs'], + dir, + ); + + assert.equal(result.status, 0, `${result.stdout}${result.stderr}`); + const summaryPath = path.join(dir, 'artifacts', 'summary.json'); + assert.ok(fs.existsSync(summaryPath), 'the artifacts belong beside the caller, not inside the package'); + const summary = JSON.parse(fs.readFileSync(summaryPath, 'utf8')) as { + runs: Array<{ runnerExitCode: number | null; collectionError: string | null }>; + }; + assert.equal(summary.runs[0]?.collectionError, null, 'the run must be readable, not written out of reach'); + assert.equal(summary.runs[0]?.runnerExitCode, 0); + assert.ok(!fs.existsSync(path.join(pkg, 'artifacts')), 'nothing may be left inside the package under test'); + } finally { + fs.rmSync(dir, { recursive: true, force: true }); + } +}); + +test('pass runner: a diagnostic report is attributed to the child that wrote it, and its body is not kept', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'sigb-report-attr-')); + try { + // Bounded by construction: an 80 MB ceiling reached through NODE_OPTIONS, and a loop that stops + // at 400 MB of intent. + const oom = path.join(dir, 'oom.test.cjs'); + fs.writeFileSync(oom, 'const held = [];\nfor (let i = 0; i < 400; i++) held.push(new Array(131072).fill(i));\n'); + const out = path.join(dir, 'out'); + const result = spawnSync( + process.execPath, + [path.join(DIAGNOSTICS_DIR, 'run-signature-b-pass.mjs'), '--runs', '1', '--max-minutes', '2', '--out', out, '--target', oom], + { cwd: dir, encoding: 'utf8', env: { ...process.env, NODE_TEST_CONTEXT: undefined, NODE_OPTIONS: '--max-old-space-size=80' } }, + ); + + const capturePath = path.join(out, 'capture.json'); + assert.ok(fs.existsSync(capturePath), `${result.stdout}${result.stderr}`); + const capture = JSON.parse(fs.readFileSync(capturePath, 'utf8')) as { + records: Array<{ child: { pid: number | null }; reportFiles: string[] }>; + }; + + // A run can hold several failures. A single flat list cannot say which child suffered the fault, + // which is the only thing the report was collected to say — and Node names each report after the + // pid that wrote it, so the join needs nothing the capture does not already hold. + const record = capture.records[0]; + assert.ok(record, 'the heap exhaustion must be captured'); + assert.equal(record.reportFiles.length, 1, `expected one report for this child, got ${JSON.stringify(record.reportFiles)}`); + assert.ok( + record.reportFiles[0]?.includes(`.${record.child.pid}.`), + `the report must name the child that died (pid ${record.child.pid}, got ${record.reportFiles[0]})`, + ); + + // A report body carries the whole command line and every environment variable — DATABASE_URL and + // its password included. The name is the evidence; the body is a credential file, and it is + // treated exactly as the raw report transcript is. + const leftBehind = fs.existsSync(path.join(out, 'reports-run1')) + ? fs.readdirSync(path.join(out, 'reports-run1')) + : []; + assert.deepEqual(leftBehind, [], 'no report body may be retained in the artifacts'); + } finally { + fs.rmSync(dir, { recursive: true, force: true }); + } +}); From 0908c18145ea3e03fb0f7d010ccf90b7c8d1e2b5 Mon Sep 17 00:00:00 2001 From: Hussein Mohamed Date: Sat, 5 Sep 2026 15:19:20 +0300 Subject: [PATCH 3/9] test(api-diagnostics): prove a report goes to the child that faulted, not to the run Falsification found the previous test powerless: with one failure in the run, a report attributed per record and a report attributed per run are the same list, so reverting the attribution changed nothing any test could see. The mutant survived, which is the only honest reading of a test that asserts a distinction it cannot make. Two files now fail in one run with the same status where Windows reports one: the first suffers a real bounded heap exhaustion, the second merely chooses to exit 134. Only the first makes Node write a report, so a run-level list would hand that report to both records and say the second faulted too -- the opposite of what the report was collected to establish. The mutant is killed. Falsification now stands at 11 of 12. The survivor is equivalent under the TAP grammar Node can emit: ending a marker region at the point line rather than at the end of the previous test's report body is unobservable, because Node indents report bodies and the diagnostic-line filter already excludes them. The stricter boundary is kept as the correct one, not as one any test distinguishes. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01CUPqvu66J5ZVyiv4nJ797r --- .../diagnostics/signature-b-correlate.test.ts | 64 +++++++++++++++++++ 1 file changed, 64 insertions(+) diff --git a/packages/api/test/diagnostics/signature-b-correlate.test.ts b/packages/api/test/diagnostics/signature-b-correlate.test.ts index 8c229553..8f2b3a69 100644 --- a/packages/api/test/diagnostics/signature-b-correlate.test.ts +++ b/packages/api/test/diagnostics/signature-b-correlate.test.ts @@ -1618,3 +1618,67 @@ test('pass runner: a diagnostic report is attributed to the child that wrote it, fs.rmSync(dir, { recursive: true, force: true }); } }); + +test('pass runner: in a run with two failures, the report goes to the child that faulted', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'sigb-two-')); + try { + // The reason per-record attribution exists. Two files die in one run, with the SAME status where + // this platform reports one: the first suffers a real heap exhaustion, the second simply chooses + // to leave. Only the first makes Node write a report. A run-level list would hand that report to + // both records and say the second faulted too, which is the opposite of what it was collected to + // establish. + // + // Bounded by construction: an 80 MB ceiling reached through NODE_OPTIONS, and a loop that stops + // at 400 MB of intent. + fs.writeFileSync( + path.join(dir, 'a-fault.test.cjs'), + 'const held = [];\nfor (let i = 0; i < 400; i++) held.push(new Array(131072).fill(i));\n', + ); + fs.writeFileSync(path.join(dir, 'b-chosen.test.cjs'), 'process.exit(134);\n'); + + const out = path.join(dir, 'out'); + const result = spawnSync( + process.execPath, + [ + path.join(DIAGNOSTICS_DIR, 'run-signature-b-pass.mjs'), + '--runs', '1', '--max-minutes', '3', '--out', out, + '--target', path.join(dir, '*.test.cjs'), + ], + { cwd: dir, encoding: 'utf8', env: { ...process.env, NODE_TEST_CONTEXT: undefined, NODE_OPTIONS: '--max-old-space-size=80' } }, + ); + + const capturePath = path.join(out, 'capture.json'); + assert.ok(fs.existsSync(capturePath), `both files must be captured:\n${result.stdout}${result.stderr}`); + const capture = JSON.parse(fs.readFileSync(capturePath, 'utf8')) as { + reportFiles: string[]; + records: Array<{ + file: string; + parent: { exitCode: number | null; signal: string | null }; + child: { pid: number | null }; + reportFiles: string[]; + }>; + }; + + assert.equal(capture.records.length, 2, `expected both files to fail, got ${capture.records.map((r) => r.file).join(', ')}`); + const faulted = capture.records.find((record) => record.file.includes('a-fault')); + const chosen = capture.records.find((record) => record.file.includes('b-chosen')); + assert.ok(faulted && chosen, 'both files must appear in the capture'); + + assert.ok(diedOfAFault(faulted.parent.exitCode, faulted.parent.signal), 'the first file must have aborted'); + assert.equal(chosen.parent.exitCode, 134, 'the second file chose the status the first was given'); + + assert.equal(faulted.reportFiles.length, 1, `the fault wrote one report, got ${JSON.stringify(faulted.reportFiles)}`); + assert.ok( + faulted.reportFiles[0]?.includes(`.${faulted.child.pid}.`), + `and it names that child (pid ${faulted.child.pid}, got ${faulted.reportFiles[0]})`, + ); + assert.deepEqual( + chosen.reportFiles, + [], + 'the file that merely chose the status wrote nothing, and must not inherit the other file\'s report', + ); + assert.equal(capture.reportFiles.length, 1, 'the run as a whole still saw exactly one report'); + } finally { + fs.rmSync(dir, { recursive: true, force: true }); + } +}); From 21ce8402d43c014f412f7dba695a11015727e811 Mon Sep 17 00:00:00 2001 From: Hussein Mohamed Date: Sat, 5 Sep 2026 15:49:38 +0300 Subject: [PATCH 4/9] fix(api-diagnostics): a run's artifacts start empty, and no report body outlives them Two further review findings, both real, and both the same class of mistake this PR set out to fix -- stale evidence, and evidence kept longer than it should be. A run's artifact directories were created but not emptied. `mkdirSync` is happy with a directory that already holds an interrupted pass's files, so a reused `--out` lent the new run the old run's diagnostic report -- and because a report is attributed by the pid in its name, a reused pid would let a file that merely chose its exit status inherit a native fault it never suffered. That is exactly the staleness the correlator refuses on the child-log side, reproduced in the artifact directory. Both directories are now emptied before a run. Report bodies were deleted only on the capture and clean-run paths. A run whose report could not be read, or a pass interrupted by SIGINT, left them on disk -- and a report body carries the whole command line and every environment variable, this suite's DATABASE_URL password included. The bodies are now deleted the moment their names are read, before any branch, so a captured run, a clean run and an unreadable run all leave the same nothing behind; the signal handler takes the in-flight directory with the process tree it already kills. Falsification: 12 of 13 killed, both new fixes among them. The one survivor remains the equivalent mutant on the marker-region boundary. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01CUPqvu66J5ZVyiv4nJ797r --- .../test/diagnostics/run-signature-b-pass.mjs | 29 +++++++++-- .../diagnostics/signature-b-correlate.test.ts | 51 +++++++++++++++++++ 2 files changed, 75 insertions(+), 5 deletions(-) diff --git a/packages/api/test/diagnostics/run-signature-b-pass.mjs b/packages/api/test/diagnostics/run-signature-b-pass.mjs index d1b176f6..365b7c2c 100644 --- a/packages/api/test/diagnostics/run-signature-b-pass.mjs +++ b/packages/api/test/diagnostics/run-signature-b-pass.mjs @@ -217,9 +217,21 @@ function killTree(child) { */ let activeChild = null; +/** + * The report directory of the run in flight, for the same reason `activeChild` exists. + * + * An interrupt reaches this process between a child writing a report and this script reading and + * deleting it, and a report body carries the whole command line and every environment variable. + * Killing the tree without this would leave that file on disk. + * + * @type {string | null} + */ +let activeReportDir = null; + for (const signal of ['SIGINT', 'SIGTERM']) { process.on(signal, () => { if (activeChild !== null) killTree(activeChild); + if (activeReportDir !== null) fs.rmSync(activeReportDir, { recursive: true, force: true }); // 128 + signal number, the conventional shell encoding for "terminated by this signal". process.exit(signal === 'SIGINT' ? 130 : 143); }); @@ -293,8 +305,16 @@ for (let run = 1; run <= maxRuns; run++) { const tapPath = path.join(outDir, `run${run}.tap`); const logDir = path.join(outDir, `child-logs-run${run}`); const reportDir = path.join(outDir, `reports-run${run}`); + // Emptied, not merely created. `mkdirSync` is happy with a directory that already holds an + // interrupted pass's artifacts, and this is the same staleness the correlator refuses on the child + // -log side: a leftover report would be counted as this run's, and since a report is attributed by + // the pid in its name, a reused pid would let a file that merely chose its exit status inherit a + // native fault it never suffered. + fs.rmSync(logDir, { recursive: true, force: true }); + fs.rmSync(reportDir, { recursive: true, force: true }); fs.mkdirSync(logDir, { recursive: true, mode: 0o700 }); fs.mkdirSync(reportDir, { recursive: true, mode: 0o700 }); + activeReportDir = reportDir; const freeBefore = freemem(); const startedRun = Date.now(); @@ -330,6 +350,10 @@ for (let run = 1; run <= maxRuns; run++) { } catch { /* nothing wrote one, which is itself the observation */ } + // Deleted here rather than on the way out of a branch, so that a run which is captured, a run + // which is clean and a run whose report could not be read all leave the same nothing behind. + fs.rmSync(reportDir, { recursive: true, force: true }); + activeReportDir = null; // Attributed to the failure whose child wrote it. A run can hold several failures, and a single // flat list cannot say which child suffered the fault — which is the only thing the report was @@ -383,10 +407,6 @@ for (let run = 1; run <= maxRuns; run++) { // suite printed, so keeping it past the normalised record retains arbitrary test output for no // diagnostic gain; the child logs stay, being enumerated lifecycle events. fs.rmSync(tapPath, { force: true }); - // Same treatment, same reason: every report name is recorded above, and a report body is a file - // of credentials. Windows cannot be given the POSIX mode this directory was created with, so - // leaving one behind would rest on ACLs this script does not control. - fs.rmSync(reportDir, { recursive: true, force: true }); console.log(`\nCAPTURED on run ${run}:\n${records.map((r) => JSON.stringify(r, null, 2)).join('\n')}`); console.log(`\nRaw TAP discarded; the normalised record is ${path.join(outDir, 'capture.json')}.`); break; @@ -394,7 +414,6 @@ for (let run = 1; run <= maxRuns; run++) { // A clean run's artifacts answer nothing and would otherwise grow without bound across the pass. fs.rmSync(logDir, { recursive: true, force: true }); - fs.rmSync(reportDir, { recursive: true, force: true }); fs.rmSync(tapPath, { force: true }); } diff --git a/packages/api/test/diagnostics/signature-b-correlate.test.ts b/packages/api/test/diagnostics/signature-b-correlate.test.ts index 8f2b3a69..4019dd9a 100644 --- a/packages/api/test/diagnostics/signature-b-correlate.test.ts +++ b/packages/api/test/diagnostics/signature-b-correlate.test.ts @@ -1682,3 +1682,54 @@ test('pass runner: in a run with two failures, the report goes to the child that fs.rmSync(dir, { recursive: true, force: true }); } }); + +test('pass runner: a reused --out cannot lend one pass a previous pass\'s diagnostic report', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'sigb-stale-report-')); + try { + // The same staleness this correlator refuses on the child-log side, on the artifact side. An + // interrupted pass leaves `reports-run1` behind; a later pass given the same `--out` created the + // directory without emptying it, so the leftover counted as this run's — and since attribution + // accepts any filename carrying the pid, a reused pid would let a file that merely chose its + // exit status inherit a native fault it never suffered. + const out = path.join(dir, 'out'); + fs.mkdirSync(path.join(out, 'reports-run1'), { recursive: true }); + fs.writeFileSync(path.join(out, 'reports-run1', 'report.19700101.000000.4242.0.001.json'), '{"stale":true}\n'); + + const result = runPass(['--runs', '1', '--max-minutes', '1', '--out', out, '--target', trivialTarget(dir)], dir); + + assert.equal(result.status, 0, `${result.stdout}${result.stderr}`); + const summary = JSON.parse(fs.readFileSync(path.join(out, 'summary.json'), 'utf8')) as { + runs: Array<{ reportFiles: number }>; + }; + assert.equal(summary.runs[0]?.reportFiles, 0, 'a leftover from an earlier pass is not this run\'s evidence'); + } finally { + fs.rmSync(dir, { recursive: true, force: true }); + } +}); + +test('pass runner: no diagnostic report body survives a run, whatever the run did', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'sigb-no-bodies-')); + try { + // A report body carries the whole command line and every environment variable, this suite's + // DATABASE_URL password included. The names are recorded; the bodies must not outlive the run + // that produced them, on the capture path or any other. Run both kinds and check the artifacts. + const clean = path.join(dir, 'clean'); + const cleanResult = runPass(['--runs', '1', '--max-minutes', '1', '--out', clean, '--target', trivialTarget(dir)], dir); + assert.equal(cleanResult.status, 0, `${cleanResult.stdout}${cleanResult.stderr}`); + + const failing = path.join(dir, 'fail-target.test.cjs'); + fs.writeFileSync(failing, 'process.exit(134);\n'); + const captured = path.join(dir, 'captured'); + runPass(['--runs', '1', '--max-minutes', '1', '--out', captured, '--target', failing], dir); + assert.ok(fs.existsSync(path.join(captured, 'capture.json')), 'the failing target must be captured'); + + for (const artifacts of [clean, captured]) { + const bodies = fs + .readdirSync(artifacts, { recursive: true, encoding: 'utf8' }) + .filter((entry) => entry.includes('report.') && entry.endsWith('.json')); + assert.deepEqual(bodies, [], `a report body survived in ${artifacts}: ${JSON.stringify(bodies)}`); + } + } finally { + fs.rmSync(dir, { recursive: true, force: true }); + } +}); From 1a96adc0b190bdde55351becacbfe4750934a10c Mon Sep 17 00:00:00 2001 From: Hussein Mohamed Date: Sat, 5 Sep 2026 16:08:49 +0300 Subject: [PATCH 5/9] fix(api-diagnostics): refuse an --out that already holds a capture Emptying each run's child logs stopped an earlier run lending a later one its evidence, and introduced the mirror problem across passes: a second pass given the same `--out` deletes the child logs the first pass captured while leaving its `capture.json` in place, so the directory ends up asserting a termination whose evidence no longer exists. Deleting the old capture would be the worse answer -- it is the rarest artifact this diagnostic produces, and a pass that silently destroys one is a pass nobody should point at a real occurrence. The runner refuses instead, before it creates or spawns anything, and says what is in the way. Falsification: 13 of 14 killed, this one among them. The single survivor is still the equivalent mutant on the marker-region boundary. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01CUPqvu66J5ZVyiv4nJ797r --- .../test/diagnostics/run-signature-b-pass.mjs | 10 +++++++ .../diagnostics/signature-b-correlate.test.ts | 29 +++++++++++++++++++ 2 files changed, 39 insertions(+) diff --git a/packages/api/test/diagnostics/run-signature-b-pass.mjs b/packages/api/test/diagnostics/run-signature-b-pass.mjs index 365b7c2c..23e79d01 100644 --- a/packages/api/test/diagnostics/run-signature-b-pass.mjs +++ b/packages/api/test/diagnostics/run-signature-b-pass.mjs @@ -161,6 +161,16 @@ function ensurePrivateDirectory(dir) { } ensurePrivateDirectory(outDir); + +// A pass owns its artifact directory: it empties each run's child logs so that no earlier run can +// lend it evidence. That is right for the logs and fatal for a capture — reusing an `--out` would +// delete the child logs an earlier pass captured while leaving its `capture.json` behind, and the +// directory would then assert a termination whose evidence no longer exists. Deleting the old +// capture instead would be worse: it is the rarest artifact this produces. So refuse, here, before +// anything is created or spawned. +if (fs.existsSync(path.join(outDir, 'capture.json'))) { + usageError(`--out ${outDir} already holds a capture from an earlier pass; move it aside or choose another directory`); +} console.log(`Signature B bounded pass — ceiling ${maxRuns} runs or ${maxMs / 60_000} minutes, stopping at first capture.`); console.log(`Artifacts: ${outDir}\n`); diff --git a/packages/api/test/diagnostics/signature-b-correlate.test.ts b/packages/api/test/diagnostics/signature-b-correlate.test.ts index 4019dd9a..8f320aac 100644 --- a/packages/api/test/diagnostics/signature-b-correlate.test.ts +++ b/packages/api/test/diagnostics/signature-b-correlate.test.ts @@ -1733,3 +1733,32 @@ test('pass runner: no diagnostic report body survives a run, whatever the run di fs.rmSync(dir, { recursive: true, force: true }); } }); + +test('pass runner: an --out that already holds a capture is refused, not quietly reused', () => { + const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'sigb-prior-capture-')); + try { + // A pass owns its artifact directory: it empties each run's child logs so no earlier run can lend + // it evidence. That is right for the logs and fatal for a capture — reusing an `--out` would + // delete the child logs a previous pass captured while leaving its `capture.json` behind, so the + // directory would end up asserting a termination whose evidence no longer exists. + // + // Deleting the old capture instead would be worse: a capture is the rarest artifact this whole + // diagnostic produces. So the pass refuses, before it creates or spawns anything. + const out = path.join(dir, 'out'); + fs.mkdirSync(out, { recursive: true }); + fs.writeFileSync(path.join(out, 'capture.json'), '{"run":1,"records":[]}\n'); + + const result = runPass(['--runs', '1', '--max-minutes', '1', '--out', out, '--target', trivialTarget(dir)], dir); + + assert.equal(result.status, 2, `${result.stdout}${result.stderr}`); + assert.match(String(result.stderr), /capture/i, 'the refusal must say what is in the way'); + assert.ok(!fs.existsSync(path.join(out, 'summary.json')), 'a refused pass must not have run anything'); + assert.deepEqual( + JSON.parse(fs.readFileSync(path.join(out, 'capture.json'), 'utf8')), + { run: 1, records: [] }, + 'and must leave the capture it refused to overwrite exactly as it found it', + ); + } finally { + fs.rmSync(dir, { recursive: true, force: true }); + } +}); From 4a05be2c5841e49353e104d7921d34f22e80270c Mon Sep 17 00:00:00 2001 From: Hussein Mohamed Date: Sat, 5 Sep 2026 16:18:54 +0300 Subject: [PATCH 6/9] test(api-diagnostics): prove a refused pass leaves the capture's own logs intact The guard against reusing an --out that already holds a capture existed, but its test only checked that the pass refused and that capture.json survived. The thing actually at risk is the evidence the capture refers to: a pass empties each run's child logs before running it, so a replacement pass would delete them and leave the capture pointing at nothing. The test now seeds that evidence and asserts it is still there after the refusal, which is what makes the orphaning scenario unreachable rather than merely unlikely. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01CUPqvu66J5ZVyiv4nJ797r --- .../test/diagnostics/signature-b-correlate.test.ts | 11 ++++++++++- 1 file changed, 10 insertions(+), 1 deletion(-) diff --git a/packages/api/test/diagnostics/signature-b-correlate.test.ts b/packages/api/test/diagnostics/signature-b-correlate.test.ts index 8f320aac..0e3568a6 100644 --- a/packages/api/test/diagnostics/signature-b-correlate.test.ts +++ b/packages/api/test/diagnostics/signature-b-correlate.test.ts @@ -1745,8 +1745,12 @@ test('pass runner: an --out that already holds a capture is refused, not quietly // Deleting the old capture instead would be worse: a capture is the rarest artifact this whole // diagnostic produces. So the pass refuses, before it creates or spawns anything. const out = path.join(dir, 'out'); - fs.mkdirSync(out, { recursive: true }); + fs.mkdirSync(path.join(out, 'child-logs-run1'), { recursive: true }); fs.writeFileSync(path.join(out, 'capture.json'), '{"run":1,"records":[]}\n'); + // The evidence the capture refers to. It is what a reused `--out` would destroy: the pass + // empties each run's child logs before running it, so a replacement pass would delete these and + // leave the capture above pointing at nothing. + fs.writeFileSync(path.join(out, 'child-logs-run1', 'run-4242-1000.jsonl'), '{"kind":"start","pid":4242}\n'); const result = runPass(['--runs', '1', '--max-minutes', '1', '--out', out, '--target', trivialTarget(dir)], dir); @@ -1758,6 +1762,11 @@ test('pass runner: an --out that already holds a capture is refused, not quietly { run: 1, records: [] }, 'and must leave the capture it refused to overwrite exactly as it found it', ); + assert.deepEqual( + fs.readdirSync(path.join(out, 'child-logs-run1')), + ['run-4242-1000.jsonl'], + 'the capture must still have the child logs it refers to; refusing early is what protects them', + ); } finally { fs.rmSync(dir, { recursive: true, force: true }); } From 9550438072fb411929f528660882b2e497a6558e Mon Sep 17 00:00:00 2001 From: Hussein Mohamed Date: Sat, 5 Sep 2026 19:34:53 +0300 Subject: [PATCH 7/9] docs: record M15 Increment 51 Signature B mechanism isolation Records the diagnostic work already on this branch in the canonical documents, now that PR #42 has merged and Increment 50 exists. Status is unchanged by this commit: Signature B remains UNRESOLVED, at root-cause acceptance Level C. 39 bounded runs produced 5 real captures across packages/api and packages/persistence, every one reporting exit status 3221226505 (0xC0000409, STATUS_STACK_BUFFER_OVERRUN) with signal null and a child lifecycle log holding only start and preload-installed. On the four captures taken with the hardened harness the fatal-marker and diagnostic-report channels were both empty, which is what excludes the measured V8/Node fatal path. What remains is a family of two - an in-process Windows fail-fast path, or an external TerminateProcess choosing that status - and the evidence does not choose between them. Avast is present and aswhook.dll was observed loaded inside a live node.exe; that is a leading candidate for a controlled A/B test, not a cause. No antivirus was disabled and no security posture was changed. Signature B occurred twice during this increment's own final validation, both on the plain spec-reporter path with no exit-status evidence. Both are recorded rather than dismissed for passing on rerun. Increment 50 is preserved unchanged. Documentation only. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01CUPqvu66J5ZVyiv4nJ797r --- docs/PROJECT_STATE.md | 191 +++++++++++++++++- docs/ROADMAP.md | 3 +- ...0140-harness-ephemeral-port-acquisition.md | 59 ++++++ 3 files changed, 251 insertions(+), 2 deletions(-) diff --git a/docs/PROJECT_STATE.md b/docs/PROJECT_STATE.md index 170017da..ab6bb48b 100644 --- a/docs/PROJECT_STATE.md +++ b/docs/PROJECT_STATE.md @@ -4,7 +4,196 @@ > to read **only this file** and continue immediately. Updated after every > milestone and every significant architectural step. -_Last updated: 2026-09-05 — M15 Increment 50: test:counts / standalone gateway host setup contract._ +_Last updated: 2026-09-05 — M15 Increment 51: Signature B mechanism isolation and diagnostic hardening._ + +Prior: _Last updated: 2026-09-05 — M15 Increment 50: test:counts / standalone gateway host setup contract._ + +## M15 Increment 51 — Signature B mechanism isolation and diagnostic hardening + +**Status: UNRESOLVED — exact terminating source unproven.** Root-cause acceptance: **LEVEL C — +narrowed to one mechanism family, exact source not established.** No production code, workflow, +migration or repository semantic changed; the increment is diagnostic and test infrastructure only. +Increment 50's setup/documentation contract remains resolved and is not altered here. + +This increment did not set out to fix Signature B. It set out to identify the process-termination +mechanism, and it is reported at the level the evidence actually reaches: a mechanism **family** is +now established by positive measurement, and the party that terminates the child is not. + +**Bounded real campaign: 39 runs, 5 real Signature B captures**, across both `packages/api` and +`packages/persistence` (six arms: API pre-fix 5 runs / 1 capture; API final 8 / 1; persistence 2 / 1 +and 1 / 1; Node 22 persistence 17 / 0; mixed-ordered persistence-then-api 6 / 1). Free memory +ranged 718–1533 MB of 16077 MB, with other agents' processes concurrently running; that +concurrent-load confounder is disclosed, not leaned on. + +**All five captures reported the identical exit status.** + +| # | Package | File | `exitCode` | `signal` | Child lifecycle | Fatal markers | Reports | +|---|---|---|---|---|---|---|---| +| 1 | api | `pg-security-ownership.integration` (641.1 ms) | `3221226505` | `null` | `start`, `preload-installed` | *not captured* | *n/a* | +| 2 | persistence | `studies.integration` (529.5 ms) | `3221226505` | `null` | `start`, `preload-installed` | `[]` | `[]` | +| 3 | persistence | `learning.integration` (523.4 ms) | `3221226505` | `null` | `start`, `preload-installed` | `[]` | `[]` | +| 4 | api | `pg-security.integration` (629.2 ms) | `3221226505` | `null` | `start`, `preload-installed` | `[]` | `[]` | +| 5 | api | `auth-signin-schema.integration` (734.5 ms) | `3221226505` | `null` | `start`, `preload-installed` | `[]` | `[]` | + +`3221226505` is `0xC0000409`, the Windows status `STATUS_STACK_BUFFER_OVERRUN` — the fail-fast-class +status. Every capture ran no normal JavaScript shutdown hook: the child log holds exactly `start` +and `preload-installed` and nothing else, including Node's own unconditional `exit`. + +**Capture 1 is deliberately not claimed as decisive.** It was taken by the pre-fix harness, whose +raw TAP and fatal-marker evidence for that run was discarded before it could be read. Its exit +status and lifecycle log stand; its marker and report channels do not exist. Captures 2–5 carry the +full evidence set and are what the exclusions below rest on. + +### What the evidence excludes + +Excluded by positive measurement on captures 2–5, not by absence of a hypothesis: + +- ordinary `process.exit` / `process.exitCode` — would have written a log line and fired `exit`; +- an ordinary uncaught JS exception — same; +- a fatal unhandled rejection — same; +- an instrumented JS `process.abort()` — the preload wraps it and records before delegating; +- **the measured V8/Node fatal path (including heap OOM)** — the synthetic fatal path produces + fatal stderr diagnostics inside the TAP report **and** a PID-attributable Node diagnostic report + under `--report-on-fatalerror`. Captures 2–5 produced **neither**: `fatalMarkers` is `[]` and + `reportFiles` is `[]` for each; +- **node:test parent cancellation in this repository's configuration** — the runner's own source + was read (`node --expose-internals`, `internal/test_runner/runner.js`): `FileTest` sets + `this.timeout = null`, so the parent enforces no file-level wall clock, and a child aborted + through `spawn(..., { signal })` sets `err` via `child.on('error')` first, so the runner throws an + `AbortError` rather than the bare fallback shape. There is no parent path in this configuration + that produces what was observed. + +### What remains possible + +Two sources remain, and the evidence does not choose between them: + +- **A.** an in-process Windows fail-fast path — a security mitigation, `RaiseFailFastException`, or + an equivalent native fail-fast source; +- **B.** an external party calling `TerminateProcess` with `0xC0000409` as the chosen exit status. + +An exit status is a 32-bit integer the terminating party picks, so the status alone cannot separate +A from B. Choosing between them requires a channel this increment deliberately did not open: ETW +`Microsoft-Windows-Kernel-Process` tracing, or WER `LocalDumps`. Both change the machine's +configuration, and neither was enabled. + +**Windows-side corroboration, recorded as measurement not conclusion.** No Windows Error Reporting +record exists for `node.exe` in the capture windows, though WER is enabled and logged other events +that day. Only two `.node` addons exist in the tree (rollup's, unused by tests). Avast Antivirus is +present, and `aswhook.dll` was observed loaded inside a live `node.exe` process. **Injection is not +causation.** This makes the AV/injection path a concrete leading candidate worth a controlled future +A/B test; it is not a root cause. No security exclusion was added, no antivirus was disabled, and no +security posture was changed. A controlled AV A/B experiment requires explicit owner authorization +and is a separate future step. + +**Node version evidence.** Node 22 and Node 24 both reproduce the fingerprint distinctions the +correlator depends on, and the diagnostics suite is identical on both. The Node 22 bounded arm was +**17 runs, 0 captures**. That does **not** prove a Node-version difference and no inference that +Node 24 causes the defect is drawn from it; 17 runs at the historically observed rate bounds little. + +### Synthetic fingerprint matrix + +Seventeen termination mechanisms were measured under the runner to build fingerprints **before** any +real capture was compared against them, with the match criteria stated in advance and multi-field +(exit status, signal, failure type, child lifecycle events, fatal stderr markers, diagnostic report +presence and PID attribution). Exactly three reproduce the historical child-side lifecycle shape — +`taskkill /F` (exit `1`), `Stop-Process -Force` (`4294967295`) and a V8 heap OOM (`134` on Windows, +`SIGABRT` on POSIX) — and each is separated from the real captures by at least one other field. Any +memory-pressure experiment ran in a child process under a strict heap limit and a bounded wall clock +with process-tree cleanup; the host's real memory was never deliberately exhausted. + +### Diagnostic improvements delivered + +Two blind spots were proven, not guessed, and only those were closed: + +- **fatal markers now come from the per-failure TAP diagnostic region** instead of the parent + runner's stderr. The runner re-emits each child stderr line as a `test:stderr` reporter event, so + a child's `FATAL ERROR:` banner lands in the TAP report and **never** on the parent's stderr — + measured at 0 bytes while the report held the banner. The previous channel could only ever read + empty; +- **lifecycle logs are isolated per process**, keyed by log file and PID, instead of merged by test + file path. The merged form produced the statement "rules out an external termination and a native + fault" — a false negative against the live hypothesis. + +Also delivered, each with tests: stale or PID-reused evidence becomes an explicit ambiguity rather +than a silent merge; diagnostic report existence and PID attribution separate fatal paths, attributed +per failure rather than per run; sensitive report bodies are never retained on any exit path, +including error and `SIGINT`; artifact paths are independent of the selected package's working +directory; a run's artifact directories start empty and an `--out` already holding a capture is +refused rather than reused; cross-package Signature B runs support `packages/api` and +`packages/persistence` via `--package`; POSIX OOM signal/status behaviour is covered alongside the +Windows encoding; and false-attribution regressions are covered directly. + +The correlator suite grew **41 → 57 tests (55 pass, 2 pre-existing POSIX-only skips, 0 fail)**, +identical on Node 22 and Node 24. Falsification killed **13 of 14 mutations**; the sole survivor is +a region-boundary mutant that is equivalent under the TAP grammar Node can emit, and it is reported +rather than hidden. All mutated sources were restored byte-identically, verified by SHA-256. + +### Validation + +Measured on 2026-09-05 on Windows 11 with Node 24.15.0 and npm 11.12.1, after merging +`origin/main` `48209e8074e00632b1f3ed29d68d3fe099eed6cf` (PR #42) into this branch with no rebase +and no conflicts. Host preparation followed Increment 50's contract exactly — `npm ci`, +`npm run build`, `npm ci --prefix services/gateway`, each exit **0**. + +- `npm run lint` **0**; `git diff --check` clean. +- `npm test` **exit 0**: **19 workspaces, 3320 tests, 3292 pass, 0 fail, 28 skipped**, with **0** + suites self-skipping for a missing `DATABASE_URL`. +- `npm run test:counts` **exit 0**: **3336 tests, 33 skipped** across the 19 root workspaces plus + the standalone gateway (`api` 1012/10, `web` 1030/0, `persistence` 186/0, `gateway-service` 16/5). +- `npm run test:scripts`, `check:ci-parity`, `check:adr-claims`, `check:variant-parity`, + `check:engine-pin-parity`, `check:observability`, `check:build-order`, `check:deploy-gates` and + `npm run test:load-harness` — each **0**. +- `services/gateway`: `build` **0**, `lint` **0**, `test` **0** (16 tests, 11 passed, 5 skipped). +- The diagnostics correlator suite: **57 tests, 55 pass, 0 fail, 2 skipped**, identical on Node + 24.15.0 and on Node 22.23.2. + +`DATABASE_URL` pointed at a dedicated PostgreSQL 16 container with `pgvector`, created for this +validation and removed afterwards; **`REDIS_URL` was unset, so Redis-backed testing was NOT RUN**. +Environment-gated skips are not passes. + +### Signature B during this increment's own final validation + +**It occurred twice, and neither occurrence is dismissed because a rerun passed.** + +| Package | File | Wall time | Node | Exit status | Child lifecycle | Fatal markers | Reports | +|---|---|---|---|---|---|---|---| +| api | `auth-signin-schema.integration.test.js` | 618.7 ms | 24.15.0 | *not captured* | *not captured* | *not captured* | *not captured* | +| api | `cookie-auth.test.js` | 707.0 ms | 24.15.0 | *not captured* | *not captured* | *not captured* | *not captured* | + +Both showed the documented shape exactly — a bare file-level `'test failed'`, no assertion, no +stack, and none of the file's own tests reported — and both occurred under the plain `npm test` / +`npm run test:counts` path, which runs the `spec` reporter with no TAP destination and no preload. +That is precisely the blind spot this increment's instrumentation exists to close, and it means +**neither occurrence has exit-status, lifecycle, marker or report evidence**: nothing here narrows +Level C, and nothing here is claimed to. A bounded instrumented pass of 3 further runs over the same +package, taken immediately after the first, produced **0 captures** — which bounds nothing at that +size and is reported only so the attempt is on the record. + +Two conditions are worth recording without being leaned on. The host was under heavy memory +pressure from unrelated concurrent work: **699 MB free of 16077 MB** at the time of the first +occurrence, against 1384–1428 MB during the instrumented pass that saw nothing, and roughly 2.5 GB +at the earliest historical captures. And an *earlier* attempt at the same full-repository run died +differently — `npm` reported exit status **3221225794** (`0xC0000142`, +`STATUS_DLL_INIT_FAILED`) for the whole `packages/api` command, a process that failed to start +rather than a file that failed to run. That is **not** Signature B and is not counted as one: the +shape, the status and the layer all differ. It is recorded because it is the same host condition, +and because omitting it would make the memory-pressure correlation look cleaner than it is. + +### Separate known defect observed during this increment — not fixed here + +**The durable analysis-cache race test `two live instances racing a cold position both compute it` +failed once and passed on rerun.** It lives in the Increment 49 suite, which this increment does not +touch, and it passed at the two prior HEADs. This is **not** Signature B: it is an ordinary assertion +failure (expected 2, actual 1) with a normal stack, not a bare file-level termination. The rerun +passing does **not** resolve it. Recorded here as a separate bounded follow-up. + +### Still open after Increment 51 + +- **Signature B — UNRESOLVED**, now at Level C: the mechanism family is established, the terminating + source is not. The controlled AV A/B test and OS-level tracing (ETW / WER `LocalDumps`) are the + named next steps, and both need owner authorization because they change machine configuration. +- **The analysis-cache race flake above — OPEN**, separate from Signature B. + ## M15 Increment 50 — test:counts / standalone gateway host setup contract diff --git a/docs/ROADMAP.md b/docs/ROADMAP.md index f8c030d1..3164c09b 100644 --- a/docs/ROADMAP.md +++ b/docs/ROADMAP.md @@ -1349,12 +1349,13 @@ Debt observed during M14. Each states what is known, not what is planned; items - **The published capability document could omit a composed feature, and the client offered controls that could not work (RESOLVED in M15 Increment 24 / ADR-0132).** Left open by Increment 23 above. `capabilitiesView` (`packages/api/src/presenters.ts`) took a hand-written `Pick`, so a feature never added to it was invisible to `GET /v1/capabilities` and nothing complained. ADR-0131 judged that "a narrower failure than a 503 — the feature works for anyone who calls the route directly", and deferred it on the grounds that closing it meant first deciding which optional dependencies are user-facing capabilities. **Both halves of that were wrong.** The failure is narrower only when the client can still reach the feature, and a live instance already existed where it could not: `GET /v1/search` served three modes from two independently-gated dependency sets behind one published `search` flag, so a deployment running the Helm chart’s `search.semanticEnabled: false` advertised search while the client offered two mode buttons whose every request answered 503. And the judgement, while genuinely underivable, can be made **unskippable**, which was the property wanted. `semanticSearch` is now published from both of its dependencies, `Exclude[0] | NotAPublishedCapability>` must be `never` so a new optional dependency cannot compile until someone gives it a flag or records why it has none, and a behavioural guard requires every capability source to change what the document publishes — because a key can sit in that parameter unread, which the compile-time half cannot see. The mutation ledger lives in ADR-0132 and is not restated here. - **The search surface was not gated on `capabilities.search`, so an absolute kill switch still showed a search box (RESOLVED in M15 Increment 24 / ADR-0132 §5).** `SEARCH_ENABLED=0` — the chart's `search.enabled: false`, an absolute kill switch per ADR-0055 — leaves `searchRepository` unconstructed and `GET /v1/search` answering 503 on every mode, keyword included. The entry point was the persistent header form in `packages/web/index.html`, present on every page; being a `
` rather than an `a[data-route]`, `NAV_CAPABILITY_MAP` could not reach it, which is why the first pass at Increment 24 gated the semantic and hybrid *modes* and left keyword ungated — the same defect class one mode over. Raised by the Qodo review of PR #155. **Fixed in the same increment rather than deferred:** the form now ships `hidden` and is revealed by `applySearchCapability` only on an explicit `search: true`; the route renders an honest unavailable notice and issues no request; and keyword search waits for the capability answer, reversing this increment's own earlier latency decision, because knowing whether a request is pointless requires having asked. A markup-contract test pins the `hidden` attribute, since the gate depends on it and every other test passes without it. - **Clicking a search mode discarded text typed since the page loaded (RESOLVED in a follow-up to M15 Increment 24 / ADR-0132).** `createModeInput` in `packages/web/src/app/search-mount.ts` closed over the query captured when the route mounted, so `navigateToSearchMode` navigated with the old term and the remount reset the input to match it. Type a new term into the header field, click **Semantic** without pressing enter, and the typed text was gone with no indication it had been discarded. Pre-existing — the closure predated Increment 24 and was untouched by it — and found by the adversarial review of PR #155 while reviewing the capability gate wrapped around the same control. **Resolved:** the query is now a `() => string` read when a mode is chosen rather than a string captured when the selector renders, matching what `main.ts`'s submit handler already does, and falling back to the mounted query only where the document has no header input. Two regression tests, one of which fails against the exact pre-fix closure. -- **`startHarness` drew ports `fetch` refuses (RESOLVED in ADR-0140); the unexplained whole-file failure is still open.** Two signatures were filed here, deliberately not as one cause, and that judgement held. **Signature A is resolved.** WHATWG Fetch blocks eighty-two ports and undici enforces the list on the port number alone, before opening a socket — so `server.listen(0)` could bind, listen and answer raw TCP while `fetch` still refused, surfacing as `TypeError: fetch failed` / `Error: bad port` at the harness’s first request rather than at the listen that caused it. Whether it can happen at all is a property of the host’s dynamic port range: a typical Linux CI range (32768–60999) contains no blocked port, while the Windows range in use here (1024–15000) contains nineteen, which is the whole of "green on CI, flaky locally". The guard that already existed was incomplete in a way that still failed — its hand-observed set of eighteen ports was the spec list intersected with one machine’s range **minus `6679`** — and it retried unboundedly, had no behaviour on exhaustion, and had been copy-pasted into `auth-signin-schema.integration.test.ts`, so the missing port had to be found twice. `packages/api/test/listen.ts` now owns port acquisition for both sites: the spec-complete eighty-two ports (verified by sweeping all 65535 through the real `fetch` on Node v24.15.0), a bounded twenty attempts, each rejected listener closed before the next is asked for, a guarded `address()` read in place of the `as AddressInfo` cast, and an exhaustion error naming the attempts and rejected ports and nothing else. **Signature B is not resolved and was not folded in.** A file still fails with `'test failed'`, no assertion, no stack and none of its own tests reported. Twenty consecutive full runs gave five failures: on pre-fix code one signature A (`auth.test.js`) and three signature B; on post-fix code one signature B and no signature A. The port fix removes A and leaves B exactly where it was, which is the evidence that they are two defects. Four different files were hit (`move-explanation-route`, `tournament-commentary-route`, `bot-detection-analyze`, `anti-cheat-analysis`), sharing no import beyond `./helpers`; each died in 589–703 ms with no test of its own reporting and no stderr. Refuted with evidence: ephemeral-port exhaustion (113 sockets in TIME_WAIT against a 13977-port range), a `Promise.race` loser becoming an unhandled rejection (`race` subscribes to every promise, confirmed on v24.15.0), a throwing `after`/`afterEach` hook (the affected files use none), and a double `close()` rejecting (awaited in a `finally`, it would be attributed to that test with a stack). A second bounded pass — twelve more full runs under the TAP reporter with a preload recording `uncaughtException`, `unhandledRejection` and any non-zero exit — produced twelve clean runs and captured nothing at the time. **A follow-up increment (`claude/node-test-signature-b`) then captured the defect directly, three more times, on three files never previously implicated** (`rate-limit-atomicity`, `dependency-parity`, `studies-api`) — seven distinct files observed with this symptom to date. Occurrences across seven distinct files make a shared or cross-cutting path more plausible and make a defect confined to one test file less likely, but do not exclude file-specific inputs or lifecycle interactions. An instrumented preload (`packages/api/test/diagnostics/signature-b-preload.cjs`) hooking process-level events — `process.exit`, `process.abort`, `process.kill`, `uncaughtExceptionMonitor` (passively observing uncaught exceptions and fatal unhandled rejections), `warning`, `beforeExit`, and Node’s own unconditional `exit` — showed **none of the hooks active at the time fired** on any of the three historical captures (though `process.abort()` was not wrapped in those initial runs and is now covered for future occurrences). A synthetic `process.exit(1)`-before-registration fixture reproduces the identical silent shape; every other synthetic mechanism tried (a post-test async throw, an emitter `'error'` with its listener removed, a synchronous module-load throw, a delayed `SIGKILL`) prints a visibly different diagnostic line, stack, or partial test output that the real defect never shows. This narrows the investigated possibilities while leaving the root cause unresolved: the per-file child process (`node --test` spawns one per file, confirmed by distinct PIDs) was not terminated by `process.exit`, uncaught exceptions, or fatal unhandled rejections, and future runs with `process.abort` instrumentation will record whether abort was called through JS; an absent record narrows in-runtime JS termination but cannot alone prove external termination without corroborating child exit status/signal data or OS-level crash evidence (e.g. distinguishing an external kill or uncatchable signal from a native C++/V8 crash). The machine had roughly 2.5 GB of 15.7 GB RAM free at capture time with several other agents’ processes concurrently running, which is circumstantially consistent with resource contention, but no crash was recorded in the Windows Application or System event logs in that window, so the exact external trigger is still not established. No fix was invented — the forbidden responses (sleeps, whole-file retries, lowering concurrency) would only hide the unresolved root cause, whose origin is not yet established. **A further increment then crossed the parent/child boundary the earlier work stopped at, and found the evidence had been there all along:** Node's runner attaches the child's `exitCode` and `signal` to the `ERR_TEST_FAILURE` it throws, and the `spec` reporter discards them — `formatError` replaces the error with `error.cause`, the bare string `'test failed'` — while the built-in `tap` reporter serializes them, so running `spec` to stdout and `tap` to a file recovers the exit status with no custom reporter and no patched internals. Exit codes were measured on this platform rather than assumed: `process.abort()` gives `134`, `Stop-Process -Force` gives `4294967295`, NTSTATUS faults surface as raw unsigned values such as `3221225477` (`0xC0000005`) — and `1` is produced alike by an uncaught exception, `process.exit(1)`, `taskkill /F` and `process.kill`, so it identifies nothing on its own and is classified `inconclusive`. `signature-b-correlate.cjs` joins the parent's TAP record to the child's JSONL log on the test file path (which also yields the child PID) and states what the pair does and does not establish; where the exit code is ambiguous, a child that reached `preload-installed` and then logged nothing still excludes `process.exit` and an uncaught exception, because both would have left a record and fired Node's `exit` event. A bounded pass of 20 runs under this instrumentation produced 0 captures — which bounds the rate and proves nothing: treating the historical ~1-in-5 as an independent per-run rate, zero captures in 20 runs has probability `(4/5)^20 ≈ 1.2%`, and independence is an assumption rather than an established fact; it ran at 3084–3834 MB free against roughly 2.5 GB at the historical captures, consistent with the resource-contention hypothesis but not evidence for it. **Signature B stays UNRESOLVED**; what changed is that the next occurrence is readable rather than silent. **During M15 Increment 46 validation, Signature B was directly observed three more times, broadening the observed scope to `packages/persistence`:** `search-backfill.integration.test.ts`, `learning.integration.test.ts`, and `test-database.integration.test.ts` each died with the documented bare whole-file `'test failed'` with zero tests reporting and no assertion or stack, and each passed standalone against the exact database state it died on. This confirms the defect is not specific to `packages/api` or its HTTP server test harness. **M15 Increment 47 hardened the diagnostic correlator across 26 bounded runs:** 1 baseline full-suite run, 20 sequential runs under `run-signature-b-pass.mjs`, and 5 concurrent repository-native load runs produced 0 captures (which bounds the rate under those conditions and does not resolve the defect). Code inspection and falsification identified and fixed four correlator blind spots in `packages/api/test/diagnostics/signature-b-correlate.cjs`: (1) cross-directory basename fallback could falsely match a child log from another directory when the target had slashes; (2) Windows path case folding was host-dependent (`process.platform === 'win32'`), breaking cross-platform correlation when logs were analyzed on a different OS; (3) signed 32-bit Windows NTSTATUS exit codes (e.g. `0xC0000005` surfacing as `-1073741819`) were not normalized to unsigned values; and (4) quoted TAP YAML scalar tokens lost exit code or duration numbers. The targeted test suite was expanded from 23 to 41 tests (39 pass, 2 skip, 0 fail), killing 4 falsification mutations. **Signature B remains UNRESOLVED;** no production fix was invented, and the next occurrence will be correlated without cross-platform or exit-code blind spots. **M15 Increment 48 recorded 0 Signature B occurrences** across its baseline reproduction, A/B/C acceptance, mutation rounds and full-repository validation — which, like Increment 47's 26 bounded runs, bounds the rate under those conditions and resolves nothing. It stays an open tracked defect. **M15 Increment 49 likewise recorded 0 occurrences** across its pre-fix reproduction on four databases, its fresh-database acceptance runs, two rounds of falsification and a full-repository run of 3304 tests. Three zero-observation increments in a row bound the rate and establish nothing about the cause; it stays an open tracked defect. See ADR-0140 §4. +- **`startHarness` drew ports `fetch` refuses (RESOLVED in ADR-0140); the unexplained whole-file failure is still open.** Two signatures were filed here, deliberately not as one cause, and that judgement held. **Signature A is resolved.** WHATWG Fetch blocks eighty-two ports and undici enforces the list on the port number alone, before opening a socket — so `server.listen(0)` could bind, listen and answer raw TCP while `fetch` still refused, surfacing as `TypeError: fetch failed` / `Error: bad port` at the harness’s first request rather than at the listen that caused it. Whether it can happen at all is a property of the host’s dynamic port range: a typical Linux CI range (32768–60999) contains no blocked port, while the Windows range in use here (1024–15000) contains nineteen, which is the whole of "green on CI, flaky locally". The guard that already existed was incomplete in a way that still failed — its hand-observed set of eighteen ports was the spec list intersected with one machine’s range **minus `6679`** — and it retried unboundedly, had no behaviour on exhaustion, and had been copy-pasted into `auth-signin-schema.integration.test.ts`, so the missing port had to be found twice. `packages/api/test/listen.ts` now owns port acquisition for both sites: the spec-complete eighty-two ports (verified by sweeping all 65535 through the real `fetch` on Node v24.15.0), a bounded twenty attempts, each rejected listener closed before the next is asked for, a guarded `address()` read in place of the `as AddressInfo` cast, and an exhaustion error naming the attempts and rejected ports and nothing else. **Signature B is not resolved and was not folded in.** A file still fails with `'test failed'`, no assertion, no stack and none of its own tests reported. Twenty consecutive full runs gave five failures: on pre-fix code one signature A (`auth.test.js`) and three signature B; on post-fix code one signature B and no signature A. The port fix removes A and leaves B exactly where it was, which is the evidence that they are two defects. Four different files were hit (`move-explanation-route`, `tournament-commentary-route`, `bot-detection-analyze`, `anti-cheat-analysis`), sharing no import beyond `./helpers`; each died in 589–703 ms with no test of its own reporting and no stderr. Refuted with evidence: ephemeral-port exhaustion (113 sockets in TIME_WAIT against a 13977-port range), a `Promise.race` loser becoming an unhandled rejection (`race` subscribes to every promise, confirmed on v24.15.0), a throwing `after`/`afterEach` hook (the affected files use none), and a double `close()` rejecting (awaited in a `finally`, it would be attributed to that test with a stack). A second bounded pass — twelve more full runs under the TAP reporter with a preload recording `uncaughtException`, `unhandledRejection` and any non-zero exit — produced twelve clean runs and captured nothing at the time. **A follow-up increment (`claude/node-test-signature-b`) then captured the defect directly, three more times, on three files never previously implicated** (`rate-limit-atomicity`, `dependency-parity`, `studies-api`) — seven distinct files observed with this symptom to date. Occurrences across seven distinct files make a shared or cross-cutting path more plausible and make a defect confined to one test file less likely, but do not exclude file-specific inputs or lifecycle interactions. An instrumented preload (`packages/api/test/diagnostics/signature-b-preload.cjs`) hooking process-level events — `process.exit`, `process.abort`, `process.kill`, `uncaughtExceptionMonitor` (passively observing uncaught exceptions and fatal unhandled rejections), `warning`, `beforeExit`, and Node’s own unconditional `exit` — showed **none of the hooks active at the time fired** on any of the three historical captures (though `process.abort()` was not wrapped in those initial runs and is now covered for future occurrences). A synthetic `process.exit(1)`-before-registration fixture reproduces the identical silent shape; every other synthetic mechanism tried (a post-test async throw, an emitter `'error'` with its listener removed, a synchronous module-load throw, a delayed `SIGKILL`) prints a visibly different diagnostic line, stack, or partial test output that the real defect never shows. This narrows the investigated possibilities while leaving the root cause unresolved: the per-file child process (`node --test` spawns one per file, confirmed by distinct PIDs) was not terminated by `process.exit`, uncaught exceptions, or fatal unhandled rejections, and future runs with `process.abort` instrumentation will record whether abort was called through JS; an absent record narrows in-runtime JS termination but cannot alone prove external termination without corroborating child exit status/signal data or OS-level crash evidence (e.g. distinguishing an external kill or uncatchable signal from a native C++/V8 crash). The machine had roughly 2.5 GB of 15.7 GB RAM free at capture time with several other agents’ processes concurrently running, which is circumstantially consistent with resource contention, but no crash was recorded in the Windows Application or System event logs in that window, so the exact external trigger is still not established. No fix was invented — the forbidden responses (sleeps, whole-file retries, lowering concurrency) would only hide the unresolved root cause, whose origin is not yet established. **A further increment then crossed the parent/child boundary the earlier work stopped at, and found the evidence had been there all along:** Node's runner attaches the child's `exitCode` and `signal` to the `ERR_TEST_FAILURE` it throws, and the `spec` reporter discards them — `formatError` replaces the error with `error.cause`, the bare string `'test failed'` — while the built-in `tap` reporter serializes them, so running `spec` to stdout and `tap` to a file recovers the exit status with no custom reporter and no patched internals. Exit codes were measured on this platform rather than assumed: `process.abort()` gives `134`, `Stop-Process -Force` gives `4294967295`, NTSTATUS faults surface as raw unsigned values such as `3221225477` (`0xC0000005`) — and `1` is produced alike by an uncaught exception, `process.exit(1)`, `taskkill /F` and `process.kill`, so it identifies nothing on its own and is classified `inconclusive`. `signature-b-correlate.cjs` joins the parent's TAP record to the child's JSONL log on the test file path (which also yields the child PID) and states what the pair does and does not establish; where the exit code is ambiguous, a child that reached `preload-installed` and then logged nothing still excludes `process.exit` and an uncaught exception, because both would have left a record and fired Node's `exit` event. A bounded pass of 20 runs under this instrumentation produced 0 captures — which bounds the rate and proves nothing: treating the historical ~1-in-5 as an independent per-run rate, zero captures in 20 runs has probability `(4/5)^20 ≈ 1.2%`, and independence is an assumption rather than an established fact; it ran at 3084–3834 MB free against roughly 2.5 GB at the historical captures, consistent with the resource-contention hypothesis but not evidence for it. **Signature B stays UNRESOLVED**; what changed is that the next occurrence is readable rather than silent. **During M15 Increment 46 validation, Signature B was directly observed three more times, broadening the observed scope to `packages/persistence`:** `search-backfill.integration.test.ts`, `learning.integration.test.ts`, and `test-database.integration.test.ts` each died with the documented bare whole-file `'test failed'` with zero tests reporting and no assertion or stack, and each passed standalone against the exact database state it died on. This confirms the defect is not specific to `packages/api` or its HTTP server test harness. **M15 Increment 47 hardened the diagnostic correlator across 26 bounded runs:** 1 baseline full-suite run, 20 sequential runs under `run-signature-b-pass.mjs`, and 5 concurrent repository-native load runs produced 0 captures (which bounds the rate under those conditions and does not resolve the defect). Code inspection and falsification identified and fixed four correlator blind spots in `packages/api/test/diagnostics/signature-b-correlate.cjs`: (1) cross-directory basename fallback could falsely match a child log from another directory when the target had slashes; (2) Windows path case folding was host-dependent (`process.platform === 'win32'`), breaking cross-platform correlation when logs were analyzed on a different OS; (3) signed 32-bit Windows NTSTATUS exit codes (e.g. `0xC0000005` surfacing as `-1073741819`) were not normalized to unsigned values; and (4) quoted TAP YAML scalar tokens lost exit code or duration numbers. The targeted test suite was expanded from 23 to 41 tests (39 pass, 2 skip, 0 fail), killing 4 falsification mutations. **Signature B remains UNRESOLVED;** no production fix was invented, and the next occurrence will be correlated without cross-platform or exit-code blind spots. **M15 Increment 48 recorded 0 Signature B occurrences** across its baseline reproduction, A/B/C acceptance, mutation rounds and full-repository validation — which, like Increment 47's 26 bounded runs, bounds the rate under those conditions and resolves nothing. It stays an open tracked defect. **M15 Increment 49 likewise recorded 0 occurrences** across its pre-fix reproduction on four databases, its fresh-database acceptance runs, two rounds of falsification and a full-repository run of 3304 tests. Three zero-observation increments in a row bound the rate and establish nothing about the cause; it stays an open tracked defect. See ADR-0140 §4. **M15 Increment 51 captured it five more times and named the mechanism family, without resolving it.** A bounded campaign of 39 runs across `packages/api` and `packages/persistence` produced 5 real captures, every one of them reporting exit status `3221226505` (`0xC0000409`, Windows `STATUS_STACK_BUFFER_OVERRUN`, the fail-fast-class status) with `signal: null`, a child lifecycle log holding exactly `start` and `preload-installed`, and no normal JavaScript shutdown hook running at all. On the four captures taken with the hardened harness, the fatal-marker channel and the Node diagnostic-report channel were both empty, which is what excludes the measured V8/Node fatal path (including heap OOM): the synthetic fatal path writes fatal stderr diagnostics into the TAP report **and** a PID-attributable report under `--report-on-fatalerror`, and these produced neither. `process.exit`, ordinary uncaught exceptions, fatal unhandled rejections and an instrumented JS `process.abort()` are excluded by the same silent lifecycle log, and node:test parent cancellation is excluded by reading the runner's own source: `FileTest` sets `this.timeout = null`, so no parent file-level wall clock exists, and an aborted child sets `err` through `child.on('error')` and is reported as an `AbortError` rather than the bare fallback. **Two sources remain and the evidence does not choose between them:** an in-process Windows fail-fast path (a security mitigation, `RaiseFailFastException`, or equivalent native source), or an external party calling `TerminateProcess` with `0xC0000409` as the chosen status — an exit status being an integer the terminating party picks, the status alone cannot separate them. Avast Antivirus is present and `aswhook.dll` was observed loaded inside a live `node.exe`; **injection is not causation**, so that is a concrete leading candidate for a controlled future A/B test and not a root cause. No antivirus was disabled, no exclusion was added and no security posture was changed. The Node 22 arm was 17 runs and 0 captures, which does not establish a Node-version difference. **Signature B stays UNRESOLVED at Level C** — mechanism family established, terminating source not — and the next steps (ETW `Microsoft-Windows-Kernel-Process` tracing or WER `LocalDumps`, and the AV A/B test) all change machine configuration and need owner authorization. - **Isolated-database test teardown dropped databases out from under connections that had not finished closing (RESOLVED in M15 Increment 45).** `withDatabase` in `packages/persistence/test/variant-migrations.integration.test.ts` ended its pool and then immediately ran `DROP DATABASE ... WITH (FORCE)`. `pool.end()` does not wait for its clients to close: in pg 8.22.0 `_pulseQueue` reaches the end callback in the same synchronous turn in which `_remove` filters the last client out of `_clients`, while `client.end()` has only queued the Terminate byte — instrumentation recorded **zero of four `remove` events fired at the moment `end()` resolved**. The drop could therefore still find a backend attached; `FORCE` terminated it, and the resulting `FATAL` arrived on a socket whose pool still had `idleListener` attached, which `pg` re-emitted as `pool.emit('error')` — an unhandled EventEmitter error that `node:test` attributed to whichever test was running rather than to the teardown that caused it. It surfaced as intermittent `terminating connection due to administrator command` failures in `postgres integration (persistence)` during M15 Increment 44, on a different test each run, which is the signature of a race rather than a broken assertion. The same shape existed in `packages/api/test/auth-signin-schema.integration.test.ts`, which had absorbed SQLSTATE 57P01 with a `pool.on('error', ...)` listener — a symptom fix for the same cause. **Resolved in Increment 45:** a shared `withTestDatabase` helper (`@chess-platform/persistence/test-support`) ends the pool under a bound, waits for `pg_stat_activity` to report the database unused, and drops it *without* `FORCE`. Measured on PostgreSQL 16.14, a plain drop against a still-attached backend fails with SQLSTATE 55006 and leaves that connection untouched, where `FORCE` succeeds by killing it — so the change trades a quiet, harmful success for a loud, harmless failure. FORCE remains only on the emergency path that guarantees the disposable database is still dropped once teardown has already failed — best effort, since that last drop runs inside a `catch` so it cannot bury the error being reported. The 57P01 absorber is deleted, because the corrected lifecycle never causes one. `createPool` and `migrate` are unchanged, and no migration was added. - **The persistence integration suite was not idempotent against a reused database (RESOLVED in M15 Increment 46).** Recorded as a known defect by Increment 45 and left open there. Against a fresh PostgreSQL 16 database the suite passed; a second run against the *same* database failed nine tests, deterministically — measured on 16.14 before any edit as **173 pass / 0 fail** then **164 pass / 9 fail**. CI provisions a fresh server per run, so it never surfaced there. One contract was being broken in two directions: a suite sharing `chess_test` must remove every row it created and remove nothing else. `achievements.integration.test.ts` broke the second half with an unqualified `DELETE FROM users` in `beforeEach` — `games.white_id` and `games.black_id` are the only references to `users` without `ON DELETE CASCADE` (thirty-one FKs point at `users`; twenty-nine cascade; the two that do not are both on `games`), so one game left behind by `pg.integration.test.ts` aborted the wipe with SQLSTATE 23503 before any assertion ran, and the same statement destroyed the bot accounts migration 0021 seeds, which nothing restores because `migrate` has already recorded 0021 as applied. `pg/identity-tokens.test.ts` and `tournaments.pg.integration.test.ts` broke the first half, leaving fixed primary keys behind and colliding on `users_pkey` and `tournaments_pkey` (the latter surfacing through the repository's compare-and-set as `VersionConflictError`). **Resolved in Increment 46:** suites that legitimately share the database delete exactly their own rows through `withSharedDatabase` (`packages/persistence/src/test-support/fixtures.ts`), the sibling of Increment 45's `withTestDatabase` and heir to its precedence rule — a cleanup failure never replaces the assertion that actually failed, and never disappears either. `pg.integration.test.ts` moved to disposable databases instead, because it cannot meet the cleanup half at all: it appends to `game_events`, which is append-only by production trigger, so cleaning up after itself would have meant weakening a production safety rule to suit a test. `users-batch`, `anti-cheat`, `bot-reports` and `analysis-cache` were corrected for the same contract though none of them ever failed — fresh `uuidv7()` ids meant their leaked rows could not collide, so the tables merely grew on every run. Serialization was not the fix and was not introduced: `--test-concurrency=1` is a pre-existing documented invariant and the second run failed identically under it. Acceptance is three consecutive runs against one database with no reset between them (186/186/186), after which only the three migration-seeded bot accounts remain, with no leaked disposable databases and no lingering backends; falsification killed 17 of 20 mutations. No production code, migration, checksum, constraint or repository conflict semantic changed, and no migration was added. - **The API pg-security integration suite leaked every row it created into the shared database (RESOLVED in M15 Increment 48).** Recorded as a known defect by Increment 46 and deliberately left open there. `packages/api/test/pg-security.integration.test.ts` created users, password credentials, roles, sessions and rate-limit buckets through real repositories and closed its pools without removing any of them. Because every identifier it mints is a fresh `uuidv7()`, no run collided with another, so all 11 tests passed indefinitely while the database grew — measured on PostgreSQL 16.14 before any edit as **11/11 passing and 25 rows leaked on the first run (4 users, 4 credentials, 4 roles, 4 sessions, 9 rate-limit buckets), then 11/11 again and an identical further 25 on a second run against the same database**. The failure mode was silent accumulation, not a failing test, which is why nothing in the file could see it. `rate_limit_buckets` is the sharpest case: no foreign key references it, so no cascade can ever reach those rows, and `PgRateLimiter.sweep` only evicts buckets that expired over an hour ago. **Resolved in Increment 48:** every test runs inside Increment 46's `withSharedDatabase` contract and names what it owns — users deleted by exact owned id, with the proven `ON DELETE CASCADE` foreign keys removing their credentials, roles and sessions (the only non-cascading references to `users` are `games.white_id` and `games.black_id`, and this file creates no games), and rate-limit buckets deleted by exact owned key. Identifiers are recorded before the statement that creates the row, so a body that throws after a commit still surrenders it. Prefix-matched cleanup, `TRUNCATE`, broad `DELETE`, random-identifier workarounds, retries, conflict suppression and serialization were all rejected; `./test-support/fixtures` was added to the persistence package's `exports` so the canonical helper could be reused rather than duplicated. A second defect was fixed in the same increment: the bucket-creation race test read backend PIDs before the `try` whose `finally` releases the leased client, and a pool with a client still checked out never settles `pool.end()`, so a failure there hung the file instead of reporting the error — both reads now sit inside the protected region. Acceptance is three consecutive runs against one migrated database with no reset between them (**11/11, 11/11, 11/11**), after which the database state is identical to its pre-run reading, migration seeds and unrelated sentinels are preserved, no `test_db_*` databases are leaked and no backends linger; falsification killed **8 of 9 mutations**, the survivor being the `backendPid` error path, which needs fault injection no passing suite performs. No production code, migration, constraint, foreign key or repository semantic changed. - **`analysis-cache-durable.integration.test.ts` depended on schema it did not establish (RESOLVED in M15 Increment 49).** Recorded as a known defect by Increment 48 and deliberately left open there. `packages/api/test/analysis-cache-durable.integration.test.ts` never called `migrate()`: it opened pools straight onto `DATABASE_URL` and assumed `engine_analysis_cache` — created by migration `0026` and indexed by `0027` — was already there. Re-proven on current `main` before any edit, on PostgreSQL 16.14: against a genuinely fresh, never-migrated database it was **10 tests, 4 pass, 6 fail**, three of them throwing SQLSTATE **42P01** from the suite's own `LOCK TABLE`, `DELETE` and `UPDATE`, and three failing as assertions because the durable row they expected was never written. The four that passed did so vacuously: `PgAnalysisCache` absorbs a database fault and returns a miss, so a suite about durability ran with no durability at all. Against an already-migrated database it was **10/10 — and left 5 rows behind every run**, `freshFen()` being collision-avoidance rather than cleanup. The masking was measured, not assumed: the whole `packages/api` package on a fresh database gave the same **6 failures**, while running `packages/persistence` first (186 pass) and then the file gave **10/10** — seventeen persistence suites migrate the shared `DATABASE_URL`, and both the root `test` script and the CI `postgres-integration` job run that package first. **Resolved in Increment 49:** the suite applies the canonical migrations itself, once per file behind a flag, asking the persistence package for the directory it ships (`migrationsDir()`) instead of assembling one from `process.cwd()`; every test runs inside Increment 46's `withSharedDatabase` contract; each FEN is recorded before the statement that creates its row; and cleanup deletes by exact FEN equality, never by the placement prefix every standard starting position shares. A disposable database per test was considered and rejected on cost and orphan risk; the shape chosen is the one the sibling suite for the same table already uses. Acceptance is three independent brand-new empty databases (**10/10, 10/10, 10/10**, 0 rows of residue each), an already-migrated database whose five unrelated rows survive untouched, and the whole API package on a brand-new database with no other package first (**996 tests, 986 pass, 0 fail, 10 skipped**, 0 residue, 0 leaked `test_db_*`). Falsification killed **7 of 9 mutations**; the two survivors are a failure-precedence contract owned by `withSharedDatabase` and killed by its own suite, and an equivalent mutant. No production code, migration or repository semantic changed. - **M15 Increment 50 — test:counts / standalone gateway host setup contract (RESOLVED: SETUP / DOCUMENTATION CONTRACT DRIFT CORRECTED).** Increments 48 and 49 recorded gateway `ioredis` compilation failures during counts runs. Controlled clean states proved that root `npm ci` does not install the intentionally standalone gateway, and gateway installation alone does not build the public workspace outputs its local `file:` dependencies need. The supported host sequence, from the root with Node.js 22+, is `npm ci`, `npm run build`, `npm ci --prefix services/gateway`, then `npm run test:counts`. The counting command includes the gateway but does not install dependencies or build workspace public outputs. A root clean install preserves an already prepared gateway dependency tree and workspace outputs, explaining why prepared and unprepared hosts gave different results. No gateway implementation defect was proven; no workspace topology, script, lockfile, CI, Docker or runtime change is required. Setup instructions and the canonical handover are now synchronized; see [M15 Increment 50](PROJECT_STATE.md#m15-increment-50--testcounts--standalone-gateway-host-setup-contract) for controlled evidence and validation. **Signature B remains unresolved and under separate investigation:** historical state D passed the gateway but failed the aggregate with a bare file-level `test failed` in `openapi.test.js`; the isolated rerun passed. No mechanism is assigned to that observation and no unmerged diagnostic findings are adopted. Skips are not passes. +- **M15 Increment 51 — Signature B mechanism isolation and diagnostic hardening (UNRESOLVED — exact terminating source unproven; root-cause acceptance LEVEL C).** The increment set out to identify the process-termination mechanism behind Signature B, not to fix it, and is reported at the level the evidence reaches. **39 bounded runs produced 5 real captures** across `packages/api` and `packages/persistence`; all five reported `3221226505` (`0xC0000409`) with `signal: null` and a child log of exactly `start` and `preload-installed`. Capture 1 was taken by the pre-fix harness and its raw TAP and fatal-marker evidence was lost, so it is not claimed as decisive; captures 2–5 carry the full evidence set and reported empty fatal markers and empty diagnostic reports. Two diagnostic blind spots were **proven, not guessed**, and only those were closed: fatal markers were read from the parent runner's stderr, where a child banner can never appear (the runner re-emits child stderr as `test:stderr` reporter events, so the banner lands in the TAP report — measured at 0 bytes on the parent stream while the report held it), and child lifecycle logs were merged by test-file path across processes, which had produced the false statement that an external termination and a native fault were ruled out. Also delivered: per-failure report attribution with PID matching, stale and PID-reused evidence surfaced as explicit ambiguity, report bodies never retained on any exit path, artifact paths independent of the selected package’s working directory, artifact directories that start empty, an `--out` already holding a capture refused rather than reused, `--package` for cross-package runs, and POSIX OOM signal/status coverage alongside the Windows encoding. A 17-mechanism synthetic fingerprint matrix was measured with multi-field match criteria fixed **before** any real capture was compared against it; the correlator suite grew **41 → 57 tests** (55 pass, 2 pre-existing POSIX-only skips, 0 fail), identical on Node 22 and Node 24, and falsification killed **13 of 14 mutations**, the survivor being an equivalent region-boundary mutant under the TAP grammar Node can emit. No production code, workflow, migration or repository semantic changed. **A separate defect was observed and deliberately not fixed here:** the Increment 49 durable analysis-cache race test `two live instances racing a cold position both compute it` failed once with an ordinary assertion (expected 2, actual 1) and passed on rerun — not Signature B, and the rerun passing does not resolve it. See [M15 Increment 51](PROJECT_STATE.md#m15-increment-51--signature-b-mechanism-isolation-and-diagnostic-hardening) and ADR-0140 §4. **Signature B itself occurred twice during this increment’s own final validation** (`auth-signin-schema.integration.test.js` at 618.7 ms, `cookie-auth.test.js` at 707.0 ms, both on Node 24.15.0), each with the documented bare file-level shape and each on the plain `spec`-reporter path with no TAP destination and no preload — so **neither has exit-status, lifecycle, marker or report evidence**, and neither narrows Level C. Both are recorded rather than dismissed for passing on rerun. The host was under heavy memory pressure from unrelated concurrent work (699 MB free of 16077 MB at the first occurrence); an earlier attempt at the same run died differently, with `0xC0000142` (`STATUS_DLL_INIT_FAILED`) for the whole package command, which is a process that failed to start and is **not** counted as Signature B. - **`ApiServer.listen` registered no `'error'` handler, so a failed bind hung and raised an uncaught event (RESOLVED in ADR-0140 §5).** `packages/api/src/server.ts` resolved its promise from the `listening` callback only and built it with no reject path. A bind failing asynchronously (`EADDRINUSE`, `EMFILE`) left the promise pending forever and, with no `'error'` listener on the `http.Server`, was re-raised as an uncaught exception. Found while investigating ADR-0140 and independently raised by the Qodo review of PR #21. **Resolved in the same increment** rather than deferred, because ADR-0140 §2's bounded, diagnosable acquisition is not true without it — the retry can only report a bind error if the listener it is handed rejects. A one-shot `'error'` listener now rejects and is removed once listening, so later server errors keep their previous semantics rather than being swallowed by a `reject` on a settled promise. The regression test fails against the exact pre-fix code, through an uncaught `ERR_UNHANDLED_ERROR`. - **`main.ts`'s controller-disposal list is manual, untested, and silently incomplete when a section is added (RESOLVED in Increment 25 / ADR-0092).** `run()` in `packages/web/src/main.ts` disposes the previous route's controllers by name, and its own comment says doing so "is what makes re-bootstrapping safe" — but adding a section to `bootstrap` and forgetting to add it there compiles, passes every gate, and leaks. Increment 23 shipped exactly that omission for `LearningController` and it was caught in PR review, not by a test. `main.ts` has no test coverage of any kind, so no section's disposal is verified. A structural fix (bootstrap returning its disposables as a collection, or a type-level exhaustiveness check keyed off the result type) would make the next omission a compile error; it is a refactor across ~15 return sites and belongs in its own increment. **Resolved in Increment 25 (ADR-0092):** extracted `createLifecycle` run loop in `lifecycle.ts`, defined `BootstrappedDisposables` and `DisposableKey` driving `DISPOSABLE_TEARDOWN_MAP: Record` for compile-time exhaustiveness, normalised `.dispose()` verb across all disposables, and cascaded `GameController.stop()` to `gameSync.stop()`. - **`stepView` sends the answers to the learner (RESOLVED in Increment 29 / ADR-0095).** `packages/api/src/presenters.ts` emits `expectedSan` on a move step and `correctIndex` on a quiz step, and `GET /v1/lessons/:id/steps` is the route the learner's own lesson page calls. Increment 23 omits both from the client-side types (`packages/web/src/api/models.ts`), so the app cannot render or grade against them and a future edit that tries becomes a compile error — but the fields are still on the wire and readable in devtools. The authoring routes legitimately need them returned to the author, so the fix is a learner-scoped step view (or a caller-dependent projection), not a deletion: an API contract decision with its own ADR. Nothing rated or rewarded depends on step progress today, so this is a wart rather than a breach. **Resolved in Increment 29 (ADR-0095):** added `LearnerStepView` / `learnerStepView` in `packages/api/src/presenters.ts` omitting `expectedSan` and `correctIndex`. `GET /v1/lessons/:id/steps` and `GET /v1/steps/:id` now check course authorship via `repo.getLesson` / `repo.getCourse`, returning full `stepView` to the author and `learnerStepView` to learners and anonymous callers. Updated OpenAPI schema and web model comments. The first attempt resolved authorship with a separate `getLesson` + `getCourse` after the step read, which doubled both routes from 3 SQL queries to 6 because `listSteps` / `getStep` had already made those reads internally and discarded the course; caught in the PR #92 review and fixed by adding `getStepWithCourse` / `listStepsWithCourse` to `LearningRepository`, which return what was already loaded. Both routes now make exactly one repository call, pinned by a counting-proxy test. diff --git a/docs/adr/0140-harness-ephemeral-port-acquisition.md b/docs/adr/0140-harness-ephemeral-port-acquisition.md index 191ba903..fa5b1c2b 100644 --- a/docs/adr/0140-harness-ephemeral-port-acquisition.md +++ b/docs/adr/0140-harness-ephemeral-port-acquisition.md @@ -367,6 +367,65 @@ names a **candidate** mechanism, while a bare `1`, a signal — which names the child but never the party that sent it — and any code outside the measured table are all reported `specific: false` and name nothing on their own. +### Follow-up: five captures name a mechanism family, and stop there + +The pass above read the exit status but caught nothing. A later increment (M15 Increment 51, +`claude/signature-b-mechanism-isolation`) captured Signature B **five times** across +`packages/api` and `packages/persistence` in 39 bounded runs, and every capture reported the same +thing: + +| Package | File | `exitCode` | `signal` | Child lifecycle | Fatal markers | Diagnostic reports | +|---|---|---|---|---|---|---| +| api | `pg-security-ownership.integration` | `3221226505` | `null` | `start`, `preload-installed` | *not captured* | *n/a* | +| persistence | `studies.integration` | `3221226505` | `null` | `start`, `preload-installed` | `[]` | `[]` | +| persistence | `learning.integration` | `3221226505` | `null` | `start`, `preload-installed` | `[]` | `[]` | +| api | `pg-security.integration` | `3221226505` | `null` | `start`, `preload-installed` | `[]` | `[]` | +| api | `auth-signin-schema.integration` | `3221226505` | `null` | `start`, `preload-installed` | `[]` | `[]` | + +`3221226505` is `0xC0000409`, `STATUS_STACK_BUFFER_OVERRUN` — the Windows fail-fast status. It is +not in the measured table above and was added by measurement, not inference. The first capture came +from the pre-fix harness, whose raw TAP and fatal-marker evidence for that run was discarded before +it could be read; its exit status and lifecycle log stand, its marker and report channels do not +exist, and it is not treated as decisive. + +**Two channels made the exclusions possible, and both had to be fixed first.** Fatal markers were +being read from the parent runner's stderr, where a child's `FATAL ERROR:` banner can never appear: +the runner attaches a readline interface to the child's stderr and re-emits each line as a +`test:stderr` reporter event, so the banner lands in the TAP report instead — measured at 0 bytes on +the parent stream while the report held it. And child lifecycle logs were merged by test-file path +across processes, so evidence from one run could be attributed to another; they are now isolated per +process, keyed by log file and PID, with a reused PID surfaced as explicit ambiguity rather than a +silent merge. Node's own `--report-on-fatalerror` was added as a third channel and measured: it +writes exactly one PID-named report for a Node/V8 fatal error and none for `process.abort()` or an +external kill, so its presence and PID attribution separate those paths. + +**What the four fully-instrumented captures exclude, by positive measurement:** `process.exit` and +`process.exitCode`, an ordinary uncaught exception, a fatal unhandled rejection and an instrumented +JS `process.abort()` — each of which leaves a lifecycle record and fires Node's `exit` event, and +none did; the measured V8/Node fatal path including heap OOM — which produces fatal stderr +diagnostics inside the TAP report **and** a PID-attributable diagnostic report, and these produced +**neither**; and node:test parent cancellation in this repository's configuration — `FileTest` sets +`this.timeout = null`, so the parent enforces no file-level wall clock, and a child aborted through +the runner's `AbortSignal` sets `err` via `child.on('error')` and is reported as an `AbortError` +rather than the bare fallback shape. + +**What remains is a family of two, and this ADR does not choose between them:** an in-process +Windows fail-fast path (a security mitigation, `RaiseFailFastException`, or an equivalent native +fail-fast source), or an external party calling `TerminateProcess` with `0xC0000409` as the chosen +exit status. The status cannot separate them — it is a 32-bit integer the terminating party picks, +which is the same limit §4 already records for `0xC0000005`. Separating them needs a channel this +increment deliberately did not open, because both change the machine's configuration: ETW +`Microsoft-Windows-Kernel-Process` tracing, or WER `LocalDumps`. No WER record exists for `node.exe` +in the capture windows, though WER is enabled and logged other events that day. Avast Antivirus is +present and `aswhook.dll` was observed loaded inside a live `node.exe` process; **injection is not +causation**, so that is a named leading candidate for a controlled A/B test under owner +authorization, not a cause. No antivirus was disabled, no exclusion was added, and no security +posture was changed. + +**Signature B remains UNRESOLVED, at acceptance Level C: the mechanism family is established, the +terminating source is not.** No fix is proposed and none is disguised. The Node 22 arm of the same +campaign was 17 runs and 0 captures, which does not establish a Node-version difference. + ## 5. `ApiServer.listen` rejects on a failed bind Raised by the Qodo review of PR #21 and **valid**. `packages/api/src/server.ts` From c9616813af1ab6ba257ee67340c6c589fa317650 Mon Sep 17 00:00:00 2001 From: Hussein Mohamed Date: Sat, 5 Sep 2026 19:41:42 +0300 Subject: [PATCH 8/9] docs: record the analysis-cache race defect's second observation The Increment 51 entry said the durable analysis-cache race test failed once and passed on rerun. It has now failed again, in CI's postgres integration (persistence) job on Linux at this branch's HEAD, with the same assertion: cross-process single-flight does not exist, expected 2, actual 1. Two observations on two operating systems make it a real intermittent defect rather than local noise, so the entry now says so. It remains owned by the Increment 49 suite, is not Signature B, is not caused by this increment, and is not fixed here. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01CUPqvu66J5ZVyiv4nJ797r --- docs/PROJECT_STATE.md | 19 ++++++++++++++----- docs/ROADMAP.md | 2 +- 2 files changed, 15 insertions(+), 6 deletions(-) diff --git a/docs/PROJECT_STATE.md b/docs/PROJECT_STATE.md index ab6bb48b..217410e7 100644 --- a/docs/PROJECT_STATE.md +++ b/docs/PROJECT_STATE.md @@ -182,17 +182,26 @@ and because omitting it would make the memory-pressure correlation look cleaner ### Separate known defect observed during this increment — not fixed here **The durable analysis-cache race test `two live instances racing a cold position both compute it` -failed once and passed on rerun.** It lives in the Increment 49 suite, which this increment does not -touch, and it passed at the two prior HEADs. This is **not** Signature B: it is an ordinary assertion -failure (expected 2, actual 1) with a normal stack, not a bare file-level termination. The rerun -passing does **not** resolve it. Recorded here as a separate bounded follow-up. +has now failed twice.** It lives in `packages/api/test/analysis-cache-durable.integration.test.ts`, +the Increment 49 suite, which this increment does not touch. First locally, during this increment's +work; then again in CI's `postgres integration (persistence)` job on Linux at this branch's HEAD, +where it failed with `cross-process single-flight does not exist` — `expected: 2, actual: 1`, +`ERR_ASSERTION`, at `analysis-cache-durable.integration.test.js:385`. It passed at the two prior +HEADs and passes on rerun. + +This is **not** Signature B: it is an ordinary assertion failure with a full stack and a named +assertion, not a bare file-level termination with no test reported. It is also not caused by this +increment, which changes only diagnostics and documentation. **A rerun passing does not resolve +it** — the second occurrence, on a different operating system from the first, makes it a real +intermittent defect rather than local noise. Recorded here as a separate bounded follow-up, not +fixed here. ### Still open after Increment 51 - **Signature B — UNRESOLVED**, now at Level C: the mechanism family is established, the terminating source is not. The controlled AV A/B test and OS-level tracing (ETW / WER `LocalDumps`) are the named next steps, and both need owner authorization because they change machine configuration. -- **The analysis-cache race flake above — OPEN**, separate from Signature B. +- **The analysis-cache race flake above — OPEN**, separate from Signature B, now observed twice (locally on Windows and in CI on Linux) and owned by the Increment 49 suite. ## M15 Increment 50 — test:counts / standalone gateway host setup contract diff --git a/docs/ROADMAP.md b/docs/ROADMAP.md index 3164c09b..57534527 100644 --- a/docs/ROADMAP.md +++ b/docs/ROADMAP.md @@ -1355,7 +1355,7 @@ Debt observed during M14. Each states what is known, not what is planned; items - **The API pg-security integration suite leaked every row it created into the shared database (RESOLVED in M15 Increment 48).** Recorded as a known defect by Increment 46 and deliberately left open there. `packages/api/test/pg-security.integration.test.ts` created users, password credentials, roles, sessions and rate-limit buckets through real repositories and closed its pools without removing any of them. Because every identifier it mints is a fresh `uuidv7()`, no run collided with another, so all 11 tests passed indefinitely while the database grew — measured on PostgreSQL 16.14 before any edit as **11/11 passing and 25 rows leaked on the first run (4 users, 4 credentials, 4 roles, 4 sessions, 9 rate-limit buckets), then 11/11 again and an identical further 25 on a second run against the same database**. The failure mode was silent accumulation, not a failing test, which is why nothing in the file could see it. `rate_limit_buckets` is the sharpest case: no foreign key references it, so no cascade can ever reach those rows, and `PgRateLimiter.sweep` only evicts buckets that expired over an hour ago. **Resolved in Increment 48:** every test runs inside Increment 46's `withSharedDatabase` contract and names what it owns — users deleted by exact owned id, with the proven `ON DELETE CASCADE` foreign keys removing their credentials, roles and sessions (the only non-cascading references to `users` are `games.white_id` and `games.black_id`, and this file creates no games), and rate-limit buckets deleted by exact owned key. Identifiers are recorded before the statement that creates the row, so a body that throws after a commit still surrenders it. Prefix-matched cleanup, `TRUNCATE`, broad `DELETE`, random-identifier workarounds, retries, conflict suppression and serialization were all rejected; `./test-support/fixtures` was added to the persistence package's `exports` so the canonical helper could be reused rather than duplicated. A second defect was fixed in the same increment: the bucket-creation race test read backend PIDs before the `try` whose `finally` releases the leased client, and a pool with a client still checked out never settles `pool.end()`, so a failure there hung the file instead of reporting the error — both reads now sit inside the protected region. Acceptance is three consecutive runs against one migrated database with no reset between them (**11/11, 11/11, 11/11**), after which the database state is identical to its pre-run reading, migration seeds and unrelated sentinels are preserved, no `test_db_*` databases are leaked and no backends linger; falsification killed **8 of 9 mutations**, the survivor being the `backendPid` error path, which needs fault injection no passing suite performs. No production code, migration, constraint, foreign key or repository semantic changed. - **`analysis-cache-durable.integration.test.ts` depended on schema it did not establish (RESOLVED in M15 Increment 49).** Recorded as a known defect by Increment 48 and deliberately left open there. `packages/api/test/analysis-cache-durable.integration.test.ts` never called `migrate()`: it opened pools straight onto `DATABASE_URL` and assumed `engine_analysis_cache` — created by migration `0026` and indexed by `0027` — was already there. Re-proven on current `main` before any edit, on PostgreSQL 16.14: against a genuinely fresh, never-migrated database it was **10 tests, 4 pass, 6 fail**, three of them throwing SQLSTATE **42P01** from the suite's own `LOCK TABLE`, `DELETE` and `UPDATE`, and three failing as assertions because the durable row they expected was never written. The four that passed did so vacuously: `PgAnalysisCache` absorbs a database fault and returns a miss, so a suite about durability ran with no durability at all. Against an already-migrated database it was **10/10 — and left 5 rows behind every run**, `freshFen()` being collision-avoidance rather than cleanup. The masking was measured, not assumed: the whole `packages/api` package on a fresh database gave the same **6 failures**, while running `packages/persistence` first (186 pass) and then the file gave **10/10** — seventeen persistence suites migrate the shared `DATABASE_URL`, and both the root `test` script and the CI `postgres-integration` job run that package first. **Resolved in Increment 49:** the suite applies the canonical migrations itself, once per file behind a flag, asking the persistence package for the directory it ships (`migrationsDir()`) instead of assembling one from `process.cwd()`; every test runs inside Increment 46's `withSharedDatabase` contract; each FEN is recorded before the statement that creates its row; and cleanup deletes by exact FEN equality, never by the placement prefix every standard starting position shares. A disposable database per test was considered and rejected on cost and orphan risk; the shape chosen is the one the sibling suite for the same table already uses. Acceptance is three independent brand-new empty databases (**10/10, 10/10, 10/10**, 0 rows of residue each), an already-migrated database whose five unrelated rows survive untouched, and the whole API package on a brand-new database with no other package first (**996 tests, 986 pass, 0 fail, 10 skipped**, 0 residue, 0 leaked `test_db_*`). Falsification killed **7 of 9 mutations**; the two survivors are a failure-precedence contract owned by `withSharedDatabase` and killed by its own suite, and an equivalent mutant. No production code, migration or repository semantic changed. - **M15 Increment 50 — test:counts / standalone gateway host setup contract (RESOLVED: SETUP / DOCUMENTATION CONTRACT DRIFT CORRECTED).** Increments 48 and 49 recorded gateway `ioredis` compilation failures during counts runs. Controlled clean states proved that root `npm ci` does not install the intentionally standalone gateway, and gateway installation alone does not build the public workspace outputs its local `file:` dependencies need. The supported host sequence, from the root with Node.js 22+, is `npm ci`, `npm run build`, `npm ci --prefix services/gateway`, then `npm run test:counts`. The counting command includes the gateway but does not install dependencies or build workspace public outputs. A root clean install preserves an already prepared gateway dependency tree and workspace outputs, explaining why prepared and unprepared hosts gave different results. No gateway implementation defect was proven; no workspace topology, script, lockfile, CI, Docker or runtime change is required. Setup instructions and the canonical handover are now synchronized; see [M15 Increment 50](PROJECT_STATE.md#m15-increment-50--testcounts--standalone-gateway-host-setup-contract) for controlled evidence and validation. **Signature B remains unresolved and under separate investigation:** historical state D passed the gateway but failed the aggregate with a bare file-level `test failed` in `openapi.test.js`; the isolated rerun passed. No mechanism is assigned to that observation and no unmerged diagnostic findings are adopted. Skips are not passes. -- **M15 Increment 51 — Signature B mechanism isolation and diagnostic hardening (UNRESOLVED — exact terminating source unproven; root-cause acceptance LEVEL C).** The increment set out to identify the process-termination mechanism behind Signature B, not to fix it, and is reported at the level the evidence reaches. **39 bounded runs produced 5 real captures** across `packages/api` and `packages/persistence`; all five reported `3221226505` (`0xC0000409`) with `signal: null` and a child log of exactly `start` and `preload-installed`. Capture 1 was taken by the pre-fix harness and its raw TAP and fatal-marker evidence was lost, so it is not claimed as decisive; captures 2–5 carry the full evidence set and reported empty fatal markers and empty diagnostic reports. Two diagnostic blind spots were **proven, not guessed**, and only those were closed: fatal markers were read from the parent runner's stderr, where a child banner can never appear (the runner re-emits child stderr as `test:stderr` reporter events, so the banner lands in the TAP report — measured at 0 bytes on the parent stream while the report held it), and child lifecycle logs were merged by test-file path across processes, which had produced the false statement that an external termination and a native fault were ruled out. Also delivered: per-failure report attribution with PID matching, stale and PID-reused evidence surfaced as explicit ambiguity, report bodies never retained on any exit path, artifact paths independent of the selected package’s working directory, artifact directories that start empty, an `--out` already holding a capture refused rather than reused, `--package` for cross-package runs, and POSIX OOM signal/status coverage alongside the Windows encoding. A 17-mechanism synthetic fingerprint matrix was measured with multi-field match criteria fixed **before** any real capture was compared against it; the correlator suite grew **41 → 57 tests** (55 pass, 2 pre-existing POSIX-only skips, 0 fail), identical on Node 22 and Node 24, and falsification killed **13 of 14 mutations**, the survivor being an equivalent region-boundary mutant under the TAP grammar Node can emit. No production code, workflow, migration or repository semantic changed. **A separate defect was observed and deliberately not fixed here:** the Increment 49 durable analysis-cache race test `two live instances racing a cold position both compute it` failed once with an ordinary assertion (expected 2, actual 1) and passed on rerun — not Signature B, and the rerun passing does not resolve it. See [M15 Increment 51](PROJECT_STATE.md#m15-increment-51--signature-b-mechanism-isolation-and-diagnostic-hardening) and ADR-0140 §4. **Signature B itself occurred twice during this increment’s own final validation** (`auth-signin-schema.integration.test.js` at 618.7 ms, `cookie-auth.test.js` at 707.0 ms, both on Node 24.15.0), each with the documented bare file-level shape and each on the plain `spec`-reporter path with no TAP destination and no preload — so **neither has exit-status, lifecycle, marker or report evidence**, and neither narrows Level C. Both are recorded rather than dismissed for passing on rerun. The host was under heavy memory pressure from unrelated concurrent work (699 MB free of 16077 MB at the first occurrence); an earlier attempt at the same run died differently, with `0xC0000142` (`STATUS_DLL_INIT_FAILED`) for the whole package command, which is a process that failed to start and is **not** counted as Signature B. +- **M15 Increment 51 — Signature B mechanism isolation and diagnostic hardening (UNRESOLVED — exact terminating source unproven; root-cause acceptance LEVEL C).** The increment set out to identify the process-termination mechanism behind Signature B, not to fix it, and is reported at the level the evidence reaches. **39 bounded runs produced 5 real captures** across `packages/api` and `packages/persistence`; all five reported `3221226505` (`0xC0000409`) with `signal: null` and a child log of exactly `start` and `preload-installed`. Capture 1 was taken by the pre-fix harness and its raw TAP and fatal-marker evidence was lost, so it is not claimed as decisive; captures 2–5 carry the full evidence set and reported empty fatal markers and empty diagnostic reports. Two diagnostic blind spots were **proven, not guessed**, and only those were closed: fatal markers were read from the parent runner's stderr, where a child banner can never appear (the runner re-emits child stderr as `test:stderr` reporter events, so the banner lands in the TAP report — measured at 0 bytes on the parent stream while the report held it), and child lifecycle logs were merged by test-file path across processes, which had produced the false statement that an external termination and a native fault were ruled out. Also delivered: per-failure report attribution with PID matching, stale and PID-reused evidence surfaced as explicit ambiguity, report bodies never retained on any exit path, artifact paths independent of the selected package’s working directory, artifact directories that start empty, an `--out` already holding a capture refused rather than reused, `--package` for cross-package runs, and POSIX OOM signal/status coverage alongside the Windows encoding. A 17-mechanism synthetic fingerprint matrix was measured with multi-field match criteria fixed **before** any real capture was compared against it; the correlator suite grew **41 → 57 tests** (55 pass, 2 pre-existing POSIX-only skips, 0 fail), identical on Node 22 and Node 24, and falsification killed **13 of 14 mutations**, the survivor being an equivalent region-boundary mutant under the TAP grammar Node can emit. No production code, workflow, migration or repository semantic changed. **A separate defect was observed and deliberately not fixed here:** the Increment 49 durable analysis-cache race test `two live instances racing a cold position both compute it` failed with an ordinary assertion (`expected: 2, actual: 1`, `cross-process single-flight does not exist`) — once locally on Windows and once in CI’s `postgres integration (persistence)` job on Linux at this branch’s HEAD. It is not Signature B (a full stack and a named assertion, not a bare file-level termination), it is not caused by this increment, and passing on rerun does not resolve it; two observations on two operating systems make it a real intermittent defect owned by the Increment 49 suite. See [M15 Increment 51](PROJECT_STATE.md#m15-increment-51--signature-b-mechanism-isolation-and-diagnostic-hardening) and ADR-0140 §4. **Signature B itself occurred twice during this increment’s own final validation** (`auth-signin-schema.integration.test.js` at 618.7 ms, `cookie-auth.test.js` at 707.0 ms, both on Node 24.15.0), each with the documented bare file-level shape and each on the plain `spec`-reporter path with no TAP destination and no preload — so **neither has exit-status, lifecycle, marker or report evidence**, and neither narrows Level C. Both are recorded rather than dismissed for passing on rerun. The host was under heavy memory pressure from unrelated concurrent work (699 MB free of 16077 MB at the first occurrence); an earlier attempt at the same run died differently, with `0xC0000142` (`STATUS_DLL_INIT_FAILED`) for the whole package command, which is a process that failed to start and is **not** counted as Signature B. - **`ApiServer.listen` registered no `'error'` handler, so a failed bind hung and raised an uncaught event (RESOLVED in ADR-0140 §5).** `packages/api/src/server.ts` resolved its promise from the `listening` callback only and built it with no reject path. A bind failing asynchronously (`EADDRINUSE`, `EMFILE`) left the promise pending forever and, with no `'error'` listener on the `http.Server`, was re-raised as an uncaught exception. Found while investigating ADR-0140 and independently raised by the Qodo review of PR #21. **Resolved in the same increment** rather than deferred, because ADR-0140 §2's bounded, diagnosable acquisition is not true without it — the retry can only report a bind error if the listener it is handed rejects. A one-shot `'error'` listener now rejects and is removed once listening, so later server errors keep their previous semantics rather than being swallowed by a `reject` on a settled promise. The regression test fails against the exact pre-fix code, through an uncaught `ERR_UNHANDLED_ERROR`. - **`main.ts`'s controller-disposal list is manual, untested, and silently incomplete when a section is added (RESOLVED in Increment 25 / ADR-0092).** `run()` in `packages/web/src/main.ts` disposes the previous route's controllers by name, and its own comment says doing so "is what makes re-bootstrapping safe" — but adding a section to `bootstrap` and forgetting to add it there compiles, passes every gate, and leaks. Increment 23 shipped exactly that omission for `LearningController` and it was caught in PR review, not by a test. `main.ts` has no test coverage of any kind, so no section's disposal is verified. A structural fix (bootstrap returning its disposables as a collection, or a type-level exhaustiveness check keyed off the result type) would make the next omission a compile error; it is a refactor across ~15 return sites and belongs in its own increment. **Resolved in Increment 25 (ADR-0092):** extracted `createLifecycle` run loop in `lifecycle.ts`, defined `BootstrappedDisposables` and `DisposableKey` driving `DISPOSABLE_TEARDOWN_MAP: Record` for compile-time exhaustiveness, normalised `.dispose()` verb across all disposables, and cascaded `GameController.stop()` to `gameSync.stop()`. - **`stepView` sends the answers to the learner (RESOLVED in Increment 29 / ADR-0095).** `packages/api/src/presenters.ts` emits `expectedSan` on a move step and `correctIndex` on a quiz step, and `GET /v1/lessons/:id/steps` is the route the learner's own lesson page calls. Increment 23 omits both from the client-side types (`packages/web/src/api/models.ts`), so the app cannot render or grade against them and a future edit that tries becomes a compile error — but the fields are still on the wire and readable in devtools. The authoring routes legitimately need them returned to the author, so the fix is a learner-scoped step view (or a caller-dependent projection), not a deletion: an API contract decision with its own ADR. Nothing rated or rewarded depends on step progress today, so this is a wart rather than a breach. **Resolved in Increment 29 (ADR-0095):** added `LearnerStepView` / `learnerStepView` in `packages/api/src/presenters.ts` omitting `expectedSan` and `correctIndex`. `GET /v1/lessons/:id/steps` and `GET /v1/steps/:id` now check course authorship via `repo.getLesson` / `repo.getCourse`, returning full `stepView` to the author and `learnerStepView` to learners and anonymous callers. Updated OpenAPI schema and web model comments. The first attempt resolved authorship with a separate `getLesson` + `getCourse` after the step read, which doubled both routes from 3 SQL queries to 6 because `listSteps` / `getStep` had already made those reads internally and discarded the course; caught in the PR #92 review and fixed by adding `getStepWithCourse` / `listStepsWithCourse` to `LearningRepository`, which return what was already loaded. Both routes now make exactly one repository call, pinned by a counting-proxy test. From 530e701ac60c04043a77b3c860e00e7f30d441f3 Mon Sep 17 00:00:00 2001 From: Hussein Mohamed Date: Sat, 5 Sep 2026 19:44:58 +0300 Subject: [PATCH 9/9] docs(adr-0140): separate process.abort from the exit-event claim Raised by the CodeRabbit review of PR #44 and valid. The exclusion list grouped an instrumented process.abort() with the paths that leave a lifecycle record and fire Node's exit event. process.abort() fires no exit event at all - which is exactly what section 4 of this ADR already establishes - so the sentence made a claim the rest of the document contradicts. What excludes abort is the record the preload writes synchronously before delegating to the original binding, and the text now says so. PROJECT_STATE.md already described it correctly; this aligns the ADR with it. No evidence, conclusion or acceptance level changes. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01CUPqvu66J5ZVyiv4nJ797r --- docs/adr/0140-harness-ephemeral-port-acquisition.md | 13 ++++++++----- 1 file changed, 8 insertions(+), 5 deletions(-) diff --git a/docs/adr/0140-harness-ephemeral-port-acquisition.md b/docs/adr/0140-harness-ephemeral-port-acquisition.md index fa5b1c2b..9a09a058 100644 --- a/docs/adr/0140-harness-ephemeral-port-acquisition.md +++ b/docs/adr/0140-harness-ephemeral-port-acquisition.md @@ -400,11 +400,14 @@ writes exactly one PID-named report for a Node/V8 fatal error and none for `proc external kill, so its presence and PID attribution separate those paths. **What the four fully-instrumented captures exclude, by positive measurement:** `process.exit` and -`process.exitCode`, an ordinary uncaught exception, a fatal unhandled rejection and an instrumented -JS `process.abort()` — each of which leaves a lifecycle record and fires Node's `exit` event, and -none did; the measured V8/Node fatal path including heap OOM — which produces fatal stderr -diagnostics inside the TAP report **and** a PID-attributable diagnostic report, and these produced -**neither**; and node:test parent cancellation in this repository's configuration — `FileTest` sets +`process.exitCode`, an ordinary uncaught exception and a fatal unhandled rejection — each of which +leaves a lifecycle record **and** fires Node's `exit` event, and none did; an instrumented JS +`process.abort()`, which is a separate case because it terminates immediately and fires no `exit` +event at all, and is excluded instead by the record the preload writes synchronously *before* +delegating to the original binding, as §4 above already sets out; the measured V8/Node fatal path +including heap OOM — which produces fatal stderr diagnostics inside the TAP report **and** a +PID-attributable diagnostic report, and these produced **neither**; and node:test parent +cancellation in this repository's configuration — `FileTest` sets `this.timeout = null`, so the parent enforces no file-level wall clock, and a child aborted through the runner's `AbortSignal` sets `err` via `child.on('error')` and is reported as an `AbortError` rather than the bare fallback shape.