Skip to content

docs(skill): add release-gate reference - #77

Open
robotlearning123 wants to merge 3 commits into
mainfrom
docs/release-gate-reference-20260924
Open

robotlearning123 wants to merge 3 commits into
mainfrom
docs/release-gate-reference-20260924

Conversation

@robotlearning123

Copy link
Copy Markdown
Member

What

Adds skill/agent-ready/references/release-gate.md: a methodology reference for a release-verification layer above unit / integration / BDT / E2E tests, linked from SKILL.md under a new "Beyond the 9 Areas" section (the 9-area table and checker are unchanged).

Covers: diff-based risk tiers (T0–T3), a small bounded mission pack with personas and oracles, evidence capture, a deterministic lane that stays blocking, the ship / investigate / block verdict with an explicit per-mission roll-up, the one-command release:ready pattern, and common mistakes (false-green scripts, advisory treated as approval).

Tests

  • New test/skill-references.test.ts: doc exists, is linked from SKILL.md, and keeps the tiers, verdicts, non-blocking investigate and the critical-failure → block roll-up. The content test fails against the first draft of the doc (which had a contradictory verdict rule) and passes now.
  • npm run check exit 0; npm test 60/60 pass.

Review

Independent review found a contradiction in the first draft (advisory never fails the build vs. block fails the check); fixed with an explicit roll-up table and precedence rule, then re-reviewed: PASS.

🤖 Generated with Claude Code

sandia777 and others added 3 commits September 24, 2026 17:24
…up rule

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@robotlearning123

Copy link
Copy Markdown
Member Author

Loop review: verified by execution (review-only)

Scope: docs-only skill reference PR. Verified at head 7fefc45 (worktree git worktree add ... origin/docs/release-gate-reference-20260924 --detach), base e704b49 == origin/main. This account is not the PR author and did not modify the branch.

1. CI statusCheckRollup at head — 8/9 SUCCESS

claude-review        FAILURE   (Claude Code Review)
Check Repository     SUCCESS
Lint & Format        SUCCESS
Validate PR          SUCCESS
Type Check           SUCCESS
Scan Agent Readiness SUCCESS
Test (Node 20.x)     SUCCESS
Test (Node 22.x)     SUCCESS
Build                SUCCESS

The only red check is claude-review, and its failure is the review bot's own harness error, not a finding about this diff:

##[error]Claude result reported subtype success with is_error:true (run did not complete successfully)
##[error]Action failed with error: Claude execution failed: result is_error:true

Pre-existing and repo-wide: gh run list --workflow="Claude Code Review" shows 10 failures in the last 12 runs across unrelated branches (chore/merge-satellites, fix/agent-ready-pr71-baseline-20260922, chore/agent-native-baseline-20260924, feat/repo-infra-dogfood, and the four 2026-09-28 loop/* PRs #89-#93 — all after this PR was opened). The first green run is #78 (2026-09-25), which re-routed the bot to GLM-5.3; this PR predates that fix. No push-event runs exist for main (workflow is pull_request-only), so cross-PR reproduction is the base-ref evidence: the PR does not make the baseline worse.

2. Docs oracle (broken-reference check) — PASS

$ grep -n "release-gate" skill/agent-ready/SKILL.md
40:| Release Gate | `references/release-gate.md` | ... |
$ ls -l skill/agent-ready/references/release-gate.md
-rw-rw-r-- 1 robot robot 9062 ... release-gate.md

Target exists at the referenced relative path, same form as the other rows in the references table.

3. Narrow test suite (test/skill-references.test.ts, added by this PR) — 3/3 PASS

$ npx tsx --test test/skill-references.test.ts
# tests 3
# pass 3
# fail 0

Tests use real oracles (existsSync, exact sentence/table-row string matches), not assert-true.

4. Supersession / overlap check

None. Sibling #76 (agent-native baseline) touches only Makefile — zero file overlap with this PR. No other open PR adds a release-gate doc.

5. Independent review (grok-4.7, full diff pasted) — FAIL with 2 doc-consistency findings

  1. (major) The roll-up table (release-gate.md lines 122-124) routes critical mission fail to block but non-critical fail / required mission not run to investigate, yet nothing operationally marks individual missions as critical/non-critical/required — the pack section only says "cover critical paths only". An evidenced pack-mission failure can be read as either block or investigate.
  2. (minor) "First match wins" leaves edge results unmapped: a pass without evidence matches no row (row 3 requires evidence); a critical fail without evidence is covered only by the trailing prose sentence, not by a table row.

The reviewer confirmed the SKILL.md link is correct and the tests are real oracles.

Verdict

Checks/tests: green at head; the single failure is a pre-existing repo-wide bot-harness failure (see above), so this PR is not worse than base. The two review findings above are docs-consistency gaps for the PR author (or a follow-up PR) to resolve — they are advisory, not execution failures, and this review lane cannot rewrite a foreign branch.

@robotlearning123

Copy link
Copy Markdown
Member Author

Independent review (grok lane, execution-based; writer was a different model) at head 7fefc45.

Verdict: FIX-FIRST (1 major, 1 minor; no scope creep, no broken references)

Findings

  1. major — skill/agent-ready/references/release-gate.md:120-126 (roll-up table gap): a required mission with result pass but a missing evidence pointer matches no roll-up row: the investigate row (line 123) covers only blocked / non-critical fail / required-not-run, and the ship row (line 124) requires pass with evidence; the precedence prose (line 126) remaps only an evidence-less fail. A literal top-to-bottom matcher on {critical: true, result: pass, evidence: false} returns UNMATCHED, yet release:ready must emit one of ship/investigate/block — an undefined case risks falling through to exit 0, i.e. the false-green anti-pattern the same doc warns against. Fix: add an explicit row/rule (e.g. evidence-less pass -> investigate) so the three-way verdict is total.
  2. minor — test/skill-references.test.ts:31: the content test locks the block row and the investigate-non-blocking phrase but does not guard the evidence-less-pass case, so finding 1 could be (re)introduced without a test failure.

Checked, does not hold

  • Diff is exactly the 3 stated files; 9-area table and checker untouched; no scope creep.
  • SKILL.md:40 links references/release-gate.md; target exists in the same PR. Not on main — not superseded; no overlap with chore: agent-native baseline (AGENTS.md, make check, CI, dev env) #76 (Makefile only).
  • Makefile fence uses a real tab (make release-ready reproduced, exit 0); prettier clean; lane semantics internally consistent (deterministic lane always blocking; investigate never fails build; no auto-merge).

Machine gates

  • tsx --test test/skill-references.test.ts at this head: 3 passed / 0 failed.
  • CI at head: Test (Node 20/22), Lint & Format, Type Check, Build, Scan Agent Readiness all SUCCESS. claude-review FAILURE is a pre-existing bot-infra failure (fails on all recent branches; being re-routed in ci: route PR review bot to GLM-5.3 via z.ai #78), not caused by this diff.

Review-only comment — merge decision stays with the owner.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants