You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A worker, reviewer prompt or model response cannot be allowed to declare its own implementation correct. Forge has manual review gates and execution evidence, but VNext needs a reusable independent Verification contract that separates deterministic checks, verifier evidence and authoritative Gate decisions. Without it, reliability/autonomy evidence can be circular self-assessment.
Desired Outcome
Forge can launch an independent Verification run against a completed Work Package/Execution or deterministic Operation result using fresh bounded context and no private reasoning transcript. It persists per-criterion structured findings/evidence under a separate identity, then a deterministic versioned Gate decides pass/rework/human-required from required evidence.
The verifier has no repository-write/merge authority.
User Story
As the Forge operator,
I want implementation/proof results checked independently from the worker that produced them,
So that completion and later earned autonomy rely on inspectable evidence rather than self-grading.
Requirements
A. Verification request contract
A versioned verification request pins:
Mission/Execution/Work Package identity where applicable;
producer Agent Run/Operation/result identity;
objective and acceptance criteria;
exact changed diff/Artifact/Resource versions to inspect;
explicitly scoped prior blocking findings for re-verification;
verifier budget/cognitive requirement/Grant.
Do not include the producer's hidden reasoning transcript or unbounded task history.
B. Deterministic checks first
Before any verifier model call, evaluate available deterministic evidence such as schema/Artifact integrity, lint/type/test/build/proof results, required evidence presence/freshness and Resource version match. If deterministic policy already establishes a hard failure/block, a model call is unnecessary.
C. Independent identity/context
Verifier Agent Run/Execution context is separate from producer identity. The producer cannot write the verifier result row/Artifact or mark itself independently verified. Context is freshly assembled under #335 budgets/egress and #336 read-only Resource Grants.
D. Criterion-level result
Every acceptance criterion reports one of:
passed;
failed;
inconclusive;
not_tested.
A verification run also has an overall evidence disposition such as passed | failed | inconclusive | blocked | needs_human_review; overall passed is impossible when required criteria/evidence are missing/inconclusive/not tested under policy.
E. Structured findings
Persist immutable findings containing stable finding identity/version, criterion/type/severity, affected Resource/file/surface where appropriate, evidence refs, verification method, verifier runtime/model snapshot, required remediation/follow-up, confidence only as advisory context, and re-verification lineage/status.
F. Trusted Gate
Verifier output is evidence. A deterministic Gate evaluator consumes required evidence + versioned policy and produces pass | rework | human_required | blocked. Neither verifier nor producer can bypass missing evidence by emitting PASS text.
G. Change-scoped rework
For implementation merge gates, a blocking finding must be caused, exposed or materially worsened by the current approved change and within acceptance/security criteria. Adjacent pre-existing improvements are preserved as separate follow-up issues/evidence, not injected into the current rework loop.
Re-verification receives only prior blocking findings plus current objective/evidence needed to judge remediation. It appends a new run/finding disposition; it never overwrites original findings.
H. High-risk verification
Policy may require a Security verifier/runtime class or human gate for specified risk/Capability/Resource classes. This is an additional evidence requirement, not permission for a model to authorize dangerous work.
I. Reliability integration
Only Gate-qualified independent verification outcomes become the verified evidence expected by #186/#189. Unverified worker success, missing verification, stale evidence or tampered Artifact identity does not count as a verified pass.
Implementation Sequence
Verification request/result/finding schemas — pure contracts + lineage/versioning.
Large / trust-critical - expected as 4-6 small PRs immediately after #336, in parallel with #355.
Technical Notes
The repository-wide orthogonal review skill can remain broad when the operator explicitly asks for a repo audit. This issue's merge-gate verifier should be bounded to the approved change so verification cannot become an endless adjacent-improvement generator.
Parent Epic: #184
Execution mode: implementation
VNext programme: #333
Depends on: #336
Consumed by: #337, #189, #191
Problem Statement
A worker, reviewer prompt or model response cannot be allowed to declare its own implementation correct. Forge has manual review gates and execution evidence, but VNext needs a reusable independent Verification contract that separates deterministic checks, verifier evidence and authoritative Gate decisions. Without it, reliability/autonomy evidence can be circular self-assessment.
Desired Outcome
Forge can launch an independent Verification run against a completed Work Package/Execution or deterministic Operation result using fresh bounded context and no private reasoning transcript. It persists per-criterion structured findings/evidence under a separate identity, then a deterministic versioned Gate decides pass/rework/human-required from required evidence.
The verifier has no repository-write/merge authority.
User Story
As the Forge operator,
I want implementation/proof results checked independently from the worker that produced them,
So that completion and later earned autonomy rely on inspectable evidence rather than self-grading.
Requirements
A. Verification request contract
A versioned verification request pins:
Do not include the producer's hidden reasoning transcript or unbounded task history.
B. Deterministic checks first
Before any verifier model call, evaluate available deterministic evidence such as schema/Artifact integrity, lint/type/test/build/proof results, required evidence presence/freshness and Resource version match. If deterministic policy already establishes a hard failure/block, a model call is unnecessary.
C. Independent identity/context
Verifier Agent Run/Execution context is separate from producer identity. The producer cannot write the verifier result row/Artifact or mark itself independently verified. Context is freshly assembled under #335 budgets/egress and #336 read-only Resource Grants.
D. Criterion-level result
Every acceptance criterion reports one of:
passed;failed;inconclusive;not_tested.A verification run also has an overall evidence disposition such as
passed | failed | inconclusive | blocked | needs_human_review; overallpassedis impossible when required criteria/evidence are missing/inconclusive/not tested under policy.E. Structured findings
Persist immutable findings containing stable finding identity/version, criterion/type/severity, affected Resource/file/surface where appropriate, evidence refs, verification method, verifier runtime/model snapshot, required remediation/follow-up, confidence only as advisory context, and re-verification lineage/status.
F. Trusted Gate
Verifier output is evidence. A deterministic Gate evaluator consumes required evidence + versioned policy and produces
pass | rework | human_required | blocked. Neither verifier nor producer can bypass missing evidence by emittingPASStext.G. Change-scoped rework
For implementation merge gates, a blocking finding must be caused, exposed or materially worsened by the current approved change and within acceptance/security criteria. Adjacent pre-existing improvements are preserved as separate follow-up issues/evidence, not injected into the current rework loop.
Re-verification receives only prior blocking findings plus current objective/evidence needed to judge remediation. It appends a new run/finding disposition; it never overwrites original findings.
H. High-risk verification
Policy may require a Security verifier/runtime class or human gate for specified risk/Capability/Resource classes. This is an additional evidence requirement, not permission for a model to authorize dangerous work.
I. Reliability integration
Only Gate-qualified independent verification outcomes become the verified evidence expected by #186/#189. Unverified worker success, missing verification, stale evidence or tampered Artifact identity does not count as a verified pass.
Implementation Sequence
Primary Code Seams To Inspect First
agent_runs,artifacts, approval/review gate storesOrthogonal Checkpoints
Acceptance Criteria
passed | failed | inconclusive | not_testedwith evidence/method.Out of Scope
Implementation Scope
Large / trust-critical - expected as 4-6 small PRs immediately after #336, in parallel with #355.
Technical Notes
The repository-wide orthogonal review skill can remain broad when the operator explicitly asks for a repo audit. This issue's merge-gate verifier should be bounded to the approved change so verification cannot become an endless adjacent-improvement generator.