Skip to content

[FEATURE] Add independent Verification Workforce execution #188

Description

@Joncallim

Parent Epic: #184
Execution mode: implementation
VNext programme: #333
Depends on: #336
Consumed by: #337, #189, #191

Problem Statement

A worker, reviewer prompt or model response cannot be allowed to declare its own implementation correct. Forge has manual review gates and execution evidence, but VNext needs a reusable independent Verification contract that separates deterministic checks, verifier evidence and authoritative Gate decisions. Without it, reliability/autonomy evidence can be circular self-assessment.

Desired Outcome

Forge can launch an independent Verification run against a completed Work Package/Execution or deterministic Operation result using fresh bounded context and no private reasoning transcript. It persists per-criterion structured findings/evidence under a separate identity, then a deterministic versioned Gate decides pass/rework/human-required from required evidence.

The verifier has no repository-write/merge authority.

User Story

As the Forge operator,
I want implementation/proof results checked independently from the worker that produced them,
So that completion and later earned autonomy rely on inspectable evidence rather than self-grading.

Requirements

A. Verification request contract

A versioned verification request pins:

  • Mission/Execution/Work Package identity where applicable;
  • producer Agent Run/Operation/result identity;
  • objective and acceptance criteria;
  • exact changed diff/Artifact/Resource versions to inspect;
  • deterministic validation evidence already produced;
  • required verification methods/policies;
  • relevant repository/Workforce rules;
  • risk classification;
  • explicitly scoped prior blocking findings for re-verification;
  • verifier budget/cognitive requirement/Grant.

Do not include the producer's hidden reasoning transcript or unbounded task history.

B. Deterministic checks first

Before any verifier model call, evaluate available deterministic evidence such as schema/Artifact integrity, lint/type/test/build/proof results, required evidence presence/freshness and Resource version match. If deterministic policy already establishes a hard failure/block, a model call is unnecessary.

C. Independent identity/context

Verifier Agent Run/Execution context is separate from producer identity. The producer cannot write the verifier result row/Artifact or mark itself independently verified. Context is freshly assembled under #335 budgets/egress and #336 read-only Resource Grants.

D. Criterion-level result

Every acceptance criterion reports one of:

  • passed;
  • failed;
  • inconclusive;
  • not_tested.

A verification run also has an overall evidence disposition such as passed | failed | inconclusive | blocked | needs_human_review; overall passed is impossible when required criteria/evidence are missing/inconclusive/not tested under policy.

E. Structured findings

Persist immutable findings containing stable finding identity/version, criterion/type/severity, affected Resource/file/surface where appropriate, evidence refs, verification method, verifier runtime/model snapshot, required remediation/follow-up, confidence only as advisory context, and re-verification lineage/status.

F. Trusted Gate

Verifier output is evidence. A deterministic Gate evaluator consumes required evidence + versioned policy and produces pass | rework | human_required | blocked. Neither verifier nor producer can bypass missing evidence by emitting PASS text.

G. Change-scoped rework

For implementation merge gates, a blocking finding must be caused, exposed or materially worsened by the current approved change and within acceptance/security criteria. Adjacent pre-existing improvements are preserved as separate follow-up issues/evidence, not injected into the current rework loop.

Re-verification receives only prior blocking findings plus current objective/evidence needed to judge remediation. It appends a new run/finding disposition; it never overwrites original findings.

H. High-risk verification

Policy may require a Security verifier/runtime class or human gate for specified risk/Capability/Resource classes. This is an additional evidence requirement, not permission for a model to authorize dangerous work.

I. Reliability integration

Only Gate-qualified independent verification outcomes become the verified evidence expected by #186/#189. Unverified worker success, missing verification, stale evidence or tampered Artifact identity does not count as a verified pass.

Implementation Sequence

  1. Verification request/result/finding schemas — pure contracts + lineage/versioning.
  2. Deterministic evidence preflight — missing/stale/tampered evidence fail-closed matrix.
  3. Separate verifier Agent Run path — fresh bounded context, read-only Grant, [FEATURE] VNext Phase 1 — deterministic budget, routing, and context economics #335 broker.
  4. Criterion parser/Artifact contract — structured results/findings; no prose-only authority.
  5. Trusted Gate evaluator — deterministic pass/rework/human/block policy.
  6. Re-verification lineage — immutable original finding + scoped rework evidence.
  7. Change-scope policy — port reviewed stale-branch test cases onto current VNext contracts.
  8. High-risk Security/human requirement — policy fixtures, no authority widening.
  9. [FEATURE] Add capability reliability ledger #186 integration — only trusted independently verified evidence counts.
  10. Standalone verification fixtures — verify one synthetic repository-change result and one generic deterministic Operation/proof result using current [FEATURE] VNext Phase 2 — secure generic execution envelope and side-effect recovery #336 contracts. Integration with the actual [FEATURE] Complete verification-goal on-demand proof execution through VNext #355 verification-goal runner happens later in [FEATURE] VNext Phase 3 — complete Software Engineering as the first generic Workforce #337, so [FEATURE] Add independent Verification Workforce execution #188 and [FEATURE] Complete verification-goal on-demand proof execution through VNext #355 remain parallel rather than cyclic.

Primary Code Seams To Inspect First

Orthogonal Checkpoints

  1. Independence: producer tries to self-verify, shared context/state leakage, same identity mutation.
  2. Evidence completeness: missing/stale/tampered diff/tests/Artifact/criterion, false pass attempts.
  3. Scope: unrelated pre-existing bug, malicious broad reviewer, repeated rework scope expansion.
  4. Security: hostile changed/source text, verifier tool request, high-risk policy bypass.
  5. Lineage: re-verification, finding overwrite/reuse, Resource version drift.
  6. Budget/failure: verifier provider failure, timeout, budget exhaustion, deterministic failure avoiding unnecessary model call.
  7. Reliability: only trusted Gate result flows to [FEATURE] Add capability reliability ledger #186/[FEATURE] Add evidence-based earned autonomy policy engine #189.
  8. Dependency isolation: implementation/tests must not import or require future [FEATURE] Complete verification-goal on-demand proof execution through VNext #355 code; [FEATURE] VNext Phase 3 — complete Software Engineering as the first generic Workforce #337 is the explicit integration point.

Acceptance Criteria

Out of Scope

Implementation Scope

Large / trust-critical - expected as 4-6 small PRs immediately after #336, in parallel with #355.

Technical Notes

The repository-wide orthogonal review skill can remain broad when the operator explicitly asks for a repo audit. This issue's merge-gate verifier should be bounded to the approved change so verification cannot become an endless adjacent-improvement generator.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    dependency-blockedREADINESS PROJECTION — Issue is blocked by unresolved dependencies. This label is a cache.enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions