You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[FEATURE] Complete verification-goal on-demand proof execution through VNext #355
The verification-goal registry/import/revision foundation from #187 is already merged on main through PRs #328-#330, but later executable proof-run work accumulated in stale monolithic PR #331 before VNext. Continuing that branch would introduce project-specific execution/scheduling machinery beside the generic Mission/Execution/Operation runtime.
Forge still needs a safe on-demand proof execution slice before Software Engineering and later scheduled verification can rely on verification goals.
Desired Outcome
A current verification-goal revision can be run on demand as one bounded VNext Execution using typed Operation bindings, explicit Resource/Grant/budget admission and #336 confinement. The run persists canonical outcome/evidence compatible with #185/#186 and safely recovers from restart/duplicate delivery. It owns no recurring scheduler and makes zero model calls by default.
User Story
As the Forge operator,
I want to run a registered “what still works” assertion on demand through the same governed runtime as other work,
So that proof evidence is trustworthy without reviving the stale project-specific runner architecture.
Requirements
Reuse the verification-goal schema/registry/import/revision code already merged on main; do not replace it.
Default proof execution uses no LLM. A verification goal is deterministic proof unless a separately admitted verifier later requires cognition.
Persist run lifecycle and canonical result at least passed | failed | inconclusive | blocked, plus [FEATURE] Normalize execution outcomes and stop reasons #185 canonical outcome, duration, evidence refs, Resource/commit fingerprint and stable failure reason.
Run contract + persistence — versioned proof-run identity/state/evidence linked to exact goal revision and generic Execution/Principal/Resource/Operation refs; immutable terminal history and rebuildable projections.
Confined deterministic runner — execute supported command/file proof operation with bounded environment/output/filesystem/process/time; default no model/provider path.
Canonical result mapper — distinguish assertion failure, runner/infrastructure failure, policy block, cancellation and inconclusive evidence; emit [FEATURE] Normalize execution outcomes and stop reasons #185 outcome without collapsing unknown state.
Evidence + reliability integration — persist deterministic evidence refs/fingerprints and feed only comparable completed evidence to [FEATURE] Add capability reliability ledger #186; last-green/first-failing projection derives from immutable runs.
Idempotency/recovery — stable request/run identity, lease fencing, duplicate request/restart/cancel tests and no double terminalization/confirmed side effects.
API/operator path — authenticated on-demand run request, current run/status/result/recovery; do not mix scheduling or reporting subsystem work.
PR331 requirement-port review — inspect stale PR feat: implement project verification goals and proof runs (#187) #331 solely for edge-case tests (redaction, output bounds, restart/idempotency, command policy) and prove every retained behavior is either implemented here or explicitly out of scope.
Primary Code Seams To Inspect First
web/lib/verification-goals/** and merged registry/import/revision tests
web/db/schema.ts / migrations for goal/run evidence
Large - expected as 3-5 small PRs: run contract/persistence; preflight+confined runner; outcome/evidence/reliability; recovery/API; final adversarial regression.
Technical Notes
New implementation branches start from current main. Treat PR #331 as reviewed historical test evidence only. After each schema/runner/recovery checkpoint, perform fresh contract, state, failure, security/privacy, regression and evidence-readiness passes before proceeding.
Parent tracking issue: #187
Parent programme: #184 / #333
Execution mode: implementation
Depends on: #336
Consumed by: #337, #356
Problem Statement
The verification-goal registry/import/revision foundation from #187 is already merged on
mainthrough PRs #328-#330, but later executable proof-run work accumulated in stale monolithic PR #331 before VNext. Continuing that branch would introduce project-specific execution/scheduling machinery beside the generic Mission/Execution/Operation runtime.Forge still needs a safe on-demand proof execution slice before Software Engineering and later scheduled verification can rely on verification goals.
Desired Outcome
A current verification-goal revision can be run on demand as one bounded VNext Execution using typed Operation bindings, explicit Resource/Grant/budget admission and #336 confinement. The run persists canonical outcome/evidence compatible with #185/#186 and safely recovers from restart/duplicate delivery. It owns no recurring scheduler and makes zero model calls by default.
User Story
As the Forge operator,
I want to run a registered “what still works” assertion on demand through the same governed runtime as other work,
So that proof evidence is trustworthy without reviving the stale project-specific runner architecture.
Requirements
main; do not replace it.passed | failed | inconclusive | blocked, plus [FEATURE] Normalize execution outcomes and stop reasons #185 canonical outcome, duration, evidence refs, Resource/commit fingerprint and stable failure reason.Implementation Sequence
Primary Code Seams To Inspect First
web/lib/verification-goals/**and merged registry/import/revision testsweb/db/schema.ts/ migrations for goal/run evidenceOrthogonal Checkpoints
Acceptance Criteria
Out of Scope
Implementation Scope
Large - expected as 3-5 small PRs: run contract/persistence; preflight+confined runner; outcome/evidence/reliability; recovery/API; final adversarial regression.
Technical Notes
New implementation branches start from current
main. Treat PR #331 as reviewed historical test evidence only. After each schema/runner/recovery checkpoint, perform fresh contract, state, failure, security/privacy, regression and evidence-readiness passes before proceeding.