Skip to content

[FEATURE] VNext Phase 6 — persistent Missions, checkpoints, leases, and bounded autonomy #340

Description

@Joncallim

Parent: #333
Execution mode: implementation
Depends on: #339, #189
Spec references: SPEC-0002 (Mission/Execution lifecycle, outcomes), SPEC-0003 (Grants, authority), SPEC-0008 (conformance), SPEC-0014 (migration)

Problem Statement

Forge's current Task model is primarily a finite request/attempt workflow. Long-lived responsibilities need durable intent that survives process restarts and can wait quiescently without preserving a model conversation. Without a first-class Mission/Execution state machine, recurring work would be forced into cron prompts, in-memory parent agents or duplicate task records, recreating the failure modes VNext is designed to remove.

Persistent autonomy also needs the evidence-based ceilings from #189; a long-lived Mission must never be able to raise its own budget/authority because it has been running for a long time.

Desired Outcome

Forge can maintain a durable Mission across multiple bounded Executions, checkpoints and idle periods. All status, budgets, Grants, pinned Workforce revisions, blockers and completion state are reconstructable from PostgreSQL/evidence. Stale workers are fenced, cancellation/revocation is durable, and an idle healthy Mission consumes zero model tokens.

User Story

As the Forge operator,
I want Forge to own ongoing responsibilities without a permanent parent-agent conversation,
So that long-running Workforces can survive restarts, remain bounded by policy and wake only when there is real work.

Requirements

A. Durable Mission model

Persist versioned Mission intent with at least:

  • stable Mission identity/version;
  • objective and operator constraints;
  • pinned Workforce/package/Workflow revision;
  • Resource bindings/classification snapshots or versioned refs;
  • maximum Grant/authority envelope;
  • [FEATURE] Add evidence-based earned autonomy policy engine #189 autonomy decision/ceiling reference;
  • budget policy/ceiling;
  • verification/completion policy;
  • deadline/stop conditions;
  • current lifecycle state per SPEC-0002 R3 (draft, active, waiting, paused, terminal) and separate Mission outcome per SPEC-0002 R5a (succeeded, failed, cancelled); Execution outcome per SPEC-0002 R5 (succeeded, failed, cancelled, blocked, indeterminate);
  • created/updated/paused/cancelled/completed timestamps and actor/evidence.

Lifecycle vs outcome: Per SPEC-0002 R2, lifecycle state tracks progression (draft→active→waiting→paused→terminal); outcome records the terminal result. These are separate fields. The issue's earlier language referring to "completed, cancelled/revoked, terminal failure" as Mission states is superseded by the canonical SPEC-0002 lifecycle — those are outcome values, not lifecycle states. Implementation MUST use the canonical lifecycle/outcome separation.

A Mission is not a transcript and does not store hidden chain-of-thought as durable runtime state.

B. Execution model

An Execution is one bounded attempt/cycle under a Mission. It pins effective inputs, policy, package/workflow, Resource versions, Grant, budget reservation and causal trigger/manual source. Every Work Package/Agent Run/Operation is attributable to one Execution.

Use explicit states and transition guards; no free-form status strings at behavior boundaries.

Execution lifecycle per SPEC-0002 R4: created, admitted, queued, leased, running, waiting, terminal. Execution outcome per SPEC-0002 R5: succeeded, failed, cancelled, blocked, indeterminate (SPEC-0002 R5).

C. Checkpoint model

An Execution can persist lightweight bounded checkpoints.

Checkpoints contain only validated bounded state needed to resume: completed package/output Artifact refs, pending graph nodes, Gate decisions, confirmed/uncertain Operation state, budget settlement and deterministic next action. They do not require replaying entire prior prompts/conversations.

D. Lease model

A leased Execution fences stale workers using database-clock-authoritative leases (SPEC-0002 R9). Expired leases abort the worker; the Execution re-queues or transitions to blocked/indeterminate depending on side-effect evidence.

E. Cancellation and revocation

Operator cancellation sets Mission lifecycle to terminal with outcome cancelled (SPEC-0002 R3/R5a). Current security revocation (SPEC-0003 R9) must block new Operations regardless of Mission lifecycle state.

F. Zero-token idle

A persistent Mission in waiting state consumes zero model tokens (SPEC-0002 R11). Trigger occurrences may transition waiting→active (SPEC-0009 R7-R8).

G. Finite vs recurring distinction

A recurring responsibility can remain active after a successful Execution. A finite Mission reaches terminal lifecycle only through its trusted completion Gate/policy.

H. Escalation

Blockers, indeterminate outcomes, and hard invariant violations produce operator-visible status and recovery actions.

Orthogonal Checkpoints

  1. State machine: invalid transitions, duplicate terminalization, version drift, finite vs recurring semantics.
  2. Lease/fencing: clock authority, stale worker, concurrent lease.
  3. Cancellation vs revocation: cancellation ends the Mission; revocation blocks new authority regardless of lifecycle.
  4. Evidence reconstruction: Mission status/blockers are reconstructable from PostgreSQL/evidence without Redis/in-memory state.
  5. Idempotent restart: checkpoint replay does not duplicate side effects.
  6. Budget continuity: budget reservation survives restart and is reconciled correctly.

Acceptance Criteria

  • Mission lifecycle follows SPEC-0002 R3 (draft→active→waiting→paused→terminal).
  • Mission outcome is separate from lifecycle (SPEC-0002 R5a).
  • Execution lifecycle follows SPEC-0002 R4.
  • Leases use database clock authority.
  • Cancellation and revocation are independently enforceable.
  • A waiting Mission consumes zero model tokens.
  • Mission state is reconstructable from PostgreSQL alone.
  • Checkpoint replay does not duplicate Operations.
  • Finite Missions terminalize deterministically.
  • Recurring Missions survive restart and continue schedule.

Out of Scope

Side-effect reconciliation dependency: #340 does NOT implement side-effect reconciliation — that belongs to #336 (Phase 2). However, #340's checkpoint/restart recovery MUST consume #336's reconciliation semantics: on restart, the system MUST reconcile any submission_uncertain Operations before proceeding. The generic reconciliation mechanism is defined by #336/SPEC-0004; #340 proves it works across restart boundaries. The Infrastructure Ops Workforce (#343) later proves it in a persistent production-like scenario.

Implementation Scope

Very Large - durable lifecycle/persistence/fencing/runtime work, expected as 6-9 small PRs.

Technical Notes

Prefer database-derived current state plus append-only/immutable evidence over a giant event-sourced rewrite unless the existing schema truly cannot represent the required lifecycle. Keep the first Mission APIs intentionally small; Phase 7 owns wakeup sources.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    dependency-blockedREADINESS PROJECTION — Issue is blocked by unresolved dependencies. This label is a cache.enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions