Skip to content

[FEATURE][RELIABILITY] Add restore quarantine and external-state reconciliation for backup recovery #366

Description

@Joncallim

Parent programme: #333
Execution mode: implementation
Depends on: #342
Blocks: #343, #344
Spec references: SPEC-0002 (Mission/Execution terminal/recovery semantics), SPEC-0003 (current security/revocation), SPEC-0004 (side-effect uncertainty/reconciliation), SPEC-0008 (restart/failure conformance), SPEC-0012 (authoritative audit/evidence), SPEC-0014 (migration/evidence preservation)

Problem Statement

Forge VNext is designed to recover from ordinary process/worker/Redis restart using PostgreSQL as durable truth. That is not sufficient for backup restore / database point-in-time rollback.

If PostgreSQL is restored to an older snapshot, Forge can forget real events that happened after the backup while external Resources retain them: completed GitHub/service/email/notification mutations, Trigger occurrences, credential revocations/rotations, autonomy demotions, budget usage, Mission/Execution progress and Operation evidence. Starting normal workers against the old snapshot could therefore replay schedules, recreate already-completed external effects, resurrect stale Grants/authority, double-spend budget, or treat an already-handled incident as new.

A backup is locally consistent historical truth, not automatically current world truth.

Desired Outcome

Forge has an explicit backup/export + restore protocol that restores durable data without automatically resuming consequential autonomy. Any restored instance enters a restore quarantine before Triggers, model calls or external mutations can resume. The operator/recovery system establishes a new restore epoch, revalidates current security/credentials/resources, reconciles externally persistent side effects and establishes fresh Trigger/Resource watermarks before normal Mission execution resumes.

Ordinary process restart remains automatic under #340/#347; backup restore is a distinct disaster-recovery authority transition.

User Story

As the Forge operator,
I want to restore Forge from a backup without Forge repeating actions or resurrecting stale authority,
So that backup recovery is safer than losing the database and does not create a second incident in external systems.

Core Invariants

  • A restored database snapshot MUST NOT immediately resume autonomous external effects, model invocations or Trigger-created Executions.
  • A restored snapshot MUST NOT be assumed current merely because its internal tables are transactionally consistent.
  • External effects that occurred after the backup MUST NOT be replayed blindly.
  • Historical evidence in the backup remains immutable under its original schema/policy; restore metadata does not rewrite it as newly observed truth.
  • Current security/credential/autonomy state is revalidated before new consequential admission; old stored allow state never overrides a newer real revoke/rotation merely because the database rolled back.
  • Trigger catch-up after restore is explicit and bounded; elapsed time during the lost interval is not silently burst-replayed.
  • Restore quarantine is durable and fail-closed across process reboot until explicitly completed/aborted by the trusted recovery path.

Requirements

A. Backup/export contract

Provide a supported Forge backup/export procedure/manifest that records at minimum:

  • Forge instance identity;
  • backup/export identity and schema version;
  • database schema/migration version;
  • backup creation DB timestamp;
  • current runtime/migration/authority epoch/version;
  • evidence/artifact storage coverage and retention status;
  • protected key/reference generations needed to interpret encrypted/digested records without embedding raw secrets;
  • expected external Resource/adapters/Trigger classes requiring restore reconciliation;
  • content digest/checksum of the backup/manifest where appropriate.

Backup must be transactionally consistent across the PostgreSQL authoritative state it claims to cover. If Artifact/evidence bytes live outside PostgreSQL, the backup protocol must define a consistent snapshot/manifest relationship and explicitly report missing/unbacked content.

Raw credentials/private keys are handled through the existing secure store/backup policy and are not copied into ordinary backup metadata, logs, GitHub artifacts or issue comments.

B. Restore detection and quarantine

A supported restore command/procedure creates a new restore epoch and durable quarantine state before ordinary application/worker execution is allowed.

Quarantine blocks at least:

  • new Trigger-created/manual autonomous Executions;
  • new model invocations;
  • new external/consequential Operations;
  • elevated autonomy use;
  • compatibility Task mutations that would resume work automatically;
  • catch-up scheduling/outbox publication whose world-state assumptions are not yet reconciled.

Read-only operator inspection and explicit recovery/reconciliation Operations may remain available under a narrow recovery Principal/Grant.

Do not rely on an in-memory flag or “remember to stop the worker.” Startup reads durable quarantine state before starting processing loops.

C. Restore epoch and stale-work fencing

Restore creates a monotonic current runtime/restore generation that invalidates pre-restore ephemeral ownership/admission state.

  • old worker/lease/admission tokens are invalid;
  • stale Redis/outbox deliveries from the pre-restore runtime cannot execute under the new epoch;
  • current continuation/Trigger/Execution admission includes the restore/runtime epoch in its fencing context;
  • Redis is treated as disposable/reconstructable and should normally be cleared/reinitialized or fully fenced against the restored database state according to the protocol;
  • old signed/current projections whose applicability depends on post-backup state are re-evaluated rather than assumed current.

A monotonic epoch stored only inside the restored snapshot cannot by itself detect an arbitrary unsignalled storage rollback. Therefore normal production support must require restoration through the explicit recovery procedure or an external/current installation marker/epoch where practical. Document that copying an old raw database over a running installation is unsupported and unsafe.

D. External uncertainty horizon

The interval from backup_created_at to restore_started_at is a lost-current-state horizon. Forge may not know which external effects/events occurred in that interval.

For every configured consequential adapter/Resource domain with active/recent Missions/Operations, classify recovery using current external truth where supported:

  • exact stable external Operation/idempotency/audit identity proves an effect already occurred;
  • current Resource state can prove the intended effect did not/does exist according to the Operation's reconciliation contract;
  • external state is insufficient/ambiguous -> operator/human-required block;
  • no relevant consequential work was possible in the horizon -> explicit no-effect evidence.

Do not invent Forge Operation provenance for an external effect that exists but cannot be tied to a backed-up Forge Operation identity. It is current external state and may be recorded as restore reconciliation evidence, not rewritten history.

E. Operation recovery

For Operations present in the backup:

For possible Operations after the backup that are absent from restored DB, Resource/domain reconciliation establishes a safe new baseline before that Resource can receive autonomous writes again.

F. Trigger/schedule/event recovery

Consume #341 semantics rather than creating a backup-specific scheduler.

For each Trigger binding establish a new restore activation/watermark:

  • schedule catch-up policy is explicitly skip | latest | bounded and defaults conservatively; never replay the entire outage by default;
  • historical schedule instants are not claimed as verified/processed unless exact preserved occurrence/Resource evidence exists;
  • webhook/event source resumes from a supported current external/update watermark where the protocol provides one, otherwise uses a documented restore floor and replay policy;
  • occurrences before the restore floor are not automatically reissued because the restored DB forgot them;
  • Trigger definition/enabled/current security state is revalidated before reactivation.

G. Persistent Mission/Execution recovery

Restore does not blanket resume every backed-up non-terminal Execution.

For each active Mission/Execution:

  • pin/verify exact Mission spec/Workflow/Resource revisions available after restore;
  • inspect checkpoints only as hints/references to canonical restored evidence;
  • inspect externally consequential Operation state and lost horizon;
  • decide resume | wait/reconcile | supersede/restart with new Execution | human_required | cancel through a closed trusted recovery policy;
  • no terminal identity is reopened;
  • a new Execution created because the previous current world state is unknowable has new identity and preserves the old restored Execution as historical/indeterminate evidence according to policy.

H. Current security, autonomy and credentials are revalidated

Before leaving quarantine:

  • current Forge/system security policy/configuration is reloaded from trusted current installation sources;
  • credential bindings are revalidated for existence/generation/audience/expiry where possible;
  • provider/adapter readiness becomes unknown/stale until current evidence exists;
  • explicit operator caps/revocations and current adverse-authority blocks are re-established/confirmed;
  • elevated [FEATURE] Add evidence-based earned autonomy policy engine #189 earned-autonomy decisions from the restored snapshot are not automatically effective. Default to a safe baseline/requalification until policy proves the exact decision/evidence is still applicable under the new restore epoch/current security state;
  • unknown/missing current authority state fails closed.

A restore cannot resurrect a revoked credential/Grant merely because the snapshot predates revocation.

I. Budget/cost reconciliation

The restored budget ledger may omit provider calls/Operations that incurred cost after the backup.

  • do not assume restored remaining = ceiling - consumed is current spend truth;
  • provider/account reconciliation uses available external billing/quota evidence where supported;
  • unknown lost-horizon model spend is conservatively held/accounted according to [FEATURE] VNext Phase 1 — deterministic budget, routing, and context economics #335 hard-budget policy;
  • if current spend cannot be bounded under a hard monetary ceiling, cost-incurring autonomy remains blocked until explicit operator reconciliation/policy reset;
  • resetting/rebasing a budget after disaster recovery is an explicit audited operator action and does not rewrite historical usage.

J. Evidence/audit integrity after restore

Append a restore record that states:

  • backup identity/time/schema;
  • restore epoch/start/completion;
  • operator/recovery Principal;
  • reconciliation manifest digest;
  • Resources/Missions/Triggers/credentials/budgets reviewed;
  • unresolved uncertainty;
  • activation floors/watermarks;
  • resulting safe state.

Do not fabricate transition/audit events for the lost interval. The audit should honestly show a gap/horizon after the backup.

If immutable evidence/artifact bytes required to interpret backed-up records are missing/corrupt, affected decisions remain unverifiable/blocked; absence is not success.

K. Restore completion gate

Quarantine may be lifted only when a trusted deterministic restore gate proves:

  • database schema/current runtime migration valid;
  • stale workers/Redis deliveries fenced;
  • required Resource domains reconciled or explicitly blocked from writes;
  • active Missions have recovery dispositions;
  • Triggers have explicit activation/catch-up floors;
  • current credentials/security/autonomy/budget state is safe;
  • no unresolved side-effect uncertainty can be blindly replayed;
  • post-restore conformance canaries pass.

Quarantine may be lifted partially by Resource/Capability domain only if the architecture implements explicit scoped recovery fences; otherwise use one global gate initially. Do not allow a model to decide restore completion.

L. Backup retention/security

Backup files contain sensitive orchestration/evidence state.

  • storage access is least privilege;
  • encryption/integrity according to deployment policy;
  • retention/deletion is explicit;
  • restore tooling never logs raw secrets/protected payloads;
  • test fixtures use synthetic credentials/data only;
  • backup metadata exposed to UI/logs is sanitized/classification-aware.

Implementation Sequence

  1. Backup/restore state contract + threat model — backup manifest, restore epoch/quarantine, unsupported raw-rollback statement.
  2. Startup/write fencing — worker/Trigger/model/Operation admission blocked under quarantine; stale Redis/lease/admission tokens invalid.
  3. Restore reconciliation manifest — Resource/Mission/Trigger/credential/budget categories and operator workflow.
  4. Operation/external-resource reconciliation — [FEATURE] VNext Phase 2 — secure generic execution envelope and side-effect recovery #336/[FEATURE] VNext Phase 8 — general Resource/Capability adapter ecosystem #342 adapters, uncertain/lost-horizon handling.
  5. Trigger schedule/event floors — [FEATURE] VNext Phase 7 — Trigger/Event runtime with dedupe, causality, and zero-token idle #341 catch-up/watermark integration.
  6. Mission recovery dispositions — [FEATURE] VNext Phase 6 — persistent Missions, checkpoints, leases, and bounded autonomy #340 checkpoint/Execution semantics without reopening terminal rows.
  7. Security/autonomy/credential/budget revalidation — [FEATURE] Add evidence-based earned autonomy policy engine #189/[FEATURE] VNext Phase 1 — deterministic budget, routing, and context economics #335/[FEATURE] VNext Phase 8 — general Resource/Capability adapter ecosystem #342 current state.
  8. Completion gate + scoped/current activation — deterministic proof, post-restore canaries.
  9. Backup artifact/security/retention — supported operator procedure and sanitized evidence.
  10. Disaster-recovery release gate — populated realistic backup, external side-effect fixtures and repeated restore/reboot tests.

Orthogonal Checkpoints

  1. State rollback: backup is internally valid but older than real external world; no replay/fake current truth.
  2. Authority: old Grant/autonomy/credential state tries to regain access after current revoke/rotation.
  3. Side effects: operation occurred after backup but local DB forgot it; no duplicate.
  4. Triggers: missed schedules/webhooks across backup horizon; no burst/replay without explicit policy.
  5. Budgets: missing post-backup spend; unknown never becomes zero/free.
  6. Evidence: lost audit interval remains explicit; no synthetic history.
  7. Startup: quarantine survives reboot and blocks all ordinary autonomous processing.
  8. Storage/security: backup secrets/evidence protected and test fixtures synthetic.
  9. Reconciliation failure: inability to establish external truth leaves bounded human-required state rather than “restore successful.”
  10. Repeatability: restore procedure is deterministic/idempotent enough to retry after failure without progressively widening authority.

Acceptance Criteria

  • [FEATURE] VNext Phase 8 — general Resource/Capability adapter ecosystem #342 and its transitive Mission/Trigger/Operation foundations are closed before implementation starts.
  • A supported backup has a versioned integrity manifest and explicit evidence-storage coverage.
  • Restoring a backup enters durable quarantine before ordinary workers/Triggers/model/external Operations can act.
  • Stale Redis/lease/admission state cannot execute after restore epoch change.
  • A confirmed external mutation performed after the backup but before restore is not repeated when current Resource truth can prove it exists.
  • Ambiguous lost-horizon external effects block/human-reconcile rather than blind retry.
  • Triggers resume from explicit restore floors/catch-up policy and do not replay the whole lost interval by default.
  • Non-terminal Missions/Executions are individually classified for safe recovery; terminal identities never reopen.
  • Restored elevated autonomy/credentials/Grants do not become current until revalidated under current security/restore epoch.
  • Unknown post-backup model spend remains conservatively budget-blocking rather than zero.
  • Audit/evidence clearly records the backup horizon and restore/reconciliation instead of inventing missing history.
  • Quarantine completion is deterministic, evidence-backed and fail-closed on unresolved required domains.
  • Reboot during an incomplete restore remains quarantined and restart-safe.
  • Hostile fixtures cover stale credentials, forgotten external effects, missed schedules, unknown spend, missing evidence and repeated restore attempts.

Out of Scope

  • Transparent active-active disaster recovery across Forge instances.
  • Automatic cross-region failover.
  • Importing arbitrary unsupported raw database snapshots and pretending rollback detection is infallible.
  • Reconstructing external events/effects that neither Forge nor the external system retained enough evidence to determine.
  • Replacing external service backup/disaster-recovery policy.

Implementation Scope

Large / trust-critical disaster-recovery slice. Expected as 4-6 small PRs plus operator recovery fixtures after generic adapters land.

Technical Notes

This issue exists because process restart and storage rollback are different failure classes. PostgreSQL is Forge's durable authority during normal operation, but restoring an older PostgreSQL snapshot moves local truth backward while the external world does not. Safe recovery therefore requires a new explicit authority epoch/quarantine and external-state reconciliation before autonomy resumes.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    dependency-blockedREADINESS PROJECTION — Issue is blocked by unresolved dependencies. This label is a cache.enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions