You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Forge VNext is designed to recover from ordinary process/worker/Redis restart using PostgreSQL as durable truth. That is not sufficient for backup restore / database point-in-time rollback.
If PostgreSQL is restored to an older snapshot, Forge can forget real events that happened after the backup while external Resources retain them: completed GitHub/service/email/notification mutations, Trigger occurrences, credential revocations/rotations, autonomy demotions, budget usage, Mission/Execution progress and Operation evidence. Starting normal workers against the old snapshot could therefore replay schedules, recreate already-completed external effects, resurrect stale Grants/authority, double-spend budget, or treat an already-handled incident as new.
A backup is locally consistent historical truth, not automatically current world truth.
Desired Outcome
Forge has an explicit backup/export + restore protocol that restores durable data without automatically resuming consequential autonomy. Any restored instance enters a restore quarantine before Triggers, model calls or external mutations can resume. The operator/recovery system establishes a new restore epoch, revalidates current security/credentials/resources, reconciles externally persistent side effects and establishes fresh Trigger/Resource watermarks before normal Mission execution resumes.
Ordinary process restart remains automatic under #340/#347; backup restore is a distinct disaster-recovery authority transition.
User Story
As the Forge operator,
I want to restore Forge from a backup without Forge repeating actions or resurrecting stale authority,
So that backup recovery is safer than losing the database and does not create a second incident in external systems.
Core Invariants
A restored database snapshot MUST NOT immediately resume autonomous external effects, model invocations or Trigger-created Executions.
A restored snapshot MUST NOT be assumed current merely because its internal tables are transactionally consistent.
External effects that occurred after the backup MUST NOT be replayed blindly.
Historical evidence in the backup remains immutable under its original schema/policy; restore metadata does not rewrite it as newly observed truth.
Current security/credential/autonomy state is revalidated before new consequential admission; old stored allow state never overrides a newer real revoke/rotation merely because the database rolled back.
Trigger catch-up after restore is explicit and bounded; elapsed time during the lost interval is not silently burst-replayed.
Restore quarantine is durable and fail-closed across process reboot until explicitly completed/aborted by the trusted recovery path.
Requirements
A. Backup/export contract
Provide a supported Forge backup/export procedure/manifest that records at minimum:
Forge instance identity;
backup/export identity and schema version;
database schema/migration version;
backup creation DB timestamp;
current runtime/migration/authority epoch/version;
evidence/artifact storage coverage and retention status;
protected key/reference generations needed to interpret encrypted/digested records without embedding raw secrets;
content digest/checksum of the backup/manifest where appropriate.
Backup must be transactionally consistent across the PostgreSQL authoritative state it claims to cover. If Artifact/evidence bytes live outside PostgreSQL, the backup protocol must define a consistent snapshot/manifest relationship and explicitly report missing/unbacked content.
Raw credentials/private keys are handled through the existing secure store/backup policy and are not copied into ordinary backup metadata, logs, GitHub artifacts or issue comments.
B. Restore detection and quarantine
A supported restore command/procedure creates a new restore epoch and durable quarantine state before ordinary application/worker execution is allowed.
Quarantine blocks at least:
new Trigger-created/manual autonomous Executions;
new model invocations;
new external/consequential Operations;
elevated autonomy use;
compatibility Task mutations that would resume work automatically;
catch-up scheduling/outbox publication whose world-state assumptions are not yet reconciled.
Read-only operator inspection and explicit recovery/reconciliation Operations may remain available under a narrow recovery Principal/Grant.
Do not rely on an in-memory flag or “remember to stop the worker.” Startup reads durable quarantine state before starting processing loops.
C. Restore epoch and stale-work fencing
Restore creates a monotonic current runtime/restore generation that invalidates pre-restore ephemeral ownership/admission state.
old worker/lease/admission tokens are invalid;
stale Redis/outbox deliveries from the pre-restore runtime cannot execute under the new epoch;
current continuation/Trigger/Execution admission includes the restore/runtime epoch in its fencing context;
Redis is treated as disposable/reconstructable and should normally be cleared/reinitialized or fully fenced against the restored database state according to the protocol;
old signed/current projections whose applicability depends on post-backup state are re-evaluated rather than assumed current.
A monotonic epoch stored only inside the restored snapshot cannot by itself detect an arbitrary unsignalled storage rollback. Therefore normal production support must require restoration through the explicit recovery procedure or an external/current installation marker/epoch where practical. Document that copying an old raw database over a running installation is unsupported and unsafe.
D. External uncertainty horizon
The interval from backup_created_at to restore_started_at is a lost-current-state horizon. Forge may not know which external effects/events occurred in that interval.
For every configured consequential adapter/Resource domain with active/recent Missions/Operations, classify recovery using current external truth where supported:
exact stable external Operation/idempotency/audit identity proves an effect already occurred;
current Resource state can prove the intended effect did not/does exist according to the Operation's reconciliation contract;
external state is insufficient/ambiguous -> operator/human-required block;
no relevant consequential work was possible in the horizon -> explicit no-effect evidence.
Do not invent Forge Operation provenance for an external effect that exists but cannot be tied to a backed-up Forge Operation identity. It is current external state and may be recorded as restore reconciliation evidence, not rewritten history.
E. Operation recovery
For Operations present in the backup:
terminal confirmed historical rows remain historical, subject to current Resource/security revalidation for future actions;
non-terminal rows whose exact external status cannot be established remain blocked/human-required;
no Operation is blindly retried because the local snapshot predates its possible external acknowledgement;
idempotency keys/external refs remain exact; restore does not mint a fresh key to “try again.”
For possible Operations after the backup that are absent from restored DB, Resource/domain reconciliation establishes a safe new baseline before that Resource can receive autonomous writes again.
F. Trigger/schedule/event recovery
Consume #341 semantics rather than creating a backup-specific scheduler.
For each Trigger binding establish a new restore activation/watermark:
schedule catch-up policy is explicitly skip | latest | bounded and defaults conservatively; never replay the entire outage by default;
historical schedule instants are not claimed as verified/processed unless exact preserved occurrence/Resource evidence exists;
webhook/event source resumes from a supported current external/update watermark where the protocol provides one, otherwise uses a documented restore floor and replay policy;
occurrences before the restore floor are not automatically reissued because the restored DB forgot them;
Trigger definition/enabled/current security state is revalidated before reactivation.
G. Persistent Mission/Execution recovery
Restore does not blanket resume every backed-up non-terminal Execution.
For each active Mission/Execution:
pin/verify exact Mission spec/Workflow/Resource revisions available after restore;
inspect checkpoints only as hints/references to canonical restored evidence;
inspect externally consequential Operation state and lost horizon;
decide resume | wait/reconcile | supersede/restart with new Execution | human_required | cancel through a closed trusted recovery policy;
no terminal identity is reopened;
a new Execution created because the previous current world state is unknowable has new identity and preserves the old restored Execution as historical/indeterminate evidence according to policy.
H. Current security, autonomy and credentials are revalidated
Before leaving quarantine:
current Forge/system security policy/configuration is reloaded from trusted current installation sources;
credential bindings are revalidated for existence/generation/audience/expiry where possible;
provider/adapter readiness becomes unknown/stale until current evidence exists;
explicit operator caps/revocations and current adverse-authority blocks are re-established/confirmed;
elevated [FEATURE] Add evidence-based earned autonomy policy engine #189 earned-autonomy decisions from the restored snapshot are not automatically effective. Default to a safe baseline/requalification until policy proves the exact decision/evidence is still applicable under the new restore epoch/current security state;
unknown/missing current authority state fails closed.
A restore cannot resurrect a revoked credential/Grant merely because the snapshot predates revocation.
I. Budget/cost reconciliation
The restored budget ledger may omit provider calls/Operations that incurred cost after the backup.
do not assume restored remaining = ceiling - consumed is current spend truth;
provider/account reconciliation uses available external billing/quota evidence where supported;
if current spend cannot be bounded under a hard monetary ceiling, cost-incurring autonomy remains blocked until explicit operator reconciliation/policy reset;
resetting/rebasing a budget after disaster recovery is an explicit audited operator action and does not rewrite historical usage.
Do not fabricate transition/audit events for the lost interval. The audit should honestly show a gap/horizon after the backup.
If immutable evidence/artifact bytes required to interpret backed-up records are missing/corrupt, affected decisions remain unverifiable/blocked; absence is not success.
K. Restore completion gate
Quarantine may be lifted only when a trusted deterministic restore gate proves:
database schema/current runtime migration valid;
stale workers/Redis deliveries fenced;
required Resource domains reconciled or explicitly blocked from writes;
active Missions have recovery dispositions;
Triggers have explicit activation/catch-up floors;
current credentials/security/autonomy/budget state is safe;
no unresolved side-effect uncertainty can be blindly replayed;
post-restore conformance canaries pass.
Quarantine may be lifted partially by Resource/Capability domain only if the architecture implements explicit scoped recovery fences; otherwise use one global gate initially. Do not allow a model to decide restore completion.
A supported backup has a versioned integrity manifest and explicit evidence-storage coverage.
Restoring a backup enters durable quarantine before ordinary workers/Triggers/model/external Operations can act.
Stale Redis/lease/admission state cannot execute after restore epoch change.
A confirmed external mutation performed after the backup but before restore is not repeated when current Resource truth can prove it exists.
Ambiguous lost-horizon external effects block/human-reconcile rather than blind retry.
Triggers resume from explicit restore floors/catch-up policy and do not replay the whole lost interval by default.
Non-terminal Missions/Executions are individually classified for safe recovery; terminal identities never reopen.
Restored elevated autonomy/credentials/Grants do not become current until revalidated under current security/restore epoch.
Unknown post-backup model spend remains conservatively budget-blocking rather than zero.
Audit/evidence clearly records the backup horizon and restore/reconciliation instead of inventing missing history.
Quarantine completion is deterministic, evidence-backed and fail-closed on unresolved required domains.
Reboot during an incomplete restore remains quarantined and restart-safe.
Hostile fixtures cover stale credentials, forgotten external effects, missed schedules, unknown spend, missing evidence and repeated restore attempts.
Out of Scope
Transparent active-active disaster recovery across Forge instances.
Automatic cross-region failover.
Importing arbitrary unsupported raw database snapshots and pretending rollback detection is infallible.
Reconstructing external events/effects that neither Forge nor the external system retained enough evidence to determine.
Replacing external service backup/disaster-recovery policy.
Implementation Scope
Large / trust-critical disaster-recovery slice. Expected as 4-6 small PRs plus operator recovery fixtures after generic adapters land.
Technical Notes
This issue exists because process restart and storage rollback are different failure classes. PostgreSQL is Forge's durable authority during normal operation, but restoring an older PostgreSQL snapshot moves local truth backward while the external world does not. Safe recovery therefore requires a new explicit authority epoch/quarantine and external-state reconciliation before autonomy resumes.
Parent programme: #333
Execution mode: implementation
Depends on: #342
Blocks: #343, #344
Spec references: SPEC-0002 (Mission/Execution terminal/recovery semantics), SPEC-0003 (current security/revocation), SPEC-0004 (side-effect uncertainty/reconciliation), SPEC-0008 (restart/failure conformance), SPEC-0012 (authoritative audit/evidence), SPEC-0014 (migration/evidence preservation)
Problem Statement
Forge VNext is designed to recover from ordinary process/worker/Redis restart using PostgreSQL as durable truth. That is not sufficient for backup restore / database point-in-time rollback.
If PostgreSQL is restored to an older snapshot, Forge can forget real events that happened after the backup while external Resources retain them: completed GitHub/service/email/notification mutations, Trigger occurrences, credential revocations/rotations, autonomy demotions, budget usage, Mission/Execution progress and Operation evidence. Starting normal workers against the old snapshot could therefore replay schedules, recreate already-completed external effects, resurrect stale Grants/authority, double-spend budget, or treat an already-handled incident as new.
A backup is locally consistent historical truth, not automatically current world truth.
Desired Outcome
Forge has an explicit backup/export + restore protocol that restores durable data without automatically resuming consequential autonomy. Any restored instance enters a restore quarantine before Triggers, model calls or external mutations can resume. The operator/recovery system establishes a new restore epoch, revalidates current security/credentials/resources, reconciles externally persistent side effects and establishes fresh Trigger/Resource watermarks before normal Mission execution resumes.
Ordinary process restart remains automatic under #340/#347; backup restore is a distinct disaster-recovery authority transition.
User Story
As the Forge operator,
I want to restore Forge from a backup without Forge repeating actions or resurrecting stale authority,
So that backup recovery is safer than losing the database and does not create a second incident in external systems.
Core Invariants
Requirements
A. Backup/export contract
Provide a supported Forge backup/export procedure/manifest that records at minimum:
Backup must be transactionally consistent across the PostgreSQL authoritative state it claims to cover. If Artifact/evidence bytes live outside PostgreSQL, the backup protocol must define a consistent snapshot/manifest relationship and explicitly report missing/unbacked content.
Raw credentials/private keys are handled through the existing secure store/backup policy and are not copied into ordinary backup metadata, logs, GitHub artifacts or issue comments.
B. Restore detection and quarantine
A supported restore command/procedure creates a new restore epoch and durable quarantine state before ordinary application/worker execution is allowed.
Quarantine blocks at least:
Read-only operator inspection and explicit recovery/reconciliation Operations may remain available under a narrow recovery Principal/Grant.
Do not rely on an in-memory flag or “remember to stop the worker.” Startup reads durable quarantine state before starting processing loops.
C. Restore epoch and stale-work fencing
Restore creates a monotonic current runtime/restore generation that invalidates pre-restore ephemeral ownership/admission state.
A monotonic epoch stored only inside the restored snapshot cannot by itself detect an arbitrary unsignalled storage rollback. Therefore normal production support must require restoration through the explicit recovery procedure or an external/current installation marker/epoch where practical. Document that copying an old raw database over a running installation is unsupported and unsafe.
D. External uncertainty horizon
The interval from
backup_created_attorestore_started_atis a lost-current-state horizon. Forge may not know which external effects/events occurred in that interval.For every configured consequential adapter/Resource domain with active/recent Missions/Operations, classify recovery using current external truth where supported:
Do not invent Forge Operation provenance for an external effect that exists but cannot be tied to a backed-up Forge Operation identity. It is current external state and may be recorded as restore reconciliation evidence, not rewritten history.
E. Operation recovery
For Operations present in the backup:
submission_uncertain/reconciling rows continue through [FEATURE] VNext Phase 2 — secure generic execution envelope and side-effect recovery #336/[FEATURE] VNext Phase 8 — general Resource/Capability adapter ecosystem #342 reconciliation under quarantine;For possible Operations after the backup that are absent from restored DB, Resource/domain reconciliation establishes a safe new baseline before that Resource can receive autonomous writes again.
F. Trigger/schedule/event recovery
Consume #341 semantics rather than creating a backup-specific scheduler.
For each Trigger binding establish a new restore activation/watermark:
skip | latest | boundedand defaults conservatively; never replay the entire outage by default;G. Persistent Mission/Execution recovery
Restore does not blanket resume every backed-up non-terminal Execution.
For each active Mission/Execution:
resume | wait/reconcile | supersede/restart with new Execution | human_required | cancelthrough a closed trusted recovery policy;H. Current security, autonomy and credentials are revalidated
Before leaving quarantine:
A restore cannot resurrect a revoked credential/Grant merely because the snapshot predates revocation.
I. Budget/cost reconciliation
The restored budget ledger may omit provider calls/Operations that incurred cost after the backup.
remaining = ceiling - consumedis current spend truth;J. Evidence/audit integrity after restore
Append a restore record that states:
Do not fabricate transition/audit events for the lost interval. The audit should honestly show a gap/horizon after the backup.
If immutable evidence/artifact bytes required to interpret backed-up records are missing/corrupt, affected decisions remain unverifiable/blocked; absence is not success.
K. Restore completion gate
Quarantine may be lifted only when a trusted deterministic restore gate proves:
Quarantine may be lifted partially by Resource/Capability domain only if the architecture implements explicit scoped recovery fences; otherwise use one global gate initially. Do not allow a model to decide restore completion.
L. Backup retention/security
Backup files contain sensitive orchestration/evidence state.
Implementation Sequence
Orthogonal Checkpoints
Acceptance Criteria
Out of Scope
Implementation Scope
Large / trust-critical disaster-recovery slice. Expected as 4-6 small PRs plus operator recovery fixtures after generic adapters land.
Technical Notes
This issue exists because process restart and storage rollback are different failure classes. PostgreSQL is Forge's durable authority during normal operation, but restoring an older PostgreSQL snapshot moves local truth backward while the external world does not. Safe recovery therefore requires a new explicit authority epoch/quarantine and external-state reconciliation before autonomy resumes.