Skip to content

[FEATURE] VNext Phase 9 — Infrastructure Ops Workforce and persistent side-effect proof #343

Description

@Joncallim

Parent: #333
Execution mode: implementation
Depends on: #342, #190, #366
Uses: #189 autonomy ceilings, #340 Missions, #341 Triggers
Spec references: SPEC-0004 (Operation side-effect lifecycle — branching state graph, submission_uncertain, reconciliation), SPEC-0012 (audit evidence for recovery), SPEC-0008 (conformance, failure injection)

Problem Statement

Forge is not proven as an ongoing agentic runtime until it can own a persistent responsibility that observes real state, stays silent/cheap when healthy, performs a bounded reversible side effect when policy allows, verifies the result, survives restart and stops/escalates safely on uncertainty or repeated failure.

A coding/research success is insufficient because those are mostly finite workflows. Infrastructure Ops provides a technically verifiable persistent side-effect proof before higher-risk Personal Ops automation.

Desired Outcome

An official Infrastructure Ops Workforce maintains a bounded set of configured service Resources under an explicit Mission, Trigger, Grant, autonomy ceiling and budget. Deterministic health checks run continuously/periodically with zero model calls while unchanged; supported reversible remediations execute only as typed Operations and are independently verified. Failures, flapping and uncertainty stop/escalate rather than looping.

User Story

As the Forge operator,
I want Forge to maintain a narrowly scoped service-health responsibility safely and quietly,
So that I can trust persistent event-driven automation before moving additional operational workflows away from Hermes.

Requirements

A. Official Workforce package

Ship Infrastructure Ops as a normal #338 package. It declares the Workflow, roles, requested service-read/restart/notification Capabilities, budget/cognitive defaults and verification policy; it carries no host credentials or executable hooks.

B. Service Resource binding

Bind an explicit allowlisted set of service Resources through #342. Each Resource specifies supported health/read and optional reversible action Capabilities. Do not expose broad host/SSH/shell authority.

Initial real or fixture-backed Operations should be deliberately narrow, e.g.:

  • service status/read health;
  • restart one named scoped service;
  • read bounded recent structured status/log excerpt if policy allows;
  • operator notification.

Use a supported adapter/runtime appropriate to the host; arbitrary model-authored commands are prohibited.

C. Persistent Mission

Create one recurring Mission with:

Mission remains quiescent between occurrences and retains no permanent model conversation.

D. Trigger and deterministic health

Use #341 schedule/event Trigger occurrences. Deterministic health adapter executes first. If state is healthy/unchanged, record compact evidence/update and terminate with zero model calls.

Only a meaningful state transition/failure creates an incident Execution requiring additional diagnosis/remediation.

E. Deterministic-first diagnosis

Use stable fault classification where evidence is sufficient (service down, known unhealthy status, transient probe failure, stale data, flapping). Invoke an economy/standard technical Agent Run only when the deterministic evidence is insufficient for the next bounded decision; frontier escalation follows #335 policy only.

Log/status text is untrusted data and cannot request tools or change policy.

F. Remediation admission

Before every action re-check Mission state, current #189 autonomy ceiling, Grant, Resource scope/version, cooldown/retry/budget and side-effect state.

The first autonomous action family must be reversible/bounded. Anything outside the current ceiling requires explicit operator Gate/approval.

G. Side-effect recovery

Use #336/#342 idempotency/reconciliation lifecycle. A restart/reboot/API call with uncertain submission is reconciled by observing service state and operation identity; never blindly reissue.

H. Independent verification

After a remediation, run deterministic verification until a bounded terminal decision:

  • recovered/healthy;
  • still unhealthy;
  • observation inconclusive/stale;
  • remediation submission uncertain;
  • policy/Grant/budget blocked.

Optional model review does not self-authorize completion. Trusted Gate policy decides whether incident returns to quiescent, retries one allowed step, escalates or revokes autonomy.

I. Flapping/repeated failure

Detect repeated state changes/failed remediations deterministically. Enforce cooldown/max attempts and stop model/token burn. High-severity repeated failure creates/updates a #190 Sentinel finding and can trigger #189 demotion/revocation before any additional autonomous action.

J. Notification

Notify only for actionable transitions: remediation performed/failed, repeated/flapping failure, policy block, budget exhaustion or human decision required. Healthy unchanged checks remain silent. Notifications are typed adapter Operations with dedupe/idempotency evidence.

K. Evidence

Persist compact incident/health fingerprints, Operations, verification results, routing/budget receipts, autonomy decision refs and last-known validated state. Do not store unbounded raw logs or conversations as Mission memory.

Implementation Sequence

  1. Infrastructure Ops package + fixture service adapter — no real host action first.
  2. Persistent health Mission + schedule Trigger — healthy fixture observation proves zero-token idle.
  3. Deterministic detector/fingerprint layer — down/stale/flapping states; integrate [FEATURE] Add Project Sentinel detection and escalation flow #190 finding contract.
  4. Bounded service restart Operation — explicit Resource/Grant, [FEATURE] VNext Phase 2 — secure generic execution envelope and side-effect recovery #336 confinement and idempotency.
  5. Post-action deterministic verification + Gate — recovered/failure/inconclusive.
  6. Fault/retry/cooldown policy — no infinite loops; [FEATURE] Add evidence-based earned autonomy policy engine #189 revocation path.
  7. Provider escalation lane — model used only for fixture diagnosis requiring cognition, fully budgeted.
  8. Notification Operation — dedupe/actionable-only behavior.
  9. Real supported service pilot — one narrowly scoped Resource, explicit operator Grant and rollback.
  10. Restart/recovery/fault-injection release gate — Forge process restart during incident plus ambiguous action fixture.

Primary Code Seams To Inspect First

Orthogonal Checkpoints

  1. Zero-token idle: extended healthy window, no hidden model/provider health/search calls.
  2. Authority: wrong service Resource, stale Grant, demotion/revoke, package prompt requesting extra action.
  3. Side effects: duplicate/replay, ambiguous restart submission, Forge crash, service already recovered.
  4. Fault policy: flapping, false alarm, stale health, slow recovery, retry exhaustion/cooldown.
  5. Injection/secrets: hostile service/log text, credential sentinel, broad host/network attempts.
  6. Verification/notification: false success, notification storms, dedupe and operator explanation.
  7. Persistence: process/Redis restart and Mission recovery without duplicate confirmed action.

Acceptance Criteria

  • A meaningful healthy observation window runs deterministic checks with zero LLM calls/tokens.
  • A known injected service fault is detected deterministically and creates one incident/finding flow.
  • An approved bounded remediation executes as a typed Operation under explicit Principal + Grant + Resource scope.
  • Post-remediation state is independently/deterministically verified before returning to healthy.
  • Failed/inconclusive remediation obeys bounded retry/cooldown/escalation instead of looping.
  • Restarting Forge during an incident does not duplicate a confirmed remediation.
  • An uncertain external/service submission enters reconciliation rather than blind retry.
  • Service-control credentials/broad host authority are not exposed to the model process.
  • Hostile log/status content cannot request Capabilities or alter policy/routing.
  • [FEATURE] Add Project Sentinel detection and escalation flow #190 high/critical findings can trigger [FEATURE] Add evidence-based earned autonomy policy engine #189 autonomy review/revocation before further action.
  • Healthy unchanged checks do not notify; actionable state changes do, with dedupe evidence.
  • Cost/token/retry evidence clearly separates deterministic idle monitoring from reasoning-active incidents.
  • A supported real-service pilot and fault-injection fixture both pass the release gate.

Out of Scope

  • General autonomous system administration.
  • Arbitrary shell/SSH commands authored by a model.
  • Unlimited unattended destructive remediation.
  • Personal email/calendar automation as the proof workload.
  • Model ensembles.

Implementation Scope

Very Large / operationally sensitive - reference Workforce plus bounded service/notification adapters and persistent release proof, expected as 6-9 small PRs.

Technical Notes

Keep the real pilot narrower than the generic architecture. Prove one safe reversible action exhaustively before adding more service-control Operations.

Activity

  1. github-actions commented on Sep 2, 2026

    @github-actions

    FORGE issue validation

    This issue is not semantically dispatchable.

    Reason Detail
    queue.issue_dependency_open A declared dependency is still open. Dependency: #342.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #190.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #366.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #341.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #340.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #339.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #338.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #337.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #188.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #336.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #334.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #335.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #346.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #353.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #355.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #189.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #347.
    queue.issue_dependency_open A declared dependency is still open. Dependency: #356.

    Readiness labels are projections, not authority. Command, dispatch, and handoff always re-resolve current semantic truth.

  2. added
    needs-clarificationREADINESS PROJECTION — Issue is missing required structure or decisions. Author correction needed.
    ready-for-agentREADINESS PROJECTION — Issue is semantically ready. This label is a cache, not authority.
    and removed
    needs-clarificationREADINESS PROJECTION — Issue is missing required structure or decisions. Author correction needed.
    on Sep 2, 2026
  3. removed
    ready-for-agentREADINESS PROJECTION — Issue is semantically ready. This label is a cache, not authority.
    on Sep 3, 2026
  4. added
    ready-for-agentREADINESS PROJECTION — Issue is semantically ready. This label is a cache, not authority.
    on Sep 4, 2026
  5. removed
    ready-for-agentREADINESS PROJECTION — Issue is semantically ready. This label is a cache, not authority.
    on Sep 5, 2026
  6. added
    dependency-blockedREADINESS PROJECTION — Issue is blocked by unresolved dependencies. This label is a cache.
    on Sep 8, 2026
  7. Joncallim commented on Sep 9, 2026

    @Joncallim
    OwnerAuthor

    Hostile Infrastructure Ops addendum — incident identity, pre-action revalidation, and restart reconciliation without fabricated causality

    A fresh persistent-side-effect pass found several correctness traps that matter even for the deliberately narrow first remediation family.

    1. Do not assume restart service is idempotent

    Repeated restarts can cause material extra disruption even when the final state is the same. The initial restart Operation must declare its real retry class based on the selected service manager/adapter contract.

    • if the external manager supports a stable request/idempotency identity and can confirm it, use idempotent_with_key;
    • if Forge can determine whether this exact requested restart happened from service-manager job/audit/generation evidence, use reconcile_before_retry;
    • if neither is available, treat restart as at_most_once/human-required after possible submission;
    • never mark it replay-safe merely because calling restart twice often ends with a running service.

    Disable hidden SDK/CLI retries per #342. Every external submission is owned by the Operation journal.

    2. “Service is healthy now” is not automatically proof that Forge's restart succeeded

    The service may have:

    • self-recovered;
    • been restarted by a human/another system;
    • failed over to another instance;
    • become reachable despite Forge's ambiguous request never arriving.

    Keep two evidence questions separate:

    1. incident objective state — is the service currently healthy enough to stop remediation/escalate differently?
    2. Operation outcome — did this exact restart Operation execute/confirm according to its declared confirmation policy?

    A healthy post-check can stop Forge from issuing another restart even while the prior Operation remains submission_uncertain/indeterminate. Do not forge reconciled_success unless evidence links the state change to the exact Operation according to the adapter contract.

    This distinction is essential for reliability/autonomy evidence later.

    3. Use Sentinel finding identity as the incident/detection truth; do not invent a parallel Infra incident database

    #343 depends on #190. Reuse its finding/observation lifecycle for “this Resource is currently unhealthy/flapping/critical” where semantically applicable.

    • Trigger occurrences are observations/wakeups;
    • Sentinel finding is the durable detection/incident-like identity and severity/escalation state;
    • Mission/Execution owns the bounded diagnosis/remediation attempt;
    • Operation owns each side effect;
    • verification/Gate owns recovery evidence.

    Do not add a second infra_incidents state machine that can disagree with #190 findings. If #190's finding contract lacks a needed field, extend it generically or add a narrow linked remediation record rather than duplicate current severity/resolution truth.

    4. Multiple health occurrences for one ongoing condition coalesce into one active finding/remediation policy

    A schedule firing every minute during a 20-minute outage must not create 20 independent restart Executions.

    Define deterministic incident/finding correlation keyed by the exact detector + logical service Resource + problem class/version according to #190 policy.

    • repeated same-state observations update/append observation evidence to the active finding;
    • an already-running remediation Execution suppresses/coalesces duplicate action requests unless policy explicitly allows another after cooldown;
    • a materially different fault/problem may create a new finding;
    • recovery closes/resolves the finding only after required verification policy;
    • recurrence after resolution is a new occurrence/episode with lineage, not reopening/overwriting old history invisibly.

    5. Revalidate current state immediately before every remediation submission

    Diagnosis evidence can become stale while a model/Gate waits.

    Before external restart submission:

    1. re-check Mission/Trigger/current autonomy/Grant/budget/cooldown as already required;
    2. take a fresh bounded deterministic health observation of the exact Resource or validate the declared freshness horizon;
    3. verify the fault/remediation precondition still holds;
    4. bind that observation/version into Operation admission;
    5. if service recovered, record remediation_not_needed/superseded evidence and do not restart a healthy service.

    This pre-action observation is not permission to bypass the original finding history; it is a stale-action guard.

    6. Service Resource identity must be canonical and non-injectable

    For the first real pilot, adapter-specific Resource binding identifies one exact logical service (e.g. exact systemd unit/container/service id) independently of display text/model suggestions.

    • no wildcard/unit glob;
    • no user/model-authored shell fragment;
    • normalize/validate service identifiers through the adapter schema;
    • symlink/alias/templated-unit/container-name semantics are resolved under reviewed policy;
    • Resource binding declares host/manager/account identity so service X on host A is not confused with host B;
    • a restart action cannot be widened to dependent/all services because the model says the dependency “probably needs it.”

    7. Model diagnosis proposes evidence/advisory classifications, never an arbitrary command or Capability

    When cognition is used, its structured output is bounded to an allowlisted diagnosis/remediation-candidate schema.

    Example:

    observed fault class
    hypothesis/evidence refs
    recommended candidate Operation id from allowed set | none
    confidence/advisory rationale
    

    The deterministic policy/Gate then decides whether the candidate Operation is permitted and appropriate under current evidence/authority.

    A model cannot emit systemctl ..., SSH commands, URLs, credentials or a new Capability that becomes executable merely because it is in the diagnosis.

    8. Health detector semantics need hysteresis/freshness, not one-sample automation

    Define deterministic policy per detector:

    • observation schema/fingerprint;
    • required freshness;
    • healthy/unhealthy/unknown/stale classification;
    • consecutive/sample-window thresholds where appropriate;
    • recovery threshold/hysteresis;
    • flapping definition/window;
    • cooldown/max-remediation counts;
    • detector version.

    One transient timeout must not become “service down” unless the policy explicitly says one failure is sufficient. Conversely, missing/stale observation is not healthy.

    All counters/windows use durable DB evidence/time so restart cannot reset flapping/retry history.

    9. Flapping cooldown and remediation budget survive process/reboot

    Persist authoritative attempt/cooldown state as projections from finding/Operation/evidence rather than local timers only.

    • process restart cannot reset max-restarts-per-window;
    • worker clock changes cannot shorten cooldown;
    • two workers/triggers cannot concurrently pass the same restart-rate limit;
    • cooldown admission uses DB time/CAS;
    • manual operator remediation, when observable, may affect state/freshness but does not silently count as Forge's autonomous successful Operation.

    10. Post-remediation verification must be independent of the write acknowledgement

    The adapter response “restart accepted” is not the health verification.

    Run a separately identified read/health proof against the exact Resource after policy-defined settling semantics. The verification result is one of healthy / unhealthy / stale-inconclusive / inaccessible, with exact evidence.

    Where possible, use evidence independent from the mutation response (e.g. fresh health/status generation or service-manager state) and never let model prose alone close the finding.

    If verification cannot establish health, Gate stays retry/human/blocked according to bounds. Unknown is not recovered.

    11. Dependency/blast-radius policy is explicit for the real pilot

    Even restarting one named service can disrupt dependents.

    The Operation definition/pilot review should document known effect/risk scope and any required preconditions, e.g.:

    • service is independently restartable under this proof;
    • restart does not intentionally restart the host/all services;
    • no critical dependent Resource is outside the granted risk envelope;
    • expected interruption/timeout bounds;
    • compensation/recovery options.

    Do not build a generic dependency graph in #343, but do not select a “reversible” pilot whose real blast radius is unbounded or not understood.

    12. Notification is independent side-effect state

    A notification failure does not make the restart/health action fail and must never cause the Mission to rerun remediation blindly.

    • notification Operation has its own idempotency/retry/reconciliation identity;
    • content classification/egress is admitted separately and minimizes logs/secrets;
    • retries only the notification Operation according to its adapter policy;
    • healthy state/remediation outcome remains whatever the underlying evidence says;
    • duplicate incident observations coalesce notification policy so storms are bounded.

    13. Current autonomy demotion/revocation wins races with queued remediation

    When #190/#189 demotes or revokes autonomy:

    • outstanding action admission tokens under the higher ceiling become invalid for new submission;
    • queued/diagnosed candidate restart rechecks current autonomy immediately before adapter submission;
    • stale worker cannot act because it passed the check earlier;
    • already-submitted Operation follows reconciliation; revocation cannot pretend the external effect did not happen;
    • no model/Workforce may restore its own autonomy after the incident.

    14. Real pilot needs an operator-visible safe-stop state

    Define the state reached when Forge cannot safely continue automatically, including:

    • Operation submission uncertain;
    • repeated failure/cooldown exhausted;
    • conflicting/stale health evidence;
    • Grant/autonomy revoked;
    • verification inconclusive;
    • adapter/credential unavailable.

    Safe stop is healthy quiescence only if health is actually proven. Otherwise Mission/finding remains attention/human-required without token-burning loops.

    15. Hostile release fixture additions

    Add:

    • restart request times out after possible submission; service becomes healthy independently -> no second restart, Operation remains uncertain until exact reconciliation/human policy;
    • service self-recovers between diagnosis and action -> pre-action check suppresses restart;
    • human restarts service concurrently -> Forge does not claim that external action as its Operation;
    • 20 repeated outage Trigger occurrences -> one active finding/remediation flow under policy;
    • transient one-sample timeout below unhealthy threshold -> no restart;
    • flapping persists across Forge reboot -> cooldown/restart count preserved;
    • two workers race restart cooldown -> one admitted;
    • model outputs arbitrary shell/restart-all -> cannot become Operation;
    • exact service name on wrong host/tenant -> Resource binding denies;
    • restart accepted but health verification stale -> not recovered;
    • autonomy revoked after diagnosis before submit -> restart denied;
    • notification fails after successful recovery -> only notification retried, no remediation repeat;
    • real pilot proves bounded blast radius/rollback and no host-wide shell authority.

    The strengthened invariant is:

    Infrastructure Ops automates one exact Resource effect only when the current fault still exists and current policy permits it. Current health, finding state and the exact side-effect outcome remain distinct evidence, so recovery never becomes fabricated causality or duplicate remediation.

  8. Joncallim commented on Oct 1, 2026

    @Joncallim
    OwnerAuthor

    Architecture reconciliation — bounded recovery, not a claim that restart is reversible

    Reviewed 1 October 2026 against this issue and its existing incident/restart addendum. That addendum already supplies the essential identity, pre-action recheck, uncertainty and notification separation. Preserve it. No service, host, account or pilot was accessed or changed.

    Resolve the remaining wording and safety gaps

    1. Restart is bounded disruption, not automatically reversible. The issue repeatedly describes the first action as reversible. A restart can lose in-memory work or interrupt dependents, and “restart again” does not undo it. Before accepting a pilot, document recoverability, data-loss/availability exposure, preconditions, maximum interruption and the actual compensation/manual recovery route. If the selected service cannot meet the approved envelope, choose a safer fixture/action rather than relabel the risk.
    2. Keep the pilot strictly service-scoped. The side-effect section's “restart/reboot/API call” wording must not be interpreted as host-reboot authorization. Prove one explicit service action. Broad SSH/shell, container-manager sockets, host reboot, Forge's own control plane, its database and its credential broker are not implicitly included in the first service pilot.
    3. Do not let monitoring destroy the observer. Select a target whose restart cannot remove Forge's durable journal, recovery adapter or observation path. If an eventual later control-plane maintenance action is required, it needs an explicit external recovery/observer design and separate authority. The initial fixture should prove safe-stop even when the target's health endpoint and action endpoint fail independently.
    4. Recovery needs a defined settling budget. A healthy response immediately after restart may be a transient start or a load balancer serving another instance. Bind target instance/generation where the service contract exposes it, define startup grace and recovery observation criteria, and cap total incident wall time as well as attempt count. Grace periods cannot restart indefinitely with each wake-up.
    5. Maintenance and deliberate shutdown are first-class inputs. A service intentionally stopped by an operator must not be restarted merely because the health detector sees “down.” Current desired-state/maintenance policy, exact Resource generation and applicable human hold must be checked before remediation. Unknown intent means observation/escalation, not automatic restart.
    6. Failure to notify is not permission to keep trying. If the safe-stop alert itself is blocked/uncertain, retain the incident's attention state and the notification Operation separately. A lack of delivered human notification must not extend the remediation budget or fabricate acknowledgement.
    7. Consume the corrected generic deny path. [FEATURE] Add Project Sentinel detection and escalation flow #190 publishes applicable adverse evidence through [FEATURE] Add evidence-based earned autonomy policy engine #189's generic current-block interface before further privileged admission. Do not depend on eventual reevaluation or clear authority because a finding was acknowledged/suppressed.

    Additional release fixtures

    • Intentional operator stop/maintenance hold: healthy monitoring continues, no automatic restart.
    • Service startup briefly reports healthy then fails inside the recovery window: no premature incident closure.
    • Load balancer returns a different healthy instance: exact-target uncertainty remains explicit.
    • Repeated events cannot reset grace period, cooldown or total incident deadline.
    • Target failure also removes its health endpoint: bounded unknown/safe-stop, not “down therefore restart forever.”
    • Notification fails after safe-stop: no resumed remediation and no false “operator informed.”
    • Pilot selection rejects a service whose disruption can remove Forge's journal/recovery control plane or exceed the approved blast radius.

    Closeout

    Keep fixtures first and the separate normal/uncertain/restart release proofs. Record incident objective health separately from exact Operation confirmation; self-recovery never fabricates Forge-caused success. A real pilot requires an explicitly selected supported Resource and its actual operator-authorized disruption envelope. This architecture review supplies no such host/action authorization and does not waive #342/#190/#366 or their transitive prerequisites.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    dependency-blockedREADINESS PROJECTION — Issue is blocked by unresolved dependencies. This label is a cache.enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions