Repository navigation
[FEATURE] VNext Phase 9 — Infrastructure Ops Workforce and persistent side-effect proof #343
Description
Activity
FORGE issue validation
This issue is not semantically dispatchable.
Reason Detail queue.issue_dependency_openA declared dependency is still open. Dependency: #342. queue.issue_dependency_openA declared dependency is still open. Dependency: #190. queue.issue_dependency_openA declared dependency is still open. Dependency: #366. queue.issue_dependency_openA declared dependency is still open. Dependency: #341. queue.issue_dependency_openA declared dependency is still open. Dependency: #340. queue.issue_dependency_openA declared dependency is still open. Dependency: #339. queue.issue_dependency_openA declared dependency is still open. Dependency: #338. queue.issue_dependency_openA declared dependency is still open. Dependency: #337. queue.issue_dependency_openA declared dependency is still open. Dependency: #188. queue.issue_dependency_openA declared dependency is still open. Dependency: #336. queue.issue_dependency_openA declared dependency is still open. Dependency: #334. queue.issue_dependency_openA declared dependency is still open. Dependency: #335. queue.issue_dependency_openA declared dependency is still open. Dependency: #346. queue.issue_dependency_openA declared dependency is still open. Dependency: #353. queue.issue_dependency_openA declared dependency is still open. Dependency: #355. queue.issue_dependency_openA declared dependency is still open. Dependency: #189. queue.issue_dependency_openA declared dependency is still open. Dependency: #347. queue.issue_dependency_openA declared dependency is still open. Dependency: #356. Readiness labels are projections, not authority. Command, dispatch, and handoff always re-resolve current semantic truth.
- addedneeds-clarificationREADINESS PROJECTION — Issue is missing required structure or decisions. Author correction needed.READINESS PROJECTION — Issue is missing required structure or decisions. Author correction needed.ready-for-agentREADINESS PROJECTION — Issue is semantically ready. This label is a cache, not authority.READINESS PROJECTION — Issue is semantically ready. This label is a cache, not authority.and removedneeds-clarificationREADINESS PROJECTION — Issue is missing required structure or decisions. Author correction needed.READINESS PROJECTION — Issue is missing required structure or decisions. Author correction needed.
on Sep 2, 2026 - removedready-for-agentREADINESS PROJECTION — Issue is semantically ready. This label is a cache, not authority.READINESS PROJECTION — Issue is semantically ready. This label is a cache, not authority.
on Sep 3, 2026 - addedready-for-agentREADINESS PROJECTION — Issue is semantically ready. This label is a cache, not authority.READINESS PROJECTION — Issue is semantically ready. This label is a cache, not authority.
on Sep 4, 2026 - removedready-for-agentREADINESS PROJECTION — Issue is semantically ready. This label is a cache, not authority.READINESS PROJECTION — Issue is semantically ready. This label is a cache, not authority.
on Sep 5, 2026 - addeddependency-blockedREADINESS PROJECTION — Issue is blocked by unresolved dependencies. This label is a cache.READINESS PROJECTION — Issue is blocked by unresolved dependencies. This label is a cache.
on Sep 8, 2026 Joncallim commented
on Sep 9, 2026 OwnerAuthorMore actionsHostile Infrastructure Ops addendum — incident identity, pre-action revalidation, and restart reconciliation without fabricated causality
A fresh persistent-side-effect pass found several correctness traps that matter even for the deliberately narrow first remediation family.
1. Do not assume
restart serviceis idempotentRepeated restarts can cause material extra disruption even when the final state is the same. The initial restart Operation must declare its real retry class based on the selected service manager/adapter contract.
- if the external manager supports a stable request/idempotency identity and can confirm it, use
idempotent_with_key; - if Forge can determine whether this exact requested restart happened from service-manager job/audit/generation evidence, use
reconcile_before_retry; - if neither is available, treat restart as
at_most_once/human-required after possible submission; - never mark it replay-safe merely because calling restart twice often ends with a running service.
Disable hidden SDK/CLI retries per #342. Every external submission is owned by the Operation journal.
2. “Service is healthy now” is not automatically proof that Forge's restart succeeded
The service may have:
- self-recovered;
- been restarted by a human/another system;
- failed over to another instance;
- become reachable despite Forge's ambiguous request never arriving.
Keep two evidence questions separate:
- incident objective state — is the service currently healthy enough to stop remediation/escalate differently?
- Operation outcome — did this exact restart Operation execute/confirm according to its declared confirmation policy?
A healthy post-check can stop Forge from issuing another restart even while the prior Operation remains
submission_uncertain/indeterminate. Do not forgereconciled_successunless evidence links the state change to the exact Operation according to the adapter contract.This distinction is essential for reliability/autonomy evidence later.
3. Use Sentinel finding identity as the incident/detection truth; do not invent a parallel Infra incident database
#343 depends on #190. Reuse its finding/observation lifecycle for “this Resource is currently unhealthy/flapping/critical” where semantically applicable.
- Trigger occurrences are observations/wakeups;
- Sentinel finding is the durable detection/incident-like identity and severity/escalation state;
- Mission/Execution owns the bounded diagnosis/remediation attempt;
- Operation owns each side effect;
- verification/Gate owns recovery evidence.
Do not add a second
infra_incidentsstate machine that can disagree with #190 findings. If #190's finding contract lacks a needed field, extend it generically or add a narrow linked remediation record rather than duplicate current severity/resolution truth.4. Multiple health occurrences for one ongoing condition coalesce into one active finding/remediation policy
A schedule firing every minute during a 20-minute outage must not create 20 independent restart Executions.
Define deterministic incident/finding correlation keyed by the exact detector + logical service Resource + problem class/version according to #190 policy.
- repeated same-state observations update/append observation evidence to the active finding;
- an already-running remediation Execution suppresses/coalesces duplicate action requests unless policy explicitly allows another after cooldown;
- a materially different fault/problem may create a new finding;
- recovery closes/resolves the finding only after required verification policy;
- recurrence after resolution is a new occurrence/episode with lineage, not reopening/overwriting old history invisibly.
5. Revalidate current state immediately before every remediation submission
Diagnosis evidence can become stale while a model/Gate waits.
Before external restart submission:
- re-check Mission/Trigger/current autonomy/Grant/budget/cooldown as already required;
- take a fresh bounded deterministic health observation of the exact Resource or validate the declared freshness horizon;
- verify the fault/remediation precondition still holds;
- bind that observation/version into Operation admission;
- if service recovered, record
remediation_not_needed/superseded evidence and do not restart a healthy service.
This pre-action observation is not permission to bypass the original finding history; it is a stale-action guard.
6. Service Resource identity must be canonical and non-injectable
For the first real pilot, adapter-specific Resource binding identifies one exact logical service (e.g. exact systemd unit/container/service id) independently of display text/model suggestions.
- no wildcard/unit glob;
- no user/model-authored shell fragment;
- normalize/validate service identifiers through the adapter schema;
- symlink/alias/templated-unit/container-name semantics are resolved under reviewed policy;
- Resource binding declares host/manager/account identity so
service Xon host A is not confused with host B; - a restart action cannot be widened to dependent/all services because the model says the dependency “probably needs it.”
7. Model diagnosis proposes evidence/advisory classifications, never an arbitrary command or Capability
When cognition is used, its structured output is bounded to an allowlisted diagnosis/remediation-candidate schema.
Example:
observed fault class hypothesis/evidence refs recommended candidate Operation id from allowed set | none confidence/advisory rationaleThe deterministic policy/Gate then decides whether the candidate Operation is permitted and appropriate under current evidence/authority.
A model cannot emit
systemctl ..., SSH commands, URLs, credentials or a new Capability that becomes executable merely because it is in the diagnosis.8. Health detector semantics need hysteresis/freshness, not one-sample automation
Define deterministic policy per detector:
- observation schema/fingerprint;
- required freshness;
- healthy/unhealthy/unknown/stale classification;
- consecutive/sample-window thresholds where appropriate;
- recovery threshold/hysteresis;
- flapping definition/window;
- cooldown/max-remediation counts;
- detector version.
One transient timeout must not become “service down” unless the policy explicitly says one failure is sufficient. Conversely, missing/stale observation is not healthy.
All counters/windows use durable DB evidence/time so restart cannot reset flapping/retry history.
9. Flapping cooldown and remediation budget survive process/reboot
Persist authoritative attempt/cooldown state as projections from finding/Operation/evidence rather than local timers only.
- process restart cannot reset max-restarts-per-window;
- worker clock changes cannot shorten cooldown;
- two workers/triggers cannot concurrently pass the same restart-rate limit;
- cooldown admission uses DB time/CAS;
- manual operator remediation, when observable, may affect state/freshness but does not silently count as Forge's autonomous successful Operation.
10. Post-remediation verification must be independent of the write acknowledgement
The adapter response “restart accepted” is not the health verification.
Run a separately identified read/health proof against the exact Resource after policy-defined settling semantics. The verification result is one of healthy / unhealthy / stale-inconclusive / inaccessible, with exact evidence.
Where possible, use evidence independent from the mutation response (e.g. fresh health/status generation or service-manager state) and never let model prose alone close the finding.
If verification cannot establish health, Gate stays retry/human/blocked according to bounds. Unknown is not recovered.
11. Dependency/blast-radius policy is explicit for the real pilot
Even restarting one named service can disrupt dependents.
The Operation definition/pilot review should document known effect/risk scope and any required preconditions, e.g.:
- service is independently restartable under this proof;
- restart does not intentionally restart the host/all services;
- no critical dependent Resource is outside the granted risk envelope;
- expected interruption/timeout bounds;
- compensation/recovery options.
Do not build a generic dependency graph in #343, but do not select a “reversible” pilot whose real blast radius is unbounded or not understood.
12. Notification is independent side-effect state
A notification failure does not make the restart/health action fail and must never cause the Mission to rerun remediation blindly.
- notification Operation has its own idempotency/retry/reconciliation identity;
- content classification/egress is admitted separately and minimizes logs/secrets;
- retries only the notification Operation according to its adapter policy;
- healthy state/remediation outcome remains whatever the underlying evidence says;
- duplicate incident observations coalesce notification policy so storms are bounded.
13. Current autonomy demotion/revocation wins races with queued remediation
When #190/#189 demotes or revokes autonomy:
- outstanding action admission tokens under the higher ceiling become invalid for new submission;
- queued/diagnosed candidate restart rechecks current autonomy immediately before adapter submission;
- stale worker cannot act because it passed the check earlier;
- already-submitted Operation follows reconciliation; revocation cannot pretend the external effect did not happen;
- no model/Workforce may restore its own autonomy after the incident.
14. Real pilot needs an operator-visible safe-stop state
Define the state reached when Forge cannot safely continue automatically, including:
- Operation submission uncertain;
- repeated failure/cooldown exhausted;
- conflicting/stale health evidence;
- Grant/autonomy revoked;
- verification inconclusive;
- adapter/credential unavailable.
Safe stop is healthy quiescence only if health is actually proven. Otherwise Mission/finding remains attention/human-required without token-burning loops.
15. Hostile release fixture additions
Add:
- restart request times out after possible submission; service becomes healthy independently -> no second restart, Operation remains uncertain until exact reconciliation/human policy;
- service self-recovers between diagnosis and action -> pre-action check suppresses restart;
- human restarts service concurrently -> Forge does not claim that external action as its Operation;
- 20 repeated outage Trigger occurrences -> one active finding/remediation flow under policy;
- transient one-sample timeout below unhealthy threshold -> no restart;
- flapping persists across Forge reboot -> cooldown/restart count preserved;
- two workers race restart cooldown -> one admitted;
- model outputs arbitrary shell/restart-all -> cannot become Operation;
- exact service name on wrong host/tenant -> Resource binding denies;
- restart accepted but health verification stale -> not recovered;
- autonomy revoked after diagnosis before submit -> restart denied;
- notification fails after successful recovery -> only notification retried, no remediation repeat;
- real pilot proves bounded blast radius/rollback and no host-wide shell authority.
The strengthened invariant is:
Infrastructure Ops automates one exact Resource effect only when the current fault still exists and current policy permits it. Current health, finding state and the exact side-effect outcome remain distinct evidence, so recovery never becomes fabricated causality or duplicate remediation.
- if the external manager supports a stable request/idempotency identity and can confirm it, use
Joncallim commented
on Oct 1, 2026 OwnerAuthorMore actionsArchitecture reconciliation — bounded recovery, not a claim that restart is reversible
Reviewed 1 October 2026 against this issue and its existing incident/restart addendum. That addendum already supplies the essential identity, pre-action recheck, uncertainty and notification separation. Preserve it. No service, host, account or pilot was accessed or changed.
Resolve the remaining wording and safety gaps
- Restart is bounded disruption, not automatically reversible. The issue repeatedly describes the first action as reversible. A restart can lose in-memory work or interrupt dependents, and “restart again” does not undo it. Before accepting a pilot, document recoverability, data-loss/availability exposure, preconditions, maximum interruption and the actual compensation/manual recovery route. If the selected service cannot meet the approved envelope, choose a safer fixture/action rather than relabel the risk.
- Keep the pilot strictly service-scoped. The side-effect section's “restart/reboot/API call” wording must not be interpreted as host-reboot authorization. Prove one explicit service action. Broad SSH/shell, container-manager sockets, host reboot, Forge's own control plane, its database and its credential broker are not implicitly included in the first service pilot.
- Do not let monitoring destroy the observer. Select a target whose restart cannot remove Forge's durable journal, recovery adapter or observation path. If an eventual later control-plane maintenance action is required, it needs an explicit external recovery/observer design and separate authority. The initial fixture should prove safe-stop even when the target's health endpoint and action endpoint fail independently.
- Recovery needs a defined settling budget. A healthy response immediately after restart may be a transient start or a load balancer serving another instance. Bind target instance/generation where the service contract exposes it, define startup grace and recovery observation criteria, and cap total incident wall time as well as attempt count. Grace periods cannot restart indefinitely with each wake-up.
- Maintenance and deliberate shutdown are first-class inputs. A service intentionally stopped by an operator must not be restarted merely because the health detector sees “down.” Current desired-state/maintenance policy, exact Resource generation and applicable human hold must be checked before remediation. Unknown intent means observation/escalation, not automatic restart.
- Failure to notify is not permission to keep trying. If the safe-stop alert itself is blocked/uncertain, retain the incident's attention state and the notification Operation separately. A lack of delivered human notification must not extend the remediation budget or fabricate acknowledgement.
- Consume the corrected generic deny path. [FEATURE] Add Project Sentinel detection and escalation flow #190 publishes applicable adverse evidence through [FEATURE] Add evidence-based earned autonomy policy engine #189's generic current-block interface before further privileged admission. Do not depend on eventual reevaluation or clear authority because a finding was acknowledged/suppressed.
Additional release fixtures
- Intentional operator stop/maintenance hold: healthy monitoring continues, no automatic restart.
- Service startup briefly reports healthy then fails inside the recovery window: no premature incident closure.
- Load balancer returns a different healthy instance: exact-target uncertainty remains explicit.
- Repeated events cannot reset grace period, cooldown or total incident deadline.
- Target failure also removes its health endpoint: bounded unknown/safe-stop, not “down therefore restart forever.”
- Notification fails after safe-stop: no resumed remediation and no false “operator informed.”
- Pilot selection rejects a service whose disruption can remove Forge's journal/recovery control plane or exceed the approved blast radius.
Closeout
Keep fixtures first and the separate normal/uncertain/restart release proofs. Record incident objective health separately from exact Operation confirmation; self-recovery never fabricates Forge-caused success. A real pilot requires an explicitly selected supported Resource and its actual operator-authorized disruption envelope. This architecture review supplies no such host/action authorization and does not waive #342/#190/#366 or their transitive prerequisites.
Parent: #333
Execution mode: implementation
Depends on: #342, #190, #366
Uses: #189 autonomy ceilings, #340 Missions, #341 Triggers
Spec references: SPEC-0004 (Operation side-effect lifecycle — branching state graph, submission_uncertain, reconciliation), SPEC-0012 (audit evidence for recovery), SPEC-0008 (conformance, failure injection)
Problem Statement
Forge is not proven as an ongoing agentic runtime until it can own a persistent responsibility that observes real state, stays silent/cheap when healthy, performs a bounded reversible side effect when policy allows, verifies the result, survives restart and stops/escalates safely on uncertainty or repeated failure.
A coding/research success is insufficient because those are mostly finite workflows. Infrastructure Ops provides a technically verifiable persistent side-effect proof before higher-risk Personal Ops automation.
Desired Outcome
An official Infrastructure Ops Workforce maintains a bounded set of configured service Resources under an explicit Mission, Trigger, Grant, autonomy ceiling and budget. Deterministic health checks run continuously/periodically with zero model calls while unchanged; supported reversible remediations execute only as typed Operations and are independently verified. Failures, flapping and uncertainty stop/escalate rather than looping.
User Story
As the Forge operator,
I want Forge to maintain a narrowly scoped service-health responsibility safely and quietly,
So that I can trust persistent event-driven automation before moving additional operational workflows away from Hermes.
Requirements
A. Official Workforce package
Ship Infrastructure Ops as a normal #338 package. It declares the Workflow, roles, requested service-read/restart/notification Capabilities, budget/cognitive defaults and verification policy; it carries no host credentials or executable hooks.
B. Service Resource binding
Bind an explicit allowlisted set of service Resources through #342. Each Resource specifies supported health/read and optional reversible action Capabilities. Do not expose broad host/SSH/shell authority.
Initial real or fixture-backed Operations should be deliberately narrow, e.g.:
Use a supported adapter/runtime appropriate to the host; arbitrary model-authored commands are prohibited.
C. Persistent Mission
Create one recurring Mission with:
Mission remains quiescent between occurrences and retains no permanent model conversation.
D. Trigger and deterministic health
Use #341 schedule/event Trigger occurrences. Deterministic health adapter executes first. If state is healthy/unchanged, record compact evidence/update and terminate with zero model calls.
Only a meaningful state transition/failure creates an incident Execution requiring additional diagnosis/remediation.
E. Deterministic-first diagnosis
Use stable fault classification where evidence is sufficient (service down, known unhealthy status, transient probe failure, stale data, flapping). Invoke an economy/standard technical Agent Run only when the deterministic evidence is insufficient for the next bounded decision; frontier escalation follows #335 policy only.
Log/status text is untrusted data and cannot request tools or change policy.
F. Remediation admission
Before every action re-check Mission state, current #189 autonomy ceiling, Grant, Resource scope/version, cooldown/retry/budget and side-effect state.
The first autonomous action family must be reversible/bounded. Anything outside the current ceiling requires explicit operator Gate/approval.
G. Side-effect recovery
Use #336/#342 idempotency/reconciliation lifecycle. A restart/reboot/API call with uncertain submission is reconciled by observing service state and operation identity; never blindly reissue.
H. Independent verification
After a remediation, run deterministic verification until a bounded terminal decision:
Optional model review does not self-authorize completion. Trusted Gate policy decides whether incident returns to quiescent, retries one allowed step, escalates or revokes autonomy.
I. Flapping/repeated failure
Detect repeated state changes/failed remediations deterministically. Enforce cooldown/max attempts and stop model/token burn. High-severity repeated failure creates/updates a #190 Sentinel finding and can trigger #189 demotion/revocation before any additional autonomous action.
J. Notification
Notify only for actionable transitions: remediation performed/failed, repeated/flapping failure, policy block, budget exhaustion or human decision required. Healthy unchanged checks remain silent. Notifications are typed adapter Operations with dedupe/idempotency evidence.
K. Evidence
Persist compact incident/health fingerprints, Operations, verification results, routing/budget receipts, autonomy decision refs and last-known validated state. Do not store unbounded raw logs or conversations as Mission memory.
Implementation Sequence
Primary Code Seams To Inspect First
Orthogonal Checkpoints
Acceptance Criteria
Out of Scope
Implementation Scope
Very Large / operationally sensitive - reference Workforce plus bounded service/notification adapters and persistent release proof, expected as 6-9 small PRs.
Technical Notes
Keep the real pilot narrower than the generic architecture. Prove one safe reversible action exhaustively before adding more service-control Operations.