Skip to content

[FEATURE] Add reliability, autonomy, and regression reporting #191

Description

@Joncallim

Parent Epic: #184
Execution mode: implementation
VNext programme: #333
Depends on: #186, #188, #189, #190, #356
Consumed by: #344

Problem Statement

The reliability/autonomy programme is not operationally useful if evidence remains scattered across low-level tables/logs or collapsed into opaque scores. Operators need to see what still works, which exact capability/resource cohort has earned or lost autonomy, what critical findings are active, and which evidence/policy produced each decision. Reporting must remain a projection over authoritative evidence, not a new workflow/state engine.

Desired Outcome

Forge provides evidence-first APIs, a compact operator surface and machine-readable exports covering verification goals/proof history, comparable capability reliability, independent verification, autonomy decisions and Sentinel findings. Every summary is scoped, explains uncertainty/freshness/sample size, links to the underlying immutable evidence and exposes only explicit audited actions.

No global “agent trust score” or hidden grade is introduced.

User Story

As the Forge operator,
I want one clear explanation of what is healthy, regressing, independently verified and allowed to act autonomously,
So that I can supervise Forge from evidence rather than reading raw logs or trusting model confidence.

Requirements

A. Read model / projection rule

Build reporting selectors/projections over canonical sources from #185/#186/#188/#189/#190/#355/#356. Do not create a second mutable truth for Mission, reliability, autonomy or finding state.

Every displayed/exported value identifies the source data/version or can be reconstructed from it. Projection refresh/caching failures cannot silently show stale data as current.

B. Verification-goal view

For each visible goal/revision show at least:

Distinguish “not run”, “inconclusive”, “blocked”, “failed assertion” and “runner/Trigger failure”.

C. Reliability cohort view

Show exact comparable cohort dimensions, including relevant:

  • Capability/Operation;
  • Resource scope fingerprint/classification;
  • Workforce/package/workflow/harness/policy lineage;
  • runtime/provider/model compatibility bucket;
  • sample window/count;
  • verified passes/failures/inconclusive/critical events;
  • freshness and exclusion reason counts.

Never aggregate materially incompatible cohorts into a single percentage.

D. Independent verification view

Expose #188 criterion/finding/Gate state and re-verification lineage. Make deterministic test/proof evidence distinguishable from model verifier findings. Missing/not-tested/inconclusive criteria remain visible and cannot be hidden by an overall friendly label.

E. Autonomy view

For each scoped policy/cohort show:

  • current effective autonomy level/ceiling;
  • operator/system maximum;
  • latest decision (promote|hold|maintain|demote|revoke|requalification);
  • policy/version/threshold snapshot;
  • stable reason codes;
  • evidence window/sample/freshness;
  • critical failure/revocation cause;
  • Resource/Capability scope;
  • next requalification/review requirement.

Explain plainly that the displayed level is a ceiling and actual Grants may be stricter.

F. Sentinel view

Show active/acknowledged/suppressed/recently resolved findings with severity, first/last observation, detector/source evidence, linked autonomy effect, escalation state and recovery action. Critical/high findings remain prominent even when aggregate reliability is high.

G. Cross-links

Where available, link evidence to Mission/Execution/Work Package/Agent Run/Operation/Artifact/Gate, proof goal/run, GitHub issue/PR, Trigger occurrence and relevant Resource identity without exposing secrets/raw protected payloads.

H. Filtering/pagination/performance

Support practical filtering by Resource/project/Mission, Capability, status/severity, autonomy level, runtime/model cohort and date/freshness window. Queries are paginated/bounded and use indexed/rebuildable projections where necessary; do not load all evidence history into a client page.

I. Machine-readable export

Provide a versioned export containing scoped metrics, reason codes, policy/evidence refs and staleness metadata suitable for audit or later Workspace/dashboard integration. Export is redaction/classification aware and does not dump protected prompts/secrets/raw logs.

J. Explicit audited actions

Reporting may expose actions only through existing authoritative service boundaries:

  • acknowledge/suppress finding;
  • cap/revoke autonomy;
  • request on-demand proof;
  • request re-verification;
  • follow an allowed recovery/escalation action.

The reporting layer itself does not mutate underlying evidence or autonomously restore authority. Every write uses the normal authenticated audited API/Operation.

K. Degraded/partial history

Define explicit states for unavailable source, stale projection, partial migrated history, evidence retention/redaction and incompatible old cohort data. Never substitute missing history with zero or a green state.

L. Accessibility/responsiveness

Any UI must remain understandable without color alone, support keyboard/screen-reader semantics and work on desktop/mobile. Keep the first view compact: “needs attention / what still works / earned autonomy / recent evidence”, then drill down.

Implementation Sequence

  1. Reporting API schemas + source map — versioned read/export contracts and explicit canonical source table/service for each field.
  2. Goal/proof selectors — [FEATURE] Complete verification-goal on-demand proof execution through VNext #355/[FEATURE] Schedule verification-goal proof runs through generic Triggers #356 status/freshness/last-green projections.
  3. Reliability cohort selectors — [FEATURE] Add capability reliability ledger #186 comparison/exclusion/sample dimensions.
  4. Verification selectors — [FEATURE] Add independent Verification Workforce execution #188 criterion/findings/Gate/re-verification lineage.
  5. Autonomy selectors — [FEATURE] Add evidence-based earned autonomy policy engine #189 current decision/effective ceiling/reason/evidence.
  6. Sentinel selectors — [FEATURE] Add Project Sentinel detection and escalation flow #190 active/resolved/suppressed/recovery links.
  7. Compact operator page — needs-attention first, evidence-backed drilldown; bounded server queries.
  8. Explicit actions — wire existing cap/revoke/proof/reverify/finding APIs with confirmation/audit state.
  9. Machine-readable export — version/redaction/staleness and evidence refs.
  10. Performance/accessibility/degraded-state gate — large fixture histories, partial migration, stale sources, mobile/keyboard/screen reader.

Primary Code Seams To Inspect First

Orthogonal Checkpoints

  1. Truth/projection: source mismatch, stale caches, duplicate rows, partial history, rebuildability.
  2. Cohort integrity: accidental aggregation across Resource/model/package/policy dimensions, misleading percentages.
  3. Critical prominence: aggregate success cannot hide critical failure/revoke/insufficient evidence.
  4. Actions/authority: reporting route tries to mutate evidence/restore autonomy or bypass audited service.
  5. Privacy/export: protected prompt/path/secret/raw payload leakage, retention/redaction state.
  6. Performance: large evidence history, pagination/indexes, N+1/client overfetch.
  7. UX/accessibility: empty/stale/degraded states, mobile, keyboard, screen-reader, plain-English reason explanations.

Acceptance Criteria

  • Operators can see every current verification goal and distinguish not-run/pass/fail/inconclusive/blocked/runner failure with evidence/freshness.
  • Reliability views identify exact comparable cohort dimensions, sample size/window and exclusions.
  • Verification views expose per-criterion and deterministic-vs-model evidence; missing/not-tested/inconclusive state remains visible.
  • Autonomy views show current level/ceiling, policy/version, reason codes, scope, sample/freshness and critical revocation cause.
  • Critical failures/unresolved verification gaps/Sentinel findings cannot be hidden by high aggregate performance.
  • Active/acknowledged/suppressed/resolved Sentinel findings are distinguishable and traceable.
  • Every material metric/decision links to or identifies its authoritative evidence inputs.
  • Operators can cap/revoke autonomy, request proof/reverification and manage findings only through explicit audited authoritative actions.
  • Machine-readable export preserves scope/reason/evidence/staleness and obeys redaction/classification policy.
  • Missing/stale/partial history is explicit and never forged into zero/green.
  • Queries are bounded/paginated and remain usable against large fixture history.
  • UI passes responsive, keyboard, screen-reader and non-color-only status checks.

Out of Scope

  • Replacing the general Mission/Workforce/task UI.
  • Opaque global letter grades or one agent trust score.
  • Automatically resolving findings/restoring autonomy from reporting.
  • Editing policy thresholds without the [FEATURE] Add evidence-based earned autonomy policy engine #189 audited policy path.
  • Building the entire long-term Workspace shell.

Implementation Scope

Large - read models/APIs, compact operator UI, audited action wiring and export; expected as 4-7 small PRs.

Technical Notes

Prefer server-side bounded selectors/projections over a giant client dashboard that downloads raw histories. Friendly labels may summarize, but machine-readable reason codes/evidence refs remain the underlying truth.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    dependency-blockedREADINESS PROJECTION — Issue is blocked by unresolved dependencies. This label is a cache.enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions