You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[FEATURE] Add reliability, autonomy, and regression reporting #191
The reliability/autonomy programme is not operationally useful if evidence remains scattered across low-level tables/logs or collapsed into opaque scores. Operators need to see what still works, which exact capability/resource cohort has earned or lost autonomy, what critical findings are active, and which evidence/policy produced each decision. Reporting must remain a projection over authoritative evidence, not a new workflow/state engine.
Desired Outcome
Forge provides evidence-first APIs, a compact operator surface and machine-readable exports covering verification goals/proof history, comparable capability reliability, independent verification, autonomy decisions and Sentinel findings. Every summary is scoped, explains uncertainty/freshness/sample size, links to the underlying immutable evidence and exposes only explicit audited actions.
No global “agent trust score” or hidden grade is introduced.
User Story
As the Forge operator,
I want one clear explanation of what is healthy, regressing, independently verified and allowed to act autonomously,
So that I can supervise Forge from evidence rather than reading raw logs or trusting model confidence.
Requirements
A. Read model / projection rule
Build reporting selectors/projections over canonical sources from #185/#186/#188/#189/#190/#355/#356. Do not create a second mutable truth for Mission, reliability, autonomy or finding state.
Every displayed/exported value identifies the source data/version or can be reconstructed from it. Projection refresh/caching failures cannot silently show stale data as current.
B. Verification-goal view
For each visible goal/revision show at least:
enabled/archived/current revision state;
requested cadence/Trigger state where applicable;
latest proof result and canonical outcome;
last-known-green and first-observed-failing Resource/ref/time when known;
Never aggregate materially incompatible cohorts into a single percentage.
D. Independent verification view
Expose #188 criterion/finding/Gate state and re-verification lineage. Make deterministic test/proof evidence distinguishable from model verifier findings. Missing/not-tested/inconclusive criteria remain visible and cannot be hidden by an overall friendly label.
Explain plainly that the displayed level is a ceiling and actual Grants may be stricter.
F. Sentinel view
Show active/acknowledged/suppressed/recently resolved findings with severity, first/last observation, detector/source evidence, linked autonomy effect, escalation state and recovery action. Critical/high findings remain prominent even when aggregate reliability is high.
G. Cross-links
Where available, link evidence to Mission/Execution/Work Package/Agent Run/Operation/Artifact/Gate, proof goal/run, GitHub issue/PR, Trigger occurrence and relevant Resource identity without exposing secrets/raw protected payloads.
H. Filtering/pagination/performance
Support practical filtering by Resource/project/Mission, Capability, status/severity, autonomy level, runtime/model cohort and date/freshness window. Queries are paginated/bounded and use indexed/rebuildable projections where necessary; do not load all evidence history into a client page.
I. Machine-readable export
Provide a versioned export containing scoped metrics, reason codes, policy/evidence refs and staleness metadata suitable for audit or later Workspace/dashboard integration. Export is redaction/classification aware and does not dump protected prompts/secrets/raw logs.
J. Explicit audited actions
Reporting may expose actions only through existing authoritative service boundaries:
acknowledge/suppress finding;
cap/revoke autonomy;
request on-demand proof;
request re-verification;
follow an allowed recovery/escalation action.
The reporting layer itself does not mutate underlying evidence or autonomously restore authority. Every write uses the normal authenticated audited API/Operation.
K. Degraded/partial history
Define explicit states for unavailable source, stale projection, partial migrated history, evidence retention/redaction and incompatible old cohort data. Never substitute missing history with zero or a green state.
L. Accessibility/responsiveness
Any UI must remain understandable without color alone, support keyboard/screen-reader semantics and work on desktop/mobile. Keep the first view compact: “needs attention / what still works / earned autonomy / recent evidence”, then drill down.
Implementation Sequence
Reporting API schemas + source map — versioned read/export contracts and explicit canonical source table/service for each field.
Large - read models/APIs, compact operator UI, audited action wiring and export; expected as 4-7 small PRs.
Technical Notes
Prefer server-side bounded selectors/projections over a giant client dashboard that downloads raw histories. Friendly labels may summarize, but machine-readable reason codes/evidence refs remain the underlying truth.
Parent Epic: #184
Execution mode: implementation
VNext programme: #333
Depends on: #186, #188, #189, #190, #356
Consumed by: #344
Problem Statement
The reliability/autonomy programme is not operationally useful if evidence remains scattered across low-level tables/logs or collapsed into opaque scores. Operators need to see what still works, which exact capability/resource cohort has earned or lost autonomy, what critical findings are active, and which evidence/policy produced each decision. Reporting must remain a projection over authoritative evidence, not a new workflow/state engine.
Desired Outcome
Forge provides evidence-first APIs, a compact operator surface and machine-readable exports covering verification goals/proof history, comparable capability reliability, independent verification, autonomy decisions and Sentinel findings. Every summary is scoped, explains uncertainty/freshness/sample size, links to the underlying immutable evidence and exposes only explicit audited actions.
No global “agent trust score” or hidden grade is introduced.
User Story
As the Forge operator,
I want one clear explanation of what is healthy, regressing, independently verified and allowed to act autonomously,
So that I can supervise Forge from evidence rather than reading raw logs or trusting model confidence.
Requirements
A. Read model / projection rule
Build reporting selectors/projections over canonical sources from #185/#186/#188/#189/#190/#355/#356. Do not create a second mutable truth for Mission, reliability, autonomy or finding state.
Every displayed/exported value identifies the source data/version or can be reconstructed from it. Projection refresh/caching failures cannot silently show stale data as current.
B. Verification-goal view
For each visible goal/revision show at least:
Distinguish “not run”, “inconclusive”, “blocked”, “failed assertion” and “runner/Trigger failure”.
C. Reliability cohort view
Show exact comparable cohort dimensions, including relevant:
Never aggregate materially incompatible cohorts into a single percentage.
D. Independent verification view
Expose #188 criterion/finding/Gate state and re-verification lineage. Make deterministic test/proof evidence distinguishable from model verifier findings. Missing/not-tested/inconclusive criteria remain visible and cannot be hidden by an overall friendly label.
E. Autonomy view
For each scoped policy/cohort show:
promote|hold|maintain|demote|revoke|requalification);Explain plainly that the displayed level is a ceiling and actual Grants may be stricter.
F. Sentinel view
Show active/acknowledged/suppressed/recently resolved findings with severity, first/last observation, detector/source evidence, linked autonomy effect, escalation state and recovery action. Critical/high findings remain prominent even when aggregate reliability is high.
G. Cross-links
Where available, link evidence to Mission/Execution/Work Package/Agent Run/Operation/Artifact/Gate, proof goal/run, GitHub issue/PR, Trigger occurrence and relevant Resource identity without exposing secrets/raw protected payloads.
H. Filtering/pagination/performance
Support practical filtering by Resource/project/Mission, Capability, status/severity, autonomy level, runtime/model cohort and date/freshness window. Queries are paginated/bounded and use indexed/rebuildable projections where necessary; do not load all evidence history into a client page.
I. Machine-readable export
Provide a versioned export containing scoped metrics, reason codes, policy/evidence refs and staleness metadata suitable for audit or later Workspace/dashboard integration. Export is redaction/classification aware and does not dump protected prompts/secrets/raw logs.
J. Explicit audited actions
Reporting may expose actions only through existing authoritative service boundaries:
The reporting layer itself does not mutate underlying evidence or autonomously restore authority. Every write uses the normal authenticated audited API/Operation.
K. Degraded/partial history
Define explicit states for unavailable source, stale projection, partial migrated history, evidence retention/redaction and incompatible old cohort data. Never substitute missing history with zero or a green state.
L. Accessibility/responsiveness
Any UI must remain understandable without color alone, support keyboard/screen-reader semantics and work on desktop/mobile. Keep the first view compact: “needs attention / what still works / earned autonomy / recent evidence”, then drill down.
Implementation Sequence
Primary Code Seams To Inspect First
Orthogonal Checkpoints
Acceptance Criteria
Out of Scope
Implementation Scope
Large - read models/APIs, compact operator UI, audited action wiring and export; expected as 4-7 small PRs.
Technical Notes
Prefer server-side bounded selectors/projections over a giant client dashboard that downloads raw histories. Friendly labels may summarize, but machine-readable reason codes/evidence refs remain the underlying truth.