Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,12 @@ python -m orchestrator.dashboard.server --port 8765
See [persistent delivery and operations](docs/delivery.md) for contracts, controls,
authority boundaries, migration, private access and deployment. The default planning
policy manages accepted commitments rather than generating speculative growth work.

The [general reliability contract](docs/reliability.md) adds scoped roles and
retained memory, measured usage and traces, isolated execution, orchestration
evaluations, release/drift gates, service targets and tested recovery procedures.
`/reliability` shows what is qualified and what is missing. These controls are
opt-in; passing engineering tests is not proof of sustained business performance.
This is not a claim of unrestricted autonomy or a completed non-coding production pilot.

**Public proof — everything is auditable:**
Expand Down
54 changes: 54 additions & 0 deletions bin/evaluate_orchestration.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,54 @@
#!/usr/bin/env python3
"""Run real controller fault scenarios, emitting a bounded engineering-eval report."""

import json
import subprocess
import sys
import tempfile
import xml.etree.ElementTree as ET
from pathlib import Path

ROOT = Path(__file__).resolve().parents[1]
sys.path.insert(0, str(ROOT))
from orchestrator.reliability import SCENARIOS


def main():
with tempfile.TemporaryDirectory() as directory:
report = Path(directory) / "result.xml"
subprocess.run(
[
sys.executable,
"-m",
"pytest",
"tests/test_reliability_scenarios.py",
"-q",
"--junitxml=" + str(report),
],
cwd=ROOT,
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL,
timeout=90,
check=False,
)
cases = ET.parse(report).getroot().findall(".//testcase")
outcomes = {
case.attrib["name"].removeprefix("test_"): not any(
case.find(tag) is not None for tag in ("failure", "error", "skipped")
)
for case in cases
}
results = {name: outcomes.get(name, False) for name in sorted(SCENARIOS)}
print(
json.dumps(
{
"samples": len(results),
"quality": sum(results.values()) / len(results),
"scenarios": results,
}
)
)


if __name__ == "__main__":
main()
9 changes: 6 additions & 3 deletions docs/delivery.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,9 +117,12 @@ exactly-once guarantee is claimed for a provider without idempotency support.
There is no preinstalled video recording/upload adapter in this change. Workers
must inspect available capabilities, preserve intermediate work and ask a specific
access/approval question when needed. A script alone cannot pass a recording check.
General CLI workers still run as trusted host processes; these action gates are
**not an OS sandbox** against a malicious CLI or shell escape. Do not delegate
unrestricted accounts or untrusted tasks under a stronger security assumption.
Legacy CLI workers in observation mode still run as trusted host processes;
action gates alone are **not an OS sandbox**. The opt-in
[general reliability contract](reliability.md) adds a measured model gateway,
network-off isolated tools and verifiers, scoped operator roles, release
qualification, drift and service-target gates. Missing qualification fails closed.
Do not treat a merged implementation as an activated or proven production release.

## Controls And Recovery

Expand Down
71 changes: 71 additions & 0 deletions docs/reliability-policy.example.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,71 @@
# Merge into private operator config after substituting absolute paths and identities.
# This is a template, NOT evidence or authorization. Never commit credentials.
reliability:
mode: observe
notify: true
recovery_operators: ["local:1000"]
planning_model_adapter: measured-planner
principals:
"local:1000":
tenants: [example, controller]
roles: [operator, memory, billing, administrator]
"telegram:123456789":
tenants: [example]
roles: [approver, reader]
tenants:
example:
repos: [owner/workspace]
profiles:
implementation: example-code
architecture: example-code
research: example-research
external_action: example-action
profiles:
example-code: &base
release: REPLACE_WITH_REVIEWED_40_CHARACTER_GIT_SHA
python: /path/to/controller/.venv/bin/python
model_adapter: measured-worker
evaluator: /path/to/controller/bin/evaluate_orchestration.py
rollback_plan: /path/to/controller/docs/runbooks/recovery.md
incident_playbook: /path/to/controller/docs/runbooks/incidents.md
artifacts: []
scenarios: [end_to_end, restart, duplicate_effect, denied_action, stale_revision, provider_failure, restore, isolation, approval_expiry, qualified_worker]
min_samples: 10
min_quality: 1.0
max_quality_drop: 0.0
evaluation_timeout_seconds: 120
evaluation_ttl_seconds: 86400
evaluation_interval_seconds: 3600
auto_evaluate: false
require_sealed_usage: true
slo_window_seconds: 2592000
slo_min_samples: 20
min_success_rate: 0.95
max_delivery_seconds: 14400
memory_ttl_seconds: 2592000
sandbox:
backend: bubblewrap
network: none
readonly: []
environment: {}
env_keys: []
example-research:
<<: *base
max_delivery_seconds: 86400
example-action:
<<: *base
max_delivery_seconds: 3600
# Model adapters are operator-owned executables, outside worker-writable directories.
# model_adapters:
# measured-worker:
# argv: [/absolute/python, -I, /operator-owned/worker_provider_adapter.py]
# cwd: /operator-owned
# tenants: [example]
# env_keys: [PROVIDER_API_KEY]
# timeout_seconds: 120
# measured-planner:
# argv: [/absolute/python, -I, /operator-owned/provider_adapter.py]
# cwd: /operator-owned
# tenants: [example]
# env_keys: [PROVIDER_API_KEY]
# timeout_seconds: 120
198 changes: 198 additions & 0 deletions docs/reliability.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,198 @@
# General Reliability Contract

This layer makes general engineering requirements executable. It does not
certify business impact, an individual customer's quality bar, or months of
successful operation. Release eligibility and production service evidence are
different claims. The live view is `/reliability`, linked from the Proof dashboard;
machine consumers use `/api/reliability` and `/api/traces/<goal-id>`.

## Adoption And Authority

Start from [the policy template](reliability-policy.example.yaml). It is invalid
until the operator assigns real repositories, identities, absolute artifact paths
and an immutable reviewed controller release. No configuration is silently
installed into the live host. The existing execution deployment pin remains an
independent approval. Enforced mode forces dispatcher-only operation, even if a
legacy project override requests full automation. Self-directed review/deployment
jobs are not part of the qualified worker boundary.

`observe` preserves the authenticated legacy single-operator installation and
reports that enforcement is absent. `enforce` denies unmanaged mailbox work,
unmapped repositories, unqualified releases, unavailable isolation, exhausted
service targets and unreconciled prior usage. It does not fall back to trusted
host execution. Legacy model formatting is disabled in enforced intake; original
human intent is retained. Programs use a registered, measured planning adapter,
not an unmetered host CLI fallback.

Repository-to-tenant and task-type-to-profile mappings come only from operator
config, never issue bodies or model output. Parent/child work cannot cross tenant
boundaries in enforced mode; neither can declared dependencies. Reuse YAML templates for profiles, but give each
tenant its own profile names and release approvals. Customer-specific external
tool targets still need exact grants in the goal ancestry.

Principals have explicit tenant scopes and roles: reader, operator, approver,
billing, memory, administrator. Unknown identities fail closed when principals
are configured, even in observe mode. CLI identity is the actual `local:<uid>`;
Telegram uses its authenticated immutable `telegram:<user-id>`, not a mutable
username or shared group ID. Legacy global Telegram commands and callback buttons
require the controller administrator role. Goal commands remain tenant scoped.
The shared private dashboard is an **operator-wide view**, not a customer portal.
Do not give its bearer credential or host access to tenant-only users.

## Traces And Measured Usage

Worker attempts, model-gateway calls, isolated tools, registered actions and acceptance verification
have persisted spans with common goal/revision trace IDs and parent span IDs.
Worker restarts do not lose those records. Errors retain their type, not raw
provider messages. Allowed attributes exclude prompts, tool inputs, credentials
and model outputs. A running span after a crash stays unfinished, not successful.
The model adapter receives a W3C-format `traceparent` for downstream propagation.
These are local traces, not an installed distributed tracing collector.

`model_gateway.call_model` accepts a controller-owned active attempt and invokes
an operator-owned adapter using fixed argv and structured JSON stdin. The adapter
receives `{attempt_id, traceparent, request}` and returns `{output, usage}`.
Only named environment keys reach it. Place adapter code outside worker-writable
paths; Python adapters should use `-I` to exclude workspace import paths. Trusted
adapters must return provider metadata, never ask a model to estimate its cost.

Usage schema:

```json
{"provider":"provider-name","account":"nonsecret-account-reference","request_id":"provider-request-id","model":"model-version","input_tokens":123,"output_tokens":45,"cost_nano_usd":123456,"final":true}
```

Costs use integer billionths of USD in receipts. Provisional receipts may have
null cost. Exact replays are idempotent, final receipts are immutable, and a
provider/account/request identity cannot be charged to two attempts. Token counts
are measured, not estimated from characters. Include caching or other billed
categories in the actual total charge; the two token counters are not a price
calculator. Account references must never be credentials.

A finished attempt's cost remains unknown until a trusted collector seals the
**complete list** of receipt keys. One observed call is not full coverage. Sealing
updates the existing delivery cost and budget accounting atomically. Additional
requests cannot be appended after sealing. Incorrect final receipts require an
audited reconciliation, not silent overwriting. Existing CLI agents that do not
expose a complete measured request manifest remain unmetered; use an instrumented
adapter or an authorized billing collector. The implementation does not invent
provider billing access or invoice data.

```bash
python -m orchestrator.reliability_ops usage --tenant example --attempt ATTEMPT --file receipt.json
python -m orchestrator.reliability_ops seal --tenant example --attempt ATTEMPT --file receipt-keys.json
python -m orchestrator.reliability_ops traces --tenant example --goal GOAL
```

Adapter calls have a bounded timeout and output size; descendants are killed on
exit or timeout. A timed-out effect is uncertain and cannot be blindly repeated.
There is no claim of provider-side exactly-once delivery. Runtime reservations
are admission controls, not absolute provider spending caps.

## Memory And Isolation

Tenant facts have provenance, writer identity, revision and expiry. Updates use
compare-and-swap; stale writers cannot overwrite newer facts. Reads are tenant
scoped and expired facts are removed. The coordinator also purges expired facts.
Only a memory-role principal may change retained facts. Prompt injection labels
memory as untrusted evidence, never as authority. Intent and decision history
remain separately owned by the goal controller.

```bash
python -m orchestrator.reliability_ops memory-put --tenant example --key fact --value 'Reviewed fact' --source 'source-reference' --revision 0 --ttl 86400
python -m orchestrator.reliability_ops memory-list --tenant example
python -m orchestrator.reliability_ops memory-delete --tenant example --key fact
```

Memory retention cannot exceed the tenant profiles' configured limit. Live row
deletion does not erase previous backups; backup retention is a separate operator
responsibility. Secrets must not be stored as facts. Redaction is defense in depth,
not a guarantee that arbitrary sensitive personal information is detected.

Enforced workers run in Bubblewrap user/mount/PID namespaces, with capabilities
dropped, a new home and temporary directory, and only their worktree writable.
Host home, controller DB, other workspaces and inherited credentials are absent.
Git metadata is masked so workers cannot redirect later controller Git commands.
Repository tests and configured-command verifiers use the same boundary. Git
publication disables hooks/fsmonitor and refuses executable filters/included
config at any Git configuration level. Handoff files reject symlinks, hard links,
special files and oversized content before privileged reads or writes. A missing
or kernel-disabled sandbox is a failure, not a fallback.

Enforced profiles require network-off tools with no provider credentials. A
controller-owned agent loop calls the measured model adapter outside the tool
sandbox, then executes only structured `argv` tool proposals inside it. The model
cannot run a privileged host shell or connect to the operator dashboard. Tool
turns, request/output sizes and total attempt time are bounded; pause/cancel also
stops an active model adapter or tool process group. Complete observed usage
manifests are sealed on attempt completion; crashes remain unreconciled.

Adapter `output` for workers must be exactly `{"tool_calls":[{"argv":[...]}]}`
or `{"final":{"status":"complete","summary":"...","blocker_code":"none"}}`.
Planning adapters return a structured work-package plan instead. Provider-specific
adapters translate this request/response protocol; raw legacy CLI agents are not
quietly substituted in enforced mode. Install executable dependencies through
explicit read-only mounts. Keep provider credentials only in the trusted adapter's
named private environment, never argv. This is not a sandbox against a compromised host
kernel, malicious host administrator or malicious operator-owned adapter. Do not
run legacy unsandboxed review/deployment automation over untrusted artifacts;
dispatcher-only mode keeps those separate from this managed execution boundary.

Queue intake and execution select the profile's adapter, not whichever legacy
CLI happens to be installed. The adapter's name is reported as the worker
identity. Failures do not silently switch to an unqualified CLI. Reconcile any
uncertain billed request before changing providers or retrying.

## Evaluation And Release Gates

Every profile binds a controller Git SHA, evaluator, test artifacts, rollback
plan and incident playbook. Their hashes and the adapter registries form the
evidence fingerprint. Changing them invalidates approval. The executing checkout
must match the SHA and have no tracked modifications. A config string cannot
pretend to be the actual deployed release.

The evaluator is a bounded operator-owned executable. It returns numeric quality,
sample count and boolean results for named scenarios. Required scenarios include
end-to-end delivery, restart recovery, duplicate effects, denied actions, stale
revisions, provider failure, restore, isolation, approval expiry and the complete
measured-model/isolated-tool/independent-verification flow. Missing,
failed, timed-out or skipped scenarios cannot pass. `bin/evaluate_orchestration.py`
runs actual controller tests and emits an **engineering** score, not a business
quality score. Its providers are local fixtures and never enter live metrics.
Ordinary test runs may skip unavailable host isolation; the qualification report
maps that skip to a failed mandatory scenario, never an eligible release.
Add task-specific scenarios and reference datasets through the evaluator/artifact
contract when qualifying a business workflow.

```bash
python -m orchestrator.reliability_ops evaluate --tenant example --profile example-code
python -m orchestrator.reliability_ops approve --tenant example --profile example-code --evaluation RUN_ID --reason 'Reviewed evidence and recovery plan'
```

The approver must differ from the evaluation initiator. Only the latest passing,
fresh evaluation of unchanged artifacts can be approved. New failed evaluations
override older passes for eligibility. Quality loss from the approved baseline
blocks admission even when the absolute minimum still passes. Evidence expires.
Optional periodic evaluation uses the existing coordinator cadence and runs only
when explicitly configured; no cron is installed by this change. Running managed
workers recheck their release gate every five seconds.

## Service Commitments

The template gives separate code, research and external-action profiles explicit
success-rate and end-to-end latency targets, a rolling window, sample minimum and
memory retention. These values are configurable operating targets, not a signed
customer SLA. Define business quality and contractual obligations with the client.

Success uses independently accepted terminal tasks, not model completion claims.
Cancelled tasks and historical imports are excluded; the denominator and sample
count are visible. Latency includes waits. Open overdue work is an immediate
breach. Too few completed samples means insufficient data, not 100% reliability.
Service breaches stop admission of new goals. Already-started goals can finish or
recover within their original authority and budgets, rather than being trapped
forever by their own overdue status. Readiness transitions route through the
existing persistent incident router and its acknowledgment/escalation policy.

See [incident response](runbooks/incidents.md) and [recovery](runbooks/recovery.md).
Distributed fleet consensus, customer connectors/evals, host-specific egress
controls and sustained production evidence are not supplied by a generic library.
Loading
Loading