diff --git a/CHANGELOG.md b/CHANGELOG.md index d90fa0c899..9f2c9e0426 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -165,6 +165,7 @@ - Raised `hourly-review-repair.yml`'s discovery ceiling from 50 to 200 while rotating deterministic 50-PR deep-inspection windows by hourly run number. The scheduler hydrates only the selected window and stops immediately after its single dispatch, preserving access to newer PRs without quadrupling expensive review/check/comment work. See `docs/doctoring/hourly-review-repair-single-file-consolidation.md`'s 2026-09-03 follow-up. ## [Unreleased] +- **Keep OpenCode same-model checkpoint evidence bounded, provider-neutral, and authority-driven.** PR #2284 reads only declared log prefixes, counts required markers only from assistant text parts, rebuilds the base prompt before each retry, and creates/consumes continuation state only for `contextual-orchestrator/orchestrator/free`. Missing or negative calibrated budget authority fails closed and injects no appendix. Provider/model/status route details remain inside CO and cannot enter the leaf checkpoint or prompt; the mutable consumer-side CO parser and fixtures were removed. Explicit budgets remain Proposed until A/B, fast-mlsirm, and allocator evidence exist. - **Bind GitHub REST redirect evidence to both production opener chains.** `.github#2279` now feeds a synthetic same-authority 302 through the CodeQL identity and Strix evidence clients' real module-level openers, proving the redirect target is never contacted and the bearer header is never forwarded. Removing `_RejectRedirects` from either opener makes the contract fail on the forbidden second request. Four stale Strix HTTP/transport/JSON fixtures now patch that same production seam; direct handler unit cases and standalone CodeQL materialization remain unchanged. - **Define an evidence-backed repository README quality standard.** Added `docs/repository-readme-quality-standard.md` as the shared review contract for product-first structure, code-current onboarding, authority boundaries, durable quality signals, and repository/source/dependency license due diligence. Product repositories continue to own their own README prose; the standard is linked from the root documentation map and does not centralize or generate product claims. - Include merge-scheduler entrypoint, core, and regression-test changes in diff --git a/docs/doctoring/opencode-same-model-midabort-20260919.md b/docs/doctoring/opencode-same-model-midabort-20260919.md new file mode 100644 index 0000000000..44f2468444 --- /dev/null +++ b/docs/doctoring/opencode-same-model-midabort-20260919.md @@ -0,0 +1,143 @@ +# OpenCode same-model mid-abort and context loss (2026-09-19) + +## Cross-file Gap baseline preservation repair + +Documentation successor `33e86ea6becb754e6e3b8a98299640220ecf27d5` accidentally treated a truncated GitHub contents response as the complete `docs/product-technical-gap-baseline.md`: the diff was `+7/-1,505`, the literal truncation diagnostic entered the file, and protected level-two authority fell from 59 sections to 36. That was not an intentional supersession and was unrelated to the checkpoint runtime repair. + +RED `117c1bf07ab07b8a7262da3d2da9d1f437c36cf0` integrates the canonical-owner preservation contract already established on #2281 and fails on five missing representative security, runtime, compliance, APA, and credential-lifetime sections. GREEN `e55c34159b43bd32fa7e039d69faa2b5f0ab5814` reconstructs the exact `7c5844ad…` parent baseline from bounded line ranges, then applies only the three intended OpenCode control-row updates. The repaired file has 59/59 protected level-two sections, no truncation marker, one instance of every OpenCode control row, and a `+3/-3` baseline diff relative to `7c5844ad…`. + +This test is a cross-file authority guard, not checkpoint behavior evidence. It prevents a future documentation-only successor from erasing unrelated PRD/TRD, Context Map, security, runtime, compliance, or APA decisions while preserving the checkpoint lane's valid delta. + + +## Scope + +Improve OpenCode stopping mid-work or misunderstanding required outputs on +**same-model retries** under the pinned pool +`contextual-orchestrator/orchestrator/free` (`opencode.jsonc`), without model +changes. Out of scope: inflight same-head dispatch dedupe / pg-erd handshake +waste (`ContextualWisdomLab/.github#2283`). + +## Reproduction cases (live Actions) + +| Run ID | Repo | Termination | Context loss | Retry/resume | Completion judgment | +| --- | --- | --- | --- | --- | --- | +| `35452307646` | `.github` PR `#2040` | `cancelled` (`cancel-superseded-opencode-review-runs`) | In-flight review discarded on superseded head; no exported session checkpoint | Replacement dispatch queued; prior partial work not reused | Job `cancelled`, not success; no formal verdict on cancelled head | +| `35401977816` | `.github` PR `#2278` | `failure` (`opencode-review` job) | Model pool cycled without host checkpoint; partial assistant export dropped between attempts | Same-model retries restarted from full prompt only | `review_status=exhausted`; fail-closed, no synthetic APPROVE | + +Pinned model for measurement: `contextual-orchestrator/orchestrator/free` +(resolved upstream ids recorded from sidecar route evidence when present). + +## Reused surfaces + +- Provider-specific route evidence remains inside ContextualWisdomLab/contextual-orchestrator. + The leaf does not consume mutable `attempts[]`, provider, model, phase, or status + fields from open CO PR #1205; a released provider-neutral projection is required + before any owner telemetry can enter continuation control context. +- `scripts/ci/opencode_review_session_checkpoint.py` — host-managed checkpoint + ledger (digest-only partial work, missing required outputs, route telemetry). +- `scripts/ci/run_opencode_review_model_pool.sh` — injects bounded same-model + continuation appendix on retry only when explicit calibrated authority supplies + `OPENCODE_SESSION_CONTINUATION_BUDGET`; missing authority injects no appendix. + +## Before / after (pinned model, fixture-backed) + +| Metric | Before | After (this change) | +| --- | --- | --- | +| Same-model retry carries prior termination reason | No | Yes (checkpoint) | +| Same-model retry carries missing required outputs | No | Yes | +| Provider-specific route telemetry on retry | Discarded | Still discarded at the leaf; CO remains the observability owner | +| Continuation budget | Unbounded prompt replay cycles only | Missing authority fails closed; explicit budget path is tested but remains Proposed pending calibration | +| False success on incomplete control | Fail-closed already | Unchanged fail-closed | +| Partial provider body replayed into prompt | N/A | Forbidden (digest only) | + +Live completion-rate deltas require a controlled replay harness on org runners; +fixture tests prove the host contract; production measurement remains open. + +## Remaining + +- Executable fresh-session loop with durable SQLite ledger (`#2068`) still needs + read-only agent boundary preserved. +- Inflight dedupe cancellation waste (`#2283`) remains lead-owned. +- Production before/after completion rate on long tool-heavy reviews is not yet + measured on live `orchestrator/free` traffic. + + +## Exact-head review repair (2026-09-20) + +Independent review of `9fdfddfaa3d792ad0a2de4d1884df14ad999ffd5` +found four checkpoint-integrity defects: + +- bounded log helpers loaded the complete file before slicing; +- user-prompt markers could satisfy assistant-output requirements; +- later retries accumulated every earlier checkpoint appendix; +- checkpoint state and continuations applied to candidates outside + `contextual-orchestrator/orchestrator/free`. + +Test-only commits `cfeda192f68af43077cae3b3cd61ef8f11eaea20` and +`afe1420d3175d81f51c091b3499d792037cc304e` reproduce all four defects. +Fresh execution at the test-only head reported exactly **4 failed**. Minimal +source commits `36fb3907eaa229f8d2292feae3f25e010643693a` and +`b400ad5d993dcc8d5efcdbcb13aafa2f6d295d2c` bound file reads, filter only +assistant text parts, rebuild the base prompt on every retry, and scope +checkpoint read/write to the pinned orchestrator route. Fresh +warnings-as-errors execution passed **43 tests** across the checkpoint, route +evidence, and model-pool suites; Bash syntax, Python compilation, and diff +whitespace checks also passed. Hosted exact-head gates remain separately +required and no predecessor receipt transfers. + + +### Continuation budget authority repair (2026-09-20) + +Review of exact `bed37694c191f13bf18a595af672bb4d63e811af` found that +`DEFAULT_CONTINUATION_BUDGET = 2` and +`${OPENCODE_SESSION_CONTINUATION_BUDGET:-2}` silently selected two +decision-affecting continuations while the controlled A/B and allocator +evidence remained open. This violated the no-heuristics contract. + +Test-only commits `f0775fd48d31bf033279701f747698310aa6d7c6` and +`c4165352bab48123f9333bcfa11a78d060cb9cb0` require the runner and direct +CLI to reject missing budget authority. Minimal source commits +`3e290447b80c3964599bbb811f5cada2435ce28f` and +`4070c161685298a37554a551e7e3f33b2b6e49ad` remove both numeric defaults: +the runner injects no checkpoint appendix without an explicit non-negative +budget, and both CLI commands require `--budget`. + +Fresh materialization of exact `4070c161…` passed Python compilation, full +runner Bash syntax, and four direct authority probes (missing CLI authority +fails, the required argument is diagnosed, an explicit budget remains +accepted, and the shell numeric fallback is absent). The execution image does +not contain pytest, so no fresh pytest count is claimed. An explicit budget is +still Proposed rather than calibrated production authority until controlled +completion/time/token evidence and the selected fast-mlsirm/Fugu/Conductor/ +TRINITY-compatible allocator receipt are integrated. Hosted exact-head gates +and independent review remain required. + + +### Provider-neutral owner-boundary repair (2026-09-20) + +The existing P0 review showed that the leaf parser read the mutable CO #1205 +`model`, `provider_name`, phase, attempt number, and HTTP status fields and +formatted them into the next model prompt. That made two envelopes with the +same provider-neutral outcome produce different control context and duplicated +an unreleased owner schema inside ContextualWisdomLab/.github. + +RED `7c5c6a7548f5836d956cb196427581060951021b` and +`3a8c056bc2b4e0f5b012f900b24bfc54636e24ba` require provider/model/phase/ +status values to have zero effect on the appendix. GREEN +`866cc6a4a1285936da5110ea0a28071548da09e2` and +`8c04d1284f96eb0a702a24287290e30ba3f1bb9d` remove route parsing, storage, +formatting, CLI plumbing, and runner input. Commits +`51a531886962f35eb091b97cf8ec2d6e05dcec81` and +`b7e8256f4c8a9cf7cf1022312ee86ba6f985f9a4` retire the consumer-owned parser +and fixtures entirely. + +Fresh exact materialization at `7c5844ad…` (tree `53908068…`) passed Python +compilation, Bash syntax, checkpoint/runner **65 tests**, and the full +warnings-as-errors suite (**3,428 passed / 5 skipped / 40 subtests passed**). +The provider-neutral probe injects hostile legacy `openrouter`, phase, HTTP 429, +and served-model fields into a checkpoint and proves none reaches the +continuation appendix. RED `68459817…` additionally proves both direct CLI paths +previously accepted negative budget authority and exited zero; GREEN +`7c5844ad…` rejects it at the parser boundary. CO issue #1106 records the +required immutable, provider-neutral, allocator/fast-mlsirm receipt before any +owner telemetry or valid budget can be consumed. diff --git a/docs/product-technical-gap-baseline.md b/docs/product-technical-gap-baseline.md index 6e5f1c549a..abf45aff59 100644 --- a/docs/product-technical-gap-baseline.md +++ b/docs/product-technical-gap-baseline.md @@ -7,6 +7,16 @@ 이 문서는 제품·기술·운영 Gap을 현재 문서와 현재 GitHub 상태에 묶어 두는 기준선이다. 새 작업은 먼저 이 문서의 Gap ID를 PR 설명과 테스트 증거에 연결하고, PR의 정확한 exact HEAD·Checks·리뷰를 다시 수집한 뒤 구현한다. 표의 상태는 작성 시점의 관측값이므로, 병합 판단에는 재사용하지 않는다. 이 인벤토리는 스냅샷이며 merge authorization이 아니다. +### 2026-09-20 current-head incident delta + +| Gap ID | 상태 | exact-head evidence | causal owner / next gate | +|---|---|---|---| +| CONTROL-OPENCODE-CHECKPOINT-INTEGRITY-01 | **Source repaired on PR #2284; hosted exact-head acceptance pending** | Review of `#2284@9fdfddfa` found complete-file reads before slicing, user-prompt marker laundering, accumulated retry appendices, and checkpoint application outside `contextual-orchestrator/orchestrator/free`. Test-only `afe1420d` produced exactly 4 failures; source `b400ad5d` produced 43 focused warnings-as-errors passes. Later test expansion at `9012eac2` hid an arithmetically unreachable `used < 0` decision from coverage while its named test exercised only `used == 0`. RED `802a4fa5` makes that vacuous oracle executable; GREEN `0aa9b902` removes only the impossible clamp. Exact `7c5844ad` (tree `53908068`) passes checkpoint/runner **65 tests** and the full warnings-as-errors suite **3,428 passed / 5 skipped / 40 subtests passed**. | ContextualWisdomLab/.github owns the trusted OpenCode host checkpoint boundary. Keep #2284 Draft until fresh terminal hosted security/quality evidence and qualifying independent review exist; no predecessor result transfers. | +| CONTROL-OPENCODE-CONTINUATION-AUTHORITY-02 | **Missing or negative authority now fails closed on PR #2284; calibrated admission remains Proposed** | Exact `bed37694` silently selected budget `2` although controlled completion/time/token evidence was still pending. RED `f0775fd4` and `c4165352` require the runner and direct CLI to reject absent authority; GREEN `3e290447` and `4070c161` remove the Python and shell defaults. Exact-head RED `68459817` then proves both direct CLI paths accepted negative authority, exited zero, and emitted `0`; GREEN `7c5844ad` validates non-negative authority at the parser boundary. Focused **65 passed**, full suite **3,428 passed / 5 skipped / 40 subtests passed**, compileall, Bash syntax, and diff check bind the repair to tree `53908068`. | ContextualWisdomLab/.github owns host enforcement. A budget may be enabled only after a versioned fast-mlsirm/Fugu/Conductor/TRINITY-compatible allocator receipt and controlled A/B evidence are integrated; absent or invalid authority keeps checkpoint injection disabled. Hosted exact-head GREEN and independent review remain required. | +| CONTROL-OPENCODE-PROVIDER-NEUTRAL-03 | **Consumer schema copy removed on PR #2284; released CO projection pending** | Exact `bed37694` parsed CO model/provider/phase/status fields and formatted them into continuation prompts. RED `7c5c6a75`/`3a8c056b` requires byte-neutral handling of hostile provider details. GREEN `866cc6a4`/`8c04d128` removes route parsing and runner plumbing; `51a53188`/`b7e8256f` deletes the mutable parser and fixtures. Exact `7c5844ad` retains that provider-neutral boundary and passes checkpoint/runner **65 tests** plus the full warnings-as-errors suite **3,428 passed / 5 skipped / 40 subtests passed**. | ContextualWisdomLab/contextual-orchestrator issue #1106 owns the released provider-neutral allocation receipt. ContextualWisdomLab/.github must know only `orchestrator/free` and the gateway token; provider identities remain CO observability data. Keep Draft until immutable owner release/pin, exact-head GREEN, and independent review. | +| CONTROL-OPENCODE-CHECKPOINT-CAUSE-04 | **Source repaired on PR #2284; hosted exact-head acceptance pending** | Exact `3f05c1fc` proves two causal-context failures: provider-like words in assistant prose selected a trusted provider termination label, and an invalid control result reached the continuation as generic `nonzero-exit` although the host already held wrapper status `3`. GREEN `f57319d1` restricts classification to structured OpenCode `type=error` events, CLI stderr, and fixed host hints; GREEN `9dc05e7d` maps wrapper status `3` to `invalid-control-output`. Exact-source assertions confirm assistant export is no longer causal while export presence remains checked, and the runner continuation contract requires `Termination reason: invalid-control`. | ContextualWisdomLab/.github owns checkpoint cause identity. Keep #2284 Draft until exact-head Python/runner/security gates are terminal GREEN and an independent review qualifies; no assistant prose, prior head, or predecessor check may author the cause. | +| CONTROL-OPENCODE-CHECKPOINT-BOUNDS-05 | **Source repaired on PR #2284; hosted exact-head acceptance pending** | RED `c8680c56` executes a valid provider session export above the existing 2 MiB OpenCode evidence bound and proves that both checkpoint summary and required-output extraction consumed it in full. GREEN `5c49d4f8` introduces `MAX_SESSION_EXPORT_BYTES`, reads at most bound + 1 byte, rejects oversized or non-UTF-8 exports before JSON parsing, and replaces both unbounded `read_text()` paths. Exact remote Python compilation, the oversized/small-export behavior probe, and 31 directly executable checkpoint cases pass; three fixture-dependent cases were not represented as hosted evidence. | ContextualWisdomLab/.github owns the review-host memory boundary. Keep #2284 Draft until exact-head hosted quality/security gates and qualifying independent review are terminal; oversized provider artifacts cannot become continuation evidence. | + ### 2026-09-19 exact-head incident delta | Gap ID | 상태 | exact-head evidence | causal owner / next gate | diff --git a/scripts/ci/opencode_review_session_checkpoint.py b/scripts/ci/opencode_review_session_checkpoint.py new file mode 100644 index 0000000000..760e421346 --- /dev/null +++ b/scripts/ci/opencode_review_session_checkpoint.py @@ -0,0 +1,412 @@ +#!/usr/bin/env python3 +"""Host-managed OpenCode same-model session checkpoints and continuations. + +The trusted review host records partial attempt state and injects bounded +resume context on same-model retries. Checkpoints never grant approval +authority and never replay provider-controlled bodies into prompts. +""" + +from __future__ import annotations + +import argparse +import hashlib +import json +import re +import sys +from collections.abc import Mapping, Sequence +from pathlib import Path +from typing import Any + +_REPO_ROOT = Path(__file__).resolve().parents[2] +if str(_REPO_ROOT) not in sys.path: # pragma: no cover - bootstrap only on direct script execution + sys.path.insert(0, str(_REPO_ROOT)) + +CHECKPOINT_SCHEMA = 1 +CONTROL_SENTINEL = "opencode-review-control-v1" +REQUIRED_OUTPUT_MARKERS = ( + CONTROL_SENTINEL, + "adversarial_validation", + '"result"', + "Developer experience:", + "User experience:", +) +TERMINATION_PATTERNS: tuple[tuple[str, re.Pattern[str]], ...] = ( + ("provider-fatal", re.compile(r"contextoverflowerror|tokens_limit_reached|model_not_found|no endpoints", re.I)), + ("provider-timeout", re.compile(r"timed?\ ?out|timeout", re.I)), + ("provider-rate-limit", re.compile(r"rate.?limit|too many requests|\b429\b", re.I)), + ("provider-error", re.compile(r'"type"\s*:\s*"error"', re.I)), + ("export-empty", re.compile(r"assistant-empty-export", re.I)), + ("invalid-control", re.compile(r"invalid-control-output", re.I)), + ("sessionless", re.compile(r"sessionless-json", re.I)), + ("nonzero-exit", re.compile(r"exit", re.I)), +) +MAX_PARTIAL_DIGEST_CHARS = 64 +MAX_CONTINUATION_BYTES = 8192 +MAX_SESSION_EXPORT_BYTES = 2 * 1024 * 1024 + + +def _read_bounded_text(path: Path, max_bytes: int) -> str: + """Read at most ``max_bytes`` from a file as UTF-8 replacement text.""" + if not path.is_file(): + return "" + try: + with path.open("rb") as bounded_stream: + data = bounded_stream.read(max_bytes) + except OSError: + return "" + return data.decode("utf-8", errors="replace") + + +def _read_bounded_session_export(path: Path) -> str: + """Read a complete session export within the owner evidence byte bound.""" + if not path.is_file(): + return "" + try: + with path.open("rb") as bounded_stream: + data = bounded_stream.read(MAX_SESSION_EXPORT_BYTES + 1) + except OSError: + return "" + if len(data) > MAX_SESSION_EXPORT_BYTES: + return "" + try: + return data.decode("utf-8") + except UnicodeDecodeError: + return "" + + +def classify_termination( + *, + json_path: Path, + export_path: Path, + exit_code: int, + stderr_path: Path | None = None, + log_hint: str = "", +) -> str: + """Return a bounded termination reason for one OpenCode attempt.""" + structured_error_events: list[str] = [] + for line in _read_bounded_text(json_path, 65536).splitlines(): + try: + event = json.loads(line) + except json.JSONDecodeError: + continue + if isinstance(event, Mapping) and event.get("type") == "error": + structured_error_events.append(line) + combined = "\n".join( + part + for part in ( + "\n".join(structured_error_events), + _read_bounded_text(stderr_path, 65536) if stderr_path else "", + log_hint, + ) + if part + ) + for label, pattern in TERMINATION_PATTERNS: + if pattern.search(combined): + return label + if exit_code != 0: + return "nonzero-exit" + if not summarize_partial_assistant(export_path).get("assistant_text_present"): + return "export-empty" + return "incomplete-control" + + +def summarize_partial_assistant(export_path: Path) -> dict[str, str | int | bool]: + """Return bounded metadata about partial assistant output, never raw text.""" + encoded_export = _read_bounded_session_export(export_path) + if not encoded_export: + return { + "assistant_text_present": False, + "assistant_line_count": 0, + "assistant_sha256": "", + "has_control_sentinel": False, + } + try: + payload = json.loads(encoded_export) + except json.JSONDecodeError: + return { + "assistant_text_present": False, + "assistant_line_count": 0, + "assistant_sha256": "", + "has_control_sentinel": False, + } + texts: list[str] = [] + if isinstance(payload, dict): + messages = payload.get("messages") + if isinstance(messages, list): + for message in messages: + if not isinstance(message, dict): + continue + info = message.get("info") + if not isinstance(info, dict) or info.get("role") != "assistant": + continue + parts = message.get("parts") + if not isinstance(parts, list): + continue + for part in parts: + if isinstance(part, dict) and part.get("type") == "text": + text = part.get("text") + if isinstance(text, str) and text.strip(): + texts.append(text) + joined = "\n".join(texts) + digest = hashlib.sha256(joined.encode("utf-8")).hexdigest() if joined else "" + return { + "assistant_text_present": bool(joined.strip()), + "assistant_line_count": len(joined.splitlines()) if joined else 0, + "assistant_sha256": digest[:MAX_PARTIAL_DIGEST_CHARS], + "has_control_sentinel": CONTROL_SENTINEL in joined, + } + + +def missing_required_outputs(partial_text: str) -> list[str]: + """Return required review output markers absent from partial assistant text.""" + missing: list[str] = [] + for marker in REQUIRED_OUTPUT_MARKERS: + if marker not in partial_text: + missing.append(marker) + return missing + + +def _load_checkpoint(path: Path) -> dict[str, Any]: + """Load an existing checkpoint or return an empty document.""" + if not path.is_file(): + return {"schema": CHECKPOINT_SCHEMA, "attempts": []} + try: + loaded = json.loads(path.read_text(encoding="utf-8")) + except (OSError, json.JSONDecodeError, UnicodeDecodeError): + return {"schema": CHECKPOINT_SCHEMA, "attempts": []} + if not isinstance(loaded, dict): + return {"schema": CHECKPOINT_SCHEMA, "attempts": []} + attempts = loaded.get("attempts") + if not isinstance(attempts, list): + loaded["attempts"] = [] + loaded.setdefault("schema", CHECKPOINT_SCHEMA) + return loaded + + +def record_attempt_checkpoint( + *, + checkpoint_path: Path, + model_candidate: str, + attempt: int, + head_sha: str, + run_id: str, + run_attempt: str, + json_path: Path, + export_path: Path, + exit_code: int, + stderr_path: Path | None = None, + log_hint: str = "", +) -> dict[str, Any]: + """Append one bounded attempt record to the host checkpoint ledger.""" + partial_summary = summarize_partial_assistant(export_path) + partial_text = "" + encoded_export = _read_bounded_session_export(export_path) + if encoded_export: + try: + payload = json.loads(encoded_export) + if isinstance(payload, dict): + messages = payload.get("messages") + if isinstance(messages, list): + chunks: list[str] = [] + for message in messages: + if not isinstance(message, dict): + continue + info = message.get("info") + if not isinstance(info, dict) or info.get("role") != "assistant": + continue + parts = message.get("parts") + if not isinstance(parts, list): + continue + for part in parts: + if ( + isinstance(part, dict) + and part.get("type") == "text" + and isinstance(part.get("text"), str) + ): + chunks.append(part["text"]) + partial_text = "\n".join(chunks) + except json.JSONDecodeError: + partial_text = "" + entry = { + "attempt": attempt, + "model_candidate": model_candidate, + "head_sha": head_sha, + "run_id": run_id, + "run_attempt": run_attempt, + "termination_reason": classify_termination( + json_path=json_path, + export_path=export_path, + exit_code=exit_code, + stderr_path=stderr_path, + log_hint=log_hint, + ), + "exit_code": exit_code, + "partial_summary": partial_summary, + "missing_required_outputs": missing_required_outputs(partial_text), + } + document = _load_checkpoint(checkpoint_path) + attempts = document.setdefault("attempts", []) + if isinstance(attempts, list): + attempts.append(entry) + document["pinned_model"] = model_candidate + document["head_sha"] = head_sha + checkpoint_path.parent.mkdir(parents=True, exist_ok=True) + checkpoint_path.write_text( + json.dumps(document, indent=2, sort_keys=True) + "\n", + encoding="utf-8", + ) + return entry + + +def continuation_budget_remaining(checkpoint_path: Path, *, budget: int) -> int: + """Return how many same-model continuations remain for this checkpoint.""" + document = _load_checkpoint(checkpoint_path) + attempts = document.get("attempts") + used_continuations = ( + len(attempts) - 1 if isinstance(attempts, list) and attempts else 0 + ) + return max(0, budget - used_continuations) + + +def build_continuation_appendix(checkpoint_path: Path, *, budget: int) -> str: + """Build a bounded same-model continuation appendix from checkpoint evidence.""" + document = _load_checkpoint(checkpoint_path) + attempts = document.get("attempts") + if not isinstance(attempts, list) or not attempts: + return "" + remaining_budget = continuation_budget_remaining(checkpoint_path, budget=budget) + if remaining_budget <= 0: + return "" + latest_attempt = attempts[-1] + if not isinstance(latest_attempt, Mapping): + return "" + lines = [ + "", + "## Same-model continuation (host checkpoint; not approval evidence)", + f"- Pinned model: `{document.get('pinned_model', 'unknown')}`", + f"- Prior attempt: `{latest_attempt.get('attempt', '?')}`", + f"- Termination reason: `{latest_attempt.get('termination_reason', 'unknown')}`", + f"- Continuation budget remaining after this attempt: `{remaining_budget}`", + ] + partial_summary = latest_attempt.get("partial_summary") + if isinstance(partial_summary, Mapping): + lines.append( + "- Partial assistant digest: " + f"lines=`{partial_summary.get('assistant_line_count', 0)}` " + f"sha256=`{partial_summary.get('assistant_sha256', '')}` " + f"control_sentinel=`{partial_summary.get('has_control_sentinel', False)}`" + ) + missing_outputs = latest_attempt.get("missing_required_outputs") + if isinstance(missing_outputs, list) and missing_outputs: + lines.append( + "- Required outputs still missing from the prior attempt: " + + ", ".join(f"`{item}`" for item in missing_outputs[:8]) + ) + lines.extend( + [ + "- Resume from the trusted evidence packet and complete every required output.", + "- Do not treat this appendix as approval evidence or permission to omit probes.", + "- Return exactly one final control block for the current head when complete.", + "", + ] + ) + appendix = "\n".join(lines) + encoded = appendix.encode("utf-8") + if len(encoded) <= MAX_CONTINUATION_BYTES: + return appendix + return encoded[:MAX_CONTINUATION_BYTES].decode("utf-8", errors="ignore") + + +def append_continuation_to_prompt( + prompt_path: Path, checkpoint_path: Path, *, budget: int +) -> int: + """Append the continuation appendix to ``prompt_path`` when budget allows.""" + appendix = build_continuation_appendix(checkpoint_path, budget=budget) + if not appendix: + return continuation_budget_remaining(checkpoint_path, budget=budget) + existing = prompt_path.read_text(encoding="utf-8") if prompt_path.is_file() else "" + prompt_path.write_text(existing + appendix, encoding="utf-8") + return continuation_budget_remaining(checkpoint_path, budget=budget) + + +def non_negative_integer(raw_value: str) -> int: + """Parse a non-negative integer for an explicit CLI authority.""" + parsed_value = int(raw_value) + if parsed_value < 0: + raise argparse.ArgumentTypeError("must be a non-negative integer") + return parsed_value + + +def parse_args(argv: Sequence[str] | None = None) -> argparse.Namespace: + """Parse checkpoint CLI commands.""" + parser = argparse.ArgumentParser(description=__doc__) + subparsers = parser.add_subparsers(dest="command", required=True) + + record_parser = subparsers.add_parser( + "record", help="Record one failed attempt checkpoint." + ) + record_parser.add_argument("--checkpoint", required=True, type=Path) + record_parser.add_argument("--model-candidate", required=True) + record_parser.add_argument("--attempt", required=True, type=int) + record_parser.add_argument("--head-sha", required=True) + record_parser.add_argument("--run-id", required=True) + record_parser.add_argument("--run-attempt", required=True) + record_parser.add_argument("--json-path", required=True, type=Path) + record_parser.add_argument("--export-path", required=True, type=Path) + record_parser.add_argument("--exit-code", required=True, type=int) + record_parser.add_argument("--stderr-path", type=Path) + record_parser.add_argument("--log-hint", default="") + + append_parser = subparsers.add_parser( + "append-continuation", help="Append a bounded continuation appendix to a prompt." + ) + append_parser.add_argument("--prompt", required=True, type=Path) + append_parser.add_argument("--checkpoint", required=True, type=Path) + append_parser.add_argument("--budget", required=True, type=non_negative_integer) + + budget_parser = subparsers.add_parser( + "budget-remaining", help="Print remaining same-model continuation budget." + ) + budget_parser.add_argument("--checkpoint", required=True, type=Path) + budget_parser.add_argument("--budget", required=True, type=non_negative_integer) + return parser.parse_args(argv) + + +def main(argv: Sequence[str] | None = None) -> int: + """Execute one checkpoint subcommand.""" + args = parse_args(argv) + if args.command == "record": + entry = record_attempt_checkpoint( + checkpoint_path=args.checkpoint, + model_candidate=args.model_candidate, + attempt=args.attempt, + head_sha=args.head_sha, + run_id=args.run_id, + run_attempt=args.run_attempt, + json_path=args.json_path, + export_path=args.export_path, + exit_code=args.exit_code, + stderr_path=args.stderr_path, + log_hint=args.log_hint, + ) + print( + json.dumps( + {"termination_reason": entry["termination_reason"]}, + separators=(",", ":"), + ) + ) + return 0 + if args.command == "append-continuation": + remaining_budget = append_continuation_to_prompt( + args.prompt, args.checkpoint, budget=args.budget + ) + print(remaining_budget) + return 0 + if args.command == "budget-remaining": + print(continuation_budget_remaining(args.checkpoint, budget=args.budget)) + return 0 + raise SystemExit(f"unknown command: {args.command}") + + +if __name__ == "__main__": # pragma: no cover - exercised via subprocess in tests + raise SystemExit(main()) diff --git a/scripts/ci/run_opencode_review_model_pool.sh b/scripts/ci/run_opencode_review_model_pool.sh index 80f57d1d43..0737bc68b3 100644 --- a/scripts/ci/run_opencode_review_model_pool.sh +++ b/scripts/ci/run_opencode_review_model_pool.sh @@ -190,6 +190,54 @@ PY } >"$prompt_file" } +record_session_checkpoint() { + local model_candidate="$1" + local attempt="$2" + local checkpoint_file="$3" + local json_file="$4" + local export_file="$5" + local exit_code="$6" + local stderr_file="$7" + local log_hint="" + if [ "$exit_code" -eq 3 ]; then + log_hint="invalid-control-output" + fi + PYTHONPATH="$GITHUB_WORKSPACE${PYTHONPATH:+:$PYTHONPATH}" python3 "$GITHUB_WORKSPACE/scripts/ci/opencode_review_session_checkpoint.py" record \ + --checkpoint "$checkpoint_file" \ + --model-candidate "$model_candidate" \ + --attempt "$attempt" \ + --head-sha "$HEAD_SHA" \ + --run-id "$RUN_ID" \ + --run-attempt "$RUN_ATTEMPT" \ + --json-path "$json_file" \ + --export-path "$export_file" \ + --exit-code "$exit_code" \ + --log-hint "$log_hint" \ + ${stderr_file:+--stderr-path "$stderr_file"} \ + || true +} + +append_same_model_continuation() { + local prompt_file="$1" + local checkpoint_file="$2" + local budget="${OPENCODE_SESSION_CONTINUATION_BUDGET:-}" + local remaining + + if ! is_non_negative_integer "$budget"; then + printf 'OpenCode continuation budget authority is not configured; checkpoint appendix injection fails closed.\n' + return 1 + fi + remaining="$(python3 "$GITHUB_WORKSPACE/scripts/ci/opencode_review_session_checkpoint.py" append-continuation \ + --prompt "$prompt_file" \ + --checkpoint "$checkpoint_file" \ + --budget "$budget" 2>/dev/null || printf '0\n')" + if [ "${remaining:-0}" -le 0 ]; then + return 1 + fi + printf 'OpenCode same-model continuation appended; budget remaining after this attempt=%s.\n' "$remaining" + return 0 +} + write_schema_repair_prompt() { local model_candidate="$1" local prompt_file="$2" @@ -536,6 +584,7 @@ main() { candidate_output_file="${RUNNER_TEMP}/opencode-review-${safe_model}.md" opencode_json_file="${candidate_output_file}.jsonl" opencode_export_file="${candidate_output_file}.session.json" + checkpoint_file="${RUNNER_TEMP}/opencode-checkpoint-${safe_model}.json" write_prompt "$model_candidate" "$prompt_file" effective_attempts="$attempts" if is_schema_repair_candidate "$model_candidate"; then @@ -546,6 +595,18 @@ main() { write_schema_repair_prompt "$model_candidate" "$prompt_file" printf 'OpenCode %s schema-repair attempt %s/%s will re-review from trusted evidence with a non-replayable control checklist.\n' \ "$model_candidate" "$attempt" "$effective_attempts" + elif [ "$attempt" -gt 1 ]; then + write_prompt "$model_candidate" "$prompt_file" + if [ "$model_candidate" = "contextual-orchestrator/orchestrator/free" ] \ + && [ -f "$checkpoint_file" ]; then + if append_same_model_continuation "$prompt_file" "$checkpoint_file"; then + printf 'OpenCode %s same-model continuation attempt %s/%s reuses host checkpoint evidence under contextual-orchestrator/orchestrator/free.\n' \ + "$model_candidate" "$attempt" "$effective_attempts" + else + printf 'OpenCode %s same-model continuation budget exhausted; retry proceeds without checkpoint appendix.\n' \ + "$model_candidate" + fi + fi fi if [ "$max_total_attempts" -gt 0 ] && [ "$total_attempts" -ge "$max_total_attempts" ]; then printf 'OpenCode model pool reached the per-run provider attempt ceiling of %s attempts; ending the pool to bound provider spend. Set OPENCODE_POOL_MAX_TOTAL_ATTEMPTS=0 to disable.\n' "$max_total_attempts" @@ -569,6 +630,12 @@ main() { else run_status=$? fi + if [ "$model_candidate" = "contextual-orchestrator/orchestrator/free" ] \ + && [ "$run_status" -ne 0 ] && [ "$run_status" -ne 2 ]; then + record_session_checkpoint "$model_candidate" "$attempt" "$checkpoint_file" \ + "$opencode_json_file" "$opencode_export_file" "$run_status" \ + "${opencode_json_file}.stderr" + fi if [ "$run_status" -ne 3 ] && is_credit_exhausted_failure "$opencode_json_file" "${opencode_json_file}.stderr"; then dead_candidate_reasons[$model_candidate]="provider credits exhausted (HTTP 402 / payment required)" printf 'OpenCode %s provider credits are exhausted; marking this candidate failed for the rest of the run so retries cannot accrue further spend.\n' "$model_candidate" diff --git a/tests/test_opencode_model_pool_runner.py b/tests/test_opencode_model_pool_runner.py index 2965d4c55c..2d50f7800e 100644 --- a/tests/test_opencode_model_pool_runner.py +++ b/tests/test_opencode_model_pool_runner.py @@ -981,3 +981,146 @@ def test_paid_provider_does_not_gain_an_implicit_schema_repair_attempt( assert "attempt 1/1" in result.stdout assert "schema-repair attempt" not in result.stdout assert "attempt 2/" not in result.stdout + + +def test_non_orchestrator_candidate_does_not_receive_checkpoint_continuation( + tmp_path: Path, +) -> None: + """Checkpoint continuation remains scoped to the pinned orchestrator/free route.""" + prompt_capture = tmp_path / "prompt.md" + result = run_failed_model( + tmp_path, + json_line='{"type":"step_start","sessionID":"session-1"}', + model_candidates="github-models/openai/gpt-5", + prompt_capture=prompt_capture, + extra_env={ + "FAKE_OPENCODE_RUN_EXIT": "0", + "FAKE_OPENCODE_EXPORT": json.dumps( + { + "messages": [ + { + "info": {"role": "assistant"}, + "parts": [{"type": "text", "text": "partial review"}], + } + ] + } + ), + "OPENCODE_MODEL_ATTEMPTS": "2", + "OPENCODE_BACKOFF_INITIAL_SECONDS": "0", + "OPENCODE_POOL_MAX_CYCLES": "1", + "OPENCODE_POOL_CYCLE_SLEEP_SECONDS": "0", + }, + ) + + assert result.returncode == 1 + assert "same-model continuation attempt" not in result.stdout + assert "Same-model continuation (host checkpoint" not in prompt_capture.read_text( + encoding="utf-8" + ) + + +def test_latest_retry_replaces_prior_checkpoint_appendix(tmp_path: Path) -> None: + """Each retry prompt carries one latest checkpoint appendix, not accumulated history.""" + prompt_capture = tmp_path / "prompt.md" + result = run_failed_model( + tmp_path, + json_line='{"type":"step_start","sessionID":"session-1"}', + model_candidates="contextual-orchestrator/orchestrator/free", + prompt_capture=prompt_capture, + extra_env={ + "FAKE_OPENCODE_RUN_EXIT": "0", + "FAKE_OPENCODE_EXPORT": json.dumps( + { + "messages": [ + { + "info": {"role": "assistant"}, + "parts": [{"type": "text", "text": "partial review"}], + } + ] + } + ), + "OPENCODE_MODEL_ATTEMPTS": "3", + "OPENCODE_BACKOFF_INITIAL_SECONDS": "0", + "OPENCODE_POOL_MAX_CYCLES": "1", + "OPENCODE_POOL_CYCLE_SLEEP_SECONDS": "0", + "OPENCODE_SESSION_CONTINUATION_BUDGET": "3", + }, + ) + + assert result.returncode == 1 + assert prompt_capture.read_text(encoding="utf-8").count( + "Same-model continuation (host checkpoint" + ) == 1 + + +def test_missing_continuation_budget_fails_closed_without_appendix(tmp_path: Path) -> None: + """Missing calibrated budget authority cannot select a continuation count.""" + prompt_capture = tmp_path / "prompt.md" + result = run_failed_model( + tmp_path, + json_line='{"type":"step_start","sessionID":"session-1"}', + model_candidates="contextual-orchestrator/orchestrator/free", + prompt_capture=prompt_capture, + extra_env={ + "FAKE_OPENCODE_RUN_EXIT": "0", + "FAKE_OPENCODE_EXPORT": json.dumps( + { + "messages": [ + { + "info": {"role": "assistant"}, + "parts": [{"type": "text", "text": "partial review"}], + } + ] + } + ), + "OPENCODE_MODEL_ATTEMPTS": "2", + "OPENCODE_BACKOFF_INITIAL_SECONDS": "0", + "OPENCODE_POOL_MAX_CYCLES": "1", + "OPENCODE_POOL_CYCLE_SLEEP_SECONDS": "0", + }, + ) + + assert result.returncode == 1 + assert "continuation budget authority is not configured" in result.stdout + assert "Same-model continuation (host checkpoint" not in prompt_capture.read_text( + encoding="utf-8" + ) + + +def test_same_model_retry_appends_host_checkpoint_continuation(tmp_path: Path) -> None: + """A second same-model attempt reuses bounded checkpoint context instead of restarting blind.""" + prompt_capture = tmp_path / "prompt.md" + result = run_failed_model( + tmp_path, + json_line='{"type":"step_start","sessionID":"session-1"}', + model_candidates="contextual-orchestrator/orchestrator/free", + prompt_capture=prompt_capture, + extra_env={ + "FAKE_OPENCODE_RUN_EXIT": "0", + "FAKE_OPENCODE_EXPORT": json.dumps( + { + "messages": [ + { + "info": {"role": "assistant"}, + "parts": [ + {"type": "text", "text": "partial review without control block"} + ], + } + ] + } + ), + "OPENCODE_MODEL_ATTEMPTS": "2", + "OPENCODE_BACKOFF_INITIAL_SECONDS": "0", + "OPENCODE_POOL_MAX_CYCLES": "1", + "OPENCODE_POOL_CYCLE_SLEEP_SECONDS": "0", + "OPENCODE_SESSION_CONTINUATION_BUDGET": "2", + }, + ) + + assert result.returncode == 1 + assert "same-model continuation attempt 2/2" in result.stdout + prompt_text = prompt_capture.read_text(encoding="utf-8") + assert "Same-model continuation (host checkpoint" in prompt_text + assert "partial review without control block" not in prompt_text + assert "contextual-orchestrator/orchestrator/free" in prompt_text + assert "Termination reason: `invalid-control`" in prompt_text diff --git a/tests/test_opencode_review_session_checkpoint.py b/tests/test_opencode_review_session_checkpoint.py new file mode 100644 index 0000000000..e361f4cdb6 --- /dev/null +++ b/tests/test_opencode_review_session_checkpoint.py @@ -0,0 +1,831 @@ +"""Tests for host-managed OpenCode same-model session checkpoints.""" + +from __future__ import annotations + +import inspect +import json +import os +import subprocess +import sys +from pathlib import Path + +from scripts.ci.opencode_review_session_checkpoint import ( + append_continuation_to_prompt, + build_continuation_appendix, + classify_termination, + continuation_budget_remaining, + missing_required_outputs, + record_attempt_checkpoint, + summarize_partial_assistant, +) + + +def _export_with_text(text: str) -> str: + return json.dumps( + { + "messages": [ + { + "info": {"role": "assistant"}, + "parts": [{"type": "text", "text": text}], + } + ] + } + ) + + +def test_summarize_partial_assistant_never_returns_raw_text(tmp_path: Path) -> None: + """Checkpoint metadata stays digest-only.""" + export_path = tmp_path / "export.json" + export_path.write_text( + _export_with_text("secret partial body\nopencode-review-control-v1"), + encoding="utf-8", + ) + summary = summarize_partial_assistant(export_path) + assert summary["assistant_text_present"] is True + assert summary["has_control_sentinel"] is True + assert "secret" not in json.dumps(summary) + + +def test_oversized_session_export_fails_closed_without_full_parse( + tmp_path: Path, +) -> None: + """Provider session exports above the owner evidence bound are not parsed.""" + export_path = tmp_path / "oversized-export.json" + oversized_text = "x" * (2 * 1024 * 1024 + 1) + "opencode-review-control-v1" + export_path.write_text(_export_with_text(oversized_text), encoding="utf-8") + + summary = summarize_partial_assistant(export_path) + entry = record_attempt_checkpoint( + checkpoint_path=tmp_path / "checkpoint.json", + model_candidate="contextual-orchestrator/orchestrator/free", + attempt=1, + head_sha="a" * 40, + run_id="1", + run_attempt="1", + json_path=tmp_path / "missing.jsonl", + export_path=export_path, + exit_code=1, + ) + + assert summary["assistant_text_present"] is False + assert entry["partial_summary"]["assistant_text_present"] is False + assert entry["missing_required_outputs"] == [ + "opencode-review-control-v1", + "adversarial_validation", + '"result"', + "Developer experience:", + "User experience:", + ] + + +def test_missing_required_outputs_lists_absent_markers() -> None: + """Incomplete control output records which contract markers are still missing.""" + missing = missing_required_outputs("partial progress only") + assert "opencode-review-control-v1" in missing + assert "adversarial_validation" in missing + complete = "\n".join( + [ + "opencode-review-control-v1", + "adversarial_validation", + '"result"', + "Developer experience:", + "User experience:", + ] + ) + assert missing_required_outputs(complete) == [] + + +def test_classify_termination_detects_provider_fatal(tmp_path: Path) -> None: + """Fatal provider signatures classify as bounded termination reasons.""" + json_path = tmp_path / "run.jsonl" + json_path.write_text( + '{"type":"error","error":{"name":"ContextOverflowError","data":{}}}\n', + encoding="utf-8", + ) + reason = classify_termination( + json_path=json_path, + export_path=tmp_path / "missing.json", + exit_code=1, + ) + assert reason == "provider-fatal" + + +def test_classify_termination_ignores_assistant_prose_for_cause( + tmp_path: Path, +) -> None: + """Assistant prose cannot author a trusted provider termination label.""" + json_path = tmp_path / "run.jsonl" + json_path.write_text( + '{"type":"step_start","sessionID":"session-1"}\n', + encoding="utf-8", + ) + export_path = tmp_path / "export.json" + export_path.write_text( + _export_with_text( + "Review finding mentions ContextOverflowError, timeout, rate limit, " + "model_not_found, and invalid-control-output." + ), + encoding="utf-8", + ) + + reason = classify_termination( + json_path=json_path, + export_path=export_path, + exit_code=3, + log_hint="invalid-control-output", + ) + + assert reason == "invalid-control" + + +def test_record_and_continue_excludes_provider_identity_from_prompt(tmp_path: Path) -> None: + """Leaf continuation prompts cannot consume provider-specific route telemetry.""" + export_path = tmp_path / "export.json" + export_path.write_text(_export_with_text("in progress"), encoding="utf-8") + checkpoint_path = tmp_path / "checkpoint.json" + record_attempt_checkpoint( + checkpoint_path=checkpoint_path, + model_candidate="contextual-orchestrator/orchestrator/free", + attempt=1, + head_sha="a" * 40, + run_id="35401977816", + run_attempt="1", + json_path=tmp_path / "run.jsonl", + export_path=export_path, + exit_code=1, + ) + document = json.loads(checkpoint_path.read_text(encoding="utf-8")) + document["attempts"][-1]["route_telemetry"] = { + "provider_attempt_count": 1, + "provider_name": "openrouter", + "upstream_phase": "connecting", + "upstream_status": 429, + "served_model": "orchestrator/free", + } + checkpoint_path.write_text(json.dumps(document), encoding="utf-8") + appendix = build_continuation_appendix(checkpoint_path, budget=2) + assert "contextual-orchestrator/orchestrator/free" in appendix + assert "termination reason" in appendix.casefold() + assert "route evidence" not in appendix.casefold() + assert "openrouter" not in appendix.casefold() + assert "connecting" not in appendix.casefold() + assert "429" not in appendix + assert "in progress" not in appendix + + +def test_append_continuation_respects_budget(tmp_path: Path) -> None: + """Continuation budget is explicit and fails closed when exhausted.""" + export_path = tmp_path / "export.json" + export_path.write_text(_export_with_text("partial"), encoding="utf-8") + checkpoint_path = tmp_path / "checkpoint.json" + prompt_path = tmp_path / "prompt.md" + prompt_path.write_text("base prompt\n", encoding="utf-8") + for attempt in (1, 2, 3): + record_attempt_checkpoint( + checkpoint_path=checkpoint_path, + model_candidate="contextual-orchestrator/orchestrator/free", + attempt=attempt, + head_sha="b" * 40, + run_id="35452307646", + run_attempt="1", + json_path=tmp_path / f"run-{attempt}.jsonl", + export_path=export_path, + exit_code=1, + ) + assert continuation_budget_remaining(checkpoint_path, budget=2) == 0 + before = prompt_path.read_text(encoding="utf-8") + remaining = append_continuation_to_prompt(prompt_path, checkpoint_path, budget=2) + assert remaining == 0 + assert prompt_path.read_text(encoding="utf-8") == before + + +def test_classify_termination_covers_remaining_branches(tmp_path: Path) -> None: + """Every bounded termination class is reachable from host evidence.""" + export_path = tmp_path / "export.json" + export_path.write_text(_export_with_text("partial only"), encoding="utf-8") + assert ( + classify_termination( + json_path=tmp_path / "missing.jsonl", + export_path=export_path, + exit_code=0, + ) + == "incomplete-control" + ) + stderr = tmp_path / "stderr.txt" + stderr.write_text("request timed out", encoding="utf-8") + assert ( + classify_termination( + json_path=tmp_path / "missing.jsonl", + export_path=tmp_path / "missing-export.json", + exit_code=1, + stderr_path=stderr, + ) + == "provider-timeout" + ) + + +def test_summarize_partial_assistant_handles_malformed_export(tmp_path: Path) -> None: + """Malformed exports fail closed to empty metadata.""" + missing = summarize_partial_assistant(tmp_path / "missing.json") + assert missing["assistant_text_present"] is False + bad = tmp_path / "bad.json" + bad.write_text("{", encoding="utf-8") + assert summarize_partial_assistant(bad)["assistant_text_present"] is False + weird = tmp_path / "weird.json" + weird.write_text( + json.dumps({"messages": ["not-a-dict", {"info": "x", "parts": "y"}]}), + encoding="utf-8", + ) + assert summarize_partial_assistant(weird)["assistant_text_present"] is False + skipped = tmp_path / "skipped.json" + skipped.write_text( + json.dumps( + { + "messages": [ + {"info": {"role": "user"}, "parts": [{"type": "text", "text": "x"}]}, + {"info": {"role": "assistant"}, "parts": [{"type": "text", "text": " "}]}, + {"info": {"role": "assistant"}, "parts": "not-a-list"}, + ] + } + ), + encoding="utf-8", + ) + assert summarize_partial_assistant(skipped)["assistant_text_present"] is False + no_messages = tmp_path / "no-messages.json" + no_messages.write_text(json.dumps({"messages": "not-a-list"}), encoding="utf-8") + assert summarize_partial_assistant(no_messages)["assistant_text_present"] is False + non_dict = tmp_path / "non-dict.json" + non_dict.write_text(json.dumps(["not-a-dict"]), encoding="utf-8") + assert summarize_partial_assistant(non_dict)["assistant_text_present"] is False + + +def test_load_checkpoint_repairs_invalid_documents(tmp_path: Path) -> None: + """Invalid checkpoint files reset to an empty ledger.""" + from scripts.ci.opencode_review_session_checkpoint import _load_checkpoint + + path = tmp_path / "checkpoint.json" + path.write_text("[]", encoding="utf-8") + loaded = _load_checkpoint(path) + assert loaded["attempts"] == [] + path.write_text("{", encoding="utf-8") + assert _load_checkpoint(path)["schema"] == 1 + path.write_text(json.dumps({"attempts": "not-a-list"}), encoding="utf-8") + assert _load_checkpoint(path)["attempts"] == [] + + +def test_record_attempt_checkpoint_repairs_corrupt_history(tmp_path: Path) -> None: + """Corrupt attempt history is replaced instead of crashing the host ledger.""" + export_path = tmp_path / "export.json" + export_path.write_text(_export_with_text("partial"), encoding="utf-8") + checkpoint_path = tmp_path / "checkpoint.json" + checkpoint_path.write_text(json.dumps({"attempts": "bad-history"}), encoding="utf-8") + record_attempt_checkpoint( + checkpoint_path=checkpoint_path, + model_candidate="contextual-orchestrator/orchestrator/free", + attempt=1, + head_sha="g" * 40, + run_id="1", + run_attempt="1", + json_path=tmp_path / "run.jsonl", + export_path=export_path, + exit_code=1, + ) + document = json.loads(checkpoint_path.read_text(encoding="utf-8")) + assert len(document["attempts"]) == 1 + + +def test_build_continuation_appendix_truncates_large_payload(tmp_path: Path) -> None: + """Continuation appendix stays within the configured byte budget.""" + export_path = tmp_path / "export.json" + export_path.write_text(_export_with_text("partial"), encoding="utf-8") + checkpoint_path = tmp_path / "checkpoint.json" + record_attempt_checkpoint( + checkpoint_path=checkpoint_path, + model_candidate="contextual-orchestrator/orchestrator/free", + attempt=1, + head_sha="d" * 40, + run_id="1", + run_attempt="1", + json_path=tmp_path / "run.jsonl", + export_path=export_path, + exit_code=1, + ) + document = json.loads(checkpoint_path.read_text(encoding="utf-8")) + document["attempts"][0]["termination_reason"] = "x" * 9000 + document["attempts"][0]["missing_required_outputs"] = [ + f"marker-{index}-{'x' * 200}" for index in range(200) + ] + checkpoint_path.write_text(json.dumps(document), encoding="utf-8") + appendix = build_continuation_appendix(checkpoint_path, budget=5) + assert 0 < len(appendix.encode("utf-8")) <= 8192 + + +def test_classify_termination_export_empty(tmp_path: Path) -> None: + """Missing assistant export is classified separately from incomplete control.""" + assert ( + classify_termination( + json_path=tmp_path / "run.jsonl", + export_path=tmp_path / "missing.json", + exit_code=0, + ) + == "export-empty" + ) + + +def test_record_attempt_checkpoint_handles_non_list_parts(tmp_path: Path) -> None: + """Partial text extraction skips malformed message parts safely.""" + export_path = tmp_path / "export.json" + export_path.write_text( + json.dumps( + { + "messages": [ + {"parts": "not-a-list"}, + {"info": {"role": "assistant"}, "parts": [{"text": "x"}]}, + ] + } + ), + encoding="utf-8", + ) + checkpoint_path = tmp_path / "checkpoint.json" + entry = record_attempt_checkpoint( + checkpoint_path=checkpoint_path, + model_candidate="contextual-orchestrator/orchestrator/free", + attempt=1, + head_sha="f" * 40, + run_id="1", + run_attempt="1", + json_path=tmp_path / "run.jsonl", + export_path=export_path, + exit_code=1, + ) + assert entry["missing_required_outputs"] + + +def test_record_attempt_checkpoint_handles_non_dict_export(tmp_path: Path) -> None: + """Non-object export payloads do not leak partial text into checkpoints.""" + export_path = tmp_path / "export.json" + export_path.write_text(json.dumps(["not-an-object"]), encoding="utf-8") + checkpoint_path = tmp_path / "checkpoint.json" + entry = record_attempt_checkpoint( + checkpoint_path=checkpoint_path, + model_candidate="contextual-orchestrator/orchestrator/free", + attempt=1, + head_sha="h" * 40, + run_id="1", + run_attempt="1", + json_path=tmp_path / "run.jsonl", + export_path=export_path, + exit_code=1, + ) + assert entry["missing_required_outputs"] + + +def test_record_attempt_checkpoint_skips_malformed_message_rows(tmp_path: Path) -> None: + """Malformed export messages never become partial prompt replay.""" + export_path = tmp_path / "export.json" + export_path.write_text( + json.dumps( + { + "messages": [ + "not-a-message", + {"parts": [{"type": "text", "text": 123}]}, + {"info": {"role": "assistant"}, "parts": [{"type": "text"}]}, + ] + } + ), + encoding="utf-8", + ) + checkpoint_path = tmp_path / "checkpoint.json" + entry = record_attempt_checkpoint( + checkpoint_path=checkpoint_path, + model_candidate="contextual-orchestrator/orchestrator/free", + attempt=1, + head_sha="i" * 40, + run_id="1", + run_attempt="1", + json_path=tmp_path / "run.jsonl", + export_path=export_path, + exit_code=1, + ) + assert entry["missing_required_outputs"] + + +def test_record_attempt_checkpoint_without_export_file(tmp_path: Path) -> None: + """Missing export files still produce a bounded checkpoint record.""" + checkpoint_path = tmp_path / "checkpoint.json" + entry = record_attempt_checkpoint( + checkpoint_path=checkpoint_path, + model_candidate="contextual-orchestrator/orchestrator/free", + attempt=1, + head_sha="j" * 40, + run_id="1", + run_attempt="1", + json_path=tmp_path / "run.jsonl", + export_path=tmp_path / "missing-export.json", + exit_code=1, + ) + assert entry["partial_summary"]["assistant_text_present"] is False + + +def test_record_attempt_checkpoint_handles_unreadable_export(tmp_path: Path) -> None: + """Unreadable export payloads fail closed during partial-text extraction.""" + export_path = tmp_path / "export.json" + export_path.write_text("{", encoding="utf-8") + checkpoint_path = tmp_path / "checkpoint.json" + entry = record_attempt_checkpoint( + checkpoint_path=checkpoint_path, + model_candidate="contextual-orchestrator/orchestrator/free", + attempt=1, + head_sha="k" * 40, + run_id="1", + run_attempt="1", + json_path=tmp_path / "run.jsonl", + export_path=export_path, + exit_code=1, + ) + assert entry["missing_required_outputs"] + + +def test_record_attempt_checkpoint_handles_messages_not_list(tmp_path: Path) -> None: + """Export objects without message lists do not produce partial replay text.""" + export_path = tmp_path / "export.json" + export_path.write_text(json.dumps({"messages": "bad"}), encoding="utf-8") + checkpoint_path = tmp_path / "checkpoint.json" + entry = record_attempt_checkpoint( + checkpoint_path=checkpoint_path, + model_candidate="contextual-orchestrator/orchestrator/free", + attempt=1, + head_sha="l" * 40, + run_id="1", + run_attempt="1", + json_path=tmp_path / "run.jsonl", + export_path=export_path, + exit_code=1, + ) + assert entry["missing_required_outputs"] + + +def test_build_continuation_appendix_empty_and_invalid_last_entry(tmp_path: Path) -> None: + """Empty or malformed checkpoint documents produce no appendix.""" + checkpoint_path = tmp_path / "checkpoint.json" + assert build_continuation_appendix(checkpoint_path, budget=2) == "" + checkpoint_path.write_text(json.dumps({"attempts": ["bad-entry"]}), encoding="utf-8") + assert build_continuation_appendix(checkpoint_path, budget=2) == "" + checkpoint_path.write_text( + json.dumps( + { + "pinned_model": "contextual-orchestrator/orchestrator/free", + "attempts": [ + { + "attempt": 1, + "termination_reason": "invalid-control", + "partial_summary": "not-a-mapping", + "missing_required_outputs": "not-a-list", + "route_telemetry": {}, + } + ], + } + ), + encoding="utf-8", + ) + appendix = build_continuation_appendix(checkpoint_path, budget=2) + assert "Same-model continuation" in appendix + assert "Partial assistant digest" not in appendix + + +def test_continuation_budget_empty_history_uses_no_budget(tmp_path: Path) -> None: + """An empty history consumes no same-model continuation budget.""" + checkpoint_path = tmp_path / "checkpoint.json" + checkpoint_path.write_text(json.dumps({"attempts": []}), encoding="utf-8") + assert continuation_budget_remaining(checkpoint_path, budget=2) == 2 + + +def test_continuation_budget_has_no_unreachable_coverage_clamp() -> None: + """Budget decisions remain executable rather than hidden from coverage.""" + source = inspect.getsource(continuation_budget_remaining) + assert "used < 0" not in source + assert "pragma: no cover" not in source + + +def test_budget_cli_requires_explicit_authority(tmp_path: Path) -> None: + """The CLI cannot invent a continuation budget when authority is absent.""" + environment = dict(os.environ) + environment.pop("OPENCODE_SESSION_CONTINUATION_BUDGET", None) + completed = subprocess.run( + [ + sys.executable, + "scripts/ci/opencode_review_session_checkpoint.py", + "budget-remaining", + "--checkpoint", + str(tmp_path / "checkpoint.json"), + ], + cwd=Path(__file__).resolve().parents[1], + capture_output=True, + text=True, + check=False, + env=environment, + ) + + assert completed.returncode != 0 + assert "--budget" in completed.stderr + + +def test_budget_cli_rejects_negative_authority(tmp_path: Path) -> None: + """Every CLI path rejects negative continuation-budget authority.""" + for command_arguments in ( + ["budget-remaining"], + ["append-continuation", "--prompt", "prompt.md"], + ): + completed = subprocess.run( + [ + sys.executable, + "scripts/ci/opencode_review_session_checkpoint.py", + *command_arguments, + "--checkpoint", + str(tmp_path / "checkpoint.json"), + "--budget", + "-1", + ], + cwd=Path(__file__).resolve().parents[1], + capture_output=True, + text=True, + check=False, + ) + + assert completed.returncode != 0 + assert "non-negative integer" in completed.stderr + + +def test_main_entrypoint(tmp_path: Path) -> None: + """Module entrypoint delegates to main().""" + export_path = tmp_path / "export.json" + export_path.write_text(_export_with_text("partial"), encoding="utf-8") + checkpoint_path = tmp_path / "checkpoint.json" + completed = subprocess.run( + [ + sys.executable, + "scripts/ci/opencode_review_session_checkpoint.py", + "budget-remaining", + "--checkpoint", + str(checkpoint_path), + "--budget", + "1", + ], + cwd=Path(__file__).resolve().parents[1], + capture_output=True, + text=True, + check=False, + ) + assert completed.returncode == 0 + assert completed.stdout.strip() == "1" + + +def test_main_subcommands(tmp_path: Path) -> None: + """CLI subcommands return bounded stdout for host orchestration.""" + from scripts.ci.opencode_review_session_checkpoint import main + + export_path = tmp_path / "export.json" + export_path.write_text(_export_with_text("partial"), encoding="utf-8") + checkpoint_path = tmp_path / "checkpoint.json" + prompt_path = tmp_path / "prompt.md" + prompt_path.write_text("base\n", encoding="utf-8") + assert ( + main( + [ + "record", + "--checkpoint", + str(checkpoint_path), + "--model-candidate", + "contextual-orchestrator/orchestrator/free", + "--attempt", + "1", + "--head-sha", + "e" * 40, + "--run-id", + "1", + "--run-attempt", + "1", + "--json-path", + str(tmp_path / "run.jsonl"), + "--export-path", + str(export_path), + "--exit-code", + "1", + ] + ) + == 0 + ) + assert ( + main( + [ + "append-continuation", + "--prompt", + str(prompt_path), + "--checkpoint", + str(checkpoint_path), + "--budget", + "2", + ] + ) + == 0 + ) + assert main(["budget-remaining", "--checkpoint", str(checkpoint_path), "--budget", "2"]) == 0 + + +def test_cli_record_emits_bounded_json(tmp_path: Path) -> None: + """The checkpoint CLI prints only bounded metadata to stdout.""" + export_path = tmp_path / "export.json" + export_path.write_text(_export_with_text("partial"), encoding="utf-8") + checkpoint_path = tmp_path / "checkpoint.json" + completed = subprocess.run( + [ + sys.executable, + "scripts/ci/opencode_review_session_checkpoint.py", + "record", + "--checkpoint", + str(checkpoint_path), + "--model-candidate", + "contextual-orchestrator/orchestrator/free", + "--attempt", + "1", + "--head-sha", + "c" * 40, + "--run-id", + "35452307646", + "--run-attempt", + "1", + "--json-path", + str(tmp_path / "run.jsonl"), + "--export-path", + str(export_path), + "--exit-code", + "1", + ], + cwd=Path(__file__).resolve().parents[1], + capture_output=True, + text=True, + check=False, + ) + assert completed.returncode == 0 + payload = json.loads(completed.stdout.strip()) + assert "termination_reason" in payload + assert "partial" not in completed.stdout + + +def test_read_bounded_text_oserror_returns_empty(tmp_path: Path, monkeypatch) -> None: + """Unreadable evidence files fail closed without raising.""" + from scripts.ci.opencode_review_session_checkpoint import _read_bounded_text + + path = tmp_path / "blocked.txt" + path.write_text("x", encoding="utf-8") + + def _raise_oserror(*_args, **_kwargs): + raise OSError("blocked") + + monkeypatch.setattr(path.__class__, "open", _raise_oserror) + assert _read_bounded_text(path, 16) == "" + + +def test_record_attempt_checkpoint_keeps_assistant_text_parts_only( + tmp_path: Path, +) -> None: + """Partial text extraction ignores non-assistant and non-text parts.""" + export_path = tmp_path / "export.json" + export_path.write_text( + json.dumps( + { + "messages": [ + { + "info": {"role": "user"}, + "parts": [{"type": "text", "text": "ignore user"}], + }, + { + "info": {"role": "assistant"}, + "parts": [ + {"type": "tool", "text": "ignore tool"}, + {"type": "text", "text": "assistant partial"}, + ], + }, + ] + } + ), + encoding="utf-8", + ) + checkpoint_path = tmp_path / "checkpoint.json" + entry = record_attempt_checkpoint( + checkpoint_path=checkpoint_path, + model_candidate="contextual-orchestrator/orchestrator/free", + attempt=1, + head_sha="m" * 40, + run_id="1", + run_attempt="1", + json_path=tmp_path / "run.jsonl", + export_path=export_path, + exit_code=1, + ) + assert "opencode-review-control-v1" in entry["missing_required_outputs"] + assert "assistant partial" not in json.dumps(entry) + + +def test_main_script_entrypoint(tmp_path: Path) -> None: + """Running the script path exercises the __main__ entrypoint.""" + checkpoint_path = tmp_path / "checkpoint.json" + completed = subprocess.run( + [ + sys.executable, + "scripts/ci/opencode_review_session_checkpoint.py", + "budget-remaining", + "--checkpoint", + str(checkpoint_path), + "--budget", + "2", + ], + cwd=Path(__file__).resolve().parents[1], + capture_output=True, + text=True, + check=False, + ) + assert completed.returncode == 0 + assert completed.stdout.strip() == "2" + + +def test_record_attempt_checkpoint_skips_non_list_attempt_container( + tmp_path: Path, monkeypatch +) -> None: + """A corrupt in-memory attempt container cannot append a new record.""" + from scripts.ci import opencode_review_session_checkpoint as checkpoint + + export_path = tmp_path / "export.json" + export_path.write_text(_export_with_text("partial"), encoding="utf-8") + checkpoint_path = tmp_path / "checkpoint.json" + + def _broken_load(_path: Path) -> dict: + return {"schema": 1, "attempts": "bad-history"} + + monkeypatch.setattr(checkpoint, "_load_checkpoint", _broken_load) + entry = checkpoint.record_attempt_checkpoint( + checkpoint_path=checkpoint_path, + model_candidate="contextual-orchestrator/orchestrator/free", + attempt=1, + head_sha="n" * 40, + run_id="1", + run_attempt="1", + json_path=tmp_path / "run.jsonl", + export_path=export_path, + exit_code=1, + ) + document = json.loads(checkpoint_path.read_text(encoding="utf-8")) + assert entry["attempt"] == 1 + assert document["attempts"] == "bad-history" + + +def test_record_attempt_checkpoint_skips_assistant_without_parts_list( + tmp_path: Path, +) -> None: + """Assistant messages without part lists do not contribute partial replay text.""" + export_path = tmp_path / "export.json" + export_path.write_text( + json.dumps({"messages": [{"info": {"role": "assistant"}, "parts": "bad"}]}), + encoding="utf-8", + ) + checkpoint_path = tmp_path / "checkpoint.json" + entry = record_attempt_checkpoint( + checkpoint_path=checkpoint_path, + model_candidate="contextual-orchestrator/orchestrator/free", + attempt=1, + head_sha="o" * 40, + run_id="1", + run_attempt="1", + json_path=tmp_path / "run.jsonl", + export_path=export_path, + exit_code=1, + ) + assert entry["missing_required_outputs"] + + +def test_module_sys_path_insert_is_idempotent() -> None: + """Re-importing the checkpoint module does not duplicate sys.path entries.""" + import importlib + + import scripts.ci.opencode_review_session_checkpoint as checkpoint + + root = str(checkpoint._REPO_ROOT) + before = sys.path.count(root) + importlib.reload(checkpoint) + assert sys.path.count(root) == before + + +def test_main_unknown_command_exits(monkeypatch) -> None: + """Unknown CLI commands fail closed after argument parsing.""" + from scripts.ci import opencode_review_session_checkpoint as checkpoint + + class Args: + command = "not-a-real-command" + + monkeypatch.setattr(checkpoint, "parse_args", lambda _argv=None: Args()) + try: + checkpoint.main([]) + except SystemExit as exc: + assert "unknown command" in str(exc) + else: + raise AssertionError("expected SystemExit") diff --git a/tests/test_product_technical_gap_baseline.py b/tests/test_product_technical_gap_baseline.py index d44ffdb8e6..f8956494c9 100644 --- a/tests/test_product_technical_gap_baseline.py +++ b/tests/test_product_technical_gap_baseline.py @@ -98,3 +98,37 @@ def test_master_context_points_at_live_baseline_without_freezing_shas() -> None: assert "ContextualWisdomLab/naruon#975" in source assert "Done" in source assert "merge authorization" in source + + +def test_baseline_keeps_checkpoint_and_coverage_lock_incident_rows() -> None: + """Restack keeps this PR's checkpoint rows and main's coverage-lock row.""" + + source = BASELINE.read_text(encoding="utf-8") + assert "<<<<<<<" not in source + assert ">>>>>>>" not in source + for marker in ( + "### 2026-09-20 current-head incident delta", + "CONTROL-OPENCODE-CHECKPOINT-INTEGRITY-01", + "CONTROL-OPENCODE-CONTINUATION-AUTHORITY-02", + "CONTROL-OPENCODE-PROVIDER-NEUTRAL-03", + "CONTROL-OPENCODE-CHECKPOINT-CAUSE-04", + "CONTROL-OPENCODE-CHECKPOINT-BOUNDS-05", + "### 2026-09-19 exact-head incident delta", + "CONTROL-OPENCODE-COVERAGE-LOCK-CONTEXT-01", + ): + assert marker in source, marker + + +def test_baseline_preserves_protected_main_authority_sections() -> None: + """Partial-file replacements must not erase protected Gap evidence.""" + + source = BASELINE.read_text(encoding="utf-8") + for marker in ( + "## 2026-08-30 sidecar-preflight outage: consolidated evidence and why it is not one deterministic bug", + "## 2026-08-30 ZDR/NIM-routing architecture review (owner-directed)", + "## 2026-09-01 OpenCode contextual-orchestrator runtime ceiling", + "## 6. Compliance and data boundary", + "## 7. APA 7th references", + "## Noema reviewer credential-lifetime delta — 2026-09-01", + ): + assert marker in source, marker