Skip to content
Merged
11 changes: 11 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,17 @@ breaking changes may land in a minor release.

## [Unreleased]

### Fixed

- Ignore hook events from nested coding-CLI sessions that inherit the relay environment,
so a child's `Stop`/`SessionEnd` no longer completes or crashes the launched session:
an id that announces its own `SessionStart` after the launched session's first is
foreign, and its events are dropped, crumbed once per id as
`foreign-hook-event-ignored`, counted as `foreign_hook_events` in `heartbeat.json` and
`timeout-fired`, and excluded from the sweep diagnostic (`hook_foreign_ids`). Forward
SessionStart `source` so a `clear`/`compact` start rebinds; re-run `bmad-loop init` to
re-vendor the relay (#767). Contributed by [@Pinstack](https://github.com/Pinstack).

## [0.13.0] — 2026-09-28

### Added
Expand Down
1 change: 1 addition & 0 deletions docs/FEATURES.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,6 +80,7 @@ See [README.md](../README.md) for the narrative overview and [setup-guide.md](se
- A session parked on a human prompt is escalated instead of nudged (DW-348/DW-350). The stall wake nudge ends in `Enter`, which can _answer_ the prompt a parked CLI shows (#727: it confirmed "No, exit" on Claude Code's Bypass Permissions dialog). Two opt-in, profile-declared signals now gate it. **Hook-reported**: the relay forwards a `Notification` payload's `notification_type`, and the profile's `[hooks.notification_types]` maps the subtypes that mean "waiting on a human" onto the canonical parked kinds `PermissionPrompt` / `IdlePrompt` / `QuotaPrompt` (a profile may also map a native event straight to one). The claude profile maps `permission_prompt`, `idle_prompt`, `quota_auto_resume_stale`, `quota_auto_resume_disabled` and the MCP elicitation dialogs `elicitation_dialog` / `elicitation_url_dialog` (as `PermissionPrompt`, DW-434); `agent_needs_input` stays unmapped. One of its triggers is a different (background) session waiting while agent view is open, and the payload does not say which trigger fired. The other, a teammate setup question, needs experimental agent teams, which bmad-loop does not enable. An overlay may map it. The adapter latches the signal, and a `Stop` clears it. Because claude sends `idle_prompt` about 60 s after every finished turn, a dev, review or workflow session that ends a turn without a result and then sits idle re-latches and, at stall-grace expiry, pauses as parked instead of receiving the stall wake nudge — intended; a project overlay that drops `idle_prompt` from `[hooks.notification_types]` restores the nudge. **Pane-matched**: at every stall-grace expiry with no latch set — a nudge due, or the final stall once the nudges are spent or when none are configured (DW-433) — the visible pane (`capture-pane -p`) is matched line by line against the profile's `parked_prompt_patterns` (claude ships the two captured #727 dialog lines). Neither signal completes a session or acts on arrival. The one decision point is the stall-grace expiry: with a parked signal, no nudge is typed and the session ends `stalled` with `parked` / `parked_evidence` (a dead window still ends `crashed`). Only an adapter with a stall grace armed reaches that point — the dev/review adapter (which also drives fix and workflow sessions) with `dev_stall_grace_s > 0`; the plain triage adapter (sweep triage and migration) has no stall grace, so its sessions still time out or retry as before. Every dev, review, fix, blocking-workflow, migration and triage site PAUSEs a parked result (`parked: <role> session stalled (<evidence>; …)`), after the env-fault arm and ahead of the no-work arm. Re-arm restores the attempt. `parked` rides `dev-decision`, `fix-decision`, `workflow-end`, `migrate-decision`, `triage-decision` and (when true) `session-end`. Capture or regex failures degrade to "no match": a due nudge goes out, or the final stall ends unparked, as before. A project initialized before this has no `Notification` relay; `hooks.registered` still passes, and hook-reported signals start after the next `bmad-loop init`. The opencode HTTP adapter is unchanged.
- An idle session is visible while it sits (#680). A session idling inside a tool call (`sleep 590; cat …`) keeps its pane log growing through spinner repaints, so the stall re-arm — correctly — never fires and nothing separated it from a working session but `session_timeout_min`. The adapter now stats the live transcript's `(mtime_ns, size)` — a baseline the moment the first hook event names it, then on the heartbeat cadence (a stat, never parsed usage, so it works for `usage_parser = "none"`) — stamps the age on `heartbeat.json` as `transcript_idle_s` (`null` before a transcript is known), and — when the run's journal is attached, which the engine does for every adapter it owns — journals one `session-idle` (`task_id`, `idle_s`, `since_ts`, `threshold_s`) when the age crosses `limits.dev_stall_grace_s` and one `session-active` (`task_id`, `idle_s`) when the transcript moves again; a later stretch emits a fresh pair. The threshold is the stall grace on purpose — the event fires exactly when the session _would_ have stalled had its pane not kept repainting, so the two records are directly comparable — and `0` disables the events with no new knob. The TUI's agent line shows the open stretch as `· idle <age>`. Observability only: nothing bounds the stretch, and the session is neither nudged, stalled nor killed for it. Scope: the notice starts when a hook event names the transcript — `SessionStart` on the `claude`, `codex`, `gemini` and `copilot` profiles, so it covers a session's first turn there; `antigravity` fires no `SessionStart` and names the transcript only on its `Stop`, so its first turn is invisible to the notice (the heartbeat's `transcript_idle_s` stays `null` until then), and the `opencode-http` transport has no pane wait loop and emits neither the field nor the events.
- A session the multiplexer lost says so (#489). Sessions complete on a hook `Stop` or on window death, and a window is gone whether the CLI exited or something destroyed the whole mux session out from under the run — an external reaper, a concurrent prune or `bmad-loop stop`, an operator `kill-session`, a server crash, the host sleeping. Both are `crashed`, so the retry/defer reason an operator reads said only `dev session crashed` — pointing at the agent when the host was at fault. The crash verdict now asks whether the _session_ still exists and, when it does not, says so in the reason (`… session crashed: the multiplexer no longer reports the session, so the window's disappearance is not evidence the CLI exited`), as `session_vanished` on `dev-decision` and `fix-decision` either way, beside the routing each fed, on every role's `session-end` journal entry when it is true (the convention `env_fault` already uses there), and as a `session-vanished` breadcrumb in `session-lifecycle.jsonl`. When the probe itself cannot ask (`has_session` raises `MultiplexerError`), the verdict stays an undiagnosed `crashed` and a `session-probe-failed` breadcrumb (`session`, `error`) records that the question went unanswered (DW-382). A negative lookup is confirmed before it is recorded (DW-459): `has_session` maps every nonzero backend result to False, including psmux's `Invalid session key` and `connection timed out` from a session that is still alive, so the adapter then lists the session's windows and writes `session-vanished` only when that listing proves the session gone; a listing that raises or still finds windows writes `session-probe-failed` instead. The repair path carries it the same way: when fix attempts are exhausted the defer names the lost session instead of blaming the tree for repairs that never ran. The wording states what the evidence _withdraws_, not what it proves: a session the multiplexer no longer reports says nothing about what removed it — enough to stop an operator reading window death as a CLI exit, not enough to name a destroyer. It composes with an environment-fault pause instead of being swallowed by it. A session reaped _after_ flushing its result still scores `completed` and is not diagnosed — it produced something. Diagnosis only — the routing is unchanged, and a retry re-creates the session.
- Hook events from a nested coding CLI are attributed, not trusted (#767). A CLI process started from inside a session inherits the relay environment, so its `SessionStart`/`Stop`/`SessionEnd` land in the parent's event stream under the parent's task id and could complete or crash it. The generic adapter applies a second layer after the task-id filter (`SessionAttribution`): the first `SessionStart` is the launched session's — identified or not, so an anonymous start (a payload the relay could not read) still takes the parent's slot — and a later `SessionStart` with a new id marks that id foreign, unless its `source` is `clear` or `compact`, which is the launched session rotating its id and rebinds (`resume` and `startup` stay foreign). `source` alone does not rebind: an id already found foreign stays foreign (a child compacting), and a `clear` start right after a foreign id's `SessionEnd` is that child clearing, so its new id is foreign too. A foreign id's events are dropped before they can complete the session, re-point its transcript, spend nudges or re-arm the stall timer; the first drop per foreign id writes a `foreign-hook-event-ignored` breadcrumb (`hook_event`, `foreign_session_id`, `dropped_so_far`), every drop counts into `foreign_hook_events` on `heartbeat.json` and the `timeout-fired` crumb, and the sweep's failed-session diagnostic replays the same rule (`hook_foreign_ids`, `; foreign ignored: N` in the escalation suffix). The rule fails toward acceptance: id-less events, ids that never announced a `SessionStart` (a rotated id, a copilot `toolu_…` subagent `Stop`) and everything before the first start — including an identified `SessionEnd` from a CLI that exited before its `SessionStart` fired (#727) — are admitted, so attribution only ever drops a known child's events and never adds a completion path. Accepted limitation: a child `SessionEnd` whose child never announced a `SessionStart` is indistinguishable from the parent's own and is admitted; nested CLIs announce their start. A profile that maps no `SessionStart` (Stop-only, e.g. `antigravity`) gets no nested-CLI protection. A relay vendored before this change forwards no `source`, so there a `clear`/`compact` start with a new id reads as foreign and the session falls back to window death or its timeout — re-run `bmad-loop init` to re-vendor the relay.
- Transport faults in the generic (multiplexer-driven) adapter leave crumbs instead of healthy-looking answers (DW-447/449/453/454). Like `session-probe-failed`, a window-liveness probe that raises `MultiplexerError` never reads as death, and every verdict is unchanged; what changes is the record. `liveness-probe-failed` (`site`, `error`) is written by the `tick` site at a wait-loop streak's first failed tick only, never once per tick, and by the one-shot sites (`over-budget`, `stall`, `post-kill`) once per verdict probe that raised. `liveness-probe-recovered` (`failures`) closes a tick streak when a later tick probes cleanly; a session that ends mid-streak (Stop, SessionEnd, abort, over-budget) leaves no recovered crumb. `heartbeat.json` carries the running streak as `probe_failures`, and `timeout-fired` carries the final one. `over-budget-fired` and `kill-escalated` gain `liveness_unknown`, which is true when the probe behind that verdict (for the kill, the last poll before escalation) raised. A stall, budget, stop or contract nudge whose send raised writes `nudge-send-failed` (`nudge`, `error`) and is never reported as sent: `stall_nudges_sent` on the heartbeat counts delivered nudges, `stall_nudges_failed` counts the failed attempts, and `contract-nudge-sent` is written only after a successful send (a failed contract nudge is still not retried). The stall-nudge cap and the #727 activity window count attempts, delivered or not. A post-kill rescue abandoned on unknown liveness or an unreadable artifact writes `post-kill-rescue-abandoned` (`reason` = `liveness-unknown` / `unreadable-artifact`, `status`, plus `error` for the read fault). The opencode HTTP adapter's heartbeat carries neither new key (its `stall_nudges_sent` counts attempts), but its nudges are crumbed the same way (DW-503): its `send_text` raises a `MultiplexerError` when the prompt POST fails, so an undelivered contract nudge writes `nudge-send-failed` instead of `contract-nudge-sent`, and its own budget, stall and stop nudges write `nudge-send-failed` and carry on exactly as before. Its dev sessions do share the post-kill reconcile, so they write `post-kill-rescue-abandoned` too, on the `unreadable-artifact` arm only (their liveness probe never answers "unknown").
- Three generic-adapter observation folds are crumbed the same way, verdicts unchanged (DW-448/450/451). A stall-expiry look at the pane whose capture raises `MultiplexerError`, or whose parked-prompt search blows `PARKED_PROMPT_MATCH_TIMEOUT_S`, still reads as "not parked", and writes `parked-probe-failed` (`reason` = `capture-failed` / `match-timeout`, `pattern` for a timeout, `error`) once per expiry; a profile with no `parked_prompt_patterns` or a backend without `capture_pane` stays silent. A pane-log stat fault other than absence still leaves the #261/#727 proof-of-work signal unknown, and writes `log-evidence-failed` (`error`) once per session however many verdict sites consult it. A present `result.json` the read-back refuses (unreadable, unparseable, not an object) still reads as no result: the Stop read-back's give-up record in `resultless-stops.jsonl` says `malformed-result-json` with the refusal instead of `no-result-json`, once per give-up rather than per poll, and the exit read-back writes `result-json-refused` (`error`). The opencode HTTP adapter shares the `resultless-stops.jsonl` verdict but not the other crumbs.
- Four more generic-adapter folds are crumbed, verdicts unchanged (DW-452/455/456/457). A budget usage sample whose transcript read raises still reads as "no sample" — with a persistent fault, `token_budget_mode = "enforce"` stays off for the session — and its streak writes `usage-sample-failed` (`error`) at the first failure and `usage-sample-recovered` (`failures`) at the next clean sample, never once per tick (one torn mid-append read is one pair); `heartbeat.json` carries the running count as `usage_sample_failures`. A #727 transcript activity scan that raises still counts as no model-side evidence, with its streak crumbed the same way (`transcript-scan-failed` / `transcript-scan-recovered`). A launch-snapshot fault that turns the #276 M1 refuse gate inert — the identity check's `resolve()` raising (the 3.11 symlink-loop `RuntimeError` included, which used to escape) or the digest read raising — still reads NEUTRAL, and writes `spec-identity-unreadable` (`spec`, `snapshot`, `error`) or `spec-digest-unreadable` (`spec`, `error`) once per read-back, not per grace poll. A spec read-back that gives up on a fault says so: `stat-failed` for a stat fault the stories read-back used to file as `stale-mtime`, `unreadable-spec` for an unreadable or undecodable spec it filed as `not-terminal` (and the scan fallback as `no-artifact`), each with the error, in `resultless-stops.jsonl` on a Stop read-back and as `spec-readback-failed` (`reason`, `spec`, `error`) on a one-shot or dead-window (post-kill reconcile) read, which recorded nothing before.
Expand Down
Loading
Loading