From ad7028ff40252137bf0814a5d60f779c8e4d10af Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Fredrik=20Wollse=CC=81n?= Date: Tue, 8 Sep 2026 12:37:22 +0300 Subject: [PATCH 1/4] feat: read what a tab has learned, pin a working set, and stop claiming sends landed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Six gaps measured while driving a 13-tab fleet, each one a place where cctabs reported something it had not established. - `cctabs transcript [n]` (alias `findings`) prints a tab's last n assistant messages from its transcript. `scrollback` returns the last painted frame, so a tab mid-turn shows a spinner — which meant sessions briefed each other from stale pictures. Searches every Claude config dir, because a tab on a backend preset writes under that preset's root and looking in one root reports "no transcript" for a healthy session. - `cctabs sort --first a,b,c` pins a chosen set to the front. The plugin's /api/tabs/reorder already did exactly this; no verb exposed it. Refuses as a whole if a name doesn't resolve — half a working set in reach, with no indication which half, is worse than an error. - `sessions --json` now carries `session_lookup`, so a null `session_id` says whether we looked and found nothing, couldn't look, or never looked. A lookup that threw was previously swallowed into the same null. - `restore` verifies what it spawned: re-reads the tab list, checks each tab has a process, resolves its session, and reports N verified / N unconfirmed / N failed, exiting non-zero on failure. A tab that came back as a *different* session counts as failed — `--resume` on an id it can't find opens a fresh conversation, so the tab looks perfect and the context is gone. The line this replaces read "78 spawned, 0 failed" with one tab absent and one stripped of its context. - `send` separates three claims that were one ✔ line: nothing arrived (hard failure, and the body is NOT submitted — a fragment reads as a whole message), arrived but completeness unverified, and verified. `--verify` does the real comparison against the target's transcript. `--path` hands over a file path instead of pasting, which has no truncation surface. `--submit` presses Enter only. `--wait-for-prompt` reads the whole tail of the buffer, so a "Restart to update" banner below a ready prompt no longer defeats it. Worth recording what the measurements actually settled about that last one, since the obvious fix is wrong: a paste chip's `+N lines` does NOT track the payload. A 6,892-byte, 76-line payload was measured delivering *completely* into an idle tab while its chip read `+10 lines`, so failing a send on a short chip fails healthy sends. The screen cannot establish completeness for a collapsed paste at all — hence "unverified" as a real answer, and hence `--verify`, which reads what the session recorded receiving. Co-Authored-By: Claude Opus 5 (1M context) --- CHANGELOG.md | 14 ++ skills/cctabs/SKILL.md | 124 ++++++++++++++- src/commands/index.ts | 5 + src/commands/restore.test.ts | 86 +++++++++- src/commands/restore.ts | 243 +++++++++++++++++++++++++++- src/commands/scrollback.ts | 42 ++--- src/commands/send.test.ts | 20 +++ src/commands/send.ts | 270 ++++++++++++++++++++++++++------ src/commands/sessions.ts | 33 +++- src/commands/sort.test.ts | 63 +++++++- src/commands/sort.ts | 116 +++++++++++++- src/commands/transcript.ts | 149 ++++++++++++++++++ src/core/open-session.test.ts | 38 ++++- src/core/open-session.ts | 71 ++++++--- src/core/paste-confirm.test.ts | 128 +++++++++++++++ src/core/paste-confirm.ts | 169 ++++++++++++++++++++ src/core/session-lookup.test.ts | 59 +++++++ src/core/session-lookup.ts | 67 ++++++++ src/core/session-status.test.ts | 37 +++++ src/core/session-status.ts | 27 ++++ src/core/session.ts | 17 +- src/core/tab-target.ts | 87 ++++++++++ src/core/transcript.test.ts | 185 ++++++++++++++++++++++ src/core/transcript.ts | 196 +++++++++++++++++++++++ 24 files changed, 2125 insertions(+), 121 deletions(-) create mode 100644 src/commands/send.test.ts create mode 100644 src/commands/transcript.ts create mode 100644 src/core/paste-confirm.test.ts create mode 100644 src/core/paste-confirm.ts create mode 100644 src/core/session-lookup.test.ts create mode 100644 src/core/session-lookup.ts create mode 100644 src/core/tab-target.ts create mode 100644 src/core/transcript.test.ts create mode 100644 src/core/transcript.ts diff --git a/CHANGELOG.md b/CHANGELOG.md index b942119..4e74ec0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -3,6 +3,20 @@ All notable changes to **cctabs** are listed here. The user-facing version of this page lives at [cctabs.com/changelog](https://cctabs.com/changelog). +## Unreleased + +- **`cctabs transcript [n]` reads what a tab has SAID, not what it is painting.** `scrollback` returns the last painted frame, so a tab mid-turn shows a spinner and nothing else — which meant sessions briefed each other from stale pictures, and one such brief told a tab to go measure two things it had already measured, missing a third it had found. Prints the last n assistant messages from the transcript (`--json` for a driver; `findings` as an alias). It searches **every** Claude config dir, because a tab on a backend preset writes under that preset's own root and looking in only `~/.claude/projects` reports "no transcript" for a healthy session — which reads as dead. Tool-only and thinking-only messages are skipped. It exits non-zero and names the failure — no session titled after this tab (with a count of transcripts that do exist for its directory, which distinguishes a renamed tab from a dead one), no transcript on disk, or a lookup that threw — because "I couldn't read it" and "it hasn't said anything" are different answers. +- **`cctabs sort --first a,b,c` pins a chosen set to the front of the bar.** Activity order is close to the opposite of what a driver wants: a tab that just delivered sinks. The Tabby plugin's `POST /api/tabs/reorder` already did exactly this — unlisted tabs keep their relative order and sort after the listed ones — with no CLI verb exposing it. Naming tabs skips the activity scan entirely (~7.7s of transcript reading on a 65-tab fleet). If any name doesn't resolve to exactly one tab, nothing moves and it exits non-zero: half a working set in reach, with no indication which half, is worse than an error. +- **`cctabs send` no longer reports success it hasn't established.** A large paste collapses to a `[Pasted text #N +M lines]` chip, and the chip renders no matter how little arrived — so "a chip is on screen" was accepted as proof of delivery, which is how a 6,835-byte brief that landed as its last 756 bytes came to be reported as sent. `send` now distinguishes three claims that used to be one ✔ line: *nothing arrived* (a hard failure — and the body is **not submitted**, because pressing Enter on a fragment sends something that reads as a complete message; the text is left in the input box and the command exits non-zero), *something arrived but completeness is unverified* (a warning naming the two ways to establish it), and *verified*. +- **`cctabs send --verify` checks what the target session actually RECEIVED.** After submitting, it reads the target's own transcript — which records the user message as received — and compares the payload's front and tail fingerprints against it, failing loudly on a mismatch and naming which end went missing. This is the only reliable completeness check available, and finding that out cost a measurement: the chip's `+N lines` count does **not** track the payload. A 6,892-byte, 76-line payload was measured delivering completely into an idle tab while its chip read `+10 lines`, so treating a shortfall as truncation fails healthy sends. The screen cannot answer this question; the transcript can. +- **`cctabs send --path ` hands a tab a file path instead of pasting its contents.** No truncation surface at all: what crosses the prompt line is a path, and the receiving session reads the payload from disk. This is the fallback that had to be used to deliver the brief describing the truncation, and it is now first-class rather than a discipline. +- **`cctabs send` refuses payloads over 1 KB into a tab with a turn in flight**, with `--force` to override. Short replies into a busy tab — `yes` to a tool call, `2` to a picker — are exactly what sending into an active tab is for and still work; a multi-kilobyte brief pushed into a busy input handler is how text gets clipped. +- **`cctabs send --submit` presses Enter only**, submitting a prompt already parked in a tab's input box. An empty body with no `--submit` now says so instead of silently pressing Enter in someone's session. +- **`cctabs send --wait-for-prompt` no longer times out against tabs that are ready.** It tested only the buffer's last non-empty line, and Claude renders notices *below* its input line — a `Restart to update` banner was enough to defeat it for the full 20s. It now reads the whole tail of the window. +- **`cctabs restore` reports the count it actually achieved.** It read "78 spawned, 0 failed" while one tab was absent entirely and another had come back with no session, having lost its context — a summary computed from "did the spawn call return?", which cannot report a failure that happens after it returns. Restore now re-reads the tab list after spawning, checks each tab has a process, resolves its session from disk, and reports `N verified, N unconfirmed, N failed`, exiting non-zero when anything failed. A tab that came back as a *different* session than requested counts as failed: `claude --resume` on an id it can't find opens a fresh conversation, so the tab looks perfect and the context is gone. `unconfirmed` is a first-class outcome rather than rounded either way. +- **`cctabs sessions --json` says why a `session_id` is missing.** Every row carries `session_lookup`: `found`, `not-found` (with `sessions_in_dir`, so a tab renamed out from under a live session is distinguishable from a directory nothing has ever run in), `no-cwd`, or `lookup-failed` (with the error — a lookup that threw is *unknown*, not absent, and was previously swallowed into the same null). Observed on a fleet as a tab whose session was perfectly readable but filed under a title that no longer matched. +- Internal: `send`, `scrollback` and `transcript` now share one target resolver (`core/tab-target.ts`) instead of three verbatim copies of the tab-then-block lookup, and `scrollback [n]` accepts the bare line count the docs have always shown. + ## 0.5.3 — 2026-09-07 - **`cctabs whoami` answers "which tab am I running in?"** Prints the tab name — so `$(cctabs whoami)` drops into a PR body or commit trailer — with `--json` adding `worktree`, `session_id`, `cwd`, `backend`, `config_dir`, `color` and `via`. It exists because self-attribution had no answer: when every session's PRs carry the same GitHub author, "whose work is this, and is it in flight?" is unanswerable, and it gets asked exactly when someone else's files are being written right now. diff --git a/skills/cctabs/SKILL.md b/skills/cctabs/SKILL.md index 85bad30..f2af9c8 100644 --- a/skills/cctabs/SKILL.md +++ b/skills/cctabs/SKILL.md @@ -185,8 +185,13 @@ cctabs rename # rename the tab title + on-disk custom cctabs color # set/clear the tab colour: blue|green|orange|purple|red|yellow|none|#rrggbb cctabs whoami [--json] # which tab is THIS session in? prints the tab name, or "unknown" cctabs sort [--dry] [--reverse] # reorder the tab bar by session activity, newest first (Tabby only) -cctabs scrollback [n] # read terminal output (default: 50 lines) +cctabs sort --first a,b,c [--dry] # PIN these tabs to the front, in this order; everything else keeps its relative order +cctabs scrollback [n] # read terminal output — the last PAINTED FRAME (default: 50 lines) +cctabs transcript [n] [--json] # read what a tab has SAID: its last n assistant messages, from the transcript (alias: `findings`) cctabs send [text] # send input — arg, --file, or stdin pipe +cctabs send --path # hand the tab a file PATH to read — the safe way to deliver anything large +cctabs send --file --verify # check the target's transcript for what it actually RECEIVED +cctabs send --submit # press Enter only, submitting a prompt already parked in the box cctabs export [--out path] # bundle a tab + its claude session into a tarball cctabs export --all [-w workspace] # bundle every tab in a workspace cctabs import [--dry-run] [-f] # restore tabs + sessions from a tarball @@ -211,6 +216,50 @@ identified the tab (`via`). every tab by scanning transcripts (~7.7s on a 65-tab fleet, minutes cold), where `whoami` is ~1s and reads no transcripts. +## What has that tab worked out? — `cctabs transcript` + +⛔ **Before you draft a message to another tab, read what it has already +said.** Briefing a session from a stale picture is how you tell a tab to go +measure two things it measured an hour ago — and miss the third thing it found +that you didn't know about. + +```bash +cctabs transcript payments # last 3 assistant messages +cctabs transcript payments 10 # last 10 +cctabs transcript payments --json # for a driver: session_id, cwd, backend, turns[] +cctabs findings payments # same command, reads better when you're asking "what did it find?" +``` + +- ⚠️ **This is not `scrollback`.** `scrollback` returns the last *painted frame*, + so a tab mid-turn shows you a spinner and nothing else — its actual findings + are in the transcript, not on the screen. `transcript` reads the transcript. +- It searches **every Claude config dir**, not just `~/.claude/projects`. A tab + running under a backend preset writes beneath that preset's own + `CLAUDE_CONFIG_DIR`, and looking in one root reports "no transcript" for a + perfectly healthy session — which reads as "that tab is dead". +- It **exits non-zero and says which failure it is** rather than printing + nothing: no session titled after this tab (with a count of how many + transcripts *do* exist for its directory — usually means the tab was renamed + after Claude started), no transcript on disk, or a lookup that failed + outright. "I couldn't read it" and "it hasn't said anything" are different + answers. +- Tool-only and thinking-only messages are skipped; you get the prose. + +## Getting a working set in reach — `cctabs sort --first` + +Activity order is close to the *opposite* of what a driver wants: a tab that +just delivered sinks to the bottom. To pin a chosen set instead: + +```bash +cctabs sort --first auth,payments,billing # these three to the front, in this order +cctabs sort --first auth,payments --dry # show the plan first +``` + +Every unlisted tab keeps its current relative order and sorts after the pinned +ones. If any name doesn't resolve to exactly one tab, **nothing is moved** and +the command exits non-zero — half a working set in reach, with no indication +which half, is worse than an error. Tabby only. + ## Tab colours `-c/--color` on `new`/`resume`/`fork`, or `cctabs color ` for a tab @@ -431,6 +480,20 @@ cctabs restore --dry # preview what would be resumed without doing cctabs restore ~/Dev/myapp # restrict the search to one project dir ``` +⚠️ **Read the count at the end, and trust it — it can now fail.** After +spawning, restore re-reads the tab list, checks each new tab has a process, and +resolves its session from disk, then reports `N verified, N unconfirmed, N +failed` and **exits non-zero if anything failed**. A tab that came back as a +*different* session than the one requested counts as failed, not restored: +`claude --resume` on an id it can't find quietly opens a fresh conversation, so +the tab looks perfect and the context is gone. `unconfirmed` is its own answer — +the tab is up but its session isn't readable yet — and those are named so you +can check them with `cctabs transcript` before briefing anything from them. + +The line this replaced read "78 spawned, 0 failed" while one tab was absent +entirely and another had lost its context, because it counted spawn calls that +returned rather than tabs that worked. + If a session was started in a different `cwd` than the tab's current directory (common after `cd`-ing inside the tab), the global search still finds it via the recorded session metadata — no need to guess the right dir. The search covers **every Claude account**, not just the default one: sessions launched under a backend preset live in that preset's own `CLAUDE_CONFIG_DIR`, and restore looks there too, then relaunches each tab under the account its session came from. Nothing to pass — a mixed-account fleet restores in one command. `--dry` names the account for any tab that isn't on the default one. @@ -445,6 +508,16 @@ cctabs restore --manifest snapshot.json --dry # preview first cctabs restore --manifest snapshot.json --create-missing # spawn tabs for entries with none ``` +**`session_id: null` now says why.** Every row carries `session_lookup`: +`found`, `not-found` (searched every config dir; nothing is titled after this +tab — with `sessions_in_dir` counting the transcripts that *do* exist for its +directory, so a renamed tab is distinguishable from a directory nothing has ever +run in), `no-cwd` (the tab reported no directory, so nothing was looked up), or +`lookup-failed` (the search itself threw — `session_lookup_error` has the +reason). A bare null conflated all four, and they call for opposite responses: +one is a tab to spawn fresh, the others are problems to fix before touching the +fleet. + `--manifest -` reads from stdin, so `cctabs sessions --json | cctabs restore --manifest - --create-missing` works as a one-liner. Entries for tabs that are already running are reported as "already running, skipping" — safe to re-run. `backend` / `config_dir` are emitted only for sessions belonging to a non-default Claude account, and restore infers them anyway from wherever it finds the session, so a hand-written manifest can omit them. **Permission mode travels with the manifest.** `cctabs sessions --json` records each tab's mode as `permission_mode`, read from Claude's own footer, and restore hands it back with `claude --permission-mode ` — so a tab that was in plan mode comes back in plan mode instead of in whatever the global `claude.flags` produce. The flag is appended after those flags and wins; it composes with `--allow-dangerously-skip-permissions`, which only makes bypass *available* rather than selecting it. Entries with no recorded mode fall back to the configured flags, and restore says how many did so rather than doing it silently. @@ -530,14 +603,32 @@ cctabs new payments ~/Dev/myapp --prompt "implement the billing endpoint" cctabs new payments ~/Dev/myapp --file /tmp/task.txt ``` -If you need to send a task after the fact, poll first: +If you need to send a task after the fact, poll first — and for anything +sizeable, hand over a **path** rather than the text: ```bash cctabs new payments ~/Dev/myapp -# Poll until ❯ appears (typically 10-15s with MCP servers) -cctabs scrollback payments 5 # repeat until you see ❯ -cctabs send payments --file /tmp/task.txt -cctabs send payments "yes\n" # quick replies +cctabs send payments --wait-for-prompt --path /tmp/task.txt # waits, then hands over the path +cctabs send payments "yes\n" # quick replies go inline +``` + +⛔ **Deliver anything large with `--path`, not `--file`.** `--file` pastes the +contents through the prompt line, and a paste can arrive as a *fraction* of +itself: a measured 6,835-byte brief landed as its last 756 bytes, beginning +mid-word, and both ends reported success. `--path` has no truncation surface at +all — only the path crosses the prompt line — so use it for briefs, specs and +diffs, and keep inline text for short replies. + +⚠️ **The screen cannot tell you whether a big paste arrived whole.** Claude +collapses it into a `[Pasted text #N +M lines]` chip, and `M` does **not** track +the payload: a 6,892-byte, 76-line payload was measured arriving *complete* into +an idle tab while its chip read `+10 lines`. So `send` reports "arrived, +completeness unverified" rather than a ✔, and if you need certainty add +`--verify` — it reads the target session's own transcript, which records what it +actually received, and fails loudly naming which end went missing. + +```bash +cctabs send payments --path /tmp/brief.txt --verify # belt and braces for a brief that matters ``` **Do NOT call `cctabs send` immediately after `cctabs new`** — Claude is still starting up and the text will land as raw shell commands. @@ -570,11 +661,30 @@ cctabs scrollback auth 200 # last 200 lines ```bash cctabs send auth "yes\n" # approve a tool call cctabs send auth "\n" # press enter (confirm a prompt) +cctabs send auth --submit # press Enter only — submits a prompt already parked in the box cctabs send auth "/clear\n" # send a slash command -cctabs send auth --file ~/prompts/task.txt # send a full prompt from file +cctabs send auth --path ~/prompts/task.txt # hand over a path — the safe way for anything large +cctabs send auth --file ~/prompts/task.txt # paste the contents (short payloads only, see above) echo "do the thing" | cctabs send auth # pipe via stdin ``` +**What `send` now refuses to do, and why it matters when driving a fleet:** + +- It **distinguishes three claims that used to be one ✔ line**: nothing arrived + (a hard failure), something arrived but completeness is unverified (a + warning), and verified. "Sent" and "arrived whole" are not the same fact. +- It **will not submit a body that did not land at all.** Pressing Enter on a + fragment sends something that reads as a complete message. The text is left in + the input box instead, and the command exits non-zero. +- `--verify` **compares what the session received** against what was sent, via + the target's transcript — the only reliable completeness check there is. +- It **refuses payloads over 1 KB into a tab with a turn in flight** (`--force` + overrides). Short replies into a busy tab still work — that's what they're + for. +- `--wait-for-prompt` reads the whole tail of the buffer, not just its last + line, so a `Restart to update` banner rendered *below* a ready prompt no + longer makes it time out. + ## Workflow: Remote Control status across the fleet Claude Code's Remote Control (`/rc`, controls a session from claude.ai/code or the mobile app) is a per-process feature — cctabs doesn't manage it directly, but since it manages the tabs *running* those processes, it's the fastest way to audit or repair RC across many sessions at once. diff --git a/src/commands/index.ts b/src/commands/index.ts index 5b15dc2..ac15053 100644 --- a/src/commands/index.ts +++ b/src/commands/index.ts @@ -10,6 +10,7 @@ import { renameCommand } from './rename.js' import { colorCommand } from './color.js' import { whoamiCommand } from './whoami.js' import { scrollbackCommand } from './scrollback.js' +import { transcriptCommand } from './transcript.js' import { sendCommand } from './send.js' import { configCommand } from './config-cmd.js' import { restoreCommand } from './restore.js' @@ -43,6 +44,10 @@ const subCommands = new Map([ ['color', colorCommand], ['whoami', whoamiCommand], ['scrollback', scrollbackCommand], + ['transcript', transcriptCommand], + // `findings` reads better at a call site that is asking "what did that tab + // work out?", which is the question the command exists for. + ['findings', transcriptCommand], ['send', sendCommand], ['config', configCommand], ['restore', restoreCommand], diff --git a/src/commands/restore.test.ts b/src/commands/restore.test.ts index 44291c6..7a27dbc 100644 --- a/src/commands/restore.test.ts +++ b/src/commands/restore.test.ts @@ -8,7 +8,7 @@ import { planRestore, type PlannedEntry, type RestoreAction, type RestoreEntry } import { DEFAULT_CONFIG_ROOT, type ClaudeConfigDir } from '../core/config-dirs.js' import { launchEnvFor } from '../core/backends.js' import { pathToProjectSlug } from '../core/session.js' -import { buildPlanDeps, buildResumeCommand, colorForEntry, describeDecision, summarizeDecision } from './restore.js' +import { buildPlanDeps, buildResumeCommand, colorForEntry, describeDecision, judgeSpawn, plannedOutcome, summarizeDecision, summarizeOutcomes } from './restore.js' /** * An adapter that answers reads and throws on every mutation. @@ -366,3 +366,87 @@ describe('colorForEntry', () => { expect(colorForEntry(entry({}), cfg('chartreuse'), presets)).toBeUndefined() }) }) + +describe('plannedOutcome', () => { + it('marks the three acting actions as pending, not done', () => { + for (const a of ['attach', 'recreate', 'spawn'] as const) { + expect(plannedOutcome(a)).toBe('unverified') + } + }) + + it('marks everything it never touches as not acted on', () => { + for (const a of ['current-tab', 'already-running', 'ambiguous', 'no-terminal', 'no-session', 'unreadable', 'duplicate', 'missing'] as const) { + expect(plannedOutcome(a)).toBe('skipped') + } + }) +}) + +describe('summarizeOutcomes', () => { + // The point of the line: it has to be able to say something other than "0 + // failed". The old one was computed from "did the spawn call return?". + it('counts each category, printing the zeroes too', () => { + expect(summarizeOutcomes(['restored', 'restored', 'failed'])).toBe('2 verified, 0 unconfirmed, 1 failed') + }) + + it('mentions skipped entries only when there are some', () => { + expect(summarizeOutcomes(['restored'])).toBe('1 verified, 0 unconfirmed, 0 failed') + expect(summarizeOutcomes(['restored', 'skipped'])).toBe('1 verified, 0 unconfirmed, 0 failed, 1 not acted on') + }) + + it('handles an empty restore', () => { + expect(summarizeOutcomes([])).toBe('0 verified, 0 unconfirmed, 0 failed') + }) +}) + +describe('judgeSpawn', () => { + const base = { tabPresent: true, hasTermBlock: true, hasProcess: true as boolean | undefined } + + it('verifies a tab that is present, running, and holding the session it asked for', () => { + const v = judgeSpawn({ ...base, requestedSessionId: 'sess-1234abcd', resolvedSessionId: 'sess-1234abcd' }) + expect(v.outcome).toBe('restored') + }) + + // The tab that was "spawned, 0 failed" and simply wasn't there. + it('fails a tab that is missing from the tab list', () => { + const v = judgeSpawn({ ...base, tabPresent: false }) + expect(v.outcome).toBe('failed') + expect(v.note).toContain('not in the tab list') + }) + + it('fails a tab with no terminal in it', () => { + expect(judgeSpawn({ ...base, hasTermBlock: false }).outcome).toBe('failed') + }) + + it('fails a tab whose process is gone', () => { + expect(judgeSpawn({ ...base, hasProcess: false }).outcome).toBe('failed') + }) + + // The other half of the observed failure: the tab looks perfect and the + // conversation is gone, because --resume quietly started a fresh one. + it('fails a tab that came back as a different session than requested', () => { + const v = judgeSpawn({ ...base, requestedSessionId: 'wanted-1', resolvedSessionId: 'other-22' }) + expect(v.outcome).toBe('failed') + expect(v.note).toContain('DIFFERENT session') + }) + + it('accepts any session for an entry that asked for a fresh one', () => { + expect(judgeSpawn({ ...base, resolvedSessionId: 'brand-new-1' }).outcome).toBe('restored') + }) + + // Not a failure: a running tab whose transcript has not appeared yet. + it('leaves a running tab with no session on disk unconfirmed', () => { + const v = judgeSpawn({ ...base, requestedSessionId: 'sess-1' }) + expect(v.outcome).toBe('unverified') + expect(v.note).toContain('no session is on disk') + }) + + // A backend that cannot report pids must not turn every tab into a failure. + it('does not fail a tab just because the backend cannot report processes', () => { + const v = judgeSpawn({ tabPresent: true, hasTermBlock: true, hasProcess: undefined, resolvedSessionId: 'sess-1' }) + expect(v.outcome).toBe('restored') + }) + + it('is unconfirmed, not failed, when neither process nor session can be read', () => { + expect(judgeSpawn({ tabPresent: true, hasTermBlock: true, hasProcess: undefined }).outcome).toBe('unverified') + }) +}) diff --git a/src/commands/restore.ts b/src/commands/restore.ts index 107feae..739e7fa 100644 --- a/src/commands/restore.ts +++ b/src/commands/restore.ts @@ -6,7 +6,7 @@ import { consola } from 'consola' import { loadConfig } from '../core/config.js' import { requireAdapter, type TerminalAdapter } from '../core/adapter.js' import { openSession } from '../core/open-session.js' -import { findSessionsByNameGlobally, locateSessionById, resolveTabSession } from '../core/session.js' +import { findSessionsByNameGlobally, locateSessionById, resetTitleIndexCache, resolveTabSession } from '../core/session.js' import { launchEnvFor, resolveBackend } from '../core/backends.js' import type { ConfigDirScope } from '../core/config-dirs.js' import { parseManifest } from '../core/manifest.js' @@ -230,9 +230,12 @@ async function runRestore(req: RestoreRequest): Promise { } const results = new Map(plan.map((p) => [p, summarizeDecision(p, dryRun)])) + const outcomes = new Map( + plan.map((p) => [p, plannedOutcome(p.action)]), + ) if (!dryRun) { - await executePlan(adapter, plan, results, { + await executePlan(adapter, plan, results, outcomes, { baseOrder, blocksOf: (tabId) => (tabsById.get(tabId) ?? []).map((b) => b.blockid), }) @@ -244,6 +247,28 @@ async function runRestore(req: RestoreRequest): Promise { for (const p of plan) { console.log(` ${p.entry.name}: ${results.get(p)}`) } + + if (dryRun) return + + // The count, from what was verified rather than from what was attempted. + const acted = plan.filter((p) => plannedOutcome(p.action) !== 'skipped') + const finalOutcomes = acted.map((p) => outcomes.get(p) ?? 'unverified') + console.log(`\n${acted.length} tab(s) acted on: ${summarizeOutcomes(finalOutcomes)}`) + + const failed = acted.filter((p) => outcomes.get(p) === 'failed') + if (failed.length) { + // Non-zero exit: a caller scripting a fleet restart has to be able to + // notice, and a restore that lost a tab is not a success. + consola.error(`${failed.length} tab(s) did not come back: ${failed.map((p) => p.entry.name).join(', ')}`) + process.exitCode = 1 + } + const unconfirmed = acted.filter((p) => outcomes.get(p) === 'unverified') + if (unconfirmed.length) { + consola.warn( + `${unconfirmed.length} tab(s) could not be confirmed either way: ${unconfirmed.map((p) => p.entry.name).join(', ')}. ` + + `Read them with \`cctabs transcript \` before briefing anything from them.`, + ) + } } /** @@ -418,11 +443,127 @@ export function summarizeDecision(p: PlannedEntry, dry: boolean): string { } } +/** + * What actually became of one entry. Distinct from the display string so the + * final count is derived from facts rather than parsed back out of prose. + * + * `unverified` is a first-class outcome and not a rounding error: a tab can be + * up with its process running while its session is not yet confirmable, and + * calling that either "restored" or "failed" is a lie in one direction or the + * other. It is the honest third answer. + */ +export type RestoreOutcome = 'restored' | 'failed' | 'unverified' | 'skipped' + +/** The outcome an action implies before execution — pending, or never acted on. */ +export function plannedOutcome(action: PlannedEntry['action']): RestoreOutcome { + return action === 'attach' || action === 'recreate' || action === 'spawn' + ? 'unverified' + : 'skipped' +} + +/** + * The count line, built from outcomes rather than from hope. + * + * The line this replaces read "78 spawned, 0 failed" while one tab was absent + * entirely and another had come back with no session — a summary computed from + * "did the spawn call return?", which cannot report a failure that happens + * after it returns. Every category is printed, including the zeroes, because + * "0 failed" only means something when it was possible for it to say otherwise. + */ +export function summarizeOutcomes(outcomes: RestoreOutcome[]): string { + const n = (o: RestoreOutcome) => outcomes.filter((x) => x === o).length + const parts = [ + `${n('restored')} verified`, + `${n('unverified')} unconfirmed`, + `${n('failed')} failed`, + ] + const skipped = n('skipped') + if (skipped) parts.push(`${skipped} not acted on`) + return parts.join(', ') +} + +/** What a post-spawn check found out about one tab. */ +export interface SpawnEvidence { + /** Is the tab still in the tab list at all? */ + tabPresent: boolean + /** Does it hold a terminal block? */ + hasTermBlock: boolean + /** Whether a process is running — `undefined` when the backend can't say. */ + hasProcess: boolean | undefined + /** The session this entry asked to resume, if any. */ + requestedSessionId?: string + /** The session now resolvable for the tab's name and directory, if any. */ + resolvedSessionId?: string +} + +export interface SpawnVerdict { + outcome: RestoreOutcome + /** The per-entry line, replacing the optimistic one written at spawn time. */ + note: string +} + +/** + * Judge a spawned tab from what the terminal and the transcripts say afterwards. + * + * Pure, because these are the rules that decide whether a restore is reported + * as successful, and they should be readable and testable without a terminal or + * a 78-tab fleet. Each branch is a failure that has actually been observed and + * reported as success. + */ +export function judgeSpawn(e: SpawnEvidence): SpawnVerdict { + if (!e.tabPresent) { + return { outcome: 'failed', note: '✘ spawn returned but the tab is not in the tab list' } + } + if (!e.hasTermBlock) { + return { outcome: 'failed', note: '✘ tab exists but has no terminal in it' } + } + if (e.hasProcess === false) { + return { outcome: 'failed', note: '✘ tab exists but nothing is running in it' } + } + + // A resume that quietly started a NEW conversation is the loss that hurts: + // the tab looks perfect and the context is gone. It shows up as a different + // session id now answering to this tab's name. + if (e.requestedSessionId && e.resolvedSessionId && e.resolvedSessionId !== e.requestedSessionId) { + return { + outcome: 'failed', + note: `✘ came back as a DIFFERENT session (${e.resolvedSessionId.slice(0, 8)}…, asked for ${e.requestedSessionId.slice(0, 8)}…) — it started a fresh conversation instead of resuming, so the context is not restored`, + } + } + + if (!e.resolvedSessionId) { + return { + outcome: 'unverified', + note: e.hasProcess + ? '? running, but no session is on disk for it yet — check it before relying on its context' + : '? tab is there; neither its process nor its session could be confirmed', + } + } + + return { + outcome: 'restored', + note: `✔ verified running ${e.resolvedSessionId.slice(0, 8)}…`, + } +} + +/** + * How long to let a freshly spawned tab settle before verifying it. + * + * Two things have to have happened: the process has to exist (immediate when + * the backend advertises `spawn-waits-for-pty`, a second or so otherwise), and + * Claude has to have written its `custom-title` line, which is what makes the + * session findable by name. Too short a wait turns healthy tabs into + * `unconfirmed`, which is noise; this is deliberately generous because the + * check runs once for the whole fleet, not once per tab. + */ +const VERIFY_SETTLE_MS = 4000 + /** Carry out a plan. Only ever called for a real (non-dry) run. */ async function executePlan( adapter: TerminalAdapter, plan: PlannedEntry[], results: Map, + outcomes: Map, ctx: { /** Scan mode: the full pre-restore tab order to rebuild. */ baseOrder: string[] | undefined @@ -471,9 +612,18 @@ async function executePlan( await sleep(10_000) for (const p of attached) { const status = adapter.detectSessionStatus(p.blockId!) - if (status === 'active' || status === 'idle') results.set(p, '✔ running') - else if (status === 'unreadable') results.set(p, '? no output captured — check it yourself') - else results.set(p, '✘ may not have started') + if (status === 'active' || status === 'idle') { + results.set(p, '✔ running') + outcomes.set(p, 'restored') + } else if (status === 'unreadable') { + // Readability is not liveness (see session-status.ts) — an empty + // capture is a statement about the capture. Neither pass nor fail. + results.set(p, '? no output captured — check it yourself') + outcomes.set(p, 'unverified') + } else { + results.set(p, '✘ may not have started') + outcomes.set(p, 'failed') + } } } @@ -507,9 +657,13 @@ async function executePlan( finalTabId.set(p, newTabId) if (p.closeTabId) replacements.set(p.closeTabId, newTabId) const verb = p.action === 'recreate' ? 'recreated' : 'spawned' - results.set(p, `✔ ${verb} [${newTabId.slice(0, 8)}] (${shortId(p.sessionId)})`) + // Provisional: the spawn call returning is not the tab working, which + // is the whole reason for the verification pass below. + results.set(p, `… ${verb} [${newTabId.slice(0, 8)}] (${shortId(p.sessionId)}), not yet verified`) + outcomes.set(p, 'unverified') } catch (err) { results.set(p, `✘ ${p.action} failed: ${(err as Error).message}`) + outcomes.set(p, 'failed') } } @@ -531,6 +685,16 @@ async function executePlan( } } + // -- verify what we just spawned, before claiming any of it worked -- + if (toSpawn.length) { + await verifySpawns( + adapter, + toSpawn.filter((p) => finalTabId.has(p)).map((p) => ({ p, tabId: finalTabId.get(p)! })), + results, + outcomes, + ) + } + // -- rebuild the tab bar -- // Best-effort: adapters without reorderTabs keep the append order, and // reorderTabs leaves unlisted tabs in their relative slot, sorted after. @@ -550,6 +714,73 @@ async function executePlan( } } +/** + * Re-read the terminal after spawning and decide what actually came up. + * + * This is the fix for a restore that reported "78 spawned, 0 failed" with one + * tab missing and one stripped of its context. Nothing here trusts the spawn + * call: the tab list is fetched again, the process is looked for, and the + * session is resolved from disk — with the title-index cache dropped first, + * because the cache was built while planning and would happily confirm the + * pre-restore world. + * + * One fleet-wide settle, then one round of reads. A per-tab wait would turn a + * 78-tab restore's verification into minutes. + */ +async function verifySpawns( + adapter: TerminalAdapter, + spawned: Array<{ p: PlannedEntry; tabId: string }>, + results: Map, + outcomes: Map, +): Promise { + if (!spawned.length) return + + consola.info(`Verifying ${spawned.length} spawned tab(s)…`) + await sleep(VERIFY_SETTLE_MS) + resetTitleIndexCache() + + let tabsById: Map + try { + ({ tabsById } = await adapter.getAllData()) + } catch (err) { + // Losing the terminal at this point tells us nothing about the tabs, so + // they stay `unverified` — reporting them as failed would be as wrong as + // reporting them as restored. + consola.warn(`Could not re-read the tab list to verify the restore: ${(err as Error).message}`) + return + } + + // Same reasoning as buildPlanDeps: if any tab reports a pid, the ones that + // don't genuinely have no process. If none do, the backend can't say. + const reportsPids = [...tabsById.values()] + .some((blocks) => blocks.some((b) => typeof b.pid === 'number')) + + for (const { p, tabId } of spawned) { + const blocks = tabsById.get(tabId) ?? [] + const term = blocks.find((b) => b.view === 'term') + + let resolvedSessionId: string | undefined + if (p.dir) { + try { + resolvedSessionId = resolveTabSession(p.dir, p.entry.name)?.id + } catch { + // An unreadable projects dir leaves the session unconfirmed, which + // judgeSpawn already treats as its own answer. + } + } + + const verdict = judgeSpawn({ + tabPresent: blocks.length > 0, + hasTermBlock: !!term, + hasProcess: reportsPids ? typeof term?.pid === 'number' : undefined, + requestedSessionId: p.sessionId, + resolvedSessionId, + }) + results.set(p, `${verdict.note} [${tabId.slice(0, 8)}]`) + outcomes.set(p, verdict.outcome) + } +} + const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms)) /** diff --git a/src/commands/scrollback.ts b/src/commands/scrollback.ts index 0169550..b5acede 100644 --- a/src/commands/scrollback.ts +++ b/src/commands/scrollback.ts @@ -1,47 +1,37 @@ import { define } from 'gunshi' import { consola } from 'consola' import { requireAdapter } from '../core/adapter.js' +import { resolveTabTarget } from '../core/tab-target.js' export const scrollbackCommand = define({ name: 'scrollback', - description: 'Show terminal output for a tab or block (default: last 50 lines)', + description: 'Show terminal output for a tab or block (default: last 50 lines). This is the last PAINTED FRAME — a tab mid-turn shows a spinner and little else, so use `cctabs transcript` to read what it has actually said.', args: { target: { type: 'positional', description: 'Tab name, tab ID prefix, or block ID prefix' }, - lines: { type: 'number', description: 'Number of lines to show', default: 50 }, + // No gunshi `default` here on purpose: a default fills the value in + // unconditionally, which would make the `[n]` positional below unreachable. + lines: { type: 'number', description: 'Number of lines to show (default: 50)' }, }, async run(ctx) { const query = ctx.positionals[1] - const lines = (ctx.values.lines as number | undefined) ?? 50 + // `[n]` as a bare second positional, matching the shape the skill documents + // (`cctabs scrollback [n]`), which --lines alone never supported. + const positionalN = Number(ctx.positionals[2]) + const lines = + (ctx.values.lines as number | undefined) ?? + (Number.isFinite(positionalN) && positionalN > 0 ? positionalN : 50) if (!query) { consola.error('Tab name or block ID is required'); process.exit(1) } const adapter = requireAdapter() const { tabsById, tabNames } = await adapter.getAllData() - // Try tab name resolution first (same logic as `send`) - const tabMatches = adapter.resolveTab(query, tabsById, tabNames) - let blockId: string - - if (tabMatches.length === 1) { - const blocks = (tabsById.get(tabMatches[0]) ?? []).filter((b) => b.view === 'term') - if (!blocks.length) { consola.error(`Tab "${tabNames.get(tabMatches[0])}" has no terminal block`); process.exit(1) } - blockId = blocks[0].blockid - } else if (tabMatches.length > 1) { - consola.error(`Multiple tabs match '${query}':`) - for (const tid of tabMatches) consola.log(` "${tabNames.get(tid)}" [${tid.slice(0, 8)}]`) + const resolved = resolveTabTarget(adapter, query, tabsById, tabNames) + if (!resolved.ok) { + consola.error(resolved.message) + for (const line of resolved.lines ?? []) consola.log(line) process.exit(1) - } else { - // Fall back to block ID prefix resolution - const allBlocks = adapter.blocksList() - const blockMatches = adapter.resolveBlock(query, allBlocks) - if (!blockMatches.length) { consola.error(`No tab or block matching '${query}' (tabs in workspaces with no open window are not visible — open that workspace first)`); process.exit(1) } - if (blockMatches.length > 1) { - consola.error(`Multiple blocks match '${query}':`) - for (const b of blockMatches) consola.log(` ${b.blockid}`) - process.exit(1) - } - blockId = blockMatches[0].blockid } - process.stdout.write(adapter.scrollback(blockId, lines)) + process.stdout.write(adapter.scrollback(resolved.target.blockId, lines)) }, }) diff --git a/src/commands/send.test.ts b/src/commands/send.test.ts new file mode 100644 index 0000000..4c4bf90 --- /dev/null +++ b/src/commands/send.test.ts @@ -0,0 +1,20 @@ +import { describe, it, expect } from 'bun:test' +import { buildPathHandoff, CLIP_RISK_BYTES } from './send.js' + +describe('buildPathHandoff', () => { + it('names the absolute path and says the contents ARE the message', () => { + const msg = buildPathHandoff('/Users/x/.cctabs-prompts/brief.txt') + expect(msg).toContain('/Users/x/.cctabs-prompts/brief.txt') + expect(msg).toContain('in full') + // Without this the receiving session treats a brief as reading material + // and summarises it back instead of acting on it. + expect(msg).toContain('not a document to summarise') + }) + + // The property that makes --path the robust mode: what crosses the prompt + // line is bounded by the path length, not by the payload. + it('stays far below the size where clipping was measured', () => { + const msg = buildPathHandoff('/Users/someone/a/fairly/deep/path/to/a/brief-file.txt') + expect(msg.length).toBeLessThan(CLIP_RISK_BYTES) + }) +}) diff --git a/src/commands/send.ts b/src/commands/send.ts index b07fca0..e7f9031 100644 --- a/src/commands/send.ts +++ b/src/commands/send.ts @@ -1,8 +1,14 @@ import { define } from 'gunshi' import { consola } from 'consola' +import { readFileSync, existsSync } from 'fs' +import { resolve } from 'path' import { requireAdapter } from '../core/adapter.js' import { sendTextWithConfirmation } from '../core/open-session.js' -import { readFileSync } from 'fs' +import { resolveTabTarget } from '../core/tab-target.js' +import { classifyTerminalBuffer, promptIsReady } from '../core/session-status.js' +import { judgeDelivery } from '../core/paste-confirm.js' +import { resolveTabSession } from '../core/session.js' +import { locateTranscriptFile, readLastUserMessage } from '../core/transcript.js' function readStdin(): Promise { return new Promise((resolve) => { @@ -12,13 +18,43 @@ function readStdin(): Promise { }) } +/** + * Payload size above which sending into a busy tab is refused outright. + * + * Short replies are the whole point of sending into an active tab — `yes` to a + * tool call, `2` to a picker — and refusing those would break the flow this + * command exists for. A multi-kilobyte brief is a different act, and pushing + * one into a tab mid-turn is what produced the measured clipping. So the guard + * is on size, not merely on busyness. + */ +export const CLIP_RISK_BYTES = 1024 + +/** + * The message that hands a tab a file to read instead of pasting its contents. + * + * This is the only send shape with no truncation surface at all: what crosses + * the prompt line is a path of a hundred-odd bytes, and the payload is read + * from disk by the receiving session. It became the operator's standing + * practice for briefs after a 6,835-byte paste arrived as 756 bytes, and it is + * first-class here for that reason rather than as a convenience. + */ +export function buildPathHandoff(absPath: string): string { + return `Read the file at ${absPath} in full, and treat its entire contents as the message intended for you — it is your instructions, not a document to summarise.` +} + export const sendCommand = define({ name: 'send', - description: 'Send input to a tab or block (text arg, --file, or stdin pipe)', + description: 'Send input to a tab or block (text arg, --file, --path, or stdin pipe)', args: { target: { type: 'positional', description: 'Tab name, tab ID prefix, or block ID prefix' }, - file: { type: 'string', short: 'f', description: 'Read text from file' }, + file: { type: 'string', short: 'f', description: 'Read the text to send from a file' }, + path: { type: 'string', short: 'p', description: 'Hand the tab this file PATH and let the receiving session read it, instead of pasting the contents. The robust way to deliver anything large: nothing but a path crosses the prompt line, so there is no truncation surface.' }, + submit: { type: 'boolean', description: 'Send Enter only, submitting whatever is already parked in the tab\'s input box. Explicit alternative to guessing at an empty send.' }, enter: { type: 'boolean', short: 'e', description: 'Append newline after text (default: true)' }, + force: { type: 'boolean', description: `Send even when the target has a turn in flight. Without this, payloads over ${CLIP_RISK_BYTES} bytes are refused for a busy tab, because that is when text gets silently clipped.` }, + 'no-confirm': { type: 'boolean', description: 'Skip the did-it-land check and report only that the bytes were handed over. Faster, and honest about knowing less.' }, + verify: { type: 'boolean', description: "After submitting, read the target session's own transcript and check that what it RECEIVED matches what was sent, front and tail. The only reliable completeness check: a collapsed paste chip cannot be read for it. Fails loudly on a mismatch." }, + 'verify-timeout': { type: 'number', description: 'Seconds to wait for the target to record the message when using --verify (default: 30)' }, 'wait-for-prompt': { type: 'boolean', short: 'w', description: 'Poll the buffer until a ready prompt is visible before sending — a shell prompt ($, %, >) or a ready Claude TUI (❯ input line / "auto mode" footer). Useful for freshly-spawned tabs.' }, 'wait-timeout': { type: 'number', description: 'Timeout in seconds for --wait-for-prompt (default: 10)' }, }, @@ -28,16 +64,45 @@ export const sendCommand = define({ // (declaring it as a positional makes gunshi require it, breaking --file and stdin) const inlineText = ctx.positionals[2] const filePath = ctx.values.file as string | undefined + const handoffPath = ctx.values.path as string | undefined + const submitOnly = (ctx.values.submit as boolean | undefined) ?? false const appendEnter = (ctx.values.enter as boolean | undefined) ?? true + const force = (ctx.values.force as boolean | undefined) ?? false + const skipConfirm = (ctx.values['no-confirm'] as boolean | undefined) ?? false + const verify = (ctx.values.verify as boolean | undefined) ?? false + const verifyTimeoutSec = (ctx.values['verify-timeout'] as number | undefined) ?? 30 const waitForPrompt = (ctx.values['wait-for-prompt'] as boolean | undefined) ?? false const waitTimeoutSec = (ctx.values['wait-timeout'] as number | undefined) ?? 10 if (!query) { consola.error('Usage: cctabs send [text]'); process.exit(1) } - // Resolve text source: inline arg > --file > stdin + // Reject contradictory sources up front rather than silently preferring + // one: "I passed --path and it sent the file contents" is the kind of + // surprise this command can't afford any more of. + const sources = [ + inlineText !== undefined && 'inline text', + filePath !== undefined && '--file', + handoffPath !== undefined && '--path', + submitOnly && '--submit', + ].filter(Boolean) as string[] + if (sources.length > 1) { + consola.error(`Pick one text source — ${sources.join(', ')} were all given.`) + process.exit(1) + } + + // Resolve text source: --submit > inline arg > --path > --file > stdin let rawText: string - if (inlineText !== undefined) { + if (submitOnly) { + rawText = '' + } else if (inlineText !== undefined) { rawText = inlineText.replace(/\\n/g, '\r').replace(/\\r/g, '\r').replace(/\\t/g, '\t') + } else if (handoffPath !== undefined) { + const abs = resolve(handoffPath) + if (!existsSync(abs)) { + consola.error(`--path ${abs} does not exist. The receiving session would be sent a path to nothing.`) + process.exit(1) + } + rawText = buildPathHandoff(abs) } else if (filePath) { rawText = readFileSync(filePath, 'utf-8').replace(/\n/g, '\r') } else { @@ -49,50 +114,45 @@ export const sendCommand = define({ // whether to fire an Enter afterwards. let sendEnter = appendEnter if (rawText.endsWith('\r')) { rawText = rawText.replace(/\r+$/, ''); sendEnter = true } + if (submitOnly) sendEnter = true + + // An empty body with no explicit --submit is almost always a mistake — an + // empty file, or a pipe that produced nothing — and it presses Enter in + // someone's session either way. Say what happened instead of doing it + // silently, and point at the flag that means it on purpose. + if (!submitOnly && rawText.length === 0) { + consola.warn( + `Nothing to send${filePath ? ` — ${filePath} is empty` : ''}. ` + + `Pressing Enter only; pass --submit if that is what you meant.`, + ) + } const adapter = requireAdapter() const { tabsById, tabNames } = await adapter.getAllData() - // Try tab resolution first, fall back to block resolution - const tabMatches = adapter.resolveTab(query, tabsById, tabNames) - let blockId: string - - if (tabMatches.length === 1) { - const blocks = (tabsById.get(tabMatches[0]) ?? []).filter((b) => b.view === 'term') - if (!blocks.length) { consola.error(`Tab "${tabNames.get(tabMatches[0])}" has no terminal block`); process.exit(1) } - blockId = blocks[0].blockid - } else if (tabMatches.length > 1) { - consola.error(`Multiple tabs match '${query}':`) - for (const tid of tabMatches) consola.log(` "${tabNames.get(tid)}" [${tid.slice(0, 8)}]`) + const resolved = resolveTabTarget(adapter, query, tabsById, tabNames) + if (!resolved.ok) { + adapter.closeSocket() + consola.error(resolved.message) + for (const line of resolved.lines ?? []) consola.log(line) + process.exit(1) + } + const { blockId, name: tabName, cwd: tabCwd } = resolved.target + + if (verify && (!tabName || !tabCwd)) { + adapter.closeSocket() + consola.error('--verify needs a named tab: it reads the target session\'s transcript, and a bare block has no session to read.') process.exit(1) - } else { - // Fall back to block resolution - const allBlocks = adapter.blocksList() - const blockMatches = adapter.resolveBlock(query, allBlocks) - if (!blockMatches.length) { consola.error(`No tab or block matching '${query}' (tabs in workspaces with no open window are not visible — open that workspace first)`); process.exit(1) } - if (blockMatches.length > 1) { - consola.error(`Multiple blocks match '${query}':`) - for (const b of blockMatches) consola.log(` ${b.blockid}`) - process.exit(1) - } - blockId = blockMatches[0].blockid } if (waitForPrompt) { const deadline = Date.now() + waitTimeoutSec * 1000 let ready = false while (Date.now() < deadline) { - const tail = adapter.scrollback(blockId, 8) - const lastLine = tail.split('\n').map((l) => l.trim()).filter(Boolean).at(-1) ?? '' - const stripped = tail.replace(/\s+/g, '') - // Ready when we see EITHER a bare shell prompt at end-of-line ($ % >), - // OR Claude's TUI input line — which starts with ❯ but is usually - // followed by a "Try …" placeholder, so the glyph is NOT at end-of-line - // — OR Claude's input footer ("auto mode" / "for agents"). The old - // end-anchored `/[$%>❯]\s*$/` never matched a ready Claude TUI. - if (/[$%>]\s*$/.test(lastLine) || /^❯/.test(lastLine) || /automode|foragents/i.test(stripped)) { - ready = true; break - } + // 20 rows, not 8: Claude renders notices (`Restart to update`) BELOW + // the input line, and a short window plus a last-line-only test is + // what made this time out against tabs that were ready. + if (promptIsReady(adapter.scrollback(blockId, 20))) { ready = true; break } await new Promise((r) => setTimeout(r, 250)) } if (!ready) { @@ -102,16 +162,62 @@ export const sendCommand = define({ } } + // Refuse to push a large payload into a tab with a turn in flight. This is + // one of the two ways a send gets clipped, and unlike the other it is + // knowable beforehand. + if (!force && rawText.length > 0) { + const status = classifyTerminalBuffer(adapter.scrollback(blockId, 200)) + if (status === 'active') { + if (rawText.length > CLIP_RISK_BYTES) { + adapter.closeSocket() + consola.error( + `${blockId.slice(0, 8)} has a turn in flight, and ${rawText.length} bytes sent into a busy input handler is how text gets silently clipped. ` + + `Wait for it to finish, hand it a file with \`--path\` (no truncation surface), or pass --force if you accept the risk.`, + ) + process.exit(1) + } + consola.warn(`${blockId.slice(0, 8)} has a turn in flight — sending anyway, as this is short enough to be a reply rather than a brief.`) + } + } + // Send the body — confirming it actually landed and re-sending if not, since // text sent into a not-yet-ready input handler can be silently lost, front - // first (see sendTextWithConfirmation) — then the submit Enter as a - // SEPARATE event. A Claude TUI treats a "text + \r" burst as one paste and - // absorbs the \r as a newline in the input box instead of submitting; a lone - // \r a beat later lands as a real Enter keypress. (Harmless for a plain - // shell — same as typing then pressing return.) Skip the body send when - // it's empty (a bare-Enter send). - let landedOk = true - if (rawText.length > 0) landedOk = await sendTextWithConfirmation(adapter, blockId, rawText) + // first, or arrive as a fraction of itself behind a paste chip that looks + // identical to a complete one (see sendTextWithConfirmation) — then the + // submit Enter as a SEPARATE event. A Claude TUI treats a "text + \r" burst + // as one paste and absorbs the \r as a newline in the input box instead of + // submitting; a lone \r a beat later lands as a real Enter keypress. + // (Harmless for a plain shell — same as typing then pressing return.) + let landed = true + let confirmed = false + let unchecked = skipConfirm + let detail = 'not checked (--no-confirm)' + if (rawText.length > 0) { + if (skipConfirm) { + await adapter.sendInput(blockId, rawText) + } else { + const result = await sendTextWithConfirmation(adapter, blockId, rawText) + landed = result.landed + confirmed = result.confirmed + unchecked = !!result.unchecked + detail = result.detail + } + } + + // A body that did not land must NOT be submitted. Pressing Enter now would + // send whatever fraction arrived as though it were the whole message, and a + // truncated brief reads as a complete one to whoever receives it — which is + // exactly the failure being fixed, with the retry moved one step later. + if (!landed) { + adapter.closeSocket() + consola.error( + `Delivery to ${blockId.slice(0, 8)} could not be confirmed: ${detail}. NOT submitting — a partial message looks like a whole one. ` + + `The text is left in the tab's input box for you to inspect (ctrl+u clears it). ` + + `For anything large, \`cctabs send ${query} --path \` avoids the prompt line entirely.`, + ) + process.exit(1) + } + let resp: unknown if (sendEnter) { if (rawText.length > 0) await new Promise((r) => setTimeout(r, 200)) @@ -121,11 +227,77 @@ export const sendCommand = define({ if (resp && (resp as Record).error) { consola.error(String((resp as Record).error)); process.exit(1) } + const preview = rawText.slice(0, 80).replace(/\n/g, '↵').replace(/\t/g, '→') const label = rawText.length > 0 ? `${JSON.stringify(preview)}${rawText.length > 80 ? '…' : ''}${sendEnter ? ' ⏎' : ''}` : '⏎' - if (!landedOk) { - consola.warn(`Text may not have landed in ${blockId.slice(0, 8)} — its front can be dropped by a not-yet-ready input handler. Check the tab by hand before trusting it arrived.`) + + // --verify: ask the RECEIVING session what it got. The screen cannot answer + // this for a collapsed paste (see paste-confirm.ts), and the transcript can + // — it records the user message as received, so the payload's front and + // tail are checkable against ground truth. + if (verify && rawText.length > 0 && sendEnter) { + const outcome = await verifyDelivered(tabName!, tabCwd!, rawText, verifyTimeoutSec) + if (!outcome.delivered) { + consola.error(`Sent to ${blockId.slice(0, 8)}, but delivery does NOT check out: ${outcome.detail}`) + process.exit(1) + } + consola.success(`Sent to ${blockId.slice(0, 8)} and verified: ${outcome.detail}`) + return + } + if (verify && rawText.length > 0 && !sendEnter) { + consola.warn('--verify has nothing to check without a submit: the message is only recorded once the turn starts.') + } + + // "Sent", "landed", and "verified to have arrived whole" are three + // different claims. They used to be one ✔ line, which is how a partial + // delivery came to look like a success. + if (rawText.length === 0) { + consola.success(`Sent to ${blockId.slice(0, 8)}: ${label}`) + } else if (unchecked) { + consola.info(`Sent to ${blockId.slice(0, 8)} (unconfirmed — ${detail}): ${label}`) + } else if (!confirmed) { + consola.warn(`Sent to ${blockId.slice(0, 8)} — arrived, but completeness UNVERIFIED: ${detail}. ${label}`) + } else { + consola.success(`Sent to ${blockId.slice(0, 8)}: ${label}`) } - consola.success(`Sent to ${blockId.slice(0, 8)}: ${label}`) }, }) + +/** + * Poll the target's transcript until it records the message, then compare. + * + * Waits rather than reading once: the message is written when the turn starts, + * which is a beat after Enter. A timeout is reported as a failed verification + * rather than a pass, because "I could not check" must not read as "it arrived". + */ +async function verifyDelivered( + tabName: string, + tabCwd: string, + payload: string, + timeoutSec: number, +): Promise<{ delivered: boolean; detail: string }> { + const session = resolveTabSession(tabCwd, tabName) + if (!session) { + return { delivered: false, detail: `no session resolved for tab "${tabName}" in ${tabCwd}, so there is no transcript to check` } + } + const located = locateTranscriptFile(session.id) + if (!located) { + return { delivered: false, detail: `session ${session.id.slice(0, 8)} has no transcript on disk under any Claude config dir` } + } + + const deadline = Date.now() + timeoutSec * 1000 + let last = judgeDelivery(payload, null) + while (Date.now() < deadline) { + await new Promise((r) => setTimeout(r, 1000)) + let received: string | null = null + try { + received = readLastUserMessage(located.file) + } catch { + // A transcript being appended to mid-read; try again. + continue + } + last = judgeDelivery(payload, received) + if (last.delivered) return last + } + return last +} diff --git a/src/commands/sessions.ts b/src/commands/sessions.ts index 76bfc3b..76edb10 100644 --- a/src/commands/sessions.ts +++ b/src/commands/sessions.ts @@ -2,6 +2,8 @@ import { define } from 'gunshi' import { requireAdapter } from '../core/adapter.js' import { resolveTabSession } from '../core/session.js' import { classifyTerminalBuffer, parsePermissionMode } from '../core/session-status.js' +import { countSessionsInDir } from '../core/transcript.js' +import { classifySessionLookup, type SessionLookup } from '../core/session-lookup.js' /** * Rows of captured output to read per tab. @@ -36,6 +38,20 @@ export const sessionsCommand = define({ status: string last_line: string session_id: string | null + /** + * Whether `session_id` is absent because we looked and found nothing, + * or because we couldn't look — see {@link SessionLookup}. Always + * present, so a caller never has to infer it from a bare null. + */ + session_lookup: SessionLookup + /** + * For `session_lookup: "not-found"`: how many transcripts exist for + * this tab's directory under any title. `> 0` means the tab was + * renamed out from under a live session rather than having none. + */ + sessions_in_dir?: number + /** For `session_lookup: "lookup-failed"`: what went wrong. */ + session_lookup_error?: string /** * Permission mode read from the session's own footer, so `restore` can * put the tab back the way it was instead of in whatever the global @@ -97,6 +113,10 @@ export const sessionsCommand = define({ // CLAUDE_CONFIG_DIR can't be resumed without it. let backend: string | undefined let configDir: string | undefined + // The failure is captured rather than swallowed: a lookup that threw + // is reported as `lookup-failed`, which is not the same answer as + // "this tab has no session" and must not be flattened into it. + let lookupError: Error | undefined if (cwd) { try { const resolved = resolveTabSession(cwd, tabName) @@ -106,10 +126,16 @@ export const sessionsCommand = define({ backend = resolved.backend configDir = resolved.configDir } - } catch { - // ignore — best-effort lookup + } catch (err) { + lookupError = err as Error } } + const lookup = classifySessionLookup({ + cwd, + found: sessionId !== null, + error: lookupError, + countInDir: () => countSessionsInDir(cwd), + }) wsRow.sessions.push({ block_id: b.blockid, @@ -120,6 +146,9 @@ export const sessionsCommand = define({ status, last_line: lastLine.slice(0, 200), session_id: sessionId, + session_lookup: lookup.status, + ...(lookup.sessionsInDir !== undefined ? { sessions_in_dir: lookup.sessionsInDir } : {}), + ...(lookup.detail ? { session_lookup_error: lookup.detail } : {}), ...(b.color !== undefined ? { color: b.color } : {}), ...(permissionMode ? { permission_mode: permissionMode } : {}), ...(backend ? { backend } : {}), diff --git a/src/commands/sort.test.ts b/src/commands/sort.test.ts index 5dd88ad..a94daa6 100644 --- a/src/commands/sort.test.ts +++ b/src/commands/sort.test.ts @@ -1,5 +1,5 @@ import { describe, expect, test } from 'bun:test' -import { rankTabsByActivity } from './sort.js' +import { parseFirstList, planPinnedOrder, rankTabsByActivity } from './sort.js' const names = (m: Record) => new Map(Object.entries(m)) const times = (m: Record) => new Map(Object.entries(m)) @@ -52,3 +52,64 @@ describe('rankTabsByActivity', () => { expect(ranked.map((r) => r.name)).toEqual(['career-strategy', 'other']) }) }) + +describe('parseFirstList', () => { + test('splits on commas and tolerates spacing', () => { + expect(parseFirstList('auth, payments ,billing')).toEqual(['auth', 'payments', 'billing']) + }) + + test('drops empty entries from stray commas', () => { + expect(parseFirstList('auth,,payments,')).toEqual(['auth', 'payments']) + expect(parseFirstList('')).toEqual([]) + }) +}) + +describe('planPinnedOrder', () => { + // The tab ids each name resolves to, standing in for adapter.resolveTab. + const resolver = (map: Record) => (q: string) => map[q] ?? [] + + test('keeps the requested order, not the bar order', () => { + const plan = planPinnedOrder( + ['payments', 'auth'], + resolver({ auth: ['t-auth'], payments: ['t-pay'] }), + ) + expect(plan.order).toEqual(['t-pay', 't-auth']) + expect(plan.names).toEqual(['payments', 'auth']) + expect(plan.unresolved).toEqual([]) + }) + + test('reports a name that matches nothing', () => { + const plan = planPinnedOrder(['auth', 'ghost'], resolver({ auth: ['t-auth'] })) + expect(plan.unresolved).toEqual([{ query: 'ghost', reason: 'not-found' }]) + }) + + test('reports an ambiguous name rather than picking one', () => { + const plan = planPinnedOrder(['gap'], resolver({ gap: ['t-1', 't-2'] })) + expect(plan.unresolved).toEqual([{ query: 'gap', reason: 'ambiguous' }]) + expect(plan.order).toEqual([]) + }) + + // --first a,b,a can't mean both "a is first" and "a is third". + test('collapses a repeated name to its first mention', () => { + const plan = planPinnedOrder( + ['auth', 'payments', 'auth'], + resolver({ auth: ['t-auth'], payments: ['t-pay'] }), + ) + expect(plan.order).toEqual(['t-auth', 't-pay']) + }) + + // Two names resolving to the same tab is the same contradiction, arriving + // via a name and an id prefix instead of a repeat. + test('collapses two names that resolve to the same tab', () => { + const plan = planPinnedOrder( + ['auth', 'aabbccdd'], + resolver({ auth: ['t-auth'], aabbccdd: ['t-auth'] }), + ) + expect(plan.order).toEqual(['t-auth']) + expect(plan.names).toEqual(['auth']) + }) + + test('an empty request pins nothing', () => { + expect(planPinnedOrder([], resolver({}))).toEqual({ order: [], names: [], unresolved: [] }) + }) +}) diff --git a/src/commands/sort.ts b/src/commands/sort.ts index cb36dbd..5b9fe8c 100644 --- a/src/commands/sort.ts +++ b/src/commands/sort.ts @@ -1,6 +1,7 @@ import { define } from 'gunshi' import { consola } from 'consola' -import { requireAdapter } from '../core/adapter.js' +import { requireAdapter, type TerminalAdapter } from '../core/adapter.js' +import type { Block } from '../types/index.js' import { buildTitleActivityMap } from '../core/session.js' function relAge(ms: number): string { @@ -48,13 +49,62 @@ export function rankTabsByActivity( return ranked.map(({ tid, name, mtime }) => ({ tid, name, mtime })) } +/** What a `--first` request resolved to, before anything is applied. */ +export interface PinnedPlan { + /** Tab ids to place at the front, in the order the caller asked for. */ + order: string[] + /** The names those ids came from, parallel to `order`, for reporting. */ + names: string[] + /** Requested names that matched no tab — or matched more than one. */ + unresolved: Array<{ query: string; reason: 'not-found' | 'ambiguous' }> +} + +/** + * Resolve a `--first a,b,c` request into the tab order to POST. + * + * Only the pinned ids go in the list, and that is deliberate rather than + * lazy: the backend's reorder contract is that tabs absent from `order` keep + * their relative order and sort after the listed ones. So naming three tabs + * moves exactly those three and disturbs nothing else — which is what pinning + * means, and why this doesn't need (or want) the activity scan that `sort` + * otherwise pays ~7.7s of transcript reading for. + * + * Duplicates in the request are collapsed to their first mention, so + * `--first a,b,a` is `a,b` rather than an order that contradicts itself. + */ +export function planPinnedOrder( + queries: string[], + resolve: (query: string) => string[], +): PinnedPlan { + const order: string[] = [] + const names: string[] = [] + const unresolved: PinnedPlan['unresolved'] = [] + + for (const query of queries) { + const matches = resolve(query) + if (matches.length === 0) { unresolved.push({ query, reason: 'not-found' }); continue } + if (matches.length > 1) { unresolved.push({ query, reason: 'ambiguous' }); continue } + if (order.includes(matches[0])) continue + order.push(matches[0]) + names.push(query) + } + + return { order, names, unresolved } +} + +/** Split a `--first` value on commas, dropping empties and surrounding space. */ +export function parseFirstList(raw: string): string[] { + return raw.split(',').map((s) => s.trim()).filter(Boolean) +} + export const sortCommand = define({ name: 'sort', - description: 'Reorder tabs by Claude session activity (most-recent first).', + description: 'Reorder tabs by Claude session activity (most-recent first), or pin a chosen set to the front with --first.', args: { dry: { type: 'boolean', short: 'n', description: 'Show planned order without applying it' }, 'dry-run': { type: 'boolean', description: 'Alias for --dry' }, reverse: { type: 'boolean', short: 'r', description: 'Oldest first instead of newest' }, + first: { type: 'string', description: 'Pin these tabs to the front of the bar, in this order (comma-separated names or id prefixes). Everything else keeps its current relative order, and no activity ranking is done — use this when you want a chosen working set in reach, which is what activity order actively works against.' }, }, async run(ctx) { const dryRun = !!(ctx.values.dry || ctx.values['dry-run']) @@ -67,6 +117,14 @@ export const sortCommand = define({ } const { tabsById, workspaces, tabNames } = await adapter.getAllData() + + const firstRaw = ctx.values.first as string | undefined + if (firstRaw !== undefined) { + await pinToFront(adapter, firstRaw, tabsById, tabNames, dryRun) + adapter.closeSocket() + return + } + const titleMtimes = buildTitleActivityMap() const now = Date.now() @@ -107,3 +165,57 @@ export const sortCommand = define({ adapter.closeSocket() }, }) + +/** + * Apply a `--first` request: move the named tabs to the front of the bar. + * + * Refuses as a whole if any name doesn't resolve. Pinning is something a driver + * does to get a working set in reach, and half a working set in reach — with no + * indication which half — is worse than an error, because the tabs that failed + * are exactly the ones that would then be looked for in the wrong place. Same + * reasoning as restore's true-count reporting: don't claim what wasn't done. + */ +async function pinToFront( + adapter: TerminalAdapter, + firstRaw: string, + tabsById: Map, + tabNames: Map, + dryRun: boolean, +): Promise { + const queries = parseFirstList(firstRaw) + if (!queries.length) { + consola.error('--first needs at least one tab name, e.g. --first auth,payments') + process.exitCode = 1 + return + } + + const plan = planPinnedOrder(queries, (q) => adapter.resolveTab(q, tabsById, tabNames)) + + if (plan.unresolved.length) { + consola.error('Not pinning anything — these did not resolve to exactly one tab:') + for (const u of plan.unresolved) { + consola.log(` ${u.query} — ${u.reason === 'ambiguous' ? 'matches several tabs; use a longer name or an id prefix' : 'no such tab'}`) + } + process.exitCode = 1 + return + } + + consola.info(`Pinning ${plan.order.length} tab(s) to the front:`) + plan.names.forEach((name, i) => { + consola.log(` ${i + 1}. ${name} [${plan.order[i].slice(0, 8)}]`) + }) + consola.log(' … every other tab keeps its current relative order, after these.') + + if (dryRun) { + consola.info('Dry run — no changes applied.') + return + } + + try { + await adapter.reorderTabs!(plan.order) + consola.success(`Pinned ${plan.order.length} tab(s) to the front.`) + } catch (err) { + consola.error(`Failed to reorder tabs: ${(err as Error).message}`) + process.exitCode = 1 + } +} diff --git a/src/commands/transcript.ts b/src/commands/transcript.ts new file mode 100644 index 0000000..b40fda7 --- /dev/null +++ b/src/commands/transcript.ts @@ -0,0 +1,149 @@ +import { define } from 'gunshi' +import { consola } from 'consola' +import { requireAdapter } from '../core/adapter.js' +import { resolveTabTarget } from '../core/tab-target.js' +import { resolveTabSession } from '../core/session.js' +import { classifySessionLookup, explainSessionLookup } from '../core/session-lookup.js' +import { + countSessionsInDir, + locateTranscriptFile, + readAssistantTurns, + truncateTurn, + type AssistantTurn, +} from '../core/transcript.js' + +const DEFAULT_TURNS = 3 + +export const transcriptCommand = define({ + name: 'transcript', + description: "Print what a tab's Claude session has actually SAID — its last assistant messages, read from the transcript rather than the screen. Unlike `scrollback` this works on a tab that is mid-turn (whose screen is just a spinner).", + args: { + target: { type: 'positional', description: 'Tab name, tab ID prefix, or block ID prefix' }, + turns: { type: 'number', short: 'n', description: `How many trailing assistant messages to print (default: ${DEFAULT_TURNS})` }, + chars: { type: 'number', description: 'Truncate each message to this many characters (default: no limit)' }, + json: { type: 'boolean', short: 'j', description: 'Emit machine-readable JSON' }, + }, + async run(ctx) { + const query = ctx.positionals[1] + // `[n]` as a bare second positional, matching the documented shape + // `cctabs transcript [n]`, with --turns as the explicit form. + const positionalN = Number(ctx.positionals[2]) + const turnCount = + (ctx.values.turns as number | undefined) ?? + (Number.isFinite(positionalN) && positionalN > 0 ? positionalN : DEFAULT_TURNS) + const maxChars = ctx.values.chars as number | undefined + const asJson = (ctx.values.json as boolean | undefined) ?? false + + if (!query) { consola.error('Usage: cctabs transcript [n]'); process.exit(1) } + + const adapter = requireAdapter() + const { tabsById, tabNames } = await adapter.getAllData() + const resolved = resolveTabTarget(adapter, query, tabsById, tabNames) + if (!resolved.ok) { + adapter.closeSocket() + consola.error(resolved.message) + for (const line of resolved.lines ?? []) consola.log(line) + process.exit(1) + } + const { name, cwd, blockId } = resolved.target + adapter.closeSocket() + + // A block that isn't in a nameable tab has no session we can resolve — + // resolveTabSession keys on the tab's name, which is what Claude was + // launched with as `--name`. + if (!name || !cwd) { + fail(asJson, blockId, 'no-tab', 'Matched a block, not a named tab — no session to look up. Pass a tab name.') + } + + let session: ReturnType = null + let lookupError: Error | undefined + try { + session = resolveTabSession(cwd!, name!) + } catch (err) { + lookupError = err as Error + } + if (!session) { + // Exactly the distinction `sessions --json` now reports, made by the same + // code: "looked and found nothing" and "couldn't look" are different + // answers, and reading the second as the first is how a healthy tab gets + // written off as dead. + const lookup = classifySessionLookup({ + cwd: cwd!, + found: false, + error: lookupError, + countInDir: () => countSessionsInDir(cwd!), + }) + fail(asJson, blockId, lookup.status, `${name}: ${explainSessionLookup(lookup)}`) + } + + const located = locateTranscriptFile(session!.id) + if (!located) { + fail( + asJson, + blockId, + 'no-transcript', + `Session ${session!.id.slice(0, 8)} resolved but its transcript is not on disk under any Claude config dir. This is a broken state, not an idle tab.`, + ) + } + + let turns: AssistantTurn[] + try { + turns = readAssistantTurns(located!.file) + } catch (err) { + fail(asJson, blockId, 'unreadable', `Could not read ${located!.file}: ${(err as Error).message}`) + return + } + const shown = turns.slice(-turnCount) + + const origin = located!.backend + ? ` backend=${located!.backend}` + : located!.configDir ? ` config_dir=${located!.configDir}` : '' + + if (asJson) { + console.log(JSON.stringify({ + tab: name, + block_id: blockId, + session_id: session!.id, + cwd, + transcript: located!.file, + ...(located!.backend ? { backend: located!.backend } : {}), + ...(located!.configDir ? { config_dir: located!.configDir } : {}), + assistant_messages: turns.length, + turns: shown.map((t) => ({ + timestamp: t.timestamp, + text: truncateTurn(t.text, maxChars), + })), + }, null, 2)) + return + } + + console.log(`# ${name} session=${session!.id.slice(0, 8)} cwd=${cwd}${origin}`) + console.log(`# ${turns.length} assistant message(s) on record; showing the last ${shown.length}`) + if (!shown.length) { + console.log('# (none yet — the session has not answered anything)') + return + } + for (const t of shown) { + console.log(`\n--- ${t.timestamp ?? 'no timestamp'} ---`) + console.log(truncateTurn(t.text, maxChars)) + } + }, +}) + +/** + * Report a lookup that could not produce a transcript, and exit non-zero. + * + * Non-zero because a driver that pipes this into a briefing must not mistake + * "I could not read this session" for "this session has said nothing" — the + * whole failure this command exists to prevent is briefing from a false + * picture. `--json` still gets a parseable object so callers can branch on + * `reason` rather than on exit code alone. + */ +function fail(asJson: boolean, blockId: string, reason: string, message: string): never { + if (asJson) { + console.log(JSON.stringify({ block_id: blockId, reason, error: message }, null, 2)) + } else { + consola.error(message) + } + process.exit(1) +} diff --git a/src/core/open-session.test.ts b/src/core/open-session.test.ts index effb8d4..ee2af51 100644 --- a/src/core/open-session.test.ts +++ b/src/core/open-session.test.ts @@ -141,7 +141,7 @@ describe('sendTextWithConfirmation', () => { const ok = await sendTextWithConfirmation(adapter, 'b1', TEXT, { sleep: async () => {} }) - expect(ok).toBe(true) + expect(ok.landed).toBe(true) expect(inputs).toEqual([TEXT]) // no retry needed }) @@ -150,7 +150,7 @@ describe('sendTextWithConfirmation', () => { const ok = await sendTextWithConfirmation(adapter, 'b1', TEXT, { sleep: async () => {} }) - expect(ok).toBe(true) + expect(ok.landed).toBe(true) expect(inputs).toEqual([TEXT, '\x15', TEXT]) // cleared before the re-send, never stacked }) @@ -167,7 +167,9 @@ describe('sendTextWithConfirmation', () => { const ok = await sendTextWithConfirmation(adapter, 'b1', 'y', { sleep: async () => {} }) - expect(ok).toBe(true) + expect(ok.landed).toBe(true) + // Reported as unchecked, so a caller can't present it as verified. + expect(ok.unchecked).toBe(true) expect(inputs).toEqual(['y']) }) @@ -183,6 +185,34 @@ describe('sendTextWithConfirmation', () => { pollCount: 2, }) - expect(ok).toBe(false) + expect(ok.landed).toBe(false) + expect(ok.confirmed).toBe(false) + expect(ok.detail).toContain('nothing from the text appeared') + }) + + // A collapsed paste is landed but unconfirmed, and that gap is the point: + // measured, a 6,892-byte payload delivered COMPLETELY into an idle tab while + // its chip read "+10 lines", so a short chip cannot be treated as truncation + // — but nor can the chip's presence be treated as proof of completeness. + it('reports a collapsed paste as landed without claiming it is complete', async () => { + const big = Array.from({ length: 120 }, (_, i) => `line ${i} of the brief`).join('\n') + const adapter = { + sendInput: async () => undefined, + scrollback: () => '[Pasted text #1 +10 lines]', + } as unknown as TerminalAdapter + + const verdict = await sendTextWithConfirmation(adapter, 'b1', big, { + sleep: async () => {}, attempts: 1, pollCount: 1, + }) + + expect(verdict.landed).toBe(true) + expect(verdict.confirmed).toBe(false) + expect(verdict.detail).toContain('unverified') + }) + + it('confirms a paste it can actually see echoed', async () => { + const { adapter } = echoAdapter([0]) + const verdict = await sendTextWithConfirmation(adapter, 'b1', TEXT, { sleep: async () => {} }) + expect(verdict.confirmed).toBe(true) }) }) diff --git a/src/core/open-session.ts b/src/core/open-session.ts index 40b2f77..3322792 100644 --- a/src/core/open-session.ts +++ b/src/core/open-session.ts @@ -8,6 +8,7 @@ import { shellQuoteArg } from './shell.js' import { autoModeDialogVisible, trustDialogVisible } from './session-status.js' import { hasPriorSessions } from './session.js' import { applyTabColor, supportsTabColor } from './colors.js' +import { countPayloadLines, judgePaste, readPasteEvidence, type PasteVerdict } from './paste-confirm.js' interface OpenSessionOptions { tabName: string @@ -97,55 +98,78 @@ export interface ConfirmedSendOptions { sleep?: (ms: number) => Promise } +/** The outcome of a confirmed send, with the evidence behind it. */ +export interface ConfirmedSendResult { + /** Something from the payload reached the tab. */ + landed: boolean + /** We could establish that ALL of it reached the tab — see judgePaste. */ + confirmed: boolean + /** Why we concluded that — surfaced verbatim to the operator. */ + detail: string + /** True when the text was too short to fingerprint, so nothing was checked. */ + unchecked?: boolean +} + /** * Send `text`, confirm it actually landed in the input box, and re-send * (clearing the line first) if not. * - * The naive "just send it" is unreliable: text sent into a not-yet-ready - * input handler is silently lost — sometimes entirely, sometimes only its - * front, which reads as a message arriving with its lead-in truncated. The - * fix is the same either way: a distinctive chunk from the *front* of the - * text (the part observed to go missing) doubles as the landed-detector, so - * a front-clipped send is caught exactly like a fully-dropped one. + * The naive "just send it" is unreliable in three distinct ways, and the third + * is the one that cost real work: + * + * 1. The paste can be dropped entirely by a not-yet-ready input handler. + * 2. It can lose its FRONT, which reads as a message whose lead-in was + * trimmed rather than as a failure. A fingerprint of the front doubles as + * the detector, so a front-clip is caught exactly like a full drop. + * 3. It can lose most of itself and still LOOK delivered. Claude collapses a + * large paste into a `[Pasted text #N +M lines]` chip, and the chip + * renders no matter how little arrived — so "a chip is on screen" was + * accepted as proof while 89% of a measured 6,835-byte payload was gone. + * The chip's own line count is now compared against what was sent (see + * paste-confirm.ts), which is the only quantitative handle a collapsed + * paste offers. * * Text too short to fingerprint reliably (under 4 non-whitespace chars, e.g. * a bare "y" answering a prompt) is sent once, unconfirmed, rather than - * looping against ambiguous scrollback noise. + * looping against ambiguous scrollback noise — reported as `unchecked` so a + * caller doesn't present "not verified" as "verified". */ export async function sendTextWithConfirmation( adapter: TerminalAdapter, blockId: string, text: string, opts: ConfirmedSendOptions = {}, -): Promise { +): Promise { const sleepFn = opts.sleep ?? sleep const payload = opts.bracketedPaste ? `\x1b[200~${text}\x1b[201~` : text const sentinel = text.replace(/\s+/g, '').slice(0, 24) + const expectedLines = countPayloadLines(text) if (sentinel.length < 4) { await adapter.sendInput(blockId, payload) - return true + return { landed: true, confirmed: false, unchecked: true, detail: 'too short to fingerprint — sent without confirmation' } } const attempts = opts.attempts ?? 3 const pollCount = opts.pollCount ?? 8 const pollIntervalMs = opts.pollIntervalMs ?? 300 - const landed = (): boolean => { - const c = adapter.scrollback(blockId, 60).replace(/\s+/g, '') - return c.includes(sentinel) || c.includes('[Pastedtext') - } - - let inBox = false - for (let attempt = 0; attempt < attempts && !inBox; attempt++) { - // Clear first on retries so a re-send never stacks a second copy. + const check = (): PasteVerdict => + judgePaste(readPasteEvidence(adapter.scrollback(blockId, 60), sentinel), expectedLines) + + let verdict: PasteVerdict = { landed: false, confirmed: false, detail: 'the send was never attempted' } + for (let attempt = 0; attempt < attempts && !verdict.landed; attempt++) { + // Clear first on retries so a re-send never stacks a second copy. This also + // clears a TRUNCATED previous attempt, which is why a short-chip verdict + // must be a retry rather than an accepted success. if (attempt > 0) { await adapter.sendInput(blockId, '\x15'); await sleepFn(200) } await adapter.sendInput(blockId, payload) for (let i = 0; i < pollCount; i++) { await sleepFn(pollIntervalMs) - if (landed()) { inBox = true; break } + verdict = check() + if (verdict.landed) break } } - return inBox + return verdict } /** @@ -212,9 +236,12 @@ async function sendInitialPrompt( const prompt = readFileSync(initialPromptFile, 'utf-8').trimEnd() // Stage 1: paste, confirm it landed, re-paste if dropped. - const inBox = await sendTextWithConfirmation(adapter, blockId, prompt, { bracketedPaste: true }) - if (!inBox) { - consola.warn('Initial prompt may not have landed in the input box — switch to the tab and press Enter (re-type if the box is empty).') + const sent = await sendTextWithConfirmation(adapter, blockId, prompt, { bracketedPaste: true }) + if (!sent.landed) { + // Do not press Enter. Submitting now would send whatever fraction DID + // arrive as though it were the whole prompt, which is worse than not + // sending it: a truncated brief reads as a complete one. + consola.warn(`Initial prompt did not land in the input box (${sent.detail}) — NOT submitting it, since a partial prompt would look like the whole one. Switch to the tab, clear the box (ctrl+u) and paste it yourself: ${initialPromptFile}`) return } diff --git a/src/core/paste-confirm.test.ts b/src/core/paste-confirm.test.ts new file mode 100644 index 0000000..ffc3c9a --- /dev/null +++ b/src/core/paste-confirm.test.ts @@ -0,0 +1,128 @@ +import { describe, it, expect } from 'bun:test' +import { + countPayloadLines, + judgeDelivery, + judgePaste, + readPasteEvidence, +} from './paste-confirm.js' + +describe('countPayloadLines', () => { + it('counts line breaks, in either newline form', () => { + expect(countPayloadLines('a\nb\nc')).toBe(2) + expect(countPayloadLines('a\rb\rc')).toBe(2) + expect(countPayloadLines('single line')).toBe(0) + expect(countPayloadLines('')).toBe(0) + }) +}) + +describe('readPasteEvidence', () => { + it('reads the chip line count', () => { + const e = readPasteEvidence('[Pasted text #1 +117 lines]', 'nothing') + expect(e.chipSeen).toBe(true) + expect(e.chipLines).toBe(117) + }) + + // Tabby's buffer endpoint drops spaces between glyphs unpredictably, which is + // why everything here matches on a whitespace-stripped copy. + it('reads a chip whose spaces the buffer dropped', () => { + expect(readPasteEvidence('[Pastedtext#2+45lines]', 'x').chipLines).toBe(45) + }) + + it('reads a chip with no paste number, and the singular "1 line"', () => { + expect(readPasteEvidence('[Pasted text +9 lines]', 'x').chipLines).toBe(9) + expect(readPasteEvidence('[Pasted text #1 +1 line]', 'x').chipLines).toBe(1) + }) + + it('notices a chip it cannot get a count out of', () => { + const e = readPasteEvidence('[Pasted text #1]', 'x') + expect(e.chipSeen).toBe(true) + expect(e.chipLines).toBeUndefined() + }) + + it('finds the sentinel across dropped whitespace', () => { + expect(readPasteEvidence('he llo fr om tab', 'hellofrom').sentinelSeen).toBe(true) + }) + + it('reports nothing for an empty buffer, and never sees an empty sentinel', () => { + expect(readPasteEvidence('', 'hellofrom').sentinelSeen).toBe(false) + expect(readPasteEvidence('anything at all', '').sentinelSeen).toBe(false) + }) +}) + +describe('judgePaste', () => { + it('confirms an echoed paste on the strength of its visible front', () => { + const v = judgePaste({ sentinelSeen: true, chipSeen: false }, 0) + expect(v).toMatchObject({ landed: true, confirmed: true }) + }) + + // The correction that cost a measurement to find: a 6,892-byte, 76-line + // payload delivered COMPLETELY into an idle tab while its chip read + // "+10 lines". The chip's count does not track the payload, so a shortfall + // is not evidence of truncation and must not fail the send. + it('treats a collapsed chip as landed but NOT confirmed, whatever it counts', () => { + for (const chipLines of [10, 60, 119, undefined]) { + const v = judgePaste({ sentinelSeen: false, chipSeen: true, chipLines }, 120) + expect(v.landed).toBe(true) + expect(v.confirmed).toBe(false) + } + }) + + it('names the two ways to actually establish completeness', () => { + const v = judgePaste({ sentinelSeen: false, chipSeen: true, chipLines: 10 }, 120) + expect(v.detail).toContain('--verify') + expect(v.detail).toContain('--path') + }) + + it('rejects a buffer showing no sign of the text at all', () => { + const v = judgePaste({ sentinelSeen: false, chipSeen: false }, 40) + expect(v.landed).toBe(false) + expect(v.confirmed).toBe(false) + }) +}) + +describe('judgeDelivery', () => { + const payload = 'FRONT-MARKER: alpha-7391\nfiller filler filler filler filler\nEND-MARKER: omega-5520 do the thing now' + + it('confirms a payload whose front and tail both reached the session', () => { + // Claude appends its own context to the recorded message, so the received + // text is a superset — hence containment rather than equality. + const received = `${payload}\nsome appended context` + const v = judgeDelivery(payload, received) + expect(v.delivered).toBe(true) + }) + + // Whitespace differs on every hop: send converts LF to CR, the transcript + // stores LF, and the terminal may re-wrap. None of that is a failure. + it('ignores whitespace differences between what was sent and what was stored', () => { + const received = payload.replace(/\n/g, '\r\n ') + expect(judgeDelivery(payload, received).delivered).toBe(true) + }) + + // The observed clipping mode, and the reason each end is named separately. + it('identifies a front-clipped delivery as such', () => { + // The tail arrives, the front does not — the received text must be long + // enough to hold the tail fingerprint, as a real clipped payload is. + const v = judgeDelivery(payload, payload.slice(-60)) + expect(v.delivered).toBe(false) + expect(v.detail).toContain('END of the payload but not its FRONT') + }) + + it('identifies a delivery cut short at the end', () => { + const v = judgeDelivery(payload, payload.slice(0, 60)) + expect(v.delivered).toBe(false) + expect(v.detail).toContain('FRONT of the payload but not its end') + }) + + it('reports an unrelated message as matching neither end', () => { + const v = judgeDelivery(payload, 'something else entirely') + expect(v.delivered).toBe(false) + expect(v.detail).toContain('neither end') + }) + + // "I could not check" must never read as "it arrived". + it('reports nothing recorded as a failed verification, not a pass', () => { + const v = judgeDelivery(payload, null) + expect(v.delivered).toBe(false) + expect(v.detail).toContain('has not recorded any message') + }) +}) diff --git a/src/core/paste-confirm.ts b/src/core/paste-confirm.ts new file mode 100644 index 0000000..5839312 --- /dev/null +++ b/src/core/paste-confirm.ts @@ -0,0 +1,169 @@ +/** + * Deciding whether a pasted payload landed in a tab's input box — and being + * honest about the limits of what the screen can tell us. + * + * The history here is worth keeping, because both of the obvious answers are + * wrong and one of them was measured wrong twice: + * + * 1. "A paste chip is on screen, so it arrived." Claude Code collapses any + * sizeable paste into `[Pasted text #1 +117 lines]`, and that chip renders + * whether the whole payload arrived or a fraction did. A 6,835-byte brief + * once landed as its last 756 bytes with both ends reporting success. + * 2. "So compare the chip's line count against what was sent." Also wrong, + * and measured: a 6,892-byte, 76-line payload delivered *completely* into + * an idle tab — front marker, tail marker and all, confirmed by reading + * what the receiving session recorded — while the chip on screen read + * `+10 lines`. The chip's count is not a count of the payload. Treating a + * shortfall as truncation fails healthy sends. + * + * What is left is a real and useful conclusion: for a collapsed paste the + * screen **cannot** establish completeness. So this module distinguishes + * "something arrived" from "all of it arrived, verified", refuses to claim the + * second when it can only see the first, and leaves the actual comparison to + * the transcript (see `judgeDelivery`), which records what the session received. + */ + +/** What the tab's buffer shows about the paste we just sent. */ +export interface PasteEvidence { + /** The front of the payload is visible in the buffer (a small, echoed paste). */ + sentinelSeen: boolean + /** A paste chip is on screen. */ + chipSeen: boolean + /** + * The line count the newest chip reports, when it reports one. + * + * Informational only — reported in messages, never used to decide the + * verdict. See the note above: it does not track the payload. + */ + chipLines?: number +} + +export interface PasteVerdict { + /** Something from the payload reached the tab. */ + landed: boolean + /** We could establish that ALL of it reached the tab. */ + confirmed: boolean + /** Operator-facing explanation of the call, used verbatim in send's output. */ + detail: string +} + +/** Line breaks in a payload. Reported for context, not used as a threshold. */ +export function countPayloadLines(text: string): number { + const matches = text.match(/[\r\n]/g) + return matches ? matches.length : 0 +} + +const CHIP_WITH_COUNT = /\[Pastedtext#?\d*\+(\d+)lines?\]/g + +/** + * Read the buffer for signs of the paste. + * + * Everything is matched against a whitespace-stripped copy: Tabby's buffer + * endpoint drops spaces between glyphs unpredictably, so `[Pasted text #1 +117 + * lines]` can arrive with any subset of its spaces missing. The `sentinel` must + * already be whitespace-stripped by the caller for the same reason. + */ +export function readPasteEvidence(buffer: string, sentinel: string): PasteEvidence { + const stripped = buffer.replace(/\s+/g, '') + + let chipLines: number | undefined + CHIP_WITH_COUNT.lastIndex = 0 + let m: RegExpExecArray | null + while ((m = CHIP_WITH_COUNT.exec(stripped)) !== null) chipLines = Number(m[1]) + + return { + sentinelSeen: sentinel.length > 0 && stripped.includes(sentinel), + chipSeen: stripped.includes('[Pastedtext'), + chipLines, + } +} + +/** + * Turn the evidence into a verdict. + * + * The echoed case is the only one the screen can actually confirm: the front of + * the text is visible, so it is there. A collapsed chip is `landed` but never + * `confirmed`, and that gap is deliberate — it is what stops `send` printing a + * ✔ over a delivery it did not check, without inventing a failure it cannot + * substantiate either. + */ +export function judgePaste(e: PasteEvidence, expectedLines: number): PasteVerdict { + if (e.sentinelSeen) { + return { landed: true, confirmed: true, detail: 'the text is visible in the tab' } + } + + if (e.chipSeen) { + const chip = e.chipLines !== undefined ? ` (its chip reads +${e.chipLines} lines` : ' (chip present' + return { + landed: true, + confirmed: false, + detail: + `the paste collapsed to a chip${chip}, which does not track the ~${expectedLines}-line payload) — ` + + `the screen cannot show how much of it arrived, so completeness is unverified. ` + + `Use --verify to check what the session actually received, or --path to avoid the prompt line entirely`, + } + } + + return { landed: false, confirmed: false, detail: 'nothing from the text appeared in the tab' } +} + +/** What the receiving session recorded, compared against what was sent. */ +export interface DeliveryVerdict { + delivered: boolean + detail: string +} + +/** + * Compare a payload against the text the target session actually received. + * + * This is the only trustworthy answer to "did all of it arrive?", and it is + * available because Claude writes the user message it received to its + * transcript. Both sides are whitespace-stripped before comparing: `send` + * converts newlines to CR on the way out, the transcript stores LF, and the + * terminal is free to re-wrap in between — none of which is a delivery failure. + * + * The front and the tail are checked separately and named separately, because + * which end is missing is the diagnostic: a missing front is the observed + * clipping mode, while a missing tail would be something else entirely. + */ +export function judgeDelivery(payload: string, received: string | null): DeliveryVerdict { + if (received === null) { + return { + delivered: false, + detail: 'the target session has not recorded any message matching this send — it may not have been submitted, or the session may not have started its turn yet', + } + } + + const strip = (s: string) => s.replace(/\s+/g, '') + const want = strip(payload) + const got = strip(received) + const FINGERPRINT = 40 + + const front = want.slice(0, FINGERPRINT) + const tail = want.slice(-FINGERPRINT) + const frontOk = got.includes(front) + const tailOk = got.includes(tail) + + if (frontOk && tailOk) { + return { + delivered: true, + detail: `the session received the whole payload (front and tail both present in its transcript, ${want.length} non-whitespace chars sent)`, + } + } + if (!frontOk && tailOk) { + return { + delivered: false, + detail: 'the session received the END of the payload but not its FRONT — this is the front-clipping failure mode; re-send with --path', + } + } + if (frontOk && !tailOk) { + return { + delivered: false, + detail: 'the session received the FRONT of the payload but not its end — it was cut short; re-send with --path', + } + } + return { + delivered: false, + detail: 'the message the session recorded matches neither end of what was sent', + } +} diff --git a/src/core/session-lookup.test.ts b/src/core/session-lookup.test.ts new file mode 100644 index 0000000..3a65175 --- /dev/null +++ b/src/core/session-lookup.test.ts @@ -0,0 +1,59 @@ +import { describe, it, expect } from 'bun:test' +import { classifySessionLookup, explainSessionLookup } from './session-lookup.js' + +describe('classifySessionLookup', () => { + const never = () => { throw new Error('should not count') } + + it('reports found without paying for a count', () => { + expect(classifySessionLookup({ cwd: '/repo', found: true, countInDir: never })) + .toEqual({ status: 'found' }) + }) + + it('reports no-cwd when there was nothing to look up', () => { + expect(classifySessionLookup({ cwd: '', found: false, countInDir: never })) + .toEqual({ status: 'no-cwd' }) + }) + + // The distinction the whole type exists for: a lookup that threw is UNKNOWN, + // not absent, and a caller that treats it as absent will spawn over a live + // session or write one off as dead. + it('reports lookup-failed, with the error, in preference to not-found', () => { + const r = classifySessionLookup({ + cwd: '/repo', found: false, error: new Error('EACCES'), countInDir: never, + }) + expect(r.status).toBe('lookup-failed') + expect(r.detail).toBe('EACCES') + }) + + it('an error with no cwd is still lookup-failed, not no-cwd', () => { + expect(classifySessionLookup({ cwd: '', found: false, error: new Error('boom'), countInDir: never }).status) + .toBe('lookup-failed') + }) + + it('reports not-found with the directory transcript count', () => { + expect(classifySessionLookup({ cwd: '/repo', found: false, countInDir: () => 4 })) + .toEqual({ status: 'not-found', sessionsInDir: 4 }) + }) +}) + +describe('explainSessionLookup', () => { + it('says nothing for a resolved session', () => { + expect(explainSessionLookup({ status: 'found' })).toBeNull() + }) + + // "renamed" vs "never ran here" are the two readings of not-found, and they + // call for opposite responses. + it('distinguishes a renamed tab from a directory with no history', () => { + expect(explainSessionLookup({ status: 'not-found', sessionsInDir: 3 })).toContain('renamed') + expect(explainSessionLookup({ status: 'not-found', sessionsInDir: 0 })).toContain('has ever run') + }) + + it('marks a failed lookup as unknown rather than absent', () => { + expect(explainSessionLookup({ status: 'lookup-failed', detail: 'EACCES' })) + .toContain('unknown, not absent') + }) + + it('explains a missing cwd', () => { + expect(explainSessionLookup({ status: 'no-cwd' })).toContain('no working directory') + }) +}) diff --git a/src/core/session-lookup.ts b/src/core/session-lookup.ts new file mode 100644 index 0000000..707ff95 --- /dev/null +++ b/src/core/session-lookup.ts @@ -0,0 +1,67 @@ +/** + * Why a tab's `session_id` is null — because "null" alone has been read as + * three different things and acted on wrongly. + * + * A caller seeing `session_id: null` cannot tell whether cctabs looked and + * found nothing (the tab really has no session), looked and *couldn't* (the + * project dir was unreadable), or never looked at all (the tab reported no + * cwd). Those want opposite responses: the first is a tab to spawn fresh, the + * second and third are bugs to fix before touching the fleet. Observed on a + * real fleet as a tab whose session was perfectly readable but whose transcript + * was filed under a title that no longer matched the tab. + */ +export type SessionLookup = + /** Resolved — `session_id` is set. */ + | 'found' + /** The tab reported no working directory, so there was nothing to look up. */ + | 'no-cwd' + /** Searched every config dir; no session is titled after this tab. */ + | 'not-found' + /** The search itself threw — a permissions or filesystem problem, not an answer. */ + | 'lookup-failed' + +export interface SessionLookupResult { + status: SessionLookup + /** + * How many transcripts exist for this tab's directory regardless of title, + * for the `not-found` case. `> 0` means the directory has history and it is + * the *name* that stopped matching — a renamed tab, not a dead one. Only set + * when it's the distinction that matters. + */ + sessionsInDir?: number + /** The error text, for `lookup-failed`. */ + detail?: string +} + +/** + * Classify one tab's session lookup. Pure, so the four outcomes are testable + * without a terminal or a home directory. + */ +export function classifySessionLookup(opts: { + cwd: string + found: boolean + error?: Error + /** Injected so the count is only paid for when it's the deciding factor. */ + countInDir: () => number +}): SessionLookupResult { + if (opts.found) return { status: 'found' } + if (opts.error) return { status: 'lookup-failed', detail: opts.error.message } + if (!opts.cwd) return { status: 'no-cwd' } + return { status: 'not-found', sessionsInDir: opts.countInDir() } +} + +/** One-line explanation of a lookup that produced no id, for the human view. */ +export function explainSessionLookup(r: SessionLookupResult): string | null { + switch (r.status) { + case 'found': + return null + case 'no-cwd': + return 'no session id: the tab reports no working directory, so none could be looked up' + case 'lookup-failed': + return `no session id: the lookup failed (${r.detail}) — this is unknown, not absent` + case 'not-found': + return r.sessionsInDir + ? `no session id: nothing here is titled after this tab, but ${r.sessionsInDir} transcript(s) exist for its directory — it was probably renamed after Claude started` + : 'no session id: no Claude session has ever run in this directory' + } +} diff --git a/src/core/session-status.test.ts b/src/core/session-status.test.ts index eb7c7c9..8591440 100644 --- a/src/core/session-status.test.ts +++ b/src/core/session-status.test.ts @@ -3,6 +3,7 @@ import { autoModeDialogVisible, classifyTerminalBuffer, parsePermissionMode, + promptIsReady, toLaunchableMode, trustDialogVisible, } from './session-status.js' @@ -228,3 +229,39 @@ describe('autoModeDialogVisible', () => { expect(autoModeDialogVisible('')).toBe(false) }) }) + +describe('promptIsReady', () => { + it('accepts a bare shell prompt', () => { + expect(promptIsReady('~/Dev/cctabs %')).toBe(true) + expect(promptIsReady('user@host:~$')).toBe(true) + }) + + it("accepts Claude's input line with its placeholder", () => { + expect(promptIsReady('❯ Try "fix the failing test"')).toBe(true) + }) + + it("accepts Claude's input footer", () => { + expect(promptIsReady(' ⏵⏵ auto mode on shift+tab to cycle')).toBe(true) + }) + + // The measured false negative: --wait-for-prompt timed out at 20s against + // tabs whose prompts were ready, because Claude renders this notice BELOW + // the input line and only the last line was being tested. + it('sees a ready prompt underneath a "Restart to update" banner', () => { + const buffer = [ + '❯ Try "fix the failing test"', + '', + ' Restart to update to v2.1.4', + ].join('\n') + expect(promptIsReady(buffer)).toBe(true) + }) + + it('is false for an empty buffer', () => { + expect(promptIsReady('')).toBe(false) + expect(promptIsReady(' \n ')).toBe(false) + }) + + it('is false for a tab showing no prompt at all', () => { + expect(promptIsReady('✽ Dilly-dallying… (14m 5s · ↓34.9k tokens)')).toBe(false) + }) +}) diff --git a/src/core/session-status.ts b/src/core/session-status.ts index 69779b3..5d6210c 100644 --- a/src/core/session-status.ts +++ b/src/core/session-status.ts @@ -145,6 +145,33 @@ export function autoModeDialogVisible(buffer: string): boolean { return /Setupautomodeforyourenvironment/i.test(c) || (/1\.Setitup/i.test(c) && /2\.Notnow/i.test(c)) } +/** + * Is a tab's input ready to receive text? + * + * Read across the whole window rather than off the last line alone, and that is + * the fix rather than a preference: Claude Code renders notices BELOW its input + * line — `Restart to update` is the one that caught this — so the last non-empty + * line is the banner and the ready prompt is a line or two above it. + * `--wait-for-prompt` consequently timed out at 20s against tabs whose prompts + * were sitting there ready, which is worse than no check, because the caller + * concludes the tab is broken and stops. + * + * Accepts any of the ready shapes: a bare shell prompt at end-of-line, Claude's + * `❯` input line (usually followed by a `Try "…"` placeholder, so the glyph is + * NOT at end-of-line), or Claude's input footer. + */ +export function promptIsReady(buffer: string): boolean { + if (!buffer.trim()) return false + const lines = buffer.split('\n').map((l) => l.trim()).filter(Boolean) + + for (const line of lines) { + if (/[$%>]\s*$/.test(line)) return true + if (/^❯/.test(line)) return true + } + + return /automode|foragents/i.test(stripWhitespace(buffer)) +} + /** * Classify a tab from the text of its captured output. * diff --git a/src/core/session.ts b/src/core/session.ts index 7434704..6b8f200 100644 --- a/src/core/session.ts +++ b/src/core/session.ts @@ -209,10 +209,25 @@ interface TitleEntry { id: string; cwd: string; mtime: number } * scan each directory a single time instead of once per tab. * * Cleared implicitly per process — cctabs commands are one-shot, so a stale - * cache is never a concern within a single invocation. + * cache is never a concern within a single invocation, with one exception: + * see {@link resetTitleIndexCache}. */ const titleIndexCache = new Map>() +/** + * Drop the title index, so the next lookup reads what is on disk NOW. + * + * `restore` is the one command that both reads sessions and *causes* new ones + * to be written inside a single invocation. Verifying a spawn against the cache + * built while planning would compare the new world against a snapshot of the + * old one and always agree with itself — which is precisely the "success line + * that cannot fail" being removed. Also used by tests, which reconfigure the + * config dirs between cases. + */ +export function resetTitleIndexCache(): void { + titleIndexCache.clear() +} + function buildTitleIndex(projectDir: string): Map { const cached = titleIndexCache.get(projectDir) if (cached) return cached diff --git a/src/core/tab-target.ts b/src/core/tab-target.ts new file mode 100644 index 0000000..0d1c9f8 --- /dev/null +++ b/src/core/tab-target.ts @@ -0,0 +1,87 @@ +import type { TerminalAdapter } from './adapter.js' +import type { Block } from '../types/index.js' + +/** + * A resolved send/read target: the terminal block to talk to, plus whatever we + * know about the tab it belongs to. + * + * The tab fields are what separates this from a bare block id. `transcript` + * needs the tab's NAME and CWD to find the Claude session running in it — a + * block id says nothing about which conversation is inside — and those two are + * exactly what `resolveTabSession` takes. They're absent when the query only + * matched a block, which is the honest answer: a block that isn't in a tab we + * can name has no session we can resolve. + */ +export interface TabTarget { + blockId: string + tabId?: string + name?: string + cwd?: string +} + +export type TabTargetResult = + | { ok: true; target: TabTarget } + /** `lines` are the candidate list printed under `message` for an ambiguous query. */ + | { ok: false; message: string; lines?: string[] } + +/** + * Resolve a user-typed target — tab name, tab id prefix, or block id prefix — + * the one way, for every command that takes one. + * + * `send`, `scrollback` and `transcript` each carried a verbatim copy of this + * (tab first, block as fallback, same two ambiguity messages), which is three + * places for the resolution order to drift. The order matters: tab names are + * what a human types and what `cctabs sessions` prints, so a name that is also + * a block-id prefix must still resolve to the tab. + */ +export function resolveTabTarget( + adapter: TerminalAdapter, + query: string, + tabsById: Map, + tabNames: Map, +): TabTargetResult { + const tabMatches = adapter.resolveTab(query, tabsById, tabNames) + + if (tabMatches.length > 1) { + return { + ok: false, + message: `Multiple tabs match '${query}':`, + lines: tabMatches.map((tid) => ` "${tabNames.get(tid)}" [${tid.slice(0, 8)}]`), + } + } + + if (tabMatches.length === 1) { + const tabId = tabMatches[0] + const blocks = (tabsById.get(tabId) ?? []).filter((b) => b.view === 'term') + if (!blocks.length) { + return { ok: false, message: `Tab "${tabNames.get(tabId)}" has no terminal block` } + } + return { + ok: true, + target: { + blockId: blocks[0].blockid, + tabId, + name: tabNames.get(tabId) ?? tabId.slice(0, 8), + cwd: blocks[0].meta?.['cmd:cwd'], + }, + } + } + + const blockMatches = adapter.resolveBlock(query, adapter.blocksList()) + if (!blockMatches.length) { + return { + ok: false, + message: `No tab or block matching '${query}' (tabs in workspaces with no open window are not visible — open that workspace first)`, + } + } + if (blockMatches.length > 1) { + return { + ok: false, + message: `Multiple blocks match '${query}':`, + lines: blockMatches.map((b) => ` ${b.blockid}`), + } + } + + const b = blockMatches[0] + return { ok: true, target: { blockId: b.blockid, tabId: b.tabid, cwd: b.meta?.['cmd:cwd'] } } +} diff --git a/src/core/transcript.test.ts b/src/core/transcript.test.ts new file mode 100644 index 0000000..cb9f632 --- /dev/null +++ b/src/core/transcript.test.ts @@ -0,0 +1,185 @@ +import { describe, it, expect, beforeEach, afterEach } from 'bun:test' +import { mkdtempSync, mkdirSync, writeFileSync, rmSync, utimesSync } from 'fs' +import { tmpdir } from 'os' +import { join } from 'path' +import { + countSessionsInDir, + locateTranscriptFile, + parseAssistantTurns, + readAssistantTurns, + truncateTurn, +} from './transcript.js' +import { pathToProjectSlug } from './session.js' +import type { ClaudeConfigDir } from './config-dirs.js' + +let tmp: string + +/** A config dir as listClaudeConfigDirs would report it. */ +function cfg(root: string, backend?: string): ClaudeConfigDir { + return { root, projectsRoot: join(root, 'projects'), backend } +} + +function writeTranscript( + projectsRoot: string, + dirForSlug: string, + opts: { id: string; lines: string[]; mtimeSec?: number }, +): string { + const projectDir = join(projectsRoot, pathToProjectSlug(dirForSlug)) + mkdirSync(projectDir, { recursive: true }) + const file = join(projectDir, `${opts.id}.jsonl`) + writeFileSync(file, opts.lines.join('\n') + '\n') + if (opts.mtimeSec) utimesSync(file, opts.mtimeSec, opts.mtimeSec) + return file +} + +const assistantLine = (text: string, timestamp?: string) => + JSON.stringify({ + type: 'assistant', + ...(timestamp ? { timestamp } : {}), + message: { role: 'assistant', content: [{ type: 'text', text }] }, + }) + +beforeEach(() => { tmp = mkdtempSync(join(tmpdir(), 'cctabs-transcript-')) }) +afterEach(() => { rmSync(tmp, { recursive: true, force: true }) }) + +describe('parseAssistantTurns', () => { + it('returns assistant text messages oldest first', () => { + const jsonl = [ + JSON.stringify({ type: 'user', message: { role: 'user', content: 'go' } }), + assistantLine('first'), + assistantLine('second'), + ].join('\n') + expect(parseAssistantTurns(jsonl).map((t) => t.text)).toEqual(['first', 'second']) + }) + + it('joins multiple text blocks in one message', () => { + const jsonl = JSON.stringify({ + message: { role: 'assistant', content: [{ type: 'text', text: 'a' }, { type: 'text', text: 'b' }] }, + }) + expect(parseAssistantTurns(jsonl)[0].text).toBe('a\nb') + }) + + it('accepts a plain string content', () => { + const jsonl = JSON.stringify({ message: { role: 'assistant', content: 'plain' } }) + expect(parseAssistantTurns(jsonl).map((t) => t.text)).toEqual(['plain']) + }) + + // A turn that only called tools carries no prose. Emitting it as an empty + // message would pad the "last 3 messages" window with nothing. + it('skips tool-only and thinking-only assistant messages', () => { + const jsonl = [ + JSON.stringify({ message: { role: 'assistant', content: [{ type: 'tool_use', name: 'Bash', input: {} }] } }), + JSON.stringify({ message: { role: 'assistant', content: [{ type: 'thinking', thinking: 'hmm' }] } }), + assistantLine('the answer'), + ].join('\n') + expect(parseAssistantTurns(jsonl).map((t) => t.text)).toEqual(['the answer']) + }) + + // A live session is being appended to as we read it, so the last line can be + // half-written. That must not cost us the other messages. + it('skips a truncated trailing line rather than failing', () => { + const jsonl = [assistantLine('complete'), '{"message":{"role":"assist'].join('\n') + expect(parseAssistantTurns(jsonl).map((t) => t.text)).toEqual(['complete']) + }) + + it('carries the entry timestamp when present', () => { + expect(parseAssistantTurns(assistantLine('x', '2026-09-08T10:00:00Z'))[0].timestamp) + .toBe('2026-09-08T10:00:00Z') + }) + + it('is empty for a transcript with no assistant messages', () => { + expect(parseAssistantTurns(JSON.stringify({ type: 'user', message: { role: 'user', content: 'hi' } }))).toEqual([]) + }) +}) + +describe('locateTranscriptFile', () => { + // The requirement that makes this command usable on a real fleet: a tab + // running under a backend preset writes beneath that preset's config dir, and + // searching only ~/.claude reports "no transcript" for a healthy session. + it('finds a transcript under a NON-default config dir and reports its backend', () => { + const enterprise = cfg(join(tmp, '.claude-enterprise'), 'enterprise') + const file = writeTranscript(enterprise.projectsRoot, '/work/repo', { id: 'sid-1', lines: [assistantLine('hi')] }) + + const found = locateTranscriptFile('sid-1', [cfg(join(tmp, '.claude')), enterprise]) + expect(found?.file).toBe(file) + expect(found?.backend).toBe('enterprise') + }) + + it('finds a transcript in the default dir, with no backend attached', () => { + const dflt = cfg(join(tmp, '.claude')) + writeTranscript(dflt.projectsRoot, '/work/repo', { id: 'sid-2', lines: [assistantLine('hi')] }) + + const found = locateTranscriptFile('sid-2', [dflt]) + expect(found?.backend).toBeUndefined() + }) + + it('searches every project slug, not just one', () => { + const dflt = cfg(join(tmp, '.claude')) + writeTranscript(dflt.projectsRoot, '/a/one', { id: 'other', lines: [] }) + const wanted = writeTranscript(dflt.projectsRoot, '/b/two', { id: 'sid-3', lines: [assistantLine('hi')] }) + + expect(locateTranscriptFile('sid-3', [dflt])?.file).toBe(wanted) + }) + + it('prefers the newest when the same id exists under two roots', () => { + const a = cfg(join(tmp, '.claude')) + const b = cfg(join(tmp, '.claude-other'), 'other') + writeTranscript(a.projectsRoot, '/work/repo', { id: 'dup', lines: [assistantLine('old')], mtimeSec: 1000 }) + const newer = writeTranscript(b.projectsRoot, '/work/repo', { id: 'dup', lines: [assistantLine('new')], mtimeSec: 9000 }) + + expect(locateTranscriptFile('dup', [a, b])?.file).toBe(newer) + }) + + it('returns null for an unknown id, and for an empty one', () => { + const dflt = cfg(join(tmp, '.claude')) + mkdirSync(dflt.projectsRoot, { recursive: true }) + expect(locateTranscriptFile('nope', [dflt])).toBeNull() + expect(locateTranscriptFile('', [dflt])).toBeNull() + }) + + it('tolerates a config dir that does not exist on disk', () => { + const missing = cfg(join(tmp, 'absent')) + expect(locateTranscriptFile('any', [missing])).toBeNull() + }) +}) + +describe('readAssistantTurns', () => { + it('reads a file from disk', () => { + const dflt = cfg(join(tmp, '.claude')) + const file = writeTranscript(dflt.projectsRoot, '/work/repo', { + id: 'sid', lines: [assistantLine('one'), assistantLine('two')], + }) + expect(readAssistantTurns(file).map((t) => t.text)).toEqual(['one', 'two']) + }) +}) + +describe('countSessionsInDir', () => { + // This is what separates "the tab is dead" from "the tab was renamed after + // Claude started" — the two failures a bare null conflates. + it('counts .jsonl files for a directory across config dirs', () => { + const a = cfg(join(tmp, '.claude')) + const b = cfg(join(tmp, '.claude-other'), 'other') + writeTranscript(a.projectsRoot, '/work/repo', { id: 's1', lines: [] }) + writeTranscript(a.projectsRoot, '/work/repo', { id: 's2', lines: [] }) + writeTranscript(b.projectsRoot, '/work/repo', { id: 's3', lines: [] }) + + expect(countSessionsInDir('/work/repo', [a, b])).toBe(3) + }) + + it('is 0 when nothing has ever run there', () => { + expect(countSessionsInDir('/never/used', [cfg(join(tmp, '.claude'))])).toBe(0) + }) +}) + +describe('truncateTurn', () => { + it('leaves text alone with no limit, or under the limit', () => { + expect(truncateTurn('short')).toBe('short') + expect(truncateTurn('short', 100)).toBe('short') + }) + + it('says how much it cut rather than trimming silently', () => { + const out = truncateTurn('abcdefghij', 4) + expect(out.startsWith('abcd')).toBe(true) + expect(out).toContain('6 more characters') + }) +}) diff --git a/src/core/transcript.ts b/src/core/transcript.ts new file mode 100644 index 0000000..9202029 --- /dev/null +++ b/src/core/transcript.ts @@ -0,0 +1,196 @@ +import { existsSync, readdirSync, readFileSync, statSync } from 'fs' +import { extname, join } from 'path' +import { originOf, scopeToDirs, type ConfigDirScope, type SessionOrigin } from './config-dirs.js' +import { pathToProjectSlug } from './session.js' + +/** A session's transcript file, and which Claude account it was found under. */ +export interface LocatedTranscript extends SessionOrigin { + file: string + mtime: number +} + +/** + * Find a session's transcript by id, across EVERY Claude config dir on the + * machine and every project slug inside them. + * + * The breadth is the whole point, and it is not defensive programming: a tab + * running under a backend preset writes its transcript beneath that preset's + * `CLAUDE_CONFIG_DIR` (say `~/.claude-enterprise/projects`), invisible to + * anything that only looks in `~/.claude/projects`. A search that checks one + * root reports "no transcript" for a perfectly healthy session — which a driver + * reads as "that tab is dead" and acts on. So: all roots, and the answer + * carries the origin so callers can say which account it came from. + * + * Unlike `findSessionFileById`, this needs no candidate directories — the id + * alone is enough, which is what a tab-to-session lookup has to work from. + * Newest wins if the same id somehow exists under two roots. + */ +export function locateTranscriptFile( + sessionId: string, + scope?: ConfigDirScope, +): LocatedTranscript | null { + if (!sessionId) return null + let best: LocatedTranscript | null = null + + for (const cfg of scopeToDirs(scope)) { + if (!existsSync(cfg.projectsRoot)) continue + const origin = originOf(cfg) + + for (const slug of readdirSync(cfg.projectsRoot)) { + const file = join(cfg.projectsRoot, slug, `${sessionId}.jsonl`) + let mtime: number + try { + const st = statSync(file) + if (!st.isFile()) continue + mtime = st.mtimeMs + } catch { + continue + } + if (!best || mtime > best.mtime) best = { file, mtime, ...origin } + } + } + + return best +} + +/** One assistant message that carried text, in transcript order. */ +export interface AssistantTurn { + text: string + /** The entry's own timestamp, when it recorded one. */ + timestamp?: string +} + +/** + * The assistant's text messages in a transcript, oldest first. + * + * Text messages, not "turns" in the conversational sense: a single turn is + * written as many entries when it calls tools, and most of those carry only + * `tool_use` blocks. Those are dropped — they are the mechanics of the work, + * not what the session concluded, and a caller asking "what has this tab + * learned?" wants the prose. Thinking blocks are dropped for the same reason + * plus one more: they are not the session's stated answer. + * + * Malformed lines are skipped rather than fatal. A transcript being appended to + * right now can have a half-written last line, and refusing to read the other + * 4,000 entries because of it would make this useless on exactly the live + * sessions it exists for. + */ +export function parseAssistantTurns(jsonl: string): AssistantTurn[] { + const turns: AssistantTurn[] = [] + + for (const line of jsonl.split('\n')) { + if (!line.trim()) continue + let entry: { + message?: { role?: string; content?: unknown } + timestamp?: string + } + try { + entry = JSON.parse(line) + } catch { + continue + } + if (entry.message?.role !== 'assistant') continue + + const content = entry.message.content + const text = Array.isArray(content) + ? content + .filter( + (c): c is { type: string; text: string } => + !!c && typeof c === 'object' && (c as { type?: string }).type === 'text' && + typeof (c as { text?: unknown }).text === 'string', + ) + .map((c) => c.text) + .join('\n') + : typeof content === 'string' + ? content + : '' + + if (!text.trim()) continue + turns.push({ text: text.trim(), timestamp: entry.timestamp }) + } + + return turns +} + +/** {@link parseAssistantTurns} for a file on disk. */ +export function readAssistantTurns(file: string): AssistantTurn[] { + return parseAssistantTurns(readFileSync(file, 'utf-8')) +} + +/** + * How many sessions exist in a directory's project folder, across all config + * dirs, regardless of what they are titled. + * + * Used to tell two very different failures apart when a tab's session can't be + * resolved by name: nothing has ever run here (count 0 — the tab really has no + * session), versus sessions exist but none is titled after this tab (count > 0 + * — a title drifted, and the transcript is right there). Reported by `sessions` + * and `transcript` instead of a bare null, because acting on the wrong one of + * those two means either restoring a session that isn't there or abandoning one + * that is. + */ +export function countSessionsInDir(dir: string, scope?: ConfigDirScope): number { + let count = 0 + for (const cfg of scopeToDirs(scope)) { + const projectDir = join(cfg.projectsRoot, pathToProjectSlug(dir)) + if (!existsSync(projectDir)) continue + try { + count += readdirSync(projectDir).filter((f) => extname(f) === '.jsonl').length + } catch { + // unreadable project dir contributes nothing + } + } + return count +} + +/** Trim a turn's text for display, marking that it was cut. */ +export function truncateTurn(text: string, maxChars?: number): string { + if (!maxChars || text.length <= maxChars) return text + return `${text.slice(0, maxChars)}\n… [${text.length - maxChars} more characters — raise --chars or read the transcript directly]` +} + +/** + * The text of the most recent user message in a transcript. + * + * Used to verify a delivery against ground truth: this is what the receiving + * session actually got, as opposed to what the terminal appeared to show. Note + * that Claude appends its own context (system reminders and the like) to the + * recorded message, so this is a superset of the payload — which is why callers + * check that the payload's ends are *contained* in it rather than comparing + * lengths. + * + * Returns null when the transcript holds no user message with text, which is + * the honest answer for a turn that hasn't started yet. + */ +export function readLastUserMessage(file: string): string | null { + let latest: string | null = null + + for (const line of readFileSync(file, 'utf-8').split('\n')) { + if (!line.trim()) continue + let entry: { message?: { role?: string; content?: unknown } } + try { + entry = JSON.parse(line) + } catch { + continue + } + if (entry.message?.role !== 'user') continue + + const content = entry.message.content + const text = Array.isArray(content) + ? content + .filter( + (c): c is { type: string; text: string } => + !!c && typeof c === 'object' && (c as { type?: string }).type === 'text' && + typeof (c as { text?: unknown }).text === 'string', + ) + .map((c) => c.text) + .join('\n') + : typeof content === 'string' + ? content + : '' + + if (text.trim()) latest = text + } + + return latest +} From 8d42199aa70bc3db80364f6bbaf6b8517fd5a2f4 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Fredrik=20Wollse=CC=81n?= Date: Tue, 8 Sep 2026 22:27:31 +0300 Subject: [PATCH 2/4] =?UTF-8?q?fix(send):=20a=20message=20that=20quoted=20?= =?UTF-8?q?a=20flag=20name=20delivered=20nothing=20and=20reported=20?= =?UTF-8?q?=E2=9C=94?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two bugs reported against the new send, both measured. The serious one first. A send whose text contained `--` delivered NOTHING and printed success. The option parser silently drops any argv element containing a double dash — reproduced on "mentions --verify here", "--leading", and even "a--b"; a single dash survives. The text was absent from both the positionals and the values, so `send` fell through to stdin, stdin was empty, and it printed ✔ Sent to 5f3e853e: ⏎ with an empty preview, which was the only tell. A ~900-byte bug report was lost this way. The input is not exotic: any message quoting a flag name hits it, and for a tool whose users are agents reporting tool bugs that is the normal case. Positionals now come from process.argv directly (core/send-argv.ts) instead of from the parser that loses them, and `--` works as an explicit terminator for text that is entirely flag-shaped. An empty body is a hard failure now rather than a ✔ — reporting success for a delivery of nothing is the same defect class as the restore success line that could not fail — while a deliberate bare Enter names itself, so an empty preview can never follow a ✔ again. The second report was that `--path` still pastes file contents. It does not: measured on a 2,470-byte file, the payload sent was the 257-char handoff. What was really wrong is `--verify`. Claude records a tool's output as a `role: "user"` message, and a `--path` handoff tells the tab to read the file — so the newest user-role entry becomes the file's contents, and verify compared the handoff against the file and reported "matches neither end of what was sent". That reads exactly like a truncated paste, which is how it was reported. Verify now skips tool results (the entry's `toolUseResult` field identifies them) and searches every real message rather than only the newest. Also fixed while in there: `--verify` resolved the session once, before its poll loop, so a freshly spawned tab failed instantly with "no session resolved" a second before its transcript existed. Lookup retries for the whole window now, dropping the per-process title-index cache each round. Confirmed for the smaller note: a verify mismatch exits 1. Co-Authored-By: Claude Opus 5 (1M context) --- CHANGELOG.md | 7 +++ skills/cctabs/SKILL.md | 22 +++++++ src/commands/send.ts | 106 +++++++++++++++++++++++---------- src/core/paste-confirm.test.ts | 45 +++++++++----- src/core/paste-confirm.ts | 42 ++++++++----- src/core/send-argv.test.ts | 89 +++++++++++++++++++++++++++ src/core/send-argv.ts | 91 ++++++++++++++++++++++++++++ src/core/transcript.test.ts | 69 +++++++++++++++++++++ src/core/transcript.ts | 36 ++++++----- 9 files changed, 433 insertions(+), 74 deletions(-) create mode 100644 src/core/send-argv.test.ts create mode 100644 src/core/send-argv.ts diff --git a/CHANGELOG.md b/CHANGELOG.md index 4e74ec0..490c6d8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,13 @@ page lives at [cctabs.com/changelog](https://cctabs.com/changelog). ## Unreleased +- **Fix: a `send` whose text quoted a flag name delivered NOTHING and reported success.** The option parser silently drops any argv element containing `--` — measured on `"mentions --verify here"`, `"--leading"`, and even `"a--b"`; a single dash survives. The text vanished from the positionals *and* the values, `send` fell through to reading stdin, stdin was empty, and it printed `✔ Sent to 5f3e853e: ⏎` with an empty preview. A ~900-byte bug report was lost this way, and the input is not exotic: any message quoting a flag name hits it, which for a tool whose users are agents reporting tool bugs is the normal case. `send` now recovers its positionals from `process.argv` directly (`core/send-argv.ts`) rather than from the parser that loses them, and `--` is supported as an explicit terminator for text that is entirely flag-shaped: `cctabs send tab -- --verify is broken`. +- **Fix: an empty body is now a hard failure rather than a ✔.** Reporting success for a delivery of nothing is the same defect class as the restore success line that could not fail. An empty `--file`, or no text source at all with empty stdin, exits non-zero and says which it was; the no-source message names the `--` terminator, since a swallowed payload is the likeliest reason to land there. An *explicit* empty is still honoured — `--submit`, or a literal `""` — and now reports itself as `Submitted Enter only (no body)`, so an empty preview after a ✔ can never appear again. That preview was the operator's only tell. +- **Fix: `--verify` compared against the target's NEWEST user message, which the target overwrites with its own work.** Claude records a tool's output as a `role: "user"` message, so a `--path` handoff — which tells the tab to read a file — makes the newest user-role entry the file's contents. Verify then compared the handoff against the file and reported that the payload "matches neither end of what was sent", which reads exactly like `--path` having pasted the contents. It now skips tool results (identified by the entry's `toolUseResult`) and searches every message rather than only the last, so a payload stays findable after the session has moved on. Both ends arriving in *different* messages is still not a delivery. +- **Fix: `--verify` failed instantly against a freshly spawned tab.** It resolved the session once, before its poll loop, so a tab whose transcript was a second from existing reported "no session resolved for tab" and failed. Session lookup now retries for the whole window, dropping the per-process title-index cache each round — the session being waited for is precisely the one that appears after that cache was built. +- `send --path`'s help text no longer implies the receiving session's transcript stays clean: the session obeys the handoff by *reading* the file, so the contents land there as a tool result. That is the handoff working, and mistaking it for a paste is what the report above turned on. + + - **`cctabs transcript [n]` reads what a tab has SAID, not what it is painting.** `scrollback` returns the last painted frame, so a tab mid-turn shows a spinner and nothing else — which meant sessions briefed each other from stale pictures, and one such brief told a tab to go measure two things it had already measured, missing a third it had found. Prints the last n assistant messages from the transcript (`--json` for a driver; `findings` as an alias). It searches **every** Claude config dir, because a tab on a backend preset writes under that preset's own root and looking in only `~/.claude/projects` reports "no transcript" for a healthy session — which reads as dead. Tool-only and thinking-only messages are skipped. It exits non-zero and names the failure — no session titled after this tab (with a count of transcripts that do exist for its directory, which distinguishes a renamed tab from a dead one), no transcript on disk, or a lookup that threw — because "I couldn't read it" and "it hasn't said anything" are different answers. - **`cctabs sort --first a,b,c` pins a chosen set to the front of the bar.** Activity order is close to the opposite of what a driver wants: a tab that just delivered sinks. The Tabby plugin's `POST /api/tabs/reorder` already did exactly this — unlisted tabs keep their relative order and sort after the listed ones — with no CLI verb exposing it. Naming tabs skips the activity scan entirely (~7.7s of transcript reading on a 65-tab fleet). If any name doesn't resolve to exactly one tab, nothing moves and it exits non-zero: half a working set in reach, with no indication which half, is worse than an error. - **`cctabs send` no longer reports success it hasn't established.** A large paste collapses to a `[Pasted text #N +M lines]` chip, and the chip renders no matter how little arrived — so "a chip is on screen" was accepted as proof of delivery, which is how a 6,835-byte brief that landed as its last 756 bytes came to be reported as sent. `send` now distinguishes three claims that used to be one ✔ line: *nothing arrived* (a hard failure — and the body is **not submitted**, because pressing Enter on a fragment sends something that reads as a complete message; the text is left in the input box and the command exits non-zero), *something arrived but completeness is unverified* (a warning naming the two ways to establish it), and *verified*. diff --git a/skills/cctabs/SKILL.md b/skills/cctabs/SKILL.md index f2af9c8..6809b37 100644 --- a/skills/cctabs/SKILL.md +++ b/skills/cctabs/SKILL.md @@ -192,6 +192,7 @@ cctabs send [text] # send input — arg, --file, or stdin cctabs send --path # hand the tab a file PATH to read — the safe way to deliver anything large cctabs send --file --verify # check the target's transcript for what it actually RECEIVED cctabs send --submit # press Enter only, submitting a prompt already parked in the box +cctabs send -- # REQUIRED when your text contains `--` (e.g. names a flag) cctabs export [--out path] # bundle a tab + its claude session into a tarball cctabs export --all [-w workspace] # bundle every tab in a workspace cctabs import [--dry-run] [-f] # restore tabs + sessions from a tarball @@ -685,6 +686,27 @@ echo "do the thing" | cctabs send auth # pipe via stdin line, so a `Restart to update` banner rendered *below* a ready prompt no longer makes it time out. +⛔ **Text containing `--` needs the `--` terminator.** The option parser drops +any argv element containing a double dash, so a message that *quotes a flag +name* — the normal case when one session reports a tool bug to another — used to +vanish silently while `send` printed a ✔. `send` now recovers its positionals +from the raw command line, so this works either way, but the terminator is the +unambiguous form and the only one for text that is *entirely* flag-shaped: + +```bash +cctabs send auth -- --verify is broken and --path too +``` + +An empty body is now a **hard error**, not a ✔ — so a swallowed payload fails +loudly instead of pressing Enter and claiming success. A deliberate bare Enter +(`--submit`, or a literal `""`) reports itself as `Submitted Enter only (no +body)`. + +⚠️ **`--path` makes the receiving session READ the file, so the file's contents +appear in its transcript as a tool result.** That is the handoff working — not a +paste. (`--verify` knows the difference: it skips tool results and searches all +of the session's real messages, not just the newest.) + ## Workflow: Remote Control status across the fleet Claude Code's Remote Control (`/rc`, controls a session from claude.ai/code or the mobile app) is a per-process feature — cctabs doesn't manage it directly, but since it manages the tabs *running* those processes, it's the fastest way to audit or repair RC across many sessions at once. diff --git a/src/commands/send.ts b/src/commands/send.ts index e7f9031..6da49e0 100644 --- a/src/commands/send.ts +++ b/src/commands/send.ts @@ -7,8 +7,9 @@ import { sendTextWithConfirmation } from '../core/open-session.js' import { resolveTabTarget } from '../core/tab-target.js' import { classifyTerminalBuffer, promptIsReady } from '../core/session-status.js' import { judgeDelivery } from '../core/paste-confirm.js' -import { resolveTabSession } from '../core/session.js' -import { locateTranscriptFile, readLastUserMessage } from '../core/transcript.js' +import { resetTitleIndexCache, resolveTabSession } from '../core/session.js' +import { locateTranscriptFile, readUserMessages } from '../core/transcript.js' +import { parseSendArgv } from '../core/send-argv.js' function readStdin(): Promise { return new Promise((resolve) => { @@ -37,18 +38,29 @@ export const CLIP_RISK_BYTES = 1024 * from disk by the receiving session. It became the operator's standing * practice for briefs after a 6,835-byte paste arrived as 756 bytes, and it is * first-class here for that reason rather than as a convenience. + * + * One consequence worth knowing, because it reads like a bug: the receiving + * session obeys this by *reading the file*, so the file's contents land in its + * transcript as a tool result. Seeing the contents there means the handoff + * worked — it does not mean they were pasted. */ export function buildPathHandoff(absPath: string): string { return `Read the file at ${absPath} in full, and treat its entire contents as the message intended for you — it is your instructions, not a document to summarise.` } +/** Report a usage failure and exit, before any terminal has been touched. */ +function adapterlessExit(message: string): never { + consola.error(message) + process.exit(1) +} + export const sendCommand = define({ name: 'send', description: 'Send input to a tab or block (text arg, --file, --path, or stdin pipe)', args: { target: { type: 'positional', description: 'Tab name, tab ID prefix, or block ID prefix' }, file: { type: 'string', short: 'f', description: 'Read the text to send from a file' }, - path: { type: 'string', short: 'p', description: 'Hand the tab this file PATH and let the receiving session read it, instead of pasting the contents. The robust way to deliver anything large: nothing but a path crosses the prompt line, so there is no truncation surface.' }, + path: { type: 'string', short: 'p', description: "Hand the tab this file PATH and let the receiving session read it, instead of pasting the contents. The robust way to deliver anything large: only the path crosses the prompt line, so there is no truncation surface. Note the receiving session then READS the file, so its transcript will contain the contents as a tool result — that is the handoff working, not a paste." }, submit: { type: 'boolean', description: 'Send Enter only, submitting whatever is already parked in the tab\'s input box. Explicit alternative to guessing at an empty send.' }, enter: { type: 'boolean', short: 'e', description: 'Append newline after text (default: true)' }, force: { type: 'boolean', description: `Send even when the target has a turn in flight. Without this, payloads over ${CLIP_RISK_BYTES} bytes are refused for a busy tab, because that is when text gets silently clipped.` }, @@ -59,10 +71,14 @@ export const sendCommand = define({ 'wait-timeout': { type: 'number', description: 'Timeout in seconds for --wait-for-prompt (default: 10)' }, }, async run(ctx) { - const query = ctx.positionals[1] - // Inline text is the second positional — undeclared to keep it optional - // (declaring it as a positional makes gunshi require it, breaking --file and stdin) - const inlineText = ctx.positionals[2] + // Positionals come from the RAW command line, not from the option parser. + // The parser drops any argv element containing `--`, which silently ate the + // payload of any message that quoted a flag name and left `send` reporting + // success for a delivery of nothing. See core/send-argv.ts. + const rawArgv = process.argv.slice(2) + const cliArgs = parseSendArgv(rawArgv[0] === 'send' ? rawArgv.slice(1) : rawArgv) + const query = cliArgs.target + const inlineText = cliArgs.text const filePath = ctx.values.file as string | undefined const handoffPath = ctx.values.path as string | undefined const submitOnly = (ctx.values.submit as boolean | undefined) ?? false @@ -116,14 +132,23 @@ export const sendCommand = define({ if (rawText.endsWith('\r')) { rawText = rawText.replace(/\r+$/, ''); sendEnter = true } if (submitOnly) sendEnter = true - // An empty body with no explicit --submit is almost always a mistake — an - // empty file, or a pipe that produced nothing — and it presses Enter in - // someone's session either way. Say what happened instead of doing it - // silently, and point at the flag that means it on purpose. - if (!submitOnly && rawText.length === 0) { - consola.warn( - `Nothing to send${filePath ? ` — ${filePath} is empty` : ''}. ` + - `Pressing Enter only; pass --submit if that is what you meant.`, + // An empty body that nobody asked for is a FAILURE, not a warning. + // + // This is the tell the operator spotted: `✔ Sent to 5f3e853e: ⏎` with an + // empty preview, for a ~900-byte report that arrived nowhere. Pressing + // Enter and reporting success is the worst possible response to "I ended up + // with no payload" — it looks like a delivery. An explicit empty is still + // honoured: `--submit`, or a literal `""` argument, both mean "just submit". + const explicitlyEmpty = submitOnly || inlineText === '' || cliArgs.explicitText + if (!explicitlyEmpty && rawText.length === 0) { + adapterlessExit( + filePath + ? `${filePath} is empty, so there is nothing to send. Pass --submit if you meant to press Enter on a prompt already in the box.` + : handoffPath !== undefined + ? `${handoffPath} produced no message to send.` + : inlineText === undefined + ? `No text to send: no inline argument, no --file, no --path, and stdin was empty. If your text contains \`--\`, put it after a \`--\` terminator: cctabs send ${query ?? ''} -- ` + : 'Nothing to send.', ) } @@ -252,7 +277,10 @@ export const sendCommand = define({ // different claims. They used to be one ✔ line, which is how a partial // delivery came to look like a success. if (rawText.length === 0) { - consola.success(`Sent to ${blockId.slice(0, 8)}: ${label}`) + // Named, not shown as an empty preview. An empty preview after a ✔ was + // the tell that a ~900-byte payload had been eaten by the arg parser, so + // that shape must never appear for anything but a deliberate bare Enter. + consola.success(`Submitted Enter only (no body) to ${blockId.slice(0, 8)}`) } else if (unchecked) { consola.info(`Sent to ${blockId.slice(0, 8)} (unconfirmed — ${detail}): ${label}`) } else if (!confirmed) { @@ -266,9 +294,16 @@ export const sendCommand = define({ /** * Poll the target's transcript until it records the message, then compare. * - * Waits rather than reading once: the message is written when the turn starts, - * which is a beat after Enter. A timeout is reported as a failed verification - * rather than a pass, because "I could not check" must not read as "it arrived". + * Everything here retries until the deadline, including finding the session. + * Resolving it once up front was a real defect: a freshly spawned tab has no + * transcript on disk for a second or two, so `--verify` failed instantly with + * "no session resolved for tab" against a tab that was perfectly fine and about + * to record the message. The title index is dropped each round for the same + * reason — it is cached per process, and the session we are waiting for is + * precisely the one that appears after the cache was built. + * + * A timeout is reported as a failed verification rather than a pass, because + * "I could not check" must not read as "it arrived". */ async function verifyDelivered( tabName: string, @@ -276,28 +311,37 @@ async function verifyDelivered( payload: string, timeoutSec: number, ): Promise<{ delivered: boolean; detail: string }> { - const session = resolveTabSession(tabCwd, tabName) - if (!session) { - return { delivered: false, detail: `no session resolved for tab "${tabName}" in ${tabCwd}, so there is no transcript to check` } - } - const located = locateTranscriptFile(session.id) - if (!located) { - return { delivered: false, detail: `session ${session.id.slice(0, 8)} has no transcript on disk under any Claude config dir` } - } - const deadline = Date.now() + timeoutSec * 1000 let last = judgeDelivery(payload, null) + let sawSession = false + while (Date.now() < deadline) { await new Promise((r) => setTimeout(r, 1000)) - let received: string | null = null + + resetTitleIndexCache() + const session = resolveTabSession(tabCwd, tabName) + if (!session) continue + sawSession = true + + const located = locateTranscriptFile(session.id) + if (!located) continue + + let messages: string[] try { - received = readLastUserMessage(located.file) + messages = readUserMessages(located.file) } catch { // A transcript being appended to mid-read; try again. continue } - last = judgeDelivery(payload, received) + last = judgeDelivery(payload, messages) if (last.delivered) return last } + + if (!sawSession) { + return { + delivered: false, + detail: `no session for tab "${tabName}" in ${tabCwd} appeared within ${timeoutSec}s, so there was no transcript to check`, + } + } return last } diff --git a/src/core/paste-confirm.test.ts b/src/core/paste-confirm.test.ts index ffc3c9a..8f7cf39 100644 --- a/src/core/paste-confirm.test.ts +++ b/src/core/paste-confirm.test.ts @@ -86,43 +86,60 @@ describe('judgeDelivery', () => { it('confirms a payload whose front and tail both reached the session', () => { // Claude appends its own context to the recorded message, so the received // text is a superset — hence containment rather than equality. - const received = `${payload}\nsome appended context` - const v = judgeDelivery(payload, received) - expect(v.delivered).toBe(true) + const received = [`${payload}\nsome appended context`] + expect(judgeDelivery(payload, received).delivered).toBe(true) }) // Whitespace differs on every hop: send converts LF to CR, the transcript // stores LF, and the terminal may re-wrap. None of that is a failure. it('ignores whitespace differences between what was sent and what was stored', () => { - const received = payload.replace(/\n/g, '\r\n ') + expect(judgeDelivery(payload, [payload.replace(/\n/g, '\r\n ')]).delivered).toBe(true) + }) + + // The race that produced a false failure: a `--path` handoff tells the tab to + // read a file, so the NEWEST user-role entry becomes the file it read. The + // payload has to be found among all the messages, not just the last. + it('finds the payload even when later messages have piled on top of it', () => { + const received = [ + 'some earlier instruction', + payload, + 'BUG REPORT (test fixture)\nobservation 001: entirely unrelated file contents', + ] expect(judgeDelivery(payload, received).delivered).toBe(true) }) // The observed clipping mode, and the reason each end is named separately. it('identifies a front-clipped delivery as such', () => { - // The tail arrives, the front does not — the received text must be long - // enough to hold the tail fingerprint, as a real clipped payload is. - const v = judgeDelivery(payload, payload.slice(-60)) + const v = judgeDelivery(payload, [payload.slice(-60)]) expect(v.delivered).toBe(false) expect(v.detail).toContain('END of the payload but not its FRONT') }) it('identifies a delivery cut short at the end', () => { - const v = judgeDelivery(payload, payload.slice(0, 60)) + const v = judgeDelivery(payload, [payload.slice(0, 60)]) expect(v.delivered).toBe(false) expect(v.detail).toContain('FRONT of the payload but not its end') }) - it('reports an unrelated message as matching neither end', () => { - const v = judgeDelivery(payload, 'something else entirely') + // Both ends present but in DIFFERENT messages is not a delivery: the payload + // was never received whole by anything. + it('does not accept the two ends arriving in separate messages', () => { + const v = judgeDelivery(payload, [payload.slice(0, 60), payload.slice(-60)]) + expect(v.delivered).toBe(false) + }) + + it('reports unrelated messages as matching neither end, and says how many it saw', () => { + const v = judgeDelivery(payload, ['something else', 'and another thing']) expect(v.delivered).toBe(false) - expect(v.detail).toContain('neither end') + expect(v.detail).toContain('none of the 2 message(s)') }) // "I could not check" must never read as "it arrived". it('reports nothing recorded as a failed verification, not a pass', () => { - const v = judgeDelivery(payload, null) - expect(v.delivered).toBe(false) - expect(v.detail).toContain('has not recorded any message') + for (const empty of [null, []]) { + const v = judgeDelivery(payload, empty) + expect(v.delivered).toBe(false) + expect(v.detail).toContain('has not recorded any message') + } }) }) diff --git a/src/core/paste-confirm.ts b/src/core/paste-confirm.ts index 5839312..3a8ff04 100644 --- a/src/core/paste-confirm.ts +++ b/src/core/paste-confirm.ts @@ -114,20 +114,26 @@ export interface DeliveryVerdict { } /** - * Compare a payload against the text the target session actually received. + * Compare a payload against the messages the target session actually received. * * This is the only trustworthy answer to "did all of it arrive?", and it is - * available because Claude writes the user message it received to its + * available because Claude writes the user messages it received to its * transcript. Both sides are whitespace-stripped before comparing: `send` * converts newlines to CR on the way out, the transcript stores LF, and the * terminal is free to re-wrap in between — none of which is a delivery failure. * + * ALL the session's messages are searched, not just its newest, and that is a + * fix rather than thoroughness: checking only the newest one raced against the + * session's own work. A `--path` handoff tells the tab to read a file, so by the + * time the check ran the newest user-role entry was the file it had read, and + * the payload was reported as matching neither end of itself. + * * The front and the tail are checked separately and named separately, because * which end is missing is the diagnostic: a missing front is the observed * clipping mode, while a missing tail would be something else entirely. */ -export function judgeDelivery(payload: string, received: string | null): DeliveryVerdict { - if (received === null) { +export function judgeDelivery(payload: string, received: readonly string[] | null): DeliveryVerdict { + if (!received || received.length === 0) { return { delivered: false, detail: 'the target session has not recorded any message matching this send — it may not have been submitted, or the session may not have started its turn yet', @@ -136,27 +142,33 @@ export function judgeDelivery(payload: string, received: string | null): Deliver const strip = (s: string) => s.replace(/\s+/g, '') const want = strip(payload) - const got = strip(received) const FINGERPRINT = 40 - const front = want.slice(0, FINGERPRINT) const tail = want.slice(-FINGERPRINT) - const frontOk = got.includes(front) - const tailOk = got.includes(tail) - if (frontOk && tailOk) { - return { - delivered: true, - detail: `the session received the whole payload (front and tail both present in its transcript, ${want.length} non-whitespace chars sent)`, + let sawFront = false + let sawTail = false + for (const message of received) { + const got = strip(message) + const frontOk = got.includes(front) + const tailOk = got.includes(tail) + if (frontOk && tailOk) { + return { + delivered: true, + detail: `the session received the whole payload (front and tail both present in its transcript, ${want.length} non-whitespace chars sent)`, + } } + sawFront = sawFront || frontOk + sawTail = sawTail || tailOk } - if (!frontOk && tailOk) { + + if (!sawFront && sawTail) { return { delivered: false, detail: 'the session received the END of the payload but not its FRONT — this is the front-clipping failure mode; re-send with --path', } } - if (frontOk && !tailOk) { + if (sawFront && !sawTail) { return { delivered: false, detail: 'the session received the FRONT of the payload but not its end — it was cut short; re-send with --path', @@ -164,6 +176,6 @@ export function judgeDelivery(payload: string, received: string | null): Deliver } return { delivered: false, - detail: 'the message the session recorded matches neither end of what was sent', + detail: `none of the ${received.length} message(s) the session recorded matches either end of what was sent`, } } diff --git a/src/core/send-argv.test.ts b/src/core/send-argv.test.ts new file mode 100644 index 0000000..1f4f57f --- /dev/null +++ b/src/core/send-argv.test.ts @@ -0,0 +1,89 @@ +import { describe, it, expect } from 'bun:test' +import { parseSendArgv } from './send-argv.js' + +const parse = (...tokens: string[]) => parseSendArgv(tokens) + +describe('parseSendArgv', () => { + it('reads the target and the inline text', () => { + expect(parse('mytab', 'hello world')).toMatchObject({ target: 'mytab', text: 'hello world' }) + }) + + it('reads a target with no text', () => { + expect(parse('mytab')).toMatchObject({ target: 'mytab', text: undefined }) + }) + + // THE BUG. The option parser drops any argv element containing `--`, so a + // message that quoted a flag name vanished and `send` reported success for a + // delivery of nothing. These are the exact shapes measured as lost. + it('keeps text that mentions flag names', () => { + expect(parse('mytab', 'BUG 1 — the --path option does not hand over a path').text) + .toBe('BUG 1 — the --path option does not hand over a path') + expect(parse('mytab', 'I ran --file and --verify too').text) + .toBe('I ran --file and --verify too') + }) + + it('keeps text containing a bare double dash anywhere', () => { + expect(parse('mytab', 'a--b').text).toBe('a--b') + expect(parse('mytab', 'trailing --').text).toBe('trailing --') + }) + + it('keeps a tab name containing a double dash', () => { + expect(parse('my--tab', 'hi')).toMatchObject({ target: 'my--tab', text: 'hi' }) + }) + + it('skips flags that come after the text', () => { + expect(parse('mytab', 'hello', '--verify', '--force')) + .toMatchObject({ target: 'mytab', text: 'hello' }) + }) + + it('skips flags that come before the text', () => { + expect(parse('--verify', 'mytab', 'hello')) + .toMatchObject({ target: 'mytab', text: 'hello' }) + }) + + // A value-taking flag's value must not be mistaken for the inline text. + it('does not treat a flag value as the text', () => { + expect(parse('mytab', '--file', '/tmp/brief.txt')).toMatchObject({ target: 'mytab', text: undefined }) + expect(parse('mytab', '-f', '/tmp/brief.txt')).toMatchObject({ target: 'mytab', text: undefined }) + expect(parse('mytab', '--path', '/tmp/brief.txt')).toMatchObject({ target: 'mytab', text: undefined }) + expect(parse('mytab', '--wait-timeout', '30')).toMatchObject({ target: 'mytab', text: undefined }) + }) + + it('handles a flag that carries its own value', () => { + expect(parse('mytab', '--file=/tmp/brief.txt')).toMatchObject({ target: 'mytab', text: undefined }) + }) + + it('treats a boolean flag as consuming nothing', () => { + expect(parse('mytab', '--verify', 'hello')).toMatchObject({ target: 'mytab', text: 'hello' }) + }) + + // A lone `-` is the stdin convention, not a flag. + it('treats a lone dash as a value', () => { + expect(parse('-', 'hello')).toMatchObject({ target: '-', text: 'hello' }) + }) + + describe('the -- terminator', () => { + it('takes everything after it as text, however flag-shaped', () => { + const r = parse('mytab', '--', '--verify', 'is', 'broken') + expect(r).toMatchObject({ target: 'mytab', text: '--verify is broken', explicitText: true }) + }) + + it('lets text override an earlier positional', () => { + const r = parse('mytab', 'ignored', '--', 'the real text') + expect(r).toMatchObject({ target: 'mytab', text: 'the real text' }) + }) + + it('marks an empty terminator as an explicit empty, not a missing one', () => { + const r = parse('mytab', '--') + expect(r).toMatchObject({ target: 'mytab', text: '', explicitText: true }) + }) + + it('leaves the target missing when there is none, rather than inventing one', () => { + expect(parse('--', 'text').target).toBeUndefined() + }) + }) + + it('reports no target for an empty command line', () => { + expect(parse()).toMatchObject({ target: undefined, text: undefined, explicitText: false }) + }) +}) diff --git a/src/core/send-argv.ts b/src/core/send-argv.ts new file mode 100644 index 0000000..aa13ddb --- /dev/null +++ b/src/core/send-argv.ts @@ -0,0 +1,91 @@ +/** + * Recovering `send`'s positional arguments from the raw command line. + * + * This exists because the option parser silently EATS them. Measured: any argv + * element containing the substring `--` is dropped from gunshi's positionals + * and never appears in its values either — `"mentions --verify here"`, + * `"--leading"`, even `"a--b"`. A single dash survives; `--` anywhere does not. + * + * The consequence was a send that reported success and delivered nothing: the + * text vanished, the command fell through to reading stdin, stdin was empty, and + * the ✔ line printed with an empty preview. That the payload happened to be a + * bug report *about* `--file`, `--path` and `--verify` is not an exotic input — + * it is the normal case for a tool whose users are agents reporting tool bugs, + * and any message quoting a flag name hits it. + * + * So the positionals are read from `process.argv` directly. Flags are still + * parsed by gunshi (which handles them correctly); only the positionals, which + * it loses, are recovered here. + */ + +/** `send` options that consume the token after them, so it isn't a positional. */ +export const SEND_VALUE_FLAGS: ReadonlySet = new Set([ + '--file', '-f', + '--path', '-p', + '--wait-timeout', + '--verify-timeout', +]) + +export interface SendArgv { + /** The tab or block to send to. */ + target?: string + /** The inline text, or undefined when none was given on the command line. */ + text?: string + /** + * True when the text came after a literal `--`. + * + * The explicit escape hatch for text that is *entirely* flag-shaped — + * `cctabs send tab -- --verify is broken` — where even a positional-recovering + * walker would otherwise have to guess. + */ + explicitText: boolean +} + +/** + * Split `send`'s tokens into its two positionals, skipping flags. + * + * Only a token that *starts* with `-` is treated as a flag, which is the whole + * point: `"mentions --verify here"` is one argv element that begins with `m`, + * so it is free text, exactly as the shell delivered it. A token after `--` is + * never a flag. + */ +export function parseSendArgv( + tokens: string[], + valueFlags: ReadonlySet = SEND_VALUE_FLAGS, +): SendArgv { + const positionals: string[] = [] + let explicitText = false + + for (let i = 0; i < tokens.length; i++) { + const t = tokens[i] + + // Everything after a bare `--` is text, joined back with the single spaces + // the shell split it on. This is the documented escape for flag-shaped text. + if (t === '--') { + const rest = tokens.slice(i + 1).join(' ') + if (positionals.length === 0) { + // `send -- text` has no target; leave it missing so the caller reports + // the usage error rather than silently sending to nothing. + positionals.push('') + } + positionals[1] = rest + explicitText = true + break + } + + // A lone `-` is a value (the stdin convention), not a flag. + if (t.startsWith('-') && t.length > 1) { + // `--file=x` carries its own value; `--file x` consumes the next token. + if (!t.includes('=') && valueFlags.has(t)) i++ + continue + } + + positionals.push(t) + } + + return { + target: positionals[0] || undefined, + text: positionals[1], + explicitText, + } +} diff --git a/src/core/transcript.test.ts b/src/core/transcript.test.ts index cb9f632..29fe157 100644 --- a/src/core/transcript.test.ts +++ b/src/core/transcript.test.ts @@ -7,6 +7,7 @@ import { locateTranscriptFile, parseAssistantTurns, readAssistantTurns, + readUserMessages, truncateTurn, } from './transcript.js' import { pathToProjectSlug } from './session.js' @@ -183,3 +184,71 @@ describe('truncateTurn', () => { expect(out).toContain('6 more characters') }) }) + +describe('readUserMessages', () => { + const write = (lines: string[]): string => { + const dflt = cfg(join(tmp, '.claude')) + return writeTranscript(dflt.projectsRoot, '/work/repo', { id: 'sid-users', lines }) + } + + it('returns typed user messages oldest first', () => { + const file = write([ + JSON.stringify({ type: 'user', origin: 'cli', message: { role: 'user', content: 'first' } }), + assistantLine('an answer'), + JSON.stringify({ type: 'user', origin: 'cli', message: { role: 'user', content: 'second' } }), + ]) + expect(readUserMessages(file)).toEqual(['first', 'second']) + }) + + // THE FIX. Claude records a tool's OUTPUT as a role:"user" message, so the + // newest user-role entry in a live session is usually a tool result. A + // `--path` handoff makes the tab read a file, and treating that file as "the + // last thing sent to this tab" made the delivery check compare the handoff + // against the file and report that the payload matched neither end of itself. + // The entry shapes here are copied from a real transcript. + it('excludes tool results, which are recorded as user messages', () => { + const file = write([ + JSON.stringify({ + type: 'user', origin: 'cli', promptSource: 'text', + message: { role: 'user', content: 'Read the file at /tmp/brief.txt in full' }, + }), + JSON.stringify({ + type: 'user', + toolUseResult: { type: 'text', file: { filePath: '/tmp/brief.txt' } }, + sourceToolAssistantUUID: 'abc', + message: { role: 'user', content: [{ type: 'tool_result', content: 'THE ENTIRE FILE CONTENTS' }] }, + }), + ]) + expect(readUserMessages(file)).toEqual(['Read the file at /tmp/brief.txt in full']) + }) + + // Belt and braces: even if a tool result arrives as a plain string rather + // than a content block, `toolUseResult` still identifies it. + it('excludes a tool result whose content is a bare string', () => { + const file = write([ + JSON.stringify({ type: 'user', message: { role: 'user', content: 'the real message' } }), + JSON.stringify({ type: 'user', toolUseResult: 'anything', message: { role: 'user', content: 'file contents' } }), + ]) + expect(readUserMessages(file)).toEqual(['the real message']) + }) + + it('joins multiple text blocks and skips empty ones', () => { + const file = write([ + JSON.stringify({ type: 'user', message: { role: 'user', content: [{ type: 'text', text: 'a' }, { type: 'text', text: 'b' }] } }), + JSON.stringify({ type: 'user', message: { role: 'user', content: [{ type: 'image', source: {} }] } }), + ]) + expect(readUserMessages(file)).toEqual(['a\nb']) + }) + + it('skips a truncated trailing line rather than failing', () => { + const file = write([ + JSON.stringify({ type: 'user', message: { role: 'user', content: 'complete' } }), + '{"message":{"role":"us', + ]) + expect(readUserMessages(file)).toEqual(['complete']) + }) + + it('is empty for a transcript with no user messages', () => { + expect(readUserMessages(write([assistantLine('only me')]))).toEqual([]) + }) +}) diff --git a/src/core/transcript.ts b/src/core/transcript.ts index 9202029..e3e7e2a 100644 --- a/src/core/transcript.ts +++ b/src/core/transcript.ts @@ -150,30 +150,38 @@ export function truncateTurn(text: string, maxChars?: number): string { } /** - * The text of the most recent user message in a transcript. + * Every genuine user message in a transcript, oldest first. * - * Used to verify a delivery against ground truth: this is what the receiving - * session actually got, as opposed to what the terminal appeared to show. Note - * that Claude appends its own context (system reminders and the like) to the - * recorded message, so this is a superset of the payload — which is why callers - * check that the payload's ends are *contained* in it rather than comparing - * lengths. + * "Genuine" is doing real work here. Claude records a tool's OUTPUT as a + * `role: "user"` message too, so the newest user-role entry in a live session is + * usually a tool result, not something anyone sent. Measured: handing a tab a + * file with `send --path` produced a 284-byte user message (the handoff) and + * then a 2,469-byte user-role `tool_result` holding the file's contents — and a + * delivery check that read "the last user message" compared the handoff against + * the file and reported that the payload matched neither end of itself. * - * Returns null when the transcript holds no user message with text, which is - * the honest answer for a turn that hasn't started yet. + * Tool results are identified by the `toolUseResult` field on the entry, which + * a typed message never carries, and dropped. All of them are returned rather + * than just the newest, so a caller looking for a specific payload can find it + * even after the session has gone on to do work and appended more. */ -export function readLastUserMessage(file: string): string | null { - let latest: string | null = null +export function readUserMessages(file: string): string[] { + const messages: string[] = [] for (const line of readFileSync(file, 'utf-8').split('\n')) { if (!line.trim()) continue - let entry: { message?: { role?: string; content?: unknown } } + let entry: { + message?: { role?: string; content?: unknown } + toolUseResult?: unknown + } try { entry = JSON.parse(line) } catch { continue } if (entry.message?.role !== 'user') continue + // A tool's output, not a message. See above. + if (entry.toolUseResult !== undefined) continue const content = entry.message.content const text = Array.isArray(content) @@ -189,8 +197,8 @@ export function readLastUserMessage(file: string): string | null { ? content : '' - if (text.trim()) latest = text + if (text.trim()) messages.push(text) } - return latest + return messages } From 1904558cd53f9418f886137430954617887a54c7 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Fredrik=20Wollse=CC=81n?= Date: Wed, 9 Sep 2026 21:13:52 +0300 Subject: [PATCH 3/4] docs(skill): how a driver decides WHICH tab gets a message MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The send mechanics now report faithfully whether text arrived. They say nothing about whether it should have been sent, to that tab, at all — and that is where the remaining mis-sends live: the wrong thing, to the wrong tab, that the tab already knew. Three gates, from a day of driving a ~15-tab fleet, each with its measurement: - Resolve the owner from the BRANCH, not the tab's name. Verified against a live 92-tab fleet where all three layers disagree: a tab's name, the worktree directory it runs in, and the branch checked out there can each point at a different topic. Most pointedly, one topic's owning tab was named after something unrelated and no tab on the fleet was named after the topic at all — so a name-based router finds nothing and picks whatever sounds adjacent, which is how a tab owning a quarterly report was once sent pricing material from a different worktree. - Count what the tab already knows before drafting. `transcript` shows what a tab concluded, not what it has seen, so grep the transcript for the specific phrases about to be relayed. One of six candidate tabs had nothing new and was dropped. - Relay what was said, not your conclusions: a quoted statement is checkable where a paraphrased directive is not, and a statement that contradicts the tab's own conclusion is the highest-value relay there is. Every example is written with synthetic branch names, tab names, figures and company names — the shapes, ratios and counts are real, the identifiers are not. The skill says so explicitly and tells a driver to do the same in anything written out of a fleet, because tab and branch names read like infrastructure while describing customer work, and this file ships to npm and to the marketplace. Two corrections to the guidance as received, both checked against the CLI rather than assumed: - `cctabs sessions` has no `--all` flag. Unknown flags are silently ignored, so `--all` looks like it worked while doing nothing; the skill says `--json`. - The phrase count should resolve the transcript path via `transcript --json | jq -r .transcript`, not a `~/.claude*/projects/*` glob — the glob picks the wrong file as soon as the tab runs under a backend preset, which is exactly the "reads as dead" failure the command was built to avoid. And one caveat the guidance did not claim, measured because the gate is only useful if its range is known: branch resolution answers for 22 of 92 tabs. 43 sit on `main` in the repo root, where the branch carries no ownership signal at all, and there the phrase count is the fallback rather than an invented mapping. Also records what the new refusals mean for a driver, since both fired correctly on the fleet and neither is a case for --force: "nothing from the text appeared in the tab" identifies a tab stuck on a rendered menu — unreachable, needs a human, and NOT a --path retry, since the handoff goes through the same prompt line and is swallowed the same way. Co-Authored-By: Claude Opus 5 (1M context) --- CHANGELOG.md | 1 + skills/cctabs/SKILL.md | 117 +++++++++++++++++++++++++++++++++++++++++ 2 files changed, 118 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 490c6d8..03e6c4e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,7 @@ page lives at [cctabs.com/changelog](https://cctabs.com/changelog). ## Unreleased +- **Skill: guidance on message ROUTING — which tab gets a message, upstream of whether it arrives.** Three gates from a day of driving a ~15-tab fleet, each measured: resolve the owning tab from its **branch** rather than its name (on a 92-tab fleet all three layers disagreed — the tab owning one topic's work was named after something else entirely, and no tab was named after the topic at all); **count the phrases you are about to relay in the target's transcript** before drafting, because `transcript` shows what a tab concluded and not what it has seen (one of six candidate tabs had nothing new and was dropped); and **relay what was said rather than your own conclusions**, since a quoted statement is checkable where a paraphrased instruction is not — with a contradiction of the tab's own conclusion being the highest-value relay there is. Also documents what the two new refusals mean: `nothing from the text appeared in the tab` identifies a tab stuck on a rendered menu, which is unreachable and needs a human rather than a `--path` retry, and the 1 KB busy-tab refusal means shorten the message rather than `--force` it. - **Fix: a `send` whose text quoted a flag name delivered NOTHING and reported success.** The option parser silently drops any argv element containing `--` — measured on `"mentions --verify here"`, `"--leading"`, and even `"a--b"`; a single dash survives. The text vanished from the positionals *and* the values, `send` fell through to reading stdin, stdin was empty, and it printed `✔ Sent to 5f3e853e: ⏎` with an empty preview. A ~900-byte bug report was lost this way, and the input is not exotic: any message quoting a flag name hits it, which for a tool whose users are agents reporting tool bugs is the normal case. `send` now recovers its positionals from `process.argv` directly (`core/send-argv.ts`) rather than from the parser that loses them, and `--` is supported as an explicit terminator for text that is entirely flag-shaped: `cctabs send tab -- --verify is broken`. - **Fix: an empty body is now a hard failure rather than a ✔.** Reporting success for a delivery of nothing is the same defect class as the restore success line that could not fail. An empty `--file`, or no text source at all with empty stdin, exits non-zero and says which it was; the no-source message names the `--` terminator, since a swallowed payload is the likeliest reason to land there. An *explicit* empty is still honoured — `--submit`, or a literal `""` — and now reports itself as `Submitted Enter only (no body)`, so an empty preview after a ✔ can never appear again. That preview was the operator's only tell. - **Fix: `--verify` compared against the target's NEWEST user message, which the target overwrites with its own work.** Claude records a tool's output as a `role: "user"` message, so a `--path` handoff — which tells the tab to read a file — makes the newest user-role entry the file's contents. Verify then compared the handoff against the file and reported that the payload "matches neither end of what was sent", which reads exactly like `--path` having pasted the contents. It now skips tool results (identified by the entry's `toolUseResult`) and searches every message rather than only the last, so a payload stays findable after the session has moved on. Both ends arriving in *different* messages is still not a delivery. diff --git a/skills/cctabs/SKILL.md b/skills/cctabs/SKILL.md index 6809b37..a906cc5 100644 --- a/skills/cctabs/SKILL.md +++ b/skills/cctabs/SKILL.md @@ -657,6 +657,104 @@ cctabs scrollback auth # last 50 lines cctabs scrollback auth 200 # last 200 lines ``` +## Routing: deciding WHICH tab gets a message + +The send mechanics will tell you whether text arrived. They cannot tell you +whether it should have been sent, to that tab, at all. Three gates, in order — +they are cheap, and each one has caught a real mis-send on a live fleet. + +### 1. Resolve the owner from the BRANCH, not the tab's name + +Tab names drift from scope as work moves; branches do not. Measured on a +92-tab fleet, all three layers disagreed: + +| tab name | worktree dir | branch (the authority) | +| --- | --- | --- | +| `report-q3` | `parser-limits` | `docs/report-q3-capacity-findings` | +| `cache-latency` | `cache-latency` | `fix/audit-table-per-row-scan` | +| `probe-8842` | `probe-8842` | `fix/invalid-address-and-coupon-reset` | + +(Shapes from a real fleet, names replaced.) Read the last row: **the owner of +"coupon" work is a tab called `probe-8842`**, and no tab on that fleet was named +anything like "coupon". A name-based router finds nothing and picks whatever +sounds adjacent — which is how a tab that owned a quarterly report was once sent +pricing material belonging to a different worktree. + +```bash +cctabs sessions --json | jq -r '.workspaces[].sessions[] | "\(.name)\t\(.cwd)"' +git -C worktree list --porcelain # cwd -> branch +gh pr list --search # branch -> the PRs that own it +``` + +- ⛔ **If no tab maps to the owning branch, the finding has no home in the + fleet. Say so — send nothing.** "Closest available tab" is not a routing + decision. +- ⚠️ `cctabs sessions` has **no `--all` flag**. Unknown flags are silently + ignored, so `--all` looks like it worked while doing nothing. `--json` is the + whole interface. +- ⚠️ **This gate answers for a minority of tabs, and that's fine.** On the same + fleet: 22 of 92 tabs sat on a topic branch (resolvable this way), 43 sat on + `main` in the repo root (the branch says nothing about ownership), and 27 were + in other repos. When the branch is `main`, skip to gate 2 rather than + inventing a mapping. + +### 2. Count what the tab ALREADY KNOWS before drafting + +`cctabs transcript` shows what a tab *concluded*. That is not the same as what +it has *seen* — and a message telling a tab what it already knows costs it a +cycle to read and teaches it nothing. So count the specific phrases you are +about to relay, in the tab's own transcript: + +```bash +F=$(cctabs transcript --json | jq -r .transcript) # exact path, right account +for phrase in "42,000" "Northwind" "onboarding reminder"; do + printf '%-24s %s\n' "$phrase" "$(grep -o -i -- "$phrase" "$F" | wc -l)" +done +``` + +Resolve the path through `transcript --json` rather than globbing +`~/.claude*/projects/*`: it picks the right session id *and* the right Claude +config dir, which a glob gets wrong as soon as the tab runs under a backend +preset. + +Measured — one message, six candidate tabs: + +| tab | already knew | genuinely new | +| --- | --- | --- | +| tab A | `` ×453, `` ×1265, `` ×114 | nothing → **dropped** | +| tab B | `` ×70, `42,000` ×2 | `first 500` ×0, `onboarding reminder` ×0 | +| tab C | `Northwind` ×75, `0.07` ×29 | `` ×0 | + +One of six had nothing new and was dropped. (Counts are real; the terms they +were counted on are replaced — see the note at the end of this section.) + +- **The signal is zero vs non-zero, not the magnitude.** A count of 453 and a + count of 70 mean the same thing: it knows. Only ×0 earns a place in the draft. +- **Count short distinctive tokens** — names, figures, product names — not + sentences. The transcript is JSON-escaped, so a phrase spanning a newline + won't match and you'll read a false ×0. + +### 3. Relay what was SAID — not your conclusions + +Turning transcript statements into directives ("re-aim to…", "do X before Y") +is the most common bad draft. Quote the speaker, attribute it, and leave the +inference to the receiver. Two reasons: the tab has context the driver does not, +and **a quoted statement is checkable while a paraphrased instruction is not.** + +⭐ **When a statement CONTRADICTS what the tab concluded, that is the +highest-value relay there is — send it, flagged as a contradiction.** One tab had +concluded, from four signatures, that a suspected cause was ruled out, while the +meeting concluded the opposite. It needs both, and it needs to know they +disagree; it does not need to be told which to believe. + +> **A note on the examples above.** They come from real fleets driven against +> private repositories, so every branch name, tab name, company, person and +> figure has been replaced with a synthetic stand-in; only the shapes, ratios and +> counts are real. Do the same in anything you write out of a fleet — a routing +> note, a PR body, a commit message, an issue. Tab names and branch names are +> the two that leak most easily, because they read like infrastructure rather +> than like the customer work they describe. + ## Workflow: Sending Input to a Session ```bash @@ -707,6 +805,25 @@ appear in its transcript as a tool result.** That is the handoff working — not paste. (`--verify` knows the difference: it skips tool results and searches all of the session's real messages, not just the newest.) +### When a refusal fires, the refusal is usually right + +Both of these have fired on a live fleet and been correct every time. Neither is +a case for `--force`. + +- ⛔ **`nothing from the text appeared in the tab` means the tab is on a + RENDERED MENU, and it is unreachable — escalate to a human.** A tab sitting on + the trust dialog, the resume picker or a permission prompt swallows pasted text + into the menu, so the send genuinely delivered nothing and correctly refused to + submit. **Do not retry with `--path`**: the handoff is text through the same + prompt line and is eaten the same way. Do not drive the menu blind either — + `cctabs scrollback ` shows which menu it is, and the wrong keypress in the + resume picker silently accepts a summary instead of the session (see the + restore section). A human unblocks it; a driver reports it. +- ⚠️ **The 1 KB busy-tab refusal means shorten the message, not force it + through.** It fired twice in one day of driving and shortening was the right + response both times — a multi-kilobyte brief aimed at a tab mid-turn is nearly + always a routing or timing mistake, which is what the gates above are for. + ## Workflow: Remote Control status across the fleet Claude Code's Remote Control (`/rc`, controls a session from claude.ai/code or the mobile app) is a per-process feature — cctabs doesn't manage it directly, but since it manages the tabs *running* those processes, it's the fastest way to audit or repair RC across many sessions at once. From 71ba9d0f5a635b29a0add0261745cd6be5eb8501 Mon Sep 17 00:00:00 2001 From: Motin Date: Wed, 16 Sep 2026 23:36:25 +0300 Subject: [PATCH 4/4] docs(skill): the trust dialog ate five tabs' briefs, and the doc said it couldn't (#22) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Measured 2026-09-16: seven tabs spawned, five as `cctabs new --worktree --file `. All five landed on Claude Code's "Is this a project you trust?" dialog and **none** received its brief. One brief was worse than lost — it reached the *shell* instead and started executing, npm-downloading `playwright` and `aws-cdk-lib` before that session dropped to a bare prompt. ## Two statements in the skill were wrong for an untrusted directory - The **✅ RIGHT** worktree example passed `--prompt` to a directory that cannot accept one. - *"polls internally until Claude's `❯` prompt appears before sending — **no race condition**"* is false when a dialog is what's on screen: the poll never sees `❯`. Both fixed — the example keeps `--prompt` (it's correct for a trusted repo, and is the ergonomic path) and gains its precondition; the guarantee is scoped to the *startup* race. ## Three facts the file was missing 1. The selection marker defaults to **`No, exit`** — a bare Enter EXITS the session. 2. The working keystroke is Down-then-Enter, and because `send` appends its own Enter that is **one** call: `printf '\033[B' | cctabs send `. 3. The precondition is **not** "is it a worktree". ## The precondition, and what was actually proved A blanket *"never `--prompt` with `--worktree`"* would have been wrong and annoying. Read out of Claude Code 2.1.273's own resolver and checked against the fleet: Trust lives in `${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json` under `projects[path].hasTrustDialogAccepted`. The lookup walks *up* from the tab's directory and takes the first ancestor marked `true` — **but the walk stops at the enclosing git repo root**, so trust never leaks in from above it. For a worktree, that root resolves through `gitdir:` → `commondir` to the **main repository**. Three consistent observations: - The tab this was written in (`…/cctabs/.claude/worktrees/cctabs-trust`) got **no dialog at all** — its repo root resolves to `…/generativereality/cctabs`, which is both an ancestor and trusted. - `~/Dev` was marked trusted on 2026-08-28 and still did **not** trust the repos beneath it on 09-16 — refutes unbounded ancestor inheritance outright. - Spot check of the documented one-liner: a never-visited repo → `None` (gated), one accepted after the incident → `True`. ⇒ The rule written: **"this path's repo root is already trusted, in the config dir this tab will use."** The config-dir clause matters — backend presets set `env_CLAUDE_CONFIG_DIR`, and the two config files on this machine carry genuinely different trust sets. ⚠️ **Not black-box confirmed.** The decisive live probe (launching `claude` in a scratch dir) was refused by the auto-mode classifier as self-modification, so this is *read-from-resolver plus three consistent field observations*, and version-pinned to 2.1.273. The documented pre-spawn check reads the same key the resolver reads, so it holds either way. Recorded in `HANDOVER.md`. ## Also - **Existing stuck-tab guidance cross-referenced, not duplicated** — but its *"a human unblocks it"* is narrowed: the trust dialog is two options with a known starting position, so a driver may clear it. The resume picker is untouched, where a wrong keypress silently accepts a summary. - **The meta-bug.** The driver was reading the cached skill at `/plugins/cache/generativereality/cctabs/0.5.0/…` — **691 lines, zero occurrences of "trust"** — while the CLI on PATH was `0.5.3` and the source was 1030 lines and already covered the stuck-tab case. The file warned that the CLI can lag the docs but not that the docs can lag the CLI; they're separate channels and drift both ways. Added that check. ## Release note `SKILL.md` ships in **both** the npm tarball and the marketplace, so this needs `npm run sync-plugin` at release time to actually reach the driver that hit this. **No version bump** — matches the previous doc-only commit (`1904558`), and `CLAUDE.md` puts the bump in the release flow. Two `## Unreleased` bullets added instead. `HANDOVER.md` records what was proved, what was not, and a CLI proposal left **unimplemented**: have `cctabs new` *refuse* `--prompt` into an untrusted repo rather than auto-answer the dialog — refusing, because answering a trust prompt on the user's behalf is the one decision that dialog exists to ask. Tests: 346 pass, 0 fail. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Fredrik Wollsén Co-authored-by: Claude Opus 5 (1M context) --- CHANGELOG.md | 2 + HANDOVER.md | 112 +++++++++++++++++++++++++++++++++++++++++ skills/cctabs/SKILL.md | 112 ++++++++++++++++++++++++++++++++++++++++- 3 files changed, 224 insertions(+), 2 deletions(-) create mode 100644 HANDOVER.md diff --git a/CHANGELOG.md b/CHANGELOG.md index 03e6c4e..1097442 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,8 @@ page lives at [cctabs.com/changelog](https://cctabs.com/changelog). ## Unreleased +- **Skill: the trust dialog, which silently eats a spawned tab's `--prompt`.** Measured 2026-09-16: seven tabs spawned, five with `--worktree --file `; all five landed on Claude Code's "Is this a project you trust?" dialog and none received its brief, while one brief reached the *shell* instead and began npm-downloading packages before that session dropped to a bare prompt. Three facts the skill was missing: the selection marker defaults to **`No, exit`**, so a bare Enter kills the session and the working keystroke is Down-then-Enter (`printf '\033[B' | cctabs send ` — one call, since `send` appends the Enter itself); the `cctabs new` poll for `❯` cannot see past the dialog, so the doc's "no race condition" promise held only for already-trusted directories; and the gate's real precondition is **not** "is it a worktree". Verified against Claude Code 2.1.273's own resolver: trust lives in `${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json` under `projects[path].hasTrustDialogAccepted`, and the lookup walks *up* from the tab's directory but **stops at the enclosing git repo root** — which for a worktree resolves to the main repository. So a worktree of a trusted repo is trusted (measured: no dialog) while a never-visited repo is gated no matter how many trusted ancestors it has (measured: `~/Dev` trusted since 2026-08-28 did not trust a repo beneath it). The skill now carries a read-only pre-spawn check, the spawn-bare-then-unblock recipe, and a carve-out to the "rendered menu needs a human" rule — the trust dialog is two options with a known starting position, so a driver may clear it, unlike the resume picker. +- **Skill: how to notice that the skill text itself is stale.** The marketplace plugin and the npm CLI are separate channels, and the doc only warned about the CLI half. The driver above was reading the cached skill at `/plugins/cache/generativereality/cctabs/0.5.0/…` — 691 lines, zero occurrences of "trust" — while the CLI on PATH was 0.5.3 and the source skill was 1030 lines and already covered the stuck-tab case. The cache path carries its version; the skill now says to compare it against `cctabs --version` and to treat anything absent as possibly-just-missing. - **Skill: guidance on message ROUTING — which tab gets a message, upstream of whether it arrives.** Three gates from a day of driving a ~15-tab fleet, each measured: resolve the owning tab from its **branch** rather than its name (on a 92-tab fleet all three layers disagreed — the tab owning one topic's work was named after something else entirely, and no tab was named after the topic at all); **count the phrases you are about to relay in the target's transcript** before drafting, because `transcript` shows what a tab concluded and not what it has seen (one of six candidate tabs had nothing new and was dropped); and **relay what was said rather than your own conclusions**, since a quoted statement is checkable where a paraphrased instruction is not — with a contradiction of the tab's own conclusion being the highest-value relay there is. Also documents what the two new refusals mean: `nothing from the text appeared in the tab` identifies a tab stuck on a rendered menu, which is unreachable and needs a human rather than a `--path` retry, and the 1 KB busy-tab refusal means shorten the message rather than `--force` it. - **Fix: a `send` whose text quoted a flag name delivered NOTHING and reported success.** The option parser silently drops any argv element containing `--` — measured on `"mentions --verify here"`, `"--leading"`, and even `"a--b"`; a single dash survives. The text vanished from the positionals *and* the values, `send` fell through to reading stdin, stdin was empty, and it printed `✔ Sent to 5f3e853e: ⏎` with an empty preview. A ~900-byte bug report was lost this way, and the input is not exotic: any message quoting a flag name hits it, which for a tool whose users are agents reporting tool bugs is the normal case. `send` now recovers its positionals from `process.argv` directly (`core/send-argv.ts`) rather than from the parser that loses them, and `--` is supported as an explicit terminator for text that is entirely flag-shaped: `cctabs send tab -- --verify is broken`. - **Fix: an empty body is now a hard failure rather than a ✔.** Reporting success for a delivery of nothing is the same defect class as the restore success line that could not fail. An empty `--file`, or no text source at all with empty stdin, exits non-zero and says which it was; the no-source message names the `--` terminator, since a swallowed payload is the likeliest reason to land there. An *explicit* empty is still honoured — `--submit`, or a literal `""` — and now reports itself as `Submitted Enter only (no body)`, so an empty preview after a ✔ can never appear again. That preview was the operator's only tell. diff --git a/HANDOVER.md b/HANDOVER.md new file mode 100644 index 0000000..9b08622 --- /dev/null +++ b/HANDOVER.md @@ -0,0 +1,112 @@ +# Handover — the trust dialog gate in `skills/cctabs/SKILL.md` + +Branch: `worktree-cctabs-trust`. Local commits only — **not pushed, no PR**, left for review. + +## Why + +On 2026-09-16 a driver spawned seven tabs, five as +`cctabs new --worktree --file `. All five landed on Claude Code's +trust dialog and **none** received its brief. One brief was consumed by the *shell* +and executed as commands (it began npm-downloading `playwright` and `aws-cdk-lib`) +before that session dropped to a bare prompt. + +The skill had nothing on this at spawn time, and two of its statements were actively +wrong for an untrusted directory. + +## What changed — `skills/cctabs/SKILL.md` (+110 lines, 2 modified) + +| # | Where | Change | +|---|---|---| +| 1 | "When to Use Worktrees", after the ❌/✅ block | The ✅ RIGHT example was broken for a first-time untrusted repo. Added the precondition beside it rather than changing the example — `--worktree` is not the trigger (see below), so the example is fine once the precondition is stated. | +| 2 | New `## The trust gate: what eats a --prompt before Claude ever sees it` | The prevention content: the captured dialog, the `No, exit` default, the one-call Down-then-Enter unblock, the gating rule + table, a read-only pre-spawn check, and the spawn-bare-then-send recipe. | +| 3 | "Workflow: Spawning a Parallel Agent" | "This polls internally until Claude's `❯` prompt appears before sending — **no race condition**" was a false guarantee. Now scoped to the *startup* race, with a ⚠️ stating the poll cannot see `❯` behind the trust dialog. | +| 4 | "When a refusal fires…" (the ⛔ rendered-menu bullet) | Cross-referenced rather than duplicated, and **carved out the trust dialog** from "a human unblocks it": two options with a known starting position is deterministic, unlike the resume picker where a wrong keypress silently accepts a summary. | +| 5 | "Check the installed version isn't stale" | Added the meta-bug: the *skill text* can be the stale half, because the plugin cache and the npm CLI are separate channels. How to detect it and what to do. | + +Also `CHANGELOG.md`: two bullets under `## Unreleased`. + +## The precondition — what I actually proved + +The brief suggested trust might be inherited from any trusted ancestor. **That is not +it**, and a blanket "never `--prompt` with `--worktree`" would also have been wrong. +I read Claude Code 2.1.273's own trust resolver out of the installed binary +(`strings` over `~/.local/share/claude/versions/2.1.273`) and checked the model +against the fleet's recorded state. + +**Storage** — `${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json`, key +`projects[""].hasTrustDialogAccepted`. The binary's own error string +confirms it: *"Run Claude Code interactively here once and accept the trust dialog, +or set projects[…].hasTrustDialogAccepted: true in …"*. + +**Match** — neither exact-path nor unbounded-prefix. The resolver walks *up* from the +cwd and returns at the first ancestor marked `true`, **but the walk is bounded by the +enclosing git repo root** and stops there. A separate "advisory" entry point passes a +`null` ceiling for an unbounded walk; the real check does not. + +**Worktrees** — the repo root is resolved by parsing a worktree's `.git` pointer file +through `gitdir:` → `commondir`, so it lands on the **main repository**, not the +worktree directory. + +Three observations, all consistent with that model and with nothing else I tried: + +- This tab: `…/cctabs/.claude/worktrees/cctabs-trust`, **no dialog**. It has no entry + of its own; its repo root resolves to `~/Dev/generativereality/cctabs`, which is an + ancestor *and* trusted. Confirmed on disk — the `.git` pointer resolves to + `…/cctabs/.git`, root `…/cctabs`, and that key reads `True`. +- `~/Dev` was marked trusted on **2026-08-28** and still did not trust the pilot-team + repos on 09-16 — refutes unbounded ancestor inheritance outright, and is explained + exactly by the walk stopping at each repo's own root. +- Spot check of the documented one-liner: one never-visited repo → `None` (would be + gated), one accepted after the incident → `True`. + +So the accurate rule, and the one I wrote: **"this path's repo root is already trusted, +in the config dir this tab will use."** The config-dir clause matters here — backend +presets set `env_CLAUDE_CONFIG_DIR`, and the two config files on this machine carry +genuinely different trust sets. + +⚠️ **Not fully proven.** I could not run the decisive live probe — launching `claude` +in a scratch directory was refused by the auto-mode classifier as self-modification — +so the model is *read from the shipped resolver plus consistent with three field +observations*, not black-box confirmed. It is also version-pinned: 2.1.273. If it ever +stops matching, the doc's pre-spawn check still gives the right answer for the common +case, because it reads the same key the resolver reads. + +## What I chose NOT to change + +- **The ❌/✅ worktree example** — left as-is with a precondition beside it. Rewriting it + to drop `--prompt` would teach the wrong lesson: for a trusted repo those lines are + correct and are the ergonomic path. +- **Version / plugin.json** — not bumped. Convention checked: the previous doc-only + commit (`1904558 docs(skill): how a driver decides WHICH tab gets a message`) touched + `CHANGELOG.md` + `SKILL.md` only, no version bump. `CLAUDE.md` puts the bump in the + *release* flow, not in each change. So I added `## Unreleased` bullets and stopped + there. **Whoever releases this should note it is a `SKILL.md` change** — per + `CLAUDE.md`, the skill ships in both the npm tarball and the marketplace, so it needs + `npm run sync-plugin` at release time to actually reach the driver that hit this. +- **`src/`** — no code changed. See the proposal below. +- **Existing stuck-tab guidance** — cross-referenced, not duplicated; only the "a human + unblocks it" clause was narrowed. + +## Proposal (NOT implemented): make `cctabs new` refuse rather than mislead + +Docs cannot fix this for a driver reading a stale cached skill — which is exactly the +driver that got hurt. The CLI is the channel that was current (0.5.3). + +Suggested, in `cctabs new` when `--prompt`/`--file` is given: + +1. Resolve the target's repo root (`git rev-parse --path-format=absolute + --git-common-dir`, strip `/.git`; fall back to the dir itself). +2. Read `projects[].hasTrustDialogAccepted` from `${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json`. + `src/core/config-dirs.ts` already enumerates config dirs and already carries the + backend→config-dir inference this needs. +3. If not `true`, **refuse before spawning**, naming the path and printing both the + `printf '\033[B' | cctabs send ` unblock and the bare-spawn alternative. + +Refuse, don't auto-drive: answering a trust prompt on the user's behalf is the one +decision that dialog exists to ask, and cctabs should not make it silently. A +`--assume-trusted` escape hatch would keep scripted fleet spawns working. + +Read-only, no plugin capability needed, no Tabby-side change. I did not implement it: +it touches the spawn path, the deliverable here is the doc fix, and the detection is +version-pinned to a resolver I read rather than a documented contract — it deserves its +own review and a test, not a ride-along. diff --git a/skills/cctabs/SKILL.md b/skills/cctabs/SKILL.md index a906cc5..e59caad 100644 --- a/skills/cctabs/SKILL.md +++ b/skills/cctabs/SKILL.md @@ -52,6 +52,24 @@ On your first cctabs invocation in a session, look at the version banner cctabs Don't silently work around an outdated CLI: detection heuristics, command flags, and bug fixes diverge between versions, so misbehavior on the user's machine is often "binary on PATH lags behind the plugin docs you're reading." The Claude Code marketplace plugin update path only refreshes this skill — the npm-installed CLI binary is a separate channel and must be upgraded explicitly. +⚠️ **The drift runs the other way too: THIS TEXT can be the stale half.** Because +those are two channels, the cached skill can lag the CLI by several releases. On +2026-09-16 a driver was reading +`/plugins/cache/generativereality/cctabs/0.5.0/skills/cctabs/SKILL.md` +— **691 lines, zero occurrences of the word "trust"** — while the CLI on PATH was +`0.5.3` and the source skill was 1030 lines and already documented the failure +that then cost it five tabs. The cache path carries the version, so compare it +against `cctabs --version`: + +```bash +ls -d "${CLAUDE_CONFIG_DIR:-$HOME/.claude}"/plugins/cache/*/cctabs/*/ && cctabs --version +``` + +If the cached version is behind, ask the user to run `/plugins` → Marketplaces → +Update generativereality. Until they do, treat anything *absent* from this file as +"possibly just missing here", not "not a thing" — and prefer `cctabs --help` +from the installed binary over this text where the two could disagree. + ### A one-time plugin install is needed Tabby is the terminal cctabs supports, and it **needs a small companion plugin** that exposes a localhost HTTP API the cctabs CLI talks to. @@ -168,6 +186,86 @@ cctabs new fix-auth ~/Dev/myapp --worktree --prompt "checkout PR #101 and fix li cctabs new fix-api ~/Dev/myapp --worktree --prompt "checkout PR #102 and fix tests" ``` +⚠️ Both ✅ lines carry a precondition: `~/Dev/myapp` must be a repo **Claude Code +has already been trusted in**. If it isn't, the tab opens on the trust dialog, the +`--prompt` is never seen, and the spawn still reports fine — see **"The trust +gate"** immediately below. + +## The trust gate: what eats a `--prompt` before Claude ever sees it + +⛔ **A tab opened where Claude Code isn't yet trusted never receives its +`--prompt` or `--file`.** Measured 2026-09-16: seven tabs spawned, five of them +`cctabs new --worktree --file `. All five landed on the trust +dialog and **none** got its brief. One brief was worse than lost — it reached the +*shell* instead and started executing, npm-downloading `playwright` and +`aws-cdk-lib` before that session dropped to a bare prompt. + +``` +Accessing workspace: +Quick safety check: Is this a project you created or one you trust? ... +Claude Code'll be able to read, edit, and execute files here. + > No, exit + Yes, I trust this folder +Enter to confirm Esc to cancel +``` + +⚠️ **The marker starts on `No, exit`, so a bare Enter EXITS the session.** The +working keystroke is Down-then-Enter — and `cctabs send` appends the Enter itself +(it logs `sent "\u001b[B" ⏎`), so this is **one** call, never two: + +```bash +printf '\033[B' | cctabs send # ✅ Down + send's own Enter → "Yes, I trust this folder" +cctabs send --submit # ⛔ bare Enter confirms "No, exit" and kills the tab +``` + +### Which directories are gated — it is NOT "is it a worktree" + +Trust is recorded per path in **`${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json`** as +`projects[""].hasTrustDialogAccepted` (that default really is +`~/.claude.json`, a sibling of `~/.claude/` and not inside it). The check walks *up* from the +tab's directory and takes the first ancestor marked `true` — but the walk **stops +at the enclosing git repo root**, so trust never leaks in from above that root. +For a `--worktree` path, that root resolves to the **main repository**, not the +worktree directory. + +| Directory | Gated? | +|---|---| +| A repo already trusted at its root | no — **including any `--worktree` of it** | +| A subdirectory of a trusted repo | no | +| A repo Claude Code has never run in | **yes**, however many trusted ancestors it has | +| Same repo, but under a different backend preset | **yes** — each `CLAUDE_CONFIG_DIR` has its own trust list | + +⇒ Two measurements from 2026-09-16 that pin this down. `~/Dev` had been marked +trusted on 2026-08-28 and still did **not** trust `~/Dev//` — the walk +stopped at ``'s own root. Meanwhile a `--worktree` tab cut from an +already-trusted repo got no dialog at all. So "never use `--prompt` with +`--worktree`" would be the wrong rule; the real precondition is **"this path's +repo root is already trusted, in the config dir this tab will use"**. + +⭐ **Check it before you spawn** — cheap, read-only, and answers the question +exactly: + +```bash +root=$(git -C ~/Dev/myapp rev-parse --path-format=absolute --git-common-dir); root=${root%/.git} +python3 -c 'import json,sys;print(json.load(open(sys.argv[1]))["projects"].get(sys.argv[2],{}).get("hasTrustDialogAccepted"))' \ + "${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json" "$root" +# True -> --prompt/--file will land. None/False -> it will not; spawn bare. +``` + +Anything but `True` ⇒ **spawn bare, clear the gate, then send**, so no brief is in +flight while a menu is on screen: + +```bash +cctabs new fix-auth ~/Dev/myapp --worktree # no --prompt/--file yet +cctabs scrollback fix-auth # confirm it IS the trust dialog +printf '\033[B' | cctabs send fix-auth # "Yes, I trust this folder" +cctabs send fix-auth --wait-for-prompt --path /tmp/brief.txt +``` + +A human can also pre-clear a repo once (run Claude Code in it and accept, or set +`hasTrustDialogAccepted` for that root by hand) — that is their call to make, not +a driver's, since it is the decision the dialog exists to ask. + ## Quick Reference ```bash @@ -597,13 +695,18 @@ The forked session shares full conversation history up to the fork point, then d As a Claude Code session, you can spawn a sibling session for a **genuinely independent** parallel task: -**Preferred: pass the initial task directly to `cctabs new`** using `--prompt` or `--file`. This polls internally until Claude's `❯` prompt appears before sending — no race condition: +**Preferred: pass the initial task directly to `cctabs new`** using `--prompt` or `--file`. This polls internally until Claude's `❯` prompt appears before sending, which closes the *startup* race: ```bash cctabs new payments ~/Dev/myapp --prompt "implement the billing endpoint" cctabs new payments ~/Dev/myapp --file /tmp/task.txt ``` +⚠️ **That guarantee holds only in an already-trusted directory.** The poll cannot +see `❯` behind the trust dialog, so in an untrusted repo the brief is not +delivered at all — and has been measured reaching the *shell* instead. Check the +precondition first: see **"The trust gate"** above. + If you need to send a task after the fact, poll first — and for anything sizeable, hand over a **path** rather than the text: @@ -818,7 +921,12 @@ a case for `--force`. prompt line and is eaten the same way. Do not drive the menu blind either — `cctabs scrollback ` shows which menu it is, and the wrong keypress in the resume picker silently accepts a summary instead of the session (see the - restore section). A human unblocks it; a driver reports it. + restore section). For the resume picker and permission prompts, a human + unblocks it; a driver reports it. **The trust dialog is the one exception** — + two options with the marker parked on `No, exit`, so + `printf '\033[B' | cctabs send ` is deterministic rather than a guess. Best + of all, don't arrive here: the gate is preventable at spawn time, see **"The + trust gate"**. - ⚠️ **The 1 KB busy-tab refusal means shorten the message, not force it through.** It fired twice in one day of driving and shortening was the right response both times — a multi-kilobyte brief aimed at a tab mid-turn is nearly