Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,27 @@ page lives at [cctabs.com/changelog](https://cctabs.com/changelog).

- **The skill description was being silently truncated, and the part cut off was the rule that matters.** Claude Code caps a skill's description at 1535 characters in the available-skills listing; ours ran 1760, so every session ever shown this skill saw it cut mid-sentence at `"do this in parallel without a new tab", or w…` — losing the tail of the anti-substitution rule and the whole `NOT for:` line. The description now runs 1432 characters, and the decisive sentence (*a subagent is NOT a tab*) has moved **ahead** of the TRIGGER list so it can no longer be the thing that gets dropped. Every trigger phrase is preserved verbatim. The opening clause now leads with the literal tokens `cctab` / `cctabs` / `terminal tabs`, because the first clause is what registers in a listing where this skill can sit 75% of the way down 82 entries.
- **`cctab` is now a real command.** The skill has always advertised "cctab" as a singular alias, but only `cctabs` was ever installed — so a session that went looking for the CLI under the name the docs use got `command not found` and concluded the tool did not exist. `cctab` is now a second bin pointing at the same entry point. This matters precisely in the case the listing cannot help with: Claude Code injects the full skill listing once per session and does **not** re-inject it after a compaction, so in a long session probing the shell is the only route left back to cctabs.
- **Skill: the trust dialog, which silently eats a spawned tab's `--prompt`.** Measured 2026-09-16: seven tabs spawned, five with `--worktree --file <brief>`; all five landed on Claude Code's "Is this a project you trust?" dialog and none received its brief, while one brief reached the *shell* instead and began npm-downloading packages before that session dropped to a bare prompt. Three facts the skill was missing: the selection marker defaults to **`No, exit`**, so a bare Enter kills the session and the working keystroke is Down-then-Enter (`printf '\033[B' | cctabs send <tab>` — one call, since `send` appends the Enter itself); the `cctabs new` poll for `❯` cannot see past the dialog, so the doc's "no race condition" promise held only for already-trusted directories; and the gate's real precondition is **not** "is it a worktree". Verified against Claude Code 2.1.273's own resolver: trust lives in `${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json` under `projects[path].hasTrustDialogAccepted`, and the lookup walks *up* from the tab's directory but **stops at the enclosing git repo root** — which for a worktree resolves to the main repository. So a worktree of a trusted repo is trusted (measured: no dialog) while a never-visited repo is gated no matter how many trusted ancestors it has (measured: `~/Dev` trusted since 2026-08-28 did not trust a repo beneath it). The skill now carries a read-only pre-spawn check, the spawn-bare-then-unblock recipe, and a carve-out to the "rendered menu needs a human" rule — the trust dialog is two options with a known starting position, so a driver may clear it, unlike the resume picker.
- **Skill: how to notice that the skill text itself is stale.** The marketplace plugin and the npm CLI are separate channels, and the doc only warned about the CLI half. The driver above was reading the cached skill at `<config-dir>/plugins/cache/generativereality/cctabs/0.5.0/…` — 691 lines, zero occurrences of "trust" — while the CLI on PATH was 0.5.3 and the source skill was 1030 lines and already covered the stuck-tab case. The cache path carries its version; the skill now says to compare it against `cctabs --version` and to treat anything absent as possibly-just-missing.
- **Skill: guidance on message ROUTING — which tab gets a message, upstream of whether it arrives.** Three gates from a day of driving a ~15-tab fleet, each measured: resolve the owning tab from its **branch** rather than its name (on a 92-tab fleet all three layers disagreed — the tab owning one topic's work was named after something else entirely, and no tab was named after the topic at all); **count the phrases you are about to relay in the target's transcript** before drafting, because `transcript` shows what a tab concluded and not what it has seen (one of six candidate tabs had nothing new and was dropped); and **relay what was said rather than your own conclusions**, since a quoted statement is checkable where a paraphrased instruction is not — with a contradiction of the tab's own conclusion being the highest-value relay there is. Also documents what the two new refusals mean: `nothing from the text appeared in the tab` identifies a tab stuck on a rendered menu, which is unreachable and needs a human rather than a `--path` retry, and the 1 KB busy-tab refusal means shorten the message rather than `--force` it.
- **Fix: a `send` whose text quoted a flag name delivered NOTHING and reported success.** The option parser silently drops any argv element containing `--` — measured on `"mentions --verify here"`, `"--leading"`, and even `"a--b"`; a single dash survives. The text vanished from the positionals *and* the values, `send` fell through to reading stdin, stdin was empty, and it printed `✔ Sent to 5f3e853e: ⏎` with an empty preview. A ~900-byte bug report was lost this way, and the input is not exotic: any message quoting a flag name hits it, which for a tool whose users are agents reporting tool bugs is the normal case. `send` now recovers its positionals from `process.argv` directly (`core/send-argv.ts`) rather than from the parser that loses them, and `--` is supported as an explicit terminator for text that is entirely flag-shaped: `cctabs send tab -- --verify is broken`.
- **Fix: an empty body is now a hard failure rather than a ✔.** Reporting success for a delivery of nothing is the same defect class as the restore success line that could not fail. An empty `--file`, or no text source at all with empty stdin, exits non-zero and says which it was; the no-source message names the `--` terminator, since a swallowed payload is the likeliest reason to land there. An *explicit* empty is still honoured — `--submit`, or a literal `""` — and now reports itself as `Submitted Enter only (no body)`, so an empty preview after a ✔ can never appear again. That preview was the operator's only tell.
- **Fix: `--verify` compared against the target's NEWEST user message, which the target overwrites with its own work.** Claude records a tool's output as a `role: "user"` message, so a `--path` handoff — which tells the tab to read a file — makes the newest user-role entry the file's contents. Verify then compared the handoff against the file and reported that the payload "matches neither end of what was sent", which reads exactly like `--path` having pasted the contents. It now skips tool results (identified by the entry's `toolUseResult`) and searches every message rather than only the last, so a payload stays findable after the session has moved on. Both ends arriving in *different* messages is still not a delivery.
- **Fix: `--verify` failed instantly against a freshly spawned tab.** It resolved the session once, before its poll loop, so a tab whose transcript was a second from existing reported "no session resolved for tab" and failed. Session lookup now retries for the whole window, dropping the per-process title-index cache each round — the session being waited for is precisely the one that appears after that cache was built.
- `send --path`'s help text no longer implies the receiving session's transcript stays clean: the session obeys the handoff by *reading* the file, so the contents land there as a tool result. That is the handoff working, and mistaking it for a paste is what the report above turned on.


- **`cctabs transcript <tab> [n]` reads what a tab has SAID, not what it is painting.** `scrollback` returns the last painted frame, so a tab mid-turn shows a spinner and nothing else — which meant sessions briefed each other from stale pictures, and one such brief told a tab to go measure two things it had already measured, missing a third it had found. Prints the last n assistant messages from the transcript (`--json` for a driver; `findings` as an alias). It searches **every** Claude config dir, because a tab on a backend preset writes under that preset's own root and looking in only `~/.claude/projects` reports "no transcript" for a healthy session — which reads as dead. Tool-only and thinking-only messages are skipped. It exits non-zero and names the failure — no session titled after this tab (with a count of transcripts that do exist for its directory, which distinguishes a renamed tab from a dead one), no transcript on disk, or a lookup that threw — because "I couldn't read it" and "it hasn't said anything" are different answers.
- **`cctabs sort --first a,b,c` pins a chosen set to the front of the bar.** Activity order is close to the opposite of what a driver wants: a tab that just delivered sinks. The Tabby plugin's `POST /api/tabs/reorder` already did exactly this — unlisted tabs keep their relative order and sort after the listed ones — with no CLI verb exposing it. Naming tabs skips the activity scan entirely (~7.7s of transcript reading on a 65-tab fleet). If any name doesn't resolve to exactly one tab, nothing moves and it exits non-zero: half a working set in reach, with no indication which half, is worse than an error.
- **`cctabs send` no longer reports success it hasn't established.** A large paste collapses to a `[Pasted text #N +M lines]` chip, and the chip renders no matter how little arrived — so "a chip is on screen" was accepted as proof of delivery, which is how a 6,835-byte brief that landed as its last 756 bytes came to be reported as sent. `send` now distinguishes three claims that used to be one ✔ line: *nothing arrived* (a hard failure — and the body is **not submitted**, because pressing Enter on a fragment sends something that reads as a complete message; the text is left in the input box and the command exits non-zero), *something arrived but completeness is unverified* (a warning naming the two ways to establish it), and *verified*.
- **`cctabs send --verify` checks what the target session actually RECEIVED.** After submitting, it reads the target's own transcript — which records the user message as received — and compares the payload's front and tail fingerprints against it, failing loudly on a mismatch and naming which end went missing. This is the only reliable completeness check available, and finding that out cost a measurement: the chip's `+N lines` count does **not** track the payload. A 6,892-byte, 76-line payload was measured delivering completely into an idle tab while its chip read `+10 lines`, so treating a shortfall as truncation fails healthy sends. The screen cannot answer this question; the transcript can.
- **`cctabs send --path <file>` hands a tab a file path instead of pasting its contents.** No truncation surface at all: what crosses the prompt line is a path, and the receiving session reads the payload from disk. This is the fallback that had to be used to deliver the brief describing the truncation, and it is now first-class rather than a discipline.
- **`cctabs send` refuses payloads over 1 KB into a tab with a turn in flight**, with `--force` to override. Short replies into a busy tab — `yes` to a tool call, `2` to a picker — are exactly what sending into an active tab is for and still work; a multi-kilobyte brief pushed into a busy input handler is how text gets clipped.
- **`cctabs send --submit` presses Enter only**, submitting a prompt already parked in a tab's input box. An empty body with no `--submit` now says so instead of silently pressing Enter in someone's session.
- **`cctabs send --wait-for-prompt` no longer times out against tabs that are ready.** It tested only the buffer's last non-empty line, and Claude renders notices *below* its input line — a `Restart to update` banner was enough to defeat it for the full 20s. It now reads the whole tail of the window.
- **`cctabs restore` reports the count it actually achieved.** It read "78 spawned, 0 failed" while one tab was absent entirely and another had come back with no session, having lost its context — a summary computed from "did the spawn call return?", which cannot report a failure that happens after it returns. Restore now re-reads the tab list after spawning, checks each tab has a process, resolves its session from disk, and reports `N verified, N unconfirmed, N failed`, exiting non-zero when anything failed. A tab that came back as a *different* session than requested counts as failed: `claude --resume` on an id it can't find opens a fresh conversation, so the tab looks perfect and the context is gone. `unconfirmed` is a first-class outcome rather than rounded either way.
- **`cctabs sessions --json` says why a `session_id` is missing.** Every row carries `session_lookup`: `found`, `not-found` (with `sessions_in_dir`, so a tab renamed out from under a live session is distinguishable from a directory nothing has ever run in), `no-cwd`, or `lookup-failed` (with the error — a lookup that threw is *unknown*, not absent, and was previously swallowed into the same null). Observed on a fleet as a tab whose session was perfectly readable but filed under a title that no longer matched.
- Internal: `send`, `scrollback` and `transcript` now share one target resolver (`core/tab-target.ts`) instead of three verbatim copies of the tab-then-block lookup, and `scrollback <tab> [n]` accepts the bare line count the docs have always shown.

## 0.5.3 — 2026-09-07

Expand Down
Loading