feat: read what a tab has learned, pin a working set, and stop claiming sends landed - #21
Conversation
…ng sends landed Six gaps measured while driving a 13-tab fleet, each one a place where cctabs reported something it had not established. - `cctabs transcript <tab> [n]` (alias `findings`) prints a tab's last n assistant messages from its transcript. `scrollback` returns the last painted frame, so a tab mid-turn shows a spinner — which meant sessions briefed each other from stale pictures. Searches every Claude config dir, because a tab on a backend preset writes under that preset's root and looking in one root reports "no transcript" for a healthy session. - `cctabs sort --first a,b,c` pins a chosen set to the front. The plugin's /api/tabs/reorder already did exactly this; no verb exposed it. Refuses as a whole if a name doesn't resolve — half a working set in reach, with no indication which half, is worse than an error. - `sessions --json` now carries `session_lookup`, so a null `session_id` says whether we looked and found nothing, couldn't look, or never looked. A lookup that threw was previously swallowed into the same null. - `restore` verifies what it spawned: re-reads the tab list, checks each tab has a process, resolves its session, and reports N verified / N unconfirmed / N failed, exiting non-zero on failure. A tab that came back as a *different* session counts as failed — `--resume` on an id it can't find opens a fresh conversation, so the tab looks perfect and the context is gone. The line this replaces read "78 spawned, 0 failed" with one tab absent and one stripped of its context. - `send` separates three claims that were one ✔ line: nothing arrived (hard failure, and the body is NOT submitted — a fragment reads as a whole message), arrived but completeness unverified, and verified. `--verify` does the real comparison against the target's transcript. `--path` hands over a file path instead of pasting, which has no truncation surface. `--submit` presses Enter only. `--wait-for-prompt` reads the whole tail of the buffer, so a "Restart to update" banner below a ready prompt no longer defeats it. Worth recording what the measurements actually settled about that last one, since the obvious fix is wrong: a paste chip's `+N lines` does NOT track the payload. A 6,892-byte, 76-line payload was measured delivering *completely* into an idle tab while its chip read `+10 lines`, so failing a send on a short chip fails healthy sends. The screen cannot establish completeness for a collapsed paste at all — hence "unverified" as a real answer, and hence `--verify`, which reads what the session recorded receiving. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ported ✔
Two bugs reported against the new send, both measured. The serious one first.
A send whose text contained `--` delivered NOTHING and printed success. The
option parser silently drops any argv element containing a double dash —
reproduced on "mentions --verify here", "--leading", and even "a--b"; a single
dash survives. The text was absent from both the positionals and the values, so
`send` fell through to stdin, stdin was empty, and it printed
✔ Sent to 5f3e853e: ⏎
with an empty preview, which was the only tell. A ~900-byte bug report was lost
this way. The input is not exotic: any message quoting a flag name hits it, and
for a tool whose users are agents reporting tool bugs that is the normal case.
Positionals now come from process.argv directly (core/send-argv.ts) instead of
from the parser that loses them, and `--` works as an explicit terminator for
text that is entirely flag-shaped. An empty body is a hard failure now rather
than a ✔ — reporting success for a delivery of nothing is the same defect class
as the restore success line that could not fail — while a deliberate bare Enter
names itself, so an empty preview can never follow a ✔ again.
The second report was that `--path` still pastes file contents. It does not:
measured on a 2,470-byte file, the payload sent was the 257-char handoff. What
was really wrong is `--verify`. Claude records a tool's output as a
`role: "user"` message, and a `--path` handoff tells the tab to read the file —
so the newest user-role entry becomes the file's contents, and verify compared
the handoff against the file and reported "matches neither end of what was
sent". That reads exactly like a truncated paste, which is how it was reported.
Verify now skips tool results (the entry's `toolUseResult` field identifies
them) and searches every real message rather than only the newest.
Also fixed while in there: `--verify` resolved the session once, before its poll
loop, so a freshly spawned tab failed instantly with "no session resolved" a
second before its transcript existed. Lookup retries for the whole window now,
dropping the per-process title-index cache each round.
Confirmed for the smaller note: a verify mismatch exits 1.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both reports fixed and pushed (
|
| argument | result |
|---|---|
"hello world" |
kept |
"hello -x" |
kept |
"mentions --verify here" |
dropped |
"--leading" |
dropped |
"a--b" |
dropped |
"trailing --" |
dropped |
So it isn't that a flag name gets consumed as an option — nothing is set in
values either. A single dash survives; -- anywhere does not. Your text
contained --file, --path and --verify, so the whole argument vanished,
send fell through to reading stdin, stdin was empty, and the empty preview
followed. Exactly the two symptoms you predicted, and exactly the diagnosis you
proposed as the fix.
Fixed as you suggested: positionals now come from process.argv directly
(core/send-argv.ts) rather than from the parser that loses them, and -- is
supported as an explicit terminator for text that is entirely flag-shaped.
Both forms verified end to end against a live tab:
$ cctabs send <tab> "BUG 1 — the --path option does not hand over a path... --file ... --verify" --verify
✔ Sent and verified: the session received the whole payload (109 non-whitespace chars sent)
$ cctabs send <tab> -- --verify is broken and --path too
✔ Sent to cbfb77e9: "--verify is broken and --path too" ⏎
16 test cases cover it, including a--b and a tab name containing --.
And the class of failure, not just the instance: an empty body is now a hard
error (exit 1) instead of a ✔, because reporting success for a delivery of
nothing is the same defect as the restore success line that could not fail. A
deliberate bare Enter reports Submitted Enter only (no body), so an empty
preview can never follow a ✔ again — your tell is now unambiguous by
construction.
BUG 1 — --path is not pasting; --verify was lying to you
I could not reproduce a paste. On a 2,470-byte fixture, the payload sent was the
257-char handoff, and the file was never read by send.
What actually went wrong is --verify. Claude records a tool's output as a
role: "user" message, and a --path handoff instructs the tab to read the
file — so the newest user-role entry becomes the file's contents. Verify
compared the handoff against the file. Straight from the transcript:
user[str] len=284 'Read the file at /private/tmp/.../bug1-fixture.txt...'
user[TOOL_RESULT] len=2469 'BUG REPORT (test fixture, 2421 bytes target)...'
That 2,469 is your 2,651. It is the file being read, which is --path working —
and comparing against it produced "matches neither end of what was sent", which
reads precisely like a front-clipped paste. Your inference was the only
reasonable one from the evidence the tool gave you; the tool gave you the wrong
evidence.
Verify now skips tool results (the entry's toolUseResult field identifies
them) and searches every real message instead of only the newest, so a payload
stays findable after the session has moved on. Both ends turning up in
different messages is still not a delivery.
I also fixed a second verify defect found while reproducing: it resolved the
session once, before its poll loop, so a freshly spawned tab failed
instantly with no session resolved for tab a second before its transcript
existed. Lookup now retries for the whole window.
--path's help text also overclaimed in the way that set this up — it implied
nothing of the file crosses into the session's record. It now says the receiving
session reads the file, so the contents appear in its transcript as a tool
result, and that this is the handoff working rather than a paste.
Smaller note — confirmed non-zero
$ cctabs send <tab> "payload" --verify --verify-timeout 0
ERROR ... delivery does NOT check out: no session ... appeared within 0s
$ echo $?
1
A verify mismatch exits 1. (Worth flagging that I made your exact mistake twice
in this session — cmd | tail reports the pipe's status, so I re-checked every
exit code without a pipe.)
npm run check: 346 pass, 0 fail. Reinstalled globally from the working
tree. Both scratch tabs closed; the fleet is back as it was.
The send mechanics now report faithfully whether text arrived. They say nothing about whether it should have been sent, to that tab, at all — and that is where the remaining mis-sends live: the wrong thing, to the wrong tab, that the tab already knew. Three gates, from a day of driving a ~15-tab fleet, each with its measurement: - Resolve the owner from the BRANCH, not the tab's name. Verified against a live 92-tab fleet where all three layers disagree: a tab's name, the worktree directory it runs in, and the branch checked out there can each point at a different topic. Most pointedly, one topic's owning tab was named after something unrelated and no tab on the fleet was named after the topic at all — so a name-based router finds nothing and picks whatever sounds adjacent, which is how a tab owning a quarterly report was once sent pricing material from a different worktree. - Count what the tab already knows before drafting. `transcript` shows what a tab concluded, not what it has seen, so grep the transcript for the specific phrases about to be relayed. One of six candidate tabs had nothing new and was dropped. - Relay what was said, not your conclusions: a quoted statement is checkable where a paraphrased directive is not, and a statement that contradicts the tab's own conclusion is the highest-value relay there is. Every example is written with synthetic branch names, tab names, figures and company names — the shapes, ratios and counts are real, the identifiers are not. The skill says so explicitly and tells a driver to do the same in anything written out of a fleet, because tab and branch names read like infrastructure while describing customer work, and this file ships to npm and to the marketplace. Two corrections to the guidance as received, both checked against the CLI rather than assumed: - `cctabs sessions` has no `--all` flag. Unknown flags are silently ignored, so `--all` looks like it worked while doing nothing; the skill says `--json`. - The phrase count should resolve the transcript path via `transcript --json | jq -r .transcript`, not a `~/.claude*/projects/*` glob — the glob picks the wrong file as soon as the tab runs under a backend preset, which is exactly the "reads as dead" failure the command was built to avoid. And one caveat the guidance did not claim, measured because the gate is only useful if its range is known: branch resolution answers for 22 of 92 tabs. 43 sit on `main` in the repo root, where the branch carries no ownership signal at all, and there the phrase count is the fallback rather than an invented mapping. Also records what the new refusals mean for a driver, since both fired correctly on the fleet and neither is a case for --force: "nothing from the text appeared in the tab" identifies a tab stuck on a rendered menu — unreachable, needs a human, and NOT a --path retry, since the handoff goes through the same prompt line and is swallowed the same way. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
acc34c4 to
1904558
Compare
… it couldn't (#22) Measured 2026-09-16: seven tabs spawned, five as `cctabs new <name> <dir> --worktree --file <brief>`. All five landed on Claude Code's "Is this a project you trust?" dialog and **none** received its brief. One brief was worse than lost — it reached the *shell* instead and started executing, npm-downloading `playwright` and `aws-cdk-lib` before that session dropped to a bare prompt. ## Two statements in the skill were wrong for an untrusted directory - The **✅ RIGHT** worktree example passed `--prompt` to a directory that cannot accept one. - *"polls internally until Claude's `❯` prompt appears before sending — **no race condition**"* is false when a dialog is what's on screen: the poll never sees `❯`. Both fixed — the example keeps `--prompt` (it's correct for a trusted repo, and is the ergonomic path) and gains its precondition; the guarantee is scoped to the *startup* race. ## Three facts the file was missing 1. The selection marker defaults to **`No, exit`** — a bare Enter EXITS the session. 2. The working keystroke is Down-then-Enter, and because `send` appends its own Enter that is **one** call: `printf '\033[B' | cctabs send <tab>`. 3. The precondition is **not** "is it a worktree". ## The precondition, and what was actually proved A blanket *"never `--prompt` with `--worktree`"* would have been wrong and annoying. Read out of Claude Code 2.1.273's own resolver and checked against the fleet: Trust lives in `${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json` under `projects[path].hasTrustDialogAccepted`. The lookup walks *up* from the tab's directory and takes the first ancestor marked `true` — **but the walk stops at the enclosing git repo root**, so trust never leaks in from above it. For a worktree, that root resolves through `gitdir:` → `commondir` to the **main repository**. Three consistent observations: - The tab this was written in (`…/cctabs/.claude/worktrees/cctabs-trust`) got **no dialog at all** — its repo root resolves to `…/generativereality/cctabs`, which is both an ancestor and trusted. - `~/Dev` was marked trusted on 2026-08-28 and still did **not** trust the repos beneath it on 09-16 — refutes unbounded ancestor inheritance outright. - Spot check of the documented one-liner: a never-visited repo → `None` (gated), one accepted after the incident → `True`. ⇒ The rule written: **"this path's repo root is already trusted, in the config dir this tab will use."** The config-dir clause matters — backend presets set `env_CLAUDE_CONFIG_DIR`, and the two config files on this machine carry genuinely different trust sets.⚠️ **Not black-box confirmed.** The decisive live probe (launching `claude` in a scratch dir) was refused by the auto-mode classifier as self-modification, so this is *read-from-resolver plus three consistent field observations*, and version-pinned to 2.1.273. The documented pre-spawn check reads the same key the resolver reads, so it holds either way. Recorded in `HANDOVER.md`. ## Also - **Existing stuck-tab guidance cross-referenced, not duplicated** — but its *"a human unblocks it"* is narrowed: the trust dialog is two options with a known starting position, so a driver may clear it. The resume picker is untouched, where a wrong keypress silently accepts a summary. - **The meta-bug.** The driver was reading the cached skill at `<config-dir>/plugins/cache/generativereality/cctabs/0.5.0/…` — **691 lines, zero occurrences of "trust"** — while the CLI on PATH was `0.5.3` and the source was 1030 lines and already covered the stuck-tab case. The file warned that the CLI can lag the docs but not that the docs can lag the CLI; they're separate channels and drift both ways. Added that check. ## Release note `SKILL.md` ships in **both** the npm tarball and the marketplace, so this needs `npm run sync-plugin` at release time to actually reach the driver that hit this. **No version bump** — matches the previous doc-only commit (`1904558`), and `CLAUDE.md` puts the bump in the release flow. Two `## Unreleased` bullets added instead. `HANDOVER.md` records what was proved, what was not, and a CLI proposal left **unimplemented**: have `cctabs new` *refuse* `--prompt` into an untrusted repo rather than auto-answer the dialog — refusing, because answering a trust prompt on the user's behalf is the one decision that dialog exists to ask. Tests: 346 pass, 0 fail. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Fredrik Wollsén <fredrik.wollsen@f-secure.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
# Conflicts: # CHANGELOG.md
…d Windows Three merged PRs: #16 (the skill description was being truncated past its key rule, plus a real cctab alias), #21 (cctabs transcript, sort --first, and a send that establishes delivery instead of asserting it), #23 (Windows - the spawn, the pid walk and the plugin-install path, each of which reported success while failing). #23 shipped without touching CHANGELOG.md, so its two entries were written for this release rather than left undocumented.
Six gaps measured while driving a 13-tab fleet. The common thread: cctabs
reported outcomes it had not established — a spawn that returned counted as a
tab that worked, a paste chip counted as a delivery, a null
session_idmeantthree different things.
Branched fresh off
origin/main(v0.5.3).feat/whoamiis untouched — see"On the two local commits" below.
Items 1 and 3 — the two that change how a fleet is driven
cctabs transcript <tab> [n](aliasfindings) prints a tab's last nassistant messages, read from the transcript rather than the screen.
scrollbackreturns the last painted frame, so a tab mid-turn shows a spinnerand nothing else — which is why sessions were briefing each other from stale
pictures. It searches every Claude config dir: a tab on a backend preset
writes beneath that preset's own root, and looking in only
~/.claude/projectsreports "no transcript" for a healthy session, which reads as dead. It exits
non-zero and names which failure it hit, including a count of transcripts that
do exist for the tab's directory — that number distinguishes a tab renamed out
from under a live session from a directory nothing has ever run in.
cctabs sort --first a,b,cpins a chosen set to the front. The plugin'sPOST /api/tabs/reorderalready had exactly the needed semantics (unlisted tabskeep their relative order and sort after the listed ones); no CLI verb exposed
it. Pinning skips the activity scan entirely, which is ~7.7s of transcript
reading on a 65-tab fleet. If any name doesn't resolve to exactly one tab,
nothing moves and it exits non-zero.
Items 2, 5, 6
sessions --jsonsays why asession_idis missing —session_lookupisone of
found/not-found(withsessions_in_dir) /no-cwd/lookup-failed(with the error). A lookup that threw was previouslyswallowed into the same null as "searched and found nothing", and those call
for opposite responses.
restorereports the count it achieved. It re-reads the tab list afterspawning, checks each tab has a process, resolves its session from disk, and
reports
N verified, N unconfirmed, N failed, exiting non-zero on failure. Atab that came back as a different session counts as failed:
--resumeon anid it can't find opens a fresh conversation, so the tab looks perfect and the
context is gone.
unconfirmedis a first-class outcome rather than roundedeither way. (This needed a title-index cache reset — the cache was built while
planning and would have cheerfully confirmed the pre-restore world.)
send --submitpresses Enter only. Also: an empty body with no--submitnow says so instead of silently pressing Enter in someone'ssession, and contradictory text sources are rejected rather than silently
ranked.
sendrefuses payloads over 1 KB into a tab with a turn in flight(
--forceoverrides). Short replies into a busy tab —yes,2— are whatsending into an active tab is for and still work.
Item 4 — and a correction worth reading
The measured incident: 6,835 bytes sent into an idle tab, 756 bytes
recorded, front-clipped, both ends reporting success. Root cause found:
sendTextWithConfirmationaccepted[Pasted textappearing anywhere in thebuffer as proof, and that chip renders no matter how little arrived.
My first fix compared the chip's
+N linesagainst the payload. That fix waswrong, and the verification the brief asked for is what caught it. Sending
6,892 bytes into an idle tab and then asking the receiving session what it got:
Front and tail both arrived — the payload delivered complete — while the
chip on screen read
+10 lines. So the chip's count does not track the payload,and failing a send on a shortfall fails healthy sends. I could not reproduce
front-truncation at 6.9 KB at all.
What that leaves is a real conclusion: for a collapsed paste the screen cannot
establish completeness. So
sendnow separates three claims that used to beone check-mark line —
pressing Enter on a fragment sends something that reads as a complete message.
Text is left in the input box; exit 1.
— and
--verifydoes the comparison item 4(b) actually asked for, againstground truth: it reads the target's own transcript, which records the user
message as received, and checks the payload's front and tail fingerprints,
naming which end went missing.
--path(item 4c) is first-class: it hands the tab a file path and lets thereceiving session read it, so only a path crosses the prompt line and there is
no truncation surface. Verified end to end.
--wait-for-prompttested only the buffer's last non-empty line, and Clauderenders notices below its input line — a
Restart to updatebanner defeatedit for the full timeout. It now reads the whole tail of the window.
Verified against the live 86-tab fleet, with the new build installed
transcripton a tab whose session lives under~/.claude-enterprise—resolved it, reported
backend=enterprise, printed the turn.sessions --json: 82found, 4not-found. One of those,ltv-cac-ent,reports 99 transcripts in its directory — previously an indistinguishable
session_id: null.sort --firstapplied for real (two tabs to the front, rest kept relativeorder), then the original 86-tab order restored exactly.
sendat 6,892 bytes into a confirmed-idle tab, with--verify;--pathhandoff; the nonexistent-
--pathand contradictory-source refusals.On the two local commits
The brief said to preserve
feat/whoami's two local-only commits.git ls-remote origin feat/whoamiis indeed empty, so they are unpushed as commits — buttheir content is already on
main:f9d8331landed squash-merged asbb9a05c(#20), and the send-confirm work plus a lateroptionalAdapterrefinement shipped in v0.5.3.
git diff origin/main feat/whoamishows thebranch removing work main has, so rebasing it would revert those refinements.
I branched fresh off
origin/mainand leftfeat/whoamiuntouched ate352987.Not done, and why
Untouched.
npm run sync-pluginnot run: it writes to../plugins, and the repo'sown
prepacksync check is already failing onmain(plugins repo at 0.5.1vs package.json 0.5.3) independently of this branch. That's a release step.
from a packed tarball of the working tree, not a registry install.
@types/bunwas missing from the stale localnode_modules, sonpm run typecheckfailed onmaintoo;npm cifixed it. No source changeinvolved.
npm run check: 322 pass, 0 fail, typecheck and build clean.🤖 Generated with Claude Code