Skip to content

feat: read what a tab has learned, pin a working set, and stop claiming sends landed - #21

Merged
motin merged 5 commits into
mainfrom
feat/fleet-driving
Sep 17, 2026
Merged

motin merged 5 commits into
mainfrom
feat/fleet-driving

Conversation

@motin

@motin motin commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

Six gaps measured while driving a 13-tab fleet. The common thread: cctabs
reported outcomes it had not established — a spawn that returned counted as a
tab that worked, a paste chip counted as a delivery, a null session_id meant
three different things.

Branched fresh off origin/main (v0.5.3). feat/whoami is untouched — see
"On the two local commits" below.

Items 1 and 3 — the two that change how a fleet is driven

cctabs transcript <tab> [n] (alias findings) prints a tab's last n
assistant messages, read from the transcript rather than the screen.
scrollback returns the last painted frame, so a tab mid-turn shows a spinner
and nothing else — which is why sessions were briefing each other from stale
pictures. It searches every Claude config dir: a tab on a backend preset
writes beneath that preset's own root, and looking in only ~/.claude/projects
reports "no transcript" for a healthy session, which reads as dead. It exits
non-zero and names which failure it hit, including a count of transcripts that
do exist for the tab's directory — that number distinguishes a tab renamed out
from under a live session from a directory nothing has ever run in.

cctabs sort --first a,b,c pins a chosen set to the front. The plugin's
POST /api/tabs/reorder already had exactly the needed semantics (unlisted tabs
keep their relative order and sort after the listed ones); no CLI verb exposed
it. Pinning skips the activity scan entirely, which is ~7.7s of transcript
reading on a 65-tab fleet. If any name doesn't resolve to exactly one tab,
nothing moves and it exits non-zero.

Items 2, 5, 6

  • sessions --json says why a session_id is missing — session_lookup is
    one of found / not-found (with sessions_in_dir) / no-cwd /
    lookup-failed (with the error). A lookup that threw was previously
    swallowed into the same null as "searched and found nothing", and those call
    for opposite responses.
  • restore reports the count it achieved. It re-reads the tab list after
    spawning, checks each tab has a process, resolves its session from disk, and
    reports N verified, N unconfirmed, N failed, exiting non-zero on failure. A
    tab that came back as a different session counts as failed: --resume on an
    id it can't find opens a fresh conversation, so the tab looks perfect and the
    context is gone. unconfirmed is a first-class outcome rather than rounded
    either way. (This needed a title-index cache reset — the cache was built while
    planning and would have cheerfully confirmed the pre-restore world.)
  • send --submit presses Enter only. Also: an empty body with no
    --submit now says so instead of silently pressing Enter in someone's
    session, and contradictory text sources are rejected rather than silently
    ranked.
  • send refuses payloads over 1 KB into a tab with a turn in flight
    (--force overrides). Short replies into a busy tab — yes, 2 — are what
    sending into an active tab is for and still work.

Item 4 — and a correction worth reading

The measured incident: 6,835 bytes sent into an idle tab, 756 bytes
recorded, front-clipped, both ends reporting success. Root cause found:
sendTextWithConfirmation accepted [Pasted text appearing anywhere in the
buffer as proof, and that chip renders no matter how little arrived.

My first fix compared the chip's +N lines against the payload. That fix was
wrong, and the verification the brief asked for is what caught it.
Sending
6,892 bytes into an idle tab and then asking the receiving session what it got:

ACK front=alpha-7391 end=omega-5520

Front and tail both arrived — the payload delivered complete — while the
chip on screen read +10 lines. So the chip's count does not track the payload,
and failing a send on a shortfall fails healthy sends. I could not reproduce
front-truncation at 6.9 KB at all.

What that leaves is a real conclusion: for a collapsed paste the screen cannot
establish completeness.
So send now separates three claims that used to be
one check-mark line —

  • nothing arrived → hard failure, and the body is not submitted, because
    pressing Enter on a fragment sends something that reads as a complete message.
    Text is left in the input box; exit 1.
  • arrived, completeness unverified → a warning naming the two remedies.
  • verified → success.

— and --verify does the comparison item 4(b) actually asked for, against
ground truth: it reads the target's own transcript, which records the user
message as received, and checks the payload's front and tail fingerprints,
naming which end went missing.

✔ Sent to a92fc9bf and verified: the session received the whole payload
  (front and tail both present in its transcript, 5748 non-whitespace chars sent)

--path (item 4c) is first-class: it hands the tab a file path and lets the
receiving session read it, so only a path crosses the prompt line and there is
no truncation surface. Verified end to end.

--wait-for-prompt tested only the buffer's last non-empty line, and Claude
renders notices below its input line — a Restart to update banner defeated
it for the full timeout. It now reads the whole tail of the window.

Verified against the live 86-tab fleet, with the new build installed

  • transcript on a tab whose session lives under ~/.claude-enterprise —
    resolved it, reported backend=enterprise, printed the turn.
  • sessions --json: 82 found, 4 not-found. One of those, ltv-cac-ent,
    reports 99 transcripts in its directory — previously an indistinguishable
    session_id: null.
  • sort --first applied for real (two tabs to the front, rest kept relative
    order), then the original 86-tab order restored exactly.
  • send at 6,892 bytes into a confirmed-idle tab, with --verify; --path
    handoff; the nonexistent---path and contradictory-source refusals.
  • Exit codes checked directly, not inferred from output.

On the two local commits

The brief said to preserve feat/whoami's two local-only commits. git ls-remote origin feat/whoami is indeed empty, so they are unpushed as commits — but
their content is already on main: f9d8331 landed squash-merged as
bb9a05c (#20), and the send-confirm work plus a later optionalAdapter
refinement shipped in v0.5.3. git diff origin/main feat/whoami shows the
branch removing work main has, so rebasing it would revert those refinements.
I branched fresh off origin/main and left feat/whoami untouched at
e352987.

Not done, and why

  • Item 7 (the trust-dialog fall-through) was outside the assigned set.
    Untouched.
  • npm run sync-plugin not run: it writes to ../plugins, and the repo's
    own prepack sync check is already failing on main (plugins repo at 0.5.1
    vs package.json 0.5.3) independently of this branch. That's a release step.
  • No version bump, no tag, no publish. The local global install was done
    from a packed tarball of the working tree, not a registry install.
  • @types/bun was missing from the stale local node_modules, so
    npm run typecheck failed on main too; npm ci fixed it. No source change
    involved.

npm run check: 322 pass, 0 fail, typecheck and build clean.

🤖 Generated with Claude Code

Fredrik Wollsén and others added 2 commits September 8, 2026 12:37
…ng sends landed

Six gaps measured while driving a 13-tab fleet, each one a place where cctabs
reported something it had not established.

- `cctabs transcript <tab> [n]` (alias `findings`) prints a tab's last n
  assistant messages from its transcript. `scrollback` returns the last painted
  frame, so a tab mid-turn shows a spinner — which meant sessions briefed each
  other from stale pictures. Searches every Claude config dir, because a tab on
  a backend preset writes under that preset's root and looking in one root
  reports "no transcript" for a healthy session.

- `cctabs sort --first a,b,c` pins a chosen set to the front. The plugin's
  /api/tabs/reorder already did exactly this; no verb exposed it. Refuses as a
  whole if a name doesn't resolve — half a working set in reach, with no
  indication which half, is worse than an error.

- `sessions --json` now carries `session_lookup`, so a null `session_id` says
  whether we looked and found nothing, couldn't look, or never looked. A lookup
  that threw was previously swallowed into the same null.

- `restore` verifies what it spawned: re-reads the tab list, checks each tab has
  a process, resolves its session, and reports N verified / N unconfirmed /
  N failed, exiting non-zero on failure. A tab that came back as a *different*
  session counts as failed — `--resume` on an id it can't find opens a fresh
  conversation, so the tab looks perfect and the context is gone. The line this
  replaces read "78 spawned, 0 failed" with one tab absent and one stripped of
  its context.

- `send` separates three claims that were one ✔ line: nothing arrived (hard
  failure, and the body is NOT submitted — a fragment reads as a whole message),
  arrived but completeness unverified, and verified. `--verify` does the real
  comparison against the target's transcript. `--path` hands over a file path
  instead of pasting, which has no truncation surface. `--submit` presses Enter
  only. `--wait-for-prompt` reads the whole tail of the buffer, so a "Restart to
  update" banner below a ready prompt no longer defeats it.

Worth recording what the measurements actually settled about that last one,
since the obvious fix is wrong: a paste chip's `+N lines` does NOT track the
payload. A 6,892-byte, 76-line payload was measured delivering *completely*
into an idle tab while its chip read `+10 lines`, so failing a send on a short
chip fails healthy sends. The screen cannot establish completeness for a
collapsed paste at all — hence "unverified" as a real answer, and hence
`--verify`, which reads what the session recorded receiving.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ported ✔

Two bugs reported against the new send, both measured. The serious one first.

A send whose text contained `--` delivered NOTHING and printed success. The
option parser silently drops any argv element containing a double dash —
reproduced on "mentions --verify here", "--leading", and even "a--b"; a single
dash survives. The text was absent from both the positionals and the values, so
`send` fell through to stdin, stdin was empty, and it printed

    ✔ Sent to 5f3e853e: ⏎

with an empty preview, which was the only tell. A ~900-byte bug report was lost
this way. The input is not exotic: any message quoting a flag name hits it, and
for a tool whose users are agents reporting tool bugs that is the normal case.

Positionals now come from process.argv directly (core/send-argv.ts) instead of
from the parser that loses them, and `--` works as an explicit terminator for
text that is entirely flag-shaped. An empty body is a hard failure now rather
than a ✔ — reporting success for a delivery of nothing is the same defect class
as the restore success line that could not fail — while a deliberate bare Enter
names itself, so an empty preview can never follow a ✔ again.

The second report was that `--path` still pastes file contents. It does not:
measured on a 2,470-byte file, the payload sent was the 257-char handoff. What
was really wrong is `--verify`. Claude records a tool's output as a
`role: "user"` message, and a `--path` handoff tells the tab to read the file —
so the newest user-role entry becomes the file's contents, and verify compared
the handoff against the file and reported "matches neither end of what was
sent". That reads exactly like a truncated paste, which is how it was reported.
Verify now skips tool results (the entry's `toolUseResult` field identifies
them) and searches every real message rather than only the newest.

Also fixed while in there: `--verify` resolved the session once, before its poll
loop, so a freshly spawned tab failed instantly with "no session resolved" a
second before its transcript existed. Lookup retries for the whole window now,
dropping the per-process title-index cache each round.

Confirmed for the smaller note: a verify mismatch exits 1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Deploying cctabs with  Cloudflare Pages  Cloudflare Pages

Latest commit: 42fd4da
Status:⚡️  Build in progress...

View logs

@motin

motin commented Sep 8, 2026

Copy link
Copy Markdown
Contributor Author

Both reports fixed and pushed (8d42199)

Thank you for the empty-preview observation — that detail is what made BUG 2
findable in one probe rather than by bisecting a send.

BUG 2 — confirmed, and your hypothesis was right about the cause

The mechanism is slightly broader than option-consumption: the parser drops
any argv element containing --, and the token is lost from positionals
and values alike. Probed directly:

argument result
"hello world" kept
"hello -x" kept
"mentions --verify here" dropped
"--leading" dropped
"a--b" dropped
"trailing --" dropped

So it isn't that a flag name gets consumed as an option — nothing is set in
values either. A single dash survives; -- anywhere does not. Your text
contained --file, --path and --verify, so the whole argument vanished,
send fell through to reading stdin, stdin was empty, and the empty preview
followed. Exactly the two symptoms you predicted, and exactly the diagnosis you
proposed as the fix.

Fixed as you suggested: positionals now come from process.argv directly
(core/send-argv.ts) rather than from the parser that loses them, and -- is
supported as an explicit terminator for text that is entirely flag-shaped.
Both forms verified end to end against a live tab:

$ cctabs send <tab> "BUG 1 — the --path option does not hand over a path... --file ... --verify" --verify
✔ Sent and verified: the session received the whole payload (109 non-whitespace chars sent)

$ cctabs send <tab> -- --verify is broken and --path too
✔ Sent to cbfb77e9: "--verify is broken and --path too" ⏎

16 test cases cover it, including a--b and a tab name containing --.

And the class of failure, not just the instance: an empty body is now a hard
error (exit 1) instead of a ✔, because reporting success for a delivery of
nothing is the same defect as the restore success line that could not fail. A
deliberate bare Enter reports Submitted Enter only (no body), so an empty
preview can never follow a ✔ again — your tell is now unambiguous by
construction.

BUG 1 — --path is not pasting; --verify was lying to you

I could not reproduce a paste. On a 2,470-byte fixture, the payload sent was the
257-char handoff, and the file was never read by send.

What actually went wrong is --verify. Claude records a tool's output as a
role: "user" message, and a --path handoff instructs the tab to read the
file
— so the newest user-role entry becomes the file's contents. Verify
compared the handoff against the file. Straight from the transcript:

user[str]         len=284   'Read the file at /private/tmp/.../bug1-fixture.txt...'
user[TOOL_RESULT] len=2469  'BUG REPORT (test fixture, 2421 bytes target)...'

That 2,469 is your 2,651. It is the file being read, which is --path working —
and comparing against it produced "matches neither end of what was sent", which
reads precisely like a front-clipped paste. Your inference was the only
reasonable one from the evidence the tool gave you; the tool gave you the wrong
evidence.

Verify now skips tool results (the entry's toolUseResult field identifies
them) and searches every real message instead of only the newest, so a payload
stays findable after the session has moved on. Both ends turning up in
different messages is still not a delivery.

I also fixed a second verify defect found while reproducing: it resolved the
session once, before its poll loop, so a freshly spawned tab failed
instantly with no session resolved for tab a second before its transcript
existed. Lookup now retries for the whole window.

--path's help text also overclaimed in the way that set this up — it implied
nothing of the file crosses into the session's record. It now says the receiving
session reads the file, so the contents appear in its transcript as a tool
result, and that this is the handoff working rather than a paste.

Smaller note — confirmed non-zero

$ cctabs send <tab> "payload" --verify --verify-timeout 0
 ERROR  ... delivery does NOT check out: no session ... appeared within 0s
$ echo $?
1

A verify mismatch exits 1. (Worth flagging that I made your exact mistake twice
in this session — cmd | tail reports the pipe's status, so I re-checked every
exit code without a pipe.)

npm run check: 346 pass, 0 fail. Reinstalled globally from the working
tree. Both scratch tabs closed; the fleet is back as it was.

The send mechanics now report faithfully whether text arrived. They say nothing
about whether it should have been sent, to that tab, at all — and that is where
the remaining mis-sends live: the wrong thing, to the wrong tab, that the tab
already knew.

Three gates, from a day of driving a ~15-tab fleet, each with its measurement:

- Resolve the owner from the BRANCH, not the tab's name. Verified against a
  live 92-tab fleet where all three layers disagree: a tab's name, the worktree
  directory it runs in, and the branch checked out there can each point at a
  different topic. Most pointedly, one topic's owning tab was named after
  something unrelated and no tab on the fleet was named after the topic at all —
  so a name-based router finds nothing and picks whatever sounds adjacent, which
  is how a tab owning a quarterly report was once sent pricing material from a
  different worktree.

- Count what the tab already knows before drafting. `transcript` shows what a
  tab concluded, not what it has seen, so grep the transcript for the specific
  phrases about to be relayed. One of six candidate tabs had nothing new and was
  dropped.

- Relay what was said, not your conclusions: a quoted statement is checkable
  where a paraphrased directive is not, and a statement that contradicts the
  tab's own conclusion is the highest-value relay there is.

Every example is written with synthetic branch names, tab names, figures and
company names — the shapes, ratios and counts are real, the identifiers are not.
The skill says so explicitly and tells a driver to do the same in anything
written out of a fleet, because tab and branch names read like infrastructure
while describing customer work, and this file ships to npm and to the
marketplace.

Two corrections to the guidance as received, both checked against the CLI rather
than assumed:

- `cctabs sessions` has no `--all` flag. Unknown flags are silently ignored, so
  `--all` looks like it worked while doing nothing; the skill says `--json`.
- The phrase count should resolve the transcript path via
  `transcript --json | jq -r .transcript`, not a `~/.claude*/projects/*` glob —
  the glob picks the wrong file as soon as the tab runs under a backend preset,
  which is exactly the "reads as dead" failure the command was built to avoid.

And one caveat the guidance did not claim, measured because the gate is only
useful if its range is known: branch resolution answers for 22 of 92 tabs. 43
sit on `main` in the repo root, where the branch carries no ownership signal at
all, and there the phrase count is the fallback rather than an invented mapping.

Also records what the new refusals mean for a driver, since both fired correctly
on the fleet and neither is a case for --force: "nothing from the text appeared
in the tab" identifies a tab stuck on a rendered menu — unreachable, needs a
human, and NOT a --path retry, since the handoff goes through the same prompt
line and is swallowed the same way.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@motin
motin force-pushed the feat/fleet-driving branch from acc34c4 to 1904558 Compare September 10, 2026 08:16
motin and others added 2 commits September 16, 2026 23:36
… it couldn't (#22)

Measured 2026-09-16: seven tabs spawned, five as `cctabs new <name>
<dir> --worktree --file <brief>`. All five landed on Claude Code's "Is
this a project you trust?" dialog and **none** received its brief. One
brief was worse than lost — it reached the *shell* instead and started
executing, npm-downloading `playwright` and `aws-cdk-lib` before that
session dropped to a bare prompt.

## Two statements in the skill were wrong for an untrusted directory

- The **✅ RIGHT** worktree example passed `--prompt` to a directory that
cannot accept one.
- *"polls internally until Claude's `❯` prompt appears before sending —
**no race condition**"* is false when a dialog is what's on screen: the
poll never sees `❯`.

Both fixed — the example keeps `--prompt` (it's correct for a trusted
repo, and is the ergonomic path) and gains its precondition; the
guarantee is scoped to the *startup* race.

## Three facts the file was missing

1. The selection marker defaults to **`No, exit`** — a bare Enter EXITS
the session.
2. The working keystroke is Down-then-Enter, and because `send` appends
its own Enter that is **one** call: `printf '\033[B' | cctabs send
<tab>`.
3. The precondition is **not** "is it a worktree".

## The precondition, and what was actually proved

A blanket *"never `--prompt` with `--worktree`"* would have been wrong
and annoying. Read out of Claude Code 2.1.273's own resolver and checked
against the fleet:

Trust lives in `${CLAUDE_CONFIG_DIR:-$HOME}/.claude.json` under
`projects[path].hasTrustDialogAccepted`. The lookup walks *up* from the
tab's directory and takes the first ancestor marked `true` — **but the
walk stops at the enclosing git repo root**, so trust never leaks in
from above it. For a worktree, that root resolves through `gitdir:` →
`commondir` to the **main repository**.

Three consistent observations:

- The tab this was written in
(`…/cctabs/.claude/worktrees/cctabs-trust`) got **no dialog at all** —
its repo root resolves to `…/generativereality/cctabs`, which is both an
ancestor and trusted.
- `~/Dev` was marked trusted on 2026-08-28 and still did **not** trust
the repos beneath it on 09-16 — refutes unbounded ancestor inheritance
outright.
- Spot check of the documented one-liner: a never-visited repo → `None`
(gated), one accepted after the incident → `True`.

⇒ The rule written: **"this path's repo root is already trusted, in the
config dir this tab will use."** The config-dir clause matters — backend
presets set `env_CLAUDE_CONFIG_DIR`, and the two config files on this
machine carry genuinely different trust sets.

⚠️ **Not black-box confirmed.** The decisive live probe (launching
`claude` in a scratch dir) was refused by the auto-mode classifier as
self-modification, so this is *read-from-resolver plus three consistent
field observations*, and version-pinned to 2.1.273. The documented
pre-spawn check reads the same key the resolver reads, so it holds
either way. Recorded in `HANDOVER.md`.

## Also

- **Existing stuck-tab guidance cross-referenced, not duplicated** — but
its *"a human unblocks it"* is narrowed: the trust dialog is two options
with a known starting position, so a driver may clear it. The resume
picker is untouched, where a wrong keypress silently accepts a summary.
- **The meta-bug.** The driver was reading the cached skill at
`<config-dir>/plugins/cache/generativereality/cctabs/0.5.0/…` — **691
lines, zero occurrences of "trust"** — while the CLI on PATH was `0.5.3`
and the source was 1030 lines and already covered the stuck-tab case.
The file warned that the CLI can lag the docs but not that the docs can
lag the CLI; they're separate channels and drift both ways. Added that
check.

## Release note

`SKILL.md` ships in **both** the npm tarball and the marketplace, so
this needs `npm run sync-plugin` at release time to actually reach the
driver that hit this.

**No version bump** — matches the previous doc-only commit (`1904558`),
and `CLAUDE.md` puts the bump in the release flow. Two `## Unreleased`
bullets added instead.

`HANDOVER.md` records what was proved, what was not, and a CLI proposal
left **unimplemented**: have `cctabs new` *refuse* `--prompt` into an
untrusted repo rather than auto-answer the dialog — refusing, because
answering a trust prompt on the user's behalf is the one decision that
dialog exists to ask.

Tests: 346 pass, 0 fail.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Fredrik Wollsén <fredrik.wollsen@f-secure.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
# Conflicts:
#	CHANGELOG.md
@motin
motin marked this pull request as ready for review September 17, 2026 06:08
@motin
motin merged commit 4aa94e6 into main Sep 17, 2026
1 check was pending
@motin
motin deleted the feat/fleet-driving branch September 17, 2026 06:08
motin added a commit that referenced this pull request Sep 17, 2026
…d Windows

Three merged PRs: #16 (the skill description was being truncated past its key rule, plus a real cctab alias), #21 (cctabs transcript, sort --first, and a send that establishes delivery instead of asserting it), #23 (Windows - the spawn, the pid walk and the plugin-install path, each of which reported success while failing).

#23 shipped without touching CHANGELOG.md, so its two entries were written for this release rather than left undocumented.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant