Skip to content

fix(chat): keep interrupted turns in history and hand waits over with a message - #73

Merged
albanm merged 13 commits into
mainfrom
fix-agents-sims
Sep 30, 2026
Merged

albanm merged 13 commits into
mainfrom
fix-agents-sims

Conversation

@albanm

@albanm albanm commented Sep 30, 2026

Copy link
Copy Markdown
Member

Fixes found by the data-fair judged simulation baseline, in the chat and in the simulation harness.

  • Chat: a turn cut short by the person speaking during a wait, or by Stop, now keeps its finished steps in history; the step still open is recorded with an "interrupted" result naming what the wait was for and asking to wait again if still relevant. Since fix(chat): treat an aborted stream as an abort, not a finished turn #70 such turns were dropped entirely, so the assistant denied work it had done and redid it.
  • Chat: wait_for_user_action takes a required message, shown as the reply to the person (after the step's own text unless it already says it); models used to read the waiting chip as their message and hand buttons over in silence.
  • lib-sim: the persona's page outline is pruned (rows by name, tables capped, toolbars folded) and cut on line boundaries with an optional per-root budget; a send the chat never took voids the run; the persona is told what it actually did.
  • Runner: waits for the turn a person's action resumes, lets the persona read it before stopping; new speak-during-wait case.

Why: in data-fair's baseline three of six cases failed on these defects; with them fixed, all seven data-fair cases pass (Sonnet, one pass).

Heads-up:

  • message is now required on wait_for_user_action: hosts that tell the model to "tell the user, then wait" should point at the message instead (data-fair does so in its companion change).
  • After Stop, the stopped turn's steps now stay in history, answered as interrupted.
  • toolResultOutput mirrors the SDK's unexported createToolModelOutput; recheck on ai upgrades.
  • Simulation verdicts are not comparable with runs from before this change (fuller page view, unsent messages now void a run).

albanm and others added 13 commits September 30, 2026 12:04
Since the abort is rethrown (#70), a turn cut short by the person speaking
during a wait, or by Stop, left before its messages reached history: the
next request carried two user messages and nothing between, and the
assistant denied work it had done and redid it. Track the messages of
finished steps and commit them before the next turn pushes its own, with
the still-open step recorded and its pending calls answered as interrupted.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The wait tool's only argument was described as shown to the user, and
models took the waiting chip for their reply: judged runs handed buttons
over with no text, and the person had to ask whether they could press.
The tool now requires the message itself, shown as the step's text when
the step wrote none, and kept in history as the call's argument.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On a dataset page a data table and a rich-text toolbar filled the head of
the capped outline, the metadata fields below landed in the cut, and the
persona told the assistant a field it had just filled did not exist. Rows
now keep their name without repeating their cells, tables list 20 rows and
count the rest, toolbars fold to one line, and unnamed images, link targets
and dividers are dropped.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
sendMessage returned once Send was clicked. In a judged run three persona
messages stayed in the composer while a wait was armed, no request
followed, and the run still counted nine turns and was judged. The chat
empties its composer when it takes a message, so a send now waits for
that and throws otherwise, voiding the run instead.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Each persona turn is a fresh query remembering only the chat, so it could
not tell its own actions from the assistant's instructions: judged runs had
it report reloading a page it never reloaded. The prompt now lists its
clicks and typing so far and says nothing else was done.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…t stopped

onStepFinish runs on the SDK's side of the stream and can report a step
finished before the loop has read its parts, so the open step could hold
calls already in the finished messages and put the same call id in
history twice. Reset the open step on the finish-step part the loop reads
and drop calls already finished. Results that arrived in the open step go
through their tool's toModelOutput as the SDK would, and Stop gets its own
interruption text.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ady says it

A step that wrote a few words ("Voilà.") before declaring the wait hid its
message, and the handover went unsaid again.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The persona asks a question before pressing the button it was handed, so
each run interrupts a declared wait and checks that the assistant answers
from what it already did instead of rebuilding it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… wait

An interrupted wait's result said only that the person had written. A
judged run answered the question and never waited again, so it never
learned the list it had prepared was created. The result now names what
the wait was for and asks to declare it again if it is still to come.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The runner ignored whether a turn ended or was holding a declared wait, so
a persona that clicked the handed-over button and stopped ended the run
before the assistant reacted. Port data-fair's hand-over handling, let the
persona read the resumed reply before stopping, and have speak-during-wait
expect to be told of the creation unprompted.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The cap cut at a fixed character count, and in a judged run it fell inside
a filled textbox: the persona read the start of the field's value as all
of it, told the assistant its terms were missing and never saved. Lines
are now shown whole or not at all, the marker counts the lines left out,
and a root can take its own budget.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added the fix label Sep 30, 2026
@albanm
albanm merged commit a0f401f into main Sep 30, 2026
3 checks passed
@albanm
albanm deleted the fix-agents-sims branch September 30, 2026 15:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant