fix(chat): keep interrupted turns in history and hand waits over with a message - #73
Merged
Merged
Conversation
Since the abort is rethrown (#70), a turn cut short by the person speaking during a wait, or by Stop, left before its messages reached history: the next request carried two user messages and nothing between, and the assistant denied work it had done and redid it. Track the messages of finished steps and commit them before the next turn pushes its own, with the still-open step recorded and its pending calls answered as interrupted. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The wait tool's only argument was described as shown to the user, and models took the waiting chip for their reply: judged runs handed buttons over with no text, and the person had to ask whether they could press. The tool now requires the message itself, shown as the step's text when the step wrote none, and kept in history as the call's argument. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On a dataset page a data table and a rich-text toolbar filled the head of the capped outline, the metadata fields below landed in the cut, and the persona told the assistant a field it had just filled did not exist. Rows now keep their name without repeating their cells, tables list 20 rows and count the rest, toolbars fold to one line, and unnamed images, link targets and dividers are dropped. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
sendMessage returned once Send was clicked. In a judged run three persona messages stayed in the composer while a wait was armed, no request followed, and the run still counted nine turns and was judged. The chat empties its composer when it takes a message, so a send now waits for that and throws otherwise, voiding the run instead. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Each persona turn is a fresh query remembering only the chat, so it could not tell its own actions from the assistant's instructions: judged runs had it report reloading a page it never reloaded. The prompt now lists its clicks and typing so far and says nothing else was done. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…t stopped onStepFinish runs on the SDK's side of the stream and can report a step finished before the loop has read its parts, so the open step could hold calls already in the finished messages and put the same call id in history twice. Reset the open step on the finish-step part the loop reads and drop calls already finished. Results that arrived in the open step go through their tool's toModelOutput as the SDK would, and Stop gets its own interruption text. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ady says it
A step that wrote a few words ("Voilà.") before declaring the wait hid its
message, and the handover went unsaid again.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The persona asks a question before pressing the button it was handed, so each run interrupts a declared wait and checks that the assistant answers from what it already did instead of rebuilding it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… wait An interrupted wait's result said only that the person had written. A judged run answered the question and never waited again, so it never learned the list it had prepared was created. The result now names what the wait was for and asks to declare it again if it is still to come. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The runner ignored whether a turn ended or was holding a declared wait, so a persona that clicked the handed-over button and stopped ended the run before the assistant reacted. Port data-fair's hand-over handling, let the persona read the resumed reply before stopping, and have speak-during-wait expect to be told of the creation unprompted. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The cap cut at a fixed character count, and in a judged run it fell inside a filled textbox: the persona read the start of the field's value as all of it, told the assistant its terms were missing and never saved. Lines are now shown whole or not at all, the marker counts the lines left out, and a root can take its own budget. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes found by the data-fair judged simulation baseline, in the chat and in the simulation harness.
wait_for_user_actiontakes a requiredmessage, shown as the reply to the person (after the step's own text unless it already says it); models used to read the waiting chip as their message and hand buttons over in silence.speak-during-waitcase.Why: in data-fair's baseline three of six cases failed on these defects; with them fixed, all seven data-fair cases pass (Sonnet, one pass).
Heads-up:
messageis now required onwait_for_user_action: hosts that tell the model to "tell the user, then wait" should point at the message instead (data-fair does so in its companion change).toolResultOutputmirrors the SDK's unexportedcreateToolModelOutput; recheck onaiupgrades.