Improve session context reliability and workspace workflows - #109
Merged
Merged
Conversation
Reasoning effort none sent thinking.type disabled, which Claude Opus 5.5, Sonnet 5.5, and the Fable and Mythos lines reject with a 400, so every turn of such a session failed. It now lowers to between_tools on Sonnet 5.5 and to adaptive thinking at low effort where thinking cannot be turned off; older models keep disabled. Display and block_binding are only defaulted when thinking is on, since thinking-off types reject both. Compaction requests replay earlier thinking but sent no thinking config, so they carried no block_binding and failed on enforced accounts after an image edit. On models that think by default they now send an explicit adaptive config carrying the binding policy. Live-verified on claude-opus-5-5 and claude-sonnet-5-5, and through the hosted sessions and runs suites with the strict binding policy. Documents LIGHTSPEED_ANTHROPIC_THINKING_PREFIX_MISMATCH and records slice 1 progress in the P186 roadmap doc. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A provider refusing a request (an invalid or oversized request, classified InvalidRequest or ContextLength) failed the run as a generic model_failure whose message wrapped the provider's text in runtime prefixes. The classification now crosses the I/O boundary: - CoreAgentIoError::Rejected carries the provider's message unchanged. - The hosted activity and the test runner turn it into a Rejected generation whose failure blob is that message. - The engine records TurnOutcome::Rejected and fails the run as RunFailureKind::RequestRejected. It only records the classification. - The public RunFailureKindView gains request_rejected; the contract and TypeScript client are regenerated. - The web transcript and CLI chat say the provider rejected the request. - Each adapter logs, on a rejection, which context entries each provider message or input item holds, since provider errors cite positions. Verified by unit and replay tests, an Anthropic runtime live test, and a hosted live test where OpenAI rejects an undecodable image and the run fails as request_rejected with OpenAI's message verbatim. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Run-appended context (tool results, run input, tool media) has no key,
so no context method could reach an entry the provider rejects, and a
keyed replace moves the entry to the tail. Every later run resent it.
session/context/replace { sessionId, entries: [{ entryId, item }] }
replaces active entries where they sit. It follows session/context/append:
items are InputItems converted by the same path, and results report
replaced, unchanged, absent, or failed with an admission failure per
entry. The engine command and event mirror ReplaceContextPrefix with ids
instead of a key prefix, reusing ContextEntryInput and ContextEntry:
ReplaceContextEntries / EntriesReplaced, projected as
contextEntriesReplaced.
A replacement keeps the entry's id, position, key, source, and kind, so a
tool result stays paired with its call; only tool results and user
messages qualify, a tool result takes only text, and replacement is
refused during a run. Adapters need no changes: a replaced image is plain
text everywhere.
The CLI gains `session context list` (from session/read), `replace`, and
a `redact` shortcut that sends standard "removed by operator"
placeholders. A hosted live test now continues past a provider rejection:
it replaces the undecodable image and runs the same session again.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
After normalization every image meets the per-image limits, but a long session still accumulates media until a request exceeds the provider's image count or body size, and every later request fails the same way. Each request now carries at most 100 media items and 24 MiB of encoded media, the same budget for every provider and model. RequestMedia prepares every image and PDF once, omits the oldest when the request is over budget, and hands the prepared copies to the adapter, so the budget adds no blob reads. Omitted media is sent as "[image · media:… · omitted from this request to stay within provider limits]". The omission count is rounded up to chunks of 10, so the cut point moves once per chunk of new media instead of every turn, but never drops below the four newest items that fit. It is a pure function of the context, so retries send identical requests. All three adapters and compaction requests share it. Live-verified on Claude Opus 5.5: crossing the budget after a thinking turn fails under the strict binding policy and continues under drop_block. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lukebuehler
enabled auto-merge (squash)
October 3, 2026 16:03
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Improves long-running session reliability and everyday workspace and chat workflows. Sessions gain context repair and bounded context-limit recovery, while workspace files become easier to transfer, preview, and reference in replies.
Validation passed locally:
cargo test --workspace --locked --no-fail-fastcargo clippy --workspace --all-targets --locked -- -D warningscargo fmt --all -- --checkscripts/release/verify-metadata.shnpm run check(generated-file checks, typechecking, consumer tests, and production/demo builds)The npm gate emitted non-failing React test and Rollup annotation/chunk-size warnings. Live and credentialed suites were not run.