Skip to content

Improve session context reliability and workspace workflows - #109

Merged
lukebuehler merged 28 commits into
mainfrom
qol1
Oct 3, 2026
Merged

lukebuehler merged 28 commits into
mainfrom
qol1

Conversation

@lukebuehler

@lukebuehler lukebuehler commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Improves long-running session reliability and everyday workspace and chat workflows. Sessions gain context repair and bounded context-limit recovery, while workspace files become easier to transfer, preview, and reference in replies.

  • Add context-entry replacement, compaction defaults and recovery, stable context catalogs, provider-safe media budgets, and clearer provider rejection reporting across the engine, runtime, API, and clients.
  • Add workspace uploads, downloads, file and folder operations, PDF/blob previews, and persistent VFS file references in transcripts.
  • Improve dictation submission, copy replies as Markdown or formatted text, universe appearance, environment activation, and runtime diagnostics.
  • Update generated contracts and provider-request fixtures, demo behavior, regression coverage (including delayed workspace file loading), and design documentation; include the Platform universe-appearance migration.

Validation passed locally:

  • cargo test --workspace --locked --no-fail-fast
  • cargo clippy --workspace --all-targets --locked -- -D warnings
  • cargo fmt --all -- --check
  • scripts/release/verify-metadata.sh
  • npm run check (generated-file checks, typechecking, consumer tests, and production/demo builds)

The npm gate emitted non-failing React test and Rollup annotation/chunk-size warnings. Live and credentialed suites were not run.

lukebuehler and others added 26 commits October 1, 2026 16:07
Reasoning effort none sent thinking.type disabled, which Claude Opus 5.5,
Sonnet 5.5, and the Fable and Mythos lines reject with a 400, so every
turn of such a session failed. It now lowers to between_tools on Sonnet
5.5 and to adaptive thinking at low effort where thinking cannot be
turned off; older models keep disabled. Display and block_binding are
only defaulted when thinking is on, since thinking-off types reject both.

Compaction requests replay earlier thinking but sent no thinking config,
so they carried no block_binding and failed on enforced accounts after an
image edit. On models that think by default they now send an explicit
adaptive config carrying the binding policy.

Live-verified on claude-opus-5-5 and claude-sonnet-5-5, and through the
hosted sessions and runs suites with the strict binding policy.

Documents LIGHTSPEED_ANTHROPIC_THINKING_PREFIX_MISMATCH and records slice
1 progress in the P186 roadmap doc.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A provider refusing a request (an invalid or oversized request,
classified InvalidRequest or ContextLength) failed the run as a generic
model_failure whose message wrapped the provider's text in runtime
prefixes. The classification now crosses the I/O boundary:

- CoreAgentIoError::Rejected carries the provider's message unchanged.
- The hosted activity and the test runner turn it into a Rejected
  generation whose failure blob is that message.
- The engine records TurnOutcome::Rejected and fails the run as
  RunFailureKind::RequestRejected. It only records the classification.
- The public RunFailureKindView gains request_rejected; the contract and
  TypeScript client are regenerated.
- The web transcript and CLI chat say the provider rejected the request.
- Each adapter logs, on a rejection, which context entries each provider
  message or input item holds, since provider errors cite positions.

Verified by unit and replay tests, an Anthropic runtime live test, and a
hosted live test where OpenAI rejects an undecodable image and the run
fails as request_rejected with OpenAI's message verbatim.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Run-appended context (tool results, run input, tool media) has no key,
so no context method could reach an entry the provider rejects, and a
keyed replace moves the entry to the tail. Every later run resent it.

session/context/replace { sessionId, entries: [{ entryId, item }] }
replaces active entries where they sit. It follows session/context/append:
items are InputItems converted by the same path, and results report
replaced, unchanged, absent, or failed with an admission failure per
entry. The engine command and event mirror ReplaceContextPrefix with ids
instead of a key prefix, reusing ContextEntryInput and ContextEntry:
ReplaceContextEntries / EntriesReplaced, projected as
contextEntriesReplaced.

A replacement keeps the entry's id, position, key, source, and kind, so a
tool result stays paired with its call; only tool results and user
messages qualify, a tool result takes only text, and replacement is
refused during a run. Adapters need no changes: a replaced image is plain
text everywhere.

The CLI gains `session context list` (from session/read), `replace`, and
a `redact` shortcut that sends standard "removed by operator"
placeholders. A hosted live test now continues past a provider rejection:
it replaces the undecodable image and runs the same session again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
After normalization every image meets the per-image limits, but a long
session still accumulates media until a request exceeds the provider's
image count or body size, and every later request fails the same way.

Each request now carries at most 100 media items and 24 MiB of encoded
media, the same budget for every provider and model. RequestMedia
prepares every image and PDF once, omits the oldest when the request is
over budget, and hands the prepared copies to the adapter, so the budget
adds no blob reads. Omitted media is sent as "[image · media:… · omitted
from this request to stay within provider limits]".

The omission count is rounded up to chunks of 10, so the cut point moves
once per chunk of new media instead of every turn, but never drops below
the four newest items that fit. It is a pure function of the context, so
retries send identical requests. All three adapters and compaction
requests share it.

Live-verified on Claude Opus 5.5: crossing the budget after a thinking
turn fails under the strict binding policy and continues under
drop_block.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lukebuehler
lukebuehler enabled auto-merge (squash) October 3, 2026 16:03
@lukebuehler
lukebuehler merged commit 07f1cf8 into main Oct 3, 2026
5 checks passed
@lukebuehler
lukebuehler deleted the qol1 branch October 4, 2026 11:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant