Skip to content

fix: recover interrupted model response streams - #39

Merged
minixalpha merged 2 commits into
mainfrom
fix/stream-recovery
Sep 25, 2026
Merged

minixalpha merged 2 commits into
mainfrom
fix/stream-recovery

Conversation

@minixalpha

@minixalpha minixalpha commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Summary

Recover interrupted model response streams instead of ending the run.

The SDK retries failures that happen before a response stream opens, but an
interruption during response-body iteration (ReadError, ReadTimeout, or
RemoteProtocolError) previously propagated out of the agent loop. SDK 1.5.0
uses httpx2, while the CLI only caught httpx.HTTPError; the two exception
families share no base class, so a mid-stream failure ended the session with no
retry.

Changes

  • Bounded stream recovery (agent.py, transport.py): retry response-body
    read errors, read timeouts, and remote protocol errors at most twice, waiting
    1s and 2s. Retries reuse the last committed conversation, never execute tools
    from an interrupted attempt, and do not consume the turn budget. A 300-second
    window starting at the first interruption bounds new retry scheduling; an
    in-flight request keeps the SDK timeout and external run deadline.
  • Transport compatibility: catch both httpx and httpx2 exception
    families through a small transport module, fixing the CLI's SDK 1.5.0 gap.
  • Journal v3 (event_journal.py, event-journal-protocol-v3.md): record
    each attempt as a model.failed event with the original error, retry intent,
    duration, and the generation ID captured before reading the body. Readers
    still accept v1/v2; v1/v2 records containing model.failed are rejected.
  • ATIF projection and cost reconciliation (atif.py, agent.py): project
    failed attempts as chronological incomplete steps, reconcile billing for
    interrupted generations, and mark token and cost totals partial when usage is
    missing.
  • Docs: bilingual CLI reference section on interrupted responses, the v3
    protocol source plus English version, and the full 20-task live retest in the
    0.8.x dev notes (zh + en).
  • CI: run the suite against Anthropic SDK 1.5.0 / httpx2 in addition to the
    locked dependencies.

Verification

Validation Result
Complete suite, locked SDK / httpx 292 passed, 4 skipped
Complete suite, SDK 1.5.0 / httpx2 2.12.0 296 passed
Harbor adapter / ATIF compatibility tests 24 passed
Historical interrupted-response replay, SDK 1.5.0 4/4 recovered
tests/test_stream_recovery.py (run during the live experiment) 17 passed

Offline replay consumes the original captured response bytes through SDK 1.5.0,
raises the recorded transport error, and supplies a synthetic successful retry.
These replays make no external requests and add no task score.

Live experiment (20 tasks)

A full 20-task Terminal-Bench 2.1 run on this branch finished 20/20 trials:
11 reward 1, 6 reward 0, 3 unscored. All 576 model calls returned completely;
model.failed and interrupted HTTP response bodies were 0, so the recovery
branch was not exercised by real traffic. The run can only say "no stream
interruption occurred this run", not "live recovery succeeded"; the mechanism is
covered by the automated tests. Details are in docs/dev_notes/{zh-CN,en}/0.8.x.md
and the Git-ignored job record jobs/tb21-stream-recovery-pilot20-20260921-record/.

Notes

  • No docs/research/ changes.
  • No bilingual README changes, so no README sync is needed.
  • Changelog: Fixed entry added under [Unreleased].

Retry response-body read errors, read timeouts, and remote protocol
errors at most twice (1s, 2s) before committing a reply to history or
running any of its tools. Handle both httpx and httpx2 transport
families; the CLI previously only caught httpx.HTTPError, so SDK 1.5.0
failures on httpx2 ended the run.

Retries reuse the last complete conversation, never execute tools from
an interrupted attempt, and do not consume the turn budget. A 300-second
window starting at the first interruption bounds new retry scheduling,
while in-flight requests keep the SDK timeout and external deadline.

Record each attempt as a journal v3 `model.failed` event and project it
as an incomplete ATIF step. Reconcile generation costs for interrupted
attempts and flag token and cost totals as partial when usage is
missing.
Add the full 20-task live retest to the 0.8.x dev notes, with its changelog entry and the regenerated English dev notes.
@minixalpha minixalpha changed the title fix/stream recovery fix: recover interrupted model response streams Sep 25, 2026
@minixalpha
minixalpha merged commit 1386df4 into main Sep 25, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant