Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -39,3 +39,5 @@ jobs:
# --frozen installs exactly what uv.lock pins and fails if the lock has
# drifted from pyproject.toml, so CI tests the declared dependencies.
run: uv run --frozen pytest
- name: Test the httpx2 SDK transport
run: uv run --with anthropic==1.5.0 pytest
139 changes: 139 additions & 0 deletions benchmarks/harbor/reports/stream-recovery-20260915.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,139 @@
# Interrupted response recovery: implementation and targeted validation

## Status

The transport compatibility fix, bounded stream recovery, offline tests, and
independent review are complete. Live targeted trials are prepared and await
explicit approval for OpenRouter data transfer and API charges. No live model
requests have been made by this experiment yet.

## Problem and implementation

The September 12 and September 14 experiments recorded four interrupted model
responses across three distinct tasks. A `RemoteProtocolError` during response
iteration ended the whole agent run. SDK 1.5.0 used `httpx2`, while the CLI
caught `httpx.HTTPError`; these exception families do not share that base class.

The change handles both HTTP families and retries response-body `ReadError`,
`ReadTimeout`, and `RemoteProtocolError` at most twice, with delays of 1 and 2
seconds. It retains the last complete conversation and executes tools only
after a complete reply is received. Interrupted replies do not consume the
completed-reply turn budget. Authentication, invalid requests, programming
errors, and user cancellation are not retried by this loop. Pre-stream errors
retain the SDK retry policy without a second outer retry layer.

No further retry is scheduled beyond 300 seconds after a reply's first
interruption. This scheduling window does not cancel an in-flight request;
the SDK timeout and external supervisor deadline remain independent limits.

Journal v3 adds `model.failed`, with a separate ID per attempt, the original
error, retry intent, duration, and generation ID captured before reading the
body. ATIF keeps failed attempts in chronological order. Interrupted-generation
billing can be reconciled without treating missing token usage as complete.
Readers continue to support v1/v2 journals.

## Automated and offline evidence

| Validation | Result |
| --- | --- |
| Complete suite, locked SDK 0.112.0 / httpx | 292 passed, 4 skipped |
| Complete suite, SDK 1.5.0 / httpx2 2.12.0 | 296 passed |
| Harbor adapter and official ATIF compatibility tests | 24 passed |
| Historical interrupted response replay, SDK 1.5.0 | 4/4 recovered |

The locked environment skips four tests that explicitly require the absent
`httpx2` package. The SDK 1.5.0 run exercises both exception families. Both
full suites retain the existing Pydantic warning in the malformed-tool-input
test. CI now runs both the locked dependencies and an SDK 1.5.0 environment.

The real SDK transport test interrupts a response after complete-looking tool
JSON arrives, verifies that the discarded tool is never executed, and checks
that an earlier append operation occurs only once. Further tests cover retry
exhaustion, the scheduling window, preserved conversation history, interrupts,
permanent errors, billing reconciliation, and journal version rejection.

Offline replay consumes the original captured response bytes through SDK
1.5.0, raises the recorded transport error at EOF, and supplies a synthetic
successful reply to the retry. All four cases closed the interrupted response,
retried the same request, preserved the failed attempt, and executed no tools:

| Historical trial | Captured bytes | Outcome |
| --- | ---: | --- |
| `torch-tensor-parallelism__EEWJenL` | 280,943 | Recovered |
| `schemelike-metacircular-eval__Xuw2yU4` | 33,440 | Recovered |
| `schemelike-metacircular-eval__CJJvMtU` | 1,057,796 | Recovered |
| `llm-inference-batching-scheduler__S95o8Jv` | 34,318 | Recovered |

These replays make no external requests and provide no new task score. They
verify response handling; the retry's success is deliberately synthetic.

## Independent review

A Codex agent named `stream-reviewer` ran in a sibling Herdr split using the
[review-agent skill](/home/minix/.codex/skills/.system/review-agent/SKILL.md).
It reviewed the complete uncommitted diff and new files, relevant call sites,
tests, retry boundaries, exception compatibility, journals, ATIF, and costs.
Its final result was **No findings**. It independently reported 76 passed /
4 skipped with locked dependencies and 80 passed with cached SDK 1.5.0.
The review split was closed before creating the benchmark split.

## Targeted live experiment

The selected tasks are every distinct task explicitly classified as
`RemoteProtocolError` in the committed historical result summaries:

| Task | Relevant prior attempt | Prior result |
| --- | --- | --- |
| `schemelike-metacircular-eval` | September 14, request 18 | Reward 0; 0/63 subcases; no `eval.scm` |
| `llm-inference-batching-scheduler` | September 14, request 3 | Reward 0; interrupted model stream |
| `torch-tensor-parallelism` | September 12, request 1 | No official score; interrupted stream and later verifier dependency timeout |

Prepared job: `tb21-stream-recovery-targeted3-20260915`.

- Model: `openrouter/deepseek/deepseek-v4-flash-0731`.
- Dataset: `terminal-bench/terminal-bench-2-1`, pinned to
`sha256:7d7bdc1cbedad549fc1140404bd4dc45e5fd0ea7c4186773687d177ad3a0699a`.
- Preserve each task reference and cached image digest from the baseline.
- 50 completed replies, 65,536 tokens per reply, concurrency 2, one attempt per
task, no automatic whole-task Harbor retries.
- 3,600 seconds per agent run, 120 seconds of finalization grace, 3,780 seconds
for the outer Harbor deadline; unchanged native verifier limits.
- SDK 1.5.0, httpx 0.28.1, httpx2 2.12.0, Pydantic 2.13.5.
- Reuse the cached uv bootstrap and passive HTTP recorder; periodic stack
dumping stays disabled. Live trials do not inject transport faults.
- Install a wheel verified against every current package source file. Preserve
the uncommitted source snapshot and file hashes so results can be compared
with the final commit without pretending the snapshot was already committed.

Local reproduction command from the repository root:

```bash
PYTHONPATH=src benchmarks/harbor/.venv/bin/python \
jobs/tb21-stream-recovery-targeted3-20260915-record/workflow/run.py
```

The runner refuses to overwrite an existing job, checks the wheel and image
digests, loads the existing endpoint credentials without writing them into the
report, and saves a source manifest, raw logs, HTTP records, journals,
trajectories, checkpoints, and verifier output under `jobs/`.

## Interpretation limits

The recovery fix cannot identify which upstream component closed the original
connections and does not establish that the model can solve these tasks. The
September 12 Scheme interpreter passed only 4/63 subcases before its stream
failed. Sampling, routing, cache state, and provider backends are uncontrolled.
This targeted experiment is not a full benchmark score or a controlled measure
of pass-rate improvement. A live run without an interruption does not exercise
the retry branch.

## Evidence locations

- [September 12 results](../results/tb21-generation-budget-65536-20260912.json)
- [September 14 results](../results/tb21-tool-input-recovery-20260914.json)
- [Journal v3 protocol](../../../docs/dev_docs/en/event-journal-protocol-v3.md)
- Local record: `jobs/tb21-stream-recovery-targeted3-20260915-record/`
- Local offline replay: `offline_replay.py` and `offline-replay.json` in that record

Raw task/model data remains in the Git-ignored local job directories. Bilingual
CLI documentation and the v2/v3 protocol sources and English versions are in sync.
29 changes: 29 additions & 0 deletions benchmarks/harbor/tests/test_atif_compatibility.py
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,35 @@ def test_projector_output_passes_harbor_atif_validator():
assert validator.validate(trajectory), validator.get_errors()


@pytest.mark.parametrize("resolved", [False, True])
def test_recovered_stream_v3_passes_harbor_atif_validator(tmp_path, resolved):
fixture = Path(__file__).parent / "fixtures" / "atif-journal-v1.jsonl"
with EventJournal.create("run-recovered", directory=tmp_path) as journal:
for entry in EventJournal.replay(fixture):
if entry.type == "model.started":
journal.append(NativeEvent("model.started", entry.payload | {
"model_call_id": "failed-attempt",
}))
journal.append(NativeEvent("model.failed", {
"model_call_id": "failed-attempt", "error_type": "RemoteProtocolError",
"message": "interrupted", "generation_id": "gen-interrupted",
"duration_ms": 10, "will_retry": True, "retry_delay_seconds": 1,
"source_timestamp": entry.payload["source_timestamp"],
}))
if entry.type == "run.completed" and resolved:
journal.append(NativeEvent("model.cost_resolved", {
"generation_id": "gen-interrupted", "amount": "0.02",
"currency": "USD", "source": "test",
"source_timestamp": entry.payload["source_timestamp"],
}))
journal.append(NativeEvent(entry.type, entry.payload))
trajectory = project_atif(EventJournal.replay(journal.path))
validator = TrajectoryValidator()
assert validator.validate(trajectory), validator.get_errors()
assert trajectory["steps"][1]["extra"]["incomplete"] is True
assert trajectory["final_metrics"]["extra"]["usage_complete"] is False


@pytest.mark.parametrize("content", [
[{"type": "text", "text": "Partial answer"}],
[{"type": "extension", "namespace": "anthropic", "source_type": "thinking",
Expand Down
6 changes: 6 additions & 0 deletions docs/changelogs/0.8.x.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,12 @@ All notable changes in the **0.8.x** release series are documented here.
completion. Preserve partial output, usage, and cost accounting; skip tools
from the truncated reply and keep subsequent interactive requests valid.
New journals use schema v2, with v1 replay still supported.
- Recover interrupted model response streams. Retry response-body read errors,
read timeouts, and remote protocol errors up to twice (1s, 2s) before
committing a reply to history or running any of its tools, handling both the
`httpx` and `httpx2` transport families. Record each attempt as a Journal v3
`model.failed` event, project it as an incomplete ATIF step, and mark token
and cost totals as partial when usage is missing.

### Changed
- Increase the default per-reply generation limit from 8192 to 32768 tokens.
Expand Down
6 changes: 6 additions & 0 deletions docs/dev_docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,3 +30,9 @@ They complement, rather than replace:
- Event Journal Protocol v1:
[English](en/event-journal-protocol-v1.md) |
[Chinese](zh-CN/event-journal-protocol-v1.md)
- Event Journal Protocol v2 (response truncation):
[English](en/event-journal-protocol-v2.md) |
[Chinese](zh-CN/event-journal-protocol-v2.md)
- Event Journal Protocol v3 (current; failed model attempts):
[English](en/event-journal-protocol-v3.md) |
[Chinese](zh-CN/event-journal-protocol-v3.md)
5 changes: 3 additions & 2 deletions docs/dev_docs/en/event-journal-protocol-v2.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,9 @@
> [`../zh-CN/event-journal-protocol-v2.md`](../zh-CN/event-journal-protocol-v2.md).
> Do not edit by hand.

v2 is implemented and is the internal Journal protocol used by the current
writer, with `schema_version = 2`. This document defines all changes relative
v2 is implemented with `schema_version = 2`. The current writer has moved to
[v3](event-journal-protocol-v3.md); this document preserves the v2 protocol.
This document defines all changes relative
to [v1](event-journal-protocol-v1.md). Envelope, event types, fields, validation,
ordering, persistence, and projection rules not listed here follow v1. Public
trajectories remain ATIF-v1.7.
Expand Down
59 changes: 59 additions & 0 deletions docs/dev_docs/en/event-journal-protocol-v3.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,59 @@
# Event Journal Implementation Protocol v3

> Generated from the Chinese source
> [`../zh-CN/event-journal-protocol-v3.md`](../zh-CN/event-journal-protocol-v3.md).
> Do not edit by hand.

The current writer emits `schema_version = 3`. Readers and the ATIF projector
continue to accept v1 and v2. Public trajectories remain ATIF-v1.7.
All contracts from [v2](event-journal-protocol-v2.md) still apply except for the
additional event and projection behavior described here.

## Failed model attempts

`model.failed` terminates an API attempt without claiming a complete message or
complete usage. It has these required payload fields:

| Field | Type | Meaning |
| --- | --- | --- |
| `model_call_id` | nonempty string | The corresponding `model.started` identifier. |
| `error_type` | nonempty string | Original exception class name. |
| `message` | string | Original exception message. |
| `generation_id` | nonempty string or null | Provider header captured when the stream opens. |
| `duration_ms` | nonnegative number | Attempt duration including stream consumption. |
| `will_retry` | boolean | Whether the agent scheduled another attempt. |
| `retry_delay_seconds` | nonnegative number | Scheduled delay; zero when no retry is scheduled. |
| `source_timestamp` | RFC 3339 UTC or null | Event time. |

The event is emitted for caught SDK and transport errors, including final
failures. Unexpected exceptions and user interrupts retain the previous
`model.started` / `run.failed` representation. Each retry gets a new
`model_call_id`; it is not another completed reply and does not consume the
turn budget. Retries retain the last committed conversation and never execute
tools from the interrupted attempt. `will_retry` records intent: cancellation
or a journal failure may prevent the next attempt from starting.

The recovery policy retries only response-body read errors, read timeouts,
and remote protocol errors, with delays of 1 and 2 seconds. No new retry is
scheduled after a 300-second window starting at the first interruption of the
reply. An in-flight request can outlast that scheduling window; SDK timeouts
and an external supervisor's deadline remain independent limits. Failures
before the stream opens retain SDK retries without another outer retry layer.

## Projection and costs

Failed attempts become chronological agent steps, containing any visible text
deltas, `llm_call_count = 1`, and `extra.incomplete = true`. Error and retry
details remain in `extra`. They contain no tool calls or observations because
no tools from that response were executed. A later successful retry may end
the run normally without making the earlier attempt complete.

Generation IDs are captured before reading the response body so finalization
can reconcile interrupted generations too. Resolved costs contribute to the
total; unresolved attempts keep cost totals explicitly partial. Missing token
usage remains unknown even when billing is resolved. Legacy incomplete model
starts still project as before.

v1/v2 records containing `model.failed` are rejected. Old readers reject v3
rather than silently dropping failed attempts. Historical journals are not
rewritten.
3 changes: 2 additions & 1 deletion docs/dev_docs/zh-CN/event-journal-protocol-v2.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,8 @@
> 本文件为**中文源文件**(source of truth);英文版
> [`../en/event-journal-protocol-v2.md`](../en/event-journal-protocol-v2.md) 由其生成。

v2 已实现,是当前 writer 使用的内部 Journal 协议,`schema_version = 2`。
v2 已实现,使用 `schema_version = 2`。当前 writer 已升级到
[v3](../en/event-journal-protocol-v3.md),本文保留 v2 的协议定义。
本文完整定义相对 [v1](event-journal-protocol-v1.md) 的变化;未列出的 envelope、
事件类型、字段、校验、排序、持久化与投影规则沿用 v1。公开 trajectory 仍为 ATIF-v1.7。

Expand Down
48 changes: 48 additions & 0 deletions docs/dev_docs/zh-CN/event-journal-protocol-v3.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
# Event Journal 实现协议 v3

> 本文件为中文源;英文版本
> [`../en/event-journal-protocol-v3.md`](../en/event-journal-protocol-v3.md) 由其生成。

当前 writer 写入 `schema_version = 3`。Reader 和 ATIF projector 继续兼容
v1、v2;公开轨迹仍为 ATIF-v1.7。除下述新增事件和投影行为外,
[v2](event-journal-protocol-v2.md) 的其他契约保持有效。

## 模型尝试失败

`model.failed` 结束一次 API 尝试,不宣称收到完整消息或完整用量。必填字段如下:

| 字段 | 类型 | 含义 |
| --- | --- | --- |
| `model_call_id` | 非空字符串 | 对应 `model.started` 的标识。 |
| `error_type` | 非空字符串 | 原始异常类名。 |
| `message` | 字符串 | 原始异常消息。 |
| `generation_id` | 非空字符串或 null | 响应流打开时取得的服务商响应头。 |
| `duration_ms` | 非负数 | 包含响应流读取的尝试耗时。 |
| `will_retry` | 布尔值 | 是否已安排另一次尝试。 |
| `retry_delay_seconds` | 非负数 | 安排的等待时长;不重试时为零。 |
| `source_timestamp` | RFC 3339 UTC 或 null | 事件时间。 |

捕获到 SDK 或传输错误时写入该事件,包括最终失败。意外异常和用户中断沿用
`model.started` / `run.failed` 的表示。每次重试分配新的 `model_call_id`;
重试不是已完成回复,不消耗轮数预算。重试保留最后一次已提交的对话,绝不执行
中断回复中的工具。`will_retry` 记录安排意图;取消或 Journal 写入失败仍可能
阻止下一次尝试启动。

恢复策略只重试响应体读取错误、读取超时和远端协议错误,等待时间分别为 1 秒和
2 秒。从本轮首次中断开始,超过 300 秒窗口后不再安排新重试。进行中的请求
可以超过这个调度窗口;SDK 超时和外部监督器的截止时间仍是独立限制。响应流
打开前的失败沿用 SDK 的重试,不再叠加外层重试。

## 投影与费用

失败尝试按顺序成为 agent step,保留可见文本增量,设置 `llm_call_count = 1`
和 `extra.incomplete = true`,错误及重试信息写入 `extra`。这些 step 不包含
工具调用或 observation,因为没有执行该回复的工具。后续重试成功可以使 run
正常结束,但不会使此前的失败尝试变为完整。

在读取响应体前取得 generation ID,使收尾阶段也能查询中断生成的费用。已确认
费用计入总额;存在未确认的尝试时,费用总额显式保持不完整。即使账单已经补全,
缺失的 token 用量仍然未知。旧版未完成的模型开始事件沿用原来的投影方式。

包含 `model.failed` 的 v1/v2 记录会被拒绝。旧 reader 拒绝 v3,避免静默丢弃
失败尝试。历史 Journal 不会被改写。
Loading
Loading