Skip to content

Anthropic adapter drops ping events, so long pre-output thinking on Opus 5.5 hits upstream_stall_timeout at 300 s #5707

Description

@daehwanahn

Client or integration

Other: Hermes Agent (Python OpenAI SDK) calling /v1/chat/completions on a local OpenCodex.

Area

Streaming

Summary

Streaming requests to claude-opus-5-5 with a large context and high reasoning effort fail with upstream_stall_timeout at almost exactly 300 s, before any visible output. The model is still thinking. On Opus 5.5 display defaults to "omitted", so no thinking text streams during that phase. The Anthropic adapter yields a heartbeat only for SSE comments and drops event: ping, so no liveness Anthropic sends in that phase reaches the bridge watchdog. The client retries at the same effort and hits the same wall each time, so the turn is lost.

Expected: a model that is still thinking inside its documented limits is not cut off by the idle watchdog, or the failure can be told apart from a hung upstream so a client knows that retrying at the same effort will fail again.

Reproduction

  1. Route anthropic/claude-opus-5-5 through the Anthropic provider (OAuth) with the default stallTimeoutSec (300).
  2. Send a streaming /v1/chat/completions request with about 190k–215k input tokens and reasoning_effort: "max" on a task that needs long reasoning.
  3. The response ends after 304–313 s with upstream stream ended early (upstream_stall_timeout). The usage.jsonl row shows status 502, failureStage: "protocol-prelude", adapterEvents: 1, semanticBytes: 0.

Observed on 2026-09-23 in usage.jsonl:

  • 10 requests failed this way, all at 304–313 s: 9 at effort max, 1 at xhigh. Every one had adapterEvents: 1 and semanticBytes: 0.
  • 13 requests on the same model succeeded after more than 240 s. They finished at 243–296 s with 22.5k–35.1k output tokens, about 115 tokens/s. Successful long turns cluster just under 300 s, so the window is acting as a cap on thinking time, not only catching dead streams.
  • The same conversation continues normally after dropping to xhigh.

Code path, checked on main at 4cb43cb (same in installed 2.63.0):

  • src/adapters/anthropic.ts:1224-1227 yields { type: "heartbeat" } only for SSE comment records.
  • The event switch (src/adapters/anthropic.ts:1249-1358) handles message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop and error. event: ping / {"type":"ping"} matches none of them and yields nothing.
  • src/bridge/sse.ts:811 resets stallTicks only on adapter events. src/bridge/sse.ts:1407-1425 emits response.incomplete with upstream_stall_timeout once stallTimeoutSec (src/stall-timeout.ts:8, 300) runs out.

Not verified: I have not captured the raw upstream SSE, so I cannot give the ping cadence during the silent phase. adapterEvents: 1 shows the adapter yielded exactly one event and then nothing it recognized for 300 s. I can attach an ocx debug provider capture if that helps.

Version

2.63.0 (code path unchanged on main at 4cb43cb)

Operating system

Windows 11

Provider and model

anthropic (OAuth) / claude-opus-5-5, reasoning_effort max (once at xhigh)

Logs or error output

# usage.jsonl, one failed row (trimmed)
{"status":502,"durationMs":304908,"model":"claude-opus-5-5","inboundProtocol":"chat","requestedEffort":"max","errorCode":"upstream_server_error","failureStage":"protocol-prelude","failureCause":"upstream-fault","upstreamError":"Upstream stalled: no data for the stall-timeout window (upstream_stall_timeout)","attempts":[{"deliverySummary":{"adapterEvents":1,"relayedEvents":9,"semanticBytes":0,"sideEffectEvents":0,"terminalEvents":1}}]}

# client side
openai.APIError: upstream stream ended early (upstream_stall_timeout)
# retried 3 times, each cut at ~300 s; the turn was abandoned after ~67 min

Suggested direction

Upstream documentation:

I am not asking for a larger default or a synthetic keepalive; I read the #2210 decision against both. Possible fixes in the spirit of the transport-level clocks from #2210:

  1. Map Anthropic ping to an adapter heartbeat, paired with a separate cap on heartbeat-only time (like the Cursor heartbeat-only clock) sized for the thinking phase instead of 90 s. max_tokens bounds thinking plus output, so the cap could come from the request's max_tokens or effort.
  2. At minimum, mark a pre-output stall as not retryable at the same effort. Right now it surfaces as a 502 upstream-fault, which invites retries that are certain to fail the same way.

Related: #2210 (same watchdog, Cursor, fixed at the transport), #2528 (keepalives on the Claude outbound side, the opposite direction), #3736 (600 s default proposal, not adopted).

Local workaround: stallTimeoutSec set to 900 in config; it takes effect on the next restart and has not been verified here yet.

Redacted configuration

{ "stallTimeoutSec": 300 }

Checks

  • I searched existing issues and documentation.
  • I removed secrets, tokens, account details, request credentials, and personal data.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    account-poolOAuth, credentials, Codex pool, quota, failover, plansbugSomething isn't workingstreamingSSE, WebSocket, terminal stream frames

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions