Skip to content

Run every model call under the AI response-timeout setting - #1092

Merged
MBombeck merged 7 commits into
mainfrom
fix/ai-timeout-everywhere
Oct 2, 2026
Merged

MBombeck merged 7 commits into
mainfrom
fix/ai-timeout-everywhere

Conversation

@MBombeck

@MBombeck MBombeck commented Oct 2, 2026

Copy link
Copy Markdown
Owner

Refs #1090.

Document reading, summaries, extraction (and with it the automatic lab staging), lab OCR, filing suggestions, medication extraction, the About-me questions, the document chat and the provider test called the model without the per-user response timeout and fell back to a hard 60 s in every client; the browser added its own 90–120 s aborts. A slow local model never had a chance.

  • The setting is bound to the provider instance: resolveProvider, resolveProviderChain and resolveProviderForTest read aiResponseTimeoutSeconds with the user lookup they already make and bind it on every client; each client takes its timeout only through callTimeoutMs() (user setting, else the surface's value, else 60 s). No call site passes or can forget it; workers use the record owner's setting. Upper bound stays 600 s.
  • Deliberate exceptions with a declared surface-ceiling: the Coach nudge (9 s) and the reaction line (12 s), which run account by account in short windows and have a deterministic fallback text.
  • /api/auth/me publishes ai.provider.responseTimeoutMs (additive, OpenAPI regenerated); the browser aborts derive from it: per-route model calls × timeout + 30 s (read-with-AI 3 calls, lab OCR 2, summary and suggestions 1). The lab OCR text mode no longer runs into apiPost's 15 s default.
  • The settings placeholder said "Default (~120)"; it now shows the real default (60 s) and that the briefing allows 180 s, with both numbers taken from the server constants.
  • Guard ai-call-timeout-guard.test.ts: every safeFetch client resolves through callTimeoutMs, the resolvers bind the timeout, clients are only constructed in provider.ts (and the offline eval harness), and the surface-ceiling exceptions are frozen; four mutations turn it red.

Not in this change: a reverse proxy in front (Cloudflare cuts at 100 s) still ends one long request; running long document reads as a background job the page polls is planned as its own change.

MBombeck and others added 7 commits October 2, 2026 23:01
The response-timeout setting reached the briefing, the status notes and
the Coach, which threaded it by hand. Document reads, document
extraction, lab OCR and lab staging called the provider without a
timeout, so every client fell back to its own 60 s literal and a slow
self-hosted model was cut off however high the setting stood.

The setting now rides on the provider instance. resolveProvider,
resolveProviderChain and resolveProviderForTest stamp the record's
aiResponseTimeoutSeconds on everything they return, and each wire
client reads its ceiling through callTimeoutMs: the setting, else the
surface's own timeoutMs, else 60 s. A call site no longer has to pass
it, so a new one cannot forget it. The connection test runs under the
same ceiling a read does.

The Coach nudge tick and the reaction line keep their short ceilings
(timeoutPolicy: "surface-ceiling"): the tick runs accounts one after
another under a shared budget, the reaction line's claim lease is
shorter than the largest setting, and both have deterministic copy
that stands in.

/api/auth/me publishes the effective value as
ai.provider.responseTimeoutMs, so a client can size its own request
timeout from it instead of guessing.

Refs #1090
…lab OCR

The browser gave up on its own: 120 s for the document routes, 90 s for
a lab scan, and the 15 s default on the lab text path, while a slow
self-hosted model with a raised response timeout was still allowed to
answer. A five-page report read by a local model failed on the client
side with the server still working.

Each request now sizes its abort from ai.provider.responseTimeoutMs on
the account payload, times the model calls the route can make, plus
30 s: one call for suggest and summary, three for Read with AI (the
transcription, then the lab staging read and its corrective retry), two
for a lab scan (the read and its retry). With the setting unset that is
90 s, 210 s and 150 s.

Refs #1090
… setting

Pins the four ways the setting could be skipped again: a wire client
reading params.timeoutMs itself instead of callTimeoutMs, an exported
resolver in provider.ts returning a provider it did not bind, a client
constructed outside provider.ts (the offline evaluation harness is the
named exception), and a surface ceiling not listed with its reason.
Each matcher asserts a non-zero match count.

Refs #1090
The placeholder said "Default (~120)" while most surfaces fall back to
60 s and the briefing to 180 s. The placeholder and the help line now
interpolate both numbers from the constants the server applies
(PROVIDER_DEFAULT_TIMEOUT_MS, AI_BUDGETS.comprehensive.timeoutMs), in
all seven locales, and a test fails if a locale writes a number of its
own.

Refs #1090
@MBombeck
MBombeck merged commit 2cd8b28 into main Oct 2, 2026
25 checks passed
@MBombeck
MBombeck deleted the fix/ai-timeout-everywhere branch October 3, 2026 00:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant