Run every model call under the AI response-timeout setting - #1092
Merged
Merged
Conversation
The response-timeout setting reached the briefing, the status notes and the Coach, which threaded it by hand. Document reads, document extraction, lab OCR and lab staging called the provider without a timeout, so every client fell back to its own 60 s literal and a slow self-hosted model was cut off however high the setting stood. The setting now rides on the provider instance. resolveProvider, resolveProviderChain and resolveProviderForTest stamp the record's aiResponseTimeoutSeconds on everything they return, and each wire client reads its ceiling through callTimeoutMs: the setting, else the surface's own timeoutMs, else 60 s. A call site no longer has to pass it, so a new one cannot forget it. The connection test runs under the same ceiling a read does. The Coach nudge tick and the reaction line keep their short ceilings (timeoutPolicy: "surface-ceiling"): the tick runs accounts one after another under a shared budget, the reaction line's claim lease is shorter than the largest setting, and both have deterministic copy that stands in. /api/auth/me publishes the effective value as ai.provider.responseTimeoutMs, so a client can size its own request timeout from it instead of guessing. Refs #1090
…lab OCR The browser gave up on its own: 120 s for the document routes, 90 s for a lab scan, and the 15 s default on the lab text path, while a slow self-hosted model with a raised response timeout was still allowed to answer. A five-page report read by a local model failed on the client side with the server still working. Each request now sizes its abort from ai.provider.responseTimeoutMs on the account payload, times the model calls the route can make, plus 30 s: one call for suggest and summary, three for Read with AI (the transcription, then the lab staging read and its corrective retry), two for a lab scan (the read and its retry). With the setting unset that is 90 s, 210 s and 150 s. Refs #1090
… setting Pins the four ways the setting could be skipped again: a wire client reading params.timeoutMs itself instead of callTimeoutMs, an exported resolver in provider.ts returning a provider it did not bind, a client constructed outside provider.ts (the offline evaluation harness is the named exception), and a surface ceiling not listed with its reason. Each matcher asserts a non-zero match count. Refs #1090
The placeholder said "Default (~120)" while most surfaces fall back to 60 s and the briefing to 180 s. The placeholder and the help line now interpolate both numbers from the constants the server applies (PROVIDER_DEFAULT_TIMEOUT_MS, AI_BUDGETS.comprehensive.timeoutMs), in all seven locales, and a test fails if a locale writes a number of its own. Refs #1090
…fix/ai-timeout-everywhere
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #1090.
Document reading, summaries, extraction (and with it the automatic lab staging), lab OCR, filing suggestions, medication extraction, the About-me questions, the document chat and the provider test called the model without the per-user response timeout and fell back to a hard 60 s in every client; the browser added its own 90–120 s aborts. A slow local model never had a chance.
resolveProvider,resolveProviderChainandresolveProviderForTestreadaiResponseTimeoutSecondswith the user lookup they already make and bind it on every client; each client takes its timeout only throughcallTimeoutMs()(user setting, else the surface's value, else 60 s). No call site passes or can forget it; workers use the record owner's setting. Upper bound stays 600 s.surface-ceiling: the Coach nudge (9 s) and the reaction line (12 s), which run account by account in short windows and have a deterministic fallback text./api/auth/mepublishesai.provider.responseTimeoutMs(additive, OpenAPI regenerated); the browser aborts derive from it: per-route model calls × timeout + 30 s (read-with-AI 3 calls, lab OCR 2, summary and suggestions 1). The lab OCR text mode no longer runs intoapiPost's 15 s default.ai-call-timeout-guard.test.ts: every safeFetch client resolves throughcallTimeoutMs, the resolvers bind the timeout, clients are only constructed inprovider.ts(and the offline eval harness), and the surface-ceiling exceptions are frozen; four mutations turn it red.Not in this change: a reverse proxy in front (Cloudflare cuts at 100 s) still ends one long request; running long document reads as a background job the page polls is planned as its own change.