diff --git a/docs-site/src/content/docs/guides/codex-integration.md b/docs-site/src/content/docs/guides/codex-integration.md index 671c63c1940..ba4da1ac432 100644 --- a/docs-site/src/content/docs/guides/codex-integration.md +++ b/docs-site/src/content/docs/guides/codex-integration.md @@ -131,6 +131,14 @@ echo `default`, so request logs show the response tier as an observation with co `assumed`. For latency-sensitive work, compare observed first-output times across the providers you actually use rather than assuming any particular channel is faster. +For a gateway forwarding to a backend with the same metadata limitation, explicitly declare +[`responseTierAuthoritative: false`](/reference/configuration/providers/#response-service-tier-authority) +on that provider. The request still sends priority and records the raw echo, while actual Fast +scheduling remains unconfirmed. Undeclared gateways and the official API keep their existing +response-based interpretation. Updating OpenCodex alone does not add this declaration to existing +gateway entries: without it, an eligible priority request followed by a `default` echo still records +`response-declined`. + The proxy listens on port `10100` by default and serves `POST /v1/responses`, `POST /v1/responses/compact`, `POST /v1/images/generations`, `POST /v1/images/edits`, `GET /v1/models`, `GET /healthz`, and the `/api/*` management surface. diff --git a/docs-site/src/content/docs/reference/configuration/providers.md b/docs-site/src/content/docs/reference/configuration/providers.md index a440ec05c02..367dd54abbc 100644 --- a/docs-site/src/content/docs/reference/configuration/providers.md +++ b/docs-site/src/content/docs/reference/configuration/providers.md @@ -205,6 +205,7 @@ Providers can expose a built-in shorthand, such as `agy` for `google-antigravity | `allowEncryptedV2AgentTasks?` | `boolean` | Disabled by default. Trust a direct key-auth `openai-responses` provider to consume or relay opaque encrypted V2 sub-agent tasks unchanged. Eligible routes skip `agentTaskRecovery`; all other routes keep the existing recovery or fail-closed behavior. OpenCodex does not decrypt, translate, or recover tasks sent through this opt-in. | | `upstreamWebsocket?` | `boolean` | Opt-in upstream Responses WebSocket transport for `openai-responses` requests (default false). Honored only for the first-party `https://api.openai.com/v1` upstream; custom-provider endpoints always use bounded HTTP/SSE because Bun cannot enforce an inbound WebSocket message limit before allocating the complete message. For the canonical ChatGPT `openai` provider, omitting it keeps the upstream WebSocket on eligible turns, an explicit `false` sends streaming turns over HTTP/SSE, and provider management rejects `true`; with `false`, native mid-turn steering and injection are unavailable. The field is independent of the client-facing `websockets` setting and changes neither the endpoint nor the credential. Plain HTTP remains on SSE; non-Responses paths and `openai-chat` requests stay on HTTP. | | `supportsServiceTier?` | `boolean` | Tri-state canonical Fast capability fallback. `true` publishes Fast in the catalog, satisfies service-tier routing requirements, contributes a supported fingerprint, and lets fast mode inject the provider's canonical wire value on a compatible final adapter. `false` strips the field and never injects, and exact model declarations cannot reopen it. Absent leaves the provider unclassified: fast mode does not inject or normalize a canonical caller value, and caller values obey the final wire's forwarding permission (`chatServiceTier` on Chat; passthrough on Responses). The registry classifies canonical OpenAI (`true`), DeepSeek, and Volcengine Ark (`false`); set it explicitly only for custom gateways that genuinely support tiers. | +| `responseTierAuthoritative?` | `boolean` | Whether the response tier can confirm or deny Fast. Set `false` explicitly for routes with non-authoritative response metadata; omission preserves existing behavior. This does not enable Fast or change outbound parameters. See [Response service-tier authority](#response-service-tier-authority). | | `fastEnabled?` | `boolean` | Operator switch for the provider's Fast lane. `false` turns Fast off (no Fast toggle, no `--fast` row, no fast wire field) and overrides `supportsServiceTier`. `true` enables a lane the registry marks opt-in. Absent keeps the registry default: off for `anthropic` and `anthropic-apikey`, whose fast mode spends usage credits at 2x price, and unchanged for every other provider. The dashboard Models page shows an Off/On row for opt-in providers. | | `modelSupportsServiceTier?` | `Record` | Exact upstream model capability overrides. Exact `true` enables canonical Fast for that model; exact `false` narrows provider defaults. An explicit provider-level `supportsServiceTier: false` remains fail-closed and cannot be reopened. Exact `true` does not authorize foreign caller-tier forwarding on Chat. Undeclared models fall back to provider-wide behavior. Management `PATCH /api/providers` merges entries and accepts `null` to clear one. | | `chatServiceTier?` | `boolean` | Provider-wide Chat-wire opt-in for forwarding caller `service_tier` values. On a classified route it governs foreign values such as `flex`, not proxy-owned canonical Fast after capability validation; on an unclassified route it governs every caller value because no Fast capability has been validated. Exact model capability does not authorize foreign forwarding. Responses routes retain their capability-based caller forwarding behavior. | @@ -565,6 +566,44 @@ cleared or unresolved. The persisted catalog field is read by Codex for the curr model, which is why a valid configured selector is copied to each applicable entry. Provider-scoped selectors (above) are applied before this root fallback and win on routed rows. +### Response service-tier authority + +Set `providers..responseTierAuthoritative` to `false` in `config.json` when that +provider's response `service_tier` cannot establish whether Fast was granted. This is an +operator declaration about the entire provider route, independent of `supportsServiceTier` and +`fastWire`. It neither enables Fast nor changes the outgoing `service_tier`. + +For example, add this field to an existing gateway provider whose Codex backend has +non-authoritative response metadata: + +```json +{ "responseTierAuthoritative": false } +``` + +**Gateway support is opt-in.** Updating OpenCodex alone does not add this declaration to +existing providers. Without it, an eligible priority request followed by `service_tier: "default"` +retains the legacy `response-declined` interpretation. Configure `false` only when the +route's response metadata is known to be non-authoritative. + +For an eligible request serialized as `priority`, both a `default` and a `priority` echo remain +observations. Logs preserve `responseServiceTier` and record +`tierOutcome.responseTierAuthoritative: false`, `fastOutcome: "applied"`, and +`confirmation: "assumed"`, without `response-declined`. Here **applied means the request parameter +was sent**, and **assumed means the actual Fast effect is unconfirmed**. The model tooltip shows +the request and raw response separately with that confirmation. Neither this setting nor these +records prove an acceleration or a billed tier. + +Cost estimates use the existing requested-tier fallback instead of treating the raw echo as a +confirmed price tier; pricing rules requiring response confirmation cannot use that echo. Historical records +without an authority flag retain their previous interpretation. + +The field accepts only booleans. Omission and explicit `true` retain response-based confirmation +for other destinations, including the official API and undeclared gateways. Canonical +`https://chatgpt.com/backend-api/codex` with `authMode: "forward"` remains automatically +non-authoritative, even if `true` is configured. Gateway names and URLs are never inferred. +If one gateway mixes response contracts, use separate provider entries for those routes and +apply the declaration only to the relevant entry. + ### FastWire B1 capability migration Fast capability and arbitrary Chat caller-tier forwarding are independent after FastWire B1. The diff --git a/docs-site/src/content/docs/zh-cn/reference/configuration/providers.md b/docs-site/src/content/docs/zh-cn/reference/configuration/providers.md index a7464eab8e9..6960bb12da3 100644 --- a/docs-site/src/content/docs/zh-cn/reference/configuration/providers.md +++ b/docs-site/src/content/docs/zh-cn/reference/configuration/providers.md @@ -85,6 +85,7 @@ selector,而不是分配一个新名称。 | `chatCompletionsPath?` | `string` | 用于 `openai-chat` 请求的相对资源路径,是 `responsesPath` 的对应项,适用相同的路径规则。当同一上游以不同前缀提供 Chat Completions 和 Responses 时需要此配置:按模型的 wire override 只更换适配器而不改动 `baseUrl`,否则已启用的 Chat 请求会被发送到 Responses base。随附示例为 Z.AI。 | | `upstreamWebsocket?` | `boolean` | 为 `openai-responses` 请求选择性启用上游 Responses WebSocket 传输(默认 `false`)。仅对第一方 `https://api.openai.com/v1` 上游生效;自定义提供者端点始终使用有界 HTTP/SSE,因为 Bun 无法在分配完整消息之前对入站 WebSocket 消息实施大小限制。对于规范 ChatGPT `openai` 提供商,省略该字段会在符合条件的轮次使用上游 WebSocket,`false` 通过 HTTP/SSE 发送流式轮次,`true` 会被拒绝;设为 `false` 时,原生轮次中操控与注入不可用。该字段独立于客户端侧的 `websockets` 设置,且不改变端点或凭据。普通 HTTP 仍使用 SSE;非 Responses 路径和 `openai-chat` 请求仍使用 HTTP。 | | `supportsServiceTier?` | `boolean` | `service_tier` 能力的三态。`true`:fast 模式可以注入,调用方提供的值也会被保留。`false`:剥离该字段且绝不注入(已明确不支持的上游不会收到它)。未设置:未分类——调用方提供的值原样保留,fast 模式绝不注入。注册表已对官方 OpenAI(`true`)、DeepSeek 和 Volcengine Ark(`false`)分类;仅对真正支持分层的自定义网关显式设置。 | +| `responseTierAuthoritative?` | `boolean` | 响应等级是否足以确认或否定 Fast。对响应元数据不具权威性的链路显式设为 `false`;未设置时保持原有行为。不会启用 Fast 或修改出站参数。详见[响应服务等级的可信度](#响应服务等级的可信度)。 | | `preserveResponsesReasoningContent?` | `boolean` | 在重放的 Responses reasoning 项中保留明文 reasoning 内容,而不是清空(清空是 ChatGPT 后端的规则)。对接受 reasoning 重放的上游(如 DeepSeek)启用。代理生成的 `ocxr1` 信封始终会被剥离。 | | `disabled?` | `boolean` | 将提供者保留在磁盘上,但从路由和模型/目录列表中排除。 | | `apiKey?` | `string` | API key,或在请求时解析的 `${ENV_VAR}` / `$ENV_VAR` 引用。 | @@ -172,6 +173,37 @@ API key 提供者可以持有字面量 key,或环境引用。OAuth 提供者 `PATCH /api/providers?name=` 只修改它指定的字段,无论目的地如何都保留其他所有已存储字段。它接受全部五项设置,`null` 表示清除。对于两个推理列表,空数组会作为显式退出选项保存,而不会被删除。 +### 响应服务等级的可信度 + +当提供者响应中的 `service_tier` 无法确定 Fast 是否生效时,在 `config.json` 中将 +`providers..responseTierAuthoritative` 设为 `false`。这是运营者对整个提供者链路的声明, +独立于 `supportsServiceTier` 和 `fastWire`;它不会启用 Fast,也不会改变出站 `service_tier`。 + +例如,对于已知 Codex 后端响应元数据不具权威性的网关,在现有 provider 对象中加入: + +```json +{ "responseTierAuthoritative": false } +``` + +**网关需要显式配置。** 仅升级 OpenCodex 不会为现有 provider 自动添加该声明。未配置时, +符合 Fast 条件且已发送 priority 的请求收到 `service_tier: "default"`,仍会按原逻辑记录为 `response-declined`。 +仅对已知响应元数据无法确认实际等级的链路设置 `false`。 + +对于符合 Fast 条件且已发送 `priority` 的请求,`default` 和 `priority` 响应都只作为观察值。 +日志保留 `responseServiceTier`,并记录 `tierOutcome.responseTierAuthoritative: false`、 +`fastOutcome: "applied"` 和 `confirmation: "assumed"`,不会仅凭响应回显判为 `response-declined`。 +这里 **applied 表示请求参数已发出**,**assumed 表示 Fast 的实际效果尚未确认**。 +模型提示会分别显示请求等级、原始响应等级和确认状态。这些记录不能证明实际加速或计费等级。 + +费用估算沿用请求等级的回退逻辑,不将原始回显当成已确认的价格等级;要求响应确认的定价规则 +不能使用该回显作为证据。没有可信度标记的历史记录保留原解释。 + +该字段只接受布尔值。对于官方 API、未声明的网关等其他目的地,省略或显式设为 `true` +都保留基于响应的判定。标准 `https://chatgpt.com/backend-api/codex` 地址配合 +`authMode: "forward"` 始终自动按非权威处理,即使配置为 `true` 也不例外。 +OpenCodex 不会依据网关名称或 URL 猜测可信度。如果同一网关混合不同的响应契约, +应拆成独立 provider 条目,只为相关条目配置此声明。 + ## 提供者诊断出站安全性 仪表板连接测试和实时模型发现使用受限的、仅 GET 传输。没有出站代理时,opencodex 只会解析一次主机名,并仅连接到该已验证地址。HTTPS 仍会保留原始 Host、SNI 和证书验证;提供者配置不能关闭证书检查。 diff --git a/scripts/test-layout/layout.json b/scripts/test-layout/layout.json index 0cca11009a2..368c0ba55e6 100644 --- a/scripts/test-layout/layout.json +++ b/scripts/test-layout/layout.json @@ -966,6 +966,7 @@ "fastwire-characterization-routing.test.ts": "routing", "fastwire-characterization-wire.test.ts": "routing", "fastwire-observability.test.ts": "routing", + "fastwire-response-authority.test.ts": "routing", "fastwire-policy.test.ts": "routing", "featherless-provider.test.ts": "providers", "fetch-header-timeout.test.ts": "server", diff --git a/src/config/schema/leaf-validators.ts b/src/config/schema/leaf-validators.ts index 8de6060f5b2..9560a863d7b 100644 --- a/src/config/schema/leaf-validators.ts +++ b/src/config/schema/leaf-validators.ts @@ -297,6 +297,7 @@ export const providerConfigSchema = z.object({ annotateEmptyToolOutputs: z.boolean().optional(), foldDeveloperRoleToSystem: z.boolean().optional(), fastWire: fastWireSchema.nullable().optional(), + responseTierAuthoritative: z.boolean().optional(), fastEnabled: z.boolean().optional(), supportsServiceTier: z.boolean().optional(), modelSupportsServiceTier: z.record(z.string().min(1), z.boolean()).optional(), diff --git a/src/providers/fastwire.ts b/src/providers/fastwire.ts index 6064179ab99..d7c10488f04 100644 --- a/src/providers/fastwire.ts +++ b/src/providers/fastwire.ts @@ -340,6 +340,9 @@ export function createAdapterTierMetadata( const loggedWireValue = wireValue === null ? null : sanitizeLogMetadataString(wireValue); const outcome: AttemptTierOutcome = { wireKind, + ...(context.responseTierAuthoritative !== undefined + ? { responseTierAuthoritative: context.responseTierAuthoritative } + : {}), ...(wireValue === null ? { wireValue: null } : loggedWireValue ? { wireValue: loggedWireValue } : {}), diff --git a/src/providers/model-rename-fields.ts b/src/providers/model-rename-fields.ts index 171a20f7539..deda3f076a8 100644 --- a/src/providers/model-rename-fields.ts +++ b/src/providers/model-rename-fields.ts @@ -16,6 +16,7 @@ export const PROVIDER_MODEL_RENAME_ROLES = { mcpMaxResultBytes: "none", modelAdapters: "record", fastWire: "none", + responseTierAuthoritative: "none", baseUrl: "none", responsesPath: "none", chatCompletionsPath: "none", diff --git a/src/providers/openai-tiers-destination.ts b/src/providers/openai-tiers-destination.ts index 5c6124f29df..a4a8e3f7f97 100644 --- a/src/providers/openai-tiers-destination.ts +++ b/src/providers/openai-tiers-destination.ts @@ -25,6 +25,15 @@ export function isCanonicalOpenAiForwardProvider(provider: OcxProviderConfig): b && normalizedBaseUrl(provider.baseUrl) === CODEX_FORWARD_BASE_URL; } +/** + * Response evidence is separate from Fast capability and request serialization. Preserve the + * known Codex exception (#2558); other destinations may declare their response contract without + * a gateway-name or URL heuristic. Absence retains the authoritative legacy default. + */ +export function responseTierAuthorityForProvider(provider: OcxProviderConfig): boolean | undefined { + return isCanonicalOpenAiForwardProvider(provider) ? false : provider.responseTierAuthoritative; +} + const OPENAI_API_ORIGIN = "https://api.openai.com"; const OPENAI_API_BASE_URL = `${OPENAI_API_ORIGIN}/v1`; const OPENAI_API_RESPONSES_URL = `${OPENAI_API_BASE_URL}/responses`; diff --git a/src/server/auth-cors.ts b/src/server/auth-cors.ts index a4176df5834..3a527eb6aa4 100644 --- a/src/server/auth-cors.ts +++ b/src/server/auth-cors.ts @@ -973,6 +973,7 @@ const PROVIDER_CONFIG_FIELD_POLICY = { mcpMaxResultBytes: "editor", modelAdapters: "editor", fastWire: "editor", + responseTierAuthoritative: "editor", fastEnabled: "editor", baseUrl: "editor", responsesPath: "editor", diff --git a/src/server/responses/core-normalize.ts b/src/server/responses/core-normalize.ts index 86b0ae546ed..cea4798908f 100644 --- a/src/server/responses/core-normalize.ts +++ b/src/server/responses/core-normalize.ts @@ -17,6 +17,7 @@ import { resolveOpenCodeGoTransport } from "../../providers/opencode-go-transpor import { getOrAllocateRequestSessionLane } from "../request-log-conversation"; import { shouldPreparePlaintextV2AgentMessages } from "../../responses/plaintext-v2-agent-messages"; import { hasValidatedActiveReasoningEffort } from "../../responses/parser"; +import { responseTierAuthorityForProvider } from "../../providers/openai-tiers-destination"; import { isCanonicalOpenAiForwardProvider } from "../../providers/openai-tiers"; import { applyOpenAiVirtualModel } from "../../providers/openai-virtual-models"; import { renameRoutedIdentityInContext } from "../../adapters/identity"; @@ -220,14 +221,12 @@ export async function applyFinalRouteRequestNormalization(args: { ); const modelServiceTierSupport = serviceTierSupportFromPolicy(fastPolicy); const callerTier = parsed.options.serviceTier; - // The ChatGPT-internal Codex backend echoes `service_tier: "default"` even on turns it - // scheduled as priority, so its echo cannot confirm OR deny Fast. Believing it reported every - // Fast request as `response-declined` (#2558). The public API's echo stays authoritative. + // Capture destination evidence policy separately from the unchanged outbound Fast decision. parsed.options.tierObservation = tierObservationContext( fastPolicy, config.fastMode, callerTier, - isCanonicalOpenAiForwardProvider(route.provider) ? false : undefined, + responseTierAuthorityForProvider(route.provider), ); parsed.options.tierDecision = decideTier(fastPolicy, config.fastMode, callerTier); parsed.options.serviceTier = tierValueAfterDecision(parsed.options.tierDecision, callerTier); diff --git a/src/types/provider.ts b/src/types/provider.ts index 55ed3bf2c71..5b314eb8302 100644 --- a/src/types/provider.ts +++ b/src/types/provider.ts @@ -218,6 +218,8 @@ export interface AttemptTierOutcome { callerFastSuppressedByConfig?: boolean; confirmation: "confirmed" | "assumed" | "downgraded" | "unknown"; responseServiceTier?: string; + /** False retains the raw echo as evidence only, including during cost estimation. */ + responseTierAuthoritative?: boolean; } /** @@ -307,6 +309,13 @@ export interface OcxProviderConfig { * absence derives from the final model adapter. */ fastWire?: FastWire | null; + /** + * Whether echoed service_tier can confirm or deny Fast. Set false for a relay whose + * response metadata cannot establish the granted tier. Absence keeps legacy authority; + * canonical ChatGPT Codex forwarding always treats the echo as non-authoritative. + * Observation only: this does not enable Fast or change request serialization. + */ + responseTierAuthoritative?: boolean; baseUrl: string; /** * Optional relative resource path for key-auth openai-responses requests. Must start with `/` diff --git a/src/usage/cost.ts b/src/usage/cost.ts index 790ba444303..19695ec3cec 100644 --- a/src/usage/cost.ts +++ b/src/usage/cost.ts @@ -439,13 +439,16 @@ export function serviceTierContext(entry: ServiceTierContext): ServiceTierContex /** Convert one adapter-observed attempt outcome into the existing pricing provenance shape. */ export function serviceTierContextFromOutcome(outcome: AttemptTierOutcome): ServiceTierContext { - if (outcome.canonical === "priority" && outcome.confirmation === "confirmed") { + const responseAuthoritative = outcome.responseTierAuthoritative !== false; + if (responseAuthoritative && outcome.canonical === "priority" && outcome.confirmation === "confirmed") { return { responseServiceTier: "priority" }; } - if (outcome.responseServiceTier !== undefined) { + // Non-authoritative echoes remain in the log, but cannot become pricing confirmation either. + if (responseAuthoritative && outcome.responseServiceTier !== undefined) { return { responseServiceTier: outcome.responseServiceTier }; } - if (outcome.canonical === "priority" && outcome.confirmation === "assumed") { + if (outcome.canonical === "priority" + && (outcome.confirmation === "assumed" || (!responseAuthoritative && outcome.confirmation === "confirmed"))) { return { requestedServiceTier: "priority" }; } // An unclassified route makes no canonical Fast claim, but its adapter can still prove that diff --git a/src/usage/log.ts b/src/usage/log.ts index 0c3b37e5b89..65aec2769cd 100644 --- a/src/usage/log.ts +++ b/src/usage/log.ts @@ -656,6 +656,7 @@ function normalizeAttemptTierOutcome(raw: unknown): AttemptTierOutcome | null { if ("callerFastSuppressedByConfig" in outcome && typeof outcome.callerFastSuppressedByConfig !== "boolean") return null; if ("responseServiceTier" in outcome && typeof outcome.responseServiceTier !== "string") return null; + if ("responseTierAuthoritative" in outcome && typeof outcome.responseTierAuthoritative !== "boolean") return null; const wireValue = sanitizeLogMetadataString(outcome.wireValue); const responseServiceTier = sanitizeLogMetadataString(outcome.responseServiceTier); return { @@ -678,6 +679,9 @@ function normalizeAttemptTierOutcome(raw: unknown): AttemptTierOutcome | null { ? { callerFastSuppressedByConfig: outcome.callerFastSuppressedByConfig } : {}), confirmation: outcome.confirmation as AttemptTierOutcome["confirmation"], + ...(typeof outcome.responseTierAuthoritative === "boolean" + ? { responseTierAuthoritative: outcome.responseTierAuthoritative } + : {}), ...(responseServiceTier ? { responseServiceTier } : {}), }; } diff --git a/structure/config.md b/structure/config.md index 2182ce0b0a2..79deb53630f 100644 --- a/structure/config.md +++ b/structure/config.md @@ -382,7 +382,7 @@ report is additive. When opencodex owns routing, it also writes `$CODEX_HOME/opencodex.config.toml` as an explicit profile target. Codex config uses `service_tier = "fast"` and `[features].fast_mode = true`; -catalog/request tier metadata may use `priority`. Do not collapse these spellings into one value. +catalog/request tier metadata may use `priority`. Do not collapse these spellings into one value. Provider `responseTierAuthoritative` is an optional strict boolean validated by `src/config/schema/leaf-validators.ts`; it changes response evidence only, as defined in the [response-tier observation contract](transports/responses.md#response-tier-observation-authority). ## Provider output defaults diff --git a/structure/gui-and-management-api.md b/structure/gui-and-management-api.md index 56ea1763b01..a30ec1ff58b 100644 --- a/structure/gui-and-management-api.md +++ b/structure/gui-and-management-api.md @@ -228,7 +228,7 @@ per-request first-party callback reads that live object; a failed write leaves i | Subagents | Read/write the featured `subagentModels` list capped at five ids. `GET/PUT /api/injection-model` manages the shared delegation model/effort selection, the independent OpenCodex guidance switch, and the default-off `syncCodexSubagentDefaults` opt-in for native Codex subagent defaults. When OpenCodex owns the active Codex routing, native `[agents]` defaults apply to newly created Codex tasks after sync/restart; external user-managed provider configs remain untouched. The defaults do not cause delegation and preserve existing user-owned defaults rather than overwriting them. PUT is partial-update: absent keys are unchanged, `null` clears, and non-object bodies are rejected with 400 before field validation. `syncCodexSubagentDefaults: true` requires a nonblank `model` and a supported Codex reasoning effort when effort is set; clearing `model` (null/empty) always clears effort and disables native-default sync even when the stored effort was invalid. | | V2 / Multi-agent mode | `GET/PUT /api/v2` — reports/sets the codex `multi_agent_v2` feature flag, the 3-state `multiAgentMode` override (`v1`/`default`/`v2`), the `keepNativeChatGptOnV1` hybrid pin, and the logical maximum thread count. Selecting `v2` normally enables the native flag; with the hybrid pin it disables that global override so native rows can resolve to v1 while routed rows resolve to v2. Selecting `v1` disables the flag; `default` leaves it unchanged. PUT rejects an explicit enabled flag that conflicts with the selected mode or hybrid pin. Every transition preserves the logical thread limit, is rollback-safe, and resyncs the catalog. GET and successful PUT also return stored `multiAgentModeHintText` plus response-only `multiAgentModeHintRecommendation: { text, revision }`; the recommendation is not a writable or persisted config field. Both also return response-only `multiAgentSurfaceAdvisory: { required, mode, recommended, version, docsUrl }`, true while the resolved mode is not v1 and the stored acknowledgement version is behind; PUT accepts `multiAgentSurfaceAdvisoryAcknowledged`, where only `true` stores the current version and `false` is an explicit no-op, and it composes with a `multiAgentMode` write in the same body so the dialog's recommended answer is one request. | | Logs & Debug | One sidebar entry (`/#logs`) with two tabs. Logs tab: request/runtime logs for local diagnosis. `LogsFilterBar` owns controls over the shared `LogFilterState`; `filterLogs` composes filters over the loaded ring. The logs envelope adds `generatedAt` (proxy epoch milliseconds); the page advances that sample with monotonic elapsed time and retains a browser-clock fallback for older proxies. Reset returns focus to the stable All surface radio. Provider/model options include attempts, model choices match normalized complete identities, and relative-time filtering refreshes every 30 seconds while the Logs tab is active, independently of network auto-refresh. Debug tab (`/#logs/debug`; legacy `/#debug` deep links redirect there): provider + usage toggles, refresh/follow log viewer. `GET/PUT /api/debug`; `GET /api/debug/logs` and `GET /api/debug/usage-logs` (monotonic `after` cursor, legacy `since` accepted). CLI: `ocx debug provider|usage …` (both streams via running proxy API). | -| Usage | `GET /api/usage` read-only aggregates of readable rows from `~/.opencodex/usage.jsonl`; the ledger is streamed in fixed 1 MiB chunks, so the former read-byte and parsed-row caps cannot omit its prefix. Oversized skipped rows produce positive `usageIncomplete` metadata. The response includes measured / reported / unreported / unsupported / estimated counts, a daily zero-filled grid, and model and provider breakdowns. `GET /api/usage/timeline` uses the same ledger and canonical attribution helpers for bounded bucketed model series. Never exposes prompts. | +| Usage | `GET /api/usage` read-only aggregates of readable rows from `~/.opencodex/usage.jsonl`; the ledger is streamed in fixed 1 MiB chunks, so the former read-byte and parsed-row caps cannot omit its prefix. Oversized skipped rows produce positive `usageIncomplete` metadata. The response includes measured / reported / unreported / unsupported / estimated counts, a daily zero-filled grid, and model and provider breakdowns. `GET /api/usage/timeline` uses the same ledger and canonical attribution helpers for bounded bucketed model series. Never exposes prompts. Fast observations and cost provenance follow the [response-tier authority contract](transports/responses.md#response-tier-observation-authority): the provider editor classifies `responseTierAuthoritative` as an operator-owned boolean under the existing redaction and write boundaries, `src/usage/log.ts` retains it on final outcomes and attempts (old records keep their legacy meaning), and `src/usage/cost.ts` never prices a non-authoritative echo or confirmed label as response-confirmed. | | Request metrics | `GET /api/metrics` exposes process-local Prometheus text format v0.0.4 only when `metricsExport.enabled` was true at startup. The ordinary management gate applies; data-plane credentials do not grant access, and disabled mode is 404. `src/server/request-metrics.ts` owns fixed counters/histograms and receives a narrow final-request fact from `src/server/request-log.ts`; `src/server/index/serve-options.ts` creates one owner and injects the recorder and read-only snapshot into the request and management paths. | | System | `POST /api/system/restart` restarts the proxy in place. Local CLI/tray callers first attest the exact runtime PID and port, then send a process-scoped HMAC capability bound to that method, path, PID, and port; the capability authorizes no other management route and is invalid after replacement. The caller observes one absolute deadline and accepts success only after a different runtime PID is healthy on the same port. `GET /api/system/health` is the authenticated scalar-only identity used by shared-plane Dashboard status and restart reconnect polling; its `spendLedger` block reports only ownership held/unheld, initialized/configured/degraded booleans and bounded persistence/corruption counters. Reading it never constructs, replays or prunes the ledger. Paths, scopes, accounts and request ids are absent, and the block never moves to unauthenticated `/healthz`. `GET /api/system/memory` — service-process runtime/memory identity (pid, Bun version/revision, optional `bunRuntimeSource` provenance, platform, RSS/heap/external/ArrayBuffers scalars, observed memory = max(RSS, external, ArrayBuffers), `bun:jsc` heap context, streamMode + eager-relay gate decision, watchdog snapshot sliced to the last 60 samples) plus privacy-safe `appOwnedBytes` retained-store totals/counters under static store ids. Its response-state block also reports spill-write `initial`/`healthy`/`degraded` status, a consecutive-failure streak, fixed error class, and failure/success timestamps. A successful publication clears the streak in the same process; raw error text and paths never enter this surface. Scalar-only payload; dashboard/admin callers use the standard management gate, while `ocx doctor` may use only the exact process-scoped local-read capability. It must never move to unauthenticated `/healthz`. | | Stop | `POST /api/stop` — restore native Codex, stop any installed service, and exit the proxy. A sibling instance restores nothing and answers `sharedTeardown: "not-owned"` ([Codex home](codex-home.md#codex-home)). | diff --git a/structure/providers-and-adapters.md b/structure/providers-and-adapters.md index ad719ac37ec..f92c4efc563 100644 --- a/structure/providers-and-adapters.md +++ b/structure/providers-and-adapters.md @@ -113,7 +113,7 @@ rewrite rules and the routed-id settlement. | `src/providers/registry.ts` | Compatibility facade; canonical provider presets for CLI, dashboard, OAuth, key providers, and metadata live in `src/providers/registry/entries-core.ts` and `entries-extended.ts`, with model seeds in `model-seeds.ts`. | | `src/providers/registry/model-ids.ts` | Classifies every `ProviderRegistryEntry` field by what its KEYS mean for selector decoding, and derives the native model ids an entry names. The classification is exhaustive by construction: a new registry field fails typecheck until its keys are given a meaning, which is what stops an identity-bearing map from being silently left out of decoding. Imported directly rather than through the facade, which is at its file-size cap. | | `src/providers/derive.ts` | Enrichment from provider presets into user config. | -| `src/providers/model-rename-fields.ts`, `src/providers/model-rename-migration.ts` | Classifies every provider config field for a declared model rename. Exact-model records, lists and nested request-pacing keys follow the replacement; an already saved replacement entry wins. Provider-wide settings and credential fields are not model identities. | +| `src/providers/model-rename-fields.ts`, `src/providers/model-rename-migration.ts` | Classifies every provider config field for a declared model rename. Exact-model records, lists and nested request-pacing keys follow the replacement; an already saved replacement entry wins. Provider-wide settings, including response-tier authority, and credential fields are not model identities. | | `src/providers/resolved-model-policy.ts`, `src/providers/resolved-model-policy-merge.ts` | Static provider/model policy resolution for the final upstream wire model, plus its pure clone/merge/URL/family helpers. The resolver detaches and freezes registry defaults, operator overrides, exact explicit input-modality declarations, provider-scoped hard wire pins (including Command Code's `claude-` prefix), aliases, and explicit false/empty values with field-level provenance. Provider derivation, routing, catalog hints, gather admission, and adapter selection consume its detached frozen result. Callers supply transport match, the exact capability row, and a credential-free effective auth decision; credential bytes, usability evidence, account/quota/health state, and observed limits remain outside the result. | | `src/oauth/` | OAuth providers, token storage, refresh, and auth-token resolution. Meta Muse device authorization, polling, and key-mint JSON responses share the 64 KiB bounded-body ceiling and the request's deadline; oversized declared or streamed bodies are rejected before JSON parsing. The login callback listener binds a per-provider FIXED loopback port, so consecutive logins reuse the same number; every response it sends ends its connection (`Connection: close`, including non-callback paths such as a stray `/favicon.ico` 404). Stopping the listener does not close an established socket, so without that a pooled client would deliver the next login's callback to the retired flow, which rejects the unknown state as a CSRF mismatch while the live flow waits. Command Code manual callback JSON remains opaque to the shared `code#state` parser and is state-validated by its provider parser. A raw Command Code paste with an explicit `#state` suffix must match the flow state on the direct prompt as well. Kiro add-account identity prefers same-session `whoami` over a leftover SQLite state profile, and never persists the Builder ID service profile ARN as `accountId`. | @@ -147,6 +147,10 @@ response. `src/bridge.ts` is the compatibility facade that re-exports both. `src/adapters/run-turn-queue.ts` preflight callers may supply an optional wait bound; timeout hands the outstanding iterator read to replay once, while callers without a bound keep the existing wait. +Fast response evidence follows the [response-tier observation contract](transports/responses.md#response-tier-observation-authority): +an explicit provider declaration can mark an intermediary's echo non-authoritative without +changing its capability or adapter wire mapping. + The image/video loop bounds each hidden iteration before replay or fulfillment; see [media iteration retention](transports/inventory.md#media-iteration-retention). diff --git a/structure/runtime.md b/structure/runtime.md index 483c2331ee1..0b710812245 100644 --- a/structure/runtime.md +++ b/structure/runtime.md @@ -4,7 +4,7 @@ The minute sweep checks persisted activation deadlines locally; only missing dea ## Resolved static model policy -`src/router.ts` attaches one frozen `ResolvedModelPolicy` to every `RouteResult`. Policy/combo +`src/router.ts` attaches one frozen `ResolvedModelPolicy` to every `RouteResult`. Fast observation, persistence and cost provenance follow the [response-tier authority contract](transports/responses.md#response-tier-observation-authority); outbound Fast policy is unchanged. Policy/combo route spreads retain that object. Every initial, fallback, and recovery route is recaptured for the request's original inbound protocol before route-dependent normalization, and all adapter rebuilds consume its recorded adapter. A translated Chat or Anthropic replay therefore cannot inherit a diff --git a/structure/transports/responses.md b/structure/transports/responses.md index 33b1a10fe3c..41d2d6b8cc7 100644 --- a/structure/transports/responses.md +++ b/structure/transports/responses.md @@ -275,11 +275,8 @@ cancellation terminates dispatch before rotation can persist another key. ### Routed service-tier capability -OpenAI-compatible service-tier support is resolved only after the final provider/model wire is -known. `supportsServiceTier` remains the provider fallback, while the exact -`modelSupportsServiceTier` map can override it per upstream model, including an explicit `false`. -The catalog and request path share this decision: a routed row publishes `service_tiers` only when -the resolved policy is eligible, and the final-route normalizer applies the same gate to +OpenAI-compatible service-tier support is resolved only after the final provider/model wire is known. `supportsServiceTier` remains the provider fallback, while the exact `modelSupportsServiceTier` map can override it per upstream model, including an explicit `false`. +The catalog and request path share this decision: a routed row publishes `service_tiers` only when the resolved policy is eligible, and the final-route normalizer applies the same gate to `service_tier`. Both `openai-responses` and `openai-chat` use the resolved provider/model capability for catalog publication, routing evidence, and fingerprints. Canonical Fast injection additionally requires a compatible FastWire mapping on the final adapter and an eligible policy. Setting @@ -288,10 +285,13 @@ foreign caller values; an exact-model `true` does not grant that forwarding perm unclassified Chat routes it gates every caller tier because no canonical Fast capability has been validated. An object-form registry wire default may also set `forwardCallerServiceTier: false` to close a known subscription gateway while leaving generic unclassified Responses passthrough -unchanged. Exact `false` -narrows provider defaults, and provider-level `supportsServiceTier: false` cannot be reopened. +unchanged. Exact `false` narrows provider defaults, and provider-level `supportsServiceTier: false` cannot be reopened. Capability is namespaced by the selected provider and model; model-name similarity and adapter type -alone never opt a gateway in. +alone never opt a gateway in. Response-tier evidence is governed by `src/providers/openai-tiers-destination.ts`: canonical ChatGPT Codex forwarding is non-authoritative (#2558); other destinations honor the optional provider boolean `responseTierAuthoritative`, with omission retaining legacy response authority. No gateway name or address infers a declaration. The final route captures this in `TierObservationContext` through `src/server/responses/core-normalize.ts`, without changing capability, `TierDecision`, or outgoing parameters. + +### Response-tier observation authority + +`src/providers/fastwire.ts` copies a defined authority flag into `AttemptTierOutcome`. With false, an eligible serialized priority request remains `fastOutcome: applied` and `confirmation: assumed`: the parameter was sent, while actual scheduling is unconfirmed. Neither a `default` nor a `priority` echo can confirm or deny Fast; sanitized `responseServiceTier` remains for inspection, while local capability and wire failures still downgrade. Logs and persisted attempts retain the flag, and `src/usage/cost.ts` uses requested-tier estimation without promoting an untrusted echo or a confirmed label to response-confirmed pricing. Official API and undeclared destinations retain legacy response judgments. Anthropic Fast eligibility and downgrade recovery use the [Responses failover contract](responses-failover.md#anthropic-fast-downgrade-recovery). diff --git a/tests/fixtures/test-layout-expected.json b/tests/fixtures/test-layout-expected.json index fd22fa5ccca..18354bba8ef 100644 --- a/tests/fixtures/test-layout-expected.json +++ b/tests/fixtures/test-layout-expected.json @@ -805,6 +805,7 @@ "fastwire-characterization-routing.test.ts": "routing", "fastwire-characterization-wire.test.ts": "routing", "fastwire-observability.test.ts": "routing", + "fastwire-response-authority.test.ts": "routing", "fastwire-policy.test.ts": "routing", "featherless-provider.test.ts": "providers", "fetch-header-timeout.test.ts": "server", diff --git a/tests/providers/model-rename-migration.test.ts b/tests/providers/model-rename-migration.test.ts index 46d1a102230..f235921d28a 100644 --- a/tests/providers/model-rename-migration.test.ts +++ b/tests/providers/model-rename-migration.test.ts @@ -57,6 +57,15 @@ function staleConfig(): OcxConfig { } describe("registry model rename migration (#1610)", () => { + test.each([false, true])("preserves provider response-tier authority %s across a model rename", authority => { + const stale = staleConfig(); + stale.providers[RENAME.provider]!.responseTierAuthoritative = authority; + const { config, changed } = projectModelRenames(stale, [RENAME]); + expect(changed).toBe(true); + expect(config.providers[RENAME.provider]!.models).toContain(RENAME.to); + expect(config.providers[RENAME.provider]!.responseTierAuthoritative).toBe(authority); + }); + test("rewrites every model-keyed field, preserving list order", () => { const { config, changed, warnings } = projectModelRenames(staleConfig(), [RENAME]); const prov = config.providers["alibaba-token-plan-intl"]!; diff --git a/tests/routing/fastwire-response-authority.test.ts b/tests/routing/fastwire-response-authority.test.ts new file mode 100644 index 00000000000..6afaba8c79e --- /dev/null +++ b/tests/routing/fastwire-response-authority.test.ts @@ -0,0 +1,212 @@ +import { afterEach, describe, expect, test } from "bun:test"; +import { modelTitle } from "../../gui/src/pages/logs-model-title"; +import { validateConfigCandidate } from "../../src/config"; +import { createAdapterTierMetadata } from "../../src/providers/fastwire"; +import { responseTierAuthorityForProvider } from "../../src/providers/openai-tiers-destination"; +import { parseProviderEditorConfigDTO, providerEditorConfigDTO } from "../../src/server/auth-cors"; +import { addFinalRequestLog, type RequestLogContext, type RequestLogEntry } from "../../src/server/request-log"; +import { handleResponses } from "../../src/server/responses/core"; +import type { OcxConfig, OcxProviderConfig } from "../../src/types"; +import { estimateComboCost, serviceTierContextFromOutcome } from "../../src/usage/cost"; +import { normalizeUsageEntryForTest } from "../../src/usage/log"; +import { acquireOwnedSpendHome } from "../helpers/owned-spend-home"; + +const originalFetch = globalThis.fetch; +let releaseSpendHome: (() => void) | undefined; +afterEach(() => { + releaseSpendHome?.(); + releaseSpendHome = undefined; + globalThis.fetch = originalFetch; +}); + +const gateway: OcxProviderConfig = { + adapter: "openai-responses", + baseUrl: "https://relay.example.test/v1", + authMode: "key", + apiKey: "sk-test", + supportsServiceTier: true, +}; + +function config(provider: OcxProviderConfig): OcxConfig { + return { port: 0, defaultProvider: "relay", providers: { relay: provider } }; +} + +async function drive(provider: OcxProviderConfig, stream: boolean, responseTier: unknown) { + const sent: Record[] = []; + const upstream = { + id: "resp_tier", object: "response", status: "completed", model: "gpt-5.6-sol", + output: [], usage: { input_tokens: 10, output_tokens: 2 }, + ...(responseTier === undefined ? {} : { service_tier: responseTier }), + }; + globalThis.fetch = (async (_url: unknown, init?: RequestInit) => { + sent.push(JSON.parse(String(init?.body))); + return new Response(stream + ? `data: ${JSON.stringify({ type: "response.completed", response: upstream })}\n\ndata: [DONE]\n\n` + : JSON.stringify(upstream), { + headers: { "content-type": stream ? "text/event-stream" : "application/json" }, + }); + }) as typeof fetch; + releaseSpendHome = acquireOwnedSpendHome(); + const log: RequestLogContext = { model: "", provider: "" }; + const response = await handleResponses(new Request("http://localhost/v1/responses", { + method: "POST", headers: { "content-type": "application/json" }, + body: JSON.stringify({ + model: "relay/gpt-5.6-sol", input: "ping", stream, service_tier: "priority", + }), + }), config(provider), log, {}); + const downstream = await response.text(); + expect(response.status).toBe(200); + expect(sent).toHaveLength(1); + expect(sent[0]?.service_tier).toBe("priority"); + expect(sent[0]).not.toHaveProperty("responseTierAuthoritative"); + if (typeof responseTier === "string") { + expect(downstream).toContain(`"service_tier":"${responseTier}"`); + } + return log; +} + +describe("response-tier authority on the final Responses route", () => { + const routes = [ + { name: "declared relay", provider: { ...gateway, responseTierAuthoritative: false }, authoritative: false }, + { name: "unconfigured relay", provider: gateway, authoritative: true }, + { name: "explicitly authoritative relay", provider: { ...gateway, responseTierAuthoritative: true }, authoritative: true }, + { name: "official API", provider: { ...gateway, baseUrl: "https://api.openai.com/v1" }, authoritative: true }, + { name: "direct Codex under a custom name", provider: { + ...gateway, baseUrl: "https://chatgpt.com/backend-api/codex", authMode: "forward" as const, + responseTierAuthoritative: true, + }, authoritative: false }, + ]; + for (const { name, provider, authoritative } of routes) { + test.each([true, false])(`${name}, stream=%s preserves wire and distinguishes evidence`, async stream => { + const log = await drive(provider, stream, "default"); + expect(log.responseServiceTier).toBe("default"); + expect(log.activeAttempt?.tierOutcome).toMatchObject({ + wireKind: "service-tier", wireValue: "priority", responseServiceTier: "default", + fastOutcome: authoritative ? "downgraded" : "applied", + confirmation: authoritative ? "downgraded" : "assumed", + }); + if (authoritative) { + expect(log.activeAttempt?.tierOutcome?.fastDowngradeReason).toBe("response-declined"); + } else { + expect(log.activeAttempt?.tierOutcome?.fastDowngradeReason).toBeUndefined(); + expect(log.activeAttempt?.tierOutcome?.responseTierAuthoritative).toBe(false); + } + }); + } + + test.each(["priority", undefined, null])("non-authoritative echo %s cannot confirm Fast", async tier => { + const log = await drive({ ...gateway, responseTierAuthoritative: false }, true, tier); + expect(log.activeAttempt?.tierOutcome).toMatchObject({ + canonical: "priority", fastOutcome: "applied", confirmation: "assumed", + responseTierAuthoritative: false, + }); + expect(log.activeAttempt?.tierOutcome?.fastDowngradeReason).toBeUndefined(); + }); + + test("official priority remains confirmed", async () => { + const log = await drive({ ...gateway, baseUrl: "https://api.openai.com/v1" }, false, "priority"); + expect(log.activeAttempt?.tierOutcome?.confirmation).toBe("confirmed"); + }); + + test("authority and raw echo survive live logs, persistence and tooltip rendering", async () => { + const log = await drive({ ...gateway, responseTierAuthoritative: false }, false, "default"); + let entry: RequestLogEntry | undefined; + addFinalRequestLog("ocx-tier-authority", Date.now(), log, 200, undefined, value => { entry = value; }); + const restored = normalizeUsageEntryForTest(JSON.parse(JSON.stringify(entry))); + expect(restored.tierOutcome).toEqual(log.activeAttempt?.tierOutcome); + expect(restored.attempts?.[0]?.tierOutcome).toEqual(restored.tierOutcome); + expect(restored.tierOutcome?.responseTierAuthoritative).toBe(false); + expect(restored.responseServiceTier).toBe("default"); + const title = modelTitle(restored, ((key: string) => key.split(".").pop()!) as never); + expect(title).toContain("requestedTier=priority"); + expect(title).toContain("responseTier=default (assumed)"); + expect(title).not.toContain("confirmed"); + expect(title).not.toContain("downgraded"); + const oldEntry = JSON.parse(JSON.stringify(entry)); + delete oldEntry.tierOutcome.responseTierAuthoritative; + delete oldEntry.attempts[0].tierOutcome.responseTierAuthoritative; + const oldRestored = normalizeUsageEntryForTest(oldEntry); + expect(oldRestored.tierOutcome?.responseTierAuthoritative).toBeUndefined(); + expect(oldRestored.attempts?.[0]?.tierOutcome?.responseTierAuthoritative).toBeUndefined(); + }); +}); + +describe("response-tier authority configuration", () => { + test("canonical Codex stays non-authoritative, while an undeclared forward relay keeps legacy authority", () => { + expect(responseTierAuthorityForProvider({ ...gateway, authMode: "forward" })).toBeUndefined(); + expect(responseTierAuthorityForProvider({ + ...gateway, authMode: "forward", baseUrl: "https://chatgpt.com/backend-api/codex/", + responseTierAuthoritative: true, + })).toBe(false); + }); + + test.each([true, false])("retains boolean declaration %s", responseTierAuthoritative => { + const result = validateConfigCandidate(config({ ...gateway, responseTierAuthoritative })); + expect(result.ok).toBe(true); + if (result.ok) expect(result.config.providers.relay?.responseTierAuthoritative).toBe(responseTierAuthoritative); + }); + + test("the provider editor round-trips the declaration without exposing credentials", () => { + const dto = providerEditorConfigDTO(config({ ...gateway, responseTierAuthoritative: false })); + expect(dto.providers.relay?.responseTierAuthoritative).toBe(false); + expect(dto.providers.relay).not.toHaveProperty("apiKey"); + expect(parseProviderEditorConfigDTO(dto)).toMatchObject({ + ok: true, value: { providers: { relay: { responseTierAuthoritative: false } } }, + }); + }); + + test.each(["false", 0, null, {}])("rejects malformed declaration %j", value => { + expect(validateConfigCandidate(config({ ...gateway, responseTierAuthoritative: value } as OcxProviderConfig)).ok) + .toBe(false); + }); +}); + +describe("non-authoritative tier cost provenance", () => { + test("persisted false authority cannot turn a confirmed label into confirmed pricing", async () => { + const log = await drive({ ...gateway, responseTierAuthoritative: false }, false, "priority"); + let entry: RequestLogEntry | undefined; + addFinalRequestLog("ocx-tier-authority", Date.now(), log, 200, undefined, value => { entry = value; }); + const stored = JSON.parse(JSON.stringify(entry)); + stored.tierOutcome.confirmation = "confirmed"; + stored.attempts[0].tierOutcome.confirmation = "confirmed"; + const restored = normalizeUsageEntryForTest(stored); + expect(restored.tierOutcome?.responseTierAuthoritative).toBe(false); + expect(serviceTierContextFromOutcome(restored.tierOutcome!)).toEqual({ requestedServiceTier: "priority" }); + expect(serviceTierContextFromOutcome(restored.attempts![0]!.tierOutcome!)).toEqual({ requestedServiceTier: "priority" }); + }); + + test("the declaration cannot hide a local unsupported-route downgrade", () => { + const tracker = createAdapterTierMetadata({ + capability: false, eligibility: "capability-unsupported", demandDecision: "inherit", + callerTier: "priority", fastWire: null, responseTierAuthoritative: false, + }, { kind: "drop" }, null, null)!; + tracker.observeResponseServiceTier("priority"); + expect(tracker.outcome).toMatchObject({ + fastOutcome: "downgraded", confirmation: "downgraded", fastDowngradeReason: "route-unsupported", + responseServiceTier: "priority", responseTierAuthoritative: false, + }); + expect(serviceTierContextFromOutcome(tracker.outcome)).toEqual({}); + }); + + test.each(["default", "priority"])("echo %s remains observational through pricing", responseTier => { + const tracker = createAdapterTierMetadata({ + capability: true, eligibility: "eligible", demandDecision: "inherit", callerTier: "priority", + fastWire: { kind: "service-tier", canonicalToWire: { priority: "priority" }, foreignCallerTiers: "verbatim" }, + responseTierAuthoritative: false, + }, { kind: "set", value: "priority" }, "service-tier", "priority")!; + tracker.observeResponseServiceTier(responseTier); + expect(serviceTierContextFromOutcome(tracker.outcome)).toEqual({ requestedServiceTier: "priority" }); + const estimate = estimateComboCost([{ + ordinal: 1, provider: "openai", model: "gpt-5.6-sol", usageStatus: "reported", + usage: { inputTokens: 200_000, outputTokens: 20_000 }, tierOutcome: tracker.outcome, + }], [{ + provider: "openai", modelId: "gpt-5.6-sol", + cost4: { input: 5, output: 30, cacheRead: 0.5, cacheWrite: 6.25 }, + source: "test", verifiedAt: "2026-08-17", status: "verified", + }])!; + expect(estimate.priorityMultiplier).toBe(2); + expect(estimate.cost.total).toBeCloseTo(3.2, 9); + expect(tracker.outcome.confirmation).toBe("assumed"); + expect(tracker.outcome.responseServiceTier).toBe(responseTier); + }); +});