diff --git a/devlog/_plan/260926_claude_input_estimate/000_overview.md b/devlog/_plan/260926_claude_input_estimate/000_overview.md new file mode 100644 index 00000000000..9475efddc24 --- /dev/null +++ b/devlog/_plan/260926_claude_input_estimate/000_overview.md @@ -0,0 +1,16 @@ +# A Claude input estimate the settled route actually sends + +Unit opened 2026-09-26. + +- [010_estimation.md](010_estimation.md) — the Messages ingress. **DONE.** + +Origin: a Paseo agent on a DeepSeek V4.1 Flash conversation through this proxy drew its context +meter at 221%. The rate implied the agent had run far past the window without compacting, so the +report read as a broken compaction loop. Compaction was healthy; the number above it was not. Same +body, same upstream: the proxy published 432,068 input tokens on `message_start` for a prompt the +upstream billed at 131,907 — **3.28x**. + +This is the `#4857` family (the floor `message_start` publishes when the upstream has sent no +confirmed usage before the first frame), but not a recurrence: `#4891` and `#5057` fixed *when* the +floor is used and *whose* count it reports. The floor faithfully reports +`estimateClaudeRequestTokens`, and that estimate was measuring the wrong body. diff --git a/devlog/_plan/260926_claude_input_estimate/010_estimation.md b/devlog/_plan/260926_claude_input_estimate/010_estimation.md new file mode 100644 index 00000000000..84933550b3c --- /dev/null +++ b/devlog/_plan/260926_claude_input_estimate/010_estimation.md @@ -0,0 +1,65 @@ +# A Claude input estimate the settled route actually sends + +Defect: `src/server/claude-messages.ts` `estimateClaudeRequestTokens`. The estimator measured the +body the **caller sent**, while `message_start` publishes it as the floor for the prompt this proxy +**actually forwarded**. On the Anthropic wire those are the same body. On every other wire they are +not, and the gap is whatever the target adapter drops. + +Measured on the live path with a real Claude Code conversation replayed byte-identical (260 +messages, 23 tools, 1,734,433 B) to `thehive/deepseek-ai/deepseek-v4.1-flash` +(`adapter: openai-chat`): `message_start` **432,068** against the upstream's `message_delta` +**131,907** — **3.28x**. The contract the estimator's own doc cites allows `>2x` drift +(`devlog/_fin/260711_claude_inbound/040_phase4_hardening.md` §3); this is outside it. + +Why the two differ. Claude Code replays its own thinking blocks, and they dominate a long body: +80 blocks, 550,930 thinking chars and 750,284 `signature` chars. In the captured request the +thinking JSON is **78.8%** of all message JSON and the base64 `signature` is **56.7%** of the +thinking JSON. The `openai-chat` wire forwards that text only when the model is listed in +`preserveReasoningContentModels`, and it has no `signature` field at all — zero `signature` +references exist anywhere under `src/adapters/openai-chat*`. The `thehive` provider config declares +no reasoning policy keys, so it drops both, and upstreams do not bill replayed reasoning they never +receive. The estimator was counting 432,068 tokens of a prompt whose real size was 131,907. + +The gap is a property of the route, not a constant, so no divisor can absorb it: a route that does +replay thinking must keep counting it. Char-per-token calibration and the `thehive/` alias prefix +were both ruled out as causes — CJK is 0.07% of the body, and moving 4 to 3.5 chars/token is ±14%, +while dropping the alias alone makes the number 14% *worse*. + +Change: + +- New leaf `src/lib/claude-request-projection.ts`: `ClaudeThinkingProjection`, the native + `{text:true, signature:true}`, `projectBlock`, and `projectClaudeRequest` — a pure, idempotent + projection that returns message content with the blocks a wire does not carry emptied. It never + mutates its input and survives a message whose content is not an array. +- `src/adapters/openai-chat/messages.ts`: new exported `openAIChatSerializesThinking(provider, + modelId)`, the single source of truth for whether that wire carries a replayed thinking block. + The `messagesToChatFormat` conversion reads the same answer once per request + (`wireSerializesThinking`) instead of re-deriving it per message, so the estimator and the + serializer cannot drift apart — they are one rule. +- `src/server/claude-messages.ts`: `estimateClaudeRequestTokens` takes an optional + `thinking: ClaudeThinkingProjection` defaulting to the native pair, and projects the messages + before measuring. `claudeRequestTokenFloor` passes the projection for the route the turn settled + on (`settledRoute`, recorded where routing resolves), and `handleClaudeCountTokens` resolves its + own route through the read-only `previewRouteModel`. A route that is not `openai-chat` keeps the + native projection, so the Anthropic-native lane stays byte-exact and pre-existing two-argument + callers keep their behavior. + +The measurement only. The caller's body is never rewritten on its way to the adapter — the +projection exists to price the prompt, not to alter it, and every reader of the floor shares the one +estimate, so no number is double-counted. + +Rejected: excluding replayed thinking unconditionally (wrong for the native lane and any adapter +that does replay it); a `billsReplayedReasoning?: boolean` flag on `ProviderAdapter` (states a +billing policy on an interface whose business is wire shape, and cannot express "text yes, +signature no", which is the shape that actually occurs). + +Tests (new `tests/claude-integration/claude-estimate-projection.test.ts`): a replayed-thinking body +projects away exactly the unserialized fields; `projectClaudeRequest` is pure, idempotent, and keeps +a message it emptied; a body with no thinking is projection-invariant; and end-to-end, the +`message_start` floor for a real captured body lands within the confirmed usage instead of 3.28x +above it. Verified to fail when the projection is disabled — the end-to-end case reports 6946 where +it expects <42, so the assertion is load-bearing rather than decorative. + +Live verification, same captured body against the real upstream (104,318 / 131,907 = **0.79x**, and +`message_delta` byte-identical at 131,907, so the upstream result is untouched and only the +published estimate moved). The Paseo meter that opened this unit reads 240% before and 58% after. diff --git a/src/adapters/openai-chat/messages.ts b/src/adapters/openai-chat/messages.ts index a712c83158a..bce315a2a19 100644 --- a/src/adapters/openai-chat/messages.ts +++ b/src/adapters/openai-chat/messages.ts @@ -7,6 +7,7 @@ import { EMPTY_TOOL_OUTPUT_ANNOTATION, isWhitespaceOnlyTextPartArray } from "../ import { identifyRoutedModel } from "../identity"; import { buildNonOpenAIToolCatalogNudgeForTools, shouldInjectNonOpenAIToolCatalogNudge } from "../tool-catalog-nudge"; import { peekReasoningForCall } from "../../responses/reasoning-replay-cache"; +import type { ClaudeThinkingProjection } from "../../lib/claude-request-projection"; import { inlineDocumentDataUrl } from "../../responses/inline-document"; import type { OcxAssistantMessage, OcxContentPart, OcxParsedRequest, OcxProviderConfig, OcxTextContent, OcxThinkingContent, OcxToolCall } from "../../types"; import { modelInList, namespacedToolName } from "../../types"; @@ -61,10 +62,36 @@ export function toolResultImageChatParts(content: string | OcxContentPart[]): un return parts; } +/** + * Whether this wire serializes a replayed thinking block for the given model. + * + * The single source of truth for that question on the Chat wire: the assistant branch below uses + * it to decide `reasoning_content`, and the Messages ingress uses it to decide how much of a + * prompt is worth counting. Two copies of this rule would drift, and the count would then + * describe a body this adapter does not send. + */ +export function openAIChatSerializesThinking( + provider: OcxProviderConfig, + modelId: string, +): ClaudeThinkingProjection { + return { + text: modelInList(provider.preserveReasoningContentModels, modelId), + // The wire has no `signature` field — base64 replay tokens are Anthropic-wire-only. + signature: false, + // A `redacted_thinking` block is opaque provider data with no Chat representation: the + // assistant branch below reads only `type: "thinking"` parts, so the encrypted form is + // dropped whatever the preserve list says. + redacted: false, + }; +} + export function messagesToChatFormat(parsed: OcxParsedRequest, provider: OcxProviderConfig): unknown[] { const out: unknown[] = []; const { context, options } = parsed; const replayCacheScope = parsed._reasoningReplayScope; + // One question, one answer for the whole conversion: does this wire carry replayed thinking + // back for this model? Asked once so a long history cannot pay for it per message. + const wireSerializesThinking = openAIChatSerializesThinking(provider, parsed.modelId).text; interface PendingToolCall { id: string; name: string } let pendingToolCalls: PendingToolCall[] = []; @@ -225,7 +252,7 @@ export function messagesToChatFormat(parsed: OcxParsedRequest, provider: OcxProv if ( reasoningContent.length === 0 && (toolCalls.length > 0 || thinkingParts.length > 0) - && modelInList(provider.preserveReasoningContentModels, parsed.modelId) + && wireSerializesThinking ) { const cached = toolCalls .map(tc => (tc.id ? peekReasoningForCall(tc.id, replayCacheScope) : undefined)) @@ -248,7 +275,7 @@ export function messagesToChatFormat(parsed: OcxParsedRequest, provider: OcxProv reasoningContent = " "; } } - if (reasoningContent.length > 0 && modelInList(provider.preserveReasoningContentModels, parsed.modelId)) { + if (reasoningContent.length > 0 && wireSerializesThinking) { // MiniMax's interleaved-thinking contract requires the structured // reasoning_details array back on the next turn; a reasoning_content // string is the native-format pass-back the docs mark unsupported. @@ -302,7 +329,7 @@ export function messagesToChatFormat(parsed: OcxParsedRequest, provider: OcxProv flushPendingToolCalls(); const name = safeToolName(msg.toolName); const cachedReasoning = - toolCallId && modelInList(provider.preserveReasoningContentModels, parsed.modelId) + toolCallId && wireSerializesThinking ? peekReasoningForCall(toolCallId, replayCacheScope) : undefined; // Same fallback as the main-assistant path: never emit a bare orphan @@ -316,7 +343,7 @@ export function messagesToChatFormat(parsed: OcxParsedRequest, provider: OcxProv // falsy hit as a miss so the placeholder still fires. const orphanReasoning = cachedReasoning - || (modelInList(provider.preserveReasoningContentModels, parsed.modelId) + || (wireSerializesThinking && modelInList(provider.requiresReasoningPlaceholderModels ?? provider.preserveReasoningContentModels, parsed.modelId) ? " " : undefined); diff --git a/src/lib/claude-request-projection.ts b/src/lib/claude-request-projection.ts new file mode 100644 index 00000000000..cb10fb5ae03 --- /dev/null +++ b/src/lib/claude-request-projection.ts @@ -0,0 +1,101 @@ +/** + * Project a Messages request body onto the content a settled route actually serializes. + * + * `estimateClaudeRequestTokens` measures the caller's body as JSON. That is exactly right for the + * Anthropic-native wire, where a replayed `thinking` block — signature included — is forwarded + * verbatim. A routed wire may serialize far less, and on the Chat wire it does: replayed thinking + * becomes an optional `reasoning_content` string, the signature is never sent at all + * (`src/adapters/openai-chat/messages.ts` has no `signature` reference), and the text is dropped + * outright unless the model is on the provider's `preserveReasoningContentModels` list. + * + * The gap that opens is not a rounding error. On a captured 260-message Claude Code turn, replayed + * thinking blocks were 78.8% of the body's JSON characters and 56.7% of those characters were + * base64 signatures — bytes that are not prompt text under any tokenizer. Counting them made the + * published `message_start.usage.input_tokens` 3.28x the count the upstream actually reported + * (#4857 + the Paseo context meter it feeds), well past the >2x drift bound this estimator is held + * to (devlog 260711_claude_inbound 040 §3). + * + * The projection is therefore applied to the ESTIMATE only. The body the caller sent is never + * rewritten; this is a measurement that describes the route, not a transformation of the request. + */ + +/** Which replayed reasoning blocks and fields a wire serializes. */ +export interface ClaudeThinkingProjection { + /** Serialize `thinking.thinking`, the model's own replayed text. */ + text: boolean; + /** + * Serialize `thinking.signature`, the provider's base64 replay token. Real prompt size for + * wires that carry it; pure overhead for wires that do not. + */ + signature: boolean; + /** + * Serialize a `redacted_thinking` block, whose `data` is an opaque provider blob. It rides + * alongside `thinking` rather than inside it: a wire can reconstruct the model's reasoning + * without carrying the encrypted form, and the Chat wire does exactly that. + */ + redacted: boolean; +} + +/** + * The Anthropic-native wire forwards a replayed thinking block verbatim, signature included, so + * nothing is projected away. This is the default: an unknown route keeps the measured body. + */ +export const CLAUDE_NATIVE_THINKING: ClaudeThinkingProjection = { text: true, signature: true, redacted: true }; + +interface ProjectableBody { + system?: unknown; + messages?: unknown; + tools?: unknown; +} + +/** + * One content block as the given wire would carry it. + * + * `undefined` means the wire sends nothing for it, and the caller drops the entry. A `thinking` + * block that keeps its text but not its signature is returned as a copy rather than edited: the + * block object is shared with the outbound request builder, which must stay untouched. + */ +function projectBlock(block: unknown, thinking: ClaudeThinkingProjection): unknown { + if (!block || typeof block !== "object") return block; + if (!("type" in block)) return block; + if (block.type === "redacted_thinking") return thinking.redacted ? block : undefined; + if (block.type !== "thinking") return block; + if (!thinking.text) return undefined; + if (thinking.signature || !("signature" in block)) return block; + // Copy without the signature: the caller's block object is shared with the outbound request + // builder, so it must never be edited in place. + const { signature: _dropped, ...rest } = block; + return rest; +} + +/** + * A copy of `raw` whose message content carries only the replayed thinking this route serializes. + * + * Pure and idempotent; the input is never mutated. Blocks are dropped only inside array-valued + * message `content`, which is the sole protocol position a replayed thinking block occupies. + * A message whose content array empties out is left as an empty array rather than deleted: the + * adapter's own "nothing left to send" rule is keyed on text, tool calls and reasoning together, + * and re-deriving it here would be a second copy of that rule to keep in sync. + */ +export function projectClaudeRequest( + raw: ProjectableBody, + thinking: ClaudeThinkingProjection, +): ProjectableBody { + if (thinking.text && thinking.signature && thinking.redacted) return raw; + if (!Array.isArray(raw.messages)) return raw; + const messages = raw.messages as unknown[]; + return { + ...raw, + messages: messages.map(message => { + if (!message || typeof message !== "object") return message; + if (!("content" in message) || !Array.isArray(message.content)) return message; + const content = message.content as unknown[]; + return { + ...message, + content: content + .map(block => projectBlock(block, thinking)) + .filter(block => block !== undefined), + }; + }), + }; +} diff --git a/src/server/claude-messages.ts b/src/server/claude-messages.ts index 4bcb558cde3..33e8f36d776 100644 --- a/src/server/claude-messages.ts +++ b/src/server/claude-messages.ts @@ -18,6 +18,7 @@ import { sseFieldValue } from "../lib/sse-decoder"; import { enforceAnthropicImageLimits, sniffImageDimensions } from "../adapters/anthropic-image-guard"; import { normalizeAnthropicImages } from "../adapters/anthropic-image-normalize"; import { createToolCallIdAllocator } from "../adapters/tool-call-id"; +import { openAIChatSerializesThinking } from "../adapters/openai-chat/messages"; import { messagesToResponsesTranslation } from "../protocols/codecs/messages"; import { AnthropicRequestError, DesktopModelMappingUnavailableError, extractOcxEffortDirective, extractOcxRouteDirective, resolveInboundModel, type ClaudeCacheKeySource } from "../claude/inbound"; import { isKnownDesktop3pModelId, resolveDesktop3pAlias } from "../claude/desktop-3p"; @@ -45,7 +46,12 @@ import { } from "../claude/outbound"; import { clearableDeadline, idleDeadline } from "../lib/abort"; import { estimateTokens } from "../lib/token-estimate"; -import { captureRouteStaticPolicy, NoEligiblePolicyCandidateError, UnknownRoutingPolicyError, routeModel } from "../router"; +import { + CLAUDE_NATIVE_THINKING, + projectClaudeRequest, + type ClaudeThinkingProjection, +} from "../lib/claude-request-projection"; +import { captureRouteStaticPolicy, NoEligiblePolicyCandidateError, previewRouteModel, routedProviderConfig, UnknownRoutingPolicyError, routeModel, type RouteResult } from "../router"; import { evidenceFromBody } from "../routing/request-evidence"; import { resolveWireProtocolOverride } from "./adapter-resolve"; import type { OcxConfig } from "../types"; @@ -904,7 +910,7 @@ async function handleClaudeMessagesWithBudget( if (!requestedModel) requestedModel = (anthropicBody as Rec).model as string; const stream = internalBody.stream === true; /** - * This proxy's count of the prompt it is about to forward, computed at most once. + * This proxy's count of the prompt it is about to forward, computed at most once per wire. * * Two readers want it and they want it under different rules. The usage log takes it as a * floor only for estimated-usage adapters, because its merge is `max(reported, estimate)` and @@ -913,12 +919,17 @@ async function handleClaudeMessagesWithBudget( * `message_delta` still corrects it (#4857). */ let requestTokenFloor: number | undefined; - const claudeRequestTokenFloor = (): number => { - if (requestTokenFloor === undefined) { - requestTokenFloor = estimateClaudeRequestTokens(anthropicBody as Rec, requestedModel); - } - return requestTokenFloor; - }; + /** + * The `text|signature|redacted` triple `requestTokenFloor` was measured under. + * + * The count is not a constant for the request: a combo re-picks its child at dispatch and a + * retry can rotate the wire, so the pre-dispatch callers and the post-dispatch translator can + * legitimately want different projections. Keying the memo re-measures when they disagree + * instead of handing the translator the pre-dispatch measurement, and the estimator still runs + * at most twice for one request. At most, because a request has one settled wire per phase and + * the key collapses every repeat of the same answer. + */ + let requestTokenFloorKey: string | undefined; // Routed adapters only support streamed turns; always stream internally and fold // the translated Anthropic SSE into a message JSON for non-streaming clients. internalBody.stream = true; @@ -927,7 +938,25 @@ async function handleClaudeMessagesWithBudget( // Native ChatGPT passthrough (openai-responses forward) accepts only Codex-shaped // bodies: it 400s on sampling params ("Unsupported parameter: max_output_tokens", // verified live 2026-07-11). Strip them for that route; routed providers keep them. - let settledRoute: ReturnType | undefined; + let settledRoute: RouteResult | undefined; + /** + * Which replayed thinking fields the send that actually happened will serialize. + * + * Read lazily: routing settles the ingress wire, core may then re-pick it (a combo child, a + * rotated retry), and this is called both before and after that happens. It therefore reads the + * physical attempt when one exists and the ingress route only until then. + */ + const claudeThinkingProjection = (): ClaudeThinkingProjection => + thinkingProjectionForDispatch(config, settledRoute, logCtx); + const claudeRequestTokenFloor = (): number => { + const thinking = claudeThinkingProjection(); + const key = `${thinking.text}|${thinking.signature}|${thinking.redacted}`; + if (requestTokenFloor === undefined || requestTokenFloorKey !== key) { + requestTokenFloor = estimateClaudeRequestTokens(anthropicBody as Rec, requestedModel, thinking); + requestTokenFloorKey = key; + } + return requestTokenFloor; + }; try { const route = routeModel(config, internalBody.model as string, evidenceFromBody(internalBody)); // Same reason as the native Chat lane: this route can be sent from here, so @@ -1330,12 +1359,20 @@ function estimateBase64AttachmentTokens(data: string): number { * characters: one 2MB screenshot is ~2.7M base64 chars, which the plain chars/token * divide reports as hundreds of thousands of tokens versus a real cost around 1.6k. * That breaks the >2x drift bound the estimator is held to (devlog 260711_claude_inbound - * 040 §3). Text and url sources are left in place and counted as characters, as is + * 040 §3); a live 260-message turn whose replayed thinking was 78.8% of the body breached + * it at 3.28x, which is why the estimate is projected onto the settled route. Text and url + * sources are left in place and counted as characters, as is * anything outside protocol content positions (tool_use.input, tool schemas). + * + * `thinking` selects which replayed thinking fields the SETTLED route serializes, so the measure + * describes the prompt this proxy forwards rather than the one the caller typed. Omitted, the + * whole body counts — correct for the Anthropic-native wire, where nothing is projected away. + * See `claude-request-projection.ts` for why a routed wire must project it out. */ export function estimateClaudeRequestTokens( raw: { system?: unknown; messages?: unknown; tools?: unknown }, modelId: string | undefined, + thinking: ClaudeThinkingProjection = CLAUDE_NATIVE_THINKING, ): number { let attachmentTokens = 0; // Blank base64 payloads ONLY in protocol content positions: message content blocks and @@ -1370,11 +1407,93 @@ export function estimateClaudeRequestTokens( : messages; const parts: string[] = []; if (raw.system !== undefined) parts.push(typeof raw.system === "string" ? raw.system : JSON.stringify(raw.system)); - if (raw.messages !== undefined) parts.push(JSON.stringify(sanitizedMessages(raw.messages))); + if (raw.messages !== undefined) { + const projected = projectClaudeRequest(raw, thinking); + parts.push(JSON.stringify(sanitizedMessages(projected.messages))); + } if (raw.tools !== undefined) parts.push(JSON.stringify(raw.tools)); return Math.max(1, estimateTokens(parts.join("\n"), modelId) + attachmentTokens); } +/** + * The projection for a route that has already settled. + * + * Only the OpenAI-shaped Chat adapter discards replayed thinking; every other settled wire + * forwards the body it was given. An unknown route keeps the full body. + */ +function thinkingProjectionForRoute(route: RouteResult | undefined): ClaudeThinkingProjection { + if (!route || route.provider.adapter !== "openai-chat") return CLAUDE_NATIVE_THINKING; + return openAIChatSerializesThinking(route.provider, route.modelId); +} + +/** + * The projection for the wire that will physically carry the request. + * + * The ingress route is not the last word on that wire. A combo re-picks its child at dispatch, + * and a retry can rotate the adapter mid-turn, so the ingress pick can name a different provider + * — and therefore a different body — than the one that is sent. `logCtx.activeAttempt` is the + * send that actually happened (`sealRequestAttemptIdentity` keeps its adapter current), so it + * wins once it exists; before the first send the ingress route is the only authority there is. + * + * A Chat identity is re-derived through `routedProviderConfig`, not read off the raw config row: + * `preserveReasoningContentModels` is registry-merged, so a row that omits it would otherwise + * price a preserve-listed model as if the wire dropped its reasoning. + */ +function thinkingProjectionForDispatch( + config: OcxConfig, + route: RouteResult | undefined, + logCtx: Pick, +): ClaudeThinkingProjection { + const attempt = logCtx.activeAttempt; + const adapter = attempt?.adapter ?? logCtx.providerAdapter; + // No send to describe yet: the ingress route is the best available answer, and for the + // Anthropic-native wire (which drops nothing) it is already the right one. + if (adapter === undefined) return thinkingProjectionForRoute(route); + if (adapter !== "openai-chat") return CLAUDE_NATIVE_THINKING; + const providerName = attempt?.provider; + const modelId = attempt?.model ?? route?.modelId; + const provider = providerName !== undefined && Object.hasOwn(config.providers, providerName) + ? config.providers[providerName] + : undefined; + // A Chat wire whose destination cannot be named keeps the route's own answer: over-counting on + // a path that publishes nothing is harmless, under-counting a real prompt is not. + if (providerName === undefined || modelId === undefined || provider === undefined) { + return thinkingProjectionForRoute(route); + } + try { + return openAIChatSerializesThinking(routedProviderConfig(providerName, provider), modelId); + } catch { + return thinkingProjectionForRoute(route); + } +} + +/** + * The projection for a route resolved only to MEASURE a body this handler never sends. + * + * `previewRouteModel` is the read-only resolver: it advances no combo round-robin state, so a + * count request cannot steer where the next real turn goes. An unresolvable model keeps the + * full body, matching the old behavior for models routing cannot place. + * + * The wire is settled exactly as the turn path settles it. Routing fills in the provider's + * registry adapter, and a per-model `modelAdapters` override or a pinned wire is applied later by + * `resolveWireProtocolOverride` — so skipping it here would price the count against a body the + * routed adapter never sends. + */ +function thinkingProjectionForPreview(config: OcxConfig, modelId: string): ClaudeThinkingProjection { + try { + const route = previewRouteModel(config, modelId); + route.staticPolicy = captureRouteStaticPolicy( + route.providerName, route.modelId, route.provider, route.staticPolicy.effectiveAlias, "anthropic", + ); + route.provider = resolveWireProtocolOverride( + route.providerName, route.modelId, route.provider, "anthropic", route.staticPolicy, + ); + return thinkingProjectionForRoute(route); + } catch { + return CLAUDE_NATIVE_THINKING; + } +} + export async function handleClaudeCountTokens( req: Request, config: OcxConfig, @@ -1435,7 +1554,14 @@ export async function handleClaudeCountTokens( const nativeCountBody = resolveProtocolSettings(config).rollout.managedMessagesNative ? (await import("./messages-native")).nativeMessagesCountBody(config, cc, raw, { fastRow: countFastRow !== null }) : undefined; - const inputTokens = estimateClaudeRequestTokens(nativeCountBody ?? raw, model); + // A count answers for the prompt a real turn from this model would forward, so it projects + // the same unserialized content that turn's `message_start` floor does. Counting the raw + // caller body instead reported replayed thinking this route never sends (#4857 family). + const inputTokens = estimateClaudeRequestTokens( + nativeCountBody ?? raw, + model, + thinkingProjectionForPreview(config, model), + ); return new Response(JSON.stringify({ input_tokens: inputTokens }), { status: 200, headers: { "Content-Type": "application/json" }, diff --git a/src/server/responses/core-combo.ts b/src/server/responses/core-combo.ts index 20b4cc20994..26d4b731c3f 100644 --- a/src/server/responses/core-combo.ts +++ b/src/server/responses/core-combo.ts @@ -667,7 +667,11 @@ export async function executeComboResponses( const attempt = beginRequestAttempt( (logCtx.attempts?.length ?? 0) + 1, pick.target.provider, - pick.target.model, + // The id the child wire will actually send, not the selector the combo named. A target may + // be an alias, and `routeConcreteModel` above is where it becomes the provider's native id; + // recording the alias here would describe a request that never left (the adapter resolves + // the id before it reads any per-model list). + targetRoute.modelId, config.providers[pick.target.provider]!.adapter, ); childLog.activeAttempt = attempt; diff --git a/tests/claude-integration/claude-estimate-projection.test.ts b/tests/claude-integration/claude-estimate-projection.test.ts new file mode 100644 index 00000000000..5459d011dbb --- /dev/null +++ b/tests/claude-integration/claude-estimate-projection.test.ts @@ -0,0 +1,493 @@ +/** + * The Claude Messages request-token estimate, projected onto the route that will carry it. + * + * A Claude Code turn replays its own thinking blocks, and on a long session those blocks dominate + * the body: on a captured 260-message turn they were 78.8% of the messages JSON, 56.7% of that + * being base64 signatures. A routed OpenAI Chat wire serializes almost none of it — the signature + * never, the text only for preserve-listed models — so counting the caller's own blocks published + * a `message_start.usage.input_tokens` 3.28x the prompt the upstream actually received. Paseo's + * context meter reads that frame, so it showed 221% of a 180k window while compaction was healthy. + * + * These cases pin the measure itself and the wiring that feeds it. The estimator's + * attachment-pricing behavior lives with the endpoint suite. + */ +import { afterEach, expect, test } from "bun:test"; +import { mkdtempSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { saveConfig } from "../../src/config"; +import { startServer } from "../../src/server"; +import { estimateClaudeRequestTokens, handleClaudeCountTokens, handleClaudeMessages } from "../../src/server/claude-messages"; +import { estimateTokens } from "../../src/lib/token-estimate"; +import { CLAUDE_NATIVE_THINKING, projectClaudeRequest } from "../../src/lib/claude-request-projection"; +import { openAIChatSerializesThinking } from "../../src/adapters/openai-chat/messages"; +import { clearComboSelectionState, clearComboTargetCooldowns } from "../../src/combos"; +import type { OcxConfig, OcxProviderConfig } from "../../src/types"; +import { removeTreeWithRetry } from "../helpers/remove-tree"; +import { acquireOwnedSpendHome } from "../helpers/owned-spend-home"; +import { SERVER_BUDGET_MS } from "../helpers/test-budget"; + +let testDir = ""; +let previousHome: string | undefined; +let releaseSpendHome: (() => void) | undefined; +let isolatedHomeActive = false; + +/** Only the end-to-end case needs a home; the estimator cases below are pure. */ +function setUpIsolatedHome(): void { + previousHome = process.env.OPENCODEX_HOME; + testDir = mkdtempSync(join(tmpdir(), "ocx-claude-estimate-")); + process.env.OPENCODEX_HOME = testDir; + releaseSpendHome = acquireOwnedSpendHome(); + isolatedHomeActive = true; +} + +function restoreIsolatedHome(): void { + // The preload arms OPENCODEX_HOME for the whole process, so an unpaired restore must not + // touch it: deleting it here left sibling files running in the same process to write the + // real home, which the preload guard then refused. + if (!isolatedHomeActive) return; + isolatedHomeActive = false; + releaseSpendHome?.(); + releaseSpendHome = undefined; + if (previousHome === undefined) delete process.env.OPENCODEX_HOME; + else process.env.OPENCODEX_HOME = previousHome; + if (testDir) removeTreeWithRetry(testDir); + testDir = ""; +} + +afterEach(restoreIsolatedHome); + +/** A Chat-completions upstream that records what the proxy actually sent it. */ +function mockChatUpstreamCapturing(): { server: ReturnType; captured: Array> } { + const captured: Array> = []; + const server = Bun.serve({ + port: 0, + async fetch(req) { + try { captured.push(await req.json() as Record); } catch { /* keep streaming */ } + const frames = [ + `data: ${JSON.stringify({ choices: [{ index: 0, delta: { role: "assistant", content: "Hello" } }] })}\n\n`, + `data: ${JSON.stringify({ choices: [{ index: 0, delta: {}, finish_reason: "stop" }], usage: { prompt_tokens: 12, completion_tokens: 3 } })}\n\n`, + "data: [DONE]\n\n", + ]; + return new Response(frames.join(""), { headers: { "Content-Type": "text/event-stream" } }); + }, + }); + return { server, captured }; +} + +function mockConfig(baseUrl: string): OcxConfig { + return { + port: 0, + defaultProvider: "mock", + providers: { + mock: { adapter: "openai-chat", baseUrl, apiKey: "k", allowPrivateNetwork: true }, + }, + } as OcxConfig; +} + +test("estimateClaudeRequestTokens drops replayed thinking the settled Chat wire does not send", () => { + // The defect this pins (#4857 family): a routed openai-chat turn serializes replayed + // thinking only as `reasoning_content`, and only for preserve-listed models — the + // signature is never sent at all. Counting the caller's own blocks made the published + // message_start floor 3.28x the upstream's reported prompt on a live 260-message turn. + const thinking = { + type: "thinking", + thinking: "T".repeat(60_000), + signature: "S".repeat(90_000), + }; + const raw = { + messages: [ + { role: "assistant", content: [thinking, { type: "text", text: "answer" }] }, + { role: "user", content: "next" }, + ], + }; + const partsWithoutThinking = [ + JSON.stringify([{ role: "assistant", content: [{ type: "text", text: "answer" }] }, { role: "user", content: "next" }]), + ]; + + // A route whose wire serializes no replayed thinking counts only what it would send. + const dropped = estimateClaudeRequestTokens(raw, "m", { text: false, signature: false }); + expect(dropped).toBe(Math.max(1, estimateTokens(partsWithoutThinking.join("\n"), "m"))); + // Signature-only projection still prices the model's own replayed text. + const textOnly = estimateClaudeRequestTokens(raw, "m", { text: true, signature: false }); + expect(textOnly).toBeGreaterThan(dropped); + // The native wire forwards the block verbatim, which is the default for an unknown route. + const native = estimateClaudeRequestTokens(raw, "m", { text: true, signature: true }); + expect(native).toBe(Math.max(1, estimateTokens(JSON.stringify(raw.messages), "m"))); + expect(native).toBe(estimateClaudeRequestTokens(raw, "m")); + // The dropped measure must not still be carrying the signature's bytes. + expect(dropped).toBeLessThan(native / 3); +}); + +test("a redacted_thinking blob is priced where the wire carries it and dropped where it cannot", () => { + // `redacted_thinking` rides in the same content array as `thinking` but is its own axis: the + // Anthropic-native lane replays the opaque blob verbatim, and the Chat wire has no + // representation for it at all. Folding it into the text/signature axes would either lose the + // blob's real bytes on the native lane or keep pricing it on a wire that never sends it, which + // is the same 40x over-count this file exists to prevent. + const data = "R".repeat(90_000); + const redactedOnly = { messages: [{ role: "assistant", content: [{ type: "redacted_thinking", data }] }] }; + const kept = estimateClaudeRequestTokens(redactedOnly, "m", { text: true, signature: true, redacted: true }); + const dropped = estimateClaudeRequestTokens(redactedOnly, "m", { text: true, signature: true, redacted: false }); + // The all-true projection is the identity fast path, so the native default must agree exactly. + expect(kept).toBe(estimateClaudeRequestTokens(redactedOnly, "m")); + expect(kept).toBeGreaterThan(10_000); + // What is left is the JSON envelope around an emptied block, not 90k characters of blob. + expect(dropped).toBeLessThan(kept / 20); + // A signature-less thinking block beside it still answers to the axes that own it. + const both = { + messages: [{ + role: "assistant", + content: [{ type: "thinking", thinking: "T".repeat(4_000), signature: "S".repeat(4_000) }, { type: "redacted_thinking", data }], + }], + }; + expect(estimateClaudeRequestTokens(both, "m", { text: true, signature: true, redacted: true })) + .toBe(estimateClaudeRequestTokens(both, "m")); + expect(estimateClaudeRequestTokens(both, "m", { text: false, signature: false, redacted: true })) + .toBeGreaterThan(estimateClaudeRequestTokens(both, "m", { text: false, signature: false, redacted: false }) * 20); +}); + +test("a preserve-listed Chat model still serializes no redacted_thinking", () => { + // The preserve list buys `reasoning_content`, nothing more: the assistant branch builds that + // string from `type: "thinking"` parts alone. The shared helper therefore reports the blob's + // axis false for every Chat provider, listed or not, and the estimate follows it. + const listed: OcxProviderConfig = { + adapter: "openai-chat", + baseUrl: "http://127.0.0.1:1/v1", + apiKey: "k", + preserveReasoningContentModels: ["m"], + allowPrivateNetwork: true, + }; + expect(openAIChatSerializesThinking(listed, "m")).toEqual({ text: true, signature: false, redacted: false }); + const raw = { messages: [{ role: "assistant", content: [{ type: "redacted_thinking", data: "R".repeat(90_000) }] }] }; + const onChat = estimateClaudeRequestTokens(raw, "m", openAIChatSerializesThinking(listed, "m")); + expect(onChat).toBeLessThan(estimateClaudeRequestTokens(raw, "m") / 20); +}); + +test("projectClaudeRequest is pure, idempotent, and keeps emptied messages", () => { + const thinking = { type: "thinking", thinking: "replayed", signature: "sig" }; + const raw = { messages: [{ role: "assistant", content: [thinking] }] }; + + const projected = projectClaudeRequest(raw, { text: false, signature: false }); + expect(projected).not.toBe(raw); + expect(projected.messages).toEqual([{ role: "assistant", content: [] }]); + // The caller's body is shared with the outbound request builder, so it must be untouched. + expect(raw.messages).toEqual([{ role: "assistant", content: [thinking] }]); + // Re-running changes nothing: an emptied content array has no thinking left to drop. + expect(projectClaudeRequest(projected, { text: false, signature: false })).toEqual(projected); +}); + +test("estimateClaudeRequestTokens counts a body with no thinking identically under every projection", () => { + // Guards the default path the pre-existing estimator tests rely on: with nothing to + // project away, the projection cannot change the answer. + const raw = { + system: "be brief", + messages: [{ role: "user", content: [{ type: "text", text: "no thinking here" }] }], + tools: [{ name: "Read", input_schema: { type: "object" } }], + }; + const native = estimateClaudeRequestTokens(raw, "m"); + expect(estimateClaudeRequestTokens(raw, "m", { text: false, signature: false })).toBe(native); + expect(estimateClaudeRequestTokens(raw, "m", { text: true, signature: false })).toBe(native); +}); + +test("count_tokens prices the wire the modelAdapters override selects, in both directions", async () => { + // A count is a promise about the prompt a real turn would forward, so it has to settle the + // wire the same way that turn does. Routing fills in the provider's registry adapter, and a + // per-model override lands afterwards — pricing the provider-wide adapter instead gets the + // answer backwards whenever the two disagree, which is precisely the case overrides exist for. + const body = { + model: "mock/test-model", + messages: [ + { role: "user", content: "u" }, + { + role: "assistant", + content: [ + { type: "thinking", thinking: "T".repeat(4_000), signature: "S".repeat(4_000) }, + { type: "text", text: "a" }, + ], + }, + ], + }; + const chatWire = estimateClaudeRequestTokens(body, "mock/test-model", { text: false, signature: false, redacted: false }); + const nativeWire = estimateClaudeRequestTokens(body, "mock/test-model", CLAUDE_NATIVE_THINKING); + const countFor = async (providerAdapter: string, override: string): Promise => { + const config = { + port: 0, + defaultProvider: "mock", + providers: { + mock: { + adapter: providerAdapter, + baseUrl: "http://127.0.0.1:1/v1", + apiKey: "k", + allowPrivateNetwork: true, + modelAdapters: { "test-model": override }, + }, + }, + } as unknown as OcxConfig; + const response = await handleClaudeCountTokens(new Request("http://localhost/v1/messages/count_tokens", { + method: "POST", + headers: { "content-type": "application/json" }, + body: JSON.stringify(body), + }), config); + expect(response.status).toBe(200); + return ((await response.json()) as { input_tokens: number }).input_tokens; + }; + // Provider says Responses, the model says Chat: the Chat body is the one that gets sent. + expect(await countFor("openai-responses", "openai-chat")).toBe(chatWire); + // And the reverse: a Chat provider whose model speaks Responses forwards the thinking verbatim. + expect(await countFor("openai-chat", "openai-responses")).toBe(nativeWire); + // The two directions must stay distinguishable, or the assertions above are vacuous. + expect(chatWire).toBeLessThan(nativeWire / 20); +}); + +test("message_start floor describes the prompt the Chat wire actually sent, not the replayed thinking", async () => { + // End-to-end pin for the route-aware projection: the estimator must read the SETTLED route, + // not just accept a projection when handed one. A routed openai-chat turn with no + // preserveReasoningContentModels entry serializes no replayed thinking, so a floor that + // still counts it overstates the prompt — the live defect that published 3.28x. + setUpIsolatedHome(); + const upstream = mockChatUpstreamCapturing(); + saveConfig(mockConfig(`${upstream.server.url.toString().replace(/\/$/, "")}/v1`)); + const server = startServer(0); + try { + const thinking = { + type: "thinking", + thinking: "replayed reasoning ".repeat(400), + signature: "S".repeat(20_000), + }; + const response = await fetch(new URL("/v1/messages", server.url), { + method: "POST", + headers: { "content-type": "application/json" }, + body: JSON.stringify({ + model: "mock/test-model", + max_tokens: 128, + stream: true, + messages: [ + { role: "user", content: "first" }, + { role: "assistant", content: [thinking, { type: "text", text: "answer" }] }, + { role: "user", content: "second" }, + ], + }), + }); + expect(response.status).toBe(200); + const text = await response.text(); + const startFrame = text.slice(text.indexOf("event: message_start")); + const published = (JSON.parse(startFrame.slice(startFrame.indexOf("data: ") + 6, startFrame.indexOf("\n\n"))) + .message.usage.input_tokens) as number; + + expect(upstream.captured).toHaveLength(1); + const sent = upstream.captured[0]!; + // What the wire actually carried: no signature field, and no reasoning_content because + // this model is not on a preserve list. + const serialized = JSON.stringify(sent.messages); + expect(serialized).not.toContain("S".repeat(64)); + expect(serialized).not.toContain("reasoning_content"); + + // The floor must therefore land near the serialized prompt, not near the caller's body. + const sentEstimate = estimateTokens(JSON.stringify(sent.messages), "mock/test-model"); + expect(published).toBeLessThan(sentEstimate * 1.5); + expect(published).toBeGreaterThan(sentEstimate * 0.5); + // And it must be far below what counting the caller's own thinking would produce. + expect(published).toBeLessThan(estimateClaudeRequestTokens({ messages: JSON.parse(JSON.stringify([thinking])) }, "mock/test-model") / 2); + } finally { + await server.stop(true); + upstream.server.stop(true); + restoreIsolatedHome(); + } +}, { timeout: SERVER_BUDGET_MS }); + +/** Frames an Anthropic-wire upstream answers with, including the usage frame a client reads. */ +const ANTHROPIC_SSE_FRAMES = [ + 'event: message_start\ndata: {"type":"message_start","message":{"id":"msg_b","type":"message","role":"assistant","model":"m2","content":[],"stop_reason":null,"stop_sequence":null,"usage":{"input_tokens":7,"output_tokens":0}}}\n\n', + 'event: content_block_start\ndata: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}\n\n', + 'event: content_block_delta\ndata: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"sunny"}}\n\n', + 'event: content_block_stop\ndata: {"type":"content_block_stop","index":0}\n\n', + 'event: message_delta\ndata: {"type":"message_delta","delta":{"stop_reason":"end_turn"},"usage":{"output_tokens":5}}\n\n', + 'event: message_stop\ndata: {"type":"message_stop"}\n\n', +].join(""); + +/** Chat-wire frames for the upstream that answers the combo's second target in the Chat case. */ +const CHAT_SSE_FRAMES = [ + `data: ${JSON.stringify({ choices: [{ index: 0, delta: { role: "assistant", content: "sunny" } }] })}\n\n`, + `data: ${JSON.stringify({ choices: [{ index: 0, delta: {}, finish_reason: "stop" }], usage: { prompt_tokens: 12, completion_tokens: 3 } })}\n\n`, + "data: [DONE]\n\n", +].join(""); + +const COMBO_THINKING = { + type: "thinking", + thinking: "replayed reasoning ".repeat(400), + signature: "S".repeat(20_000), +}; +const COMBO_MESSAGES = [ + { role: "user", content: "first" }, + { role: "assistant", content: [COMBO_THINKING, { type: "text", text: "answer" }] }, + { role: "user", content: "second" }, +]; + +/** + * One `combo/pair` turn whose first target refuses so the combo hops to `second`. + * + * The floor is read from the translated stream's own `message_start` frame, which is where a + * Claude client — and Paseo's context meter — reads it. + */ +async function comboFailoverFloor(second: { + provider: string; + model: string; + adapter: string; + frames: string; + /** Native ids the destination advertises, for a target the caller names by alias. */ + models?: string[]; + modelAliases?: Record; + preserveReasoningContentModels?: string[]; +}): Promise<{ published: number; secondBodies: Array> }> { + setUpIsolatedHome(); + clearComboSelectionState(); + clearComboTargetCooldowns(); + const firstBodies: Array> = []; + const first = Bun.serve({ + hostname: "127.0.0.1", + port: 0, + async fetch(req) { + firstBodies.push(await req.json() as Record); + return Response.json({ error: { message: "fixture outage" } }, { status: 503 }); + }, + }); + const secondBodies: Array> = []; + const secondServer = Bun.serve({ + hostname: "127.0.0.1", + port: 0, + async fetch(req) { + secondBodies.push(await req.json() as Record); + return new Response(second.frames, { headers: { "content-type": "text/event-stream" } }); + }, + }); + const loopback = (url: URL): string => url.toString().replace(/\/$/, ""); + const config = { + port: 0, + defaultProvider: "first", + providers: { + first: { adapter: "openai-chat", baseUrl: `${loopback(first.url)}/v1`, apiKey: "k", allowPrivateNetwork: true }, + [second.provider]: { + adapter: second.adapter, + baseUrl: second.adapter === "anthropic" ? loopback(secondServer.url) : `${loopback(secondServer.url)}/v1`, + apiKey: "k", + allowPrivateNetwork: true, + ...(second.models ? { models: second.models } : {}), + ...(second.modelAliases ? { modelAliases: second.modelAliases } : {}), + ...(second.preserveReasoningContentModels + ? { preserveReasoningContentModels: second.preserveReasoningContentModels } + : {}), + }, + }, + combos: { + pair: { + strategy: "failover", + targets: [{ provider: "first", model: "m1" }, { provider: second.provider, model: second.model }], + }, + }, + } as unknown as OcxConfig; + try { + const response = await handleClaudeMessages(new Request("http://localhost/v1/messages", { + method: "POST", + headers: { "content-type": "application/json" }, + body: JSON.stringify({ model: "combo/pair", max_tokens: 128, stream: true, messages: COMBO_MESSAGES }), + }), config, { model: "", provider: "" }, { requestId: `combo-floor-${crypto.randomUUID()}`, start: Date.now() }); + expect(response.status).toBe(200); + const text = await response.text(); + const startFrame = text.slice(text.indexOf("event: message_start")); + const published = (JSON.parse(startFrame.slice(startFrame.indexOf("data: ") + 6, startFrame.indexOf("\n\n"))) + .message.usage.input_tokens) as number; + // The hop really happened: the first target refused once, the second served once. + expect(firstBodies).toHaveLength(1); + expect(secondBodies).toHaveLength(1); + return { published, secondBodies }; + } finally { + first.stop(true); + secondServer.stop(true); + clearComboSelectionState(); + clearComboTargetCooldowns(); + restoreIsolatedHome(); + } +} + +test("a combo failover publishes the floor of the target that answered, not the ingress pick", async () => { + // The ingress route names the combo's first target; the physical send names whichever target + // answered. Those are different bodies when the targets sit on different wires, and a memo read + // from the ingress pick prices the wrong one: here target A is a Chat wire that would drop the + // replayed thinking entirely, while target B is an Anthropic wire that forwards it verbatim. + // A floor left at the ingress pick therefore understates a prompt B really received — the + // mirror image of the over-count this file pins, and just as wrong for a context meter. + const { published, secondBodies } = await comboFailoverFloor({ + provider: "b", model: "m2", adapter: "anthropic", frames: ANTHROPIC_SSE_FRAMES, + }); + // B's body is the caller's, replayed thinking and signature included. + const forwarded = JSON.stringify(secondBodies[0]!.messages); + expect(forwarded).toContain("replayed reasoning"); + expect(forwarded).toContain("S".repeat(64)); + + const nativeWire = estimateClaudeRequestTokens({ messages: COMBO_MESSAGES }, "combo/pair", CLAUDE_NATIVE_THINKING); + const chatWire = estimateClaudeRequestTokens({ messages: COMBO_MESSAGES }, "combo/pair", { text: false, signature: false, redacted: false }); + // The ablation: the ingress pick is the Chat target, whose projection prices this body at a + // rounding error next to what B received. Holding the floor above it is what a memo keyed on + // the settled wire buys. + expect(chatWire).toBeLessThan(nativeWire / 100); + expect(published).toBeGreaterThan(chatWire * 20); + expect(published).toBeLessThan(nativeWire * 1.5); + expect(published).toBeGreaterThan(nativeWire * 0.5); +}, { timeout: SERVER_BUDGET_MS }); + +test("a combo failover to a registry provider prices its merged preserve list", async () => { + // `preserveReasoningContentModels` is not a config-row field on most providers: it is merged + // in from the registry by `routedProviderConfig`. A dispatch that reads the raw row therefore + // prices a preserve-listed model as if its reasoning were dropped, which is the under-count + // this pair of cases exists to prevent. `moonshot` is a real registry entry whose endpoint a + // user may override, so a loopback row reaches the same merge path production does. + const { published, secondBodies } = await comboFailoverFloor({ + provider: "moonshot", model: "kimi-k3", adapter: "openai-chat", frames: CHAT_SSE_FRAMES, + }); + // The merged list is what made the adapter serialize the replayed text at all. + const forwarded = JSON.stringify(secondBodies[0]!.messages); + expect(forwarded).toContain("reasoning_content"); + expect(forwarded).toContain("replayed reasoning"); + // The signature has no Chat representation, so it is the one field the list does not buy. + expect(forwarded).not.toContain("S".repeat(64)); + + const textKept = estimateClaudeRequestTokens({ messages: COMBO_MESSAGES }, "combo/pair", { text: true, signature: false, redacted: false }); + const textDropped = estimateClaudeRequestTokens({ messages: COMBO_MESSAGES }, "combo/pair", { text: false, signature: false, redacted: false }); + expect(textKept).toBeGreaterThan(textDropped * 20); + expect(published).toBeGreaterThan(textDropped * 20); + expect(published).toBeLessThan(textKept * 1.5); + expect(published).toBeGreaterThan(textKept * 0.5); +}, { timeout: SERVER_BUDGET_MS }); + +test("a combo target named by alias prices the preserve list under its resolved id", async () => { + // A combo target may name a model by alias (`am`), and the Chat adapter resolves that to the + // provider's native id before it decides whether the wire serializes replayed thinking — its + // preserve list holds native ids, and matching is exact. An attempt that records the alias + // therefore reads the preserve list under a name that is not in it, prices the replayed text as + // dropped, and understates a prompt the upstream really received. The attempt row has to carry + // the id the adapter will actually send, which is what the routing result already holds. + const { published, secondBodies } = await comboFailoverFloor({ + provider: "aliased", + model: "am", + adapter: "openai-chat", + frames: CHAT_SSE_FRAMES, + models: ["aliased-model"], + modelAliases: { "aliased-model": "am" }, + preserveReasoningContentModels: ["aliased-model"], + }); + // The wire received the resolved id, and with it the replayed text the preserve list buys. + expect(secondBodies[0]!.model).toBe("aliased-model"); + const forwarded = JSON.stringify(secondBodies[0]!.messages); + expect(forwarded).toContain("reasoning_content"); + expect(forwarded).toContain("replayed reasoning"); + + const textKept = estimateClaudeRequestTokens({ messages: COMBO_MESSAGES }, "combo/pair", { text: true, signature: false, redacted: false }); + const textDropped = estimateClaudeRequestTokens({ messages: COMBO_MESSAGES }, "combo/pair", { text: false, signature: false, redacted: false }); + expect(textKept).toBeGreaterThan(textDropped * 20); + // The alias must not cost the projection the preserve list's verdict. + expect(published).toBeGreaterThan(textDropped * 20); + expect(published).toBeLessThan(textKept * 1.5); + expect(published).toBeGreaterThan(textKept * 0.5); +}, { timeout: SERVER_BUDGET_MS }); +