Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 16 additions & 0 deletions devlog/_plan/260926_claude_input_estimate/000_overview.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
# A Claude input estimate the settled route actually sends

Unit opened 2026-09-26.

- [010_estimation.md](010_estimation.md) — the Messages ingress. **DONE.**

Origin: a Paseo agent on a DeepSeek V4.1 Flash conversation through this proxy drew its context
meter at 221%. The rate implied the agent had run far past the window without compacting, so the
report read as a broken compaction loop. Compaction was healthy; the number above it was not. Same
body, same upstream: the proxy published 432,068 input tokens on `message_start` for a prompt the
upstream billed at 131,907 — **3.28x**.

This is the `#4857` family (the floor `message_start` publishes when the upstream has sent no
confirmed usage before the first frame), but not a recurrence: `#4891` and `#5057` fixed *when* the
floor is used and *whose* count it reports. The floor faithfully reports
`estimateClaudeRequestTokens`, and that estimate was measuring the wrong body.
65 changes: 65 additions & 0 deletions devlog/_plan/260926_claude_input_estimate/010_estimation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,65 @@
# A Claude input estimate the settled route actually sends

Defect: `src/server/claude-messages.ts` `estimateClaudeRequestTokens`. The estimator measured the
body the **caller sent**, while `message_start` publishes it as the floor for the prompt this proxy
**actually forwarded**. On the Anthropic wire those are the same body. On every other wire they are
not, and the gap is whatever the target adapter drops.

Measured on the live path with a real Claude Code conversation replayed byte-identical (260
messages, 23 tools, 1,734,433 B) to `thehive/deepseek-ai/deepseek-v4.1-flash`
(`adapter: openai-chat`): `message_start` **432,068** against the upstream's `message_delta`
**131,907** — **3.28x**. The contract the estimator's own doc cites allows `>2x` drift
(`devlog/_fin/260711_claude_inbound/040_phase4_hardening.md` §3); this is outside it.

Why the two differ. Claude Code replays its own thinking blocks, and they dominate a long body:
80 blocks, 550,930 thinking chars and 750,284 `signature` chars. In the captured request the
thinking JSON is **78.8%** of all message JSON and the base64 `signature` is **56.7%** of the
thinking JSON. The `openai-chat` wire forwards that text only when the model is listed in
`preserveReasoningContentModels`, and it has no `signature` field at all — zero `signature`
references exist anywhere under `src/adapters/openai-chat*`. The `thehive` provider config declares
no reasoning policy keys, so it drops both, and upstreams do not bill replayed reasoning they never
receive. The estimator was counting 432,068 tokens of a prompt whose real size was 131,907.

The gap is a property of the route, not a constant, so no divisor can absorb it: a route that does
replay thinking must keep counting it. Char-per-token calibration and the `thehive/` alias prefix
were both ruled out as causes — CJK is 0.07% of the body, and moving 4 to 3.5 chars/token is ±14%,
while dropping the alias alone makes the number 14% *worse*.

Change:

- New leaf `src/lib/claude-request-projection.ts`: `ClaudeThinkingProjection`, the native
`{text:true, signature:true}`, `projectBlock`, and `projectClaudeRequest` — a pure, idempotent
projection that returns message content with the blocks a wire does not carry emptied. It never
mutates its input and survives a message whose content is not an array.
- `src/adapters/openai-chat/messages.ts`: new exported `openAIChatSerializesThinking(provider,
modelId)`, the single source of truth for whether that wire carries a replayed thinking block.
The `messagesToChatFormat` conversion reads the same answer once per request
(`wireSerializesThinking`) instead of re-deriving it per message, so the estimator and the
serializer cannot drift apart — they are one rule.
- `src/server/claude-messages.ts`: `estimateClaudeRequestTokens` takes an optional
`thinking: ClaudeThinkingProjection` defaulting to the native pair, and projects the messages
before measuring. `claudeRequestTokenFloor` passes the projection for the route the turn settled
on (`settledRoute`, recorded where routing resolves), and `handleClaudeCountTokens` resolves its
own route through the read-only `previewRouteModel`. A route that is not `openai-chat` keeps the
native projection, so the Anthropic-native lane stays byte-exact and pre-existing two-argument
callers keep their behavior.

The measurement only. The caller's body is never rewritten on its way to the adapter — the
projection exists to price the prompt, not to alter it, and every reader of the floor shares the one
estimate, so no number is double-counted.

Rejected: excluding replayed thinking unconditionally (wrong for the native lane and any adapter
that does replay it); a `billsReplayedReasoning?: boolean` flag on `ProviderAdapter` (states a
billing policy on an interface whose business is wire shape, and cannot express "text yes,
signature no", which is the shape that actually occurs).

Tests (new `tests/claude-integration/claude-estimate-projection.test.ts`): a replayed-thinking body
projects away exactly the unserialized fields; `projectClaudeRequest` is pure, idempotent, and keeps
a message it emptied; a body with no thinking is projection-invariant; and end-to-end, the
`message_start` floor for a real captured body lands within the confirmed usage instead of 3.28x
above it. Verified to fail when the projection is disabled — the end-to-end case reports 6946 where
it expects <42, so the assertion is load-bearing rather than decorative.

Live verification, same captured body against the real upstream (104,318 / 131,907 = **0.79x**, and
`message_delta` byte-identical at 131,907, so the upstream result is untouched and only the
published estimate moved). The Paseo meter that opened this unit reads 240% before and 58% after.
35 changes: 31 additions & 4 deletions src/adapters/openai-chat/messages.ts
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@ import { EMPTY_TOOL_OUTPUT_ANNOTATION, isWhitespaceOnlyTextPartArray } from "../
import { identifyRoutedModel } from "../identity";
import { buildNonOpenAIToolCatalogNudgeForTools, shouldInjectNonOpenAIToolCatalogNudge } from "../tool-catalog-nudge";
import { peekReasoningForCall } from "../../responses/reasoning-replay-cache";
import type { ClaudeThinkingProjection } from "../../lib/claude-request-projection";
import { inlineDocumentDataUrl } from "../../responses/inline-document";
import type { OcxAssistantMessage, OcxContentPart, OcxParsedRequest, OcxProviderConfig, OcxTextContent, OcxThinkingContent, OcxToolCall } from "../../types";
import { modelInList, namespacedToolName } from "../../types";
Expand Down Expand Up @@ -61,10 +62,36 @@ export function toolResultImageChatParts(content: string | OcxContentPart[]): un
return parts;
}

/**
* Whether this wire serializes a replayed thinking block for the given model.
*
* The single source of truth for that question on the Chat wire: the assistant branch below uses
* it to decide `reasoning_content`, and the Messages ingress uses it to decide how much of a
* prompt is worth counting. Two copies of this rule would drift, and the count would then
* describe a body this adapter does not send.
*/
export function openAIChatSerializesThinking(
provider: OcxProviderConfig,
modelId: string,
): ClaudeThinkingProjection {
return {
text: modelInList(provider.preserveReasoningContentModels, modelId),
// The wire has no `signature` field — base64 replay tokens are Anthropic-wire-only.
signature: false,
// A `redacted_thinking` block is opaque provider data with no Chat representation: the
// assistant branch below reads only `type: "thinking"` parts, so the encrypted form is
// dropped whatever the preserve list says.
redacted: false,
};
}

export function messagesToChatFormat(parsed: OcxParsedRequest, provider: OcxProviderConfig): unknown[] {
const out: unknown[] = [];
const { context, options } = parsed;
const replayCacheScope = parsed._reasoningReplayScope;
// One question, one answer for the whole conversion: does this wire carry replayed thinking
// back for this model? Asked once so a long history cannot pay for it per message.
const wireSerializesThinking = openAIChatSerializesThinking(provider, parsed.modelId).text;

interface PendingToolCall { id: string; name: string }
let pendingToolCalls: PendingToolCall[] = [];
Expand Down Expand Up @@ -225,7 +252,7 @@ export function messagesToChatFormat(parsed: OcxParsedRequest, provider: OcxProv
if (
reasoningContent.length === 0
&& (toolCalls.length > 0 || thinkingParts.length > 0)
&& modelInList(provider.preserveReasoningContentModels, parsed.modelId)
&& wireSerializesThinking
) {
const cached = toolCalls
.map(tc => (tc.id ? peekReasoningForCall(tc.id, replayCacheScope) : undefined))
Expand All @@ -248,7 +275,7 @@ export function messagesToChatFormat(parsed: OcxParsedRequest, provider: OcxProv
reasoningContent = " ";
}
}
if (reasoningContent.length > 0 && modelInList(provider.preserveReasoningContentModels, parsed.modelId)) {
if (reasoningContent.length > 0 && wireSerializesThinking) {
// MiniMax's interleaved-thinking contract requires the structured
// reasoning_details array back on the next turn; a reasoning_content
// string is the native-format pass-back the docs mark unsupported.
Expand Down Expand Up @@ -302,7 +329,7 @@ export function messagesToChatFormat(parsed: OcxParsedRequest, provider: OcxProv
flushPendingToolCalls();
const name = safeToolName(msg.toolName);
const cachedReasoning =
toolCallId && modelInList(provider.preserveReasoningContentModels, parsed.modelId)
toolCallId && wireSerializesThinking
? peekReasoningForCall(toolCallId, replayCacheScope)
: undefined;
// Same fallback as the main-assistant path: never emit a bare orphan
Expand All @@ -316,7 +343,7 @@ export function messagesToChatFormat(parsed: OcxParsedRequest, provider: OcxProv
// falsy hit as a miss so the placeholder still fires.
const orphanReasoning =
cachedReasoning
|| (modelInList(provider.preserveReasoningContentModels, parsed.modelId)
|| (wireSerializesThinking
&& modelInList(provider.requiresReasoningPlaceholderModels ?? provider.preserveReasoningContentModels, parsed.modelId)
? " "
: undefined);
Expand Down
101 changes: 101 additions & 0 deletions src/lib/claude-request-projection.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,101 @@
/**
* Project a Messages request body onto the content a settled route actually serializes.
*
* `estimateClaudeRequestTokens` measures the caller's body as JSON. That is exactly right for the
* Anthropic-native wire, where a replayed `thinking` block — signature included — is forwarded
* verbatim. A routed wire may serialize far less, and on the Chat wire it does: replayed thinking
* becomes an optional `reasoning_content` string, the signature is never sent at all
* (`src/adapters/openai-chat/messages.ts` has no `signature` reference), and the text is dropped
* outright unless the model is on the provider's `preserveReasoningContentModels` list.
*
* The gap that opens is not a rounding error. On a captured 260-message Claude Code turn, replayed
* thinking blocks were 78.8% of the body's JSON characters and 56.7% of those characters were
* base64 signatures — bytes that are not prompt text under any tokenizer. Counting them made the
* published `message_start.usage.input_tokens` 3.28x the count the upstream actually reported
* (#4857 + the Paseo context meter it feeds), well past the >2x drift bound this estimator is held
* to (devlog 260711_claude_inbound 040 §3).
*
* The projection is therefore applied to the ESTIMATE only. The body the caller sent is never
* rewritten; this is a measurement that describes the route, not a transformation of the request.
*/

/** Which replayed reasoning blocks and fields a wire serializes. */
export interface ClaudeThinkingProjection {
/** Serialize `thinking.thinking`, the model's own replayed text. */
text: boolean;
/**
* Serialize `thinking.signature`, the provider's base64 replay token. Real prompt size for
* wires that carry it; pure overhead for wires that do not.
*/
signature: boolean;
/**
* Serialize a `redacted_thinking` block, whose `data` is an opaque provider blob. It rides
* alongside `thinking` rather than inside it: a wire can reconstruct the model's reasoning
* without carrying the encrypted form, and the Chat wire does exactly that.
*/
redacted: boolean;
}

/**
* The Anthropic-native wire forwards a replayed thinking block verbatim, signature included, so
* nothing is projected away. This is the default: an unknown route keeps the measured body.
*/
export const CLAUDE_NATIVE_THINKING: ClaudeThinkingProjection = { text: true, signature: true, redacted: true };

interface ProjectableBody {
system?: unknown;
messages?: unknown;
tools?: unknown;
}

/**
* One content block as the given wire would carry it.
*
* `undefined` means the wire sends nothing for it, and the caller drops the entry. A `thinking`
* block that keeps its text but not its signature is returned as a copy rather than edited: the
* block object is shared with the outbound request builder, which must stay untouched.
*/
function projectBlock(block: unknown, thinking: ClaudeThinkingProjection): unknown {
if (!block || typeof block !== "object") return block;
if (!("type" in block)) return block;
if (block.type === "redacted_thinking") return thinking.redacted ? block : undefined;
if (block.type !== "thinking") return block;
if (!thinking.text) return undefined;
if (thinking.signature || !("signature" in block)) return block;
// Copy without the signature: the caller's block object is shared with the outbound request
// builder, so it must never be edited in place.
const { signature: _dropped, ...rest } = block;
return rest;
}

/**
* A copy of `raw` whose message content carries only the replayed thinking this route serializes.
*
* Pure and idempotent; the input is never mutated. Blocks are dropped only inside array-valued
* message `content`, which is the sole protocol position a replayed thinking block occupies.
* A message whose content array empties out is left as an empty array rather than deleted: the
* adapter's own "nothing left to send" rule is keyed on text, tool calls and reasoning together,
* and re-deriving it here would be a second copy of that rule to keep in sync.
*/
export function projectClaudeRequest(
raw: ProjectableBody,
thinking: ClaudeThinkingProjection,
): ProjectableBody {
if (thinking.text && thinking.signature && thinking.redacted) return raw;
if (!Array.isArray(raw.messages)) return raw;
const messages = raw.messages as unknown[];
return {
...raw,
messages: messages.map(message => {
if (!message || typeof message !== "object") return message;
if (!("content" in message) || !Array.isArray(message.content)) return message;
const content = message.content as unknown[];
return {
...message,
content: content
.map(block => projectBlock(block, thinking))
.filter(block => block !== undefined),
};
}),
};
}
Loading
Loading