Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 14 additions & 4 deletions docs-site/src/content/docs/guides/providers.md
Original file line number Diff line number Diff line change
Expand Up @@ -296,13 +296,23 @@ and 30-day observations are local usage estimates, not live remaining quota or b
authoritative limit event is shown only after the upstream reports a concrete limit event (and, when
provided, its reset).

The provider pins the current Zen Go lineup (25 model ids) as its static catalog, with live
`/v1/models` discovery authoritative on the canonical host — ids outside the trusted set are
quarantined rather than routed. Transports follow the official endpoint table: Qwen and MiniMax
models go over Anthropic Messages, `gpt-5.6-luna` and `grok-4.5` over OpenAI Responses, and the
The provider pins 28 verified Zen Go model ids, including `deepseek-v4.1-flash`,
`glm-5.3-flash`, and `qwen3.8-flash`, as
its static catalog. Live `/v1/models` discovery is authoritative on the canonical host; ids
outside the trusted set are quarantined rather than routed. Transports follow the official
endpoint table: Qwen and MiniMax models go over Anthropic Messages, `gpt-5.6-luna` and `grok-4.5`
over OpenAI Responses, and the
remaining models over OpenAI Chat Completions. These trust facts attach only to the canonical
`https://opencode.ai/zen/go/v1` destination; a same-named custom provider keeps its own behavior.

OpenCode Go requires a session identifier for each conversation. CodexCommander derives an opaque
`x-opencode-session` from the Codex task or session header and keeps it stable through that
conversation, including subagent turns. Claude Code turns use their per-session metadata when available.
An explicit `x-opencode-session` from a Responses, Chat Completions, or Messages client is used when no Codex identity is available. Requests without either identifier receive a distinct temporary
session, and a configured provider header takes precedence. This behavior applies only to the
canonical OpenCode Go destination. The proxy identifies itself with a CodexCommander user agent
unless the provider configuration sets one explicitly.

The built-in preset is key-based, so Add Provider groups it under **Paid**, not account-login
providers, and CodexCommander does not offer an OpenCode Go OAuth flow. It is also separate from both the
**OpenCode** client under Client Apps and the no-key **OpenCode Free** provider. Add Provider search
Expand Down
2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "codexcommander",
"version": "0.1.18",
"version": "0.1.19",
"private": true,
"description": "CodexCommander — universal provider proxy for OpenAI Codex & Claude Code — use any LLM with Codex CLI/App/SDK and Claude Code",
"type": "module",
Expand Down
51 changes: 51 additions & 0 deletions src/providers/opencode-go-transport.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
import { createHash, randomUUID } from "node:crypto";
import { resolveCodexTaskIdentity } from "../codex/task-identity";
import type { CodexCommanderProviderConfig } from "../types";
import { providerMatchesRegistryTransport } from "./registry";

const SESSION_HEADER = "x-opencode-session";
const USER_AGENT_HEADER = "user-agent";

function hasHeader(headers: Record<string, string> | undefined, name: string): boolean {
return Object.keys(headers ?? {}).some(key => key.toLowerCase() === name);
}

function validSession(value: string | null): string | undefined {
return value && value.length <= 512 && !/[\x00-\x20\x7f,]/.test(value) ? value : undefined;
}

/** Give OpenCode Go a stable, opaque lane for each Codex conversation. */
export function resolveOpenCodeGoTransport(
providerName: string,
provider: CodexCommanderProviderConfig,
headers: Headers,
clientMetadata?: unknown,
): CodexCommanderProviderConfig {
// The router may already have pinned this model to Go's Anthropic or Responses
// wire. Verify the canonical key destination against the registry's base adapter,
// while accepting only Go's three documented wire adapters.
if (providerName !== "opencode-go"
|| !["openai-chat", "anthropic", "openai-responses"].includes(provider.adapter)
|| !providerMatchesRegistryTransport(providerName, { ...provider, adapter: "openai-chat" })) return provider;
const hasSession = hasHeader(provider.headers, SESSION_HEADER);
const hasUserAgent = hasHeader(provider.headers, USER_AGENT_HEADER);
if (hasSession && hasUserAgent) return provider;

const outboundHeaders = { ...provider.headers };
if (!hasSession) {
const lane = resolveCodexTaskIdentity(headers, clientMetadata).taskId
?? validSession(headers.get(SESSION_HEADER))
?? randomUUID();
const session = createHash("sha256")
.update("codexcommander/opencode-go/session/v1\0")
.update(lane)
.digest("hex")
.slice(0, 32);
outboundHeaders[SESSION_HEADER] = `ccx_${session}`;
}
if (!hasUserAgent) outboundHeaders["User-Agent"] = "CodexCommander";
return {
...provider,
headers: outboundHeaders,
};
}
38 changes: 24 additions & 14 deletions src/providers/registry.ts
Original file line number Diff line number Diff line change
Expand Up @@ -359,8 +359,9 @@ const THINKING_BUDGET_MODELS = [
];
const OPENCODE_GO_THINKING_BUDGET_MODELS = ["qwen3.5-plus", "qwen3.6-plus", "qwen3.7-max", "qwen3.7-plus", "qwen3.8-max"];
/**
* Pinned last-known-good OpenCode Go lineup (25 ids): the exact id set advertised by
* `GET https://opencode.ai/zen/go/v1/models`, verified 2026-08-05. That endpoint is
* Pinned OpenCode Go lineup (28 ids): the 25-id snapshot from 2026-08-05 plus
* DeepSeek V4.1 Flash, GLM-5.3 Flash, and Qwen3.8 Flash, verified 2026-09-27.
* `GET https://opencode.ai/zen/go/v1/models` is
* existence-only — it returns ids without context/output/pricing metadata — so this list is
* the catalog seed, and the registry-only discovery filter below admits exactly these ids:
* any other model upstream starts (or stops) advertising is quarantined rather than guessed
Expand All @@ -373,16 +374,17 @@ const OPENCODE_GO_THINKING_BUDGET_MODELS = ["qwen3.5-plus", "qwen3.6-plus", "qwe
* Transport split per the official endpoint table (https://opencode.ai/docs/go/#endpoints):
* Qwen and MiniMax rows serve Anthropic Messages (`/zen/go/v1/messages`), GPT-5.6 Luna and
* Grok 4.5 serve OpenAI Responses (`/zen/go/v1/responses`), and the remaining rows serve
* OpenAI Chat Completions (`/zen/go/v1/chat/completions`). The Anthropic subset is a hard
* OpenAI Chat Completions (`/zen/go/v1/chat/completions`). DeepSeek V4.1 Flash is included
* on that Chat wire. The Anthropic subset is a hard
* wire pin owned by types.ts (OPENCODE_GO_ANTHROPIC_WIRE_MODEL_IDS); the two OpenAI-shaped
* wires are registry `modelWireDefaults` on the entry below.
*/
const OPENCODE_GO_MODELS = [
"minimax-m3", "minimax-m2.7", "minimax-m2.5",
"kimi-k3", "kimi-k2.7-code", "kimi-k2.6", "kimi-k2.5",
"glm-5.2", "glm-5.1", "glm-5",
"deepseek-v4-pro", "deepseek-v4-flash",
"qwen3.8-max", "qwen3.7-max", "qwen3.7-plus", "qwen3.6-plus", "qwen3.5-plus",
"glm-5.3-flash", "glm-5.2", "glm-5.1", "glm-5",
"deepseek-v4-pro", "deepseek-v4-flash", "deepseek-v4.1-flash",
"qwen3.8-max", "qwen3.8-flash", "qwen3.7-max", "qwen3.7-plus", "qwen3.6-plus", "qwen3.5-plus",
"mimo-v2-pro", "mimo-v2-omni", "mimo-v2.5-pro", "mimo-v2.5",
"hy3", "hy3-preview",
"gpt-5.6-luna", "grok-4.5",
Expand All @@ -392,6 +394,7 @@ const OPENCODE_GO_RESPONSES_WIRE_MODELS = ["gpt-5.6-luna", "grok-4.5"];
// ladder on the `xai` entry); GPT-5.6 Luna serves the OpenAI API GPT-5.6 ladder.
const OPENCODE_GO_GROK45_REASONING_EFFORTS = ["low", "medium", "high"];
const DEEPSEEK_THINKING_MODELS = ["deepseek-v4-pro", "deepseek-v4-flash"];
const OPENCODE_GO_DEEPSEEK_THINKING_MODELS = [...DEEPSEEK_THINKING_MODELS, "deepseek-v4.1-flash"];
const OPENCODE_FREE_DEEPSEEK_MODELS = ["deepseek-v4-flash-free"];
/*
* Zen free models that reject `image_url` upstream (#1043, and the reproducible
Expand Down Expand Up @@ -1159,7 +1162,7 @@ export const PROVIDER_REGISTRY: readonly ProviderRegistryEntry[] = [
jawcodeBundle: "opencode-go", note: "GLM, DeepSeek, Kimi, Qwen, MiMo…",
models: [...OPENCODE_GO_MODELS],
// Live /v1/models is the authoritative lineup; the static list above is the last-good
// fallback seed. The registry-only filter quarantines any id outside the trusted 25, and
// fallback seed. The registry-only filter quarantines any id outside the trusted set, and
// `preserveCustomDestination` keeps the whole trusted transport registry (this filter, the
// wire defaults, and every registry metadata backfill) attached to the canonical Zen Go
// host only — a same-named row pointed elsewhere is a custom provider and gets none of it.
Expand All @@ -1177,20 +1180,27 @@ export const PROVIDER_REGISTRY: readonly ProviderRegistryEntry[] = [
modelWireDefaults: {
...Object.fromEntries(OPENCODE_GO_RESPONSES_WIRE_MODELS.map(id => [id, "openai-responses"])),
},
// Zen Go context windows not covered by the generated jawcode bundle (qwen3.8-max and
// gpt-5.6-luna have no bundle row yet): official data pages
// https://opencode.ai/data/qwen/qwen3-8-max (1M) and
// https://opencode.ai/data/openai/gpt-5-6-luna (1.1M — the OpenAI API value 1,050,000).
// Zen Go context windows not covered by the generated jawcode bundle:
// https://opencode.ai/data/deepseek/deepseek-v4-1-flash,
// https://stats.opencode.ai/data/zhipu/glm-5-3-flash, and
// https://stats.opencode.ai/data/qwen/qwen3-8-flash (all 1M), plus the
// previously pinned Qwen3.8 Max and GPT-5.6 Luna values.
modelContextWindows: {
"kimi-k3": KIMI_K3_STANDARD_CONTEXT_WINDOW,
"deepseek-v4.1-flash": 1_000_000,
"glm-5.3-flash": 1_000_000,
"qwen3.8-max": 1_000_000,
"qwen3.8-flash": 1_000_000,
"gpt-5.6-luna": 1_050_000,
},
// qwen3.8-max (text/image/video) and gpt-5.6-luna (text/image/pdf) are multimodal upstream;
// the jawcode type can only represent text+image, so video/pdf stay source facts.
modelInputModalities: {
"kimi-k3": ["text", "image"],
"deepseek-v4.1-flash": ["text", "image"],
"glm-5.3-flash": ["text", "image"],
"qwen3.8-max": ["text", "image"],
"qwen3.8-flash": ["text", "image"],
"gpt-5.6-luna": ["text", "image"],
},
modelReasoningEfforts: {
Expand All @@ -1202,15 +1212,15 @@ export const PROVIDER_REGISTRY: readonly ProviderRegistryEntry[] = [
"gpt-5.6-luna": OPENAI_API_GPT56_REASONING_EFFORTS,
...Object.fromEntries(OPENCODE_GO_THINKING_TOGGLE_MODELS.map(id => [id, THINKING_TOGGLE_EFFORTS])),
...Object.fromEntries(OPENCODE_GO_THINKING_BUDGET_MODELS.map(id => [id, THINKING_BUDGET_EFFORTS])),
...Object.fromEntries(DEEPSEEK_THINKING_MODELS.map(id => [id, deepseekThinkingEffortsFor(id)])),
...Object.fromEntries(OPENCODE_GO_DEEPSEEK_THINKING_MODELS.map(id => [id, deepseekThinkingEffortsFor(id)])),
},
modelDefaultReasoningEfforts: { "kimi-k3": "max" },
// glm-5.2 uses identity labels now that `max` is a native Codex level (no alias map);
// the thinking-toggle map is a REAL wire alias (effort -> enabled/disabled) and stays.
modelReasoningEffortMap: {
"kimi-k3": KIMI_CODING_K3_REASONING_EFFORT_MAP,
...Object.fromEntries(OPENCODE_GO_THINKING_TOGGLE_MODELS.map(id => [id, THINKING_TOGGLE_MAP])),
...Object.fromEntries(DEEPSEEK_THINKING_MODELS.map(id => [id, deepseekReasoningMapFor(id)])),
...Object.fromEntries(OPENCODE_GO_DEEPSEEK_THINKING_MODELS.map(id => [id, deepseekReasoningMapFor(id)])),
},
thinkingToggleModels: OPENCODE_GO_THINKING_TOGGLE_MODELS,
thinkingBudgetModels: THINKING_BUDGET_MODELS,
Expand All @@ -1230,7 +1240,7 @@ export const PROVIDER_REGISTRY: readonly ProviderRegistryEntry[] = [
noPenaltyModels: ["kimi-k3", "kimi-k2.7-code", "kimi-k2.7-code-highspeed"],
autoToolChoiceOnlyModels: ["kimi-k2.7-code", "kimi-k2.7-code-highspeed"],
// Issue #78: DeepSeek V4 thinking mode requires reasoning_content replay on tool-call turns.
preserveReasoningContentModels: ["glm-5.2", "kimi-k3", "kimi-k2.7-code", "kimi-k2.7-code-highspeed", ...DEEPSEEK_THINKING_MODELS],
preserveReasoningContentModels: ["glm-5.2", "kimi-k3", "kimi-k2.7-code", "kimi-k2.7-code-highspeed", ...OPENCODE_GO_DEEPSEEK_THINKING_MODELS],
},
{
id: "neuralwatt",
Expand Down
4 changes: 4 additions & 0 deletions src/server/chat-completions.ts
Original file line number Diff line number Diff line change
Expand Up @@ -167,6 +167,10 @@ async function handleChatCompletionsWithBudget(
const value = req.headers.get(name);
if (value) headers.set(name, value);
}
// Keep an explicit Go session through the Chat-to-Responses bridge. Only the
// canonical Go destination turns it into an upstream header.
const goSession = req.headers.get("x-opencode-session");
if (goSession) headers.set("x-opencode-session", goSession);
// Prefer main ChatGPT auth so OpenAI-backed sidecars remain reachable on routed turns.
if (!directRoute) {
// This enrichment is optional for routed/non-main providers. If native main
Expand Down
16 changes: 15 additions & 1 deletion src/server/claude-messages.ts
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,7 @@ import {
import { clearableDeadline, idleDeadline } from "../lib/abort";
import { estimateTokens } from "../lib/token-estimate";
import { NoEligiblePolicyCandidateError, routeModel } from "../router";
import { providerMatchesRegistryTransport } from "../providers/registry";
import { evidenceFromBody } from "../routing/request-evidence";
import { resolveWireProtocolOverride } from "./adapter-resolve";
import type { CodexCommanderConfig } from "../types";
Expand Down Expand Up @@ -706,11 +707,15 @@ async function handleClaudeMessagesWithBudget(
// bodies: it 400s on sampling params ("Unsupported parameter: max_output_tokens",
// verified live 2026-07-11). Strip them for that route; routed providers keep them.
let nativeRoute = false;
let openCodeGoRoute = false;
try {
const route = routeModel(config, internalBody.model as string, evidenceFromBody(internalBody));
// Settle the wire once so the sampling decision below reads the effective
// adapter rather than the provider-wide default (#404).
route.provider = resolveWireProtocolOverride(route.providerName, route.modelId, route.provider, "anthropic");
openCodeGoRoute = route.providerName === "opencode-go"
&& ["openai-chat", "anthropic", "openai-responses"].includes(route.provider.adapter)
&& providerMatchesRegistryTransport(route.providerName, { ...route.provider, adapter: "openai-chat" });
logCtx.routeDecision = route.routeDecision;
if (route.provider.adapter === "openai-responses") {
nativeRoute = true;
Expand Down Expand Up @@ -754,6 +759,13 @@ async function handleClaudeMessagesWithBudget(
const value = req.headers.get(name);
if (value) headers.set(name, value);
}
// The Messages replay uses Responses routing; carry an explicit Go lane across
// that internal boundary only for the canonical Go destination. Its transport
// hashes the value before sending it upstream.
if (openCodeGoRoute) {
const goSession = req.headers.get("x-opencode-session");
if (goSession) headers.set("x-opencode-session", goSession);
}
// Routed replays need main ChatGPT auth so OpenAI-backed sidecars remain reachable;
// native replays have no caller ChatGPT credential. This enrichment is optional:
// auth-context later rejects a real physical-main selection, while routed/pool
Expand All @@ -766,14 +778,16 @@ async function handleClaudeMessagesWithBudget(
headers.set("chatgpt-account-id", token.chatgptAccountId);
}
}
if (nativeRoute) {
if (nativeRoute || (openCodeGoRoute && !headers.has("x-opencode-session"))) {
// ChatGPT-backend prompt-cache affinity rides the session_id HEADER (codex
// clients always send their session uuid; implementation contract follow-up: body-level
// prompt_cache_key alone still yielded cached_tokens:0). Claude Code never sends
// the header, so synthesize a stable per-session uuid from the same cache key —
// but ONLY for a real per-session key (metadata.user_id). The system-hash fallback
// key is shared across Desktop conversations, and a shared session_id's backend
// semantics are unproven (audit 133 R2#3): body prompt_cache_key only there.
// Go also needs a per-session identity on routed Claude Code turns. Never
// derive it from the shared system fallback or override an explicit Go lane.
if (cacheKeySource === "metadata" && !headers.has("session_id") && typeof internalBody.prompt_cache_key === "string") {
headers.set("session_id", uuidFromHex(internalBody.prompt_cache_key));
}
Expand Down
7 changes: 7 additions & 0 deletions src/server/responses/core.ts
Original file line number Diff line number Diff line change
Expand Up @@ -126,6 +126,7 @@ import {
import { noteProviderCredentialVerified } from "../../providers/credential-verification";
import { shouldAttemptImageTierRetry } from "../image-retry";
import { resolveProviderTransport } from "../../providers/xai-transport";
import { resolveOpenCodeGoTransport } from "../../providers/opencode-go-transport";
import type { WsData } from "../ws-bridge";
import { codexAccountSelectionForTurn, registerTurn, trackStreamLifetime, unregisterTurn, type ActiveTurnLease } from "../lifecycle";
import { redactSecretString } from "../../lib/redact";
Expand Down Expand Up @@ -1789,6 +1790,12 @@ async function handleResponsesInner(
parsed.options.promptCacheKey,
route.providerName === "github-copilot" ? getOAuthCredentialApiBaseUrl(route.providerName) : undefined,
);
route.provider = resolveOpenCodeGoTransport(
route.providerName,
route.provider,
req.headers,
nativeClientMetadata(parsed._rawBody),
);
requestDispatchContext(logCtx, options.abortSignal ?? req.signal, undefined, config.providers[route.providerName]);
const adapterProvider = resolveWireProtocolOverride(route.providerName, route.modelId, route.provider, inboundWire);
const adapter = resolveAdapter(adapterProvider, config.cacheRetention);
Expand Down
2 changes: 1 addition & 1 deletion src/types.ts
Original file line number Diff line number Diff line change
Expand Up @@ -1355,7 +1355,7 @@ export const MODEL_ADAPTER_OVERRIDE_ALLOWED: ReadonlySet<string> = new Set([
*/
export const OPENCODE_GO_ANTHROPIC_WIRE_MODEL_IDS = [
"minimax-m2.5", "minimax-m2.7", "minimax-m3",
"qwen3.5-plus", "qwen3.6-plus", "qwen3.7-max", "qwen3.7-plus", "qwen3.8-max",
"qwen3.5-plus", "qwen3.6-plus", "qwen3.7-max", "qwen3.7-plus", "qwen3.8-max", "qwen3.8-flash",
] as const;

const ANTHROPIC_WIRE_MODELS: Record<string, ReadonlySet<string>> = {
Expand Down
Loading
Loading