Skip to content

[Bug]: client config export sets maxTokens to the 32000 schema stand-in for every model, truncating max-effort thinking turns (and over-reporting grok-4.20) #5828

Description

@thisisjun786

Client or integration

OMO (senpi) through ocx integration client enable --client omo. The same export path serves Pi and OMP.

Area

Client config export (src/clients/config-export.ts, config-export/model-metadata.ts)

Summary

Every model in the exported providers.opencodex block gets maxTokens: 32000. The value comes from outputBudgetFor(context) = min(SCHEMA_REQUIRED_OUTPUT_BUDGET, context). The constant's own comment says it is "a ceiling for schema validity, NOT a claim about any specific model's true maximum".

Clients treat maxTokens as the real output ceiling and send it as max_tokens on every request. That causes two failures:

  1. Too low for thinking models. Claude Opus 5.5 at effort: max on a ~220k-token context spent the whole 32,000-token budget on adaptive thinking. The turn ended stop_reason: max_tokens with no text and no tool call (usage.jsonl: status 200, outputTokens 32000, "Output reached the requested token limit"). The client got an empty assistant message after 319 s. The real Opus 5.5 limit is 128K, and ocx's own model-seeds.ts says so.
  2. Too high for some models. xai/grok-4.20-0309-non-reasoning is exported at 32,000, but its limit is 30,000 (ocx generated/model-metadata.ts).

ocx already has authoritative per-model output limits for most exported models. They are in generated/model-metadata.ts (third column) and the Anthropic registry seeds:

exported id exported ocx metadata
anthropic/claude-opus-5-5, claude-fable-5-1, claude-sonnet-5 32000 128000
anthropic/claude-haiku-4-5 32000 64000
gpt-5.6-{luna,sol,terra}, gpt-6-{luna,sol} (+ --fast) 32000 128000
google/gemini-3.8-flash 32000 65536
ollama-cloud/glm-5.3, glm-5.3-flash 32000 131072
kimi/k3[1m] (kimi-k3) 32000 131072
xai/grok-4.7 32000 500000
xai/grok-4.20-0309-non-reasoning 32000 30000

Expected

  • When the catalog knows a model's output limit, the export uses it, clamped to the exported context as it is today.
  • The 32K stand-in applies only to models with no known limit.
  • A limit known to be lower than 32K is never raised to 32K.

Reproduction

  1. ocx integration client enable --client omo
  2. Inspect ~/.omo/agent/models.json → providers.opencodex.models[*].maxTokens: all 31 entries are 32000.
  3. Run a long max-effort Opus 5.5 turn. It ends max_tokens with empty output.

Version

opencodex 2.65.0

Operating system

Linux x64

Provider and model

anthropic (OAuth) claude-opus-5-5; affects every exported model

Checks

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingpriority: P2Medium: provider/client-specific bug with a workaround, bounded enhancement tied to a tracked issue,toolstool_calls, MCP, web-search / sidecar tools

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions