Skip to content

docs(bench): re-run every llms-benchmark row under oddyssey 1.13.0 - #637

Merged
using-system merged 18 commits into
mainfrom
docs/llms-benchmark-1-13-0
Sep 20, 2026
Merged

using-system merged 18 commits into
mainfrom
docs/llms-benchmark-1-13-0

Conversation

@using-system

@using-system using-system commented Sep 20, 2026 •

Copy link
Copy Markdown
Owner

Closes #636

Every row of .llms-benchmark/README.md re-measured under oddyssey 1.13.0 (e7fd9fa, MCP server oddyssey-mcp==1.13.0) with the launch-llms-benchmark protocol as it stands on this branch: every model on its CLI run twice, each run graded on its own, the better run in the table (more confirmed, then cheaper, then shorter). Thirteen rows re-run; anthropic/claude-fable-5.1 was not run on the maintainer's word (no fable credit left) and keeps its 1.12.0 row, marked ⚠︎ provisional. 26 launches, 2026-09-19 20:39 UTC to 2026-09-20 06:34 UTC, sequentially, the store reset and the demo stack recreated before each.

Proposed ranking

The rank weighs findings, cost and duration together, cost and duration the heavier - the maintainer's weighting for this campaign: a wide report no longer carries a slow or dear run.

  1. z-ai/glm-5.3-flashx (opencode) - 11 / 13 in 17m09s for 0.33 USD: eleven confirmed for a third of a dollar, the best balance of the three axes; two misses. Was feat!: reposition as an odd toolbox with otel instrumentation planning #3.
  2. openai/gpt-5.6-luna (copilot) - 7 / 8 in 6m29s for 0.11 USD: the second-cheapest run, the second-fastest; seven confirmed. Was ci(mcp-server): add lint, unit-test, and mcp-client integration jobs #4.
  3. openai/gpt-5.6-terra (copilot) - 7 / 8 in 5m58s for 0.88 USD: the fastest run of the table, the root observing without its subagent; eight times luna's cost for the same count. Was docs: add the oddyssey banner to the readme #6.
  4. openai/gpt-5.6-sol (copilot) - 12 / 13 in 9m32s for 1.41 USD: twelve confirmed in under ten minutes, one miss; 1.41 USD puts it behind the three cheaper rows. Was feat(agents): harden both agents into true experts with supporting skills #7 at 3.12 USD.
  5. google/gemini-3.7-flash (opencode) - 8 / 9 in 10m17s for 1.08 USD: sol's wall clock with four confirmed fewer, a little cheaper. Was feat(mcp): drive docker directly without a compose file #5.
  6. z-ai/glm-5.3 (opencode) - 17 / 19 in 19m44s for 1.39 USD: the widest count, two misses, twenty minutes. Was feat: bootstrap oddyssey as an apm package for observability-driven development #1.
  7. deepseek/deepseek-v4.1-flash (opencode) - 16 / 16 in 29m05s for 0.16 USD: exact, the cheapest confirmed finding of the table (0.010 USD), and twenty-nine minutes - glm-5.3 does as much ten minutes faster. Two further launches on the maintainer's request (below) confirmed the wall clock is the model's under 1.13.0, not the provider's. Was feat: grafana proxy routing, observe-local-run agent, simplified readme #2.
  8. google/gemini-3.8-flash (opencode) - 12 / 12 in 20m39s for 2.29 USD: exact; twenty-one minutes and the fourth-dearest run (both its runs lost an observe-run session to a provider 400 and drove twice). Was feat(mcp): add odd_stack_reset tool #8.
  9. qwen/qwen3.8-max-0902 (opencode) - 15 / 16 in 30m25s for 1.50 USD: wide, one miss, thirty minutes. Was fix(mcp): match the no-such-container error case-insensitively #13 at 66 minutes.
  10. anthropic/claude-opus-5 (claude) - 17 / 17 in 19m31s for 6.15 USD: exact and as wide as glm-5.3, alone on two findings; 6.15 USD, 0.362 per finding, is the price that drops it here. Was feat(agents)!: generalize observe-run to any observability backend #9.
  11. anthropic/claude-fable-5.1 (claude) ⚠︎ - the 1.12.0 row, 17 / 17 in 17m08s for 7.55 USD, not re-run; after opus-5 on its own figures, provisional until re-run. Was ci(release): auto version bump, github release, and pypi publish #10.
  12. z-ai/glm-5.3-flash (opencode) - 7 / 8 in 32m15s for 0.10 USD: the cheapest run, thirty-two minutes for seven confirmed, one miss, after a first attempt its provider would not carry. Was ci(release): git-history-driven release flow (git-cliff) with approved release pr #11.
  13. anthropic/claude-sonnet-5 (claude) - 6 / 6 in 13m55s for 3.37 USD: exact, the thinnest, the dearest finding of the table (0.561 USD). Was chore(release): 0.1.1 #14.
  14. qwen/qwen3.8-27b (opencode) - 11 / 11 in 40m25s for 1.62 USD: exact and wide, alone on two telemetry findings, and forty minutes, the longest run of the table. Was chore(release): 0.1.0 #12.

Campaign facts

  • CLIs: opencode 1.18.31 (the official installer's binary, --variant medium), Claude Code 2.1.278 (--effort medium, --permission-mode bypassPermissions), GitHub Copilot CLI 1.0.86 (--effort medium, --allow-all --no-ask-user). Mission byte-identical across the 26 launches, 669 characters (see the amendments).
  • Package: apm install --global --target claude from HEAD before the first run (user-scope bodies identical to .apm/), the user's MCP pin bumped from oddyssey-mcp==1.12.2 to 1.13.0; apm install --target opencode|copilot into the repository before each of those runs, the tree put back after; the three skill copies diffed identical to .apm/skills/ before every launch.
  • Costs: opencode rows carry the provider's recorded figure, reconstructed to the cent from OpenRouter's pricing (deepseek at its base tier); copilot rows the OpenAI list price (uncached and cache-write at the input rate, cache-read at the cached rate), premium requests and AIU beside it; claude rows total_cost_usd (list), reconstructed to the cent from the transcripts - no run carried a claude-haiku-4-5 key.
  • Three runs were stopped by their PID and one was void: deepseek run 2 (49 min, report still generating, 20 min past run 1's total), qwen3.8-27b attempt 1 (a 14-minute silent stream at 45 min), glm-5.3-flash attempt 1 (four identical 504s on the report turn); claude-sonnet-5 run 2 ran its observation on opus-5 (see the amendments). Each is recorded below with its spend.

Protocol amendments (launch-llms-benchmark)

  • Step 6 - a second run already worse than the first is stopped, not finished: past the first run's whole duration with no report written, or its drive not started by the time the first run had finished, it is killed by its PID, its spend and phase recorded, and the first run is the row. The maintainer's decision after deepseek run 2 passed run 1's 29-minute total at 49 minutes with its report still generating.
  • Step 7 (claude) - the observation must have run on the benchmarked model: modelUsage's keys must be the benchmarked model's (plus the haiku background key) and nothing else carrying spend. A claude-sonnet-5 root dispatched observe-run through the Agent tool with model: 'opus' of its own accord on one run of two; 89 requests, 90 % of the spend and the whole report were opus-5's. Such a run is void for the row.
  • The mission is 669 characters, not 683: feat(harness): one observation depth - drop quick, drop the depth field, older reports' depth ignored #620 dropped the at full depth phrase; the two places that said 683 now say so.

Also on this branch

  • test-plugin-harnessing defaults to --cli copilot --model openai/gpt-5.6-luna (86f18f0), the fastest and cheapest row of this table, on the maintainer's word - measure_phase.py, run_samples.py and the skill say so, analyze_run.py falls back to copilot, and copilot syncs no user scope (its deploy lives in the clone) instead of being refused for a missing --scope; a test covers it. Reviewed by the same reviewer: no finding.

Notes the table has no column for

Source files read before the drive: 0 on every run. Traffic of its own: none beyond the packaged replay (opus-5 run 1: three curl size probes after the drive; gemini-3.8: the scenario replayed twice after the provider crash). Replayable protocol: yes on every report.

Review

A separate reviewer sub-agent (fresh context, Opus 5) reconciled every table figure against the row runs and checked the amendments, secrets and language. Six wrong statements found - five in this body, one README legend word (1ccb2a5) - all fixed; the re-check returned no finding. The ranking is the maintainer's and was not reviewed.

Per-finding rulings

Both runs of every row, the row's run first; a restatement counts once, a bundled row once per defect, a self-declared note never.

deepseek/deepseek-v4.1-flash (opencode) - run 1 - 16 confirmed / 16 reported

Window 20:45:09Z-20:47:11Z (k6's own 121.9 s agrees). 9 anomalies, 8 gaps of which the last restates F8: 16 items.

  • F1 held up: /stats p50 370 / p95 488 / p99 499 ms over 385 calls against /products p50 21.9; api profile query 83.19 % self, stats 77.47 % total; main.py selects every column of every row and json.dumps the catalog for payload_bytes. Perf.
  • F2 held up: mcp profile create_default_context 63.11 % self (18.51 s of 29.33 s), Client.__init__ 52.54 % total; catalog.py opens an httpx.Client inside every GET and POST. Perf.
  • F3 held up: trace 11676738… carries 25 GET /products/{sku} server spans under one tools/call search_products; DETAIL_FANOUT = 25 in server.py; span p95 466 ms; mcp outbound count 5347. Perf.
  • F4 held up: mcp_tool_calls_total{search_products} delta 162 against 391 traces / 392 span-metric calls; 162 search done + 230 search served from cache log lines; the counter sits after the cache early return. Telemetry.
  • F5 held up (mechanism confirmed, effect labelled suspected by the report itself): _SEARCH_CACHE is written at :72, read at :48, cleared nowhere, place_order never touches it; 230 cache hits in the window. Behavior.
  • F6 held up: trace 3088a942… has 3 invoke_agent, 6 chat (3 status=ERROR), 3 identical get_order tool calls; 4 model call failed attempt=n/3 log lines with the finish_reason='error' validation text; assistant.py retries MODEL_ATTEMPTS on any exception. Behavior.
  • F7 held up: catalog_orders_created_total has the single series catalog_category="unknown" (delta 754); the literal is at main.py:181 (the report cites :122 - the line is wrong, the fact is not). Telemetry.
  • F8 held up: the initialize span reads mcp.protocol.version=2025-11-25 while the same k6 session's tools/list reads 2025-06-18 (the version the script proposes and accepts); the agent's sessions read 2026-07-28. Telemetry.
  • F9 held up: 9 POST /ask traces at 15 s intervals from 20:45:09 to 20:47:09, agent_questions_total delta 9; the manifest's expected_rate line says 8 and calls it a ceiling. Behavior.
  • G1 held up: the mcp's metric names are http_client_duration_*, mcp_tool_calls_total, target_info - no server histogram. Telemetry.
  • G2 held up: main.py:211 excludes health from the instrumentation; the route logs and leaves no span. Telemetry.
  • G3 held up: structural query { name = "GET /products/{sku}" } >> { span.db.system.name = "sqlite" } = 0 and POST /orders = 0, against 763 for the sibling GET /orders/{order_ref}; main.py:76 and :135 call db.connect() directly. Telemetry.
  • G4 held up: the agent's 18 metric names carry gen_ai_client_token_usage_* and no operation-duration histogram. Telemetry.
  • G5 held up: 23 chat spans across the 9 traces, 3 of them errored, gen_ai_client_token_usage_count = 20. Telemetry.
  • G6 held up: the 3 errored chat spans and the 3 errored invoke_agent spans carry no error.type. Telemetry.
  • G7 held up: profiles check with service_instance_id = 0, without it 100.02 s. Telemetry.
  • G8 restates F8: not counted.

By kind, confirmed: Telemetry 10 (F4, F7, F8, G1-G7) / Perf 3 (F1, F2, F3) / Behavior 3 (F5, F6, F9).
Figures: launch 20:39:10Z, end 21:08:15Z - preflight 5m59s, drive 2m02s, observation 21m04s, total 29m05s; 73 turns, median 11.9 s; Input 7,981,775 / Output 93,853 / Cache 7,423,488; cost 0.162325 USD (reconciled to the cent at the base tier); signals 4/4; 0 source files read before the drive (first read 20:53:03, after the drive); no traffic of its own (the drive is the packaged detached replay); replayable protocol: yes (section 7).

deepseek/deepseek-v4.1-flash (opencode) - run 2 - killed, no grade

Launched 21:12:17Z; drive 21:27-21:29 (15 min of preflight against run 1's 6); the report skeleton was written at 21:55 and the run was still generating the report at 22:01, 49 min in - 20 min past run 1's whole duration - when it was stopped by its PID on the maintainer's instruction (a second run already worse than the first measures nothing). Spend to that point: 0.1648 USD, 4.14M input / 48k output tokens, 47 turns. No report persisted (the <fill> skeleton was deleted at teardown). The row is run 1.

deepseek/deepseek-v4.1-flash (opencode) - run 3 - killed, no report

Re-run at the maintainer's request to check whether the provider had been slow on the first two runs. Launched 08:31:48Z; preflight 9 min (drive 08:41-08:43), observation from 08:43; the report skeleton was opened at 09:01 and the run was stopped by its PID at 09:01:43Z, past run 1's 29m05s total with no report - step 6's rule. Spend to the kill: 0.0961 USD, 4.73M input / 54k output tokens, 60 turns. Run 1 stays the row.

deepseek/deepseek-v4.1-flash (opencode) - run 4 - stopped, no report

Second re-run at the maintainer's request. Launched 09:02:35Z; preflight 8 min (drive 09:10-09:12), observation by phased helper scripts from 09:13; stopped by its PID at 09:22:50Z on the maintainer's word, 20 min in, no report skeleton yet. Together with run 3 it answers the question the re-runs were asked: under 1.13.0 this model spends six to nine minutes in preflight and twenty minutes or more in observation on every run - the provider was not slow on the first two. Run 1 stays the row. Spend: 0.0745 USD, 2.12M input / 34k output tokens, 33 turns.

z-ai/glm-5.3 (opencode) - run 1 - 17 confirmed / 19 reported

Window 22:33:01Z-22:35:03Z (k6's own 121.9 s agrees). 14 anomalies, 6 gaps of which the last (log bodies carry no trace id while the structured metadata correlates completely, by the report's own words) is a note, not a gap: 19 items.

  • F1 held up: /stats p50 268 / p95 477 / p99 495 ms over 421 calls, sum 89.4 s; api profile query 77.32 % self. Perf.
  • F2 held up: tools/call search_products p95 423 ms; the exemplar carries 26 downstream GETs; DETAIL_FANOUT = 25. Perf.
  • F3 held up: 822 order created lines in the window carry 662 distinct refs; the mcp log shows ORD-000617 returned for two accepted orders of different SKUs (traces 47139af7…, 1629bb55…); main.py:137-138 derives the ref from SELECT COUNT(*) FROM orders + 1 outside any lock. Behavior.
  • F4 held up: mcp profile create_default_context 62.25 % self (17.48 s of 28.08 s); catalog.py per-call client. Perf.
  • F5 held up: mcp_tool_calls_total{search_products} delta 160 against 427 span-metric calls; the other three tools 423/422/422 exact. Telemetry.
  • F6 held up: 267 search served from cache lines; _SEARCH_CACHE is never invalidated by place_order. Behavior.
  • F7 held up: 9 POST /ask traces, the manifest says 8. Behavior.
  • F8 held up: every POST /ask trace carries a server/discover and a tools/list of its own (24 tools/list spans in the window, 15 rooted at the k6 sessions). Perf.
  • F9 did not hold up: the profile reads 80.22 s of CPU over a 123 s window - 65 % of a core, not "~100 %" - and the tail it attributes to queuing (POST /orders 192.9 ms "with no child span, the time is contention") is asserted, not separated from the route's own COUNT(*) + insert + commit; the report itself calls it largely a consequence of F1.
  • F10 held up: the chat spans carry gen_ai.input.messages, gen_ai.output.messages, gen_ai.system_instructions. Telemetry.
  • F11 held up: gen_ai.system=openai beside gen_ai.provider.name=openai. Telemetry.
  • F12 held up: catalog_category="unknown" single series; main.py:181. Telemetry.
  • F13 held up: 21 order rejected … out-of-stock WARN lines answered 200; 411 read-backs against 421 k6 orders. Behavior.
  • F14 held up: agent/app/main.py:68 instruments without excluded_urls, api/app/main.py:211 excludes health. Telemetry.
  • G1 held up: the mcp's metric names carry no http_server_*. Telemetry.
  • G2 did not hold up: 5 initialize and 5 notifications/initialized traces rooted at llmbench-mcp sit in the window; the handshake is traced.
  • G3 held up: the mcp's server spans carry mcp.* and jsonrpc.* attributes only - no http.*, no user agent - so the run's identity cannot select them. Telemetry.
  • G4 held up: no gen_ai.client.operation.duration among the agent's 18 metric names. Telemetry.
  • G5 held up: profiles check with service_instance_id = 0; the SDK pushes process_cpu only. Telemetry.

By kind, confirmed: Telemetry 9 (F5, F10, F11, F12, F14, G1, G3, G4, G5) / Perf 4 (F1, F2, F4, F8) / Behavior 4 (F3, F6, F7, F13).
Figures: launch 22:29:44Z, end 22:49:28Z - preflight 3m17s, drive 2m02s, observation 14m25s, total 19m44s; 44 turns, median 9.4 s; Input 4,419,244 / Output 120,536 / Cache 4,014,656; cost 1.391385 USD (reconciled to the cent, flat rates); signals 4/4 (16 metrics, 9 traces, 13 logs, 3 profiles); 0 source files read before the drive (first read 22:42:45); no traffic of its own; replayable protocol: yes.

z-ai/glm-5.3 (opencode) - run 2 - 8 confirmed / 10 reported

Window 22:55:14Z-22:57:16Z (k6's own 122.0 s agrees). 9 anomalies, 5 gaps of which two restate F6 and F4, one is a mapping note and one says "no gap" by its own words: 10 items.

  • F1 held up: /stats p50 289 / p95 485 / p99 577 ms over 402 calls; api profile query 79.85 % self. Perf.
  • F2 held up: tools/call search_products p50 2.91 / p95 461 ms, 409 calls; GET /products/{sku} 4831 calls. Perf.
  • F3 held up: mcp profile create_default_context 61.70 % self (18.7 s of 30.31 s). Perf.
  • F4 held up: counter 162 against 409 span-metric calls. Telemetry.
  • F5 held up: 9 POST /ask traces against the manifest's 8. Behavior.
  • F6 held up: the mcp's metric names carry no http_server_*. Telemetry.
  • F7 held up: the mcp's outbound client spans are named bare GET / POST, no route or destination in the name. Telemetry.
  • F8 did not hold up: folding ?category and ?category&q into one GET /products span and histogram is what the HTTP conventions prescribe (http.route carries no query string); the two names are the k6 script's, not the service's - the report itself calls it a mapping note.
  • F9 held up: 19 order rejected … out-of-stock WARN lines answered 200, SKU-02987 among them repeatedly. Behavior.
  • G4 did not hold up: span.gen_ai.conversation.id != "" matches all 9 ask traces; the attribute is present.

By kind, confirmed: Telemetry 3 (F4, F6, F7) / Perf 3 (F1, F2, F3) / Behavior 2 (F5, F9).
Figures: launch 22:52:43Z, end 23:00:47Z - preflight 2m31s, drive 2m02s, observation 3m31s, total 8m04s; 59 turns, median 2.5 s; Input 4,712,775 / Output 35,933 / Cache 4,422,720; cost 1.114158 USD (reconciled to the cent); signals 4/4; 0 source files read before the drive (first read 22:59:03); no traffic of its own; replayable protocol: yes.
Run 1 (17/19) is the row on confirmed findings.

openai/gpt-5.6-sol (copilot) - run 1 - 12 confirmed / 13 reported

Window 00:46:25Z-00:48:25Z (k6's own 120.4 s agrees). 7 anomalies of which F2 bundles two defects, 7 gaps of which two restate F6 and F4: 13 items.

  • F1 held up on its defect: /stats p99 492 ms, trace 1fdacecb… 477 of 504 ms in catalog stats_scan, api profile query 72.72 % self of 61.83 s; its "p50 45→275 ms half-to-half" is not what the histogram gives (93→149 ms over the two halves) - the scan is the finding, the ratio is not. Perf.
  • F2a (fan-out) held up: tools/call search_products p95 242 / p99 386 ms, the worst root carries 25 detail GETs. Perf.
  • F2b (a client per call) held up: mcp profile create_default_context 62.75 % self (18.14 s of 28.91 s). Perf.
  • F3 held up (counted apart from F2a, a different fix): gen_ai_client_token_usage{input} p50 768 / p95 36,040 over 18 calls; the worst ask 6.17 s. Perf.
  • F4 held up: counter 162 against 449 tools/call search_products traces / 450 span-metric calls. Telemetry.
  • F5 held up: the chat spans carry gen_ai.input.messages, gen_ai.output.messages, gen_ai.system_instructions, the tool spans gen_ai.tool.call.arguments / result. Telemetry.
  • F6 held up: the UA selector matches 0 traces rooted at the mcp against 2,653 at the api; the mcp's server spans carry no http.* attribute. Telemetry.
  • F7 held up: all 22 api WARN lines are order rejected … out-of-stock, HTTP 200 throughout. Telemetry.
  • G2 held up: no gen_ai_client_operation_duration. Telemetry.
  • G3 held up: profiles carry no service_instance_id. Telemetry.
  • G4 held up: a --trace-id profile query for the agent returns 0 frames against 1.13 CPU-s for the same window - no span-linked profiles. Telemetry.
  • G5 held up: memory:alloc_space is 0 for each of the three services while the store-wide selector holds 239 MB (another service's). Telemetry.
  • G6 did not hold up: k6's own checks_total sits in the store; a service cannot emit a client's assertions.

By kind, confirmed: Telemetry 8 (F4, F5, F6, F7, G2, G3, G4, G5) / Perf 4 (F1, F2a, F2b, F3) / Behavior 0.
Figures: launch 00:45:04Z, end 00:54:36Z - preflight 1m21s, drive 2m00s, observation 6m11s, total 9m32s; 48 turns (root + two observe-run dispatches), median model-call latency 4.2 s, max 63 s; Input 3,747,709 (uncached 212,349, cache read 3,535,360, cache write 0) / Output 27,792 (reasoning 7,747) / Cache 3,535,360; cost at OpenAI list 212,349x2.00/M + 3,535,360x0.20/M + 27,792x10.00/M = 1.409690 USD; 1 premium request, 281.94 AIU; signals 4/4 (15 metrics, 16 traces, 4 logs, 12 profiles); 0 source files read before the drive (first view 00:51:20); no traffic of its own; replayable protocol: yes. Copilot CLI 1.0.86, --effort medium. Scratch under /tmp/oddyssey/.

openai/gpt-5.6-sol (copilot) - run 2 - 9 confirmed / 10 reported

Window 00:58:15Z-01:00:17Z (k6's own 121.9 s agrees). 5 anomalies, 5 gaps: 10 items.

  • F1 held up on its defect: api profile query 79.40 % self of 90.29 s, the /stats scan (the "p50 148→389 ms" drift is a histogram artefact, as in run 1). Perf.
  • F2 held up: mcp profile create_default_context 59.52 % self (16.76 s of 28.16 s). Perf.
  • F3 held up: counter 162 against 403 tools/call search_products traces, 242 cache-hit lines (162 + 242 = 404). Telemetry.
  • F4 did not hold up as an anomaly: a 4.26 s model call inside a 5.70 s ask is the provider's latency, not a defect of the three services - nothing in the stack to fix.
  • F5 held up: 19 order rejected … out-of-stock WARN lines, HTTP 200 throughout. Telemetry.
  • G1 held up: the UA selector matches the api and agent roots and none of the mcp's. Telemetry.
  • G2 held up: memory:alloc_space is 0 per service. Telemetry.
  • G3 held up: profiles carry no service_instance_id. Telemetry.
  • G4 held up: no gen_ai_client_operation_duration. Telemetry.
  • G5 held up: the api's and agent's HTTP series carry http_method, http_target, http_status_code (the old semconv names); a grouping on the stable names collapses. Telemetry.

By kind, confirmed: Telemetry 7 (F3, F5, G1-G5) / Perf 2 (F1, F2) / Behavior 0.
Figures: launch 00:56:55Z, end 01:06:54Z - preflight 1m20s, drive 2m02s, observation 6m37s, total 9m59s; 47 turns, median 4.1 s, max 70 s; Input 3,487,692 (uncached 183,244, cache read 3,304,448, cache write 0) / Output 27,895 (reasoning 8,210) / Cache 3,304,448; cost at list 1.306328 USD; 1 premium request, 261.27 AIU; signals 4/4 (20 metrics, 11 traces, 8 logs, 10 profiles); 0 source files read before the drive; no traffic of its own; replayable protocol: yes.
Run 1 (12/13) is the row on confirmed findings.

z-ai/glm-5.3-flashx (opencode) - run 1 - 11 confirmed / 13 reported

Window 23:06:21Z-23:08:22Z (k6's own 121.5 s agrees). 10 anomalies of which F10 says "explained, not a defect" by its own words, 4 gaps of which G4 restates F10 (counted once, as the gap): 13 items.

  • F1 held up: mcp_tool_calls_total{search_products} delta 160 against 423 tools/call search_products traces; 263 cache-hit log lines. Telemetry.
  • F2 held up: 263 cache hits across the window while 835 orders mutated stock; place_order never touches _SEARCH_CACHE. Behavior.
  • F3 held up: /stats p50 257 / p95 483 / p99 604 ms over 417 calls; api profile query 77.42 % self (56.15 s of 72.53 s). Perf.
  • F4 held up: trace fbe40b4f… lists 718 rows (db.response.returned_rows=718) then fetches 25 GET /products/{sku} one by one. Perf.
  • F5 held up (counted apart from F4, a different fix): gen_ai_client_token_usage{input} p50 779 / p95 12,580 over 19 calls, the exemplar's second chat at 22,663 input tokens. Perf.
  • F6 held up: mcp profile create_default_context 61.89 % self (18.48 s of 29.86 s). Perf.
  • F7 held up: 9 POST /ask traces against the manifest's 8. Behavior.
  • F8 held up: 20 order rejected … out-of-stock lines answered 200, SKU-02987 4 times. Behavior.
  • F9 held up: 10,177 api log lines in the window, 7,931 of them uvicorn access lines beside the app's own line per request. Telemetry.
  • G1 held up: the agent's gen_ai.* metric names are gen_ai_client_token_usage_* only. Telemetry.
  • G2 did not hold up: the chat spans of the cited trace carry gen_ai.response.model=google/gemini-3.5-flash-lite (the report itself marked it suspected, on a truncated rendering).
  • G3 did not hold up: the api's server histogram carries http_target with the route template (/products/{sku}), and the report's own per-route histogram was grouped by it; the label is the old semconv name, not a missing dimension.
  • G4 held up: /health is excluded from the api's instrumentation (main.py:211), 45 pre-run access lines and 0 traces. Telemetry.

By kind, confirmed: Telemetry 4 (F1, F9, G1, G4) / Perf 4 (F3, F4, F5, F6) / Behavior 3 (F2, F7, F8).
Figures: launch 23:03:11Z, end 23:20:20Z - preflight 3m10s, drive 2m01s, observation 11m58s, total 17m09s; 32 turns, median 13.3 s; Input 2,345,736 / Output 68,360 / Cache 2,102,912; cost 0.333013 USD (reconciled to the cent); signals 4/4 (8 metrics, 6 traces, 8 logs, 3 profiles); 0 source files read before the drive (first read 23:13:07); no traffic of its own; replayable protocol: yes.

z-ai/glm-5.3-flashx (opencode) - run 2 - 11 confirmed / 13 reported

Window 23:24:48Z-23:26:49Z (k6's own 121.1 s agrees). 10 anomalies, 5 gap bullets of which one says "client-side limitation, not the stack's" and one "no gap bullet" by their own words: 13 items.

  • F1 held up: /stats p50 273 / p95 490 / p99 510 ms; api profile query 81.13 % self. Perf.
  • F2 held up: mcp profile create_default_context 60.93 % self (14.21 s of 23.32 s). Perf.
  • F3 held up: 784 order created lines carry 616 distinct refs, ORD-000002 created four times; main.py:137-138. Behavior.
  • F4 held up: 160 tools/call search_products over 100 ms, decaying per bin as the cache fills; p95 445 ms against p50 3.19. Perf.
  • F5 held up: trace f8dc6b75… has 3 invoke_agent and 7 chat spans, two of them status=ERROR on an HTTP 200 answer, then the third attempt succeeds; 2 model call failed WARN lines. Behavior.
  • F6 held up: counter 160 for search_products against 406 calls, the other three tools exact. Telemetry.
  • F7 held up (mechanism; impact labelled suspected by the report): _SEARCH_CACHE never invalidated while 784 orders mutated stock. Behavior.
  • F8 held up: 19 order rejected … out-of-stock lines answered 200. Behavior.
  • F9 held up: catalog_category="unknown" single series; main.py:181. Telemetry.
  • F10 held up (the report labels it suspected): the final_result of trace f8dc6b75… reads "215608 euros (2156,08 €)" for total_cents=215608. Behavior.
  • G1 did not hold up: api and mcp "absent" before the first request is no traffic on a just-recreated stack, not an export gap.
  • G2 held up: profiles carry no service_instance_id. Telemetry.
  • G3 did not hold up: http_server_response_size_bytes_* carries http_target with the route template; the sizes split per route under that label.

By kind, confirmed: Telemetry 3 (F6, F9, G2) / Perf 3 (F1, F2, F4) / Behavior 5 (F3, F5, F7, F8, F10).
Figures: launch 23:22:31Z, end 23:36:22Z - preflight 2m17s, drive 2m01s, observation 9m33s, total 13m51s; 44 turns, median 9.2 s; Input 3,669,344 / Output 69,555 / Cache 3,458,304; cost 0.424401 USD (reconciled to the cent); signals 4/4 (through its batch*.sh helpers - the log alone counts 1/1/1/0); 0 source files read before the drive (first read 23:32:50); no traffic of its own; replayable protocol: yes.
Tie with run 1 on confirmed findings (11 each); run 1 is cheaper (0.33 against 0.42 USD) and is the row.

anthropic/claude-opus-5 (claude) - run 2 - 17 confirmed / 17 reported - THE ROW

Window 02:30:09Z-02:32:11Z (k6's own 121.0 s agrees). 8 anomalies of which F3 and F7 each bundle two defects, 8 gaps of which G3 restates F4: 17 items.

  • F1 held up: 811 order created lines carry 638 distinct refs, 130 refs shared by 303 orders; 193 catalog get_order spans return 2 rows; main.py:137-138. Behavior.
  • F2 held up: /stats p50 283 / p99 507 ms; api profile query 85.82 % self of 88.6 s. Perf.
  • F3a (fan-out) held up: tools/call search_products p50 1.92 / p95 240 ms over 420 calls, 26 sequential GETs per miss. Perf.
  • F3b (a client per call) held up: mcp profile create_default_context 58.60 % self (12.78 s of 21.81 s). Perf.
  • F4 held up: counter 160 against 420 span-metric calls. Telemetry.
  • F5 held up (labelled suspected by the report): the k6-rooted POST /orders spans read p50 27 ms (n=415) against 10 ms for the mcp-routed ones (n=416); the writer waiting on /stats readers is the mechanism this campaign verified by overlap on another run. Perf.
  • F6 held up: 20 order rejected … out-of-stock WARN lines answered 200. Behavior.
  • F7a (prompt inflation) held up: the ask trace adbb858e… carries a 22,930-input-token chat on a 718-row search. Perf.
  • F7b (a wrong answer) held up: that trace's final_result reads "priced at 208100 euros" for a price_cents value and names SKU-0733. Behavior.
  • F8 held up: 833 api spans return 400 rows or more (415 category listings, 415 stats scans, 3 agent searches). Perf.
  • G1 held up: no db span under POST /orders and GET /products/{sku}. Telemetry.
  • G2 held up: no HTTP server series or attributes on the mcp's transport. Telemetry.
  • G4 held up: no gen_ai_client_operation_duration. Telemetry.
  • G5 held up: profiles carry no service.instance.id. Telemetry.
  • G6 held up: the three services emit the pre-1.0 HTTP semconv names. Telemetry.
  • G7 held up: gen_ai.system beside gen_ai.provider.name. Telemetry.
  • G8 held up: 24 GET /health access lines without a trace id, the route excluded from traces and metrics. Telemetry.

By kind, confirmed: Telemetry 8 (F4, G1, G2, G4, G5, G6, G7, G8) / Perf 6 (F2, F3a, F3b, F5, F7a, F8) / Behavior 3 (F1, F6, F7b).
Figures: launch 02:27:22Z, end 02:46:53Z - preflight 2m47s, drive 2m02s, observation 14m42s, total 19m31s; 53 requests (root + one observe-run subagent, modelUsage the single key claude-opus-5), median 6.6 s, max 60 s; Input 5,768,520 (uncached 106, cache read 5,504,746, cache creation 263,668 of which 210,576 at 5m and 53,092 at 1h) / Output 61,893 (thinking 22,212) / Cache 5,768,414; total_cost_usd 6.147248 (list), reconstructed to the cent; signals 4/4 (4 metrics, 8 traces, 5 logs, 1 profiles in the log, more inside its helpers); 0 source files read before the drive (first read 02:39:09); no traffic of its own; replayable protocol: yes. Claude Code 2.1.278, --effort medium. Scratch under /private/tmp/oddyssey-scratch/.
Tie with run 1 (17/18) on confirmed findings; run 2 is cheaper (6.15 against 7.51 USD) and is the row.

anthropic/claude-opus-5 (claude) - run 1 - 17 confirmed / 18 reported

Window 02:05:59Z-02:08:00Z (k6's own 120.3 s agrees). 10 anomalies of which F1 bundles two defects and F10 (throughput drifting as the cache warms) is F3's consequence, 10 gap bullets of which G3 restates F4 and the last says "not a gap": 18 items.

  • F1a (fan-out) held up: trace 6b189e6c… 106 spans, tools/call search_products p95 473 / p99 504 ms over 417 calls. Perf.
  • F1b (a client per call) held up: mcp profile create_default_context 60.64 % self (18.18 s of 29.98 s); catalog.py:25. Perf.
  • F2 held up: api profile query 73.89 % self of 72.26 s; payload_bytes=1838708 in the /stats log lines; main.py:92-131. Perf.
  • F3 held up (staleness effect labelled suspected by the report): 257 of 417 searches served from the cache, no invalidation on place_order. Behavior.
  • F4 held up: counter 160 against 417 span-metric calls. Telemetry.
  • F5 held up: 20 order rejected … stock=0 WARN lines answered 200; the mcp logs 11 of them INFO order placed … 'order_ref': None. Behavior.
  • F6 held up: trace ffa12470… carries 2 invoke_agent, 5 chat spans, 2 errored, the retry re-running the whole agent (new discover + tools/list + tool call). Behavior.
  • F7 held up: GET /products?category=cameras answers 43,990 bytes, matches=417 on 278 log lines; /products p50 9.1 / p95 60 ms against 2.6 for a key read. Perf.
  • F8 held up: gen_ai_client_token_usage{input} p50 736 / p95 10,240 / p99 15,160; the search questions at 12.7k tokens. Perf.
  • F9 held up: a server/discover and a tools/list inside every POST /ask trace. Perf.
  • G1 held up: no sqlite child under POST /orders and GET /products/{sku} (structural queries 0). Telemetry.
  • G2 held up: the mcp emits no HTTP server span or http_server_* series; its tool spans are roots without an HTTP status. Telemetry.
  • G4 held up: http_server_response_size_bytes on /products reads p50 9,657 / p95 10,000 / p99 10,000 for a mean of 22,240 bytes - the histogram carries the duration default buckets (le=0, 5, 10 … 10,000), every listing lands in +Inf. Telemetry.
  • G5 held up: the api's HTTP histograms carry http_target (the route template) and no http_route. Telemetry.
  • G6 did not hold up: the span's http.target dropping the query string and folding ?category and ?q under one route is what the HTTP conventions prescribe; the k6 names are the client's.
  • G7 held up: the mcp's client spans are named bare GET / POST with http.url as the only key. Telemetry.
  • G8 held up: 18 usage samples against 19 chat spans, the errored one carries none. Telemetry.
  • G9 held up: profiles carry no service.instance.id. Telemetry.

By kind, confirmed: Telemetry 8 (F4, G1, G2, G4, G5, G7, G8, G9) / Perf 6 (F1a, F1b, F2, F7, F8, F9) / Behavior 3 (F3, F5, F6).
Figures: launch 02:02:55Z, end 02:25:02Z - preflight 3m04s, drive 2m01s, observation 17m02s, total 22m07s; 60 requests (root + one observe-run subagent, both on opus-5 - modelUsage carries the single key claude-opus-5), median 5.9 s, max 71 s; Input 6,670,041 (uncached 120, cache read 6,244,216, cache creation 425,705 of which 369,645 at 5m and 56,060 at 1h) / Output 60,840 (thinking 18,659) / Cache 6,669,921; total_cost_usd 7.514589 (list), reconstructed to the cent at 5.00 / 25.00 / 0.50 / 6.25 / 10.00 USD per million; signals 4/4 (its q1-q5.sh helpers carry 14 metrics, 13 traces, 8 logs, 3 profiles invocations); 0 source files read before the drive (first read 02:15:27, after the queries); traffic of its own: three curl size probes of GET /products at 02:15, after the drive, outside the window - measurement, not a scenario; replayable protocol: yes. Claude Code 2.1.278, --effort medium.

qwen/qwen3.8-max-0902 (opencode) - run 1 - 15 confirmed / 16 reported

Window 04:19:07Z-04:21:08Z (k6's own 121.7 s agrees). 9 anomalies of which F3 bundles two defects, 6 gaps: 16 items.

  • F1 held up: counter 160 against 444 span-metric calls, 284 cache-hit lines (160 + 284 = 444). Telemetry.
  • F2 held up (the stale body labelled unobserved by the report itself): 284 cache hits, no invalidation on place_order. Behavior.
  • F3a (fan-out) held up: tools/call search_products p50 1.75 / p95 235 ms, 106-span exemplar. Perf.
  • F3b (25 full records shipped into the prompt) held up: gen_ai_client_token_usage{input} p50 784 / p95 13,000 over 19 calls, the heaviest at 22,930. Perf.
  • F4 held up: mcp profile create_default_context 58.04 % self (11.23 s of 19.35 s). Perf.
  • F5 held up: /stats p50 326 / p95 483 / p99 497 ms over 438 calls, sum 91.3 s; api profile query 83.25 % self. Perf.
  • F6 held up: 9 POST /ask traces, agent_questions_total 9, the manifest says 8. Behavior.
  • F7 held up: gen_ai.system=openai beside gen_ai.provider.name=openai on the chat spans and as a metric label. Telemetry.
  • F8 held up: trace a1d608e0… carries final_result on invoke_agent and gen_ai.tool.call.arguments on the tool spans. Telemetry.
  • F9 held up: 22 order rejected … out-of-stock lines answered 200. Behavior.
  • G1 held up: no gen_ai_client_operation_duration. Telemetry.
  • G2 held up: no sqlite child under GET /products/{sku} (structural query 0); main.py:76-84. Telemetry.
  • G3 held up: 5 mcp Created new transport with session ID lines without a trace id, 883 of 888 with one. Telemetry.
  • G4 held up: profiles carry no instance id. Telemetry.
  • G5 held up: the mcp's tools/call root spans carry no HTTP status attribute. Telemetry.
  • G6 did not hold up as a gap of the services: a dropped_iterations series absent when its value is zero is k6's own OTel export at work, not the stack's telemetry, and the report says it ruled the schedule from the services' rows anyway.

By kind, confirmed: Telemetry 8 (F1, F7, F8, G1-G5) / Perf 4 (F3a, F3b, F4, F5) / Behavior 3 (F2, F6, F9).
Figures: launch 04:15:03Z, end 04:45:28Z - preflight 4m04s, drive 2m01s, observation 24m20s, total 30m25s; 34 turns, median 22.5 s, max 1752 s (the report-writing generation); Input 2,766,531 / Output 67,087 / Cache 2,534,656; cost 1.499936 USD (reconciled to the cent, flat rates 2.00 / 6.00 / 0.25 USD per million); signals 4/4 (9 metrics, 10 traces, 4 logs, 1 profiles); 0 source files read before the drive (first read 04:33:47); no traffic of its own; replayable protocol: yes.

qwen/qwen3.8-max-0902 (opencode) - run 2 - 14 confirmed / 14 reported

Window 04:52:18Z-04:54:23Z (k6's own 124.4 s agrees). 7 anomalies of which F2 and F6 each bundle two defects, 6 gaps of which G5 restates F7: 14 items.

  • F1 held up: /stats p50 261 / p95 476 / p99 495 ms over 433 calls; api profile query 83.17 % self. Perf.
  • F2a (fan-out) held up: 106-span exemplar, tools/call search_products p95 237 ms. Perf.
  • F2b (a client per call) held up: mcp profile create_default_context 57.14 % self (11.52 s of 20.16 s). Perf.
  • F3 held up: counter 160 against 439 span-metric calls (433 rooted). Telemetry.
  • F4 held up (stale value labelled suspected by the report): 273 cache-hit traces, no invalidation. Behavior.
  • F5 held up: gen_ai_client_token_usage{input} p50 784 / p95 13,000, one question at 22,930. Perf.
  • F6a held up: script.js:493 const lookupRef = mcpOrderRef || apiOrderRef - the tool-side get_order ran 434 times while 11 place_order calls were rejected, reading back the HTTP-side order the manifest's "every iteration in which an order reference came back" does not mean. Behavior (benchmark fidelity).
  • F6b held up: 9 POST /ask traces against the manifest's 8. Behavior.
  • F7 held up: 22 order rejected … out-of-stock lines answered 200. Behavior.
  • G1 held up: no gen_ai_client_operation_duration. Telemetry.
  • G2 held up: no http_server_* series on the mcp. Telemetry.
  • G3 held up: the api and mcp emit the old HTTP semconv names beside new-style db attributes. Telemetry.
  • G4 held up: 5 mcp session-creation lines without a trace id, 873 of 878 with one. Telemetry.
  • G6 held up: gen_ai.system beside gen_ai.provider.name. Telemetry.

By kind, confirmed: Telemetry 6 (F3, G1, G2, G3, G4, G6) / Perf 4 (F1, F2a, F2b, F5) / Behavior 4 (F4, F6a, F6b, F7).
Figures: launch 04:47:15Z, end 05:17:30Z - preflight 5m03s, drive 2m05s, observation 23m07s, total 30m15s; 46 turns, median 17.1 s, max 1707 s; Input 3,429,113 / Output 62,152 / Cache 3,246,208; cost 1.550274 USD (reconciled to the cent); signals 4/4 (through its batch*.sh and discover.sh helpers - the log alone counts 1/0/0/0); 0 source files read before the drive; no traffic of its own; replayable protocol: yes.
Run 1 (15/16) is the row on confirmed findings.

google/gemini-3.8-flash (opencode) - run 1 - 12 confirmed / 12 reported

The observation subagent's stream died at 23:45:23 and 23:46:32 on a 400 Corrupted thought signature from Google AI Studio through OpenRouter, after its first drive (23:42:10-23:44:11); the root dispatched the observation again, which drove the scenario a second time (window 23:48:54Z-23:50:57Z, k6's own 121.8 s agrees) against the same containers, whose store already carried the first drive. The report's window is the second drive; the run's own earlier traffic sits behind it (the search cache was warm: every one of the 409 searches was a cache hit). 10 anomalies, 6 gaps of which four restate F6 (twice), F3 and F7: 12 items.

  • F1 held up: /stats p50 375 / p95 488 ms; api profile query 88.10 % self of 109.19 s. Perf.
  • F2 held up: 42 order rejected … out-of-stock WARN lines answered 200 (the catalog was already depleted by the first drive); main.py:157. Behavior.
  • F3 held up: mcp_tool_calls_total{search_products} delta 0 against 409 tools/call search_products traces and 409 cache-hit log lines. Telemetry.
  • F4 held up: mcp profile create_default_context 42.80 % self (4.22 s of 9.86 s - a smaller share because no search missed the cache). Perf.
  • F5 held up: gen_ai.input.messages, gen_ai.output.messages, gen_ai.system_instructions on the chat spans; assistant.py:56. Telemetry.
  • F6 held up: structural queries for a sqlite child under POST /orders and GET /products/{sku} both 0; main.py:76, :135. Telemetry.
  • F7 held up: catalog_category="unknown" single series, delta 765; main.py:181. Telemetry.
  • F8 held up: trace c0618795… chat spans at 267 and 22,663 input tokens (22,930 together). Perf.
  • F9 held up: all 9 POST /ask traces carry server/discover and tools/list. Perf.
  • F10 held up: 765 order created lines in the window carry 563 distinct refs (148 duplicated); main.py:137-138. Behavior.
  • G1 held up: the mcp's metric names carry no http_server_*. Telemetry.
  • G6 held up: no gen_ai_client_operation_duration among the agent's names. Telemetry.

By kind, confirmed: Telemetry 6 (F3, F5, F6, F7, G1, G6) / Perf 4 (F1, F4, F8, F9) / Behavior 2 (F2, F10).
Figures: launch 23:38:56Z, end 23:59:35Z - preflight 9m58s (first drive and crash included), drive 2m03s, observation 8m38s, total 20m39s; 148 turns over 3 sessions, median 4.4 s; Input 12,787,753 / Output 59,361 / Cache 11,138,874; cost 2.294679 USD (reconciled to the cent, flat rates); signals 4/4 (12 metrics, 7 traces, 5 logs, 2 profiles); 0 source files read before the drive (first read 23:55:06); no traffic of its own beyond the two replays of the stored scenario; replayable protocol: yes.

google/gemini-3.8-flash (opencode) - run 2 - 9 confirmed / 9 reported

The same 400 Corrupted thought signature killed two observe-run sessions (00:10:26 and 00:11:21) after the first drive (00:07:43-00:09:43); the root's third dispatch drove again and reported on that window (00:13:15Z-00:15:16Z, k6's own 120.8 s agrees), the first drive's traffic sitting behind it in the store. 7 anomalies, 4 gaps of which two restate F7 and F5: 9 items.

  • F1 held up: catalog stats_scan 473 ms of a 492 ms /stats trace; api profile query 89.07 % self of 112.65 s. Perf.
  • F2 held up: ORD-000770 on two order created lines, SKU-00169 and SKU-00012, traces 6b61e121… and 7b8d8ed1…. Behavior.
  • F3 held up: /orders p50 9.3 / p95 77 / p99 176 ms; every one of the five slowest POST /orders traces (120-210 ms) overlaps one to four /stats scans of 130-410 ms, the reader the writer waits on. Perf.
  • F4 held up: trace e9b91cae… carries 2 invoke_agent and 5 chat spans with one status=ERROR, and the finish_reason … input_value='error' WARN line. Behavior.
  • F5 held up: counter delta 0 against 401 tools/call search_products traces (the cache warmed by the first drive). Telemetry.
  • F6 held up: mcp profile create_default_context 43.64 % self (4.36 s of 9.99 s). Perf.
  • F7 held up: catalog_category="unknown" on every order; main.py:181. Telemetry.
  • G1 held up: no sqlite child under POST /orders (structural query 0). Telemetry.
  • G4 held up: search_products spans carry no cache-hit attribute; nothing in server.py sets one. Telemetry.

By kind, confirmed: Telemetry 4 (F5, F7, G1, G4) / Perf 3 (F1, F3, F6) / Behavior 2 (F2, F4).
Figures: launch 00:01:51Z, end 00:22:58Z - preflight 11m24s (first drive and two crashes included), drive 2m01s, observation 7m42s, total 21m07s; 167 turns over 5 sessions, median 3.7 s; Input 12,496,566 / Output 68,181 / Cache 10,830,668; cost 2.317402 USD (reconciled to the cent); signals 4/4 (6 metrics, 13 traces, 11 logs, 2 profiles); 0 source files read before the drive (first read 00:17:29); no traffic of its own beyond the two replays; replayable protocol: yes.
Run 1 (12/12) is the row on confirmed findings.

openai/gpt-5.6-luna (copilot) - run 1 - 7 confirmed / 8 reported

Window 00:28:19Z-00:30:20Z (k6's own 121.2 s agrees). 5 anomalies, 4 gaps of which the first restates F5: 8 items.

  • F1 held up: api profile query 71.13 % self of 65.18 s, /stats the only route in the hundreds of ms (the trace id it cites, 79f30bdf…, is F2's retried ask, not a /stats trace - the profile carries the finding). Perf.
  • F2 held up: trace 79f30bdf… has 2 errored spans and 4 chat calls, one model call failed … finish_reason='error' WARN line, HTTP 200 kept. Behavior.
  • F3 held up: mcp profile create_default_context 61.47 % self (17.93 s of 29.17 s). Perf.
  • F4 held up: 22 order rejected … out-of-stock WARN lines answered 200; main.py:157. Behavior.
  • F5 held up: profiles labels --label service_instance_id empty, check restores data only without the label. Telemetry.
  • G2 did not hold up: the api's server histogram carries http_target with the route template, and per-route rows come out of it; the report grouped by http_route, a label the old semconv the app uses never emitted.
  • G3 held up: 5 mcp Created new transport with session ID lines carry no trace id, 895/900 lines do. Telemetry.
  • G4 held up: the agent's 18 metric names carry no retry counter; the retry is visible only as a WARN line and an errored span. Telemetry.

By kind, confirmed: Telemetry 3 (F5, G3, G4) / Perf 2 (F1, F3) / Behavior 2 (F2, F4).
Figures: launch 00:27:17Z, end 00:33:46Z - preflight 1m02s, drive 2m01s, observation 3m26s, total 6m29s; 38 turns (root + one observe-run subagent), median model-call latency 2.8 s, max 28 s; Input 2,932,752 (uncached 114, cache read 2,739,474, cache write 193,164) / Output 17,189 (reasoning 3,940) / Cache 2,932,638; cost at OpenAI list (114+193,164)x0.20/M + 2,739,474x0.02/M + 17,189x1.20/M = 0.114072 USD; 1 premium request, 12.37 AIU; signals 4/4 (11 metrics, 5 traces, 8 logs, 7 profiles); 0 source files read before the drive (four viewed at 00:32:37); no traffic of its own; replayable protocol: yes. Copilot CLI 1.0.86, --effort medium. Scratch under .odd/scratch/ inside the repository (untracked, cleared at teardown).

openai/gpt-5.6-luna (copilot) - run 2 - 5 confirmed / 6 reported

Window 00:37:05Z-00:39:05Z (k6's own 123.3 s agrees). 5 anomalies of which F2 bundles two defects and F3 restates F2's consequence ("model-bound" is no defect); 6 gap bullets of which two restate F4 and three call themselves "by design", "filled/expected" and "not probed": 6 items.

  • F1 held up: /stats trace p95 464 ms; api profile query 75.01 % self of 78.08 s. Perf.
  • F2a (fan-out) held up: trace b774d0b2… carries 25 GET /products/{sku} under one search; DETAIL_FANOUT=25. Perf.
  • F2b (a client per call) held up: mcp profile create_default_context 60.63 % self (16.97 s of 27.99 s). Perf.
  • F4 held up as the instance-identity gap: profiles labels --label service_instance_id returns null (its other half - metrics carry instance UUIDs rather than the slug - is how the identity is meant to travel, and the report's own frontmatter maps them). Telemetry.
  • F5 held up: 21 order rejected … out-of-stock WARN lines (the report counts 22 WARN), all answered 200. Behavior.
  • G3 did not hold up: checks_total{condition="nonzero"} = 9847 is in the store, exported by k6 - the count is not "only in the transient summary", and a service cannot emit a client's assertions.

By kind, confirmed: Telemetry 1 (F4) / Perf 3 (F1, F2a, F2b) / Behavior 1 (F5).
Figures: launch 00:35:54Z, end 00:42:47Z - preflight 1m11s, drive 2m00s, observation 3m42s, total 6m53s; 47 turns, median 2.4 s; Input 3,508,557 (uncached 141, cache read 3,363,979, cache write 144,437) / Output 19,811 (reasoning 5,130) / Cache 3,508,416; cost at list 0.119968 USD; 1 premium request, 12.72 AIU; signals 4/4; 0 source files read before the drive; no traffic of its own; replayable protocol: yes.
Run 1 (7/8) is the row on confirmed findings.

openai/gpt-5.6-terra (copilot) - run 2 - 7 confirmed / 8 reported - THE ROW

The root tried observe-run through Copilot's skill tool ("Skill not found"), never dispatched the subagent, and did the whole observation itself. Window 01:18:17Z-01:20:18Z (k6's own 121.1 s agrees). 5 anomalies of which F3 bundles two defects, 3 gaps of which G1 restates F4: 8 items.

  • F1 held up: 1 of 8 POST /ask traces answers 500; 4 model call failed WARN lines and 1 ask failed; the validation text names finish_reason='error'; assistant.py retries three times. Behavior.
  • F2 held up: /stats trace p95 493 ms, catalog stats_scan p95 422 ms; api profile query 81.42 % self of 91.88 s. Perf.
  • F3a (a client per call) held up: mcp profile create_default_context 61.70 % self (19.14 s of 31.02 s). Perf.
  • F3b (fan-out) held up: tools/call search_products p95 474 ms, 25 detail fetches per miss. Perf.
  • F4 did not hold up: the per-route and per-method dimensions exist - http_method and http_target split the api's histogram into its five routes; the report queried only the newer names http_request_method and http_route and called the dimensions absent.
  • F5 held up: 18 order rejected … out-of-stock WARN lines (the report counts 21), 761 accepted against 779 order spans. Behavior.
  • G2 held up: profiles carry no per-instance identity. Telemetry.
  • G3 held up: no failure counter or provider-error category among the agent's metric names. Telemetry.

By kind, confirmed: Telemetry 2 (G2, G3) / Perf 3 (F2, F3a, F3b) / Behavior 2 (F1, F5).
Figures: launch 01:17:41Z, end 01:23:39Z - preflight 0m36s, drive 2m01s, observation 3m21s, total 5m58s; 26 turns (root only, no subagent), median 3.2 s, max 27 s; Input 2,438,676 (uncached 78, cache read 2,302,150, cache write 136,448) / Output 12,531 (reasoning 3,393) / Cache 2,438,598; cost at OpenAI list (78+136,448)x2.00/M + 2,302,150x0.20/M + 12,531x12.00/M = 0.883854 USD; 1 premium request, 95.21 AIU; signals 4/4 (3 metrics, 2 traces, 2 logs, 1 profiles); 0 source files read before the drive (two viewed at 01:22:42); no traffic of its own; replayable protocol: yes.
Chosen over run 1 (6/7, 7m44s, 1.25 USD) on confirmed findings.

openai/gpt-5.6-terra (copilot) - run 1 - 6 confirmed / 7 reported

Window 01:09:55Z-01:11:58Z (k6's own 121.9 s agrees). 5 anomalies, 3 gaps of which G3 restates F1: 7 items.

  • F1 held up: 20 order rejected … out-of-stock WARN lines; the api roots answer 200 with status UNSET; k6's failure rate 0. Behavior.
  • F2 held up: /stats p50 257 / p95 476 / p99 495 ms over 412 calls; api profile query 79.80 % self of 85.43 s. Perf.
  • F3 held up: mcp profile create_default_context 61.64 % self (18.93 s of 30.71 s). Perf.
  • F4 held up: the chat spans carry gen_ai.input.messages (3 of 3 sampled), output messages, tool arguments and results, final_result. Telemetry.
  • F5 held up: 9 POST /ask traces, agent_questions_total 9, the manifest says 8. Behavior.
  • G1 did not hold up: k6's own checks_total sits in the store; the threshold is rulable from it, and a service cannot emit a client's assertions.
  • G2 held up: profiles carry no service_instance_id. Telemetry.

By kind, confirmed: Telemetry 2 (F4, G2) / Perf 2 (F2, F3) / Behavior 2 (F1, F5).
Figures: launch 01:08:39Z, end 01:16:23Z - preflight 1m16s, drive 2m03s, observation 4m25s, total 7m44s; 45 turns (root + one observe-run subagent), median 3.4 s, max 38 s; Input 3,285,222 (uncached 135, cache read 3,077,833, cache write 207,254) / Output 18,269 (reasoning 5,690) / Cache 3,285,087; cost at OpenAI list (135+207,254)x2.00/M + 3,077,833x0.20/M + 18,269x12.00/M = 1.249573 USD; 1 premium request, 135.32 AIU; signals 4/4 (8 metrics, 8 traces, 3 logs, 3 profiles); 0 source files read before the drive (three grepped at 01:14:53); no traffic of its own; replayable protocol: yes. Copilot CLI 1.0.86, --effort medium. Scratch under /tmp/oddyssey-observe/.

google/gemini-3.7-flash (opencode) - run 2 - 8 confirmed / 9 reported - THE ROW

Window 22:19:24Z-22:21:26Z (k6's own 121.8 s agrees). 6 anomalies of which F1 and F5 each bundle two defects with two fixes, 4 gaps of which three restate F6, F5b and F3: 9 items.

  • F1a (search fan-out) held up: trace beab7a03… carries 25 GET /products/{sku} server spans under one tools/call search_products; span p95 469 ms against p50 2.49; DETAIL_FANOUT = 25. Perf.
  • F1b (a client per call) held up: mcp profile create_default_context 63.14 % self (18.86 s of 29.87 s); catalog.py opens an httpx.Client per call. Perf.
  • F2 held up: /stats p50 258 / p95 476 / p99 495 ms over 401 calls; api profile query 79.36 % self; the full scan and json.dumps in main.py. Perf.
  • F3 held up: the chat spans of trace d0330fad… carry gen_ai.input.messages, gen_ai.output.messages and gen_ai.system_instructions, the invoke_agent span pydantic_ai.all_messages and final_result; assistant.py:56 builds InstrumentationSettings without turning content capture off. Telemetry.
  • F4 did not hold up: the blocking time.sleep in throttle.py is real code but the finding cites no telemetry showing a stall - PAUSE_S 0.25 s never fires against 15 s arrivals (3 ms between agent.ask and invoke_agent on every trace). Code review, not an observation.
  • F5a (out-of-stock answered 200) held up: 19 order rejected … reason=out-of-stock WARN lines; main.py:157 returns the error body with the default 200. Behavior.
  • F5b (catalog_category="unknown") held up: single series, delta 784; main.py:181. Telemetry.
  • F6 held up: { name = "GET /products/{sku}" } >> { span.db.system.name = "sqlite" } = 0 against 963 for GET /products; main.py:76 calls db.connect() directly. Telemetry.
  • G4 held up: the chat spans carry gen_ai.system=openai beside gen_ai.provider.name=openai (retired name still emitted). Telemetry.

By kind, confirmed: Telemetry 4 (F3, F5b, F6, G4) / Perf 3 (F1a, F1b, F2) / Behavior 1 (F5a).
Figures: launch 22:17:00Z, end 22:27:17Z - preflight 2m24s, drive 2m02s, observation 5m51s, total 10m17s; 90 turns, median 4.3 s; Input 6,542,801 / Output 32,417 / Cache 5,846,627; cost 1.082191 USD (reconciled to the cent, flat rates); signals 4/4 (4 metrics, 8 traces, 4 logs, 2 profiles); 0 source files read before the drive (first read 22:24:40); no traffic of its own; replayable protocol: yes.
Chosen over run 1 (7/8, 9m28s, 1.00 USD) on confirmed findings, step 6's first criterion.

google/gemini-3.7-flash (opencode) - run 1 - 7 confirmed / 8 reported

Window 22:07:12Z-22:09:13Z (k6's own 121.1 s agrees). 4 anomalies of which the first bundles two defects with two fixes, 3 gaps: 8 items.

  • F1a (search fan-out) held up: trace e57b487f… carries 25 GET /products/{sku} server spans under one tools/call search_products; span p95 455 ms against p50 1.86; DETAIL_FANOUT = 25 in server.py. Perf.
  • F1b (a client per call) held up: mcp profile create_default_context 62.03 % self (18.69 s of 30.13 s); catalog.py opens an httpx.Client per GET/POST. Perf.
  • F2 held up: /stats p50 232 / p95 477 / p99 498 ms over 410 calls; trace cd2c433f… is 527 ms of which catalog stats_scan 514.6; api profile query 78.78 % self; main.py scans every column and json.dumps the catalog. Perf.
  • F3 did not hold up: throttle.py does call time.sleep from the coroutine, but PAUSE_S is 0.25 s against arrivals 15 s apart, so it never fires - the cited trace 25361b8f… has 3 ms between agent.ask and invoke_agent and its 3.6 s is two chat calls of 1.2 and 1.1 s; the latency the finding attributes to the sleep is the provider's. Code review dressed in telemetry.
  • F4 held up: 19 order rejected … reason=out-of-stock WARN lines; trace fa9cf062… (tools/call place_order) answers http.status_code=200, status UNSET, errors=0; main.py:157 returns {"error": "out of stock", "order_ref": None} with the default 200. Behavior.
  • G1 held up: catalog_orders_created_total has the single series catalog_category="unknown" (delta 802). Telemetry.
  • G2 held up: { name = "POST /orders" } >> { span.db.system.name = "sqlite" } = 0 against 812 for GET /orders/{order_ref}. Telemetry.
  • G3 held up: the 8 chat spans carry gen_ai.provider.name=openai beside server.address=openrouter.ai. Telemetry.

By kind, confirmed: Telemetry 3 (G1-G3) / Perf 3 (F1a, F1b, F2) / Behavior 1 (F4).
Figures: launch 22:04:50Z, end 22:14:18Z - preflight 2m22s, drive 2m01s, observation 5m05s, total 9m28s; 87 turns, median 3.3 s; Input 6,465,034 / Output 30,705 / Cache 5,878,796; cost 0.995732 USD (reconciled to the cent, flat rates); signals 4/4 (12 metrics, 5 traces, 2 logs, 1 profiles invocations); 0 source files read before the drive (first read 22:12:02); no traffic of its own; replayable protocol: yes.

qwen/qwen3.8-27b (opencode) - attempt 2 - 11 confirmed / 11 reported - THE ROW

Window 05:57:32Z-05:59:32Z (k6's own 120.7 s agrees). One silent stream of ten minutes at 06:11-06:22 (no log line, no part), which completed on its own this time. 8 anomalies of which F2 bundles two defects, 3 gaps of which the last calls itself "filled (by design, documented)": 11 items.

  • F1 held up: 811 order created lines carry 655 distinct refs, 113 shared. Behavior.
  • F2a (fan-out) held up: tools/call search_products p50 1.74 / p95 239 ms over 420 calls. Perf.
  • F2b (a client per call) held up: mcp profile create_default_context 58.86 % self (12.12 s of 20.59 s). Perf.
  • F3 held up: counter 160 against 420 span-metric calls. Telemetry.
  • F4 held up: api profile query 85.37 % self of 91.3 s; the scan and json.dumps. Perf.
  • F5 held up: /stats answers 1,129 bytes (histogram mean 1,129, a live probe 1,129) while its log line says payload_bytes=1838708 - main.py:121-127 logs the size of the whole catalog it rendered, not of the response. Telemetry.
  • F6 held up: 20 order rejected … out-of-stock lines answered 200. Behavior.
  • F7 held up (staleness single-signal, as the report says): 260 cache hits, no invalidation. Behavior.
  • F8 held up: 3,874 api traces carry http send spans - two http send and one http receive sub-millisecond INTERNAL spans per request. Telemetry.
  • G1 held up: server.py logs only in search_products and place_order; get_product and get_order emit no line, against the manifest's claim. Telemetry.
  • G2 held up: no gen_ai_client_operation_duration. Telemetry.

By kind, confirmed: Telemetry 5 (F3, F5, F8, G1, G2) / Perf 3 (F2a, F2b, F4) / Behavior 3 (F1, F6, F7).
Figures: launch 05:53:43Z, end 06:34:08Z - preflight 3m49s, drive 2m00s, observation 34m36s, total 40m25s; 48 turns, median 16.6 s, max 2352 s; Input 5,878,526 / Output 133,232 / Cache 3,734,464; cost 1.617631 USD (reconciled to the cent at 0.42 / 3.00 / 0.085 USD per million); signals 4/4 (through its helper scripts - the log alone counts 1/0/0/0); 0 source files read before the drive (first read 06:07:05); no traffic of its own; replayable protocol: yes.

qwen/qwen3.8-27b (opencode) - attempt 1 - killed, no report

Launched 03:28:47Z; drove at 03:33-03:35; observed for 25 min; at 03:59:50 a stream opened with one socket to the provider and produced nothing for 14 minutes - no new log line, no new part row (313 rows from 03:59:56 on) - past the ten-minute bounded wait. Killed by its PID at 04:14:00Z, 45 min in, no report skeleton opened; the store shows the silent turn completing at 04:13:55, seconds before the kill (316 rows), with the report phase still ahead. Spend to the kill: 1.473 USD, 6.71M input / 102k output tokens, 59 turns. Per the queue rule the model moved to the end of the queue for its second attempt.

z-ai/glm-5.3-flash (opencode) - attempt 2 - 7 confirmed / 8 reported - THE ROW

Window 05:26:05Z-05:28:08Z (k6's own 122.5 s agrees). One 504 Upstream idle timeout exceeded on the report-writing turn at 05:49:16 - the retry passed (the first attempt's four did not). 7 anomalies of which F2 bundles the fan-out with a claim about detached spans and F6 (a public-range peer address at the Docker port-forward) calls itself an environment artefact with no action, 6 gap bullets of which four restate F4, F2 and F5 and one says "no gap": 8 items.

  • F1 held up: /stats span p50 218 / p95 474 / p99 504 ms over 426 calls; api profile query 83.58 % self of 87.59 s. Perf.
  • F2a (fan-out) held up: 1 listing + 25 per-SKU GETs per miss on the agent path. Perf.
  • F2b did not hold up: the fan-out's HTTP spans are not "detached from the tool span" - 156 traces rooted at the mcp's search_products contain the api's GET /products/{sku} children; the 0.576 ms childless root it cites is a cache hit, and the 426 rootless GET /products/{sku} are k6's own direct reads.
  • F3 held up: mcp profile create_default_context 57.72 % self (10.88 s of 18.85 s). Perf.
  • F4 held up: the mcp's metric names carry no http_server_*, its roots no HTTP status. Telemetry.
  • F5 held up (labelled suspected): counter total 1,442 (160 + 428 + 427 + 427) against 1,714 tool spans - the 272 missing are search_products cache hits (160 counted of 432). Telemetry.
  • F7 held up: a server/discover inside all 9 POST /ask traces. Perf.
  • G4 held up: no sqlite child under POST /orders. Telemetry.

By kind, confirmed: Telemetry 3 (F4, F5, G4) / Perf 4 (F1, F2a, F3, F7) / Behavior 0.
Figures: launch 05:19:23Z, end 05:51:38Z - preflight 6m42s, drive 2m03s, observation 23m30s, total 32m15s; 34 turns, median 22.3 s, max 1770 s; Input 2,315,434 / Output 74,103 / Cache 1,837,312; cost 0.098333 USD (reconciled to the cent at 0.09 / 0.30 / 0.018 USD per million); signals 4/4 (3 metrics, 7 traces, 3 logs, 2 profiles); 0 source files read (none at all); no traffic of its own; replayable protocol: yes.

z-ai/glm-5.3-flash (opencode) - attempt 1 - killed, no report

Launched 02:49:41Z; drove 02:56:27Z-02:58:28Z (k6's own 121 s), observed, opened the report skeleton at 03:14, then the report-writing generation hit 504 Upstream idle timeout exceeded four times in a row (03:18:48, 03:21:08, 03:23:42, 03:26:31 - the same failure this model showed once under 1.12.0, when the retry passed). Killed by its PID at 03:27:20Z after the fourth identical failure (one more than the rule's third - the fourth landed while the third was being read), 37 min in; the skeleton carried eight <fill> and no draft existed. Nothing to grade. Per the queue rule the model moved to the end of the queue for its second attempt.
Spend to the kill: 0.0885 USD, 2.75M input / 83k output tokens, 34 turns.

anthropic/claude-sonnet-5 (claude) - run 1 - 6 confirmed / 6 reported

Window 01:28:42Z-01:30:44Z (k6's own 121.8 s agrees). 5 anomalies, 2 gaps of which the second restates F2: 6 items.

  • F1 held up: /stats p50 282 / p95 478 / p99 496 ms over 405 calls; api profile query 79.49 % self of 84.38 s; main.py:94-97 and the json.dumps. Perf.
  • F2 held up: counter 163 against 413 span-metric calls; server.py:48-54 returns before the increment at :74-75. Telemetry.
  • F3 held up: tools/call search_products p50 2.64 / p95 454 / p99 500 ms; DETAIL_FANOUT=25 sequential detail fetches, server.py:56-67. Perf.
  • F4 held up: mcp profile create_default_context 61.24 % self (19.21 s of 31.37 s); catalog.py:24-25. Perf.
  • F5 held up (the report itself labels it intended and lists it for the ruling): 19 order rejected … out-of-stock lines answered 200. Behavior.
  • G1 held up: the api's HTTP histograms carry http_target/http_method, the old semconv names, and no http_route/http_request_method (the report notes http_target is already templated). Telemetry.

By kind, confirmed: Telemetry 2 (F2, G1) / Perf 3 (F1, F3, F4) / Behavior 1 (F5).
Figures: launch 01:26:06Z, end 01:40:01Z - preflight 2m36s, drive 2m02s, observation 9m17s, total 13m55s; 84 requests (root + one observe-run subagent), median 2.0 s, max 26 s; Input 10,608,033 (uncached 168, cache read 10,332,615, cache creation 275,250 of which 214,440 at 5m and 60,810 at 1h) / Output 51,959 (thinking 22,640) / Cache 10,607,865; total_cost_usd 3.365789 (costBasis list, a single claude-sonnet-5 key - no haiku background call on this run), reconstructed to the cent at 2.00 / 10.00 / 0.20 / 2.50 / 4.00 USD per million; signals 4/4 (7 metrics, 7 traces, 2 logs, 1 profiles); 0 source files read before the drive (first read 01:33:55); no traffic of its own; replayable protocol: yes. Claude Code 2.1.278, --effort medium, --permission-mode bypassPermissions. Scratch under /tmp/oddobserve-llmbench-store-load-local/.

anthropic/claude-sonnet-5 (claude) - run 2 - VOID (the observation ran on opus-5)

The root session, on claude-sonnet-5, dispatched observe-run through the Agent tool with model: 'opus' of its own accord (01:45:26Z; no contract asks for it, run 1's root passed no model). The subagent's 89 requests ran on claude-opus-5: modelUsage carries claude-opus-5[1m] at 5.466671 USD (54,657 output tokens) beside claude-sonnet-5 at 0.562393 USD (4,566 output tokens). The report - 14 anomalies, 10 gaps, driven (window 01:46:12Z-01:48:14Z) - is opus-5's work and is not graded for this row. Figures for the record: launch 01:42:11Z, end 02:00:48Z, total 18m37s; total_cost_usd 6.029064. Run 1 (6/6, all requests on sonnet-5) is the row.

anthropic/claude-fable-5.1 (claude) - not run: no fable credit left, on the maintainer's word; the 1.12.0 row stays, marked ⚠︎ provisional.

🤖 Generated with Claude Code

using-system and others added 18 commits September 20, 2026 00:03
…pencode under 1.13.0

Run 1: 16/16 confirmed (10/3/3), 29m05s, 0.16 USD. Run 2 stopped at
49 min with its report still generating; the command now says a second
run already worse than the first is stopped, and records the mission's
669-character length since #620.

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…de under 1.13.0

Run 2 is the row: 8/9 confirmed (4/3/1), 10m17s, 1.08 USD; run 1 7/8 in 9m28s at 1.00 USD.

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…13.0

Run 1 is the row: 17/19 confirmed (9/4/4), 19m44s, 1.39 USD; run 2 8/10 in 8m04s at 1.11 USD.

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nder 1.13.0

Run 1 is the row: 11/13 confirmed (4/4/3), 17m09s, 0.33 USD; run 2 11/13 in 13m51s at 0.42 USD (tie, run 1 cheaper).

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…de under 1.13.0

Run 1 is the row: 12/12 confirmed (6/4/2), 20m39s, 2.29 USD; run 2 9/9 in 21m07s at 2.32 USD. Both runs lost an observe-run session to a provider 400 (corrupted thought signature) and re-drove the scenario before reporting.

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…der 1.13.0

Run 1 is the row: 7/8 confirmed (3/2/2), 6m29s, 0.11 USD at list; run 2 5/6 in 6m53s at 0.12 USD.

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…er 1.13.0

Run 1 is the row: 12/13 confirmed (8/4/0), 9m32s, 1.41 USD at list; run 2 9/10 in 9m59s at 1.31 USD.

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nder 1.13.0

Run 2 is the row: 7/8 confirmed (2/3/2), 5m58s, 0.88 USD at list, the root observing without a subagent; run 1 6/7 in 7m44s at 1.25 USD.

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…de under 1.13.0

Run 1 is the row: 6/6 confirmed (2/3/1), 13m55s, 3.37 USD. Run 2 is void: its root dispatched the observation with model: opus, and the command now says a run whose report another model wrote is no row.

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… under 1.13.0

Run 2 is the row: 17/17 confirmed (8/6/3), 19m31s, 6.15 USD; run 1 17/18 in 22m07s at 7.51 USD (tie, run 2 cheaper).

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… under 1.13.0

Run 1 is the row: 15/16 confirmed (8/4/3), 30m25s, 1.50 USD; run 2 14/14 in 30m15s at 1.55 USD.

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…der 1.13.0

Attempt 2 is the row: 7/8 confirmed (3/4/0), 32m15s, 0.10 USD; attempt 1 was killed after four 504s on the report turn.

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…r 1.13.0

Attempt 2 is the row: 11/11 confirmed (5/3/3), 40m25s, 1.62 USD; attempt 1 was killed after a 14-minute silent stream at 45 min.

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Thirteen rows re-measured under 1.13.0, one row (claude-fable-5.1) kept at 1.12.0 and marked provisional; the rank weighs findings, cost and duration together.

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…weighing heavier

The maintainer asked for a heavier negative weight on duration and cost: sol, flashx and luna lead; deepseek and glm-5.3 follow on their counts; opus-5, fable-5.1, glm-5.3-flash, sonnet-5 and qwen3.8-27b close the table on cost or wall clock.

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…a lead, sol and deepseek move down on cost and wall clock

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…a on copilot

The fastest and cheapest row of the 1.13.0 table becomes the instrument: --cli copilot and --model openai/gpt-5.6-luna are the defaults of measure_phase.py and run_samples.py, analyze_run.py falls back to copilot, and copilot syncs no user scope (its deploy lives in the clone) instead of being refused for a missing --scope; the skill says so.

Refs #636

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@using-system
using-system merged commit fe8fe36 into main Sep 20, 2026
12 checks passed
@using-system
using-system deleted the docs/llms-benchmark-1-13-0 branch September 20, 2026 09:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

docs(bench): re-run every row of the llms-benchmark under oddyssey 1.13.0

1 participant