docs(bench): re-run every llms-benchmark row under oddyssey 1.13.0 - #637
Merged
Merged
Conversation
…pencode under 1.13.0 Run 1: 16/16 confirmed (10/3/3), 29m05s, 0.16 USD. Run 2 stopped at 49 min with its report still generating; the command now says a second run already worse than the first is stopped, and records the mission's 669-character length since #620. Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…de under 1.13.0 Run 2 is the row: 8/9 confirmed (4/3/1), 10m17s, 1.08 USD; run 1 7/8 in 9m28s at 1.00 USD. Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…13.0 Run 1 is the row: 17/19 confirmed (9/4/4), 19m44s, 1.39 USD; run 2 8/10 in 8m04s at 1.11 USD. Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nder 1.13.0 Run 1 is the row: 11/13 confirmed (4/4/3), 17m09s, 0.33 USD; run 2 11/13 in 13m51s at 0.42 USD (tie, run 1 cheaper). Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…de under 1.13.0 Run 1 is the row: 12/12 confirmed (6/4/2), 20m39s, 2.29 USD; run 2 9/9 in 21m07s at 2.32 USD. Both runs lost an observe-run session to a provider 400 (corrupted thought signature) and re-drove the scenario before reporting. Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…der 1.13.0 Run 1 is the row: 7/8 confirmed (3/2/2), 6m29s, 0.11 USD at list; run 2 5/6 in 6m53s at 0.12 USD. Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…er 1.13.0 Run 1 is the row: 12/13 confirmed (8/4/0), 9m32s, 1.41 USD at list; run 2 9/10 in 9m59s at 1.31 USD. Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nder 1.13.0 Run 2 is the row: 7/8 confirmed (2/3/2), 5m58s, 0.88 USD at list, the root observing without a subagent; run 1 6/7 in 7m44s at 1.25 USD. Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…de under 1.13.0 Run 1 is the row: 6/6 confirmed (2/3/1), 13m55s, 3.37 USD. Run 2 is void: its root dispatched the observation with model: opus, and the command now says a run whose report another model wrote is no row. Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… under 1.13.0 Run 2 is the row: 17/17 confirmed (8/6/3), 19m31s, 6.15 USD; run 1 17/18 in 22m07s at 7.51 USD (tie, run 2 cheaper). Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… under 1.13.0 Run 1 is the row: 15/16 confirmed (8/4/3), 30m25s, 1.50 USD; run 2 14/14 in 30m15s at 1.55 USD. Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…der 1.13.0 Attempt 2 is the row: 7/8 confirmed (3/4/0), 32m15s, 0.10 USD; attempt 1 was killed after four 504s on the report turn. Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…r 1.13.0 Attempt 2 is the row: 11/11 confirmed (5/3/3), 40m25s, 1.62 USD; attempt 1 was killed after a 14-minute silent stream at 45 min. Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Thirteen rows re-measured under 1.13.0, one row (claude-fable-5.1) kept at 1.12.0 and marked provisional; the rank weighs findings, cost and duration together. Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…weighing heavier The maintainer asked for a heavier negative weight on duration and cost: sol, flashx and luna lead; deepseek and glm-5.3 follow on their counts; opus-5, fable-5.1, glm-5.3-flash, sonnet-5 and qwen3.8-27b close the table on cost or wall clock. Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…a lead, sol and deepseek move down on cost and wall clock Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…a on copilot The fastest and cheapest row of the 1.13.0 table becomes the instrument: --cli copilot and --model openai/gpt-5.6-luna are the defaults of measure_phase.py and run_samples.py, analyze_run.py falls back to copilot, and copilot syncs no user scope (its deploy lives in the clone) instead of being refused for a missing --scope; the skill says so. Refs #636 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #636
Every row of
.llms-benchmark/README.mdre-measured under oddyssey 1.13.0 (e7fd9fa, MCP serveroddyssey-mcp==1.13.0) with thelaunch-llms-benchmarkprotocol as it stands on this branch: every model on its CLI run twice, each run graded on its own, the better run in the table (more confirmed, then cheaper, then shorter). Thirteen rows re-run;anthropic/claude-fable-5.1was not run on the maintainer's word (no fable credit left) and keeps its 1.12.0 row, marked ⚠︎ provisional. 26 launches, 2026-09-19 20:39 UTC to 2026-09-20 06:34 UTC, sequentially, the store reset and the demo stack recreated before each.Proposed ranking
The rank weighs findings, cost and duration together, cost and duration the heavier - the maintainer's weighting for this campaign: a wide report no longer carries a slow or dear run.
z-ai/glm-5.3-flashx(opencode) - 11 / 13 in 17m09s for 0.33 USD: eleven confirmed for a third of a dollar, the best balance of the three axes; two misses. Was feat!: reposition as an odd toolbox with otel instrumentation planning #3.openai/gpt-5.6-luna(copilot) - 7 / 8 in 6m29s for 0.11 USD: the second-cheapest run, the second-fastest; seven confirmed. Was ci(mcp-server): add lint, unit-test, and mcp-client integration jobs #4.openai/gpt-5.6-terra(copilot) - 7 / 8 in 5m58s for 0.88 USD: the fastest run of the table, the root observing without its subagent; eight times luna's cost for the same count. Was docs: add the oddyssey banner to the readme #6.openai/gpt-5.6-sol(copilot) - 12 / 13 in 9m32s for 1.41 USD: twelve confirmed in under ten minutes, one miss; 1.41 USD puts it behind the three cheaper rows. Was feat(agents): harden both agents into true experts with supporting skills #7 at 3.12 USD.google/gemini-3.7-flash(opencode) - 8 / 9 in 10m17s for 1.08 USD: sol's wall clock with four confirmed fewer, a little cheaper. Was feat(mcp): drive docker directly without a compose file #5.z-ai/glm-5.3(opencode) - 17 / 19 in 19m44s for 1.39 USD: the widest count, two misses, twenty minutes. Was feat: bootstrap oddyssey as an apm package for observability-driven development #1.deepseek/deepseek-v4.1-flash(opencode) - 16 / 16 in 29m05s for 0.16 USD: exact, the cheapest confirmed finding of the table (0.010 USD), and twenty-nine minutes - glm-5.3 does as much ten minutes faster. Two further launches on the maintainer's request (below) confirmed the wall clock is the model's under 1.13.0, not the provider's. Was feat: grafana proxy routing, observe-local-run agent, simplified readme #2.google/gemini-3.8-flash(opencode) - 12 / 12 in 20m39s for 2.29 USD: exact; twenty-one minutes and the fourth-dearest run (both its runs lost anobserve-runsession to a provider 400 and drove twice). Was feat(mcp): add odd_stack_reset tool #8.qwen/qwen3.8-max-0902(opencode) - 15 / 16 in 30m25s for 1.50 USD: wide, one miss, thirty minutes. Was fix(mcp): match the no-such-container error case-insensitively #13 at 66 minutes.anthropic/claude-opus-5(claude) - 17 / 17 in 19m31s for 6.15 USD: exact and as wide as glm-5.3, alone on two findings; 6.15 USD, 0.362 per finding, is the price that drops it here. Was feat(agents)!: generalize observe-run to any observability backend #9.anthropic/claude-fable-5.1(claude) ⚠︎ - the 1.12.0 row, 17 / 17 in 17m08s for 7.55 USD, not re-run; after opus-5 on its own figures, provisional until re-run. Was ci(release): auto version bump, github release, and pypi publish #10.z-ai/glm-5.3-flash(opencode) - 7 / 8 in 32m15s for 0.10 USD: the cheapest run, thirty-two minutes for seven confirmed, one miss, after a first attempt its provider would not carry. Was ci(release): git-history-driven release flow (git-cliff) with approved release pr #11.anthropic/claude-sonnet-5(claude) - 6 / 6 in 13m55s for 3.37 USD: exact, the thinnest, the dearest finding of the table (0.561 USD). Was chore(release): 0.1.1 #14.qwen/qwen3.8-27b(opencode) - 11 / 11 in 40m25s for 1.62 USD: exact and wide, alone on two telemetry findings, and forty minutes, the longest run of the table. Was chore(release): 0.1.0 #12.Campaign facts
--variant medium), Claude Code 2.1.278 (--effort medium,--permission-mode bypassPermissions), GitHub Copilot CLI 1.0.86 (--effort medium,--allow-all --no-ask-user). Mission byte-identical across the 26 launches, 669 characters (see the amendments).apm install --global --target claudefrom HEAD before the first run (user-scope bodies identical to.apm/), the user's MCP pin bumped fromoddyssey-mcp==1.12.2to1.13.0;apm install --target opencode|copilotinto the repository before each of those runs, the tree put back after; the three skill copies diffed identical to.apm/skills/before every launch.total_cost_usd(list), reconstructed to the cent from the transcripts - no run carried aclaude-haiku-4-5key.Protocol amendments (
launch-llms-benchmark)modelUsage's keys must be the benchmarked model's (plus the haiku background key) and nothing else carrying spend. Aclaude-sonnet-5root dispatchedobserve-runthrough the Agent tool withmodel: 'opus'of its own accord on one run of two; 89 requests, 90 % of the spend and the whole report were opus-5's. Such a run is void for the row.at full depthphrase; the two places that said 683 now say so.Also on this branch
test-plugin-harnessingdefaults to--cli copilot --model openai/gpt-5.6-luna(86f18f0), the fastest and cheapest row of this table, on the maintainer's word -measure_phase.py,run_samples.pyand the skill say so,analyze_run.pyfalls back to copilot, and copilot syncs no user scope (its deploy lives in the clone) instead of being refused for a missing--scope; a test covers it. Reviewed by the same reviewer: no finding.Notes the table has no column for
Source files read before the drive: 0 on every run. Traffic of its own: none beyond the packaged replay (opus-5 run 1: three
curlsize probes after the drive; gemini-3.8: the scenario replayed twice after the provider crash). Replayable protocol: yes on every report.Review
A separate reviewer sub-agent (fresh context, Opus 5) reconciled every table figure against the row runs and checked the amendments, secrets and language. Six wrong statements found - five in this body, one README legend word (1ccb2a5) - all fixed; the re-check returned no finding. The ranking is the maintainer's and was not reviewed.
Per-finding rulings
Both runs of every row, the row's run first; a restatement counts once, a bundled row once per defect, a self-declared note never.
deepseek/deepseek-v4.1-flash(opencode) - run 1 - 16 confirmed / 16 reportedWindow 20:45:09Z-20:47:11Z (k6's own 121.9 s agrees). 9 anomalies, 8 gaps of which the last restates F8: 16 items.
/statsp50 370 / p95 488 / p99 499 ms over 385 calls against/productsp50 21.9; api profilequery83.19 % self,stats77.47 % total;main.pyselects every column of every row andjson.dumpsthe catalog forpayload_bytes. Perf.create_default_context63.11 % self (18.51 s of 29.33 s),Client.__init__52.54 % total;catalog.pyopens anhttpx.Clientinside every GET and POST. Perf.11676738…carries 25GET /products/{sku}server spans under onetools/call search_products;DETAIL_FANOUT = 25inserver.py; span p95 466 ms; mcp outbound count 5347. Perf.mcp_tool_calls_total{search_products}delta 162 against 391 traces / 392 span-metric calls; 162search done+ 230search served from cachelog lines; the counter sits after the cache early return. Telemetry._SEARCH_CACHEis written at :72, read at :48, cleared nowhere,place_ordernever touches it; 230 cache hits in the window. Behavior.3088a942…has 3invoke_agent, 6chat(3status=ERROR), 3 identicalget_ordertool calls; 4model call failed attempt=n/3log lines with thefinish_reason='error'validation text;assistant.pyretriesMODEL_ATTEMPTSon any exception. Behavior.catalog_orders_created_totalhas the single seriescatalog_category="unknown"(delta 754); the literal is atmain.py:181(the report cites :122 - the line is wrong, the fact is not). Telemetry.initializespan readsmcp.protocol.version=2025-11-25while the same k6 session'stools/listreads2025-06-18(the version the script proposes and accepts); the agent's sessions read2026-07-28. Telemetry.POST /asktraces at 15 s intervals from 20:45:09 to 20:47:09,agent_questions_totaldelta 9; the manifest'sexpected_rateline says 8 and calls it a ceiling. Behavior.http_client_duration_*,mcp_tool_calls_total,target_info- no server histogram. Telemetry.main.py:211excludeshealthfrom the instrumentation; the route logs and leaves no span. Telemetry.{ name = "GET /products/{sku}" } >> { span.db.system.name = "sqlite" }= 0 andPOST /orders= 0, against 763 for the siblingGET /orders/{order_ref};main.py:76and:135calldb.connect()directly. Telemetry.gen_ai_client_token_usage_*and no operation-duration histogram. Telemetry.chatspans across the 9 traces, 3 of them errored,gen_ai_client_token_usage_count= 20. Telemetry.chatspans and the 3 erroredinvoke_agentspans carry noerror.type. Telemetry.profiles checkwithservice_instance_id= 0, without it 100.02 s. Telemetry.By kind, confirmed: Telemetry 10 (F4, F7, F8, G1-G7) / Perf 3 (F1, F2, F3) / Behavior 3 (F5, F6, F9).
Figures: launch 20:39:10Z, end 21:08:15Z - preflight 5m59s, drive 2m02s, observation 21m04s, total 29m05s; 73 turns, median 11.9 s; Input 7,981,775 / Output 93,853 / Cache 7,423,488; cost 0.162325 USD (reconciled to the cent at the base tier); signals 4/4; 0 source files read before the drive (first read 20:53:03, after the drive); no traffic of its own (the drive is the packaged detached replay); replayable protocol: yes (section 7).
deepseek/deepseek-v4.1-flash(opencode) - run 2 - killed, no gradeLaunched 21:12:17Z; drive 21:27-21:29 (15 min of preflight against run 1's 6); the report skeleton was written at 21:55 and the run was still generating the report at 22:01, 49 min in - 20 min past run 1's whole duration - when it was stopped by its PID on the maintainer's instruction (a second run already worse than the first measures nothing). Spend to that point: 0.1648 USD, 4.14M input / 48k output tokens, 47 turns. No report persisted (the
<fill>skeleton was deleted at teardown). The row is run 1.deepseek/deepseek-v4.1-flash(opencode) - run 3 - killed, no reportRe-run at the maintainer's request to check whether the provider had been slow on the first two runs. Launched 08:31:48Z; preflight 9 min (drive 08:41-08:43), observation from 08:43; the report skeleton was opened at 09:01 and the run was stopped by its PID at 09:01:43Z, past run 1's 29m05s total with no report - step 6's rule. Spend to the kill: 0.0961 USD, 4.73M input / 54k output tokens, 60 turns. Run 1 stays the row.
deepseek/deepseek-v4.1-flash(opencode) - run 4 - stopped, no reportSecond re-run at the maintainer's request. Launched 09:02:35Z; preflight 8 min (drive 09:10-09:12), observation by phased helper scripts from 09:13; stopped by its PID at 09:22:50Z on the maintainer's word, 20 min in, no report skeleton yet. Together with run 3 it answers the question the re-runs were asked: under 1.13.0 this model spends six to nine minutes in preflight and twenty minutes or more in observation on every run - the provider was not slow on the first two. Run 1 stays the row. Spend: 0.0745 USD, 2.12M input / 34k output tokens, 33 turns.
z-ai/glm-5.3(opencode) - run 1 - 17 confirmed / 19 reportedWindow 22:33:01Z-22:35:03Z (k6's own 121.9 s agrees). 14 anomalies, 6 gaps of which the last (log bodies carry no trace id while the structured metadata correlates completely, by the report's own words) is a note, not a gap: 19 items.
/statsp50 268 / p95 477 / p99 495 ms over 421 calls, sum 89.4 s; api profilequery77.32 % self. Perf.tools/call search_productsp95 423 ms; the exemplar carries 26 downstream GETs;DETAIL_FANOUT = 25. Perf.order createdlines in the window carry 662 distinct refs; the mcp log showsORD-000617returned for two accepted orders of different SKUs (traces47139af7…,1629bb55…);main.py:137-138derives the ref fromSELECT COUNT(*) FROM orders+ 1 outside any lock. Behavior.create_default_context62.25 % self (17.48 s of 28.08 s);catalog.pyper-call client. Perf.mcp_tool_calls_total{search_products}delta 160 against 427 span-metric calls; the other three tools 423/422/422 exact. Telemetry.search served from cachelines;_SEARCH_CACHEis never invalidated byplace_order. Behavior.POST /asktraces, the manifest says 8. Behavior.POST /asktrace carries aserver/discoverand atools/listof its own (24tools/listspans in the window, 15 rooted at the k6 sessions). Perf.POST /orders192.9 ms "with no child span, the time is contention") is asserted, not separated from the route's ownCOUNT(*)+ insert + commit; the report itself calls it largely a consequence of F1.chatspans carrygen_ai.input.messages,gen_ai.output.messages,gen_ai.system_instructions. Telemetry.gen_ai.system=openaibesidegen_ai.provider.name=openai. Telemetry.catalog_category="unknown"single series;main.py:181. Telemetry.order rejected … out-of-stockWARN lines answered 200; 411 read-backs against 421 k6 orders. Behavior.agent/app/main.py:68instruments withoutexcluded_urls,api/app/main.py:211excludeshealth. Telemetry.http_server_*. Telemetry.initializeand 5notifications/initializedtraces rooted atllmbench-mcpsit in the window; the handshake is traced.mcp.*andjsonrpc.*attributes only - nohttp.*, no user agent - so the run's identity cannot select them. Telemetry.gen_ai.client.operation.durationamong the agent's 18 metric names. Telemetry.profiles checkwithservice_instance_id= 0; the SDK pushesprocess_cpuonly. Telemetry.By kind, confirmed: Telemetry 9 (F5, F10, F11, F12, F14, G1, G3, G4, G5) / Perf 4 (F1, F2, F4, F8) / Behavior 4 (F3, F6, F7, F13).
Figures: launch 22:29:44Z, end 22:49:28Z - preflight 3m17s, drive 2m02s, observation 14m25s, total 19m44s; 44 turns, median 9.4 s; Input 4,419,244 / Output 120,536 / Cache 4,014,656; cost 1.391385 USD (reconciled to the cent, flat rates); signals 4/4 (16 metrics, 9 traces, 13 logs, 3 profiles); 0 source files read before the drive (first read 22:42:45); no traffic of its own; replayable protocol: yes.
z-ai/glm-5.3(opencode) - run 2 - 8 confirmed / 10 reportedWindow 22:55:14Z-22:57:16Z (k6's own 122.0 s agrees). 9 anomalies, 5 gaps of which two restate F6 and F4, one is a mapping note and one says "no gap" by its own words: 10 items.
/statsp50 289 / p95 485 / p99 577 ms over 402 calls; api profilequery79.85 % self. Perf.tools/call search_productsp50 2.91 / p95 461 ms, 409 calls;GET /products/{sku}4831 calls. Perf.create_default_context61.70 % self (18.7 s of 30.31 s). Perf.POST /asktraces against the manifest's 8. Behavior.http_server_*. Telemetry.GET/POST, no route or destination in the name. Telemetry.?categoryand?category&qinto oneGET /productsspan and histogram is what the HTTP conventions prescribe (http.routecarries no query string); the two names are the k6 script's, not the service's - the report itself calls it a mapping note.order rejected … out-of-stockWARN lines answered 200,SKU-02987among them repeatedly. Behavior.span.gen_ai.conversation.id != ""matches all 9 ask traces; the attribute is present.By kind, confirmed: Telemetry 3 (F4, F6, F7) / Perf 3 (F1, F2, F3) / Behavior 2 (F5, F9).
Figures: launch 22:52:43Z, end 23:00:47Z - preflight 2m31s, drive 2m02s, observation 3m31s, total 8m04s; 59 turns, median 2.5 s; Input 4,712,775 / Output 35,933 / Cache 4,422,720; cost 1.114158 USD (reconciled to the cent); signals 4/4; 0 source files read before the drive (first read 22:59:03); no traffic of its own; replayable protocol: yes.
Run 1 (17/19) is the row on confirmed findings.
openai/gpt-5.6-sol(copilot) - run 1 - 12 confirmed / 13 reportedWindow 00:46:25Z-00:48:25Z (k6's own 120.4 s agrees). 7 anomalies of which F2 bundles two defects, 7 gaps of which two restate F6 and F4: 13 items.
/statsp99 492 ms, trace1fdacecb…477 of 504 ms incatalog stats_scan, api profilequery72.72 % self of 61.83 s; its "p50 45→275 ms half-to-half" is not what the histogram gives (93→149 ms over the two halves) - the scan is the finding, the ratio is not. Perf.tools/call search_productsp95 242 / p99 386 ms, the worst root carries 25 detail GETs. Perf.create_default_context62.75 % self (18.14 s of 28.91 s). Perf.gen_ai_client_token_usage{input}p50 768 / p95 36,040 over 18 calls; the worst ask 6.17 s. Perf.tools/call search_productstraces / 450 span-metric calls. Telemetry.gen_ai.input.messages,gen_ai.output.messages,gen_ai.system_instructions, the tool spansgen_ai.tool.call.arguments/result. Telemetry.http.*attribute. Telemetry.order rejected … out-of-stock, HTTP 200 throughout. Telemetry.gen_ai_client_operation_duration. Telemetry.service_instance_id. Telemetry.--trace-idprofile query for the agent returns 0 frames against 1.13 CPU-s for the same window - no span-linked profiles. Telemetry.memory:alloc_spaceis 0 for each of the three services while the store-wide selector holds 239 MB (another service's). Telemetry.checks_totalsits in the store; a service cannot emit a client's assertions.By kind, confirmed: Telemetry 8 (F4, F5, F6, F7, G2, G3, G4, G5) / Perf 4 (F1, F2a, F2b, F3) / Behavior 0.
Figures: launch 00:45:04Z, end 00:54:36Z - preflight 1m21s, drive 2m00s, observation 6m11s, total 9m32s; 48 turns (root + two observe-run dispatches), median model-call latency 4.2 s, max 63 s; Input 3,747,709 (uncached 212,349, cache read 3,535,360, cache write 0) / Output 27,792 (reasoning 7,747) / Cache 3,535,360; cost at OpenAI list 212,349x2.00/M + 3,535,360x0.20/M + 27,792x10.00/M = 1.409690 USD; 1 premium request, 281.94 AIU; signals 4/4 (15 metrics, 16 traces, 4 logs, 12 profiles); 0 source files read before the drive (first view 00:51:20); no traffic of its own; replayable protocol: yes. Copilot CLI 1.0.86, --effort medium. Scratch under
/tmp/oddyssey/.openai/gpt-5.6-sol(copilot) - run 2 - 9 confirmed / 10 reportedWindow 00:58:15Z-01:00:17Z (k6's own 121.9 s agrees). 5 anomalies, 5 gaps: 10 items.
query79.40 % self of 90.29 s, the/statsscan (the "p50 148→389 ms" drift is a histogram artefact, as in run 1). Perf.create_default_context59.52 % self (16.76 s of 28.16 s). Perf.tools/call search_productstraces, 242 cache-hit lines (162 + 242 = 404). Telemetry.order rejected … out-of-stockWARN lines, HTTP 200 throughout. Telemetry.memory:alloc_spaceis 0 per service. Telemetry.service_instance_id. Telemetry.gen_ai_client_operation_duration. Telemetry.http_method,http_target,http_status_code(the old semconv names); a grouping on the stable names collapses. Telemetry.By kind, confirmed: Telemetry 7 (F3, F5, G1-G5) / Perf 2 (F1, F2) / Behavior 0.
Figures: launch 00:56:55Z, end 01:06:54Z - preflight 1m20s, drive 2m02s, observation 6m37s, total 9m59s; 47 turns, median 4.1 s, max 70 s; Input 3,487,692 (uncached 183,244, cache read 3,304,448, cache write 0) / Output 27,895 (reasoning 8,210) / Cache 3,304,448; cost at list 1.306328 USD; 1 premium request, 261.27 AIU; signals 4/4 (20 metrics, 11 traces, 8 logs, 10 profiles); 0 source files read before the drive; no traffic of its own; replayable protocol: yes.
Run 1 (12/13) is the row on confirmed findings.
z-ai/glm-5.3-flashx(opencode) - run 1 - 11 confirmed / 13 reportedWindow 23:06:21Z-23:08:22Z (k6's own 121.5 s agrees). 10 anomalies of which F10 says "explained, not a defect" by its own words, 4 gaps of which G4 restates F10 (counted once, as the gap): 13 items.
mcp_tool_calls_total{search_products}delta 160 against 423tools/call search_productstraces; 263 cache-hit log lines. Telemetry.place_ordernever touches_SEARCH_CACHE. Behavior./statsp50 257 / p95 483 / p99 604 ms over 417 calls; api profilequery77.42 % self (56.15 s of 72.53 s). Perf.fbe40b4f…lists 718 rows (db.response.returned_rows=718) then fetches 25GET /products/{sku}one by one. Perf.gen_ai_client_token_usage{input}p50 779 / p95 12,580 over 19 calls, the exemplar's second chat at 22,663 input tokens. Perf.create_default_context61.89 % self (18.48 s of 29.86 s). Perf.POST /asktraces against the manifest's 8. Behavior.order rejected … out-of-stocklines answered 200,SKU-029874 times. Behavior.gen_ai.*metric names aregen_ai_client_token_usage_*only. Telemetry.gen_ai.response.model=google/gemini-3.5-flash-lite(the report itself marked it suspected, on a truncated rendering).http_targetwith the route template (/products/{sku}), and the report's own per-route histogram was grouped by it; the label is the old semconv name, not a missing dimension./healthis excluded from the api's instrumentation (main.py:211), 45 pre-run access lines and 0 traces. Telemetry.By kind, confirmed: Telemetry 4 (F1, F9, G1, G4) / Perf 4 (F3, F4, F5, F6) / Behavior 3 (F2, F7, F8).
Figures: launch 23:03:11Z, end 23:20:20Z - preflight 3m10s, drive 2m01s, observation 11m58s, total 17m09s; 32 turns, median 13.3 s; Input 2,345,736 / Output 68,360 / Cache 2,102,912; cost 0.333013 USD (reconciled to the cent); signals 4/4 (8 metrics, 6 traces, 8 logs, 3 profiles); 0 source files read before the drive (first read 23:13:07); no traffic of its own; replayable protocol: yes.
z-ai/glm-5.3-flashx(opencode) - run 2 - 11 confirmed / 13 reportedWindow 23:24:48Z-23:26:49Z (k6's own 121.1 s agrees). 10 anomalies, 5 gap bullets of which one says "client-side limitation, not the stack's" and one "no gap bullet" by their own words: 13 items.
/statsp50 273 / p95 490 / p99 510 ms; api profilequery81.13 % self. Perf.create_default_context60.93 % self (14.21 s of 23.32 s). Perf.order createdlines carry 616 distinct refs,ORD-000002created four times;main.py:137-138. Behavior.tools/call search_productsover 100 ms, decaying per bin as the cache fills; p95 445 ms against p50 3.19. Perf.f8dc6b75…has 3invoke_agentand 7chatspans, two of themstatus=ERRORon an HTTP 200 answer, then the third attempt succeeds; 2model call failedWARN lines. Behavior.search_productsagainst 406 calls, the other three tools exact. Telemetry._SEARCH_CACHEnever invalidated while 784 orders mutated stock. Behavior.order rejected … out-of-stocklines answered 200. Behavior.catalog_category="unknown"single series;main.py:181. Telemetry.final_resultof tracef8dc6b75…reads "215608 euros (2156,08 €)" fortotal_cents=215608. Behavior.service_instance_id. Telemetry.http_server_response_size_bytes_*carrieshttp_targetwith the route template; the sizes split per route under that label.By kind, confirmed: Telemetry 3 (F6, F9, G2) / Perf 3 (F1, F2, F4) / Behavior 5 (F3, F5, F7, F8, F10).
Figures: launch 23:22:31Z, end 23:36:22Z - preflight 2m17s, drive 2m01s, observation 9m33s, total 13m51s; 44 turns, median 9.2 s; Input 3,669,344 / Output 69,555 / Cache 3,458,304; cost 0.424401 USD (reconciled to the cent); signals 4/4 (through its
batch*.shhelpers - the log alone counts 1/1/1/0); 0 source files read before the drive (first read 23:32:50); no traffic of its own; replayable protocol: yes.Tie with run 1 on confirmed findings (11 each); run 1 is cheaper (0.33 against 0.42 USD) and is the row.
anthropic/claude-opus-5(claude) - run 2 - 17 confirmed / 17 reported - THE ROWWindow 02:30:09Z-02:32:11Z (k6's own 121.0 s agrees). 8 anomalies of which F3 and F7 each bundle two defects, 8 gaps of which G3 restates F4: 17 items.
order createdlines carry 638 distinct refs, 130 refs shared by 303 orders; 193catalog get_orderspans return 2 rows;main.py:137-138. Behavior./statsp50 283 / p99 507 ms; api profilequery85.82 % self of 88.6 s. Perf.tools/call search_productsp50 1.92 / p95 240 ms over 420 calls, 26 sequential GETs per miss. Perf.create_default_context58.60 % self (12.78 s of 21.81 s). Perf.POST /ordersspans read p50 27 ms (n=415) against 10 ms for the mcp-routed ones (n=416); the writer waiting on/statsreaders is the mechanism this campaign verified by overlap on another run. Perf.order rejected … out-of-stockWARN lines answered 200. Behavior.adbb858e…carries a 22,930-input-token chat on a 718-row search. Perf.final_resultreads "priced at 208100 euros" for aprice_centsvalue and namesSKU-0733. Behavior.POST /ordersandGET /products/{sku}. Telemetry.gen_ai_client_operation_duration. Telemetry.service.instance.id. Telemetry.gen_ai.systembesidegen_ai.provider.name. Telemetry.GET /healthaccess lines without a trace id, the route excluded from traces and metrics. Telemetry.By kind, confirmed: Telemetry 8 (F4, G1, G2, G4, G5, G6, G7, G8) / Perf 6 (F2, F3a, F3b, F5, F7a, F8) / Behavior 3 (F1, F6, F7b).
Figures: launch 02:27:22Z, end 02:46:53Z - preflight 2m47s, drive 2m02s, observation 14m42s, total 19m31s; 53 requests (root + one observe-run subagent,
modelUsagethe single keyclaude-opus-5), median 6.6 s, max 60 s; Input 5,768,520 (uncached 106, cache read 5,504,746, cache creation 263,668 of which 210,576 at 5m and 53,092 at 1h) / Output 61,893 (thinking 22,212) / Cache 5,768,414;total_cost_usd6.147248 (list), reconstructed to the cent; signals 4/4 (4 metrics, 8 traces, 5 logs, 1 profiles in the log, more inside its helpers); 0 source files read before the drive (first read 02:39:09); no traffic of its own; replayable protocol: yes. Claude Code 2.1.278, --effort medium. Scratch under/private/tmp/oddyssey-scratch/.Tie with run 1 (17/18) on confirmed findings; run 2 is cheaper (6.15 against 7.51 USD) and is the row.
anthropic/claude-opus-5(claude) - run 1 - 17 confirmed / 18 reportedWindow 02:05:59Z-02:08:00Z (k6's own 120.3 s agrees). 10 anomalies of which F1 bundles two defects and F10 (throughput drifting as the cache warms) is F3's consequence, 10 gap bullets of which G3 restates F4 and the last says "not a gap": 18 items.
6b189e6c…106 spans,tools/call search_productsp95 473 / p99 504 ms over 417 calls. Perf.create_default_context60.64 % self (18.18 s of 29.98 s);catalog.py:25. Perf.query73.89 % self of 72.26 s;payload_bytes=1838708in the/statslog lines;main.py:92-131. Perf.place_order. Behavior.order rejected … stock=0WARN lines answered 200; the mcp logs 11 of them INFOorder placed … 'order_ref': None. Behavior.ffa12470…carries 2invoke_agent, 5chatspans, 2 errored, the retry re-running the whole agent (new discover + tools/list + tool call). Behavior.GET /products?category=camerasanswers 43,990 bytes,matches=417on 278 log lines;/productsp50 9.1 / p95 60 ms against 2.6 for a key read. Perf.gen_ai_client_token_usage{input}p50 736 / p95 10,240 / p99 15,160; the search questions at 12.7k tokens. Perf.server/discoverand atools/listinside everyPOST /asktrace. Perf.POST /ordersandGET /products/{sku}(structural queries 0). Telemetry.http_server_*series; its tool spans are roots without an HTTP status. Telemetry.http_server_response_size_byteson/productsreads p50 9,657 / p95 10,000 / p99 10,000 for a mean of 22,240 bytes - the histogram carries the duration default buckets (le=0, 5, 10 … 10,000), every listing lands in +Inf. Telemetry.http_target(the route template) and nohttp_route. Telemetry.http.targetdropping the query string and folding?categoryand?qunder one route is what the HTTP conventions prescribe; the k6 names are the client's.GET/POSTwithhttp.urlas the only key. Telemetry.chatspans, the errored one carries none. Telemetry.service.instance.id. Telemetry.By kind, confirmed: Telemetry 8 (F4, G1, G2, G4, G5, G7, G8, G9) / Perf 6 (F1a, F1b, F2, F7, F8, F9) / Behavior 3 (F3, F5, F6).
Figures: launch 02:02:55Z, end 02:25:02Z - preflight 3m04s, drive 2m01s, observation 17m02s, total 22m07s; 60 requests (root + one observe-run subagent, both on opus-5 -
modelUsagecarries the single keyclaude-opus-5), median 5.9 s, max 71 s; Input 6,670,041 (uncached 120, cache read 6,244,216, cache creation 425,705 of which 369,645 at 5m and 56,060 at 1h) / Output 60,840 (thinking 18,659) / Cache 6,669,921;total_cost_usd7.514589 (list), reconstructed to the cent at 5.00 / 25.00 / 0.50 / 6.25 / 10.00 USD per million; signals 4/4 (itsq1-q5.shhelpers carry 14 metrics, 13 traces, 8 logs, 3 profiles invocations); 0 source files read before the drive (first read 02:15:27, after the queries); traffic of its own: threecurlsize probes ofGET /productsat 02:15, after the drive, outside the window - measurement, not a scenario; replayable protocol: yes. Claude Code 2.1.278, --effort medium.qwen/qwen3.8-max-0902(opencode) - run 1 - 15 confirmed / 16 reportedWindow 04:19:07Z-04:21:08Z (k6's own 121.7 s agrees). 9 anomalies of which F3 bundles two defects, 6 gaps: 16 items.
place_order. Behavior.tools/call search_productsp50 1.75 / p95 235 ms, 106-span exemplar. Perf.gen_ai_client_token_usage{input}p50 784 / p95 13,000 over 19 calls, the heaviest at 22,930. Perf.create_default_context58.04 % self (11.23 s of 19.35 s). Perf./statsp50 326 / p95 483 / p99 497 ms over 438 calls, sum 91.3 s; api profilequery83.25 % self. Perf.POST /asktraces,agent_questions_total9, the manifest says 8. Behavior.gen_ai.system=openaibesidegen_ai.provider.name=openaion the chat spans and as a metric label. Telemetry.a1d608e0…carriesfinal_resultoninvoke_agentandgen_ai.tool.call.argumentson the tool spans. Telemetry.order rejected … out-of-stocklines answered 200. Behavior.gen_ai_client_operation_duration. Telemetry.GET /products/{sku}(structural query 0);main.py:76-84. Telemetry.Created new transport with session IDlines without a trace id, 883 of 888 with one. Telemetry.tools/callroot spans carry no HTTP status attribute. Telemetry.dropped_iterationsseries absent when its value is zero is k6's own OTel export at work, not the stack's telemetry, and the report says it ruled the schedule from the services' rows anyway.By kind, confirmed: Telemetry 8 (F1, F7, F8, G1-G5) / Perf 4 (F3a, F3b, F4, F5) / Behavior 3 (F2, F6, F9).
Figures: launch 04:15:03Z, end 04:45:28Z - preflight 4m04s, drive 2m01s, observation 24m20s, total 30m25s; 34 turns, median 22.5 s, max 1752 s (the report-writing generation); Input 2,766,531 / Output 67,087 / Cache 2,534,656; cost 1.499936 USD (reconciled to the cent, flat rates 2.00 / 6.00 / 0.25 USD per million); signals 4/4 (9 metrics, 10 traces, 4 logs, 1 profiles); 0 source files read before the drive (first read 04:33:47); no traffic of its own; replayable protocol: yes.
qwen/qwen3.8-max-0902(opencode) - run 2 - 14 confirmed / 14 reportedWindow 04:52:18Z-04:54:23Z (k6's own 124.4 s agrees). 7 anomalies of which F2 and F6 each bundle two defects, 6 gaps of which G5 restates F7: 14 items.
/statsp50 261 / p95 476 / p99 495 ms over 433 calls; api profilequery83.17 % self. Perf.tools/call search_productsp95 237 ms. Perf.create_default_context57.14 % self (11.52 s of 20.16 s). Perf.gen_ai_client_token_usage{input}p50 784 / p95 13,000, one question at 22,930. Perf.script.js:493const lookupRef = mcpOrderRef || apiOrderRef- the tool-sideget_orderran 434 times while 11place_ordercalls were rejected, reading back the HTTP-side order the manifest's "every iteration in which an order reference came back" does not mean. Behavior (benchmark fidelity).POST /asktraces against the manifest's 8. Behavior.order rejected … out-of-stocklines answered 200. Behavior.gen_ai_client_operation_duration. Telemetry.http_server_*series on the mcp. Telemetry.gen_ai.systembesidegen_ai.provider.name. Telemetry.By kind, confirmed: Telemetry 6 (F3, G1, G2, G3, G4, G6) / Perf 4 (F1, F2a, F2b, F5) / Behavior 4 (F4, F6a, F6b, F7).
Figures: launch 04:47:15Z, end 05:17:30Z - preflight 5m03s, drive 2m05s, observation 23m07s, total 30m15s; 46 turns, median 17.1 s, max 1707 s; Input 3,429,113 / Output 62,152 / Cache 3,246,208; cost 1.550274 USD (reconciled to the cent); signals 4/4 (through its
batch*.shanddiscover.shhelpers - the log alone counts 1/0/0/0); 0 source files read before the drive; no traffic of its own; replayable protocol: yes.Run 1 (15/16) is the row on confirmed findings.
google/gemini-3.8-flash(opencode) - run 1 - 12 confirmed / 12 reportedThe observation subagent's stream died at 23:45:23 and 23:46:32 on a 400
Corrupted thought signaturefrom Google AI Studio through OpenRouter, after its first drive (23:42:10-23:44:11); the root dispatched the observation again, which drove the scenario a second time (window 23:48:54Z-23:50:57Z, k6's own 121.8 s agrees) against the same containers, whose store already carried the first drive. The report's window is the second drive; the run's own earlier traffic sits behind it (the search cache was warm: every one of the 409 searches was a cache hit). 10 anomalies, 6 gaps of which four restate F6 (twice), F3 and F7: 12 items./statsp50 375 / p95 488 ms; api profilequery88.10 % self of 109.19 s. Perf.order rejected … out-of-stockWARN lines answered 200 (the catalog was already depleted by the first drive);main.py:157. Behavior.mcp_tool_calls_total{search_products}delta 0 against 409tools/call search_productstraces and 409 cache-hit log lines. Telemetry.create_default_context42.80 % self (4.22 s of 9.86 s - a smaller share because no search missed the cache). Perf.gen_ai.input.messages,gen_ai.output.messages,gen_ai.system_instructionson the chat spans;assistant.py:56. Telemetry.POST /ordersandGET /products/{sku}both 0;main.py:76,:135. Telemetry.catalog_category="unknown"single series, delta 765;main.py:181. Telemetry.c0618795…chat spans at 267 and 22,663 input tokens (22,930 together). Perf.POST /asktraces carryserver/discoverandtools/list. Perf.order createdlines in the window carry 563 distinct refs (148 duplicated);main.py:137-138. Behavior.http_server_*. Telemetry.gen_ai_client_operation_durationamong the agent's names. Telemetry.By kind, confirmed: Telemetry 6 (F3, F5, F6, F7, G1, G6) / Perf 4 (F1, F4, F8, F9) / Behavior 2 (F2, F10).
Figures: launch 23:38:56Z, end 23:59:35Z - preflight 9m58s (first drive and crash included), drive 2m03s, observation 8m38s, total 20m39s; 148 turns over 3 sessions, median 4.4 s; Input 12,787,753 / Output 59,361 / Cache 11,138,874; cost 2.294679 USD (reconciled to the cent, flat rates); signals 4/4 (12 metrics, 7 traces, 5 logs, 2 profiles); 0 source files read before the drive (first read 23:55:06); no traffic of its own beyond the two replays of the stored scenario; replayable protocol: yes.
google/gemini-3.8-flash(opencode) - run 2 - 9 confirmed / 9 reportedThe same 400
Corrupted thought signaturekilled twoobserve-runsessions (00:10:26 and 00:11:21) after the first drive (00:07:43-00:09:43); the root's third dispatch drove again and reported on that window (00:13:15Z-00:15:16Z, k6's own 120.8 s agrees), the first drive's traffic sitting behind it in the store. 7 anomalies, 4 gaps of which two restate F7 and F5: 9 items.catalog stats_scan473 ms of a 492 ms/statstrace; api profilequery89.07 % self of 112.65 s. Perf.ORD-000770on twoorder createdlines,SKU-00169andSKU-00012, traces6b61e121…and7b8d8ed1…. Behavior./ordersp50 9.3 / p95 77 / p99 176 ms; every one of the five slowestPOST /orderstraces (120-210 ms) overlaps one to four/statsscans of 130-410 ms, the reader the writer waits on. Perf.e9b91cae…carries 2invoke_agentand 5chatspans with onestatus=ERROR, and thefinish_reason … input_value='error'WARN line. Behavior.tools/call search_productstraces (the cache warmed by the first drive). Telemetry.create_default_context43.64 % self (4.36 s of 9.99 s). Perf.catalog_category="unknown"on every order;main.py:181. Telemetry.POST /orders(structural query 0). Telemetry.search_productsspans carry no cache-hit attribute; nothing inserver.pysets one. Telemetry.By kind, confirmed: Telemetry 4 (F5, F7, G1, G4) / Perf 3 (F1, F3, F6) / Behavior 2 (F2, F4).
Figures: launch 00:01:51Z, end 00:22:58Z - preflight 11m24s (first drive and two crashes included), drive 2m01s, observation 7m42s, total 21m07s; 167 turns over 5 sessions, median 3.7 s; Input 12,496,566 / Output 68,181 / Cache 10,830,668; cost 2.317402 USD (reconciled to the cent); signals 4/4 (6 metrics, 13 traces, 11 logs, 2 profiles); 0 source files read before the drive (first read 00:17:29); no traffic of its own beyond the two replays; replayable protocol: yes.
Run 1 (12/12) is the row on confirmed findings.
openai/gpt-5.6-luna(copilot) - run 1 - 7 confirmed / 8 reportedWindow 00:28:19Z-00:30:20Z (k6's own 121.2 s agrees). 5 anomalies, 4 gaps of which the first restates F5: 8 items.
query71.13 % self of 65.18 s,/statsthe only route in the hundreds of ms (the trace id it cites,79f30bdf…, is F2's retried ask, not a/statstrace - the profile carries the finding). Perf.79f30bdf…has 2 errored spans and 4chatcalls, onemodel call failed … finish_reason='error'WARN line, HTTP 200 kept. Behavior.create_default_context61.47 % self (17.93 s of 29.17 s). Perf.order rejected … out-of-stockWARN lines answered 200;main.py:157. Behavior.profiles labels --label service_instance_idempty,checkrestores data only without the label. Telemetry.http_targetwith the route template, and per-route rows come out of it; the report grouped byhttp_route, a label the old semconv the app uses never emitted.Created new transport with session IDlines carry no trace id, 895/900 lines do. Telemetry.By kind, confirmed: Telemetry 3 (F5, G3, G4) / Perf 2 (F1, F3) / Behavior 2 (F2, F4).
Figures: launch 00:27:17Z, end 00:33:46Z - preflight 1m02s, drive 2m01s, observation 3m26s, total 6m29s; 38 turns (root + one observe-run subagent), median model-call latency 2.8 s, max 28 s; Input 2,932,752 (uncached 114, cache read 2,739,474, cache write 193,164) / Output 17,189 (reasoning 3,940) / Cache 2,932,638; cost at OpenAI list (114+193,164)x0.20/M + 2,739,474x0.02/M + 17,189x1.20/M = 0.114072 USD; 1 premium request, 12.37 AIU; signals 4/4 (11 metrics, 5 traces, 8 logs, 7 profiles); 0 source files read before the drive (four viewed at 00:32:37); no traffic of its own; replayable protocol: yes. Copilot CLI 1.0.86, --effort medium. Scratch under
.odd/scratch/inside the repository (untracked, cleared at teardown).openai/gpt-5.6-luna(copilot) - run 2 - 5 confirmed / 6 reportedWindow 00:37:05Z-00:39:05Z (k6's own 123.3 s agrees). 5 anomalies of which F2 bundles two defects and F3 restates F2's consequence ("model-bound" is no defect); 6 gap bullets of which two restate F4 and three call themselves "by design", "filled/expected" and "not probed": 6 items.
/statstrace p95 464 ms; api profilequery75.01 % self of 78.08 s. Perf.b774d0b2…carries 25GET /products/{sku}under one search;DETAIL_FANOUT=25. Perf.create_default_context60.63 % self (16.97 s of 27.99 s). Perf.profiles labels --label service_instance_idreturns null (its other half - metrics carry instance UUIDs rather than the slug - is how the identity is meant to travel, and the report's own frontmatter maps them). Telemetry.order rejected … out-of-stockWARN lines (the report counts 22 WARN), all answered 200. Behavior.checks_total{condition="nonzero"}= 9847 is in the store, exported by k6 - the count is not "only in the transient summary", and a service cannot emit a client's assertions.By kind, confirmed: Telemetry 1 (F4) / Perf 3 (F1, F2a, F2b) / Behavior 1 (F5).
Figures: launch 00:35:54Z, end 00:42:47Z - preflight 1m11s, drive 2m00s, observation 3m42s, total 6m53s; 47 turns, median 2.4 s; Input 3,508,557 (uncached 141, cache read 3,363,979, cache write 144,437) / Output 19,811 (reasoning 5,130) / Cache 3,508,416; cost at list 0.119968 USD; 1 premium request, 12.72 AIU; signals 4/4; 0 source files read before the drive; no traffic of its own; replayable protocol: yes.
Run 1 (7/8) is the row on confirmed findings.
openai/gpt-5.6-terra(copilot) - run 2 - 7 confirmed / 8 reported - THE ROWThe root tried
observe-runthrough Copilot'sskilltool ("Skill not found"), never dispatched the subagent, and did the whole observation itself. Window 01:18:17Z-01:20:18Z (k6's own 121.1 s agrees). 5 anomalies of which F3 bundles two defects, 3 gaps of which G1 restates F4: 8 items.POST /asktraces answers 500; 4model call failedWARN lines and 1ask failed; the validation text namesfinish_reason='error';assistant.pyretries three times. Behavior./statstrace p95 493 ms,catalog stats_scanp95 422 ms; api profilequery81.42 % self of 91.88 s. Perf.create_default_context61.70 % self (19.14 s of 31.02 s). Perf.tools/call search_productsp95 474 ms, 25 detail fetches per miss. Perf.http_methodandhttp_targetsplit the api's histogram into its five routes; the report queried only the newer nameshttp_request_methodandhttp_routeand called the dimensions absent.order rejected … out-of-stockWARN lines (the report counts 21), 761 accepted against 779 order spans. Behavior.By kind, confirmed: Telemetry 2 (G2, G3) / Perf 3 (F2, F3a, F3b) / Behavior 2 (F1, F5).
Figures: launch 01:17:41Z, end 01:23:39Z - preflight 0m36s, drive 2m01s, observation 3m21s, total 5m58s; 26 turns (root only, no subagent), median 3.2 s, max 27 s; Input 2,438,676 (uncached 78, cache read 2,302,150, cache write 136,448) / Output 12,531 (reasoning 3,393) / Cache 2,438,598; cost at OpenAI list (78+136,448)x2.00/M + 2,302,150x0.20/M + 12,531x12.00/M = 0.883854 USD; 1 premium request, 95.21 AIU; signals 4/4 (3 metrics, 2 traces, 2 logs, 1 profiles); 0 source files read before the drive (two viewed at 01:22:42); no traffic of its own; replayable protocol: yes.
Chosen over run 1 (6/7, 7m44s, 1.25 USD) on confirmed findings.
openai/gpt-5.6-terra(copilot) - run 1 - 6 confirmed / 7 reportedWindow 01:09:55Z-01:11:58Z (k6's own 121.9 s agrees). 5 anomalies, 3 gaps of which G3 restates F1: 7 items.
order rejected … out-of-stockWARN lines; the api roots answer 200 with status UNSET; k6's failure rate 0. Behavior./statsp50 257 / p95 476 / p99 495 ms over 412 calls; api profilequery79.80 % self of 85.43 s. Perf.create_default_context61.64 % self (18.93 s of 30.71 s). Perf.gen_ai.input.messages(3 of 3 sampled), output messages, tool arguments and results,final_result. Telemetry.POST /asktraces,agent_questions_total9, the manifest says 8. Behavior.checks_totalsits in the store; the threshold is rulable from it, and a service cannot emit a client's assertions.service_instance_id. Telemetry.By kind, confirmed: Telemetry 2 (F4, G2) / Perf 2 (F2, F3) / Behavior 2 (F1, F5).
Figures: launch 01:08:39Z, end 01:16:23Z - preflight 1m16s, drive 2m03s, observation 4m25s, total 7m44s; 45 turns (root + one observe-run subagent), median 3.4 s, max 38 s; Input 3,285,222 (uncached 135, cache read 3,077,833, cache write 207,254) / Output 18,269 (reasoning 5,690) / Cache 3,285,087; cost at OpenAI list (135+207,254)x2.00/M + 3,077,833x0.20/M + 18,269x12.00/M = 1.249573 USD; 1 premium request, 135.32 AIU; signals 4/4 (8 metrics, 8 traces, 3 logs, 3 profiles); 0 source files read before the drive (three grepped at 01:14:53); no traffic of its own; replayable protocol: yes. Copilot CLI 1.0.86, --effort medium. Scratch under
/tmp/oddyssey-observe/.google/gemini-3.7-flash(opencode) - run 2 - 8 confirmed / 9 reported - THE ROWWindow 22:19:24Z-22:21:26Z (k6's own 121.8 s agrees). 6 anomalies of which F1 and F5 each bundle two defects with two fixes, 4 gaps of which three restate F6, F5b and F3: 9 items.
beab7a03…carries 25GET /products/{sku}server spans under onetools/call search_products; span p95 469 ms against p50 2.49;DETAIL_FANOUT = 25. Perf.create_default_context63.14 % self (18.86 s of 29.87 s);catalog.pyopens anhttpx.Clientper call. Perf./statsp50 258 / p95 476 / p99 495 ms over 401 calls; api profilequery79.36 % self; the full scan andjson.dumpsinmain.py. Perf.chatspans of traced0330fad…carrygen_ai.input.messages,gen_ai.output.messagesandgen_ai.system_instructions, theinvoke_agentspanpydantic_ai.all_messagesandfinal_result;assistant.py:56buildsInstrumentationSettingswithout turning content capture off. Telemetry.time.sleepinthrottle.pyis real code but the finding cites no telemetry showing a stall -PAUSE_S0.25 s never fires against 15 s arrivals (3 ms betweenagent.askandinvoke_agenton every trace). Code review, not an observation.order rejected … reason=out-of-stockWARN lines;main.py:157returns the error body with the default 200. Behavior.catalog_category="unknown") held up: single series, delta 784;main.py:181. Telemetry.{ name = "GET /products/{sku}" } >> { span.db.system.name = "sqlite" }= 0 against 963 forGET /products;main.py:76callsdb.connect()directly. Telemetry.chatspans carrygen_ai.system=openaibesidegen_ai.provider.name=openai(retired name still emitted). Telemetry.By kind, confirmed: Telemetry 4 (F3, F5b, F6, G4) / Perf 3 (F1a, F1b, F2) / Behavior 1 (F5a).
Figures: launch 22:17:00Z, end 22:27:17Z - preflight 2m24s, drive 2m02s, observation 5m51s, total 10m17s; 90 turns, median 4.3 s; Input 6,542,801 / Output 32,417 / Cache 5,846,627; cost 1.082191 USD (reconciled to the cent, flat rates); signals 4/4 (4 metrics, 8 traces, 4 logs, 2 profiles); 0 source files read before the drive (first read 22:24:40); no traffic of its own; replayable protocol: yes.
Chosen over run 1 (7/8, 9m28s, 1.00 USD) on confirmed findings, step 6's first criterion.
google/gemini-3.7-flash(opencode) - run 1 - 7 confirmed / 8 reportedWindow 22:07:12Z-22:09:13Z (k6's own 121.1 s agrees). 4 anomalies of which the first bundles two defects with two fixes, 3 gaps: 8 items.
e57b487f…carries 25GET /products/{sku}server spans under onetools/call search_products; span p95 455 ms against p50 1.86;DETAIL_FANOUT = 25inserver.py. Perf.create_default_context62.03 % self (18.69 s of 30.13 s);catalog.pyopens anhttpx.Clientper GET/POST. Perf./statsp50 232 / p95 477 / p99 498 ms over 410 calls; tracecd2c433f…is 527 ms of whichcatalog stats_scan514.6; api profilequery78.78 % self;main.pyscans every column andjson.dumpsthe catalog. Perf.throttle.pydoes calltime.sleepfrom the coroutine, butPAUSE_Sis 0.25 s against arrivals 15 s apart, so it never fires - the cited trace25361b8f…has 3 ms betweenagent.askandinvoke_agentand its 3.6 s is twochatcalls of 1.2 and 1.1 s; the latency the finding attributes to the sleep is the provider's. Code review dressed in telemetry.order rejected … reason=out-of-stockWARN lines; tracefa9cf062…(tools/call place_order) answershttp.status_code=200, status UNSET, errors=0;main.py:157returns{"error": "out of stock", "order_ref": None}with the default 200. Behavior.catalog_orders_created_totalhas the single seriescatalog_category="unknown"(delta 802). Telemetry.{ name = "POST /orders" } >> { span.db.system.name = "sqlite" }= 0 against 812 forGET /orders/{order_ref}. Telemetry.chatspans carrygen_ai.provider.name=openaibesideserver.address=openrouter.ai. Telemetry.By kind, confirmed: Telemetry 3 (G1-G3) / Perf 3 (F1a, F1b, F2) / Behavior 1 (F4).
Figures: launch 22:04:50Z, end 22:14:18Z - preflight 2m22s, drive 2m01s, observation 5m05s, total 9m28s; 87 turns, median 3.3 s; Input 6,465,034 / Output 30,705 / Cache 5,878,796; cost 0.995732 USD (reconciled to the cent, flat rates); signals 4/4 (12 metrics, 5 traces, 2 logs, 1 profiles invocations); 0 source files read before the drive (first read 22:12:02); no traffic of its own; replayable protocol: yes.
qwen/qwen3.8-27b(opencode) - attempt 2 - 11 confirmed / 11 reported - THE ROWWindow 05:57:32Z-05:59:32Z (k6's own 120.7 s agrees). One silent stream of ten minutes at 06:11-06:22 (no log line, no part), which completed on its own this time. 8 anomalies of which F2 bundles two defects, 3 gaps of which the last calls itself "filled (by design, documented)": 11 items.
order createdlines carry 655 distinct refs, 113 shared. Behavior.tools/call search_productsp50 1.74 / p95 239 ms over 420 calls. Perf.create_default_context58.86 % self (12.12 s of 20.59 s). Perf.query85.37 % self of 91.3 s; the scan andjson.dumps. Perf./statsanswers 1,129 bytes (histogram mean 1,129, a live probe 1,129) while its log line sayspayload_bytes=1838708-main.py:121-127logs the size of the whole catalog it rendered, not of the response. Telemetry.order rejected … out-of-stocklines answered 200. Behavior.http sendspans - twohttp sendand onehttp receivesub-millisecond INTERNAL spans per request. Telemetry.server.pylogs only insearch_productsandplace_order;get_productandget_orderemit no line, against the manifest's claim. Telemetry.gen_ai_client_operation_duration. Telemetry.By kind, confirmed: Telemetry 5 (F3, F5, F8, G1, G2) / Perf 3 (F2a, F2b, F4) / Behavior 3 (F1, F6, F7).
Figures: launch 05:53:43Z, end 06:34:08Z - preflight 3m49s, drive 2m00s, observation 34m36s, total 40m25s; 48 turns, median 16.6 s, max 2352 s; Input 5,878,526 / Output 133,232 / Cache 3,734,464; cost 1.617631 USD (reconciled to the cent at 0.42 / 3.00 / 0.085 USD per million); signals 4/4 (through its helper scripts - the log alone counts 1/0/0/0); 0 source files read before the drive (first read 06:07:05); no traffic of its own; replayable protocol: yes.
qwen/qwen3.8-27b(opencode) - attempt 1 - killed, no reportLaunched 03:28:47Z; drove at 03:33-03:35; observed for 25 min; at 03:59:50 a stream opened with one socket to the provider and produced nothing for 14 minutes - no new log line, no new
partrow (313 rows from 03:59:56 on) - past the ten-minute bounded wait. Killed by its PID at 04:14:00Z, 45 min in, no report skeleton opened; the store shows the silent turn completing at 04:13:55, seconds before the kill (316 rows), with the report phase still ahead. Spend to the kill: 1.473 USD, 6.71M input / 102k output tokens, 59 turns. Per the queue rule the model moved to the end of the queue for its second attempt.z-ai/glm-5.3-flash(opencode) - attempt 2 - 7 confirmed / 8 reported - THE ROWWindow 05:26:05Z-05:28:08Z (k6's own 122.5 s agrees). One
504 Upstream idle timeout exceededon the report-writing turn at 05:49:16 - the retry passed (the first attempt's four did not). 7 anomalies of which F2 bundles the fan-out with a claim about detached spans and F6 (a public-range peer address at the Docker port-forward) calls itself an environment artefact with no action, 6 gap bullets of which four restate F4, F2 and F5 and one says "no gap": 8 items./statsspan p50 218 / p95 474 / p99 504 ms over 426 calls; api profilequery83.58 % self of 87.59 s. Perf.search_productscontain the api'sGET /products/{sku}children; the 0.576 ms childless root it cites is a cache hit, and the 426 rootlessGET /products/{sku}are k6's own direct reads.create_default_context57.72 % self (10.88 s of 18.85 s). Perf.http_server_*, its roots no HTTP status. Telemetry.search_productscache hits (160 counted of 432). Telemetry.server/discoverinside all 9POST /asktraces. Perf.POST /orders. Telemetry.By kind, confirmed: Telemetry 3 (F4, F5, G4) / Perf 4 (F1, F2a, F3, F7) / Behavior 0.
Figures: launch 05:19:23Z, end 05:51:38Z - preflight 6m42s, drive 2m03s, observation 23m30s, total 32m15s; 34 turns, median 22.3 s, max 1770 s; Input 2,315,434 / Output 74,103 / Cache 1,837,312; cost 0.098333 USD (reconciled to the cent at 0.09 / 0.30 / 0.018 USD per million); signals 4/4 (3 metrics, 7 traces, 3 logs, 2 profiles); 0 source files read (none at all); no traffic of its own; replayable protocol: yes.
z-ai/glm-5.3-flash(opencode) - attempt 1 - killed, no reportLaunched 02:49:41Z; drove 02:56:27Z-02:58:28Z (k6's own 121 s), observed, opened the report skeleton at 03:14, then the report-writing generation hit
504 Upstream idle timeout exceededfour times in a row (03:18:48, 03:21:08, 03:23:42, 03:26:31 - the same failure this model showed once under 1.12.0, when the retry passed). Killed by its PID at 03:27:20Z after the fourth identical failure (one more than the rule's third - the fourth landed while the third was being read), 37 min in; the skeleton carried eight<fill>and no draft existed. Nothing to grade. Per the queue rule the model moved to the end of the queue for its second attempt.Spend to the kill: 0.0885 USD, 2.75M input / 83k output tokens, 34 turns.
anthropic/claude-sonnet-5(claude) - run 1 - 6 confirmed / 6 reportedWindow 01:28:42Z-01:30:44Z (k6's own 121.8 s agrees). 5 anomalies, 2 gaps of which the second restates F2: 6 items.
/statsp50 282 / p95 478 / p99 496 ms over 405 calls; api profilequery79.49 % self of 84.38 s;main.py:94-97and thejson.dumps. Perf.server.py:48-54returns before the increment at :74-75. Telemetry.tools/call search_productsp50 2.64 / p95 454 / p99 500 ms;DETAIL_FANOUT=25sequential detail fetches,server.py:56-67. Perf.create_default_context61.24 % self (19.21 s of 31.37 s);catalog.py:24-25. Perf.order rejected … out-of-stocklines answered 200. Behavior.http_target/http_method, the old semconv names, and nohttp_route/http_request_method(the report noteshttp_targetis already templated). Telemetry.By kind, confirmed: Telemetry 2 (F2, G1) / Perf 3 (F1, F3, F4) / Behavior 1 (F5).
Figures: launch 01:26:06Z, end 01:40:01Z - preflight 2m36s, drive 2m02s, observation 9m17s, total 13m55s; 84 requests (root + one observe-run subagent), median 2.0 s, max 26 s; Input 10,608,033 (uncached 168, cache read 10,332,615, cache creation 275,250 of which 214,440 at 5m and 60,810 at 1h) / Output 51,959 (thinking 22,640) / Cache 10,607,865;
total_cost_usd3.365789 (costBasislist, a singleclaude-sonnet-5key - no haiku background call on this run), reconstructed to the cent at 2.00 / 10.00 / 0.20 / 2.50 / 4.00 USD per million; signals 4/4 (7 metrics, 7 traces, 2 logs, 1 profiles); 0 source files read before the drive (first read 01:33:55); no traffic of its own; replayable protocol: yes. Claude Code 2.1.278, --effort medium, --permission-mode bypassPermissions. Scratch under/tmp/oddobserve-llmbench-store-load-local/.anthropic/claude-sonnet-5(claude) - run 2 - VOID (the observation ran on opus-5)The root session, on
claude-sonnet-5, dispatchedobserve-runthrough the Agent tool withmodel: 'opus'of its own accord (01:45:26Z; no contract asks for it, run 1's root passed no model). The subagent's 89 requests ran onclaude-opus-5:modelUsagecarriesclaude-opus-5[1m]at 5.466671 USD (54,657 output tokens) besideclaude-sonnet-5at 0.562393 USD (4,566 output tokens). The report - 14 anomalies, 10 gaps, driven (window 01:46:12Z-01:48:14Z) - is opus-5's work and is not graded for this row. Figures for the record: launch 01:42:11Z, end 02:00:48Z, total 18m37s;total_cost_usd6.029064. Run 1 (6/6, all requests on sonnet-5) is the row.anthropic/claude-fable-5.1(claude) - not run: no fable credit left, on the maintainer's word; the 1.12.0 row stays, marked ⚠︎ provisional.🤖 Generated with Claude Code