From 17af1c4714e181678a8da6474c59b2279ee66aed Mon Sep 17 00:00:00 2001 From: Dmitry Rybakov <160598371+comandeo-mongo@users.noreply.github.com> Date: Mon, 21 Sep 2026 11:55:06 +0200 Subject: [PATCH 01/11] DRIVERS-3620 Add Performance Implications section to the OpenTelemetry spec --- source/open-telemetry/open-telemetry.md | 106 +++++++++++++++++++++++- 1 file changed, 104 insertions(+), 2 deletions(-) diff --git a/source/open-telemetry/open-telemetry.md b/source/open-telemetry/open-telemetry.md index 4059a1c35b..7cd48497b6 100644 --- a/source/open-telemetry/open-telemetry.md +++ b/source/open-telemetry/open-telemetry.md @@ -488,8 +488,9 @@ The OpenTelemetry specification covers all driver operations including but not l ## Backwards Compatibility Introduction of OpenTelemetry in new driver versions should not significantly affect existing applications that do not -enable OpenTelemetry. However, since the no-op tracing operation may introduce some performance degradation (though it -should be negligible), customers should be informed of this feature and how to disable it completely. +enable OpenTelemetry. However, since the no-op tracing operation may introduce some performance degradation (see +[Performance Implications section](#performance-implications)), customers should be informed of this feature and how to +disable it completely. If a driver is used in an application that has OpenTelemetry enabled, customers will see traces from the driver in their OpenTelemetry backends. This may be unexpected and MAY cause negative effects in some cases (e.g., the OpenTelemetry @@ -503,6 +504,104 @@ SHOULD follow the [Security](https://github.com/mongodb/specifications/blob/master/source/command-logging-and-monitoring/command-logging-and-monitoring.md#security) guidance of the Command Logging and Monitoring spec. +## Performance Implications + +Drivers MUST measure the performance impact of their OpenTelemetry implementations and SHOULD keep it within the +[Performance Targets](#performance-targets) defined below. Exceeding a target does not by itself make an implementation +non-compliant, but drivers MUST investigate the exceedance and document why it cannot be avoided. + +### Implementation Guidelines + +Not every trace is going to be recorded: samplers decide whether a span is recorded, and the OpenTelemetry API exposes +this decision via `isRecording` — a method that requires the span to be already created. Only the attributes provided at +span creation are visible to the sampler; attributes set later are not. Building every attribute before creating the +span therefore pays the full cost even for spans that will never be recorded. + +Therefore, drivers SHOULD: + +- Create a span with a minimum set of attributes that are cheap to create. The recommended set for command spans is + `db.system.name`, `db.namespace`, `db.collection.name`, `db.command.name`, `server.address`, `server.port`, + `network.transport`, `db.mongodb.server_connection_id`, and `db.mongodb.driver_connection_id`. The recommended set + for operation spans is `db.system.name`, `db.namespace`, `db.collection.name`, `db.operation.name`, and + `db.operation.summary`. +- After creating the span, check whether the span is being recorded. If not, no more attributes are added to the span. + If yes, add the rest of the attributes to the span: `db.query.summary`, `db.query.text`, `db.mongodb.lsid`, + `db.mongodb.cursor_id`, and `db.mongodb.txn_number` for command spans, `db.mongodb.cursor_id` for operation spans. + +Attributes added after span creation are invisible to the sampler. This is an acceptable trade-off: the built-in +samplers do not read attributes at all, and `db.query.text` is disabled by default. A host application whose custom +sampler keys on a deferred attribute will not see it. + +Drivers SHOULD compute values that do not change during the life of an object once, and reuse them: connection +attributes for the life of a connection, operation names per operation class, the formatted session id per session. Such +caches MUST be safe for concurrent use. + +Drivers MAY skip all span work — making the span current, adding attributes, processing the result, ending the span — +when the created span's context is not valid, i.e. its trace id or span id is all zeroes (see +[IsValid](https://opentelemetry.io/docs/specs/otel/trace/api/#isvalid)). An invalid context has no trace identity: it +cannot be propagated, continued, or correlated with anything downstream, so any work on such a span is wasted. This is a +state check on the span at hand, not the prohibited detection of whether the OpenTelemetry SDK is available: a custom +API-only tracer provider that returns spans with valid contexts gets the full tracing path. Drivers MUST NOT use +`isRecording` for this short-circuit: a valid but unsampled span MUST still be made current, so that its context is +propagated (see [Propagating Trace Context to the Server](#propagating-trace-context-to-the-server)). + +### Benchmarking + +Drivers MUST benchmark the following driver versions in the following configurations, and record the overhead of every +configuration relative to the baseline: + +| Configuration | Driver version | OpenTelemetry setup | What it measures | +| :------------ | :----------------------------------------- | :----------------------------------------------------------------------------------------- | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `pre-otel` | last version without OpenTelemetry support | none | The reference for `off`: proves that merely shipping the disabled instrumentation costs nothing. Compared once, when OpenTelemetry support is first released. | +| `off` | version with OpenTelemetry | tracing disabled | The baseline. All other configurations MUST be compared to it. | +| `api-only` | version with OpenTelemetry | tracing enabled; OpenTelemetry API available, no SDK | The overhead of the driver's own implementation: without an SDK every OpenTelemetry API method is a no-op, so everything measured here is the driver's code. | +| `sdk-never` | version with OpenTelemetry | tracing enabled; SDK installed; sampler drops every trace | The overhead of the OpenTelemetry SDK regardless of the sampling decision. Also a direct check of the guidelines above: if this configuration costs nearly as much as `sdk-always`, attribute building is not gated on the sampling decision. | +| `sdk-parent` | version with OpenTelemetry | tracing enabled; SDK installed; parent-based sampler with a representative ratio (e.g. 1%) | A realistic production setting. | +| `sdk-always` | version with OpenTelemetry | tracing enabled; SDK installed; every trace sampled | The full overhead of the OpenTelemetry implementation; the ceiling. | + +The ladder is chosen so that the differences between consecutive configurations are meaningful: `off` → `api-only` is +the driver-side cost, `api-only` → `sdk-never` is the SDK bookkeeping, and `sdk-never` → `sdk-always` is the cost of +actually recording. + +Drivers SHOULD use their standardized performance testing infrastructure (see +[Performance Benchmarking](../benchmarking/benchmarking.md)) rather than a purpose-built OpenTelemetry benchmark: the +configurations above are just more configurations of the same tasks. The following rules apply to the setup: + +- Each configuration MUST run in its own process. OpenTelemetry cannot be reconfigured once its SDK has been installed + into a process, and `api-only` requires that the SDK was never loaded. +- Configurations MUST be run interleaved — the whole set, then the whole set again — rather than one configuration to + completion and then the next. The quantity being measured is a difference between two configurations, and a machine + that slows down halfway through a run would otherwise report the difference as overhead. +- The overhead relative to the baseline SHOULD be recorded as a metric of its own, so that it can be watched for + regressions directly, and so that host-to-host variation cancels out of it. +- Benchmark tasks that never talk to a server (e.g. BSON micro-benchmarks) MUST be excluded: they create no spans and + only add noise to the comparison. +- SDK configurations SHOULD NOT install an exporter or a span processor. The cost to attribute to the driver is creating + and recording spans; a processor charges the SDK's export machinery to the driver's account and adds variance. + Sampled spans are still fully recorded without one. +- A task is only a useful signal if it creates spans, and how many spans a task creates is driver-specific. Drivers + SHOULD report the number of spans per operation for each task, or at least name the tasks that create no spans, so + that a measured zero overhead is never read as evidence of an efficient implementation. + +Additionally, drivers MAY guard the tracing hot path with an allocation-based unit test, asserting the number of objects +allocated per traced command: allocation counts are near-deterministic and catch this class of regression in seconds +rather than in a multi-hour benchmark. + +### Performance Targets + +The targets below apply to the worst small-document task of the benchmarking suite (e.g. `Small doc insertOne`, +`findOne by ID`). Such tasks against a local server are the pessimistic upper bound: there is no network latency, so the +driver's own cost is the largest share of each operation; the percentages shrink with realistic latency and concurrency. +The numbers are anchored to measured results in Ruby: implementations that built every attribute eagerly measured 17–32% +overhead in the non-recording configurations; after applying the +[Implementation Guidelines](#implementation-guidelines), 4–8%. + +- `off` SHOULD show no measurable overhead against `pre-otel`: the disabled instrumentation is a few branch checks. +- `api-only` SHOULD stay within 5% of the baseline. Anything above that is avoidable driver-side cost. +- `sdk-never` and `sdk-parent` SHOULD stay within 10% of the baseline. +- `sdk-always` SHOULD stay within 15% of the baseline. This is the budgeted cost of the feature itself; it cannot be + zero. + ## Future Work ### Query Parametrization @@ -552,6 +651,9 @@ redesigning the payload format. ## Changelog +- 2026-09-21: Added the Performance Implications section: sampling-aware implementation guidelines, the required + benchmark configurations, and per-configuration performance targets (DRIVERS-3620). + - 2026-08-19: Specified the `error.type` attribute on command spans, which drivers MUST add when a command fails and which matches `db.response.status_code` when the command failed with a server error and is otherwise the name of the exception class associated with that command's failure. Specified that drivers MUST NOT set it when the command From 01cc75a042fb3e9f9f28047c1b1b0bbcc7c22f89 Mon Sep 17 00:00:00 2001 From: Dmitry Rybakov <160598371+comandeo-mongo@users.noreply.github.com> Date: Thu, 1 Oct 2026 13:24:11 +0200 Subject: [PATCH 02/11] DRIVERS-3620 Replace OTel performance targets with achieved results Other specs set no numeric performance limits; measured results live in rationale sections. The 5/10/15% targets are removed. The figures move to a Design Rationale entry as what an implementation following the guidelines achieved in Ruby, including with YJIT. The benchmarking rules gain what the measurements showed: which tasks to use, rotating the configuration order, comparing tasks on one host, running with the production runtime configuration (e.g. a JIT), recording CPU time per operation with and without GC, and not judging a single run against a fixed threshold. --- source/open-telemetry/open-telemetry.md | 86 ++++++++++++++++++------- 1 file changed, 62 insertions(+), 24 deletions(-) diff --git a/source/open-telemetry/open-telemetry.md b/source/open-telemetry/open-telemetry.md index 7cd48497b6..a5e24d4af2 100644 --- a/source/open-telemetry/open-telemetry.md +++ b/source/open-telemetry/open-telemetry.md @@ -506,9 +506,11 @@ guidance of the Command Logging and Monitoring spec. ## Performance Implications -Drivers MUST measure the performance impact of their OpenTelemetry implementations and SHOULD keep it within the -[Performance Targets](#performance-targets) defined below. Exceeding a target does not by itself make an implementation -non-compliant, but drivers MUST investigate the exceedance and document why it cannot be avoided. +Drivers MUST measure the performance impact of their OpenTelemetry implementations as described in +[Benchmarking](#benchmarking), and SHOULD follow the [Implementation Guidelines](#implementation-guidelines). This +specification sets no numeric limit: the cost depends on the language runtime, the host and the workload. For reference, +[What overhead is achievable?](#what-overhead-is-achievable) describes the results of an implementation that follows the +guidelines. ### Implementation Guidelines @@ -567,15 +569,32 @@ Drivers SHOULD use their standardized performance testing infrastructure (see [Performance Benchmarking](../benchmarking/benchmarking.md)) rather than a purpose-built OpenTelemetry benchmark: the configurations above are just more configurations of the same tasks. The following rules apply to the setup: +- Drivers SHOULD measure the `Small doc insertOne` and `Find one by ID` tasks. They perform one small operation per + command against the server, so the driver's own cost, and with it the cost of tracing, is the largest share of each + operation. Other tasks are a poor signal: `Run command` sends `hello`, which drivers do not trace; + `Find many and empty the cursor` creates few spans per document; large-document and bulk tasks are dominated by the + server. +- Benchmark tasks that never talk to a server (e.g. BSON micro-benchmarks) MUST be excluded: they create no spans and + only add noise to the comparison. - Each configuration MUST run in its own process. OpenTelemetry cannot be reconfigured once its SDK has been installed into a process, and `api-only` requires that the SDK was never loaded. - Configurations MUST be run interleaved — the whole set, then the whole set again — rather than one configuration to completion and then the next. The quantity being measured is a difference between two configurations, and a machine that slows down halfway through a run would otherwise report the difference as overhead. -- The overhead relative to the baseline SHOULD be recorded as a metric of its own, so that it can be watched for - regressions directly, and so that host-to-host variation cancels out of it. -- Benchmark tasks that never talk to a server (e.g. BSON micro-benchmarks) MUST be excluded: they create no spans and - only add noise to the comparison. +- The order of configurations SHOULD be rotated between repetitions, so that each configuration runs in each position + equally often. With a fixed order, an effect tied to the position within a repetition shows up as a difference + between configurations. +- Tasks whose overheads are compared with each other MUST run on the same host. Hosts differ in speed by more than the + overhead of tracing does. +- Drivers SHOULD run the benchmarks with the runtime configuration typical for production, e.g. with a JIT compiler + enabled where the language runtime offers one. The cost of tracing is mostly the cost of calling into the + OpenTelemetry API, which an interpreter can make several times higher than a JIT does. +- Drivers SHOULD record, in addition to the throughput score, the CPU time per operation. CPU time leaves out the time + spent waiting on the server, so it shows the cost of tracing with less noise than throughput. Where the runtime + reports time spent in garbage collection, drivers SHOULD also record CPU time per operation excluding it: when and + for how long the collector runs can vary between processes by more than the cost of tracing. +- The overhead of each configuration relative to the baseline SHOULD be recorded as a metric of its own, so that it can + be watched for regressions directly, and so that host-to-host variation cancels out of it. - SDK configurations SHOULD NOT install an exporter or a span processor. The cost to attribute to the driver is creating and recording spans; a processor charges the SDK's export machinery to the driver's account and adds variance. Sampled spans are still fully recorded without one. @@ -583,25 +602,14 @@ configurations above are just more configurations of the same tasks. The followi SHOULD report the number of spans per operation for each task, or at least name the tasks that create no spans, so that a measured zero overhead is never read as evidence of an efficient implementation. +The overheads measured are small differences between two noisy numbers. A single run on shared CI hosts can be off by a +large share of the overhead itself, so drivers SHOULD NOT compare a single run against a fixed threshold, and SHOULD +look at the trend over several runs instead. + Additionally, drivers MAY guard the tracing hot path with an allocation-based unit test, asserting the number of objects allocated per traced command: allocation counts are near-deterministic and catch this class of regression in seconds rather than in a multi-hour benchmark. -### Performance Targets - -The targets below apply to the worst small-document task of the benchmarking suite (e.g. `Small doc insertOne`, -`findOne by ID`). Such tasks against a local server are the pessimistic upper bound: there is no network latency, so the -driver's own cost is the largest share of each operation; the percentages shrink with realistic latency and concurrency. -The numbers are anchored to measured results in Ruby: implementations that built every attribute eagerly measured 17–32% -overhead in the non-recording configurations; after applying the -[Implementation Guidelines](#implementation-guidelines), 4–8%. - -- `off` SHOULD show no measurable overhead against `pre-otel`: the disabled instrumentation is a few branch checks. -- `api-only` SHOULD stay within 5% of the baseline. Anything above that is avoidable driver-side cost. -- `sdk-never` and `sdk-parent` SHOULD stay within 10% of the baseline. -- `sdk-always` SHOULD stay within 15% of the baseline. This is the budgeted cost of the feature itself; it cannot be - zero. - ## Future Work ### Query Parametrization @@ -649,10 +657,40 @@ negotiation mechanism. Carrying the traceparent inside a BSON document allows future propagation fields to be added to the same section without redesigning the payload format. +### What overhead is achievable? + +The figures below were measured with the Ruby driver, on the `Small doc insertOne` and `Find one by ID` tasks against a +standalone server on the same host. They are a reference for what an implementation can achieve, not requirements: the +cost of tracing depends on the language runtime, the host and the workload, and it shrinks as a share of each operation +once there is network latency or concurrency. + +An implementation that built every attribute eagerly, before the [Implementation Guidelines](#implementation-guidelines) +were applied, gave up 9–17% of throughput in `api-only`, 13–22% in `sdk-never` and 20–26% in `sdk-always`. Applying the +guidelines cut this to 5–7%, 9–11% and 16% respectively, with Ruby's interpreter. + +With the YJIT compiler enabled, as Ruby applications commonly run in production, the CPU time added per operation was: + +| Configuration | Added CPU time per operation | Share of the operation's CPU time | +| :------------ | :--------------------------- | :-------------------------------- | +| `api-only` | 2–3 µs | 3–4% | +| `sdk-never` | 4 µs | 5–6% | +| `sdk-parent` | 5 µs | 6–7% | +| `sdk-always` | 8 µs | 9–11% | + +Each of these operations creates two spans, an operation span and a command span, so the cost per span ranged from about +1–1.5 µs without an SDK to about 4 µs when every span was recorded. On shared CI hosts, which were 2–3 times slower, the +absolute figures were correspondingly higher and the shares similar, but a single run varied by several microseconds per +operation. + +As a rule of thumb, an implementation that follows the guidelines and runs under the runtime's production configuration +can keep the overhead on these tasks to about 5% without an SDK, about 10% with an SDK that samples few traces, and +about 15% when every trace is recorded. + ## Changelog -- 2026-09-21: Added the Performance Implications section: sampling-aware implementation guidelines, the required - benchmark configurations, and per-configuration performance targets (DRIVERS-3620). +- 2026-10-01: Added the Performance Implications section: sampling-aware implementation guidelines and the required + benchmark configurations and setup, and a Design Rationale entry with the overhead achieved by an implementation + that follows the guidelines (DRIVERS-3620). - 2026-08-19: Specified the `error.type` attribute on command spans, which drivers MUST add when a command fails and which matches `db.response.status_code` when the command failed with a server error and is otherwise the name of the From 64d563a84f995f5b14d7514d8c4c982398244e38 Mon Sep 17 00:00:00 2001 From: Dmitry Rybakov <160598371+comandeo-mongo@users.noreply.github.com> Date: Thu, 1 Oct 2026 13:58:44 +0200 Subject: [PATCH 03/11] DRIVERS-3620 Address review of the OTel performance section - Drop the claim that drivers do not trace hello: only sensitive commands are excluded. Run command is a poor signal because its spans have no namespace or collection. - Make the pre-OTel comparison an explicit one-time requirement for new implementations, and move the per-configuration explanations out of the table. - State the goal (robust to host drift) as the requirement and interleaving, rotation, same-host runs and trend-over-single-run as SHOULDs, rather than dictating the CI layout. - Require no active parent span, and state that the harness, not the driver, configures the SDK and sampler. - Move the MUST about making unsampled spans current into the propagation section, next to the rule on propagating unsampled contexts. - Benchmark in the runtime's default production configuration; GC-excluded CPU time is a MAY. - Rationale figures name no driver or runtime, name the metric in each sentence, and the rule of thumb is dropped. --- source/open-telemetry/open-telemetry.md | 177 +++++++++++++----------- 1 file changed, 96 insertions(+), 81 deletions(-) diff --git a/source/open-telemetry/open-telemetry.md b/source/open-telemetry/open-telemetry.md index a5e24d4af2..7122f0f385 100644 --- a/source/open-telemetry/open-telemetry.md +++ b/source/open-telemetry/open-telemetry.md @@ -429,7 +429,9 @@ those of the [W3C traceparent header](https://www.w3.org/TR/trace-context/#trace `parent-id` carries the command span's own span id — the parent of the spans the server creates) and neither the trace-id nor the parent-id is all zeroes. This mirrors the server-side validation. If no valid value is available, drivers MUST omit the section entirely rather than send an invalid or truncated value. Drivers MUST propagate unsampled -trace contexts (trace-flags `00`); the sampling decision MUST NOT affect whether the section is attached. +trace contexts (trace-flags `00`); the sampling decision MUST NOT affect whether the section is attached. In particular, +a span that is not recording but has a valid context MUST still be made current, so that its context is propagated; +drivers MUST NOT use `isRecording` to skip that work (see [Implementation Guidelines](#implementation-guidelines)). A message MUST NOT contain more than one telemetry section. Commands that carry no command span (for example server monitoring, authentication, and security-sensitive commands) naturally send no section. @@ -514,10 +516,10 @@ guidelines. ### Implementation Guidelines -Not every trace is going to be recorded: samplers decide whether a span is recorded, and the OpenTelemetry API exposes -this decision via `isRecording` — a method that requires the span to be already created. Only the attributes provided at -span creation are visible to the sampler; attributes set later are not. Building every attribute before creating the -span therefore pays the full cost even for spans that will never be recorded. +Not every span is recorded: samplers decide whether a span is recorded, and the OpenTelemetry API exposes this decision +via `isRecording` — a method that requires the span to be already created. Only the attributes provided at span creation +are visible to the sampler; attributes set later are not. Building every attribute before creating the span therefore +pays the full cost even for spans that will never be recorded. Therefore, drivers SHOULD: @@ -543,72 +545,86 @@ when the created span's context is not valid, i.e. its trace id or span id is al [IsValid](https://opentelemetry.io/docs/specs/otel/trace/api/#isvalid)). An invalid context has no trace identity: it cannot be propagated, continued, or correlated with anything downstream, so any work on such a span is wasted. This is a state check on the span at hand, not the prohibited detection of whether the OpenTelemetry SDK is available: a custom -API-only tracer provider that returns spans with valid contexts gets the full tracing path. Drivers MUST NOT use -`isRecording` for this short-circuit: a valid but unsampled span MUST still be made current, so that its context is -propagated (see [Propagating Trace Context to the Server](#propagating-trace-context-to-the-server)). +API-only tracer provider that returns spans with valid contexts gets the full tracing path. The short-circuit keys on +context validity, not on `isRecording`: a valid but unsampled span still has to be made current (see +[Propagating Trace Context to the Server](#propagating-trace-context-to-the-server)). ### Benchmarking -Drivers MUST benchmark the following driver versions in the following configurations, and record the overhead of every -configuration relative to the baseline: - -| Configuration | Driver version | OpenTelemetry setup | What it measures | -| :------------ | :----------------------------------------- | :----------------------------------------------------------------------------------------- | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `pre-otel` | last version without OpenTelemetry support | none | The reference for `off`: proves that merely shipping the disabled instrumentation costs nothing. Compared once, when OpenTelemetry support is first released. | -| `off` | version with OpenTelemetry | tracing disabled | The baseline. All other configurations MUST be compared to it. | -| `api-only` | version with OpenTelemetry | tracing enabled; OpenTelemetry API available, no SDK | The overhead of the driver's own implementation: without an SDK every OpenTelemetry API method is a no-op, so everything measured here is the driver's code. | -| `sdk-never` | version with OpenTelemetry | tracing enabled; SDK installed; sampler drops every trace | The overhead of the OpenTelemetry SDK regardless of the sampling decision. Also a direct check of the guidelines above: if this configuration costs nearly as much as `sdk-always`, attribute building is not gated on the sampling decision. | -| `sdk-parent` | version with OpenTelemetry | tracing enabled; SDK installed; parent-based sampler with a representative ratio (e.g. 1%) | A realistic production setting. | -| `sdk-always` | version with OpenTelemetry | tracing enabled; SDK installed; every trace sampled | The full overhead of the OpenTelemetry implementation; the ceiling. | - -The ladder is chosen so that the differences between consecutive configurations are meaningful: `off` → `api-only` is -the driver-side cost, `api-only` → `sdk-never` is the SDK bookkeeping, and `sdk-never` → `sdk-always` is the cost of -actually recording. +Drivers MUST benchmark the following configurations, and record the overhead of every configuration relative to the +baseline. The benchmark harness, not the driver, installs and configures the OpenTelemetry SDK and its sampler for the +SDK configurations; this does not conflict with the rule that drivers MUST NOT configure OpenTelemetry on the host +application level. No configuration has an active parent span: with a valid parent, a tracer without an SDK returns +spans with the parent's valid context, and `api-only` would no longer measure the same thing across drivers. + +| Configuration | OpenTelemetry setup | +| :------------ | :----------------------------------------------------------------------------------------- | +| `off` | tracing disabled; the baseline | +| `api-only` | tracing enabled; OpenTelemetry API available, no SDK | +| `sdk-never` | tracing enabled; SDK installed; sampler drops every trace | +| `sdk-parent` | tracing enabled; SDK installed; parent-based sampler with a representative ratio (e.g. 1%) | +| `sdk-always` | tracing enabled; SDK installed; every trace sampled | + +`off` is the baseline that all other configurations are compared to. `api-only` measures the driver's own +implementation: without an SDK every OpenTelemetry API method is a no-op. `sdk-never` adds the SDK's cost regardless of +the sampling decision, and is a direct check of the guidelines above: if it costs nearly as much as `sdk-always`, +attribute building is not gated on the sampling decision. `sdk-parent` is a realistic production setting, and +`sdk-always` the full cost of the feature. The differences between consecutive configurations are meaningful: `off` → +`api-only` is the driver-side cost, `api-only` → `sdk-never` is the SDK bookkeeping, and `sdk-never` → `sdk-always` is +the cost of recording. + +Additionally, when a driver first releases OpenTelemetry support, it MUST compare `off` once against the last driver +version without OpenTelemetry support, to show that the disabled instrumentation has no measurable overhead. Drivers +that already shipped OpenTelemetry support before this requirement was added need not do so. Drivers SHOULD use their standardized performance testing infrastructure (see [Performance Benchmarking](../benchmarking/benchmarking.md)) rather than a purpose-built OpenTelemetry benchmark: the -configurations above are just more configurations of the same tasks. The following rules apply to the setup: - -- Drivers SHOULD measure the `Small doc insertOne` and `Find one by ID` tasks. They perform one small operation per - command against the server, so the driver's own cost, and with it the cost of tracing, is the largest share of each - operation. Other tasks are a poor signal: `Run command` sends `hello`, which drivers do not trace; - `Find many and empty the cursor` creates few spans per document; large-document and bulk tasks are dominated by the - server. -- Benchmark tasks that never talk to a server (e.g. BSON micro-benchmarks) MUST be excluded: they create no spans and - only add noise to the comparison. -- Each configuration MUST run in its own process. OpenTelemetry cannot be reconfigured once its SDK has been installed - into a process, and `api-only` requires that the SDK was never loaded. -- Configurations MUST be run interleaved — the whole set, then the whole set again — rather than one configuration to - completion and then the next. The quantity being measured is a difference between two configurations, and a machine - that slows down halfway through a run would otherwise report the difference as overhead. -- The order of configurations SHOULD be rotated between repetitions, so that each configuration runs in each position - equally often. With a fixed order, an effect tied to the position within a repetition shows up as a difference - between configurations. -- Tasks whose overheads are compared with each other MUST run on the same host. Hosts differ in speed by more than the +configurations above are just more configurations of the same tasks. + +Drivers SHOULD measure the `Small doc insertOne` and `Find one by ID` tasks. They perform one small operation per +command against the server, so the driver's own cost, and with it the cost of tracing, is the largest share of each +operation. Other tasks are a poor signal: `Run command` has no namespace or collection, so its spans are not +representative; `Find many and empty the cursor` creates few spans relative to the documents it processes; +large-document and bulk tasks are dominated by the server. Benchmark tasks that never talk to a server (e.g. BSON +micro-benchmarks) MUST be excluded: they create no spans. + +Each configuration MUST run in its own process: OpenTelemetry cannot be reconfigured once its SDK has been installed +into a process, and `api-only` requires that the SDK was never loaded. + +The overheads measured are small differences between two noisy numbers, so the comparison MUST be robust to the host +changing speed during a run and between runs. Drivers SHOULD: + +- Run the configurations interleaved — the whole set, then the whole set again — rather than one configuration to + completion and then the next, so that a host that slows down halfway through a run is not reported as overhead. +- Rotate the order of configurations between repetitions, so that each configuration runs in each position equally + often; with a fixed order, an effect tied to the position shows up as a difference between configurations. +- Run tasks whose overheads are compared with each other on the same host: hosts differ in speed by more than the overhead of tracing does. -- Drivers SHOULD run the benchmarks with the runtime configuration typical for production, e.g. with a JIT compiler - enabled where the language runtime offers one. The cost of tracing is mostly the cost of calling into the - OpenTelemetry API, which an interpreter can make several times higher than a JIT does. -- Drivers SHOULD record, in addition to the throughput score, the CPU time per operation. CPU time leaves out the time - spent waiting on the server, so it shows the cost of tracing with less noise than throughput. Where the runtime - reports time spent in garbage collection, drivers SHOULD also record CPU time per operation excluding it: when and - for how long the collector runs can vary between processes by more than the cost of tracing. -- The overhead of each configuration relative to the baseline SHOULD be recorded as a metric of its own, so that it can - be watched for regressions directly, and so that host-to-host variation cancels out of it. -- SDK configurations SHOULD NOT install an exporter or a span processor. The cost to attribute to the driver is creating - and recording spans; a processor charges the SDK's export machinery to the driver's account and adds variance. - Sampled spans are still fully recorded without one. -- A task is only a useful signal if it creates spans, and how many spans a task creates is driver-specific. Drivers - SHOULD report the number of spans per operation for each task, or at least name the tasks that create no spans, so - that a measured zero overhead is never read as evidence of an efficient implementation. - -The overheads measured are small differences between two noisy numbers. A single run on shared CI hosts can be off by a -large share of the overhead itself, so drivers SHOULD NOT compare a single run against a fixed threshold, and SHOULD -look at the trend over several runs instead. - -Additionally, drivers MAY guard the tracing hot path with an allocation-based unit test, asserting the number of objects -allocated per traced command: allocation counts are near-deterministic and catch this class of regression in seconds -rather than in a multi-hour benchmark. +- Not compare a single run against a fixed threshold, but look at the trend over several runs: a single run on shared CI + hosts can be off by a large share of the overhead itself. + +Drivers SHOULD run the benchmarks in the runtime's default production configuration. The cost of tracing is mostly the +cost of calling into the OpenTelemetry API, and its size depends heavily on how the runtime executes code. + +Drivers SHOULD record, in addition to the throughput score: + +- The CPU time per operation. CPU time leaves out the time spent waiting on the server, so it shows the cost of tracing + with less noise than throughput. Where the runtime reports time spent in garbage collection, drivers MAY also record + CPU time per operation excluding it: when and for how long the collector runs can vary between processes by more + than the cost of tracing. +- The overhead of each configuration relative to the baseline as a metric of its own, so that it can be watched for + regressions directly and so that host-to-host variation cancels out of it. +- The number of spans per operation for each task, or at least the names of the tasks that create no spans, so that a + measured zero overhead is never read as evidence of an efficient implementation. How many spans a task creates is + driver-specific. + +SDK configurations SHOULD NOT install an exporter or a span processor. The cost to attribute to the driver is creating +and recording spans; a processor charges the SDK's export machinery to the driver's account and adds variance. Sampled +spans are still fully recorded without one. + +Drivers MAY guard the tracing hot path with an allocation-based unit test, asserting the number of objects allocated per +traced command: allocation counts are near-deterministic and catch this class of regression in seconds rather than in a +multi-hour benchmark. ## Future Work @@ -659,16 +675,18 @@ redesigning the payload format. ### What overhead is achievable? -The figures below were measured with the Ruby driver, on the `Small doc insertOne` and `Find one by ID` tasks against a -standalone server on the same host. They are a reference for what an implementation can achieve, not requirements: the -cost of tracing depends on the language runtime, the host and the workload, and it shrinks as a share of each operation -once there is network latency or concurrency. +The figures below were measured with one driver's implementation, on the `Small doc insertOne` and `Find one by ID` +tasks against a standalone server on the same host. They are a reference for what an implementation can achieve, not +requirements and not targets: the cost of tracing depends on the language runtime, the host and the workload, and it +shrinks as a share of each operation once there is network latency or concurrency. Other runtimes can be expected to +land elsewhere. -An implementation that built every attribute eagerly, before the [Implementation Guidelines](#implementation-guidelines) -were applied, gave up 9–17% of throughput in `api-only`, 13–22% in `sdk-never` and 20–26% in `sdk-always`. Applying the -guidelines cut this to 5–7%, 9–11% and 16% respectively, with Ruby's interpreter. +With an interpreting runtime, an implementation that built every attribute eagerly, before the +[Implementation Guidelines](#implementation-guidelines) were applied, lost 9–17% of throughput in `api-only`, 13–22% in +`sdk-never` and 20–26% in `sdk-always`. Applying the guidelines reduced the throughput loss to 5–7%, 9–11% and 16%. -With the YJIT compiler enabled, as Ruby applications commonly run in production, the CPU time added per operation was: +With the same runtime's JIT compiler enabled, which is its production configuration, the CPU time added per operation +was: | Configuration | Added CPU time per operation | Share of the operation's CPU time | | :------------ | :--------------------------- | :-------------------------------- | @@ -677,20 +695,17 @@ With the YJIT compiler enabled, as Ruby applications commonly run in production, | `sdk-parent` | 5 µs | 6–7% | | `sdk-always` | 8 µs | 9–11% | -Each of these operations creates two spans, an operation span and a command span, so the cost per span ranged from about -1–1.5 µs without an SDK to about 4 µs when every span was recorded. On shared CI hosts, which were 2–3 times slower, the -absolute figures were correspondingly higher and the shares similar, but a single run varied by several microseconds per -operation. - -As a rule of thumb, an implementation that follows the guidelines and runs under the runtime's production configuration -can keep the overhead on these tasks to about 5% without an SDK, about 10% with an SDK that samples few traces, and -about 15% when every trace is recorded. +Each of these operations creates two spans, an operation span and a command span, so the CPU time per span ranged from +about 1–1.5 µs without an SDK to about 4 µs when every span was recorded. On shared CI hosts, which were 2–3 times +slower, the added CPU time was correspondingly higher and its share similar, but a single run varied by several +microseconds per operation. ## Changelog - 2026-10-01: Added the Performance Implications section: sampling-aware implementation guidelines and the required benchmark configurations and setup, and a Design Rationale entry with the overhead achieved by an implementation - that follows the guidelines (DRIVERS-3620). + that follows the guidelines. Specified that a valid but unsampled span MUST still be made current so that its + context is propagated (DRIVERS-3620). - 2026-08-19: Specified the `error.type` attribute on command spans, which drivers MUST add when a command fails and which matches `db.response.status_code` when the command failed with a server error and is otherwise the name of the From 3305c68d7c85cb3d3d3cf821e0a5d66e8e03d054 Mon Sep 17 00:00:00 2001 From: Dmitry Rybakov <160598371+comandeo-mongo@users.noreply.github.com> Date: Thu, 1 Oct 2026 14:02:37 +0200 Subject: [PATCH 04/11] DRIVERS-3620 Drop a language-specific example from the OTel spec The Host Application Level option cited Ruby as an example of enabling all instrumentations, which says nothing a reader of other languages can use. --- source/open-telemetry/open-telemetry.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/source/open-telemetry/open-telemetry.md b/source/open-telemetry/open-telemetry.md index 7122f0f385..3f0336ad53 100644 --- a/source/open-telemetry/open-telemetry.md +++ b/source/open-telemetry/open-telemetry.md @@ -72,8 +72,8 @@ Drivers SHOULD support configuring OpenTelemetry on multiple levels. environment variable `OTEL_#{LANG}_INSTRUMENTATION_MONGODB_ENABLED`. Drivers MAY provide other means to globally disable OpenTelemetry that are more suitable for their language ecosystem. This option MUST override settings on the higher level. -- **Host Application Level**: If the host application enables OpenTelemetry for all available instrumentations (e.g., - Ruby), and a driver can detect this, OpenTelemetry SHOULD be enabled in the driver. +- **Host Application Level**: If the host application enables OpenTelemetry for all available instrumentations, and a + driver can detect this, OpenTelemetry SHOULD be enabled in the driver. Drivers MUST NOT try to detect whether the OpenTelemetry SDK library is available, and enable tracing based on this. Drivers MUST NOT add means that configure OpenTelemetry SDK (e.g., setting a specific exporter). Drivers MUST NOT add From 4f31667cf66d942dcdcd2b0b695dccdf1398afdc Mon Sep 17 00:00:00 2001 From: Dmitry Rybakov <160598371+comandeo-mongo@users.noreply.github.com> Date: Thu, 1 Oct 2026 16:02:48 +0200 Subject: [PATCH 05/11] DRIVERS-3620 Shorten the OTel performance section and make it language-neutral - Drop the rule that unsampled spans must be made current: the nesting and propagation rules already require the outcome. - Define "SDK" once, neutrally, and describe isRecording as the tracing API's equivalent. - Rename sdk-parent to sdk-ratio: with no parent span it is a ratio sampler. - Scope CPU time to the threads executing the operations, or process CPU time with background work held constant; benchmark both sync and async APIs where a driver has both. - Cut repeated rationale, the allocation test and the interpreting-runtime figures. --- source/open-telemetry/open-telemetry.md | 149 ++++++++---------------- 1 file changed, 51 insertions(+), 98 deletions(-) diff --git a/source/open-telemetry/open-telemetry.md b/source/open-telemetry/open-telemetry.md index 3f0336ad53..ac071b3ba4 100644 --- a/source/open-telemetry/open-telemetry.md +++ b/source/open-telemetry/open-telemetry.md @@ -429,9 +429,7 @@ those of the [W3C traceparent header](https://www.w3.org/TR/trace-context/#trace `parent-id` carries the command span's own span id — the parent of the spans the server creates) and neither the trace-id nor the parent-id is all zeroes. This mirrors the server-side validation. If no valid value is available, drivers MUST omit the section entirely rather than send an invalid or truncated value. Drivers MUST propagate unsampled -trace contexts (trace-flags `00`); the sampling decision MUST NOT affect whether the section is attached. In particular, -a span that is not recording but has a valid context MUST still be made current, so that its context is propagated; -drivers MUST NOT use `isRecording` to skip that work (see [Implementation Guidelines](#implementation-guidelines)). +trace contexts (trace-flags `00`); the sampling decision MUST NOT affect whether the section is attached. A message MUST NOT contain more than one telemetry section. Commands that carry no command span (for example server monitoring, authentication, and security-sensitive commands) naturally send no section. @@ -512,14 +510,14 @@ Drivers MUST measure the performance impact of their OpenTelemetry implementatio [Benchmarking](#benchmarking), and SHOULD follow the [Implementation Guidelines](#implementation-guidelines). This specification sets no numeric limit: the cost depends on the language runtime, the host and the workload. For reference, [What overhead is achievable?](#what-overhead-is-achievable) describes the results of an implementation that follows the -guidelines. +guidelines. In this section, "SDK" means the tracer implementation a host application installs; where tracing is built +into the runtime, it means the component that subscribes to and records the runtime's spans. ### Implementation Guidelines -Not every span is recorded: samplers decide whether a span is recorded, and the OpenTelemetry API exposes this decision -via `isRecording` — a method that requires the span to be already created. Only the attributes provided at span creation -are visible to the sampler; attributes set later are not. Building every attribute before creating the span therefore -pays the full cost even for spans that will never be recorded. +Not every span is recorded: samplers decide whether a span is recorded, and the tracing API reports this decision only +on a span that has already been created (`isRecording` in the OpenTelemetry API, or the tracing API's equivalent). Only +the attributes provided at span creation are visible to the sampler. Therefore, drivers SHOULD: @@ -532,99 +530,67 @@ Therefore, drivers SHOULD: If yes, add the rest of the attributes to the span: `db.query.summary`, `db.query.text`, `db.mongodb.lsid`, `db.mongodb.cursor_id`, and `db.mongodb.txn_number` for command spans, `db.mongodb.cursor_id` for operation spans. -Attributes added after span creation are invisible to the sampler. This is an acceptable trade-off: the built-in -samplers do not read attributes at all, and `db.query.text` is disabled by default. A host application whose custom -sampler keys on a deferred attribute will not see it. +Deferring attributes is an acceptable trade-off: only custom samplers keying on deferred attributes are affected. Drivers SHOULD compute values that do not change during the life of an object once, and reuse them: connection -attributes for the life of a connection, operation names per operation class, the formatted session id per session. Such +attributes for the life of a connection, operation names per operation type, the formatted session id per session. Such caches MUST be safe for concurrent use. -Drivers MAY skip all span work — making the span current, adding attributes, processing the result, ending the span — +Drivers MAY skip all span work — making the span current, adding attributes, recording the outcome, ending the span — when the created span's context is not valid, i.e. its trace id or span id is all zeroes (see -[IsValid](https://opentelemetry.io/docs/specs/otel/trace/api/#isvalid)). An invalid context has no trace identity: it -cannot be propagated, continued, or correlated with anything downstream, so any work on such a span is wasted. This is a -state check on the span at hand, not the prohibited detection of whether the OpenTelemetry SDK is available: a custom -API-only tracer provider that returns spans with valid contexts gets the full tracing path. The short-circuit keys on -context validity, not on `isRecording`: a valid but unsampled span still has to be made current (see -[Propagating Trace Context to the Server](#propagating-trace-context-to-the-server)). +[IsValid](https://opentelemetry.io/docs/specs/otel/trace/api/#isvalid)): such a span cannot be propagated or correlated, +so any work on it is wasted. This permission does not apply to a valid but unsampled span, which remains subject to the +nesting and [propagation](#propagating-trace-context-to-the-server) rules. ### Benchmarking -Drivers MUST benchmark the following configurations, and record the overhead of every configuration relative to the -baseline. The benchmark harness, not the driver, installs and configures the OpenTelemetry SDK and its sampler for the -SDK configurations; this does not conflict with the rule that drivers MUST NOT configure OpenTelemetry on the host -application level. No configuration has an active parent span: with a valid parent, a tracer without an SDK returns -spans with the parent's valid context, and `api-only` would no longer measure the same thing across drivers. - -| Configuration | OpenTelemetry setup | -| :------------ | :----------------------------------------------------------------------------------------- | -| `off` | tracing disabled; the baseline | -| `api-only` | tracing enabled; OpenTelemetry API available, no SDK | -| `sdk-never` | tracing enabled; SDK installed; sampler drops every trace | -| `sdk-parent` | tracing enabled; SDK installed; parent-based sampler with a representative ratio (e.g. 1%) | -| `sdk-always` | tracing enabled; SDK installed; every trace sampled | - -`off` is the baseline that all other configurations are compared to. `api-only` measures the driver's own -implementation: without an SDK every OpenTelemetry API method is a no-op. `sdk-never` adds the SDK's cost regardless of -the sampling decision, and is a direct check of the guidelines above: if it costs nearly as much as `sdk-always`, -attribute building is not gated on the sampling decision. `sdk-parent` is a realistic production setting, and -`sdk-always` the full cost of the feature. The differences between consecutive configurations are meaningful: `off` → -`api-only` is the driver-side cost, `api-only` → `sdk-never` is the SDK bookkeeping, and `sdk-never` → `sdk-always` is -the cost of recording. +Drivers MUST benchmark the following configurations, and record the overhead of each relative to the baseline as a +metric of its own. The benchmark harness, not the driver, installs and configures the SDK and its sampler for the SDK +configurations. No configuration has an active parent span, so that `api-only` produces no valid span in any driver. + +| Configuration | Tracing setup | +| :------------ | :---------------------------------------------------------------------------------- | +| `off` | tracing disabled; the baseline | +| `api-only` | tracing enabled; no SDK, so every tracing call is a no-op | +| `sdk-never` | tracing enabled; SDK installed; sampler drops every span | +| `sdk-ratio` | tracing enabled; SDK installed; ratio sampler with a representative ratio (e.g. 1%) | +| `sdk-always` | tracing enabled; SDK installed; every span sampled | + +The differences between consecutive configurations are meaningful: `off` → `api-only` is the driver-side cost, +`api-only` → `sdk-never` is the SDK bookkeeping, and `sdk-never` → `sdk-always` is the cost of recording. If `sdk-never` +costs nearly as much as `sdk-always`, attribute building is not gated on the sampling decision. Additionally, when a driver first releases OpenTelemetry support, it MUST compare `off` once against the last driver -version without OpenTelemetry support, to show that the disabled instrumentation has no measurable overhead. Drivers -that already shipped OpenTelemetry support before this requirement was added need not do so. +version without OpenTelemetry support, to show that the disabled instrumentation has no measurable overhead. Drivers SHOULD use their standardized performance testing infrastructure (see -[Performance Benchmarking](../benchmarking/benchmarking.md)) rather than a purpose-built OpenTelemetry benchmark: the -configurations above are just more configurations of the same tasks. +[Performance Benchmarking](../benchmarking/benchmarking.md)) rather than a purpose-built OpenTelemetry benchmark. -Drivers SHOULD measure the `Small doc insertOne` and `Find one by ID` tasks. They perform one small operation per -command against the server, so the driver's own cost, and with it the cost of tracing, is the largest share of each -operation. Other tasks are a poor signal: `Run command` has no namespace or collection, so its spans are not -representative; `Find many and empty the cursor` creates few spans relative to the documents it processes; -large-document and bulk tasks are dominated by the server. Benchmark tasks that never talk to a server (e.g. BSON -micro-benchmarks) MUST be excluded: they create no spans. +Drivers SHOULD measure the `Small doc insertOne` and `Find one by ID` tasks, through both the synchronous and the +asynchronous API where a driver has both. Tasks that never talk to a server (e.g. BSON micro-benchmarks) MUST be +excluded: they create no spans. -Each configuration MUST run in its own process: OpenTelemetry cannot be reconfigured once its SDK has been installed -into a process, and `api-only` requires that the SDK was never loaded. +Each configuration MUST run in its own process, because an SDK cannot be reliably removed once installed, and `off` and +`api-only` require that none was. The overheads measured are small differences between two noisy numbers, so the comparison MUST be robust to the host -changing speed during a run and between runs. Drivers SHOULD: - -- Run the configurations interleaved — the whole set, then the whole set again — rather than one configuration to - completion and then the next, so that a host that slows down halfway through a run is not reported as overhead. -- Rotate the order of configurations between repetitions, so that each configuration runs in each position equally - often; with a fixed order, an effect tied to the position shows up as a difference between configurations. -- Run tasks whose overheads are compared with each other on the same host: hosts differ in speed by more than the - overhead of tracing does. -- Not compare a single run against a fixed threshold, but look at the trend over several runs: a single run on shared CI - hosts can be off by a large share of the overhead itself. +changing speed during a run and between runs. Drivers SHOULD run the whole set of configurations interleaved and rotate +their order between repetitions, run compared configurations on the same host, and judge the trend over several runs +rather than a single run against a threshold. -Drivers SHOULD run the benchmarks in the runtime's default production configuration. The cost of tracing is mostly the -cost of calling into the OpenTelemetry API, and its size depends heavily on how the runtime executes code. +Drivers SHOULD run the benchmarks in the runtime's default production configuration, with warm-up and iterations as in +[Performance Benchmarking](../benchmarking/benchmarking.md). Drivers SHOULD record, in addition to the throughput score: -- The CPU time per operation. CPU time leaves out the time spent waiting on the server, so it shows the cost of tracing - with less noise than throughput. Where the runtime reports time spent in garbage collection, drivers MAY also record - CPU time per operation excluding it: when and for how long the collector runs can vary between processes by more - than the cost of tracing. -- The overhead of each configuration relative to the baseline as a metric of its own, so that it can be watched for - regressions directly and so that host-to-host variation cancels out of it. -- The number of spans per operation for each task, or at least the names of the tasks that create no spans, so that a - measured zero overhead is never read as evidence of an efficient implementation. How many spans a task creates is - driver-specific. +- The CPU time per operation of the threads executing the operations, or process CPU time with background work held + constant. It excludes waiting on the server and so is less noisy than throughput. Drivers MAY also record it + excluding garbage collection time, where the runtime reports it. +- The number of spans per operation for each task, so that a zero overhead is not mistaken for an efficient + implementation. -SDK configurations SHOULD NOT install an exporter or a span processor. The cost to attribute to the driver is creating -and recording spans; a processor charges the SDK's export machinery to the driver's account and adds variance. Sampled -spans are still fully recorded without one. - -Drivers MAY guard the tracing hot path with an allocation-based unit test, asserting the number of objects allocated per -traced command: allocation counts are near-deterministic and catch this class of regression in seconds rather than in a -multi-hour benchmark. +SDK configurations SHOULD NOT install an exporter or a span processor. Spans are still sampled and recorded without +them; processing cost depends on the host application's choice of processor. ## Future Work @@ -677,35 +643,22 @@ redesigning the payload format. The figures below were measured with one driver's implementation, on the `Small doc insertOne` and `Find one by ID` tasks against a standalone server on the same host. They are a reference for what an implementation can achieve, not -requirements and not targets: the cost of tracing depends on the language runtime, the host and the workload, and it -shrinks as a share of each operation once there is network latency or concurrency. Other runtimes can be expected to -land elsewhere. - -With an interpreting runtime, an implementation that built every attribute eagerly, before the -[Implementation Guidelines](#implementation-guidelines) were applied, lost 9–17% of throughput in `api-only`, 13–22% in -`sdk-never` and 20–26% in `sdk-always`. Applying the guidelines reduced the throughput loss to 5–7%, 9–11% and 16%. +requirements and not targets: other runtimes can be expected to land elsewhere. -With the same runtime's JIT compiler enabled, which is its production configuration, the CPU time added per operation -was: +With the measured driver's runtime in its production configuration, the CPU time added per operation was: | Configuration | Added CPU time per operation | Share of the operation's CPU time | | :------------ | :--------------------------- | :-------------------------------- | | `api-only` | 2–3 µs | 3–4% | | `sdk-never` | 4 µs | 5–6% | -| `sdk-parent` | 5 µs | 6–7% | +| `sdk-ratio` | 5 µs | 6–7% | | `sdk-always` | 8 µs | 9–11% | -Each of these operations creates two spans, an operation span and a command span, so the CPU time per span ranged from -about 1–1.5 µs without an SDK to about 4 µs when every span was recorded. On shared CI hosts, which were 2–3 times -slower, the added CPU time was correspondingly higher and its share similar, but a single run varied by several -microseconds per operation. +Each of these operations creates two spans, an operation span and a command span. ## Changelog -- 2026-10-01: Added the Performance Implications section: sampling-aware implementation guidelines and the required - benchmark configurations and setup, and a Design Rationale entry with the overhead achieved by an implementation - that follows the guidelines. Specified that a valid but unsampled span MUST still be made current so that its - context is propagated (DRIVERS-3620). +- 2026-10-01: Added the Performance Implications section (DRIVERS-3620). - 2026-08-19: Specified the `error.type` attribute on command spans, which drivers MUST add when a command fails and which matches `db.response.status_code` when the command failed with a server error and is otherwise the name of the From a0fb7370fb6327a1a81e1b5e792c4d20c929e758 Mon Sep 17 00:00:00 2001 From: Dmitry Rybakov <160598371+comandeo-mongo@users.noreply.github.com> Date: Thu, 1 Oct 2026 16:24:32 +0200 Subject: [PATCH 06/11] DRIVERS-3620 Drop the sdk-never benchmark configuration api-only isolates the cost of the driver's own code, which drivers control. sdk-ratio already exercises the unsampled path, and comparing it with sdk-always checks that attribute building is gated on the sampling decision. --- source/open-telemetry/open-telemetry.md | 8 +++----- 1 file changed, 3 insertions(+), 5 deletions(-) diff --git a/source/open-telemetry/open-telemetry.md b/source/open-telemetry/open-telemetry.md index ac071b3ba4..e5c36ddc25 100644 --- a/source/open-telemetry/open-telemetry.md +++ b/source/open-telemetry/open-telemetry.md @@ -552,13 +552,12 @@ configurations. No configuration has an active parent span, so that `api-only` p | :------------ | :---------------------------------------------------------------------------------- | | `off` | tracing disabled; the baseline | | `api-only` | tracing enabled; no SDK, so every tracing call is a no-op | -| `sdk-never` | tracing enabled; SDK installed; sampler drops every span | | `sdk-ratio` | tracing enabled; SDK installed; ratio sampler with a representative ratio (e.g. 1%) | | `sdk-always` | tracing enabled; SDK installed; every span sampled | -The differences between consecutive configurations are meaningful: `off` → `api-only` is the driver-side cost, -`api-only` → `sdk-never` is the SDK bookkeeping, and `sdk-never` → `sdk-always` is the cost of recording. If `sdk-never` -costs nearly as much as `sdk-always`, attribute building is not gated on the sampling decision. +The differences between consecutive configurations are meaningful: `off` → `api-only` is the cost of the driver's own +code, which drivers control, and `api-only` → `sdk-ratio` → `sdk-always` adds the cost of the SDK and of recording. If +`sdk-ratio` costs nearly as much as `sdk-always`, attribute building is not gated on the sampling decision. Additionally, when a driver first releases OpenTelemetry support, it MUST compare `off` once against the last driver version without OpenTelemetry support, to show that the disabled instrumentation has no measurable overhead. @@ -650,7 +649,6 @@ With the measured driver's runtime in its production configuration, the CPU time | Configuration | Added CPU time per operation | Share of the operation's CPU time | | :------------ | :--------------------------- | :-------------------------------- | | `api-only` | 2–3 µs | 3–4% | -| `sdk-never` | 4 µs | 5–6% | | `sdk-ratio` | 5 µs | 6–7% | | `sdk-always` | 8 µs | 9–11% | From 15ed3d6ac01dd50a5667536709325339d0e554a6 Mon Sep 17 00:00:00 2001 From: Dmitry Rybakov <160598371+comandeo-mongo@users.noreply.github.com> Date: Thu, 1 Oct 2026 16:46:36 +0200 Subject: [PATCH 07/11] DRIVERS-3620 Simplify the OTel benchmarking requirements after driver review Reviewers from the .NET, Python, and Node drivers found several benchmarking requirements hard to meet or verify in their runtimes: - Drop the per-process MUST, the span count metric, the untestable robustness MUST, and the BSON exclusion. - Downgrade the one-time pre-OTel comparison to SHOULD and compare the driver's default configuration against the commit before OTel landed. - Make CPU time per operation an optional metric. - Name the TraceIdRatioBased sampler, cache the lsid per server session, and refer to the disabled instrumentation in Backwards Compatibility. --- source/open-telemetry/open-telemetry.md | 44 ++++++++++--------------- 1 file changed, 17 insertions(+), 27 deletions(-) diff --git a/source/open-telemetry/open-telemetry.md b/source/open-telemetry/open-telemetry.md index e5c36ddc25..1e4379cca5 100644 --- a/source/open-telemetry/open-telemetry.md +++ b/source/open-telemetry/open-telemetry.md @@ -488,7 +488,7 @@ The OpenTelemetry specification covers all driver operations including but not l ## Backwards Compatibility Introduction of OpenTelemetry in new driver versions should not significantly affect existing applications that do not -enable OpenTelemetry. However, since the no-op tracing operation may introduce some performance degradation (see +enable OpenTelemetry. However, since the disabled instrumentation may introduce some performance degradation (see [Performance Implications section](#performance-implications)), customers should be informed of this feature and how to disable it completely. @@ -533,8 +533,8 @@ Therefore, drivers SHOULD: Deferring attributes is an acceptable trade-off: only custom samplers keying on deferred attributes are affected. Drivers SHOULD compute values that do not change during the life of an object once, and reuse them: connection -attributes for the life of a connection, operation names per operation type, the formatted session id per session. Such -caches MUST be safe for concurrent use. +attributes for the life of a connection, operation names per operation type, the formatted session id per server +session. Such caches MUST be safe for concurrent use. Drivers MAY skip all span work — making the span current, adding attributes, recording the outcome, ending the span — when the created span's context is not valid, i.e. its trace id or span id is all zeroes (see @@ -548,45 +548,35 @@ Drivers MUST benchmark the following configurations, and record the overhead of metric of its own. The benchmark harness, not the driver, installs and configures the SDK and its sampler for the SDK configurations. No configuration has an active parent span, so that `api-only` produces no valid span in any driver. -| Configuration | Tracing setup | -| :------------ | :---------------------------------------------------------------------------------- | -| `off` | tracing disabled; the baseline | -| `api-only` | tracing enabled; no SDK, so every tracing call is a no-op | -| `sdk-ratio` | tracing enabled; SDK installed; ratio sampler with a representative ratio (e.g. 1%) | -| `sdk-always` | tracing enabled; SDK installed; every span sampled | +| Configuration | Tracing setup | +| :------------ | :------------------------------------------------------------------------------------------------ | +| `off` | tracing disabled; the baseline | +| `api-only` | tracing enabled; no SDK, so every tracing call is a no-op | +| `sdk-ratio` | tracing enabled; SDK installed; `TraceIdRatioBased` sampler with a representative ratio (e.g. 1%) | +| `sdk-always` | tracing enabled; SDK installed; every span sampled | The differences between consecutive configurations are meaningful: `off` → `api-only` is the cost of the driver's own code, which drivers control, and `api-only` → `sdk-ratio` → `sdk-always` adds the cost of the SDK and of recording. If `sdk-ratio` costs nearly as much as `sdk-always`, attribute building is not gated on the sampling decision. -Additionally, when a driver first releases OpenTelemetry support, it MUST compare `off` once against the last driver -version without OpenTelemetry support, to show that the disabled instrumentation has no measurable overhead. +Additionally, when a driver first releases OpenTelemetry support, it SHOULD compare its default configuration (normally +`off`) once against the commit before OpenTelemetry support was added, and record the overhead it measures. Drivers SHOULD use their standardized performance testing infrastructure (see [Performance Benchmarking](../benchmarking/benchmarking.md)) rather than a purpose-built OpenTelemetry benchmark. Drivers SHOULD measure the `Small doc insertOne` and `Find one by ID` tasks, through both the synchronous and the -asynchronous API where a driver has both. Tasks that never talk to a server (e.g. BSON micro-benchmarks) MUST be -excluded: they create no spans. +asynchronous API where a driver has both. -Each configuration MUST run in its own process, because an SDK cannot be reliably removed once installed, and `off` and -`api-only` require that none was. - -The overheads measured are small differences between two noisy numbers, so the comparison MUST be robust to the host -changing speed during a run and between runs. Drivers SHOULD run the whole set of configurations interleaved and rotate -their order between repetitions, run compared configurations on the same host, and judge the trend over several runs -rather than a single run against a threshold. +The overheads measured are small differences between two noisy numbers. Drivers SHOULD run the whole set of +configurations interleaved and rotate their order between repetitions, run compared configurations on the same host, and +judge the trend over several runs rather than a single run against a threshold. Drivers SHOULD run the benchmarks in the runtime's default production configuration, with warm-up and iterations as in [Performance Benchmarking](../benchmarking/benchmarking.md). -Drivers SHOULD record, in addition to the throughput score: - -- The CPU time per operation of the threads executing the operations, or process CPU time with background work held - constant. It excludes waiting on the server and so is less noisy than throughput. Drivers MAY also record it - excluding garbage collection time, where the runtime reports it. -- The number of spans per operation for each task, so that a zero overhead is not mistaken for an efficient - implementation. +Drivers MAY also record the CPU time per operation alongside the throughput score: it excludes waiting on the server, so +it is usually less noisy than throughput. SDK configurations SHOULD NOT install an exporter or a span processor. Spans are still sampled and recorded without them; processing cost depends on the host application's choice of processor. From 49a54177ad89d25409b2fc44e50ef9f16b26bcf0 Mon Sep 17 00:00:00 2001 From: Dmitry Rybakov <160598371+comandeo-mongo@users.noreply.github.com> Date: Fri, 2 Oct 2026 10:44:19 +0200 Subject: [PATCH 08/11] DRIVERS-3620 Do not create command spans under an unrecorded span A command span created under an operation span that is not being recorded is either dropped as well or, with a sampler that ignores the parent's decision, exported without its parent. Require drivers to skip it, and to propagate the current span's context in its place so that unsampled contexts still reach the server. Add a prose test that uses a sampler which drops operation spans and records command spans. --- source/open-telemetry/open-telemetry.md | 19 +++++++++++++++---- source/open-telemetry/tests/README.md | 15 +++++++++++++++ 2 files changed, 30 insertions(+), 4 deletions(-) diff --git a/source/open-telemetry/open-telemetry.md b/source/open-telemetry/open-telemetry.md index 1e4379cca5..33598216be 100644 --- a/source/open-telemetry/open-telemetry.md +++ b/source/open-telemetry/open-telemetry.md @@ -262,7 +262,12 @@ application, the same value as the operation span's `exception.type` attribute a #### Instrumenting Server Commands Drivers MUST create a span for every server command sent to the server as a result of a public API call, except for -sensitive commands as listed in the command logging and monitoring specification. +sensitive commands as listed in the command logging and monitoring specification, and except as described below. + +Drivers MUST NOT create a command span when the current span, normally the operation span, has a valid context but is +not being recorded (`isRecording` in the OpenTelemetry API, or the tracing API's equivalent). A command span created +there is either not recorded as well, or, with a sampler that ignores the parent's decision, exported without its +parent. A command sent with no current span is not affected. Spans for commands MUST be nested to the span for the corresponding driver operation span. If the command is being retried, the driver MUST create a separate span for each retry; all the retries MUST be nested to the same operation @@ -414,14 +419,17 @@ single BSON document with the following schema: `traceparent` is the [W3C traceparent](https://www.w3.org/TR/trace-context/#traceparent-header) value of the **command span**: the propagated context MUST be that of the command span for the command being sent, so server spans join the -trace as children of the exact command (and retry attempt) that produced them. +trace as children of the exact command (and retry attempt) that produced them. If no command span was created because +the current span is not being recorded (see [Instrumenting Server Commands](#instrumenting-server-commands)), the +propagated context MUST be that of the current span. Drivers MUST attach the section to a command if and only if all of the following hold: 1. Tracing is enabled for the `MongoClient` (see [Enabling, Disabling, and Configuring OpenTelemetry](#enabling-disabling-and-configuring-opentelemetry)). 2. The connection's `maxWireVersion` is greater than or equal to 29 (MongoDB 9.0). -3. A valid `traceparent` value is available from the command span for the command being sent. +3. A valid `traceparent` value is available from the command span for the command being sent, or from the current span + when no command span was created because it is not being recorded. A `traceparent` value is valid if and only if it is exactly 55 characters of the form `00---` (the field names are @@ -431,7 +439,7 @@ trace-id nor the parent-id is all zeroes. This mirrors the server-side validatio drivers MUST omit the section entirely rather than send an invalid or truncated value. Drivers MUST propagate unsampled trace contexts (trace-flags `00`); the sampling decision MUST NOT affect whether the section is attached. -A message MUST NOT contain more than one telemetry section. Commands that carry no command span (for example server +A message MUST NOT contain more than one telemetry section. Commands that are never traced (for example server monitoring, authentication, and security-sensitive commands) naturally send no section. No tracing data is returned in server responses as part of this feature. @@ -646,6 +654,9 @@ Each of these operations creates two spans, an operation span and a command span ## Changelog +- 2026-10-02: Specified that drivers MUST NOT create a command span when the current span is not being recorded, and + propagate the current span's context instead (DRIVERS-3620). + - 2026-10-01: Added the Performance Implications section (DRIVERS-3620). - 2026-08-19: Specified the `error.type` attribute on command spans, which drivers MUST add when a command fails and diff --git a/source/open-telemetry/tests/README.md b/source/open-telemetry/tests/README.md index d4e74f7f46..d540580ce7 100644 --- a/source/open-telemetry/tests/README.md +++ b/source/open-telemetry/tests/README.md @@ -202,3 +202,18 @@ replica set or sharded cluster). > The server may create spans for commands the driver did not trace (head-based sampling applies server-side too). > Assertions are therefore always made on the join between driver `traceId`s and server spans, never on the raw contents > of the trace directory. + +#### Unrecorded Parent Spans + +*Test 10: No command span under an unrecorded operation span* + +This test verifies that drivers do not create a command span when the operation span is not being recorded (see +[Instrumenting Server Commands](../open-telemetry.md#instrumenting-server-commands)). It needs a sampler that ignores +the parent's decision: with a parent-based sampler, a command span under an unrecorded operation span is not recorded +either, so the test could not tell whether it was created. + +1. Install a tracer provider whose sampler drops spans that carry the `db.operation.name` attribute (operation spans) + and records all other spans. +2. Create a `MongoClient` with tracing enabled that uses this tracer provider. +3. Perform an `insertOne` operation on a test collection. +4. Assert that no recorded span carries the `db.command.name` attribute with the value `insert`. From 947b7b7ec6563e28005745f89e1b8b6e19d20183 Mon Sep 17 00:00:00 2001 From: Dmitry Rybakov <160598371+comandeo-mongo@users.noreply.github.com> Date: Fri, 2 Oct 2026 10:45:34 +0200 Subject: [PATCH 09/11] Revert "DRIVERS-3620 Do not create command spans under an unrecorded span" This reverts commit 49a54177ad89d25409b2fc44e50ef9f16b26bcf0. --- source/open-telemetry/open-telemetry.md | 19 ++++--------------- source/open-telemetry/tests/README.md | 15 --------------- 2 files changed, 4 insertions(+), 30 deletions(-) diff --git a/source/open-telemetry/open-telemetry.md b/source/open-telemetry/open-telemetry.md index 33598216be..1e4379cca5 100644 --- a/source/open-telemetry/open-telemetry.md +++ b/source/open-telemetry/open-telemetry.md @@ -262,12 +262,7 @@ application, the same value as the operation span's `exception.type` attribute a #### Instrumenting Server Commands Drivers MUST create a span for every server command sent to the server as a result of a public API call, except for -sensitive commands as listed in the command logging and monitoring specification, and except as described below. - -Drivers MUST NOT create a command span when the current span, normally the operation span, has a valid context but is -not being recorded (`isRecording` in the OpenTelemetry API, or the tracing API's equivalent). A command span created -there is either not recorded as well, or, with a sampler that ignores the parent's decision, exported without its -parent. A command sent with no current span is not affected. +sensitive commands as listed in the command logging and monitoring specification. Spans for commands MUST be nested to the span for the corresponding driver operation span. If the command is being retried, the driver MUST create a separate span for each retry; all the retries MUST be nested to the same operation @@ -419,17 +414,14 @@ single BSON document with the following schema: `traceparent` is the [W3C traceparent](https://www.w3.org/TR/trace-context/#traceparent-header) value of the **command span**: the propagated context MUST be that of the command span for the command being sent, so server spans join the -trace as children of the exact command (and retry attempt) that produced them. If no command span was created because -the current span is not being recorded (see [Instrumenting Server Commands](#instrumenting-server-commands)), the -propagated context MUST be that of the current span. +trace as children of the exact command (and retry attempt) that produced them. Drivers MUST attach the section to a command if and only if all of the following hold: 1. Tracing is enabled for the `MongoClient` (see [Enabling, Disabling, and Configuring OpenTelemetry](#enabling-disabling-and-configuring-opentelemetry)). 2. The connection's `maxWireVersion` is greater than or equal to 29 (MongoDB 9.0). -3. A valid `traceparent` value is available from the command span for the command being sent, or from the current span - when no command span was created because it is not being recorded. +3. A valid `traceparent` value is available from the command span for the command being sent. A `traceparent` value is valid if and only if it is exactly 55 characters of the form `00---` (the field names are @@ -439,7 +431,7 @@ trace-id nor the parent-id is all zeroes. This mirrors the server-side validatio drivers MUST omit the section entirely rather than send an invalid or truncated value. Drivers MUST propagate unsampled trace contexts (trace-flags `00`); the sampling decision MUST NOT affect whether the section is attached. -A message MUST NOT contain more than one telemetry section. Commands that are never traced (for example server +A message MUST NOT contain more than one telemetry section. Commands that carry no command span (for example server monitoring, authentication, and security-sensitive commands) naturally send no section. No tracing data is returned in server responses as part of this feature. @@ -654,9 +646,6 @@ Each of these operations creates two spans, an operation span and a command span ## Changelog -- 2026-10-02: Specified that drivers MUST NOT create a command span when the current span is not being recorded, and - propagate the current span's context instead (DRIVERS-3620). - - 2026-10-01: Added the Performance Implications section (DRIVERS-3620). - 2026-08-19: Specified the `error.type` attribute on command spans, which drivers MUST add when a command fails and diff --git a/source/open-telemetry/tests/README.md b/source/open-telemetry/tests/README.md index d540580ce7..d4e74f7f46 100644 --- a/source/open-telemetry/tests/README.md +++ b/source/open-telemetry/tests/README.md @@ -202,18 +202,3 @@ replica set or sharded cluster). > The server may create spans for commands the driver did not trace (head-based sampling applies server-side too). > Assertions are therefore always made on the join between driver `traceId`s and server spans, never on the raw contents > of the trace directory. - -#### Unrecorded Parent Spans - -*Test 10: No command span under an unrecorded operation span* - -This test verifies that drivers do not create a command span when the operation span is not being recorded (see -[Instrumenting Server Commands](../open-telemetry.md#instrumenting-server-commands)). It needs a sampler that ignores -the parent's decision: with a parent-based sampler, a command span under an unrecorded operation span is not recorded -either, so the test could not tell whether it was created. - -1. Install a tracer provider whose sampler drops spans that carry the `db.operation.name` attribute (operation spans) - and records all other spans. -2. Create a `MongoClient` with tracing enabled that uses this tracer provider. -3. Perform an `insertOne` operation on a test collection. -4. Assert that no recorded span carries the `db.command.name` attribute with the value `insert`. From 530837d147800deedb1d652035799bb52a8a7fca Mon Sep 17 00:00:00 2001 From: Dmitry Rybakov <160598371+comandeo-mongo@users.noreply.github.com> Date: Mon, 5 Oct 2026 11:25:41 +0200 Subject: [PATCH 10/11] DRIVERS-3620 Point the OTel overhead section at the implementation PR The reference table came from a single local run and no longer matches the current measurements in the linked Ruby implementation PR, so readers could not reproduce it. Drop the figures and point to that PR for the current numbers and the environment they were measured in. --- source/open-telemetry/open-telemetry.md | 20 +++++++------------- 1 file changed, 7 insertions(+), 13 deletions(-) diff --git a/source/open-telemetry/open-telemetry.md b/source/open-telemetry/open-telemetry.md index 1e4379cca5..ca04a14e16 100644 --- a/source/open-telemetry/open-telemetry.md +++ b/source/open-telemetry/open-telemetry.md @@ -630,19 +630,13 @@ redesigning the payload format. ### What overhead is achievable? -The figures below were measured with one driver's implementation, on the `Small doc insertOne` and `Find one by ID` -tasks against a standalone server on the same host. They are a reference for what an implementation can achieve, not -requirements and not targets: other runtimes can be expected to land elsewhere. - -With the measured driver's runtime in its production configuration, the CPU time added per operation was: - -| Configuration | Added CPU time per operation | Share of the operation's CPU time | -| :------------ | :--------------------------- | :-------------------------------- | -| `api-only` | 2–3 µs | 3–4% | -| `sdk-ratio` | 5 µs | 6–7% | -| `sdk-always` | 8 µs | 9–11% | - -Each of these operations creates two spans, an operation span and a command span. +The overhead of tracing depends on the language runtime, the host and the workload, so this specification sets no +numeric target. A reference implementation that follows the [Implementation Guidelines](#implementation-guidelines) +measured the CPU time it adds per operation on the `Small doc insertOne` and `Find one by ID` tasks, each of which +creates an operation span and a command span. The current figures, together with the environment they were measured in, +are reported in that implementation's pull request: +[mongo-ruby-driver#3108](https://github.com/mongodb/mongo-ruby-driver/pull/3108). They are a reference for what an +implementation can achieve, not requirements: other runtimes can be expected to land elsewhere. ## Changelog From 523dec6990efa4b4b3d37732e93a698f6d43d81e Mon Sep 17 00:00:00 2001 From: Dmitry Rybakov <160598371+comandeo-mongo@users.noreply.github.com> Date: Mon, 5 Oct 2026 11:43:11 +0200 Subject: [PATCH 11/11] Remove references to Ruby numbers --- source/open-telemetry/open-telemetry.md | 17 +++-------------- 1 file changed, 3 insertions(+), 14 deletions(-) diff --git a/source/open-telemetry/open-telemetry.md b/source/open-telemetry/open-telemetry.md index ca04a14e16..75ecd8347d 100644 --- a/source/open-telemetry/open-telemetry.md +++ b/source/open-telemetry/open-telemetry.md @@ -508,10 +508,9 @@ guidance of the Command Logging and Monitoring spec. Drivers MUST measure the performance impact of their OpenTelemetry implementations as described in [Benchmarking](#benchmarking), and SHOULD follow the [Implementation Guidelines](#implementation-guidelines). This -specification sets no numeric limit: the cost depends on the language runtime, the host and the workload. For reference, -[What overhead is achievable?](#what-overhead-is-achievable) describes the results of an implementation that follows the -guidelines. In this section, "SDK" means the tracer implementation a host application installs; where tracing is built -into the runtime, it means the component that subscribes to and records the runtime's spans. +specification sets no numeric limit: the cost depends on the language runtime, the host and the workload. In this +section, "SDK" means the tracer implementation a host application installs; where tracing is built into the runtime, it +means the component that subscribes to and records the runtime's spans. ### Implementation Guidelines @@ -628,16 +627,6 @@ negotiation mechanism. Carrying the traceparent inside a BSON document allows future propagation fields to be added to the same section without redesigning the payload format. -### What overhead is achievable? - -The overhead of tracing depends on the language runtime, the host and the workload, so this specification sets no -numeric target. A reference implementation that follows the [Implementation Guidelines](#implementation-guidelines) -measured the CPU time it adds per operation on the `Small doc insertOne` and `Find one by ID` tasks, each of which -creates an operation span and a command span. The current figures, together with the environment they were measured in, -are reported in that implementation's pull request: -[mongo-ruby-driver#3108](https://github.com/mongodb/mongo-ruby-driver/pull/3108). They are a reference for what an -implementation can achieve, not requirements: other runtimes can be expected to land elsewhere. - ## Changelog - 2026-10-01: Added the Performance Implications section (DRIVERS-3620).