Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
87 changes: 83 additions & 4 deletions source/open-telemetry/open-telemetry.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,8 +72,8 @@ Drivers SHOULD support configuring OpenTelemetry on multiple levels.
environment variable `OTEL_#{LANG}_INSTRUMENTATION_MONGODB_ENABLED`. Drivers MAY provide other means to globally
disable OpenTelemetry that are more suitable for their language ecosystem. This option MUST override settings on the
higher level.
- **Host Application Level**: If the host application enables OpenTelemetry for all available instrumentations (e.g.,
Ruby), and a driver can detect this, OpenTelemetry SHOULD be enabled in the driver.
- **Host Application Level**: If the host application enables OpenTelemetry for all available instrumentations, and a
driver can detect this, OpenTelemetry SHOULD be enabled in the driver.

Drivers MUST NOT try to detect whether the OpenTelemetry SDK library is available, and enable tracing based on this.
Drivers MUST NOT add means that configure OpenTelemetry SDK (e.g., setting a specific exporter). Drivers MUST NOT add
Expand Down Expand Up @@ -488,8 +488,9 @@ The OpenTelemetry specification covers all driver operations including but not l
## Backwards Compatibility

Introduction of OpenTelemetry in new driver versions should not significantly affect existing applications that do not
enable OpenTelemetry. However, since the no-op tracing operation may introduce some performance degradation (though it
should be negligible), customers should be informed of this feature and how to disable it completely.
enable OpenTelemetry. However, since the disabled instrumentation may introduce some performance degradation (see
[Performance Implications section](#performance-implications)), customers should be informed of this feature and how to
disable it completely.

If a driver is used in an application that has OpenTelemetry enabled, customers will see traces from the driver in their
OpenTelemetry backends. This may be unexpected and MAY cause negative effects in some cases (e.g., the OpenTelemetry
Expand All @@ -503,6 +504,82 @@ SHOULD follow the
[Security](https://github.com/mongodb/specifications/blob/master/source/command-logging-and-monitoring/command-logging-and-monitoring.md#security)
guidance of the Command Logging and Monitoring spec.

## Performance Implications

Drivers MUST measure the performance impact of their OpenTelemetry implementations as described in
[Benchmarking](#benchmarking), and SHOULD follow the [Implementation Guidelines](#implementation-guidelines). This
specification sets no numeric limit: the cost depends on the language runtime, the host and the workload. In this
section, "SDK" means the tracer implementation a host application installs; where tracing is built into the runtime, it
means the component that subscribes to and records the runtime's spans.

### Implementation Guidelines

Not every span is recorded: samplers decide whether a span is recorded, and the tracing API reports this decision only
on a span that has already been created (`isRecording` in the OpenTelemetry API, or the tracing API's equivalent). Only
the attributes provided at span creation are visible to the sampler.

Therefore, drivers SHOULD:

- Create a span with a minimum set of attributes that are cheap to create. The recommended set for command spans is

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Have we benchmarked what performance difference (if any) each attribute adds? Can we create a span with zero attributes to absolutely minimize performance impact and then add all of them once we know the span will be recorded?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Attributes that are added after span creation are not visible to samplers. OpenTelemetry recommends "Prefer adding attributes at span creation to make the attributes available to SDK sampling."

Good idea about benchmarking every attribute, I'll do it.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Benchmarked it. All figures below are the Ruby driver with OpenTelemetry SDK 1.13.0, Ruby 4.0.1 + YJIT, on Evergreen; median of 20 reps x 50,000 spans per profile. Each profile is timed against the same span created with an empty attribute hash, the cheapest span the driver can create: ~1.38 µs/span when recording. The relative columns use the untraced DriverBench operations measured on the same host (find ≈ 214 µs, insert ≈ 189 µs, CPU excluding GC).

Each attribute, added at span creation (one attribute at a time):

Attribute Type +µs/span % of empty span % of untraced find % of untraced insert
db.system.name string +0.120 +8.7% 0.056% 0.064%
db.namespace string +0.117 +8.5% 0.055% 0.062%
db.command.name string +0.122 +8.8% 0.057% 0.064%
db.collection.name string +0.120 +8.7% 0.056% 0.063%
server.address string +0.123 +8.9% 0.058% 0.065%
network.transport string +0.122 +8.8% 0.057% 0.064%
db.query.summary string +0.121 +8.7% 0.056% 0.064%
db.mongodb.lsid string +0.122 +8.8% 0.057% 0.064%
server.port integer +0.284 +20.5% 0.133% 0.150%
db.mongodb.server_connection_id integer +0.286 +20.7% 0.134% 0.151%
db.mongodb.driver_connection_id integer +0.283 +20.5% 0.132% 0.150%
db.mongodb.cursor_id integer +0.288 +20.8% 0.135% 0.152%
db.mongodb.txn_number integer +0.287 +20.8% 0.134% 0.152%

Each attribute costs roughly 0.12 µs/span, about 0.05-0.07% of the untraced operation; none stands out as individually expensive. The integer rows read ~0.28 µs, ~2.4x the strings. That is an artifact of the sweep, not the SDK: timing single-attribute profiles one after another in one process makes the SDK's attribute-validation call sites flip between String and Integer receivers and get recompiled under YJIT. With one profile per process the gap disappears, and it does not appear in a realistic multi-attribute hash: the 9-attribute creation set adds 1.04 µs total, i.e. ~0.12 µs per attribute regardless of type.

Adding an attribute after creation costs ~0.33 µs/span (set_attribute, measured over the 4 deferred attributes), about 2.8x an attribute added at creation.

"Zero attributes at creation, add all of them once we know the span is recorded" is worse for recording spans. Per command span, against the empty-attribute span:

Span shape +µs/span % of untraced find
all 13 attributes at creation +1.45 0.68%
current shape (cheap set at creation, rest deferred behind isRecording) +2.21 1.03%
zero at creation, all 13 via set_attribute +3.81 1.78%

Creating the span empty and adding everything later costs ~1.6 µs/span more when the span is recorded, because set_attribute is the expensive path. It does avoid building attributes when the span is dropped, but that only pays off when nearly all spans are dropped, and it makes every attribute invisible to samplers.

Two things the numbers also show:

  • Span creation, not attributes, is the bulk of the cost. A recorded span with no attributes is already ~1.38 µs at the SDK boundary, and the full 13-attribute set is under 1% of the untraced operation.
  • On a span the sampler dropped, attributes are essentially free: every profile is within ~0.01 µs of the empty-attribute span. That is the property the "cheap set at creation, rest behind the recording check" structure is built around.

Raw output: Evergreen patch 1996.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like it's not worth it to delay adding the basic attributes until we know it'll be sampled, thanks for the benchmarking!

`db.system.name`, `db.namespace`, `db.collection.name`, `db.command.name`, `server.address`, `server.port`,
`network.transport`, `db.mongodb.server_connection_id`, and `db.mongodb.driver_connection_id`. The recommended set
for operation spans is `db.system.name`, `db.namespace`, `db.collection.name`, `db.operation.name`, and
`db.operation.summary`.
- After creating the span, check whether the span is being recorded. If not, no more attributes are added to the span.
If yes, add the rest of the attributes to the span: `db.query.summary`, `db.query.text`, `db.mongodb.lsid`,
`db.mongodb.cursor_id`, and `db.mongodb.txn_number` for command spans, `db.mongodb.cursor_id` for operation spans.

Deferring attributes is an acceptable trade-off: only custom samplers keying on deferred attributes are affected.

Drivers SHOULD compute values that do not change during the life of an object once, and reuse them: connection
attributes for the life of a connection, operation names per operation type, the formatted session id per server
session. Such caches MUST be safe for concurrent use.

Drivers MAY skip all span work — making the span current, adding attributes, recording the outcome, ending the span —
when the created span's context is not valid, i.e. its trace id or span id is all zeroes (see
[IsValid](https://opentelemetry.io/docs/specs/otel/trace/api/#isvalid)): such a span cannot be propagated or correlated,
so any work on it is wasted. This permission does not apply to a valid but unsampled span, which remains subject to the

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What does this imply? Can we have an unsampled parent span, but then a child span is sampled and requires the parent to be valid even though it wasn't initially sampled? Or is this only about propagation to the server?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I raised https://jira.mongodb.org/browse/DRIVERS-3669 to clarify this in the specification.

nesting and [propagation](#propagating-trace-context-to-the-server) rules.

### Benchmarking

Drivers MUST benchmark the following configurations, and record the overhead of each relative to the baseline as a
metric of its own. The benchmark harness, not the driver, installs and configures the SDK and its sampler for the SDK
configurations. No configuration has an active parent span, so that `api-only` produces no valid span in any driver.

| Configuration | Tracing setup |
| :------------ | :------------------------------------------------------------------------------------------------ |
| `off` | tracing disabled; the baseline |
| `api-only` | tracing enabled; no SDK, so every tracing call is a no-op |
| `sdk-ratio` | tracing enabled; SDK installed; `TraceIdRatioBased` sampler with a representative ratio (e.g. 1%) |
| `sdk-always` | tracing enabled; SDK installed; every span sampled |

The differences between consecutive configurations are meaningful: `off` → `api-only` is the cost of the driver's own
code, which drivers control, and `api-only` → `sdk-ratio` → `sdk-always` adds the cost of the SDK and of recording. If
`sdk-ratio` costs nearly as much as `sdk-always`, attribute building is not gated on the sampling decision.

Additionally, when a driver first releases OpenTelemetry support, it SHOULD compare its default configuration (normally
`off`) once against the commit before OpenTelemetry support was added, and record the overhead it measures.

Drivers SHOULD use their standardized performance testing infrastructure (see

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should this be a "MUST"? Why would a driver ever want to not re-use their existing infra?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tend to avoid using "MUST" when we telling other teams how to achieve things. Like here, we MUST benchmark, and here is how we recommend to do so. If you find a way that works better in your ecosystem and still achieve the benchmarking goals, it's okay.

However, if you feel that MUST suits here better, I'm totally fine with that.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's fair. My main concern is drift between drivers resulting in tangible differences: if one driver uses their existing infrastructure that results in overall better performance for OTel than another driver that uses a bespoke testing setup, which one is more valid? If everyone is required to use the same rough benchmarking setup that we already have, it reduces the risk of such a situation.

[Performance Benchmarking](../benchmarking/benchmarking.md)) rather than a purpose-built OpenTelemetry benchmark.

Drivers SHOULD measure the `Small doc insertOne` and `Find one by ID` tasks, through both the synchronous and the

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same here, what context would cause a driver to choose to measure different tasks?

asynchronous API where a driver has both.

The overheads measured are small differences between two noisy numbers. Drivers SHOULD run the whole set of
configurations interleaved and rotate their order between repetitions, run compared configurations on the same host, and
judge the trend over several runs rather than a single run against a threshold.

Drivers SHOULD run the benchmarks in the runtime's default production configuration, with warm-up and iterations as in
[Performance Benchmarking](../benchmarking/benchmarking.md).

Drivers MAY also record the CPU time per operation alongside the throughput score: it excludes waiting on the server, so
it is usually less noisy than throughput.

SDK configurations SHOULD NOT install an exporter or a span processor. Spans are still sampled and recorded without
them; processing cost depends on the host application's choice of processor.

## Future Work

### Query Parametrization
Expand Down Expand Up @@ -552,6 +629,8 @@ redesigning the payload format.

## Changelog

- 2026-10-01: Added the Performance Implications section (DRIVERS-3620).

- 2026-08-19: Specified the `error.type` attribute on command spans, which drivers MUST add when a command fails and
which matches `db.response.status_code` when the command failed with a server error and is otherwise the name of the
exception class associated with that command's failure. Specified that drivers MUST NOT set it when the command
Expand Down
Loading