Skip to content

DRIVERS-3620 Add Performance Implications section to the OpenTelemetry spec - #1986

Open
comandeo-mongo wants to merge 11 commits into
mongodb:masterfrom
comandeo-mongo:DRIVERS-3620
Open

comandeo-mongo wants to merge 11 commits into
mongodb:masterfrom
comandeo-mongo:DRIVERS-3620

Conversation

@comandeo-mongo

@comandeo-mongo comandeo-mongo commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

DRIVERS-3620

Preliminary benchmarks for DRIVERS-719 showed 15-50% throughput loss with OpenTelemetry enabled. Benchmarking in the Ruby driver located the cost: almost all of it was driver code building span attributes before the sampling decision was known, so dropping 100% of traces recovered almost nothing. After restructuring the Ruby implementation, the non-recording configurations went from 17-26% overhead to 4-8%.

This PR adds a Performance Implications section to the OpenTelemetry spec so other drivers apply the same structure.

Ruby implementation: mongodb/mongo-ruby-driver#3108

Please complete the following before merging:

  • Is the relevant DRIVERS ticket in the PR title?
  • Update changelog.
  • Test changes in at least one language driver.
  • Test these changes against all server versions and topologies (including standalone, replica set, and sharded
    clusters).

Other specs set no numeric performance limits; measured results live in
rationale sections. The 5/10/15% targets are removed. The figures move to
a Design Rationale entry as what an implementation following the
guidelines achieved in Ruby, including with YJIT.

The benchmarking rules gain what the measurements showed: which tasks to
use, rotating the configuration order, comparing tasks on one host,
running with the production runtime configuration (e.g. a JIT), recording
CPU time per operation with and without GC, and not judging a single run
against a fixed threshold.
- Drop the claim that drivers do not trace hello: only sensitive commands
  are excluded. Run command is a poor signal because its spans have no
  namespace or collection.
- Make the pre-OTel comparison an explicit one-time requirement for new
  implementations, and move the per-configuration explanations out of the
  table.
- State the goal (robust to host drift) as the requirement and interleaving,
  rotation, same-host runs and trend-over-single-run as SHOULDs, rather than
  dictating the CI layout.
- Require no active parent span, and state that the harness, not the driver,
  configures the SDK and sampler.
- Move the MUST about making unsampled spans current into the propagation
  section, next to the rule on propagating unsampled contexts.
- Benchmark in the runtime's default production configuration; GC-excluded
  CPU time is a MAY.
- Rationale figures name no driver or runtime, name the metric in each
  sentence, and the rule of thumb is dropped.
The Host Application Level option cited Ruby as an example of enabling
all instrumentations, which says nothing a reader of other languages can
use.
…e-neutral

- Drop the rule that unsampled spans must be made current: the nesting and
  propagation rules already require the outcome.
- Define "SDK" once, neutrally, and describe isRecording as the tracing
  API's equivalent.
- Rename sdk-parent to sdk-ratio: with no parent span it is a ratio sampler.
- Scope CPU time to the threads executing the operations, or process CPU
  time with background work held constant; benchmark both sync and async
  APIs where a driver has both.
- Cut repeated rationale, the allocation test and the interpreting-runtime
  figures.
api-only isolates the cost of the driver's own code, which drivers
control. sdk-ratio already exercises the unsampled path, and comparing
it with sdk-always checks that attribute building is gated on the
sampling decision.
… review

Reviewers from the .NET, Python, and Node drivers found several benchmarking
requirements hard to meet or verify in their runtimes:

- Drop the per-process MUST, the span count metric, the untestable
  robustness MUST, and the BSON exclusion.
- Downgrade the one-time pre-OTel comparison to SHOULD and compare the
  driver's default configuration against the commit before OTel landed.
- Make CPU time per operation an optional metric.
- Name the TraceIdRatioBased sampler, cache the lsid per server session,
  and refer to the disabled instrumentation in Backwards Compatibility.
A command span created under an operation span that is not being
recorded is either dropped as well or, with a sampler that ignores the
parent's decision, exported without its parent. Require drivers to skip
it, and to propagate the current span's context in its place so that
unsampled contexts still reach the server.

Add a prose test that uses a sampler which drops operation spans and
records command spans.
@comandeo-mongo
comandeo-mongo marked this pull request as ready for review October 5, 2026 08:44
@comandeo-mongo
comandeo-mongo requested a review from a team as a code owner October 5, 2026 08:44
@comandeo-mongo
comandeo-mongo requested review from a team, NoahStapp, jyemin and nhachicha and a balanced review from Copilot and removed request for a team October 5, 2026 08:44

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The reference measurements conflict with the linked implementation, and one benchmark diagnostic overstates its conclusion.

Review effort: Balanced
Findings: 2 Low severity

Open (2)
What changed in this PR

Adds performance guidance and benchmarking requirements to the OpenTelemetry specification.

Changes:

  • Recommends deferring expensive span attributes until sampling is known.
  • Defines benchmark configurations and methodology.
  • Adds reference overhead measurements and updates the changelog.
File Description
source/​open-telemetry/​open-telemetry.md Documents OpenTelemetry performance practices and benchmarks.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.


The differences between consecutive configurations are meaningful: `off` → `api-only` is the cost of the driver's own
code, which drivers control, and `api-only` → `sdk-ratio` → `sdk-always` adds the cost of the SDK and of recording. If
`sdk-ratio` costs nearly as much as `sdk-always`, attribute building is not gated on the sampling decision.
Comment thread source/open-telemetry/open-telemetry.md Outdated
The reference table came from a single local run and no longer matches
the current measurements in the linked Ruby implementation PR, so
readers could not reproduce it. Drop the figures and point to that PR
for the current numbers and the environment they were measured in.

Therefore, drivers SHOULD:

- Create a span with a minimum set of attributes that are cheap to create. The recommended set for command spans is

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Have we benchmarked what performance difference (if any) each attribute adds? Can we create a span with zero attributes to absolutely minimize performance impact and then add all of them once we know the span will be recorded?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Attributes that are added after span creation are not visible to samplers. OpenTelemetry recommends "Prefer adding attributes at span creation to make the attributes available to SDK sampling."

Good idea about benchmarking every attribute, I'll do it.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Benchmarked it. All figures below are the Ruby driver with OpenTelemetry SDK 1.13.0, Ruby 4.0.1 + YJIT, on Evergreen; median of 20 reps x 50,000 spans per profile. Each profile is timed against the same span created with an empty attribute hash, the cheapest span the driver can create: ~1.38 µs/span when recording. The relative columns use the untraced DriverBench operations measured on the same host (find ≈ 214 µs, insert ≈ 189 µs, CPU excluding GC).

Each attribute, added at span creation (one attribute at a time):

Attribute Type +µs/span % of empty span % of untraced find % of untraced insert
db.system.name string +0.120 +8.7% 0.056% 0.064%
db.namespace string +0.117 +8.5% 0.055% 0.062%
db.command.name string +0.122 +8.8% 0.057% 0.064%
db.collection.name string +0.120 +8.7% 0.056% 0.063%
server.address string +0.123 +8.9% 0.058% 0.065%
network.transport string +0.122 +8.8% 0.057% 0.064%
db.query.summary string +0.121 +8.7% 0.056% 0.064%
db.mongodb.lsid string +0.122 +8.8% 0.057% 0.064%
server.port integer +0.284 +20.5% 0.133% 0.150%
db.mongodb.server_connection_id integer +0.286 +20.7% 0.134% 0.151%
db.mongodb.driver_connection_id integer +0.283 +20.5% 0.132% 0.150%
db.mongodb.cursor_id integer +0.288 +20.8% 0.135% 0.152%
db.mongodb.txn_number integer +0.287 +20.8% 0.134% 0.152%

Each attribute costs roughly 0.12 µs/span, about 0.05-0.07% of the untraced operation; none stands out as individually expensive. The integer rows read ~0.28 µs, ~2.4x the strings. That is an artifact of the sweep, not the SDK: timing single-attribute profiles one after another in one process makes the SDK's attribute-validation call sites flip between String and Integer receivers and get recompiled under YJIT. With one profile per process the gap disappears, and it does not appear in a realistic multi-attribute hash: the 9-attribute creation set adds 1.04 µs total, i.e. ~0.12 µs per attribute regardless of type.

Adding an attribute after creation costs ~0.33 µs/span (set_attribute, measured over the 4 deferred attributes), about 2.8x an attribute added at creation.

"Zero attributes at creation, add all of them once we know the span is recorded" is worse for recording spans. Per command span, against the empty-attribute span:

Span shape +µs/span % of untraced find
all 13 attributes at creation +1.45 0.68%
current shape (cheap set at creation, rest deferred behind isRecording) +2.21 1.03%
zero at creation, all 13 via set_attribute +3.81 1.78%

Creating the span empty and adding everything later costs ~1.6 µs/span more when the span is recorded, because set_attribute is the expensive path. It does avoid building attributes when the span is dropped, but that only pays off when nearly all spans are dropped, and it makes every attribute invisible to samplers.

Two things the numbers also show:

  • Span creation, not attributes, is the bulk of the cost. A recorded span with no attributes is already ~1.38 µs at the SDK boundary, and the full 13-attribute set is under 1% of the untraced operation.
  • On a span the sampler dropped, attributes are essentially free: every profile is within ~0.01 µs of the empty-attribute span. That is the property the "cheap set at creation, rest behind the recording check" structure is built around.

Raw output: Evergreen patch 1996.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like it's not worth it to delay adding the basic attributes until we know it'll be sampled, thanks for the benchmarking!

Drivers MAY skip all span work — making the span current, adding attributes, recording the outcome, ending the span —
when the created span's context is not valid, i.e. its trace id or span id is all zeroes (see
[IsValid](https://opentelemetry.io/docs/specs/otel/trace/api/#isvalid)): such a span cannot be propagated or correlated,
so any work on it is wasted. This permission does not apply to a valid but unsampled span, which remains subject to the

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What does this imply? Can we have an unsampled parent span, but then a child span is sampled and requires the parent to be valid even though it wasn't initially sampled? Or is this only about propagation to the server?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I raised https://jira.mongodb.org/browse/DRIVERS-3669 to clarify this in the specification.

Additionally, when a driver first releases OpenTelemetry support, it SHOULD compare its default configuration (normally
`off`) once against the commit before OpenTelemetry support was added, and record the overhead it measures.

Drivers SHOULD use their standardized performance testing infrastructure (see

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should this be a "MUST"? Why would a driver ever want to not re-use their existing infra?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tend to avoid using "MUST" when we telling other teams how to achieve things. Like here, we MUST benchmark, and here is how we recommend to do so. If you find a way that works better in your ecosystem and still achieve the benchmarking goals, it's okay.

However, if you feel that MUST suits here better, I'm totally fine with that.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's fair. My main concern is drift between drivers resulting in tangible differences: if one driver uses their existing infrastructure that results in overall better performance for OTel than another driver that uses a bespoke testing setup, which one is more valid? If everyone is required to use the same rough benchmarking setup that we already have, it reduces the risk of such a situation.

Drivers SHOULD use their standardized performance testing infrastructure (see
[Performance Benchmarking](../benchmarking/benchmarking.md)) rather than a purpose-built OpenTelemetry benchmark.

Drivers SHOULD measure the `Small doc insertOne` and `Find one by ID` tasks, through both the synchronous and the

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same here, what context would cause a driver to choose to measure different tasks?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants