Repository navigation
DRIVERS-3620 Add Performance Implications section to the OpenTelemetry spec - #1986
comandeo-mongo wants to merge 11 commits into
Conversation
Other specs set no numeric performance limits; measured results live in rationale sections. The 5/10/15% targets are removed. The figures move to a Design Rationale entry as what an implementation following the guidelines achieved in Ruby, including with YJIT. The benchmarking rules gain what the measurements showed: which tasks to use, rotating the configuration order, comparing tasks on one host, running with the production runtime configuration (e.g. a JIT), recording CPU time per operation with and without GC, and not judging a single run against a fixed threshold.
- Drop the claim that drivers do not trace hello: only sensitive commands are excluded. Run command is a poor signal because its spans have no namespace or collection. - Make the pre-OTel comparison an explicit one-time requirement for new implementations, and move the per-configuration explanations out of the table. - State the goal (robust to host drift) as the requirement and interleaving, rotation, same-host runs and trend-over-single-run as SHOULDs, rather than dictating the CI layout. - Require no active parent span, and state that the harness, not the driver, configures the SDK and sampler. - Move the MUST about making unsampled spans current into the propagation section, next to the rule on propagating unsampled contexts. - Benchmark in the runtime's default production configuration; GC-excluded CPU time is a MAY. - Rationale figures name no driver or runtime, name the metric in each sentence, and the rule of thumb is dropped.
The Host Application Level option cited Ruby as an example of enabling all instrumentations, which says nothing a reader of other languages can use.
…e-neutral - Drop the rule that unsampled spans must be made current: the nesting and propagation rules already require the outcome. - Define "SDK" once, neutrally, and describe isRecording as the tracing API's equivalent. - Rename sdk-parent to sdk-ratio: with no parent span it is a ratio sampler. - Scope CPU time to the threads executing the operations, or process CPU time with background work held constant; benchmark both sync and async APIs where a driver has both. - Cut repeated rationale, the allocation test and the interpreting-runtime figures.
api-only isolates the cost of the driver's own code, which drivers control. sdk-ratio already exercises the unsampled path, and comparing it with sdk-always checks that attribute building is gated on the sampling decision.
… review Reviewers from the .NET, Python, and Node drivers found several benchmarking requirements hard to meet or verify in their runtimes: - Drop the per-process MUST, the span count metric, the untestable robustness MUST, and the BSON exclusion. - Downgrade the one-time pre-OTel comparison to SHOULD and compare the driver's default configuration against the commit before OTel landed. - Make CPU time per operation an optional metric. - Name the TraceIdRatioBased sampler, cache the lsid per server session, and refer to the disabled instrumentation in Backwards Compatibility.
A command span created under an operation span that is not being recorded is either dropped as well or, with a sampler that ignores the parent's decision, exported without its parent. Require drivers to skip it, and to propagate the current span's context in its place so that unsampled contexts still reach the server. Add a prose test that uses a sampler which drops operation spans and records command spans.
…span" This reverts commit 49a5417.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The reference measurements conflict with the linked implementation, and one benchmark diagnostic overstates its conclusion.
Review effort: Balanced
Findings: 2
Open (2)
What changed in this PR
Adds performance guidance and benchmarking requirements to the OpenTelemetry specification.
Changes:
- Recommends deferring expensive span attributes until sampling is known.
- Defines benchmark configurations and methodology.
- Adds reference overhead measurements and updates the changelog.
| File | Description |
|---|---|
source/open-telemetry/open-telemetry.md |
Documents OpenTelemetry performance practices and benchmarks. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
|
||
| The differences between consecutive configurations are meaningful: `off` → `api-only` is the cost of the driver's own | ||
| code, which drivers control, and `api-only` → `sdk-ratio` → `sdk-always` adds the cost of the SDK and of recording. If | ||
| `sdk-ratio` costs nearly as much as `sdk-always`, attribute building is not gated on the sampling decision. |
The reference table came from a single local run and no longer matches the current measurements in the linked Ruby implementation PR, so readers could not reproduce it. Drop the figures and point to that PR for the current numbers and the environment they were measured in.
|
|
||
| Therefore, drivers SHOULD: | ||
|
|
||
| - Create a span with a minimum set of attributes that are cheap to create. The recommended set for command spans is |
There was a problem hiding this comment.
Have we benchmarked what performance difference (if any) each attribute adds? Can we create a span with zero attributes to absolutely minimize performance impact and then add all of them once we know the span will be recorded?
There was a problem hiding this comment.
Attributes that are added after span creation are not visible to samplers. OpenTelemetry recommends "Prefer adding attributes at span creation to make the attributes available to SDK sampling."
Good idea about benchmarking every attribute, I'll do it.
There was a problem hiding this comment.
Benchmarked it. All figures below are the Ruby driver with OpenTelemetry SDK 1.13.0, Ruby 4.0.1 + YJIT, on Evergreen; median of 20 reps x 50,000 spans per profile. Each profile is timed against the same span created with an empty attribute hash, the cheapest span the driver can create: ~1.38 µs/span when recording. The relative columns use the untraced DriverBench operations measured on the same host (find ≈ 214 µs, insert ≈ 189 µs, CPU excluding GC).
Each attribute, added at span creation (one attribute at a time):
| Attribute | Type | +µs/span | % of empty span | % of untraced find | % of untraced insert |
|---|---|---|---|---|---|
db.system.name |
string | +0.120 | +8.7% | 0.056% | 0.064% |
db.namespace |
string | +0.117 | +8.5% | 0.055% | 0.062% |
db.command.name |
string | +0.122 | +8.8% | 0.057% | 0.064% |
db.collection.name |
string | +0.120 | +8.7% | 0.056% | 0.063% |
server.address |
string | +0.123 | +8.9% | 0.058% | 0.065% |
network.transport |
string | +0.122 | +8.8% | 0.057% | 0.064% |
db.query.summary |
string | +0.121 | +8.7% | 0.056% | 0.064% |
db.mongodb.lsid |
string | +0.122 | +8.8% | 0.057% | 0.064% |
server.port |
integer | +0.284 | +20.5% | 0.133% | 0.150% |
db.mongodb.server_connection_id |
integer | +0.286 | +20.7% | 0.134% | 0.151% |
db.mongodb.driver_connection_id |
integer | +0.283 | +20.5% | 0.132% | 0.150% |
db.mongodb.cursor_id |
integer | +0.288 | +20.8% | 0.135% | 0.152% |
db.mongodb.txn_number |
integer | +0.287 | +20.8% | 0.134% | 0.152% |
Each attribute costs roughly 0.12 µs/span, about 0.05-0.07% of the untraced operation; none stands out as individually expensive. The integer rows read ~0.28 µs, ~2.4x the strings. That is an artifact of the sweep, not the SDK: timing single-attribute profiles one after another in one process makes the SDK's attribute-validation call sites flip between String and Integer receivers and get recompiled under YJIT. With one profile per process the gap disappears, and it does not appear in a realistic multi-attribute hash: the 9-attribute creation set adds 1.04 µs total, i.e. ~0.12 µs per attribute regardless of type.
Adding an attribute after creation costs ~0.33 µs/span (set_attribute, measured over the 4 deferred attributes), about 2.8x an attribute added at creation.
"Zero attributes at creation, add all of them once we know the span is recorded" is worse for recording spans. Per command span, against the empty-attribute span:
| Span shape | +µs/span | % of untraced find |
|---|---|---|
| all 13 attributes at creation | +1.45 | 0.68% |
current shape (cheap set at creation, rest deferred behind isRecording) |
+2.21 | 1.03% |
zero at creation, all 13 via set_attribute |
+3.81 | 1.78% |
Creating the span empty and adding everything later costs ~1.6 µs/span more when the span is recorded, because set_attribute is the expensive path. It does avoid building attributes when the span is dropped, but that only pays off when nearly all spans are dropped, and it makes every attribute invisible to samplers.
Two things the numbers also show:
- Span creation, not attributes, is the bulk of the cost. A recorded span with no attributes is already ~1.38 µs at the SDK boundary, and the full 13-attribute set is under 1% of the untraced operation.
- On a span the sampler dropped, attributes are essentially free: every profile is within ~0.01 µs of the empty-attribute span. That is the property the "cheap set at creation, rest behind the recording check" structure is built around.
Raw output: Evergreen patch 1996.
There was a problem hiding this comment.
Looks like it's not worth it to delay adding the basic attributes until we know it'll be sampled, thanks for the benchmarking!
| Drivers MAY skip all span work — making the span current, adding attributes, recording the outcome, ending the span — | ||
| when the created span's context is not valid, i.e. its trace id or span id is all zeroes (see | ||
| [IsValid](https://opentelemetry.io/docs/specs/otel/trace/api/#isvalid)): such a span cannot be propagated or correlated, | ||
| so any work on it is wasted. This permission does not apply to a valid but unsampled span, which remains subject to the |
There was a problem hiding this comment.
What does this imply? Can we have an unsampled parent span, but then a child span is sampled and requires the parent to be valid even though it wasn't initially sampled? Or is this only about propagation to the server?
There was a problem hiding this comment.
I raised https://jira.mongodb.org/browse/DRIVERS-3669 to clarify this in the specification.
| Additionally, when a driver first releases OpenTelemetry support, it SHOULD compare its default configuration (normally | ||
| `off`) once against the commit before OpenTelemetry support was added, and record the overhead it measures. | ||
|
|
||
| Drivers SHOULD use their standardized performance testing infrastructure (see |
There was a problem hiding this comment.
Should this be a "MUST"? Why would a driver ever want to not re-use their existing infra?
There was a problem hiding this comment.
I tend to avoid using "MUST" when we telling other teams how to achieve things. Like here, we MUST benchmark, and here is how we recommend to do so. If you find a way that works better in your ecosystem and still achieve the benchmarking goals, it's okay.
However, if you feel that MUST suits here better, I'm totally fine with that.
There was a problem hiding this comment.
That's fair. My main concern is drift between drivers resulting in tangible differences: if one driver uses their existing infrastructure that results in overall better performance for OTel than another driver that uses a bespoke testing setup, which one is more valid? If everyone is required to use the same rough benchmarking setup that we already have, it reduces the risk of such a situation.
| Drivers SHOULD use their standardized performance testing infrastructure (see | ||
| [Performance Benchmarking](../benchmarking/benchmarking.md)) rather than a purpose-built OpenTelemetry benchmark. | ||
|
|
||
| Drivers SHOULD measure the `Small doc insertOne` and `Find one by ID` tasks, through both the synchronous and the |
There was a problem hiding this comment.
Same here, what context would cause a driver to choose to measure different tasks?

DRIVERS-3620
Preliminary benchmarks for DRIVERS-719 showed 15-50% throughput loss with OpenTelemetry enabled. Benchmarking in the Ruby driver located the cost: almost all of it was driver code building span attributes before the sampling decision was known, so dropping 100% of traces recovered almost nothing. After restructuring the Ruby implementation, the non-recording configurations went from 17-26% overhead to 4-8%.
This PR adds a Performance Implications section to the OpenTelemetry spec so other drivers apply the same structure.
Ruby implementation: mongodb/mongo-ruby-driver#3108
Please complete the following before merging:
clusters).