Repository navigation
DRIVERS-3620 Reduce OpenTelemetry tracing overhead, benchmark driver configurations - #3108
Draft
comandeo-mongo wants to merge 25 commits into
Draft
comandeo-mongo wants to merge 25 commits into
comandeo-mongo wants to merge 25 commits into
Conversation
DriverBench only ever measured the default configuration, so no automated benchmark executed the OpenTelemetry code path at all and a large regression in it went unnoticed. Make the driver configuration an axis of the benchmark instead of a property of the environment. A result is now identified by the pair (task, configuration), so every configuration becomes its own time series that can be watched for regressions independently. Comparison runs the tasks under each configuration and records the throughput given up relative to the baseline as a metric of its own, rather than leaving it to be derived by comparing two time series later: recording it directly makes it watchable, and computing it within a run cancels host-to-host variation. Configurations are measured one per subprocess, because OpenTelemetry cannot be reconfigured once its SDK has been installed and the "api-only" configuration requires that the SDK was never loaded. They are run interleaved rather than one to completion and then the next, so that a machine which slows down partway through a run does not put one configuration's samples in the fast half and another's in the slow half and report the difference as overhead. The SDK configurations install no span processor. The cost to attribute to the driver is that of creating and recording spans; a processor would additionally charge the SDK's SpanData conversion, background thread and queue to the driver's account, and add variance that hides small driver-side changes. Sampled spans are still fully recorded without one. Task filters are included because the full suite spends at least a minute per micro-benchmark per configuration, which is too slow for an optimize-and-remeasure loop. Composites are skipped when the task list is filtered, since they average a fixed list of micro-benchmarks.
The comparison is too long to run on a laptop: it runs the tasks once per configuration, several times over, and each micro-benchmark has a 60 second floor. It also has to run undisturbed, since the quantity being measured is a difference between configurations and a machine that sleeps or throttles partway through reports that as overhead. Runs a bounded set of micro-benchmarks rather than the whole suite so that the task fits in a sensible wall clock: the four high-signal single- and multi-doc tasks, which create one operation span and one command span per operation and so show the per-span cost most clearly. Parallel tasks are excluded as disk- and concurrency-bound, and BSON tasks are excluded by Comparison itself since they never reach a server and create no spans. All three of the task list, the configuration list and the repetition count are expansions, so a patch can widen them without a config change. The results are uploaded as a task artifact rather than sent to the performance store. What is wanted here is a baseline to compare later runs against by hand, not a trend series. The task needs its own exec_timeout_secs; the global 5400 is not enough.
…to OperationTracer
3 of 4 tasks
The command span read db.mongodb.server_connection_id from the server description, which comes from the monitoring connection's hello, not from the connection the command ran on. Since connection attributes are now cached per connection, that wrong value was also frozen for the connection's lifetime. Use the connection's own description, as command monitoring already does.
…Tel comparison Throughput alone could not tell tracing cost from server noise: on small-document inserts the loss moved by tens of percent between reps. Each iteration now also records process CPU time and allocated objects, and single-doc tasks report both per operation. Under sdk-always, one extra iteration counts spans per operation with a counting processor, outside the timed runs. The comparison takes the median of every metric across reps, records the spread, the CPU and allocation overhead, and each configuration's target from the OpenTelemetry spec, and can fail the run when a target is missed. The Evergreen task now runs only the two small-document tasks the targets are defined on, with five shorter reps, and sends the results to the performance store as well as uploading them.
…ix perf.send
perf.send rejected the results because args only take integers and the
configuration was passed as a string. The configuration now goes into the
test name ("Find one by ID [api-only]"); the baseline keeps the bare name,
so its series continues the one recorded before configurations existed.
The comparison ran about two hours. It is now one Evergreen task per
micro-benchmark, which run in parallel on separate hosts and still
interleave every configuration on one host, with at most 10 iterations per
run instead of 30.
Both benchmark tasks now run with RUBY_YJIT_ENABLE=1, since without YJIT
the interpreter's per-call cost makes tracing look several times more
expensive than in production. A ruby built without YJIT only warns, so the
suite fails instead of quietly measuring the interpreter.
Under YJIT the time spent in GC varied between processes by more than the whole cost of tracing, and the first YJIT run showed sdk-parent-1pct cheaper than sdk-never. Each iteration now also records GC time where the runtime reports it, and the comparison reports CPU per operation both with and without GC; the summary table shows the latter and GC separately. Also raise the iteration cap from 10 to 20: with ten, the throughput spread across reps was 7-10%, wider than most of what is measured.
…arison With a fixed order, sdk-parent-1pct measured cheaper than sdk-never in every run on both tasks, though it does strictly more work and allocates more. Interleaving cancels drift across a run but not an effect tied to a slot within each rep. Rotating the order by one each rep puts every configuration in every slot exactly once when there are as many reps as configurations, as on Evergreen.
Throughput loss moved by 2-3 points between runs and ranked sdk-never below api-only on inserts, which the work they do rules out. Targets are now checked against the CPU overhead excluding GC, which ranked the configurations correctly. Ten reps instead of five halve the run-to-run spread's influence on the median and keep the order rotation balanced, since every configuration still runs in every slot equally often. The task timeout goes to two hours to fit them.
…chmarks on one host The split tasks run on different hosts, and host speed differs by more than the insert-vs-find overhead difference being investigated. This task runs both micro-benchmarks in one comparison so they can be compared with each other. It is not activated on the waterfall; request it in a patch.
The split tasks ran on different hosts, which differ in speed by more than the overheads being compared: the insert and find results swapped places between runs. One task, driver-bench-otel, now runs both small-document benchmarks on one host, daily. Review fixes in the comparison: the median across reps averages the middle two of an even count instead of always taking the lower one, and the rotation and environment comments match what the code does.
Production deployments of Ruby commonly run with YJIT, and the toolchain now builds MRI 3.2+ with it. Every run script now exports RUBY_YJIT_ENABLE=1 after choosing its ruby, but only when YJIT is actually available: JRuby and older MRI would just print a warning. A value set by the caller, such as RUBY_YJIT_ENABLE=0, is kept.
Reviewers of the OpenTelemetry spec asked whether the cost of each span attribute has been measured, and whether a span could be created with no attributes and given all of them once it is known to be recorded. Add a harness under profile/otel_attributes that answers both. One half measures the SDK's handling of an attribute, with the values fixed as literals: the cost of carrying an attribute through start_span and the cost of setting one after creation. The other half measures the driver's own cost of producing each value, calling the tracer helpers directly with a realistic command document and connection, caches warmed. Every profile runs in one process with the order rotated each repetition, the same control the DriverBench comparison uses, and the sweep runs under a recording sampler and a dropping one because set_attribute discards on a span that is not recording. rake otel_attributes:span and rake otel_attributes:construction run them.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
DRIVERS-3620
OpenTelemetry tracing previously cost up to 26% of throughput on small-document tasks, mostly in work performed before the sampling decision. This PR reduces that overhead and adds configuration benchmarks to catch regressions.
Evergreen results from 2026-10-01 (10 repetitions, medians; throughput loss and added CPU/GC time and allocations vs
off):CPU excludes GC; throughput ranges are min–max across repetitions. The comparison reports 2 spans for every row. Seven of eight CPU targets passed;
sdk-parent-1pctinsertOne exceeded its 10% target by 0.23 percentage points. The benchmark command exited successfully (test_status=0).Related spec: mongodb/specifications#1986