Skip to content

runtime: avoid user cleanup reentrancy in BDWGC callbacks - #2587

Merged
xushiwei merged 2 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/compat-runtime-20260914
Sep 17, 2026
Merged

xushiwei merged 2 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/compat-runtime-20260914

Conversation

@zhouguangyuan0718

@zhouguangyuan0718 zhouguangyuan0718 commented Sep 14, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Queue user AddCleanup callbacks instead of running them synchronously inside BDWGC finalizers.
  • Reuse one persistent cleanup worker, executing callbacks serially outside the queue lock. Hosted LLGo goroutines use OS threads, so there is no per-callback goroutine/thread creation.
  • Keep internal weak-handle invalidation synchronous and preserve Cleanup.Stop cancellation until the worker claims the callback.
  • Add regressions for collection while holding a mutex needed by a cleanup, and serial execution of a cleanup batch.

Root cause and design

client-go/rest can deadlock when a transport-cache cleanup holds the cache mutex and allocates. BDWGC recursively invokes another cleanup, which attempts to acquire the same mutex. The timer thread can also block in cleanup, preventing the test timeout from firing. This remains after the internal weak cleanup fix in #2575.

The intrusive queue is initialized before user cleanup registration and does not allocate while holding its native mutex. The single worker detaches a batch, releases the queue lock, then runs callbacks directly. Callback allocations can enqueue more work but cannot synchronously reenter another user cleanup. A long-running or blocked callback delays subsequent callbacks; the runtime does not create additional threads to compensate.

This change targets hosted BDWGC builds; wasm, baremetal and nogc are unchanged.

Validation

Current single-worker revision, local macOS ARM64:

  • Deadlock regression passes 30 repetitions. The original baseline timed out after 10 seconds; the native Go control passed 30 repetitions.
  • k8s.io/client-go/rest v0.36.2: full-package run passes (Go 1.26.5).
  • test/std/runtime and test/std/weak pass.
  • TestRuntimeCleanupSerialWorker, TestRuntimeAddCleanupStop, TestRuntimeGCDrainsBDWGCFinalizersOnDemand and TestRuntimeConcurrentGCFinalizers pass.
  • Weak cleanup reentrancy stress test passes with LLGO_STRESS_PROFILE=quick.
  • git diff --check passes.

Linux AMD64 CI validation is pending. This PR does not claim to resolve the separate client-go/transport GC retention assertions or semver test portability issue.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: run user cleanups outside BDWGC finalizer callbacks

The core change is sound. Moving cancelable cleanup callbacks off the BDWGC collecting goroutine (via the intrusive cleanupQueue + worker) correctly fixes the deadlock where a finalizer runs synchronously while an application lock is held. I verified the important paths:

  • The fix covers the API the test exercises. runtime.AddCleanup → AddCancelableCleanupPtr → the queued path, so TestCleanupDoesNotBlockCollector genuinely validates the new behavior.
  • The remaining synchronous id==0 path is intentional and safe. Its only caller is weak_llgo.go's retireWeakHandle, which by design neither allocates nor takes a lock (see its comment at weak_llgo.go:50-52) — so leaving it inline in the finalizer is correct.
  • Slot reuse + generation and the lock-free free-list are correct. The generation<<32 | index+1 id defeats stale Stop calls, and slot pops are serialized under cleanupSlots.mu, so no ABA on the free-list head. StopCleanupPtr correctly cancels a queued-but-unclaimed entry via the cleanupActive→cleanupStopped CAS racing the worker's cleanupActive→cleanupRunning CAS.
  • Comments accurately describe the new behavior.

The findings below are design/robustness considerations, not correctness blockers.

Test coverage. The new test covers the deadlock-avoidance scenario well, but the PR adds substantial machinery (queue, worker fan-out, slot reuse, cancel-after-queue) that is otherwise untested. Consider adding: (1) a Stop-after-object-unreachable-but-before-worker-runs test — the exact "cancel a queued cleanup" case the StopCleanupPtr doc now claims to support; (2) a many-cleanups stress test to exercise popCleanupSlot/freeCleanupSlot reuse and concurrent execution.

e.next = nil
// Execute outside the queue lock so allocations made by a callback
// can enqueue more cleanups without reentering user code. Hosted
// goroutines each own an OS thread: reuse this worker rather than

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Unbounded goroutine fan-out under GC pressure. runCleanups starts one go runCleanup(e) per drained entry with no ceiling. A GC cycle that finalizes many objects carrying Cleanup handles will spawn one goroutine per cleanup at once — precisely when scheduler/memory pressure is highest. The comment's goal ("a blocked cleanup must not prevent other cleanups from running") can be met with a bounded worker pool or a semaphore, which would cap concurrent goroutines and avoid paying a goroutine spawn+teardown for each typically-short callback. Worth considering as a follow-up if high cleanup churn is expected.

cleanupQueue.mu.Lock()
e.next = cleanupQueue.head
cleanupQueue.head = e
cleanupQueue.ready.Signal()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor: ready.Signal() fires on every enqueue, including bursts where the worker is already draining, so a burst of finalizers issues one cond-signal syscall per entry on the collecting goroutine's path. Signaling only on the empty→non-empty transition (i.e. when cleanupQueue.head == nil before the push) would coalesce these. Not a correctness issue — the current form is safe.

@codecov

codecov Bot commented Sep 14, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

4552029a9058 | workflow run | long-term charts

WebAssembly output sizes

Profile and compiler Wasm module vs base Generated JS glue vs base
ec32/LLGo 113243 B 0 B / +0.0% 70736 B 0 B / +0.0%
ec64/LLGo 118688 B 0 B / +0.0% 74033 B 0 B / +0.0%
js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
js/LLGo 66660 B 0 B / +0.0% 68511 B 0 B / +0.0%
wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
wasip1/LLGo 72917 B 0 B / +0.0% 0 B 0 B / 0.0%
wc32/LLGo 117359 B 0 B / +0.0% 0 B 0 B / 0.0%

LLGo WebAssembly build measurements

Profile Build vs base
ec32 5.609 s -245.2 ms / -4.2% (better)
ec64 5.700 s +579.7 ms / +11.3% (worse)
js 4.508 s -125.6 ms / -2.7% (better)
wasip1 3.037 s -294.2 ms / -8.8% (better)
wc32 4.377 s +553.9 ms / +14.5% (worse)

Compared with 5c5874359c1e measured in the same runner job.

@github-actions

Copy link
Copy Markdown

LLGo baseline benchmarks

4552029a9058 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 19880 B +48 B / +0.2% (worse) 387 B 0 B / +0.0% 594.634 ms +128.3 ms / +27.5% (worse) 1.382 ms +130 us / +10.4% (worse)
Linux cprintf-lto 19632 B +48 B / +0.2% (worse) 368 B 0 B / +0.0% 594.550 ms +126.8 ms / +27.1% (worse) 1.370 ms +96.14 us / +7.5% (worse)
Linux fmtprintf 1627112 B +1232 B / +0.1% (worse) 498326 B +364 B / +0.1% (worse) 4.494 s +519.9 ms / +13.1% (worse) 4.054 ms +801.9 us / +24.7% (worse)
Linux fmtprintf-lto 1481320 B +1144 B / +0.1% (worse) 438196 B +375 B / +0.1% (worse) 11.469 s +349 ms / +3.1% (worse) 2.947 ms -22.39 us / -0.8% (better)
Linux println 62272 B +48 B / +0.1% (worse) 14901 B 0 B / +0.0% 686.414 ms +224.3 ms / +48.5% (worse) 1.622 ms +76.92 us / +5.0% (worse)
Linux println-lto 54328 B +48 B / +0.1% (worse) 12335 B 0 B / +0.0% 806.678 ms +128.9 ms / +19.0% (worse) 1.629 ms +45.52 us / +2.9% (worse)
macOS cprintf 84480 B 0 B / +0.0% 17149 B +48 B / +0.3% (worse) 572.231 ms -277.4 ms / -32.6% (better) 4.929 ms +100.7 us / +2.1% (worse)
macOS cprintf-lto 84288 B 0 B / +0.0% 12913 B +48 B / +0.4% (worse) 561.476 ms -392.6 ms / -41.1% (better) 3.037 ms -2.772 ms / -47.7% (better)
macOS fmtprintf 1473168 B 0 B / +0.0% 872188 B +176 B / +0.02018% (worse) 2.832 s -701.8 ms / -19.9% (better) 4.641 ms -1.955 ms / -29.6% (better)
macOS fmtprintf-lto 1159424 B 0 B / +0.0% 847956 B +204 B / +0.02406% (worse) 10.795 s +1.566 s / +17.0% (worse) 11.055 ms +872.7 us / +8.6% (worse)
macOS println 114672 B 0 B / +0.0% 34860 B +48 B / +0.1% (worse) 546.595 ms -322.2 ms / -37.1% (better) 3.722 ms -2.864 ms / -43.5% (better)
macOS println-lto 118720 B 0 B / +0.0% 32296 B +48 B / +0.1% (worse) 958.195 ms +73.32 ms / +8.3% (worse) 6.101 ms +1.62 ms / +36.2% (worse)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.176 s -20.84 ms / -1.7% (better) 4.022 ms +284.6 us / +7.6% (worse)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.225 s -10.56 ms / -0.9% (better) 3.449 ms -241.5 us / -6.5% (better)
Windows MinGW fmtprintf 1895424 B 0 B / +0.0% 601110 B 0 B / +0.0% 4.229 s +39.08 ms / +0.9% (worse) 8.186 ms -629.4 us / -7.1% (better)
Windows MinGW fmtprintf-lto 1936896 B 0 B / +0.0% 550566 B 0 B / +0.0% 11.257 s +636.7 ms / +6.0% (worse) 8.505 ms +115.6 us / +1.4% (worse)
Windows MinGW println 71168 B 0 B / +0.0% 24054 B 0 B / +0.0% 1.169 s -5.885 ms / -0.5% (better) 6.856 ms -181.2 us / -2.6% (better)
Windows MinGW println-lto 65536 B 0 B / +0.0% 20678 B 0 B / +0.0% 1.383 s -26.82 ms / -1.9% (better) 6.949 ms -208.3 us / -2.9% (better)
Windows MinGW 386 cprintf 42496 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.187 s -56.19 ms / -4.5% (better) 5.208 ms -1.407 ms / -21.3% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.195 s -77.73 ms / -6.1% (better) 6.950 ms +580.4 us / +9.1% (worse)
Windows MinGW 386 fmtprintf 1858560 B 0 B / +0.0% 472430 B 0 B / +0.0% 4.304 s -231 ms / -5.1% (better) 11.807 ms -178.4 us / -1.5% (better)
Windows MinGW 386 fmtprintf-lto 2165760 B 0 B / +0.0% 452522 B 0 B / +0.0% 9.738 s -827.9 ms / -7.8% (better) 10.295 ms -2.363 ms / -18.7% (better)
Windows MinGW 386 println 91136 B 0 B / +0.0% 20038 B 0 B / +0.0% 1.079 s -156.3 ms / -12.7% (better) 8.750 ms -2.117 ms / -19.5% (better)
Windows MinGW 386 println-lto 69632 B 0 B / +0.0% 17954 B 0 B / +0.0% 1.280 s -154.6 ms / -10.8% (better) 8.504 ms -1.865 ms / -18.0% (better)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.404 s -39.88 ms / -2.8% (better) 6.087 ms -384.9 us / -5.9% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.455 s -17.02 ms / -1.2% (better) 6.220 ms +31.8 us / +0.5% (worse)
Windows MinGW ARM64 fmtprintf 1781248 B +512 B / +0.02875% (worse) 510696 B 0 B / +0.0% 4.218 s +164.4 ms / +4.1% (worse) 13.665 ms +1.363 ms / +11.1% (worse)
Windows MinGW ARM64 fmtprintf-lto 1860096 B 0 B / +0.0% 478908 B -4 B / -0.0008352% (better) 10.065 s +233.9 ms / +2.4% (worse) 13.191 ms +622 us / +4.9% (worse)
Windows MinGW ARM64 println 68608 B 0 B / +0.0% 22544 B 0 B / +0.0% 1.393 s -69.03 ms / -4.7% (better) 10.625 ms -307.7 us / -2.8% (better)
Windows MinGW ARM64 println-lto 64000 B 0 B / +0.0% 19604 B 0 B / +0.0% 1.610 s -15.56 ms / -1.0% (better) 10.949 ms -25.2 us / -0.2% (better)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65782 B 0 B / +0.0% 752.087 ms -22.23 ms / -2.9% (better) 2.694 ms -153.4 us / -5.4% (better)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65718 B 0 B / +0.0% 781.268 ms +45.63 ms / +6.2% (worse) 2.717 ms +338.1 us / +14.2% (worse)
Windows MSVC fmtprintf 1626112 B 0 B / +0.0% 696630 B 0 B / +0.0% 3.056 s +14.8 ms / +0.5% (worse) 6.756 ms -195.2 us / -2.8% (better)
Windows MSVC fmtprintf-lto 1625088 B 0 B / +0.0% 652886 B 0 B / +0.0% 7.210 s -109.7 ms / -1.5% (better) 7.452 ms -190.1 us / -2.5% (better)
Windows MSVC println 193024 B 0 B / +0.0% 119478 B 0 B / +0.0% 738.816 ms +9.484 ms / +1.3% (worse) 5.578 ms -73.8 us / -1.3% (better)
Windows MSVC println-lto 189952 B 0 B / +0.0% 116614 B 0 B / +0.0% 974.708 ms +103.3 ms / +11.9% (worse) 6.162 ms +603.7 us / +10.9% (worse)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 922.350 ms -170.5 ms / -15.6% (better) 5.312 ms -1.867 ms / -26.0% (better)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 955.572 ms -16.79 ms / -1.7% (better) 5.303 ms -962.6 us / -15.4% (better)
Windows MSVC 386 fmtprintf 1190912 B 0 B / +0.0% 455797 B 0 B / +0.0% 3.728 s -190.1 ms / -4.9% (better) 12.082 ms +366.2 us / +3.1% (worse)
Windows MSVC 386 fmtprintf-lto 1231872 B 0 B / +0.0% 431417 B 0 B / +0.0% 8.718 s +2.972 ms / +0.0341% (worse) 12.116 ms -2.036 ms / -14.4% (better)
Windows MSVC 386 println 34304 B 0 B / +0.0% 18865 B 0 B / +0.0% 931.842 ms -14.68 ms / -1.6% (better) 9.021 ms -534.8 us / -5.6% (better)
Windows MSVC 386 println-lto 32768 B 0 B / +0.0% 17015 B 0 B / +0.0% 1.098 s -33.02 ms / -2.9% (better) 9.367 ms -673.1 us / -6.7% (better)
Windows MSVC ARM64 cprintf 11264 B 0 B / +0.0% 3976 B 0 B / +0.0% 2.060 s -8.749 ms / -0.4% (better) 7.590 ms +476.9 us / +6.7% (worse)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 3868 B 0 B / +0.0% 2.050 s +37.59 ms / +1.9% (worse) 7.133 ms +553.9 us / +8.4% (worse)
Windows MSVC ARM64 fmtprintf 1371648 B 0 B / +0.0% 511524 B 0 B / +0.0% 6.642 s +352.4 ms / +5.6% (worse) 14.367 ms -103.2 us / -0.7% (better)
Windows MSVC ARM64 fmtprintf-lto 1395712 B 0 B / +0.0% 481700 B -16 B / -0.003321% (better) 15.849 s +515.9 ms / +3.4% (worse) 14.882 ms +1.144 ms / +8.3% (worse)
Windows MSVC ARM64 println 41472 B 0 B / +0.0% 21784 B 0 B / +0.0% 1.945 s +8.79 ms / +0.5% (worse) 12.960 ms +973.3 us / +8.1% (worse)
Windows MSVC ARM64 println-lto 39936 B 0 B / +0.0% 19532 B 0 B / +0.0% 2.299 s +30.39 ms / +1.3% (worse) 12.373 ms +188.1 us / +1.5% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.680 ns/op +0.17 ns/op / +1.2% (worse)
Linux BenchmarkMergeCompilerFlags 199.200 ns/op +3.4 ns/op / +1.7% (worse)
Linux BenchmarkMergeLinkerFlags 125.700 ns/op -3.1 ns/op / -2.4% (better)
Linux BenchmarkChannelBuffered 55.770 ns/op -0.42 ns/op / -0.7% (better)
Linux BenchmarkChannelHandoff 17354 ns/op -1839 ns/op / -9.6% (better)
Linux BenchmarkDefer 51.330 ns/op +4.68 ns/op / +10.0% (worse)
Linux BenchmarkDirectCall 1.626 ns/op +0.071 ns/op / +4.6% (worse)
Linux BenchmarkGlobalRead 1.177 ns/op 0 ns/op / +0.0%
Linux BenchmarkGlobalWrite 7.851 ns/op +0.041 ns/op / +0.5% (worse)
Linux BenchmarkGoroutine 22317 ns/op -8897 ns/op / -28.5% (better)
Linux BenchmarkInterfaceCall 5.860 ns/op -0.406 ns/op / -6.5% (better)
Linux BenchmarkRuntimeGetG 2.385 ns/op -0.068 ns/op / -2.8% (better)
macOS BenchmarkLookupPCRandom 11.120 ns/op -3.04 ns/op / -21.5% (better)
macOS BenchmarkMergeCompilerFlags 101.700 ns/op -23.5 ns/op / -18.8% (better)
macOS BenchmarkMergeLinkerFlags 56.790 ns/op -23.95 ns/op / -29.7% (better)
macOS BenchmarkChannelBuffered 26.840 ns/op -2.51 ns/op / -8.6% (better)
macOS BenchmarkChannelHandoff 6532 ns/op -2975 ns/op / -31.3% (better)
macOS BenchmarkDefer 37.720 ns/op +5.11 ns/op / +15.7% (worse)
macOS BenchmarkDirectCall 1.087 ns/op +0.028 ns/op / +2.6% (worse)
macOS BenchmarkGlobalRead 1.095 ns/op +0.006 ns/op / +0.6% (worse)
macOS BenchmarkGlobalWrite 1.374 ns/op -0.115 ns/op / -7.7% (better)
macOS BenchmarkGoroutine 31033 ns/op -6879 ns/op / -18.1% (better)
macOS BenchmarkInterfaceCall 3.973 ns/op +0.019 ns/op / +0.5% (worse)
macOS BenchmarkRuntimeGetG 2.188 ns/op +0.024 ns/op / +1.1% (worse)
Windows MinGW BenchmarkLookupPCRandom 12.370 ns/op -0.17 ns/op / -1.4% (better)
Windows MinGW BenchmarkMergeCompilerFlags 545.600 ns/op +11 ns/op / +2.1% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 463 ns/op +1.3 ns/op / +0.3% (worse)
Windows MinGW BenchmarkChannelBuffered 33.630 ns/op -1.82 ns/op / -5.1% (better)
Windows MinGW BenchmarkChannelHandoff 1400 ns/op +6 ns/op / +0.4% (worse)
Windows MinGW BenchmarkDefer 56.080 ns/op +1.75 ns/op / +3.2% (worse)
Windows MinGW BenchmarkDirectCall 1.747 ns/op +0.001 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGlobalRead 1.748 ns/op +0.001 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGlobalWrite 2.787 ns/op +0.001 ns/op / +0.03589% (worse)
Windows MinGW BenchmarkGoroutine 72023 ns/op -932 ns/op / -1.3% (better)
Windows MinGW BenchmarkInterfaceCall 9.093 ns/op +0.355 ns/op / +4.1% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.098 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.640 ns/op -1.05 ns/op / -3.8% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 793.700 ns/op +7.1 ns/op / +0.9% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 719.500 ns/op -14.8 ns/op / -2.0% (better)
Windows MinGW 386 BenchmarkChannelBuffered 41.460 ns/op -0.53 ns/op / -1.3% (better)
Windows MinGW 386 BenchmarkChannelHandoff 1096 ns/op -19 ns/op / -1.7% (better)
Windows MinGW 386 BenchmarkDefer 44.910 ns/op -0.12 ns/op / -0.3% (better)
Windows MinGW 386 BenchmarkDirectCall 1.549 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.860 ns/op -0.003 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkGlobalWrite 7.782 ns/op -0.006 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGoroutine 84155 ns/op -2928 ns/op / -3.4% (better)
Windows MinGW 386 BenchmarkInterfaceCall 8.434 ns/op +0.025 ns/op / +0.3% (worse)
Windows MinGW 386 BenchmarkRuntimeGetG 2.169 ns/op -0.003 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.100 ns/op +0.04 ns/op / +0.3% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 578.100 ns/op +9.1 ns/op / +1.6% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 538.500 ns/op +2.4 ns/op / +0.4% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 39.050 ns/op +0.21 ns/op / +0.5% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 2106 ns/op +128 ns/op / +6.5% (worse)
Windows MinGW ARM64 BenchmarkDefer 53.630 ns/op -1 ns/op / -1.8% (better)
Windows MinGW ARM64 BenchmarkDirectCall 0.589 ns/op -0.0004 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.663 ns/op -0.0004 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.885 ns/op -0.0002 ns/op / -0.0226% (better)
Windows MinGW ARM64 BenchmarkGoroutine 56177 ns/op +4226 ns/op / +8.1% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.305 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.805 ns/op +0.036 ns/op / +2.0% (worse)
Windows MSVC BenchmarkLookupPCRandom 9.677 ns/op -0.049 ns/op / -0.5% (better)
Windows MSVC BenchmarkMergeCompilerFlags 369.800 ns/op -35.4 ns/op / -8.7% (better)
Windows MSVC BenchmarkMergeLinkerFlags 331.700 ns/op -16.6 ns/op / -4.8% (better)
Windows MSVC BenchmarkChannelBuffered 27.180 ns/op +1.07 ns/op / +4.1% (worse)
Windows MSVC BenchmarkChannelHandoff 1028 ns/op -130 ns/op / -11.2% (better)
Windows MSVC BenchmarkDefer 44.620 ns/op +0.02 ns/op / +0.04484% (worse)
Windows MSVC BenchmarkDirectCall 1.356 ns/op 0 ns/op / +0.0%
Windows MSVC BenchmarkGlobalRead 1.370 ns/op -0.258 ns/op / -15.8% (better)
Windows MSVC BenchmarkGlobalWrite 2.166 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC BenchmarkGoroutine 49668 ns/op +62 ns/op / +0.1% (worse)
Windows MSVC BenchmarkInterfaceCall 6.794 ns/op +0.009 ns/op / +0.1% (worse)
Windows MSVC BenchmarkRuntimeGetG 1.629 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkLookupPCRandom 26.520 ns/op +0.05 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 733.900 ns/op -17.5 ns/op / -2.3% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 686.900 ns/op +5.3 ns/op / +0.8% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 42.060 ns/op -5.04 ns/op / -10.7% (better)
Windows MSVC 386 BenchmarkChannelHandoff 981.100 ns/op -12.8 ns/op / -1.3% (better)
Windows MSVC 386 BenchmarkDefer 45 ns/op -2.6 ns/op / -5.5% (better)
Windows MSVC 386 BenchmarkDirectCall 1.546 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGlobalRead 1.549 ns/op -0.003 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkGlobalWrite 7.784 ns/op -0.015 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkGoroutine 88096 ns/op +1874 ns/op / +2.2% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 8.133 ns/op -0.208 ns/op / -2.5% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 1.926 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.060 ns/op +0.05 ns/op / +0.4% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 577.900 ns/op -56.4 ns/op / -8.9% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 539.900 ns/op -32 ns/op / -5.6% (better)
Windows MSVC ARM64 BenchmarkChannelBuffered 39.610 ns/op +0.37 ns/op / +0.9% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 2439 ns/op +313 ns/op / +14.7% (worse)
Windows MSVC ARM64 BenchmarkDefer 63.380 ns/op +1.93 ns/op / +3.1% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.589 ns/op -0.0005 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.666 ns/op +0.0004 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.751 ns/op -0.005 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkGoroutine 55645 ns/op -2333 ns/op / -4.0% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.170 ns/op +0.004 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 2.252 ns/op +0.045 ns/op / +2.0% (worse)

Timer runtime benchmarks

Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 919.200 ns/op +12.8 ns/op / +1.4% (worse)
Linux AfterFuncZeroDelivery/LLGo 31417 ns/op -8447 ns/op / -21.2% (better)
Linux CreateStop/Go 289.300 ns/op 0 ns/op / +0.0%
Linux CreateStop/LLGo 1598 ns/op -164 ns/op / -9.3% (better)
Linux RearmStopped/Go 116.600 ns/op +1.4 ns/op / +1.2% (worse)
Linux RearmStopped/LLGo 1123 ns/op -93 ns/op / -7.6% (better)
Linux ResetActive/Go 68.990 ns/op +1.24 ns/op / +1.8% (worse)
Linux ResetActive/LLGo 721.400 ns/op +27.8 ns/op / +4.0% (worse)
Linux ResetHeap1024/Go 67.080 ns/op -0.24 ns/op / -0.4% (better)
Linux ResetHeap1024/LLGo 177.800 ns/op -8.8 ns/op / -4.7% (better)
macOS AfterFuncZeroDelivery/Go 441.200 ns/op -30.7 ns/op / -6.5% (better)
macOS AfterFuncZeroDelivery/LLGo 77991 ns/op -17095 ns/op / -18.0% (better)
macOS CreateStop/Go 125.300 ns/op -50.6 ns/op / -28.8% (better)
macOS CreateStop/LLGo 546.800 ns/op -439.2 ns/op / -44.5% (better)
macOS RearmStopped/Go 56.170 ns/op -19.65 ns/op / -25.9% (better)
macOS RearmStopped/LLGo 340.800 ns/op -193.7 ns/op / -36.2% (better)
macOS ResetActive/Go 46.870 ns/op -3.18 ns/op / -6.4% (better)
macOS ResetActive/LLGo 166.500 ns/op -32 ns/op / -16.1% (better)
macOS ResetHeap1024/Go 46.160 ns/op -1.9 ns/op / -4.0% (better)
macOS ResetHeap1024/LLGo 80.570 ns/op -16.46 ns/op / -17.0% (better)
Windows MinGW AfterFuncZeroDelivery/Go 486.800 ns/op -5.5 ns/op / -1.1% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 127798 ns/op -1738 ns/op / -1.3% (better)
Windows MinGW CreateStop/Go 117 ns/op +1.2 ns/op / +1.0% (worse)
Windows MinGW CreateStop/LLGo 454.800 ns/op -7 ns/op / -1.5% (better)
Windows MinGW RearmStopped/Go 31.570 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW RearmStopped/LLGo 281.400 ns/op +5.4 ns/op / +2.0% (worse)
Windows MinGW ResetActive/Go 19.070 ns/op -0.29 ns/op / -1.5% (better)
Windows MinGW ResetActive/LLGo 156.300 ns/op -34.5 ns/op / -18.1% (better)
Windows MinGW ResetHeap1024/Go 19.210 ns/op +0.05 ns/op / +0.3% (worse)
Windows MinGW ResetHeap1024/LLGo 136.300 ns/op -2.6 ns/op / -1.9% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 965.400 ns/op -11.8 ns/op / -1.2% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 181078 ns/op -2131 ns/op / -1.2% (better)
Windows MinGW 386 CreateStop/Go 192.500 ns/op -21.4 ns/op / -10.0% (better)
Windows MinGW 386 CreateStop/LLGo 1757 ns/op -108 ns/op / -5.8% (better)
Windows MinGW 386 RearmStopped/Go 63.530 ns/op -2.64 ns/op / -4.0% (better)
Windows MinGW 386 RearmStopped/LLGo 355.400 ns/op +12.5 ns/op / +3.6% (worse)
Windows MinGW 386 ResetActive/Go 39.290 ns/op -1.09 ns/op / -2.7% (better)
Windows MinGW 386 ResetActive/LLGo 942.700 ns/op +43.8 ns/op / +4.9% (worse)
Windows MinGW 386 ResetHeap1024/Go 39.480 ns/op -0.37 ns/op / -0.9% (better)
Windows MinGW 386 ResetHeap1024/LLGo 185.700 ns/op -2.6 ns/op / -1.4% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 660.400 ns/op -7.3 ns/op / -1.1% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 135140 ns/op -9821 ns/op / -6.8% (better)
Windows MinGW ARM64 CreateStop/Go 198.300 ns/op -2.4 ns/op / -1.2% (better)
Windows MinGW ARM64 CreateStop/LLGo 358.200 ns/op -19.3 ns/op / -5.1% (better)
Windows MinGW ARM64 RearmStopped/Go 70.530 ns/op -0.08 ns/op / -0.1% (better)
Windows MinGW ARM64 RearmStopped/LLGo 253.500 ns/op +0.8 ns/op / +0.3% (worse)
Windows MinGW ARM64 ResetActive/Go 31.010 ns/op +0.01 ns/op / +0.03226% (worse)
Windows MinGW ARM64 ResetActive/LLGo 124.900 ns/op -1.4 ns/op / -1.1% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.080 ns/op +0.03 ns/op / +0.1% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 127.500 ns/op -0.4 ns/op / -0.3% (better)
Windows MSVC AfterFuncZeroDelivery/Go 372 ns/op +0.8 ns/op / +0.2% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 96924 ns/op -6253 ns/op / -6.1% (better)
Windows MSVC CreateStop/Go 90.370 ns/op -0.56 ns/op / -0.6% (better)
Windows MSVC CreateStop/LLGo 346.700 ns/op -2.8 ns/op / -0.8% (better)
Windows MSVC RearmStopped/Go 24.430 ns/op -0.05 ns/op / -0.2% (better)
Windows MSVC RearmStopped/LLGo 222.200 ns/op -20.4 ns/op / -8.4% (better)
Windows MSVC ResetActive/Go 14.780 ns/op +0.02 ns/op / +0.1% (worse)
Windows MSVC ResetActive/LLGo 117.700 ns/op -0.5 ns/op / -0.4% (better)
Windows MSVC ResetHeap1024/Go 14.760 ns/op -0.02 ns/op / -0.1% (better)
Windows MSVC ResetHeap1024/LLGo 105.400 ns/op -0.3 ns/op / -0.3% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 960.100 ns/op +6 ns/op / +0.6% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 189768 ns/op -887 ns/op / -0.5% (better)
Windows MSVC 386 CreateStop/Go 191.500 ns/op -0.2 ns/op / -0.1% (better)
Windows MSVC 386 CreateStop/LLGo 1649 ns/op -152 ns/op / -8.4% (better)
Windows MSVC 386 RearmStopped/Go 63.230 ns/op -0.08 ns/op / -0.1% (better)
Windows MSVC 386 RearmStopped/LLGo 336.900 ns/op +13.9 ns/op / +4.3% (worse)
Windows MSVC 386 ResetActive/Go 39.050 ns/op -0.17 ns/op / -0.4% (better)
Windows MSVC 386 ResetActive/LLGo 987.600 ns/op +2.6 ns/op / +0.3% (worse)
Windows MSVC 386 ResetHeap1024/Go 39.490 ns/op +0.27 ns/op / +0.7% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 172.800 ns/op +1.5 ns/op / +0.9% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 679.700 ns/op +9.9 ns/op / +1.5% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 144530 ns/op +18502 ns/op / +14.7% (worse)
Windows MSVC ARM64 CreateStop/Go 200.400 ns/op -0.1 ns/op / -0.04988% (better)
Windows MSVC ARM64 CreateStop/LLGo 394.500 ns/op +1.5 ns/op / +0.4% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.590 ns/op -0.02 ns/op / -0.02832% (better)
Windows MSVC ARM64 RearmStopped/LLGo 273.700 ns/op -1 ns/op / -0.4% (better)
Windows MSVC ARM64 ResetActive/Go 30.790 ns/op -0.2 ns/op / -0.6% (better)
Windows MSVC ARM64 ResetActive/LLGo 131.900 ns/op +3.7 ns/op / +2.9% (worse)
Windows MSVC ARM64 ResetHeap1024/Go 31.080 ns/op +0.02 ns/op / +0.1% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 137.700 ns/op -1.6 ns/op / -1.1% (better)

Compared with 5c5874359c1e measured in the same runner job.

@xushiwei
xushiwei merged commit 9754758 into xgo-dev:main Sep 17, 2026
65 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants