Skip to content

fix(build): fingerprint injected Go source patch bodies - #2527

Merged
xushiwei merged 4 commits into
xgo-dev:mainfrom
cpunion:codex/fix-source-patch-cache-20260908
Sep 8, 2026
Merged

xushiwei merged 4 commits into
xgo-dev:mainfrom
cpunion:codex/fix-source-patch-cache-20260908

Conversation

@cpunion

@cpunion cpunion commented Sep 8, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

Injected Go source patches are added to CompiledGoFiles after go list, but package cache fingerprints previously considered only GoFiles. Changing an injected function body could silently reuse a stale archive. This common build-cache defect was found while validating a wasm syscall timestamp fix.

Fix

  • Include selected overlaid CompiledGoFiles in Go source inputs without counting original files twice.
  • Apply the same rule to alternate packages.
  • Keep unselected overlays out of the cache key.
  • Add a regression that changes only an injected patch body and verifies cache invalidation.

This is independent of the WASM R4 runtime changes; it changes only build fingerprinting and its tests.

Validation

  • Focused build/fingerprint tests pass on LLVM 22 against upstream main be23e488a.
  • The cache-enabled wide-stat regression in R4 passes with this fix.
  • Fork PR cpunion/llgo#242 completed 62 successful CI checks and one skipped check at b1082dd3590a25c86efd4606cfd91da408713d20, with no failed or pending checks before this upstream contribution.
  • Upstream's subsequent float-conversion changes through 3dce98b91 do not overlap this patch.

Why earlier CI did not expose this

A clean build is unaffected: the injected source is compiled correctly on the first cache miss. The stale result requires two builds sharing one package cache where only the selected injected patch body changes. Earlier clean-checkout and ordinary cache-hit lanes did not exercise that transition. The new regression performs that exact second build, so prior green clean-build results are not evidence against this fix.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: source patch cache fix

The goal — folding source-patch bodies into the package fingerprint so a patch-body edit invalidates the build cache — is the right fix. But as written the change does not take effect in production because of a path-key mismatch, and the added test masks it. Details inline.

Summary of findings

  • [P0] packageGoSourceInputs matches overlay[file] against the _patch/ paths in CompiledGoFiles, but the overlay is keyed by z_llgo_patch_* paths — the lookup never succeeds for real builds, so patch bodies still don't enter the fingerprint.
  • [P1] The new test uses identical synthetic paths for CompiledGoFiles and the overlay key, so it passes without exercising the real path relationship — giving false confidence.
  • [P3] The doc comment overstates that "source patches ... are not present in GoFiles"; only appended patch files fit that description.

Comment thread internal/build/collect_test.go Outdated
Comment on lines +43 to +46
pkg := &aPackage{Package: &packages.Package{
ID: "example.test", PkgPath: "example.test",
GoFiles: []string{original}, CompiledGoFiles: []string{original, patch},
}}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Test uses identical synthetic paths, masking the P0 mismatch

This test constructs CompiledGoFiles: []string{original, patch} and keys the overlay by the same patch path, so overlay[file] succeeds and the test passes. In production these are distinct paths (a _patch/ source path in CompiledGoFiles vs. a z_llgo_patch_* overlay key), so the guarded branch never fires. As written the test asserts the intended property on an unrealistic path relationship and therefore does not catch the P0 defect.

Recommend modeling production: set the CompiledGoFiles entry to the _patch/ path and key the overlay by the injected z_llgo_patch_* path, so the test reproduces the real mismatch and fails until the helper is corrected.

@cpunion cpunion Sep 8, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in f2387a8. The regression now models distinct production paths: _patch/llgo_patch.go is the selected source file, while z_llgo_patch_llgo_patch.go is the injected overlay key. It mutates only the selected source body and still verifies that an unselected overlay does not change the manifest.

Comment thread internal/build/collect.go Outdated
Comment on lines +194 to +195
// Source patches are appended to CompiledGoFiles after go list; they are not
// present in GoFiles. Their overlaid bodies must participate in cache keys too.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] Comment overstates that source patches are absent from GoFiles

The comment says source patches "are not present in GoFiles." That holds only for the appended patch files (z_llgo_patch_*). The source-patch machinery also overlays files that are in GoFiles under their original names (asm stub replacement, packageStubSource skip-all, filterSourcePatchFile) — those bodies are captured via the slices.Clone(p.GoFiles) + digestFilesWithOverlay path, not this loop. Consider narrowing the wording to "appended source-patch files" to avoid implying this loop covers all patch kinds.

@cpunion cpunion Sep 8, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated in f2387a8. The comment now says specifically that appended source-patch files are absent from GoFiles; it no longer describes all source-patch forms that way.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review addendum: P0 detail

The most important finding (P0) could not attach inline in the prior review event due to a title-length limit. Re-posting it here inline. See the prior review for the P1 (test) and P3 (comment) findings.

Comment thread internal/build/collect.go Outdated
Comment on lines +198 to +202
for _, file := range p.CompiledGoFiles {
if _, ok := overlay[file]; ok && !slices.Contains(files, file) {
files = append(files, file)
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P0] Overlay key mismatch: patch bodies excluded from fingerprint

Overlay key mismatch: patch bodies never enter the fingerprint in production.

In production the appended CompiledGoFiles entries and the overlay keys are different paths, so overlay[file] never matches and no patch file is added here.

  • The slice appended to CompiledGoFiles (load.go:345) comes from applySourcePatchForPkg, which appends src.filename — the _patch/ source path filepath.Join(patchDir, name) (source_patch.go:132, :261).
  • The overlay is keyed by the injected GOROOT target filepath.Join(srcDir, "z_llgo_patch_"+name) (source_patch.go:258). There is no out[src.filename] entry.

Since a _patch/ path is never a key in the overlay, the if _, ok := overlay[file]; ok guard is false for every real source-patch file, so it is never appended and its body never reaches the fingerprint — the same stale-cache defect this PR intends to fix: editing a registered source-patch body will not invalidate the cached archive for the affected stdlib package.

Suggested fix: fingerprint by the content actually keyed in the overlay — e.g. match/append the injected z_llgo_patch_* target path (or digest the _patch/ file recorded in CompiledGoFiles), and update the test to use the real distinct paths so it reproduces the mismatch.

@cpunion cpunion Sep 8, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in f2387a8. Fingerprinting now consumes the exact selected source-patch file list from context.patchFiles[p.PkgPath], so it no longer relies on matching _patch/... CompiledGoFiles paths against z_llgo_patch_... overlay keys. The selected source body is digested directly.

@codecov

codecov Bot commented Sep 8, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@cpunion
cpunion force-pushed the codex/fix-source-patch-cache-20260908 branch from b1082dd to f2387a8 Compare September 8, 2026 01:22
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown

LLGo baseline benchmarks

484c74c24635 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 19848 B 0 B / +0.0% 387 B 0 B / +0.0% 895.673 ms +5.525 ms / +0.6% (worse) 2.064 ms -187.8 us / -8.3% (better)
Linux cprintf-lto 19600 B 0 B / +0.0% 368 B 0 B / +0.0% 905.856 ms -5.582 ms / -0.6% (better) 2.123 ms -382.1 us / -15.3% (better)
Linux fmtprintf 1617304 B 0 B / +0.0% 492562 B 0 B / +0.0% 6.869 s -47.77 ms / -0.7% (better) 5.567 ms -744.1 us / -11.8% (better)
Linux fmtprintf-lto 1467792 B 0 B / +0.0% 434301 B 0 B / +0.0% 18.128 s -257.1 ms / -1.4% (better) 5.701 ms +66.01 us / +1.2% (worse)
Linux println 62752 B 0 B / +0.0% 15039 B 0 B / +0.0% 890.787 ms +11.71 ms / +1.3% (worse) 2.699 ms -97.06 us / -3.5% (better)
Linux println-lto 54312 B 0 B / +0.0% 12431 B 0 B / +0.0% 1.248 s -5.219 ms / -0.4% (better) 2.916 ms -88.62 us / -2.9% (better)
macOS cprintf 84480 B 0 B / +0.0% 17117 B 0 B / +0.0% 1.114 s +332.6 ms / +42.6% (worse) 5.565 ms -602.5 us / -9.8% (better)
macOS cprintf-lto 84288 B 0 B / +0.0% 12881 B 0 B / +0.0% 970.814 ms +244.1 ms / +33.6% (worse) 5.793 ms +2.248 ms / +63.4% (worse)
macOS fmtprintf 1473744 B 0 B / +0.0% 867044 B 0 B / +0.0% 3.470 s -23.79 ms / -0.7% (better) 6.797 ms +1.925 ms / +39.5% (worse)
macOS fmtprintf-lto 1159552 B 0 B / +0.0% 840436 B 0 B / +0.0% 8.501 s +229.9 ms / +2.8% (worse) 5.476 ms +1.331 ms / +32.1% (worse)
macOS println 114864 B 0 B / +0.0% 35069 B 0 B / +0.0% 829.686 ms +238.9 ms / +40.4% (worse) 5.806 ms +1.41 ms / +32.1% (worse)
macOS println-lto 118736 B 0 B / +0.0% 32729 B 0 B / +0.0% 933.357 ms +5.803 ms / +0.6% (worse) 4.304 ms +495 us / +13.0% (worse)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.078 s -41.5 ms / -3.7% (better) 3.404 ms -455.3 us / -11.8% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.127 s -4.431 ms / -0.4% (better) 3.519 ms -362.7 us / -9.3% (better)
Windows MinGW fmtprintf 1886720 B 0 B / +0.0% 593350 B 0 B / +0.0% 3.988 s +48.05 ms / +1.2% (worse) 8.128 ms -318.2 us / -3.8% (better)
Windows MinGW fmtprintf-lto 1927680 B 0 B / +0.0% 541606 B 0 B / +0.0% 9.655 s +6.114 ms / +0.1% (worse) 9.377 ms +1.009 ms / +12.1% (worse)
Windows MinGW println 71680 B 0 B / +0.0% 23974 B 0 B / +0.0% 1.089 s -18.84 ms / -1.7% (better) 6.437 ms -1.017 ms / -13.6% (better)
Windows MinGW println-lto 65536 B 0 B / +0.0% 20726 B 0 B / +0.0% 1.298 s -7.428 ms / -0.6% (better) 6.603 ms -638.8 us / -8.8% (better)
Windows MinGW 386 cprintf 37888 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.198 s -27.95 ms / -2.3% (better) 5.144 ms -762.5 us / -12.9% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.225 s -23.59 ms / -1.9% (better) 5.621 ms -445.4 us / -7.3% (better)
Windows MinGW 386 fmtprintf 1843712 B 0 B / +0.0% 467434 B 0 B / +0.0% 4.444 s +86.88 ms / +2.0% (worse) 11.965 ms -593.1 us / -4.7% (better)
Windows MinGW 386 fmtprintf-lto 2163712 B 0 B / +0.0% 447162 B 0 B / +0.0% 10.425 s -71.31 ms / -0.7% (better) 12.589 ms +237.6 us / +1.9% (worse)
Windows MinGW 386 println 86528 B 0 B / +0.0% 20374 B 0 B / +0.0% 1.179 s -31.74 ms / -2.6% (better) 9.444 ms -459.8 us / -4.6% (better)
Windows MinGW 386 println-lto 70656 B 0 B / +0.0% 18262 B 0 B / +0.0% 1.414 s -35.2 ms / -2.4% (better) 9.771 ms -772.3 us / -7.3% (better)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4436 B 0 B / +0.0% 1.528 s +120.2 ms / +8.5% (worse) 7.695 ms +1.319 ms / +20.7% (worse)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4368 B 0 B / +0.0% 1.707 s +241.2 ms / +16.5% (worse) 8.394 ms +1.846 ms / +28.2% (worse)
Windows MinGW ARM64 fmtprintf 1775104 B 0 B / +0.0% 506200 B 0 B / +0.0% 4.507 s +248.8 ms / +5.8% (worse) 14.401 ms +1.188 ms / +9.0% (worse)
Windows MinGW ARM64 fmtprintf-lto 1855488 B 0 B / +0.0% 473204 B 0 B / +0.0% 10.391 s +377.5 ms / +3.8% (worse) 14.435 ms -563.4 us / -3.8% (better)
Windows MinGW ARM64 println 68608 B 0 B / +0.0% 22880 B 0 B / +0.0% 1.448 s +5.49 ms / +0.4% (worse) 10.393 ms -1.285 ms / -11.0% (better)
Windows MinGW ARM64 println-lto 65024 B 0 B / +0.0% 20160 B 0 B / +0.0% 1.744 s +78.93 ms / +4.7% (worse) 11.158 ms -671.1 us / -5.7% (better)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65782 B 0 B / +0.0% 1.007 s +118.9 ms / +13.4% (worse) 3.508 ms +150.9 us / +4.5% (worse)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65718 B 0 B / +0.0% 902.893 ms -7.111 ms / -0.8% (better) 3.475 ms +37.5 us / +1.1% (worse)
Windows MSVC fmtprintf 1616896 B 0 B / +0.0% 688870 B 0 B / +0.0% 3.733 s +116.2 ms / +3.2% (worse) 9.079 ms +143 us / +1.6% (worse)
Windows MSVC fmtprintf-lto 1615872 B 0 B / +0.0% 644326 B 0 B / +0.0% 8.953 s +132.5 ms / +1.5% (worse) 9.896 ms +1.113 ms / +12.7% (worse)
Windows MSVC println 193024 B 0 B / +0.0% 119382 B 0 B / +0.0% 877.374 ms -7.272 ms / -0.8% (better) 7.078 ms +169.8 us / +2.5% (worse)
Windows MSVC println-lto 189952 B 0 B / +0.0% 116678 B 0 B / +0.0% 1.068 s -1.452 ms / -0.1% (better) 6.967 ms -32.2 us / -0.5% (better)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 821.027 ms +62.82 ms / +8.3% (worse) 5.043 ms +588.7 us / +13.2% (worse)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 723.777 ms -60 ms / -7.7% (better) 4.374 ms -226.9 us / -4.9% (better)
Windows MSVC 386 fmtprintf 1186304 B 0 B / +0.0% 450801 B 0 B / +0.0% 2.966 s +21.37 ms / +0.7% (worse) 10.272 ms +1.134 ms / +12.4% (worse)
Windows MSVC 386 fmtprintf-lto 1228288 B 0 B / +0.0% 426629 B 0 B / +0.0% 6.838 s -24.78 ms / -0.4% (better) 9.139 ms +302.4 us / +3.4% (worse)
Windows MSVC 386 println 34816 B 0 B / +0.0% 19233 B 0 B / +0.0% 714.156 ms -170.5 ms / -19.3% (better) 7.734 ms +524.2 us / +7.3% (worse)
Windows MSVC 386 println-lto 32768 B 0 B / +0.0% 17351 B 0 B / +0.0% 844.425 ms -9.906 ms / -1.2% (better) 7.303 ms +131.9 us / +1.8% (worse)
Windows MSVC ARM64 cprintf 11264 B 0 B / +0.0% 3976 B 0 B / +0.0% 1.994 s +13.71 ms / +0.7% (worse) 6.330 ms -293 us / -4.4% (better)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 3868 B 0 B / +0.0% 1.999 s +30.12 ms / +1.5% (worse) 7.081 ms +528.5 us / +8.1% (worse)
Windows MSVC ARM64 fmtprintf 1363456 B 0 B / +0.0% 505740 B 0 B / +0.0% 6.293 s -51.34 ms / -0.8% (better) 13.549 ms -1.341 ms / -9.0% (better)
Windows MSVC ARM64 fmtprintf-lto 1386496 B 0 B / +0.0% 473916 B 0 B / +0.0% 15.347 s -491.1 ms / -3.1% (better) 13.945 ms -951.8 us / -6.4% (better)
Windows MSVC ARM64 println 41984 B 0 B / +0.0% 22216 B 0 B / +0.0% 1.924 s +23.36 ms / +1.2% (worse) 11.230 ms +44.7 us / +0.4% (worse)
Windows MSVC ARM64 println-lto 40448 B 0 B / +0.0% 20044 B 0 B / +0.0% 2.263 s +50.47 ms / +2.3% (worse) 11.170 ms -134.1 us / -1.2% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 42.690 ns/op -0.16 ns/op / -0.4% (better)
Linux BenchmarkMergeCompilerFlags 472.100 ns/op -27.3 ns/op / -5.5% (better)
Linux BenchmarkMergeLinkerFlags 322.300 ns/op -12.1 ns/op / -3.6% (better)
Linux BenchmarkChannelBuffered 162.400 ns/op -2.1 ns/op / -1.3% (better)
Linux BenchmarkChannelHandoff 27178 ns/op +52 ns/op / +0.2% (worse)
Linux BenchmarkDefer 151.600 ns/op -7.8 ns/op / -4.9% (better)
Linux BenchmarkDirectCall 4.295 ns/op +0.264 ns/op / +6.5% (worse)
Linux BenchmarkGlobalRead 5.122 ns/op -0.046 ns/op / -0.9% (better)
Linux BenchmarkGlobalWrite 6.950 ns/op -0.021 ns/op / -0.3% (better)
Linux BenchmarkGoroutine 40564 ns/op +834 ns/op / +2.1% (worse)
Linux BenchmarkInterfaceCall 31 ns/op +0.18 ns/op / +0.6% (worse)
Linux BenchmarkRuntimeGetG 5.875 ns/op +0.308 ns/op / +5.5% (worse)
macOS BenchmarkLookupPCRandom 15.060 ns/op -2.07 ns/op / -12.1% (better)
macOS BenchmarkMergeCompilerFlags 160.300 ns/op +25.1 ns/op / +18.6% (worse)
macOS BenchmarkMergeLinkerFlags 83.880 ns/op +9.03 ns/op / +12.1% (worse)
macOS BenchmarkChannelBuffered 30.310 ns/op +0.09 ns/op / +0.3% (worse)
macOS BenchmarkChannelHandoff 10174 ns/op +487 ns/op / +5.0% (worse)
macOS BenchmarkDefer 37.040 ns/op -9.44 ns/op / -20.3% (better)
macOS BenchmarkDirectCall 1.071 ns/op -0.192 ns/op / -15.2% (better)
macOS BenchmarkGlobalRead 1.048 ns/op -0.562 ns/op / -34.9% (better)
macOS BenchmarkGlobalWrite 1.057 ns/op -0.585 ns/op / -35.6% (better)
macOS BenchmarkGoroutine 30649 ns/op -25508 ns/op / -45.4% (better)
macOS BenchmarkInterfaceCall 4.900 ns/op -0.946 ns/op / -16.2% (better)
macOS BenchmarkRuntimeGetG 2.154 ns/op -0.096 ns/op / -4.3% (better)
Windows MinGW BenchmarkLookupPCRandom 13.080 ns/op -0.09 ns/op / -0.7% (better)
Windows MinGW BenchmarkMergeCompilerFlags 632.800 ns/op +25.3 ns/op / +4.2% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 536.400 ns/op +6.8 ns/op / +1.3% (worse)
Windows MinGW BenchmarkChannelBuffered 36.900 ns/op -0.25 ns/op / -0.7% (better)
Windows MinGW BenchmarkChannelHandoff 989.200 ns/op -17.8 ns/op / -1.8% (better)
Windows MinGW BenchmarkDefer 56.690 ns/op -0.35 ns/op / -0.6% (better)
Windows MinGW BenchmarkDirectCall 1.546 ns/op -0.004 ns/op / -0.3% (better)
Windows MinGW BenchmarkGlobalRead 1.857 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW BenchmarkGlobalWrite 2.458 ns/op -0.005 ns/op / -0.2% (better)
Windows MinGW BenchmarkGoroutine 91716 ns/op +11329 ns/op / +14.1% (worse)
Windows MinGW BenchmarkInterfaceCall 8.680 ns/op -0.002 ns/op / -0.02304% (better)
Windows MinGW BenchmarkRuntimeGetG 2.170 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkLookupPCRandom 26.610 ns/op +0.04 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkMergeCompilerFlags 764.600 ns/op +5.2 ns/op / +0.7% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 705 ns/op -8.1 ns/op / -1.1% (better)
Windows MinGW 386 BenchmarkChannelBuffered 42.320 ns/op -0.1 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkChannelHandoff 1036 ns/op -1 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkDefer 44.450 ns/op +0.65 ns/op / +1.5% (worse)
Windows MinGW 386 BenchmarkDirectCall 1.856 ns/op -0.004 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.869 ns/op +0.01 ns/op / +0.5% (worse)
Windows MinGW 386 BenchmarkGlobalWrite 7.768 ns/op -0.01 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGoroutine 89093 ns/op -1352 ns/op / -1.5% (better)
Windows MinGW 386 BenchmarkInterfaceCall 9.605 ns/op -0.017 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 2.174 ns/op +0.006 ns/op / +0.3% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.080 ns/op +0.05 ns/op / +0.4% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 573.400 ns/op +7 ns/op / +1.2% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 548.800 ns/op +18.6 ns/op / +3.5% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 45.850 ns/op +2.27 ns/op / +5.2% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 1696 ns/op -168 ns/op / -9.0% (better)
Windows MinGW ARM64 BenchmarkDefer 60.840 ns/op +3.33 ns/op / +5.8% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.589 ns/op -0.0009 ns/op / -0.2% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.663 ns/op +0.0001 ns/op / +0.01507% (worse)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.737 ns/op -0.0003 ns/op / -0.04071% (better)
Windows MinGW ARM64 BenchmarkGoroutine 58284 ns/op -3814 ns/op / -6.1% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.720 ns/op +0.002 ns/op / +0.04239% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.768 ns/op -0.044 ns/op / -2.4% (better)
Windows MSVC BenchmarkLookupPCRandom 12.590 ns/op +0.21 ns/op / +1.7% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 560.100 ns/op +7.6 ns/op / +1.4% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 453.300 ns/op -13.2 ns/op / -2.8% (better)
Windows MSVC BenchmarkChannelBuffered 39.710 ns/op -0.21 ns/op / -0.5% (better)
Windows MSVC BenchmarkChannelHandoff 1320 ns/op -45 ns/op / -3.3% (better)
Windows MSVC BenchmarkDefer 54.940 ns/op -0.52 ns/op / -0.9% (better)
Windows MSVC BenchmarkDirectCall 1.751 ns/op +0.005 ns/op / +0.3% (worse)
Windows MSVC BenchmarkGlobalRead 1.752 ns/op -0.018 ns/op / -1.0% (better)
Windows MSVC BenchmarkGlobalWrite 2.790 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC BenchmarkGoroutine 64711 ns/op +120 ns/op / +0.2% (worse)
Windows MSVC BenchmarkInterfaceCall 10.140 ns/op 0 ns/op / +0.0%
Windows MSVC BenchmarkRuntimeGetG 2.245 ns/op +0.148 ns/op / +7.1% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 21.520 ns/op -0.01 ns/op / -0.04645% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 563.500 ns/op +4 ns/op / +0.7% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 519.300 ns/op -0.9 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkChannelBuffered 36.400 ns/op -0.06 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkChannelHandoff 607.900 ns/op -0.9 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkDefer 38.300 ns/op -1.85 ns/op / -4.6% (better)
Windows MSVC 386 BenchmarkDirectCall 1.356 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGlobalRead 1.356 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGlobalWrite 6.983 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkGoroutine 57603 ns/op +1039 ns/op / +1.8% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 8.414 ns/op -0.004 ns/op / -0.04752% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 1.631 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.090 ns/op +0.01 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 563 ns/op -13 ns/op / -2.3% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 544.700 ns/op -42.3 ns/op / -7.2% (better)
Windows MSVC ARM64 BenchmarkChannelBuffered 43.460 ns/op -3.11 ns/op / -6.7% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 1892 ns/op +26 ns/op / +1.4% (worse)
Windows MSVC ARM64 BenchmarkDefer 62.140 ns/op +1.59 ns/op / +2.6% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op +0.0001 ns/op / +0.01697% (worse)
Windows MSVC ARM64 BenchmarkGlobalRead 0.663 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkGlobalWrite 3.756 ns/op +0.01 ns/op / +0.3% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 55380 ns/op +1175 ns/op / +2.2% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.717 ns/op -0.002 ns/op / -0.04238% (better)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.804 ns/op +0.034 ns/op / +1.9% (worse)

Timer runtime benchmarks

Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 1536 ns/op -101 ns/op / -6.2% (better)
Linux AfterFuncZeroDelivery/LLGo 59877 ns/op +2150 ns/op / +3.7% (worse)
Linux CreateStop/Go 386.300 ns/op -16.1 ns/op / -4.0% (better)
Linux CreateStop/LLGo 3412 ns/op +176 ns/op / +5.4% (worse)
Linux RearmStopped/Go 138.900 ns/op -0.5 ns/op / -0.4% (better)
Linux RearmStopped/LLGo 3262 ns/op +621 ns/op / +23.5% (worse)
Linux ResetActive/Go 97.480 ns/op -0.13 ns/op / -0.1% (better)
Linux ResetActive/LLGo 2026 ns/op +692 ns/op / +51.9% (worse)
Linux ResetHeap1024/Go 96.890 ns/op +0.24 ns/op / +0.2% (worse)
Linux ResetHeap1024/LLGo 435.800 ns/op -2 ns/op / -0.5% (better)
macOS AfterFuncZeroDelivery/Go 525.500 ns/op -29.2 ns/op / -5.3% (better)
macOS AfterFuncZeroDelivery/LLGo 75700 ns/op -24158 ns/op / -24.2% (better)
macOS CreateStop/Go 165.600 ns/op -49.9 ns/op / -23.2% (better)
macOS CreateStop/LLGo 507.300 ns/op -544.7 ns/op / -51.8% (better)
macOS RearmStopped/Go 61.380 ns/op -9.7 ns/op / -13.6% (better)
macOS RearmStopped/LLGo 583.300 ns/op +225.3 ns/op / +62.9% (worse)
macOS ResetActive/Go 49.340 ns/op -5.77 ns/op / -10.5% (better)
macOS ResetActive/LLGo 172.100 ns/op -6.1 ns/op / -3.4% (better)
macOS ResetHeap1024/Go 48.810 ns/op -5.27 ns/op / -9.7% (better)
macOS ResetHeap1024/LLGo 100 ns/op +5.24 ns/op / +5.5% (worse)
Windows MinGW AfterFuncZeroDelivery/Go 552.900 ns/op -32.4 ns/op / -5.5% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 164675 ns/op +1105 ns/op / +0.7% (worse)
Windows MinGW CreateStop/Go 116.200 ns/op +0.3 ns/op / +0.3% (worse)
Windows MinGW CreateStop/LLGo 498.200 ns/op -23.7 ns/op / -4.5% (better)
Windows MinGW RearmStopped/Go 31.530 ns/op +0.22 ns/op / +0.7% (worse)
Windows MinGW RearmStopped/LLGo 301.900 ns/op -25.7 ns/op / -7.8% (better)
Windows MinGW ResetActive/Go 20.070 ns/op +0.01 ns/op / +0.04985% (worse)
Windows MinGW ResetActive/LLGo 150.700 ns/op -20.6 ns/op / -12.0% (better)
Windows MinGW ResetHeap1024/Go 20.370 ns/op -0.04 ns/op / -0.2% (better)
Windows MinGW ResetHeap1024/LLGo 142 ns/op -1.4 ns/op / -1.0% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 964.900 ns/op -10.5 ns/op / -1.1% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 190721 ns/op +2228 ns/op / +1.2% (worse)
Windows MinGW 386 CreateStop/Go 205.600 ns/op +9.6 ns/op / +4.9% (worse)
Windows MinGW 386 CreateStop/LLGo 2156 ns/op -60 ns/op / -2.7% (better)
Windows MinGW 386 RearmStopped/Go 63.480 ns/op +0.17 ns/op / +0.3% (worse)
Windows MinGW 386 RearmStopped/LLGo 368.500 ns/op -11.1 ns/op / -2.9% (better)
Windows MinGW 386 ResetActive/Go 39.180 ns/op +0.01 ns/op / +0.02553% (worse)
Windows MinGW 386 ResetActive/LLGo 840.300 ns/op -158.3 ns/op / -15.9% (better)
Windows MinGW 386 ResetHeap1024/Go 39.460 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW 386 ResetHeap1024/LLGo 197.500 ns/op +2.3 ns/op / +1.2% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 667.800 ns/op +0.2 ns/op / +0.02996% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 133610 ns/op -984 ns/op / -0.7% (better)
Windows MinGW ARM64 CreateStop/Go 195.800 ns/op +3.1 ns/op / +1.6% (worse)
Windows MinGW ARM64 CreateStop/LLGo 442.400 ns/op -46.3 ns/op / -9.5% (better)
Windows MinGW ARM64 RearmStopped/Go 70.540 ns/op -0.06 ns/op / -0.1% (better)
Windows MinGW ARM64 RearmStopped/LLGo 304.600 ns/op +0.9 ns/op / +0.3% (worse)
Windows MinGW ARM64 ResetActive/Go 31.010 ns/op -0.07 ns/op / -0.2% (better)
Windows MinGW ARM64 ResetActive/LLGo 146.800 ns/op +2.4 ns/op / +1.7% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 30.960 ns/op -4.99 ns/op / -13.9% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 139.100 ns/op -0.6 ns/op / -0.4% (better)
Windows MSVC AfterFuncZeroDelivery/Go 491.400 ns/op -6.6 ns/op / -1.3% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 126117 ns/op -7233 ns/op / -5.4% (better)
Windows MSVC CreateStop/Go 119.500 ns/op +2.5 ns/op / +2.1% (worse)
Windows MSVC CreateStop/LLGo 503.900 ns/op +29.3 ns/op / +6.2% (worse)
Windows MSVC RearmStopped/Go 32.550 ns/op +0.87 ns/op / +2.7% (worse)
Windows MSVC RearmStopped/LLGo 308.600 ns/op +0.9 ns/op / +0.3% (worse)
Windows MSVC ResetActive/Go 23.050 ns/op +4 ns/op / +21.0% (worse)
Windows MSVC ResetActive/LLGo 150.800 ns/op -10.2 ns/op / -6.3% (better)
Windows MSVC ResetHeap1024/Go 19.750 ns/op +0.55 ns/op / +2.9% (worse)
Windows MSVC ResetHeap1024/LLGo 152.800 ns/op -1.7 ns/op / -1.1% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 770.200 ns/op +1.1 ns/op / +0.1% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 124906 ns/op +723 ns/op / +0.6% (worse)
Windows MSVC 386 CreateStop/Go 167.800 ns/op -3.2 ns/op / -1.9% (better)
Windows MSVC 386 CreateStop/LLGo 1945 ns/op -41 ns/op / -2.1% (better)
Windows MSVC 386 RearmStopped/Go 56.810 ns/op +0.08 ns/op / +0.1% (worse)
Windows MSVC 386 RearmStopped/LLGo 286.600 ns/op -20.6 ns/op / -6.7% (better)
Windows MSVC 386 ResetActive/Go 32.840 ns/op +0.21 ns/op / +0.6% (worse)
Windows MSVC 386 ResetActive/LLGo 208.800 ns/op 0 ns/op / +0.0%
Windows MSVC 386 ResetHeap1024/Go 33.270 ns/op +0.4 ns/op / +1.2% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 150.800 ns/op -0.3 ns/op / -0.2% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 662.800 ns/op -2.3 ns/op / -0.3% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 124619 ns/op -498 ns/op / -0.4% (better)
Windows MSVC ARM64 CreateStop/Go 197.100 ns/op -14.4 ns/op / -6.8% (better)
Windows MSVC ARM64 CreateStop/LLGo 498.800 ns/op -2.3 ns/op / -0.5% (better)
Windows MSVC ARM64 RearmStopped/Go 70.590 ns/op -0.01 ns/op / -0.01416% (better)
Windows MSVC ARM64 RearmStopped/LLGo 311.900 ns/op -9 ns/op / -2.8% (better)
Windows MSVC ARM64 ResetActive/Go 30.920 ns/op +0.06 ns/op / +0.2% (worse)
Windows MSVC ARM64 ResetActive/LLGo 165.400 ns/op -2.1 ns/op / -1.3% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.130 ns/op +0.11 ns/op / +0.4% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 143.800 ns/op -0.9 ns/op / -0.6% (better)

Compared with 3dce98b9191e measured in the same runner job.

@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

484c74c24635 | workflow run | long-term charts

WebAssembly output sizes

Profile and compiler Wasm module vs base Generated JS glue vs base
ec32/LLGo 112483 B 0 B / +0.0% 70724 B 0 B / +0.0%
ec64/LLGo 118028 B 0 B / +0.0% 74021 B 0 B / +0.0%
js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
js/LLGo 66072 B 0 B / +0.0% 68499 B 0 B / +0.0%
wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
wasip1/LLGo 72535 B 0 B / +0.0% 0 B 0 B / 0.0%
wc32/LLGo 116428 B 0 B / +0.0% 0 B 0 B / 0.0%

LLGo WebAssembly build measurements

Profile Build vs base
ec32 5.434 s -172.7 ms / -3.1% (better)
ec64 5.002 s -155.2 ms / -3.0% (better)
js 4.567 s -129.8 ms / -2.8% (better)
wasip1 3.051 s -77.72 ms / -2.5% (better)
wc32 3.802 s +7.47 ms / +0.2% (worse)

Compared with 3dce98b9191e measured in the same runner job.

@xushiwei
xushiwei merged commit f6c8d67 into xgo-dev:main Sep 8, 2026
64 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants