Skip to content

fix(build): stabilize wasm test flows - #2511

Merged
xushiwei merged 4 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-r3-toolchain-20260904
Sep 7, 2026
Merged

xushiwei merged 4 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-r3-toolchain-20260904

Conversation

@cpunion

@cpunion cpunion commented Sep 6, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This starts R3 after the merged R2.1 lifecycle work and focuses on WebAssembly toolchain completeness.

  • Route named-target llgo test -c through the non-executing emulator path after linking, instead of falling through to device flash or serial discovery.
  • Cover compile-only artifacts for Emscripten wasm32, Emscripten Memory64, and WASI in the single-worker integration script.
  • Disable clang implicit Binaryen discovery with clang 19+ --no-wasm-opt when LLGo owns the Asyncify pipeline.
  • Explicitly preserve the established optimized-build order: optimization after wasm-ld, then Asyncify/EH translation, then post-Asyncify optimization. This avoids environment-dependent double processing without increasing wasm output size.

This PR intentionally contains no multi-worker runtime work. The fork-runner CI fallback used during development remains isolated in cpunion#236 and is not part of this upstream contribution.

Base

Rebased onto the current xgo-dev/main at 9317592bb, including the merged R2.1 and Windows/CI fixes.

Validation

  • go test ./internal/build ./internal/crosscompile ./cmd/internal/test ./cmd/internal/flags -count=1
  • bash -n dev/test_wasm_single_worker.sh
  • shellcheck dev/test_wasm_single_worker.sh
  • LLVM 22 full dev/test_wasm_single_worker.sh run with LLGO_BUILD_CACHE=off
  • Same-root base/head WebAssembly benchmark across js, wasip1, Emscripten wasm32, Emscripten Memory64, and WASI: all module and glue sizes are byte-for-byte unchanged

The integration run passed scheduler, timer, GC, lifecycle, host-callback, expected deadlock/Goexit failures, public llgo test execution, and llgo test -c artifact validation across Emscripten wasm32, Emscripten Memory64, and WASI.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: hide clang's implicit wasm-opt for Asyncify wasm links

The core design is sound. When clang drives a wasm Asyncify link it can implicitly run any wasm-opt it finds on PATH, which would optimize the module before LLGo's deliberate Asyncify post-link pass. Filtering standalone wasm-opt directories out of the linker's PATH while leaving postLinkWasm's independent exec.LookPath untouched (wasm_postlink.go) is the right approach, and the feature is well-guarded (shouldHideClangImplicitWasmOpt: wasm + Asyncify + no explicit -extld) and unit-tested for the clang/emcc/non-wasm cases.

A few findings below. Security review found no issues (PATH is only filtered, never broadened; symlinks are canonicalized; the shell helper quotes all expansions and invokes tools as argv[0]). Performance is a non-concern — the scan runs once per produced executable, immediately before a far more expensive linker spawn.

Additional non-blocking notes:

  • Several new helpers (shouldHideClangImplicitWasmOpt, clangDriverMayRunWasmOpt, withoutClangImplicitWasmOpt, resolveToolDirInPath) lack doc comments, unlike the surrounding functions in this file. The why is load-bearing and currently only captured in test names — notably why emcc is deliberately excluded (it manages its own wasm-opt pipeline). A brief comment on each would help future maintainers.
  • executableNames (build.go:1722) appends both name+ext and name+strings.ToLower(ext) for each PATHEXT entry, doubling os.Stat calls on a case-insensitive filesystem; the lower-cased duplicates can be dropped.

Comment thread internal/build/build.go Outdated
Comment thread internal/build/build.go Outdated
Comment thread dev/test_wasm_single_worker.sh Outdated
@codecov

codecov Bot commented Sep 6, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Sep 6, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

58c37709a4e1 | workflow run | long-term charts

WebAssembly output sizes

Profile and compiler Wasm module vs base Generated JS glue vs base
ec32/LLGo 112737 B 0 B / +0.0% 70724 B 0 B / +0.0%
ec64/LLGo 118288 B 0 B / +0.0% 74021 B 0 B / +0.0%
js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
js/LLGo 65512 B 0 B / +0.0% 68499 B 0 B / +0.0%
wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
wasip1/LLGo 71802 B 0 B / +0.0% 0 B 0 B / 0.0%
wc32/LLGo 116701 B 0 B / +0.0% 0 B 0 B / 0.0%

LLGo WebAssembly build measurements

Profile Build vs base
ec32 5.184 s -511 ms / -9.0% (better)
ec64 4.854 s -282.4 ms / -5.5% (better)
js 4.564 s -159.8 ms / -3.4% (better)
wasip1 3.083 s -152.3 ms / -4.7% (better)
wc32 3.620 s -268.1 ms / -6.9% (better)

Compared with 9317592bb30a measured in the same runner job.

@cpunion
cpunion force-pushed the codex/wasm-r3-toolchain-20260904 branch from b3a87cd to 90192e1 Compare September 6, 2026 08:39
@cpunion

cpunion commented Sep 6, 2026

Copy link
Copy Markdown
Collaborator Author

Also addressed the non-blocking review notes in 283c2b2: documented the PATH-filtering helpers and why emcc is excluded, and removed redundant lowercase PATHEXT probes. LLVM 22 internal/build tests pass (309.071s), targeted review regressions pass, and the shell script passes syntax and ShellCheck validation.

@github-actions

github-actions Bot commented Sep 6, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

58c37709a4e1 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 19848 B 0 B / +0.0% 387 B 0 B / +0.0% 454.098 ms +1.848 ms / +0.4% (worse) 1.280 ms -36.83 us / -2.8% (better)
Linux cprintf-lto 19600 B 0 B / +0.0% 368 B 0 B / +0.0% 444.597 ms -16.89 ms / -3.7% (better) 1.269 ms +5.95 us / +0.5% (worse)
Linux fmtprintf 1617008 B +8 B / +0.0004947% (worse) 493198 B 0 B / +0.0% 3.529 s -56.41 ms / -1.6% (better) 3.061 ms -15.44 us / -0.5% (better)
Linux fmtprintf-lto 1478744 B 0 B / +0.0% 442979 B 0 B / +0.0% 10.565 s -11.77 ms / -0.1% (better) 3.084 ms +185.9 us / +6.4% (worse)
Linux println 62600 B 0 B / +0.0% 14972 B 0 B / +0.0% 436.739 ms -19.34 ms / -4.2% (better) 1.569 ms +3.826 us / +0.2% (worse)
Linux println-lto 54656 B 0 B / +0.0% 12537 B 0 B / +0.0% 654.497 ms -16.95 ms / -2.5% (better) 1.607 ms +39.07 us / +2.5% (worse)
macOS cprintf 84480 B 0 B / +0.0% 17117 B 0 B / +0.0% 656.872 ms -91.86 ms / -12.3% (better) 3.247 ms +304.2 us / +10.3% (worse)
macOS cprintf-lto 84288 B 0 B / +0.0% 12881 B 0 B / +0.0% 657.542 ms -70.33 ms / -9.7% (better) 2.382 ms -1.579 ms / -39.9% (better)
macOS fmtprintf 1473408 B 0 B / +0.0% 867160 B 0 B / +0.0% 2.648 s -574.4 ms / -17.8% (better) 3.805 ms -5.298 ms / -58.2% (better)
macOS fmtprintf-lto 1159568 B 0 B / +0.0% 848480 B 0 B / +0.0% 8.901 s -1.294 s / -12.7% (better) 5.149 ms -392.2 us / -7.1% (better)
macOS println 114800 B 0 B / +0.0% 34929 B 0 B / +0.0% 572.260 ms -73.31 ms / -11.4% (better) 3.764 ms -455.8 us / -10.8% (better)
macOS println-lto 118736 B 0 B / +0.0% 32881 B 0 B / +0.0% 748.845 ms -185.3 ms / -19.8% (better) 3.825 ms +52.29 us / +1.4% (worse)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.062 s +44.28 ms / +4.4% (worse) 3.737 ms -43.1 us / -1.1% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.113 s +72.91 ms / +7.0% (worse) 3.687 ms -129.8 us / -3.4% (better)
Windows MinGW fmtprintf 1886208 B 0 B / +0.0% 593766 B 0 B / +0.0% 3.764 s +97.94 ms / +2.7% (worse) 8.408 ms +99.4 us / +1.2% (worse)
Windows MinGW fmtprintf-lto 1939456 B 0 B / +0.0% 557286 B 0 B / +0.0% 9.052 s +372.1 ms / +4.3% (worse) 8.082 ms -165.3 us / -2.0% (better)
Windows MinGW println 71680 B 0 B / +0.0% 24038 B 0 B / +0.0% 1.046 s -11.09 ms / -1.0% (better) 6.447 ms -725.8 us / -10.1% (better)
Windows MinGW println-lto 66048 B 0 B / +0.0% 21014 B 0 B / +0.0% 1.221 s +4.726 ms / +0.4% (worse) 6.341 ms -635 us / -9.1% (better)
Windows MinGW 386 cprintf 37888 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.137 s -13.41 ms / -1.2% (better) 5.129 ms -751.8 us / -12.8% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.165 s +18.31 ms / +1.6% (worse) 5.154 ms +62.3 us / +1.2% (worse)
Windows MinGW 386 fmtprintf 1843200 B 0 B / +0.0% 468394 B 0 B / +0.0% 4.136 s +7.779 ms / +0.2% (worse) 11.125 ms +366.9 us / +3.4% (worse)
Windows MinGW 386 fmtprintf-lto 2166272 B 0 B / +0.0% 463406 B 0 B / +0.0% 9.837 s +197.7 ms / +2.1% (worse) 12.332 ms +922.1 us / +8.1% (worse)
Windows MinGW 386 println 86528 B 0 B / +0.0% 20454 B 0 B / +0.0% 1.167 s +34.13 ms / +3.0% (worse) 9.259 ms +240.9 us / +2.7% (worse)
Windows MinGW 386 println-lto 71168 B 0 B / +0.0% 18398 B 0 B / +0.0% 1.329 s -7.048 ms / -0.5% (better) 9.308 ms +229.7 us / +2.5% (worse)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4436 B 0 B / +0.0% 1.446 s +41.63 ms / +3.0% (worse) 6.615 ms +205.6 us / +3.2% (worse)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4368 B 0 B / +0.0% 1.470 s +27 ms / +1.9% (worse) 6.156 ms -113.2 us / -1.8% (better)
Windows MinGW ARM64 fmtprintf 1774080 B 0 B / +0.0% 506492 B 0 B / +0.0% 4.225 s +59.98 ms / +1.4% (worse) 13.126 ms +451.4 us / +3.6% (worse)
Windows MinGW ARM64 fmtprintf-lto 1866752 B 0 B / +0.0% 487084 B 0 B / +0.0% 9.790 s +58.45 ms / +0.6% (worse) 12.556 ms -131.1 us / -1.0% (better)
Windows MinGW ARM64 println 69120 B 0 B / +0.0% 22688 B 0 B / +0.0% 1.440 s +34.15 ms / +2.4% (worse) 11.104 ms +195.3 us / +1.8% (worse)
Windows MinGW ARM64 println-lto 65024 B 0 B / +0.0% 20260 B 0 B / +0.0% 1.633 s +37.8 ms / +2.4% (worse) 11.210 ms +439.5 us / +4.1% (worse)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65782 B 0 B / +0.0% 961.233 ms -4.455 ms / -0.5% (better) 3.309 ms -131.4 us / -3.8% (better)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65718 B 0 B / +0.0% 1.047 s +82.8 ms / +8.6% (worse) 3.404 ms +70.7 us / +2.1% (worse)
Windows MSVC fmtprintf 1616896 B 0 B / +0.0% 689286 B 0 B / +0.0% 4.502 s +642.9 ms / +16.7% (worse) 13.936 ms +4.59 ms / +49.1% (worse)
Windows MSVC fmtprintf-lto 1627136 B 0 B / +0.0% 659686 B 0 B / +0.0% 9.543 s +457.3 ms / +5.0% (worse) 13.886 ms +4.655 ms / +50.4% (worse)
Windows MSVC println 193024 B 0 B / +0.0% 119446 B 0 B / +0.0% 979.316 ms +26.74 ms / +2.8% (worse) 7.326 ms +74.5 us / +1.0% (worse)
Windows MSVC println-lto 190464 B 0 B / +0.0% 116822 B 0 B / +0.0% 1.318 s +184.4 ms / +16.3% (worse) 9.289 ms +2.295 ms / +32.8% (worse)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 964.083 ms +3.759 ms / +0.4% (worse) 5.559 ms +16.4 us / +0.3% (worse)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.049 s +77.16 ms / +7.9% (worse) 6.192 ms +1.03 ms / +20.0% (worse)
Windows MSVC 386 fmtprintf 1186304 B 0 B / +0.0% 451761 B 0 B / +0.0% 3.959 s +59.52 ms / +1.5% (worse) 12.328 ms -718.1 us / -5.5% (better)
Windows MSVC 386 fmtprintf-lto 1241600 B 0 B / +0.0% 443525 B 0 B / +0.0% 8.657 s -99.95 ms / -1.1% (better) 11.664 ms +97.7 us / +0.8% (worse)
Windows MSVC 386 println 34816 B 0 B / +0.0% 19313 B 0 B / +0.0% 975.747 ms +34.19 ms / +3.6% (worse) 10.680 ms +456.1 us / +4.5% (worse)
Windows MSVC 386 println-lto 33792 B 0 B / +0.0% 17463 B 0 B / +0.0% 1.132 s +3.259 ms / +0.3% (worse) 9.836 ms +70.9 us / +0.7% (worse)
Windows MSVC ARM64 cprintf 11264 B 0 B / +0.0% 3976 B 0 B / +0.0% 1.965 s -12.64 ms / -0.6% (better) 6.315 ms -92.9 us / -1.4% (better)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 3868 B 0 B / +0.0% 1.944 s -54.01 ms / -2.7% (better) 6.079 ms -280.6 us / -4.4% (better)
Windows MSVC ARM64 fmtprintf 1362944 B 0 B / +0.0% 506316 B 0 B / +0.0% 6.111 s -37.12 ms / -0.6% (better) 12.790 ms -33 us / -0.3% (better)
Windows MSVC ARM64 fmtprintf-lto 1398784 B 0 B / +0.0% 487836 B 0 B / +0.0% 14.758 s +3.675 ms / +0.02491% (worse) 12.648 ms -1.188 ms / -8.6% (better)
Windows MSVC ARM64 println 42496 B 0 B / +0.0% 22264 B 0 B / +0.0% 1.917 s -15.08 ms / -0.8% (better) 11.109 ms -162.7 us / -1.4% (better)
Windows MSVC ARM64 println-lto 40448 B 0 B / +0.0% 20156 B 0 B / +0.0% 2.218 s -66.42 ms / -2.9% (better) 11.483 ms +117.2 us / +1.0% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.720 ns/op +0.2 ns/op / +1.4% (worse)
Linux BenchmarkMergeCompilerFlags 199 ns/op +3.5 ns/op / +1.8% (worse)
Linux BenchmarkMergeLinkerFlags 143.300 ns/op +17 ns/op / +13.5% (worse)
Linux BenchmarkChannelBuffered 70.020 ns/op -0.08 ns/op / -0.1% (better)
Linux BenchmarkChannelHandoff 15324 ns/op +2500 ns/op / +19.5% (worse)
Linux BenchmarkDefer 47.730 ns/op +1.03 ns/op / +2.2% (worse)
Linux BenchmarkDirectCall 1.169 ns/op +0.005 ns/op / +0.4% (worse)
Linux BenchmarkGlobalRead 1.168 ns/op -0.001 ns/op / -0.1% (better)
Linux BenchmarkGlobalWrite 7.756 ns/op -0.104 ns/op / -1.3% (better)
Linux BenchmarkGoroutine 28833 ns/op +8587 ns/op / +42.4% (worse)
Linux BenchmarkInterfaceCall 6.623 ns/op -0.015 ns/op / -0.2% (better)
Linux BenchmarkRuntimeGetG 2.447 ns/op -0.004 ns/op / -0.2% (better)
macOS BenchmarkLookupPCRandom 18.790 ns/op +1.94 ns/op / +11.5% (worse)
macOS BenchmarkMergeCompilerFlags 181.300 ns/op +71.6 ns/op / +65.3% (worse)
macOS BenchmarkMergeLinkerFlags 110.800 ns/op +40.96 ns/op / +58.6% (worse)
macOS BenchmarkChannelBuffered 30.380 ns/op +0.67 ns/op / +2.3% (worse)
macOS BenchmarkChannelHandoff 10746 ns/op +281 ns/op / +2.7% (worse)
macOS BenchmarkDefer 43.450 ns/op +0.78 ns/op / +1.8% (worse)
macOS BenchmarkDirectCall 1.088 ns/op -0.174 ns/op / -13.8% (better)
macOS BenchmarkGlobalRead 1.214 ns/op -0.026 ns/op / -2.1% (better)
macOS BenchmarkGlobalWrite 1.063 ns/op -0.334 ns/op / -23.9% (better)
macOS BenchmarkGoroutine 33134 ns/op +6812 ns/op / +25.9% (worse)
macOS BenchmarkInterfaceCall 5.370 ns/op -0.889 ns/op / -14.2% (better)
macOS BenchmarkRuntimeGetG 2.578 ns/op +0.353 ns/op / +15.9% (worse)
Windows MinGW BenchmarkLookupPCRandom 10.430 ns/op -0.15 ns/op / -1.4% (better)
Windows MinGW BenchmarkMergeCompilerFlags 512.200 ns/op -0.3 ns/op / -0.1% (better)
Windows MinGW BenchmarkMergeLinkerFlags 454.500 ns/op +9.6 ns/op / +2.2% (worse)
Windows MinGW BenchmarkChannelBuffered 51.960 ns/op +0.23 ns/op / +0.4% (worse)
Windows MinGW BenchmarkChannelHandoff 1296 ns/op -358 ns/op / -21.6% (better)
Windows MinGW BenchmarkDefer 51.860 ns/op +4.04 ns/op / +8.4% (worse)
Windows MinGW BenchmarkDirectCall 1.197 ns/op +0.005 ns/op / +0.4% (worse)
Windows MinGW BenchmarkGlobalRead 1.210 ns/op -0.045 ns/op / -3.6% (better)
Windows MinGW BenchmarkGlobalWrite 7.687 ns/op -0.136 ns/op / -1.7% (better)
Windows MinGW BenchmarkGoroutine 75930 ns/op +6619 ns/op / +9.5% (worse)
Windows MinGW BenchmarkInterfaceCall 5.814 ns/op +0.074 ns/op / +1.3% (worse)
Windows MinGW BenchmarkRuntimeGetG 1.221 ns/op +0.001 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkLookupPCRandom 26.540 ns/op 0 ns/op / +0.0%
Windows MinGW 386 BenchmarkMergeCompilerFlags 767.900 ns/op +30.2 ns/op / +4.1% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 689.600 ns/op -1.1 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkChannelBuffered 42.890 ns/op -0.07 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkChannelHandoff 855 ns/op -196 ns/op / -18.6% (better)
Windows MinGW 386 BenchmarkDefer 44.510 ns/op +1.76 ns/op / +4.1% (worse)
Windows MinGW 386 BenchmarkDirectCall 1.548 ns/op +0.001 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGlobalRead 1.552 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGlobalWrite 7.772 ns/op -0.002 ns/op / -0.02573% (better)
Windows MinGW 386 BenchmarkGoroutine 89017 ns/op +6766 ns/op / +8.2% (worse)
Windows MinGW 386 BenchmarkInterfaceCall 9.604 ns/op +0.003 ns/op / +0.03125% (worse)
Windows MinGW 386 BenchmarkRuntimeGetG 1.859 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 11.960 ns/op -0.09 ns/op / -0.7% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 567 ns/op +6 ns/op / +1.1% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 531.900 ns/op +17.2 ns/op / +3.3% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 43.850 ns/op -2.6 ns/op / -5.6% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 2613 ns/op -27 ns/op / -1.0% (better)
Windows MinGW ARM64 BenchmarkDefer 55.560 ns/op +1.84 ns/op / +3.4% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.590 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGlobalRead 0.665 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGlobalWrite 0.589 ns/op -0.0001 ns/op / -0.01697% (better)
Windows MinGW ARM64 BenchmarkGoroutine 65943 ns/op +4538 ns/op / +7.4% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.721 ns/op +0.004 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.769 ns/op -0.033 ns/op / -1.8% (better)
Windows MSVC BenchmarkLookupPCRandom 12.930 ns/op -0.25 ns/op / -1.9% (better)
Windows MSVC BenchmarkMergeCompilerFlags 639.800 ns/op +21.8 ns/op / +3.5% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 542.100 ns/op -11 ns/op / -2.0% (better)
Windows MSVC BenchmarkChannelBuffered 37.560 ns/op +0.01 ns/op / +0.02663% (worse)
Windows MSVC BenchmarkChannelHandoff 1256 ns/op +30 ns/op / +2.4% (worse)
Windows MSVC BenchmarkDefer 55.080 ns/op -1.22 ns/op / -2.2% (better)
Windows MSVC BenchmarkDirectCall 1.548 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalRead 1.549 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC BenchmarkGlobalWrite 2.468 ns/op -0.001 ns/op / -0.0405% (better)
Windows MSVC BenchmarkGoroutine 83445 ns/op -1799 ns/op / -2.1% (better)
Windows MSVC BenchmarkInterfaceCall 9.312 ns/op +0.014 ns/op / +0.2% (worse)
Windows MSVC BenchmarkRuntimeGetG 1.862 ns/op -0.005 ns/op / -0.3% (better)
Windows MSVC 386 BenchmarkLookupPCRandom 26.500 ns/op -0.08 ns/op / -0.3% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 705.800 ns/op -33.1 ns/op / -4.5% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 657.100 ns/op -17 ns/op / -2.5% (better)
Windows MSVC 386 BenchmarkChannelBuffered 42.720 ns/op -0.04 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkChannelHandoff 784.900 ns/op -219.1 ns/op / -21.8% (better)
Windows MSVC 386 BenchmarkDefer 48.710 ns/op -0.22 ns/op / -0.4% (better)
Windows MSVC 386 BenchmarkDirectCall 1.546 ns/op -0.003 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkGlobalRead 1.858 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGlobalWrite 7.774 ns/op -0.001 ns/op / -0.01286% (better)
Windows MSVC 386 BenchmarkGoroutine 86442 ns/op +581 ns/op / +0.7% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 9.610 ns/op +0.014 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 2.166 ns/op -0.001 ns/op / -0.04615% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.010 ns/op -0.04 ns/op / -0.3% (better)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 577 ns/op +6.6 ns/op / +1.2% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 534.700 ns/op +6.9 ns/op / +1.3% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 47.550 ns/op +3.29 ns/op / +7.4% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 2067 ns/op -98 ns/op / -4.5% (better)
Windows MSVC ARM64 BenchmarkDefer 63.020 ns/op +2.65 ns/op / +4.4% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op +0.0004 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGlobalRead 0.664 ns/op +0.0005 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.795 ns/op -0.001 ns/op / -0.02634% (better)
Windows MSVC ARM64 BenchmarkGoroutine 54680 ns/op +170 ns/op / +0.3% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.729 ns/op +0.007 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.770 ns/op -0.041 ns/op / -2.3% (better)

Timer runtime benchmarks

Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 899.500 ns/op -3.8 ns/op / -0.4% (better)
Linux AfterFuncZeroDelivery/LLGo 32776 ns/op +1544 ns/op / +4.9% (worse)
Linux CreateStop/Go 290 ns/op -0.2 ns/op / -0.1% (better)
Linux CreateStop/LLGo 1845 ns/op -15 ns/op / -0.8% (better)
Linux RearmStopped/Go 115.900 ns/op +0.1 ns/op / +0.1% (worse)
Linux RearmStopped/LLGo 1180 ns/op -189 ns/op / -13.8% (better)
Linux ResetActive/Go 68.860 ns/op +0.23 ns/op / +0.3% (worse)
Linux ResetActive/LLGo 783.300 ns/op +24.5 ns/op / +3.2% (worse)
Linux ResetHeap1024/Go 67.100 ns/op +0.01 ns/op / +0.01491% (worse)
Linux ResetHeap1024/LLGo 188.800 ns/op -2.6 ns/op / -1.4% (better)
macOS AfterFuncZeroDelivery/Go 678.800 ns/op +166.4 ns/op / +32.5% (worse)
macOS AfterFuncZeroDelivery/LLGo 113733 ns/op +38754 ns/op / +51.7% (worse)
macOS CreateStop/Go 268 ns/op +104.3 ns/op / +63.7% (worse)
macOS CreateStop/LLGo 1218 ns/op +646 ns/op / +112.9% (worse)
macOS RearmStopped/Go 88.490 ns/op +16.29 ns/op / +22.6% (worse)
macOS RearmStopped/LLGo 398.500 ns/op +69.2 ns/op / +21.0% (worse)
macOS ResetActive/Go 67.070 ns/op +16.33 ns/op / +32.2% (worse)
macOS ResetActive/LLGo 221.900 ns/op +88.5 ns/op / +66.3% (worse)
macOS ResetHeap1024/Go 65.780 ns/op +13.55 ns/op / +25.9% (worse)
macOS ResetHeap1024/LLGo 110 ns/op +13.09 ns/op / +13.5% (worse)
Windows MinGW AfterFuncZeroDelivery/Go 619.600 ns/op -20 ns/op / -3.1% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 127498 ns/op +701 ns/op / +0.6% (worse)
Windows MinGW CreateStop/Go 171.800 ns/op -6.8 ns/op / -3.8% (better)
Windows MinGW CreateStop/LLGo 677 ns/op +82.9 ns/op / +14.0% (worse)
Windows MinGW RearmStopped/Go 63.680 ns/op +0.02 ns/op / +0.03142% (worse)
Windows MinGW RearmStopped/LLGo 339.800 ns/op -8.4 ns/op / -2.4% (better)
Windows MinGW ResetActive/Go 28.450 ns/op +0.05 ns/op / +0.2% (worse)
Windows MinGW ResetActive/LLGo 189.300 ns/op +3.1 ns/op / +1.7% (worse)
Windows MinGW ResetHeap1024/Go 28.570 ns/op +0.45 ns/op / +1.6% (worse)
Windows MinGW ResetHeap1024/LLGo 127.400 ns/op +6.6 ns/op / +5.5% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 967.500 ns/op +16.4 ns/op / +1.7% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 188371 ns/op -6269 ns/op / -3.2% (better)
Windows MinGW 386 CreateStop/Go 193.100 ns/op 0 ns/op / +0.0%
Windows MinGW 386 CreateStop/LLGo 2035 ns/op -39 ns/op / -1.9% (better)
Windows MinGW 386 RearmStopped/Go 63.420 ns/op -0.18 ns/op / -0.3% (better)
Windows MinGW 386 RearmStopped/LLGo 365.800 ns/op -7.7 ns/op / -2.1% (better)
Windows MinGW 386 ResetActive/Go 39.030 ns/op -0.1 ns/op / -0.3% (better)
Windows MinGW 386 ResetActive/LLGo 915.100 ns/op +99.9 ns/op / +12.3% (worse)
Windows MinGW 386 ResetHeap1024/Go 39.320 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW 386 ResetHeap1024/LLGo 193.100 ns/op -0.1 ns/op / -0.1% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 665.300 ns/op -6 ns/op / -0.9% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 133513 ns/op +7860 ns/op / +6.3% (worse)
Windows MinGW ARM64 CreateStop/Go 200.100 ns/op +5.7 ns/op / +2.9% (worse)
Windows MinGW ARM64 CreateStop/LLGo 392.400 ns/op +1.6 ns/op / +0.4% (worse)
Windows MinGW ARM64 RearmStopped/Go 70.580 ns/op +0.01 ns/op / +0.01417% (worse)
Windows MinGW ARM64 RearmStopped/LLGo 277.900 ns/op -2.6 ns/op / -0.9% (better)
Windows MinGW ARM64 ResetActive/Go 30.990 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW ARM64 ResetActive/LLGo 131.200 ns/op -0.5 ns/op / -0.4% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.090 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 138.400 ns/op -0.3 ns/op / -0.2% (better)
Windows MSVC AfterFuncZeroDelivery/Go 580 ns/op +9.7 ns/op / +1.7% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 152911 ns/op -10626 ns/op / -6.5% (better)
Windows MSVC CreateStop/Go 113.500 ns/op -1.9 ns/op / -1.6% (better)
Windows MSVC CreateStop/LLGo 438.100 ns/op +11.1 ns/op / +2.6% (worse)
Windows MSVC RearmStopped/Go 31.440 ns/op -0.05 ns/op / -0.2% (better)
Windows MSVC RearmStopped/LLGo 293.300 ns/op -13 ns/op / -4.2% (better)
Windows MSVC ResetActive/Go 20.180 ns/op -0.03 ns/op / -0.1% (better)
Windows MSVC ResetActive/LLGo 154.200 ns/op -13.6 ns/op / -8.1% (better)
Windows MSVC ResetHeap1024/Go 20.440 ns/op -0.04 ns/op / -0.2% (better)
Windows MSVC ResetHeap1024/LLGo 137.200 ns/op +0.5 ns/op / +0.4% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 952.900 ns/op -4.8 ns/op / -0.5% (better)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 186291 ns/op -6165 ns/op / -3.2% (better)
Windows MSVC 386 CreateStop/Go 193.500 ns/op +1.8 ns/op / +0.9% (worse)
Windows MSVC 386 CreateStop/LLGo 1687 ns/op +55 ns/op / +3.4% (worse)
Windows MSVC 386 RearmStopped/Go 63.490 ns/op +0.15 ns/op / +0.2% (worse)
Windows MSVC 386 RearmStopped/LLGo 327.500 ns/op -6.2 ns/op / -1.9% (better)
Windows MSVC 386 ResetActive/Go 39.060 ns/op +0.05 ns/op / +0.1% (worse)
Windows MSVC 386 ResetActive/LLGo 938.800 ns/op +8.9 ns/op / +1.0% (worse)
Windows MSVC 386 ResetHeap1024/Go 39.440 ns/op +0.12 ns/op / +0.3% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 175.100 ns/op -0.1 ns/op / -0.1% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 669.400 ns/op +3.9 ns/op / +0.6% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 133982 ns/op +16108 ns/op / +13.7% (worse)
Windows MSVC ARM64 CreateStop/Go 204.800 ns/op +7.8 ns/op / +4.0% (worse)
Windows MSVC ARM64 CreateStop/LLGo 423.400 ns/op +6.8 ns/op / +1.6% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.560 ns/op -0.04 ns/op / -0.1% (better)
Windows MSVC ARM64 RearmStopped/LLGo 291.600 ns/op -2.1 ns/op / -0.7% (better)
Windows MSVC ARM64 ResetActive/Go 30.970 ns/op -0.08 ns/op / -0.3% (better)
Windows MSVC ARM64 ResetActive/LLGo 136.400 ns/op +4.8 ns/op / +3.6% (worse)
Windows MSVC ARM64 ResetHeap1024/Go 31.080 ns/op -0.06 ns/op / -0.2% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 142.400 ns/op +0.3 ns/op / +0.2% (worse)

Compared with 9317592bb30a measured in the same runner job.

@cpunion

cpunion commented Sep 6, 2026

Copy link
Copy Markdown
Collaborator Author

Follow-up after the review and size/coverage audit:

  • Replaced PATH/PATHEXT filtering with clang 19+ official --no-wasm-opt, so the behavior is deterministic even if clang and Binaryen share a directory.
  • Made the former implicit pre-Asyncify optimization explicit before LLGo instrumentation. Same-root base/head outputs are byte-identical for every WebAssembly benchmark profile; the previous ~20 KiB WASI regressions are eliminated.
  • Removed the now-unused platform PATH helpers, including the lines responsible for the prior patch-coverage gap.
  • Added unit coverage for both clang and clang++ driver paths and both Binaryen phases.

Full local single-worker integration and the focused Go package suite pass. Waiting for the refreshed CI, Codecov, and benchmark checks before treating the PR as complete.

@cpunion cpunion changed the title [WASM R3] build: stabilize wasm test flows (wasm)build: stabilize wasm test flows Sep 7, 2026
@cpunion cpunion changed the title (wasm)build: stabilize wasm test flows fix(build): stabilize wasm test flows Sep 7, 2026
@xushiwei
xushiwei merged commit d98438f into xgo-dev:main Sep 7, 2026
75 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants