Skip to content

[WASM W2-A] Complete host integration and reflection - #2601

Merged
xushiwei merged 35 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-w2-host-reflect-20260913
Sep 17, 2026
Merged

xushiwei merged 35 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-w2-host-reflect-20260913

Conversation

@cpunion

@cpunion cpunion commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Tracks #2152. Depends on W1 #2578. Replaces the closed #2579 with a clean review thread; its fixes remain in the branch.

Scope

  • Complete JavaScript host boundaries: exceptions, embedded NUL strings, constructor errors, synchronous callback continuations, memory growth, exit, and asynchronous filesystem operations.
  • Reuse GOROOT syscall/js semantics for J32/GoJS (wasm32 with the Go-compatible JavaScript host); retain the adapters for J32/Emscripten (wasm32 with the Emscripten JavaScript host) and J64/Emscripten (wasm64 with the Emscripten JavaScript host). Exercise Node and real Chrome.
  • Complete reflection calls, method values, variadics, multiple/aggregate results, MakeFunc, and runtime-created function descriptors, with GC-root and Asyncify lifetime handling.
  • JavaScript providers continue to use libffi and emit no typed reflection bridges. Only W32/WASI (wasm32 with WASI Preview 1) uses demand-driven Call_/Make_ bridges, deduplicated by lowered signature and GC-root shape, with compact names and cache-aware backend selection.

Review boundary

Head 1eb979a9515f, based on W1 cb7214ae45fa and main 2db247e43848. The dedicated range is cb7214ae45fa..1eb979a9515f: 35 commits, 83 files, +4,360/-369. The original 33 feature commits are unchanged; follow-ups add the review-requested license header and move the already-validated browser completion driver from W3 into this host layer. The full diff against main also includes W1; single-worker runtime/GC and broad compatibility acceptance belong to the subsequent PRs.

Validation

The previous CI failure was a build-time integration error: newly merged llgo env referenced the removed Config.WasmABI field. W1 now reports WasmProfile and WasmProvider, with regression tests for all named profiles, both aliases, and a non-WASM target. This branch inherits that fix.

The runtime job on the previous head passed the Node scheduler, timers, GC, lifecycle, and callback fixtures, then timed out waiting for Chrome. The browser driver now uses the page-result protocol already present in W3, rather than coupling test completion to --dump-dom and browser shutdown. The 60-second deadline remains unchanged; explicit failures and missing results still fail. A new negative check requires a missing module to produce a page-load failure, not a timeout. No compiler/runtime implementation or WASM fixture code changes in this follow-up.

Current-head local compiler construction, reflection regressions, real Chrome callback checks on J32/Emscripten (wasm32 with the Emscripten JavaScript host) and J32/GoJS (wasm32 with the Go-compatible JavaScript host), and the explicit page-load failure check pass. The fork-only final-candidate validation passed the complete runtime gate on LLVM 22, Emscripten 4.0.21, Node 24, and Chrome for Testing 153.0.8010.36, including the negative load check and both real-browser providers. Its candidate differs from this head only by the validation-only workflow; that workflow is not included in this PR.

The upstream current-head runtime job and test-command job also pass. Remaining current-head CI, patch coverage, review, and paired size/performance benchmarks remain required gates. Previous measurements and resolved review history remain in #2579; they are not treated as new-head passes. Benchmarks include cprintf, println, fmtprintf, and active reflection, with particular attention to keeping WASI bridge costs out of JavaScript providers.

@codecov

codecov Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 99.29245% with 3 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
internal/build/wasm_reflect.go 97.34% 3 Missing ⚠️

📢 Thoughts on this report? Let us know!

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: [WASM W2-A] Complete host integration and reflection

Reviewed the full diff (~12.5k lines, 165 files) across code quality, performance, security, and documentation. This is a large, carefully written change. The tricky new invariants all hold up under close reading:

  • Wide Go word storage (ssa/storage.go, ssa/memory.go): static struct form {i8*, i32} vs. dynamic i64 access are internally consistent (both 8 bytes, struct-typed alloca aliases correctly); integer narrowing inserts runtime range assertions / compile-time panics.
  • Reflection bridges (ssa/reflect_bridge.go, internal/build/wasm_reflect.go): bijective base-62 IDs and shape-based dedup are correct; bridge generation is memoized by lowered-signature shape key. WASI-only gating matches the docs.
  • JS/WASM host boundary (syscall/js, emval.cpp, host.c): the CopyBytes/Length rewrites re-acquire heap views after growth, clamp copy lengths with min, and the new panic-on-non-array behavior matches upstream Go semantics. Embedded-NUL strings are handled by byte-length rather than C-string scanning.
  • Reflect call bridges (call_bridge_wasm.go, makefunc*.go): arg-count validated before the call; result copies use min(ffiType.Size, typ.Size_); indirect args are deep-copied off the transient frame.

No blocking correctness, security, or performance issues found.

Notes (non-blocking, no action required):

  • runtime/internal/ffi/pointer_array_wasm32.go: the wasm32 valuePointerArray adds a per-call []uint32 allocation on the FFI call path (Go 8-byte pointers → 4-byte libffi slots). This is functionally required; flagging only in case it shows up in profiles.
  • runtime/internal/wasmjs/_wrap/host.c (hostFinalize): decrements _goRefCounts[id] with no bounds/underflow guard. Mirrors upstream wasm_exec.js and id comes from the trusted Go runtime, so not attacker-reachable; a defensive id < length guard would match the other new bounds handling.

One minor consistency item is inline below.

Comment thread ssa/reflect_bridge.go
@github-actions

github-actions Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

5659f7f35bb4 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 19848 B +16 B / +0.1% (worse) 387 B 0 B / +0.0% 584.598 ms +15.06 ms / +2.6% (worse) 1.344 ms -29.54 us / -2.2% (better)
Linux cprintf-lto 19600 B +16 B / +0.1% (worse) 368 B 0 B / +0.0% 620.105 ms +45.48 ms / +7.9% (worse) 1.426 ms +16.5 us / +1.2% (worse)
Linux fmtprintf 1638112 B +1856 B / +0.1% (worse) 486657 B +430 B / +0.1% (worse) 4.869 s -233.5 ms / -4.6% (better) 3.338 ms -477.9 us / -12.5% (better)
Linux fmtprintf-lto 1476560 B +1184 B / +0.1% (worse) 422326 B +328 B / +0.1% (worse) 12.253 s -575.7 ms / -4.5% (better) 3.142 ms -215.3 us / -6.4% (better)
Linux println 61984 B +16 B / +0.02582% (worse) 14668 B 0 B / +0.0% 623.828 ms +29.2 ms / +4.9% (worse) 1.789 ms +53.45 us / +3.1% (worse)
Linux println-lto 54056 B +16 B / +0.02961% (worse) 12115 B 0 B / +0.0% 793.322 ms -24.94 ms / -3.0% (better) 1.693 ms +3.454 us / +0.2% (worse)
macOS cprintf 84480 B 0 B / +0.0% 17117 B +16 B / +0.1% (worse) 543.392 ms +34.2 ms / +6.7% (worse) 3.388 ms +1.179 ms / +53.4% (worse)
macOS cprintf-lto 84288 B 0 B / +0.0% 12881 B +16 B / +0.1% (worse) 1.002 s +470.2 ms / +88.4% (worse) 3.791 ms +1.604 ms / +73.4% (worse)
macOS fmtprintf 1483824 B +336 B / +0.02265% (worse) 860864 B +1240 B / +0.1% (worse) 3.729 s +1.068 s / +40.1% (worse) 9.604 ms +5.902 ms / +159.4% (worse)
macOS fmtprintf-lto 1175808 B 0 B / +0.0% 831832 B +968 B / +0.1% (worse) 7.231 s +785.5 ms / +12.2% (worse) 6.167 ms +1.377 ms / +28.7% (worse)
macOS println 114672 B 0 B / +0.0% 34684 B +16 B / +0.04615% (worse) 838.924 ms +334.2 ms / +66.2% (worse) 5.841 ms +2.982 ms / +104.3% (worse)
macOS println-lto 118720 B 0 B / +0.0% 32128 B +16 B / +0.04983% (worse) 965.155 ms +317.2 ms / +48.9% (worse) 5.686 ms +2.701 ms / +90.5% (worse)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 875.865 ms +13.96 ms / +1.6% (worse) 2.720 ms -14 us / -0.5% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 897.330 ms -8.521 ms / -0.9% (better) 2.717 ms -2.9 us / -0.1% (better)
Windows MinGW fmtprintf 1911808 B +1536 B / +0.1% (worse) 588790 B +480 B / +0.1% (worse) 3.206 s -18.56 ms / -0.6% (better) 6.192 ms +103.9 us / +1.7% (worse)
Windows MinGW fmtprintf-lto 1932288 B +1024 B / +0.1% (worse) 534790 B +256 B / +0.04789% (worse) 7.935 s -213.7 ms / -2.6% (better) 6.241 ms +201.6 us / +3.3% (worse)
Windows MinGW println 71168 B 0 B / +0.0% 23718 B 0 B / +0.0% 870.960 ms +4.307 ms / +0.5% (worse) 5.001 ms +11.7 us / +0.2% (worse)
Windows MinGW println-lto 65024 B 0 B / +0.0% 20406 B 0 B / +0.0% 1.027 s -45.2 ms / -4.2% (better) 4.952 ms -155.9 us / -3.1% (better)
Windows MinGW 386 cprintf 42496 B 0 B / +0.0% 5326 B 0 B / +0.0% 964.949 ms +99.09 ms / +11.4% (worse) 4.492 ms +202.6 us / +4.7% (worse)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 954.627 ms +74.45 ms / +8.5% (worse) 4.535 ms +99.4 us / +2.2% (worse)
Windows MinGW 386 fmtprintf 1874944 B +3072 B / +0.2% (worse) 462910 B +352 B / +0.1% (worse) 3.456 s +165.9 ms / +5.0% (worse) 9.329 ms +431.7 us / +4.9% (worse)
Windows MinGW 386 fmtprintf-lto 2151424 B +2048 B / +0.1% (worse) 440454 B +180 B / +0.04088% (worse) 7.904 s +600.3 ms / +8.2% (worse) 10.181 ms +624.2 us / +6.5% (worse)
Windows MinGW 386 println 90624 B 0 B / +0.0% 19830 B 0 B / +0.0% 934.689 ms +94.18 ms / +11.2% (worse) 8.660 ms +1.349 ms / +18.5% (worse)
Windows MinGW 386 println-lto 69120 B 0 B / +0.0% 17742 B 0 B / +0.0% 1.095 s +90.34 ms / +9.0% (worse) 7.752 ms +496.5 us / +6.8% (worse)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.518 s +20.02 ms / +1.3% (worse) 7.494 ms +391 us / +5.5% (worse)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.541 s +22.72 ms / +1.5% (worse) 7.336 ms +225.4 us / +3.2% (worse)
Windows MinGW ARM64 fmtprintf 1799680 B +1536 B / +0.1% (worse) 501284 B +504 B / +0.1% (worse) 4.420 s +80.64 ms / +1.9% (worse) 13.573 ms -381.2 us / -2.7% (better)
Windows MinGW ARM64 fmtprintf-lto 1857024 B +2048 B / +0.1% (worse) 465500 B +208 B / +0.0447% (worse) 9.800 s -144.8 ms / -1.5% (better) 13.680 ms +338.4 us / +2.5% (worse)
Windows MinGW ARM64 println 68096 B 0 B / +0.0% 22376 B 0 B / +0.0% 1.506 s +10.24 ms / +0.7% (worse) 12.118 ms -279.7 us / -2.3% (better)
Windows MinGW ARM64 println-lto 63488 B 0 B / +0.0% 19444 B 0 B / +0.0% 1.705 s +23.22 ms / +1.4% (worse) 12.390 ms +539.4 us / +4.6% (worse)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65782 B 0 B / +0.0% 943.061 ms -195.1 ms / -17.1% (better) 3.595 ms -211.1 us / -5.5% (better)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65718 B 0 B / +0.0% 1.041 s +62.96 ms / +6.4% (worse) 3.417 ms -8.4 us / -0.2% (better)
Windows MSVC fmtprintf 1626112 B +1536 B / +0.1% (worse) 684310 B +496 B / +0.1% (worse) 3.688 s -146.6 ms / -3.8% (better) 8.914 ms -538.4 us / -5.7% (better)
Windows MSVC fmtprintf-lto 1615360 B +1024 B / +0.1% (worse) 635414 B +240 B / +0.03778% (worse) 9.063 s +188 ms / +2.1% (worse) 10.491 ms +1.415 ms / +15.6% (worse)
Windows MSVC println 192512 B 0 B / +0.0% 119142 B 0 B / +0.0% 937.295 ms -18.13 ms / -1.9% (better) 7.527 ms +327.8 us / +4.6% (worse)
Windows MSVC println-lto 189952 B 0 B / +0.0% 116374 B 0 B / +0.0% 1.120 s -177.3 ms / -13.7% (better) 7.363 ms -401.1 us / -5.2% (better)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 968.128 ms +12.24 ms / +1.3% (worse) 8.848 ms +3.055 ms / +52.7% (worse)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.011 s -119.7 ms / -10.6% (better) 5.981 ms -2.417 ms / -28.8% (better)
Windows MSVC 386 fmtprintf 1188352 B +1536 B / +0.1% (worse) 446261 B +336 B / +0.1% (worse) 3.971 s +121.7 ms / +3.2% (worse) 12.041 ms +906.7 us / +8.1% (worse)
Windows MSVC 386 fmtprintf-lto 1223168 B +1024 B / +0.1% (worse) 417065 B +192 B / +0.04606% (worse) 8.897 s +256.6 ms / +3.0% (worse) 12.630 ms -221 us / -1.7% (better)
Windows MSVC 386 println 34304 B 0 B / +0.0% 18673 B 0 B / +0.0% 959.394 ms -762.9 us / -0.1% (better) 9.217 ms -30 us / -0.3% (better)
Windows MSVC 386 println-lto 32256 B 0 B / +0.0% 16855 B 0 B / +0.0% 1.159 s +67.18 ms / +6.2% (worse) 9.838 ms +566.1 us / +6.1% (worse)
Windows MSVC ARM64 cprintf 11264 B 0 B / +0.0% 3976 B 0 B / +0.0% 1.858 s -21.32 ms / -1.1% (better) 6.282 ms +266.4 us / +4.4% (worse)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 3868 B 0 B / +0.0% 1.861 s -18.5 ms / -1.0% (better) 6.385 ms -39 us / -0.6% (better)
Windows MSVC ARM64 fmtprintf 1373184 B +1024 B / +0.1% (worse) 502228 B +496 B / +0.1% (worse) 6.100 s +19.66 ms / +0.3% (worse) 12.716 ms +437.2 us / +3.6% (worse)
Windows MSVC ARM64 fmtprintf-lto 1389568 B +1024 B / +0.1% (worse) 468420 B +224 B / +0.04784% (worse) 14.426 s -28.12 ms / -0.2% (better) 12.820 ms +163.3 us / +1.3% (worse)
Windows MSVC ARM64 println 41472 B 0 B / +0.0% 21624 B 0 B / +0.0% 1.833 s -11.55 ms / -0.6% (better) 10.770 ms +15.3 us / +0.1% (worse)
Windows MSVC ARM64 println-lto 39424 B 0 B / +0.0% 19372 B 0 B / +0.0% 2.137 s -12.2 ms / -0.6% (better) 10.998 ms +139.2 us / +1.3% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 15.160 ns/op +0.35 ns/op / +2.4% (worse)
Linux BenchmarkMergeCompilerFlags 203.800 ns/op -12.1 ns/op / -5.6% (better)
Linux BenchmarkMergeLinkerFlags 135.700 ns/op -18.6 ns/op / -12.1% (better)
Linux BenchmarkChannelBuffered 55.330 ns/op -0.3 ns/op / -0.5% (better)
Linux BenchmarkChannelHandoff 13901 ns/op +414 ns/op / +3.1% (worse)
Linux BenchmarkDefer 48.030 ns/op -5.01 ns/op / -9.4% (better)
Linux BenchmarkDirectCall 1.167 ns/op -0.016 ns/op / -1.4% (better)
Linux BenchmarkGlobalRead 1.553 ns/op -0.013 ns/op / -0.8% (better)
Linux BenchmarkGlobalWrite 7.762 ns/op -0.057 ns/op / -0.7% (better)
Linux BenchmarkGoroutine 23268 ns/op -346 ns/op / -1.5% (better)
Linux BenchmarkInterfaceCall 5.904 ns/op +0.018 ns/op / +0.3% (worse)
Linux BenchmarkRuntimeGetG 2.439 ns/op -0.003 ns/op / -0.1% (better)
macOS BenchmarkLookupPCRandom 13.520 ns/op +1.21 ns/op / +9.8% (worse)
macOS BenchmarkMergeCompilerFlags 92.760 ns/op -5.21 ns/op / -5.3% (better)
macOS BenchmarkMergeLinkerFlags 59.990 ns/op -2.37 ns/op / -3.8% (better)
macOS BenchmarkChannelBuffered 24.700 ns/op -0.83 ns/op / -3.3% (better)
macOS BenchmarkChannelHandoff 7107 ns/op -2169 ns/op / -23.4% (better)
macOS BenchmarkDefer 35.800 ns/op +1.08 ns/op / +3.1% (worse)
macOS BenchmarkDirectCall 1.030 ns/op -0.027 ns/op / -2.6% (better)
macOS BenchmarkGlobalRead 1.060 ns/op -0.018 ns/op / -1.7% (better)
macOS BenchmarkGlobalWrite 1.123 ns/op +0.078 ns/op / +7.5% (worse)
macOS BenchmarkGoroutine 53070 ns/op +19963 ns/op / +60.3% (worse)
macOS BenchmarkInterfaceCall 3.792 ns/op -0.103 ns/op / -2.6% (better)
macOS BenchmarkRuntimeGetG 2.021 ns/op -0.132 ns/op / -6.1% (better)
Windows MinGW BenchmarkLookupPCRandom 9.601 ns/op -0.174 ns/op / -1.8% (better)
Windows MinGW BenchmarkMergeCompilerFlags 396.700 ns/op +1 ns/op / +0.3% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 333.200 ns/op +4.2 ns/op / +1.3% (worse)
Windows MinGW BenchmarkChannelBuffered 24.040 ns/op +0.58 ns/op / +2.5% (worse)
Windows MinGW BenchmarkChannelHandoff 1206 ns/op +79 ns/op / +7.0% (worse)
Windows MinGW BenchmarkDefer 41.940 ns/op +0.05 ns/op / +0.1% (worse)
Windows MinGW BenchmarkDirectCall 1.357 ns/op 0 ns/op / +0.0%
Windows MinGW BenchmarkGlobalRead 1.364 ns/op +0.001 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGlobalWrite 2.167 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW BenchmarkGoroutine 57465 ns/op +2633 ns/op / +4.8% (worse)
Windows MinGW BenchmarkInterfaceCall 7.061 ns/op -0.267 ns/op / -3.6% (better)
Windows MinGW BenchmarkRuntimeGetG 1.408 ns/op -0.223 ns/op / -13.7% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 51.860 ns/op +0.03 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkMergeCompilerFlags 534.500 ns/op +9.8 ns/op / +1.9% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 526.700 ns/op +20.1 ns/op / +4.0% (worse)
Windows MinGW 386 BenchmarkChannelBuffered 36.060 ns/op -1.02 ns/op / -2.8% (better)
Windows MinGW 386 BenchmarkChannelHandoff 4268 ns/op -236 ns/op / -5.2% (better)
Windows MinGW 386 BenchmarkDefer 34.210 ns/op -9.23 ns/op / -21.2% (better)
Windows MinGW 386 BenchmarkDirectCall 0.354 ns/op -0.0228 ns/op / -6.0% (better)
Windows MinGW 386 BenchmarkGlobalRead 0.692 ns/op -0.0128 ns/op / -1.8% (better)
Windows MinGW 386 BenchmarkGlobalWrite 12.700 ns/op -0.32 ns/op / -2.5% (better)
Windows MinGW 386 BenchmarkGoroutine 68770 ns/op -8196 ns/op / -10.6% (better)
Windows MinGW 386 BenchmarkInterfaceCall 4.138 ns/op -0.114 ns/op / -2.7% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 0.962 ns/op -0.0211 ns/op / -2.1% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.080 ns/op -0.04 ns/op / -0.3% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 567.300 ns/op -0.8 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 536.700 ns/op -8.3 ns/op / -1.5% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 38.560 ns/op -1.22 ns/op / -3.1% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 2327 ns/op +160 ns/op / +7.4% (worse)
Windows MinGW ARM64 BenchmarkDefer 57.220 ns/op -0.11 ns/op / -0.2% (better)
Windows MinGW ARM64 BenchmarkDirectCall 0.663 ns/op -0.0001 ns/op / -0.01507% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.590 ns/op -0.0026 ns/op / -0.4% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.663 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGoroutine 61948 ns/op +2489 ns/op / +4.2% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.146 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.804 ns/op +0.035 ns/op / +2.0% (worse)
Windows MSVC BenchmarkLookupPCRandom 13.190 ns/op -0.02 ns/op / -0.2% (better)
Windows MSVC BenchmarkMergeCompilerFlags 613.900 ns/op +24.2 ns/op / +4.1% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 545.100 ns/op +21.4 ns/op / +4.1% (worse)
Windows MSVC BenchmarkChannelBuffered 28.580 ns/op -0.62 ns/op / -2.1% (better)
Windows MSVC BenchmarkChannelHandoff 1190 ns/op +69 ns/op / +6.2% (worse)
Windows MSVC BenchmarkDefer 54.320 ns/op +0.81 ns/op / +1.5% (worse)
Windows MSVC BenchmarkDirectCall 1.855 ns/op -0.008 ns/op / -0.4% (better)
Windows MSVC BenchmarkGlobalRead 1.857 ns/op -0.003 ns/op / -0.2% (better)
Windows MSVC BenchmarkGlobalWrite 2.470 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC BenchmarkGoroutine 80302 ns/op +1461 ns/op / +1.9% (worse)
Windows MSVC BenchmarkInterfaceCall 8.382 ns/op +0.267 ns/op / +3.3% (worse)
Windows MSVC BenchmarkRuntimeGetG 2.480 ns/op +0.284 ns/op / +12.9% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 26.840 ns/op +0.3 ns/op / +1.1% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 757.600 ns/op +31.2 ns/op / +4.3% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 705.800 ns/op +30.2 ns/op / +4.5% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 38.710 ns/op -0.77 ns/op / -2.0% (better)
Windows MSVC 386 BenchmarkChannelHandoff 828.900 ns/op -82.4 ns/op / -9.0% (better)
Windows MSVC 386 BenchmarkDefer 45.460 ns/op +0.1 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkDirectCall 1.550 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.860 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGlobalWrite 7.775 ns/op -0.021 ns/op / -0.3% (better)
Windows MSVC 386 BenchmarkGoroutine 87467 ns/op -1069 ns/op / -1.2% (better)
Windows MSVC 386 BenchmarkInterfaceCall 8.385 ns/op +0.001 ns/op / +0.01193% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 2.179 ns/op +0.251 ns/op / +13.0% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 11.980 ns/op -0.09 ns/op / -0.7% (better)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 587.300 ns/op +20.3 ns/op / +3.6% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 539 ns/op +9.9 ns/op / +1.9% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 37.500 ns/op +0.17 ns/op / +0.5% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 2519 ns/op -283 ns/op / -10.1% (better)
Windows MSVC ARM64 BenchmarkDefer 59.090 ns/op -0.06 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.589 ns/op -0.0002 ns/op / -0.03393% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.664 ns/op +0.0005 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.794 ns/op -0.001 ns/op / -0.02635% (better)
Windows MSVC ARM64 BenchmarkGoroutine 52895 ns/op -1120 ns/op / -2.1% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.165 ns/op +0.009 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.812 ns/op -0.447 ns/op / -19.8% (better)

Timer runtime benchmarks

Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 964.200 ns/op -75.8 ns/op / -7.3% (better)
Linux AfterFuncZeroDelivery/LLGo 42045 ns/op +7130 ns/op / +20.4% (worse)
Linux CreateStop/Go 334.500 ns/op +12.6 ns/op / +3.9% (worse)
Linux CreateStop/LLGo 1809 ns/op +121 ns/op / +7.2% (worse)
Linux RearmStopped/Go 117.500 ns/op 0 ns/op / +0.0%
Linux RearmStopped/LLGo 1393 ns/op +58 ns/op / +4.3% (worse)
Linux ResetActive/Go 69.640 ns/op -1.72 ns/op / -2.4% (better)
Linux ResetActive/LLGo 766.400 ns/op +5.6 ns/op / +0.7% (worse)
Linux ResetHeap1024/Go 67.060 ns/op -0.43 ns/op / -0.6% (better)
Linux ResetHeap1024/LLGo 181.200 ns/op -2 ns/op / -1.1% (better)
macOS AfterFuncZeroDelivery/Go 420.500 ns/op -18.7 ns/op / -4.3% (better)
macOS AfterFuncZeroDelivery/LLGo 83055 ns/op +21472 ns/op / +34.9% (worse)
macOS CreateStop/Go 149.400 ns/op +14 ns/op / +10.3% (worse)
macOS CreateStop/LLGo 515.200 ns/op +134.6 ns/op / +35.4% (worse)
macOS RearmStopped/Go 52.080 ns/op -4.9 ns/op / -8.6% (better)
macOS RearmStopped/LLGo 324.200 ns/op -5.2 ns/op / -1.6% (better)
macOS ResetActive/Go 50.470 ns/op +8.53 ns/op / +20.3% (worse)
macOS ResetActive/LLGo 159.600 ns/op -1.5 ns/op / -0.9% (better)
macOS ResetHeap1024/Go 42.880 ns/op +0.34 ns/op / +0.8% (worse)
macOS ResetHeap1024/LLGo 90.600 ns/op +5.52 ns/op / +6.5% (worse)
Windows MinGW AfterFuncZeroDelivery/Go 373.300 ns/op +5.8 ns/op / +1.6% (worse)
Windows MinGW AfterFuncZeroDelivery/LLGo 101241 ns/op +1748 ns/op / +1.8% (worse)
Windows MinGW CreateStop/Go 89.430 ns/op -0.29 ns/op / -0.3% (better)
Windows MinGW CreateStop/LLGo 347.800 ns/op +2.9 ns/op / +0.8% (worse)
Windows MinGW RearmStopped/Go 24.460 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW RearmStopped/LLGo 217.800 ns/op +4.7 ns/op / +2.2% (worse)
Windows MinGW ResetActive/Go 14.770 ns/op -0.04 ns/op / -0.3% (better)
Windows MinGW ResetActive/LLGo 125.900 ns/op +4.3 ns/op / +3.5% (worse)
Windows MinGW ResetHeap1024/Go 14.780 ns/op -0.05 ns/op / -0.3% (better)
Windows MinGW ResetHeap1024/LLGo 104.100 ns/op -1.3 ns/op / -1.2% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 695.200 ns/op -17.5 ns/op / -2.5% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 149196 ns/op -2850 ns/op / -1.9% (better)
Windows MinGW 386 CreateStop/Go 181 ns/op +0.1 ns/op / +0.1% (worse)
Windows MinGW 386 CreateStop/LLGo 818.600 ns/op +25.8 ns/op / +3.3% (worse)
Windows MinGW 386 RearmStopped/Go 64.920 ns/op -1.47 ns/op / -2.2% (better)
Windows MinGW 386 RearmStopped/LLGo 448.200 ns/op -67.3 ns/op / -13.1% (better)
Windows MinGW 386 ResetActive/Go 31.160 ns/op -0.62 ns/op / -2.0% (better)
Windows MinGW 386 ResetActive/LLGo 217.800 ns/op -34.6 ns/op / -13.7% (better)
Windows MinGW 386 ResetHeap1024/Go 31.430 ns/op +0.06 ns/op / +0.2% (worse)
Windows MinGW 386 ResetHeap1024/LLGo 124.900 ns/op -3 ns/op / -2.3% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 658.100 ns/op -5.6 ns/op / -0.8% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 136456 ns/op +4674 ns/op / +3.5% (worse)
Windows MinGW ARM64 CreateStop/Go 198.800 ns/op +0.3 ns/op / +0.2% (worse)
Windows MinGW ARM64 CreateStop/LLGo 383.500 ns/op +5.5 ns/op / +1.5% (worse)
Windows MinGW ARM64 RearmStopped/Go 70.640 ns/op +0.11 ns/op / +0.2% (worse)
Windows MinGW ARM64 RearmStopped/LLGo 255.400 ns/op -0.4 ns/op / -0.2% (better)
Windows MinGW ARM64 ResetActive/Go 30.960 ns/op -0.09 ns/op / -0.3% (better)
Windows MinGW ARM64 ResetActive/LLGo 127.300 ns/op +2.9 ns/op / +2.3% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.090 ns/op -0.04 ns/op / -0.1% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 127.100 ns/op -1.2 ns/op / -0.9% (better)
Windows MSVC AfterFuncZeroDelivery/Go 549.800 ns/op -2.4 ns/op / -0.4% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 154452 ns/op -764 ns/op / -0.5% (better)
Windows MSVC CreateStop/Go 115.400 ns/op -0.8 ns/op / -0.7% (better)
Windows MSVC CreateStop/LLGo 444.400 ns/op +12.2 ns/op / +2.8% (worse)
Windows MSVC RearmStopped/Go 31.210 ns/op -0.52 ns/op / -1.6% (better)
Windows MSVC RearmStopped/LLGo 261.800 ns/op -4.8 ns/op / -1.8% (better)
Windows MSVC ResetActive/Go 20.060 ns/op -0.28 ns/op / -1.4% (better)
Windows MSVC ResetActive/LLGo 134.800 ns/op -12.2 ns/op / -8.3% (better)
Windows MSVC ResetHeap1024/Go 20.470 ns/op -0.19 ns/op / -0.9% (better)
Windows MSVC ResetHeap1024/LLGo 125.300 ns/op -1.3 ns/op / -1.0% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 972.200 ns/op +17.6 ns/op / +1.8% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 192715 ns/op -2238 ns/op / -1.1% (better)
Windows MSVC 386 CreateStop/Go 196.600 ns/op +2.5 ns/op / +1.3% (worse)
Windows MSVC 386 CreateStop/LLGo 513.600 ns/op +43.8 ns/op / +9.3% (worse)
Windows MSVC 386 RearmStopped/Go 63.560 ns/op +0.06 ns/op / +0.1% (worse)
Windows MSVC 386 RearmStopped/LLGo 328.600 ns/op +4.6 ns/op / +1.4% (worse)
Windows MSVC 386 ResetActive/Go 39.160 ns/op +0.05 ns/op / +0.1% (worse)
Windows MSVC 386 ResetActive/LLGo 966.500 ns/op +4.7 ns/op / +0.5% (worse)
Windows MSVC 386 ResetHeap1024/Go 39.560 ns/op -0.01 ns/op / -0.02527% (better)
Windows MSVC 386 ResetHeap1024/LLGo 174.300 ns/op +2.2 ns/op / +1.3% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 677 ns/op +7.1 ns/op / +1.1% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 129166 ns/op +13145 ns/op / +11.3% (worse)
Windows MSVC ARM64 CreateStop/Go 197.800 ns/op -4.8 ns/op / -2.4% (better)
Windows MSVC ARM64 CreateStop/LLGo 370.100 ns/op +3.8 ns/op / +1.0% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.690 ns/op +0.03 ns/op / +0.04246% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 267.700 ns/op -1.4 ns/op / -0.5% (better)
Windows MSVC ARM64 ResetActive/Go 30.860 ns/op -0.16 ns/op / -0.5% (better)
Windows MSVC ARM64 ResetActive/LLGo 133.200 ns/op -1 ns/op / -0.7% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.050 ns/op +0.08 ns/op / +0.3% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 134.900 ns/op -1.2 ns/op / -0.9% (better)

Compared with f5208c189ede measured in the same runner job.

@cpunion
cpunion force-pushed the codex/wasm-w2-host-reflect-20260913 branch from 1eb979a to 5659f7f Compare September 17, 2026 00:31
@github-actions

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

5659f7f35bb4 | workflow run | long-term charts

WebAssembly output sizes

Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 130903 B -5 B / -0.003819% (better) 70786 B -35 B / -0.04942% (better)
cprintf/j32-goos-js/LLGo 129212 B +6 B / +0.004644% (worse) 69150 B -35 B / -0.1% (better)
cprintf/j64-emscripten-memory64/LLGo 119512 B -2 B / -0.001673% (better) 73998 B -35 B / -0.04728% (better)
cprintf/w32-goos-wasip1/LLGo 133478 B +2 B / +0.001498% (worse) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 133353 B +2 B / +0.0015% (worse) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3224244 B -4958 B / -0.2% (better) 114540 B +861 B / +0.8% (worse)
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3216656 B -8345 B / -0.3% (better) 98239 B -13804 B / -12.3% (better)
fmtprintf/j64-emscripten-memory64/LLGo 2925049 B -5418 B / -0.2% (better) 121521 B +664 B / +0.5% (worse)
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 3072421 B +34757 B / +1.1% (worse) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2932138 B +19370 B / +0.7% (worse) 0 B 0 B / 0.0%
j32-emscripten/LLGo 130035 B +1 B / +0.000769% (worse) 70786 B -35 B / -0.04942% (better)
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 128573 B +12 B / +0.009334% (worse) 69150 B -35 B / -0.1% (better)
j64-emscripten-memory64/LLGo 118746 B +1 B / +0.0008421% (worse) 73998 B -35 B / -0.04728% (better)
reflectcall/j32-emscripten/LLGo 1554294 B +723 B / +0.04654% (worse) 88949 B -35 B / -0.03933% (better)
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1555434 B +1161 B / +0.1% (worse) 87313 B -35 B / -0.04007% (better)
reflectcall/j64-emscripten-memory64/LLGo 1419569 B +648 B / +0.04567% (worse) 94158 B -35 B / -0.03716% (better)
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1646213 B +80914 B / +5.2% (worse) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1574579 B +76612 B / +5.1% (worse) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 132620 B +1 B / +0.000754% (worse) 0 B 0 B / 0.0%
w32-wasi/LLGo 132584 B +1 B / +0.0007542% (worse) 0 B 0 B / 0.0%

LLGo WebAssembly build measurements

Example and profile Build vs base
j32-emscripten 6.387 s -174.2 ms / -2.7% (better)
j32-goos-js 6.542 s -239.7 ms / -3.5% (better)
j64-emscripten-memory64 5.492 s -138.2 ms / -2.5% (better)
reflectcall/w32-wasi 31.137 s +688.4 ms / +2.3% (worse)
w32-goos-wasip1 4.877 s -101.5 ms / -2.0% (better)
w32-wasi 4.832 s -153.9 ms / -3.1% (better)

Compared with f5208c189ede measured in the same runner job.

@xushiwei
xushiwei merged commit cc16553 into xgo-dev:main Sep 17, 2026
65 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants