Skip to content

[WASM R4] Audit official profiles and standard-library completeness - #235

Closed
cpunion wants to merge 217 commits into
mainfrom
codex/wasm-r4-stdlib-text-io-20260905
Closed

cpunion wants to merge 217 commits into
mainfrom
codex/wasm-r4-stdlib-text-io-20260905

Conversation

@cpunion

@cpunion cpunion commented Sep 5, 2026 •

Copy link
Copy Markdown
Owner

Scope

Consolidated R4 development PR for xgo-dev#2152. R4 raises WebAssembly completeness in two high-priority areas:

  1. llgo run, build, and test toolchain completeness.
  2. Compatibility with the official Go js/wasm and wasip1/wasm profiles, standard-library behavior, and executable test//GOROOT acceptance.

The required LLGo profiles are EC32, EC64, WC32, GJS, and GWASI. Official-Go GJS/GWASI rows are used as reference executions. Multi-worker/parallel goroutines, WasmGC, exception handling, and JSPI/stack switching remain later work and are not claimed by this PR.

Design references: parent proposal xgo-dev/llgo#2152, typed reflection bridge sub-proposal #2556, and Emscripten host sub-proposal #2557.

Current R4 head: 05e962ddf, rebased on upstream main 309f75c5e (LLVM 22).

What this adds

Toolchain and profile completeness

  • Raw GOOS=js GOARCH=wasm and single-worker GOOS=wasip1 GOARCH=wasm builds select LLGo's linear-memory collector by default. Unsupported raw WASI threads are rejected instead of silently selecting nogc.
  • GJS reuses the selected GOROOT's public syscall/js implementation; LLGo supplies the host import/event bridge and follows the same wasm_exec.js contract.
  • Official-Go GWASI references use GOROOT's go_wasip1_wasm_exec. LLGo GWASI/WC32 run through Wasmtime with matching filesystem, working-directory, and feature configuration.
  • Source-patch inputs, Wasm ABI, native toolchain identity, test-package identity, synthetic debug locations, and goroutine stack sizing participate correctly in package/cache invalidation.
  • run, build, and test are exercised through real profile acceptance, not compile-only probes.

Runtime and standard-library completeness

  • Raw GJS/GWASI support finalizers and weak pointers with the linear collector.
  • GJS covers nested/interleaved callbacks, channel and timer waits, panic recovery, typed byte copies, JS errors, reference accounting, and exactly-once side effects under both LLGo and official Go.
  • WASI polling covers readiness, timeout, deadline changes, cancellation, and close wakeups, with explicit C32-compatible stat/poll layouts.
  • Wasm panic/caller handling preserves logical Go frames across defer, recovery, and the Wasm shadow stack, including the previously failing issue14646, issue5856, and issue22662 cases.
  • Standard-library DNS behavior returns the appropriate net.DNSError for optional localhost records rather than skipping the supported test.
  • RSA/TLS fixtures reuse generated keys where key generation itself is not under test, while retaining coverage of real key-generation paths.

Linear GC correctness

  • Runtime continuation storage can be kept live without conservatively scanning its entire capacity. This avoids retaining arbitrary objects through stale words in multi-megabyte Fiber/Asyncify buffers.
  • Every execution context publishes its compiler root chain and the used logical stack range. The active range is bounded by the current SP; suspended contexts use their saved SP.
  • Emscripten Fiber layout is checked at C compile time, including the saved stack-pointer offset.
  • The main-stack regression suspends with the only live reference encoded as a uintptr, lets another goroutine perform two collections, then verifies the object survives. This specifically covers WASI's heap-allocated logical main stack.
  • Static-data root discovery uses a one-link marker before mutable data, so immutable packed .rodata is not scanned. The marker adds only one pointer-sized boundary object and preserves LTO/GlobalDCE.
  • Native TLS package blocks are published as one exact context root while their owning Fiber is suspended. The collector does not conservatively scan the complete Emscripten TLS or Fiber allocation; five-profile stress tests force two collections before reusing a mixed scalar/pointer initializer.
  • The collector remains a single-mutator design. Embedded targets reuse the unchanged default allocation behavior; separate ESP32/ESP32-C3 gates guard against Wasm-only regressions.

Acceptance policy

  • Five startup sentinels gate the full matrices; browser execution is independently required.
  • Compatible test/ cases run in two shards per profile. Runnable GOROOT directives run in four shards per profile, with retained logs and JSON reports.
  • Exact reviewed notapplicable classifications are distinct from expected failures, failures, incomplete execution, source exclusion, and cases that never ran. Top-level test skips fail acceptance; only intentional skipped subtests inside a passing test are accepted.
  • test/cgo remains explicitly outside GOARCH=wasm because Go's cgo frontend rejects that architecture, even for LLGo's named C ABI targets. In contrast, every LLGo Wasm profile now executes the compatible runtime/cgo.Handle tests; only official-Go reference profiles exclude them.
  • The pre-existing Go 1.27 genmeth1.go SSA assertion is tracked separately in go/ssa: Go 1.27 genmeth1.go panics while instantiating generic method expressions xgo-dev/llgo#2526 with one exact xfail entry.
  • Missing required tools such as llvm-nm are errors, not silent skips.

Current validation

  • Final-head independent GC: run 34444621899 — passed Linux x86-64/ARM64, macOS, Windows, ESP32/Xtensa, and ESP32-C3/RISC-V. The collector remains explicitly single-mutator; unsupported concurrent entry is rejected.
  • Rebased-head full Wasm acceptance: run 34460966471 — browser acceptance and all five profile sentinels passed; the complete compatible test/ and GOROOT matrices are running with no observed failure.
  • Rebased-head standard-library/reference matrix: run 34460966559 — all five rows passed.
  • Rebased-head/latest-main size and build benchmark: run 34460966769 — passed. The size gate remains open: versus 309f75c5e, cprintf adds 35,526 bytes on EC32 and 37,096 bytes on EC64; fmtprintf shrinks by 136–175 KiB on the C-ABI profiles, while GJS/GWASI include the larger Go-compatible runtime and host surface.

Local post-rebase checks pass:

go test ./test/goroot -count=1
go test ./internal/build -run 'TestGC(Independent|Finalizer)Arena|TestWasmStaticRootMarker|TestWasmFuncInfoNoScanLinking' -count=1
(cd runtime && go test ./internal/gcroot ./internal/clite/emscripten ./internal/wasmcontext ./internal/runtime/tinygogc -count=1)
go build ./cmd/llgo
node dev/test_wasm_fs.mjs
git diff --check

Remaining merge gates

  • Complete all five final-head Wasm acceptance and GOROOT matrices, including the two previously borderline issue30116* timeout cases.
  • Review final-head stdlib/reference reports and browser acceptance.
  • Compare final-head cprintf, println, and fmtprintf sizes against the rebased main baseline; investigate any material increase.
  • Run the general coverage workflow on the final head and require patch/project coverage to pass. Acceptance-only code without CI execution is not considered complete.
  • After the focused Wasm gates are green, run the remaining general repository workflows and resolve review comments.

@cpunion
cpunion force-pushed the codex/wasm-r4-stdlib-reference-20260905 branch from 9ba1e90 to e79eba5 Compare September 5, 2026 14:47
@cpunion
cpunion force-pushed the codex/wasm-r4-stdlib-text-io-20260905 branch from 65b7d19 to 5ae575d Compare September 5, 2026 14:47
@cpunion
cpunion force-pushed the codex/wasm-r4-stdlib-reference-20260905 branch from e79eba5 to ecd2e88 Compare September 5, 2026 15:22
@cpunion
cpunion force-pushed the codex/wasm-r4-stdlib-text-io-20260905 branch from 5ae575d to 671e903 Compare September 5, 2026 15:22
@cpunion
cpunion force-pushed the codex/wasm-r4-stdlib-text-io-20260905 branch from 671e903 to 622addb Compare September 7, 2026 03:28
@cpunion cpunion changed the title [WASM R4] Expand stdlib acceptance to formatting, conversion and pipe IO [WASM R4] Consolidate stdlib acceptance and three-workload benchmarks Sep 7, 2026
@cpunion
cpunion changed the base branch from codex/wasm-r4-stdlib-reference-20260905 to codex/wasm-r4-main-base-20260907 September 7, 2026 03:28
@cpunion
cpunion force-pushed the codex/wasm-r4-stdlib-text-io-20260905 branch from f3ed85d to 7ccc420 Compare September 7, 2026 06:16
@cpunion cpunion changed the title [WASM R4] Consolidate stdlib acceptance and three-workload benchmarks [WASM R4] Complete single-worker profiles and full acceptance Sep 7, 2026
@cpunion cpunion changed the title [WASM R4] Complete single-worker profiles and full acceptance [WASM R4] Audit official profiles and standard-library completeness Sep 7, 2026
@cpunion
cpunion force-pushed the codex/wasm-r4-stdlib-text-io-20260905 branch from 60f1d3b to 41bde95 Compare September 7, 2026 14:42
@cpunion
cpunion force-pushed the codex/wasm-r4-stdlib-text-io-20260905 branch 3 times, most recently from 5f659cb to 03abd35 Compare September 8, 2026 15:42
@cpunion

cpunion commented Sep 8, 2026

Copy link
Copy Markdown
Owner Author

4685f7d validation update: all five profile sentinels, browser acceptance, and the standard-library reference workflow passed. Full test/ and GOROOT shards are still running. The two completed reference failures are test/go/memprofile: the assertion incorrectly applied LLGo’s pointer-width-scaled block size to official Go wasm, whose tiny allocator still uses 16-byte blocks. Local e305bad restricts that implementation-specific assertion to runtime.Compiler == "llgo"; official Go GJS/GWASI, native Go, and the LLGo GJS regression all pass. Holding this small test-only fix for the next batch while the current matrix collects other evidence. The new Wasm size benchmark completed: versus 03abd35, cprintf is approximately +0.6 KiB and fmtprintf +2 KiB; the total R4 size target is not yet met. Compile-phase diagnostics for the two known resource failures are isolated in fork run34260202320 (one job, unchanged per-case limits), currently waiting for a runner. SetGCPercent/GOGC collector integration is being validated separately; no claim of full R4 completion or full-matrix success.

@cpunion

cpunion commented Sep 8, 2026

Copy link
Copy Markdown
Owner Author

Moving the consolidated R4 contribution and further CI to xgo-dev/llgo, as requested. The active fork acceptance run has been cancelled; completed evidence and discussion remain here. The fork branch is retained, and the replacement upstream PR will be linked shortly. This is a migration, not a claim of completed R4 acceptance.

@cpunion cpunion closed this Sep 8, 2026
@cpunion

cpunion commented Sep 8, 2026

Copy link
Copy Markdown
Owner Author

Continuing R4 CI in cpunion/llgo per the updated runner-capacity decision. No replacement R4 PR was created upstream. The branch now includes e305bad, which fixes the official-Go wasm tiny-allocation assertion; focused official-Go GJS/GWASI, native Go, and LLGo GJS checks pass. The previous full acceptance run was cancelled during the attempted migration and is not claimed as completed. A fresh fork CI run will validate this head; pending GC-pacing and caller-stack experiments are not included yet.

@cpunion cpunion reopened this Sep 8, 2026
cpunion and others added 25 commits September 10, 2026 17:28
js.FuncOf queued every callback and ran it later on a new goroutine.
syscall.fsCall waits on a buffered channel after fs.write, so that
path parked, fiber-swapped, and aborted with Asyncify unreachable.

Keep host events such as setTimeout queued, but run the callback on
the current goroutine when Go is still inside Value.Call/Invoke.
Copy targets/wasm_fs.js next to emcc glue for .html, .js, and .mjs
executables, and insert the script tag into generated HTML so
syscall/fs_js.go has a browser fs/process/path host.

Map Emscripten errno 13 to ECONNABORTED. Skip a same-path copy and
close the source before replace so Windows rebuilds do not fail with
Access is denied. Do not os.Remove a directory after a failed rename;
that replaced occupied/ with a regular file and broke
TestWindowsLinkObjFilesExactOutput/copy_error. Skip the 0555 CreateTemp
probe on Windows, where directory mode bits do not block new files.
@cpunion
cpunion force-pushed the codex/wasm-r4-stdlib-text-io-20260905 branch from a4afd6c to 05e962d Compare September 10, 2026 09:29
@github-actions

Copy link
Copy Markdown

LLGo baseline benchmarks

05e962ddf030 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 20264 B +416 B / +2.1% (worse) 387 B 0 B / +0.0% 441.532 ms +28.24 ms / +6.8% (worse) 1.294 ms -21.03 us / -1.6% (better)
Linux cprintf-lto 20016 B +416 B / +2.1% (worse) 368 B 0 B / +0.0% 428.812 ms +18.11 ms / +4.4% (worse) 1.415 ms +126.6 us / +9.8% (worse)
Linux fmtprintf 1645312 B +28928 B / +1.8% (worse) 510635 B +18328 B / +3.7% (worse) 3.013 s +168.6 ms / +5.9% (worse) 3.026 ms +111.2 us / +3.8% (worse)
Linux fmtprintf-lto 1496600 B +29408 B / +2.0% (worse) 452298 B +18186 B / +4.2% (worse) 8.815 s +201.6 ms / +2.3% (worse) 2.860 ms +6.501 us / +0.2% (worse)
Linux println 62912 B +160 B / +0.3% (worse) 14806 B -233 B / -1.5% (better) 419.115 ms +13.32 ms / +3.3% (worse) 1.631 ms -82.44 us / -4.8% (better)
Linux println-lto 54488 B +176 B / +0.3% (worse) 12211 B -220 B / -1.8% (better) 603.763 ms +26.54 ms / +4.6% (worse) 1.608 ms -52.83 us / -3.2% (better)
macOS cprintf 84480 B 0 B / +0.0% 17533 B +416 B / +2.4% (worse) 606.097 ms -489.5 ms / -44.7% (better) 2.509 ms -4.806 ms / -65.7% (better)
macOS cprintf-lto 84288 B 0 B / +0.0% 13297 B +416 B / +3.2% (worse) 623.849 ms -215.4 ms / -25.7% (better) 2.626 ms -1.498 ms / -36.3% (better)
macOS fmtprintf 1490672 B +17184 B / +1.2% (worse) 889816 B +23188 B / +2.7% (worse) 2.992 s -174.1 ms / -5.5% (better) 6.740 ms +622 us / +10.2% (worse)
macOS fmtprintf-lto 1175888 B +16448 B / +1.4% (worse) 861852 B +21784 B / +2.6% (worse) 7.052 s -1.482 s / -17.4% (better) 5.314 ms +454 us / +9.3% (worse)
macOS println 114864 B 0 B / +0.0% 35373 B +272 B / +0.8% (worse) 638.278 ms +18.66 ms / +3.0% (worse) 3.930 ms +458 us / +13.2% (worse)
macOS println-lto 118736 B 0 B / +0.0% 33009 B +280 B / +0.9% (worse) 673.411 ms -230.7 ms / -25.5% (better) 3.651 ms -292.9 us / -7.4% (better)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 915.798 ms -37.64 ms / -3.9% (better) 2.816 ms +22.2 us / +0.8% (worse)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 948.457 ms -22.33 ms / -2.3% (better) 2.699 ms -139.7 us / -4.9% (better)
Windows MinGW fmtprintf 1910784 B +24576 B / +1.3% (worse) 608694 B +15664 B / +2.6% (worse) 3.375 s +33.03 ms / +1.0% (worse) 7.304 ms +59.5 us / +0.8% (worse)
Windows MinGW fmtprintf-lto 1958912 B +31232 B / +1.6% (worse) 556358 B +14992 B / +2.8% (worse) 8.518 s +190.5 ms / +2.3% (worse) 6.856 ms +263 us / +4.0% (worse)
Windows MinGW println 71680 B 0 B / +0.0% 23638 B -336 B / -1.4% (better) 915.135 ms -10.07 ms / -1.1% (better) 5.501 ms -88.8 us / -1.6% (better)
Windows MinGW println-lto 65024 B -512 B / -0.8% (better) 20454 B -272 B / -1.3% (better) 1.094 s +4.377 ms / +0.4% (worse) 5.521 ms -297.5 us / -5.1% (better)
Windows MinGW 386 cprintf 43008 B +5120 B / +13.5% (worse) 5326 B 0 B / +0.0% 1.182 s +60.53 ms / +5.4% (worse) 5.195 ms +211.4 us / +4.2% (worse)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.159 s +26.24 ms / +2.3% (worse) 5.139 ms +92 us / +1.8% (worse)
Windows MinGW 386 fmtprintf 1880576 B +37376 B / +2.0% (worse) 485454 B +18208 B / +3.9% (worse) 4.345 s +264.9 ms / +6.5% (worse) 11.854 ms +328.2 us / +2.8% (worse)
Windows MinGW 386 fmtprintf-lto 2217472 B +53760 B / +2.5% (worse) 460242 B +13240 B / +3.0% (worse) 10.314 s +540.8 ms / +5.5% (worse) 11.142 ms -1.026 ms / -8.4% (better)
Windows MinGW 386 println 92160 B +5632 B / +6.5% (worse) 20182 B -192 B / -0.9% (better) 1.112 s +18.46 ms / +1.7% (worse) 8.680 ms -232 us / -2.6% (better)
Windows MinGW 386 println-lto 70656 B 0 B / +0.0% 18050 B -212 B / -1.2% (better) 1.310 s -12.85 ms / -1.0% (better) 8.592 ms +24.3 us / +0.3% (worse)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4436 B 0 B / +0.0% 1.450 s -52.71 ms / -3.5% (better) 6.106 ms -1.128 ms / -15.6% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4368 B 0 B / +0.0% 1.467 s -29.65 ms / -2.0% (better) 6.432 ms -230.9 us / -3.5% (better)
Windows MinGW ARM64 fmtprintf 1803776 B +29696 B / +1.7% (worse) 526184 B +20272 B / +4.0% (worse) 4.240 s +28.14 ms / +0.7% (worse) 13.354 ms +569.6 us / +4.5% (worse)
Windows MinGW ARM64 fmtprintf-lto 1890304 B +36352 B / +2.0% (worse) 492308 B +19284 B / +4.1% (worse) 10.029 s +299.4 ms / +3.1% (worse) 12.405 ms -686.8 us / -5.2% (better)
Windows MinGW ARM64 println 68608 B 0 B / +0.0% 22712 B -168 B / -0.7% (better) 1.414 s -36.87 ms / -2.5% (better) 11.256 ms -99.7 us / -0.9% (better)
Windows MinGW ARM64 println-lto 65024 B 0 B / +0.0% 20000 B -160 B / -0.8% (better) 1.616 s -34.31 ms / -2.1% (better) 10.586 ms -310.9 us / -2.9% (better)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65782 B 0 B / +0.0% 928.965 ms -170.6 ms / -15.5% (better) 3.382 ms -570.6 us / -14.4% (better)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65718 B 0 B / +0.0% 1.128 s +192.9 ms / +20.6% (worse) 4.031 ms +663.2 us / +19.7% (worse)
Windows MSVC fmtprintf 1640448 B +24576 B / +1.5% (worse) 704198 B +15664 B / +2.3% (worse) 3.711 s +31.82 ms / +0.9% (worse) 10.078 ms +1.154 ms / +12.9% (worse)
Windows MSVC fmtprintf-lto 1647616 B +32256 B / +2.0% (worse) 659126 B +15040 B / +2.3% (worse) 8.889 s -100.7 ms / -1.1% (better) 10.180 ms +865.7 us / +9.3% (worse)
Windows MSVC println 192512 B -512 B / -0.3% (better) 119046 B -336 B / -0.3% (better) 922.909 ms -1.322 ms / -0.1% (better) 6.953 ms -123.8 us / -1.7% (better)
Windows MSVC println-lto 189952 B 0 B / +0.0% 116438 B -240 B / -0.2% (better) 1.108 s +1.387 ms / +0.1% (worse) 7.045 ms -212.2 us / -2.9% (better)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 963.145 ms -51.73 ms / -5.1% (better) 6.487 ms +170.1 us / +2.7% (worse)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 988.649 ms -212 ms / -17.7% (better) 5.425 ms -840 us / -13.4% (better)
Windows MSVC 386 fmtprintf 1207296 B +21504 B / +1.8% (worse) 468789 B +18192 B / +4.0% (worse) 4.357 s +225.4 ms / +5.5% (worse) 12.014 ms -1.129 ms / -8.6% (better)
Windows MSVC 386 fmtprintf-lto 1251840 B +24576 B / +2.0% (worse) 439481 B +13008 B / +3.1% (worse) 9.498 s +785.1 ms / +9.0% (worse) 12.556 ms +903.7 us / +7.8% (worse)
Windows MSVC 386 println 34816 B 0 B / +0.0% 19025 B -208 B / -1.1% (better) 954.640 ms -31.86 ms / -3.2% (better) 9.446 ms -1.072 ms / -10.2% (better)
Windows MSVC 386 println-lto 32768 B 0 B / +0.0% 17191 B -160 B / -0.9% (better) 1.147 s -33.82 ms / -2.9% (better) 9.754 ms -1.196 ms / -10.9% (better)
Windows MSVC ARM64 cprintf 11264 B 0 B / +0.0% 3976 B 0 B / +0.0% 2.062 s +42.47 ms / +2.1% (worse) 7.060 ms +569.2 us / +8.8% (worse)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 3868 B 0 B / +0.0% 2.103 s +73.34 ms / +3.6% (worse) 7.586 ms +793.1 us / +11.7% (worse)
Windows MSVC ARM64 fmtprintf 1392128 B +29184 B / +2.1% (worse) 525380 B +19920 B / +3.9% (worse) 6.571 s +208.7 ms / +3.3% (worse) 14.136 ms +386 us / +2.8% (worse)
Windows MSVC ARM64 fmtprintf-lto 1421824 B +35328 B / +2.5% (worse) 493028 B +19296 B / +4.1% (worse) 15.799 s +268.3 ms / +1.7% (worse) 13.765 ms -745.2 us / -5.1% (better)
Windows MSVC ARM64 println 41984 B 0 B / +0.0% 22072 B -144 B / -0.6% (better) 2.011 s +17.35 ms / +0.9% (worse) 12.093 ms +409.1 us / +3.5% (worse)
Windows MSVC ARM64 println-lto 39936 B -512 B / -1.3% (better) 19884 B -160 B / -0.8% (better) 2.340 s -6.579 ms / -0.3% (better) 11.794 ms -545.6 us / -4.4% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 13.370 ns/op 0 ns/op / +0.0%
Linux BenchmarkMergeCompilerFlags 141.200 ns/op +1.3 ns/op / +0.9% (worse)
Linux BenchmarkMergeLinkerFlags 89.460 ns/op +0.43 ns/op / +0.5% (worse)
Linux BenchmarkChannelBuffered 52.950 ns/op +13.97 ns/op / +35.8% (worse)
Linux BenchmarkChannelHandoff 26300 ns/op -1187 ns/op / -4.3% (better)
Linux BenchmarkDefer 50.730 ns/op +1.31 ns/op / +2.7% (worse)
Linux BenchmarkDirectCall 1.557 ns/op -0.005 ns/op / -0.3% (better)
Linux BenchmarkGlobalRead 1.866 ns/op -0.002 ns/op / -0.1% (better)
Linux BenchmarkGlobalWrite 2.480 ns/op 0 ns/op / +0.0%
Linux BenchmarkGoroutine 31861 ns/op -568 ns/op / -1.8% (better)
Linux BenchmarkInterfaceCall 8.404 ns/op +0.3 ns/op / +3.7% (worse)
Linux BenchmarkRuntimeGetG 1.867 ns/op -0.655 ns/op / -26.0% (better)
macOS BenchmarkLookupPCRandom 16.560 ns/op +4.13 ns/op / +33.2% (worse)
macOS BenchmarkMergeCompilerFlags 129.800 ns/op +6.8 ns/op / +5.5% (worse)
macOS BenchmarkMergeLinkerFlags 104.100 ns/op +29.93 ns/op / +40.4% (worse)
macOS BenchmarkChannelBuffered 57.130 ns/op +32 ns/op / +127.3% (worse)
macOS BenchmarkChannelHandoff 8560 ns/op +58 ns/op / +0.7% (worse)
macOS BenchmarkDefer 64.650 ns/op +31.84 ns/op / +97.0% (worse)
macOS BenchmarkDirectCall 1.398 ns/op +0.331 ns/op / +31.0% (worse)
macOS BenchmarkGlobalRead 1.442 ns/op +0.284 ns/op / +24.5% (worse)
macOS BenchmarkGlobalWrite 1.777 ns/op +0.695 ns/op / +64.2% (worse)
macOS BenchmarkGoroutine 62622 ns/op +30650 ns/op / +95.9% (worse)
macOS BenchmarkInterfaceCall 6.733 ns/op +2.079 ns/op / +44.7% (worse)
macOS BenchmarkRuntimeGetG 3.297 ns/op +1.062 ns/op / +47.5% (worse)
Windows MinGW BenchmarkLookupPCRandom 9.576 ns/op +0.032 ns/op / +0.3% (worse)
Windows MinGW BenchmarkMergeCompilerFlags 385.700 ns/op -2.9 ns/op / -0.7% (better)
Windows MinGW BenchmarkMergeLinkerFlags 349.700 ns/op +16.5 ns/op / +5.0% (worse)
Windows MinGW BenchmarkChannelBuffered 36.840 ns/op +6.45 ns/op / +21.2% (worse)
Windows MinGW BenchmarkChannelHandoff 1136 ns/op +19 ns/op / +1.7% (worse)
Windows MinGW BenchmarkDefer 46.300 ns/op +2.18 ns/op / +4.9% (worse)
Windows MinGW BenchmarkDirectCall 1.357 ns/op 0 ns/op / +0.0%
Windows MinGW BenchmarkGlobalRead 1.358 ns/op 0 ns/op / +0.0%
Windows MinGW BenchmarkGlobalWrite 2.165 ns/op -0.001 ns/op / -0.04617% (better)
Windows MinGW BenchmarkGoroutine 56820 ns/op -1116 ns/op / -1.9% (better)
Windows MinGW BenchmarkInterfaceCall 7.988 ns/op -0.01 ns/op / -0.1% (better)
Windows MinGW BenchmarkRuntimeGetG 1.694 ns/op +0.285 ns/op / +20.2% (worse)
Windows MinGW 386 BenchmarkLookupPCRandom 26.620 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 774.600 ns/op +12.4 ns/op / +1.6% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 702.100 ns/op +15.3 ns/op / +2.2% (worse)
Windows MinGW 386 BenchmarkChannelBuffered 45.530 ns/op +1.43 ns/op / +3.2% (worse)
Windows MinGW 386 BenchmarkChannelHandoff 965.900 ns/op -52.1 ns/op / -5.1% (better)
Windows MinGW 386 BenchmarkDefer 43.460 ns/op -1.04 ns/op / -2.3% (better)
Windows MinGW 386 BenchmarkDirectCall 1.549 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGlobalRead 1.548 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalWrite 7.779 ns/op -0.003 ns/op / -0.03855% (better)
Windows MinGW 386 BenchmarkGoroutine 88288 ns/op +1153 ns/op / +1.3% (worse)
Windows MinGW 386 BenchmarkInterfaceCall 9.630 ns/op +0.02 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkRuntimeGetG 2.486 ns/op +0.314 ns/op / +14.5% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.050 ns/op -0.05 ns/op / -0.4% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 569.700 ns/op -3.6 ns/op / -0.6% (better)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 532.300 ns/op -11.9 ns/op / -2.2% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 44.790 ns/op -0.99 ns/op / -2.2% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 2321 ns/op +164 ns/op / +7.6% (worse)
Windows MinGW ARM64 BenchmarkDefer 54.910 ns/op +1.23 ns/op / +2.3% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.589 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGlobalRead 0.884 ns/op +0.2204 ns/op / +33.2% (worse)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.590 ns/op +0.0001 ns/op / +0.01696% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 58978 ns/op +917 ns/op / +1.6% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.715 ns/op -0.003 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.800 ns/op -0.004 ns/op / -0.2% (better)
Windows MSVC BenchmarkLookupPCRandom 13.170 ns/op +0.26 ns/op / +2.0% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 633.200 ns/op +14.7 ns/op / +2.4% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 563.400 ns/op +34.7 ns/op / +6.6% (worse)
Windows MSVC BenchmarkChannelBuffered 47.240 ns/op +10.98 ns/op / +30.3% (worse)
Windows MSVC BenchmarkChannelHandoff 1064 ns/op -55 ns/op / -4.9% (better)
Windows MSVC BenchmarkDefer 53.580 ns/op -1.52 ns/op / -2.8% (better)
Windows MSVC BenchmarkDirectCall 1.548 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC BenchmarkGlobalRead 1.546 ns/op -0.315 ns/op / -16.9% (better)
Windows MSVC BenchmarkGlobalWrite 2.470 ns/op +0.021 ns/op / +0.9% (worse)
Windows MSVC BenchmarkGoroutine 83905 ns/op +543 ns/op / +0.7% (worse)
Windows MSVC BenchmarkInterfaceCall 9.289 ns/op -0.318 ns/op / -3.3% (better)
Windows MSVC BenchmarkRuntimeGetG 2.167 ns/op -0.005 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkLookupPCRandom 26.590 ns/op +0.09 ns/op / +0.3% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 775.800 ns/op +67 ns/op / +9.5% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 715.900 ns/op +36.4 ns/op / +5.4% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 44.540 ns/op +1.79 ns/op / +4.2% (worse)
Windows MSVC 386 BenchmarkChannelHandoff 1004 ns/op -51 ns/op / -4.8% (better)
Windows MSVC 386 BenchmarkDefer 46.040 ns/op -2.01 ns/op / -4.2% (better)
Windows MSVC 386 BenchmarkDirectCall 1.547 ns/op -0.003 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkGlobalRead 1.550 ns/op -0.312 ns/op / -16.8% (better)
Windows MSVC 386 BenchmarkGlobalWrite 7.790 ns/op +0.009 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGoroutine 88210 ns/op +4744 ns/op / +5.7% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 9.669 ns/op +0.06 ns/op / +0.6% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 2.165 ns/op +0.238 ns/op / +12.4% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.060 ns/op -0.01 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 559.300 ns/op +10.9 ns/op / +2.0% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 536 ns/op +14.4 ns/op / +2.8% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 43.660 ns/op +0.19 ns/op / +0.4% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 3120 ns/op -190 ns/op / -5.7% (better)
Windows MSVC ARM64 BenchmarkDefer 60.550 ns/op -0.44 ns/op / -0.7% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op -0.0001 ns/op / -0.01696% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.664 ns/op +0.0003 ns/op / +0.04521% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.794 ns/op +0.031 ns/op / +0.8% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 56768 ns/op -371 ns/op / -0.6% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.728 ns/op +0.01 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.769 ns/op -0.001 ns/op / -0.1% (better)

Timer runtime benchmarks

Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 599.400 ns/op -6.3 ns/op / -1.0% (better)
Linux AfterFuncZeroDelivery/LLGo 53579 ns/op +408 ns/op / +0.8% (worse)
Linux CreateStop/Go 148.600 ns/op -1.1 ns/op / -0.7% (better)
Linux CreateStop/LLGo 1130 ns/op +261.2 ns/op / +30.1% (worse)
Linux RearmStopped/Go 54.190 ns/op -1.12 ns/op / -2.0% (better)
Linux RearmStopped/LLGo 743.700 ns/op -343.3 ns/op / -31.6% (better)
Linux ResetActive/Go 41.540 ns/op -0.03 ns/op / -0.1% (better)
Linux ResetActive/LLGo 352.700 ns/op -25.8 ns/op / -6.8% (better)
Linux ResetHeap1024/Go 41.410 ns/op -0.03 ns/op / -0.1% (better)
Linux ResetHeap1024/LLGo 152.300 ns/op +7 ns/op / +4.8% (worse)
macOS AfterFuncZeroDelivery/Go 700.600 ns/op +184.6 ns/op / +35.8% (worse)
macOS AfterFuncZeroDelivery/LLGo 106465 ns/op +22151 ns/op / +26.3% (worse)
macOS CreateStop/Go 171.800 ns/op -5.8 ns/op / -3.3% (better)
macOS CreateStop/LLGo 889.500 ns/op -43.5 ns/op / -4.7% (better)
macOS RearmStopped/Go 75.700 ns/op +7.1 ns/op / +10.3% (worse)
macOS RearmStopped/LLGo 481.100 ns/op -124.6 ns/op / -20.6% (better)
macOS ResetActive/Go 52.910 ns/op -10.26 ns/op / -16.2% (better)
macOS ResetActive/LLGo 177.200 ns/op -0.8 ns/op / -0.4% (better)
macOS ResetHeap1024/Go 75.040 ns/op +26.88 ns/op / +55.8% (worse)
macOS ResetHeap1024/LLGo 139.400 ns/op +37.3 ns/op / +36.5% (worse)
Windows MinGW AfterFuncZeroDelivery/Go 372.900 ns/op -18.3 ns/op / -4.7% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 99310 ns/op -1335 ns/op / -1.3% (better)
Windows MinGW CreateStop/Go 89.640 ns/op -1.7 ns/op / -1.9% (better)
Windows MinGW CreateStop/LLGo 390.900 ns/op +20.6 ns/op / +5.6% (worse)
Windows MinGW RearmStopped/Go 24.480 ns/op +0.02 ns/op / +0.1% (worse)
Windows MinGW RearmStopped/LLGo 247.700 ns/op -6.9 ns/op / -2.7% (better)
Windows MinGW ResetActive/Go 14.770 ns/op -0.1 ns/op / -0.7% (better)
Windows MinGW ResetActive/LLGo 126.700 ns/op -2.3 ns/op / -1.8% (better)
Windows MinGW ResetHeap1024/Go 14.770 ns/op -0.12 ns/op / -0.8% (better)
Windows MinGW ResetHeap1024/LLGo 117.100 ns/op -0.6 ns/op / -0.5% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 960.200 ns/op -1 ns/op / -0.1% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 190286 ns/op +705 ns/op / +0.4% (worse)
Windows MinGW 386 CreateStop/Go 192.200 ns/op -1 ns/op / -0.5% (better)
Windows MinGW 386 CreateStop/LLGo 2146 ns/op -30 ns/op / -1.4% (better)
Windows MinGW 386 RearmStopped/Go 63.540 ns/op +0.13 ns/op / +0.2% (worse)
Windows MinGW 386 RearmStopped/LLGo 382.900 ns/op +9.3 ns/op / +2.5% (worse)
Windows MinGW 386 ResetActive/Go 39.070 ns/op +0.01 ns/op / +0.0256% (worse)
Windows MinGW 386 ResetActive/LLGo 946.400 ns/op -54.6 ns/op / -5.5% (better)
Windows MinGW 386 ResetHeap1024/Go 39.460 ns/op 0 ns/op / +0.0%
Windows MinGW 386 ResetHeap1024/LLGo 195.400 ns/op -0.4 ns/op / -0.2% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 663.400 ns/op -1 ns/op / -0.2% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 126531 ns/op -8486 ns/op / -6.3% (better)
Windows MinGW ARM64 CreateStop/Go 200.100 ns/op +0.7 ns/op / +0.4% (worse)
Windows MinGW ARM64 CreateStop/LLGo 412.400 ns/op +15.2 ns/op / +3.8% (worse)
Windows MinGW ARM64 RearmStopped/Go 70.560 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 RearmStopped/LLGo 288.600 ns/op +5.6 ns/op / +2.0% (worse)
Windows MinGW ARM64 ResetActive/Go 31.040 ns/op -0.04 ns/op / -0.1% (better)
Windows MinGW ARM64 ResetActive/LLGo 129.100 ns/op -4.6 ns/op / -3.4% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.150 ns/op +0.09 ns/op / +0.3% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 140.100 ns/op +1.5 ns/op / +1.1% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 563.100 ns/op +6.7 ns/op / +1.2% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 157350 ns/op +1165 ns/op / +0.7% (worse)
Windows MSVC CreateStop/Go 117.700 ns/op +0.9 ns/op / +0.8% (worse)
Windows MSVC CreateStop/LLGo 460 ns/op +13.9 ns/op / +3.1% (worse)
Windows MSVC RearmStopped/Go 31.310 ns/op -0.58 ns/op / -1.8% (better)
Windows MSVC RearmStopped/LLGo 299.500 ns/op -4.3 ns/op / -1.4% (better)
Windows MSVC ResetActive/Go 20.220 ns/op +0.06 ns/op / +0.3% (worse)
Windows MSVC ResetActive/LLGo 169.600 ns/op +10.4 ns/op / +6.5% (worse)
Windows MSVC ResetHeap1024/Go 20.460 ns/op +0.04 ns/op / +0.2% (worse)
Windows MSVC ResetHeap1024/LLGo 138.900 ns/op -2.8 ns/op / -2.0% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 957.300 ns/op +9.4 ns/op / +1.0% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 180994 ns/op +9070 ns/op / +5.3% (worse)
Windows MSVC 386 CreateStop/Go 195.400 ns/op +3.8 ns/op / +2.0% (worse)
Windows MSVC 386 CreateStop/LLGo 1756 ns/op +137 ns/op / +8.5% (worse)
Windows MSVC 386 RearmStopped/Go 63.610 ns/op +0.2 ns/op / +0.3% (worse)
Windows MSVC 386 RearmStopped/LLGo 337.300 ns/op -7.6 ns/op / -2.2% (better)
Windows MSVC 386 ResetActive/Go 38.980 ns/op -0.08 ns/op / -0.2% (better)
Windows MSVC 386 ResetActive/LLGo 949 ns/op -21.5 ns/op / -2.2% (better)
Windows MSVC 386 ResetHeap1024/Go 39.490 ns/op +0.06 ns/op / +0.2% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 179.800 ns/op +1.2 ns/op / +0.7% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 668.400 ns/op -6.4 ns/op / -0.9% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 145633 ns/op +3516 ns/op / +2.5% (worse)
Windows MSVC ARM64 CreateStop/Go 199.600 ns/op -14.6 ns/op / -6.8% (better)
Windows MSVC ARM64 CreateStop/LLGo 407.400 ns/op -1.7 ns/op / -0.4% (better)
Windows MSVC ARM64 RearmStopped/Go 70.620 ns/op -0.05 ns/op / -0.1% (better)
Windows MSVC ARM64 RearmStopped/LLGo 295.900 ns/op +5.7 ns/op / +2.0% (worse)
Windows MSVC ARM64 ResetActive/Go 31 ns/op -0.11 ns/op / -0.4% (better)
Windows MSVC ARM64 ResetActive/LLGo 136.500 ns/op -7.1 ns/op / -4.9% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.080 ns/op -0.05 ns/op / -0.2% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 144.500 ns/op +1 ns/op / +0.7% (worse)

Compared with 309f75c5e5b3 measured in the same runner job.

@github-actions

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

05e962ddf030 | workflow run | long-term charts

WebAssembly output sizes

Profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/ec32/LLGo 148741 B +35526 B / +31.4% (worse) 80425 B +9701 B / +13.7% (worse)
cprintf/ec64/LLGo 155891 B +37096 B / +31.2% (worse) 84170 B +10149 B / +13.7% (worse)
cprintf/js/LLGo 146203 B +79453 B / +119.0% (worse) 80011 B +11512 B / +16.8% (worse)
cprintf/wasip1/LLGo 149164 B +75576 B / +102.7% (worse) 0 B 0 B / 0.0%
cprintf/wc32/LLGo 149386 B +32282 B / +27.6% (worse) 0 B 0 B / 0.0%
ec32/LLGo 148113 B +35630 B / +31.7% (worse) 80425 B +9701 B / +13.7% (worse)
ec64/LLGo 155226 B +37198 B / +31.5% (worse) 84170 B +10149 B / +13.7% (worse)
fmtprintf/ec32/LLGo 2752823 B -146485 B / -5.1% (better) 106316 B +8889 B / +9.1% (worse)
fmtprintf/ec64/LLGo 2912524 B -174870 B / -5.7% (better) 117510 B +9074 B / +8.4% (worse)
fmtprintf/js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/js/LLGo 2750155 B +427717 B / +18.4% (worse) 91271 B -4091 B / -4.3% (better)
fmtprintf/wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/wasip1/LLGo 2593551 B +478302 B / +22.6% (worse) 0 B 0 B / 0.0%
fmtprintf/wc32/LLGo 2569011 B -136082 B / -5.0% (better) 0 B 0 B / 0.0%
js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
js/LLGo 145767 B +79695 B / +120.6% (worse) 80011 B +11512 B / +16.8% (worse)
wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
wasip1/LLGo 148584 B +76049 B / +104.8% (worse) 0 B 0 B / 0.0%
wc32/LLGo 148803 B +32375 B / +27.8% (worse) 0 B 0 B / 0.0%

LLGo WebAssembly build measurements

Profile Build vs base
cprintf/ec32 5.975 s +1.148 s / +23.8% (worse)
cprintf/ec64 6.019 s +1.535 s / +34.2% (worse)
cprintf/js 6.063 s +2.178 s / +56.1% (worse)
cprintf/wasip1 5.029 s +2.382 s / +89.9% (worse)
cprintf/wc32 5.035 s +1.685 s / +50.3% (worse)
ec32 6.153 s +1.305 s / +26.9% (worse)
ec64 6.235 s +1.765 s / +39.5% (worse)
fmtprintf/ec32 40.814 s +3.526 s / +9.5% (worse)
fmtprintf/ec64 36.579 s -147.7 ms / -0.4% (better)
fmtprintf/js 38.755 s +8.873 s / +29.7% (worse)
fmtprintf/wasip1 35.496 s +9.524 s / +36.7% (worse)
fmtprintf/wc32 35.319 s -3.228 s / -8.4% (better)
js 6.277 s +2.336 s / +59.3% (worse)
wasip1 5.114 s +2.368 s / +86.2% (worse)
wc32 4.993 s +1.643 s / +49.0% (worse)

Compared with 309f75c5e5b3 measured in the same runner job.

@cpunion

cpunion commented Sep 10, 2026

Copy link
Copy Markdown
Owner Author

R1-R3 are now merged and the profile model has been simplified. This oversized R4 aggregate is superseded by W1 in #246. The branch is retained as a migration source; only re-audited necessary changes will be ported.

@cpunion cpunion closed this Sep 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants