Skip to content

[WASM W2-B] Complete single-worker runtime and aggregate lowering - #2602

Merged
xushiwei merged 33 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-w2-runtime-lowering-20260913
Sep 17, 2026
Merged

xushiwei merged 33 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-w2-runtime-lowering-20260913

Conversation

@cpunion

@cpunion cpunion commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Tracks #2152. Depends on W2-A #2601. Replaces the closed #2580 with a clean review thread; its fixes remain in the branch.

Scope

  • Complete single-worker logical-goroutine state, recover/repanic tracebacks, caller identity, recoverable nil checks, and physical Memory32 overflow checks.
  • Scan registered tinygogc finalizer candidates with linear ready-queue traversal; preserve finalizer/cleanup ordering and collection accounting.
  • Root fixed Fiber stacks for their owning logical goroutines, with acceptance-derived Memory32/Memory64 budgets.
  • Lower large aggregate snapshots through bounded rooted copies and localize stack-slot addresses before SjLj expansion. This avoids LLVM scalarization/liveness blowups without raising container limits.
  • Add focused standard-library acceptance for J32/GoJS (wasm32 with the Go-compatible JavaScript host), J32/Emscripten (wasm32 with the Emscripten JavaScript host), J64/Emscripten (wasm64 with the Emscripten JavaScript host), and W32/WASI (wasm32 with WASI Preview 1), with official Go host references.

This does not add multi-worker scheduling or parallel goroutines. Complete applicable test/** and GOROOT inventory accounting belong to W3.

Review boundary

Head 3b3a5e599051, based on W2-A 1eb979a9515f and main 2db247e43848. The dedicated range is 1eb979a9515f..3b3a5e599051: 33 commits, 96 files, +4,392/-292. The original 28 feature commits remain patch-equivalent; follow-ups document event-stack ordering, test aggregate convergence at 1/8/32 nesting levels, correct conservative-GC assumptions in lifecycle acceptance, and use the hosted documentation-check runner proposed independently in #2604. The latest follow-up withdraws the caller-literal size optimization and adds a focused allocation-overhead regression; it changes only two files. The full diff against main includes the prerequisite profile, host, and reflection layers.

Validation

Downstream full GOROOT acceptance in W3 exposed two allocation-count regressions in the previous head: closure.go and fixedbugs/issue4667.go. Allocation stack tracing identifies two 16-byte heap allocations per RecordPanicLocationWasm call: address-taken temporary string headers introduced by the caller-literal size optimization escape during lowering. This follow-up restores unsafe.String without changing the six-scalar ABI or noinline boundary. Its regression compares wrapper calls with the same underlying operations: a separate trace confirms that the push already allocates a 72-byte frame snapshot, so the contract is no additional wrapper allocation, not zero total shadow-stack allocation. No GOROOT xfail is added.

The corrected isolated GOROOT validation passes both original allocation cases in 30 fresh processes on each of four providers: all 240 execution records have exit status zero and all 240 logs have the expected empty output. Each unmodified source is staged separately, as in the formal GOROOT runner. An earlier raw GoJS probe was rejected because it picked up GOROOT's neighboring cmplxdivide.c; its zero exit status was not used as acceptance evidence. The same run also reproduced the existing init1.go Memory32/WASI xfail, which remains unchanged. The final public test-command gate passes: the corrected overhead regression passes on all four paths, as do complete test, compile-only test -c, and raw JavaScript/WASI run checks. The old implementation fails the same regression (38 allocations through wrappers versus 20 through underlying helpers). Current head 3b3a5e599051 has now completed all 71 checks successfully, with 98.24% patch coverage against a 94.38% target and no unresolved review threads. Only the conditional non-PR release publication job is skipped. W3 is rebased afterward for full-corpus acceptance.

The runtime and size follow-up passes 300 lifecycle executions, the complete runtime/browser gate, and paired output/size checks. Its over-strict zero-total-allocation test was subsequently replaced by the direct-helper comparison above. Restoring correct allocation behavior adds 1,215 B to GoJS and Memory32 Emscripten cprintf, 1,544 B to Memory64 Emscripten, and 1,260 B to WASI, relative to the withdrawn optimization. println adds 1,202–1,555 B and fmtprintf 998–1,164 B in the same paired builds. No native implementation changes in this follow-up.

The branch inherits W1's llgo env profile/provider fix and W2-A's explicit browser-result protocol. LLVM 22 compiler construction, the complete internal/abi suite, scalar caller-ABI regressions, and git diff --check pass locally.

The finalizer-focused validation passes 200 fresh lifecycle executions each for J32/Emscripten (wasm32 with the Emscripten JavaScript host), J64/Emscripten (wasm64 with the Emscripten JavaScript host), and W32/WASI (wasm32 with WASI Preview 1), followed by the complete runtime gate including panic/repanic tracebacks and both real-browser providers. All 600 retained logs have the expected completion marker and no unexpected panic/fatal output. After restoring the caller helpers, all 300 additional lifecycle logs were checked again. No diagnostic runtime instrumentation is included in the contribution or its clean validation candidates.

The earlier caller-literal validation measured approximately 1.2 KB/1.5 KB savings for Memory32/Memory64. That optimization is being withdrawn because it introduces heap allocations on the instrumentation path; passing runtime output checks did not establish its allocation behavior. Correct allocation semantics take precedence over these byte savings. A fresh paired size comparison is included in the focused follow-up validation.

Native size guardrail

Native size growth and reproducible performance regressions remain acceptance concerns. Native cprintf should stay unchanged where possible, and other native growth should remain small and explained. The current-head paired benchmark at 3b3a5e599051 measures these executable-file changes against main 2db247e43848 (default, non-LTO builds):

Platform cprintf println fmtprintf
Linux +160 B +144 B +3,376 B
macOS 0 B 0 B +416 B
Windows MinGW amd64 0 B 0 B +3,584 B
Windows MSVC amd64 0 B 0 B +3,072 B
Windows MSVC arm64 0 B 0 B +3,072 B
Windows MinGW arm64 0 B 0 B +3,584 B
Windows MSVC 386 0 B 0 B +4,096 B
Windows MinGW 386 0 B 0 B +4,096 B

All eight native configurations have completed at the current head. All 192 file/text/data/BSS measurements across the six programs per platform, including LTO, match the previous fcce9aaf1c2f head exactly. LTO cprintf has the same +160 B/zero pattern. Retained Linux ELF files identify the 160 B as ten net additional 16-byte function-information index records: .rodata grows, while executable code and the actual .bss section do not. The existing pre-link metadata index can retain records for dead runtime helpers; this does not mean those helpers' machine code is linked.

The paired native symbol audit attributes all 767 B of Linux fmtprintf text growth: 385 B net in reflect argument/result conversion, 224 B in libffi signature/element handling, and 158 B in caller/panic frame handling. The 1,760 B .rodata increase is predominantly function metadata: 22 additional table/index entries and 31 additional strings account for 1,728 B through the table, string pool, offsets, and index; the remaining 32 B is other read-only content/alignment. Unwind metadata adds 320 B and function-entry metadata 56 B, totaling the measured 2,136 B data growth. LTO reduces the text delta to 539 B and data delta to 1,852 B. Native file growth is therefore bounded and attributed, not evidence of a bulk WASM runtime being linked. Single-sample program run timings are not evidence of performance equivalence; the new formal benchmark round also checks earlier noisy timer/channel medians.

WASM growth accounting

WASM growth is recorded and evaluated against the capabilities it adds, rather than an unconditional zero-growth requirement. Further Wasm-specific size optimization is deferred to follow-up work and will not reopen these prerequisite PRs for small byte savings. The current-head paired benchmark at 3b3a5e599051 versus main 2db247e43848 measures these module-byte changes:

Provider cprintf println fmtprintf
J32/Emscripten (wasm32 with the Emscripten JavaScript host) +22,109 B +22,115 B +336,585 B
J64/Emscripten (wasm64 with the Emscripten JavaScript host) +4,810 B +4,906 B -99,671 B
W32/WASI (wasm32 with WASI Preview 1) +20,584 B +20,614 B +186,370 B

Compared with the previous fcce9aaf1c2f head, all 20 LLGo modules increase by 1,202–1,553 B after restoring allocation-free string construction; glue and official-Go reference sizes are unchanged. The earlier focused probe uses different fixture/output paths, so its byte deltas are retained separately rather than substituted for these formal measurements.

For J32/Emscripten cprintf at the previous head, about 17.3 KB enters in W1's Go64-on-Memory32 model; W2-A is effectively unchanged, and the remaining approximately 3.6 KB enters in W2-B. Symbol inspection associates the Memory32 increase with widened Go operations and checked Go/C boundaries, while W2-B adds panic/caller tracking. These are measured stage boundaries and an attribution based on symbol inspection, not proof that every byte is unavoidable. Raw GOOS=js/wasip1 baselines also selected different compatibility runtime/provider paths, so their larger differences are not like-for-like ABI comparisons. W3's published fe7f99a301c3 module-size results match the previous fcce9aaf1c2f head exactly; its acceptance layer added no module bytes before the pending sequential rebase.

The previous head fcce9aaf1c2f completed all 71 checks, with no unresolved review threads. Those narrower green gates did not detect the allocation-count regression subsequently found by W3's full corpus. Historical acceptance and measurements remain in #2580; previous green runs are not reused as current-head evidence.

At that previous head, the four LLGo standard-library slices, both official Go reference slices, runtime/browser integration, and public test-command checks passed. Patch coverage was 98.24% against a 94.38% target. All 52 Wasm byte metrics matched 291adee7373a. The second native benchmark round did not reproduce the earlier macOS timer/channel increases: their medians were 22.5%/24.2% lower than the paired base. Linux timer samples had wide overlapping ranges, so isolated median changes are not treated as proof of either regression or equivalence. W3's full compatibility corpus is validated separately in #2603.

The earlier W3 runtime failure exposed an intermittent typed finalizers did not complete. Isolated first-mark tracing identified an initialized static-data word, 16779018, conservatively treated as an interior pointer into the pending object at 16778976. The same bytes are present in the original Wasm data segment: this is conservative retention, not a damaged queue or a return-signature-specific ABI failure.

The fixture now registers three independent objects per signature, requires all 12 signatures to execute, and keeps per-registration invalid/duplicate event checks using a separate channel for each copy. Noncapturing callbacks obtain their channel from the target object so late callbacks cannot encounter a cleared or reassigned global channel. A deterministic collector-registry test separately verifies that a marked object keeps its finalizer pending without blocking other candidates, and becomes eligible when its retaining root disappears. No GC safety rules, skips, or collection deadlines are changed. A separate optional size experiment remains excluded.

A subsequent W3 benchmark round exposed a reproducible sub-nanosecond MinGW getg difference. An isolated same-source ABBA comparison measured roughly +0.3 ns for both getg and a trivial direct call. After normalizing relocation addresses, disassembly is identical for both benchmark loops (17/19 instructions), the direct-call body (2), and the complete getg function (61), including the TLS offsets. Code placement differs. This records a layout-sensitive microbenchmark result, not performance equivalence; it does not justify adding runtime checks or globally increasing alignment and binary size. The probe and its workflow are not included in this PR.

An additional ARM64 MinGW same-source comparison at the current head investigates the repeatable full-suite direct-call/atomic-write differences. Eight fresh-process samples per revision, interleaved in ABBA order, give main/head medians of 0.59010/0.59005 ns for a direct call, 0.66375/0.66395 ns for an atomic write, and 1.8055/1.774 ns for getg. The earlier approximately 0.073/0.22 ns differences do not reproduce in this isolated fixture. Both 17-instruction loops, the two-instruction call body, and the complete four-instruction atomic-write body are identical; the getg hot path retains the same operations and TLS offset. Its cold-path addresses move. This supports tracking the full-suite result as layout-sensitive, not asserting universal native performance equivalence. The validation-only source and workflow are excluded from this PR.

The decisive full-package fixture control separates the two variables using 24 interleaved fresh processes, eight per variant:

Native ARM64 variant Direct call Atomic write getg Channel handoff
Main compiler/runtime, main test sources 0.59005 ns 0.66400 ns 1.7990 ns 1,852.5 ns
Current compiler/runtime, same main test sources 0.58995 ns 0.66370 ns 1.7995 ns 1,813.5 ns
Current compiler/runtime, current test sources 0.66395 ns 0.88555 ns 1.7965 ns 1,936.5 ns

The repeated direct-call/write differences follow the test-source change, not the compiler/runtime change under fixed sources. The loops retain 17 instructions and their helpers retain 2/4 instructions, with only symbol/data relocations differing; the current fixture shifts the benchmark functions by four bytes and also relocates their callees. This does not identify a particular CPU cache mechanism or prove equivalence for every workload. Other native timing variations remain explicitly monitored: macOS interface-call/defer/timer medians changed direction across successive rounds. No benchmark or runtime code is changed merely to tune these positions, and no diagnostic code/workflow is included in the PR.

A separate 32-bit MinGW control completes the same three-variant experiment (24 fresh processes). Its main / current-with-main-fixture / current-with-current-fixture medians are 0.93645 / 1.11100 / 0.94240 ns for getg, and 0.88815 / 1.11250 / 0.88860 ns for a direct call. Atomic writes stay at about 7.98 ns and channel handoff at about 593–604 ns. Unlike ARM64, the fixed-fixture head is slower in the two tiny calls, while the current fixture returns close to the baseline. Complete disassembly, including internal function-info/PC-site anchors, confirms 17 instructions in each loop, 3 in the direct callee, and 109 in getg for all variants; changes are symbol/data relocations, not additional operations or a changed TLS slot. This records a real layout-sensitive sub-nanosecond effect, not a zero-regression guarantee. It does not justify globally changing function alignment, growing native programs, or changing runtime algorithms to fit one benchmark layout. The formal 386 results and this different controlled-run result are retained separately rather than conflated.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review summary

Comprehensive review across code quality, performance, security/memory-safety, and documentation accuracy. This is a large, carefully engineered change (WASM W2 runtime lowering) with strong test coverage.

No blocking defects found. The highest-risk areas are implemented soundly:

  • wasm32 physical-address overflow (ssa/expr.go, checkUnsafeBuiltinBounds): the new remaining = maxAddr - addr non-wrapping comparison correctly traps at the real 32-bit boundary even when the semantic uintptr is i64.
  • Wide-vs-physical pointer split (ssa/storage.go): 64-bit Go indices truncate to the 32-bit physical width before GEP; integer ABI conversions get an explicit runtime range assertion (fitLLVMValue).
  • GC roots for lowering-created allocations (internal/abi/large.go, internal/abi/gcroot.go, cl/uintptr_escapes.go, cl/gcroot.go): aggregate-copy snapshots and //go:uintptrescapes pointers are re-rooted, preventing use-after-free.
  • JS/host boundary (runtime/internal/wasmjs/_wrap/host.c, embind/_wrap/emval.cpp): copies are length-clamped, memory views re-acquired after growth, and Emscripten control-flow re-thrown correctly.

Verified performance claims (both hold):

  • finalizer.go preserveFinalizableObjects replaces an O(heapBlocks × finalizers) whole-heap scan with a linear pass over the finalizer list; TestCandidateTraversalScalesLinearly locks it in.
  • wasm_stack_addresses.go localizeWasmStackAddresses rematerializes constant alloca-GEP addresses only for functions that call setjmp, avoiding spilling live addresses across throwing calls.
  • Bonus: ssa/gcroot.go appendGCRootPointers now guards recursion with GCRootCount, so pointer-free aggregates emit zero per-element extractvalue instructions (a 128 KiB byte array drops from 131072 to 0).

The findings below are minor robustness/maintainability suggestions, not correctness bugs.

Notes without a single inline location:

  • runtime/internal/runtime/tinygogc/finalizer.go (markFinalizerObjectBlocked, ~line 212): the remaining per-heap-edge linear scan over the finalizer list is the one non-linear part left in the finalizer path. Not a regression (pre-existing behavior), and fine for small finalizer sets — flagging only because this PR advertises linearizing finalizer traversal. An object-keyed index built once per cycle would remove the quadratic factor for finalizer-heavy workloads.
  • finalizer.go (preserveFinalizableObjects, ~lines 128-190): removing earlierFinalizerForObject de-dup and queueing all eligible candidates in one pass is memory-safe (the subsequent re-mark keeps objects alive another cycle), but the ordering/dedup semantics for multiple finalizers on one object changed — worth confirming against the intended finalizer-ordering contract.
  • ssa/target.go / internal/build/wasm_reflect.go: two same-named usesWasmReflectBridges methods on different receivers (*Target vs *wasmProgramUse) with different meanings — legal but easy to conflate; consider distinct names.
  • doc/wasm-proposal.md (Reflection section): the detected bridge entry points also include Seq/Seq2 (see isWasmReflectBridgeName), which the prose omits. Defensible for a scope summary.

Comment thread internal/abi/large.go
Comment thread runtime/internal/runtime/host_events_js.go
@codecov

codecov Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 99.13545% with 6 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
internal/build/build.go 89.09% 6 Missing ⚠️

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

f28418ac274b | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 20040 B +144 B / +0.7% (worse) 387 B 0 B / +0.0% 477.852 ms -4.687 ms / -1.0% (better) 1.351 ms +103.5 us / +8.3% (worse)
Linux cprintf-lto 19792 B +144 B / +0.7% (worse) 368 B 0 B / +0.0% 482.879 ms +2.919 ms / +0.6% (worse) 1.245 ms -47.48 us / -3.7% (better)
Linux fmtprintf 1640168 B +880 B / +0.1% (worse) 487155 B +134 B / +0.02751% (worse) 3.818 s -133.2 ms / -3.4% (better) 3.125 ms +115.3 us / +3.8% (worse)
Linux fmtprintf-lto 1478456 B +784 B / +0.1% (worse) 422763 B +62 B / +0.01467% (worse) 11.379 s +366.8 ms / +3.3% (worse) 2.825 ms -27.73 us / -1.0% (better)
Linux println 62160 B +128 B / +0.2% (worse) 14647 B -21 B / -0.1% (better) 470.763 ms +6.817 ms / +1.5% (worse) 1.522 ms -22.41 us / -1.5% (better)
Linux println-lto 54248 B +144 B / +0.3% (worse) 12115 B 0 B / +0.0% 675.830 ms -16.97 ms / -2.4% (better) 1.616 ms +67.47 us / +4.4% (worse)
macOS cprintf 84480 B 0 B / +0.0% 17309 B +144 B / +0.8% (worse) 642.242 ms +44.73 ms / +7.5% (worse) 4.100 ms +1.331 ms / +48.1% (worse)
macOS cprintf-lto 84288 B 0 B / +0.0% 13073 B +144 B / +1.1% (worse) 1.106 s +322.7 ms / +41.2% (worse) 5.683 ms +1.829 ms / +47.4% (worse)
macOS fmtprintf 1483824 B 0 B / +0.0% 861924 B +884 B / +0.1% (worse) 3.100 s -888.3 ms / -22.3% (better) 3.928 ms -3.204 ms / -44.9% (better)
macOS fmtprintf-lto 1175808 B 0 B / +0.0% 832884 B +872 B / +0.1% (worse) 7.387 s -1.933 s / -20.7% (better) 4.502 ms -928.4 us / -17.1% (better)
macOS println 114672 B 0 B / +0.0% 34882 B +150 B / +0.4% (worse) 831.919 ms +234.1 ms / +39.2% (worse) 3.418 ms -2.999 ms / -46.7% (better)
macOS println-lto 118720 B 0 B / +0.0% 32320 B +144 B / +0.4% (worse) 803.839 ms -144.9 ms / -15.3% (better) 4.414 ms -859.4 us / -16.3% (better)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 921.620 ms +51.44 ms / +5.9% (worse) 3.041 ms -276.7 us / -8.3% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 955.119 ms +40.76 ms / +4.5% (worse) 2.828 ms -609.1 us / -17.7% (better)
Windows MinGW fmtprintf 1913344 B +1536 B / +0.1% (worse) 589158 B +368 B / +0.1% (worse) 3.383 s +72.67 ms / +2.2% (worse) 6.969 ms -54.5 us / -0.8% (better)
Windows MinGW fmtprintf-lto 1933824 B +1024 B / +0.1% (worse) 535110 B +320 B / +0.1% (worse) 7.788 s +119.5 ms / +1.6% (worse) 7.416 ms +405 us / +5.8% (worse)
Windows MinGW println 71168 B 0 B / +0.0% 23702 B -16 B / -0.1% (better) 885.994 ms -12.08 ms / -1.3% (better) 5.846 ms -22.1 us / -0.4% (better)
Windows MinGW println-lto 65024 B 0 B / +0.0% 20406 B 0 B / +0.0% 1.071 s +21.91 ms / +2.1% (worse) 5.664 ms -350.9 us / -5.8% (better)
Windows MinGW 386 cprintf 43008 B +512 B / +1.2% (worse) 5326 B 0 B / +0.0% 887.752 ms -70.54 ms / -7.4% (better) 3.813 ms -170.7 us / -4.3% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 893.194 ms -45.61 ms / -4.9% (better) 3.833 ms -98.4 us / -2.5% (better)
Windows MinGW 386 fmtprintf 1876992 B +2048 B / +0.1% (worse) 463342 B +432 B / +0.1% (worse) 3.394 s +130.9 ms / +4.0% (worse) 7.923 ms +73.6 us / +0.9% (worse)
Windows MinGW 386 fmtprintf-lto 2152448 B +512 B / +0.02379% (worse) 440730 B +276 B / +0.1% (worse) 7.724 s +69.9 ms / +0.9% (worse) 7.883 ms +102.9 us / +1.3% (worse)
Windows MinGW 386 println 91136 B +512 B / +0.6% (worse) 19814 B -16 B / -0.1% (better) 890.178 ms -6.681 ms / -0.7% (better) 6.524 ms -322.5 us / -4.7% (better)
Windows MinGW 386 println-lto 69120 B 0 B / +0.0% 17742 B 0 B / +0.0% 1.031 s -32.07 ms / -3.0% (better) 6.555 ms -192.6 us / -2.9% (better)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.460 s +17.55 ms / +1.2% (worse) 6.447 ms +34.7 us / +0.5% (worse)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.499 s +29.56 ms / +2.0% (worse) 6.538 ms -158 us / -2.4% (better)
Windows MinGW ARM64 fmtprintf 1801216 B +1536 B / +0.1% (worse) 501720 B +436 B / +0.1% (worse) 4.303 s +63.88 ms / +1.5% (worse) 13.499 ms +741.2 us / +5.8% (worse)
Windows MinGW ARM64 fmtprintf-lto 1858048 B +512 B / +0.02756% (worse) 465864 B +336 B / +0.1% (worse) 9.635 s +132.6 ms / +1.4% (worse) 12.887 ms -550.5 us / -4.1% (better)
Windows MinGW ARM64 println 68096 B 0 B / +0.0% 22376 B 0 B / +0.0% 1.466 s +26.13 ms / +1.8% (worse) 10.875 ms -348.2 us / -3.1% (better)
Windows MinGW ARM64 println-lto 63488 B 0 B / +0.0% 19444 B 0 B / +0.0% 1.644 s +6.203 ms / +0.4% (worse) 11.619 ms +89.3 us / +0.8% (worse)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65782 B 0 B / +0.0% 862.373 ms -40.83 ms / -4.5% (better) 3.599 ms +152.7 us / +4.4% (worse)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65718 B 0 B / +0.0% 891.929 ms -161.2 ms / -15.3% (better) 3.448 ms -919 us / -21.0% (better)
Windows MSVC fmtprintf 1627648 B +1536 B / +0.1% (worse) 684678 B +368 B / +0.1% (worse) 3.636 s -3.379 ms / -0.1% (better) 9.783 ms +614.8 us / +6.7% (worse)
Windows MSVC fmtprintf-lto 1616384 B +1024 B / +0.1% (worse) 635750 B +336 B / +0.1% (worse) 8.751 s +326.4 ms / +3.9% (worse) 10.920 ms +1.552 ms / +16.6% (worse)
Windows MSVC println 192512 B 0 B / +0.0% 119126 B -16 B / -0.01343% (better) 881.678 ms -4.382 ms / -0.5% (better) 7.123 ms -410.5 us / -5.4% (better)
Windows MSVC println-lto 189952 B 0 B / +0.0% 116374 B 0 B / +0.0% 1.069 s -12.61 ms / -1.2% (better) 8.102 ms +276.4 us / +3.5% (worse)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 951.929 ms -41.75 ms / -4.2% (better) 5.196 ms -837.5 us / -13.9% (better)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.117 s +130.4 ms / +13.2% (worse) 7.835 ms +1.994 ms / +34.1% (worse)
Windows MSVC 386 fmtprintf 1190400 B +1536 B / +0.1% (worse) 446693 B +432 B / +0.1% (worse) 3.752 s -298.7 ms / -7.4% (better) 11.886 ms +184.3 us / +1.6% (worse)
Windows MSVC 386 fmtprintf-lto 1224704 B +1024 B / +0.1% (worse) 417385 B +320 B / +0.1% (worse) 8.603 s -154.3 ms / -1.8% (better) 11.605 ms -204.4 us / -1.7% (better)
Windows MSVC 386 println 34304 B 0 B / +0.0% 18641 B -32 B / -0.2% (better) 927.305 ms -52.55 ms / -5.4% (better) 9.412 ms -1.531 ms / -14.0% (better)
Windows MSVC 386 println-lto 32256 B 0 B / +0.0% 16855 B 0 B / +0.0% 1.087 s -41.54 ms / -3.7% (better) 8.813 ms -990.8 us / -10.1% (better)
Windows MSVC ARM64 cprintf 11264 B 0 B / +0.0% 3976 B 0 B / +0.0% 2.091 s +15.06 ms / +0.7% (worse) 7.312 ms +308.4 us / +4.4% (worse)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 3868 B 0 B / +0.0% 2.094 s -10.9 ms / -0.5% (better) 7.452 ms +181.6 us / +2.5% (worse)
Windows MSVC ARM64 fmtprintf 1374720 B +1024 B / +0.1% (worse) 502660 B +432 B / +0.1% (worse) 6.679 s +81.26 ms / +1.2% (worse) 14.945 ms +1.024 ms / +7.4% (worse)
Windows MSVC ARM64 fmtprintf-lto 1391104 B +1024 B / +0.1% (worse) 468772 B +336 B / +0.1% (worse) 15.706 s +627 ms / +4.2% (worse) 14.713 ms +1.23 ms / +9.1% (worse)
Windows MSVC ARM64 println 41472 B 0 B / +0.0% 21624 B 0 B / +0.0% 2.028 s +4.955 ms / +0.2% (worse) 12.723 ms +259.1 us / +2.1% (worse)
Windows MSVC ARM64 println-lto 39424 B 0 B / +0.0% 19372 B 0 B / +0.0% 2.394 s +38.41 ms / +1.6% (worse) 16.491 ms +1.845 ms / +12.6% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.500 ns/op -0.01 ns/op / -0.1% (better)
Linux BenchmarkMergeCompilerFlags 206.800 ns/op +8.1 ns/op / +4.1% (worse)
Linux BenchmarkMergeLinkerFlags 138.400 ns/op +9.5 ns/op / +7.4% (worse)
Linux BenchmarkChannelBuffered 55.520 ns/op +0.4 ns/op / +0.7% (worse)
Linux BenchmarkChannelHandoff 13145 ns/op -303 ns/op / -2.3% (better)
Linux BenchmarkDefer 54.500 ns/op +6.24 ns/op / +12.9% (worse)
Linux BenchmarkDirectCall 1.172 ns/op +0.007 ns/op / +0.6% (worse)
Linux BenchmarkGlobalRead 1.170 ns/op -0.395 ns/op / -25.2% (better)
Linux BenchmarkGlobalWrite 7.792 ns/op +0.013 ns/op / +0.2% (worse)
Linux BenchmarkGoroutine 25747 ns/op +4721 ns/op / +22.5% (worse)
Linux BenchmarkInterfaceCall 6.005 ns/op +0.174 ns/op / +3.0% (worse)
Linux BenchmarkRuntimeGetG 2.439 ns/op -0.013 ns/op / -0.5% (better)
macOS BenchmarkLookupPCRandom 14.480 ns/op -8.43 ns/op / -36.8% (better)
macOS BenchmarkMergeCompilerFlags 124.700 ns/op -58.8 ns/op / -32.0% (better)
macOS BenchmarkMergeLinkerFlags 83.680 ns/op -50.82 ns/op / -37.8% (better)
macOS BenchmarkChannelBuffered 40.730 ns/op +11.52 ns/op / +39.4% (worse)
macOS BenchmarkChannelHandoff 12125 ns/op +1851 ns/op / +18.0% (worse)
macOS BenchmarkDefer 60.680 ns/op +16.89 ns/op / +38.6% (worse)
macOS BenchmarkDirectCall 1.675 ns/op +0.436 ns/op / +35.2% (worse)
macOS BenchmarkGlobalRead 1.372 ns/op +0.1 ns/op / +7.9% (worse)
macOS BenchmarkGlobalWrite 1.589 ns/op +0.241 ns/op / +17.9% (worse)
macOS BenchmarkGoroutine 83200 ns/op +30665 ns/op / +58.4% (worse)
macOS BenchmarkInterfaceCall 7.524 ns/op +2.606 ns/op / +53.0% (worse)
macOS BenchmarkRuntimeGetG 3.761 ns/op +1.228 ns/op / +48.5% (worse)
Windows MinGW BenchmarkLookupPCRandom 8.081 ns/op -0.206 ns/op / -2.5% (better)
Windows MinGW BenchmarkMergeCompilerFlags 393.500 ns/op +19.8 ns/op / +5.3% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 349.600 ns/op -0.9 ns/op / -0.3% (better)
Windows MinGW BenchmarkChannelBuffered 28.380 ns/op -2.18 ns/op / -7.1% (better)
Windows MinGW BenchmarkChannelHandoff 6186 ns/op +2430 ns/op / +64.7% (worse)
Windows MinGW BenchmarkDefer 41.760 ns/op -5.41 ns/op / -11.5% (better)
Windows MinGW BenchmarkDirectCall 0.273 ns/op -0.0104 ns/op / -3.7% (better)
Windows MinGW BenchmarkGlobalRead 0.310 ns/op -0.071 ns/op / -18.6% (better)
Windows MinGW BenchmarkGlobalWrite 6.932 ns/op -0.155 ns/op / -2.2% (better)
Windows MinGW BenchmarkGoroutine 69146 ns/op +988 ns/op / +1.4% (worse)
Windows MinGW BenchmarkInterfaceCall 4.071 ns/op -0.343 ns/op / -7.8% (better)
Windows MinGW BenchmarkRuntimeGetG 0.891 ns/op +0.0047 ns/op / +0.5% (worse)
Windows MinGW 386 BenchmarkLookupPCRandom 21.530 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 573.600 ns/op -22.1 ns/op / -3.7% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 539.600 ns/op +12.5 ns/op / +2.4% (worse)
Windows MinGW 386 BenchmarkChannelBuffered 35.740 ns/op +1.88 ns/op / +5.6% (worse)
Windows MinGW 386 BenchmarkChannelHandoff 2412 ns/op -27 ns/op / -1.1% (better)
Windows MinGW 386 BenchmarkDefer 34.640 ns/op -0.46 ns/op / -1.3% (better)
Windows MinGW 386 BenchmarkDirectCall 1.358 ns/op -0.271 ns/op / -16.6% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.630 ns/op +0.274 ns/op / +20.2% (worse)
Windows MinGW 386 BenchmarkGlobalWrite 6.985 ns/op +0.008 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGoroutine 57079 ns/op -1698 ns/op / -2.9% (better)
Windows MinGW 386 BenchmarkInterfaceCall 7.334 ns/op +0.006 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkRuntimeGetG 1.630 ns/op -0.271 ns/op / -14.3% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.070 ns/op +0.1 ns/op / +0.8% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 578 ns/op +8.5 ns/op / +1.5% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 535.900 ns/op +0.9 ns/op / +0.2% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 37.700 ns/op +0.49 ns/op / +1.3% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 2127 ns/op -80 ns/op / -3.6% (better)
Windows MinGW ARM64 BenchmarkDefer 54.360 ns/op -0.48 ns/op / -0.9% (better)
Windows MinGW ARM64 BenchmarkDirectCall 0.663 ns/op +0.0002 ns/op / +0.03017% (worse)
Windows MinGW ARM64 BenchmarkGlobalRead 0.590 ns/op -0.0002 ns/op / -0.03392% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.884 ns/op +0.2211 ns/op / +33.3% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 60123 ns/op +3689 ns/op / +6.5% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.299 ns/op +0.133 ns/op / +3.2% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.772 ns/op -0.04 ns/op / -2.2% (better)
Windows MSVC BenchmarkLookupPCRandom 10.940 ns/op +0.948 ns/op / +9.5% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 560.900 ns/op +67.4 ns/op / +13.7% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 486.700 ns/op +48.9 ns/op / +11.2% (worse)
Windows MSVC BenchmarkChannelBuffered 42.100 ns/op -1.71 ns/op / -3.9% (better)
Windows MSVC BenchmarkChannelHandoff 1833 ns/op -71 ns/op / -3.7% (better)
Windows MSVC BenchmarkDefer 64.420 ns/op +4.39 ns/op / +7.3% (worse)
Windows MSVC BenchmarkDirectCall 1.087 ns/op -0.197 ns/op / -15.3% (better)
Windows MSVC BenchmarkGlobalRead 1.313 ns/op -0.064 ns/op / -4.6% (better)
Windows MSVC BenchmarkGlobalWrite 8.148 ns/op -0.179 ns/op / -2.1% (better)
Windows MSVC BenchmarkGoroutine 66266 ns/op +1457 ns/op / +2.2% (worse)
Windows MSVC BenchmarkInterfaceCall 5.671 ns/op -0.195 ns/op / -3.3% (better)
Windows MSVC BenchmarkRuntimeGetG 1.335 ns/op -0.559 ns/op / -29.5% (better)
Windows MSVC 386 BenchmarkLookupPCRandom 26.590 ns/op +0.07 ns/op / +0.3% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 767.100 ns/op +45.8 ns/op / +6.3% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 681.800 ns/op -13.1 ns/op / -1.9% (better)
Windows MSVC 386 BenchmarkChannelBuffered 38.410 ns/op -7.29 ns/op / -16.0% (better)
Windows MSVC 386 BenchmarkChannelHandoff 755.300 ns/op -137.2 ns/op / -15.4% (better)
Windows MSVC 386 BenchmarkDefer 48.290 ns/op +3.03 ns/op / +6.7% (worse)
Windows MSVC 386 BenchmarkDirectCall 1.549 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.857 ns/op -0.008 ns/op / -0.4% (better)
Windows MSVC 386 BenchmarkGlobalWrite 7.777 ns/op -0.004 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGoroutine 87983 ns/op -2735 ns/op / -3.0% (better)
Windows MSVC 386 BenchmarkInterfaceCall 8.356 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkRuntimeGetG 2.168 ns/op +0.233 ns/op / +12.0% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.060 ns/op +0.01 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 580.100 ns/op +0.3 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 542 ns/op -9.2 ns/op / -1.7% (better)
Windows MSVC ARM64 BenchmarkChannelBuffered 39.380 ns/op +0.65 ns/op / +1.7% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 2945 ns/op +506 ns/op / +20.7% (worse)
Windows MSVC ARM64 BenchmarkDefer 63.040 ns/op -2.44 ns/op / -3.7% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op +0.0002 ns/op / +0.03393% (worse)
Windows MSVC ARM64 BenchmarkGlobalRead 0.663 ns/op -0.0002 ns/op / -0.03015% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.760 ns/op -0.036 ns/op / -0.9% (better)
Windows MSVC ARM64 BenchmarkGoroutine 58097 ns/op +368 ns/op / +0.6% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.146 ns/op -0.023 ns/op / -0.6% (better)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.789 ns/op -0.007 ns/op / -0.4% (better)

Timer runtime benchmarks

Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 913.800 ns/op -1.4 ns/op / -0.2% (better)
Linux AfterFuncZeroDelivery/LLGo 33989 ns/op +1271 ns/op / +3.9% (worse)
Linux CreateStop/Go 290.700 ns/op +0.9 ns/op / +0.3% (worse)
Linux CreateStop/LLGo 1716 ns/op -769 ns/op / -30.9% (better)
Linux RearmStopped/Go 114.800 ns/op -1.5 ns/op / -1.3% (better)
Linux RearmStopped/LLGo 1385 ns/op +267 ns/op / +23.9% (worse)
Linux ResetActive/Go 67.540 ns/op -2.21 ns/op / -3.2% (better)
Linux ResetActive/LLGo 619.100 ns/op -324.3 ns/op / -34.4% (better)
Linux ResetHeap1024/Go 67.240 ns/op -0.54 ns/op / -0.8% (better)
Linux ResetHeap1024/LLGo 291 ns/op +111.5 ns/op / +62.1% (worse)
macOS AfterFuncZeroDelivery/Go 562.700 ns/op -79.9 ns/op / -12.4% (better)
macOS AfterFuncZeroDelivery/LLGo 84100 ns/op -18814 ns/op / -18.3% (better)
macOS CreateStop/Go 175 ns/op -53.1 ns/op / -23.3% (better)
macOS CreateStop/LLGo 923.200 ns/op +153.8 ns/op / +20.0% (worse)
macOS RearmStopped/Go 73.270 ns/op -25.61 ns/op / -25.9% (better)
macOS RearmStopped/LLGo 681.200 ns/op +18.8 ns/op / +2.8% (worse)
macOS ResetActive/Go 49.750 ns/op -16.97 ns/op / -25.4% (better)
macOS ResetActive/LLGo 275.200 ns/op +68.1 ns/op / +32.9% (worse)
macOS ResetHeap1024/Go 52.320 ns/op -15.23 ns/op / -22.5% (better)
macOS ResetHeap1024/LLGo 96.550 ns/op -14.25 ns/op / -12.9% (better)
Windows MinGW AfterFuncZeroDelivery/Go 467.900 ns/op -21.4 ns/op / -4.4% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 126482 ns/op +272 ns/op / +0.2% (worse)
Windows MinGW CreateStop/Go 133.700 ns/op +0.2 ns/op / +0.1% (worse)
Windows MinGW CreateStop/LLGo 714.500 ns/op +95.4 ns/op / +15.4% (worse)
Windows MinGW RearmStopped/Go 49.560 ns/op -0.17 ns/op / -0.3% (better)
Windows MinGW RearmStopped/LLGo 284.300 ns/op +40.8 ns/op / +16.8% (worse)
Windows MinGW ResetActive/Go 22.400 ns/op +0.31 ns/op / +1.4% (worse)
Windows MinGW ResetActive/LLGo 245.100 ns/op +77.3 ns/op / +46.1% (worse)
Windows MinGW ResetHeap1024/Go 22.100 ns/op +0.07 ns/op / +0.3% (worse)
Windows MinGW ResetHeap1024/LLGo 93.810 ns/op -2.09 ns/op / -2.2% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 774 ns/op +16.8 ns/op / +2.2% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 123507 ns/op +1062 ns/op / +0.9% (worse)
Windows MinGW 386 CreateStop/Go 166.300 ns/op -2.2 ns/op / -1.3% (better)
Windows MinGW 386 CreateStop/LLGo 512.500 ns/op +4.4 ns/op / +0.9% (worse)
Windows MinGW 386 RearmStopped/Go 56.560 ns/op -2.81 ns/op / -4.7% (better)
Windows MinGW 386 RearmStopped/LLGo 295.400 ns/op -5.9 ns/op / -2.0% (better)
Windows MinGW 386 ResetActive/Go 32.560 ns/op -0.29 ns/op / -0.9% (better)
Windows MinGW 386 ResetActive/LLGo 720.100 ns/op +405 ns/op / +128.5% (worse)
Windows MinGW 386 ResetHeap1024/Go 32.790 ns/op -0.49 ns/op / -1.5% (better)
Windows MinGW 386 ResetHeap1024/LLGo 150.200 ns/op +0.4 ns/op / +0.3% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 656.800 ns/op -5.5 ns/op / -0.8% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 136330 ns/op -5038 ns/op / -3.6% (better)
Windows MinGW ARM64 CreateStop/Go 200 ns/op +1.1 ns/op / +0.6% (worse)
Windows MinGW ARM64 CreateStop/LLGo 364.300 ns/op +10 ns/op / +2.8% (worse)
Windows MinGW ARM64 RearmStopped/Go 70.550 ns/op -0.01 ns/op / -0.01417% (better)
Windows MinGW ARM64 RearmStopped/LLGo 254.300 ns/op +1.9 ns/op / +0.8% (worse)
Windows MinGW ARM64 ResetActive/Go 30.920 ns/op -0.15 ns/op / -0.5% (better)
Windows MinGW ARM64 ResetActive/LLGo 127.500 ns/op +4.9 ns/op / +4.0% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.110 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 ResetHeap1024/LLGo 127.300 ns/op -0.4 ns/op / -0.3% (better)
Windows MSVC AfterFuncZeroDelivery/Go 669.300 ns/op +66.3 ns/op / +11.0% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 134678 ns/op +16509 ns/op / +14.0% (worse)
Windows MSVC CreateStop/Go 183.800 ns/op +16.3 ns/op / +9.7% (worse)
Windows MSVC CreateStop/LLGo 649.300 ns/op +171.1 ns/op / +35.8% (worse)
Windows MSVC RearmStopped/Go 66.110 ns/op +4.92 ns/op / +8.0% (worse)
Windows MSVC RearmStopped/LLGo 301 ns/op -10.5 ns/op / -3.4% (better)
Windows MSVC ResetActive/Go 30.120 ns/op +3.02 ns/op / +11.1% (worse)
Windows MSVC ResetActive/LLGo 187 ns/op +25 ns/op / +15.4% (worse)
Windows MSVC ResetHeap1024/Go 29.750 ns/op +2.21 ns/op / +8.0% (worse)
Windows MSVC ResetHeap1024/LLGo 127.700 ns/op +13 ns/op / +11.3% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 956.400 ns/op -2.5 ns/op / -0.3% (better)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 187702 ns/op -6653 ns/op / -3.4% (better)
Windows MSVC 386 CreateStop/Go 194.200 ns/op +0.6 ns/op / +0.3% (worse)
Windows MSVC 386 CreateStop/LLGo 442.500 ns/op -25.7 ns/op / -5.5% (better)
Windows MSVC 386 RearmStopped/Go 63.150 ns/op -0.26 ns/op / -0.4% (better)
Windows MSVC 386 RearmStopped/LLGo 325.900 ns/op -1.5 ns/op / -0.5% (better)
Windows MSVC 386 ResetActive/Go 39 ns/op -0.06 ns/op / -0.2% (better)
Windows MSVC 386 ResetActive/LLGo 287.400 ns/op -5 ns/op / -1.7% (better)
Windows MSVC 386 ResetHeap1024/Go 39.440 ns/op +0.05 ns/op / +0.1% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 171 ns/op +0.6 ns/op / +0.4% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 663.400 ns/op -5.7 ns/op / -0.9% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 130895 ns/op -12122 ns/op / -8.5% (better)
Windows MSVC ARM64 CreateStop/Go 200.500 ns/op -5.5 ns/op / -2.7% (better)
Windows MSVC ARM64 CreateStop/LLGo 382.700 ns/op -1 ns/op / -0.3% (better)
Windows MSVC ARM64 RearmStopped/Go 70.540 ns/op +0.01 ns/op / +0.01418% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 268.600 ns/op -1.5 ns/op / -0.6% (better)
Windows MSVC ARM64 ResetActive/Go 31.110 ns/op -0.03 ns/op / -0.1% (better)
Windows MSVC ARM64 ResetActive/LLGo 130.600 ns/op -2.7 ns/op / -2.0% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.200 ns/op +0.11 ns/op / +0.4% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 139.300 ns/op +1.8 ns/op / +1.3% (worse)

Compared with cc16553ba9d0 measured in the same runner job.

@cpunion
cpunion force-pushed the codex/wasm-w2-runtime-lowering-20260913 branch from c8b6e7d to 291adee Compare September 15, 2026 08:24
@github-actions

github-actions Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

f28418ac274b | workflow run | long-term charts

WebAssembly output sizes

Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 135746 B +4843 B / +3.7% (worse) 70786 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 133995 B +4783 B / +3.7% (worse) 69150 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 124243 B +4731 B / +4.0% (worse) 73998 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 138261 B +4783 B / +3.6% (worse) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 138265 B +4912 B / +3.7% (worse) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3090543 B -133701 B / -4.1% (better) 114540 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3082582 B -134074 B / -4.2% (better) 98239 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2829088 B -95961 B / -3.3% (better) 121521 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2828843 B -243578 B / -7.9% (better) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2697398 B -234740 B / -8.0% (better) 0 B 0 B / 0.0%
j32-emscripten/LLGo 135020 B +4985 B / +3.8% (worse) 70786 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 133441 B +4868 B / +3.8% (worse) 69150 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 123573 B +4827 B / +4.1% (worse) 73998 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1481254 B -73040 B / -4.7% (better) 88949 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1482975 B -72459 B / -4.7% (better) 87313 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1368281 B -51288 B / -3.6% (better) 94158 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1534123 B -112090 B / -6.8% (better) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1460764 B -113815 B / -7.2% (better) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 137550 B +4930 B / +3.7% (worse) 0 B 0 B / 0.0%
w32-wasi/LLGo 137621 B +5037 B / +3.8% (worse) 0 B 0 B / 0.0%

LLGo WebAssembly build measurements

Example and profile Build vs base
j32-emscripten 5.896 s +11.33 ms / +0.2% (worse)
j32-goos-js 5.834 s -131.6 ms / -2.2% (better)
j64-emscripten-memory64 4.973 s -72.95 ms / -1.4% (better)
reflectcall/w32-wasi 26.380 s -3.671 s / -12.2% (better)
w32-goos-wasip1 4.449 s +68.44 ms / +1.6% (worse)
w32-wasi 4.400 s +32.42 ms / +0.7% (worse)

Compared with cc16553ba9d0 measured in the same runner job.

…p object

Keep the existing allocation-free registry. Restart after queue removal and clear remaining candidates for the same object so cleanup still waits until a later collection.

Validate with the existing lifecycle suite on wasm32, wasm64 and WASI (five repetitions each), plus repeated JS callback and filesystem tests. No new skip or relaxed lifecycle assertion.
Exercise the production registry operations with isolated collector state. Count heap metadata reads instead of timing collection, and check interleaved finalizer/cleanup records and queue removal. Run the check in the existing Wasm runtime job alongside integration fixtures.

The old heap scan fails the operation-count checks. Mutations removing same-object candidate clearing or following detached next links fail queue-order checks. The fixed implementation passes ten repetitions.
@cpunion
cpunion force-pushed the codex/wasm-w2-runtime-lowering-20260913 branch from 3b3a5e5 to f28418a Compare September 17, 2026 11:07
@xushiwei
xushiwei merged commit 1e828ce into xgo-dev:main Sep 17, 2026
71 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants