Skip to content

[WASM W2-B] Complete single-worker runtime and aggregate lowering - #2580

Closed
cpunion wants to merge 77 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-w2-runtime-lowering-20260913
Closed

cpunion wants to merge 77 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-w2-runtime-lowering-20260913

Conversation

@cpunion

@cpunion cpunion commented Sep 13, 2026 •

Copy link
Copy Markdown
Collaborator

Tracks #2152. Stacked on #2579; review this PR's 28-commit W2-B range after W2-A.

This combines the single-worker runtime/standard-library completeness layer with the aggregate/SjLj lowering it requires for realistic WASM packages. It deliberately excludes the broad test/** and GOROOT acceptance inventory, which remains W3.

Scope

  • Bind goroutine-local state to logical WebAssembly goroutines, including initialization-failure caching and per-goroutine cleanup, while preserving single-worker scheduling semantics.
  • Preserve unrecovered and same-value repanic tracebacks, keep synthetic caller PCs unique across goroutine-local stores, symbolize fatal WebAssembly traces from the shadow stack, and bound Memory32 caller traversal by the real stack extent.
  • Add recoverable nil-dereference checks for user code in linear memory without retaining the full panic implementation for every runtime memory access.
  • Detect wasm32 physical-address overflow at the actual memory width and stabilize closure/locality lowering across supported profiles.
  • Scan only registered tinygogc finalizer candidates, preserve finalizer-before-cleanup ordering, expose completed collections through runtime.MemStats.NumGC, and schedule pending cleanup work after an explicit collection.
  • Use acceptance-derived Fiber stack budgets of 128 KiB for Memory32 and 256 KiB for Memory64, with stack storage rooted and released by its owning logical goroutine.
  • Run six complete repository standard-library packages across J32/GoJS (wasm32 with the Go-compatible JavaScript host), J32/Emscripten (wasm32 with the Emscripten JavaScript host), J64/Emscripten Memory64 (wasm64 with the Emscripten JavaScript host), W32/WASI (wasm32 with WASI Preview 1), official Go js/wasm, and official Go wasip1/wasm; account for every other test/std package as unvalidated or source-excluded.
  • Lower medium and large WebAssembly aggregate snapshots to bounded rooted memory copies before LLVM scalarization, preserve Go assignment timing and volatile semantics, and localize constant stack-slot addresses before SjLj expansion.

Root cause and boundary

The runtime half closes observable single-worker correctness gaps found by host, GC, panic, and standard-library acceptance. The lowering half is required by that acceptance: W32 go/types previously produced a roughly 14.2 MB go/types.(*Checker).builtin function after LLVM scalarized aggregate loads and SjLj extended stack-slot liveness, exceeding Wasmtime's function-body limit. The final lowering shares its 4 KiB threshold with frontend safepoint analysis, roots the source before allocation and the snapshot afterwards, and does not raise a host limit.

The post-rebase W3 acceptance run also exposed two caller-state bugs in newly merged main tests. A recovered same-value panic could lose its original WASM shadow-stack prefix after nested longjmps, while goroutine-local synthetic PC sequences collided in the public runtime's process-wide FuncForPC cache. The follow-up freezes the existing shadow-stack prefix without copying it, retains it only for the matching recover activation, and assigns process-unique synthetic PCs on hosted LLGo targets; bare-metal and host-tool builds keep the non-atomic single-store path.

Current head is 748483792502, based on W2-A head 8a7f17a8a227 (28 commits, 92 files, +4,216/-272). Twenty-five original dedicated commits retain exact one-to-one range-diff equivalence; one commit reconciles main’s broader native panic-site recording with W2-B’s explicit WebAssembly nil guard, the caller follow-up addresses the runtime regressions found by stacked acceptance, and the final test-only follow-up keeps subprocess-based fatal checks out of WASM guests that correctly return ENOSYS for pipe.

Validation

  • The runtime staging range passed its six-profile standard-library gate, WASM runtime gate, public llgo test gate, Ubuntu instrumented coverage, and WASM benchmark. Codecov reported 100% coverage of modified coverable lines.
  • The lowering staging range passed its six-profile standard-library gate, WASM runtime gate, public llgo test gate, Ubuntu instrumented coverage, and WASM benchmark. Codecov again reported 100% coverage of modified coverable lines.
  • W32 go/types passes; the pre-Asyncify module falls from about 36 MB to 12 MB and its largest function falls from about 14.2 MB to 400,916 bytes.
  • After the latest stack rebase, the W2-B-specific cl and internal/build tests, upstream native panic regression tests, ssa, internal/abi, internal/locality/layout, dev/wasmstdlib, relevant runtime packages, workflow actionlint, shell syntax, and git diff --check pass.
  • The caller follow-up passes the runtime module's complete host test set, focused cl/ssa/internal/build tests, the complete caller/statement-line order on J32 Emscripten, and normal plus nested same-value-repanic scheduler fixtures on J32 Emscripten (wasm32 with Emscripten/JavaScript), J64 Emscripten Memory64 (wasm64 with Emscripten/JavaScript), and W32 WASI (wasm32 with WASI Preview 1).
  • Review follow-up replaces the superlinear finalizer ready-queue scan with deterministic O(N) traversal coverage, passes the real finalizer/cleanup/weak-reference lifecycle on J32 Emscripten, J64 Emscripten Memory64, and W32 WASI, and covers an actual 8 KiB aggregate C-export wrapper.
  • The final range keeps subprocess-based fatal tests as native wrappers while the WASM scheduler exercises repanic and caller-cache behavior through the host; it adds no t.Skip, TODO/FIXME, diagnostic probes, temporary dependency changes, or undocumented package exclusions.
  • This consolidated PR's own CI, coverage, review, size, and performance results are the final gate; the staging runs establish that both unchanged component ranges were already exercised independently.

Size and build-cost audit

The runtime layer adds about 4 KiB (3.0%-3.5%) to optimized tiny cprintf/println WASM modules for source-aware panic/caller support and logical goroutine ownership; generated Emscripten glue is unchanged. Larger programs improve: fmtprintf shrinks 7.3%-9.1%, reflectcall shrinks 11.4%-13.6%, and the measured W32 reflection cold build falls from 32.03 s to 16.81 s.

Against the previous W2-B head, the caller follow-up changes an unchanged J32 Emscripten wasm-runtime module from 407,190 to 409,817 bytes (+2,627, +0.65%); gzip size changes from 154,947 to 155,919 bytes (+972, +0.63%), and generated JS changes by 4 bytes. It allocates no panic snapshot and adds no atomic dependency to bare-metal builds.

The aggregate/SjLj layer adds zero bytes to tiny programs and makes J32/J64 large modules effectively unchanged or up to 100 bytes smaller. W32 fmtprintf shrinks 3.86%-3.97% and W32 reflectcall shrinks 2.42%-2.62%; every paired cold-build sample improved by 0.50%-5.30%. The consolidated benchmark will refresh these measurements on the final stack.

Review boundary

The intended review range is 8a7f17a8a227..748483792502. It covers single-worker scheduler/runtime semantics, tinygogc finalizer accounting, fixed-stack sizing established by acceptance, the focused standard-library slice, bounded aggregate/SjLj lowering, and the final panic/caller identity corrections. It does not claim the complete applicable test/** or GOROOT corpus; W3 owns that final acceptance and classification.

@cpunion

cpunion commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai Please review the intended W2-B range 7d63156a1fa5..a82726740fba. Please focus on logical-goroutine state ownership, single-worker scheduler/GC/finalizer behavior, panic/caller fidelity, Memory32 address and stack bounds, aggregate-copy root timing, SjLj locality, the focused standard-library accounting, and whether the reported size/build improvements match the implementation boundary.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FennoAI Review

Comprehensive review of the WASM (W2) runtime lowering PR across code quality, performance, security, and documentation. Overall the change is high quality: the wide-Go-word storage abstraction is coherent, load/store/atomic/struct-init paths are consistent, the JS/host byte-copy boundary is bounds-checked and type-validated, integer narrowing on the wasm32 ABI is guarded by runtime/compile-time checks, and the documentation (doc/wasm-proposal.md, benchmark/stdlib READMEs) matches the implementation. No security or documentation issues found.

The findings below are one performance concern in the per-GC-cycle finalizer scan (inline) plus a few maintainability/consistency notes.

Additional notes (no reliable inline location)

  • ssa/reflect_bridge.go (wasmReflectSlot/wasmReflectFrame) — GEP index constants use prog.Int().ll (i64 on the wide-storage profiles) as the index type, while dynamic pointer indices elsewhere are deliberately narrowed via physicalPointerIndex to avoid leaving i64 addresses on memory32 loads at -O0. These indices are compile-time constants so the backend folds them, but a comment (or using a physical-width constant) would keep the invariant uniform.
  • internal/abi/large.go (LowerWasmAggregateCopies) — The fixed-point for {} loop relies on an implicit termination argument (each pass strictly reduces remaining large aggregate loads as projections descend nested types). A brief comment asserting monotonic progress would make the termination explicit.
  • cl/gcroot.go — The //go:uintptrescapes special case (forcing GCRootCount = 1 in prepareGCRoots and storing the raw uintptr via SetGCRoot in publishGCRoot, bypassing GCRootPointers) couples two functions implicitly. A cross-referencing comment or shared helper would reduce drift risk.

Comment thread runtime/internal/runtime/tinygogc/finalizer.go Outdated
@fennoai

fennoai Bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Reviewing PR #2580 (W2-B range)

Reviewing the intended W2-B range 7d63156a1fa5..a82726740fba (84 files, +3867/-166).

  • Prepare workspace and lock diff range
  • Run code-quality, performance, security, documentation reviewers
  • Deduplicate and validate findings
  • Submit PR review

View job run

@codecov

codecov Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

W2-B range review (7d63156a1fa5..a82726740fba, 84 files, +3867/-166)

I reviewed the intended W2-B range across all four dimensions (code quality, performance, security, documentation), focusing on the areas you called out. Overall this is careful, well-tested work — I found no blocking correctness or security defects, and the reported size/build claims match the implementation. Two small items below (one worth confirming, one comment nit), plus a few non-blocking observations.

Focus-area verification (all confirmed sound):

  • Logical-goroutine ownership — releaseGoroutineLocalBlocks clears only block.cacheSlot = nil and never writes through *cacheSlot, so it can't corrupt the shared process-global GLS key (a real hazard that releaseLocalBlocks avoids). GLS context is released on the scheduler after the G is _Gdead in both proc_wasm.go and proc_wasip1.go — the correct point.
  • Finalizer/GC — the new registry-driven scan drops the Phase-3 walk from O(heap) to O(registered callbacks); preserveFinalizableObjects visits each object once per cycle and defers cleanups-after-finalizers to a later collection. NumGC exposure is consistent across gc.go/gc_tinygo.go/runtime_gc_nonmoving.go.
  • Panic/caller fidelity & Memory32 bounds — traceback and Caller/Callers iterate the logical shadow stack (store.stack), bounded by the live instrumented frame count; PopCallerLocationFrame validates mark against len(store.stack) before slicing. No OOB.
  • wasm32 physical-address overflow — checkUnsafeBuiltinBounds (ssa/expr.go) computes remaining = maxAddr - addr in physical-width space with maxAddr clamped to the pointer width; the unsigned subtraction can't underflow and the byte-size multiply can't wrap. Sound for J32-over-wasm32.
  • Aggregate-copy root timing / SjLj locality — the 4 KiB MinWasmAggregateCopySize threshold is applied consistently in abi/large.go, cl/gcroot.go (root safepoint accounting), and validated by the 4095→false / 4096→true boundary tests. The LowerWasmAggregateCopies fixpoint terminates (each pass exposes only strictly-nested loads). localizeWasmStackAddresses only rewrites constant-index GEPs rooted at allocas and preserves volatility/ordering/alignment.
  • Fixed Fiber budgets — 128 KiB Memory32 / 256 KiB Memory64 match stack_default*.go, with the Asyncify buffer tied to the same budget.
  • Std-lib accounting — the six profiles agree across dev/wasmstdlib/README.md, main.go, and .github/workflows/wasm-stdlib.yml; all six acceptance witnesses exist in their test/std packages.

Non-blocking observations (no change required):

  • cl/gcroot.go (isGCSafepoint): every aggregate load ≥ 4 KiB on Wasm is treated as a GC safepoint, but the single adjacent load/store pair lowers to a memmove with no AllocU (no new safepoint), so those loads get root frames the lowering never needs. Safe (extra roots, not missing), minor shadow-frame overhead.
  • internal/abi/large.go (allocResult): reuses an existing runtimeAllocU decl but builds the call against a locally-derived intType width. Safe because goWordSize is fixed per module; an assertion or fn.GlobalValueType() reuse would make the coupling explicit.
  • dev/wasmstdlib/main.go: not-run packages carry two different reason strings ("outside this acceptance slice; not validated" vs. the preflight "source selection not completed; not validated") for the same status — cosmetic.
  • panic_traceback_wasm.go: printWasmPanicTraceback reads getg().goid after guarding only store != nil; no realistic nil-getg() path on the fatal-panic path, noted for completeness.

Comment thread internal/build/build.go
Comment thread runtime/internal/runtime/tinygogc/finalizer.go Outdated
@github-actions

github-actions Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

748483792502 | workflow run | long-term charts

WebAssembly output sizes

Profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 135633 B +21658 B / +19.0% (worse) 70786 B +50 B / +0.1% (worse)
cprintf/j32-goos-js/LLGo 133904 B +66566 B / +98.9% (worse) 69150 B +639 B / +0.9% (worse)
cprintf/j64-emscripten-memory64/LLGo 124275 B +4821 B / +4.0% (worse) 73998 B -35 B / -0.04728% (better)
cprintf/w32-goos-wasip1/LLGo 138059 B +64084 B / +86.6% (worse) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 138268 B +20235 B / +17.1% (worse) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3195716 B +221541 B / +7.4% (worse) 114540 B +5834 B / +5.4% (worse)
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3188314 B +825227 B / +34.9% (worse) 98239 B -8402 B / -7.9% (better)
fmtprintf/j64-emscripten-memory64/LLGo 2927596 B -236126 B / -7.5% (better) 121334 B +663 B / +0.5% (worse)
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2953757 B +802783 B / +37.3% (worse) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2818531 B +73681 B / +2.7% (worse) 0 B 0 B / 0.0%
j32-emscripten/LLGo 134887 B +21644 B / +19.1% (worse) 70786 B +50 B / +0.1% (worse)
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 133372 B +66712 B / +100.1% (worse) 69150 B +639 B / +0.9% (worse)
j64-emscripten-memory64/LLGo 123605 B +4917 B / +4.1% (worse) 73998 B -35 B / -0.04728% (better)
reflectcall/j32-emscripten/LLGo 1580370 B +27808 B / +1.8% (worse) 88948 B +1375 B / +1.6% (worse)
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1581809 B +321816 B / +25.5% (worse) 87312 B +2099 B / +2.5% (worse)
reflectcall/j64-emscripten-memory64/LLGo 1464062 B -185554 B / -11.2% (better) 94487 B -35 B / -0.03703% (better)
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1635842 B +381005 B / +30.4% (worse) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1562853 B +43743 B / +2.9% (worse) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 137288 B +64371 B / +88.3% (worse) 0 B 0 B / 0.0%
w32-wasi/LLGo 137593 B +20234 B / +17.2% (worse) 0 B 0 B / 0.0%

LLGo WebAssembly build measurements

Profile Build vs base
j32-emscripten 5.750 s +447.1 ms / +8.4% (worse)
j32-goos-js 5.741 s +1.393 s / +32.0% (worse)
j64-emscripten-memory64 5.068 s +177.7 ms / +3.6% (worse)
reflectcall/w32-wasi 28.687 s -7.241 s / -20.2% (better)
w32-goos-wasip1 4.494 s +1.53 s / +51.6% (worse)
w32-wasi 4.384 s +672 ms / +18.1% (worse)

Compared with 4564a01aa520 measured in the same runner job.

@github-actions

github-actions Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

748483792502 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 19992 B +160 B / +0.8% (worse) 387 B 0 B / +0.0% 448.751 ms -19.81 ms / -4.2% (better) 1.296 ms +31.32 us / +2.5% (worse)
Linux cprintf-lto 19744 B +160 B / +0.8% (worse) 368 B 0 B / +0.0% 460.038 ms -8.288 ms / -1.8% (better) 1.215 ms -70.99 us / -5.5% (better)
Linux fmtprintf 1629232 B +3352 B / +0.2% (worse) 498716 B +754 B / +0.2% (worse) 3.716 s -37.66 ms / -1.0% (better) 3.048 ms +46.08 us / +1.5% (worse)
Linux fmtprintf-lto 1482200 B +2024 B / +0.1% (worse) 438154 B +333 B / +0.1% (worse) 11.041 s +198.1 ms / +1.8% (worse) 2.882 ms -96.73 us / -3.2% (better)
Linux println 62368 B +144 B / +0.2% (worse) 14880 B -21 B / -0.1% (better) 469.270 ms +23.9 ms / +5.4% (worse) 1.553 ms -15.08 us / -1.0% (better)
Linux println-lto 54440 B +160 B / +0.3% (worse) 12335 B 0 B / +0.0% 668.347 ms -65.73 ms / -9.0% (better) 1.504 ms -110.7 us / -6.9% (better)
macOS cprintf 84480 B 0 B / +0.0% 17261 B +160 B / +0.9% (worse) 783.263 ms +99.13 ms / +14.5% (worse) 2.968 ms -1.204 ms / -28.9% (better)
macOS cprintf-lto 84288 B 0 B / +0.0% 13025 B +160 B / +1.2% (worse) 610.946 ms -770.4 ms / -55.8% (better) 2.359 ms -7.373 ms / -75.8% (better)
macOS fmtprintf 1473584 B +416 B / +0.02824% (worse) 874528 B +2516 B / +0.3% (worse) 3.654 s -285.2 ms / -7.2% (better) 4.095 ms -3.055 ms / -42.7% (better)
macOS fmtprintf-lto 1175824 B +16400 B / +1.4% (worse) 849772 B +2020 B / +0.2% (worse) 8.633 s -632.5 ms / -6.8% (better) 5.407 ms +38.08 us / +0.7% (worse)
macOS println 114672 B 0 B / +0.0% 34978 B +166 B / +0.5% (worse) 607.948 ms -426.1 ms / -41.2% (better) 3.847 ms -1.345 ms / -25.9% (better)
macOS println-lto 118720 B 0 B / +0.0% 32408 B +160 B / +0.5% (worse) 946.286 ms -192.9 ms / -16.9% (better) 6.515 ms +2.411 ms / +58.7% (worse)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.119 s +21.78 ms / +2.0% (worse) 3.399 ms +14.7 us / +0.4% (worse)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.170 s +23.52 ms / +2.1% (worse) 3.385 ms -53.8 us / -1.6% (better)
Windows MinGW fmtprintf 1897984 B +2560 B / +0.1% (worse) 602102 B +992 B / +0.2% (worse) 4.080 s +7.699 ms / +0.2% (worse) 8.131 ms +58.4 us / +0.7% (worse)
Windows MinGW fmtprintf-lto 1939968 B +3072 B / +0.2% (worse) 551062 B +496 B / +0.1% (worse) 10.129 s +48.86 ms / +0.5% (worse) 7.841 ms +434.7 us / +5.9% (worse)
Windows MinGW println 71168 B 0 B / +0.0% 24022 B -32 B / -0.1% (better) 1.147 s +43.23 ms / +3.9% (worse) 6.511 ms +15.9 us / +0.2% (worse)
Windows MinGW println-lto 65536 B 0 B / +0.0% 20678 B 0 B / +0.0% 1.312 s -11.34 ms / -0.9% (better) 6.461 ms -47.9 us / -0.7% (better)
Windows MinGW 386 cprintf 42496 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.099 s -5.35 ms / -0.5% (better) 5.012 ms -576 us / -10.3% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.167 s +27.3 ms / +2.4% (worse) 5.385 ms +22.4 us / +0.4% (worse)
Windows MinGW 386 fmtprintf 1863168 B +4608 B / +0.2% (worse) 473294 B +864 B / +0.2% (worse) 4.151 s +24.52 ms / +0.6% (worse) 10.827 ms +87.9 us / +0.8% (worse)
Windows MinGW 386 fmtprintf-lto 2169344 B +3584 B / +0.2% (worse) 453178 B +656 B / +0.1% (worse) 9.675 s -24.65 ms / -0.3% (better) 10.603 ms -325.9 us / -3.0% (better)
Windows MinGW 386 println 91136 B 0 B / +0.0% 20006 B -32 B / -0.2% (better) 1.098 s +8.283 ms / +0.8% (worse) 8.758 ms -569.1 us / -6.1% (better)
Windows MinGW 386 println-lto 69632 B 0 B / +0.0% 17954 B 0 B / +0.0% 1.304 s -881.2 us / -0.1% (better) 8.437 ms -302.8 us / -3.5% (better)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.490 s +30.31 ms / +2.1% (worse) 7.534 ms +1.228 ms / +19.5% (worse)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.510 s +38.97 ms / +2.6% (worse) 6.653 ms -34.8 us / -0.5% (better)
Windows MinGW ARM64 fmtprintf 1784320 B +3584 B / +0.2% (worse) 511704 B +1008 B / +0.2% (worse) 4.334 s +11.24 ms / +0.3% (worse) 13.592 ms -66.2 us / -0.5% (better)
Windows MinGW ARM64 fmtprintf-lto 1862656 B +2560 B / +0.1% (worse) 479308 B +396 B / +0.1% (worse) 10.114 s -45.23 ms / -0.4% (better) 13.729 ms +47.8 us / +0.3% (worse)
Windows MinGW ARM64 println 68608 B 0 B / +0.0% 22544 B 0 B / +0.0% 1.479 s +47.89 ms / +3.3% (worse) 10.946 ms +190.1 us / +1.8% (worse)
Windows MinGW ARM64 println-lto 64000 B 0 B / +0.0% 19604 B 0 B / +0.0% 1.677 s +43.48 ms / +2.7% (worse) 11.637 ms +433.7 us / +3.9% (worse)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65782 B 0 B / +0.0% 1.143 s -18.48 ms / -1.6% (better) 3.354 ms -753.1 us / -18.3% (better)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65718 B 0 B / +0.0% 1.010 s -31.72 ms / -3.0% (better) 3.454 ms -431.2 us / -11.1% (better)
Windows MSVC fmtprintf 1628672 B +2560 B / +0.2% (worse) 697622 B +992 B / +0.1% (worse) 4.078 s -144.2 ms / -3.4% (better) 11.406 ms +992.4 us / +9.5% (worse)
Windows MSVC fmtprintf-lto 1627136 B +2048 B / +0.1% (worse) 653398 B +512 B / +0.1% (worse) 9.901 s -65.94 ms / -0.7% (better) 10.551 ms +180.7 us / +1.7% (worse)
Windows MSVC println 193024 B 0 B / +0.0% 119446 B -32 B / -0.02678% (better) 1.006 s -9.803 ms / -1.0% (better) 7.920 ms -872.5 us / -9.9% (better)
Windows MSVC println-lto 189952 B 0 B / +0.0% 116614 B 0 B / +0.0% 1.187 s -22.71 ms / -1.9% (better) 8.050 ms -311.3 us / -3.7% (better)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.027 s +13.09 ms / +1.3% (worse) 6.038 ms -34.3 us / -0.6% (better)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.235 s +213.6 ms / +20.9% (worse) 6.080 ms +447.1 us / +7.9% (worse)
Windows MSVC 386 fmtprintf 1193472 B +2560 B / +0.2% (worse) 456661 B +864 B / +0.2% (worse) 4.092 s -51.46 ms / -1.2% (better) 13.174 ms -811.7 us / -5.8% (better)
Windows MSVC 386 fmtprintf-lto 1234432 B +2560 B / +0.2% (worse) 431993 B +576 B / +0.1% (worse) 9.203 s +40.48 ms / +0.4% (worse) 13.389 ms +1.107 ms / +9.0% (worse)
Windows MSVC 386 println 34304 B 0 B / +0.0% 18849 B -16 B / -0.1% (better) 1.024 s +28.77 ms / +2.9% (worse) 10.884 ms +723.4 us / +7.1% (worse)
Windows MSVC 386 println-lto 32768 B 0 B / +0.0% 17015 B 0 B / +0.0% 1.183 s +9.511 ms / +0.8% (worse) 9.477 ms -575.4 us / -5.7% (better)
Windows MSVC ARM64 cprintf 11264 B 0 B / +0.0% 3976 B 0 B / +0.0% 2.141 s +3.165 ms / +0.1% (worse) 7.970 ms +545.5 us / +7.3% (worse)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 3868 B 0 B / +0.0% 2.122 s -25.92 ms / -1.2% (better) 7.941 ms +581.9 us / +7.9% (worse)
Windows MSVC ARM64 fmtprintf 1374720 B +3072 B / +0.2% (worse) 512532 B +1008 B / +0.2% (worse) 6.594 s -8.736 ms / -0.1% (better) 16.235 ms +1.171 ms / +7.8% (worse)
Windows MSVC ARM64 fmtprintf-lto 1397760 B +2048 B / +0.1% (worse) 482100 B +384 B / +0.1% (worse) 17.092 s +207.7 ms / +1.2% (worse) 15.326 ms -780.2 us / -4.8% (better)
Windows MSVC ARM64 println 41472 B 0 B / +0.0% 21784 B 0 B / +0.0% 2.099 s +56.98 ms / +2.8% (worse) 14.225 ms +693.8 us / +5.1% (worse)
Windows MSVC ARM64 println-lto 39936 B 0 B / +0.0% 19532 B 0 B / +0.0% 2.447 s +106.1 ms / +4.5% (worse) 12.179 ms -298.1 us / -2.4% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.520 ns/op -0.2 ns/op / -1.4% (better)
Linux BenchmarkMergeCompilerFlags 195.900 ns/op -4.4 ns/op / -2.2% (better)
Linux BenchmarkMergeLinkerFlags 142.300 ns/op +11 ns/op / +8.4% (worse)
Linux BenchmarkChannelBuffered 56.050 ns/op +0.09 ns/op / +0.2% (worse)
Linux BenchmarkChannelHandoff 13792 ns/op +857 ns/op / +6.6% (worse)
Linux BenchmarkDefer 49.010 ns/op +3.18 ns/op / +6.9% (worse)
Linux BenchmarkDirectCall 1.165 ns/op -0.391 ns/op / -25.1% (better)
Linux BenchmarkGlobalRead 1.941 ns/op +0.774 ns/op / +66.3% (worse)
Linux BenchmarkGlobalWrite 7.775 ns/op +0.013 ns/op / +0.2% (worse)
Linux BenchmarkGoroutine 29299 ns/op +5084 ns/op / +21.0% (worse)
Linux BenchmarkInterfaceCall 5.892 ns/op -0.401 ns/op / -6.4% (better)
Linux BenchmarkRuntimeGetG 2.859 ns/op +0.358 ns/op / +14.3% (worse)
macOS BenchmarkLookupPCRandom 15.560 ns/op +2.97 ns/op / +23.6% (worse)
macOS BenchmarkMergeCompilerFlags 171.400 ns/op +65.8 ns/op / +62.3% (worse)
macOS BenchmarkMergeLinkerFlags 96.300 ns/op +28.56 ns/op / +42.2% (worse)
macOS BenchmarkChannelBuffered 34.540 ns/op +8.17 ns/op / +31.0% (worse)
macOS BenchmarkChannelHandoff 9817 ns/op +588 ns/op / +6.4% (worse)
macOS BenchmarkDefer 47.490 ns/op +12.9 ns/op / +37.3% (worse)
macOS BenchmarkDirectCall 1.289 ns/op +0.17 ns/op / +15.2% (worse)
macOS BenchmarkGlobalRead 1.376 ns/op +0.228 ns/op / +19.9% (worse)
macOS BenchmarkGlobalWrite 1.412 ns/op -0.04 ns/op / -2.8% (better)
macOS BenchmarkGoroutine 50941 ns/op +6797 ns/op / +15.4% (worse)
macOS BenchmarkInterfaceCall 4.829 ns/op +0.575 ns/op / +13.5% (worse)
macOS BenchmarkRuntimeGetG 2.294 ns/op -0.05 ns/op / -2.1% (better)
Windows MinGW BenchmarkLookupPCRandom 13.180 ns/op +0.46 ns/op / +3.6% (worse)
Windows MinGW BenchmarkMergeCompilerFlags 606.400 ns/op -1 ns/op / -0.2% (better)
Windows MinGW BenchmarkMergeLinkerFlags 547.800 ns/op +14.2 ns/op / +2.7% (worse)
Windows MinGW BenchmarkChannelBuffered 36.990 ns/op +2.42 ns/op / +7.0% (worse)
Windows MinGW BenchmarkChannelHandoff 938 ns/op +26.8 ns/op / +2.9% (worse)
Windows MinGW BenchmarkDefer 59.440 ns/op +3.62 ns/op / +6.5% (worse)
Windows MinGW BenchmarkDirectCall 1.551 ns/op +0.005 ns/op / +0.3% (worse)
Windows MinGW BenchmarkGlobalRead 1.857 ns/op +0.31 ns/op / +20.0% (worse)
Windows MinGW BenchmarkGlobalWrite 2.453 ns/op -0.016 ns/op / -0.6% (better)
Windows MinGW BenchmarkGoroutine 81888 ns/op -3674 ns/op / -4.3% (better)
Windows MinGW BenchmarkInterfaceCall 8.690 ns/op +0.32 ns/op / +3.8% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.171 ns/op -0.008 ns/op / -0.4% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.540 ns/op +0.07 ns/op / +0.3% (worse)
Windows MinGW 386 BenchmarkMergeCompilerFlags 687.300 ns/op -74 ns/op / -9.7% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 636.100 ns/op -46 ns/op / -6.7% (better)
Windows MinGW 386 BenchmarkChannelBuffered 40.480 ns/op -1.28 ns/op / -3.1% (better)
Windows MinGW 386 BenchmarkChannelHandoff 967.500 ns/op -69.5 ns/op / -6.7% (better)
Windows MinGW 386 BenchmarkDefer 44.400 ns/op +2.55 ns/op / +6.1% (worse)
Windows MinGW 386 BenchmarkDirectCall 2.169 ns/op +0.623 ns/op / +40.3% (worse)
Windows MinGW 386 BenchmarkGlobalRead 1.550 ns/op -0.31 ns/op / -16.7% (better)
Windows MinGW 386 BenchmarkGlobalWrite 7.769 ns/op -0.001 ns/op / -0.01287% (better)
Windows MinGW 386 BenchmarkGoroutine 83962 ns/op -1263 ns/op / -1.5% (better)
Windows MinGW 386 BenchmarkInterfaceCall 8.347 ns/op -0.051 ns/op / -0.6% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 2.477 ns/op +0.311 ns/op / +14.4% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12 ns/op -0.02 ns/op / -0.2% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 584.500 ns/op -3.3 ns/op / -0.6% (better)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 547.500 ns/op +0.6 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 39.080 ns/op +0.26 ns/op / +0.7% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 1849 ns/op +42 ns/op / +2.3% (worse)
Windows MinGW ARM64 BenchmarkDefer 55.730 ns/op +0.24 ns/op / +0.4% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.590 ns/op +0.0001 ns/op / +0.01695% (worse)
Windows MinGW ARM64 BenchmarkGlobalRead 0.884 ns/op +0.22 ns/op / +33.1% (worse)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.590 ns/op -0.294 ns/op / -33.3% (better)
Windows MinGW ARM64 BenchmarkGoroutine 59714 ns/op -1974 ns/op / -3.2% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.324 ns/op +0.016 ns/op / +0.4% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.772 ns/op -0.031 ns/op / -1.7% (better)
Windows MSVC BenchmarkLookupPCRandom 13.220 ns/op +0.06 ns/op / +0.5% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 628.400 ns/op +5.3 ns/op / +0.9% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 533.300 ns/op -5.1 ns/op / -0.9% (better)
Windows MSVC BenchmarkChannelBuffered 32.860 ns/op +0.27 ns/op / +0.8% (worse)
Windows MSVC BenchmarkChannelHandoff 1163 ns/op +26 ns/op / +2.3% (worse)
Windows MSVC BenchmarkDefer 56.510 ns/op +0.27 ns/op / +0.5% (worse)
Windows MSVC BenchmarkDirectCall 1.858 ns/op +0.311 ns/op / +20.1% (worse)
Windows MSVC BenchmarkGlobalRead 1.549 ns/op -0.312 ns/op / -16.8% (better)
Windows MSVC BenchmarkGlobalWrite 2.471 ns/op +0.01 ns/op / +0.4% (worse)
Windows MSVC BenchmarkGoroutine 81358 ns/op +236 ns/op / +0.3% (worse)
Windows MSVC BenchmarkInterfaceCall 8.379 ns/op +0.005 ns/op / +0.1% (worse)
Windows MSVC BenchmarkRuntimeGetG 1.863 ns/op -0.309 ns/op / -14.2% (better)
Windows MSVC 386 BenchmarkLookupPCRandom 26.540 ns/op -0.03 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 783.200 ns/op -1.6 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 692.900 ns/op -4.8 ns/op / -0.7% (better)
Windows MSVC 386 BenchmarkChannelBuffered 41.110 ns/op -4.5 ns/op / -9.9% (better)
Windows MSVC 386 BenchmarkChannelHandoff 955.500 ns/op -59.5 ns/op / -5.9% (better)
Windows MSVC 386 BenchmarkDefer 49.310 ns/op -0.81 ns/op / -1.6% (better)
Windows MSVC 386 BenchmarkDirectCall 1.550 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalRead 2.172 ns/op +0.623 ns/op / +40.2% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 7.770 ns/op -0.014 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkGoroutine 87697 ns/op -3827 ns/op / -4.2% (better)
Windows MSVC 386 BenchmarkInterfaceCall 8.163 ns/op -0.222 ns/op / -2.6% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 2.175 ns/op +0.247 ns/op / +12.8% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.020 ns/op -0.04 ns/op / -0.3% (better)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 581.500 ns/op -41.5 ns/op / -6.7% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 551.200 ns/op -31.3 ns/op / -5.4% (better)
Windows MSVC ARM64 BenchmarkChannelBuffered 39.570 ns/op +0.27 ns/op / +0.7% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 2193 ns/op +41 ns/op / +1.9% (worse)
Windows MSVC ARM64 BenchmarkDefer 67.380 ns/op -1.27 ns/op / -1.8% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.663 ns/op +0.0736 ns/op / +12.5% (worse)
Windows MSVC ARM64 BenchmarkGlobalRead 0.589 ns/op -0.076 ns/op / -11.4% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.747 ns/op -0.007 ns/op / -0.2% (better)
Windows MSVC ARM64 BenchmarkGoroutine 53907 ns/op -884 ns/op / -1.6% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.154 ns/op -0.028 ns/op / -0.7% (better)
Windows MSVC ARM64 BenchmarkRuntimeGetG 2.102 ns/op +0.315 ns/op / +17.6% (worse)

Timer runtime benchmarks

Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 908.800 ns/op +7.1 ns/op / +0.8% (worse)
Linux AfterFuncZeroDelivery/LLGo 39588 ns/op +6689 ns/op / +20.3% (worse)
Linux CreateStop/Go 290.200 ns/op +1.1 ns/op / +0.4% (worse)
Linux CreateStop/LLGo 1788 ns/op +515 ns/op / +40.5% (worse)
Linux RearmStopped/Go 116 ns/op +0.2 ns/op / +0.2% (worse)
Linux RearmStopped/LLGo 1158 ns/op -56 ns/op / -4.6% (better)
Linux ResetActive/Go 68.590 ns/op +0.03 ns/op / +0.04376% (worse)
Linux ResetActive/LLGo 659.400 ns/op -150 ns/op / -18.5% (better)
Linux ResetHeap1024/Go 67.160 ns/op 0 ns/op / +0.0%
Linux ResetHeap1024/LLGo 181.300 ns/op -7.9 ns/op / -4.2% (better)
macOS AfterFuncZeroDelivery/Go 573.100 ns/op +113.6 ns/op / +24.7% (worse)
macOS AfterFuncZeroDelivery/LLGo 73384 ns/op +8057 ns/op / +12.3% (worse)
macOS CreateStop/Go 208.800 ns/op +62.8 ns/op / +43.0% (worse)
macOS CreateStop/LLGo 608.300 ns/op +153.9 ns/op / +33.9% (worse)
macOS RearmStopped/Go 84.050 ns/op +23.92 ns/op / +39.8% (worse)
macOS RearmStopped/LLGo 323.700 ns/op -22.1 ns/op / -6.4% (better)
macOS ResetActive/Go 51.680 ns/op +8.39 ns/op / +19.4% (worse)
macOS ResetActive/LLGo 147 ns/op -20.5 ns/op / -12.2% (better)
macOS ResetHeap1024/Go 51.480 ns/op +7.5 ns/op / +17.1% (worse)
macOS ResetHeap1024/LLGo 91.950 ns/op +6.95 ns/op / +8.2% (worse)
Windows MinGW AfterFuncZeroDelivery/Go 544.400 ns/op -27.7 ns/op / -4.8% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 163330 ns/op +2668 ns/op / +1.7% (worse)
Windows MinGW CreateStop/Go 115.300 ns/op +0.2 ns/op / +0.2% (worse)
Windows MinGW CreateStop/LLGo 438.200 ns/op +10.1 ns/op / +2.4% (worse)
Windows MinGW RearmStopped/Go 31.580 ns/op +0.25 ns/op / +0.8% (worse)
Windows MinGW RearmStopped/LLGo 271.300 ns/op -3.4 ns/op / -1.2% (better)
Windows MinGW ResetActive/Go 20.150 ns/op +0.09 ns/op / +0.4% (worse)
Windows MinGW ResetActive/LLGo 171.600 ns/op -86.8 ns/op / -33.6% (better)
Windows MinGW ResetHeap1024/Go 20.560 ns/op +0.11 ns/op / +0.5% (worse)
Windows MinGW ResetHeap1024/LLGo 125.800 ns/op 0 ns/op / +0.0%
Windows MinGW 386 AfterFuncZeroDelivery/Go 953.400 ns/op -10.2 ns/op / -1.1% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 189097 ns/op +552 ns/op / +0.3% (worse)
Windows MinGW 386 CreateStop/Go 190.700 ns/op -1.2 ns/op / -0.6% (better)
Windows MinGW 386 CreateStop/LLGo 1613 ns/op -3 ns/op / -0.2% (better)
Windows MinGW 386 RearmStopped/Go 63.180 ns/op -0.28 ns/op / -0.4% (better)
Windows MinGW 386 RearmStopped/LLGo 338.500 ns/op -4.4 ns/op / -1.3% (better)
Windows MinGW 386 ResetActive/Go 39.160 ns/op +0.19 ns/op / +0.5% (worse)
Windows MinGW 386 ResetActive/LLGo 972.100 ns/op +626.6 ns/op / +181.4% (worse)
Windows MinGW 386 ResetHeap1024/Go 39.550 ns/op +0.15 ns/op / +0.4% (worse)
Windows MinGW 386 ResetHeap1024/LLGo 186.300 ns/op -6.2 ns/op / -3.2% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 676 ns/op +6.4 ns/op / +1.0% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 130437 ns/op -10638 ns/op / -7.5% (better)
Windows MinGW ARM64 CreateStop/Go 206.400 ns/op +4.6 ns/op / +2.3% (worse)
Windows MinGW ARM64 CreateStop/LLGo 429.600 ns/op -7.6 ns/op / -1.7% (better)
Windows MinGW ARM64 RearmStopped/Go 70.650 ns/op +0.06 ns/op / +0.1% (worse)
Windows MinGW ARM64 RearmStopped/LLGo 265.700 ns/op +3 ns/op / +1.1% (worse)
Windows MinGW ARM64 ResetActive/Go 30.960 ns/op -0.09 ns/op / -0.3% (better)
Windows MinGW ARM64 ResetActive/LLGo 151.400 ns/op +11.1 ns/op / +7.9% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.190 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 127.200 ns/op -0.8 ns/op / -0.6% (better)
Windows MSVC AfterFuncZeroDelivery/Go 573 ns/op +8.3 ns/op / +1.5% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 162340 ns/op +8547 ns/op / +5.6% (worse)
Windows MSVC CreateStop/Go 118.400 ns/op -2.4 ns/op / -2.0% (better)
Windows MSVC CreateStop/LLGo 462.500 ns/op +1.2 ns/op / +0.3% (worse)
Windows MSVC RearmStopped/Go 31.480 ns/op -0.06 ns/op / -0.2% (better)
Windows MSVC RearmStopped/LLGo 267.300 ns/op +3.4 ns/op / +1.3% (worse)
Windows MSVC ResetActive/Go 20.110 ns/op +0.04 ns/op / +0.2% (worse)
Windows MSVC ResetActive/LLGo 141.500 ns/op +2.6 ns/op / +1.9% (worse)
Windows MSVC ResetHeap1024/Go 20.380 ns/op -0.05 ns/op / -0.2% (better)
Windows MSVC ResetHeap1024/LLGo 125.100 ns/op -0.4 ns/op / -0.3% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 955.200 ns/op -16.9 ns/op / -1.7% (better)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 196356 ns/op -5307 ns/op / -2.6% (better)
Windows MSVC 386 CreateStop/Go 199.700 ns/op -8.9 ns/op / -4.3% (better)
Windows MSVC 386 CreateStop/LLGo 1889 ns/op -21 ns/op / -1.1% (better)
Windows MSVC 386 RearmStopped/Go 63.680 ns/op -0.05 ns/op / -0.1% (better)
Windows MSVC 386 RearmStopped/LLGo 334.700 ns/op -15.9 ns/op / -4.5% (better)
Windows MSVC 386 ResetActive/Go 39.020 ns/op -0.13 ns/op / -0.3% (better)
Windows MSVC 386 ResetActive/LLGo 802.100 ns/op -118 ns/op / -12.8% (better)
Windows MSVC 386 ResetHeap1024/Go 39.460 ns/op 0 ns/op / +0.0%
Windows MSVC 386 ResetHeap1024/LLGo 176 ns/op +4.3 ns/op / +2.5% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 684.500 ns/op +26.5 ns/op / +4.0% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 129338 ns/op +3582 ns/op / +2.8% (worse)
Windows MSVC ARM64 CreateStop/Go 196.200 ns/op -5.3 ns/op / -2.6% (better)
Windows MSVC ARM64 CreateStop/LLGo 376 ns/op -8 ns/op / -2.1% (better)
Windows MSVC ARM64 RearmStopped/Go 70.580 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 RearmStopped/LLGo 276 ns/op +5.3 ns/op / +2.0% (worse)
Windows MSVC ARM64 ResetActive/Go 31.070 ns/op +0.09 ns/op / +0.3% (worse)
Windows MSVC ARM64 ResetActive/LLGo 133.800 ns/op -0.8 ns/op / -0.6% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.140 ns/op +0.12 ns/op / +0.4% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 139.400 ns/op +2.1 ns/op / +1.5% (worse)

Compared with 4564a01aa520 measured in the same runner job.

@cpunion

cpunion commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai Please re-review the incremental range a82726740fba..311a615707b9. The three inline findings are addressed: finalizer ready-queue processing is now linear with a deterministic 1,024-record scaling test; the stale queueing comment was removed; and Wasm C-export wrapper modules now run post-C-ABI aggregate-copy lowering, optimization, and stack-address localization with an actual 8 KiB wrapper regression test. The reflect constant-GEP, aggregate fixed-point termination, and uintptr-root coupling notes also now have explicit invariants/cross-references. Focused lifecycle tests pass on J32 Emscripten (wasm32 with Emscripten/JavaScript), J64 Emscripten Memory64 (wasm64 with Emscripten/JavaScript), and W32 WASI (wasm32 with WASI Preview 1).

@fennoai

fennoai Bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Re-reviewing PR #2580 (incremental)

Re-reviewing the incremental range a82726740fba..311a615707b9 (7 files) against my prior findings.

  • Prepare workspace and lock diff range
  • Verify each prior finding is addressed
  • Submit follow-up review

View job run

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review: incremental range a82726740fba..311a615707b9 (7 files)

All prior findings are resolved. I read the full incremental diff and traced the semantics of each change; the fixes are correct and backed by targeted tests. No new issues found. I could not run the tests locally (this sandbox has no LLVM C headers, so the cgo-bound packages don't build); the review below is from source inspection, and the PR reports the focused lifecycle tests passing on J32/J64/W32 CI.

[P2] Wasm C-export wrapper aggregate lowering — resolved.
buildMainLink now calls the new lowerMainCExportModule, which for GOARCH == "wasm" runs lowerWasmAggregateCopies, applySizeOptimizationAttributes, the LLVM optimization pipeline, and localizeWasmStackAddresses — matching compilePackageModule. The opt pipeline is nicely factored into a shared optimizeLLVMModule used by both paths, and errors now propagate instead of being dropped. TestLowerMainCExportModuleWasmCopies builds an 8 KiB-array export wrapper and asserts the load/store [8192 x i8] are replaced by @llvm.mem*, i.e. the copy pass actually ran on the wrapper module.

[P3] Stale finalizer comment — resolved, and the loop is now genuinely linear.
The rewrite replaces the O(D·N) nested scan (and its record = finalizers restart) with a single-pass unlink over finalizers, deleting candidateForObject/earlierFinalizerForObject/hasCandidateFinalizer/finalizerObjectBlocked/queueCallbacksForObject. I checked the behavior is preserved:

  • Dropping earlierFinalizerForObject in the dependency scan is safe — startMark operates on the object block and is idempotent under the blockStateHead gate, so starting from every candidate finalizer of an object equals starting once.
  • noteFinalizerReference → markFinalizerObjectBlocked now gates on record.candidate; since blocked is only ever read on candidate records, this is equivalent to the old outer candidateForObject != nil guard.
  • The two-phase "unlink all eligible, then startMark" split correctly resurrects only this pass's newly-queued records — readyFinalizers is a LIFO stack and the record != readyBefore bound walks exactly the records pushed in this pass. candidate implies finalizerActive (set at the top), so the old state == finalizerActive filter is preserved.

TestCandidateTraversalScalesLinearly (1,024 records, instrumented registryReads) makes the linear-scaling claim deterministic rather than timeout-based — a good approach.

Note items — addressed. The uintptr-root single-slot reservation now has paired invariant comments in cl/gcroot.go (prepareGCRoots ↔ publishGCRoot), the LowerWasmAggregateCopies fixed-point termination is documented as monotonic in internal/abi/large.go, and the reflect constant-GEP folding is cross-referenced in ssa/reflect_bridge.go (wasmReflectSlot/wasmReflectFrame).

LGTM.

@cpunion
cpunion force-pushed the codex/wasm-w2-runtime-lowering-20260913 branch from 311a615 to a50fd7c Compare September 13, 2026 18:49
@cpunion

cpunion commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai Please re-review the rebased intended W2-B range 8a7f17a8a227..a50fd7c6aee5. Twenty-five of 26 dedicated commits are exact range-diff matches. The only reconciled commit preserves main’s native hardware-fault panic-site recording while using W2-B’s explicit WebAssembly nil guard exactly once; its W2-B-specific and upstream native panic regression tests pass. Please also verify that the prior finalizer-complexity, aggregate-wrapper, root-timing, and documentation findings remain resolved.

@fennoai

fennoai Bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Re-reviewing PR #2580 (rebased W2-B)

Re-reviewing rebased range 8a7f17a8a227..a50fd7c6aee5. Range-diff confirms 25/26 commits are exact matches; verifying the one reconciled commit and prior findings.

  • Prepare workspace and lock diff range
  • Confirm range-diff equivalence vs previously-reviewed range
  • Review the reconciled panic-site / nil-guard commit
  • Confirm prior findings remain resolved
  • Submit review

View job run

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review: rebased W2-B range 8a7f17a8a227..a50fd7c6aee5

Verified via git range-diff against the range I reviewed and approved last round. 25 of 26 commits are exact range-diff matches (=); the only changed commit is 7827393e2 fix(wasm): preserve single-worker Go semantics, and its sole textual delta is the panic-site/nil-guard reconciliation described in the request. No new issues found. (As before, I couldn't execute tests locally — this sandbox lacks LLVM C headers so the cgo-bound packages don't build; the reconciled commit is small and verified by reading. The PR reports its W2-B-specific and upstream native panic regression tests passing.)

Reconciled commit — correct. In the field-store path (cl/compile.go:2270-2284) the change computes wasmNilGuard := p.needsWasmNilGuard(va) once, then:

  • WASM (wasmNilGuard true): records the panic site once, emits the base check + explicit AssertNilDeref; the native branch is skipped by !wasmNilGuard.
  • Native (wasmNilGuard always false, since needsWasmNilGuard is gated on GOARCH == "wasm"): records the panic site iff !isKnownNonNilAddr && !isWrapNilCheckCall, preserving main's hardware-fault panic-site behavior; no explicit guard.

The two recordPanicSite branches are mutually exclusive, so the naive union's double-record can't occur. I checked all four combinations: on WASM with a provably-non-nil address, needsWasmNilGuard is false and the native branch's !isKnownNonNilAddr && !isWrapNilCheckCall is false, so nothing is recorded — correct, since such a store can't fault. Logic is sound across the matrix.

Prior findings remain resolved. git diff 311a615707b9 a50fd7c6aee5 over runtime/.../tinygogc/finalizer.go, internal/build/build.go, internal/abi/large.go, cl/gcroot.go, and ssa/reflect_bridge.go is empty — these are byte-for-byte identical to the versions confirmed resolved last round:

  • Finalizer ready traversal is the single-pass linear form (the O(D·N) helpers queueCallbacksForObject/candidateForObject/hasCandidateFinalizer/earlierFinalizerForObject/finalizerObjectBlocked are gone).
  • lowerMainCExportModule still runs lowerWasmAggregateCopies + opt pipeline + localizeWasmStackAddresses on wasm export wrappers.
  • Medium-aggregate root timing and the documentation/invariant comments are unchanged.

LGTM. The rebase is clean and the single reconciliation is correct.

@cpunion

cpunion commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai Please re-review the incremental range a50fd7c6aee5..dff40ac32f8b. W3 acceptance against newly merged main exposed two WASM caller-state regressions: same-value repanic lost the original shadow-stack prefix, and goroutine-local synthetic PCs collided in the process-wide runtime.FuncForPC cache. The fix freezes the existing prefix without copying, uses raw interface identity plus the recover activation to retain it only for a true repanic, assigns process-unique synthetic PCs only for hosted LLGo builds, and keeps bare-metal on the non-atomic single-store path. New scheduler coverage runs normal caller-cache isolation and nested repanic tracebacks on J32 Emscripten (wasm32 with Emscripten/JavaScript), J64 Emscripten Memory64 (wasm64 with Emscripten/JavaScript), and W32 WASI (wasm32 with WASI Preview 1). Please focus on prefix lifetime, replacement/later-same-value behavior, PC sequence wrap/lookup, bare-metal isolation, and the measured +0.65% raw/+0.63% gzip incremental module size.

@fennoai

fennoai Bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Re-reviewing PR #2580 (incremental)

Re-reviewing incremental range a50fd7c6aee5..dff40ac32f8b (1 commit, 10 files) — the two WASM caller-state regression fixes.

  • Prepare workspace and lock diff range
  • Review prefix lifetime + repanic identity
  • Review synthetic-PC sequence / lookup / bare-metal isolation
  • Submit review

View job run

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review: incremental range a50fd7c6aee5..dff40ac32f8b (1 commit, 10 files)

Reviewed the single commit dff40ac32 fix(wasm): preserve panic caller identity in full. Both regression fixes are correct and well-covered. No new issues found. (As in prior rounds I couldn't execute tests locally — no LLVM C headers in this sandbox; verified by reading. The PR reports the new scheduler coverage passing on J32/J64/W32.)

Prefix lifetime — correct. capturePanicCallerFrames freezes the prefix in place by recording panicDepth = len(store.stack) (no copy). PopCallerLocationFrame (caller.go:87-91) only raises mark up to panicDepth, so a repanic longjmp can't truncate below the frozen prefix. The prefix is released by clearPanicCallerSnapshot() (panicDepth = 0) on every recover-completion path — EndRecoverFrame, EndRecoverFrameAlias, and abortPanics (the latter also covers Goexit/teardown). Both PopCallerLocationFrame and activePanicCallerFrames/panicCallerSnapshotAvailable additionally guard panicDepth <= len(store.stack), so a stale depth exceeding a shortened stack is safely ignored rather than reading OOB.

Replacement / later-same-value — correct. Retention requires raw interface identity (_type + data, not Go equality, so uncomparable values still match) and p.recovered.frame == gp.recoverFrame != nil. Since p.recovered is armed only in Recover and zeroed in EndRecoverFrame*/abortPanics, a later panic with a coincidentally-identical value after the recovering activation has returned sees recovered.frame == nil → treated as new → fresh panicDepth. A different value during an active repanic window replaces the prefix. Both behaviors match Go's same-value repanic semantics.

PC sequence wrap / lookup — correct. The GLS nextCallerPCBase (caller_gls.go) draws a process-unique value from an atomic counter, so goroutine-local synthetic PCs no longer collide in the process-wide FuncForPC cache — the root cause of the second regression. seq << 2 with the seq == 0 || seq > ^uintptr(0)>>2 guard fatals on exhaustion instead of wrapping into a colliding PC (correct on 32-bit wasm uintptr). Because sequences are now sparse per store, syntheticFrameForPC correctly switches from the old dense seq-1 index to a binary search over PC &^ callerPCMask; the slice stays sorted since assignment is monotonic even under goroutine interleaving. capturePC/captureFrameAt consistently use the local frame.captured-1 index while encoding the global base from s.synthetic[idx].PC.

Bare-metal isolation — correct. caller_global.go provides a non-atomic single-store nextCallerPCBase (len+1), and panic_snapshot_default.go (//go:build !wasm) keeps capturePanicCallerFrames/clearPanicCallerSnapshot as no-ops with panicCallerSnapshotAvailable reading the native panicPCs.n. Native same-value repanic detection stays in SavePanicCallerFrames, unchanged. On WASM, capturePanicCallerFrames runs first in Panic and authoritatively consumes p.recovered for the shadow-stack mechanism the WASM traceback actually uses; the native snapshot path (only linked with unwind_llgo) then sees a zeroed recovered but that path is unused for WASM tracebacks, so there's no conflict.

Coverage. testCallerCacheIsolation asserts FuncForPC(Caller(0)).Name() resolves per-goroutine across two goroutines (exercises the collision fix); repanicTracebackCaller drives a nested same-value repanic and asserts the traceback origin is the original main.repanicTracebackOrigin site (exercises the prefix fix). Both wired across the three profiles via LLGO_WASM_SCHEDULER_REPANIC_TRACEBACK.

The reported +0.65% raw / +0.63% gzip incremental size is consistent with what the diff adds (one atomic global, a binary-search lookup, and small snapshot bookkeeping; the prefix is frozen in place with no per-panic copy).

LGTM.

@cpunion

cpunion commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai Please review the final test-only incremental range dff40ac32f8b..748483792502. Stacked W3 full-package acceptance exposed that TestCallerPanicTraceback and TestCallerRepanicTraceback used exec.Cmd.CombinedOutput inside a WASM guest, where the documented Go-compatible process contract returns ENOSYS for pipe. This change keeps both subprocess harnesses as !wasm wrappers; WASM runtime behavior remains exercised through the host-driven J32 Emscripten, J64 Emscripten Memory64, and W32 WASI scheduler fixtures without adding a skip. Please check the build-tag boundary and that native coverage remains intact.

@fennoai

fennoai Bot commented Sep 13, 2026

Copy link
Copy Markdown
Contributor

FennoAI is processing your request, please wait a moment...

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review: test-only incremental range dff40ac32f8b..748483792502 (1 commit, 2 files, +17/-2)

Reviewed 748483792 test(wasm): leave fatal subprocesses to the host. Clean, correct, minimal. No issues found — and this round the change is pure-Go test code with no LLVM dependency, so I was able to run the tests: both pass on the native target.

Build-tag boundary — correct. The two subprocess-driven entry points are renamed to unconstrained helpers testCallerPanicTraceback/testCallerRepanicTraceback in caller_runtime_test.go (no //go:build line, compiles on every target), and re-exported as TestCallerPanicTraceback/TestCallerRepanicTraceback from the new caller_panic_native_test.go gated //go:build !wasm. So on native the tests run exactly as before; on the WASM guest the TestXxx entry points simply don't exist — no t.Skip, no runtime-behavior change, matching the stated approach. Package name (gotest) matches and all imports (os, os/exec, …) remain used by the retained helpers.

Self-re-exec stays consistent. testCallerRepanicTraceback re-launches the test binary with -test.run=^TestCallerRepanicTraceback$ (caller_runtime_test.go:175) to drive its child panic modes. Since that whole path only executes on !wasm and the entry point exists there, the child filter resolves correctly in the same binary.

Native coverage intact. The helper bodies are unchanged (only the signatures were renamed), and go test ./test/go/ -run '^TestCallerPanicTraceback$|^TestCallerRepanicTraceback$' passes both. go vet ./test/go/ compiles the package with no new diagnostics on the two touched files (the pre-existing vet notes are in unrelated files). WASM traceback behavior remains exercised by the host-driven J32 Emscripten / J64 Memory64 / W32 WASI scheduler fixtures (the panic/repanic traceback fixtures added earlier in this stack), so gating the in-guest subprocess harnesses loses no WASM coverage — it removes an invalid pipe/ENOSYS call the Go-compatible WASM process contract can't satisfy.

LGTM. This completes my review of the stack.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant