Skip to content

fix(wasi): size initial memory from static data and heap budget - #2541

Merged
cpunion merged 4 commits into
xgo-dev:mainfrom
cpunion:codex/fix-wasi-initial-heap-20260909
Sep 22, 2026
Merged

cpunion merged 4 commits into
xgo-dev:mainfrom
cpunion:codex/fix-wasi-initial-heap-20260909

Conversation

@cpunion

@cpunion cpunion commented Sep 8, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Fix a main-branch WASI linker limit exposed by expanded WebAssembly acceptance: a legal program with 100 MiB of static data cannot link because the toolchain forces an exact 64 MiB initial memory size.

For single-worker WASI, reserve the existing 54 MiB heap budget after static data and the 10 MiB process stack with --initial-heap, allowing wasm-ld to size initial memory from the actual layout. Keep the WASI pthreads backend's explicit 64 MiB shared-memory contract unchanged. This does not add threads support or change maximum memory or stack sizes.

The final-link default is omitted when effective driver arguments already specify initial-memory, initial-heap, or a user-authored response file. Package/config/environment/linker-prefix options remain authoritative. clang.Cmd.LinkArguments is a small extraction of the existing argument merger, so inspection and execution use the same flags; no unrelated FuncForPC relinking code is included.

An explicitly empty initial heap is valid even when static data ends exactly at a Wasm page boundary. Initialize it by growing one page before GC metadata setup; normal heap growth then applies. If the host refuses growth, retain the explicit initialization failure. Baremetal is unchanged.

Validation and CI coverage

  • LLVM 22 focused build/crosscompile/clang tests pass. New/extracted functions have 100% targeted statement coverage.
  • 25 option-policy cases cover named/raw WASI, explicit user options and spellings, native/Emscripten exclusions, response files, and environment/command-prefix flags.
  • Eight lightweight linkObjFiles → linker subprocess regressions verify that defaults actually reach the production command and explicit/shared-memory flags are not lost. They run in the existing main Go coverage workflow, with no new CI job or SDK dependency.
  • The existing WASI runtime CI suite now exercises an object-only response file, a dynamically derived exact zero-heap boundary, and growth refusal (--max-memory), requiring exit 2 and the precise initialization diagnostic. All three pass locally. These links reuse package caches; changing heap flags did not rebuild the internal runtime/TinyGC archives.
  • On this main-based branch, Go 1.27's unmodified fixedbugs/issue34395.go builds and runs through raw GOOS=wasip1 GOARCH=wasm (wasm32 with WASI Preview 1 host, using that historical revision's data model) with clang/LLD 22 (7.77s for the case). This is local focused evidence, not a claim that the new PR's CI has finished.

Response-file policy

User-authored response files own the memory policy: do not inject the 54 MiB default alongside an opaque file that may specify its own exact initial memory or zero heap, including in nested files. Contents and quoting remain the external driver's responsibility. An object-only response file therefore starts with page-rounding slack (possibly zero) and grows memory when needed. LLGo's automatic response-file creation happens after inspection and does not opt out of the default.

The underlying failure exists on main and is not waived with a GOROOT exclusion or a larger timeout.

Rebase validation (2026-09-14)

Rebased the three WASI commits onto current main 5c5874359. Main now contains the descendant-probe corrections, so Git dropped both obsolete goroot dependency commits. The branch still carries the independently reviewed #2562 malformed-LTO-metadata fix as one dependency commit; its patch-id is identical to the rebased #2562 head and it can be dropped from this PR once #2562 lands on main.

The dev/test_wasm_single_worker.sh conflict was resolved by preserving main’s TinyGC host test inside the runtime/all suite before the target fixtures; the new gc-heap suite remains limited to the dedicated empty-heap checks. Focused WASI heap/default/linker-argument tests, the complete internal/clang and internal/crosscompile packages, and the nested TinyGC tests pass locally. Fresh CI is running on the rebased head.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: fix WASI initial heap

Solid, well-tested change. The switch from a fixed --initial-memory=67108864 cap for the single-worker WASI C profile to a --initial-heap=56623104 (54 MiB) default — appended only when the user has not already set --initial-memory/--initial-heap — is a sound approach: it stops static data from eating into the previous 64 MiB total while preserving the effective budget. The Link/LinkArguments/linkArguments split is clean and lets the default selection inspect the effective linker command (prefix + env/config flags + args) without executing it.

I verified the load-bearing details:

  • Heap math is correct and consistent. 56623104 = 54 MiB, and the wasip1/wasm link path always sets -z stack-size=10485760 (10 MiB) at internal/crosscompile/crosscompile.go:686, so the comment's "64 MiB total minus 10 MiB stack" (54 MiB heap) holds for every path where the default heap is injected.
  • wasi_threads keeps its --initial-memory=67108864 via config LDFLAGS, so defaultWASIHeapArgs correctly detects it and does not add the heap default. Covered by TestWASIHeapReachesLinker (shared-memory-contract case) and the crosscompile tests.
  • Detection covers -Wl,, comma-separated, -Xlinker, direct, CCFLAGS, and LDFLAGS forms, with good false-positive tests (data-initial-memory.o, -Wl,-Map,initial-heap.map).

One known limitation worth documenting (not blocking): the detection deliberately does not expand @response-file contents (asserted by the "response contents not expanded" test). If a user passes --initial-memory=<N> inside a response file, the default --initial-heap is still appended, which — per wasm-ld's initial-memory = stack + static + heap default — could make an explicit --initial-memory too small and fail the link. The code comment mentions response files are left to the driver; a one-line note that explicit memory/heap flags must be passed directly (not via @file) to be honored would help future users.

Minor inline notes below. No blocking issues.

Comment thread internal/build/build.go Outdated
Comment thread internal/build/wasm_memory.go Outdated
@codecov

codecov Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@cpunion

cpunion commented Sep 8, 2026

Copy link
Copy Markdown
Collaborator Author

CI triage: the new Windows x64 MinGW LLGo failure is a dependency-download transport error, not a heap-policy regression. The Windows runtime, stdlib, FFI and network smoke tests completed; the earlier nil-pointer panic is explicitly the expected unrecovered-fault smoke case. The next step failed while downloading github.com/goplus/lib@v0.3.1 from proxy.golang.org: HTTP/2 stream INTERNAL_ERROR from the peer, before test/buildcache could compile. See job https://github.com/xgo-dev/llgo/actions/runs/34255224036/job/102159518324 . No source change or blanket retry is being added for this failure. An upstream maintainer can rerun the failed job; upstream Actions have not been modified.

@github-actions

github-actions Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

ea838b5761dd | workflow run | long-term charts

WebAssembly output sizes

Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 135933 B +187 B / +0.1% (worse) 70786 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 134182 B +187 B / +0.1% (worse) 69150 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 124293 B +50 B / +0.04024% (worse) 73998 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 138465 B +204 B / +0.1% (worse) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 138469 B +204 B / +0.1% (worse) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3090694 B +189 B / +0.006116% (worse) 114540 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3082733 B +189 B / +0.006131% (worse) 98239 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2829095 B +51 B / +0.001803% (worse) 121521 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2828971 B +196 B / +0.006929% (worse) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2697526 B +196 B / +0.007266% (worse) 0 B 0 B / 0.0%
j32-emscripten/LLGo 135207 B +187 B / +0.1% (worse) 70786 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 133628 B +187 B / +0.1% (worse) 69150 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 123623 B +50 B / +0.04046% (worse) 73998 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1481443 B +189 B / +0.01276% (worse) 88949 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1483164 B +189 B / +0.01274% (worse) 87313 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1368332 B +51 B / +0.003727% (worse) 94158 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1534319 B +196 B / +0.01278% (worse) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1460960 B +196 B / +0.01342% (worse) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 137746 B +196 B / +0.1% (worse) 0 B 0 B / 0.0%
w32-wasi/LLGo 137817 B +196 B / +0.1% (worse) 0 B 0 B / 0.0%

LLGo WebAssembly build measurements

Example and profile Build vs base
j32-emscripten 5.787 s -67.55 ms / -1.2% (better)
j32-goos-js 5.858 s -219.2 ms / -3.6% (better)
j64-emscripten-memory64 5.046 s -53.5 ms / -1.0% (better)
reflectcall/w32-wasi 26.111 s +280.8 ms / +1.1% (worse)
w32-goos-wasip1 4.511 s -27.87 ms / -0.6% (better)
w32-wasi 4.473 s -170.7 ms / -3.7% (better)

Compared with a1a44a6e765c measured in the same runner job.

@cpunion
cpunion force-pushed the codex/fix-wasi-initial-heap-20260909 branch from 865c7c5 to 93a0584 Compare September 11, 2026 05:01
@github-actions

github-actions Bot commented Sep 11, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

ea838b5761dd | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 20040 B 0 B / +0.0% 387 B 0 B / +0.0% 479.750 ms +26.64 ms / +5.9% (worse) 1.243 ms +13.1 us / +1.1% (worse)
Linux cprintf-lto 19792 B 0 B / +0.0% 368 B 0 B / +0.0% 479.963 ms +22.49 ms / +4.9% (worse) 1.228 ms -225 ns / -0.01832% (better)
Linux fmtprintf 1640296 B 0 B / +0.0% 487292 B 0 B / +0.0% 3.815 s -112.5 ms / -2.9% (better) 2.962 ms -33.38 us / -1.1% (better)
Linux fmtprintf-lto 1478552 B 0 B / +0.0% 422855 B 0 B / +0.0% 10.534 s -156.5 ms / -1.5% (better) 3.038 ms +152.1 us / +5.3% (worse)
Linux println 62160 B 0 B / +0.0% 14647 B 0 B / +0.0% 465.747 ms +11.76 ms / +2.6% (worse) 1.556 ms +39.03 us / +2.6% (worse)
Linux println-lto 54248 B 0 B / +0.0% 12115 B 0 B / +0.0% 671.734 ms +14.67 ms / +2.2% (worse) 1.521 ms -47.88 us / -3.1% (better)
macOS cprintf 84480 B 0 B / +0.0% 17309 B 0 B / +0.0% 553.270 ms -57.03 ms / -9.3% (better) 2.841 ms +444.4 us / +18.5% (worse)
macOS cprintf-lto 84288 B 0 B / +0.0% 13073 B 0 B / +0.0% 644.601 ms +62.06 ms / +10.7% (worse) 2.421 ms +57.33 us / +2.4% (worse)
macOS fmtprintf 1483824 B 0 B / +0.0% 862012 B 0 B / +0.0% 2.801 s -275.4 ms / -9.0% (better) 3.808 ms -881 us / -18.8% (better)
macOS fmtprintf-lto 1175808 B 0 B / +0.0% 832972 B 0 B / +0.0% 5.491 s -1.998 s / -26.7% (better) 3.618 ms -37.42 us / -1.0% (better)
macOS println 114672 B 0 B / +0.0% 34882 B 0 B / +0.0% 512.718 ms -107.6 ms / -17.3% (better) 2.990 ms -1.706 ms / -36.3% (better)
macOS println-lto 118720 B 0 B / +0.0% 32320 B 0 B / +0.0% 647.473 ms -400.9 ms / -38.2% (better) 2.930 ms -2.03 ms / -40.9% (better)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.055 s -104.5 ms / -9.0% (better) 3.367 ms -278.8 us / -7.6% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.096 s -113.5 ms / -9.4% (better) 3.443 ms -38.5 us / -1.1% (better)
Windows MinGW fmtprintf 1913856 B 0 B / +0.0% 589718 B 0 B / +0.0% 4.033 s -157.4 ms / -3.8% (better) 7.728 ms -788 us / -9.3% (better)
Windows MinGW fmtprintf-lto 1934336 B 0 B / +0.0% 535846 B 0 B / +0.0% 9.753 s -371 ms / -3.7% (better) 7.369 ms -866.5 us / -10.5% (better)
Windows MinGW println 71168 B 0 B / +0.0% 23702 B 0 B / +0.0% 1.078 s -96.83 ms / -8.2% (better) 6.376 ms -661.3 us / -9.4% (better)
Windows MinGW println-lto 65024 B 0 B / +0.0% 20406 B 0 B / +0.0% 1.289 s -90.49 ms / -6.6% (better) 6.161 ms -1.013 ms / -14.1% (better)
Windows MinGW 386 cprintf 43008 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.097 s +13.32 ms / +1.2% (worse) 5.130 ms +32.1 us / +0.6% (worse)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.180 s +70.02 ms / +6.3% (worse) 5.366 ms +208.5 us / +4.0% (worse)
Windows MinGW 386 fmtprintf 1876992 B 0 B / +0.0% 463342 B 0 B / +0.0% 4.182 s +122.8 ms / +3.0% (worse) 10.759 ms +615 us / +6.1% (worse)
Windows MinGW 386 fmtprintf-lto 2152448 B 0 B / +0.0% 440730 B 0 B / +0.0% 10.712 s -104 ms / -1.0% (better) 12.185 ms +1.306 ms / +12.0% (worse)
Windows MinGW 386 println 91136 B 0 B / +0.0% 19814 B 0 B / +0.0% 1.120 s +30.37 ms / +2.8% (worse) 9.474 ms +671.6 us / +7.6% (worse)
Windows MinGW 386 println-lto 69120 B 0 B / +0.0% 17742 B 0 B / +0.0% 1.295 s +18.17 ms / +1.4% (worse) 8.780 ms +177.3 us / +2.1% (worse)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.373 s -31.38 ms / -2.2% (better) 6.095 ms -35.6 us / -0.6% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.405 s -30.66 ms / -2.1% (better) 5.940 ms -230.1 us / -3.7% (better)
Windows MinGW ARM64 fmtprintf 1801728 B 0 B / +0.0% 502004 B 0 B / +0.0% 4.206 s -16.78 ms / -0.4% (better) 11.751 ms -1.553 ms / -11.7% (better)
Windows MinGW ARM64 fmtprintf-lto 1858560 B 0 B / +0.0% 466144 B 0 B / +0.0% 9.358 s -174.2 ms / -1.8% (better) 12.310 ms -774.1 us / -5.9% (better)
Windows MinGW ARM64 println 68096 B 0 B / +0.0% 22376 B 0 B / +0.0% 1.421 s +27.43 ms / +2.0% (worse) 10.224 ms -405.1 us / -3.8% (better)
Windows MinGW ARM64 println-lto 63488 B 0 B / +0.0% 19444 B 0 B / +0.0% 1.538 s -69.67 ms / -4.3% (better) 10.236 ms -73.8 us / -0.7% (better)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65782 B 0 B / +0.0% 972.551 ms +22.88 ms / +2.4% (worse) 3.458 ms +81.6 us / +2.4% (worse)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65718 B 0 B / +0.0% 988.769 ms -21.79 ms / -2.2% (better) 3.468 ms -1.089 ms / -23.9% (better)
Windows MSVC fmtprintf 1628160 B 0 B / +0.0% 685238 B 0 B / +0.0% 3.892 s +109.3 ms / +2.9% (worse) 10.235 ms +995.6 us / +10.8% (worse)
Windows MSVC fmtprintf-lto 1616896 B 0 B / +0.0% 636262 B 0 B / +0.0% 9.202 s -18.78 ms / -0.2% (better) 9.716 ms -5 us / -0.1% (better)
Windows MSVC println 192512 B 0 B / +0.0% 119126 B 0 B / +0.0% 954.902 ms -126.5 ms / -11.7% (better) 7.665 ms -757.8 us / -9.0% (better)
Windows MSVC println-lto 189952 B 0 B / +0.0% 116374 B 0 B / +0.0% 1.331 s +181.1 ms / +15.8% (worse) 7.038 ms -1.314 ms / -15.7% (better)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.006 s -3.042 ms / -0.3% (better) 5.376 ms -367.1 us / -6.4% (better)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.034 s +11.8 ms / +1.2% (worse) 8.388 ms +2.625 ms / +45.5% (worse)
Windows MSVC 386 fmtprintf 1190400 B 0 B / +0.0% 446693 B 0 B / +0.0% 4.124 s +130.1 ms / +3.3% (worse) 13.047 ms +1.338 ms / +11.4% (worse)
Windows MSVC 386 fmtprintf-lto 1224704 B 0 B / +0.0% 417385 B 0 B / +0.0% 9.209 s -9.956 ms / -0.1% (better) 12.632 ms -713.5 us / -5.3% (better)
Windows MSVC 386 println 34304 B 0 B / +0.0% 18641 B 0 B / +0.0% 1.030 s +31.69 ms / +3.2% (worse) 10.706 ms +420.7 us / +4.1% (worse)
Windows MSVC 386 println-lto 32256 B 0 B / +0.0% 16855 B 0 B / +0.0% 1.187 s -146.6 ms / -11.0% (better) 10.626 ms +709 us / +7.1% (worse)
Windows MSVC ARM64 cprintf 11264 B 0 B / +0.0% 3976 B 0 B / +0.0% 2.045 s -24.32 ms / -1.2% (better) 6.956 ms -252.8 us / -3.5% (better)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 3868 B 0 B / +0.0% 2.072 s -18.11 ms / -0.9% (better) 7.189 ms +349.6 us / +5.1% (worse)
Windows MSVC ARM64 fmtprintf 1375232 B 0 B / +0.0% 502948 B 0 B / +0.0% 6.723 s -122.3 ms / -1.8% (better) 14.987 ms -217.5 us / -1.4% (better)
Windows MSVC ARM64 fmtprintf-lto 1391616 B 0 B / +0.0% 469044 B 0 B / +0.0% 15.528 s +190.3 ms / +1.2% (worse) 15.722 ms +129.5 us / +0.8% (worse)
Windows MSVC ARM64 println 41472 B 0 B / +0.0% 21624 B 0 B / +0.0% 2.031 s -23.07 ms / -1.1% (better) 12.379 ms -770.7 us / -5.9% (better)
Windows MSVC ARM64 println-lto 39424 B 0 B / +0.0% 19372 B 0 B / +0.0% 2.334 s -70.24 ms / -2.9% (better) 11.988 ms -2.363 ms / -16.5% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.620 ns/op +0.04 ns/op / +0.3% (worse)
Linux BenchmarkMergeCompilerFlags 208.900 ns/op +13.7 ns/op / +7.0% (worse)
Linux BenchmarkMergeLinkerFlags 132.500 ns/op -6.7 ns/op / -4.8% (better)
Linux BenchmarkChannelBuffered 55.020 ns/op -0.1 ns/op / -0.2% (better)
Linux BenchmarkChannelHandoff 13006 ns/op -2226 ns/op / -14.6% (better)
Linux BenchmarkDefer 53.060 ns/op +1.37 ns/op / +2.7% (worse)
Linux BenchmarkDirectCall 1.164 ns/op -0.014 ns/op / -1.2% (better)
Linux BenchmarkGlobalRead 1.164 ns/op -0.005 ns/op / -0.4% (better)
Linux BenchmarkGlobalWrite 7.761 ns/op -0.002 ns/op / -0.02576% (better)
Linux BenchmarkGoroutine 21088 ns/op +258 ns/op / +1.2% (worse)
Linux BenchmarkInterfaceCall 5.820 ns/op -0.006 ns/op / -0.1% (better)
Linux BenchmarkRuntimeGetG 2.443 ns/op -0.001 ns/op / -0.04092% (better)
macOS BenchmarkLookupPCRandom 10.740 ns/op -3.16 ns/op / -22.7% (better)
macOS BenchmarkMergeCompilerFlags 95.380 ns/op -1.5 ns/op / -1.5% (better)
macOS BenchmarkMergeLinkerFlags 54.600 ns/op -9.31 ns/op / -14.6% (better)
macOS BenchmarkChannelBuffered 22.600 ns/op -3.46 ns/op / -13.3% (better)
macOS BenchmarkChannelHandoff 6506 ns/op -2497 ns/op / -27.7% (better)
macOS BenchmarkDefer 28.520 ns/op -3.77 ns/op / -11.7% (better)
macOS BenchmarkDirectCall 0.974 ns/op -0.2068 ns/op / -17.5% (better)
macOS BenchmarkGlobalRead 0.942 ns/op -0.1542 ns/op / -14.1% (better)
macOS BenchmarkGlobalWrite 0.941 ns/op -0.2409 ns/op / -20.4% (better)
macOS BenchmarkGoroutine 26117 ns/op -15696 ns/op / -37.5% (better)
macOS BenchmarkInterfaceCall 3.621 ns/op -0.24 ns/op / -6.2% (better)
macOS BenchmarkRuntimeGetG 1.973 ns/op -0.178 ns/op / -8.3% (better)
Windows MinGW BenchmarkLookupPCRandom 13.230 ns/op +0.09 ns/op / +0.7% (worse)
Windows MinGW BenchmarkMergeCompilerFlags 601.900 ns/op -6.6 ns/op / -1.1% (better)
Windows MinGW BenchmarkMergeLinkerFlags 535.400 ns/op +5.2 ns/op / +1.0% (worse)
Windows MinGW BenchmarkChannelBuffered 31.250 ns/op +0.01 ns/op / +0.03201% (worse)
Windows MinGW BenchmarkChannelHandoff 976.800 ns/op -13.1 ns/op / -1.3% (better)
Windows MinGW BenchmarkDefer 55.330 ns/op -1.27 ns/op / -2.2% (better)
Windows MinGW BenchmarkDirectCall 1.547 ns/op -0.005 ns/op / -0.3% (better)
Windows MinGW BenchmarkGlobalRead 1.547 ns/op -0.003 ns/op / -0.2% (better)
Windows MinGW BenchmarkGlobalWrite 2.470 ns/op -0.003 ns/op / -0.1% (better)
Windows MinGW BenchmarkGoroutine 81173 ns/op +4379 ns/op / +5.7% (worse)
Windows MinGW BenchmarkInterfaceCall 8.392 ns/op +0.025 ns/op / +0.3% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.167 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.670 ns/op +0.11 ns/op / +0.4% (worse)
Windows MinGW 386 BenchmarkMergeCompilerFlags 785.400 ns/op +18 ns/op / +2.3% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 721.400 ns/op +7.5 ns/op / +1.1% (worse)
Windows MinGW 386 BenchmarkChannelBuffered 39.450 ns/op +0.14 ns/op / +0.4% (worse)
Windows MinGW 386 BenchmarkChannelHandoff 867.200 ns/op -82.4 ns/op / -8.7% (better)
Windows MinGW 386 BenchmarkDefer 44.990 ns/op +0.3 ns/op / +0.7% (worse)
Windows MinGW 386 BenchmarkDirectCall 1.551 ns/op 0 ns/op / +0.0%
Windows MinGW 386 BenchmarkGlobalRead 1.551 ns/op 0 ns/op / +0.0%
Windows MinGW 386 BenchmarkGlobalWrite 7.783 ns/op +0.01 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGoroutine 86284 ns/op -1816 ns/op / -2.1% (better)
Windows MinGW 386 BenchmarkInterfaceCall 8.368 ns/op -0.017 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 2.478 ns/op +0.003 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.160 ns/op +0.11 ns/op / +0.9% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 631.900 ns/op +50.1 ns/op / +8.6% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 589.900 ns/op +60.6 ns/op / +11.4% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 39.180 ns/op +1.38 ns/op / +3.7% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 2021 ns/op -324 ns/op / -13.8% (better)
Windows MinGW ARM64 BenchmarkDefer 52.160 ns/op +0.4 ns/op / +0.8% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.589 ns/op -0.0003 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.590 ns/op -0.0001 ns/op / -0.01695% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.590 ns/op +0.0001 ns/op / +0.01696% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 60321 ns/op +921 ns/op / +1.6% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.239 ns/op +0.002 ns/op / +0.0472% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.797 ns/op +0.027 ns/op / +1.5% (worse)
Windows MSVC BenchmarkLookupPCRandom 13.080 ns/op +0.02 ns/op / +0.2% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 615 ns/op -9.2 ns/op / -1.5% (better)
Windows MSVC BenchmarkMergeLinkerFlags 528.600 ns/op -2.8 ns/op / -0.5% (better)
Windows MSVC BenchmarkChannelBuffered 29.120 ns/op -0.02 ns/op / -0.1% (better)
Windows MSVC BenchmarkChannelHandoff 1079 ns/op -49 ns/op / -4.3% (better)
Windows MSVC BenchmarkDefer 56.080 ns/op +0.28 ns/op / +0.5% (worse)
Windows MSVC BenchmarkDirectCall 1.549 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalRead 1.549 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC BenchmarkGlobalWrite 2.460 ns/op +0.005 ns/op / +0.2% (worse)
Windows MSVC BenchmarkGoroutine 85602 ns/op +434 ns/op / +0.5% (worse)
Windows MSVC BenchmarkInterfaceCall 9.282 ns/op -0.016 ns/op / -0.2% (better)
Windows MSVC BenchmarkRuntimeGetG 2.172 ns/op +0.004 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 26.650 ns/op +0.04 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 751 ns/op +11.5 ns/op / +1.6% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 700.400 ns/op +14.3 ns/op / +2.1% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 38.310 ns/op -0.01 ns/op / -0.0261% (better)
Windows MSVC 386 BenchmarkChannelHandoff 919.600 ns/op +46.7 ns/op / +5.3% (worse)
Windows MSVC 386 BenchmarkDefer 47.150 ns/op -1.95 ns/op / -4.0% (better)
Windows MSVC 386 BenchmarkDirectCall 1.550 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.858 ns/op -0.004 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkGlobalWrite 7.772 ns/op -0.008 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGoroutine 85150 ns/op -2684 ns/op / -3.1% (better)
Windows MSVC 386 BenchmarkInterfaceCall 8.395 ns/op +0.012 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 2.171 ns/op +0.001 ns/op / +0.04608% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.040 ns/op -0.03 ns/op / -0.2% (better)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 573 ns/op +19 ns/op / +3.4% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 568 ns/op +47 ns/op / +9.0% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 39.610 ns/op +1.9 ns/op / +5.0% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 2426 ns/op -34 ns/op / -1.4% (better)
Windows MSVC ARM64 BenchmarkDefer 65.240 ns/op +2.13 ns/op / +3.4% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.663 ns/op +0.0002 ns/op / +0.03015% (worse)
Windows MSVC ARM64 BenchmarkGlobalRead 0.663 ns/op -0.0002 ns/op / -0.03013% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.795 ns/op +0.001 ns/op / +0.02636% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 53816 ns/op -1222 ns/op / -2.2% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.139 ns/op +0.002 ns/op / +0.04834% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 2.251 ns/op +0.003 ns/op / +0.1% (worse)

Timer runtime benchmarks

Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 906.900 ns/op -14.2 ns/op / -1.5% (better)
Linux AfterFuncZeroDelivery/LLGo 31268 ns/op -4884 ns/op / -13.5% (better)
Linux CreateStop/Go 294.300 ns/op +4.9 ns/op / +1.7% (worse)
Linux CreateStop/LLGo 1687 ns/op +91 ns/op / +5.7% (worse)
Linux RearmStopped/Go 114.500 ns/op -0.5 ns/op / -0.4% (better)
Linux RearmStopped/LLGo 1279 ns/op -58 ns/op / -4.3% (better)
Linux ResetActive/Go 67.370 ns/op -0.12 ns/op / -0.2% (better)
Linux ResetActive/LLGo 788.500 ns/op -27 ns/op / -3.3% (better)
Linux ResetHeap1024/Go 67.070 ns/op -0.22 ns/op / -0.3% (better)
Linux ResetHeap1024/LLGo 176.100 ns/op +0.6 ns/op / +0.3% (worse)
macOS AfterFuncZeroDelivery/Go 372.700 ns/op -55.2 ns/op / -12.9% (better)
macOS AfterFuncZeroDelivery/LLGo 58585 ns/op -8185 ns/op / -12.3% (better)
macOS CreateStop/Go 115.200 ns/op -43.1 ns/op / -27.2% (better)
macOS CreateStop/LLGo 443.500 ns/op -9.3 ns/op / -2.1% (better)
macOS RearmStopped/Go 50.940 ns/op -7.97 ns/op / -13.5% (better)
macOS RearmStopped/LLGo 320.300 ns/op +130 ns/op / +68.3% (worse)
macOS ResetActive/Go 37.510 ns/op -6.95 ns/op / -15.6% (better)
macOS ResetActive/LLGo 153.100 ns/op +49.5 ns/op / +47.8% (worse)
macOS ResetHeap1024/Go 40.710 ns/op -3.39 ns/op / -7.7% (better)
macOS ResetHeap1024/LLGo 78.100 ns/op -16.24 ns/op / -17.2% (better)
Windows MinGW AfterFuncZeroDelivery/Go 564.600 ns/op +10.8 ns/op / +2.0% (worse)
Windows MinGW AfterFuncZeroDelivery/LLGo 160949 ns/op -204 ns/op / -0.1% (better)
Windows MinGW CreateStop/Go 113.700 ns/op -2.9 ns/op / -2.5% (better)
Windows MinGW CreateStop/LLGo 447.900 ns/op -31.4 ns/op / -6.6% (better)
Windows MinGW RearmStopped/Go 31.270 ns/op -0.09 ns/op / -0.3% (better)
Windows MinGW RearmStopped/LLGo 275.700 ns/op +4 ns/op / +1.5% (worse)
Windows MinGW ResetActive/Go 20.050 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW ResetActive/LLGo 179.700 ns/op +0.4 ns/op / +0.2% (worse)
Windows MinGW ResetHeap1024/Go 20.330 ns/op -0.19 ns/op / -0.9% (better)
Windows MinGW ResetHeap1024/LLGo 127.500 ns/op -0.9 ns/op / -0.7% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 970.300 ns/op +13.6 ns/op / +1.4% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 187479 ns/op +6937 ns/op / +3.8% (worse)
Windows MinGW 386 CreateStop/Go 198.300 ns/op +4.3 ns/op / +2.2% (worse)
Windows MinGW 386 CreateStop/LLGo 535.800 ns/op +9.3 ns/op / +1.8% (worse)
Windows MinGW 386 RearmStopped/Go 64.810 ns/op +1.47 ns/op / +2.3% (worse)
Windows MinGW 386 RearmStopped/LLGo 348.900 ns/op +4.3 ns/op / +1.2% (worse)
Windows MinGW 386 ResetActive/Go 39.170 ns/op +0.22 ns/op / +0.6% (worse)
Windows MinGW 386 ResetActive/LLGo 964.400 ns/op -47.6 ns/op / -4.7% (better)
Windows MinGW 386 ResetHeap1024/Go 39.510 ns/op -0.06 ns/op / -0.2% (better)
Windows MinGW 386 ResetHeap1024/LLGo 185.700 ns/op -0.8 ns/op / -0.4% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 667 ns/op -4.2 ns/op / -0.6% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 134780 ns/op -5191 ns/op / -3.7% (better)
Windows MinGW ARM64 CreateStop/Go 198.700 ns/op -0.3 ns/op / -0.2% (better)
Windows MinGW ARM64 CreateStop/LLGo 355.400 ns/op -14.9 ns/op / -4.0% (better)
Windows MinGW ARM64 RearmStopped/Go 70.530 ns/op +0.04 ns/op / +0.1% (worse)
Windows MinGW ARM64 RearmStopped/LLGo 250.900 ns/op -1.3 ns/op / -0.5% (better)
Windows MinGW ARM64 ResetActive/Go 30.900 ns/op -0.13 ns/op / -0.4% (better)
Windows MinGW ARM64 ResetActive/LLGo 125.700 ns/op +2.5 ns/op / +2.0% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 30.980 ns/op -0.06 ns/op / -0.2% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 126.200 ns/op -0.4 ns/op / -0.3% (better)
Windows MSVC AfterFuncZeroDelivery/Go 552.500 ns/op -2.1 ns/op / -0.4% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 169471 ns/op +9571 ns/op / +6.0% (worse)
Windows MSVC CreateStop/Go 115.900 ns/op -1.4 ns/op / -1.2% (better)
Windows MSVC CreateStop/LLGo 435.700 ns/op -43.2 ns/op / -9.0% (better)
Windows MSVC RearmStopped/Go 31.400 ns/op -0.32 ns/op / -1.0% (better)
Windows MSVC RearmStopped/LLGo 273.400 ns/op -7.4 ns/op / -2.6% (better)
Windows MSVC ResetActive/Go 20.020 ns/op +0.01 ns/op / +0.04998% (worse)
Windows MSVC ResetActive/LLGo 160.600 ns/op +14.8 ns/op / +10.2% (worse)
Windows MSVC ResetHeap1024/Go 20.520 ns/op 0 ns/op / +0.0%
Windows MSVC ResetHeap1024/LLGo 125.600 ns/op -1.8 ns/op / -1.4% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 980.100 ns/op +15.9 ns/op / +1.6% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 179787 ns/op -6363 ns/op / -3.4% (better)
Windows MSVC 386 CreateStop/Go 195.600 ns/op -2.9 ns/op / -1.5% (better)
Windows MSVC 386 CreateStop/LLGo 465.800 ns/op -14.4 ns/op / -3.0% (better)
Windows MSVC 386 RearmStopped/Go 63.770 ns/op +0.19 ns/op / +0.3% (worse)
Windows MSVC 386 RearmStopped/LLGo 333 ns/op +14.3 ns/op / +4.5% (worse)
Windows MSVC 386 ResetActive/Go 39.030 ns/op -0.08 ns/op / -0.2% (better)
Windows MSVC 386 ResetActive/LLGo 916.200 ns/op +5 ns/op / +0.5% (worse)
Windows MSVC 386 ResetHeap1024/Go 39.410 ns/op +0.1 ns/op / +0.3% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 171.100 ns/op -1.9 ns/op / -1.1% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 651.100 ns/op -7.9 ns/op / -1.2% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 124438 ns/op -2993 ns/op / -2.3% (better)
Windows MSVC ARM64 CreateStop/Go 206 ns/op +9.8 ns/op / +5.0% (worse)
Windows MSVC ARM64 CreateStop/LLGo 390.500 ns/op +4.6 ns/op / +1.2% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.660 ns/op +0.1 ns/op / +0.1% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 274.500 ns/op +0.4 ns/op / +0.1% (worse)
Windows MSVC ARM64 ResetActive/Go 31.020 ns/op -0.12 ns/op / -0.4% (better)
Windows MSVC ARM64 ResetActive/LLGo 138.200 ns/op +5.9 ns/op / +4.5% (worse)
Windows MSVC ARM64 ResetHeap1024/Go 31.130 ns/op +0.05 ns/op / +0.2% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 136.500 ns/op -0.7 ns/op / -0.5% (better)

Compared with a1a44a6e765c measured in the same runner job.

@cpunion
cpunion force-pushed the codex/fix-wasi-initial-heap-20260909 branch 2 times, most recently from c4d3700 to a7997fc Compare September 14, 2026 05:32
@cpunion

cpunion commented Sep 14, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai Review again after rebasing onto the latest main.

@fennoai

fennoai Bot commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

FennoAI is processing your request, please wait a moment...

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review (post-rebase)

Re-reviewed at head a7997fc on the fixed range 5c58743...a7997fc. The core WASI initial-memory fix is sound and the change is well-reasoned, with unusually thorough test coverage (nested response files, -Xlinker split options, CCFLAGS/LDFLAGS parsing, and the zero-heap boundary are all exercised).

Verified as correct (no action needed):

  • The 56623104 heap constant is exactly 54 MiB (64 − 10 MiB stack); with --initial-heap and no explicit --initial-memory, wasm-ld sizes initial memory as stack + static data + heap, matching the "64 MiB plus static data" comment.
  • The gcWasmGrowMemory return-value polarity is consistent across both call sites: _wrap/gc_wasm.c returns 1 on success / 0 on failure, so the bootstrap branch (!= 0 → success) and gcGrowMemory (== 0 → failure) both read it correctly. The grow-failure path is intentionally caught downstream by initGC's "gc: invalid heap range" panic (asserted by the shell test).
  • Response-file contents are deliberately not parsed; any leading @ token cedes memory policy to the user, avoiding a nested-response-file trust-boundary problem. Overflow guards in gcGrowMemory are sound. LinkArguments inspection performs no file I/O for WASI targets.

Findings below are minor quality/maintainability items — none block merge.

Additional findings

  • cl/compile_test.go:40: [P3] Unrelated LTO-plugin change bundled into a WASI-memory PR: This context-timeout + exit-code-1 assertion (paired with the report_fatal_error → reportFatalUsageError swap in LLGOInterfaceMethodTypeIDPass.cpp) is good hardening, but it is unrelated to WASI initial-memory sizing. Bundling an independent LTO-plugin behavior change hurts reviewability and bisection. Consider splitting it out, or (as the PR description does note the dependency corrections) calling it out explicitly as a separate concern.

Comment thread internal/build/wasm_memory.go
Comment thread dev/test_wasm_single_worker.sh
@cpunion

cpunion commented Sep 14, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai Review again after the two latest review-comment updates.

@fennoai

fennoai Bot commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

FennoAI is processing your request, please wait a moment...

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review (head 63b2d5c)

Re-reviewed after the two review-comment updates. The delta since the last reviewed head (a7997fc) is exactly two cosmetic changes on the files I commented on — both directly address the prior findings:

  • ✅ wasm_memory.go:30 — Added the cross-reference // The 10 MiB value matches -z stack-size=10485760 in crosscompile.go., making the cross-package coupling of the 54 MiB heap default to the 10 MiB stack size discoverable. (Comment-only; the 56623104 = 54 MiB value is unchanged and still correct.)
  • ✅ test_wasm_single_worker.sh — Extracted a named wasm_page_size=65536 local and used it in the boundary check, replacing the bare magic number.

Both changes are documentation/refactor only, with no behavior change, so the core WASI initial-memory fix verified at the prior head still holds. No new issues.

Remaining (optional, non-blocking, unchanged from prior review): the LTO-plugin hardening in cl/compile_test.go + LLGOInterfaceMethodTypeIDPass.cpp is still bundled into this WASI-memory PR. Splitting it out (or calling it out explicitly) would improve reviewability/bisection, but it is a reasonable dependency correction to carry if that's intentional.

LGTM.

@cpunion

cpunion commented Sep 14, 2026

Copy link
Copy Markdown
Collaborator Author

Final validation on 63b2d5c: all 64 checks passed, no pending or failed checks, and Codecov confirms all modified coverable lines are covered. Both wasm-runtime jobs passed, as did Linux/macOS/Windows coverage, the full LLGo matrix, benchmarks, and release-artifact execution tests. LLGo workflow: https://github.com/xgo-dev/llgo/actions/runs/34810489181 . Release workflow: https://github.com/xgo-dev/llgo/actions/runs/34810489059 . The branch is based on main 5c58743; the requested stack-size cross-reference and named Wasm page-size constant are implemented, the latest fennoai review covers the current head, and all review threads are answered and resolved. No extra coverage-only changes or workflow relaxation were necessary.

@cpunion
cpunion force-pushed the codex/fix-wasi-initial-heap-20260909 branch from 63b2d5c to ea838b5 Compare September 18, 2026 01:18
@cpunion
cpunion merged commit 71cd3ba into xgo-dev:main Sep 22, 2026
71 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants