Skip to content

[WASM W1] Define supported profiles and ABI foundation - #2578

Merged
xushiwei merged 17 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-w1-profiles-20260910
Sep 16, 2026
Merged

xushiwei merged 17 commits into
xgo-dev:mainfrom
cpunion:codex/wasm-w1-profiles-20260910

Conversation

@cpunion

@cpunion cpunion commented Sep 13, 2026 •

Copy link
Copy Markdown
Collaborator

Tracks #2152.

R1-R3 are merged. This PR implements W1: the supported hosted-WebAssembly profiles and ABI foundation. Later W2 work owns host integration, reflection, and standard-library completeness; W3 owns broad test/** and GOROOT acceptance.

Scope

  • Define J32 (wasm32 with JavaScript host), J64 (wasm64 with JavaScript host), and W32 (wasm32 with WASI Preview 1 host), with profile and provider identity included in package-cache inputs.
  • Keep stable target identifiers and provider build tags. wasm.json and wasip1.json are inheritance aliases of emscripten and wasi; -target wasm and -target wasip1 remain supported aliases. Remove unused tinygo.wasm and unsupported wasm-unknown / wasip2 target definitions.
  • Implement the official Go 64-bit int, uint, and uintptr storage model over Memory32 while retaining physical i32 addresses and ILP32 C/host layouts.
  • Add checked Go/C/host width adaptation for loads, stores, calls, multi-value results, aggregates, interfaces, maps, atomics, static relocations, GC shadow roots, scheduler context, and WASI Asyncify state.
  • Lower Memory32 pointer indexes to the physical i32 width before O0 WebAssembly code generation, without changing native, embedded, or already-physical C indexes.
  • Keep the existing Emscripten libffi backend correct under Go64 storage by retaining C-layout ffi_type / ffi_cif metadata and adapting pointer arrays to physical wasm32 slots.
  • Make raw GOOS=js GOARCH=wasm and GOOS=wasip1 GOARCH=wasm select the revised profile/provider metadata and complete single-worker runtime rather than the former compatibility runtime.
  • Preserve native min/max lowering after replacing the former target ABI switch with the resolved profile model.

The raw JavaScript entry in W1 is still Emscripten-backed. A Go-compatible JavaScript host boundary and portable WASI reflection are later work. W64, WASI Preview 2/WIT, threads, WasmGC, EH, JSPI/stack switching, multi-worker scheduling, and parallel goroutines are outside W1.

Validation

Current head cb7214ae45fa is rebased onto main 2db247e43848. The original 16 feature commits remain patch-equivalent; one integration fix updates the newly merged llgo env command to report WasmProfile and WasmProvider instead of the removed Config.WasmABI. The stale field caused the shared compiler-build failure in the previous stacked W2-A CI round. Regression tests cover all named profiles, both target aliases, and a non-WASM target; the modified target-field output function has 100% statement coverage. LLVM 22 compiler construction, complete cmd/... tests, and target resolver tests pass locally. Patch coverage is now 96.47% against a 94.38% target. Remaining full CI and the size/performance audit are still required; a passing benchmark job alone is not approval of a size increase.

The earlier native macOS test/go signal termination on 6b7350457dc8 remains unproven: paired fork runs passed on both W1 and its exact base without retries or sampling. This rebase fixes the separate deterministic llgo env build error, not that intermittent runtime termination. Earlier feature validation and benchmark results below are historical rather than current-head evidence.

Post-rebase local checks pass for internal/targets, the complete internal/crosscompile suite, focused W1 SSA/compiler regressions, git diff --check, and a clean worktree. Review-fix checks additionally cover CIF return/argument ownership, aggregate ffi type ownership, instantiation failure propagation, native-storage caching/target gating, build tags/output selection, and real J32, J64, and W32 builds. J32/Emscripten and W32/WASI lifecycle fixtures both pass after the final wasm32 ownership fix.

The equivalent pre-rebase feature range passed focused fork CI:

This upstream PR refreshes CI against the current main rather than treating the pre-rebase runs as final-head evidence.

Size and build audit

Current-head paired benchmark, comparing main 2db247e43848 with W1 cb7214ae45fa on the same runner (WASM module bytes):

Provider cprintf println fmtprintf
J32/Emscripten (wasm32 with the Emscripten JavaScript host) 113,637 → 130,908 112,905 → 130,034 2,753,958 → 3,229,202
J64/Emscripten (wasm64 with the Emscripten JavaScript host) 119,433 → 119,514 118,667 → 118,745 2,928,759 → 2,930,467
W32/WASI (wasm32 with WASI Preview 1) 117,681 → 133,351 117,007 → 132,583 2,511,028 → 2,912,768

The goal remains no cprintf growth where possible and only small justified println growth. The Memory32 figures above do not yet meet that goal. Diagnostics show growth distributed across Go64 runtime operations and checked Go64/C32 boundaries, rather than a bulk set of WASI reflection bridges; this identifies optimization targets, not proof that every added byte is unavoidable. W2-B separately reduces caller-wrapper overhead, and the final complete-stack comparison remains required.

Earlier W1 same-run benchmark output sizes (base → W1):

Entry cprintf println fmtprintf
J32/Emscripten (wasm32 with Emscripten JavaScript host) 114,075 → 130,933 (+14.8%) 113,339 → 130,079 (+14.8%) 2,971,098 → 3,505,655 (+18.0%)
J64/Emscripten (wasm64 with Emscripten Memory64) 119,587 → 119,666 (+0.07%) 118,820 → 118,900 (+0.07%) 3,160,277 → 3,162,091 (+0.06%)
W32/WASI (wasm32 with WASI Preview 1 host) 118,135 → 133,533 (+13.0%) 117,459 → 132,764 (+13.0%) 2,741,463 → 3,200,444 (+16.7%)

The explicit Memory32 increases are the measured cost of switching Go storage from the former 32-bit model to the specified 64-bit Go model; J64 remains effectively flat. Raw js/wasm and wasip1/wasm aliases have larger small-program deltas because main routed them through incomplete compatibility runtimes, while W1 makes them aliases of the complete J32/W32 implementations; those are not like-for-like regressions. The final ffi ownership review fix leaves J32/W32 cprintf and println byte-for-byte unchanged from the prior head; its largest additional change is 754 bytes in J64 fmtprintf (about 0.02%).

Diff audit

The range consists of proposal/profile metadata, target and cache selection, Go64-on-Memory32 lowering, runtime/GC storage adaptation, executable acceptance, and benchmark/coverage support needed to verify them. It contains no diagnostic probes, temporary dependency overrides, or new t.Skip. The audit removed dead provider state and redundant helpers, restored alias regressions, and avoided an unnecessary itab traversal that changed native output.

@codecov

codecov Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: WASM target profiles + storage lowering

Reviewed the full diff (base 52ae044...head 5a43f6a) across code quality, performance, security, and documentation. This is a solid, well-tested change — the logical-value vs. physical-storage abstraction is coherent, the //go:uintptrescapes GC-root derivation correctly declines to chase integer conversions, and target/profile validation is gated by a closed allow-list. Findings below are inline; a few non-inline notes follow.

Non-inline notes

  • Perf (hot compiler paths, all targets): the new storage helpers add uncached work to per-instruction type queries. GoWordSize() (ssa/storage.go:147) recomputes effectiveGOARCH() + string compares on every call (unlike the cached PointerSize()); childStorageType (ssa/storage.go:258) routes every Elem/Index/Field through hasNativeTypeLayout, which allocates a pkg.Name string and takes an RWMutex RLock per call; and withNativeStorage (ssa/storage.go:171) clones + inserts into an unbounded map keyed by fresh pointer identity, so it never dedupes and the map grows for the whole compilation. Since these fire on native (non-WASM) builds too, consider a cached "wide/native storage active" boolean on aProgram (set once in NewProgram) and an early return before the layout lookups.
  • Stale references (outside this diff): internal/build/outputs.go:325 and outputs_test.go:43 still list "wasip2"/"wasm-unknown" in defaultAppExt, but this PR deletes those target JSONs and the proposal states they're removed. Dead entries now — worth a follow-up cleanup.
  • Info (runtime): the wasm32-j32 GC now scans at a 4-byte stride (scanword_wasm32.go), doubling scan-loop iterations vs. the 8-byte word on that profile. Appears required for correctness (C/host pointers at 4-byte offsets) — flagging as a conscious trade-off, not a defect.

Comment thread runtime/internal/ffi/pointer_array_wasm32.go
Comment thread ssa/type.go Outdated
Comment thread targets/emscripten-runner.mjs Outdated
Comment thread internal/build/build.go Outdated
@github-actions

github-actions Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

cb7214ae45fa | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 19832 B 0 B / +0.0% 387 B 0 B / +0.0% 739.492 ms +134.1 ms / +22.1% (worse) 1.471 ms +56.11 us / +4.0% (worse)
Linux cprintf-lto 19584 B 0 B / +0.0% 368 B 0 B / +0.0% 725.802 ms +79.39 ms / +12.3% (worse) 1.520 ms +34.57 us / +2.3% (worse)
Linux fmtprintf 1636256 B +672 B / +0.04109% (worse) 486227 B +203 B / +0.04177% (worse) 5.561 s -104.6 ms / -1.8% (better) 3.863 ms -5.258 ms / -57.6% (better)
Linux fmtprintf-lto 1475376 B +432 B / +0.02929% (worse) 421998 B +149 B / +0.03532% (worse) 13.181 s -625.3 ms / -4.5% (better) 3.249 ms -233.6 us / -6.7% (better)
Linux println 61968 B 0 B / +0.0% 14668 B 0 B / +0.0% 656.651 ms +43.46 ms / +7.1% (worse) 1.698 ms +44.65 us / +2.7% (worse)
Linux println-lto 54040 B 0 B / +0.0% 12115 B 0 B / +0.0% 926.574 ms +15.03 ms / +1.6% (worse) 1.836 ms +34.2 us / +1.9% (worse)
macOS cprintf 84480 B 0 B / +0.0% 17101 B 0 B / +0.0% 756.599 ms -237.1 ms / -23.9% (better) 2.995 ms -3.099 ms / -50.9% (better)
macOS cprintf-lto 84288 B 0 B / +0.0% 12865 B 0 B / +0.0% 629.643 ms -234 ms / -27.1% (better) 3.273 ms -1.052 ms / -24.3% (better)
macOS fmtprintf 1483488 B +80 B / +0.005393% (worse) 859624 B +416 B / +0.04842% (worse) 3.022 s -1.273 s / -29.6% (better) 4.400 ms -789.3 us / -15.2% (better)
macOS fmtprintf-lto 1175808 B 0 B / +0.0% 830864 B +380 B / +0.04576% (worse) 6.733 s -2.288 s / -25.4% (better) 4.371 ms -6.908 ms / -61.2% (better)
macOS println 114672 B 0 B / +0.0% 34668 B 0 B / +0.0% 759.202 ms -92.95 ms / -10.9% (better) 4.192 ms -485.8 us / -10.4% (better)
macOS println-lto 118720 B 0 B / +0.0% 32112 B 0 B / +0.0% 1.047 s +82.18 ms / +8.5% (worse) 4.720 ms +622.9 us / +15.2% (worse)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.102 s +13.5 ms / +1.2% (worse) 3.478 ms +67.7 us / +2.0% (worse)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.149 s +42.47 ms / +3.8% (worse) 3.522 ms +125 us / +3.7% (worse)
Windows MinGW fmtprintf 1910272 B +1024 B / +0.1% (worse) 588310 B +176 B / +0.02993% (worse) 4.137 s +40.34 ms / +1.0% (worse) 7.533 ms -284 us / -3.6% (better)
Windows MinGW fmtprintf-lto 1931264 B +1024 B / +0.1% (worse) 534534 B +128 B / +0.02395% (worse) 10.030 s +8.83 ms / +0.1% (worse) 7.786 ms -766.4 us / -9.0% (better)
Windows MinGW println 71168 B 0 B / +0.0% 23718 B 0 B / +0.0% 1.122 s +23.03 ms / +2.1% (worse) 6.614 ms +88.2 us / +1.4% (worse)
Windows MinGW println-lto 65024 B 0 B / +0.0% 20406 B 0 B / +0.0% 1.297 s -39.62 ms / -3.0% (better) 6.288 ms -299.2 us / -4.5% (better)
Windows MinGW 386 cprintf 42496 B 0 B / +0.0% 5326 B 0 B / +0.0% 912.980 ms -71.79 ms / -7.3% (better) 3.921 ms -611.3 us / -13.5% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 930.955 ms -80.97 ms / -8.0% (better) 4.032 ms -379.8 us / -8.6% (better)
Windows MinGW 386 fmtprintf 1871872 B 0 B / +0.0% 462558 B +96 B / +0.02076% (worse) 3.302 s -124.4 ms / -3.6% (better) 8.309 ms -371.2 us / -4.3% (better)
Windows MinGW 386 fmtprintf-lto 2149376 B +512 B / +0.02383% (worse) 440274 B +132 B / +0.02999% (worse) 7.721 s -407.4 ms / -5.0% (better) 7.784 ms -376.9 us / -4.6% (better)
Windows MinGW 386 println 90624 B 0 B / +0.0% 19830 B 0 B / +0.0% 894.593 ms -89.26 ms / -9.1% (better) 6.778 ms -2.295 ms / -25.3% (better)
Windows MinGW 386 println-lto 69120 B 0 B / +0.0% 17742 B 0 B / +0.0% 1.073 s -566.7 ms / -34.6% (better) 6.827 ms -1.375 ms / -16.8% (better)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.361 s -161.1 ms / -10.6% (better) 5.920 ms -1.906 ms / -24.4% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.371 s -187.5 ms / -12.0% (better) 5.896 ms -1.981 ms / -25.1% (better)
Windows MinGW ARM64 fmtprintf 1798144 B +1024 B / +0.1% (worse) 500780 B +104 B / +0.02077% (worse) 4.083 s -244.2 ms / -5.6% (better) 11.193 ms -1.922 ms / -14.7% (better)
Windows MinGW ARM64 fmtprintf-lto 1854976 B +512 B / +0.02761% (worse) 465292 B +72 B / +0.01548% (worse) 9.170 s -268.2 ms / -2.8% (better) 12.119 ms -962.4 us / -7.4% (better)
Windows MinGW ARM64 println 68096 B 0 B / +0.0% 22376 B 0 B / +0.0% 1.372 s -179.6 ms / -11.6% (better) 10.120 ms -2.979 ms / -22.7% (better)
Windows MinGW ARM64 println-lto 63488 B 0 B / +0.0% 19444 B 0 B / +0.0% 1.544 s -255.7 ms / -14.2% (better) 10.741 ms -2.231 ms / -17.2% (better)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65782 B 0 B / +0.0% 951.431 ms +38.29 ms / +4.2% (worse) 3.470 ms +186.4 us / +5.7% (worse)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65718 B 0 B / +0.0% 984.478 ms -122.3 ms / -11.1% (better) 3.901 ms -168 us / -4.1% (better)
Windows MSVC fmtprintf 1624576 B +512 B / +0.03153% (worse) 683814 B +160 B / +0.0234% (worse) 3.823 s +97.11 ms / +2.6% (worse) 10.300 ms +1.544 ms / +17.6% (worse)
Windows MSVC fmtprintf-lto 1614336 B +512 B / +0.03173% (worse) 635174 B +128 B / +0.02016% (worse) 9.161 s +324.4 ms / +3.7% (worse) 10.057 ms +1.204 ms / +13.6% (worse)
Windows MSVC println 192512 B 0 B / +0.0% 119142 B 0 B / +0.0% 1.044 s +138.3 ms / +15.3% (worse) 7.557 ms +637.8 us / +9.2% (worse)
Windows MSVC println-lto 189952 B 0 B / +0.0% 116374 B 0 B / +0.0% 1.125 s +45.52 ms / +4.2% (worse) 7.125 ms +301.4 us / +4.4% (worse)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 782.770 ms -20.85 ms / -2.6% (better) 4.422 ms -49.3 us / -1.1% (better)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 815.202 ms +2.372 ms / +0.3% (worse) 4.470 ms -387.2 us / -8.0% (better)
Windows MSVC 386 fmtprintf 1186816 B +512 B / +0.04316% (worse) 445925 B +96 B / +0.02153% (worse) 3.282 s +133.8 ms / +4.3% (worse) 9.985 ms +997.8 us / +11.1% (worse)
Windows MSVC 386 fmtprintf-lto 1222144 B +512 B / +0.04191% (worse) 416873 B +32 B / +0.007677% (worse) 7.397 s +105.5 ms / +1.4% (worse) 9.900 ms +222.7 us / +2.3% (worse)
Windows MSVC 386 println 34304 B 0 B / +0.0% 18673 B 0 B / +0.0% 785.933 ms -3.932 ms / -0.5% (better) 8.023 ms -468 us / -5.5% (better)
Windows MSVC 386 println-lto 32256 B 0 B / +0.0% 16855 B 0 B / +0.0% 945.294 ms -49.35 ms / -5.0% (better) 7.716 ms -373.8 us / -4.6% (better)
Windows MSVC ARM64 cprintf 11264 B 0 B / +0.0% 3976 B 0 B / +0.0% 2.132 s -73.24 ms / -3.3% (better) 7.526 ms -40.3 us / -0.5% (better)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 3868 B 0 B / +0.0% 2.170 s +49.02 ms / +2.3% (worse) 7.514 ms -154.1 us / -2.0% (better)
Windows MSVC ARM64 fmtprintf 1372160 B +512 B / +0.03733% (worse) 501732 B +112 B / +0.02233% (worse) 6.891 s +233.5 ms / +3.5% (worse) 15.885 ms -72.7 us / -0.5% (better)
Windows MSVC ARM64 fmtprintf-lto 1388544 B +512 B / +0.03689% (worse) 468196 B +64 B / +0.01367% (worse) 15.842 s +255 ms / +1.6% (worse) 16.074 ms -62.4 us / -0.4% (better)
Windows MSVC ARM64 println 41472 B 0 B / +0.0% 21624 B 0 B / +0.0% 2.125 s +23.23 ms / +1.1% (worse) 13.534 ms +296.2 us / +2.2% (worse)
Windows MSVC ARM64 println-lto 39424 B 0 B / +0.0% 19372 B 0 B / +0.0% 2.458 s -3.037 ms / -0.1% (better) 14.283 ms +1.254 ms / +9.6% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.760 ns/op -0.17 ns/op / -1.1% (better)
Linux BenchmarkMergeCompilerFlags 255.300 ns/op -55.1 ns/op / -17.8% (better)
Linux BenchmarkMergeLinkerFlags 142.300 ns/op -36.7 ns/op / -20.5% (better)
Linux BenchmarkChannelBuffered 55.090 ns/op -2.66 ns/op / -4.6% (better)
Linux BenchmarkChannelHandoff 13336 ns/op +191 ns/op / +1.5% (worse)
Linux BenchmarkDefer 53.340 ns/op -9.63 ns/op / -15.3% (better)
Linux BenchmarkDirectCall 1.172 ns/op +0.005 ns/op / +0.4% (worse)
Linux BenchmarkGlobalRead 1.555 ns/op +0.378 ns/op / +32.1% (worse)
Linux BenchmarkGlobalWrite 7.769 ns/op -0.023 ns/op / -0.3% (better)
Linux BenchmarkGoroutine 25272 ns/op +1534 ns/op / +6.5% (worse)
Linux BenchmarkInterfaceCall 5.835 ns/op -0.096 ns/op / -1.6% (better)
Linux BenchmarkRuntimeGetG 2.439 ns/op -0.017 ns/op / -0.7% (better)
macOS BenchmarkLookupPCRandom 12.640 ns/op -2.05 ns/op / -14.0% (better)
macOS BenchmarkMergeCompilerFlags 99.760 ns/op -25.44 ns/op / -20.3% (better)
macOS BenchmarkMergeLinkerFlags 72.780 ns/op -5.31 ns/op / -6.8% (better)
macOS BenchmarkChannelBuffered 22.660 ns/op -3.78 ns/op / -14.3% (better)
macOS BenchmarkChannelHandoff 7657 ns/op -2880 ns/op / -27.3% (better)
macOS BenchmarkDefer 28.360 ns/op -13.13 ns/op / -31.6% (better)
macOS BenchmarkDirectCall 1.064 ns/op -0.133 ns/op / -11.1% (better)
macOS BenchmarkGlobalRead 0.944 ns/op -0.6064 ns/op / -39.1% (better)
macOS BenchmarkGlobalWrite 1.049 ns/op -0.323 ns/op / -23.5% (better)
macOS BenchmarkGoroutine 36398 ns/op +4822 ns/op / +15.3% (worse)
macOS BenchmarkInterfaceCall 3.556 ns/op -1.969 ns/op / -35.6% (better)
macOS BenchmarkRuntimeGetG 2.184 ns/op -0.607 ns/op / -21.7% (better)
Windows MinGW BenchmarkLookupPCRandom 13.220 ns/op -0.09 ns/op / -0.7% (better)
Windows MinGW BenchmarkMergeCompilerFlags 618.600 ns/op -3.5 ns/op / -0.6% (better)
Windows MinGW BenchmarkMergeLinkerFlags 527.900 ns/op -5 ns/op / -0.9% (better)
Windows MinGW BenchmarkChannelBuffered 30.900 ns/op -1.4 ns/op / -4.3% (better)
Windows MinGW BenchmarkChannelHandoff 967.500 ns/op -3 ns/op / -0.3% (better)
Windows MinGW BenchmarkDefer 57.840 ns/op +1.9 ns/op / +3.4% (worse)
Windows MinGW BenchmarkDirectCall 1.547 ns/op 0 ns/op / +0.0%
Windows MinGW BenchmarkGlobalRead 1.856 ns/op +0.306 ns/op / +19.7% (worse)
Windows MinGW BenchmarkGlobalWrite 2.460 ns/op -0.013 ns/op / -0.5% (better)
Windows MinGW BenchmarkGoroutine 90522 ns/op +6110 ns/op / +7.2% (worse)
Windows MinGW BenchmarkInterfaceCall 9.001 ns/op +0.295 ns/op / +3.4% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.183 ns/op +0.32 ns/op / +17.2% (worse)
Windows MinGW 386 BenchmarkLookupPCRandom 21.540 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 566.500 ns/op -27.9 ns/op / -4.7% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 529.900 ns/op -8.6 ns/op / -1.6% (better)
Windows MinGW 386 BenchmarkChannelBuffered 33.900 ns/op -0.21 ns/op / -0.6% (better)
Windows MinGW 386 BenchmarkChannelHandoff 692.800 ns/op -61.1 ns/op / -8.1% (better)
Windows MinGW 386 BenchmarkDefer 34.130 ns/op -2.12 ns/op / -5.8% (better)
Windows MinGW 386 BenchmarkDirectCall 1.631 ns/op +0.543 ns/op / +49.9% (worse)
Windows MinGW 386 BenchmarkGlobalRead 1.357 ns/op -0.006 ns/op / -0.4% (better)
Windows MinGW 386 BenchmarkGlobalWrite 6.983 ns/op -0.006 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGoroutine 58041 ns/op +485 ns/op / +0.8% (worse)
Windows MinGW 386 BenchmarkInterfaceCall 7.340 ns/op -0.01 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 1.634 ns/op +0.004 ns/op / +0.2% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.100 ns/op +0.02 ns/op / +0.2% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 564.300 ns/op -45.4 ns/op / -7.4% (better)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 535.900 ns/op -48.8 ns/op / -8.3% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 37.640 ns/op +0.1 ns/op / +0.3% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 2009 ns/op +34 ns/op / +1.7% (worse)
Windows MinGW ARM64 BenchmarkDefer 52.440 ns/op -0.84 ns/op / -1.6% (better)
Windows MinGW ARM64 BenchmarkDirectCall 0.663 ns/op +0.074 ns/op / +12.6% (worse)
Windows MinGW ARM64 BenchmarkGlobalRead 0.593 ns/op -0.1437 ns/op / -19.5% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.663 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGoroutine 56334 ns/op +842 ns/op / +1.5% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.144 ns/op -0.193 ns/op / -4.5% (better)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.795 ns/op -0.009 ns/op / -0.5% (better)
Windows MSVC BenchmarkLookupPCRandom 13.120 ns/op -0.03 ns/op / -0.2% (better)
Windows MSVC BenchmarkMergeCompilerFlags 609.400 ns/op +3.8 ns/op / +0.6% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 551.200 ns/op +21.3 ns/op / +4.0% (worse)
Windows MSVC BenchmarkChannelBuffered 29.150 ns/op +0.63 ns/op / +2.2% (worse)
Windows MSVC BenchmarkChannelHandoff 1052 ns/op -74 ns/op / -6.6% (better)
Windows MSVC BenchmarkDefer 53.540 ns/op -0.43 ns/op / -0.8% (better)
Windows MSVC BenchmarkDirectCall 1.860 ns/op +0.312 ns/op / +20.2% (worse)
Windows MSVC BenchmarkGlobalRead 1.857 ns/op -0.003 ns/op / -0.2% (better)
Windows MSVC BenchmarkGlobalWrite 2.469 ns/op +0.019 ns/op / +0.8% (worse)
Windows MSVC BenchmarkGoroutine 81532 ns/op +3337 ns/op / +4.3% (worse)
Windows MSVC BenchmarkInterfaceCall 8.065 ns/op -0.629 ns/op / -7.2% (better)
Windows MSVC BenchmarkRuntimeGetG 2.171 ns/op +0.306 ns/op / +16.4% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 21.540 ns/op -0.02 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 559.800 ns/op -39.8 ns/op / -6.6% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 521.900 ns/op -5.6 ns/op / -1.1% (better)
Windows MSVC 386 BenchmarkChannelBuffered 34.260 ns/op +0.39 ns/op / +1.2% (worse)
Windows MSVC 386 BenchmarkChannelHandoff 671.100 ns/op -12.3 ns/op / -1.8% (better)
Windows MSVC 386 BenchmarkDefer 40.790 ns/op -1.4 ns/op / -3.3% (better)
Windows MSVC 386 BenchmarkDirectCall 1.357 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGlobalRead 1.629 ns/op +0.27 ns/op / +19.9% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 6.977 ns/op -0.009 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGoroutine 57397 ns/op -570 ns/op / -1.0% (better)
Windows MSVC 386 BenchmarkInterfaceCall 7.330 ns/op -0.01 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 1.630 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.080 ns/op +0.01 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 581 ns/op +4.9 ns/op / +0.9% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 543.500 ns/op +8 ns/op / +1.5% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 39.430 ns/op +0.22 ns/op / +0.6% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 3113 ns/op +487 ns/op / +18.5% (worse)
Windows MSVC ARM64 BenchmarkDefer 64.150 ns/op -0.44 ns/op / -0.7% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.589 ns/op -0.0002 ns/op / -0.03392% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.663 ns/op -0.0026 ns/op / -0.4% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.795 ns/op +0.046 ns/op / +1.2% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 59152 ns/op -395 ns/op / -0.7% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.160 ns/op +0.025 ns/op / +0.6% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.797 ns/op -0.448 ns/op / -20.0% (better)

Timer runtime benchmarks

Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 964.800 ns/op -91.2 ns/op / -8.6% (better)
Linux AfterFuncZeroDelivery/LLGo 34948 ns/op -11710 ns/op / -25.1% (better)
Linux CreateStop/Go 317.800 ns/op -87.2 ns/op / -21.5% (better)
Linux CreateStop/LLGo 1702 ns/op +248 ns/op / +17.1% (worse)
Linux RearmStopped/Go 116 ns/op -0.9 ns/op / -0.8% (better)
Linux RearmStopped/LLGo 1357 ns/op -224 ns/op / -14.2% (better)
Linux ResetActive/Go 68.680 ns/op -0.01 ns/op / -0.01456% (better)
Linux ResetActive/LLGo 784.600 ns/op -42.3 ns/op / -5.1% (better)
Linux ResetHeap1024/Go 67.360 ns/op -0.56 ns/op / -0.8% (better)
Linux ResetHeap1024/LLGo 181.900 ns/op -6.7 ns/op / -3.6% (better)
macOS AfterFuncZeroDelivery/Go 366.700 ns/op -121 ns/op / -24.8% (better)
macOS AfterFuncZeroDelivery/LLGo 58935 ns/op -39472 ns/op / -40.1% (better)
macOS CreateStop/Go 138.900 ns/op -27.3 ns/op / -16.4% (better)
macOS CreateStop/LLGo 471.600 ns/op +23 ns/op / +5.1% (worse)
macOS RearmStopped/Go 57.890 ns/op -9.64 ns/op / -14.3% (better)
macOS RearmStopped/LLGo 324.500 ns/op -98 ns/op / -23.2% (better)
macOS ResetActive/Go 38.910 ns/op -16.18 ns/op / -29.4% (better)
macOS ResetActive/LLGo 157.400 ns/op +6.5 ns/op / +4.3% (worse)
macOS ResetHeap1024/Go 45.660 ns/op -1.97 ns/op / -4.1% (better)
macOS ResetHeap1024/LLGo 77.170 ns/op -25.13 ns/op / -24.6% (better)
Windows MinGW AfterFuncZeroDelivery/Go 548.100 ns/op -14.1 ns/op / -2.5% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 166152 ns/op -1217 ns/op / -0.7% (better)
Windows MinGW CreateStop/Go 117.300 ns/op +0.2 ns/op / +0.2% (worse)
Windows MinGW CreateStop/LLGo 450.400 ns/op +12.2 ns/op / +2.8% (worse)
Windows MinGW RearmStopped/Go 31.430 ns/op +0.07 ns/op / +0.2% (worse)
Windows MinGW RearmStopped/LLGo 276 ns/op -8.8 ns/op / -3.1% (better)
Windows MinGW ResetActive/Go 20.100 ns/op -0.08 ns/op / -0.4% (better)
Windows MinGW ResetActive/LLGo 157.900 ns/op +4.7 ns/op / +3.1% (worse)
Windows MinGW ResetHeap1024/Go 20.350 ns/op -0.13 ns/op / -0.6% (better)
Windows MinGW ResetHeap1024/LLGo 127.800 ns/op -0.8 ns/op / -0.6% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 775.100 ns/op +8.9 ns/op / +1.2% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 125199 ns/op -1500 ns/op / -1.2% (better)
Windows MinGW 386 CreateStop/Go 169.200 ns/op -3.6 ns/op / -2.1% (better)
Windows MinGW 386 CreateStop/LLGo 395.900 ns/op +5.4 ns/op / +1.4% (worse)
Windows MinGW 386 RearmStopped/Go 56.820 ns/op +0.06 ns/op / +0.1% (worse)
Windows MinGW 386 RearmStopped/LLGo 281.300 ns/op +6.3 ns/op / +2.3% (worse)
Windows MinGW 386 ResetActive/Go 32.700 ns/op +0.06 ns/op / +0.2% (worse)
Windows MinGW 386 ResetActive/LLGo 873.400 ns/op -67 ns/op / -7.1% (better)
Windows MinGW 386 ResetHeap1024/Go 32.860 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW 386 ResetHeap1024/LLGo 148.700 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 667.400 ns/op +5.7 ns/op / +0.9% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 119745 ns/op -19427 ns/op / -14.0% (better)
Windows MinGW ARM64 CreateStop/Go 197.700 ns/op -4.8 ns/op / -2.4% (better)
Windows MinGW ARM64 CreateStop/LLGo 364.800 ns/op -0.5 ns/op / -0.1% (better)
Windows MinGW ARM64 RearmStopped/Go 70.540 ns/op -0.02 ns/op / -0.02834% (better)
Windows MinGW ARM64 RearmStopped/LLGo 251.100 ns/op -5.4 ns/op / -2.1% (better)
Windows MinGW ARM64 ResetActive/Go 31.050 ns/op +0.04 ns/op / +0.1% (worse)
Windows MinGW ARM64 ResetActive/LLGo 123.900 ns/op -0.8 ns/op / -0.6% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.050 ns/op -0.07 ns/op / -0.2% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 125.800 ns/op -1 ns/op / -0.8% (better)
Windows MSVC AfterFuncZeroDelivery/Go 553.300 ns/op -22.7 ns/op / -3.9% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 151935 ns/op -8734 ns/op / -5.4% (better)
Windows MSVC CreateStop/Go 115.800 ns/op +0.3 ns/op / +0.3% (worse)
Windows MSVC CreateStop/LLGo 428.200 ns/op +16 ns/op / +3.9% (worse)
Windows MSVC RearmStopped/Go 31.270 ns/op -0.18 ns/op / -0.6% (better)
Windows MSVC RearmStopped/LLGo 266 ns/op +5 ns/op / +1.9% (worse)
Windows MSVC ResetActive/Go 20.040 ns/op -0.17 ns/op / -0.8% (better)
Windows MSVC ResetActive/LLGo 140.900 ns/op -27 ns/op / -16.1% (better)
Windows MSVC ResetHeap1024/Go 20.570 ns/op -0.01 ns/op / -0.04859% (better)
Windows MSVC ResetHeap1024/LLGo 127.600 ns/op +2.2 ns/op / +1.8% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 781.400 ns/op +10.2 ns/op / +1.3% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 121527 ns/op -2122 ns/op / -1.7% (better)
Windows MSVC 386 CreateStop/Go 168.200 ns/op -0.9 ns/op / -0.5% (better)
Windows MSVC 386 CreateStop/LLGo 367.500 ns/op -9.1 ns/op / -2.4% (better)
Windows MSVC 386 RearmStopped/Go 56.780 ns/op +0.12 ns/op / +0.2% (worse)
Windows MSVC 386 RearmStopped/LLGo 270.200 ns/op +3.2 ns/op / +1.2% (worse)
Windows MSVC 386 ResetActive/Go 32.600 ns/op +0.02 ns/op / +0.1% (worse)
Windows MSVC 386 ResetActive/LLGo 894.200 ns/op +44.4 ns/op / +5.2% (worse)
Windows MSVC 386 ResetHeap1024/Go 32.870 ns/op +0.26 ns/op / +0.8% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 143.200 ns/op +0.7 ns/op / +0.5% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 671.500 ns/op +3.9 ns/op / +0.6% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 151234 ns/op -2865 ns/op / -1.9% (better)
Windows MSVC ARM64 CreateStop/Go 213.200 ns/op +14.5 ns/op / +7.3% (worse)
Windows MSVC ARM64 CreateStop/LLGo 406.300 ns/op +10.2 ns/op / +2.6% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.580 ns/op -0.11 ns/op / -0.2% (better)
Windows MSVC ARM64 RearmStopped/LLGo 273.700 ns/op +3.9 ns/op / +1.4% (worse)
Windows MSVC ARM64 ResetActive/Go 31.010 ns/op -0.04 ns/op / -0.1% (better)
Windows MSVC ARM64 ResetActive/LLGo 131.600 ns/op -5 ns/op / -3.7% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.090 ns/op -0.01 ns/op / -0.03215% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 136.700 ns/op -2.1 ns/op / -1.5% (better)

Compared with 2db247e43848 measured in the same runner job.

@github-actions

github-actions Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

cb7214ae45fa | workflow run | long-term charts

WebAssembly output sizes

Profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 130908 B +17271 B / +15.2% (worse) 70821 B +85 B / +0.1% (worse)
cprintf/j32-goos-js/LLGo 129206 B +62136 B / +92.6% (worse) 69185 B +674 B / +1.0% (worse)
cprintf/j64-emscripten-memory64/LLGo 119514 B +81 B / +0.1% (worse) 74033 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 133476 B +59729 B / +81.0% (worse) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 133351 B +15670 B / +13.3% (worse) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3229202 B +475244 B / +17.3% (worse) 113679 B +4973 B / +4.6% (worse)
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3225001 B +1066106 B / +49.4% (worse) 112043 B +5402 B / +5.1% (worse)
fmtprintf/j64-emscripten-memory64/LLGo 2930467 B +1708 B / +0.1% (worse) 120857 B -1 B / -0.0008274% (better)
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 3037664 B +1095292 B / +56.4% (worse) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2912768 B +401740 B / +16.0% (worse) 0 B 0 B / 0.0%
j32-emscripten/LLGo 130034 B +17129 B / +15.2% (worse) 70821 B +85 B / +0.1% (worse)
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 128561 B +62169 B / +93.6% (worse) 69185 B +674 B / +1.0% (worse)
j64-emscripten-memory64/LLGo 118745 B +78 B / +0.1% (worse) 74033 B 0 B / +0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 132619 B +59930 B / +82.4% (worse) 0 B 0 B / 0.0%
w32-wasi/LLGo 132583 B +15576 B / +13.3% (worse) 0 B 0 B / 0.0%

LLGo WebAssembly build measurements

Profile Build vs base
j32-emscripten 6.225 s +409.4 ms / +7.0% (worse)
j32-goos-js 6.185 s +1.044 s / +20.3% (worse)
j64-emscripten-memory64 5.331 s -36.68 ms / -0.7% (better)
w32-goos-wasip1 4.781 s +1.622 s / +51.4% (worse)
w32-wasi 4.544 s +323.5 ms / +7.7% (worse)

Compared with 2db247e43848 measured in the same runner job.

@cpunion

cpunion commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed the non-inline review notes in 5943d5e as well: GoWordSize is computed once in NewProgram; native-storage propagation now returns immediately on ordinary native/J64 targets; canonical types cache and reuse one native view instead of allocating a fresh clone on each query; and a regression test checks both the W1 cache and the native fast path. Removed the explicit wasip2/wasm-unknown output entries and their stale test rows. The 4-byte J32 GC scan stride remains intentional for physical host/C pointers.

@cpunion

cpunion commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai Please re-review the latest head 5943d5eb0. It addresses all four inline findings and both non-inline actionable notes: CIF-owned wasm32 arg storage, explicit instantiation rejection propagation with a regression test, the dead assignment and unused build-tag parameters, cached/target-gated/deduplicated storage views, and stale target output entries. Please focus on the CIF lifetime/container layout and whether the native hot-path fast return is complete.

@fennoai

fennoai Bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Re-review complete

Posted a non-blocking PR review of head 5943d5eb0.

Requested focus areas — both verified correct:

  • CIF lifetime / container layout (newSignatureStorage): the CIF is the first struct field, rooting args and the physical 4-byte slots array for the CIF's lifetime; wasm32 slot width matches the index_wasm32.go read side; J64/memory64 uses the 8-byte default path via build tags.
  • Native hot-path fast return (ssa/stmt_builder.go Return): the !isNativeStorage branch reconstructs the same toLLVMTuple-based aggregate the InGo function was declared with; native codegen short-circuits all wide-storage helpers.

Prior findings confirmed addressed: instantiation rejection propagation (+ TestWasmProfileValidationErrors), removed wasip2/wasm-unknown targets, DefaultBuildTags unused params dropped, dead-assignment/closure params cleaned, native-closure return lowering fixed. internal/crosscompile + internal/targets pass locally; git diff --check clean.

Two non-blocking items raised inline: a latent wasm32-only GC-liveness gap in StructOf/ArrayOf's typePointerArray (element *Type roots not retained the way newSignatureStorage retains them), and a minor per-alloca TargetData overhead in requireStorageAlignment on native targets.

SSA-level tests need LLVM headers unavailable in this sandbox; those rely on CI per the PR description.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review of 5943d5eb0 — WASM W1 profiles and ABI foundation

I re-reviewed the fixed range 52ae044d...5943d5eb0 with the two requested focuses, and ran the four review passes (quality, performance, security, docs).

Requested focus areas — both verified correct:

  • CIF lifetime / container layout (newSignatureStorage). signatureStorage.cif is the first struct field, so &storage.cif roots the whole allocation; args (typed *Type roots) and slots (the physical 4-byte uint32 array libffi retains after PrepCif) stay reachable for the CIF's lifetime. The wasm32 4-byte slot width matches the read side in index_wasm32.go (base + i*4, read as *uint32), and J64/memory64 correctly falls through to the default 8-byte path via build tags. The !llgo-tagged TestNewSignatureStorageOwnsArgumentArray passes. ✅
  • Native hot-path fast return (ssa/stmt_builder.go Return). The !isNativeStorage branch reconstructs aggregateValue(rawType(tret), …), and rawType(*types.Tuple) lowers through the same toLLVMTuple(out) used to declare InGo function return types — so the reconstructed aggregate matches the declared type. The native branch reads Func.ll.ReturnType() element types directly, matching toLLVMNativeTypes. All storage helpers short-circuit on GoWordSize() <= PointerSize() (native targets), and isNativeStorage is only ever populated when usesWideGoStorage(), so native codegen never enters the wide-storage reconstruction. ✅

Prior findings — confirmed addressed:

  • Explicit instantiation rejection now propagates (validWasmSelection, unsupported-profile/provider errors), covered by TestWasmProfileValidationErrors; wasip2.json / wasm-unknown.json removed.
  • DefaultBuildTags dropped its unused goarch, target params; the sizes closure's unused params are now _, _; native-closure return lowering fixed (toLLVMFuncBackground(entrySig, bg) instead of forced InGo).
  • internal/crosscompile and internal/targets suites pass locally (40.6s / 2.6s); git diff --check clean; no new t.Skip, probes, or debug prints.

(SSA-level tests require LLVM headers not present in this sandbox; relying on CI for those, per the PR description.)

One latent lifetime finding (wasm32-only) and one minor perf nit are inline below. Neither blocks the two focus areas, which are sound.

Comment thread runtime/internal/ffi/pointer_array_wasm32.go Outdated

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Follow-up: one minor performance nit that did not fit inline in the prior review event.

Comment thread ssa/storage.go
@cpunion

cpunion commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai Please re-review the latest head fcafb8546. It addresses both follow-up findings, extends the lifetime audit to the CIF return type, and adds focused coverage for J32 native arrays, signed width conversion, physical pointer indexes, and synthetic aggregate fields. Please verify the aggregate/signature first-field ownership containers and the native alignment fast return.

@fennoai

fennoai Bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Re-review complete — follow-up fixes verified

Posted a non-blocking PR review of head fcafb8546. Both prior findings are correctly resolved.

  • First-field ownership containers verified. StructOf/ArrayOf/newComplexType now route through newAggregateType, wrapping element roots in aggregateTypeStorage (first field typ Type, plus retained elements/slots); &storage.typ roots the whole object, keeping the typed element roots GC-reachable — closing the wasm32 use-after-free gap. signatureStorage now also retains ret *Type, closing the return-type lifetime gap. No unrooted Type.Elements producer remains.
  • Native alignment fast return verified. requireStorageAlignment early-returns on !usesWideGoStorage(); a pure no-op removal on native targets with no behavior change, wasm path intact.
  • New coverage confirmed: ownership tests (incl. ret retention) pass locally; added J32 native-array / signed-width / physical-index / synthetic-aggregate-field SSA tests. git diff --check clean; no new skips/probes.

One non-blocking note: the ffi ownership tests are //go:build !llgo, so the wasm32 slot-encoding variant stays CI-unverified on the default target (rooting structure is identical across builds, so the design is validated). SSA-level tests rely on CI (no LLVM headers in this sandbox).

Verdict: LGTM.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review of fcafb8546 — follow-up fixes verified

Both prior findings are correctly resolved, and the lifetime audit was extended to the CIF return type as described. Ran all four review passes on the incremental diff (5943d5eb0..fcafb8546); no blocking issues.

Requested verification points — both confirmed:

  • Aggregate/signature first-field ownership containers. StructOf/ArrayOf/newComplexType now route through the new newAggregateType, which wraps element roots in aggregateTypeStorage{typ Type; elements []*Type; slots []uint32} (wasm32) / {typ Type; elements []*Type} (default). typ is the first field, so returning &storage.typ roots the whole allocation, keeping the typed elements roots (and physical slots) GC-reachable for the Type's lifetime — closing the wasm32 use-after-free gap where dynamically-built nested types (e.g. in mfinal_ffi.go) could be collected. signatureStorage now also retains ret *Type, co-rooted via &storage.cif, closing the analogous return-type lifetime gap. No unrooted Type.Elements producer remains (only the permanently-rooted typeEmpty). ✅
  • Native alignment fast return. requireStorageAlignment now early-returns on !usesWideGoStorage(). On native targets the skipped work was always a no-op (storageType(t) == t.ll, so want == have and SetAlignment was never called), so this is a pure overhead removal — three per-alloca calls including a TargetData query — with no behavior change, and the wasm path is unaffected. ✅

New coverage confirmed: TestNewAggregateTypeOwnsElementArray + extended TestNewSignatureStorageOwnsTypes (now asserts ret retention) pass locally; TestWasm32StorageIntegerConversions, TestWasm32PhysicalPointerIndexBounds, and TestSyntheticFieldOutOfRangePanics added for J32 native arrays / signed width / physical indexes / synthetic aggregate fields. git diff --check clean; no new t.Skip/probes.

One non-blocking note: the ffi ownership tests are //go:build !llgo, so they exercise the default container rather than the wasm32 slot-encoding variant (typePointerArrayFromSlots and the uint32 narrowing). The rooting structure is identical across both builds, so the design is validated — but the wasm32-specific encoding remains CI-unverified on the default target. A wasm-tagged test would close that gap if/when a harness is available.

LGTM. SSA-level tests require LLVM headers unavailable in this sandbox and rely on CI per the PR description.

@cpunion

cpunion commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai Please re-review W1 after its conflict-free rebase onto current main 4564a01. The 16-commit range moved from fcafb85 to db71cc5, and git range-diff reports all 16 commits exactly equivalent. Please focus on interactions with the newly merged C ABI aggregate-return change (#2570) and panic/Goexit fixes (#2567), plus confirmation that prior review findings remain resolved. No content changes were made during rebase.

@fennoai

fennoai Bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Rebase re-review complete

Posted a non-blocking PR review of head db71cc582. The rebase is content-preserving and integrates cleanly with both newly-merged PRs.

  • Rebase equivalence: git range-diff reports all 16 W1 commits exactly equivalent (=) — no content changes.
  • Prior findings still resolved: first-field ownership containers (incl. retained CIF ret) and the requireStorageAlignment native fast return are intact; ffi ownership + crosscompile + targets suites pass locally.
  • cabi: place aggregate return conversion storage in entry #2570 (C ABI aggregate-return) — safe: W1 doesn't touch internal/cabi/; cabi: place aggregate return conversion storage in entry #2570's return-slot rework is gated to AttrWidthType/AttrWidthType2, which the wasm classifier never emits (wasm uses the sret path). W1's wide-pointer {ptr, iN} storage is structurally excluded from C-ABI aggregate returns, so it can't corrupt the cabi bitcast/load.
  • fix(runtime): preserve recovered panic source locations #2567 (panic line / Goexit) — safe: W1 doesn't touch caller.go/unwind_llgo.go/z_rt.go; shared-file edits are in disjoint regions. W1's AddIncoming terminator-move reads only the nil-check success-block tail (never the failure block's terminator kind), and NeedsFramePointer() is false on wasm, so fix(runtime): preserve recovered panic source locations #2567's new CreateRet nil-check path is inert on the W1 profiles.

Verdict: LGTM. SSA-level codegen tests rely on CI (no LLVM headers in this sandbox).

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review of db71cc582 — post-rebase onto main 4564a01aa

Rebase verified clean and the two newly-merged PRs interact safely with W1. No new findings.

Rebase equivalence confirmed. git range-diff 52ae044d..fcafb8546 4564a01aa..db71cc582 reports all 16 W1 commits = (exactly equivalent) — no content changes during rebase, matching the description.

Prior findings remain resolved (verified at db71cc582):

  • Aggregate/signature first-field ownership containers (newAggregateType / signatureStorage with retained ret) are intact; TestNewAggregateTypeOwnsElementArray and TestNewSignatureStorageOwnsTypes pass.
  • requireStorageAlignment native fast return (!usesWideGoStorage()) is intact.
  • internal/crosscompile + internal/targets suites pass locally.

Interaction with #2570 (C ABI aggregate-return) — safe. W1 does not touch internal/cabi/ at all. #2570's reworked return-slot logic is gated to AttrWidthType/AttrWidthType2, which the wasm classifier (TypeInfoWasm) never emits — wasm aggregate returns take the AttrPointer/sret path, so that code is inert on wasm. Independently, W1's wide-pointer storage {ptr, iN} can never appear in a C-ABI aggregate return: native-storage types are excluded from wide-pointer wrapping (needsWidePointerStorage returns false when isNativeStorage), and nativeStorageLLVMType has no pointer case, so C-background pointer fields stay plain ptr — identical to the pre-W1 layout cabi expects. No path lets W1 layout corrupt the cabi return-slot bitcast/load.

Interaction with #2567 (panic line / Goexit) — safe. W1 does not touch caller.go, unwind_llgo.go, or z_rt.go. In the shared files the two changes occupy disjoint regions (W1: storage/return/phi lowering; #2567: PanicSite/recordPanicSite/nil-check-return), cleanly merged. The one plausible coupling — W1's AddIncoming terminator-move vs. #2567's new formal CreateRet(undef) nil-check failure block — is decoupled by design: AssertNilDeref always retargets b.blk.last to the success block blks[1] before returning, and AddIncoming reads only that success-block tail; it never inspects the failure block's terminator kind. Moreover NeedsFramePointer() is false on wasm (funcinfo.go), so #2567's new CreateRet path and preserveNilCheckCondition are inert on the W1 profiles, and that path bypasses Builder.Return() regardless.

Verdict: LGTM — rebase is content-preserving and integrates cleanly with #2570 and #2567. SSA-level codegen tests require LLVM headers unavailable in this sandbox and rely on CI per the PR description.

@cpunion
cpunion force-pushed the codex/wasm-w1-profiles-20260910 branch from db71cc5 to 6b73504 Compare September 14, 2026 21:32
@cpunion

cpunion commented Sep 15, 2026

Copy link
Copy Markdown
Collaborator Author

The fork-only macOS comparison completed successfully on both exact W1 head and its base, without sampling, retries, or reduced test pressure; details are now in the validation section. The original upstream signal termination has not been reproduced or fixed. The independently reproduced loss of signal diagnostics in llgo test is isolated in #2600, with a failing-before/passing-after subprocess regression and no runtime changes. W1 has not been changed or pushed for this investigation.

@cpunion
cpunion force-pushed the codex/wasm-w1-profiles-20260910 branch from 6b73504 to cb7214a Compare September 15, 2026 05:21
@xushiwei
xushiwei merged commit f5208c1 into xgo-dev:main Sep 16, 2026
66 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants