Skip to content

fix(ssa): stabilize local stack slots and synthetic debug locations - #2530

Merged
cpunion merged 10 commits into
xgo-dev:mainfrom
cpunion:codex/fix-loop-local-stack-20260908
Sep 14, 2026
Merged

cpunion merged 10 commits into
xgo-dev:mainfrom
cpunion:codex/fix-loop-local-stack-20260908

Conversation

@cpunion

@cpunion cpunion commented Sep 8, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Reserve each non-escaping Go local's stack slot once in the LLVM entry block while leaving zero-initialization at the declaration. A declaration inside a loop must reuse one slot per function call; previously, emitting alloca in the loop grew the stack on every iteration. The bounded WASI polling test exposed this after roughly 1,300 retries on a 128 KiB Fiber stack.

This PR also contains the synthetic-builder debug-location prerequisite originally reviewed in #2529. It preserves source locations across generated blocks and prevents invalid or missing debug locations in synthetic defer paths. #2529 remains review history and is not a separate merge dependency.

Correctness and review follow-up

alloca reserves storage; memset initializes that storage. Hoisting one does not remove the other. The current FileCheck fixtures assert the typed and aligned entry-block reservation, bind its address to declaration-time initialization, and follow subsequent value use. The loop regression uses two differently sized and aligned slots and verifies each initializer's destination and full byte count.

The local-allocation hot path uses a raw LLVM builder rather than allocating a full Go builder and unused debug-scope cache. The builder is disposed immediately after the entry-block allocation.

Validation

The previously validated head 826e07e5e414 was based on main 410af8e93a1c, which includes the merged #2575 weak-cleanup reentrancy fix. The original seven compiler/test patches remain unchanged according to git range-diff. The follow-up only corrects four comment lines about channel scratch-slot scope and includes the real subprocess error in two failing compatibility assertions; it does not change compiler execution or add retries. That revision changed 20 files, +474/-215, with no workflow changes.

The normal LLGo workflow for this head and all other normal checks completed with 64 successes, one expected tag-only release skip, and no failures. Patch coverage is 100%. All seven Windows LLGo jobs passed; the MinGW ARM64 test step completed in 4m22s (28m32s for the job), and the previously interrupted MSVC ARM64 shard-0 test step completed in 5m30s (30m07s for the job).

  • go test ./ssa -count=1 -timeout=5m passes locally with Go 1.27 and LLVM 22 (6.2 seconds), including the loop-local allocation and synthetic debug-location regressions. An independent coverage run hits all 30 added coverable lines.
  • The _testgo compiler/FileCheck group passed locally (123.4 seconds), covering all six changed fixtures. The five changed _testrt fixtures passed in a separately bounded invocation (170.5 seconds): go test ./cl -run '^TestRunAndTestFromTestrt$/^(float2any|gblarray|index|tpmap|unsafe)$' -count=1 -timeout=5m -v. These exercise the rebased compiler and runtime; they are not a claim of an unfiltered local ./cl pass.
  • On 826e07e5e414, go test ./ssa ./test/cmd/llgo -count=1 -timeout=5m passes locally (6.0s and 5.5s respectively).
  • git diff --check passes.

Previous Windows ARM64 failures

Two earlier Windows ARM64 jobs ended with GitHub's hosted-runner-lost-communication annotation: MSVC shard 0 on 0da6c73dc2be, then MinGW ARM64 on 0401ed90c589. The failed-job archives are unavailable. The preserved MinGW live log showed empty compiler-subprocess output and native Go cgo.exe: exit status 2 before the runner disappeared, but no panic, native stack, or OOM diagnostic. This establishes runner loss, not a specific test timeout, weak deadlock, or compiler fault.

The first single-job fork diagnosis passed the full LLGo tests plus 20 TLS repetitions, ten isolated repetitions of the affected tool-compile tests, native Go controls, and go list -export runtime/cgo. A second diagnostic used the same MinGW ARM64 job and 15-second resource sampler to run the full suite three times on the PR compiler changes and on main:

Revision Full-suite iteration times Sampled peak LLGo RSS (MiB) Minimum free physical memory (MiB)
PR diagnostic 4m27s / 3m07s / 3m09s 4,983 / 5,711 / 4,925 6,630 / 6,671 / 7,539
main control 4m18s / 3m07s / 3m04s 4,590 / 5,529 / 5,227 6,675 / 6,628 / 6,902

Both diagnostic jobs passed, including the isolated LLGo repetitions and native Go controls. The measurements came from separate runner instances and a 15-second sampler, so they are observational rather than a precise allocation benchmark; they show no progressive or PR-specific memory exhaustion in these runs. The earlier runner losses were not reproduced, and their root cause remains unconfirmed. The temporary diagnostic workflow is not part of this contribution, and no retry or failure-suppression behavior was added.

Size and performance

The pre-rebase benchmark run passed all 9 jobs and compares 826e07e5e414 with main 410af8e93a1c. None of the 48 native program variants grew in file size; cprintf is unchanged with and without LTO on every platform.

Platform fmtprintf / LTO size change println / LTO size change
Linux -2,616 / -4,072 B -256 / -240 B
macOS -96 / -16 B 0 / 0 B
Windows MSVC AMD64 -5,632 / -6,656 B 0 / 0 B
Windows MSVC 386 -1,024 / -2,048 B 0 / 0 B
Windows MSVC ARM64 -2,560 / -3,584 B 0 / -512 B
Windows MinGW AMD64 -6,144 / -7,168 B -512 / -512 B
Windows MinGW 386 -1,536 / -1,536 B 0 / 0 B
Windows MinGW ARM64 -3,072 / -3,584 B -512 / 0 B

The Memory64 ec64 module grows by 8 B; the other LLGo Wasm modules shrink by 228–350 B, and all generated glue sizes are unchanged.

Build and run timings are single samples, not a statistically established speedup or regression. The six native-build aggregates stay within -1.35% to +1.60% on Windows and are -3.52% on Linux; macOS reports +25.42%, while unchanged Go controls and compiler benchmarks on that runner also slow down. These results do not establish a PR-caused broad slowdown, and they do not establish that every individual benchmark is regression-free. All benchmark jobs completed within 25 minutes (maximum 22m19s).

This contribution supersedes fork PR cpunion#243.

Latest main rebase (2026-09-14)

Rebased onto current main 5c5874359 without conflicts. The implementation patches remain unchanged; commit 726fc6875 only updates the two receive/select comments requested in review so they accurately state that function-entry slots remain valid across StackRestore. The full ssa package passes locally. Full local cl testing reached the macOS 10-minute package limit while executing generated binaries, with no assertion failure; fresh CI is running on the rebased head. All review threads were answered and resolved.

LTO failure dependency

This branch now also includes the independently reviewed fix from #2562. The rebased CI LTO job https://github.com/xgo-dev/llgo/actions/runs/34809862200/job/103868886091 timed out in TestLTOPluginRejectsMalformedInterfaceAttributes while waiting for opt after invalid metadata triggered an abort. The dependency makes invalid metadata exit with status 1 and bounds the regression subprocess; it does not suppress errors or retry tests. With the LLVM 22 plugin built from the same fix and the dev build tag enabled, all three malformed-metadata regression cases pass locally (no skips). The combined head fc0ee5071 has now passed all 64 checks, including the formerly failing LTO job and all benchmark jobs. Codecov confirms all modified coverable lines are covered, and the current-head fennoai review has no unresolved findings. The macOS coverage job passed in 59m45s, close to the existing 60-minute limit; no timeout increase or retry behavior was added.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: fix loop-local stack growth + synthetic-builder debug locations

This is a focused, well-executed change. Two independent fixes:

  1. Hoist stack-local allocas into the entry block (ssa/memory.go) so a local declared inside a loop reserves one slot per call instead of growing the stack every iteration. The reservation moves to the entry block while zero-initialization (memset) correctly stays at the declaration site inside the loop body — verified by the new TestLocalLoopAllocReservesOneSlotPerCall, including the GC-roots variant.
  2. Track debug location on the Builder (ssa/di.go, ssa/decl.go, ssa/eh.go, ssa/stmt_builder.go) so synthetic builders inherit a valid function scope, satisfying LLVM's requirement that inlinable calls in a debug-info function carry a !dbg location.

Correctness / safety: No issues. Entry-block allocas dominate all uses and are always inserted before the first instruction (never past a terminator). Slot reuse is only taken for non-escaping (heap=false) locals — escaping loop-locals go through the heap AllocZ path — so per-call semantics are preserved. Comments accurately describe behavior. VerifyModule passes in the new tests.

Suggestions (non-blocking):

  • One inline note below on the per-alloc builder allocation cost.
  • Consider adding a debug-info-enabled case to TestLocalLoopAllocReservesOneSlotPerCall (it currently runs without InitDebug), to lock in that the hoisted alloca's !dbg interaction verifies cleanly.

Nice work.

Comment thread ssa/memory.go Outdated
@codecov

codecov Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@cpunion
cpunion force-pushed the codex/fix-loop-local-stack-20260908 branch 2 times, most recently from 85e3775 to ea25e9d Compare September 8, 2026 02:23
@github-actions

github-actions Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

fc0ee5071896 | workflow run | long-term charts

WebAssembly output sizes

Profile and compiler Wasm module vs base Generated JS glue vs base
ec32/LLGo 112905 B -338 B / -0.3% (better) 70736 B 0 B / +0.0%
ec64/LLGo 118667 B -21 B / -0.01769% (better) 74033 B 0 B / +0.0%
js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
js/LLGo 66392 B -268 B / -0.4% (better) 68511 B 0 B / +0.0%
wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
wasip1/LLGo 72689 B -228 B / -0.3% (better) 0 B 0 B / 0.0%
wc32/LLGo 117007 B -352 B / -0.3% (better) 0 B 0 B / 0.0%

LLGo WebAssembly build measurements

Profile Build vs base
ec32 5.998 s -23.36 ms / -0.4% (better)
ec64 5.447 s -464.4 ms / -7.9% (better)
js 4.835 s +186.2 ms / +4.0% (worse)
wasip1 3.109 s +40.16 ms / +1.3% (worse)
wc32 3.897 s -269.4 ms / -6.5% (better)

Compared with 5c5874359c1e measured in the same runner job.

@cpunion
cpunion force-pushed the codex/fix-loop-local-stack-20260908 branch from ea25e9d to 804d14f Compare September 11, 2026 01:48
@cpunion cpunion changed the title fix(ssa): reserve loop-local stack slots once per call fix(ssa): stabilize local stack slots and synthetic debug locations Sep 11, 2026
@github-actions

github-actions Bot commented Sep 11, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

fc0ee5071896 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 19832 B 0 B / +0.0% 387 B 0 B / +0.0% 473.123 ms -32.01 ms / -6.3% (better) 1.364 ms +65.82 us / +5.1% (worse)
Linux cprintf-lto 19584 B 0 B / +0.0% 368 B 0 B / +0.0% 510.012 ms +13.38 ms / +2.7% (worse) 1.267 ms -14.76 us / -1.2% (better)
Linux fmtprintf 1623352 B -2528 B / -0.2% (better) 495623 B -2339 B / -0.5% (better) 3.760 s -386.5 ms / -9.3% (better) 3.034 ms -277.5 us / -8.4% (better)
Linux fmtprintf-lto 1475664 B -4512 B / -0.3% (better) 434912 B -2909 B / -0.7% (better) 10.949 s -697 ms / -6.0% (better) 2.967 ms -223 us / -7.0% (better)
Linux println 61968 B -256 B / -0.4% (better) 14668 B -233 B / -1.6% (better) 496.158 ms +13.78 ms / +2.9% (worse) 1.624 ms -2.432 us / -0.1% (better)
Linux println-lto 54040 B -240 B / -0.4% (better) 12115 B -220 B / -1.8% (better) 683.124 ms -29.67 ms / -4.2% (better) 1.570 ms -177 us / -10.1% (better)
macOS cprintf 84480 B 0 B / +0.0% 17101 B 0 B / +0.0% 873.885 ms -151.7 ms / -14.8% (better) 3.554 ms -599.4 us / -14.4% (better)
macOS cprintf-lto 84288 B 0 B / +0.0% 12865 B 0 B / +0.0% 1.053 s -37.93 ms / -3.5% (better) 7.313 ms -79.79 us / -1.1% (better)
macOS fmtprintf 1473072 B -96 B / -0.006517% (better) 869620 B -2392 B / -0.3% (better) 4.824 s -413.7 ms / -7.9% (better) 8.621 ms -2.534 ms / -22.7% (better)
macOS fmtprintf-lto 1159392 B -32 B / -0.00276% (better) 843924 B -3828 B / -0.5% (better) 13.450 s +268.8 ms / +2.0% (worse) 10.182 ms +1.896 ms / +22.9% (worse)
macOS println 114672 B 0 B / +0.0% 34668 B -144 B / -0.4% (better) 913.236 ms +30.95 ms / +3.5% (worse) 7.484 ms -2.486 ms / -24.9% (better)
macOS println-lto 118720 B 0 B / +0.0% 32112 B -136 B / -0.4% (better) 1.262 s +272.3 ms / +27.5% (worse) 16.896 ms +11.24 ms / +198.7% (worse)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.116 s +5.512 ms / +0.5% (worse) 3.926 ms +260.6 us / +7.1% (worse)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.239 s +101.4 ms / +8.9% (worse) 4.085 ms +437.8 us / +12.0% (worse)
Windows MinGW fmtprintf 1889792 B -5632 B / -0.3% (better) 595574 B -5536 B / -0.9% (better) 4.345 s +307.1 ms / +7.6% (worse) 9.366 ms +1.596 ms / +20.5% (worse)
Windows MinGW fmtprintf-lto 1930240 B -6656 B / -0.3% (better) 545030 B -5536 B / -1.0% (better) 10.044 s -194.9 ms / -1.9% (better) 8.686 ms +860.8 us / +11.0% (worse)
Windows MinGW println 71168 B 0 B / +0.0% 23718 B -336 B / -1.4% (better) 1.219 s +115.9 ms / +10.5% (worse) 7.683 ms +937.4 us / +13.9% (worse)
Windows MinGW println-lto 65024 B -512 B / -0.8% (better) 20406 B -272 B / -1.3% (better) 1.442 s +124.2 ms / +9.4% (worse) 7.256 ms +653.4 us / +9.9% (worse)
Windows MinGW 386 cprintf 42496 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.096 s +13.87 ms / +1.3% (worse) 5.350 ms +170.3 us / +3.3% (worse)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.108 s +6.498 ms / +0.6% (worse) 5.989 ms +877.9 us / +17.2% (worse)
Windows MinGW 386 fmtprintf 1857536 B -1024 B / -0.1% (better) 471118 B -1312 B / -0.3% (better) 4.029 s +35.75 ms / +0.9% (worse) 10.475 ms -103.9 us / -1.0% (better)
Windows MinGW 386 fmtprintf-lto 2163712 B -2048 B / -0.1% (better) 451070 B -1452 B / -0.3% (better) 9.606 s +47.3 ms / +0.5% (worse) 10.175 ms -6.6 us / -0.1% (better)
Windows MinGW 386 println 90624 B -512 B / -0.6% (better) 19830 B -208 B / -1.0% (better) 1.077 s -24.16 ms / -2.2% (better) 8.834 ms +361.5 us / +4.3% (worse)
Windows MinGW 386 println-lto 69120 B -512 B / -0.7% (better) 17742 B -212 B / -1.2% (better) 1.295 s +25.58 ms / +2.0% (worse) 8.279 ms -601.2 us / -6.8% (better)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.474 s +4.19 ms / +0.3% (worse) 6.456 ms -115 us / -1.8% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.501 s +6.415 ms / +0.4% (worse) 6.359 ms -573.3 us / -8.3% (better)
Windows MinGW ARM64 fmtprintf 1777664 B -3072 B / -0.2% (better) 507556 B -3140 B / -0.6% (better) 4.330 s +119.2 ms / +2.8% (worse) 12.798 ms -12.7 us / -0.1% (better)
Windows MinGW ARM64 fmtprintf-lto 1855488 B -4608 B / -0.2% (better) 475164 B -3748 B / -0.8% (better) 9.902 s +93.35 ms / +1.0% (worse) 12.870 ms +275.5 us / +2.2% (worse)
Windows MinGW ARM64 println 68096 B -512 B / -0.7% (better) 22376 B -168 B / -0.7% (better) 1.475 s -3.077 ms / -0.2% (better) 11.615 ms -31 us / -0.3% (better)
Windows MinGW ARM64 println-lto 63488 B -512 B / -0.8% (better) 19444 B -160 B / -0.8% (better) 1.668 s +1.811 ms / +0.1% (worse) 11.046 ms -227.7 us / -2.0% (better)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65782 B 0 B / +0.0% 956.929 ms +34.65 ms / +3.8% (worse) 3.894 ms +521.8 us / +15.5% (worse)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65718 B 0 B / +0.0% 997.736 ms -150.2 ms / -13.1% (better) 3.484 ms -16.2 us / -0.5% (better)
Windows MSVC fmtprintf 1620480 B -5632 B / -0.3% (better) 691094 B -5536 B / -0.8% (better) 3.773 s +19.07 ms / +0.5% (worse) 9.519 ms +119.4 us / +1.3% (worse)
Windows MSVC fmtprintf-lto 1618432 B -6656 B / -0.4% (better) 647126 B -5760 B / -0.9% (better) 9.090 s -2.326 ms / -0.02558% (better) 9.860 ms +747.4 us / +8.2% (worse)
Windows MSVC println 192512 B -512 B / -0.3% (better) 119142 B -336 B / -0.3% (better) 964.949 ms +22.72 ms / +2.4% (worse) 7.676 ms +167.8 us / +2.2% (worse)
Windows MSVC println-lto 189952 B 0 B / +0.0% 116374 B -240 B / -0.2% (better) 1.147 s +26.62 ms / +2.4% (worse) 7.522 ms -168.4 us / -2.2% (better)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 942.858 ms -144.5 ms / -13.3% (better) 6.030 ms +359.5 us / +6.3% (worse)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 975.594 ms +21.79 ms / +2.3% (worse) 5.556 ms +105.6 us / +1.9% (worse)
Windows MSVC 386 fmtprintf 1189376 B -1536 B / -0.1% (better) 454469 B -1328 B / -0.3% (better) 3.926 s +66.33 ms / +1.7% (worse) 12.774 ms +499.2 us / +4.1% (worse)
Windows MSVC 386 fmtprintf-lto 1230336 B -1536 B / -0.1% (better) 429753 B -1664 B / -0.4% (better) 9.072 s +218.3 ms / +2.5% (worse) 11.584 ms -617.9 us / -5.1% (better)
Windows MSVC 386 println 34304 B 0 B / +0.0% 18673 B -192 B / -1.0% (better) 942.900 ms +26.96 ms / +2.9% (worse) 9.646 ms +423.7 us / +4.6% (worse)
Windows MSVC 386 println-lto 32256 B -512 B / -1.6% (better) 16855 B -160 B / -0.9% (better) 1.290 s +175.3 ms / +15.7% (worse) 10.999 ms +1.909 ms / +21.0% (worse)
Windows MSVC ARM64 cprintf 11264 B 0 B / +0.0% 3976 B 0 B / +0.0% 1.975 s -79.68 ms / -3.9% (better) 6.504 ms -709.6 us / -9.8% (better)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 3868 B 0 B / +0.0% 2.037 s +4.742 ms / +0.2% (worse) 6.870 ms +228.6 us / +3.4% (worse)
Windows MSVC ARM64 fmtprintf 1368576 B -3072 B / -0.2% (better) 508484 B -3040 B / -0.6% (better) 6.584 s -242.7 ms / -3.6% (better) 13.160 ms -995.6 us / -7.0% (better)
Windows MSVC ARM64 fmtprintf-lto 1391616 B -4096 B / -0.3% (better) 478004 B -3712 B / -0.8% (better) 15.566 s +311.5 ms / +2.0% (worse) 13.130 ms -1.867 ms / -12.5% (better)
Windows MSVC ARM64 println 41472 B 0 B / +0.0% 21624 B -160 B / -0.7% (better) 1.938 s -5.645 ms / -0.3% (better) 11.310 ms -639.5 us / -5.4% (better)
Windows MSVC ARM64 println-lto 39424 B -512 B / -1.3% (better) 19372 B -160 B / -0.8% (better) 2.254 s -15.29 ms / -0.7% (better) 11.551 ms +132.7 us / +1.2% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.610 ns/op +0.07 ns/op / +0.5% (worse)
Linux BenchmarkMergeCompilerFlags 208.700 ns/op +10.6 ns/op / +5.4% (worse)
Linux BenchmarkMergeLinkerFlags 127.700 ns/op +1.2 ns/op / +0.9% (worse)
Linux BenchmarkChannelBuffered 59.910 ns/op +3.67 ns/op / +6.5% (worse)
Linux BenchmarkChannelHandoff 14181 ns/op +62 ns/op / +0.4% (worse)
Linux BenchmarkDefer 56.570 ns/op +3.36 ns/op / +6.3% (worse)
Linux BenchmarkDirectCall 1.233 ns/op -0.361 ns/op / -22.6% (better)
Linux BenchmarkGlobalRead 1.179 ns/op +0.011 ns/op / +0.9% (worse)
Linux BenchmarkGlobalWrite 8.043 ns/op +0.24 ns/op / +3.1% (worse)
Linux BenchmarkGoroutine 22149 ns/op +1076 ns/op / +5.1% (worse)
Linux BenchmarkInterfaceCall 6.645 ns/op +0.293 ns/op / +4.6% (worse)
Linux BenchmarkRuntimeGetG 2.631 ns/op +0.181 ns/op / +7.4% (worse)
macOS BenchmarkLookupPCRandom 18.140 ns/op +0.32 ns/op / +1.8% (worse)
macOS BenchmarkMergeCompilerFlags 169 ns/op -5 ns/op / -2.9% (better)
macOS BenchmarkMergeLinkerFlags 109.800 ns/op +3.1 ns/op / +2.9% (worse)
macOS BenchmarkChannelBuffered 32.530 ns/op -14.64 ns/op / -31.0% (better)
macOS BenchmarkChannelHandoff 11945 ns/op +2309 ns/op / +24.0% (worse)
macOS BenchmarkDefer 50.210 ns/op -14.87 ns/op / -22.8% (better)
macOS BenchmarkDirectCall 1.320 ns/op -0.266 ns/op / -16.8% (better)
macOS BenchmarkGlobalRead 1.541 ns/op -0.13 ns/op / -7.8% (better)
macOS BenchmarkGlobalWrite 1.583 ns/op -0.839 ns/op / -34.6% (better)
macOS BenchmarkGoroutine 58078 ns/op +14207 ns/op / +32.4% (worse)
macOS BenchmarkInterfaceCall 4.754 ns/op -1.39 ns/op / -22.6% (better)
macOS BenchmarkRuntimeGetG 2.926 ns/op -0.069 ns/op / -2.3% (better)
Windows MinGW BenchmarkLookupPCRandom 13.150 ns/op -0.11 ns/op / -0.8% (better)
Windows MinGW BenchmarkMergeCompilerFlags 637 ns/op -24.3 ns/op / -3.7% (better)
Windows MinGW BenchmarkMergeLinkerFlags 530.700 ns/op -45.6 ns/op / -7.9% (better)
Windows MinGW BenchmarkChannelBuffered 29.780 ns/op -4.32 ns/op / -12.7% (better)
Windows MinGW BenchmarkChannelHandoff 1024 ns/op +156.1 ns/op / +18.0% (worse)
Windows MinGW BenchmarkDefer 58.650 ns/op +0.45 ns/op / +0.8% (worse)
Windows MinGW BenchmarkDirectCall 1.547 ns/op +0.001 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGlobalRead 1.549 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW BenchmarkGlobalWrite 2.469 ns/op -0.001 ns/op / -0.04049% (better)
Windows MinGW BenchmarkGoroutine 86627 ns/op -335 ns/op / -0.4% (better)
Windows MinGW BenchmarkInterfaceCall 8.998 ns/op +0.617 ns/op / +7.4% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.169 ns/op -0.008 ns/op / -0.4% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.490 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 734.600 ns/op -25.4 ns/op / -3.3% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 683 ns/op +0.5 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkChannelBuffered 39.620 ns/op -2.68 ns/op / -6.3% (better)
Windows MinGW 386 BenchmarkChannelHandoff 909 ns/op -29.4 ns/op / -3.1% (better)
Windows MinGW 386 BenchmarkDefer 42.070 ns/op +0.86 ns/op / +2.1% (worse)
Windows MinGW 386 BenchmarkDirectCall 1.547 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.860 ns/op 0 ns/op / +0.0%
Windows MinGW 386 BenchmarkGlobalWrite 7.781 ns/op +0.007 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGoroutine 82609 ns/op -1286 ns/op / -1.5% (better)
Windows MinGW 386 BenchmarkInterfaceCall 8.134 ns/op -0.237 ns/op / -2.8% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 2.167 ns/op -0.001 ns/op / -0.04613% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.080 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 576.400 ns/op +17.5 ns/op / +3.1% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 530.900 ns/op +2.7 ns/op / +0.5% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 39.720 ns/op -0.18 ns/op / -0.5% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 2253 ns/op +352 ns/op / +18.5% (worse)
Windows MinGW ARM64 BenchmarkDefer 56.200 ns/op +1.24 ns/op / +2.3% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.590 ns/op -0.0002 ns/op / -0.03392% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.737 ns/op +0.0737 ns/op / +11.1% (worse)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.663 ns/op -0.2206 ns/op / -25.0% (better)
Windows MinGW ARM64 BenchmarkGoroutine 60732 ns/op +2119 ns/op / +3.6% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.335 ns/op +0.028 ns/op / +0.7% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.770 ns/op -0.033 ns/op / -1.8% (better)
Windows MSVC BenchmarkLookupPCRandom 13.200 ns/op -0.01 ns/op / -0.1% (better)
Windows MSVC BenchmarkMergeCompilerFlags 597.400 ns/op -65.2 ns/op / -9.8% (better)
Windows MSVC BenchmarkMergeLinkerFlags 531.700 ns/op -30.1 ns/op / -5.4% (better)
Windows MSVC BenchmarkChannelBuffered 30.710 ns/op -2.06 ns/op / -6.3% (better)
Windows MSVC BenchmarkChannelHandoff 1151 ns/op -17 ns/op / -1.5% (better)
Windows MSVC BenchmarkDefer 55.560 ns/op -0.26 ns/op / -0.5% (better)
Windows MSVC BenchmarkDirectCall 1.547 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalRead 1.856 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalWrite 2.450 ns/op -0.014 ns/op / -0.6% (better)
Windows MSVC BenchmarkGoroutine 82707 ns/op -729 ns/op / -0.9% (better)
Windows MSVC BenchmarkInterfaceCall 8.994 ns/op +0.605 ns/op / +7.2% (worse)
Windows MSVC BenchmarkRuntimeGetG 2.479 ns/op +0.31 ns/op / +14.3% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 26.590 ns/op +0.06 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 766.200 ns/op +44.2 ns/op / +6.1% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 693 ns/op +7.2 ns/op / +1.0% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 39.190 ns/op -6.46 ns/op / -14.2% (better)
Windows MSVC 386 BenchmarkChannelHandoff 876.300 ns/op -110 ns/op / -11.2% (better)
Windows MSVC 386 BenchmarkDefer 47.720 ns/op +1.41 ns/op / +3.0% (worse)
Windows MSVC 386 BenchmarkDirectCall 1.547 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGlobalRead 1.859 ns/op +0.31 ns/op / +20.0% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 7.767 ns/op -0.027 ns/op / -0.3% (better)
Windows MSVC 386 BenchmarkGoroutine 89108 ns/op +1777 ns/op / +2.0% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 8.245 ns/op -0.151 ns/op / -1.8% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 2.488 ns/op +0.561 ns/op / +29.1% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.020 ns/op -0.04 ns/op / -0.3% (better)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 584 ns/op +12.2 ns/op / +2.1% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 558.300 ns/op +23.6 ns/op / +4.4% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 37.740 ns/op -1.28 ns/op / -3.3% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 2542 ns/op +49 ns/op / +2.0% (worse)
Windows MSVC ARM64 BenchmarkDefer 63.970 ns/op -0.78 ns/op / -1.2% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.589 ns/op -0.0005 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.666 ns/op -0.0002 ns/op / -0.03003% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.752 ns/op -0.006 ns/op / -0.2% (better)
Windows MSVC ARM64 BenchmarkGoroutine 55866 ns/op +380 ns/op / +0.7% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.259 ns/op +0.087 ns/op / +2.1% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 2.273 ns/op +0.479 ns/op / +26.7% (worse)

Timer runtime benchmarks

Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 912.100 ns/op +9.6 ns/op / +1.1% (worse)
Linux AfterFuncZeroDelivery/LLGo 31983 ns/op -219 ns/op / -0.7% (better)
Linux CreateStop/Go 291 ns/op -2.2 ns/op / -0.8% (better)
Linux CreateStop/LLGo 1385 ns/op -232 ns/op / -14.3% (better)
Linux RearmStopped/Go 115.300 ns/op -0.6 ns/op / -0.5% (better)
Linux RearmStopped/LLGo 1138 ns/op -224 ns/op / -16.4% (better)
Linux ResetActive/Go 67.640 ns/op -0.9 ns/op / -1.3% (better)
Linux ResetActive/LLGo 718.200 ns/op -33 ns/op / -4.4% (better)
Linux ResetHeap1024/Go 67.270 ns/op +0.1 ns/op / +0.1% (worse)
Linux ResetHeap1024/LLGo 181 ns/op -5.9 ns/op / -3.2% (better)
macOS AfterFuncZeroDelivery/Go 735.100 ns/op -72.5 ns/op / -9.0% (better)
macOS AfterFuncZeroDelivery/LLGo 124071 ns/op -2329 ns/op / -1.8% (better)
macOS CreateStop/Go 197.400 ns/op -7.6 ns/op / -3.7% (better)
macOS CreateStop/LLGo 683.800 ns/op -379.2 ns/op / -35.7% (better)
macOS RearmStopped/Go 80.440 ns/op -9.56 ns/op / -10.6% (better)
macOS RearmStopped/LLGo 545.400 ns/op -126.5 ns/op / -18.8% (better)
macOS ResetActive/Go 60.890 ns/op +2.53 ns/op / +4.3% (worse)
macOS ResetActive/LLGo 206.200 ns/op -89 ns/op / -30.1% (better)
macOS ResetHeap1024/Go 59.820 ns/op -5.2 ns/op / -8.0% (better)
macOS ResetHeap1024/LLGo 102 ns/op -16.8 ns/op / -14.1% (better)
Windows MinGW AfterFuncZeroDelivery/Go 556.300 ns/op +9.4 ns/op / +1.7% (worse)
Windows MinGW AfterFuncZeroDelivery/LLGo 163863 ns/op +2550 ns/op / +1.6% (worse)
Windows MinGW CreateStop/Go 116 ns/op +0.6 ns/op / +0.5% (worse)
Windows MinGW CreateStop/LLGo 434.500 ns/op +1 ns/op / +0.2% (worse)
Windows MinGW RearmStopped/Go 31.340 ns/op -0.25 ns/op / -0.8% (better)
Windows MinGW RearmStopped/LLGo 270.400 ns/op +0.7 ns/op / +0.3% (worse)
Windows MinGW ResetActive/Go 20.170 ns/op +0.05 ns/op / +0.2% (worse)
Windows MinGW ResetActive/LLGo 140.700 ns/op -13.6 ns/op / -8.8% (better)
Windows MinGW ResetHeap1024/Go 20.520 ns/op +0.04 ns/op / +0.2% (worse)
Windows MinGW ResetHeap1024/LLGo 126.500 ns/op -1.2 ns/op / -0.9% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 950.600 ns/op -1.6 ns/op / -0.2% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 182518 ns/op -4824 ns/op / -2.6% (better)
Windows MinGW 386 CreateStop/Go 190 ns/op -0.5 ns/op / -0.3% (better)
Windows MinGW 386 CreateStop/LLGo 1702 ns/op +38 ns/op / +2.3% (worse)
Windows MinGW 386 RearmStopped/Go 63.400 ns/op +0.11 ns/op / +0.2% (worse)
Windows MinGW 386 RearmStopped/LLGo 341.700 ns/op +5.2 ns/op / +1.5% (worse)
Windows MinGW 386 ResetActive/Go 38.920 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW 386 ResetActive/LLGo 992.100 ns/op +10.9 ns/op / +1.1% (worse)
Windows MinGW 386 ResetHeap1024/Go 39.260 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW 386 ResetHeap1024/LLGo 187.400 ns/op +1.3 ns/op / +0.7% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 688.100 ns/op +22.3 ns/op / +3.3% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 137921 ns/op +15343 ns/op / +12.5% (worse)
Windows MinGW ARM64 CreateStop/Go 202.800 ns/op +4 ns/op / +2.0% (worse)
Windows MinGW ARM64 CreateStop/LLGo 361.900 ns/op -2.6 ns/op / -0.7% (better)
Windows MinGW ARM64 RearmStopped/Go 70.540 ns/op -0.02 ns/op / -0.02834% (better)
Windows MinGW ARM64 RearmStopped/LLGo 251.100 ns/op -0.1 ns/op / -0.03981% (better)
Windows MinGW ARM64 ResetActive/Go 31.090 ns/op +0.01 ns/op / +0.03218% (worse)
Windows MinGW ARM64 ResetActive/LLGo 121.700 ns/op -7.1 ns/op / -5.5% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.120 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 127.800 ns/op +0.3 ns/op / +0.2% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 527.200 ns/op -27.3 ns/op / -4.9% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 155225 ns/op +2425 ns/op / +1.6% (worse)
Windows MSVC CreateStop/Go 116.400 ns/op +2.1 ns/op / +1.8% (worse)
Windows MSVC CreateStop/LLGo 417 ns/op -36.3 ns/op / -8.0% (better)
Windows MSVC RearmStopped/Go 31.760 ns/op +0.21 ns/op / +0.7% (worse)
Windows MSVC RearmStopped/LLGo 258.300 ns/op -0.2 ns/op / -0.1% (better)
Windows MSVC ResetActive/Go 20.110 ns/op +0.02 ns/op / +0.1% (worse)
Windows MSVC ResetActive/LLGo 149.800 ns/op +3.1 ns/op / +2.1% (worse)
Windows MSVC ResetHeap1024/Go 20.410 ns/op -0.24 ns/op / -1.2% (better)
Windows MSVC ResetHeap1024/LLGo 126.700 ns/op +1.3 ns/op / +1.0% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 955.700 ns/op -9.6 ns/op / -1.0% (better)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 189344 ns/op +5173 ns/op / +2.8% (worse)
Windows MSVC 386 CreateStop/Go 192.500 ns/op -2.4 ns/op / -1.2% (better)
Windows MSVC 386 CreateStop/LLGo 2030 ns/op +113 ns/op / +5.9% (worse)
Windows MSVC 386 RearmStopped/Go 63.710 ns/op +0.25 ns/op / +0.4% (worse)
Windows MSVC 386 RearmStopped/LLGo 337.100 ns/op +11.5 ns/op / +3.5% (worse)
Windows MSVC 386 ResetActive/Go 39.160 ns/op +0.15 ns/op / +0.4% (worse)
Windows MSVC 386 ResetActive/LLGo 1002 ns/op +2 ns/op / +0.2% (worse)
Windows MSVC 386 ResetHeap1024/Go 39.560 ns/op +0.23 ns/op / +0.6% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 171.900 ns/op -0.9 ns/op / -0.5% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 670.400 ns/op -2 ns/op / -0.3% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 137923 ns/op +14199 ns/op / +11.5% (worse)
Windows MSVC ARM64 CreateStop/Go 197.500 ns/op -2.1 ns/op / -1.1% (better)
Windows MSVC ARM64 CreateStop/LLGo 383.200 ns/op +7.2 ns/op / +1.9% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.600 ns/op -0.01 ns/op / -0.01416% (better)
Windows MSVC ARM64 RearmStopped/LLGo 273.800 ns/op +0.1 ns/op / +0.03654% (worse)
Windows MSVC ARM64 ResetActive/Go 30.980 ns/op -0.03 ns/op / -0.1% (better)
Windows MSVC ARM64 ResetActive/LLGo 138.600 ns/op +12.3 ns/op / +9.7% (worse)
Windows MSVC ARM64 ResetHeap1024/Go 31.020 ns/op -0.07 ns/op / -0.2% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 137.200 ns/op -0.1 ns/op / -0.1% (better)

Compared with 5c5874359c1e measured in the same runner job.

@cpunion
cpunion force-pushed the codex/fix-loop-local-stack-20260908 branch 2 times, most recently from 0da6c73 to 0401ed9 Compare September 13, 2026 02:29
@cpunion

cpunion commented Sep 14, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai Review again

@fennoai

fennoai Bot commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

FennoAI is processing your request, please wait a moment...

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: stabilize local stack slots and synthetic debug locations

Reviewed the locked range 410af8e9...826e07e5 with four subagents (code quality, performance, security, documentation accuracy). The change is well-constructed and thoroughly tested.

What holds up well

  • Alloca hoisting is correct. Reserving each non-escaping local's slot once in the entry block while keeping zeroinit/memset at the declaration preserves per-iteration zeroing semantics and fixes unbounded per-iteration stack growth for loop-body locals. Escaping locals still route to the heap path (AllocZ), so shared-slot aliasing cannot occur.
  • GC-roots interaction is safe. FirstBasicBlock() resolves to the gcroot entry block, so hoisted allocas land ahead of the terminator in a block dominating the body. Covered by roots=true in memory_local_test.go.
  • Debug-location mirroring is sound. The diLocation shadow field routes all mutations through setDebugLocation/restoreDebugLocation, correctly handling non-debug builders and builders created before DebugFunction attaches the DISubprogram. SetBlockEx reapplies the location at the right choke point.
  • Builder disposal on the Alloc path is leak-free (unconditional Dispose, no early return between create and dispose).

Minor findings (none blocking) are inline below.

Comment thread ssa/datastruct.go Outdated
Comment thread ssa/datastruct.go Outdated
Comment thread ssa/memory.go
Comment thread ssa/memory.go
@cpunion
cpunion force-pushed the codex/fix-loop-local-stack-20260908 branch from 826e07e to 726fc68 Compare September 14, 2026 05:30
@cpunion

cpunion commented Sep 14, 2026

Copy link
Copy Markdown
Collaborator Author

@fennoai Review again after the latest main rebase and review-comment updates.

@fennoai

fennoai Bot commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

FennoAI is processing your request, please wait a moment...

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review after main rebase + review-comment updates

Locked range 5c587435...726fc687. Re-ran all four subagents (code quality, performance, security, documentation).

What changed since my last review (826e07e5): only the two ssa/datastruct.go comments (channel Recv / Select) were reworded. The rest of the PR's own diff — ssa/memory.go, decl.go, di.go, eh.go, stmt_builder.go, tests, and FileCheck fixtures — is unchanged. The large delta versus the old head is the main rebase (overflow/cabi/caller_frame work), which is in both base and head and therefore outside this PR's diff.

Prior findings resolved

  • ✅ The Recv and Select comments now correctly state the receive buffers "use function-entry slots and remain valid across StackRestore" and describe the load-before-restore ordering as conservative/consistent rather than a correctness requirement. This matches the code: entry-block (static) allocas are not reclaimed by llvm.stackrestore. Both stale-comment findings from the previous review are addressed.

Re-verified clean

  • Alloca hoisting preserves per-iteration zero-init semantics; escaping locals still route to the heap path, so no cross-iteration aliasing.
  • StackSave/StackRestore lifetime for channel recv/select buffers is sound; the volatile load-before-restore is preserved, and selectOpsSlice correctly keeps its dynamic alloca stack-save scoped.
  • Builder disposal is leak-free; VerifyModule passes in the debug-enabled tests.

One optional, non-blocking performance nit remains inline. LGTM.

Comment thread ssa/memory.go
@cpunion

cpunion commented Sep 14, 2026

Copy link
Copy Markdown
Collaborator Author

Included the independently reviewed #2562 fix after the rebased LTO job timed out in TestLTOPluginRejectsMalformedInterfaceAttributes waiting for opt. All three cases pass locally with the built LLVM 22 plugin and -tags=dev (no skips). The dependency and failed-job link are documented in the body. @fennoai please review the updated head, especially the dependency integration.

@fennoai

fennoai Bot commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

FennoAI is processing your request, please wait a moment...

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review of the #2562 dependency integration

Locked range 5c587435...fc0ee507. Re-ran all four subagents (code quality, performance, security, documentation).

Delta since my last review (726fc687) is exactly the #2562 fix — two files:

  • ltoplugin/LLGOInterfaceMethodTypeIDPass.cpp (+3/-1)
  • cl/compile_test.go (+14/-3)

Everything else in the PR is byte-identical to the previously reviewed, clean state.

Assessment of the integration — looks correct

  1. invalidMetadata(): report_fatal_error → reportFatalUsageError. This is the right fix for the hang. reportFatalUsageError is LLVM's sanctioned API for invalid user input (it forwards to report_fatal_error(reason, /*GenCrashDiag=*/false)): it prints the message and exit(1)s with no crash trace, no symbolizer, and no "report an LLVM bug" prompt — which is exactly the diagnostic path that could stall opt before. Available in LLVM 21+ (repo targets LLVM 22), and llvm/Support/ErrorHandling.h is already included. The [[noreturn]] contract is preserved, so callers that rely on non-return after a validation branch (e.g. the Decl.Methods[*MethodIndex] index right after a bounds check) stay safe. The added comment accurately describes the behavior.

  2. TestLTOPluginRejectsMalformedInterfaceAttributes hardening. Wrapping opt in context.WithTimeout(30s) + cmd.WaitDelay = time.Second, failing loudly on ctx.Err(), and asserting a real *exec.ExitError with ExitCode() == 1 (not a crash/signal, not success) is a meaningfully stronger and hang-proof encoding of the fixed contract. The exact exit-1 assertion is appropriate here: the child is opt invoked directly (process-level exit code, not signal-translated), and the test is dev-gated behind the built LLVM plugin. The substring assertion matches the emitted "invalid interface type-id metadata:" prefix.

Safety / perf: The rejection path stays fail-closed and leak-free (no continued processing of invalid metadata, no attacker-controlled payload echoed in the message, deterministic termination). The change is off any compile hot path, and the timeout only costs anything on the failure path.

Non-blocking note (out of scope): the sibling ltoplugin/LLGOReflectMethodByNamePass.cpp still uses report_fatal_error at two sites, but those report internal invariant failures rather than malformed user input, so the crash-diag semantics are arguably correct there. Not part of this PR; flagging only for a possible future alignment.

No new inline findings. The dependency integration is clean and the earlier findings remain resolved. LGTM.

@cpunion

cpunion commented Sep 14, 2026

Copy link
Copy Markdown
Collaborator Author

The previously failing Dev LTO GlobalDCE and Go Method Drop job now passes on fc0ee50 after including #2562: https://github.com/xgo-dev/llgo/actions/runs/34812965081/job/103877868707 . The log shows the TestLTOPlugin* group completed successfully instead of hanging in malformed-metadata rejection. Local go test ./ssa -count=1 also passes (8.2s), and fennoai has reviewed the integrated head with no new findings. The rest of the CI matrix and ordinary coverage uploads are still running.

@cpunion

cpunion commented Sep 14, 2026

Copy link
Copy Markdown
Collaborator Author

Final validation on fc0ee50: all 64 checks passed, no pending or failed checks, and Codecov confirms all modified coverable lines are covered. The current head has a completed fennoai review with no unresolved threads. The rebased LTO failure is fixed by the explicit #2562 dependency, with the formerly failing job passing: https://github.com/xgo-dev/llgo/actions/runs/34812965081/job/103877868707 . The complete LLGo workflow is https://github.com/xgo-dev/llgo/actions/runs/34812965111 and the benchmark workflow is https://github.com/xgo-dev/llgo/actions/runs/34812965149 . Note that the macOS coverage job passed in 59m45s, close to its 60-minute limit; this run establishes correctness, not that CI duration has comfortable headroom. No timeout increase or retry behavior was added.

@cpunion
cpunion merged commit 5147277 into xgo-dev:main Sep 14, 2026
65 checks passed
@fennoai fennoai Bot mentioned this pull request Sep 16, 2026
cpunion added a commit that referenced this pull request Sep 17, 2026
Local Alloc hoists stack-slot reservations into the entry block so a
loop-local slot is reserved once per call (issue #80196, PR #2530). The
implementation created and disposed a fresh LLVM builder and re-derived
the entry-block insertion point on every Alloc. For functions with many
locals — e.g. the generated init in test/cmplxdivide.go, which builds a
~1000+-entry complex128 table — this turned local allocation into a
severe compile-time regression (~7x slower, timing out the GOROOT daily
Windows jobs; issue #2611).

Reserve the slots through a single builder cached on the function,
anchored at the entry block once and reused for every local Alloc, then
disposed in EndBuild. Allocas still land in the entry block (one slot per
call) with zeroing emitted at the declaration site each iteration, so the
issue #80196 fix and ssa/memory_local_test.go invariants are preserved.

Fixes #2611
cpunion added a commit that referenced this pull request Sep 18, 2026
- Add 90s timeouts and flake classification for fixedbugs/issue79186.go, fixedbugs/issue30041.go, and fixedbugs/issue39541.go on windows/386.
- Remove version: go1.27 constraint from index0.go timeout and flake entries so Windows 386 uses the 90s budget across all supported Go releases.
- Remove fixedbugs/issue80196.go from xfails: per-call stack-slot semantics for loops fixed the stack overflow in PR #2530 and PR #2613.
- Remove fixedbugs/issue75764.go on windows/amd64 from xfails: the deep interface tail-call test now passes reliably.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants