Skip to content

fix(runtime): preserve recovered panic source locations - #2567

Merged
xushiwei merged 8 commits into
xgo-dev:mainfrom
cpunion:codex/fix-panic-line-20260912
Sep 13, 2026
Merged

xushiwei merged 8 commits into
xgo-dev:mainfrom
cpunion:codex/fix-panic-line-20260912

Conversation

@cpunion

@cpunion cpunion commented Sep 12, 2026 •

Copy link
Copy Markdown
Collaborator

Dependency and targeted CI

Rebased onto #2575 at a7863ee05, which independently fixes a captured weak-cleanup reentrancy deadlock in the Qiniu Linux shard-0 test. The original five panic-location commits remain patch-identical after rebase. The GitHub base remains main because both contributions come from a fork; the dependency's changes will drop out of this diff once #2575 is merged. The complete affected Qiniu Ubuntu 24.04 large / LLVM 22 / Go 1.27 shard-0 job passed for this stack in run 34702086188: the first test phase completed in 136 seconds, symbol checks in 12 seconds, build-mode checks in 602 seconds, all later integration checks passed, and the new weak-runtime pressure regression passed three times in 0.16 seconds, without ptrace sampling. #2575 separately passed the same complete job plus its old-runtime negative control in run 34703427423. The ordinary matrix for 148f3d099 completed with 64 successful checks, no failure, and only the expected tag-gated release job skipped; its normal affected Qiniu job independently passed in 22m26s, and Codecov patch coverage is 100% (26/26). No diagnostic workflow or stress retry is added to this PR.

The current head cf807ffb1 rebases those same five patches onto #2575's source-comment update. git range-diff reports all five patches unchanged; compared with the previously validated 148f3d099, the source diff contains only weak-runtime comments. CI results above belong to the named earlier revisions, not a new full-matrix result for this rebased head.

Summary

  • Preserve the original panic PC snapshot when a recovering defer throws the same raw interface value again, matching the pattern used by testing.tRunner, while tying reuse to the still-active recover activation so a later panic of the same value gets a fresh traceback.
  • Keep nil-check failure blocks separate from the call-free success path, with a formal return that preserves the caller's saved return PC. Keep the noinline runtime helper's condition opaque to LTO using one empty matching-register constraint inside the helper, not barriers at every call site. Emit the source anchor inside the cold failure block so backend block placement cannot assign it a later statement's line.
  • Emit recover-observable pointer-store panic-site records on Linux and Darwin as well as Windows, and remove the now-passing fixedbugs/issue34123.go xfails.

Fixes #2566. Addresses the fixedbugs/issue17381.go and fixedbugs/issue27201.go regressions reported by #2564.

Cause

LLGo longjmp-unwinds the original panic frames and reconstructs them from a per-goroutine PC snapshot. A same-value panic(recover()) previously replaced that snapshot with the re-panic site, so go test failures stopped at testing.go:2123. Separately, the nil-check cold-block self-loop introduced by #2538 allowed LLVM to infer that a static nil-dereference caller never returns; this both exposed a return-PC/function-entry boundary misattribution in issue17381 and moved the recovered panic line in issue27201. Native pointer stores also lacked the scoped PC-line carrier already emitted on Windows, leaving issue34123 at the earlier pointer-load line.

The two recent regressions were traced to #2538 commit cdaa60ff7: both original GOROOT programs pass with its parent 88a698364 and fail with that commit on macOS/arm64 with Go 1.27. Same-value repanic retention is a longer-standing omission in the snapshot design introduced by #2026, not a new #2538 regression. The pointer-store metadata was added during #2405 development and then restricted to Windows before merging, leaving the existing Linux/macOS gap unresolved; the former Windows-only runtime assertion is now enabled on every platform.

Size and performance

The original version added 80 bytes of symbol-index records to cprintf for five new runtime helpers and approximately 11 KiB of macOS machine code to non-LTO fmtprintf by joining nil-check failures back into the normal path. This revision removes those five helpers, moves snapshot bookkeeping behind the existing public-runtime hook so tiny programs do not link it, reuses the existing recover activation token, and leaves recoverState at its original two-pointer ABI. The snapshot retains three words of per-goroutine state without allocating a separate recovery object. Returning from a defer or transparent wrapper and Goexit clear the retained value.

Local paired builds of base daada50d2270 and the updated source used the same checkout path, compiler options, and macOS/arm64 host. The code column below is the Mach-O __text section; file sizes include alignment and all metadata.

Workload File size Change vs base Machine-code change vs base
cprintf 84,480 B 0 B 0 B
cprintf, full LTO 84,288 B 0 B 0 B
println 114,848 B -80 B -108 B
println, full LTO 118,736 B 0 B 0 B
fmtprintf 1,473,264 B 0 B -1,184 B
fmtprintf, full LTO 1,159,440 B +16 B +5,664 B

The remaining full-LTO code increase is explicitly not zero: keeping the panic helper potentially returning prevents return-PC boundary misattribution even after whole-program optimization. The protection is confined to the helper, emits no assembly instructions itself, and adds no runtime call to successful nil checks. CI's paired benchmarks remain the cross-platform size and timing acceptance check; local timing samples were too sensitive to host load to support a speedup claim.

The paired CI measurements for 5daef74c420b now confirm no file or Text growth for cprintf or cprintf-lto on any native platform; println and println-lto are unchanged or smaller. Non-LTO fmtprintf file deltas range from -3,584 B to +512 B, while full-LTO Text deltas are +3,104 B to +5,632 B and full-LTO file deltas are 0 B to +6,144 B. The report's Text metric is distinct from the local Mach-O __text measurement above. The broad macOS timing slowdown also affects unchanged Go timer controls, so this single run does not isolate an LLGo performance regression or establish a speedup. The LTO size cost is a limitation of this approach, not a proven lower bound.

Validation

  • Ordinary PR CI covers all four scenarios through the existing test/go suite, compatible with both go test and llgo test. TestRuntimeStatementLineInfo/panic_caller checks raw runtime.Callers PCs still resolve to the recovering caller after a statically panicking leaf (issue17381); load_panic_line requires the nil load's exact line rather than the following observable statement (issue27201); store_panic_line requires the pointer store's line rather than the earlier pointer load (issue34123) on every platform. TestCallerRepanicTraceback uses nested testing.T.Run to reproduce llcppg's same-value repanic. These extend the existing test binary without additional build-driver tests, GOROOT subprocess builds, workflow changes, nightly opt-in, or xfail allowances.
  • The strengthened test/go cases pass with host Go and LLGo default/full LTO. A negative control using the unmodified cdaa60ff7 compiler and its runtime fails all four scenarios: the caller is missing, the nil load is reported two lines late, the store two lines early, and the nested test panic loses callerRepanicOrigin. These are behavioral checks, not merely compilation checks; the added in-process subtests take less than the test logger's 0.01-second resolution locally.
  • go test ./ssa -run '^TestAssertNilDeref(ZeroExprNoPanic|ColdCall)$' -count=1
  • go test ./cl -run '^TestCompileRuntimeCaller(StorePanicPCLineMetadataOnNativeTargets|PanicPCLineMetadata)$' -count=1
  • go test ./test/go -run '^(TestCallerRepanicTraceback|TestCallerPanicTraceback|TestRuntimeStatementLineInfo)$' -count=1
  • llgo test -v -run '^(TestCallerRepanicTraceback|TestCallerPanicTraceback|TestRuntimeStatementLineInfo)$' ./test/go
  • Go 1.27 GOROOT fixedbugs/issue17381.go, fixedbugs/issue27201.go, and fixedbugs/issue34123.go pass on macOS/arm64 with both default compilation and full LTO, without xfail allowances.
  • Full llgo test -timeout=5m ./test/go, plus full-LTO caller/recover/statement-line tests; repanic coverage includes transparent method wrappers, indirect recover attempts, uncomparable slice values, replacement panics, and later reuse of the same value.
  • go test ./ssa ./cl -skip '^TestRunAndTestFrom' -timeout=5m -count=1; targeted SSA tests model whole-program argument propagation and emit objects for Linux/amd64, Darwin/arm64, and Windows/amd64, 386, and arm64, while retaining the original WebAssembly failure path.
  • The goplus/llcppg@dev-panic-at reproduction now prints the complete application stack instead of stopping at testing.tRunner; its innermost frame is the exact nil access at go/ast/ast.go:520, followed by gogen and llcppg compile_test.go:49 and compile_test.go:77.

The open #2530 debug-location work was tested on current main and does not fix either GOROOT regression, so this PR does not depend on it.

A separate full local go test ./cl exceeded its five-minute whole-suite budget while advancing through native run fixtures; this is not listed as a full local pass. CI covers the complete fixture suites.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: fix panic line (recovered/re-panicked traceback)

The core fix is sound and well-tested. Removing the Windows-only guard around recordPanicSite is correctly scoped downstream (recordPanicSite only emits the PC-line label when p.panicSiteFuncs[p.goFn] is set), so it adds no per-store IR for programs that don't observe recovered-panic stacks. The recoverActive + interface-identity matching in takeRecovered is memory-safe: recoveredData is a real unsafe.Pointer field embedded in the GC-visible g, and it is only ever compared for identity, never dereferenced. The AssertNilDeref block rewiring to the continuation edge is correct.

One issue must be addressed before merge, plus two minor notes below (inline).

[P1] Removing the issue34123 xfail entries breaks an existing unit test

test/goroot/xfail.yaml drops the linux/amd64 and darwin/arm64 fixedbugs/issue34123.go entries, but test/goroot/runner_unit_test.go:362 still asserts that this case must match an xfail entry on linux/amd64:

{version: "go1.27.0", platform: "linux/amd64", tc: testCase{RelPath: "fixedbugs/issue34123.go", Directive: "run"}},

TestObservedFailuresHaveXFailClassifications calls cfg.Match(...) and t.Errorf when !match. With the yaml entry gone, Match returns false, so this test will now fail. That line is not part of this PR's diff, so the removal leaves the test suite inconsistent. Remove the stale assertion at runner_unit_test.go:362 as part of this change. (The complementary windows-msvc/amd64 "must not match" assertion at line 388 stays consistent with the removal.)

Comment thread runtime/internal/runtime/z_rt.go Outdated
Comment thread ssa/memory.go Outdated
Comment thread runtime/internal/runtime/caller.go Outdated
@cpunion

cpunion commented Sep 12, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed the P1 consistency issue in 2601197: removed the stale Linux observed-failure fixture for fixedbugs/issue34123.go and added Linux/macOS observed-pass assertions so the xfail removal remains covered. The focused GOROOT runner tests pass locally. The first compatibility CI run failed only on that stale fixture for each Go version; a fresh run is now in progress.

@codecov

codecov Bot commented Sep 12, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Sep 12, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

cf807ffb10ec | workflow run | long-term charts

WebAssembly output sizes

Profile and compiler Wasm module vs base Generated JS glue vs base
ec32/LLGo 113450 B -96 B / -0.1% (better) 70736 B 0 B / +0.0%
ec64/LLGo 118870 B -131 B / -0.1% (better) 74033 B 0 B / +0.0%
js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
js/LLGo 66763 B +65 B / +0.1% (worse) 68511 B 0 B / +0.0%
wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
wasip1/LLGo 73012 B +75 B / +0.1% (worse) 0 B 0 B / 0.0%
wc32/LLGo 117437 B -100 B / -0.1% (better) 0 B 0 B / 0.0%

LLGo WebAssembly build measurements

Profile Build vs base
ec32 5.331 s -224 ms / -4.0% (better)
ec64 5.016 s +68.06 ms / +1.4% (worse)
js 4.514 s +88.5 ms / +2.0% (worse)
wasip1 3.065 s +101.4 ms / +3.4% (worse)
wc32 3.748 s -80.63 ms / -2.1% (better)

Compared with b07cd12ad1ac measured in the same runner job.

@github-actions

github-actions Bot commented Sep 12, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

cf807ffb10ec | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 19848 B 0 B / +0.0% 387 B 0 B / +0.0% 445.585 ms -4.767 ms / -1.1% (better) 1.254 ms +13.8 us / +1.1% (worse)
Linux cprintf-lto 19600 B 0 B / +0.0% 368 B 0 B / +0.0% 454.899 ms +1.302 ms / +0.3% (worse) 1.235 ms -19.61 us / -1.6% (better)
Linux fmtprintf 1626368 B -2040 B / -0.1% (better) 498245 B +203 B / +0.04076% (worse) 3.559 s -36.48 ms / -1.0% (better) 3.008 ms +58.73 us / +2.0% (worse)
Linux fmtprintf-lto 1478008 B +3904 B / +0.3% (worse) 438417 B +3879 B / +0.9% (worse) 10.514 s -5.336 ms / -0.1% (better) 2.873 ms -46.19 us / -1.6% (better)
Linux println 62688 B -280 B / -0.4% (better) 15057 B -113 B / -0.7% (better) 438.989 ms +533.8 us / +0.1% (worse) 1.559 ms +55.25 us / +3.7% (worse)
Linux println-lto 54312 B -16 B / -0.02945% (better) 12356 B -15 B / -0.1% (better) 648.804 ms +2.198 ms / +0.3% (worse) 1.547 ms -13.46 us / -0.9% (better)
macOS cprintf 84480 B 0 B / +0.0% 17117 B 0 B / +0.0% 889.482 ms +168.1 ms / +23.3% (worse) 4.273 ms -89.83 us / -2.1% (better)
macOS cprintf-lto 84288 B 0 B / +0.0% 12881 B 0 B / +0.0% 877.557 ms +186.3 ms / +27.0% (worse) 5.549 ms +1.461 ms / +35.7% (worse)
macOS fmtprintf 1473264 B 0 B / +0.0% 873360 B -1056 B / -0.1% (better) 3.836 s -159.1 ms / -4.0% (better) 8.372 ms +3.458 ms / +70.4% (worse)
macOS fmtprintf-lto 1159440 B +16 B / +0.00138% (worse) 847332 B +5756 B / +0.7% (worse) 11.808 s +706 ms / +6.4% (worse) 7.781 ms -2.467 ms / -24.1% (better)
macOS println 114848 B -80 B / -0.1% (better) 35120 B -108 B / -0.3% (better) 774.792 ms -62.76 ms / -7.5% (better) 3.989 ms -2.21 ms / -35.7% (better)
macOS println-lto 118736 B 0 B / +0.0% 32714 B 0 B / +0.0% 791.539 ms -149.5 ms / -15.9% (better) 6.256 ms +2.52 ms / +67.5% (worse)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.113 s -8.195 ms / -0.7% (better) 3.414 ms -52 us / -1.5% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.139 s -3.917 ms / -0.3% (better) 3.460 ms +56.6 us / +1.7% (worse)
Windows MinGW fmtprintf 1896448 B -2048 B / -0.1% (better) 601622 B -400 B / -0.1% (better) 4.073 s +62.62 ms / +1.6% (worse) 8.048 ms +49.2 us / +0.6% (worse)
Windows MinGW fmtprintf-lto 1937920 B +3584 B / +0.2% (worse) 550998 B +3776 B / +0.7% (worse) 10.095 s +38.78 ms / +0.4% (worse) 8.233 ms +113.5 us / +1.4% (worse)
Windows MinGW println 72192 B 0 B / +0.0% 24278 B -80 B / -0.3% (better) 1.088 s -32.61 ms / -2.9% (better) 6.406 ms -247.9 us / -3.7% (better)
Windows MinGW println-lto 65536 B 0 B / +0.0% 20694 B -16 B / -0.1% (better) 1.326 s +12.65 ms / +1.0% (worse) 6.347 ms -215.9 us / -3.3% (better)
Windows MinGW 386 cprintf 42496 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.114 s +21.65 ms / +2.0% (worse) 5.442 ms -120.2 us / -2.2% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.151 s +13.03 ms / +1.1% (worse) 5.183 ms -84.3 us / -1.6% (better)
Windows MinGW 386 fmtprintf 1861632 B 0 B / +0.0% 474094 B +288 B / +0.1% (worse) 4.156 s +28.8 ms / +0.7% (worse) 17.039 ms -2.449 ms / -12.6% (better)
Windows MinGW 386 fmtprintf-lto 2169856 B +6656 B / +0.3% (worse) 454406 B +4548 B / +1.0% (worse) 9.721 s +156.4 ms / +1.6% (worse) 17.664 ms -646.7 us / -3.5% (better)
Windows MinGW 386 println 91648 B 0 B / +0.0% 20310 B -112 B / -0.5% (better) 1.096 s +3.903 ms / +0.4% (worse) 8.525 ms -484.1 us / -5.4% (better)
Windows MinGW 386 println-lto 70144 B -512 B / -0.7% (better) 18258 B -16 B / -0.1% (better) 1.304 s -44.33 ms / -3.3% (better) 9.380 ms -24.6 us / -0.3% (better)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.382 s -21.58 ms / -1.5% (better) 6.484 ms +86.7 us / +1.4% (worse)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.407 s -191.7 ms / -12.0% (better) 5.936 ms -263.4 us / -4.2% (better)
Windows MinGW ARM64 fmtprintf 1782784 B -3584 B / -0.2% (better) 512624 B -516 B / -0.1% (better) 4.167 s +20.82 ms / +0.5% (worse) 12.075 ms -515.7 us / -4.1% (better)
Windows MinGW ARM64 fmtprintf-lto 1862144 B 0 B / +0.0% 481216 B +3112 B / +0.7% (worse) 9.789 s -73.5 ms / -0.7% (better) 12.095 ms -984.3 us / -7.5% (better)
Windows MinGW ARM64 println 68608 B -512 B / -0.7% (better) 22944 B -108 B / -0.5% (better) 1.379 s -30.92 ms / -2.2% (better) 10.713 ms -130 us / -1.2% (better)
Windows MinGW ARM64 println-lto 65024 B 0 B / +0.0% 20184 B 0 B / +0.0% 1.569 s -10.72 ms / -0.7% (better) 14.797 ms +3.901 ms / +35.8% (worse)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65782 B 0 B / +0.0% 878.530 ms -32.72 ms / -3.6% (better) 3.384 ms -83.3 us / -2.4% (better)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65718 B 0 B / +0.0% 903.063 ms -7.247 ms / -0.8% (better) 3.393 ms -1.1 us / -0.03241% (better)
Windows MSVC fmtprintf 1626624 B -2560 B / -0.2% (better) 697142 B -400 B / -0.1% (better) 3.624 s -119.9 ms / -3.2% (better) 9.137 ms -591.7 us / -6.1% (better)
Windows MSVC fmtprintf-lto 1625088 B +2560 B / +0.2% (worse) 653302 B +3184 B / +0.5% (worse) 9.085 s -148.6 ms / -1.6% (better) 9.825 ms +11.4 us / +0.1% (worse)
Windows MSVC println 193024 B 0 B / +0.0% 119686 B -96 B / -0.1% (better) 889.942 ms -43.57 ms / -4.7% (better) 7.101 ms -395.2 us / -5.3% (better)
Windows MSVC println-lto 189952 B 0 B / +0.0% 116630 B -16 B / -0.01372% (better) 1.093 s -11.11 ms / -1.0% (better) 7.172 ms -259.4 us / -3.5% (better)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.011 s +35.76 ms / +3.7% (worse) 6.957 ms +186.8 us / +2.8% (worse)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.053 s +41.14 ms / +4.1% (worse) 6.519 ms -255.5 us / -3.8% (better)
Windows MSVC 386 fmtprintf 1192960 B +1024 B / +0.1% (worse) 457445 B +272 B / +0.1% (worse) 3.783 s -96.52 ms / -2.5% (better) 12.186 ms -375.9 us / -3.0% (better)
Windows MSVC 386 fmtprintf-lto 1233408 B +4608 B / +0.4% (worse) 433353 B +4240 B / +1.0% (worse) 8.584 s -793.5 ms / -8.5% (better) 11.749 ms -1.352 ms / -10.3% (better)
Windows MSVC 386 println 34816 B 0 B / +0.0% 19169 B -96 B / -0.5% (better) 1.139 s +156 ms / +15.9% (worse) 10.027 ms +578.6 us / +6.1% (worse)
Windows MSVC 386 println-lto 32768 B 0 B / +0.0% 17335 B -16 B / -0.1% (better) 1.127 s -72.46 ms / -6.0% (better) 9.441 ms -1.343 ms / -12.5% (better)
Windows MSVC ARM64 cprintf 11264 B 0 B / +0.0% 3976 B 0 B / +0.0% 2.147 s -66.69 ms / -3.0% (better) 7.817 ms -572.4 us / -6.8% (better)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 3868 B 0 B / +0.0% 2.185 s -56.63 ms / -2.5% (better) 7.969 ms -208.4 us / -2.5% (better)
Windows MSVC ARM64 fmtprintf 1373184 B -3584 B / -0.3% (better) 513460 B -576 B / -0.1% (better) 6.774 s +106 ms / +1.6% (worse) 16.257 ms +577.9 us / +3.7% (worse)
Windows MSVC ARM64 fmtprintf-lto 1397760 B 0 B / +0.0% 483972 B +3200 B / +0.7% (worse) 16.713 s -103.8 ms / -0.6% (better) 15.346 ms -1.205 ms / -7.3% (better)
Windows MSVC ARM64 println 41984 B 0 B / +0.0% 22184 B -128 B / -0.6% (better) 2.148 s -93.41 ms / -4.2% (better) 13.667 ms -756.3 us / -5.2% (better)
Windows MSVC ARM64 println-lto 40448 B 0 B / +0.0% 20092 B 0 B / +0.0% 2.328 s -258 ms / -10.0% (better) 11.227 ms -3.551 ms / -24.0% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.680 ns/op -0.04 ns/op / -0.3% (better)
Linux BenchmarkMergeCompilerFlags 197.600 ns/op -7.3 ns/op / -3.6% (better)
Linux BenchmarkMergeLinkerFlags 137.500 ns/op +7.4 ns/op / +5.7% (worse)
Linux BenchmarkChannelBuffered 63.700 ns/op -6.1 ns/op / -8.7% (better)
Linux BenchmarkChannelHandoff 12776 ns/op -953 ns/op / -6.9% (better)
Linux BenchmarkDefer 47.840 ns/op +2.33 ns/op / +5.1% (worse)
Linux BenchmarkDirectCall 1.163 ns/op -0.775 ns/op / -40.0% (better)
Linux BenchmarkGlobalRead 1.170 ns/op +0.007 ns/op / +0.6% (worse)
Linux BenchmarkGlobalWrite 7.757 ns/op +0.001 ns/op / +0.01289% (worse)
Linux BenchmarkGoroutine 20786 ns/op -225 ns/op / -1.1% (better)
Linux BenchmarkInterfaceCall 5.852 ns/op -0.038 ns/op / -0.6% (better)
Linux BenchmarkRuntimeGetG 2.655 ns/op +0.211 ns/op / +8.6% (worse)
macOS BenchmarkLookupPCRandom 17.240 ns/op -1.29 ns/op / -7.0% (better)
macOS BenchmarkMergeCompilerFlags 173 ns/op +15.6 ns/op / +9.9% (worse)
macOS BenchmarkMergeLinkerFlags 123 ns/op +17.5 ns/op / +16.6% (worse)
macOS BenchmarkChannelBuffered 33.750 ns/op +6.87 ns/op / +25.6% (worse)
macOS BenchmarkChannelHandoff 11914 ns/op +1919 ns/op / +19.2% (worse)
macOS BenchmarkDefer 49.330 ns/op +12.57 ns/op / +34.2% (worse)
macOS BenchmarkDirectCall 1.287 ns/op -0.006 ns/op / -0.5% (better)
macOS BenchmarkGlobalRead 1.453 ns/op +0.227 ns/op / +18.5% (worse)
macOS BenchmarkGlobalWrite 1.540 ns/op +0.286 ns/op / +22.8% (worse)
macOS BenchmarkGoroutine 51549 ns/op +2660 ns/op / +5.4% (worse)
macOS BenchmarkInterfaceCall 5.166 ns/op +0.515 ns/op / +11.1% (worse)
macOS BenchmarkRuntimeGetG 2.830 ns/op +0.042 ns/op / +1.5% (worse)
Windows MinGW BenchmarkLookupPCRandom 13.210 ns/op -0.04 ns/op / -0.3% (better)
Windows MinGW BenchmarkMergeCompilerFlags 611.300 ns/op +2.5 ns/op / +0.4% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 534.700 ns/op +14.8 ns/op / +2.8% (worse)
Windows MinGW BenchmarkChannelBuffered 38.220 ns/op +1.54 ns/op / +4.2% (worse)
Windows MinGW BenchmarkChannelHandoff 1016 ns/op +15 ns/op / +1.5% (worse)
Windows MinGW BenchmarkDefer 56.840 ns/op -1.45 ns/op / -2.5% (better)
Windows MinGW BenchmarkDirectCall 1.547 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW BenchmarkGlobalRead 1.546 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW BenchmarkGlobalWrite 2.471 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGoroutine 86667 ns/op -4253 ns/op / -4.7% (better)
Windows MinGW BenchmarkInterfaceCall 9.029 ns/op +0.965 ns/op / +12.0% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.169 ns/op -0.31 ns/op / -12.5% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 68.550 ns/op -0.22 ns/op / -0.3% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 837.600 ns/op -51.5 ns/op / -5.8% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 776.900 ns/op -65.4 ns/op / -7.8% (better)
Windows MinGW 386 BenchmarkChannelBuffered 56.610 ns/op -4.87 ns/op / -7.9% (better)
Windows MinGW 386 BenchmarkChannelHandoff 2900 ns/op +326 ns/op / +12.7% (worse)
Windows MinGW 386 BenchmarkDefer 52.540 ns/op -1.85 ns/op / -3.4% (better)
Windows MinGW 386 BenchmarkDirectCall 0.865 ns/op -0.0024 ns/op / -0.3% (better)
Windows MinGW 386 BenchmarkGlobalRead 0.868 ns/op +0.001 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGlobalWrite 16.180 ns/op +0.02 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGoroutine 83019 ns/op +2656 ns/op / +3.3% (worse)
Windows MinGW 386 BenchmarkInterfaceCall 6.561 ns/op -0.221 ns/op / -3.3% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 2.007 ns/op +0.027 ns/op / +1.4% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 11.970 ns/op -0.03 ns/op / -0.2% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 573.500 ns/op -31.9 ns/op / -5.3% (better)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 532.400 ns/op -36.3 ns/op / -6.4% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 42.710 ns/op -4.14 ns/op / -8.8% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 1886 ns/op +95 ns/op / +5.3% (worse)
Windows MinGW ARM64 BenchmarkDefer 52.940 ns/op -1.13 ns/op / -2.1% (better)
Windows MinGW ARM64 BenchmarkDirectCall 0.589 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGlobalRead 0.663 ns/op -0.221 ns/op / -25.0% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.664 ns/op +0.0003 ns/op / +0.04523% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 54873 ns/op -1886 ns/op / -3.3% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.321 ns/op +0.18 ns/op / +4.3% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.770 ns/op 0 ns/op / +0.0%
Windows MSVC BenchmarkLookupPCRandom 12.420 ns/op -0.05 ns/op / -0.4% (better)
Windows MSVC BenchmarkMergeCompilerFlags 524 ns/op -17.9 ns/op / -3.3% (better)
Windows MSVC BenchmarkMergeLinkerFlags 459.700 ns/op -4.1 ns/op / -0.9% (better)
Windows MSVC BenchmarkChannelBuffered 35.550 ns/op -4.15 ns/op / -10.5% (better)
Windows MSVC BenchmarkChannelHandoff 1320 ns/op -99 ns/op / -7.0% (better)
Windows MSVC BenchmarkDefer 54.250 ns/op -0.56 ns/op / -1.0% (better)
Windows MSVC BenchmarkDirectCall 1.747 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalRead 1.747 ns/op 0 ns/op / +0.0%
Windows MSVC BenchmarkGlobalWrite 2.788 ns/op +0.001 ns/op / +0.03588% (worse)
Windows MSVC BenchmarkGoroutine 64009 ns/op -4268 ns/op / -6.3% (better)
Windows MSVC BenchmarkInterfaceCall 8.735 ns/op -0.083 ns/op / -0.9% (better)
Windows MSVC BenchmarkRuntimeGetG 2.105 ns/op +0.288 ns/op / +15.9% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 26.500 ns/op -0.12 ns/op / -0.5% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 748.600 ns/op -40.1 ns/op / -5.1% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 687.500 ns/op -12.8 ns/op / -1.8% (better)
Windows MSVC 386 BenchmarkChannelBuffered 41.140 ns/op -2.01 ns/op / -4.7% (better)
Windows MSVC 386 BenchmarkChannelHandoff 942.400 ns/op -13.8 ns/op / -1.4% (better)
Windows MSVC 386 BenchmarkDefer 45.220 ns/op -5.03 ns/op / -10.0% (better)
Windows MSVC 386 BenchmarkDirectCall 1.548 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGlobalRead 1.863 ns/op +0.315 ns/op / +20.3% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 7.782 ns/op -0.007 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGoroutine 83743 ns/op -2154 ns/op / -2.5% (better)
Windows MSVC 386 BenchmarkInterfaceCall 8.357 ns/op -0.005 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 1.926 ns/op -0.245 ns/op / -11.3% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.070 ns/op +0.02 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 565.500 ns/op +0.9 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 533.900 ns/op -1.3 ns/op / -0.2% (better)
Windows MSVC ARM64 BenchmarkChannelBuffered 40.580 ns/op -5.95 ns/op / -12.8% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 2036 ns/op -326 ns/op / -13.8% (better)
Windows MSVC ARM64 BenchmarkDefer 68.780 ns/op +2.97 ns/op / +4.5% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op -0.0001 ns/op / -0.01695% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.664 ns/op +0.074 ns/op / +12.5% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.755 ns/op -0.004 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkGoroutine 54192 ns/op +33 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.138 ns/op -0.002 ns/op / -0.04831% (better)
Windows MSVC ARM64 BenchmarkRuntimeGetG 2.109 ns/op +0.321 ns/op / +18.0% (worse)

Timer runtime benchmarks

Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 905.300 ns/op 0 ns/op / +0.0%
Linux AfterFuncZeroDelivery/LLGo 34193 ns/op -262 ns/op / -0.8% (better)
Linux CreateStop/Go 289.500 ns/op +1 ns/op / +0.3% (worse)
Linux CreateStop/LLGo 1685 ns/op +81 ns/op / +5.0% (worse)
Linux RearmStopped/Go 115.800 ns/op +0.1 ns/op / +0.1% (worse)
Linux RearmStopped/LLGo 1273 ns/op -180 ns/op / -12.4% (better)
Linux ResetActive/Go 68.510 ns/op -0.02 ns/op / -0.02918% (better)
Linux ResetActive/LLGo 710.400 ns/op -133.1 ns/op / -15.8% (better)
Linux ResetHeap1024/Go 67.060 ns/op +0.02 ns/op / +0.02983% (worse)
Linux ResetHeap1024/LLGo 177.400 ns/op -3.1 ns/op / -1.7% (better)
macOS AfterFuncZeroDelivery/Go 642.600 ns/op +7 ns/op / +1.1% (worse)
macOS AfterFuncZeroDelivery/LLGo 125005 ns/op +25069 ns/op / +25.1% (worse)
macOS CreateStop/Go 201.600 ns/op -49.2 ns/op / -19.6% (better)
macOS CreateStop/LLGo 1130 ns/op +391 ns/op / +52.9% (worse)
macOS RearmStopped/Go 79.890 ns/op -1.03 ns/op / -1.3% (better)
macOS RearmStopped/LLGo 723 ns/op +356.2 ns/op / +97.1% (worse)
macOS ResetActive/Go 55.120 ns/op -4.64 ns/op / -7.8% (better)
macOS ResetActive/LLGo 254.900 ns/op +73.8 ns/op / +40.8% (worse)
macOS ResetHeap1024/Go 58.650 ns/op +2.23 ns/op / +4.0% (worse)
macOS ResetHeap1024/LLGo 98.240 ns/op +4.28 ns/op / +4.6% (worse)
Windows MinGW AfterFuncZeroDelivery/Go 560.300 ns/op +3.7 ns/op / +0.7% (worse)
Windows MinGW AfterFuncZeroDelivery/LLGo 165676 ns/op -3216 ns/op / -1.9% (better)
Windows MinGW CreateStop/Go 122.100 ns/op +0.6 ns/op / +0.5% (worse)
Windows MinGW CreateStop/LLGo 446.700 ns/op +23.7 ns/op / +5.6% (worse)
Windows MinGW RearmStopped/Go 31.390 ns/op -0.07 ns/op / -0.2% (better)
Windows MinGW RearmStopped/LLGo 263.700 ns/op -6.4 ns/op / -2.4% (better)
Windows MinGW ResetActive/Go 20.080 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW ResetActive/LLGo 160.200 ns/op +10.6 ns/op / +7.1% (worse)
Windows MinGW ResetHeap1024/Go 20.330 ns/op -0.14 ns/op / -0.7% (better)
Windows MinGW ResetHeap1024/LLGo 129.800 ns/op +3.7 ns/op / +2.9% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 1120 ns/op +1 ns/op / +0.1% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 402297 ns/op +9357 ns/op / +2.4% (worse)
Windows MinGW 386 CreateStop/Go 273.400 ns/op +4.4 ns/op / +1.6% (worse)
Windows MinGW 386 CreateStop/LLGo 14280 ns/op -5 ns/op / -0.035% (better)
Windows MinGW 386 RearmStopped/Go 91.690 ns/op -0.33 ns/op / -0.4% (better)
Windows MinGW 386 RearmStopped/LLGo 416.900 ns/op +26.7 ns/op / +6.8% (worse)
Windows MinGW 386 ResetActive/Go 45.880 ns/op -0.28 ns/op / -0.6% (better)
Windows MinGW 386 ResetActive/LLGo 310.900 ns/op -155.2 ns/op / -33.3% (better)
Windows MinGW 386 ResetHeap1024/Go 46.330 ns/op +0.05 ns/op / +0.1% (worse)
Windows MinGW 386 ResetHeap1024/LLGo 182.400 ns/op -4.8 ns/op / -2.6% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 686.500 ns/op +11 ns/op / +1.6% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 124163 ns/op -4309 ns/op / -3.4% (better)
Windows MinGW ARM64 CreateStop/Go 196.800 ns/op -3 ns/op / -1.5% (better)
Windows MinGW ARM64 CreateStop/LLGo 396.900 ns/op -15.6 ns/op / -3.8% (better)
Windows MinGW ARM64 RearmStopped/Go 70.640 ns/op +0.05 ns/op / +0.1% (worse)
Windows MinGW ARM64 RearmStopped/LLGo 256.600 ns/op -2.4 ns/op / -0.9% (better)
Windows MinGW ARM64 ResetActive/Go 30.930 ns/op -0.08 ns/op / -0.3% (better)
Windows MinGW ARM64 ResetActive/LLGo 138.400 ns/op -4.2 ns/op / -2.9% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.090 ns/op -0.09 ns/op / -0.3% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 126.200 ns/op -1.4 ns/op / -1.1% (better)
Windows MSVC AfterFuncZeroDelivery/Go 486 ns/op -12.3 ns/op / -2.5% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 128700 ns/op -7403 ns/op / -5.4% (better)
Windows MSVC CreateStop/Go 116.800 ns/op +0.5 ns/op / +0.4% (worse)
Windows MSVC CreateStop/LLGo 466.400 ns/op +10.8 ns/op / +2.4% (worse)
Windows MSVC RearmStopped/Go 31.590 ns/op +0.1 ns/op / +0.3% (worse)
Windows MSVC RearmStopped/LLGo 363.900 ns/op +8.9 ns/op / +2.5% (worse)
Windows MSVC ResetActive/Go 19.040 ns/op -0.13 ns/op / -0.7% (better)
Windows MSVC ResetActive/LLGo 205.600 ns/op +44.9 ns/op / +27.9% (worse)
Windows MSVC ResetHeap1024/Go 19.250 ns/op -0.61 ns/op / -3.1% (better)
Windows MSVC ResetHeap1024/LLGo 149.400 ns/op +16.3 ns/op / +12.2% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 954.200 ns/op -2.3 ns/op / -0.2% (better)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 184493 ns/op -670 ns/op / -0.4% (better)
Windows MSVC 386 CreateStop/Go 191 ns/op -3.9 ns/op / -2.0% (better)
Windows MSVC 386 CreateStop/LLGo 1620 ns/op -45 ns/op / -2.7% (better)
Windows MSVC 386 RearmStopped/Go 63.220 ns/op -0.12 ns/op / -0.2% (better)
Windows MSVC 386 RearmStopped/LLGo 325.100 ns/op -5.9 ns/op / -1.8% (better)
Windows MSVC 386 ResetActive/Go 38.980 ns/op -0.06 ns/op / -0.2% (better)
Windows MSVC 386 ResetActive/LLGo 953.400 ns/op -3.4 ns/op / -0.4% (better)
Windows MSVC 386 ResetHeap1024/Go 39.430 ns/op 0 ns/op / +0.0%
Windows MSVC 386 ResetHeap1024/LLGo 170.900 ns/op -1.4 ns/op / -0.8% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 670.900 ns/op +14.7 ns/op / +2.2% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 129999 ns/op -13912 ns/op / -9.7% (better)
Windows MSVC ARM64 CreateStop/Go 196.500 ns/op -0.6 ns/op / -0.3% (better)
Windows MSVC ARM64 CreateStop/LLGo 377.200 ns/op -5.7 ns/op / -1.5% (better)
Windows MSVC ARM64 RearmStopped/Go 70.550 ns/op -0.02 ns/op / -0.02834% (better)
Windows MSVC ARM64 RearmStopped/LLGo 269.900 ns/op -1.9 ns/op / -0.7% (better)
Windows MSVC ARM64 ResetActive/Go 31.050 ns/op -0.07 ns/op / -0.2% (better)
Windows MSVC ARM64 ResetActive/LLGo 138.700 ns/op -1.7 ns/op / -1.2% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.150 ns/op -0.03 ns/op / -0.1% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 138 ns/op -0.6 ns/op / -0.4% (better)

Compared with b07cd12ad1ac measured in the same runner job.

@cpunion

cpunion commented Sep 12, 2026

Copy link
Copy Markdown
Collaborator Author

Updated in 5daef74 after investigating the paired benchmark comment. The +80 B in cprintf came from five new helper-function symbol-index rows, not machine code. Those helpers are removed, snapshot bookkeeping is only linked through the public-runtime hook, and the existing two-word recover-state ABI is preserved. Nil-check failures now have a separate formal return instead of joining the normal path, preserving non-nil facts and avoiding the original non-LTO code growth. Source anchors follow the actual cold panic call. A single empty matching-register barrier inside the noinline runtime helper also protects return-PC identity from LTO argument propagation; placing barriers at every call site was tested and rejected because it increased code size.

Same-path local macOS/arm64 comparison against daada50d2270: cprintf and cprintf-lto have no file or code growth; println is -80 B file / -108 B code and println-lto is unchanged; fmtprintf has unchanged file size / -1,184 B code. fmtprintf-lto still has +16 B file / +5,664 B code, documented explicitly in the PR body. CI's paired benchmarks will determine the cross-platform size and timing result; noisy local timings are not presented as a speedup.

Validation includes the full llgo test ./test/go package, default and full-LTO runs of GOROOT issues 17381/27201/34123, full-LTO caller/recover tests, compiler unit tests, and the actual llcppg reproduction still pointing to go/ast/ast.go:520 and the application call chain. New tests cover wrapped/indirect recover and uncomparable repanics, and model LTO condition propagation with object emission for the native CI architectures. The separate full local cl fixture run reached its five-minute suite budget while still advancing; it is not claimed as a full local pass. No CI workflows or unrelated code were changed.

@fennoai Please re-review the size-conscious nil-check lowering and recovered-panic lifetime changes.

@fennoai

fennoai Bot commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

Re-review: size-conscious nil-check lowering & recovered-panic lifetime

Re-reviewed at head 5daef74c against base daada50d (git diff daada50d...5daef74c). Focus was the two areas called out in your update. No correctness defects found; the revision holds up.

Nil-check lowering (ssa/memory.go, ssa/stmt_builder.go, cl/instr.go, cl/compile.go)

  • Separate formal return instead of self-loop / success-path join — AssertNilDeref now emits a real ret (void or ret undef for value-returning callers) in the cold failure block when NeedsFramePointer(), and keeps Jump(blks[0]) on non-FP targets. This keeps the caller's saved return PC inside the caller without losing the success block's non-nil fact — the comment at ssa/memory.go:379-392 documents exactly why neither the self-loop nor the join is acceptable. The ret undef path is correct here: these results are never observed and this is deliberately not a Go return (no defer dispatch).
  • Single in-helper barrier — preserveNilCheckCondition installs one =r,0 matching-register InlineAsm in the runtime.AssertNilDeref entry, gated to that one function. The ReplaceAllUsesWith(opaque) then SetOperand(0, param) dance is the right idiom to avoid the barrier feeding itself; the comment flags it. Confining it to the noinline helper (rather than per call site) is the correct size/optimization tradeoff and matches the PR's size numbers.
  • Panic-site anchor — PanicSite is installed once per builder and invoked from inside the cold failure block (AssertNilDeref calls b.PanicSite(b) after SetBlockEx(blks[0])), so backend block placement can't attribute a later statement's line. The panicSitePos lifecycle is safe: it's reset at the top of every compileInstr/compileBlock/EndBuild, and every call site sets it via recordPanicSite immediately before the matching AssertNilDeref/Store in the same instruction, so no stale anchor can leak across instructions.
  • Native store panic-site metadata — dropping the GOOS=="windows" guard in cl/compile.go so ELF/Mach-O also emit the scoped carrier is consistent with recordPanicSite already scoping via panicSiteFuncs. Test renamed to ...OnNativeTargets and asserts presence on linux/darwin/windows.

Recovered-panic lifetime (runtime/internal/runtime/caller.go, z_rt.go, unwind_llgo.go)

  • Reuse tied to the active recover activation — SavePanicCallerFrames(v) now returns whether a fresh snapshot is needed, reusing the retained snapshot only when the raw interface words (_type+data, not Go equality — correct for uncomparable values) and p.recovered.frame == gp.recoverFrame match. Anchoring reuse to the still-live activation is the key property: a later panic(sameValue) after the activation returned correctly gets a fresh traceback.
  • Lifetime clearing is complete — the retained recovered value is cleared on the paths that end the activation: EndRecoverFrame, EndRecoverFrameAlias (guarded on the wrapper token mismatch), abortPanics/Goexit, and storePanicPCs. Recover only sets it behind the RecoverMark != nil && panicPCs.n != 0 guard, so tiny programs without the public-runtime hook don't link or exercise the bookkeeping.
  • ABI preserved — recoverState stays two pointers; the snapshot adds recoveredPanic{value, frame} to the per-goroutine panicPCStore only, no separate allocation. Panic now calls the hook PanicPCSnapshot(v) directly (nil when lib/runtime isn't linked), and capturePanicPCs(v) short-circuits on !SavePanicCallerFrames(v). This matches the "no new helpers, hook-gated bookkeeping" claim.

Tests

Coverage maps directly onto the risk surface and I'm satisfied it exercises the tricky paths:

  • TestAssertNilDerefColdCall adds darwin/arm64 + three Windows arches, a value-returning callee (exercises ret undef), asserts the callback runs in the failure block, and models whole-program LTO — reconstructs the helper as internal+noinline with the opaque barrier, runs default<O2>, and asserts the static-nil caller does not gain noreturn. That directly validates the barrier's purpose.
  • TestCompileRuntimeNilCheckCondition confirms the barrier is present only for AssertNilDeref on non-wasm and that the helper branches on the opaque result.
  • TestCallerRepanicTraceback covers same-value repanic, replacement (different-value) panic, later reuse of the same value after the activation returned, direct/indirect transparent wrappers, and uncomparable slice repanics.
  • PC-line record counts in TestCompileRuntimeCallerPanicPCLineMetadata were updated to account for the per-cold-block anchors (1→3, 2→4), which is the expected consequence of the separate failure blocks.

Notes (non-blocking)

  • I could not build or run the Go-side tests in this sandbox — the LLVM C headers / toolchain aren't installed, so go test ./ssa ./cl ./test/go fails at cgo link. This was a static re-review; CI's paired benchmarks and full fixture suites remain the acceptance gate, as your PR body already states. Nothing in the diff contradicts the reported size/behavior results.
  • The full-LTO fmtprintf +16 B file / +5,664 B code increase is the one real, intentional cost — it's the price of keeping the helper potentially-returning so return-PC identity survives whole-program optimization. It's documented in the PR body and confined to the helper (no per-call assembly, no runtime call on successful checks), which the LTO-modeling test backs up.

LGTM on the two focus areas. FennoAI cannot approve PRs, so this is advisory feedback rather than an approving review.

@cpunion

cpunion commented Sep 12, 2026

Copy link
Copy Markdown
Collaborator Author

Thanks for the re-review. Follow-up 7aa6c5f92 changes only two files under test/go, leaving the reviewed compiler/runtime implementation unchanged. The existing statement-line test now checks the caller identity after a statically panicking leaf (issue17381), the exact nil-load line before a conversion/global write (issue27201), and the exact pointer-store line on all platforms (issue34123). The existing repanic test now uses nested testing.T.Run, matching the llcppg reproduction. There is no additional GOROOT build-driver test or workflow change; these checks run through ordinary go test and llgo test CI.

All four scenarios pass locally with Go and LLGo default/full LTO. As a negative control, the unmodified cdaa60ff7 compiler/runtime fails each corresponding assertion: missing caller, nil-load line two lines late, pointer-store line two lines early, and the nested test panic missing its original application frame. The first draft of the load check did not catch the old bug, so it was tightened to preserve the original load/conversion/store shape before committing. The new in-process checks take less than 0.01 seconds each locally and reuse the existing test binary.

The previous CI round finished before this push: every other check passed, including 100% patch coverage and all paired benchmarks, but the Qiniu Linux LLGo shard hit its one-hour execution limit. GitHub currently returns no job log for that shard, so I cannot yet attribute the timeout to a particular test or claim the entire round passed. The follow-up starts a fresh CI round. The PR body now includes the cross-platform size results and the precise scope of the regular-CI regression coverage. All existing inline review threads remain resolved.

@cpunion
cpunion force-pushed the codex/fix-panic-line-20260912 branch from cab2663 to 148f3d0 Compare September 12, 2026 16:16
@cpunion
cpunion force-pushed the codex/fix-panic-line-20260912 branch from 148f3d0 to cf807ff Compare September 12, 2026 23:32
@xushiwei
xushiwei merged commit 4564a01 into xgo-dev:main Sep 13, 2026
65 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

runtime: same-value repanic overwrites the original panic traceback

2 participants