Describe the bug
Built with Go 1.27 for linux/arm64, dmaChannel.startIO (bcm283x/dma.go:552) compiles to one 64-bit store covering CS and CONBLK_AD:
go1.27.1: dma.go:552 STPW (R2, R1), (R0)
go1.24.4: dma.go:552 MOVW R1, 4(R0)
dma.go:553 MOVW R1, (R0)
The peripheral bus does not take it as two register writes. The channel goes ACTIVE without a control block and never completes, so wait() spins forever. Reached through Pin.PWM on DMA-driven pins, StreamIn and StreamOut. Go 1.24 through 1.26 emit two ordered 32-bit stores.
To Reproduce
GOTOOLCHAIN=go1.27.1 GOOS=linux GOARCH=arm64 go build -o bcm283x.a periph.io/x/host/v3/bcm283x
go tool objdump -s 'bcm283x.(*dmaChannel).startIO' bcm283x.a
- See STPW. Repeat with go1.24.4 and see two MOVW.
Expected behavior
Each register write is one 32-bit access in program order.
Platform
- OS: Raspberry Pi OS 64-bit (any arm64 build shows it in the disassembly)
- Board: Raspberry Pi 4
Additional context
Cause is cmd/compile CL 760100 (commit 996b985008, in Go 1.27): stores to one base pointer are buffered, sorted by offset and paired. Working as intended on the Go side; plain stores carry no width or ordering guarantee. Register access needs sync/atomic, which emits one ordered 32-bit access each (STLRW/LDARW on arm64, no measurable cost at register-write rates).
I hit the same failure in a vendored ledctl driver and fixed it with atomic loads and stores on every register field. I'd like to take this one. It touches every register map in bcm283x (dma, clock, pwm, pcm, timer, gpio), so before writing it: atomic.Uint32 fields with Load/Store at each access, so a missed access fails to compile, or keep the uint32 fields and go through helper functions with pointer casts, which keeps the structs reading like the datasheet? Single PR, bcm283x only; allwinner has the same pattern and would follow separately.
Describe the bug
Built with Go 1.27 for linux/arm64,
dmaChannel.startIO(bcm283x/dma.go:552) compiles to one 64-bit store covering CS and CONBLK_AD:The peripheral bus does not take it as two register writes. The channel goes ACTIVE without a control block and never completes, so
wait()spins forever. Reached throughPin.PWMon DMA-driven pins,StreamInandStreamOut. Go 1.24 through 1.26 emit two ordered 32-bit stores.To Reproduce
GOTOOLCHAIN=go1.27.1 GOOS=linux GOARCH=arm64 go build -o bcm283x.a periph.io/x/host/v3/bcm283xgo tool objdump -s 'bcm283x.(*dmaChannel).startIO' bcm283x.aExpected behavior
Each register write is one 32-bit access in program order.
Platform
Additional context
Cause is cmd/compile CL 760100 (commit 996b985008, in Go 1.27): stores to one base pointer are buffered, sorted by offset and paired. Working as intended on the Go side; plain stores carry no width or ordering guarantee. Register access needs sync/atomic, which emits one ordered 32-bit access each (STLRW/LDARW on arm64, no measurable cost at register-write rates).
I hit the same failure in a vendored ledctl driver and fixed it with atomic loads and stores on every register field. I'd like to take this one. It touches every register map in bcm283x (dma, clock, pwm, pcm, timer, gpio), so before writing it:
atomic.Uint32fields with Load/Store at each access, so a missed access fails to compile, or keep theuint32fields and go through helper functions with pointer casts, which keeps the structs reading like the datasheet? Single PR, bcm283x only; allwinner has the same pattern and would follow separately.