Skip to content

alpenglow: improve replay, leader packing, and vote latency - #279

Closed
7layermagik wants to merge 22 commits into
alpenglow-devfrom
7layer/alpenglow-validator-performance
Closed

7layermagik wants to merge 22 commits into
alpenglow-devfrom
7layer/alpenglow-validator-performance

Conversation

@7layermagik

@7layermagik 7layermagik commented Sep 14, 2026 •

Copy link
Copy Markdown

Superseded by #287 (Turbine/production), #288 (voting), #281 (execution), and #292 (replay/recovery).

Busy Alpenglow slots leave avoidable work between shred arrival, transaction execution and vote enqueue. This consolidates the streaming-verification and leader-packing branches with the subsequent certificate, vote-persistence, replay-progress and checkpoint fixes into one PR against current alpenglow-dev (33dde405).

The implementation is organized into focused commits. Start with the review guide, which maps each subsystem to its tests, operating contract and benchmark evidence. Generated benchmark output is retained but collapsed in the diff.

Changes

  • Update Narya and overlap transaction signature verification with completed entry components during shred arrival. Small ready batches dispatch immediately; larger ready requests use bounded groups without a fill timer.
  • Compare cached entry bytes directly against shred slices, reuse authenticated FEC-root inputs, reduce completion ordering work, and allocate each prepared transaction-pointer view once.
  • Prepare owned transactions before leadership, recheck bank-dependent state, reduce execution scratch/entry-root/queue overhead, expose bounded queue and completion-reserve settings, and correct slot-duration-dependent resource limits. Include deterministic near-limit block generation and shred round-trip tests.
  • Bound scheduler memory by removing consumed, evicted and expired entries from both indexed priority heaps. Preserve the existing slot-local scan/retry policy, new higher-priority arrivals, ordering and capacity. Add retention, repeated-rebuffering and concurrent-access regressions.
  • Reduce observer/VM allocation overhead and correct the Votor certificate layout for Agave/Firedancer compatibility.
  • Add opt-in durable signing reservations, --wait-to-vote-slot, and one ordered background history writer. Preserve detailed decisions, fail closed on writer faults, seal clean shutdowns, and conservatively gate voting and leadership after uncertain recovery. Default history persistence remains synchronous.
  • Preserve valid live voting opportunities within the retained pool/history window, order durable pruning behind replay, and advance trusted replay progress without waiting on the certificate-verification mutex.
  • Capture immutable transaction-status checkpoint lineage on replay and encode it on the existing fold worker. Preserve the checkpoint format, sidecar/manifest selection, failure handling and commit ordering.

FEC acceleration remains in #259; the small shared DATA_COMPLETE boundary correctness fix is included for streaming. Genesis bootstrap remains in #276, and #278's skipped-parent replay work is separate. This incorporates the existing 7layer/streaming-sigverify and 7layer/leader-block-packing branches without opening duplicate PRs.

Benchmarks and scope

Ryzen 7 9700X / Go 1.26.4. These are stage measurements, not a combined whole-validator speedup:

Work Baseline Result
Build a 48,622-transaction bank from wire bytes Current alpenglow-dev, 33dde405 209.64 → 150.30 ms
Same bank from decoded transactions Current alpenglow-dev, 33dde405 189.73 → 132.40 ms
Capture a roughly 30 MB status checkpoint on replay Original synchronous snapshot algorithm, unchanged in current base 202–207 ms → 5.92–6.06 µs; encoding moves to the worker

The bank fixture uses unique 198-byte, single-signature transactions and three alternating paired rounds. It excludes signature verification, real AccountsDB, broadcast and consensus. It also checks the applicable cost limit and round-trips emitted entries through shreds.

Separately, a synthetic tip workload of 33,760 valid 1,232-byte transactions arriving over 200 ms measured final-shred-to-ready 99.89 → 6.504 ms with overlap disabled/enabled on the same extracted implementation. This is explicitly not development-head versus PR; both sides use the new Narya/batching/completion code. Historical component comparisons retain their original intermediate baselines.

The live checkpoint trial recorded 29–34 µs replay-side captures across ten checkpoints; worker encoding still cost 179–367 ms. Live trials also include the separate FEC/producer work, and the running binary is not identical to this independent PR. Changing workload and restart conditions prevent attributing all observed FAST-score improvement to this patch. Full methods, raw outputs and exclusions are linked from the review guide.

Validation

  • Combined-source race tests passed locally and natively for voting/consensus, replay, Turbine, signature verification, leader/scheduler, accounts, config, cost model, Merkle, SBF, stats, block, transaction fixture and node packages.
  • Native vet and the complete Mithril build passed. Formatting/diff checks passed.
  • Add a regression-tests CI job running complete race suites for Alpenglow, consensus, replay, Turbine, signature verification, block production/scheduling and node startup. The exact command passed locally and in Linux CI at bcada839; scheduler vet and a complete local build also passed.
  • The broad race command is not fully green: pkg/sealevel has 19 BPF-loader test failures ending in an UpgradeableLoaderClose panic. Unchanged 33dde405 reproduces the same failures on both machines. Baseline and candidate logs are included in the validation report.

Recovery and remaining limits

Reserved voting is experimental and opt-in. Enrolling advances the history format and adds a signed reservation; preserve both safety files across restarts and AccountsDB recovery. Software crash tests and clean testnet restarts are not host power-loss or mainnet qualification. The recovery contract describes initialization, clean-marker consumption, leader gates, uncertainty handling and halted-cluster behavior.

The queue follow-up fixes heap retention: after consuming 100,000 higher-priority entries, its regression retains one entry in each heap for the one remaining active transaction. Local M4 microbenchmarks measured median removal cost of 143 → 188 ns with equal rewards and 254 → 274 ns with mixed rewards; this is the cost of immediate counterpart removal, not a net-validator speedup. Queue validation and raw evidence. The current slot-local scan/retry policy is unchanged. Effects on live block fullness and FAST scores remain unmeasured, and remaining reward/FAST omissions are follow-ups. The live validator, continuous own-leader load and both monitors were left running during PR preparation; runtime keys, faucet/controller automation and live ledgers are not included.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant