Fix/prt timers - #287
Draft
GCdePaula wants to merge 69 commits into
Draft
Fix/prt timers#287GCdePaula wants to merge 69 commits into
GCdePaula wants to merge 69 commits into
Conversation
Add deterministic FFI vectors, logged from a bare emulator, for a manual yield that the step lets outrank halt and the input budget: an RX_ACCEPTED opening on the budget's last cycle (also at the maximum mcycle) or on a preset halted machine delivers the input and differs from the idle opening, and an RX_REJECTED closing on the budget's last cycle still substitutes the revert root. SendCmioResponse and UArchReset read only the pending yield, so these are the canonical answers the node and Lua client must match.
node_metadata now records a COMMITMENT_SEMANTICS version next to the node version and schema fingerprint, and startup refuses a store stamped with another value. The crate version is frozen, so a change to leaf values or transition shapes that leaves schema.sql alone would otherwise reuse stores built under the old rules; a schema revert or backport could even restore an old fingerprint. Every commitment change bumps the stamp.
The step reads only the pending manual yield: SendCmioResponse delivers the next input to any RX_ACCEPTED yield and UArchReset reverts any RX_REJECTED one, whatever the halt flag or the input budget. The node and the Lua client checked halt and mcycle overflow first, so an input yield executed on the budget's last cycle (mcycle == imcyclemax), or preset on a halted template, was padded as terminal. At the next fed window the node's leaf then differed from the canonical one and its own proof, which logs the delivery, contradicted it: an adversary matching the node up to that leaf could win with an arbitrary suffix. Terminal now means an exception or unexpected manual yield, or halt or overflow with no manual yield pending, in both clients. The docs that stated the old rule are corrected, and COMMITMENT_SEMANTICS moves to 2.
The hand-written patch chains had stopped reaching their targets: the 2026-07 chains used 2^28 links against a level-1 stride of 27, and Lua's 64-bit integers turned stf_all's (1 << 68) into 0. A baseline run showed stf_all epochs 2 to 4 and stf_revert sealing on other closing slots than the ones they document, so the on-chain STF coverage they claim was not being exercised. Env.steering_patches now builds each chain in 256-bit arithmetic from the deployed level table, and Env.run_steered_epoch asserts that the dispute's leaf match sealed exactly on the steered transition, reading its divergence cycle from sealedMatch pinned at the LeafMatchSealed block. stf_all, stf_revert, and big_input use it; every epoch now seals where intended, including input 1's fused feed at 2^68, the first dispute past window 0.
The node compiled in the root stride (LOG2_STRIDE = 44) and refused any other deployment. It now reads the factory's whole level table at startup, validates it with engine::TournamentGeometry (the root spans the 92-bit ruler, levels tile, the leaf stride is zero, the root stride lies between one big cycle and one input window), and pins it with the consensus address in sling_config. The runner, the window-root rows, the roll's prefix check, the settlement fold, and the dispute facade all use the pinned root stride, and settlement roots stay byte-identical under the canonical [44,27,0] table. Table stability is a trust assumption of the parameters provider, so the Hero checks every tournament descriptor on its path against the pinned row for its level, and the root commitment it would join with against the settled computation hash. The frontier fold now serves only the run stride: a coarser tree samples other states and was served wrongly before, though no production level asked for one. The e2e oracle and the gc, multi_sybil, and bad_commitment scenarios take their strides and heights from chain as well.
DEVNET_GEOMETRY=two-level builds a devnet whose factory serves the selected [37,0]/[55,37] table ahead of the canonical switch. A test-only TableTournamentParametersProvider (validated on-chain by the existing table validator) and a devnet-only script that inherits the production DeploymentScript through a separate entry point replace only the parameters provider; production scripts and src/ are untouched. The bundle marker records the geometry, and every consumer verifies against the DEVNET_GEOMETRY it runs with, so a bundle of one geometry never passes for another. `just test-rollups-two-level-smoke` runs echo `simple` on it.
A two-level leaf commitment spans 2^37 transitions, and compute_and_store built it leaf by leaf: every run and every Merkle node of the span stayed in memory, and idle stretches expanded into their churn runs cycle by cycle. The two-level smoke lost its dispute by timeout while the node was still building. A stride-0 quartet tall enough that its 8 stored levels stay above big-cycle granularity now takes Ruler::collect_big_cycle_roots: each big cycle folds to its subtree root as it completes, and an idle stretch steps one captured cycle and repeats its root. The quartet's tree is built c levels up from those roots; its root and stored fanout are unchanged. Shorter quartets keep the leaf-by-leaf path, which also serves as the differential oracle. Tests pin both: the folded roots against the brute-force oracle over every big-aligned span, fresh and resumed; every stored fanout row of a tall quartet, including one exactly at the threshold height; a real machine across an active and an idle cycle; and an idle stretch costing one stepped cycle.
The Hero's dispute loop runs inside the epoch manager's async task, and its commitment builds and proofs are synchronous machine work. A leaf build of many minutes held a tokio worker the whole time, and tasks queued on that worker, the chain-facing reader included, could starve behind it. hero::machine_work wraps level_material and action preparation in block_in_place on the multi-threaded runtime, so the worker's queue moves to a replacement first; a current-thread runtime runs the work inline. A test on a one-worker runtime shows a queued task running during the wait, and it fails without the hand-off. The manager itself still waits for the result; node-architecture.md debt 7 records what that still delays.
After the per-PR e2e battery, the e2e job rebuilds the devnet with DEVNET_GEOMETRY=two-level and runs echo `simple` against it, so the node keeps winning a dispute on the selected [37,0]/[55,37] table before the canonical constants switch. It runs last because it replaces the devnet bundle. sealed_leaf_timeout_* stays canonical-only: its kill choreography depends on each level's height parity.
The e2e oracle and the node cross-check each other, but both are Dave code. The plan makes the released v0.21.0 cartesi-machine CLI the reference implementation, so every epoch the oracle computes is now recomputed by the CLI from the oracle's epoch snapshot and the chain inputs: its mcycle computation hash at period (root stride - 20), read from the deployed table, must equal the oracle's commitment, which the node's already must match. The CLI runs on a disposable clone in stored revert mode (its default fork mode needs a machine server), and a mismatch keeps the inputs, the CLI log, and a repro.sh under _oracle/cli-epoch-<n>/. The gate refuses any other CLI version. Checked on canonical echo `simple`, yield `stf_revert` (three rejected inputs), and the two-level smoke, where the period-17 CLI hash of a four-input epoch equals the node's root.
A review of the whole stack, each finding adversarially verified, found no bug in the two-level logic. It found a protocol flaw (inner commitment construction drains the correct party's clock across Sybil matches), the unmeasured scale properties of dense height-37 leaves, and several robustness and evidence gaps. The plan now tracks them as R1-R15. test-harness.md claimed every scenario runs on either table; only echo `simple` has run on two levels, and steered STF scenarios need dense leaves the devnet clock cannot host. The plan also named v0.21.1-test1's commits unrelated; one is the O_CLOEXEC fix behind the flaky node unit tests.
A correct commitment must build its child commitment before it can join a child tournament, and the join's lateness is charged to its clock. The winner's remainder then replaced the parent clock and nothing refilled it, so every sealed parent match against a new Sybil, diverging at a fresh span, cost the correct party one build. Once the censorship budget C was spent, a few Sybil bonds eliminated it: two with dense two-level leaves, and a handful on the checked-in table too, since the node's finality-gated joins cost minutes each. The rule is now that every honest action gets one inclusion G, and joining a child also gets the build T. winInnerTournament refills the returning winner by up to T + 2G (build, join, propagation), capped by the sealed pair's envelope max(r1, r2), the child's own allowance. No clock exceeds that envelope and the refill never lowers a winner, so no clock mass is created and no new elimination path opens; the adversary gains at most T + 2G of lingering per child it pays a bond for. T (commitmentBudget) is a new TournamentParameters and clone-argument field. It belongs with the geometry: ArbitrationConstants.COMMITMENT_BUDGET is 30 minutes for the checked-in table, and the two-level devnet profile uses 60. The new ClockBudgets library derives every block budget from wall-clock inputs, including maxAllowance = C + G + (L - 1)(T + 2G): the root join's inclusion plus one pending delegation per inner level. The canonical provider takes the block time, C and G, and each chain kind states only its censorship budget (mainnet one week, testnet 8 hours, devnet none). Mainnet moves from one week plus 60 minutes to one week plus 85. The test-only table validator now requires uniform budgets and a root allowance that holds the pending delegations. Tests: a correct commitment joining within T + G and propagating within G returns with its pre-seal clock, delegation after delegation, from either side of the envelope; the property fails on the old rule (193 != 200), and mutating the refill, the envelope wiring, or the inner clones' T is caught. Deployment and devnet-profile tests deploy the providers from the scripts' own encodings. Gas allocations are unchanged: the inner seal and inner win still fit theirs, with recorded headroom reduced (6k and 41k), and the leaf-win alternate now ties the selected witness. The accepted calibration under the pinned release Forge remains to be recorded. This is a new deployment generation: the factory ABI and the clone-argument encoding change, and Tournament bytecode changes (its ABI and storage do not). The e2e decoders follow.
The measurement report, its generated documents, and the harness comments quoted the old allowance literals (devnet one hour, mainnet one week plus one hour) and a deployment helper that no longer exists. They now point to ClockBudgets and state the derived values. The plan records how R1 was fixed and the follow-ups under the same rule: a G discount for the leaf proof and for a paused winner's timeout cleanup, and inner joins from the latest block.
The plan now also tracks the node's honest-action latency (proofs replay from window boundaries, a 30 s tick, one action per tick) and the unpropagated Sybil-versus-Sybil child winners, and records two decisions: joins stay finality-gated, since the Hero builds on the latest view and T + G covers max(build, finality) + inclusion; and the node does not pin or judge T, leaving that to operators. R13 notes that honeypot and the CLI gate on its image now pass. A new backlog keeps the tooling friction observed in this campaign, so a later cleanup starts from evidence.
The constants benchmark and generator named T the "inner timeout", while the contracts and the clock model now call it the commitment budget: the time every inner tournament's commitment must build within. The Lua benchmark reads DAVE_COMMITMENT_BUDGET_MINUTES, the Rust generator takes --commitment-budget-minutes, and their reports, the node's capacity warning, and the living docs follow.
Under the rule that every honest action gets G, a response already earned a non-bankable discount, but a correct party that won a leaf proof or a timeout claim paid the action's full cost, and the adversary chooses how many such wins it must make inside one child. Clock.pauseWinnerAt now charges a winner its live cost (time run plus any deferred charge) beyond one responseBudget, never above its stored balance. Eliminating both sides earns nothing. Survival is still decided on the full cost by the unchanged timeout classifier, so the node and the Lua client, which read that view, are unaffected. winMatchByTimeout checks for a finished tournament inline to decode its arguments once. The lifecycle, recursive, four-level, and bounded-delay oracles follow the rule. In the finite delay model a pair can now end with a free proof inside G, so its survivor re-pairs with a full clock: each later sequential pair may take max(g - 1, 0) blocks more, and the replayed maximum witness ends at block 20 instead of 19. The coarse per-match bound is unchanged. The four- level trace keeps one refill below the envelope so it still pins T + 2G. Tournament ABI and storage are unchanged.
The review of 2c502f6 and 935dc13 is kept verbatim as a frozen record, and the reviews index gains it and the missing 2026-08-27 recalibration entry. The plan takes over its open work. CF-01 stays an accepted limitation, with its reasoning restated: the adversary chooses how often a proof is overtaken by the opponent's expiry, but not the latencies, which cost nothing inside G and are censorship beyond it; R16's latency target now covers the fallback claim. R18 owns the STF evidence gate and R19 the honest-survival model. W2.1 and W5.1 stop overclaiming the proven transitions and the gas calibration.
The harness counted a sealed leaf transition as proved, but a timeout win leaves the same seal and settlement behind, and the node claims a timeout before retrying a rejected proof, so a broken on-chain state transition could pass stf_all, stf_revert, or big_input (R18, CF-02 in the 2026-09-29 clock refill review). The reader now returns each sealed leaf's match ID hash with its transition, and Env.assert_leaf_match_proved requires exactly one MatchDeleted with reason STEP for it. run_steered_epoch applies it after settlement, and simple opts in, so the per-PR smokes, the two-level one included, gate the transition too. run_epoch stays neutral, since kill and chaos runs may end a leaf on a timeout. sealed_leaf_timeout_winner is the negative control: its timeout-won record must correlate with the fixture's match and be refused.
The Hero's checks against the pinned tournament table and the epoch-start state returned errors that the manager logged and retried every tick, so a node that hit one stopped defending without alarm (R8). Both are invariant violations, not transient failures, so they now panic, per the node's own doctrine. The check that the root commitment equals the settled computation hash is gone. It compared two of the node's own folds, whose equality is unit-tested under both tables; the staging assert still catches a mismatch on a win; and when the settlement fold was the wrong one, it made the node refuse a dispute it would have won.
More than 2^24 inputs in one epoch is an economic bound, a flood of roughly 5e11 gas, where the contracts drop the tail and the node panics; a clearer error would still stop the node, so D5 builds nothing. Proof positioning replays one input's prefix from the boundary the dispute wrote back, bounded by the per-input contract, so R16 needs no mid-input snapshots; W4.7 notes that the snapshot gap bounds the first root response inside an input. R9 folds into W6.4, which reads yields from registers anyway, and W6.4 states that seam 1 needs a 2^48-cycle input and recommends no collector guard. The canonical switch becomes a follow-up PR that should reduce to small Solidity changes.
The test-only TournamentGeometry::canonical() hard-coded the three-level table, so the day ArbitrationConstants switches to two levels every test calling it would silently keep testing the old geometry. Tests that need a particular table now name it (three_level or two_level), and the drift guard and the devnet discovery test read the checked-in table itself through TournamentGeometry::checked_in(), which parses ArbitrationConstants.sol. The storage tests pin two levels, so stride 37 runs in this PR rather than first in the switch.
…able The deployment and DaveAppFactory tests asserted three levels and a 30-minute commitment budget as literals, so switching ArbitrationConstants would have broken them in the switch PR. They now derive the pending refills and the commitment blocks from ArbitrationConstants, and the canonical-geometry tests compare against ClockBudgets. The zero-allowance case uses a block time longer than any budget, since at T = 60 min a one-hour block still rounds the commitment budget to one. testCheckedInCanonicalTable stays the one deliberate pin of the table.
M2, the new `just measure-two-level-leaf`, built the dense two-level leaf (stride 0, height 37, the stress workload's first input) in 8,405 s against a 30-minute target, although its density matches steady state. The emulator fixes its hash-tree concurrency at load and hashes in parallel once a hash's dirty pages outnumber the host's cores; at one root hash per ustep the thread pool's fork and join then cost about 110 us against 12 us of hashing. Rulers now say how they will hash (Hashing::PerStep for a stride-0 build, Sampled otherwise), and the node loads per-step machines with serial hash-tree updates while sampled strides, which hash large dirty sets, keep the parallel path. The Lua client, whose commitments hash per ustep, loads serially throughout. M2 now takes 1,100 s with 103 MiB peak RSS, and the bounded-memory leaf build holds. The plan records M2, adds pages per hash as a second dimension to R2, and asks upstream for a work-based parallelism threshold.
The scenario asserted the three-level heights 48/17/27 and assumed the sybil seals every inner match, which holds only at an even-height root. At the two-level root (55, odd) the node joins first and seals, so the sybil would never have become the leaf's final responder. The scenario now reads the level table, requires odd heights below the root, and expects the sybil to seal an even-height root and every inner level below. When the node seals a child's parent, the node is stopped across the sybil's join of that child, as it already was around the sybil's own seals. Both scenarios pass on the canonical and the two-level devnets.
The big-cycle-root builder, which builds every stride-0 quartet of height 28 or more, had real-machine differentials only over idle padding. A span over the first 2^8 big cycles of window 1, where echo runs the input, now checks its root, children and proofs against the leaf-by-leaf prototype (R12).
The constants report printed the three-level table as the "(current)" row and warned when a derived root was taller than it, and the `--full` help named its 2^44 window as if it were every table's. The bench's pinned table is now named for what it is, the report no longer ranks derivations against it, and the bump checklist includes COMMITMENT_BUDGET.
The docs, contract guardrails and CI comments described the three-level table as the checked-in one and the two-level table as planned or not live, so the canonical switch would have had to rewrite them. They now state both tables and point at ArbitrationConstants for the deployed one, and the CLI gate note says what it samples at each root stride. The plan records the switch as a follow-up PR of small Solidity changes (ArbitrationConstants with COMMITMENT_BUDGET = 60 min, plus the one pinning test), the two-level dry run of the steered STF scenarios, which needs no dense devnet profile (R6), and R12 done.
1c6018d loaded the constants harness's one machine per-step, so the hash-cost curve, which prices the sampled root stride, would have been measured with serial hashing that production does not use there. The dense leaf rate and the curve now come from separate machines, loaded per-step and sampled respectively. The Lua constants harness hashed its leaf phase with the emulator's default parallel updates and now loads it serially, like both clients do.
Gaps 3 and 4 of the test strategy reset, for Dave-owned cases. The released CLI's uarch cycle computation hash over one mcycle period is Dave's stride-0 commitment under a root leaf at stride period + 20, computed independently and through the collect API. reference_cli_goldens_hold now also records nine such leaves on echo and yield, and leaf_commitments_match_the_reference_cli (engine gate, about 23 s) requires the node's facade to reproduce them on a fresh store. Period 8 exercises the big-cycle-root builder (height 28, what a switch to the collect API replaces); period 7 the plain collection path (height 27, the three-level leaf). The periods hold a fed window start, an accepted yield, a rejection's revert on each program, a revert crossed while positioning, and a padding window. All nine agree. Period 17 is left out: the CLI spends about 2 minutes on a mostly idle period-17 leaf and more than 13 on a dense one. The CLI invocation is shared between the root and leaf goldens.
Gap 5 of the test strategy reset. node_proof_vectors_hold pins the node's commitment proofs (the root join of a runner epoch under each table, served from the runner's rows; a seal's agree-state opening and a join's last leaf at the three-level leaf height) and the epoch's settlement validity proof. NodeProofsTest in cartesi-rollups/contracts opens each commitment proof with the tournament's Commitment library (getRoot, and getRootForLastLeaf for last leaves) and validates the settlement with LibMachineValidityProof, as DaveConsensus stages it, requiring the outputs Merkle root the reference CLI wrote after echo's last accepted input. A flipped sibling fails either test. The CLI goldens gain each program's final state (what --final-hash prints) and echo's outputs Merkle root, and the runner test now also requires the settled final state to be the CLI's. The CLI invocation is factored so the epoch-end run shares it.
A node killed during a dispute-time build leaves the strata it stored (rows commit per build, all or nothing) and the boundaries its positioning wrote back. restarted_source_resumes_a_half_built_level drops a source after a lopsided partial descent of the big-cycle-root builder's active level in window 1, restarts one over the same state and work directories, and requires a fresh store's root, last-leaf proof and agree proof. This is the unit form of the kill_commitment_build e2e scenario, which the e2e cut can then drop.
Records what now stands between the legacy collector and the collect APIs: double-witnessed leaf and root goldens, the node's witnesses and proofs checked by the contracts, snapshot and restart tests, and work-count tests that would catch the switch hashing idle stretches. Still owed: the release corpus's uarch cases, the wrapper metamorphics, and keeping the test-only legacy builder on the goldens after the switch so regeneration stays double-witnessed.
The Dave corpus test compared only the 17 mcycle cases. It now also maps each uarch case (period 9, an epoch-wide period index) to the stride-0 leaf of height 29 it denotes, the big-cycle-root builder's path, and requires the node's facade to reproduce the released hash. 16 of the 18 agree, including the shapes no Dave program reaches: exception, halt, mcycle overflow and an unexpected manual yield. uarch-overflow-tail has no released hash. uarch-near-limit-tail is out of model, as an investigation settled against the Solidity step (artifacts kept outside the repo). Its template carries custom uarch code instead of the deployed step's pristine uarch. Every implementation assumes that uarch at big-cycle boundaries: Dave's run_big and idle replay, and the v0.21 collector, which also resets only dirty uarch words. Solidity, the CLI and Dave give three different roots. The node's per-step verbs and witness bytes match Solidity there. The test excludes the case but asserts its premise (a non-pristine template, the same released hash), so a corpus or template change forces a revisit. Two comments claimed the pristine uarch is checked at load, store and resume; only uarch_cycle == 0 is. They now say the template carries it as a trusted-app bound, which computation-hash.md and dimensioning.md state. The plan's exclusion wording is corrected: upstream 22b4431 makes the case error-no-hash, not Solidity-conformant.
Every commitment shortcut the node takes (whole big cycles on the big machine, one captured idle span for every later one) assumes the deployed step's pristine uarch at big-cycle boundaries. Every closing reset restores it, so only the template can break it: custom uarch code, or an image built by an emulator with another uarch. Over such a template the node's commitments were silently wrong, and a dispute would crash-loop on the cache's collision tripwire. Storage now refuses such a template when it imports it: a uarch reset must leave the template's root unchanged. That costs one private load at initialization. A unit test refuses a tiny template with a moved uarch pc, and the corpus test asserts that the real out-of-model template (uarch-near-limit-tail) is refused, replacing its ad hoc premise check. The comments and computation-hash.md now say where the check lives.
The owner accepted the harness designed by workflow wf_ad5ec6df-6d0: a cfg(test) module driving the honest node's real workers tick by tick over the devnet bundle, every wave in its own block, receipts checked, timestamps pinned; adversaries as the production Hero over a test-only tail overlay that never writes; idle divergence points only until a cheap geometry exists for active spans. The plan records the delivery order and which e2e scenarios move, drop or stay.
At the budget's seams the v0.21 CLI departs from the Solidity step, and the node was tied to the step there only through an emulator-log replay (seam 2) or its classification alone (seam 1). seam_witness_vectors_hold, a plain unit test with no images, now pins the node's own witness bytes for an RX_ACCEPTED opening on the budget's last cycle, the same on a halted machine (seam 2), and an RX_REJECTED closing on the budget's last cycle (seam 1), whose post-state must be the revert root. No input can reach seam 1 by execution, so, like the Lua vectors, the test shrinks the budget after a physical delivery and hands the node a machine stf with its pre-feed checkpoint. NodeWitnessesTest replays these vectors through CartesiStateTransition alongside the existing ones.
The big_input e2e scenario existed to push the largest input the InputBox accepts through the on-chain state transition. node_witness_vectors_hold now pins the node's witness for that opening (a 65,120-byte payload, 65,412 bytes EvmAdvance-encoded, a 93,868-byte proof) and NodeWitnessesTest replays it, so the e2e scenario can go once the harness ingests a large input.
Step 1 of the accepted harness design (test-strategy-reset.md, item 3). src/harness drives the honest node's real reader, runner and epoch manager tick by tick against anvil over the devnet bundle. The test owns the only clock: automine is off, each round's wave mines into its own block, block timestamps advance one second per block from the bundle's (InputBox stamps them into inputs), the gas limit fits a whole wave, finality is checked to be latest - 2, and every mined receipt is checked, so an honest transaction that reverts fails the test. Restart rebuilds the workers over the same state, as a process restart would. Four tests, about a second each: two consecutive epochs settle, the second with a maximum-size input; a restart after a lost and after a mined acceptance settles the epoch exactly once (kill_settle); and with a sentry that never claims, acceptance waits out the staging period. Production code changes only in visibility: BlockchainReader::tick, EpochManager::tick and discover_deployed_tournament become pub(crate). test_utils splits anvil spawning from app deployment. The tests run serially in their own process (`just test-node-harness`, added to CI's Rust job): the v0.21 emulator flocks files it creates without close-on-exec, so an anvil spawned by a concurrent test inherits the lock and later loads fail. That root cause, also behind the test_blockchain_reader flake, joins the upstream asks.
Step 2 of the harness design, first part. The adversary is the production Hero over a test-only tail overlay in DisputeSource: every ruler leaf at or past position D reads as Z, on every stride, so the adversary diverges from the honest node at exactly transition D at every level with no steering. A fully patched quartet folds Z to its height; a straddling one joins its children, whose honest values take the normal path. Patched values are never stored and prove_transition is untouched, so the adversary cannot prove a patched transition. Release builds carry none of it: the field, the two hooks in node and children, Tail itself and Hero::with_tail are all cfg(test). tail_overlay_diverges_at_its_position_and_never_writes checks the overlay on the toy: leaves are honest before D and Z from D on, at stride 0 and a coarser stride; proofs open the patched roots; and afterwards an honest source over the shared store answers every quartet like a fresh one. Two dispute tests over the empty epoch 0, diverging in idle material: - a joiner that goes silent loses its root match by timeout (bad_commitment); - a rational adversary disputes all three levels down to the leaf, where the node wins by STEP, then wins each inner tournament and settles (simple).
gc_match and gc_tournament now run against the node in-process. Two adversaries join the root before the node acts (Node::roll ingests and executes without the epoch manager), so they pair with each other, while the node joins unpaired. - the_node_collects_an_abandoned_match: both go silent after joining; the node deletes their match by timeout with no winner and wins the root. - the_node_collects_an_abandoned_child_tournament: they take their match into a child tournament, both join it, then stop; the node deletes the child match by timeout with no winner, then the parent match by its child, and wins the root. Both assert the exact MatchDeleted events (reason and winner on the paired match), as the Lua scenarios did. A first version let the node join while rolling, so an adversary paired with it instead; World::trace, a per-block listing of mined actions by key and verb, now prints whenever a run does not finish, which made that visible.
kill_join and kill_mid_match become restart_after_a_{lost,mined}_{join,
advance}_wins_the_dispute: during simple's dispute the node plans its join
(or its next advance) and dies with the action lost, or mined just before;
it stays down for ten blocks while the adversary plays on, restarts over
the same state, and still wins by STEP and settles. A repeated action would
revert and fail the run, so each branch also pins that the node does not
redo what landed.
multi_sybil becomes three_sybils_lose_and_the_node_recovers_its_bond_first:
two sybils play and one joins and goes silent, so the root holds two
concurrent matches, one sybil against sybil. It asserts the silent sybil's
single timeout deletion (never as winner), and, after a restart following
settlement, exactly one BondRecovered paid to the node's commitment before
its join to the next root, with the root's balance drained.
sealed_leaf_timeout_winner and _both become the_longer_clock_wins_a_sealed_leaf_by_timeout and both_clocks_expire_on_a_sealed_leaf. The setup plays simple's dispute with the adversary joining every child before the node, so it is commitment one and the leaf's final responder at odd heights, and with a new Hold policy the adversary keeps its leaf seal until three blocks before its own clock expires. The seal charges its overdue time, leaving its reserve far below the node's paused one; both deadlines are read back from commitmentStanding (they landed 385 blocks apart). The node stays offline from the seal on. - Winner: classifyMatchTimeout reads none at short - 1 and two-wins at the short deadline and past the retired classifier's midpoint; the node, back there, claims the timeout in the next block, before the long deadline. - Both: two-wins at long - 1, eliminate-both at the long deadline; the node deletes the match in the next block with no winner. Each asserts the exact MatchDeleted (timeout, with the expected winner) and that the node is commitment two. The Lua fixture needed sender hooks that killed and respawned the node around every child creation; turn order does it here. The plan records step 2 as done.
A four-lens review (workflow wf_3ff996b1-35a: overlay, soundness, parity with the Lua scenarios, production changes; each finding put to a skeptic) confirmed eleven findings, five blocking the e2e cut. All are addressed: - multi_sybil: join order now pairs the node with one sybil while the second meets the silent third, so the node plays two matches (one beside the sybil-vs-sybil match) and wins two leaves by STEP; the restart lands once its bond recovery has begun, before the epoch completes. - big_input: the ingested input must be the InputBox's (65,412 bytes, its keccak equal to getInputHash, the hash DaveConsensus checks). - Sealed-leaf timeouts: the node restarts as a fresh process at the boundary, as the Lua scenarios respawned it, before its claim. - The adversary works on a private copy of the node's store (VACUUM INTO), so the node computes its own subtrees and, after a restart, harvests the downtime's events from the chain instead of rows the adversary wrote. - simple's dispute now diverges at a closing slot, so the node's leaf proof carries a ustep and the uarch reset. - The overlay's guard test also runs over a store with window roots, where the frontier fold serves root-stride quartets. - Timestamps were not pinned (the first block after loading the state took the wall clock): anvil's clock is now set to the head's first; a run's hashes repeat exactly. - The pristine-uarch check is worded as what it compares: the linked emulator's uarch, assumed to be the deployed step's.
Item 4 of the test strategy reset, for the scenarios the in-crate harness
and the L1 tests now cover (after the harness's adversarial review closed
the parity gaps it found):
- removed scenarios: bad_commitment, big_input, deposit_withdrawal,
gc_match, gc_tournament, kill_catchup, kill_commitment_build, kill_join,
kill_mid_match, kill_settle, multi_sybil, sealed_leaf_timeout_{winner,both},
simple_no_input; the sealed-leaf helper and its clock probe; the dead
idle_strategy and dummy_commitment sybil helpers;
- removed recipes: test-kill-all, test-multi-sybil,
test-sealed-leaf-timeouts, and the root test-prt-timeout-boundaries;
test-honeypot-all and test-yield-all shrink to their remaining scenarios;
- the battery is the smoke set: echo simple, chaos and
kill_catchup_batched, honeypot simple and stf_all, yield stf_revert;
- the oracle's per-epoch CLI gate is gone: the runner, leaf and corpus
goldens check the node against the release CLI below e2e.
CI's e2e lane ran only kept scenarios and is unchanged. The Lua oracle
lineage stays, since the remaining sybils build from its epoch snapshots.
docs/test-harness.md now lists the smoke set and points to the harness.
The duration log started after context assembly, which is where a cold commitment build happens (it builds the local material), so the expensive part never reached the number operators compare with their budgets; a join waiting for finality built in a tick that logged nothing. The timer now starts with the tick: an action reports the time from reading the chain through builds and proving, and a tick that took a second or more without acting reports too. The README no longer claims the node runs at the emulator's own cost: work-count tests guard the algorithms, and a comparison with the emulator is a release measurement. (External review, 2026-10-02.)
From the external review of 3e22525..0a9976a (2026-10-02): - just test-node-harness verifies the devnet bundle's fingerprint first: the loader only checked that one existed, so the harness passed 15/15 against a bundle of the other geometry. - The maximum-size input was the old Lua scenario's workaround size, 65,120 bytes. The InputBox's limit applies to the EvmAdvance encoding (a 4-byte selector, nine head words, the payload padded to a word), so the maximum payload is 65,216 bytes. The witness vector is regenerated at that size and replays through the step; the harness settles it, ties the stored 65,508 bytes to InputBox's hash, and checks the boundary both ways (65,216 accepted, 65,217 refused) with simulated calls. - Wording made precise: the half-built-level test shows recovery from partially populated committed state, not an interruption mid-build; the Lua STEP gate lost its negative control with the sealed-leaf scenarios, and the per-PR STEP evidence now also comes from the harness and NodeWitnessesTest. The plan's runbook comparison carries the review's concrete specification.
Safe wrappers for cm_collect_mcycle_root_hashes and cm_collect_uarch_cycle_root_hashes (W6.1). The results keep every hash, offset, partial bundle, break reason and console error, and parsing refuses offsets that do not partition the hashes. Tests on the yield image tie the collectors to the per-step API: unbundled uarch periods equal hand-stepped ones (the idle period is the revert tail), mcycle samples equal sampled runs, calls split off the grid continue with phase and partial bundle, a rejection reports the revert root and then the tail, and a call without the tail is refused before anything executes.
The runbook recipe for the node's no-overhead claim (test-strategy-reset.md item 2): just measure-node-vs-emulator times a cold leaf join and a deep proof, each against the emulator doing the same work in process (cm_run to the position, then cm_collect_uarch_cycle_root_hashes bundled per big cycle, or the logged step and reset). The join targets a leaf three quarters into the gap's last input on a fresh store, so the whole gap replays; the proof is the closing slot ending that leaf. Each row runs in its own process for its own peak RSS, disk is the row's free-space delta, and the node's leaf root must equal the emulator's.
W6.2, narrowed to the one path that hashes per step at scale: the active big cycles of a dense leaf build. Ruler::collect_big_cycle_roots hands them to a new optional Stf verb, which MachineStf implements with cm_collect_uarch_cycle_root_hashes bundled per big cycle, so each period ends with the cycle's root. Idle stretches, positioning, sampled strides and proving keep their paths. The revert tail is the idle period collected at feed. Seam 1 stays stepped: v0.21.0's collector keeps the physical root for a rejection on the input budget's last cycle, so bulk collection declines that cycle and the ruler steps it. A control test pins the collector's answer there, to say when the guard can go. The stepped path stays selectable (DisputeSource::on_store_with, Collector::Stepped) as the permanent reference (D7): the leaf goldens and every dense corpus leaf are checked both ways, a run-level differential localizes disagreements, the collect API's bundles reduce to the unbundled leaves at every bundle size, and the toy drives the ruler's bulk branch through chunked and declined calls.
How dense leaves are built now and why the stepped path stays (computation-hash.md), the corpus gate building dense leaves both ways (build-system.md), the new differentials (test-harness.md), the migration's evidence and what remains (collect-hashes-migration.md), and W6's status with the seam-1 guard (two-level-sling.md).
From the switch's review: a delivery that leaves the machine at a fixed point (a template preset halted with an input yield pending, which the step still feeds) made the bulk collector return its one fixed-point period without progress, and big_cycle_roots failed where stepping builds the leaves. That period is the opening cycle's root. Also from the review: a ruler-level bulk-versus-stepped test at both seam openings (the budget's last cycle, the halted template), and a period-10 echo leaf around the first automatic yield in the run-level differential, so the real collector crosses chunks and re-enters after the yield. Chunks shrink to 256 cycles to keep that leaf cheap; a call's overhead is small next to a cycle's uarch work.
The emulator opens and flocks machine files without O_CLOEXEC (W7), so an anvil spawned while another test stores or clones a machine keeps that machine's files locked, and the next load fails. New store-heavy engine tests made test_blockchain_reader's long-known flake frequent (3 of 6 suite runs); skipping the anvil tests made it vanish (0 of 6). A close-on-exec sweep before each spawn did not help: the emulator opens short-lived locked files throughout a store, so the window between the sweep and the fork stays open. The seven tests that spawn anvil join the harness's serial run in just test-node-harness, which CI's Rust job runs.
Dispute positioning wrote back every input boundary it crossed with a full machine store: about 410 MiB of disk per boundary on the stress image, 25 GiB for a cold join at the end of a 64-input snapshot gap, and the stores' page reads put the join's peak RSS at 606 MiB. Every later action of a dispute stays inside one input, so only that input's boundary is load-bearing (R16). Positioning now crosses whole windows from the nearest stored boundary on the runner's copy-on-write clone chain, reusing its swap through kernels extracted from record_accepted and record_reverted, and publishes and registers only the target's boundary (no window roots, no gap GC). The target window is then positioned as before. Measured at full scale (just measure-node-vs-emulator): disk 26.6 MiB, peak RSS 116 MiB against the emulator's 111, join time unchanged at 1.03x. From the change's review: a torn gap snapshot that positioning skips could come back as a revert checkpoint, since a feed adopts an existing content-addressed directory without rehashing it, and a rejection would then reload a wrong state silently. The checkpoint now carries its revert root, and the reload asserts it. Tests: crossings register only their final boundary as clones of the start (inodes) and leave no working clone, an all-rejected crossing reuses the floor, the gap test hides the epoch start so its floor is proved, a torn checkpoint fails loudly, and an abandoned crossing leaves no trace.
docs/measurements/node-vs-emulator.md is the full-scale run after both changes (64-input gap, height-37 leaf, stress image): the cold join at 1.03x the emulator's time, 116 MiB peak RSS against 111, 26.6 MiB of disk. The plans record the sequence: 1.18x with the per-step builder, 1.03x with the bulk collector, and the CoW crossing taking the join's disk from 25 GiB to 27 MiB.
From the external review of 0a9976a..c9c1f70. The safe wrappers checked only that a bundle exponent fits int32, but v0.21.0 computes with it before validating: a uarch exponent above 20 makes a negative shift (undefined behavior), and mcycle bundle construction does arithmetic before its range check. Exponents past 20 (uarch) and 63 (mcycle) now return INVALID_ARGUMENT, and a test checks the machine is left unchanged. The node only ever passes valid sizes.
From the same review: the report now names the source revision (git describe, dirty when uncommitted) and the workload's fingerprint line, so later runs can be compared.
From the same review: an input fed at mcycle 2^64 - 2 saturates its budget, so its first cycle is its budget's last, the seam-1 guard declines it, and the ruler's assert stops the node where stepping would proceed. Reaching it takes a template preset there or centuries of machine time, so it is documented, with the assert naming the cause, rather than handled.
The battery-cleanup regression test still expected the pre-cut battery's 25 scenarios, so since 0a9976a it failed in CI's build job, before the Rust, engine and harness steps, which have not run on the branch since. It now counts the battery's SCENARIOS list, so a cut or an addition cannot leave it stale.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.