Skip to content

Fix/prt timers - #287

Draft
GCdePaula wants to merge 69 commits into
mainfrom
fix/prt-timers
Draft

GCdePaula wants to merge 69 commits into
mainfrom
fix/prt-timers

Conversation

@GCdePaula

Copy link
Copy Markdown
Member

No description provided.

Add deterministic FFI vectors, logged from a bare emulator, for a manual
yield that the step lets outrank halt and the input budget:
an RX_ACCEPTED opening on the budget's last cycle (also at the maximum
mcycle) or on a preset halted machine delivers the input and differs from
the idle opening, and an RX_REJECTED closing on the budget's last cycle
still substitutes the revert root. SendCmioResponse and UArchReset read
only the pending yield, so these are the canonical answers the node and
Lua client must match.
node_metadata now records a COMMITMENT_SEMANTICS version next to the node
version and schema fingerprint, and startup refuses a store stamped with
another value. The crate version is frozen, so a change to leaf values or
transition shapes that leaves schema.sql alone would otherwise reuse
stores built under the old rules; a schema revert or backport could even
restore an old fingerprint. Every commitment change bumps the stamp.
The step reads only the pending manual yield: SendCmioResponse delivers the
next input to any RX_ACCEPTED yield and UArchReset reverts any RX_REJECTED
one, whatever the halt flag or the input budget. The node and the Lua
client checked halt and mcycle overflow first, so an input yield executed
on the budget's last cycle (mcycle == imcyclemax), or preset on a halted
template, was padded as terminal. At the next fed window the node's leaf
then differed from the canonical one and its own proof, which logs the
delivery, contradicted it: an adversary matching the node up to that leaf
could win with an arbitrary suffix.

Terminal now means an exception or unexpected manual yield, or halt or
overflow with no manual yield pending, in both clients. The docs that
stated the old rule are corrected, and COMMITMENT_SEMANTICS moves to 2.
The hand-written patch chains had stopped reaching their targets: the
2026-07 chains used 2^28 links against a level-1 stride of 27, and Lua's
64-bit integers turned stf_all's (1 << 68) into 0. A baseline run showed
stf_all epochs 2 to 4 and stf_revert sealing on other closing slots than
the ones they document, so the on-chain STF coverage they claim was not
being exercised.

Env.steering_patches now builds each chain in 256-bit arithmetic from the
deployed level table, and Env.run_steered_epoch asserts that the dispute's
leaf match sealed exactly on the steered transition, reading its divergence
cycle from sealedMatch pinned at the LeafMatchSealed block. stf_all,
stf_revert, and big_input use it; every epoch now seals where intended,
including input 1's fused feed at 2^68, the first dispute past window 0.
The node compiled in the root stride (LOG2_STRIDE = 44) and refused any
other deployment. It now reads the factory's whole level table at startup,
validates it with engine::TournamentGeometry (the root spans the 92-bit
ruler, levels tile, the leaf stride is zero, the root stride lies between
one big cycle and one input window), and pins it with the consensus address
in sling_config. The runner, the window-root rows, the roll's prefix check,
the settlement fold, and the dispute facade all use the pinned root stride,
and settlement roots stay byte-identical under the canonical [44,27,0]
table.

Table stability is a trust assumption of the parameters provider, so the
Hero checks every tournament descriptor on its path against the pinned row
for its level, and the root commitment it would join with against the
settled computation hash. The frontier fold now serves only the run stride:
a coarser tree samples other states and was served wrongly before, though
no production level asked for one.

The e2e oracle and the gc, multi_sybil, and bad_commitment scenarios take
their strides and heights from chain as well.
DEVNET_GEOMETRY=two-level builds a devnet whose factory serves the selected
[37,0]/[55,37] table ahead of the canonical switch. A test-only
TableTournamentParametersProvider (validated on-chain by the existing
table validator) and a devnet-only script that inherits the production
DeploymentScript through a separate entry point replace only the
parameters provider; production scripts and src/ are untouched. The bundle
marker records the geometry, and every consumer verifies against the
DEVNET_GEOMETRY it runs with, so a bundle of one geometry never passes for
another. `just test-rollups-two-level-smoke` runs echo `simple` on it.
A two-level leaf commitment spans 2^37 transitions, and compute_and_store
built it leaf by leaf: every run and every Merkle node of the span stayed
in memory, and idle stretches expanded into their churn runs cycle by
cycle. The two-level smoke lost its dispute by timeout while the node was
still building.

A stride-0 quartet tall enough that its 8 stored levels stay above
big-cycle granularity now takes Ruler::collect_big_cycle_roots: each big
cycle folds to its subtree root as it completes, and an idle stretch steps
one captured cycle and repeats its root. The quartet's tree is built c
levels up from those roots; its root and stored fanout are unchanged.
Shorter quartets keep the leaf-by-leaf path, which also serves as the
differential oracle.

Tests pin both: the folded roots against the brute-force oracle over every
big-aligned span, fresh and resumed; every stored fanout row of a tall
quartet, including one exactly at the threshold height; a real machine
across an active and an idle cycle; and an idle stretch costing one
stepped cycle.
The Hero's dispute loop runs inside the epoch manager's async task, and
its commitment builds and proofs are synchronous machine work. A leaf build
of many minutes held a tokio worker the whole time, and tasks queued on
that worker, the chain-facing reader included, could starve behind it.

hero::machine_work wraps level_material and action preparation in
block_in_place on the multi-threaded runtime, so the worker's queue moves
to a replacement first; a current-thread runtime runs the work inline. A
test on a one-worker runtime shows a queued task running during the wait,
and it fails without the hand-off. The manager itself still waits for the
result; node-architecture.md debt 7 records what that still delays.
After the per-PR e2e battery, the e2e job rebuilds the devnet with
DEVNET_GEOMETRY=two-level and runs echo `simple` against it, so the node
keeps winning a dispute on the selected [37,0]/[55,37] table before the
canonical constants switch. It runs last because it replaces the devnet
bundle. sealed_leaf_timeout_* stays canonical-only: its kill choreography
depends on each level's height parity.
The e2e oracle and the node cross-check each other, but both are Dave
code. The plan makes the released v0.21.0 cartesi-machine CLI the reference
implementation, so every epoch the oracle computes is now recomputed by the
CLI from the oracle's epoch snapshot and the chain inputs: its mcycle
computation hash at period (root stride - 20), read from the deployed
table, must equal the oracle's commitment, which the node's already must
match. The CLI runs on a disposable clone in stored revert mode (its
default fork mode needs a machine server), and a mismatch keeps the inputs,
the CLI log, and a repro.sh under _oracle/cli-epoch-<n>/. The gate refuses
any other CLI version.

Checked on canonical echo `simple`, yield `stf_revert` (three rejected
inputs), and the two-level smoke, where the period-17 CLI hash of a
four-input epoch equals the node's root.
A review of the whole stack, each finding adversarially verified, found no
bug in the two-level logic. It found a protocol flaw (inner commitment
construction drains the correct party's clock across Sybil matches), the
unmeasured scale properties of dense height-37 leaves, and several
robustness and evidence gaps. The plan now tracks them as R1-R15.

test-harness.md claimed every scenario runs on either table; only echo
`simple` has run on two levels, and steered STF scenarios need dense leaves
the devnet clock cannot host. The plan also named v0.21.1-test1's commits
unrelated; one is the O_CLOEXEC fix behind the flaky node unit tests.
A correct commitment must build its child commitment before it can join a
child tournament, and the join's lateness is charged to its clock. The
winner's remainder then replaced the parent clock and nothing refilled it,
so every sealed parent match against a new Sybil, diverging at a fresh span,
cost the correct party one build. Once the censorship budget C was spent, a
few Sybil bonds eliminated it: two with dense two-level leaves, and a
handful on the checked-in table too, since the node's finality-gated joins
cost minutes each.

The rule is now that every honest action gets one inclusion G, and joining a
child also gets the build T. winInnerTournament refills the returning winner
by up to T + 2G (build, join, propagation), capped by the sealed pair's
envelope max(r1, r2), the child's own allowance. No clock exceeds that
envelope and the refill never lowers a winner, so no clock mass is created
and no new elimination path opens; the adversary gains at most T + 2G of
lingering per child it pays a bond for.

T (commitmentBudget) is a new TournamentParameters and clone-argument field.
It belongs with the geometry: ArbitrationConstants.COMMITMENT_BUDGET is 30
minutes for the checked-in table, and the two-level devnet profile uses 60.
The new ClockBudgets library derives every block budget from wall-clock
inputs, including maxAllowance = C + G + (L - 1)(T + 2G): the root join's
inclusion plus one pending delegation per inner level. The canonical
provider takes the block time, C and G, and each chain kind states only its
censorship budget (mainnet one week, testnet 8 hours, devnet none). Mainnet
moves from one week plus 60 minutes to one week plus 85. The test-only table
validator now requires uniform budgets and a root allowance that holds the
pending delegations.

Tests: a correct commitment joining within T + G and propagating within G
returns with its pre-seal clock, delegation after delegation, from either
side of the envelope; the property fails on the old rule (193 != 200), and
mutating the refill, the envelope wiring, or the inner clones' T is caught.
Deployment and devnet-profile tests deploy the providers from the scripts'
own encodings. Gas allocations are unchanged: the inner seal and inner win
still fit theirs, with recorded headroom reduced (6k and 41k), and the
leaf-win alternate now ties the selected witness. The accepted calibration
under the pinned release Forge remains to be recorded.

This is a new deployment generation: the factory ABI and the clone-argument
encoding change, and Tournament bytecode changes (its ABI and storage do
not). The e2e decoders follow.
The measurement report, its generated documents, and the harness comments
quoted the old allowance literals (devnet one hour, mainnet one week plus one
hour) and a deployment helper that no longer exists. They now point to
ClockBudgets and state the derived values. The plan records how R1 was
fixed and the follow-ups under the same rule: a G discount for the leaf
proof and for a paused winner's timeout cleanup, and inner joins from the
latest block.
The plan now also tracks the node's honest-action latency (proofs replay
from window boundaries, a 30 s tick, one action per tick) and the unpropagated
Sybil-versus-Sybil child winners, and records two decisions: joins stay
finality-gated, since the Hero builds on the latest view and T + G covers
max(build, finality) + inclusion; and the node does not pin or judge T,
leaving that to operators. R13 notes that honeypot and the CLI gate on its
image now pass.

A new backlog keeps the tooling friction observed in this campaign, so a
later cleanup starts from evidence.
The constants benchmark and generator named T the "inner timeout", while the
contracts and the clock model now call it the commitment budget: the time
every inner tournament's commitment must build within. The Lua benchmark
reads DAVE_COMMITMENT_BUDGET_MINUTES, the Rust generator takes
--commitment-budget-minutes, and their reports, the node's capacity warning,
and the living docs follow.
Under the rule that every honest action gets G, a response already earned a
non-bankable discount, but a correct party that won a leaf proof or a timeout
claim paid the action's full cost, and the adversary chooses how many such
wins it must make inside one child. Clock.pauseWinnerAt now charges a winner
its live cost (time run plus any deferred charge) beyond one responseBudget,
never above its stored balance. Eliminating both sides earns nothing.

Survival is still decided on the full cost by the unchanged timeout
classifier, so the node and the Lua client, which read that view, are
unaffected. winMatchByTimeout checks for a finished tournament inline to
decode its arguments once.

The lifecycle, recursive, four-level, and bounded-delay oracles follow the
rule. In the finite delay model a pair can now end with a free proof inside
G, so its survivor re-pairs with a full clock: each later sequential pair
may take max(g - 1, 0) blocks more, and the replayed maximum witness ends at
block 20 instead of 19. The coarse per-match bound is unchanged. The four-
level trace keeps one refill below the envelope so it still pins T + 2G.
Tournament ABI and storage are unchanged.
The review of 2c502f6 and 935dc13 is kept verbatim as a frozen record, and
the reviews index gains it and the missing 2026-08-27 recalibration entry.

The plan takes over its open work. CF-01 stays an accepted limitation, with
its reasoning restated: the adversary chooses how often a proof is overtaken
by the opponent's expiry, but not the latencies, which cost nothing inside G
and are censorship beyond it; R16's latency target now covers the fallback
claim. R18 owns the STF evidence gate and R19 the honest-survival model. W2.1
and W5.1 stop overclaiming the proven transitions and the gas calibration.
The harness counted a sealed leaf transition as proved, but a timeout win
leaves the same seal and settlement behind, and the node claims a timeout
before retrying a rejected proof, so a broken on-chain state transition could
pass stf_all, stf_revert, or big_input (R18, CF-02 in the 2026-09-29 clock
refill review).

The reader now returns each sealed leaf's match ID hash with its transition,
and Env.assert_leaf_match_proved requires exactly one MatchDeleted with
reason STEP for it. run_steered_epoch applies it after settlement, and
simple opts in, so the per-PR smokes, the two-level one included, gate the
transition too. run_epoch stays neutral, since kill and chaos runs may end a
leaf on a timeout. sealed_leaf_timeout_winner is the negative control: its
timeout-won record must correlate with the fixture's match and be refused.
The Hero's checks against the pinned tournament table and the epoch-start
state returned errors that the manager logged and retried every tick, so a
node that hit one stopped defending without alarm (R8). Both are invariant
violations, not transient failures, so they now panic, per the node's own
doctrine.

The check that the root commitment equals the settled computation hash is
gone. It compared two of the node's own folds, whose equality is unit-tested
under both tables; the staging assert still catches a mismatch on a win; and
when the settlement fold was the wrong one, it made the node refuse a
dispute it would have won.
More than 2^24 inputs in one epoch is an economic bound, a flood of roughly
5e11 gas, where the contracts drop the tail and the node panics; a clearer
error would still stop the node, so D5 builds nothing. Proof positioning
replays one input's prefix from the boundary the dispute wrote back, bounded
by the per-input contract, so R16 needs no mid-input snapshots; W4.7 notes
that the snapshot gap bounds the first root response inside an input. R9
folds into W6.4, which reads yields from registers anyway, and W6.4 states
that seam 1 needs a 2^48-cycle input and recommends no collector guard.

The canonical switch becomes a follow-up PR that should reduce to small
Solidity changes.
The test-only TournamentGeometry::canonical() hard-coded the three-level
table, so the day ArbitrationConstants switches to two levels every test
calling it would silently keep testing the old geometry. Tests that need a
particular table now name it (three_level or two_level), and the drift guard
and the devnet discovery test read the checked-in table itself through
TournamentGeometry::checked_in(), which parses ArbitrationConstants.sol.

The storage tests pin two levels, so stride 37 runs in this PR rather than
first in the switch.
…able

The deployment and DaveAppFactory tests asserted three levels and a 30-minute
commitment budget as literals, so switching ArbitrationConstants would have
broken them in the switch PR. They now derive the pending refills and the
commitment blocks from ArbitrationConstants, and the canonical-geometry tests
compare against ClockBudgets. The zero-allowance case uses a block time
longer than any budget, since at T = 60 min a one-hour block still rounds the
commitment budget to one.

testCheckedInCanonicalTable stays the one deliberate pin of the table.
M2, the new `just measure-two-level-leaf`, built the dense two-level leaf
(stride 0, height 37, the stress workload's first input) in 8,405 s against a
30-minute target, although its density matches steady state. The emulator
fixes its hash-tree concurrency at load and hashes in parallel once a hash's
dirty pages outnumber the host's cores; at one root hash per ustep the thread
pool's fork and join then cost about 110 us against 12 us of hashing.

Rulers now say how they will hash (Hashing::PerStep for a stride-0 build,
Sampled otherwise), and the node loads per-step machines with serial
hash-tree updates while sampled strides, which hash large dirty sets, keep
the parallel path. The Lua client, whose commitments hash per ustep, loads
serially throughout. M2 now takes 1,100 s with 103 MiB peak RSS, and the
bounded-memory leaf build holds.

The plan records M2, adds pages per hash as a second dimension to R2, and
asks upstream for a work-based parallelism threshold.
The scenario asserted the three-level heights 48/17/27 and assumed the sybil
seals every inner match, which holds only at an even-height root. At the
two-level root (55, odd) the node joins first and seals, so the sybil would
never have become the leaf's final responder.

The scenario now reads the level table, requires odd heights below the root,
and expects the sybil to seal an even-height root and every inner level
below. When the node seals a child's parent, the node is stopped across the
sybil's join of that child, as it already was around the sybil's own seals.
Both scenarios pass on the canonical and the two-level devnets.
The big-cycle-root builder, which builds every stride-0 quartet of height 28
or more, had real-machine differentials only over idle padding. A span over
the first 2^8 big cycles of window 1, where echo runs the input, now checks
its root, children and proofs against the leaf-by-leaf prototype (R12).
The constants report printed the three-level table as the "(current)" row and
warned when a derived root was taller than it, and the `--full` help named
its 2^44 window as if it were every table's. The bench's pinned table is now
named for what it is, the report no longer ranks derivations against it, and
the bump checklist includes COMMITMENT_BUDGET.
The docs, contract guardrails and CI comments described the three-level table
as the checked-in one and the two-level table as planned or not live, so the
canonical switch would have had to rewrite them. They now state both tables
and point at ArbitrationConstants for the deployed one, and the CLI gate
note says what it samples at each root stride.

The plan records the switch as a follow-up PR of small Solidity changes
(ArbitrationConstants with COMMITMENT_BUDGET = 60 min, plus the one pinning
test), the two-level dry run of the steered STF scenarios, which needs no
dense devnet profile (R6), and R12 done.
1c6018d loaded the constants harness's one machine per-step, so the
hash-cost curve, which prices the sampled root stride, would have been
measured with serial hashing that production does not use there. The dense
leaf rate and the curve now come from separate machines, loaded per-step and
sampled respectively. The Lua constants harness hashed its leaf phase with
the emulator's default parallel updates and now loads it serially, like both
clients do.
Gaps 3 and 4 of the test strategy reset, for Dave-owned cases. The released
CLI's uarch cycle computation hash over one mcycle period is Dave's stride-0
commitment under a root leaf at stride period + 20, computed independently
and through the collect API. reference_cli_goldens_hold now also records nine
such leaves on echo and yield, and leaf_commitments_match_the_reference_cli
(engine gate, about 23 s) requires the node's facade to reproduce them on a
fresh store.

Period 8 exercises the big-cycle-root builder (height 28, what a switch to
the collect API replaces); period 7 the plain collection path (height 27, the
three-level leaf). The periods hold a fed window start, an accepted yield, a
rejection's revert on each program, a revert crossed while positioning, and
a padding window. All nine agree. Period 17 is left out: the CLI spends about
2 minutes on a mostly idle period-17 leaf and more than 13 on a dense one.
The CLI invocation is shared between the root and leaf goldens.
Gap 5 of the test strategy reset. node_proof_vectors_hold pins the node's
commitment proofs (the root join of a runner epoch under each table, served
from the runner's rows; a seal's agree-state opening and a join's last leaf
at the three-level leaf height) and the epoch's settlement validity proof.
NodeProofsTest in cartesi-rollups/contracts opens each commitment proof with
the tournament's Commitment library (getRoot, and getRootForLastLeaf for last
leaves) and validates the settlement with LibMachineValidityProof, as
DaveConsensus stages it, requiring the outputs Merkle root the reference CLI
wrote after echo's last accepted input. A flipped sibling fails either test.

The CLI goldens gain each program's final state (what --final-hash prints)
and echo's outputs Merkle root, and the runner test now also requires the
settled final state to be the CLI's. The CLI invocation is factored so the
epoch-end run shares it.
A node killed during a dispute-time build leaves the strata it stored (rows
commit per build, all or nothing) and the boundaries its positioning wrote
back. restarted_source_resumes_a_half_built_level drops a source after a
lopsided partial descent of the big-cycle-root builder's active level in
window 1, restarts one over the same state and work directories, and
requires a fresh store's root, last-leaf proof and agree proof. This is the
unit form of the kill_commitment_build e2e scenario, which the e2e cut can
then drop.
Records what now stands between the legacy collector and the collect APIs:
double-witnessed leaf and root goldens, the node's witnesses and proofs
checked by the contracts, snapshot and restart tests, and work-count tests
that would catch the switch hashing idle stretches. Still owed: the release
corpus's uarch cases, the wrapper metamorphics, and keeping the test-only
legacy builder on the goldens after the switch so regeneration stays
double-witnessed.
The Dave corpus test compared only the 17 mcycle cases. It now also maps
each uarch case (period 9, an epoch-wide period index) to the stride-0 leaf
of height 29 it denotes, the big-cycle-root builder's path, and requires
the node's facade to reproduce the released hash. 16 of the 18 agree,
including the shapes no Dave program reaches: exception, halt, mcycle
overflow and an unexpected manual yield. uarch-overflow-tail has no
released hash.

uarch-near-limit-tail is out of model, as an investigation settled against
the Solidity step (artifacts kept outside the repo). Its template carries
custom uarch code instead of the deployed step's pristine uarch. Every
implementation assumes that uarch at big-cycle boundaries: Dave's run_big
and idle replay, and the v0.21 collector, which also resets only dirty uarch
words. Solidity, the CLI and Dave give three different roots. The node's
per-step verbs and witness bytes match Solidity there. The test excludes
the case but asserts its premise (a non-pristine template, the same
released hash), so a corpus or template change forces a revisit.

Two comments claimed the pristine uarch is checked at load, store and
resume; only uarch_cycle == 0 is. They now say the template carries it as a
trusted-app bound, which computation-hash.md and dimensioning.md state. The
plan's exclusion wording is corrected: upstream 22b4431 makes the case
error-no-hash, not Solidity-conformant.
Every commitment shortcut the node takes (whole big cycles on the big
machine, one captured idle span for every later one) assumes the deployed
step's pristine uarch at big-cycle boundaries. Every closing reset restores
it, so only the template can break it: custom uarch code, or an image built
by an emulator with another uarch. Over such a template the node's
commitments were silently wrong, and a dispute would crash-loop on the
cache's collision tripwire.

Storage now refuses such a template when it imports it: a uarch reset must
leave the template's root unchanged. That costs one private load at
initialization. A unit test refuses a tiny template with a moved uarch pc,
and the corpus test asserts that the real out-of-model template
(uarch-near-limit-tail) is refused, replacing its ad hoc premise check. The
comments and computation-hash.md now say where the check lives.
The owner accepted the harness designed by workflow wf_ad5ec6df-6d0: a
cfg(test) module driving the honest node's real workers tick by tick over
the devnet bundle, every wave in its own block, receipts checked, timestamps
pinned; adversaries as the production Hero over a test-only tail overlay
that never writes; idle divergence points only until a cheap geometry exists
for active spans. The plan records the delivery order and which e2e
scenarios move, drop or stay.
At the budget's seams the v0.21 CLI departs from the Solidity step, and the
node was tied to the step there only through an emulator-log replay (seam 2)
or its classification alone (seam 1). seam_witness_vectors_hold, a plain
unit test with no images, now pins the node's own witness bytes for an
RX_ACCEPTED opening on the budget's last cycle, the same on a halted machine
(seam 2), and an RX_REJECTED closing on the budget's last cycle (seam 1),
whose post-state must be the revert root. No input can reach seam 1 by
execution, so, like the Lua vectors, the test shrinks the budget after a
physical delivery and hands the node a machine stf with its pre-feed
checkpoint. NodeWitnessesTest replays these vectors through
CartesiStateTransition alongside the existing ones.
The big_input e2e scenario existed to push the largest input the InputBox
accepts through the on-chain state transition. node_witness_vectors_hold now
pins the node's witness for that opening (a 65,120-byte payload, 65,412 bytes
EvmAdvance-encoded, a 93,868-byte proof) and NodeWitnessesTest replays it,
so the e2e scenario can go once the harness ingests a large input.
Step 1 of the accepted harness design (test-strategy-reset.md, item 3).
src/harness drives the honest node's real reader, runner and epoch manager
tick by tick against anvil over the devnet bundle. The test owns the only
clock: automine is off, each round's wave mines into its own block, block
timestamps advance one second per block from the bundle's (InputBox stamps
them into inputs), the gas limit fits a whole wave, finality is checked to
be latest - 2, and every mined receipt is checked, so an honest transaction
that reverts fails the test. Restart rebuilds the workers over the same
state, as a process restart would.

Four tests, about a second each: two consecutive epochs settle, the second
with a maximum-size input; a restart after a lost and after a mined
acceptance settles the epoch exactly once (kill_settle); and with a sentry
that never claims, acceptance waits out the staging period.

Production code changes only in visibility: BlockchainReader::tick,
EpochManager::tick and discover_deployed_tournament become pub(crate).
test_utils splits anvil spawning from app deployment. The tests run
serially in their own process (`just test-node-harness`, added to CI's Rust
job): the v0.21 emulator flocks files it creates without close-on-exec, so
an anvil spawned by a concurrent test inherits the lock and later loads
fail. That root cause, also behind the test_blockchain_reader flake, joins
the upstream asks.
Step 2 of the harness design, first part. The adversary is the production
Hero over a test-only tail overlay in DisputeSource: every ruler leaf at or
past position D reads as Z, on every stride, so the adversary diverges from
the honest node at exactly transition D at every level with no steering.
A fully patched quartet folds Z to its height; a straddling one joins its
children, whose honest values take the normal path. Patched values are
never stored and prove_transition is untouched, so the adversary cannot
prove a patched transition. Release builds carry none of it: the field, the
two hooks in node and children, Tail itself and Hero::with_tail are all
cfg(test).

tail_overlay_diverges_at_its_position_and_never_writes checks the overlay
on the toy: leaves are honest before D and Z from D on, at stride 0 and a
coarser stride; proofs open the patched roots; and afterwards an honest
source over the shared store answers every quartet like a fresh one.

Two dispute tests over the empty epoch 0, diverging in idle material:
- a joiner that goes silent loses its root match by timeout (bad_commitment);
- a rational adversary disputes all three levels down to the leaf, where the
  node wins by STEP, then wins each inner tournament and settles (simple).
gc_match and gc_tournament now run against the node in-process. Two
adversaries join the root before the node acts (Node::roll ingests and
executes without the epoch manager), so they pair with each other, while
the node joins unpaired.
- the_node_collects_an_abandoned_match: both go silent after joining; the
  node deletes their match by timeout with no winner and wins the root.
- the_node_collects_an_abandoned_child_tournament: they take their match
  into a child tournament, both join it, then stop; the node deletes the
  child match by timeout with no winner, then the parent match by its child,
  and wins the root.
Both assert the exact MatchDeleted events (reason and winner on the paired
match), as the Lua scenarios did. A first version let the node join while
rolling, so an adversary paired with it instead; World::trace, a per-block
listing of mined actions by key and verb, now prints whenever a run does
not finish, which made that visible.
kill_join and kill_mid_match become restart_after_a_{lost,mined}_{join,
advance}_wins_the_dispute: during simple's dispute the node plans its join
(or its next advance) and dies with the action lost, or mined just before;
it stays down for ten blocks while the adversary plays on, restarts over
the same state, and still wins by STEP and settles. A repeated action would
revert and fail the run, so each branch also pins that the node does not
redo what landed.

multi_sybil becomes three_sybils_lose_and_the_node_recovers_its_bond_first:
two sybils play and one joins and goes silent, so the root holds two
concurrent matches, one sybil against sybil. It asserts the silent sybil's
single timeout deletion (never as winner), and, after a restart following
settlement, exactly one BondRecovered paid to the node's commitment before
its join to the next root, with the root's balance drained.
sealed_leaf_timeout_winner and _both become
the_longer_clock_wins_a_sealed_leaf_by_timeout and
both_clocks_expire_on_a_sealed_leaf. The setup plays simple's dispute with
the adversary joining every child before the node, so it is commitment one
and the leaf's final responder at odd heights, and with a new Hold policy
the adversary keeps its leaf seal until three blocks before its own clock
expires. The seal charges its overdue time, leaving its reserve far below
the node's paused one; both deadlines are read back from commitmentStanding
(they landed 385 blocks apart). The node stays offline from the seal on.
- Winner: classifyMatchTimeout reads none at short - 1 and two-wins at the
  short deadline and past the retired classifier's midpoint; the node, back
  there, claims the timeout in the next block, before the long deadline.
- Both: two-wins at long - 1, eliminate-both at the long deadline; the node
  deletes the match in the next block with no winner.
Each asserts the exact MatchDeleted (timeout, with the expected winner) and
that the node is commitment two. The Lua fixture needed sender hooks that
killed and respawned the node around every child creation; turn order does
it here. The plan records step 2 as done.
A four-lens review (workflow wf_3ff996b1-35a: overlay, soundness, parity
with the Lua scenarios, production changes; each finding put to a skeptic)
confirmed eleven findings, five blocking the e2e cut. All are addressed:
- multi_sybil: join order now pairs the node with one sybil while the
  second meets the silent third, so the node plays two matches (one beside
  the sybil-vs-sybil match) and wins two leaves by STEP; the restart lands
  once its bond recovery has begun, before the epoch completes.
- big_input: the ingested input must be the InputBox's (65,412 bytes, its
  keccak equal to getInputHash, the hash DaveConsensus checks).
- Sealed-leaf timeouts: the node restarts as a fresh process at the
  boundary, as the Lua scenarios respawned it, before its claim.
- The adversary works on a private copy of the node's store (VACUUM INTO),
  so the node computes its own subtrees and, after a restart, harvests the
  downtime's events from the chain instead of rows the adversary wrote.
- simple's dispute now diverges at a closing slot, so the node's leaf
  proof carries a ustep and the uarch reset.
- The overlay's guard test also runs over a store with window roots, where
  the frontier fold serves root-stride quartets.
- Timestamps were not pinned (the first block after loading the state took
  the wall clock): anvil's clock is now set to the head's first; a run's
  hashes repeat exactly.
- The pristine-uarch check is worded as what it compares: the linked
  emulator's uarch, assumed to be the deployed step's.
Item 4 of the test strategy reset, for the scenarios the in-crate harness
and the L1 tests now cover (after the harness's adversarial review closed
the parity gaps it found):
- removed scenarios: bad_commitment, big_input, deposit_withdrawal,
  gc_match, gc_tournament, kill_catchup, kill_commitment_build, kill_join,
  kill_mid_match, kill_settle, multi_sybil, sealed_leaf_timeout_{winner,both},
  simple_no_input; the sealed-leaf helper and its clock probe; the dead
  idle_strategy and dummy_commitment sybil helpers;
- removed recipes: test-kill-all, test-multi-sybil,
  test-sealed-leaf-timeouts, and the root test-prt-timeout-boundaries;
  test-honeypot-all and test-yield-all shrink to their remaining scenarios;
- the battery is the smoke set: echo simple, chaos and
  kill_catchup_batched, honeypot simple and stf_all, yield stf_revert;
- the oracle's per-epoch CLI gate is gone: the runner, leaf and corpus
  goldens check the node against the release CLI below e2e.
CI's e2e lane ran only kept scenarios and is unchanged. The Lua oracle
lineage stays, since the remaining sybils build from its epoch snapshots.
docs/test-harness.md now lists the smoke set and points to the harness.
The duration log started after context assembly, which is where a cold
commitment build happens (it builds the local material), so the expensive
part never reached the number operators compare with their budgets; a join
waiting for finality built in a tick that logged nothing. The timer now
starts with the tick: an action reports the time from reading the chain
through builds and proving, and a tick that took a second or more without
acting reports too. The README no longer claims the node runs at the
emulator's own cost: work-count tests guard the algorithms, and a
comparison with the emulator is a release measurement. (External review,
2026-10-02.)
From the external review of 3e22525..0a9976a (2026-10-02):
- just test-node-harness verifies the devnet bundle's fingerprint first:
  the loader only checked that one existed, so the harness passed 15/15
  against a bundle of the other geometry.
- The maximum-size input was the old Lua scenario's workaround size,
  65,120 bytes. The InputBox's limit applies to the EvmAdvance encoding (a
  4-byte selector, nine head words, the payload padded to a word), so the
  maximum payload is 65,216 bytes. The witness vector is regenerated at
  that size and replays through the step; the harness settles it, ties the
  stored 65,508 bytes to InputBox's hash, and checks the boundary both ways
  (65,216 accepted, 65,217 refused) with simulated calls.
- Wording made precise: the half-built-level test shows recovery from
  partially populated committed state, not an interruption mid-build; the
  Lua STEP gate lost its negative control with the sealed-leaf scenarios,
  and the per-PR STEP evidence now also comes from the harness and
  NodeWitnessesTest. The plan's runbook comparison carries the review's
  concrete specification.
Safe wrappers for cm_collect_mcycle_root_hashes and
cm_collect_uarch_cycle_root_hashes (W6.1). The results keep every
hash, offset, partial bundle, break reason and console error, and
parsing refuses offsets that do not partition the hashes.

Tests on the yield image tie the collectors to the per-step API:
unbundled uarch periods equal hand-stepped ones (the idle period
is the revert tail), mcycle samples equal sampled runs, calls split
off the grid continue with phase and partial bundle, a rejection
reports the revert root and then the tail, and a call without the
tail is refused before anything executes.
The runbook recipe for the node's no-overhead claim
(test-strategy-reset.md item 2): just measure-node-vs-emulator times a
cold leaf join and a deep proof, each against the emulator doing the
same work in process (cm_run to the position, then
cm_collect_uarch_cycle_root_hashes bundled per big cycle, or the
logged step and reset). The join targets a leaf three quarters into
the gap's last input on a fresh store, so the whole gap replays; the
proof is the closing slot ending that leaf. Each row runs in its own
process for its own peak RSS, disk is the row's free-space delta, and
the node's leaf root must equal the emulator's.
W6.2, narrowed to the one path that hashes per step at scale: the
active big cycles of a dense leaf build. Ruler::collect_big_cycle_roots
hands them to a new optional Stf verb, which MachineStf implements with
cm_collect_uarch_cycle_root_hashes bundled per big cycle, so each period
ends with the cycle's root. Idle stretches, positioning, sampled strides
and proving keep their paths. The revert tail is the idle period
collected at feed.

Seam 1 stays stepped: v0.21.0's collector keeps the physical root for a
rejection on the input budget's last cycle, so bulk collection declines
that cycle and the ruler steps it. A control test pins the collector's
answer there, to say when the guard can go.

The stepped path stays selectable (DisputeSource::on_store_with,
Collector::Stepped) as the permanent reference (D7): the leaf goldens
and every dense corpus leaf are checked both ways, a run-level
differential localizes disagreements, the collect API's bundles reduce
to the unbundled leaves at every bundle size, and the toy drives the
ruler's bulk branch through chunked and declined calls.
How dense leaves are built now and why the stepped path stays
(computation-hash.md), the corpus gate building dense leaves both ways
(build-system.md), the new differentials (test-harness.md), the
migration's evidence and what remains (collect-hashes-migration.md),
and W6's status with the seam-1 guard (two-level-sling.md).
From the switch's review: a delivery that leaves the machine at a fixed
point (a template preset halted with an input yield pending, which the
step still feeds) made the bulk collector return its one fixed-point
period without progress, and big_cycle_roots failed where stepping
builds the leaves. That period is the opening cycle's root.

Also from the review: a ruler-level bulk-versus-stepped test at both
seam openings (the budget's last cycle, the halted template), and a
period-10 echo leaf around the first automatic yield in the run-level
differential, so the real collector crosses chunks and re-enters after
the yield. Chunks shrink to 256 cycles to keep that leaf cheap; a call's
overhead is small next to a cycle's uarch work.
The emulator opens and flocks machine files without O_CLOEXEC (W7), so
an anvil spawned while another test stores or clones a machine keeps
that machine's files locked, and the next load fails. New store-heavy
engine tests made test_blockchain_reader's long-known flake frequent
(3 of 6 suite runs); skipping the anvil tests made it vanish (0 of 6).
A close-on-exec sweep before each spawn did not help: the emulator
opens short-lived locked files throughout a store, so the window
between the sweep and the fork stays open.

The seven tests that spawn anvil join the harness's serial run in
just test-node-harness, which CI's Rust job runs.
Dispute positioning wrote back every input boundary it crossed with a
full machine store: about 410 MiB of disk per boundary on the stress
image, 25 GiB for a cold join at the end of a 64-input snapshot gap,
and the stores' page reads put the join's peak RSS at 606 MiB. Every
later action of a dispute stays inside one input, so only that input's
boundary is load-bearing (R16).

Positioning now crosses whole windows from the nearest stored boundary
on the runner's copy-on-write clone chain, reusing its swap through
kernels extracted from record_accepted and record_reverted, and
publishes and registers only the target's boundary (no window roots, no
gap GC). The target window is then positioned as before. Measured at
full scale (just measure-node-vs-emulator): disk 26.6 MiB, peak RSS
116 MiB against the emulator's 111, join time unchanged at 1.03x.

From the change's review: a torn gap snapshot that positioning skips
could come back as a revert checkpoint, since a feed adopts an existing
content-addressed directory without rehashing it, and a rejection would
then reload a wrong state silently. The checkpoint now carries its
revert root, and the reload asserts it. Tests: crossings register only
their final boundary as clones of the start (inodes) and leave no
working clone, an all-rejected crossing reuses the floor, the gap test
hides the epoch start so its floor is proved, a torn checkpoint fails
loudly, and an abandoned crossing leaves no trace.
docs/measurements/node-vs-emulator.md is the full-scale run after both
changes (64-input gap, height-37 leaf, stress image): the cold join at
1.03x the emulator's time, 116 MiB peak RSS against 111, 26.6 MiB of
disk. The plans record the sequence: 1.18x with the per-step builder,
1.03x with the bulk collector, and the CoW crossing taking the join's
disk from 25 GiB to 27 MiB.
From the external review of 0a9976a..c9c1f70. The safe wrappers
checked only that a bundle exponent fits int32, but v0.21.0 computes
with it before validating: a uarch exponent above 20 makes a negative
shift (undefined behavior), and mcycle bundle construction does
arithmetic before its range check. Exponents past 20 (uarch) and 63
(mcycle) now return INVALID_ARGUMENT, and a test checks the machine is
left unchanged. The node only ever passes valid sizes.
From the same review: the report now names the source revision
(git describe, dirty when uncommitted) and the workload's fingerprint
line, so later runs can be compared.
From the same review: an input fed at mcycle 2^64 - 2 saturates its
budget, so its first cycle is its budget's last, the seam-1 guard
declines it, and the ruler's assert stops the node where stepping
would proceed. Reaching it takes a template preset there or centuries
of machine time, so it is documented, with the assert naming the
cause, rather than handled.
The battery-cleanup regression test still expected the pre-cut
battery's 25 scenarios, so since 0a9976a it failed in CI's build job,
before the Rust, engine and harness steps, which have not run on the
branch since. It now counts the battery's SCENARIOS list, so a cut or
an addition cannot leave it stale.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant