Repository navigation
probbit 0.8.0: safety kit (reward provenance, control lines, writer lock, checkpoints) - #23
Merged
Merged
Conversation
A persona may declare reward_from: the sources (human, env) a reward-bearing input may come from, for every one (a list or any) or per input (the learning flags and goals.<id>.win). An event says who produced it with src (human[:id], env[:id], self, clock). With the key, a reward from src self, without a source or from an undeclared one is refused whole, as a bad event is; src is read (not ignored) and echoed in the stance's inputs. Without the key nothing changes: src is an undeclared input as before. fuzz and prove search over stances, not sources (turn_any_source). State::zero_credit clears the credit (used by control lines).
One writer per strand: live --strand takes STRAND.lock (created exclusively with the writer's pid and start; a lock whose process is gone is taken over) before it reads the strand and holds it to its exit; a second writer exits 4 and changes nothing; every append first checks the lock is still the writer's. probbit_live_event and live control take it for their append. probbit live control STRAND pause|resume|retire --by WHO --reason TEXT appends a chained control line. While paused or retired every event is refused (code paused / retired, exit 4, nothing written); retire is final; the credit is cleared at every control line, so feedback after a resume credits no stance from before the pause. verify replays control lines and reports controls and status; resume applies the ones after the last event. Checkpoint lines every K events (--checkpoint-every, default 1000) carry the event count, that event's stance and the state. verify checks each one, verify --from-checkpoint starts at the last one, monitor --once draws from it. Tests: provenance read and refusal, P3 (40 x 300), the lock, control lines and P10 (10,000 events after retire), P5 (50 x 1,000 with pause/resume pairs), checkpoints.
A persona without reward_from gives 0.8.0's trace, strand and state with src in its events; with it, self and clock rewards are refused. Exit codes of checkpoints, control lines and a held lock.
…nts (§5.7) persona.md: the reward_from key and src, what is refused and what is not; one writer per strand; pause, resume and retire as control lines, with no credit across them; checkpoint lines and verify --from-checkpoint; monitor --once from the last checkpoint; exit code 4. CHANGELOG under 0.9.0 (unreleased).
The lock's temporary file gets a per-process counter (two takes in one process never share it); checkpoint lines are computed only with --strand (a run without one prints the chain head of 0.8.0); with reward_from no input may be called src (it names the event's source).
A 10,000-event drives strand opens in monitor --once from its last checkpoint in 0.005 s (full verify: 192 s); 500 events past the last checkpoint take 26 s to replay at load 6-13 for that persona.
…red in the badge The first read of --follow and --serve replays a strand with checkpoint lines from the last one, as --once does, and draws its event at once. A paused or retired individual says so next to the badge.
BitmapAsset
force-pushed
the
b1-safety-kit
branch
from
October 9, 2026 07:16
9d52224 to
f5d6bf3
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Draft: round 1 of 2. Four engine-side guards for loops where an agent's own behaviour produces the events that move an individual. Personas without the new key and strands without the new lines give 0.8.0's documents and strands, byte for byte. Zero new dependencies.
What lands
src(human[:id],env[:sensor],self,clock). A persona may declarereward_from: a list ofhuman/env,any, or one per reward-bearing input (the learning flags andgoals.<id>.win). A reward fromsrc: self, without a source or from an undeclared one is refused whole, as a bad event is (one{"error"}atinputs.src, nothing changed).srcis read, echoed in the stance's inputs and logged in the strand.live --strandtakesSTRAND.lock(created exclusively with the writer's pid and start; a lock whose process is gone is taken over) before it reads the strand, and holds it to its exit. A second writer exits 4 and changes nothing. Every append first checks the lock is still the writer's.probbit_live_eventandlive controltake it for their append,--demo week --strandfor its week.probbit live control STRAND pause|resume|retire --by human:ID --reason TEXTappends a chained control line under the lock. While paused or retired every event is refused (codepaused/retired, exit 4, nothing written); retire is final; the credit is cleared at every control line, so feedback after a resume credits no stance from before the pause.verifyreplays them and reportscontrolsandstatus. Deliberately not an MCP tool.--checkpoint-every K, default 1,000; 0 = none) a checkpoint line carries the event count, that event's stance and the state.verifychecks each against the replay;verify --from-checkpointstarts at the last one;monitor(--once, and the first read of--follow/--serve) draws from it, and shows a paused or retired status.Done-when, with the numbers so far (an Apple M4, a shared machine: load given per row)
src: self) withreward_from: [human, env]: all refused, the individual ends in its initial state byte for byte (P3)retire, 10,000 random events are all refused (P10)retiredmonitor --oncefrom its last checkpoint in under 5 sverifyof the same strand 192 s. Worst case measured beside it: 500 events past the last checkpoint take 26 s to replay at load 6-13 (this persona costs ~40-50 ms an event), so the default K = 1,000 bounds the wait at K events; a smaller default is a decision for reviewcargo test --release --workspacegreen locally, twice); a persona without the key withsrcon every event: trace, strand and state equal the 0.8.0 binary's (test). CI green on the head 9d52224 (run 37856238295): tests on Ubuntu, macOS and Windows, wasm, all four release targetsA note on P5. As written, P5 (the run with pause/resume pairs equals the run without them) and "credit is cleared at pause and at resume" cannot both hold: clearing the credit means the first feedback after a resume moves nothing, where the uninterrupted run would credit the stance before the pause. This draft clears the credit (no reward crosses an interruption) and tests P5 in the exact form above. Keeping the credit across a pause instead would make P5 hold literally (one line in
Live::controland inresume); that is a decision for review.Not in this round
Never merged, tagged or released from here.