Skip to content

Publish the replication runs and fix the audit findings - #3

Open
devYRPauli wants to merge 17 commits into
mainfrom
audit-fixes-2026-10
Open

devYRPauli wants to merge 17 commits into
mainfrom
audit-fixes-2026-10

Conversation

@devYRPauli

@devYRPauli devYRPauli commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

Problem

An audit of the repo found these issues:

  • The case studies cite n=5 replication runs, but the run files were not in the repo.
  • The README and the site said that Ollama generates unconstrained text. That is true on the Go template path only.
  • Several docs called the seven passing scenarios "the negative traps". Five are negative traps. The other two are tool_choice scenarios.
  • The docs and the registry evidence published the ssh alias of the measurement host.
  • Two Python bytecode files were tracked.
  • CI did not run the Python tool tests or the path redaction ledger check.
  • The redaction migration doc named a ledger path that does not exist.
  • The llama.cpp 500s case study gave an error breakdown that fits two of its three arms.
  • CI fails on Rust 1.99, on main too. Clippy's result_large_err rejects send_recorded in client.rs.

Fix

  • Add the 130 replication run files (schema v2, 50 scenarios each) under docs/case-studies/evidence/replication/, with a README. Link them from the case studies. The transcripts are not published. Each run file records the SHA-256 of each transcript.
  • Rewrite the README in simplified technical English. Keep the facts and the commands.
  • Limit the decode sentence in the README and the site to Ollama with a Go template.
  • Describe the seven scenarios as "the scenarios where making no call is correct".
  • Replace the host alias with "the measurement host".
  • Stop tracking __pycache__ and add a tools job to CI.
  • Fix the ledger path, the 500s breakdown, CONTRIBUTING, and the spelling.
  • Box the captured turn in RecordedAttemptError. The extra allocation happens only on a timeout or a transport error.

Testing

  • python3 -m unittest discover -s tools -v: 9 tests pass.
  • python3 tools/redact_local_paths.py --check-all migrations/path-redaction --evidence-root results: 6 runs and 300 transcripts verified.
  • I checked each changed number against results/ and the replication run files.
  • The Rust changes are one sentence in the site template, the same sentence in its golden file, and the boxed error type. The Rust jobs in this PR check them.

devYRPauli and others added 17 commits October 5, 2026 14:46
The case studies claim five runs per arm, but the repo held only run 1
of the granite arms. Add the 130 run files from the M5 and M6 batches.
The transcripts are not published. Each scenario keeps its
evidence_hash, so a transcript can still be checked against its run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The README had long sentences and several wrong claims. It said the
decode difference favors llama.cpp in every row. That is false. It
described a bug that never existed. Its seed-results sentence was out
of date. It listed only exit codes 0, 1, and 3.

Split the long sentences and cut the filler. Replace the false claim
with the two same-blob rows that score 7 of 50 on both servers. Scope
the decode claim to Ollama 0.32.1 with its Go template, and to mlx-lm.
Add the Jinja template route from ollama/ollama#17274. Name the real
Unicode bug from commit 0d4aa33. Add exit codes 2 and 4. Link the
replication runs. Remove the roadmap items that are done.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Two compiled .pyc files for the redaction tool were in git. They are
build output, and each Python run can change them.

Remove the two files from the index. Delete the tools/__pycache__
directory. Ignore __pycache__/ and *.pyc.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The doc named docs/migrations/local-path-redaction-ledgers. That
directory does not exist. The ledgers are in migrations/path-redaction.
A command copied from the doc fails with "ledger path does not exist".

Use the real path in the three commands.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
CI did not run the tests for tools/redact_local_paths.py. It did not
check the path redaction ledgers. A broken tool or ledger could merge
without notice.

Add a tools job. It runs the unit tests in tools/ and runs --check-all
on migrations/path-redaction. The check reads old files with git show,
so the checkout fetches the full history. The job sets
PYTHONDONTWRITEBYTECODE, so the tests leave no bytecode behind.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Two evidence files named the private ssh alias of the measurement host
in 31 provenance strings.

Replace "ssh <alias>" with "measurement host" in those strings. No
other byte changes, and the record order stays the same. Every
provenance_ref pointer still resolves. No code reads these strings,
and nothing pins the files by hash.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Two specs named the private ssh alias of the measurement host. The M7
spec also said the n=5 data lives off-repo. That changed on
2026-10-05. Both specs used some British spellings.

Write "the measurement host" in place of the alias. Say that the run
files are now in docs/case-studies/evidence/replication/ and that the
transcripts are not published. Use American spelling.

The design spec keeps its line count. The decode-modes registry still
cites lines 248 to 267 of that file, and those lines are unchanged.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The case study named the measurement host by its private ssh alias.
The alias is a hostname, and the repo must not publish it.

Write "the measurement host" in its place and rewrap that paragraph.
No fact, number, or link changes.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Four documents said that granite and mlx 8-bit pass "the negative
traps". That is wrong. Two negative_trap scenarios need a call, and
both models fail them. Two passes are tool_choice scenarios.

The seven passes are the seven scenarios where making no call is
correct. This holds for both granite rows, all three phi4-mini rows,
and every mlx 8-bit run. No other scenario expects no call.

Say that in both case studies and in the two specs. The design spec
keeps its line count, so lines 248 to 267 stay in place.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
One sentence used "generalised". The project uses American spelling.
Write "generalized".

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The run files behind the replication claims are now in
docs/case-studies/evidence/replication/. The case studies did not point
to them.

Add a pointer to the matching folders in four case studies: granite
(m5-replication), quantization (m6-armA and m6-armS), llama.cpp 500s
(m6-armA, m6-armB, and m6-armS), and mlx 8-bit (m6-mlx-repl). The
granite study now also says that the transcripts are not published.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The scenario example had no response_requirement. The corpus test
requires one on every turn of an embedded scenario. CONTRIBUTING did
not say where a scenario file goes. It left out the evidence directory
that the README requires in a result pull request. Its failure class
text could suggest that an error scenario has a class. CI now runs two
Python checks, and CONTRIBUTING did not list them.

Add response_requirement to the example and to the rules. Add the file
location rule, the evidence directory, and the two Python commands.
Say that an error scenario has no failure class. Split two sentences
that were too long.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Two README sentences and one case study sentence had more than 20 words.
Each now has 20 words or fewer. The meaning is the same. The case study
aside repeated the sentence before it, so the edit removes the aside.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The paragraph said twice that granite passes only the scenarios where
no call is correct. The second sentence also used a colon setup. The
paragraph now gives the seven passes and the 43 failures. The numbers
match the results files.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Ollama generates unconstrained text when it renders a Go template. Its
Jinja template route uses constrained generation, so the old sentence
was too broad. The golden file changes with the template.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The old sentence fit the Q8_0 and Q4_K_M arms only. At Q3_K_M the greedy
runs have 7 errors: 4 multi_turn, 1 negative trap, 1 parallel call and
1 single call. The counts come from the m6-armA and m6-armB run files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Clippy 1.99 rejects send_recorded under result_large_err. The error type
held a CapturedTurn of at least 168 bytes, so every Result from
send_recorded was large. The turn is now boxed. The extra allocation
happens only on a timeout or a transport error.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant