Repository navigation
Publish the replication runs and fix the audit findings - #3
Open
devYRPauli wants to merge 17 commits into
Open
devYRPauli wants to merge 17 commits into
devYRPauli wants to merge 17 commits into
Conversation
The case studies claim five runs per arm, but the repo held only run 1 of the granite arms. Add the 130 run files from the M5 and M6 batches. The transcripts are not published. Each scenario keeps its evidence_hash, so a transcript can still be checked against its run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The README had long sentences and several wrong claims. It said the decode difference favors llama.cpp in every row. That is false. It described a bug that never existed. Its seed-results sentence was out of date. It listed only exit codes 0, 1, and 3. Split the long sentences and cut the filler. Replace the false claim with the two same-blob rows that score 7 of 50 on both servers. Scope the decode claim to Ollama 0.32.1 with its Go template, and to mlx-lm. Add the Jinja template route from ollama/ollama#17274. Name the real Unicode bug from commit 0d4aa33. Add exit codes 2 and 4. Link the replication runs. Remove the roadmap items that are done. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Two compiled .pyc files for the redaction tool were in git. They are build output, and each Python run can change them. Remove the two files from the index. Delete the tools/__pycache__ directory. Ignore __pycache__/ and *.pyc. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The doc named docs/migrations/local-path-redaction-ledgers. That directory does not exist. The ledgers are in migrations/path-redaction. A command copied from the doc fails with "ledger path does not exist". Use the real path in the three commands. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
CI did not run the tests for tools/redact_local_paths.py. It did not check the path redaction ledgers. A broken tool or ledger could merge without notice. Add a tools job. It runs the unit tests in tools/ and runs --check-all on migrations/path-redaction. The check reads old files with git show, so the checkout fetches the full history. The job sets PYTHONDONTWRITEBYTECODE, so the tests leave no bytecode behind. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Two evidence files named the private ssh alias of the measurement host in 31 provenance strings. Replace "ssh <alias>" with "measurement host" in those strings. No other byte changes, and the record order stays the same. Every provenance_ref pointer still resolves. No code reads these strings, and nothing pins the files by hash. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Two specs named the private ssh alias of the measurement host. The M7 spec also said the n=5 data lives off-repo. That changed on 2026-10-05. Both specs used some British spellings. Write "the measurement host" in place of the alias. Say that the run files are now in docs/case-studies/evidence/replication/ and that the transcripts are not published. Use American spelling. The design spec keeps its line count. The decode-modes registry still cites lines 248 to 267 of that file, and those lines are unchanged. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The case study named the measurement host by its private ssh alias. The alias is a hostname, and the repo must not publish it. Write "the measurement host" in its place and rewrap that paragraph. No fact, number, or link changes. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Four documents said that granite and mlx 8-bit pass "the negative traps". That is wrong. Two negative_trap scenarios need a call, and both models fail them. Two passes are tool_choice scenarios. The seven passes are the seven scenarios where making no call is correct. This holds for both granite rows, all three phi4-mini rows, and every mlx 8-bit run. No other scenario expects no call. Say that in both case studies and in the two specs. The design spec keeps its line count, so lines 248 to 267 stay in place. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
One sentence used "generalised". The project uses American spelling. Write "generalized". Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The run files behind the replication claims are now in docs/case-studies/evidence/replication/. The case studies did not point to them. Add a pointer to the matching folders in four case studies: granite (m5-replication), quantization (m6-armA and m6-armS), llama.cpp 500s (m6-armA, m6-armB, and m6-armS), and mlx 8-bit (m6-mlx-repl). The granite study now also says that the transcripts are not published. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The scenario example had no response_requirement. The corpus test requires one on every turn of an embedded scenario. CONTRIBUTING did not say where a scenario file goes. It left out the evidence directory that the README requires in a result pull request. Its failure class text could suggest that an error scenario has a class. CI now runs two Python checks, and CONTRIBUTING did not list them. Add response_requirement to the example and to the rules. Add the file location rule, the evidence directory, and the two Python commands. Say that an error scenario has no failure class. Split two sentences that were too long. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Two README sentences and one case study sentence had more than 20 words. Each now has 20 words or fewer. The meaning is the same. The case study aside repeated the sentence before it, so the edit removes the aside. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The paragraph said twice that granite passes only the scenarios where no call is correct. The second sentence also used a colon setup. The paragraph now gives the seven passes and the 43 failures. The numbers match the results files. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Ollama generates unconstrained text when it renders a Go template. Its Jinja template route uses constrained generation, so the old sentence was too broad. The golden file changes with the template. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The old sentence fit the Q8_0 and Q4_K_M arms only. At Q3_K_M the greedy runs have 7 errors: 4 multi_turn, 1 negative trap, 1 parallel call and 1 single call. The counts come from the m6-armA and m6-armB run files. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Clippy 1.99 rejects send_recorded under result_large_err. The error type held a CapturedTurn of at least 168 bytes, so every Result from send_recorded was large. The turn is now boxed. The extra allocation happens only on a timeout or a transport error. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
An audit of the repo found these issues:
result_large_errrejectssend_recordedinclient.rs.Fix
docs/case-studies/evidence/replication/, with a README. Link them from the case studies. The transcripts are not published. Each run file records the SHA-256 of each transcript.__pycache__and add atoolsjob to CI.RecordedAttemptError. The extra allocation happens only on a timeout or a transport error.Testing
python3 -m unittest discover -s tools -v: 9 tests pass.python3 tools/redact_local_paths.py --check-all migrations/path-redaction --evidence-root results: 6 runs and 300 transcripts verified.results/and the replication run files.