fix: audit nested ATIF research evidence and coverage - #85
vincentkoc wants to merge 1 commit into
Conversation
|
🦞👀 Pull request received. I will update this pull request when review starts. ClawSweeper review completeClawSweeper finished reviewing this revision. The review result is being finalized. |
|
Codex review: blocked before merge. Reviewed October 4, 2026, 2:22 PM ET / 18:22 UTC (Revision 2). ClawSweeper reviewWhat this changesExtends ShellBench’s archived agent-trace audit to include nested sessions, validate export receipts, preserve zero usage, and report independent evidence coverage. Merge readiness⛔ Blocked before merge - 2 items remain This PR addresses defects still present on current main and remains the candidate fix for the linked issue. No discrete introduced bug was found, but schema-v2 downstream upgrade compatibility remains unverified. Priority: P2 Review scores
Verification
How this fits togetherShellBench’s research audit reads archived benchmark results and agent trajectories, then exports tables for model identity, tool use, usage, and research evidence. These tables inform downstream analysis without changing execution or leaderboard eligibility. flowchart TD
A[Archived benchmark results] --> C[Research audit]
B[Nested agent traces] --> C
D[Export receipts] --> E[Integrity and coverage checks]
C --> E
E --> F[Identity and evidence statuses]
C --> G[Research CSV tables]
F --> G
Before merge
Agent review detailsSecurityNone. Review metrics
Merge-risk optionsMaintainer options:
Technical reviewBest possible solution: Keep the repair in the archive consumer and demonstrate that legacy archives and downstream analysis can transition safely to node-aware schema-v2 tables. Do we have a high-confidence way to reproduce the issue? Yes, source establishes the current-main root-only traversal and explicit-zero fallback defects; nested traces and zero-valued harness usage are concrete triggers. No execution was performed in this read-only review. Is this the best way to solve the issue? Yes, repairing the archive consumer is the appropriate boundary, and no introduced correctness defect was found. Downstream schema-v2 upgrade validation is still needed. AGENTS.md: not found in the target repository. Codex review notes: model internal, reasoning medium; reviewed against e3f8d25a01f4. LabelsLabel changes:
Label justifications:
EvidenceWhat I checked:
Likely related people:
Rank-up movesOptional improvements that raise the rating; they are not merge blockers.
Rating scale
Overall follows the weaker of proof and patch quality. Workflow
HistoryReview history (1 earlier review cycle)
|
What does this PR do?
Extends the existing native research audit to consume unique nested ATIF trajectories, preserve session/tool lineage, verify receipt integrity and node coverage, and report independent evidence facets. This is an archive-consumer change, not a new execution engine or a generic Harbor schema change.
Fixes #84. Related: #48 (task analysis and upload) and #60 (discovery telemetry). Those features and branches are unchanged. The export-loop edits may require normal reconciliation with #60 when landing.
Why?
The current audit inspects only root steps/models, falls through explicit zero usage, and calls runtime-reported costs exact without billing reconciliation. Nested child/grandchild traces can therefore disappear from the model check and research rows.
Provenance: the root-only consumer, zero fallback, and exact cost labels originated in #51, authored and merged by @vincentkoc on 2026-07-29 (
569b5c39c7831347ad37583cbb8251c5238fcfdd). The nested exporter makes the traversal gap visible. This does not claim the historical flattened converter always lost child identity.Changes
trajectory_nodes.csv,trajectory_links.csv, andmodel_usage_coverage.csv; preserve missing values versus real zero and report per-component coverage denominatorsreported_*; retain deprecated exact-count keys as zero and add reported-cost countsProduction LOC: +607/-62 (net +545). Tests: +468/-5. The added evidence module owns graph reconciliation and independent evidence checks; it does not duplicate the runner, ATIF converter, task analysis, discovery parser, or archive uploader.
Tests
python -m pytest -q— 553 passed, 5 skipped (baseline: 527 passed, 5 skipped)python -m ruff check clawbench app.py scripts testsBest-fix review and limits
Best-fix verdict: owner-boundary consumer fix. Changing generic Harbor schemas would affect unrelated agents without repairing these CSV/identity paths. Flattening the exporter would discard the hierarchy that the consumer needs. Folding this into #48 or #60 would mix distinct analysis/upload/discovery features.
Code read: research audit, aggregate acceptance boundary, native runtime/execution outcome, harness trajectory conversion, current research runbook, and public exporter mapper/schema/receipt contracts. The current exporter main at
4816624739db14ddd54de2403c59fe8e79deadd3has unchanged mapper/receipt implementations relative to the inspected public fixture source7dd7bcc6ec958e61e08e70d1c1df2d185b3a5642; the fixture probe is not a packaged/live run.This draft is not live campaign qualification. Capture is bounded to selected public exported generations; observed agent steps are not a complete provider-request ledger. Runtime costs remain billing-unreconciled. Judge/reasoning identity, authenticated delegated lifecycle, resource-host measurements, archive restore, and a new scored campaign have not been run. Resource acceptance stays explicitly unavailable until a validated resource artifact contract exists. No provider calls, paid infrastructure, release, or merge is included.