feat: collect repository metrics in an independent store - #206
Conversation
|
🦞👀 Pull request received. I will update this pull request when review starts. ClawSweeper review completeClawSweeper finished reviewing this revision. The review result is being finalized. |
|
Codex review: needs real behavior proof before merge. Reviewed September 28, 2026, 7:14 AM ET / 11:14 UTC (Revision 9). ClawSweeper reviewWhat this changesThis PR adds commands to collect GitHub repository counters and releases, import historical observations, and inspect a separate SQLite metrics database. Merge readiness⛔ Blocked before merge - 6 items remain The metrics workflow is absent from current main and v0.12.0. The latest branch addresses both earlier review findings, but its foreign-database safeguard can still change a substituted SQLite file before rejecting it. Priority: P2 Review scores
Verification
How this fits togetherGitcrawl normally stores GitHub issues and pull requests for local maintainer search. The new metrics commands take a separate config, GitHub responses or imported history, then write an independent database and report its status. flowchart LR
A[Metrics config] --> B[Metrics commands]
C[GitHub repository API] --> D[Counter and release collector]
D --> B
E[Imported history] --> B
B --> F[Separate SQLite store]
F --> G[Status output]
Decision needed
Why: The documented archive isolation guarantee and CrawlKit's current writable-open behavior meet at a shared-library boundary; an owner must approve the connection strategy. Before merge
Findings
Agent review detailsSecurityNeeds attention: The database ownership check occurs after shared-library writable connection setup can affect a substituted foreign file. Review metrics
Merge-risk optionsMaintainer options:
Technical reviewBest possible solution: Verify ownership on the exact database connection before any journal-mode, permission, or schema change, then prove a substituted archive retains its bytes, mode, and journal setting. Do we have a high-confidence way to reproduce the issue? Yes, by injecting a foreign SQLite file at the existing pre-open swap hook and comparing its journal mode and permissions before and after rejection. Source shows why the current assertions miss those effects; I did not execute that check. Is this the best way to solve the issue? Not yet. The separate store is a coherent boundary, but its writable opener must honor that boundary during connection setup. Full review comments:
Overall correctness: patch is incorrect AGENTS.md: not found in the target repository. Codex review notes: model internal, reasoning high; reviewed against 290d943b1c0c. LabelsLabel changes:
Label justifications:
EvidenceSecurity concerns:
What I checked:
Likely related people:
Rank-up movesOptional improvements that raise the rating; they are not merge blockers.
Rating scale
Overall follows the weaker of proof and patch quality. Workflow
HistoryReview history (8 earlier review cycles)
|
Retain both isolated metrics installation guidance and SIMD source-build guidance in the sole conflict, docs/installation.md.
Restore gofmt indentation in the oversized-row error path. The previous Ubuntu CI run stopped here before executing tests; no runtime behavior changes.
Recover returned first-initialization failures without removing pre-existing files, reject database hardlink aliases, compare imported timestamps and UTC days chronologically, and enforce the config size limit. Preserve original import IDs, timestamp spelling, NULLs, zeroes, and corrections.
Stop on exhausted quota while keeping completed reads, including cancellation after a response. Reject missing provider lists and emit structured partial results and explicit zero status totals. Document archive ownership inspection, permissions, recovery limits, and collection request costs.
Keep read-only prechecks, then pin the writable connection and validate ownership under BEGIN IMMEDIATE before applying schema and metadata atomically. Require new databases to remain empty, reject replaced file identities, and retain guarded cleanup. Cover file and parent swaps, in-place changes, and rollback when metadata insertion fails.
|
Thanks @hannesrudolph! I merged
The full suite and race tests passed on Linux, the Windows cross-compile passed, and autoreview is clean. |
What Problem This Solves
Repository headline history currently requires a separate collector instead of a discoverable Gitcrawl command.
User Impact
Adds
gitcrawl metrics collect|import|status --config metrics.jsonfor stars, forks, actual subscribers, open PRs/issues, optional completed-day clones, and stable releases. Operators can retain OpenClaw repository metrics in a separate private SQLite metrics database, using native GitHub authentication and JSON output.Why This Change Was Made
The metrics store preserves unknown values, zeroes, decreases, original import IDs, and daily corrections. Imports are scoped and atomic; archive/wrong-owner databases are rejected before a writable open. The commands never invoke archive refresh, embeddings, models, or schedules. Help, control metadata, and source documentation expose the full workflow.
Evidence