Skip to content

docs: state the module-size limit #853 actually enforces - #888

Merged
dgenio merged 1 commit into
mainfrom
claude/fix-module-size-doc-drift
Sep 17, 2026
Merged

dgenio merged 1 commit into
mainfrom
claude/fix-module-size-doc-drift

Conversation

@dgenio

@dgenio dgenio commented Sep 17, 2026

Copy link
Copy Markdown
Owner

Summary

scripts/check_module_size.py has enforced LIMIT = 500 since 663b0f2 ("reconcile invariants and module-size policy", #853). That commit updated AGENTS.md, docs/agent-context/invariants.md, .claude/CLAUDE.md and the script — and missed the two places that state the rule to someone looking it up. This finishes #853.

Found while checking my own PR #887 against the checklist item "every modified module stays ≤ 300 lines", which is itself the stale number.

Fixes #

Changes

  • Makefile:246 — "enforces ≤300 lines for new modules" → ≤500.
  • docs/agent-context/workflows.md:18 — "enforce the ≤300-line convention" → ≤500.
  • src/contextweaver/_schema_gen.py — drops a stale self-description claiming the module "exceeds the 300-line soft cap (~360 lines)". It is 280 lines and the cap is 500, so the sentence was false in both halves. The exemption note it sat inside is kept: the coupling argument for not splitting the file is still why it is exempt.
  • llms-full.txt — regenerated with make llms, not hand-edited. Sole delta is the workflows.md line above.

Why it matters

The gate's own docstring says the limit is "a ratchet, not a target to split cohesive modules mechanically". An agent consulting the Makefile or workflows.md would read 300 and split a cohesive 350-line module for nothing — the precise outcome both documents exist to prevent.

Checklist

  • Tests added or updated for every new/changed public function — n/a, comments and docs only; no behaviour changes
  • make ci passes locally — partially, see Notes
  • CHANGELOG.md updated under ## [Unreleased]
  • Docstrings added for all new public APIs (Google-style) — n/a
  • Public-API change? — none; gen_api_manifest.py --check reports "up to date"
  • Every modified module stays ≤ 500 lines — check_module_size.py: OK (236 modules, 8 grandfathered ≤ frozen ceiling)
  • Related issue linked in the summary above (refactor: reconcile invariants and module-size policy #853)
  • Agent-facing docs updated if pipeline, API, or conventions changed — this PR is that update

Notes for reviewers

Deliberately not changed — and I want this reviewed as a judgment call, not assumed. Eleven module docstrings under src/ still say 300. Most narrate why a past split happened:

store/_json_file_io.py:3 — "Extracted to keep the store module under the 300-line ceiling (issue #497)."

Those are accurate records of a decision taken when 300 was the limit. Rewriting them to 500 would falsify the history without making any current rule clearer, so I left them. Same reasoning the repo applies to dated audit records elsewhere.

Four read as live guidance rather than history, and are worth a follow-up:

Site Text
extras/embeddings.py:28 "so this file stays under the project's 300-line module guideline"
routing/filters.py:14 "within the ≤ 300 line per-module guideline"
routing/explanation.py:6 "does not grow further beyond the soft 300-line cap"
__main__.py:30 "Exempt from the 300-line module limit"

I did not sweep these because the fix is not purely mechanical — each needs a decision about whether the sentence is describing the rule or the reason, and __main__.py's exemption claim is correct even though its number is not. Listing them explicitly rather than leaving them for someone to rediscover; happy to do them in a follow-up if you'd rather they all move together.

What I ran: check_module_size.py (OK), drift_check.py --check (all 10 artifacts up to date, after make llms), gen_api_manifest.py --check (up to date), ruff check and ruff format --check on the changed source file. I did not run full make ci: mkdocs is not installed in this container so the docs target cannot run, and the test suite is unaffected by a comment-only change. CI is the authority on the rest.


🤖 Generated with Claude Code

https://claude.ai/code/session_01XhJp1Rnr6mF4vdSLkj14jU


Generated by Claude Code

`scripts/check_module_size.py` has enforced `LIMIT = 500` since 663b0f2
("reconcile invariants and module-size policy", #853). That change updated
AGENTS.md, docs/agent-context/invariants.md, .claude/CLAUDE.md and the script
itself, but missed two places that state the rule to anyone looking it up:

* `Makefile:246` - "enforces <=300 lines for new modules"
* `docs/agent-context/workflows.md:18` - "enforce the <=300-line convention"

An agent consulting either would split a cohesive 350-line module for no
reason, which is exactly what the gate's own docstring warns against
("not a target to split cohesive modules mechanically").

Also drops a stale self-description in `_schema_gen.py`, which claimed the
module "exceeds the 300-line soft cap (~360 lines)". It is 280 lines, and the
cap is 500, so the sentence was false in both halves. The exemption note it
sat in is kept - the coupling argument for not splitting the file is still
the reason it is exempt.

`llms-full.txt` regenerated with `make llms` rather than hand-edited; the
only delta is the workflows.md line above.

Deliberately NOT changed: eleven module docstrings under `src/` that narrate
why a past split happened ("Extracted to keep the store module under the
300-line ceiling (issue #497)"). Those are accurate records of a decision
taken when 300 was the limit, and rewriting them to 500 would falsify the
history without making any current rule clearer. Four of them read as live
guidance rather than history and are worth a follow-up; they are listed in
the PR body rather than swept silently here.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XhJp1Rnr6mF4vdSLkj14jU
@github-actions

Copy link
Copy Markdown
Contributor

Benchmark delta (vs main)

Soft regression feedback only — this comment never blocks the PR.
Latency budget: ⚠️ when head > base × 1.3. Accuracy budget: ⚠️ when head < base - 1pp.

Routing summary (single backend × catalog sizes)

size recall@k (head Δ vs base) MRR (head Δ vs base) p99 (ms)
50 ✅ 0.5649 (+0.0000) ✅ 0.4978 (+0.0000) ✅ 0.503 (base 0.650)
83 ✅ 0.3825 (+0.0000) ✅ 0.3242 (+0.0000) ✅ 0.739 (base 1.205)
1000 ✅ 0.1475 (+0.0000) ✅ 0.1456 (+0.0000) ✅ 37.639 (base 47.855)

Per-backend × per-size matrix

backend size recall@k (Δ) MRR (Δ) p99 (ms)
bm25 100 ✅ 0.3825 (+0.0000) ✅ 0.3399 (+0.0000) ✅ 6.558 (base 8.399)
bm25 500 ✅ 0.2250 (+0.0000) ✅ 0.2165 (+0.0000) ✅ 29.474 (base 41.865)
bm25 1000 ✅ 0.1575 (+0.0000) ✅ 0.1525 (+0.0000) ✅ 86.691 (base 112.387)
embedding_hashing 100 ✅ 0.5175 (+0.0000) ✅ 0.4360 (+0.0000) ✅ 7.557 (base 7.218)
embedding_hashing 500 ✅ 0.2700 (+0.0000) ✅ 0.2674 (+0.0000) ✅ 41.186 (base 43.171)
embedding_hashing 1000 ✅ 0.2000 (+0.0000) ✅ 0.1931 (+0.0000) ✅ 99.116 (base 102.979)
embedding_st 100 skipped (skipped: missing sentence-transformers) — —
embedding_st 500 skipped (skipped: missing sentence-transformers) — —
embedding_st 1000 skipped (skipped: missing sentence-transformers) — —
fuzzy 100 skipped (skipped: missing rapidfuzz) — —
fuzzy 500 skipped (skipped: missing rapidfuzz) — —
fuzzy 1000 skipped (skipped: missing rapidfuzz) — —
tfidf 100 ✅ 0.3825 (+0.0000) ✅ 0.3220 (+0.0000) ✅ 1.097 (base 1.193)
tfidf 500 ✅ 0.2325 (+0.0000) ✅ 0.2314 (+0.0000) ✅ 9.102 (base 10.803)
tfidf 1000 ✅ 0.1475 (+0.0000) ✅ 0.1456 (+0.0000) ✅ 37.180 (base 42.986)

Context pipeline (per scenario)

scenario tokens dropped dedup
large_catalog 1480 (base 1480, Δ+0) 0 (base 0, Δ+0) 0 (base 0, Δ+0)
long_conversation 2500 (base 2500, Δ+0) 0 (base 0, Δ+0) 0 (base 0, Δ+0)
mixed_payload 488 (base 488, Δ+0) 0 (base 0, Δ+0) 0 (base 0, Δ+0)
short_conversation 487 (base 487, Δ+0) 0 (base 0, Δ+0) 0 (base 0, Δ+0)
stress_conversation 6590 (base 6590, Δ+0) 11 (base 11, Δ+0) 4 (base 4, Δ+0)
tiny_payload 256 (base 256, Δ+0) 0 (base 0, Δ+0) 0 (base 0, Δ+0)

Numbers come from make benchmark / make benchmark-matrix.
Latency is hardware-dependent — treat the markers as a rough guide.
See benchmarks/scorecard.md for the full picture.

@dgenio
dgenio merged commit 101fb33 into main Sep 17, 2026
16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants