Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
22 commits
Select commit Hold shift + click to select a range
ebf21bb
feat(core): declare semantics on all shipped grammars
azaharizaman Aug 11, 2026
f8ac1f6
refactor: rename target_grammars to target_semantics
azaharizaman Aug 11, 2026
811c30f
docs(adr): add ADR-0003 semantic affinity routing and implementation …
azaharizaman Aug 11, 2026
0ea7727
feat(core): enforce Grammar.semantics at class-definition time
azaharizaman Aug 11, 2026
0f0f9d1
refactor(engine): route on semantics, not grammar name
azaharizaman Aug 11, 2026
3c592b8
refactor(capabilities): coalesce Date grammars to calendar-date seman…
azaharizaman Aug 11, 2026
9317596
refactor(capabilities): coalesce Email grammars to addr-spec semantics
azaharizaman Aug 11, 2026
fd22152
test: widen _ProbeRow notation type for multi-capability groups
azaharizaman Aug 11, 2026
abae3e4
refactor(capabilities): coalesce Phone grammars to E.164 semantics
azaharizaman Aug 11, 2026
36315dc
test: lock same-semantics field-mapping consistency
azaharizaman Aug 11, 2026
f6f7790
docs: sweep target_grammars and document semantic affinity
azaharizaman Aug 11, 2026
125c06d
docs: ruff-format code blocks in swept documentation
azaharizaman Aug 11, 2026
42fd19e
test: tighten same-semantics guard coverage (oracle review)
azaharizaman Aug 11, 2026
73b04f4
docs: clarify semantics-id opt-in in extra_grammars (oracle review)
azaharizaman Aug 11, 2026
9ce5fdb
fix: bound slash-ISO pattern, clarify activation docs (coderabbit rev…
azaharizaman Aug 11, 2026
78794cc
fix: bound remaining date grammars against embedded digits
azaharizaman Aug 11, 2026
634aef8
docs: sync slash-ISO docstring with bounded date grammar set
azaharizaman Aug 11, 2026
22d461e
docs(adr): correct target_grammars inventory to verified 54 files
azaharizaman Aug 11, 2026
e02c96d
test: guard grammar-semantics rule coverage, fix dangling-target docs…
azaharizaman Aug 11, 2026
98f4255
test: pin renamed singleton semantics ids, cover raw semantics-id opt-in
azaharizaman Aug 11, 2026
d307ae1
docs: correct ADR-0003 inventory count, note lookaround change
azaharizaman Aug 11, 2026
500e9de
docs: align plan gate scope, digit-glued exclusion, sweep proofs
azaharizaman Aug 12, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -97,7 +97,7 @@ The engine is the orchestration layer that coordinates the full pipeline. It:

The engine is capability-agnostic. It does not know what a "grammar" or "rule" does — it only knows that grammars produce span-bearing recognition matches and rules produce candidates.

Before the recognition phase, the engine composes each capability's shipped grammars and rules with any community extensions registered for that capability (see "Community Extensions" below). Composition is guarded: duplicate names fail fast, and every rule's declared `target_grammars` must resolve within the composed set.
Before the recognition phase, the engine composes each capability's shipped grammars and rules with any community extensions registered for that capability (see "Community Extensions" below). Composition is guarded: duplicate names fail fast, and every rule's declared `target_semantics` must resolve within the composed set.

### Public API

Expand Down Expand Up @@ -171,9 +171,9 @@ Capabilities are closed for modification but open for extension. Community contr

A contract opts a registered grammar in by naming it in `extra_grammars`, a base `CapabilityContract` field surfaced on every `create_contract` factory. The engine composes the shipped active set with the opted-in extras, deduplicating names while preserving order — shipped slots first, extras after (unknown extra names are silently skipped). The shipped slots are `contract.active_grammars` when the contract implements it (the gated capabilities), or every shipped grammar in `get_grammars()` order when it returns `None` (the base default). Opt-in preserves determinism: a contract that names no extras composes to exactly the shipped set, so non-opt-in behavior is byte-identical.

Community rules follow the same opt-in discipline: a registered rule runs only when the contract names one of its `target_grammars` in `extra_grammars`. An un-opted community rule — even one targeting a shipped grammar — never affects results, so a default contract resolves with shipped rules only.
Community rules follow the same opt-in discipline: a registered rule runs only when the contract's `extra_grammars` resolve to one of its `target_semantics` ids. An un-opted community rule — even one targeting a shipped grammar's semantics — never affects results, so a default contract resolves with shipped rules only.

Composition is guarded at pipeline start: a community grammar name colliding with a shipped name raises `CapabilityError`, and an opted-in community rule whose `target_grammars` names a missing grammar raises `ContractError` — failing fast rather than producing a silently wrong result. Community grammars and rules are pure functions of their inputs, and the composed set is fixed once the registries freeze, so the determinism guarantees of "Determinism by Construction" extend unchanged.
Composition is guarded at pipeline start: a community grammar name colliding with a shipped name raises `CapabilityError`, and an opted-in community rule whose `target_semantics` names an id no grammar claims raises `ContractError` — failing fast rather than producing a silently wrong result. Community grammars and rules are pure functions of their inputs, and the composed set is fixed once the registries freeze, so the determinism guarantees of "Determinism by Construction" extend unchanged.

---

Expand Down
12 changes: 8 additions & 4 deletions CONTEXT.md
Original file line number Diff line number Diff line change
Expand Up @@ -137,7 +137,7 @@ class Section341AddrSpec(Rule[EmailNotation]):
strategy = RuleStrategy.REGEX
provenance = PUBLICATION
citation = "Section 3.4.1 (addr-spec)" # Human-readable citation
target_grammars = frozenset({"standard_recognition", "obfuscated_recognition"})
target_semantics = frozenset({"rfc5322_addr_spec"})
requires_features = frozenset() # authority features this rule gates on

def matches(self, notation: EmailNotation, contract: Contract) -> bool:
Expand All @@ -152,7 +152,7 @@ class Section341AddrSpec(Rule[EmailNotation]):
return f"{notation.local_part.lower()}@{notation.domain_part.lower()}"
```

Every rule declares six metadata attrs — `name`, `strategy`, `provenance`, `citation`, `target_grammars` (non-empty `frozenset[str]`), `requires_features` (`frozenset[str]`) — enforced at import time by `Rule.__init_subclass__`. Rules never raise and never read `output_format`.
Every rule declares six metadata attrs — `name`, `strategy`, `provenance`, `citation`, `target_semantics` (non-empty `frozenset[str]`), `requires_features` (`frozenset[str]`) — enforced at import time by `Rule.__init_subclass__`. Every grammar declares a `semantics` string — the meaning id its recognized notations carry (its identity id by default; a coalesced id shared by same-meaning grammars, e.g. the standard and obfuscated Email grammars both declare `"rfc5322_addr_spec"`) — enforced as a non-empty `str` at import time by `Grammar.__init_subclass__`. Rules never raise and never read `output_format`.

### Notation Purpose
Notation exists for **placement-sensitive rules**:
Expand All @@ -173,13 +173,16 @@ The resolver **consumes notation** and outputs a canonical_value (not notation).
```python
# capabilities/Currency/rules/iso_4217_ed2015.py


class SectionCode(Rule[CurrencyNotation]):
"""ISO 4217 Section 3 - Currency and funds codes"""

name = "Section 3-code"
strategy = RuleStrategy.LOOKUP_TABLE
provenance = PUBLICATION
target_grammars = frozenset({"code_recognition", "symbol_recognition", "word_recognition"})
target_semantics = frozenset(
{"code_recognition", "symbol_recognition", "word_recognition"}
)
requires_features = frozenset()

def matches(self, notation: CurrencyNotation, contract: Contract) -> bool:
Expand All @@ -204,7 +207,7 @@ class Section431CalendarDate(Rule[DateNotation]):
name = "Section 4.3.1-calendar-date"
strategy = RuleStrategy.PARSER
provenance = PUBLICATION
target_grammars = frozenset({"iso8601_recognition"})
target_semantics = frozenset({"iso8601_calendar_date"})
requires_features = frozenset()

def matches(self, notation: DateNotation, contract: Contract) -> bool:
Expand Down Expand Up @@ -244,6 +247,7 @@ class StandardEmailGrammar(Grammar[EmailNotation]):
"""Standard email recognition: user@domain.tld"""

name = "standard_recognition"
semantics = "rfc5322_addr_spec" # coalesced — shared with obfuscated_recognition

def recognize(self, text: str) -> list[RecognitionMatch[EmailNotation]]:
"""Extract span-bearing email matches from text."""
Expand Down
17 changes: 9 additions & 8 deletions HOW_TO_ADD_NEW_CAPABILITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -169,6 +169,7 @@ class StandardMyDomainGrammar(Grammar[MyDomainNotation]):
"""Standard recognition for the MyDomain capability."""

name = "standard_recognition"
semantics = "standard_recognition"

def recognize(self, text: str) -> list[RecognitionMatch[MyDomainNotation]]:
"""Extract span-bearing matches from raw text.
Expand Down Expand Up @@ -289,7 +290,7 @@ Codebase examples: IP's IPv4 grammar is a loose regex (`\d{1,3}` octets) and `rf

3. Set `provenance` to the `PUBLICATION` constant defined above
4. Set `citation` to a human-readable citation (e.g., "Section 3.4.1 (addr-spec)")
5. Set `target_grammars` to the `frozenset[str]` of grammar names whose notations this rule validates (e.g., `frozenset({"standard_recognition"})`)
5. Set `target_semantics` to the `frozenset[str]` of grammar semantics whose notations this rule validates (e.g., `frozenset({"standard_recognition"})`)
6. Set `requires_features` to the `frozenset[str]` of contract fields that must be truthy for the rule to run (`frozenset()` when it always runs)

All six attributes are enforced by `Rule.__init_subclass__` at class-definition time; see the Rule metadata section in Step 7.
Expand Down Expand Up @@ -438,19 +439,19 @@ Capability-specific parameters come after the common block. Every capability sat

**3. Rule metadata**

Every `Rule` subclass must declare six class attributes: `name`, `strategy`, `provenance`, `citation`, `target_grammars`, and `requires_features`. `Rule.__init_subclass__` enforces this at class-definition time, and a subclass missing any of them fails to import with a `TypeError`:
Every `Rule` subclass must declare six class attributes: `name`, `strategy`, `provenance`, `citation`, `target_semantics`, and `requires_features`. `Rule.__init_subclass__` enforces this at class-definition time, and a subclass missing any of them fails to import with a `TypeError`:

```python
class SectionYourRule(Rule[YourDomainNotation]):
name = "Section 1-your-rule"
strategy = RuleStrategy.REGEX
provenance = PUBLICATION
citation = "Section 1 (your-rule)"
target_grammars = frozenset({"your_recognition"})
target_semantics = frozenset({"your_recognition"})
requires_features = frozenset()
```

- **`target_grammars: ClassVar[frozenset[str]]`** is the non-empty set of grammar names whose notations this rule validates. The engine uses it for affinity routing: each recognition is validated only by rules whose `target_grammars` includes the producing grammar's name, and a rule declaring a grammar the capability does not have fails fast with a `ContractError` before any candidate is produced. `Rule.__init_subclass__` also rejects an empty set at import time, since such a rule could never match a recognition.
- **`target_semantics: ClassVar[frozenset[str]]`** is the non-empty set of grammar `semantics` ids whose notations this rule validates. The engine uses it for affinity routing: each recognition is validated only by rules whose `target_semantics` includes the producing grammar's `semantics`, and a rule declaring a semantics id no grammar claims fails fast with a `ContractError` before any candidate is produced. `Rule.__init_subclass__` also rejects an empty set at import time, since such a rule could never match a recognition.
- **`requires_features: ClassVar[frozenset[str]]`** is the set of Contract field names that must be truthy for the rule to run. An empty set is valid and is the common case: it means the rule always runs once selected. The engine validates that every named feature exists on the contract (a missing name raises `ContractError`) and applies the final feature filter *after* pinning, exclusion, and year selection: a rule whose required feature is present but `False` is dropped.

**Feature gating has two loci, and they produce different `Resolution` statuses:**
Expand Down Expand Up @@ -569,7 +570,7 @@ The key invariant: `None`, `"default"`, and the default format string are **trea

Example — Date input `"01/02/2026"` is recognized by both the US and European grammars and validated by both rules, yielding two distinct canonical values (`2026-01-02` and `2026-02-01`). The result is `AMBIGUOUS` regardless of `output_format`. `output_format="US"` merely renders those two values as `01/02/2026` and `02/01/2026`; it cannot and must not decide which interpretation is "correct".

> Note: the grammar→rule routing decision (which rule validates which recognized notation) is an entirely separate concern from `output_format`. Routing is declared on the rule (e.g. `Rule.target_grammars`); it operates in the recognition→validation stage and never touches formatting. Keep the two orthogonal.
> Note: the grammar→rule routing decision (which rule validates which recognized notation) is an entirely separate concern from `output_format`. Routing is declared on the rule (`Rule.target_semantics`) and matched against each grammar's `semantics`; it operates in the recognition→validation stage and never touches formatting. Keep the two orthogonal.

Example wiring — inherited from `CapabilityContract`, you only set the class variables:

Expand Down Expand Up @@ -1032,7 +1033,7 @@ If your rule needs to read a capability-specific parameter (like `two_digit_base

A shipped capability is closed for modification but open for extension: you can add recognition and validation without touching the capability package.

1. **Author** a `Grammar` subclass (Step 4) and a `Rule` subclass (Step 5) for the capability's notation, exactly as you would for a new capability — the same contracts apply, including span-bearing `RecognitionMatch` output, `target_grammars`, and `requires_features`.
1. **Author** a `Grammar` subclass (Step 4) and a `Rule` subclass (Step 5) for the capability's notation, exactly as you would for a new capability — the same contracts apply, including span-bearing `RecognitionMatch` output, `semantics` on the grammar, `target_semantics` on the rule, and `requires_features`.
2. **Register** them before the first `canonicalize()` call:

```python
Expand All @@ -1053,7 +1054,7 @@ Semantics to rely on:

- The extension registries freeze with the capability registry — registration after the first pipeline run raises `CapabilityError`.
- Opt-in only: an un-named registered grammar never affects results, keeping shipped behavior byte-identical for non-opt-in contracts.
- Community rules are opt-in too: a registered rule runs only when the contract names one of its `target_grammars` in `extra_grammars`; an un-opted rule — even one targeting a shipped grammar — never affects results.
- Community rules are opt-in too: a registered rule runs only when the contract's `extra_grammars` resolve to one of its `target_semantics`; an un-opted rule — even one targeting a shipped grammar's semantics — never affects results.
- Unknown `extra_grammars` names are silently skipped; shipped names listed in `extra_grammars` are deduplicated.
- Composition is guarded: a grammar name colliding with a shipped name, or an opted-in community rule naming a missing grammar, fails fast at pipeline start.

Expand All @@ -1066,7 +1067,7 @@ Use this checklist to verify your capability is complete:
- [ ] Notation is a frozen dataclass with `as_list()` method
- [ ] Each grammar extends `Grammar[YourDomainNotation]` and implements `recognize(text) -> list[RecognitionMatch[YourDomainNotation]]`
- [ ] Each rule extends `Rule[YourDomainNotation]` and implements `matches(notation, contract) -> bool` and `normalize(notation, contract) -> str`
- [ ] Each rule declares `target_grammars` (non-empty `frozenset[str]`) and `requires_features` (`frozenset()` when the rule always runs)
- [ ] Each rule declares `target_semantics` (non-empty `frozenset[str]`) and `requires_features` (`frozenset()` when the rule always runs)
- [ ] Each rule file has a `PUBLICATION` provenance constant
- [ ] Capability extends `Capability` and implements `get_grammars()` and `get_rules()`
- [ ] Contract inherits `CapabilityContract` (frozen dataclass, no `slots=True`) and satisfies the `Contract` protocol
Expand Down
Loading
Loading