Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions HOW_TO_ADD_NEW_CAPABILITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -260,14 +260,14 @@ Recognize by syntactic *shape*: fixed widths, delimiters, or character classes.
Choose Regex when the representation has a distinctive, enumerable shape. All Date, Email, IP, ISBN, and Phone grammars, plus Country's `alpha2`/`alpha3`/`numeric` grammars, are Regex grammars. Three recurring sub-patterns:

- **Compile once, iterate with `finditer()`** — never compile inside `recognize()` (it runs for every input).
- **Sanitize the matched token** — the notation value is a *cleaned* raw token, never a canonical value. Phone strips separators (`strip_separators` in `Phone/grammar/common.py`), Country uppercases, ISBN strips separators and guards length. `raw_text` is always the original span, so `len(raw_text) == end - start` keeps holding.
- **Sanitize the matched token** — the notation value is a *cleaned* raw token, never a canonical value. Phone strips separators (a local `strip_separators` in each Phone grammar), Country uppercases, ISBN strips separators and guards length. `raw_text` is always the original span, so `len(raw_text) == end - start` keeps holding.
- **Guard boundaries against sibling grammars** — when two grammars could claim the same span, use lookbehind/lookahead so each claims only its own representation. See Phone's `national_recognition` (four stacked lookbehinds rejecting `+1…` and `tel:+…` prefixes) and `e164_recognition` (a `(?<![\w:.])` lookbehind rejecting email plus-tags and `tel:+…`).

**Strategy 2 — Lexicon (key-set membership)**

Recognize by membership in a known vocabulary, not by shape. Normalize the input (case-fold, Unicode decomposition, separator folding), test membership in a key-only table under `grammar/data/`, and emit the trimmed input token as the notation value (with a `shape` discriminator when rules route by format). Keys are syntax-normalized forms only — no token maps to a canonical value; rules own every token-to-meaning decision.

Choose Lexicon when the representation is free-form text whose recognizability *is* the vocabulary — no regex shape separates "United States" from "XYZ". Country's `name_recognition` is the exemplar: it normalizes with `normalize_name()`, tests `_KNOWN_NAME_KEYS` (the union of per-locale key sets in `grammar/data/`), and returns the trimmed token with `shape="name"`. See its `grammar/data/` modules for the key-only table pattern.
Choose Lexicon when the representation is free-form text whose recognizability *is* the vocabulary — no regex shape separates "United States" from "XYZ". Country's `name_recognition` is the exemplar: it normalizes with `normalize_name()` (in `Country/notation.py`) for lookup-key membership against `_KNOWN_NAME_KEYS` (the union of per-locale key sets in `grammar/data/`), but emits the original trimmed token (`raw_text` and `notation.value` preserve the input case, not the normalized key) with `shape="name"`. Phone's `strip_separators` (in `paxman/capabilities/Phone/grammar/_common.py`) and ISBN's digit extraction follow the same ownership model: syntax-only cleaning lives in the grammar layer, semantic mapping in rules. See its `grammar/data/` modules for the key-only table pattern.

**Decision guidance:**

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -150,7 +150,7 @@ def active_grammars(self) -> list[str]:
return [name for name, active in grammar_rules.items() if active]
```

**Benefit:** Adding a grammar = one dict entry, no method edit needed.
**Benefit:** Adding a grammar requires one dictionary entry in this property.

---

Expand All @@ -173,7 +173,7 @@ All four architectural deepening candidates have been implemented:

### Architecture Layers

```
```text
paxman/core/ Domain objects, protocols, discovery
paxman/capabilities/ Capability implementations (Email, Date, Country)
paxman/engine/ Pipeline orchestrator
Expand All @@ -182,7 +182,7 @@ paxman/api/ Public API entry points

### Import Boundaries (ADR-0001)

```
```text
paxman.core → (no imports from capabilities, engine, api)
paxman.capabilities → (can import from core)
paxman.engine → (can import from core, capabilities)
Expand All @@ -191,7 +191,7 @@ paxman.api → (can import from everything)

### Hot Spots (Recent Commits)

```
```text
05732ae refactor(core): address code review findings for Email capability
1a69f7e fix: quality gate fixes (ruff, pyright, import-linter, formatting)
5de47b7 feat(capabilities): add Email capability exports
Expand All @@ -203,7 +203,7 @@ d0a6f2f test(integration): add ambiguity detection tests

### Test Structure

```
```text
tests/unit/ Domain object immutability and protocol compliance
tests/capabilities/ Grammar recognition and rule normalization
tests/integration/ Full pipeline flow, ambiguity, temporal filtering
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -170,12 +170,15 @@ from `ValidationError` (itself a `ValueError`): `InvalidFormat`, `InvalidChecksu
normalization happens (separator stripping, Unicode folding, case folding).
This is exactly the "grammars emit raw tokens; one shared syntax seam" shape
Paxman's F3 resolution already gestured at with
`paxman/capabilities/Country/name_normalization.py`.
`paxman/capabilities/Country/notation.py:normalize_name` (lookup-key
normalization, distinct from the emitted notation which preserves the original
trimmed case).
- **Validation/formatting separation (F4).** `validate()` returns the canonical
minimal representation; `format()` is purely presentation and is per-module
parameterized (`separator`, `add_check_digit`). This matches the
already-shipped `Capability.format_value()` seam: rules emit a default
canonical form, formatting is a separate seam.
minimal representation; `format()` is purely presentation only for separator
changes (`separator` parameter), while `add_check_digit=True` appends a Luhn
check digit to a 14-digit IMEI and therefore is not purely presentation.
This matches the already-shipped `Capability.format_value()` seam: rules emit
a default canonical form, formatting is a separate seam.

### What it does not solve

Expand Down Expand Up @@ -453,7 +456,7 @@ return all trees, (3) visit the SPPF yourself
| Paxman invariant | stdnum | phonenumbers | Lark |
|------------------|--------|--------------|------|
| Deterministic / replay-safe | Yes (pure functions) | Yes (pure functions) | Yes, but `dynamic_complete` cost is input-dependent |
| No guessing | Yes (`validate` is all-or-nothing) | **No** -- matcher picks best match per span | Default is best-derivation (**no**); `ambiguity='explicit'` is **yes** |
| No guessing | Yes (`validate` is all-or-nothing) | **No** -- matcher picks first candidate accepted by the configured leniency per span (POSSIBLE may accept possible-but-invalid numbers; VALID is default) | Default is best-derivation (**no**); `ambiguity='explicit'` is **yes** |
| AMBIGUOUS preserved | N/A (no ambiguity model) | **No** -- collapses to one match | **Yes** -- retains all derivations (but syntactic) |
| Provenance-first | No provenance concept | Implicit metadata only | No provenance concept |
| No-raise rule policy | **No** -- raises typed exceptions | Returns bools | N/A |
Expand All @@ -473,11 +476,14 @@ candidates for `01/02/2026`.
Adopt a **span-first recognition contract** and a **shared syntax-normalization
seam**, regularized by an engine-enforced precedence order. Concretely:

1. **Carry source spans in `RecognizedRep`** (add `start: int`, `end: int`,
`raw_text: str`, mirroring `PhoneNumberMatch`). Keep `Grammar.recognize()`
returning `list[NotationT]` but make the span available at the seam the
engine already owns (`RecognizedRep` construction). Addresses Tier 2 #2 and
#3; without spans neither uniform dedup nor uniform ordering is possible.
1. **Carry source spans via `RecognitionMatch`** — `Grammar.recognize()`
produces span-bearing `RecognitionMatch` values (`start`, `end`, `raw_text`
alongside `notation`, mirroring `PhoneNumberMatch`), and `RecognizedRep`
preserves `start`, `end`, and `raw_text` alongside the notation. Do not let
recognition reduce results to bare `NotationT` values before this metadata
reaches the engine, so uniform deduplication and document ordering remain
possible. Addresses Tier 2 #2 and #3; without spans neither uniform dedup nor
uniform ordering is possible.
2. **Unify dedup as per-grammar span-overlap, engine-enforced.** Replace the
three divergent mechanisms (part-keyed, address-keyed, span-overlap) with one
rule: within a single grammar's output, drop a recognition whose span is
Expand All @@ -493,10 +499,18 @@ seam**, regularized by an engine-enforced precedence order. Concretely:
construction.
4. **Standardize the recognition level at "raw tokens + shared syntax seam".**
Grammars emit raw recognized tokens; a capability-level syntax-normalization
helper (the stdnum `clean()`/`compact()` idea, already partially realized as
`Country/name_normalization.py`) does case folding, Unicode cleanup, and
separator stripping; semantic mapping stays in rules (Tier 2 #1; reinforces
the F3 resolution). Optionally expose a phonenumbers-style lenient/authoritative
helper (the stdnum `clean()`/`compact()` idea) does case folding, Unicode
cleanup, and separator stripping — preserved in the notation layer as
`paxman/capabilities/Country/notation.py:normalize_name` (lookup-key
transform, distinct from the emitted notation which preserves the original
trimmed case), and in grammar layers as separator stripping in
`paxman/capabilities/Phone/grammar/_common.py:strip_separators` and digit
extraction in `paxman/capabilities/ISBN/grammar/*_recognition.py`; semantic
mapping stays in rules (Tier 2 #1; reinforces the F3 resolution). Country's
`normalize_name` remains in the notation layer (not moved into grammar
code); Phone/ISBN follow the same principle with syntax normalization in
grammar layers and semantic mapping in rules.
Optionally expose a phonenumbers-style lenient/authoritative
two-tier validation as a *diagnostic* (reason codes), not a behavior change.
5. **Keep the existing `format_value` seam and the no-raise rule policy.**
Validation/formatting separation is already shipped (F4, 2026-08-04) and
Expand Down
1 change: 0 additions & 1 deletion paxman/capabilities/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,6 @@ Every capability must conform to the same structural surface. `CapabilityContrac
- **Quality gates before merge** — `ruff check`, `ruff format --check`, `pyright` (strict), `import-linter lint`, `pytest` (95% coverage).

## ANTI-PATTERNS & LEGACY EXCEPTIONS
- Don't imitate the one-off modules: `Country/name_normalization.py` and `Phone/grammar/common.py` predate the sanctioned strategies — not patterns to copy.
- Don't force a representation into a regex that fights it — consult HOW_TO's recognition-strategy section (scanner, format-candidate, parser combinators, Unicode-property, automaton) before choosing.
- Don't invert the two-locus gating model (e.g., gating recognition on authority features or gating validation on input-shape features) — it produces the wrong `Resolution` statuses.
- Don't add `slots=True` to contracts.
Expand Down
3 changes: 1 addition & 2 deletions paxman/capabilities/Country/capability.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,8 +9,7 @@
from paxman.capabilities.Country.grammar.alpha3_recognition import Alpha3Grammar
from paxman.capabilities.Country.grammar.name_recognition import NameGrammar
from paxman.capabilities.Country.grammar.numeric_recognition import NumericGrammar
from paxman.capabilities.Country.name_normalization import normalize_name
from paxman.capabilities.Country.notation import CountryNotation
from paxman.capabilities.Country.notation import CountryNotation, normalize_name
from paxman.capabilities.Country.rules.cldr_localized_ed2025 import (
SectionLocalizedNames,
)
Expand Down
48 changes: 20 additions & 28 deletions paxman/capabilities/Country/grammar/alpha2_recognition.py
Original file line number Diff line number Diff line change
@@ -1,16 +1,28 @@
"""Alpha-2 country code recognition grammar."""
"""Alpha-2 country code recognition grammar (staged pipeline).

Recognizes exactly 2 ASCII letters as an alpha-2 country code shape. The word
boundary is supplied by BoundaryGuard.word_only() (ADR-0008 D5) so no
hard-coded lookaround literal remains in this file. Syntax only: the grammar
never resolves the code to a country.
"""

from __future__ import annotations

import re

from paxman.capabilities.Country.notation import CountryNotation
from paxman.core.domain import Grammar, RecognitionMatch
from paxman.core.grammar import BoundaryGuard, PipelineGrammar, RegexStage, StandardPre

_GUARD = BoundaryGuard.word_only()
_ALPHA2_PATTERN = _GUARD.lookbehind + r"[A-Za-z]{2}" + _GUARD.lookahead


_ALPHA2_PATTERN = re.compile(r"\b[A-Za-z]{2}\b")
def _alpha2_notation(match: re.Match[str]) -> CountryNotation:
"""Map an alpha-2 match to its upper-cased notation."""
return CountryNotation(shape="alpha2", value=match.group(0).upper())


class Alpha2Grammar(Grammar[CountryNotation]):
class Alpha2Grammar(PipelineGrammar[CountryNotation]):
"""Recognizes exactly 2 ASCII letters as alpha-2 country code shape.

Examples: "US", "GB", "us", "gB"
Expand All @@ -21,27 +33,7 @@ class Alpha2Grammar(Grammar[CountryNotation]):
semantics = "alpha2_recognition"
single_value = True

def recognize(self, text: str) -> list[RecognitionMatch[CountryNotation]]:
"""Extract alpha-2 patterns from text.

Args:
text: Raw input text.

Returns:
List of span-bearing matches with shape="alpha2" notations.
"""
if not text.strip():
return []
matches: list[RecognitionMatch[CountryNotation]] = []
for match in _ALPHA2_PATTERN.finditer(text):
matches.append(
RecognitionMatch(
notation=CountryNotation(
shape="alpha2", value=match.group(0).upper()
),
start=match.start(),
end=match.end(),
raw_text=match.group(0),
)
)
return matches
pre = StandardPre[CountryNotation](empty_guard=True)
regex = RegexStage[CountryNotation](
pattern=_ALPHA2_PATTERN, notation_fn=_alpha2_notation
)
48 changes: 20 additions & 28 deletions paxman/capabilities/Country/grammar/alpha3_recognition.py
Original file line number Diff line number Diff line change
@@ -1,16 +1,28 @@
"""Alpha-3 country code recognition grammar."""
"""Alpha-3 country code recognition grammar (staged pipeline).

Recognizes exactly 3 ASCII letters as an alpha-3 country code shape. The word
boundary is supplied by BoundaryGuard.word_only() (ADR-0008 D5) so no
hard-coded lookaround literal remains in this file. Syntax only: the grammar
never resolves the code to a country.
"""

from __future__ import annotations

import re

from paxman.capabilities.Country.notation import CountryNotation
from paxman.core.domain import Grammar, RecognitionMatch
from paxman.core.grammar import BoundaryGuard, PipelineGrammar, RegexStage, StandardPre

_GUARD = BoundaryGuard.word_only()
_ALPHA3_PATTERN = _GUARD.lookbehind + r"[A-Za-z]{3}" + _GUARD.lookahead


_ALPHA3_PATTERN = re.compile(r"\b[A-Za-z]{3}\b")
def _alpha3_notation(match: re.Match[str]) -> CountryNotation:
"""Map an alpha-3 match to its upper-cased notation."""
return CountryNotation(shape="alpha3", value=match.group(0).upper())


class Alpha3Grammar(Grammar[CountryNotation]):
class Alpha3Grammar(PipelineGrammar[CountryNotation]):
"""Recognizes exactly 3 ASCII letters as alpha-3 country code shape.

Examples: "USA", "GBR", "usa", "gbr"
Expand All @@ -21,27 +33,7 @@ class Alpha3Grammar(Grammar[CountryNotation]):
semantics = "alpha3_recognition"
single_value = True

def recognize(self, text: str) -> list[RecognitionMatch[CountryNotation]]:
"""Extract alpha-3 patterns from text.

Args:
text: Raw input text.

Returns:
List of span-bearing matches with shape="alpha3" notations.
"""
if not text.strip():
return []
matches: list[RecognitionMatch[CountryNotation]] = []
for match in _ALPHA3_PATTERN.finditer(text):
matches.append(
RecognitionMatch(
notation=CountryNotation(
shape="alpha3", value=match.group(0).upper()
),
start=match.start(),
end=match.end(),
raw_text=match.group(0),
)
)
return matches
pre = StandardPre[CountryNotation](empty_guard=True)
regex = RegexStage[CountryNotation](
pattern=_ALPHA3_PATTERN, notation_fn=_alpha3_notation
)
2 changes: 1 addition & 1 deletion paxman/capabilities/Country/grammar/data/chinese_names.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@

from __future__ import annotations

from paxman.capabilities.Country.name_normalization import normalize_name
from paxman.capabilities.Country.notation import normalize_name

CHINESE_NAME_KEYS: frozenset[str] = frozenset(
normalize_name(key)
Expand Down
2 changes: 1 addition & 1 deletion paxman/capabilities/Country/grammar/data/english_names.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@

from __future__ import annotations

from paxman.capabilities.Country.name_normalization import normalize_name
from paxman.capabilities.Country.notation import normalize_name

ENGLISH_NAME_KEYS: frozenset[str] = frozenset(
normalize_name(key)
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@

from __future__ import annotations

from paxman.capabilities.Country.name_normalization import normalize_name
from paxman.capabilities.Country.notation import normalize_name

HISTORICAL_NAME_KEYS: frozenset[str] = frozenset(
normalize_name(key)
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@

from __future__ import annotations

from paxman.capabilities.Country.name_normalization import normalize_name
from paxman.capabilities.Country.notation import normalize_name

LOCALIZED_NAME_KEYS: frozenset[str] = frozenset(
normalize_name(key)
Expand Down
44 changes: 10 additions & 34 deletions paxman/capabilities/Country/grammar/name_recognition.py
Original file line number Diff line number Diff line change
Expand Up @@ -20,17 +20,16 @@
from paxman.capabilities.Country.grammar.data.localized_names import (
LOCALIZED_NAME_KEYS,
)
from paxman.capabilities.Country.name_normalization import normalize_name
from paxman.capabilities.Country.notation import CountryNotation
from paxman.core.domain import Grammar, RecognitionMatch
from paxman.capabilities.Country.notation import CountryNotation, normalize_name
from paxman.core.grammar import PipelineGrammar, StandardPre, WholeInputLookup

# Union of every recognized name representation across locales.
_KNOWN_NAME_KEYS = (
_KNOWN_NAME_KEYS = frozenset(
ENGLISH_NAME_KEYS | HISTORICAL_NAME_KEYS | CHINESE_NAME_KEYS | LOCALIZED_NAME_KEYS
)


class NameGrammar(Grammar[CountryNotation]):
class NameGrammar(PipelineGrammar[CountryNotation]):
"""Recognizes country name representations from recognition key sets.

Decides whether an input is a known country name representation and
Expand All @@ -52,32 +51,9 @@ class NameGrammar(Grammar[CountryNotation]):
semantics = "name_recognition"
single_value = True

def recognize(self, text: str) -> list[RecognitionMatch[CountryNotation]]:
"""Extract a country name representation from text.

Args:
text: Raw input text.

Returns:
A list with a single span-bearing match of shape="name" carrying
the trimmed input token when the token is a known name
representation, or an empty list for empty/unknown input.
"""
trimmed = text.strip()
if not trimmed:
return []

normalized = normalize_name(trimmed)

if normalized in _KNOWN_NAME_KEYS:
start = len(text) - len(text.lstrip())
return [
RecognitionMatch(
notation=CountryNotation(shape="name", value=trimmed),
start=start,
end=start + len(trimmed),
raw_text=trimmed,
)
]

return []
pre = StandardPre[CountryNotation](empty_guard=True)
lexicon = WholeInputLookup[CountryNotation](
keys=_KNOWN_NAME_KEYS,
normalizer=normalize_name,
notation_fn=lambda trimmed: CountryNotation(shape="name", value=trimmed),
)
Loading
Loading