Skip to content

Release 0.4.0: alphabet, check character, namespace - #7

Merged
neosergio merged 1 commit into
mainfrom
feat/v0.4.0
Jul 13, 2026
Merged

neosergio merged 1 commit into
mainfrom
feat/v0.4.0

Conversation

@neosergio

Copy link
Copy Markdown
Owner

Three changes to how a reference is derived, landing together so callers get one migration rather than three. None of them moves a reference already stored: every new argument defaults to what earlier versions did, and the 0.1.0 strings are still pinned by test.

namespace fixes a bug the feature request led me to. prefix is only a label -- it never reached the hash. So an order and an invoice derived from one customer's ULID drew the same suffix, every time:

ORD-20260713-133083
INV-20260713-133083

That is a certainty, not the one-in-a-million the collision helpers describe, and none of them accounted for it. The namespace rides in BLAKE2b's key, beside the attempt in its salt, so neither can be spelled as part of some other source. An empty key is what BLAKE2b already hashed with, so the default reproduces every reference made before.

alphabet makes the library live up to its name. Six decimal characters hold a million values; six Crockford base32 characters hold 1.07 billion -- 32 times the volume at the same length, which is the single biggest lever on the collision problem the README spends so long apologising for. It is a trade, not a free win: digits work on a numeric keypad and cannot spell anything. Digits stay the default. Every sizing helper gained base=, because a base32 suffix sized as decimal is sized wrong by orders of magnitude.

check is the gap that was most conspicuous in a library for human-facing references: a mistyped one was merely absent, which a caller cannot tell from a record that never existed. Now it is invalid, which is a different answer and an actionable one. Luhn mod N over the alphabet, covering the date as well as the suffix, since a mistyped day lands in the wrong bucket.

I audited it rather than trusting the textbook: exhaustively, it catches every single-character error in both alphabets, and every adjacent transposition except the swap of the alphabet's first and last characters -- 09 <-> 90 in decimal, 0Z <-> Z0 in base32. That is Luhn's known blind spot, 2.7% of transpositions in decimal and 0.3% in base32, and a test pins it so nobody later claims the check is stronger than it is. Damm would close it but needs a totally anti-symmetric quasigroup of the alphabet's order, which cannot be built for an arbitrary alphabet at call time.

The check adds a character but no capacity: it is computed from the reference, not drawn from the hash. Six characters plus a check hold a million values, not ten million, and the docs say so.

Three changes to how a reference is derived, landing together so callers
get one migration rather than three. None of them moves a reference
already stored: every new argument defaults to what earlier versions did,
and the 0.1.0 strings are still pinned by test.

namespace fixes a bug the feature request led me to. prefix is only a
label -- it never reached the hash. So an order and an invoice derived
from one customer's ULID drew the same suffix, every time:

    ORD-20260713-133083
    INV-20260713-133083

That is a certainty, not the one-in-a-million the collision helpers
describe, and none of them accounted for it. The namespace rides in
BLAKE2b's key, beside the attempt in its salt, so neither can be spelled
as part of some other source. An empty key is what BLAKE2b already
hashed with, so the default reproduces every reference made before.

alphabet makes the library live up to its name. Six decimal characters
hold a million values; six Crockford base32 characters hold 1.07 billion
-- 32 times the volume at the same length, which is the single biggest
lever on the collision problem the README spends so long apologising for.
It is a trade, not a free win: digits work on a numeric keypad and cannot
spell anything. Digits stay the default. Every sizing helper gained base=,
because a base32 suffix sized as decimal is sized wrong by orders of
magnitude.

check is the gap that was most conspicuous in a library for human-facing
references: a mistyped one was merely absent, which a caller cannot tell
from a record that never existed. Now it is invalid, which is a different
answer and an actionable one. Luhn mod N over the alphabet, covering the
date as well as the suffix, since a mistyped day lands in the wrong bucket.

I audited it rather than trusting the textbook: exhaustively, it catches
every single-character error in both alphabets, and every adjacent
transposition except the swap of the alphabet's first and last characters
-- 09 <-> 90 in decimal, 0Z <-> Z0 in base32. That is Luhn's known blind
spot, 2.7% of transpositions in decimal and 0.3% in base32, and a test
pins it so nobody later claims the check is stronger than it is. Damm
would close it but needs a totally anti-symmetric quasigroup of the
alphabet's order, which cannot be built for an arbitrary alphabet at call
time.

The check adds a character but no capacity: it is computed from the
reference, not drawn from the hash. Six characters plus a check hold a
million values, not ten million, and the docs say so.
@coderabbitai

coderabbitai Bot commented Jul 13, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@neosergio, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 3 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 5c192161-b8bf-44fb-9f6c-ff1e18698ce7

📥 Commits

Reviewing files that changed from the base of the PR and between 0192033 and 98b7c1e.

📒 Files selected for processing (7)
  • CHANGELOG.md
  • README.md
  • ROADMAP.md
  • pyproject.toml
  • src/compactref/__init__.py
  • src/compactref/core.py
  • tests/test_core.py
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/v0.4.0

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@neosergio
neosergio merged commit 6631839 into main Jul 13, 2026
8 checks passed
@neosergio
neosergio deleted the feat/v0.4.0 branch July 13, 2026 18:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant