RFM segmentation, clustering, churn scoring, and CLV estimation for a customer base usually means stitching together several separate tools (scikit-learn for clustering, a hand-rolled RFM script, a separate churn model), each with its own performance ceiling on a large transaction table.
A Rust-powered customer segmentation engine with Python bindings.
RFM analysis, KMeans/K-Prototypes clustering, churn prediction, customer lifetime value, SQL export to 8 warehouse dialects, differential privacy / k-anonymity, real-time streaming segmentation, drift detection, lookalike audiences, cohort analytics, lifecycle tracking, rule-based behavioral segmentation, and segment profiling — all real, tested, and callable from Python today.
- Building a marketing segmentation pipeline on a large transaction table where scikit-learn's Python-level clustering is the bottleneck — the Rust core is rayon-parallelized and deterministic for a given seed.
- Exporting customer segments directly into a warehouse —
export_segment_sqlgenerates injection-safe SQL for 8 dialects rather than hand-writing per- warehouse export scripts. - Privacy-constrained segmentation — differential privacy (Laplace/ Gaussian noise) and k-anonymity suppression/generalization are real, tested primitives, not a compliance checkbox.
- Not yet a good fit for: anything needing K-Prototypes' categorical
support through the main
AudienceSegmenterclass (numeric-only today — see Known Issues); Linux/Windows deployment viapip install(macOS ARM64 wheel only — see Installation).
from clusteraudiencekit import AudienceSegmenter, RFMConfig, calculate_rfm
# transactions: list of (customer_id, iso8601_date, amount).
# n_clusters must be <= the number of distinct customers — use your real,
# larger transaction history here; this toy example has 3 customers.
transactions = [
("cust_1", "2026-06-01T00:00:00+00:00", 120.0),
("cust_1", "2026-07-15T00:00:00+00:00", 80.0),
("cust_2", "2026-01-10T00:00:00+00:00", 15.0),
("cust_3", "2026-08-01T00:00:00+00:00", 500.0),
]
# Real RFM scoring (recency/frequency/monetary, quintile-scored, 13-segment
# classification), not a mock.
scores = calculate_rfm(transactions, RFMConfig())
# Cluster customers by their RFM features with real KMeans (k-means++ init,
# rayon-parallelized assignment step, deterministic for a given seed).
features = [[s.recency, s.frequency, s.monetary] for s in scores]
segmenter = AudienceSegmenter(2)
segmenter.fit(features)
segments = segmenter.predict(features)Marketing/data teams need RFM segmentation, clustering, churn scoring, and
CLV estimation, usually stitched together from several tools. This package
does the core numeric work in Rust (fast, deterministic, real unit-tested
algorithms — not scikit-learn wrappers) with a Python API, so you get one
dependency instead of five, and you can inspect exactly what's real (see
docs/ROADMAP_HONEST.md — this project tracks its
own honesty about what's implemented vs. planned, on purpose).
Everything below is backed by real Rust logic with cargo test coverage
and exposed through the compiled Python extension (import clusteraudiencekit) with its own Python-level tests — not a stub, not a
mock, not aspirational documentation.
| Capability | Python entry points |
|---|---|
| RFM analysis | calculate_rfm, RFMConfig, RFMScore |
| KMeans / K-Prototypes clustering | kmeans, AudienceSegmenter |
| Cluster quality metrics | silhouette_score, davies_bouldin_score, calinski_harabasz_score, assess_cluster_quality |
| Automatic K selection | estimate_k_elbow, estimate_k_gap_statistic, estimate_k_silhouette, estimate_k_combined |
| Churn prediction (incl. real AUC-ROC) | ChurnPrediction, ChurnRiskLevel |
| Customer lifetime value | CustomerLTV, calculate_simple_ltv |
| SQL export (8 dialects, injection-safe) | export_segment_sql, export_all_segments_sql, get_supported_sql_dialects |
| Differential privacy & k-anonymity | PyPrivacyBudget, add_laplace_noise, add_gaussian_noise, check_k_anonymity, suppress_to_k_anonymous, generalize_numeric |
| Real-time streaming segmentation | PyStreamingSegmentationEngine, PyStreamingEvent, PyStreamingConfig |
| Drift detection | kolmogorov_smirnov, hellinger_distance, chi_square_drift, detect_feature_drift, detect_segment_composition_change |
| Lookalike audiences | generate_lookalike, find_similar_customers, cosine_similarity |
| Cohort analytics | create_cohort, cohort_id_for, compare_cohorts, aggregate_cohorts_by_period, cohort_retention_table |
| Lifecycle tracking | classify_lifecycle_stage, lifecycle_retention_actions, lifecycle_stage_distribution |
| Rule-based behavioral segmentation | PyBehavioralSegmenter, PyBehavioralSegment, PyBehavioralRule, PyCondition |
| Segment profiling | profile_segment |
Segmentation output: 13 named RFM segments (Champions, Loyal Customers, Potential Loyalists, At Risk, Cannot Lose Them, About to Sleep, New Customers, Promising, Need Attention, Lost, At Risk - Sleeping, Hibernating, VIP).
segment_intelligence, pattern_discovery, temporal_analytics,
price_intelligence, revenue_intelligence, and neural_networks are real,
tested Rust modules (not stubs) that are large enough we deferred wiring
them to a follow-up release rather than rush it. See
docs/ROADMAP_HONEST.md for specifics on each.
External platform activation (pushing segments to ad/CRM platforms),
B2B governance/workflow tooling, a dashboard UI, and a plugin framework are
deliberately not part of this library — see
docs/ROADMAP_HONEST.md for why.
- Python 3.8+
- NumPy, Pandas, PyArrow (see
pyproject.tomlfor exact ranges) - Precompiled Rust core (ships as a platform wheel; no local Rust toolchain needed to install) — but see the platform caveat below
pip install clusteraudiencekitThe only wheel currently published to PyPI is macOS ARM64 (cp39). There is
no source distribution and no CI pipeline building Linux/Windows wheels yet,
so pip install will fail on other platforms today — see
Known Issues. To use this on Linux/Windows/other Python
versions in the meantime, clone the repo and build locally with
maturin (pip install maturin && maturin develop --release), which does require a Rust toolchain.
- Honest roadmap — what's real, what's deferred, and why.
- Architecture — real module layout and data
flow, checked against
src/(replaces two previous architecture docs that described integrations and directory layouts that never existed — seedocs/archive/README.md). - Security audit
- SQL export reference
- Examples
No single OSS package covers everything here, so this compares the two
pieces with direct real competitors: clustering vs scikit-learn's
KMeans, and CLV vs the lifetimes package's real BG/NBD probabilistic
model. Tested against a real public e-commerce dataset — the UCI "Online
Retail" transaction log (541,909 real transactions, cleaned to 397,884
real rows / 4,338 real customers with valid quantity/price/customer ID) —
not synthetic data.
| ClusterAudienceKit | scikit-learn | |
|---|---|---|
| RFM computation (4,338 real customers) | 0.026s | — (not scikit-learn's job) |
| KMeans, K=5, on raw RFM features | 0.001s, silhouette=0.9502 | 0.025s (unscaled, matching CAK's real input), silhouette=0.8387, RuntimeWarning: overflow encountered in matmul |
| KMeans, K=5, standardized features | — (no built-in scaling) | 0.080s, silhouette=0.7744 |
The silhouette numbers above are misleading on their own, and this is
the real, important finding, not the timing. Inspecting the actual
cluster sizes on real, unscaled RFM features: 4,300 of 4,338 real
customers (99.1%) land in one giant cluster, with 4 tiny outlier
clusters (28, 5, 3, 2 customers) splitting off customers with extreme real
monetary values (one real customer spent $280,206 vs. a $2,054 mean).
That's a textbook unscaled-KMeans failure — monetary has ~450x the
numeric range of recency/frequency, so raw Euclidean distance is
dominated entirely by it, and the "high" silhouette score is an artifact
of that degenerate split (one tight blob + isolated far-away outliers),
not a sign of good segmentation. This exact failure mode is what running
the README's own "30-Second Start" example verbatim on real data
produces — AudienceSegmenter does no internal feature standardization,
and neither the quick-start example nor assess_cluster_quality's output
warns you about it. Standardize recency/frequency/monetary yourself (e.g.
sklearn.preprocessing.StandardScaler) before calling .fit() on real
data — not documented anywhere in this repo before this pass, verified as
a real, reproducible gap, not fixed here (changing AudienceSegmenter's
default behavior needs a deliberate decision this pass didn't have scope
to make).
ClusterAudienceKit calculate_simple_ltv |
lifetimes (BG/NBD + Gamma-Gamma) |
|
|---|---|---|
| Method | Deterministic: (total_spent / days_active) * 365 * lifespan_years |
Real probabilistic model fit on repeat-purchase patterns |
| Repeat-purchaser customers scored by both | 2,790 real customers | 2,790 real customers |
| Rank correlation (Spearman ρ) between the two tools' CLV rankings | 0.7111 (p < 1e-300) | — |
A substantial, real positive correlation — the two independently-computed approaches broadly agree on who a business's most valuable real customers are, despite one being a simple rate projection and the other a real statistical model fit to purchase-timing data.
Real bug found and fixed: calculate_simple_ltv — the only CLV
function exposed to Python — used to always return churn_probability: 0.15 for every customer, verified against real data (every one of 4,338
real customers in this benchmark got exactly 0.15, regardless of their
actual recency/frequency/monetary values). Root cause, src/engine/clv.rs
line ~96: this hardcoded value lived in the "Simple" CLV model's
constructor. Fixed to derive a real, per-customer churn signal from the
purchase-frequency/average-order-value data this model already computes
internally, reusing the same risk weighting calculate_probabilistic_ltv's
churn scoring uses in the same file — a low-frequency, low-spend customer
now scores measurably higher churn risk than a frequent, high-spend one
(regression test added, test_simple_ltv_churn_probability_varies_with_real_customer_data).
A separate, more complete calculate_probabilistic_ltv function (which
also uses recency/tenure inputs this simpler model isn't given) still
exists in the same file and is not exposed to Python at all (confirmed:
hasattr(clusteraudiencekit, "calculate_probabilistic_ltv") is False) —
exposing it via new PyO3 bindings is a real, separate feature addition, not
in scope for this fix. docs/ROADMAP_HONEST.md previously said CLV was
"real, shipped" with no caveat about the hardcoded-value bug; corrected.
Verified as of this audit (October 2026):
- Multi-platform wheel-building CI now exists (
.github/workflows/wheels.yml, added this pass) — builds and tests wheels on Linux, macOS (arm64 + x86_64 via cross-compilation), and Windows on every push/PR, and is wired to publish to PyPI via Trusted Publishing on tagged releases. The PyPI-side Trusted Publisher has not yet been configured for this project — the publish job will fail until that one-time setup is done on pypi.org. Until then, releases are still published by hand. The package also now builds as anabi3wheel (Python ≥3.9, one wheel per OS/arch instead of one per Python minor version). cargo audit: 1 open advisory with no fix available.anyhow1.0.102 has an unsoundness advisory inError::downcast_mut()(RUSTSEC-2026-0190). No patched version exists upstream yet; tracked in RepoIssues #34, re-check on futurecargo auditruns.v7.0.0remains installable from PyPI despite a confirmed import-crashing bug. It was never yanked. If you have it pinned, upgrade to>=7.2.0.- K-Prototypes categorical support is partial.
AudienceSegmenter.fit()only accepts a numeric feature matrix today, so selecting K-Prototypes through that class currently runs in numeric-only mode (effectively KMeans). The underlyingengine::clustering::kprototypesRust implementation does support mixed numeric/categorical data; it just isn't reachable from that Python entry point yet. Details indocs/ROADMAP_HONEST.md. - Two Rust utility functions are unimplemented stubs.
pandas_to_arrow/arrow_to_pandasinsrc/utils/conversions.rsreturnErr("Not implemented"). They are not called from anywhere else in the crate and are not exposed to Python, so they don't affect any documented functionality — noted here for completeness. - Registry check: repo version is
7.4.0(Cargo.toml/pyproject.toml); latest PyPI release at time of writing is7.3.2— this release has not been published yet as of this commit. - No open GitHub issues at the time of this audit.
- Six real, tested Rust modules are implemented but not yet wired to the
Python API (
segment_intelligence,pattern_discovery,temporal_analytics,price_intelligence,revenue_intelligence,neural_networks) — see "What's real but not yet exposed to Python" above.
Apache License 2.0. See
LICENSE for the full terms.