Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
113 changes: 113 additions & 0 deletions docs/analytics.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,113 @@
# Analytics

Admin-only commercial reporting, under `/v1/admin/analytics`.

Before this existed, the admin dashboard derived every money figure client-side
from the most recent 100 orders. That is fine for a demo and wrong for a shop:
the numbers stop being true the moment the hundred-and-first order is placed.
These endpoints compute the same figures in SQL over the whole order history.

## Endpoints

| Endpoint | Answers |
|---|---|
| `GET /summary` | What did we take, and how does that compare to last period? |
| `GET /timeseries` | How did that move day by day? |
| `GET /products` | What sold, what came back, what never sold at all? |
| `GET /funnel` | Where do orders stop? |

All four take optional `from` and `to` calendar dates, inclusive, defaulting to
the last 30 days.

## Definitions

These are choices, not facts, so they are written down rather than left implicit
in a query:

| Term | Definition |
|---|---|
| Gross revenue | Orders in `paid`, `ready_to_ship` or `shipped` |
| Refunded | Orders in `refunded` — **whole-order only** |
| Net revenue | Gross less refunded |
| Average order value | Gross divided by revenue-producing orders |
| Units | Line quantities on revenue-producing orders |

Never counted: `draft`, `pending_payment`, `cancelled`, and anything with
`deleted_at` set.

## Four things that are easy to get wrong

**Money is grouped by currency.** `orders.currency` permits more than one, and a
cross-currency total is not slightly wrong — it is meaningless. Every money
figure is therefore a list keyed by currency, which collapses to a single entry
for the usual single-currency shop.

**Order money and line money are queried separately.** Joining `orders` to
`order_items` and summing `total_amount` multiplies each order's value by its
line count. The join is only ever used for quantities and line revenue; order
totals come from their own statement and the two are merged in Python.

**Days are cut in the shop's timezone.** `SHOP_TIMEZONE` (default
`Europe/Berlin`) feeds `timezone(tz, created_at)`, Postgres' `AT TIME ZONE`.
Bucketing on raw UTC moves evening orders into the following day for any shop
east of Greenwich, which is the kind of error that looks like a data problem for
weeks.

**Checkout is read from `payments`, not from `orders.status`.** Status records
only where an order is *now*. A cancelled order is indistinguishable from one
that never reached checkout, and a shipped-then-refunded order no longer says it
shipped. A payment row is written when checkout starts and survives whatever
happens next.

## Known limits

Stated here because a reader will otherwise infer something stronger:

- **The funnel is an order funnel, not a visitor funnel.** It begins at order
creation and cannot see shoppers who browsed and never started one. Visitor
conversion needs session data the API does not collect. That is S2.
- **Partial refunds are not modelled.** `orders.status = refunded` is
all-or-nothing, so refund figures will not reconcile against a partially
refunded Stripe charge.
- **Return rate is an upper bound per SKU.** Returns are recorded per order, not
per line, so a return on a two-line order counts against both SKUs.
- **Product revenue need not equal order revenue.** Product figures sum line
values; an order total may carry shipping or adjustments belonging to no line.

## Performance

Aggregates are computed live rather than from a rollup table — always current,
no staleness, no extra moving parts. Three indexes support it:

```sql
ix_orders_created_at (created_at)
ix_orders_status_created_at (status, created_at)
ix_order_items_sku (sku)
```

The schema is created with `Base.metadata.create_all`, which creates missing
*tables* but does not alter existing ones. A database that predates this change
therefore needs the indexes applied by hand:

```sql
CREATE INDEX IF NOT EXISTS ix_orders_created_at ON orders (created_at);
CREATE INDEX IF NOT EXISTS ix_orders_status_created_at ON orders (status, created_at);
CREATE INDEX IF NOT EXISTS ix_order_items_sku ON order_items (sku);
```

Without them the endpoints still return correct figures, just with a sequential
scan. A window longer than five years is refused outright, since live
aggregation has no rollup behind it to make an unbounded range cheap.

When order volume outgrows live aggregation, the next step is a nightly rollup
table written by the ARQ worker — deliberately not built yet, because it costs a
job, a table and staleness in exchange for speed nobody needs at this size.

## Testing

- `tests/test_analytics_unit.py` — period arithmetic, timezone conversion,
gap filling, undefined-vs-infinite percentage change. No database.
- `tests/test_analytics_integration.py` — seeds a deterministic dataset inside
March 2025, a window nothing else touches, and asserts exact figures. Covers
currency isolation, soft-delete exclusion, the multi-line fan-out trap, the
23:30-UTC timezone boundary, and the payments-based funnel.
4 changes: 4 additions & 0 deletions src/app/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@

from app.chore import lifespan
from app.services.admin import admin_api_router
from app.services.analytics import analytics_api_router
from app.services.crud_item_store import router as item_store_router
from app.services.customers import customers_api_router
from app.services.fulfillment import fulfillment_api_router
Expand Down Expand Up @@ -152,6 +153,9 @@ async def generic_exception_handler(request: Request, exc: Exception) -> JSONRes
# Include admin service router (Phase 2)
app.include_router(admin_api_router, prefix="/v1")

# Include analytics service router (S1)
app.include_router(analytics_api_router, prefix="/v1")

# Include fulfillment service router (Phase 3)
app.include_router(fulfillment_api_router, prefix="/v1")

Expand Down
27 changes: 27 additions & 0 deletions src/app/services/analytics/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
"""
Analytics Service

Admin-only commercial reporting, computed in SQL over the whole order history.

Endpoints:
GET /admin/analytics/summary — headline figures vs the previous period
GET /admin/analytics/timeseries — the same figures bucketed over time
GET /admin/analytics/products — per-SKU performance and dead stock
GET /admin/analytics/funnel — where orders stop

Usage:
from app.services.analytics import analytics_api_router
app.include_router(analytics_api_router, prefix="/v1")
"""

from fastapi import APIRouter

from .routers import analytics_router

analytics_api_router = APIRouter(
prefix="/admin/analytics",
tags=["Analytics"],
)
analytics_api_router.include_router(analytics_router)

__all__ = ["analytics_api_router"]
10 changes: 10 additions & 0 deletions src/app/services/analytics/dependencies.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
"""
Analytics Dependencies

Analytics is admin-only. Re-exports the shared admin guard so the policy has one
home, matching services/admin/dependencies.py.
"""

from app.authorize import require_admin

__all__ = ["require_admin"]
19 changes: 19 additions & 0 deletions src/app/services/analytics/functions/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
"""Analytics helper functions."""

from .periods import (
MAX_PERIOD_DAYS,
Interval,
Period,
build_period,
percent_change,
resolve_timezone,
)

__all__ = [
"MAX_PERIOD_DAYS",
"Interval",
"Period",
"build_period",
"percent_change",
"resolve_timezone",
]
149 changes: 149 additions & 0 deletions src/app/services/analytics/functions/periods.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,149 @@
"""
Reporting Period Helpers

Turns the ``from``/``to`` query parameters into a validated window, and derives
the immediately preceding window of equal length so every figure can be shown
against a comparable baseline.

Buckets are cut in the shop's timezone rather than UTC. A German shop closing
at 23:00 local would otherwise see that evening's orders land on the following
day for half the year, which makes daily revenue look wrong in a way that is
hard to spot and easy to disbelieve.
"""

from __future__ import annotations

from dataclasses import dataclass
from datetime import date, datetime, time, timedelta
from enum import Enum
from zoneinfo import ZoneInfo, ZoneInfoNotFoundError

from app.shared.exceptions import ValidationError

# A window longer than this is refused. The aggregates are computed live rather
# than from a rollup table, so an unbounded range is a slow query waiting to
# happen; five years is far beyond any dashboard use and still a clear ceiling.
MAX_PERIOD_DAYS = 366 * 5


class Interval(str, Enum):
"""Bucket width for time series."""

DAY = "day"
WEEK = "week"
MONTH = "month"


@dataclass(frozen=True)
class Period:
"""
A half-open reporting window ``[start, end)`` in UTC.

Half-open matters: an order placed at 23:59:59.999 on the last day belongs
to the period, and one placed at 00:00:00 the next day does not. A closed
range on dates either drops the final day or double-counts a boundary
order when two periods are compared.
"""

start: datetime
end: datetime
timezone: str

@property
def days(self) -> int:
return max((self.end - self.start).days, 1)

def previous(self) -> "Period":
"""The window of equal length ending where this one starts."""
length = self.end - self.start
return Period(
start=self.start - length,
end=self.start,
timezone=self.timezone,
)


def resolve_timezone(name: str) -> ZoneInfo:
"""
Look up a timezone, failing with a useful message rather than a 500.

A misconfigured SHOP_TIMEZONE should say so plainly — the alternative is
every analytics endpoint returning an opaque internal error.
"""
try:
return ZoneInfo(name)
except (ZoneInfoNotFoundError, ValueError) as exc:
raise ValidationError(
message=f"Unknown shop timezone {name!r}",
context={"timezone": name},
original_exception=exc,
) from exc


def build_period(
date_from: date | None,
date_to: date | None,
timezone_name: str,
default_days: int = 30,
) -> Period:
"""
Build a reporting window from inclusive calendar dates.

Both bounds are interpreted in the shop's timezone and converted to UTC,
because that is what the database stores. ``date_to`` is inclusive to the
reader — asking for the 1st to the 31st should include the 31st — so the
exclusive end is midnight at the start of the following day.

Args:
date_from: First day to include. Defaults to ``default_days`` before
``date_to``.
date_to: Last day to include. Defaults to today in the shop timezone.
timezone_name: IANA timezone the operator runs the shop in.
default_days: Window length used when ``date_from`` is omitted.

Returns:
The resolved window, in UTC.

Raises:
ValidationError: If the timezone is unknown, the range is inverted, or
the range exceeds MAX_PERIOD_DAYS.
"""
tz = resolve_timezone(timezone_name)

last_day = date_to or datetime.now(tz).date()
first_day = date_from or (last_day - timedelta(days=default_days - 1))

if first_day > last_day:
raise ValidationError(
message="The start of the period is after its end",
context={"from": first_day.isoformat(), "to": last_day.isoformat()},
)

span = (last_day - first_day).days + 1
if span > MAX_PERIOD_DAYS:
raise ValidationError(
message=(
f"Requested period spans {span} days, "
f"more than the {MAX_PERIOD_DAYS} day maximum"
),
context={"days": span, "maximum": MAX_PERIOD_DAYS},
)

start = datetime.combine(first_day, time.min, tzinfo=tz)
# Exclusive end: midnight opening the day after the last requested day.
end = datetime.combine(last_day + timedelta(days=1), time.min, tzinfo=tz)

return Period(start=start, end=end, timezone=timezone_name)


def percent_change(current: int | float, previous: int | float) -> float | None:
"""
Percentage change from ``previous`` to ``current``.

Returns None when there is no baseline to compare against. Growth from zero
is not "infinite percent" or "100%" — it is undefined, and reporting a
number there invites a reader to trust something meaningless.
"""
if previous == 0:
return None
return round(((current - previous) / previous) * 100, 1)
33 changes: 33 additions & 0 deletions src/app/services/analytics/models/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
"""Analytics schemas."""

from .analytics_models import (
AnalyticsFunnelResponse,
AnalyticsProductsResponse,
AnalyticsSummaryResponse,
AnalyticsTimeseriesResponse,
Change,
CurrencySeries,
CurrencyTotals,
CurrencyTotalsPrevious,
FunnelStep,
NeverSoldItem,
PeriodInfo,
ProductPerformance,
SeriesPoint,
)

__all__ = [
"AnalyticsFunnelResponse",
"AnalyticsProductsResponse",
"AnalyticsSummaryResponse",
"AnalyticsTimeseriesResponse",
"Change",
"CurrencySeries",
"CurrencyTotals",
"CurrencyTotalsPrevious",
"FunnelStep",
"NeverSoldItem",
"PeriodInfo",
"ProductPerformance",
"SeriesPoint",
]
Loading
Loading