Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ retained diagnostics.

## Sequence

This is the second of four runs published on 2026-09-23:
This is the second of the runs published on 2026-09-23:

1. [Harness capacity](../2026-09-23-north-star-steady-v1-harness-capacity-33c95bf/analysis.md):
`kq-bench` saturated the host near 13,000 tasks/s, mostly on its own
Expand All @@ -22,6 +22,9 @@ This is the second of four runs published on 2026-09-23:
4. [Share group with 1000 members](../2026-09-23-adhoc-share-group-1000-members-33c95bf/analysis.md):
revisit this run's 20,000 RPS concurrency series with the member cap
raised from 200 to 1000.
5. [Shard-10 ramp](../2026-09-23-adhoc-shard-10-ramp-33c95bf/analysis.md):
ramp arrival with 1000 members x concurrency 10; 20,000 RPS is the highest
rate near the Redis reference, limited by slot capacity.

## Purpose

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ retained diagnostics.

## Sequence

This is the third of four runs published on 2026-09-23:
This is the third of the runs published on 2026-09-23:

1. [Harness capacity](../2026-09-23-north-star-steady-v1-harness-capacity-33c95bf/analysis.md):
`kq-bench` saturated the host near 13,000 tasks/s.
Expand All @@ -20,6 +20,9 @@ This is the third of four runs published on 2026-09-23:
4. [Share group with 1000 members](../2026-09-23-adhoc-share-group-1000-members-33c95bf/analysis.md):
follows suggestion 2 below by raising the member cap to 1000 and running
1000 members x concurrency 10 at 20,000 RPS.
5. [Shard-10 ramp](../2026-09-23-adhoc-shard-10-ramp-33c95bf/analysis.md):
ramp arrival with 1000 members x concurrency 10; 20,000 RPS is the highest
rate near the Redis reference, limited by slot capacity.

## Purpose

Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,107 @@
# Run Analysis

Run: [run.json](run.json)

Results: [result.json](result.json)

Kafka: [compose.yaml](../2026-09-23-adhoc-share-group-1000-members-33c95bf/compose.yaml)
(`group.share.max.size=1000`, `group.share.partition.max.record.locks=4000`)

This document is generated by AI from the run metadata, machine results, and
retained diagnostics.

## Sequence

This is the fifth of the runs published on 2026-09-23:

1. [Harness capacity](../2026-09-23-north-star-steady-v1-harness-capacity-33c95bf/analysis.md):
`kq-bench` saturated the host near 13,000 tasks/s.
2. [Lean-probe capacity](../2026-09-23-adhoc-lean-probe-capacity-33c95bf/analysis.md):
kq reached 150,000 tasks/s, but queue p50 stayed near 465 ms.
3. [Queue-latency diagnosis](../2026-09-23-adhoc-queue-latency-diagnosis-33c95bf/analysis.md):
queue time comes from the worker generation wait.
4. [Share group with 1000 members](../2026-09-23-adhoc-share-group-1000-members-33c95bf/analysis.md):
1000 members x concurrency 10 sustained 20,000 RPS with low queue time at
8 to 32 partitions and collapsed at 2 and 4.
5. **This run.** Hold the run 4 topology fixed and ramp arrival at 8, 16, 32,
and 64 partitions to find the highest rate that keeps queue time near the
Redis reference.

## Purpose

Find the highest arrival rate at which 1000 share-group members x concurrency
10 keep north-star queue time near the Redis reference (36 / 77 / 100 ms
p50 / p95 / p99), ramping arrival at 8, 16, 32, and 64 ready partitions with
the record-lock window fixed at 4000.

A point meets the target when every task completes exactly once, enqueue
reaches at least 99% of the requested rate, completions end within 3 s of
production, and queue p50 / p95 / p99 are at most 50 / 150 / 300 ms.

## Observations

| Partitions | Requested RPS | Median completions/s | Tail after production | Queue p50 / p95 / p99 / max | Outcome |
| ---: | ---: | ---: | ---: | ---: | --- |
| 8 | 20,000 | 20,013 | 1 s | 14 / 48 / 99 / 589 ms | meets target |
| 8 | 21,000 | 21,042 | 1 s | 50 / 160 / 243 / 633 ms | sustained; p95 over limit |
| 8 | 22,000 | 20,911 | 5 s | 1,431 / 3,205 / 3,813 / 3,978 ms | overloaded |
| 16 | 20,000 | 20,010 | 1 s | 24 / 70 / 101 / 837 ms | meets target |
| 16 | 21,000 | 20,938 | 4 s | 60 / 597 / 2,815 / 3,485 ms | overloaded |
| 32 | 20,000 | 19,996 | 1 s | 42 / 133 / 207 / 544 ms | meets target |
| 32 | 21,000 | 20,884 | 14 s | 94 / 465 / 11,391 / 13,639 ms | overloaded |
| 64 | 18,000 | 18,003 | 1 s | 39 / 125 / 203 / 724 ms | meets target |
| 64 | 19,000 | 18,992 | 2 s | 51 / 171 / 421 / 1,435 ms | sustained; over limit |
| 64 | 20,000 | 20,005 | 1 s | 73 / 225 / 361 / 903 ms | sustained; over limit |

- Highest rate meeting the target: 20,000 RPS at 8, 16, and 32 partitions;
18,000 RPS at 64 partitions.
- Every overloaded point plateaued at 20,884 to 20,938 median completions/s
regardless of partition count.
- At a fixed 20,000 RPS, queue p50 rose monotonically with partition count:
14, 24, 42, and 73 ms at 8, 16, 32, and 64 partitions.
- All 12,120,000 tasks across all points completed with zero duplicates and
zero missing IDs.
- Kafka used 5.7 to 6.3 cores at sustained points and 7.9 cores at the
8-partition 22,000 RPS overload. Workers used about 1.6 cores and producers
about 1.4 cores. The host was not CPU-saturated.
- The 8-partition 20,000 RPS point (14 / 48 / 99 ms) was below the Redis
reference at p50 and p95 and matched it at p99. The same configuration in
run 4 measured 14 / 104 / 203 ms, so p95 and p99 varied between otherwise
identical runs.

## Interpretation

- The ceiling is worker slot capacity, not Kafka partitions or host CPU.
About 20,900 completions/s across 10,000 slots is about 2.1 tasks/s per
slot, or about 480 ms per generation of 10 against a 250 ms average task.
Roughly half of each slot's time is spent waiting for the slowest task in
its generation.
- Because `group.share.max.size` cannot exceed 1000, slots can only grow by
raising per-member concurrency, which lengthens each generation.
- More partitions than needed increase queue time. Hypothesis: each member's
small fetches are spread across more partitions and fetch sessions. 8 to 16
partitions is the best range for this topology.
- Overload raises Kafka CPU, consistent with the collapse feedback
hypothesized in run 4, though no point here collapsed.

## Suggestions

1. Measure concurrency 20 with the same ramp to see how much capacity larger
generations add and at what latency cost.
2. Prioritize a worker that refills slots as tasks complete. At about 4
tasks/s per slot, the same 10,000 slots could plausibly support 35,000 to
40,000 tasks/s; this estimate is untested.
3. Repeat the 8- and 16-partition 20,000 RPS points several times to bound
run-to-run variation in p95 and p99.

## Caveats

- Single Dockerized Kafka broker on localhost, RF=1; one run per point with a
60 s arrival window.
- Same probe as runs 2 to 4: ready success path only, sleep-only handlers,
200-byte payloads, queue time measured from just before `Enqueue`.
- The ramp stopped at the first failing point per partition count, so some
partition counts have only two points.
- The acceptance limits are an engineering judgment of "close to" the Redis
reference, not a production SLO. The Redis figures were not measured with
this workload or harness.
Loading
Loading