Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
105 changes: 105 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,105 @@
name: Benchmark CI

on:
pull_request:
workflow_dispatch:

permissions:
contents: read

concurrency:
group: benchmark-ci-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true

env:
ARGON_ENGINE_REF: ba6e06a9a3b31124d6c37475b5667dd70ab42379
ARGON_SUITE_REF: ${{ github.event.pull_request.head.sha || github.sha }}

jobs:
workflow-smoke:
name: Unit tests and replica-set workflow smoke
runs-on: ubuntu-latest
timeout-minutes: 15
defaults:
run:
shell: bash
working-directory: suite
steps:
- name: Check out exact suite revision
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: ${{ env.ARGON_SUITE_REF }}
path: suite
fetch-depth: 0
persist-credentials: false

# Keep the engine outside the suite: run.py inventories suite source
# files, so a nested engine checkout would pollute that inventory.
- name: Check out fixed companion engine
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
repository: argon-lab/argon
ref: ${{ env.ARGON_ENGINE_REF }}
path: engine
fetch-depth: 1
persist-credentials: false

- name: Configure isolated Compose fixture
run: |
echo "ARGON_ENGINE_SOURCE=$GITHUB_WORKSPACE/engine" >> "$GITHUB_ENV"
mkdir -p results

- name: Validate and build the Go 1.26.6 container recipe
run: |
docker compose config --quiet
docker compose build bench

- name: Start MongoDB replica set
run: docker compose up --detach --wait --wait-timeout 120 mongo

# run.py runs all Go unit tests before building the executable, then
# freezes both sources and records the exact engine ref and Go version.
- name: Run unit tests and one-cell correctness smoke
run: |
docker compose run --rm --no-deps bench \
--ref "$ARGON_ENGINE_REF" --results /suite/results/ci-smoke -- \
-sizes 100 -concurrency 1 -depths 1 \
-metadata-samples 2 -workflow-samples 2 -read-samples 2 \
-divergence-docs 3 -divergence-rounds 2 -timeout 3m \
2>&1 | tee results/smoke.log

- name: Verify recorded source and toolchain
run: |
python3 - <<'PY'
import json, os
from pathlib import Path
report = json.loads(Path('results/ci-smoke/raw.json').read_text())
assert report['complete']
assert report['provenance']['engine']['git_head'] == os.environ['ARGON_ENGINE_REF']
assert report['provenance']['suite']['git_head'] == os.environ['ARGON_SUITE_REF']
assert not report['provenance']['engine']['dirty']
assert not report['provenance']['suite']['dirty']
assert report['environment']['go'] == 'go1.26.6'
assert len(report['scenarios']) == 1
scenario = report['scenarios'][0]
assert (scenario['documents'], scenario['concurrency'], scenario['ancestry_depth']) == (100, 1, 1)
assert scenario['storage']['captured_records'] == 6
print('Exact source/toolchain provenance and six captured updates verified.')
print('This correctness smoke has no latency/throughput SLA assertions.')
PY

- name: Collect fixture logs and remove containers
if: always()
run: |
mkdir -p results
docker compose logs --no-color > results/compose.log 2>&1 || true
docker compose down --volumes --remove-orphans

- name: Retain raw smoke samples, provenance, source archives and logs
if: always()
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
with:
name: benchmark-smoke-${{ github.run_id }}-${{ github.run_attempt }}
path: suite/results/
if-no-files-found: warn
retention-days: 14
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
@@ -1,3 +1,6 @@
results/
benchmarks
argonbench

__pycache__/
*.pyc
15 changes: 6 additions & 9 deletions Dockerfile
Original file line number Diff line number Diff line change
@@ -1,10 +1,7 @@
FROM golang:1.24

WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
FROM golang:1.26.6-trixie
RUN apt-get update && apt-get install -y --no-install-recommends python3 git ca-certificates && rm -rf /var/lib/apt/lists/*
WORKDIR /suite
COPY . .
RUN CGO_ENABLED=0 go build -o /usr/local/bin/argonbench .

ENTRYPOINT ["argonbench"]
CMD ["-out", "/out/report.md"]
# The explicit engine mount is frozen and replaced at run time; no historical
# module version is silently presented as the code under measurement.
ENTRYPOINT ["python3", "scripts/run.py", "--engine", "/engine"]
147 changes: 66 additions & 81 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,87 +1,72 @@
# argon-benchmarks
# Argon workflow benchmarks

Reproducible benchmarks for [Argon](https://github.com/argon-lab/argon), the
Git-like branching and time-travel engine for MongoDB.
This suite measures an explicit Argon engine checkout. It distinguishes a metadata-only fork from a sandbox with a physical MongoDB copy and capture ready, then measures the first native query, captured-write visibility and storage after divergence. Every successful run writes all timing samples, nearest-rank p50/p95/p99, configuration, environment and source provenance.

**The contract:** every performance number Argon publishes — on
[argonlabs.tech](https://www.argonlabs.tech), in the README, anywhere — must
come from a run of this suite that you can reproduce yourself. No number
without a runnable source. Numbers that predate this suite were removed from
all Argon materials in July 2026.
The workflow runner requires the engine's `StartCapture`, `SyncBranch` and `WaitAuto` APIs. The dependency in `go.mod` is the historical published baseline, **not** the code measured by this runner. Until these APIs are released, an explicit checkout containing the hardening changes is required. `scripts/run.py` uses a temporary module replacement, freezes the exact engine source, and records its Git ref, dirty status, tracked diff SHA256, complete source manifest/hash and build module information. It also archives the frozen engine and executable suite source, including untracked files; unpublished code is labelled accordingly. The suite's executable source hash identifies runner changes before a suite commit exists.

## Run it
## Run locally

```bash
git clone https://github.com/argon-lab/benchmarks
cd benchmarks
docker compose up --build --abort-on-container-exit
Requires Go 1.26.6+, Python 3.12+, Git and MongoDB 7+ configured as a replica set. Use an isolated database deployment. The suite creates unique `argonbench_meta_*` and `argonbench_src_*` databases plus physical sandbox databases, and removes its databases after each scenario. It never uses the normal `argon_wal` metadata database. Abrupt process termination can leave these fixture databases behind.

```sh
export MONGODB_URI='mongodb://localhost:27017/?replicaSet=rs0'
export GOCACHE=/tmp/argon-benchmark-go-cache
export GOTOOLCHAIN=go1.26.6
python3 scripts/run.py --engine /absolute/path/to/argon -- \
-sizes 1000,10000 -concurrency 1,4 -depths 1,4 \
-metadata-samples 100 -workflow-samples 20 -read-samples 20
```

Each run gets a fresh `results/<UTC timestamp>/` directory. Set `--results /new/output/directory` before `--` to choose it. Existing result directories are rejected. Add `--ref <exact-engine-commit>` before `--` to measure that committed tree even if the source checkout later changes. Both engine and suite source are frozen before building; later edits in the original checkouts cannot affect the measured binary. The wrapper rejects edits detected during the copy.

To run both the suite and MongoDB in containers (an explicitly selected engine is still required):

```sh
export ARGON_ENGINE_SOURCE=/absolute/path/to/argon
mkdir -p results
docker compose up --build --abort-on-container-exit --exit-code-from bench
# Removes only this Compose project's test deployment.
docker compose down
```

The report prints to stdout and lands in `./results/report.md`, including the
exact engine commit, MongoDB version, and hardware context of the run.

Tunables (edit `docker-compose.yml` command or run the binary directly):

| flag | default | meaning |
|---|---|---|
| `-docs` | 50000 | documents seeded and imported (history size) |
| `-iters` | 200 | iterations for branch-create latency |
| `-branches` | 200 | branches created for the storage suite |
| `-depths` | 1000,10000,50000 | history depths (LSNs) for time-travel reads |

## What is measured, and how

The suite seeds a plain MongoDB database, imports it through
`walcli.ImportDatabase` (the same code path the `argon` CLI drives, which
creates the project and its WAL history), then measures through
`pkg/walcli` — the same Go services the CLI itself uses. The engine
version is pinned in `go.mod` and embedded in every report automatically.

1. **Branch creation latency** — p50/p95/p99 over `-iters` creations on a
project that already has real history. Argon's claim is architectural:
a branch is one metadata document, so latency must not depend on data
size.
2. **Time-travel materialization** — latency of reconstructing collection
state at increasing history depths, with the engine's shipped defaults
(automatic snapshots on). The report also prints how many auto-snapshots
existed after import, so the effect is attributable.
3. **Snapshot at head** — the same read before an explicit snapshot, the
snapshot's creation cost, and the read after it: the bounded-replay
effect in isolation.
4. **Materialization throughput** — documents/second derived from the head
reads in (3).
5. **Storage cost per branch** — `dbStats` delta (data+index bytes) across
`-branches` creations: the metadata-only claim, measured.
6. **Bulk import throughput** — wall time of `walcli.ImportDatabase` for the
whole seeded dataset. This is the bulk-ingest path, not a per-operation
write microbenchmark (see below).

## What is deliberately NOT measured (yet)

- **Per-operation write throughput** through the driver interceptor, and
**storage amplification under branch divergence** — both need write access
from an external module, which the engine does not currently export
(`internal/driver` is not reachable through `pkg/walcli`). Tracked in
[argon-lab/argon#16](https://github.com/argon-lab/argon/issues/16); the
suites land here as soon as the surface exists.
- Anything involving a wire-protocol proxy or per-branch connection strings
(that's M3).

## Methodology notes

- The engine version is pinned in `go.mod` and read from build info at
runtime, so it appears in every report. A result without a pinned ref is
not a result.
- Runs happen inside Docker on whatever machine you have; absolute numbers
vary with hardware. The report always embeds the environment. Compare
shapes and ratios (e.g. flat branch-create latency vs history size, the
before/after-snapshot delta), not absolute values across machines.
- One warm MongoDB instance per run, fresh and empty. No connection reuse
tricks, no cache pre-warming beyond what a real user gets.
- The suite exits non-zero on any error: a partial run never yields a report.

## Publishing rules for maintainers

Official numbers quoted by Argon come from runs recorded in
[RESULTS.md](RESULTS.md), each entry with the machine, engine ref, and the
full report. Update the website/README only by linking one of those entries.
The Compose configuration pins MongoDB 7.0.14 and configures replica set `rs0`. Host-local and container measurements are separate environments and must not be compared as equivalent runs. The current local report was run on the host; the container recipe has not yet been exercised on this machine.

## Measurements and definitions

| Metric | Boundary |
|---|---|
| `metadata_fork_ms` | `CreateBranch` against inherited history, no physical checkout |
| `sandbox_capture_ready_ms` | Fork + full checkout + capture startup readiness acknowledgement |
| `first_native_query_ms` | New Mongo client/handshake + verified indexed `FindOne`, after readiness |
| `native_write_ack_ms` | One native `$inc` with majority write concern |
| `capture_ack_to_observed_ms` | Native acknowledgement → polling observes the committed WAL event; includes polling/query overhead |
| `native_write_to_observed_ms` | Native update start → that same visibility observation |
| `head_materialize_as_shipped_ms` | Full inherited collection read, before this suite creates an explicit snapshot |
| `head_materialize_after_snapshot_ms` | Same read after an explicit snapshot at imported main |
| `historical_25pct_ms`, `historical_50pct_ms` | Full collection replay at 25%/50% of seeded document count as a target LSN; control records also consume LSNs |
| `divergence_bulk_and_capture_ms` | Repeated native bulk updates + capture barrier, one sample per concurrent branch |

The imported data has deterministic string IDs and a 128-byte payload. `-depths` counts ancestry edges from imported main to the workload parent; it is distinct from historical replay LSN. Sandbox workflows and divergence run after the imported-main snapshot phase, so their inherited source has an explicit snapshot. Every first query must return the known document; captured point materialization must contain the new value. Divergence must capture exactly `workers × min(documents, divergence-docs) × divergence-rounds` WAL records; any mismatch fails the run.

Metadata storage observations use MongoDB `dbStats` and keep logical data, allocated collection bytes and allocated index bytes separate. Each checked-out database is measured too: a sandbox consumes a full physical copy in this architecture. Divergence records raw stored WAL BSON bytes and changed current-document BSON bytes, followed by explicit divergent snapshots. That ratio counts repeated updates in the numerator and each changed current document once in the denominator; it excludes snapshots, indexes and physical copies. It is not disk amplification. WiredTiger allocation deltas are quantized, can include retained freed pages and are not universal per-branch prices. Storage measurements currently require the MongoDB chunk backend.

All samples are retained, including first executions. There is no discarded warmup or coordinated-omission correction. A fixed worker count submits the next operation after the previous finishes (closed loop); scenarios and phases run sequentially. Concurrency is within a phase, not an externally sustained arrival rate. Small local sample counts describe that run; p99 with fewer than 100 samples is generally the maximum and does not establish production tail latency. Absolute rates depend on journal, write concern, CPU, RAM, storage and topology.

## Expanded scale/concurrency matrix

```sh
export ARGON_ENGINE_SOURCE=/absolute/path/to/argon
./scripts/expanded.sh
```

This runs 1k/50k/1m documents × 1/4/16 workers × 1/4/16 ancestry depth, with 1,000 fork, 200 workflow and 100 read samples per phase, 1,000 changed documents and 10 divergence rounds. It is resource intensive and has a 24-hour deadline. It is a runnable experiment plan, **not a claim that this matrix has been measured**. Override flags at the end to scope a run. For a quick smoke test use one size, one worker and two samples.

## Pull-request CI

The PR workflow checks out companion engine commit `ba6e06a9a3b31124d6c37475b5667dd70ab42379` beside the suite, builds the Go 1.26.6 Docker recipe, and starts the Compose MongoDB replica set. It runs the runner's Go unit tests followed by one 100-document / one-worker / one-level smoke cell, with two workflow samples and six required captured divergence updates. `--ref` is explicit; CI verifies the engine/suite refs and actual Go version in the generated provenance. Raw samples, reports, both source archives and logs are retained as a workflow artifact for 14 days, including available diagnostics on failure.

This job checks correctness and reproducibility of the container workflow. Its tiny sample counts and shared CI runner are unsuitable for performance SLAs or comparisons with the recorded local matrix. The historical reports and their measured suite refs remain unchanged.

## Published results

See [RESULTS.md](RESULTS.md). Historical numbers retain their original date and exact scope. New results include raw samples and source hashes; no dirty working tree is identified as a released engine version. Benchmark failure exits nonzero and does not publish a complete report.
Loading
Loading