Evidence-first benchmarking for proxy and web-retrieval workloads.
Measure usable results, not proxy count.
proxybench is a local, deterministic Python utility for comparing sanitized retrieval outcomes using metrics that matter to operators: usable success rate, requests per usable result, rotations per usable result, latency, and cost per usable result.
Status: alpha · Python 3.10+ · zero runtime dependencies · local-only · no telemetry
Built by PN Labs.
| Project | Purpose |
|---|---|
| proxy-outcome | Classify what happened without over-attributing the cause |
| proxybench | Compare usable-result efficiency between retrieval policies |
The tools are intentionally separate: classify evidence first, then benchmark whether a policy actually improves outcomes.
Adding more proxies or rotating more often does not guarantee better retrieval outcomes. A benchmark should answer a narrower question:
baseline policy
vs
candidate policy
↓
usable success rate
requests / usable result
rotations / usable result
latency distribution
cost / usable result
proxybench deliberately does not perform crawling, make network requests, accept proxy credentials, or select providers. It measures evidence you already collected from an authorized workload.
proxy-outcome answers:
What observation do we actually have, and how strong is the proxy-layer evidence?
proxybench answers:
Did policy A or policy B produce better usable outcomes and efficiency?
Classification and benchmarking stay separate.
python -m pip install .Runtime dependencies: none.
Input is JSON Lines (.jsonl). Every line is one sanitized retrieval event.
Allowed fields only:
{"usable": true, "latency_ms": 420, "cost_units": 0.0021, "rotated": false, "outcome": "SUCCESS"}usable— required boolean. Whether the result was usable for the workload.latency_ms— optional non-negative number.cost_units— optional non-negative number in any consistent cost unit.rotated— optional boolean.outcome— optional uppercase categorical token such asHTTP_RATE_LIMIT.
Unknown fields are rejected. URLs, IP addresses, proxy identifiers, provider credentials, cookies, headers, payloads, customer identifiers, and other production context are neither required nor part of the schema.
proxybench summarize examples/baseline.jsonlOutput includes:
- request count;
- usable result count;
- usable success rate + descriptive Wilson interval;
- requests per usable result;
- rotation coverage and rotations per usable result;
- latency coverage, p50, and p95;
- cost coverage and cost per usable result;
- categorical outcome counts.
proxybench compare examples/baseline.jsonl examples/candidate.jsonlThe comparison reports directional deltas without pretending that request events are necessarily independent or causal.
Optional operator-defined gates:
proxybench compare examples/baseline.jsonl examples/candidate.jsonl \
--min-requests 20 \
--min-success-uplift-pp 2 \
--max-rpu-regression-pct 5 \
--max-cost-regression-pct 5Gate verdicts are PASS, FAIL, INCONCLUSIVE, or NO_GATES_CONFIGURED.
- Usable-result first — HTTP success alone is not the buyer KPI.
- Data minimization — no URLs, IPs, credentials, raw headers, or production payloads are required.
- Local only — no network I/O or telemetry.
- Evidence before claims — descriptive intervals are not presented as causal proof.
- Explicit coverage — cost/latency/rotation metrics report how much of the input actually contained that field.
- Zero runtime dependencies — easy to embed in CI and benchmark harnesses.
This repository contains the transparent measurement baseline only. It does not contain PN Labs' private routing/scoring implementation, provider selection logic, target × egress learning, promotion/demotion intelligence, private benchmark datasets, production infrastructure, or cost-optimization control plane.
python -m pip install .
python -m unittest discover -s tests -v
python scripts/public_hygiene_check.pyCI installs the package before testing, smoke-tests the installed CLI, and runs a non-echoing vendor-neutral public repository hygiene gate.
Small, measurable improvements are welcome. See CONTRIBUTING.md before opening a pull request.
See SECURITY.md. Do not put production logs, credentials, private endpoint information, personal/customer data, or private infrastructure details into public benchmark fixtures or issues.
Use proxybench only with workloads you are authorized to run. Respect applicable target policies, rate limits, robots directives where relevant, and law.
MIT.