Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 28 additions & 0 deletions .github/ISSUE_TEMPLATE/benchmark-question.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
---
name: Benchmark question
description: Ask about a sanitized benchmark result or protocol
labels: []
assignees: []
---

> Public issue: do not paste credentials, private URLs/IPs, raw production logs, cookies, headers, customer data, or infrastructure details.

## Benchmark goal

Describe the operator metric you are trying to improve.

## Sanitized setup

Describe the baseline and candidate policy at a high level without naming private systems, endpoints, or accounts.

## Metrics

Paste only aggregate `proxybench` output or synthetic fixture data.

## Question

What interpretation or benchmark-design issue do you want help with?

## Authorization / responsible use

Confirm the workload is authorized and respects applicable target policies, rate limits, and law.
52 changes: 52 additions & 0 deletions .github/workflows/test.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
name: tests

on:
push:
pull_request:

permissions:
contents: read

jobs:
test:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
python-version: ["3.10", "3.11", "3.12"]

steps:
- name: Check out repository
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false

- name: Set up Python
uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6.2.0
with:
python-version: ${{ matrix.python-version }}

- name: Public repository hygiene check
run: python scripts/public_hygiene_check.py

- name: Install package from source
run: python -m pip install --disable-pip-version-check .

- name: Compile package
run: python -m compileall -q src

- name: Run tests against installed package
run: python -m unittest discover -s tests -v

- name: Installed CLI smoke tests
shell: bash
run: |
proxybench summarize examples/baseline.jsonl \
| python -c 'import json,sys; d=json.load(sys.stdin); assert d["requests"] == 20; assert d["usable_results"] == 13'

python -m proxybench compare examples/baseline.jsonl examples/candidate.jsonl \
--min-requests 20 \
--min-success-uplift-pp 10 \
--max-rpu-regression-pct 0 \
--max-cost-regression-pct 0 \
| python -c 'import json,sys; d=json.load(sys.stdin); assert d["gate"]["verdict"] == "PASS"'
35 changes: 35 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
# Python
__pycache__/
*.py[cod]
*.egg-info/
build/
dist/
.venv/
venv/
.coverage
.pytest_cache/
.mypy_cache/
.ruff_cache/

# Local secrets and credentials
.env
.env.*
!.env.example
*.pem
*.key
*.p12
*.pfx
credentials*
secrets*

# Potentially sensitive captures/data
*.har
*.pcap
*.pcapng
*.sqlite
*.sqlite3
*.db
*.log
private/
local-data/
raw-data/
21 changes: 21 additions & 0 deletions LICENSE
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
MIT License

Copyright (c) 2026 PN Labs

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
138 changes: 137 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
@@ -1 +1,137 @@
# proxybench
# proxybench

**Measure usable results, not proxy count.**

`proxybench` is a local, deterministic benchmark utility for proxy and web-retrieval workloads. It compares sanitized request outcomes using metrics that matter to operators: usable success rate, requests per usable result, rotations per usable result, latency, and cost per usable result.

Built by **PN Labs**.

## Why this exists

Adding more proxies or rotating more often does not guarantee better retrieval outcomes. A benchmark should answer a narrower question:

```text
baseline policy
vs
candidate policy
↓
usable success rate
requests / usable result
rotations / usable result
latency distribution
cost / usable result
```

`proxybench` deliberately does **not** perform crawling, make network requests, accept proxy credentials, or select providers. It measures evidence you already collected from an authorized workload.

## Relationship to proxy-outcome

[`proxy-outcome`](https://github.com/pnlabs-dev/proxy-outcome) answers:

> What observation do we actually have, and how strong is the proxy-layer evidence?

`proxybench` answers:

> Did policy A or policy B produce better usable outcomes and efficiency?

Classification and benchmarking stay separate.

## Install

```bash
python -m pip install .
```

Runtime dependencies: **none**.

## Input format

Input is JSON Lines (`.jsonl`). Every line is one sanitized retrieval event.

Allowed fields only:

```json
{"usable": true, "latency_ms": 420, "cost_units": 0.0021, "rotated": false, "outcome": "SUCCESS"}
```

- `usable` — required boolean. Whether the result was usable for the workload.
- `latency_ms` — optional non-negative number.
- `cost_units` — optional non-negative number in any consistent cost unit.
- `rotated` — optional boolean.
- `outcome` — optional uppercase categorical token such as `HTTP_RATE_LIMIT`.

Unknown fields are rejected. URLs, IP addresses, proxy identifiers, provider credentials, cookies, headers, payloads, customer identifiers, and other production context are neither required nor part of the schema.

## Summarize one arm

```bash
proxybench summarize examples/baseline.jsonl
```

Output includes:

- request count;
- usable result count;
- usable success rate + descriptive Wilson interval;
- requests per usable result;
- rotation coverage and rotations per usable result;
- latency coverage, p50, and p95;
- cost coverage and cost per usable result;
- categorical outcome counts.

## Compare A/B arms

```bash
proxybench compare examples/baseline.jsonl examples/candidate.jsonl
```

The comparison reports directional deltas without pretending that request events are necessarily independent or causal.

Optional operator-defined gates:

```bash
proxybench compare examples/baseline.jsonl examples/candidate.jsonl \
--min-requests 20 \
--min-success-uplift-pp 2 \
--max-rpu-regression-pct 5 \
--max-cost-regression-pct 5
```

Gate verdicts are `PASS`, `FAIL`, `INCONCLUSIVE`, or `NO_GATES_CONFIGURED`.

## Design principles

- **Usable-result first** — HTTP success alone is not the buyer KPI.
- **Data minimization** — no URLs, IPs, credentials, raw headers, or production payloads are required.
- **Local only** — no network I/O or telemetry.
- **Evidence before claims** — descriptive intervals are not presented as causal proof.
- **Explicit coverage** — cost/latency/rotation metrics report how much of the input actually contained that field.
- **Zero runtime dependencies** — easy to embed in CI and benchmark harnesses.

## Public / commercial boundary

This repository contains the transparent measurement baseline only. It does not contain PN Labs' private routing/scoring implementation, provider selection logic, target × egress learning, promotion/demotion intelligence, private benchmark datasets, production infrastructure, or cost-optimization control plane.

See [`docs/PUBLIC-BOUNDARY.md`](docs/PUBLIC-BOUNDARY.md).

## Validation

```bash
python -m pip install .
python -m unittest discover -s tests -v
python scripts/public_hygiene_check.py
```

CI installs the package before testing, smoke-tests the installed CLI, and runs a non-echoing vendor-neutral public repository hygiene gate.

## Security

See [`SECURITY.md`](SECURITY.md). Do not put production logs, credentials, private endpoint information, personal/customer data, or private infrastructure details into public benchmark fixtures or issues.

## Responsible use

Use `proxybench` only with workloads you are authorized to run. Respect applicable target policies, rate limits, robots directives where relevant, and law.

## License

MIT.
49 changes: 49 additions & 0 deletions SECURITY.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,49 @@
# Security policy

`proxybench` is intentionally local-only. It performs no network I/O, telemetry, crawling, provider login, proxy authentication, or credential storage.

## Public input boundary

The benchmark event schema is deliberately small. It accepts only:

- `usable`
- `latency_ms`
- `cost_units`
- `rotated`
- `outcome`

Unknown fields are rejected. This is a data-minimization boundary: target URLs, proxy endpoints, IP addresses, provider names, headers, cookies, credentials, payloads, customer identifiers, and infrastructure details are not required.

## Never publish sensitive material

Do not place any of the following in issues, pull requests, examples, fixtures, screenshots, benchmark artifacts, or CI logs:

- API keys, access tokens, passwords, or private keys;
- proxy credentials or credential-bearing proxy URLs;
- session cookies, authorization headers, or authenticated request dumps;
- private endpoint URLs, IP addresses, internal hostnames, or infrastructure topology;
- production HAR/PCAP captures or raw production logs;
- customer data, personal data, or private datasets;
- non-public commercial implementation details.

Use synthetic fixtures and aggregate counts instead.

## Error behavior

Input-validation errors do not intentionally echo raw JSON lines, unsupported values, or local input paths. This reduces the chance that CI output becomes a secondary disclosure path.

## Statistical safety

The Wilson interval reported for usable success rate is descriptive. `proxybench` does not claim that request events are independent, randomized, or causal. Operator-defined gates are policy checks, not scientific proof.

## Repository hygiene gate

CI runs `scripts/public_hygiene_check.py`, which searches public text files for common secret/token shapes, credential-bearing URLs, IPv4/IPv6 literals, internal-hostname shapes, sensitive filenames, and email addresses.

Detected values are never printed; only the rule class and file path are reported.

This is defense-in-depth, not a replacement for review or dedicated secret scanning.

## Reporting security issues

Do not open a public issue containing sensitive reproduction material. Reduce the problem to a synthetic reproducer or sanitized aggregate description before sharing.
Loading