diff --git a/.github/ISSUE_TEMPLATE/bug-report.md b/.github/ISSUE_TEMPLATE/bug-report.md new file mode 100644 index 0000000..e52777b --- /dev/null +++ b/.github/ISSUE_TEMPLATE/bug-report.md @@ -0,0 +1,31 @@ +--- +name: Bug report +description: Report a reproducible benchmark or CLI issue using sanitized evidence +title: "[bug] " +labels: [] +assignees: [] +--- + +> Public issue: do not paste credentials, private URLs/IPs, provider account details, raw production logs, customer data, or infrastructure details. + +## Observed behavior + +Describe what `proxybench` returned or did. + +## Expected behavior + +Describe the deterministic behavior you expected. + +## Minimal synthetic reproducer + +Provide the smallest sanitized JSONL example that reproduces the issue. + +## Environment + +- Python version: +- `proxybench` version/commit: +- OS/runtime, if relevant: + +## Additional context + +Include only non-sensitive details needed to understand the bug. diff --git a/.github/pull_request_template.md b/.github/pull_request_template.md new file mode 100644 index 0000000..81097fa --- /dev/null +++ b/.github/pull_request_template.md @@ -0,0 +1,16 @@ +## Summary + +What changed and why? + +## Evidence + +What metric, test, or reproducible case supports the change? + +## Safety / scope check + +- [ ] No credentials, private endpoints/IPs, provider account details, customer data, or production captures are included. +- [ ] Missing-data behavior remains explicit. +- [ ] Descriptive comparisons are not presented as causal proof. +- [ ] Public/private implementation boundaries are preserved. +- [ ] Tests were added or updated when behavior changed. +- [ ] Local tests and repository hygiene checks pass. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md new file mode 100644 index 0000000..0bcc5a2 --- /dev/null +++ b/CONTRIBUTING.md @@ -0,0 +1,46 @@ +# Contributing + +Thanks for considering a contribution to `proxybench`. + +## Scope + +Good contributions improve the public measurement baseline without weakening its data-minimization or evidence-first behavior. + +Useful changes include: + +- additional deterministic summary metrics; +- clearer comparison semantics; +- explicit missing-data coverage; +- regression tests for edge cases; +- CLI usability improvements; +- packaging, CI, security, or documentation improvements. + +Private routing/scoring logic, provider selection, credentials, production infrastructure, private benchmark data, and customer/workload identifiers do not belong in this repository. + +## Development + +```bash +python -m pip install . +python -m unittest discover -s tests -v +python scripts/public_hygiene_check.py +``` + +Keep runtime dependencies at zero unless there is a strong, documented reason to change that constraint. + +## Pull requests + +A good pull request should: + +1. state the metric or operator problem precisely; +2. include tests for behavior changes; +3. keep input/output semantics deterministic; +4. report missing-data coverage rather than silently imputing values; +5. avoid presenting descriptive comparisons as causal proof; +6. keep fixtures synthetic and non-sensitive; +7. keep CI green across the supported Python matrix. + +## Sensitive information + +Do not post credentials, private endpoints/IPs, provider account details, production logs, customer data, or private infrastructure details. + +See [`SECURITY.md`](SECURITY.md) for the repository security boundary. diff --git a/README.md b/README.md index 0627f01..a0911b7 100644 --- a/README.md +++ b/README.md @@ -1,11 +1,24 @@ # proxybench +> Evidence-first benchmarking for proxy and web-retrieval workloads. + **Measure usable results, not proxy count.** -`proxybench` is a local, deterministic benchmark utility for proxy and web-retrieval workloads. It compares sanitized request outcomes using metrics that matter to operators: usable success rate, requests per usable result, rotations per usable result, latency, and cost per usable result. +`proxybench` is a local, deterministic Python utility for comparing sanitized retrieval outcomes using metrics that matter to operators: usable success rate, requests per usable result, rotations per usable result, latency, and cost per usable result. + +**Status:** alpha · Python 3.10+ · zero runtime dependencies · local-only · no telemetry Built by **PN Labs**. +## PN Labs reliability toolkit + +| Project | Purpose | +| --- | --- | +| [**proxy-outcome**](https://github.com/pnlabs-dev/proxy-outcome) | Classify what happened without over-attributing the cause | +| **proxybench** | Compare usable-result efficiency between retrieval policies | + +The tools are intentionally separate: classify evidence first, then benchmark whether a policy actually improves outcomes. + ## Why this exists Adding more proxies or rotating more often does not guarantee better retrieval outcomes. A benchmark should answer a narrower question: @@ -124,6 +137,10 @@ python scripts/public_hygiene_check.py CI installs the package before testing, smoke-tests the installed CLI, and runs a non-echoing vendor-neutral public repository hygiene gate. +## Contributing + +Small, measurable improvements are welcome. See [`CONTRIBUTING.md`](CONTRIBUTING.md) before opening a pull request. + ## Security See [`SECURITY.md`](SECURITY.md). Do not put production logs, credentials, private endpoint information, personal/customer data, or private infrastructure details into public benchmark fixtures or issues.