diff --git a/.github/ISSUE_TEMPLATE/bug-report.md b/.github/ISSUE_TEMPLATE/bug-report.md new file mode 100644 index 0000000..c37445b --- /dev/null +++ b/.github/ISSUE_TEMPLATE/bug-report.md @@ -0,0 +1,31 @@ +--- +name: Bug report +description: Report a reproducible behavior issue using sanitized evidence +title: "[bug] " +labels: [] +assignees: [] +--- + +> Public issue: do not paste credentials, cookies, private URLs/IPs, internal hostnames, raw production logs, HAR/PCAP captures, customer data, or infrastructure details. + +## Observed behavior + +Describe what the library returned or did. + +## Expected behavior + +Describe the evidence-preserving behavior you expected. + +## Minimal synthetic reproducer + +Provide the smallest sanitized example that reproduces the issue. + +## Environment + +- Python version: +- `proxy-outcome` version/commit: +- OS/runtime, if relevant: + +## Additional context + +Include only non-sensitive details needed to understand the bug. diff --git a/.github/pull_request_template.md b/.github/pull_request_template.md new file mode 100644 index 0000000..b8f1eb0 --- /dev/null +++ b/.github/pull_request_template.md @@ -0,0 +1,15 @@ +## Summary + +What changed and why? + +## Evidence + +What observation, test, or reproducible case supports the change? + +## Safety / scope check + +- [ ] No credentials, private endpoints/IPs, customer data, or production captures are included. +- [ ] The change does not invent root cause when evidence is insufficient. +- [ ] Public/private implementation boundaries are preserved. +- [ ] Tests were added or updated when behavior changed. +- [ ] Local tests and repository hygiene checks pass. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md new file mode 100644 index 0000000..e68e197 --- /dev/null +++ b/CONTRIBUTING.md @@ -0,0 +1,45 @@ +# Contributing + +Thanks for considering a contribution to `proxy-outcome`. + +## Scope + +Good contributions improve the public deterministic baseline without weakening its core rule: preserve observations without inventing root cause. + +Useful changes include: + +- clearer taxonomy or documentation; +- additional conservative classification cases; +- regression tests for ambiguous outcomes; +- CLI usability improvements; +- packaging or CI improvements; +- security and data-minimization hardening. + +Private routing/scoring logic, provider-specific private configuration, credentials, production infrastructure, and non-public benchmark data do not belong in this repository. + +## Development + +```bash +python -m pip install . +python -m unittest discover -s tests -v +python scripts/public_hygiene_check.py +``` + +Keep runtime dependencies at zero unless there is a strong, documented reason to change that constraint. + +## Pull requests + +A good pull request should: + +1. explain the observed problem before proposing attribution; +2. include tests for behavior changes; +3. avoid changing unrelated code; +4. preserve fail-closed behavior when evidence is insufficient; +5. keep examples synthetic and non-sensitive; +6. keep CI green across the supported Python matrix. + +## Sensitive information + +Do not post credentials, cookies, raw authorization headers, private URLs or IPs, internal hostnames, production logs, HAR/PCAP captures, customer data, or private infrastructure details. + +See [`SECURITY.md`](SECURITY.md) for the repository security boundary. diff --git a/README.md b/README.md index 79560eb..97eacdf 100644 --- a/README.md +++ b/README.md @@ -1,13 +1,24 @@ # proxy-outcome +> Evidence-first classification for proxy and web-retrieval outcomes. + **403 ≠ bad proxy. 429 ≠ bad proxy. 451 ≠ bad proxy.** -`proxy-outcome` is a small, deterministic classification library for HTTP and proxy-path observations in scraping and web-retrieval systems. +`proxy-outcome` is a small, deterministic Python library that classifies HTTP and proxy-path observations without inventing root cause. It is designed for scraper, crawler, browser-automation, and web-retrieval systems where a failed request should not automatically poison a proxy pool. -Its job is deliberately narrow: preserve evidence without inventing a root cause. +**Status:** alpha · Python 3.10+ · zero runtime dependencies · local-only · no telemetry Built by **PN Labs**. +## PN Labs reliability toolkit + +| Project | Purpose | +| --- | --- | +| **proxy-outcome** | Classify what happened without over-attributing the cause | +| [**proxybench**](https://github.com/pnlabs-dev/proxybench) | Compare usable-result efficiency between retrieval policies | + +The tools are intentionally separate: classify evidence first, then benchmark whether a policy actually improves outcomes. + ## Why this exists A common reliability mistake is: @@ -158,6 +169,10 @@ CI installs the package before testing, runs CLI smoke tests, and executes a non See [`docs/ARCHITECTURE.md`](docs/ARCHITECTURE.md). +## Contributing + +Small, evidence-preserving improvements are welcome. See [`CONTRIBUTING.md`](CONTRIBUTING.md) before opening a pull request. + ## Design partners If you operate a legitimate scraping/web-retrieval workload and want to compare naive rotation against evidence-based classification, see [`DESIGN-PARTNER.md`](DESIGN-PARTNER.md).