Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 31 additions & 0 deletions .github/ISSUE_TEMPLATE/bug-report.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
---
name: Bug report
description: Report a reproducible behavior issue using sanitized evidence
title: "[bug] "
labels: []
assignees: []
---

> Public issue: do not paste credentials, cookies, private URLs/IPs, internal hostnames, raw production logs, HAR/PCAP captures, customer data, or infrastructure details.

## Observed behavior

Describe what the library returned or did.

## Expected behavior

Describe the evidence-preserving behavior you expected.

## Minimal synthetic reproducer

Provide the smallest sanitized example that reproduces the issue.

## Environment

- Python version:
- `proxy-outcome` version/commit:
- OS/runtime, if relevant:

## Additional context

Include only non-sensitive details needed to understand the bug.
15 changes: 15 additions & 0 deletions .github/pull_request_template.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
## Summary

What changed and why?

## Evidence

What observation, test, or reproducible case supports the change?

## Safety / scope check

- [ ] No credentials, private endpoints/IPs, customer data, or production captures are included.
- [ ] The change does not invent root cause when evidence is insufficient.
- [ ] Public/private implementation boundaries are preserved.
- [ ] Tests were added or updated when behavior changed.
- [ ] Local tests and repository hygiene checks pass.
45 changes: 45 additions & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
# Contributing

Thanks for considering a contribution to `proxy-outcome`.

## Scope

Good contributions improve the public deterministic baseline without weakening its core rule: preserve observations without inventing root cause.

Useful changes include:

- clearer taxonomy or documentation;
- additional conservative classification cases;
- regression tests for ambiguous outcomes;
- CLI usability improvements;
- packaging or CI improvements;
- security and data-minimization hardening.

Private routing/scoring logic, provider-specific private configuration, credentials, production infrastructure, and non-public benchmark data do not belong in this repository.

## Development

```bash
python -m pip install .
python -m unittest discover -s tests -v
python scripts/public_hygiene_check.py
```

Keep runtime dependencies at zero unless there is a strong, documented reason to change that constraint.

## Pull requests

A good pull request should:

1. explain the observed problem before proposing attribution;
2. include tests for behavior changes;
3. avoid changing unrelated code;
4. preserve fail-closed behavior when evidence is insufficient;
5. keep examples synthetic and non-sensitive;
6. keep CI green across the supported Python matrix.

## Sensitive information

Do not post credentials, cookies, raw authorization headers, private URLs or IPs, internal hostnames, production logs, HAR/PCAP captures, customer data, or private infrastructure details.

See [`SECURITY.md`](SECURITY.md) for the repository security boundary.
19 changes: 17 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,13 +1,24 @@
# proxy-outcome

> Evidence-first classification for proxy and web-retrieval outcomes.

**403 ≠ bad proxy. 429 ≠ bad proxy. 451 ≠ bad proxy.**

`proxy-outcome` is a small, deterministic classification library for HTTP and proxy-path observations in scraping and web-retrieval systems.
`proxy-outcome` is a small, deterministic Python library that classifies HTTP and proxy-path observations without inventing root cause. It is designed for scraper, crawler, browser-automation, and web-retrieval systems where a failed request should not automatically poison a proxy pool.

Its job is deliberately narrow: preserve evidence without inventing a root cause.
**Status:** alpha · Python 3.10+ · zero runtime dependencies · local-only · no telemetry

Built by **PN Labs**.

## PN Labs reliability toolkit

| Project | Purpose |
| --- | --- |
| **proxy-outcome** | Classify what happened without over-attributing the cause |
| [**proxybench**](https://github.com/pnlabs-dev/proxybench) | Compare usable-result efficiency between retrieval policies |

The tools are intentionally separate: classify evidence first, then benchmark whether a policy actually improves outcomes.

## Why this exists

A common reliability mistake is:
Expand Down Expand Up @@ -158,6 +169,10 @@ CI installs the package before testing, runs CLI smoke tests, and executes a non

See [`docs/ARCHITECTURE.md`](docs/ARCHITECTURE.md).

## Contributing

Small, evidence-preserving improvements are welcome. See [`CONTRIBUTING.md`](CONTRIBUTING.md) before opening a pull request.

## Design partners

If you operate a legitimate scraping/web-retrieval workload and want to compare naive rotation against evidence-based classification, see [`DESIGN-PARTNER.md`](DESIGN-PARTNER.md).
Expand Down