Skip to content

Prod hardening - #2

Merged
apayne185 merged 13 commits into
mainfrom
prod-hardening
Oct 3, 2026
Merged

apayne185 merged 13 commits into
mainfrom
prod-hardening

Conversation

@apayne185

Copy link
Copy Markdown
Owner

No description provided.

apayne185 added 13 commits October 3, 2026 13:38
QNNBENCH_LOG_FORMAT=json emits one JSON object per line with extra
fields as top-level keys, for container log shippers; text stays the
default for terminals.
…etrics

The 64-client load test showed an unbounded queue turning overload into
multi-second tails at flat throughput. Now:
- the queue is bounded; when full, requests get 503 + Retry-After
  immediately instead of waiting;
- each request has a deadline; requests that expire while queued are
  dropped before inference (504) so no compute goes to answers nobody
  is waiting for;
- on shutdown /readyz turns 503, queued work drains, then the worker
  stops; /healthz stays a pure liveness check;
- metrics move to prometheus_client histograms (latency, batch size,
  inference time), outcome-labelled counters and a queue-depth gauge;
- responses carry a model version (checkpoint content hash);
- settings are one validated dataclass read from QNNBENCH_* variables.

The load test reports shed and timed-out requests separately from errors.
Overload runs showed client p99 far above the server's 2 s request
deadline with no 504s: the waiting happened in uvicorn's connection
backlog, before a request reaches the app, where application-level
shedding cannot see it. loadtest --limit-concurrency passes uvicorn's
limit through so excess connections are refused with 503 before any
parsing work. Results now include raw status counts, and the startup
log shows the resolved device rather than 'auto'.
python -m qnnbench.profile records training steps (wait/warmup/active
schedule), exports a Chrome/Perfetto trace, a key_averages summary and
run metadata. Gates and fused blocks become named ranges, and --nvtx
also emits NVTX ranges for Nsight Systems on CUDA. Labels are off by
default so normal runs and torch.compile graphs are unchanged.
- mypy (check_untyped_defs) passes. Fixes it found: Callable instead
  of the builtin callable in train.build_model; a clear error instead
  of indexing params=None for a parameterised gate; explicit invariants
  for derivatives and shift rules; DDP code keeps separate references
  to the network and its wrapper instead of swapping one variable.
- New tests for the experiment summary, every plot, the load test
  (driving the real app in-process via an injectable transport) and
  the log formatters. Coverage 65% -> 75%; the floor is 70% because
  CUDA-only branches can't run on CPU runners.
…tions

- uv.lock pins all 94 packages; locked installs route torch to the
  PyTorch CPU index (uv sources), and GPU machines install with
  --no-sources to get the CUDA build. CI uses uv sync --locked, which
  also fails if the lock is stale.
- CI: mypy in lint, coverage floor in tests (XML artifact), a profiler
  smoke run, and pip-audit over the locked set (torch audited by its
  release version, since +cpu builds are not on PyPI).
- Every action is pinned to a full commit SHA with its tag noted.
- Docker workflow: builds on PRs that touch the image, pushes
  sha-<commit>/main/semver tags to GHCR from main and tags, and
  uploads a Trivy scan to the Security tab (report-only until the
  base image's CVE baseline is reviewed).
- Dependabot: weekly grouped updates for uv, actions and the base
  image, with PyTorch in its own PR so it can be benchmarked.
- pre-commit: large-file guard (the old committed venv), YAML/TOML
  checks, private-key detection, and ruff/mypy/lock checks run at
  the locked versions; legacy/ is excluded.
- Result JSON files end with a newline.
…PyTorchJob

- Manifests pin an immutable release tag instead of :latest; a
  kustomization lets deploys switch to a sha-<commit> build.
- Serving: 2 replicas with a PodDisruptionBudget, zero-unavailability
  rollouts, startup/readiness(/readyz)/liveness probes, a preStop pause
  so endpoints drop the pod before the app drains, uvicorn connection
  limits, JSON logs, Prometheus scrape annotations, and a non-root,
  read-only, capability-free security context.
- HPA scales on the app's queue-depth metric rather than CPU, which
  says little about a GPU-bound service.
- Multi-node DDP as a Kubeflow PyTorchJob; same training code.
- k8s/validate.sh checks everything with kubeconform -strict, including
  the PyTorchJob against the Training Operator's CRD schema; CI runs it.
--config takes a TOML file; precedence is dataclass defaults, then the
file, then explicit flags. Unknown keys are an error, so a typo can't
silently fall back to a default. Replaces the per-CLI flag loops with
one helper. configs/ holds the paper protocol for the QNN and MLP and
the hybrid runs, including a DDP + gradient accumulation example;
tests check every shipped config loads.
python -m qnnbench.experiments --suite hybrid runs the quantum, random
and learned 2x2 patch encoders over several seeds. The single-seed
results differed by a few tenths of a percent, too small to read
without seed variance.
…coders

Quantum filter 98.90% +- 0.19, random classical 98.69% +- 0.25,
learned 98.83% +- 0.23 (Welch p >= 0.17 for every pair). The earlier
single-seed run had the order reversed; both are inside seed noise.
@github-advanced-security

Copy link
Copy Markdown

You are seeing this message because GitHub Code Scanning has recently been set up for this repository, or this pull request contains the workflow file for the Code Scanning tool.

What Enabling Code Scanning Means:

  • The 'Security' tab will display more code scanning analysis results (e.g., for the default branch).
  • Depending on your configuration and choice of analysis tool, future pull requests will be annotated with code scanning analysis results.
  • You will be able to see the analysis results for the pull request's branch on this overview once the scans have completed and the checks have passed.

For more information about GitHub Code Scanning, check out the documentation.

@apayne185
apayne185 merged commit ebd6b54 into main Oct 3, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants