Repository navigation
Prod hardening - #2
Merged
Merged
Conversation
added 13 commits
October 3, 2026 13:38
QNNBENCH_LOG_FORMAT=json emits one JSON object per line with extra fields as top-level keys, for container log shippers; text stays the default for terminals.
…etrics The 64-client load test showed an unbounded queue turning overload into multi-second tails at flat throughput. Now: - the queue is bounded; when full, requests get 503 + Retry-After immediately instead of waiting; - each request has a deadline; requests that expire while queued are dropped before inference (504) so no compute goes to answers nobody is waiting for; - on shutdown /readyz turns 503, queued work drains, then the worker stops; /healthz stays a pure liveness check; - metrics move to prometheus_client histograms (latency, batch size, inference time), outcome-labelled counters and a queue-depth gauge; - responses carry a model version (checkpoint content hash); - settings are one validated dataclass read from QNNBENCH_* variables. The load test reports shed and timed-out requests separately from errors.
Overload runs showed client p99 far above the server's 2 s request deadline with no 504s: the waiting happened in uvicorn's connection backlog, before a request reaches the app, where application-level shedding cannot see it. loadtest --limit-concurrency passes uvicorn's limit through so excess connections are refused with 503 before any parsing work. Results now include raw status counts, and the startup log shows the resolved device rather than 'auto'.
python -m qnnbench.profile records training steps (wait/warmup/active schedule), exports a Chrome/Perfetto trace, a key_averages summary and run metadata. Gates and fused blocks become named ranges, and --nvtx also emits NVTX ranges for Nsight Systems on CUDA. Labels are off by default so normal runs and torch.compile graphs are unchanged.
- mypy (check_untyped_defs) passes. Fixes it found: Callable instead of the builtin callable in train.build_model; a clear error instead of indexing params=None for a parameterised gate; explicit invariants for derivatives and shift rules; DDP code keeps separate references to the network and its wrapper instead of swapping one variable. - New tests for the experiment summary, every plot, the load test (driving the real app in-process via an injectable transport) and the log formatters. Coverage 65% -> 75%; the floor is 70% because CUDA-only branches can't run on CPU runners.
…tions - uv.lock pins all 94 packages; locked installs route torch to the PyTorch CPU index (uv sources), and GPU machines install with --no-sources to get the CUDA build. CI uses uv sync --locked, which also fails if the lock is stale. - CI: mypy in lint, coverage floor in tests (XML artifact), a profiler smoke run, and pip-audit over the locked set (torch audited by its release version, since +cpu builds are not on PyPI). - Every action is pinned to a full commit SHA with its tag noted. - Docker workflow: builds on PRs that touch the image, pushes sha-<commit>/main/semver tags to GHCR from main and tags, and uploads a Trivy scan to the Security tab (report-only until the base image's CVE baseline is reviewed). - Dependabot: weekly grouped updates for uv, actions and the base image, with PyTorch in its own PR so it can be benchmarked. - pre-commit: large-file guard (the old committed venv), YAML/TOML checks, private-key detection, and ruff/mypy/lock checks run at the locked versions; legacy/ is excluded. - Result JSON files end with a newline.
…PyTorchJob - Manifests pin an immutable release tag instead of :latest; a kustomization lets deploys switch to a sha-<commit> build. - Serving: 2 replicas with a PodDisruptionBudget, zero-unavailability rollouts, startup/readiness(/readyz)/liveness probes, a preStop pause so endpoints drop the pod before the app drains, uvicorn connection limits, JSON logs, Prometheus scrape annotations, and a non-root, read-only, capability-free security context. - HPA scales on the app's queue-depth metric rather than CPU, which says little about a GPU-bound service. - Multi-node DDP as a Kubeflow PyTorchJob; same training code. - k8s/validate.sh checks everything with kubeconform -strict, including the PyTorchJob against the Training Operator's CRD schema; CI runs it.
--config takes a TOML file; precedence is dataclass defaults, then the file, then explicit flags. Unknown keys are an error, so a typo can't silently fall back to a default. Replaces the per-CLI flag loops with one helper. configs/ holds the paper protocol for the QNN and MLP and the hybrid runs, including a DDP + gradient accumulation example; tests check every shipped config loads.
python -m qnnbench.experiments --suite hybrid runs the quantum, random and learned 2x2 patch encoders over several seeds. The single-seed results differed by a few tenths of a percent, too small to read without seed variance.
…onfigs and CI gates
…coders Quantum filter 98.90% +- 0.19, random classical 98.69% +- 0.25, learned 98.83% +- 0.23 (Welch p >= 0.17 for every pair). The earlier single-seed run had the order reversed; both are inside seed noise.
|
You are seeing this message because GitHub Code Scanning has recently been set up for this repository, or this pull request contains the workflow file for the Code Scanning tool. What Enabling Code Scanning Means:
For more information about GitHub Code Scanning, check out the documentation. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.