Skip to content

test: harden the drain-timeout upgrade timing assertion - #19

Merged
novelKR merged 1 commit into
mainfrom
codex/upgrade-drain-timeout-timing
Sep 28, 2026
Merged

novelKR merged 1 commit into
mainfrom
codex/upgrade-drain-timeout-timing

Conversation

@novelKR

@novelKR novelKR commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Summary

Harden the timing-sensitive drain-timeout upgrade test without changing product upgrade behavior or weakening its existing 10-second guard.

This is a test-only, independent CI-hardening change. It is not CS-RG implementation, changes no DevGuard resource semantics, and does not modify or amend #17.

Base: 9e21cc8f707f16b7da98490293b5ea7e4548916d
Head: 92a0611e442baf277b9b578c087c8028984da1f1

Changed file: crates/daemon/tests/upgrade.rs only.

Motivation

#17 is a one-file historical-document change. Its pull-request macOS run 36408760157 nevertheless ran the full workspace and failed only here:

assertion failed: started.elapsed() < Duration::from_secs(10)

The expected ResourceUnavailable result and did not finish message had already been produced. The same branch tree passed push run 36408700221 on the same macos-14-arm64 image. The failed workspace test prevented all later DG1 suites from running.

This PR does not treat that as a product regression. It makes the assertion diagnostic and ensures a busy-host timing failure no longer suppresses the test's later safety assertions and qualification receipt.

What changes

  • Save the total duration as the immediate next operation after upgrade() returns.
  • Keep the existing elapsed < 10 s guard unchanged and run it last.
  • Run all existing functional/safety assertions first: expected error, current PID/release, no backup, admission reopened, charge preserved, later admission, settlement, and successful retry/replacement.
  • Write drain-timeout.json before the timing assertion, then print the exact captured Duration, milliseconds and phase diagnostics in any timing panic.
  • Add no retry around the timed operation. The existing successful upgrade after settlement remains the test's original recovery assertion, not a retry used to make timing green.

Diagnostic marker sampling

A test-only watcher samples the admission-closure marker every 10 ms. It is explicitly not an exact latency clock: server marker creation precedes the complete in-memory close operation, marker removal precedes the complete reopen response, and scheduling can widen sampling gaps.

The receipt records:

  • configured interval, sample count, observed maximum sample gap and effective measurement resolution;
  • initial marker state;
  • total upgrade() duration;
  • first sampled marker presence after the call began (coarse pre-marker boundary);
  • first sampled absence after presence, the observed marker-present interval, and any sampled tail to return;
  • null plus a field-specific reason for every missed/unorderable transition; and
  • read/start/stop diagnostics.

Startup waits at most one second for the first sample. Shutdown sets a stop flag, waits at most one second, joins only a finished thread, and otherwise detaches after recording the failure; Drop uses the same bounded path. A post-stop observation reduces the final transition race.

The marker-present interval combines the one-second drain/polling period with the early part of reopen. It distinguishes that work from pre-close package/path/hash/selection/recovery-copy preparation; it does not claim to measure logical admission latency exactly.

Baseline and local evidence

Read-only baseline preserved both #17 runs before this branch was created. The failing run was in validate.py's cargo test --workspace, not dg1-upgrade; later suites were skipped.

At unmodified base 9e21cc8…, ten sequential warm exact-test runs all passed. Whole-test time was 21.99–22.26 s (median 22.055 s); the original test did not expose the first upgrade() duration, so those runs cannot localize the delay.

At exact head 92a0611…:

  • A final direct probe passed: total 6.142 s, marker first present at 5.057 s, marker observed present for 1.086 s; 434 samples, 61 ms observed maximum gap/effective resolution.
  • A temporary local zero threshold (never committed) failed last after the receipt had recorded the expected timeout, preserved charge and successful replacement retry. The panic printed 5.970 s and phase data. The 10-second source bound was restored before commit.
  • python3 scripts/qualify.py dg1-upgrade --offline --output target/qualification/pr2-dg1-upgrade-92a0611: passed all stages (3 + 1 + 3 + 3 + 1 + 15 tests). Its timing receipt: total 5.948 s, marker first present 4.882 s, marker observed present 1.070 s, 425 samples, 15 ms effective resolution.
  • Rust 1.95.0 python3 scripts/validate.py --offline --output target/qualification/pr2-full-92a0611: passed source/docs/dependency, fmt, workspace Clippy and 318 workspace test results.
  • git diff --check origin/main HEAD: passed.

Frozen hosted-observation protocol and decision rule

This protocol is recorded before reading any hosted result from this head.

  1. Collect three macOS jobs tied to this exact branch head/tree: the automatic branch-push run, the pull-request run, and one rerun of the pull-request macOS job. Preserve every run/attempt, logs and artifacts separately before rerunning anything.
  2. Extend the same-head sample to at most five macOS jobs if any hosted job fails or essential phase evidence is ambiguous. Essential evidence is unambiguous when the expected safety outcome passes; initial_marker_present is false; marker presence, later absence and the marker-present interval are non-null; watcher diagnostics are empty; and effective resolution is at most 250 ms. A null return-tail caused only by the final absence sample following the captured return is acceptable because it has an explicit reason and does not prevent phase separation.
  3. Treat a failed functional/safety assertion, admission not reopened, charge/release change, or marker-present interval over 3 seconds as a possible product/operability defect. Preserve it and stop; do not change the product or timing threshold without a separate decision.
  4. If an over-10-second total has valid marker-present time near the one-second drain and the excess is before marker presence, the total guard is conflating preparation with drain/reopen. Do not merely raise 10 seconds: separate a marker-phase operability bound from a broader total hang guard, justify both from the preserved observations, then repeat this protocol on the new head.
  5. If the same-head samples remain below 10 seconds with valid phases and no distinct slow phase, retain the unchanged 10-second guard and this diagnostic evidence. If evidence remains ambiguous after five jobs, retain the guard and report the unresolved limitation rather than inventing a value.

No internal retry, average or rerun can turn an individual failed job into a pass.

Hosted results and decision

All three prescribed macOS jobs completed successfully on macos-14-arm64 image 20260831.0302.1 with Rust 1.95.0:

Sample Run / attempt / job Source Total Marker first present Marker present interval Effective resolution
Push 36452785046, 1, 109031443286 branch head 92a0611… 6.956 s 5.797 s 1.173 s 100 ms
PR 36452815437, 1, 109031546203 merge source c58950d… 6.935 s 5.895 s 1.073 s 96 ms
PR rerun 36452815437, 2, 109043451981 merge source c58950d… 7.399 s 6.176 s 1.239 s 100 ms

The merge source's parents are the base and exact branch head, and its Git tree equals the branch head's tree. In every sample the complete workspace validator and dg1-upgrade passed; initial marker state was false; the essential presence/absence fields were present; diagnostics were empty; and resolution was under the frozen 250 ms ceiling. The nullable return-tail had only the pre-accepted, field-specific reason that its final marker-absence sample followed the already-captured return instant.

No job failed, no essential evidence was ambiguous, all totals were below 10 seconds, and all marker-present intervals were below 3 seconds. Therefore the protocol does not escalate to five jobs. It retains the unchanged elapsed < 10 s guard. No product defect or evidence for a new timing value was found; the new diagnostic remains so a future outlier preserves whether time accumulated before the marker or during drain/reopen.

Compatibility and scope

No product source, wire/capability, manifest, lock, dependency, service, credential, journal, release or host setting changes. The installed user service is not involved; tests use an isolated fixture authority and fake service manager. Existing internal idempotent control-call behavior is unchanged.

Rollback

Revert this test-only commit. Product behavior and persisted state require no rollback.

Evidence

Ignored local evidence: evidence/ci-hardening-2026-09-28/, including the two #17 runs, 10-run base observations, forced-final-assertion probe, exact-head qualification and full-validation hashes, all three hosted attempts and pr2/hosted-observations.json. Every hosted attempt was preserved before a rerun; no earlier failure was erased.

Capture the total upgrade duration immediately on return, keep every existing functional and safety assertion ahead of the unchanged ten-second guard, and include the actual duration and diagnostic phase data in any timing panic.

A test-only watcher samples the admission marker at a configured ten-millisecond interval. It records initial state, sample count, observed maximum gap, effective resolution, marker-presence phases and explicit reasons for every missing value. Shutdown waits at most one second and joins only a finished thread; Drop uses the same bounded path.

Qualification receipts now retain the timing data before the final assertion. No product upgrade code, retry behavior, threshold or resource semantics change.
@novelKR
novelKR merged commit 7b1548a into main Sep 28, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant