Skip to content

fix(deploy): make readiness waits configurable (default 90s) - #24

Merged
echoomegaprime merged 1 commit into
mainfrom
agent/deploy-readiness-timeout
Sep 25, 2026
Merged

echoomegaprime merged 1 commit into
mainfrom
agent/deploy-readiness-timeout

Conversation

@echoomegaprime

Copy link
Copy Markdown
Owner

What

deploy/deploy_forge.sh waited only 20s (40 × 0.5s) for /healthz in three places:

Wait Failure mode when the service is merely slow
Staging boot Release aborts with "staging never became healthy"
Rollback restore A healthy prior release is reported as ROLLBACK FAILED
Production health (8/9) A good release gets rolled back

On 2026-09-24 the d8ad9fd security release aborted at staging twice for exactly this reason. The same candidate passed staging and prod live-smoke once the window was widened.

Change

  • New CERTFORGE_READY_TIMEOUT_S (default 90) drives all three waits. Success still returns on the first healthy probe, so healthy deploys are no slower.
  • Staging readiness failure now prints the /healthz probe result and keeps service.log under $STATE_ROOT/deploy-logs/staging-<release>.log (the EXIT trap previously deleted it, so these failures left no evidence).
  • Existing failure messages are unchanged (test_deploy_gate.py pins one of them).

Verification

  • bash -n deploy/deploy_forge.sh: OK
  • pytest tests/test_deploy_gate.py: 10 passed

🤖 Generated with Claude Code

https://claude.ai/code/session_0152mN1wMwj9vE2YmV2xZF4F

The staging boot, rollback-restore and production health waits each polled
/healthz for only 20s (40 x 0.5s). A cold uvicorn boot on a loaded FORGE
exceeds that: it aborted the d8ad9fd security release at staging, and the
same budget can report a healthy rollback as failed or roll back a good
production release.

- CERTFORGE_READY_TIMEOUT_S (default 90) drives all three waits; a healthy
  service still returns on the first successful probe.
- On a staging readiness failure, print the /healthz probe result and keep
  service.log under $STATE_ROOT/deploy-logs/ before the EXIT trap removes
  the scratch directory.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0152mN1wMwj9vE2YmV2xZF4F
@echoomegaprime
echoomegaprime merged commit 4f7d99e into main Sep 25, 2026
2 checks passed
@echoomegaprime
echoomegaprime deleted the agent/deploy-readiness-timeout branch September 25, 2026 22:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant