@fuze the deployment-freeze watchdog detected stuck-containercreating.
fuzepicker/fuzepicker-picker-bcd66cdd7-wz4rn container picker has been PodInitializing for 9.2d (threshold 30m) — likely a stuck volume attach.
Diagnostics (verbatim from the cluster)
{
"container": "picker",
"namespace": "fuzepicker",
"node": "mendys-worker-1",
"pending_minutes": 13188.7,
"pod": "fuzepicker-picker-bcd66cdd7-wz4rn",
"threshold_minutes": 30.0,
"waiting_message": null,
"waiting_reason": "PodInitializing"
}
Automated action taken
- None. This watchdog is read-only against the cluster; prod is GitOps under Argo
selfHeal, so any fix must land in git.
Watchdog run: https://github.com/izzywdev/FuzeInfra/actions/runs/33512624463
This watchdog exists because of the 2026-08-30..09-01 fleet-deployment freeze: a full Loki PVC crash-looped Loki 694 times, the permanently-unhealthy StatefulSet wedged one Argo sync operation in phase=Running for 37h, and every prod change queued silently behind it for 2d11h. Every component reported "running"; nothing alerted.
Thresholds: governance/watchdog-thresholds.json. Detection: scripts-tools/deployment_watchdog.py. This issue is deduplicated on the watchdog-key marker above — it will not be re-filed while it stays open, so close it once the condition is actually resolved.
@fuze the deployment-freeze watchdog detected stuck-containercreating.
fuzepicker/fuzepicker-picker-bcd66cdd7-wz4rncontainerpickerhas been PodInitializing for 9.2d (threshold 30m) — likely a stuck volume attach.Diagnostics (verbatim from the cluster)
{ "container": "picker", "namespace": "fuzepicker", "node": "mendys-worker-1", "pending_minutes": 13188.7, "pod": "fuzepicker-picker-bcd66cdd7-wz4rn", "threshold_minutes": 30.0, "waiting_message": null, "waiting_reason": "PodInitializing" }Automated action taken
selfHeal, so any fix must land in git.Watchdog run: https://github.com/izzywdev/FuzeInfra/actions/runs/33512624463
This watchdog exists because of the 2026-08-30..09-01 fleet-deployment freeze: a full Loki PVC crash-looped Loki 694 times, the permanently-unhealthy StatefulSet wedged one Argo sync operation in
phase=Runningfor 37h, and every prod change queued silently behind it for 2d11h. Every component reported "running"; nothing alerted.Thresholds:
governance/watchdog-thresholds.json. Detection:scripts-tools/deployment_watchdog.py. This issue is deduplicated on thewatchdog-keymarker above — it will not be re-filed while it stays open, so close it once the condition is actually resolved.