Skip to content

fix: keep polling intervals when a provider update throws - #322

Merged
matt-edmondson merged 3 commits into
mainfrom
fix/299-failed-update-keeps-interval
Sep 30, 2026
Merged

matt-edmondson merged 3 commits into
mainfrom
fix/299-failed-update-keeps-interval

Conversation

@matt-edmondson

@matt-edmondson matt-edmondson commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #299

What was wrong

MakeGitHubRequestAsync rethrows 5xx ApiExceptions. When one arrived:

  • BuildSync.UpdateAsync and RunSync.UpdateAsync threw before UpdateTimer.Restart(), so ShouldUpdate stayed true.
  • The Task.WhenAll in BuildMonitor.UpdateAsync faulted, which skipped ProviderRefreshTimer.Restart() and PruneOrphanedAndCompletedSyncs.
  • OnRender then restarted the update loop immediately. Discovery and polling ran back to back instead of every 120 s / 300 s.

Change

  • New SyncGuard.RunAsync runs one item of a batch and logs its exception with Log.Error instead of letting it fault the batch.
  • BuildSync.UpdateAsync and RunSync.UpdateAsync run their provider call through SyncGuard and restart their timer in finally.
  • In BuildMonitor.UpdateAsync, every per-owner and per-repository discovery call is now guarded. One bad repository no longer stops discovery for the others, and the ProviderRefreshTimer.Restart() after the batch always runs.
  • The build and run batches can no longer fault, so pruning runs every cycle again.

Tests

SyncFailureTests uses a fake provider whose update throws HttpRequestException:

  • AFailedBuildUpdateStillRestartsItsTimer: after a failed update, TimeRemaining is back at the full interval and ShouldUpdate is false.
  • AFailedBuildUpdateDoesNotStopTheOthersInTheBatch: a WhenAll over one failing build and one healthy build completes without faulting, and the healthy build is still polled.
  • AFailedRunUpdateDoesNotStopTheOthersInTheBatch: the same check for runs. The failed run is also no longer due.

Each test fails with its production change reverted and passes with it. The full suite passes, 80/80.

🤖 Generated with Claude Code

https://claude.ai/code/session_01F4JLDZ8V6BpUZBB9ZnsaNt

A persistent 5xx from one repository made BuildSync/RunSync.UpdateAsync
throw before restarting their timers, and faulted the Task.WhenAll over
the batch, which skipped ProviderRefreshTimer.Restart and pruning. The
update loop then re-ran discovery and re-polled back to back.

Each sync and each discovery item now runs through SyncGuard, which logs
a failure instead of faulting the batch; the sync timers and the
provider refresh timer restart in finally blocks.

Fixes #299

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F4JLDZ8V6BpUZBB9ZnsaNt
Narrow the BuildMonitor.cs change to guarding each discovery item: once
no item can fault the batch, the ProviderRefreshTimer restart after it
is always reached, so the method no longer needs moving. Add a test
that a failing run update neither faults its batch nor stays due.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F4JLDZ8V6BpUZBB9ZnsaNt
Moves the per-item guarding out of the static update loop into a helper
that can be tested directly, and covers it with a test that a failing
discovery item is logged while the rest still run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F4JLDZ8V6BpUZBB9ZnsaNt
@sonarqubecloud

Copy link
Copy Markdown

@matt-edmondson
matt-edmondson merged commit 78d99e4 into main Sep 30, 2026
14 checks passed
@matt-edmondson
matt-edmondson deleted the fix/299-failed-update-keeps-interval branch September 30, 2026 06:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

One persistent GitHub 5xx makes the app re-run discovery and re-poll back-to-back instead of on its interval

2 participants