What's wrong
Tick calls FetchAllReposIfStale() on every frame (ProjectDirector/ProjectDirector.cs:690). A cloned repository is stale when its LastFetchTime is more than MinFetchIntervalSeconds old. The default is 60 (GitRepository.cs:16), and the setting that would change it is commented out (:741-750). For each stale repository, FetchRepo (:436-447):
- sets
LastFetchTime = UtcNow
- queues an options save
- starts a new
Task that runs git fetch origin, with no limit on how many run at once
All stale repositories are handled in the same frame, which leads to three problems:
- On startup, every cloned repository is stale, because
LastFetchTime was saved on the previous run. N git fetch processes start in the first frame.
- After that they stay in lockstep. All N got the same
LastFetchTime, so they all go stale on the same frame 60 seconds later and start together again. That is a burst of N concurrent network processes every minute for as long as the app is open.
- Nothing limits concurrency or checks for a fetch already running. If a fetch takes longer than the interval (a slow network, a credential-helper prompt, a large repository), the next tick starts a second
git fetch in the same repository while the first is still running. The two then compete for ref locks (cannot lock ref 'refs/remotes/origin/...').
Failure scenario
With 40 repositories cloned over SSH, 40 SSH connections to github.com open within milliseconds every minute. GitHub throttles concurrent unauthenticated SSH handshakes (sshd MaxStartups), so some of them fail with kex_exchange_identification: Connection closed by remote host. Those fetches then show up as failures in the log on every cycle.
Over HTTPS, 40 credential-helper invocations can all prompt at the same time. A manual Pull that happens during the burst also races the background fetch in the same repository.
Whether throttling actually happens depends on the network. The lockstep schedule and the unbounded fan-out follow directly from the code above (traced at HEAD 011730d).
Suggested fix / acceptance criteria
- Run background fetches through a small bounded worker, such as a
SemaphoreSlim of 4, or a single queue drained a few at a time.
- Skip a repository whose previous fetch is still running, tracking in-flight fetches per
LocalPath.
- Spread the schedule so the repositories drift out of phase instead of refiring together. Either add jitter to the next fetch time, or fetch one stale repository per tick.
- With 40 cloned repositories, the number of concurrent
git fetch processes never exceeds the bound, and one repository never has two fetches running at once.
Related but distinct: #438 also involves a fetch failing every minute, but its cause is a wrong LocalPath after the owner scan.
What's wrong
TickcallsFetchAllReposIfStale()on every frame (ProjectDirector/ProjectDirector.cs:690). A cloned repository is stale when itsLastFetchTimeis more thanMinFetchIntervalSecondsold. The default is 60 (GitRepository.cs:16), and the setting that would change it is commented out (:741-750). For each stale repository,FetchRepo(:436-447):LastFetchTime = UtcNowTaskthat runsgit fetch origin, with no limit on how many run at onceAll stale repositories are handled in the same frame, which leads to three problems:
LastFetchTimewas saved on the previous run. Ngit fetchprocesses start in the first frame.LastFetchTime, so they all go stale on the same frame 60 seconds later and start together again. That is a burst of N concurrent network processes every minute for as long as the app is open.git fetchin the same repository while the first is still running. The two then compete for ref locks (cannot lock ref 'refs/remotes/origin/...').Failure scenario
With 40 repositories cloned over SSH, 40 SSH connections to github.com open within milliseconds every minute. GitHub throttles concurrent unauthenticated SSH handshakes (sshd
MaxStartups), so some of them fail withkex_exchange_identification: Connection closed by remote host. Those fetches then show up as failures in the log on every cycle.Over HTTPS, 40 credential-helper invocations can all prompt at the same time. A manual Pull that happens during the burst also races the background fetch in the same repository.
Whether throttling actually happens depends on the network. The lockstep schedule and the unbounded fan-out follow directly from the code above (traced at HEAD 011730d).
Suggested fix / acceptance criteria
SemaphoreSlimof 4, or a single queue drained a few at a time.LocalPath.git fetchprocesses never exceeds the bound, and one repository never has two fetches running at once.Related but distinct: #438 also involves a fetch failing every minute, but its cause is a wrong
LocalPathafter the owner scan.