What's wrong
LockListRefresher.RefreshAsync walks the upstream lock listing page by page. Once the count exceeds Locks.MaxSnapshotLocks, it returns LockRefreshResult.TooLarge (GitLfsCache/Locks/LockListRefresher.cs:120-123).
LockListService handles TooLarge the same as any other failure (LockListService.cs:166-170): it records a refresh failure and relays. Nothing remembers the outcome. TooLarge has no other production reference and no test.
The next GET locks for that repository finds no snapshot, starts a new refresh, walks to the ceiling again, and then relays again.
The locks spec says such a repository "falls back to relaying GET locks for that repository, logging once" (docs/superpowers/specs/2026-08-19-locks-subsystem-design.md:232). There is no per-repository fallback and no log.
Why it matters
Take a repository over the ceiling (default 100,000 locks) that editors poll every 30 s. On GitHub's page size, each request costs about ceiling ÷ page-size upstream page fetches, around 1,000, followed by the relayed call. That is far more upstream traffic than having no proxy at all, and it eats the token's rate limit. The failure is also silent: only a generic refresh-failure metric moves.
Suggested fix
- On
TooLarge, record the LockSnapshotKey as over-ceiling, with a TTL such as a few multiples of ListTtl, or for the life of the process.
- While the mark is set,
ResolveAsync returns Relay() straight away without walking.
- Log once per key when it is marked, and add a distinct metric or outcome tag so the over-ceiling state is visible.
Acceptance criteria
- Against a stub upstream whose lock count exceeds
MaxSnapshotLocks, the first GET locks walks to the ceiling and relays.
- The next N
GET locks within the TTL relay without any page walk; the stub counts GET locks page requests to check this.
- One warning is logged per key.
What's wrong
LockListRefresher.RefreshAsyncwalks the upstream lock listing page by page. Once the count exceedsLocks.MaxSnapshotLocks, it returnsLockRefreshResult.TooLarge(GitLfsCache/Locks/LockListRefresher.cs:120-123).LockListServicehandlesTooLargethe same as any other failure (LockListService.cs:166-170): it records a refresh failure and relays. Nothing remembers the outcome.TooLargehas no other production reference and no test.The next
GET locksfor that repository finds no snapshot, starts a new refresh, walks to the ceiling again, and then relays again.The locks spec says such a repository "falls back to relaying
GET locksfor that repository, logging once" (docs/superpowers/specs/2026-08-19-locks-subsystem-design.md:232). There is no per-repository fallback and no log.Why it matters
Take a repository over the ceiling (default 100,000 locks) that editors poll every 30 s. On GitHub's page size, each request costs about ceiling ÷ page-size upstream page fetches, around 1,000, followed by the relayed call. That is far more upstream traffic than having no proxy at all, and it eats the token's rate limit. The failure is also silent: only a generic refresh-failure metric moves.
Suggested fix
TooLarge, record theLockSnapshotKeyas over-ceiling, with a TTL such as a few multiples ofListTtl, or for the life of the process.ResolveAsyncreturnsRelay()straight away without walking.Acceptance criteria
MaxSnapshotLocks, the firstGET lockswalks to the ceiling and relays.GET lockswithin the TTL relay without any page walk; the stub countsGET lockspage requests to check this.