Skip to content

A repository with more locks than MaxSnapshotLocks is re-walked up to the ceiling on every GET locks, then relayed anyway #65

Description

@matt-edmondson

What's wrong

LockListRefresher.RefreshAsync walks the upstream lock listing page by page. Once the count exceeds Locks.MaxSnapshotLocks, it returns LockRefreshResult.TooLarge (GitLfsCache/Locks/LockListRefresher.cs:120-123).

LockListService handles TooLarge the same as any other failure (LockListService.cs:166-170): it records a refresh failure and relays. Nothing remembers the outcome. TooLarge has no other production reference and no test.

The next GET locks for that repository finds no snapshot, starts a new refresh, walks to the ceiling again, and then relays again.

The locks spec says such a repository "falls back to relaying GET locks for that repository, logging once" (docs/superpowers/specs/2026-08-19-locks-subsystem-design.md:232). There is no per-repository fallback and no log.

Why it matters

Take a repository over the ceiling (default 100,000 locks) that editors poll every 30 s. On GitHub's page size, each request costs about ceiling ÷ page-size upstream page fetches, around 1,000, followed by the relayed call. That is far more upstream traffic than having no proxy at all, and it eats the token's rate limit. The failure is also silent: only a generic refresh-failure metric moves.

Suggested fix

  • On TooLarge, record the LockSnapshotKey as over-ceiling, with a TTL such as a few multiples of ListTtl, or for the life of the process.
  • While the mark is set, ResolveAsync returns Relay() straight away without walking.
  • Log once per key when it is marked, and add a distinct metric or outcome tag so the over-ceiling state is visible.

Acceptance criteria

  • Against a stub upstream whose lock count exceeds MaxSnapshotLocks, the first GET locks walks to the ceiling and relays.
  • The next N GET locks within the TTL relay without any page walk; the stub counts GET locks page requests to check this.
  • One warning is logged per key.

Activity

  1. matt-edmondson commented on Sep 28, 2026

    @matt-edmondson
    ContributorAuthor

    Triage

    • Category: Bug
    • Priority: Medium. Only repositories above MaxSnapshotLocks (default 100,000) are affected. For those, every polled GET locks costs about 1,000 upstream page fetches and then relays anyway, which can exhaust the token's rate limit. The spec's "fall back, log once" is not implemented.
    • Area / suggested assignment: Locks, in GitLfsCache/Locks/LockListService.cs (lines 166-170) and LockListRefresher.cs (lines 120-123)
    • Duplicates: none found among open issues in the org
    • In progress: no direct match. Open PRs Invalidate every ref's lock snapshot when a relayed lock changes #60 (snapshot invalidation) and Forward the client's refspec on cached lock walks and probes #61 (refspec on lock walks and probes) touch the same LockListRefresher/LockListService code, so rebase onto them.

    Notes: A per-LockSnapshotKey over-ceiling mark with a TTL, plus a one-time warning and a distinct metric tag, is enough. Consider fixing it together with #63, which is also about lock-list walks that bypass throttling.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions