What's wrong
Single-flight coalescing doesn't survive the leader going away:
Fetching/SingleFlight.cs (Ticket.Dispose): a leader that never reported an outcome completes its followers with false.
Endpoints/ObjectRouteHandler.cs (around lines 163–201): a follower whose WaitForLeaderAsync returns false logs LeaderDidNotFinish and then calls StreamFromUpstreamAsync(..., storeLocally: true, ...) itself. It never calls coalescer.Acquire again.
The leader's upstream fetch runs on its own client's RequestAborted token. So when that one client disconnects mid-transfer (a killed CI pod, a client timeout, Ctrl-C), all waiting followers start their own upstream fetch of the same object at the same moment. They also race to stage and publish it.
This is the thundering herd the README's "one upstream fetch per object" promise is meant to prevent. It happens exactly when many CI jobs start together and pull the same large object.
Reproduction
Test with a gated upstream, 1 leader and 5 followers:
- Queue the 5 followers behind the leader.
- Cancel the leader's request.
- Release the upstream gate.
Result: all 5 followers returned 200, but upstream saw 6 object fetches (1 before the abort, 5 after). The expected count is 2 at most: the aborted one plus a single replacement.
Suggested fix
Either of these, or both:
- A follower released with
false, including after a FollowerTimeout, calls coalescer.Acquire(upstream, oid) again. One follower becomes the new leader and the rest keep waiting, with a bound on how many times this can repeat.
- Run the leader's upstream fetch-and-store on a token that isn't tied to the leader's client. The object is then still fetched and published for the followers when the leader disconnects; only the leader's copy to its own response stops.
Acceptance: a test with N followers and an aborted leader sees at most one additional upstream fetch.
What's wrong
Single-flight coalescing doesn't survive the leader going away:
Fetching/SingleFlight.cs(Ticket.Dispose): a leader that never reported an outcome completes its followers withfalse.Endpoints/ObjectRouteHandler.cs(around lines 163–201): a follower whoseWaitForLeaderAsyncreturnsfalselogsLeaderDidNotFinishand then callsStreamFromUpstreamAsync(..., storeLocally: true, ...)itself. It never callscoalescer.Acquireagain.The leader's upstream fetch runs on its own client's
RequestAbortedtoken. So when that one client disconnects mid-transfer (a killed CI pod, a client timeout, Ctrl-C), all waiting followers start their own upstream fetch of the same object at the same moment. They also race to stage and publish it.This is the thundering herd the README's "one upstream fetch per object" promise is meant to prevent. It happens exactly when many CI jobs start together and pull the same large object.
Reproduction
Test with a gated upstream, 1 leader and 5 followers:
Result: all 5 followers returned 200, but upstream saw 6 object fetches (1 before the abort, 5 after). The expected count is 2 at most: the aborted one plus a single replacement.
Suggested fix
Either of these, or both:
false, including after aFollowerTimeout, callscoalescer.Acquire(upstream, oid)again. One follower becomes the new leader and the rest keep waiting, with a bound on how many times this can repeat.Acceptance: a test with N followers and an aborted leader sees at most one additional upstream fetch.