Skip to content

Hand an aborted leader's fetch to one follower instead of all of them [patch] - #110

Merged
matt-edmondson merged 1 commit into
mainfrom
fix/58-reacquire-after-aborted-leader
Oct 9, 2026
Merged

matt-edmondson merged 1 commit into
mainfrom
fix/58-reacquire-after-aborted-leader

Conversation

@matt-edmondson

Copy link
Copy Markdown
Contributor

Fixes #58

What

If a leader's client disconnected mid-transfer, Ticket.Dispose released its followers with false. Each follower then fetched the object from upstream on its own, so 1 leader and 5 followers produced 6 upstream fetches.

This implements the issue's first option.

  • ObjectRouteHandler.DownloadAsync: a follower released without the object (an aborted or failed leader, or a FollowerTimeout) disposes its ticket and calls coalescer.Acquire again. One released follower becomes the new leader and the others wait for it. The loop is bounded by MaxFollowerAttempts = 3; after that a follower fetches for itself, as it did before, so a run of failing leaders can't hold a request forever. The coalesced_waits metric is still recorded once per request.
  • SingleFlight: Complete and Dispose now retire the entry before they set its result. Before, a released follower that re-acquired quickly could find the finished entry still registered and get a follower ticket whose result was already false. That wasted one of its attempts.

I did not do the second option, which runs the leader's fetch on a token that isn't tied to its client. It would mean detaching the upstream fetch-and-store from the response stream the leader is writing. That is a larger change, and this fix already meets the acceptance criterion without it.

Tests

  • ProxyFlowTests.AbortedLeader_HandsTheFetchToOneFollowerInsteadOfEveryFollowerFetching is the issue's repro. It holds the upstream behind a gate, queues 5 followers behind a leader, cancels the leader's request and then opens the gate. It asserts that every follower gets the right bytes and that upstream saw exactly 2 fetches. A coalesced_waits meter listener confirms the followers are queued before the abort.
  • StubUpstream has a new ObjectGate that holds object fetches in flight. A fetch whose request is cancelled while it waits is abandoned.

Verification

  • The new test fails on unchanged main at Assert.AreEqual(2, FetchCount), and passes with the fix.
  • The full suite passes: 355 of 355.
  • The new test passed 8 runs in a row, alongside ConcurrentMissesForOneObject_ProduceASingleUpstreamFetch.

#107 also edits ObjectRouteHandler, but in the store-hit path, so any conflict should be small.

🤖 Generated with Claude Code

https://claude.ai/code/session_017ft9gJVQx6JFNQ4VRoAwd9


Generated by Claude Code

… [patch]

When a leader's client disconnected mid-transfer, its followers were released
with a failure and each fetched the object from upstream itself: the aborted
fetch plus one per follower. A follower released without the object now
acquires again, so one of them becomes the new leader and the rest wait for
it, bounded at three leaders before a follower fetches for itself. The single
flight retires an entry before releasing its followers, so a follower that
acquires again gets a fresh entry rather than the finished one.

Fixes #58

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_017ft9gJVQx6JFNQ4VRoAwd9
@sonarqubecloud

sonarqubecloud Bot commented Oct 8, 2026

Copy link
Copy Markdown

@matt-edmondson
matt-edmondson merged commit 76bea7b into main Oct 9, 2026
14 checks passed
@matt-edmondson
matt-edmondson deleted the fix/58-reacquire-after-aborted-leader branch October 9, 2026 08:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

When the leader request is aborted, every waiting follower fetches the same object from upstream separately (N+1 fetches instead of one)

2 participants