Repository navigation
fix(ws): keep heads and filter subscriptions on primaries while they are up - #31
Merged
Merged
Conversation
The hedge keeper kept an ErrUpstreamsExhausted whose causes were all missing-data, on the premise that no sibling leg could do better. A sibling may be on an upstream this leg never tried, e.g. a leg the fallback escape sent to a slow fallback while another leg re-swept the primaries, so the fast miss cancelled the leg about to succeed. Such a result now keeps the race going; if every leg ends that way the hedge still returns the last one to the retry layer. Co-authored-by: Vitas Spokas <[email protected]> Co-Authored-By: Claude Opus 5.5 <[email protected]>
newHeads is subscribed on every WebSocket upstream and each head goes out from whichever source announces it first. With failover on, a fallback-tier upstream whose WebSocket is faster told clients about a block no primary had yet; their follow-up reads then missed on the primaries, or escaped to the fallback while the primaries were healthy. A fallback's head now reaches clients, and feeds the delivered-head floor, only while no upstream outside that tier which the selection policy routes to has a live newHeads subscription of its own. "Down" is the policy's verdict, as for reads, so custom policies that keep fallbacks in the ordered list behave the same. A network whose only WebSocket upstreams are fallbacks, or whose primaries' subscriptions are dead, keeps receiving fallback heads. A held-back head still updates that fallback's state poller. NetworkHandle.SuggestLatestBlock now reports whether the head may be delivered; the indexer drops it before dedup otherwise. Co-Authored-By: Claude Opus 5.5 <[email protected]>
logs and newPendingTransactions picked their upstreams once, when the first client subscribed: on every primary whose eth_subscribe breaker was closed, on the fallbacks only if every primary rejected the subscribe. Nothing revisited that for the filter's lifetime, so a filter went silent when its primaries dropped it, stayed on the fallbacks after the primaries recovered, and never reached a primary whose first subscribe failed. A filter now belongs on every primary the selection policy routes to, and on the fallback tier while none of those has it live; the indexer rechecks this on every head, adding and removing the filter as that changes. The fallbacks hand a filter back only once a primary has it live, so there is no gap. An upstream whose subscribe fails keeps the filter and retries it in the background, like its other subscriptions. Co-Authored-By: Claude Opus 5.5 <[email protected]>
This was referenced Oct 6, 2026
This was referenced Oct 7, 2026
Draft
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Supersedes #29. Same problem, fixed where it starts instead of on the read path.
Problem
newHeadsis subscribed on every WebSocket upstream and each head goes out from whichever source announces it first. Withfailover.onDefaultsExhaustedon, atier:fallbackupstream with a faster WebSocket told clients about a block no primary had yet. Their follow-up reads then missed on the primaries ("block not found") or escaped to the fallback while the primaries were healthy.#29 routed those reads to the fallback that had the block. That sends traffic to fallbacks while the primaries are up, and fallbacks should only serve when the primaries are down. So this PR stops eRPC advertising a block only a fallback has, and applies the same tier rule to every WebSocket subscription.
Changes
erpc/network_executor.go). AnErrUpstreamsExhaustedwhose causes are all missing-data no longer ends the hedge race. A sibling leg may be on an upstream this leg never tried, for example a slow fallback reached through the escape. This bug is independent of the rest: bothTestFailover_Hedge*tests fail on05126d01with no tip leader involved. (Tests from fix(network): route block-pinned requests to fallbacks that have the block #29.)erpc/subscription_manager.go,indexer/). A fallback's head reaches clients, and feeds the delivered-head floor, only while no non-fallback upstream that the selection policy routes to has a livenewHeadssubscription. A held-back head still updates that fallback's state poller.indexer/).logsandnewPendingTransactionsused to pick their upstreams once, at first subscribe, and never revisit them. Now a filter is kept on every eligible primary, and on the fallback tier only while none of those has it live. This is rechecked on every head.eth_subscribecircuit breaker.Behaviour to know
evalIntervaltick. WithevalInterval: 1mthat is about 30–90s. LowerevalIntervalto shorten it.indexer.NetworkHandle.SuggestLatestBlockreturns whether to deliver the head, andindexer.EventIngressgainsFilterLive.Test plan
go test ./erpc/ ./indexer/... ./common/ ./upstream/ ./telemetry/. The only failures areTestEvmJsonRpcCache_{DynamoDB,Redis}, which need Docker (testcontainers).-race.🤖 Generated with Claude Code