Skip to content

fix(ws): keep heads and filter subscriptions on primaries while they are up - #31

Merged
jleeh merged 3 commits into
feat/websocket-supportfrom
fix/fallback-tier-subscriptions
Oct 7, 2026
Merged

jleeh merged 3 commits into
feat/websocket-supportfrom
fix/fallback-tier-subscriptions

Conversation

@jleeh

@jleeh jleeh commented Oct 6, 2026

Copy link
Copy Markdown

Supersedes #29. Same problem, fixed where it starts instead of on the read path.

Problem

newHeads is subscribed on every WebSocket upstream and each head goes out from whichever source announces it first. With failover.onDefaultsExhausted on, a tier:fallback upstream with a faster WebSocket told clients about a block no primary had yet. Their follow-up reads then missed on the primaries ("block not found") or escaped to the fallback while the primaries were healthy.

#29 routed those reads to the fallback that had the block. That sends traffic to fallbacks while the primaries are up, and fallbacks should only serve when the primaries are down. So this PR stops eRPC advertising a block only a fallback has, and applies the same tier rule to every WebSocket subscription.

Changes

  1. Hedge keeper (erpc/network_executor.go). An ErrUpstreamsExhausted whose causes are all missing-data no longer ends the hedge race. A sibling leg may be on an upstream this leg never tried, for example a slow fallback reached through the escape. This bug is independent of the rest: both TestFailover_Hedge* tests fail on 05126d01 with no tip leader involved. (Tests from fix(network): route block-pinned requests to fallbacks that have the block #29.)
  2. Fallback heads (erpc/subscription_manager.go, indexer/). A fallback's head reaches clients, and feeds the delivered-head floor, only while no non-fallback upstream that the selection policy routes to has a live newHeads subscription. A held-back head still updates that fallback's state poller.
    • Fallback heads still flow when the network has no WebSocket primary, or when the primaries' subscriptions are dead.
  3. Filter subscriptions (indexer/). logs and newPendingTransactions used to pick their upstreams once, at first subscribe, and never revisit them. Now a filter is kept on every eligible primary, and on the fallback tier only while none of those has it live. This is rechecked on every head.
    • Fallbacks hand a filter back only once a primary has it live, so there is no gap.
    • An upstream whose subscribe fails keeps the filter and retries it with the adapter's backoff. Before, the filter was dropped and that upstream never carried it.
    • "Down" is now the selection policy's verdict for heads, filters and reads alike. Filters used to go by the eth_subscribe circuit breaker.

Behaviour to know

  • When primaries fail, fallback heads and filters take over once the policy excludes the primaries: error-rate and lag predicates, then the next evalInterval tick. With evalInterval: 1m that is about 30–90s. Lower evalInterval to shorten it.
  • A primary whose subscription is live but silent holds the fallbacks back until the policy excludes it for lag.
  • Interface change: indexer.NetworkHandle.SuggestLatestBlock returns whether to deliver the head, and indexer.EventIngress gains FilterLive.

Test plan

  • go test ./erpc/ ./indexer/... ./common/ ./upstream/ ./telemetry/. The only failures are TestEvmJsonRpcCache_{DynamoDB,Redis}, which need Docker (testcontainers).
  • WebSocket, gate, failover and indexer tests under -race.
  • Canary on a network with WebSocket on both tiers

🤖 Generated with Claude Code

jleeh and others added 3 commits October 6, 2026 09:01
The hedge keeper kept an ErrUpstreamsExhausted whose causes were all
missing-data, on the premise that no sibling leg could do better. A
sibling may be on an upstream this leg never tried, e.g. a leg the
fallback escape sent to a slow fallback while another leg re-swept the
primaries, so the fast miss cancelled the leg about to succeed. Such a
result now keeps the race going; if every leg ends that way the hedge
still returns the last one to the retry layer.

Co-authored-by: Vitas Spokas <[email protected]>
Co-Authored-By: Claude Opus 5.5 <[email protected]>
newHeads is subscribed on every WebSocket upstream and each head goes
out from whichever source announces it first. With failover on, a
fallback-tier upstream whose WebSocket is faster told clients about a
block no primary had yet; their follow-up reads then missed on the
primaries, or escaped to the fallback while the primaries were healthy.

A fallback's head now reaches clients, and feeds the delivered-head
floor, only while no upstream outside that tier which the selection
policy routes to has a live newHeads subscription of its own. "Down"
is the policy's verdict, as for reads, so custom policies that keep
fallbacks in the ordered list behave the same. A network whose only
WebSocket upstreams are fallbacks, or whose primaries' subscriptions
are dead, keeps receiving fallback heads. A held-back head still
updates that fallback's state poller.

NetworkHandle.SuggestLatestBlock now reports whether the head may be
delivered; the indexer drops it before dedup otherwise.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
logs and newPendingTransactions picked their upstreams once, when the
first client subscribed: on every primary whose eth_subscribe breaker
was closed, on the fallbacks only if every primary rejected the
subscribe. Nothing revisited that for the filter's lifetime, so a
filter went silent when its primaries dropped it, stayed on the
fallbacks after the primaries recovered, and never reached a primary
whose first subscribe failed.

A filter now belongs on every primary the selection policy routes to,
and on the fallback tier while none of those has it live; the indexer
rechecks this on every head, adding and removing the filter as that
changes. The fallbacks hand a filter back only once a primary has it
live, so there is no gap. An upstream whose subscribe fails keeps the
filter and retries it in the background, like its other subscriptions.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant