Skip to content

fix(network): stop sweeping upstreams for a block above every head - #32

Closed
snowkide wants to merge 1 commit into
fix/fallback-tier-subscriptionsfrom
fix/hedge-keeper-empty-fanout
Closed

snowkide wants to merge 1 commit into
fix/fallback-tier-subscriptionsfrom
fix/hedge-keeper-empty-fanout

Conversation

@snowkide

@snowkide snowkide commented Oct 7, 2026

Copy link
Copy Markdown

Stacked on #31 (fix/fallback-tier-subscriptions @ 361a84a, the image running on all five internal-erpc instances).

Problem

Since the #31 roll, internal-erpc eth-mainnet sends ~17 r/s of eth_getBlockByNumber to paid fallbacks (Chainstack/Infura). ~97% of those come back null.

Clients poll eth_getBlockByNumber for the block after the tip, about 9s before it is produced (each future block is requested ~5 times). No upstream has it. eRPC already knows this: EmptyResultBeyondConfidence says the block is above the network's highest head, and emptyIsMiss uses that to skip the fallback escape. But two places ignore it:

  • the upstream loop treats the null as "try the next upstream" unless the method is in emptyResultAccept
  • the hedge keep rejects the emptyish result, so the race keeps going

So each request walks every routed upstream, and the retry layer repeats the walk. On eth-mainnet, which routes every tier in one ordered list, that means the paid fallbacks.

The pre-rewrite lineage (the old ws-eth-call-lag-test1 image) avoided this through #14, which pinned near-tip getBlockByNumber to the tip leader. The rewritten feat/websocket-support keeps only an ordering hint (preferTipLeaderForNearTipGetBlock). ws-tip-leader-8e98cb2 was reverted on Oct 2 for the same symptom.

The hedge-keeper change in #31 (46c6a83) is not the cause. With a null result the leg returns the response, not ErrUpstreamsExhausted, so the removed rule never applied. The repro gives the same numbers with and without it.

Change

When the requested block is beyond confidence, an empty result is the answer:

  • erpc/networks.go: the upstream loop accepts it, so no sweep and no escape
  • erpc/network_executor.go: the hedge keep accepts it, so sibling legs are cancelled

A block at or below the head that a primary misses still sweeps and escapes as before.

Tests

erpc/networks_failover_hedge_empty_test.go:

  • TestFailover_EmptyBeyondHeadDoesNotSweepToFallbacks

    • all four upstreams answer null for head+1, with an eth-mainnet-like failsafe (retry 5, hedge maxCount 1)
    • AllTiersRouted (eth-mainnet shape): 0 fallback hits, at most 1 primary hit per request
    • FallbacksCordonedByPolicy: no fallback escape. The policy's probeExcluded mirrors still reach cordoned fallbacks; that predates fix(ws): keep heads and filter subscriptions on primaries while they are up #31 and is unchanged.
    • Without the fix: 4 fallback + 4 primary hits per request; fails
  • TestFailover_EmptyAtHeadStillEscapesToFallbacks: primaries null at the head and fallbacks have the block, so the request still escapes and returns it

  • new tests pass, and fail without the fix

  • go test ./erpc/ ./indexer/... ./common/ ./upstream/ ./telemetry/ ./architecture/evm/ (running)

  • canary, then internal-erpc: paid eth_getBlockByNumber on evm:1 drops; erpc_network_hedged_request_total{network="evm:1"} back near ~34 r/s

Rollout note

#30 (feat/cll-on-websocket-support) is rebased on #31, so it needs another rebase and rebuild after this lands. Draft services-argocd-deployments#1721 pins cll-ws-tier-subs-11108fd, which has this bug.

🤖 Generated with Claude Code

Clients poll eth_getBlockByNumber for the block after the tip before it
is produced. Every upstream answers null, and the request already knows
that (EmptyResultBeyondConfidence: the block is above the network's
highest head), but that only stopped the fallback escape. The upstream
loop and the hedge keeper still treated the null as a miss, so each
request walked every routed upstream, including the paid fallback tier
on networks that route all tiers in one list, and the retry layer
repeated the walk.

Accept the empty result as the answer in both places when the block is
beyond confidence. A block at or below the head that a primary misses
still sweeps and escapes as before.

On internal-erpc eth-mainnet this is the ~17 r/s of
eth_getBlockByNumber reaching Chainstack/Infura since the #31 roll, ~97%
of which came back null. The pre-rewrite lineage avoided it by pinning
near-tip getBlockByNumber to the tip leader (#14).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@snowkide snowkide closed this Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant