Skip to content

fix(network): route block-pinned requests to fallbacks that have the block - #29

Open
shpookas wants to merge 2 commits into
feat/websocket-supportfrom
fix/fallback-tip-leader
Open

shpookas wants to merge 2 commits into
feat/websocket-supportfrom
fix/fallback-tip-leader

Conversation

@shpookas

Copy link
Copy Markdown

Summary

  • Tip-leader routing: a request pinned to a block above every routed upstream's known head is tried first on the tier:fallback upstreams whose head already reached it, then on the routed list. When a fallback's newHeads leads, clients call at blocks the routed upstreams don't have yet, and today those calls fail on the routed upstreams first. Counted in erpc_network_tip_leader_route_total.
  • Hedge keeper: once a request has escalated, an all-missing ErrUpstreamsExhausted is no longer kept, so a hedge leg that only re-swept the routed upstreams can't cancel the leg still waiting on a fallback.
  • Future-block short-circuit (served-tip): no synthetic null for a block a reachable fallback already has.

Still one escalation per request. No-op for consensus, with failover off, or when a routed head is unknown or at the block. Leaders respect use-upstream and their availability bounds.

Test plan

  • TestFailover_TipLeaderRouting and related tests (15 scenarios), each failing without its fix; TestFailover_* passes under -race
  • go test ./erpc/ ./common/ ./upstream/ ./telemetry/: only TestNetwork_Forward/ForwardLlamaRPCEndpointRateLimitResponseSingle fails, identically on 05126d0
  • Canary

…block

- Tip-leader routing: a request pinned to a block above every routed
  upstream's known head is tried first on the tier:fallback upstreams whose
  head already reached it (their newHeads keep their pollers current), then
  on the routed list. It takes the per-request escalation, so other hedge
  legs and retries do not escape; the sweep that took it still escapes once
  to the fallbacks it has not tried. No-op when a routed head is unknown or
  at the block, for consensus, and with failover off. Leaders must match the
  use-upstream selector and their enforced availability bounds. Counted in
  erpc_network_tip_leader_route_total.
- Hedge keeper: once a request has escalated, an all-missing
  ErrUpstreamsExhausted is not kept, so a leg that only re-swept the routed
  upstreams cannot cancel a leg still waiting on a fallback. If every leg
  misses, the hedge returns the last result.
- Future-block short-circuit (served-tip): skip the synthetic null when a
  reachable fallback already has the block, except for consensus.
…tion

Same outcome as the previous commit, with fewer moving parts:

- Fallbacks whose polled head has the block join the request's upstream
  list, and tiering now runs before the has-the-block partition, so an
  upstream that has the block goes first whatever its tier. This is
  ordering, not an escalation: drops leaderRouted, the shared
  MarkEscalatedToFallbacks handling, the maxLoopIterations bump and
  NormalizedRequest.EscalatedToFallbacks. The once-per-request escape is
  unchanged; use-upstream is enforced by NextUpstream as for any upstream.
- The future-block short-circuit is skipped when a tip leader was added,
  instead of rescanning the fallbacks.
- The hedge keeper no longer keeps an all-missing ErrUpstreamsExhausted
  at all, rather than only after escalation: a sibling leg may be on an
  upstream this leg never tried, which already happens in the plain
  escape path with no tip leader (TestFailover_HedgeKeepsEscalatedSibling
  and TestFailover_HedgeKeepsSlowFallbackAfterFastFallbackMiss fail on
  05126d0). If every leg ends that way the hedge still returns the last
  one to the retry layer.

Tests unchanged; TestFailover_* pass, including under -race.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants