System Design
B2

Load balancing: when, which kind, which algorithm

Add a load balancer when you run two or more copies of a service. Choose L4 or L7 by what it must read, and choose the algorithm by how much the servers differ.

Not startedSaved in this browser only.
  1. 1One instance needs no balancer. Two or more stateless copies need one, with health checks and draining.
  2. 2L4 picks per connection and reads no HTTP. L7 ends TLS and picks per request by path, header or cookie.
  3. 3Pick by live load, not by turn: least connections or two random choices. Round robin fails when servers differ.
  4. 4The balancer fails too. Run two or more, and add DNS or anycast for regions.
B2
    A

    When to add one, and which kind

    decision flow
    question: no ↓ · yes →what to addIs one instance enough?load fits, and a restart may cause downtimeyesNo load balancerDNS points at the one instanceDoes an instance keep user state?sessions, carts, uploads in its memory or diskyesMove the state out firstsessions to Redis; files to blob storageMust you survive a balancer failure?one balancer is a single point of failureyesRun 2 or more balancersfloating IP pair, or managed across zonesUsers in many regions, or a region to lose?latency from far away, or a whole region downyesAdd global routingDNS with health checks (GSLB), or anycastMillions of connections, or not HTTP?raw TCP or UDP, TLS passthrough, a static IPyesPut an L4 balancer in frontpicks per connection, reads no HTTPDo your services call each other?east-west traffic inside the clusteryesBalance inside, per callerclient-side library or a sidecar proxynonothenAdd an L7 load balancer2+ stateless copies · health checks · drainingthen ask, as you growquestion: yes ↳ what to add · no ↓Is one instance enough?load fits, and a restart may cause downtimeyesNo load balancerDNS points at the one instanceDoes an instance keep user state?sessions, carts, uploads in its memory or diskyesMove the state out firstsessions to Redis; files to blob storageMust you survive a balancer failure?one balancer is a single point of failureyesRun 2 or more balancersfloating IP pair, or managed across zonesUsers in many regions, or a region to lose?latency from far away, or a whole region downyesAdd global routingDNS with health checks (GSLB), or anycastMillions of connections, or not HTTP?raw TCP or UDP, TLS passthrough, a static IPyesPut an L4 balancer in frontpicks per connection, reads no HTTPDo your services call each other?east-west traffic inside the clusteryesBalance inside, per callerclient-side library or a sidecar proxynonothenAdd an L7 load balancer2+ stateless copies · health checksthen ask, as you grownonono

    I add a balancer when I run a second copy, for load or for deploys without downtime. First I make the copies stateless.

    B

    L4 or L7

    what each can read and do
    L4 readsone pick per connectionIPsrc, dstTCPportsTLSencryptedHTTPmethod, path, headers, cookiesbodyL7 readsends TLS, then one pick per requestAn L4 balancer sees encrypted bytes after the TCP header, so it cannot route on a path.
    capabilityL4 load balancerL7 proxy
    BalancesConnectionsRequests
    End TLS, hold the certificateOptionalYes
    Route by path or hostNoYes
    Route by header or cookieNoYes
    Retry a failed read elsewhereNoYes
    WebSocketsBytes passUpgrade
    Any TCP or UDP protocolYesHTTP, gRPC
    Cost per requestLowerHigher: parses HTTP

    L4 is cheap and protocol-blind. L7 reads the request, so it can route, retry and end TLS.

    C

    Where it sits

    edge, internal, in the caller
    Edgeone publicentryUsersEdge balancerTLS, routesWeb service × NInternalone IPper serviceService AInternal balancerone extra hopService B × NIn the callersidecar orlibraryService A + balancerbackend list from service discoveryService B × N
    • Edge: one public entry point. It ends TLS and routes by path.
    • Internal: one virtual IP per service. Simple, but every call takes one extra hop.
    • In the caller: a client library or a sidecar proxy reads the backend list from service discovery. No extra hop.
    • Each caller sees only its own traffic, so it balances with partial information.

    At the edge I use a managed L7 balancer. Between services I balance in the caller, so there is no extra hop.

    D

    The request path

    click a step; its path lights up
    Clientbrowser or appDNShealth-checkedL4 balancerper connectionL7 proxyTLS, routes, picksService × Nstateless copiesService, drainingno new requestsRedissessions

    Step 1: Resolve

    • DNS returns the address of a healthy region, often the nearest one.
    • With anycast, every site announces the same IP. The network delivers to the nearest site.

    If it fails

    A region fails its health checks: DNS stops returning it. Clients that cached the old answer keep failing until the TTL ends.

    DNS picks the region, L4 spreads connections, L7 picks a backend per request, and health checks and draining keep dead copies out.

    E

    Capabilities used

    what each tool gives you
    toolcapabilitywhat it gives this designalso used for
    L4 load balancerFlow hash on addresses and portsOne pick per connection. It reads no payload, so one managed L4 balancer handles millions of requests a second.Databases, MQTT, game servers
    L4 load balancerStatic IP, TLS passthroughClients can allow-list one IP. The certificate stays on the backends.Partner integrations
    L7 proxyTLS terminationOne place holds certificates. The proxy can then read the request.HTTP/2 to clients, HTTP/1.1 to backends
    L7 proxyRouting by path, host, header, cookieSend /api and /static to different pools. Send 5% of users to a canary.Blue-green deploys, A/B tests
    L7 proxyRetry an unanswered read on another backendA crashed backend costs no client error. The lab proxy lost 0 requests when one of 3 backends died.Hedged requests
    L7 proxyActive and passive health checksProbes remove a bad backend in interval × failures. Errors on live traffic remove it at once.Outlier detection
    L7 proxyConnection drainingA deploy finishes the requests in flight before it stops a copy.Scale-in, node upgrades
    L7 proxyIn-flight count per backendLeast connections and two random choices read it. They adapt to slow servers and bursts.Concurrency limits
    DNSHealth-checked, weighted or latency recordsSend users to a healthy, near region. Failover waits for the TTL in client caches.Blue-green across regions
    AnycastOne IP announced from many sites (BGP)The network delivers to the nearest site. A failed site withdraws its route.Public DNS resolvers, CDNs
    Client libraryBackend list from service discoveryThe caller picks the backend. No extra hop, no central balancer to fail.gRPC, service mesh sidecars
    RedisShared session store with TTLCopies stay stateless, so no sticky sessions are necessary.Carts, rate limits

    The L7 proxy gives me per-request routing, retries, health checks and draining. DNS and anycast give me regions. Redis lets any copy serve any user.

    F

    Six ways to pick

    what each reads, when to use it
    algorithmreadsstatus
    Round robinNothingEqual servers
    Weighted round robinFixed weightsSteady traffic
    Least connectionsIn-flight countsApproved
    Least response timeAverage latency × in flightFresh averages
    Two random choicesIn-flight of 2 samplesApproved
    Consistent hashingA key: user, cache keyCache affinity
    RandomNothingEqual servers
    Source IP hashClient IPNot approved

    Source IP hash sends every user behind one corporate NAT to one server. A server that gets slow keeps its full share under any blind algorithm.

    • Servers differ in speed, or requests differ in cost: least connections.
    • Many balancers, each with its own view: two random choices.
    • Servers cache by key: consistent hashing, with bounded loads.
    • Equal servers, equal requests, one balancer: round robin is enough.

    My default is least connections, or two random choices when there are many balancers. I use consistent hashing only for cache affinity.

    G

    Pick, in pseudo code

    pseudo code
    four pickerspseudo code
    least_connections():           // ties: rotate the start
      RETURN the up backend with the fewest requests in flight1
    
    two_random_choices():
      a, b = two different up backends, at random2
      RETURN the one of a, b with fewer requests in flight
    
    weighted_round_robin():                // smooth, as in nginx
      FOR EACH up backend: current += weight
      best = the backend with the largest current
      best.current -= sum of all weights3
      RETURN best
    
    consistent_hash(key):
      FOR EACH point clockwise4 from hash(key) on the ring:
        IF the point's backend is up: RETURN it
    1. 1The count is local to one balancer. With 3 balancers, each one sees only its own third of the traffic.
    2. 2Random samples stop many balancers from all choosing the same idle server at the same moment.
    3. 3Weights 5, 1, 1 give a a b a c a a. The heavy server never gets 5 requests in a row.
    4. 4A backend that is down passes its keys to the next point. Only its own keys move.
    Tested source Go: least connections · Go: two random choices · Go: weighted round robin · Go: consistent hashing
    Go: least connectionsgo
    // LeastConnections picks the backend with the fewest requests in flight. The scan starts one
    // place later on every pick, so ties spread over the backends.
    type LeastConnections struct{ start int }
    
    // Pick implements Picker.
    func (p *LeastConnections) Pick(v View, _ uint64) (int, bool) {
      return p.min(v, func(i int) int64 { return int64(v.InFlight(i)) })
    }
    
    func (p *LeastConnections) min(v View, score func(i int) int64) (int, bool) {
      n := v.Len()
      best, bestScore := -1, int64(0)
      for k := range n {
        i := (p.start + k) % n
        if !v.Up(i) {
          continue
        }
        if s := score(i); best < 0 || s < bestScore {
          best, bestScore = i, s
        }
      }
      p.start = (p.start + 1) % n
      return best, best >= 0
    }
    Go: two random choicesgo
    // PowerOfTwo samples two different backends at random and sends the request to the one with fewer
    // requests in flight.
    type PowerOfTwo struct {
      rng *rand.Rand
      up  []int
    }
    
    // NewPowerOfTwo returns a power-of-two-choices picker with a seeded random source.
    func NewPowerOfTwo(seed uint64) *PowerOfTwo {
      return &PowerOfTwo{rng: rand.New(rand.NewPCG(seed, 0x9e3779b97f4a7c15))}
    }
    
    // Pick implements Picker.
    func (p *PowerOfTwo) Pick(v View, _ uint64) (int, bool) {
      p.up = p.up[:0]
      for i := range v.Len() {
        if v.Up(i) {
          p.up = append(p.up, i)
        }
      }
      switch len(p.up) {
      case 0:
        return 0, false
      case 1:
        return p.up[0], true
      }
      a := p.rng.IntN(len(p.up))
      b := p.rng.IntN(len(p.up) - 1)
      if b >= a {
        b++
      }
      if v.InFlight(p.up[b]) < v.InFlight(p.up[a]) {
        return p.up[b], true
      }
      return p.up[a], true
    }
    Go: weighted round robingo
    // WeightedRoundRobin is the smooth weighted round robin of nginx: over one cycle each backend gets
    // requests in proportion to its weight, and the picks of a heavy backend are spread out.
    type WeightedRoundRobin struct{ current []int }
    
    // Pick implements Picker.
    func (p *WeightedRoundRobin) Pick(v View, _ uint64) (int, bool) {
      if len(p.current) != v.Len() {
        p.current = make([]int, v.Len())
      }
      total, best := 0, -1
      for i := range v.Len() {
        if !v.Up(i) {
          continue
        }
        p.current[i] += v.Weight(i)
        total += v.Weight(i)
        if best < 0 || p.current[i] > p.current[best] {
          best = i
        }
      }
      if best < 0 {
        return 0, false
      }
      p.current[best] -= total
      return best, true
    }
    Go: consistent hashinggo
    // Pick implements Picker. A backend that is down is skipped; its keys go to the next backend
    // clockwise, and only those keys move.
    func (c *ConsistentHash) Pick(v View, key uint64) (int, bool) {
      h := Mix(key)
      at, _ := slices.BinarySearchFunc(c.ring, h, func(p point, h uint64) int {
        switch {
        case p.hash < h:
          return -1
        case p.hash > h:
          return 1
        }
        return 0
      })
      for k := range len(c.ring) {
        p := c.ring[(at+k)%len(c.ring)]
        if p.backend < v.Len() && v.Up(p.backend) {
          return p.backend, true
        }
      }
      return 0, false
    }

    Least connections reads a live count. Two random choices reads two counts. Consistent hashing reads only the key.

    H

    Try it: six algorithms, one cluster

    recorded from a seeded simulation
    traffic
    algorithm

    Steady: 400 a second, bursts of 800. Speeds 4, 4, 2, 2, 2, 2, 1, 1. A cache hit costs 10 ms at speed 1, a miss 30 ms.

    load = work sent ÷ time · longest queuefull02S1speed 40.35q 9S2speed 40.35q 9S3speed 20.70q 44S4speed 20.69q 49S5speed 20.71q 46S6speed 20.72q 34S7speed 11.40q 1096S8speed 11.42q 1148
    p99 latency, every algorithm (log scale)10 ms100 ms1.0 s10.0 s100.0 sRound robin24.8 sWeighted round robin382 msLeast connections335 msLeast response time332 msTwo random choices412 msConsistent hashing311 ms
    p50 34 ms · p99 24.8 s · p99.9 25.6 s · cache hits 29% · longest queue 1148 on S8S7 and S8 get more work than they can do. Their queues grow for the whole run.
    arrivals per 250 ms219S1S1 at 0.00 s: 0 in flightS1 at 0.25 s: 0 in flightS1 at 0.50 s: 2 in flightS1 at 0.75 s: 0 in flightS1 at 1.00 s: 3 in flightS1 at 1.25 s: 0 in flightS1 at 1.50 s: 0 in flightS1 at 1.75 s: 1 in flightS1 at 2.00 s: 1 in flightS1 at 2.25 s: 0 in flightS1 at 2.50 s: 1 in flightS1 at 2.75 s: 0 in flightS1 at 3.00 s: 0 in flightS1 at 3.25 s: 0 in flightS1 at 3.50 s: 0 in flightS1 at 3.75 s: 1 in flightS1 at 4.00 s: 1 in flightS1 at 4.25 s: 0 in flightS1 at 4.50 s: 0 in flightS1 at 4.75 s: 0 in flightS1 at 5.00 s: 0 in flightS1 at 5.25 s: 1 in flightS1 at 5.50 s: 2 in flightS1 at 5.75 s: 1 in flightS1 at 6.00 s: 1 in flightS1 at 6.25 s: 0 in flightS1 at 6.50 s: 0 in flightS1 at 6.75 s: 0 in flightS1 at 7.00 s: 0 in flightS1 at 7.25 s: 0 in flightS1 at 7.50 s: 0 in flightS1 at 7.75 s: 1 in flightS1 at 8.00 s: 0 in flightS1 at 8.25 s: 0 in flightS1 at 8.50 s: 0 in flightS1 at 8.75 s: 0 in flightS1 at 9.00 s: 0 in flightS1 at 9.25 s: 0 in flightS1 at 9.50 s: 0 in flightS1 at 9.75 s: 0 in flightS2S2 at 0.00 s: 0 in flightS2 at 0.25 s: 2 in flightS2 at 0.50 s: 2 in flightS2 at 0.75 s: 3 in flightS2 at 1.00 s: 1 in flightS2 at 1.25 s: 0 in flightS2 at 1.50 s: 0 in flightS2 at 1.75 s: 0 in flightS2 at 2.00 s: 0 in flightS2 at 2.25 s: 1 in flightS2 at 2.50 s: 1 in flightS2 at 2.75 s: 0 in flightS2 at 3.00 s: 0 in flightS2 at 3.25 s: 0 in flightS2 at 3.50 s: 0 in flightS2 at 3.75 s: 0 in flightS2 at 4.00 s: 0 in flightS2 at 4.25 s: 0 in flightS2 at 4.50 s: 1 in flightS2 at 4.75 s: 0 in flightS2 at 5.00 s: 0 in flightS2 at 5.25 s: 1 in flightS2 at 5.50 s: 4 in flightS2 at 5.75 s: 1 in flightS2 at 6.00 s: 1 in flightS2 at 6.25 s: 1 in flightS2 at 6.50 s: 0 in flightS2 at 6.75 s: 1 in flightS2 at 7.00 s: 0 in flightS2 at 7.25 s: 1 in flightS2 at 7.50 s: 0 in flightS2 at 7.75 s: 2 in flightS2 at 8.00 s: 0 in flightS2 at 8.25 s: 0 in flightS2 at 8.50 s: 0 in flightS2 at 8.75 s: 0 in flightS2 at 9.00 s: 0 in flightS2 at 9.25 s: 0 in flightS2 at 9.50 s: 1 in flightS2 at 9.75 s: 0 in flightS3S3 at 0.00 s: 0 in flightS3 at 0.25 s: 19 in flightS3 at 0.50 s: 31 in flightS3 at 0.75 s: 38 in flightS3 at 1.00 s: 42 in flightS3 at 1.25 s: 35 in flightS3 at 1.50 s: 29 in flightS3 at 1.75 s: 25 in flightS3 at 2.00 s: 21 in flightS3 at 2.25 s: 18 in flightS3 at 2.50 s: 14 in flightS3 at 2.75 s: 0 in flightS3 at 3.00 s: 1 in flightS3 at 3.25 s: 1 in flightS3 at 3.50 s: 0 in flightS3 at 3.75 s: 4 in flightS3 at 4.00 s: 0 in flightS3 at 4.25 s: 2 in flightS3 at 4.50 s: 1 in flightS3 at 4.75 s: 1 in flightS3 at 5.00 s: 1 in flightS3 at 5.25 s: 8 in flightS3 at 5.50 s: 17 in flightS3 at 5.75 s: 22 in flightS3 at 6.00 s: 38 in flightS3 at 6.25 s: 28 in flightS3 at 6.50 s: 18 in flightS3 at 6.75 s: 13 in flightS3 at 7.00 s: 10 in flightS3 at 7.25 s: 5 in flightS3 at 7.50 s: 6 in flightS3 at 7.75 s: 1 in flightS3 at 8.00 s: 2 in flightS3 at 8.25 s: 2 in flightS3 at 8.50 s: 0 in flightS3 at 8.75 s: 1 in flightS3 at 9.00 s: 3 in flightS3 at 9.25 s: 1 in flightS3 at 9.50 s: 1 in flightS3 at 9.75 s: 1 in flightS4S4 at 0.00 s: 0 in flightS4 at 0.25 s: 7 in flightS4 at 0.50 s: 10 in flightS4 at 0.75 s: 16 in flightS4 at 1.00 s: 26 in flightS4 at 1.25 s: 15 in flightS4 at 1.50 s: 13 in flightS4 at 1.75 s: 5 in flightS4 at 2.00 s: 0 in flightS4 at 2.25 s: 1 in flightS4 at 2.50 s: 2 in flightS4 at 2.75 s: 1 in flightS4 at 3.00 s: 1 in flightS4 at 3.25 s: 1 in flightS4 at 3.50 s: 1 in flightS4 at 3.75 s: 3 in flightS4 at 4.00 s: 2 in flightS4 at 4.25 s: 0 in flightS4 at 4.50 s: 1 in flightS4 at 4.75 s: 2 in flightS4 at 5.00 s: 1 in flightS4 at 5.25 s: 5 in flightS4 at 5.50 s: 20 in flightS4 at 5.75 s: 33 in flightS4 at 6.00 s: 43 in flightS4 at 6.25 s: 46 in flightS4 at 6.50 s: 40 in flightS4 at 6.75 s: 29 in flightS4 at 7.00 s: 29 in flightS4 at 7.25 s: 28 in flightS4 at 7.50 s: 20 in flightS4 at 7.75 s: 13 in flightS4 at 8.00 s: 7 in flightS4 at 8.25 s: 1 in flightS4 at 8.50 s: 0 in flightS4 at 8.75 s: 2 in flightS4 at 9.00 s: 0 in flightS4 at 9.25 s: 1 in flightS4 at 9.50 s: 0 in flightS4 at 9.75 s: 2 in flightS5S5 at 0.00 s: 0 in flightS5 at 0.25 s: 15 in flightS5 at 0.50 s: 26 in flightS5 at 0.75 s: 38 in flightS5 at 1.00 s: 45 in flightS5 at 1.25 s: 41 in flightS5 at 1.50 s: 36 in flightS5 at 1.75 s: 35 in flightS5 at 2.00 s: 34 in flightS5 at 2.25 s: 33 in flightS5 at 2.50 s: 29 in flightS5 at 2.75 s: 25 in flightS5 at 3.00 s: 20 in flightS5 at 3.25 s: 9 in flightS5 at 3.50 s: 13 in flightS5 at 3.75 s: 5 in flightS5 at 4.00 s: 1 in flightS5 at 4.25 s: 0 in flightS5 at 4.50 s: 1 in flightS5 at 4.75 s: 1 in flightS5 at 5.00 s: 0 in flightS5 at 5.25 s: 6 in flightS5 at 5.50 s: 8 in flightS5 at 5.75 s: 14 in flightS5 at 6.00 s: 21 in flightS5 at 6.25 s: 22 in flightS5 at 6.50 s: 16 in flightS5 at 6.75 s: 13 in flightS5 at 7.00 s: 11 in flightS5 at 7.25 s: 6 in flightS5 at 7.50 s: 8 in flightS5 at 7.75 s: 0 in flightS5 at 8.00 s: 1 in flightS5 at 8.25 s: 1 in flightS5 at 8.50 s: 1 in flightS5 at 8.75 s: 0 in flightS5 at 9.00 s: 0 in flightS5 at 9.25 s: 1 in flightS5 at 9.50 s: 2 in flightS5 at 9.75 s: 1 in flightS6S6 at 0.00 s: 0 in flightS6 at 0.25 s: 6 in flightS6 at 0.50 s: 12 in flightS6 at 0.75 s: 16 in flightS6 at 1.00 s: 28 in flightS6 at 1.25 s: 24 in flightS6 at 1.50 s: 21 in flightS6 at 1.75 s: 13 in flightS6 at 2.00 s: 7 in flightS6 at 2.25 s: 2 in flightS6 at 2.50 s: 3 in flightS6 at 2.75 s: 0 in flightS6 at 3.00 s: 1 in flightS6 at 3.25 s: 1 in flightS6 at 3.50 s: 0 in flightS6 at 3.75 s: 2 in flightS6 at 4.00 s: 3 in flightS6 at 4.25 s: 2 in flightS6 at 4.50 s: 3 in flightS6 at 4.75 s: 0 in flightS6 at 5.00 s: 2 in flightS6 at 5.25 s: 9 in flightS6 at 5.50 s: 15 in flightS6 at 5.75 s: 22 in flightS6 at 6.00 s: 30 in flightS6 at 6.25 s: 27 in flightS6 at 6.50 s: 18 in flightS6 at 6.75 s: 12 in flightS6 at 7.00 s: 8 in flightS6 at 7.25 s: 3 in flightS6 at 7.50 s: 1 in flightS6 at 7.75 s: 1 in flightS6 at 8.00 s: 3 in flightS6 at 8.25 s: 0 in flightS6 at 8.50 s: 0 in flightS6 at 8.75 s: 1 in flightS6 at 9.00 s: 0 in flightS6 at 9.25 s: 1 in flightS6 at 9.50 s: 2 in flightS6 at 9.75 s: 0 in flightS7S7 at 0.00 s: 0 in flightS7 at 0.25 s: 17 in flightS7 at 0.50 s: 32 in flightS7 at 0.75 s: 50 in flightS7 at 1.00 s: 70 in flightS7 at 1.25 s: 70 in flightS7 at 1.50 s: 69 in flightS7 at 1.75 s: 71 in flightS7 at 2.00 s: 72 in flightS7 at 2.25 s: 72 in flightS7 at 2.50 s: 82 in flightS7 at 2.75 s: 87 in flightS7 at 3.00 s: 87 in flightS7 at 3.25 s: 89 in flightS7 at 3.50 s: 85 in flightS7 at 3.75 s: 83 in flightS7 at 4.00 s: 86 in flightS7 at 4.25 s: 87 in flightS7 at 4.50 s: 95 in flightS7 at 4.75 s: 100 in flightS7 at 5.00 s: 99 in flightS7 at 5.25 s: 114 in flightS7 at 5.50 s: 131 in flightS7 at 5.75 s: 142 in flightS7 at 6.00 s: 157 in flightS7 at 6.25 s: 162 in flightS7 at 6.50 s: 164 in flightS7 at 6.75 s: 163 in flightS7 at 7.00 s: 169 in flightS7 at 7.25 s: 169 in flightS7 at 7.50 s: 169 in flightS7 at 7.75 s: 168 in flightS7 at 8.00 s: 174 in flightS7 at 8.25 s: 177 in flightS7 at 8.50 s: 181 in flightS7 at 8.75 s: 181 in flightS7 at 9.00 s: 192 in flightS7 at 9.25 s: 192 in flightS7 at 9.50 s: 198 in flightS7 at 9.75 s: 195 in flightS8S8 at 0.00 s: 0 in flightS8 at 0.25 s: 17 in flightS8 at 0.50 s: 33 in flightS8 at 0.75 s: 46 in flightS8 at 1.00 s: 65 in flightS8 at 1.25 s: 70 in flightS8 at 1.50 s: 78 in flightS8 at 1.75 s: 86 in flightS8 at 2.00 s: 91 in flightS8 at 2.25 s: 97 in flightS8 at 2.50 s: 105 in flightS8 at 2.75 s: 109 in flightS8 at 3.00 s: 109 in flightS8 at 3.25 s: 111 in flightS8 at 3.50 s: 118 in flightS8 at 3.75 s: 121 in flightS8 at 4.00 s: 121 in flightS8 at 4.25 s: 125 in flightS8 at 4.50 s: 131 in flightS8 at 4.75 s: 131 in flightS8 at 5.00 s: 132 in flightS8 at 5.25 s: 149 in flightS8 at 5.50 s: 167 in flightS8 at 5.75 s: 181 in flightS8 at 6.00 s: 198 in flightS8 at 6.25 s: 208 in flightS8 at 6.50 s: 209 in flightS8 at 6.75 s: 216 in flightS8 at 7.00 s: 220 in flightS8 at 7.25 s: 230 in flightS8 at 7.50 s: 237 in flightS8 at 7.75 s: 239 in flightS8 at 8.00 s: 245 in flightS8 at 8.25 s: 249 in flightS8 at 8.50 s: 249 in flightS8 at 8.75 s: 249 in flightS8 at 9.00 s: 254 in flightS8 at 9.25 s: 257 in flightS8 at 9.50 s: 267 in flightS8 at 9.75 s: 265 in flight0 s1 s2 s3 s4 s5 s6 s7 s8 s9 s10 s
    0 1 10 100+requests in flight per server, sampled every 250 ms. Bursts run from 0 to 1 s and from 5 to 6 s.

    Round robin ignores speed, so slow servers fall behind for good. Least connections follows the live queue. Consistent hashing wins the cache but loses to a hot key.

    I

    Health checks, retries, draining

    pseudo code
    keep dead copies outpseudo code
    every interval:                        // active health check
      FOR EACH backend:
        ok = GET /healthz returns 200 before the timeout
        IF NOT ok 'fall' times in a row1: mark it down
        IF ok 'rise' times in a row: mark it up
    
    on a request error, before any response:   // passive check
      mark the backend down
      IF the request is a read with no body2:
        send it to the next backend
      ELSE: RETURN 502                         // never repeat a write3
    
    drain(backend):                        // before a deploy or a scale-in
      send no new requests to it
      WAIT until its requests in flight = 04
      THEN stop or remove it
    1. 1One lost probe is noise. With a 30 s interval and 2 failures, removal takes about 60 s.
    2. 2GET, HEAD and OPTIONS change nothing, so a second attempt is safe. The lab proxy lost 0 requests when a backend died.
    3. 3The first backend may have charged the card before it died. Return the error; the client retries with an idempotency key.
    4. 4Set the longest wait above the slowest request. One managed balancer waits 300 s by default.
    Tested source Go: health check · Go: retry a read on another backend · Go: drain
    Go: health checkgo
    // HealthCheck probes every backend's /healthz each interval until ctx ends. A backend goes down
    // after fall failed probes in a row, and up again after rise good probes in a row.
    func (p *Proxy) HealthCheck(ctx context.Context, interval time.Duration, fall, rise int) {
      client := &http.Client{Transport: p.transport, Timeout: interval}
      bad := make([]int, len(p.backends))
      good := make([]int, len(p.backends))
      t := time.NewTicker(interval)
      defer t.Stop()
      for {
        select {
        case <-ctx.Done():
          return
        case <-t.C:
        }
        for i, b := range p.backends {
          if p.probe(ctx, client, b) {
            bad[i], good[i] = 0, good[i]+1
            if good[i] >= rise && b.down.Swap(false) {
              p.log.Info("backend up", "backend", b.URL.Host)
            }
            continue
          }
          good[i], bad[i] = 0, bad[i]+1
          if bad[i] >= fall && !b.down.Swap(true) {
            p.log.Warn("backend down", "backend", b.URL.Host)
          }
        }
      }
    }
    Go: retry a read on another backendgo
    // RoundTrip sends one request to a backend. When the backend cannot be reached and the request is
    // safe to repeat, it marks the backend down (a passive health check) and tries the next one.
    func (p *Proxy) RoundTrip(req *http.Request) (*http.Response, error) {
      v := view{b: p.backends, tried: make([]bool, len(p.backends))}
      key := routeKey(req)
      for {
        p.mu.Lock()
        i, ok := p.picker.Pick(v, key)
        p.mu.Unlock()
        if !ok {
          return nil, ErrNoBackend
        }
        b := p.backends[i]
        out := req.Clone(req.Context())
        out.URL.Scheme, out.URL.Host, out.Host = b.URL.Scheme, b.URL.Host, b.URL.Host
        b.inFlight.Add(1)
        start := time.Now()
        resp, err := p.transport.RoundTrip(out)
        if err != nil {
          b.inFlight.Add(-1)
          if !idempotent(req) || req.Context().Err() != nil {
            return nil, fmt.Errorf("backend %s: %w", b.URL.Host, err)
          }
          p.log.Warn("backend unreachable; marking it down and retrying", "backend", b.URL.Host, "err", err)
          b.down.Store(true)
          v.tried[i] = true
          continue
        }
        resp.Body = &trackedBody{ReadCloser: resp.Body, done: func() {
          b.inFlight.Add(-1)
          b.served.Add(1)
          l := time.Since(start).Microseconds()
          old := b.ewma.Load()
          b.ewma.Store(old + (l-old)/8)
        }}
        return resp, nil
      }
    }
    Go: draingo
    // Drain stops new requests to backend i and waits until its requests in flight finish. Remove
    // the backend, or restart it, after Drain returns.
    func (p *Proxy) Drain(ctx context.Context, i int) error {
      b := p.backends[i]
      b.draining.Store(true)
      t := time.NewTicker(5 * time.Millisecond)
      defer t.Stop()
      for b.inFlight.Load() > 0 {
        select {
        case <-ctx.Done():
          return fmt.Errorf("drain %s: %d requests still in flight: %w", b.URL.Host, b.inFlight.Load(), ctx.Err())
        case <-t.C:
        }
      }
      return nil
    }

    Active checks remove a dead backend within interval times failures. Passive checks and a retry of reads cover the gap. Drain before every restart.

    J

    Consistent hashing

    for cache affinity
    ABCDADBCADBCEEk1k2k3k4hash spaceclockwise →Add server Ek1: C → Ek2: D, staysk3: A → Ek4: C, staysarcs that move to EOnly keys on these arcs move:about 1/n of all keys.hash mod n moves aboutn/(n+1) of all keys.
    • Use it when a server caches by key: sessions, a local cache, a shard of data.
    • Each server takes many points (virtual nodes), in proportion to its weight. Shares then stay within about 30% of the weight.
    • It ignores load. A hot key overloads its one server. Cap each server's share (bounded loads), or copy the hot key.

    Consistent hashing keeps each key on one server, so its cache stays warm. Adding a server moves about one key in n.

    K

    Sticky sessions

    and why to avoid them
    where the session liveswherestatus
    In server memory, no stickinessServiceNot approved
    In server memory, sticky cookieL7 proxyStopgap
    Source IP stickinessL4 load balancerNot approved
    Shared store, key from a cookieRedisApproved
    Signed token in the clientServiceSmall state
    • A sticky user stays on one server even when that server is busy.
    • A failed server loses every session it held.
    • Scale-in and deploys break sessions, unless you drain for the full session length.
    • WebSockets are sticky by nature: one connection stays on one backend.

    I keep sessions out of the server, in Redis or a signed token, so I do not need sticky sessions.

    L

    The balancer fails too

    no single point of failure
    Active and standbyone floating IP, moved on failureIP 203.0.113.10Balancer 1active, owns IPBalancer 2standbyheartbeatBalancer 1 stops: balancer 2 takes the IP.Half the capacity sits idle.Open connections break; clients reconnect.All activeseveral IPs in DNS, or anycastapi.example.comLB 1zone aLB 2zone bLB 3zone cOne stops: health checks remove its IP,or its route is withdrawn (anycast).Clients that cache DNS past the TTL still fail.
    • A managed balancer runs nodes in every zone you enable. A failed node leaves DNS.
    • Self-run: two proxies share a floating IP through a heartbeat (for example VRRP).
    • Across regions: DNS health checks or anycast route withdrawal.

    I run two or more balancers: a floating IP pair, or a managed balancer in each zone, with DNS or anycast in front.

    M

    Failure cases

    what breaks, and what saves you
    eventresultwhy it stays safesaved by
    A backend crashesNew connections to it are refused.The proxy marks it down and sends reads to the next backend. Clients see no error.Retry, passive check
    A backend hangs but accepts connectionsRequests to it time out.Probes time out too. After N failures in a row it leaves the pool.Active check
    A backend gets slowIts queue grows.Least connections and two random choices send it fewer new requests.Least connections
    Every backend fails its checkThe pool is empty.One managed balancer then sends to all targets anyway (fail open). Keep checks shallow.Fail open
    A deploy restarts a copyIts requests in flight would break.Draining stops new requests and waits for the old ones.Draining
    A balancer node diesIts open connections break.The standby takes the floating IP, or DNS drops the node. Clients reconnect.Redundant balancers
    A region goes downIts users get errors.DNS health checks stop returning it. Clients that ignore the TTL keep failing.DNS
    One key gets 30% of requestsUnder consistent hashing, one server queues.Bounded loads move the overflow to the next server. Or copy the hot key.Bounded loads
    A user's session server diesWith sticky sessions, the user logs in again.With sessions in Redis, any copy serves the next request.Redis
    N

    Scale ladder

    start simple; climb only on a signal
    Each step adds one component1One instance2+ L7 balancer3+ 2nd balancer4+ DNS, anycast5+ in the callermore load →
    Capacity against demand1001k10k100k1M10MDemand, average: 1,157 small requests per secondDemand, average1,157Demand, 3× peak: 3,472 small requests per secondDemand, 3× peak3,472Go L7 proxy, lab: 25,000 small requests per secondGo L7 proxy, lab25,000Go L4 proxy, lab: 61,000 small requests per secondGo L4 proxy, lab61,000Managed L4, docs: millionsManaged L4, docsmillionssmall requests per second, log scale
    stepaddit handlesmove up when you see
    1One instance. DNS points at it.What one machine can serve. A restart is a short outage.You need a second copy, for load or for deploys without downtime.
    2An L7 balancer and 2 or more stateless copies. Sessions in Redis. Health checks and draining.A small Go proxy forwarded about 25,000 small requests a second in the lab. A managed one scales out by itself.The balancer is now the single point of failure.
    3Redundant balancers: a floating IP pair, or a managed balancer across zones.One balancer node or one zone can fail.Users far away see high latency, or the target says survive a region outage.
    4Global routing: DNS with health checks, or anycast, over 2 or more regions.A whole region can fail. Users reach a near region.Many services call each other, and the central balancer adds a hop to every call.
    5Balance in the caller: a client library or a sidecar per service. L4 in front of the L7 tier.East-west traffic with no central hop. L4 absorbs connection volume.Top of the ladder.

    Demand example: 100 million requests a day ÷ 86,400 s = 1,157 a second. Lab numbers come from one laptop that also ran the clients, so read them as orders of magnitude. Do not start at step 3. Each step adds a component to run.

    I start with one instance and no balancer. I add an L7 balancer when I need a second copy, and redundancy and regions only when the availability target asks for them.

    O

    Drill

    predict, then reveal

    0 of 9 known

    1. You run one instance and want deploys without downtime. Do you need a load balancer?

    2. Round robin over servers of speed 4, 4, 2, 2, 2, 2, 1, 1. What happens at about 475 requests a second?

    3. Weights already match the server speeds. Why does least connections still win at the tail?

    4. You run 10 balancers. Why use two random choices instead of least connections?

    5. gRPC clients keep one HTTP/2 connection open. Why is the load uneven behind an L4 balancer?

    6. A ring has 8 servers. You add a 9th with the same weight. How many keys move? With hash mod n?

    7. Why avoid sticky sessions?

    8. With the defaults of one managed L7 balancer, how long until a dead server leaves the pool?

    9. The health check opens a database connection. The database is down. What happens?

    P

    Numbers to say

    measured, derived, cited
    detect
    One managed L7 balancer: a check every 30 s, 2 failures, so about 60 s. Kubernetes readiness: 10 s × 3 = 30 s.
    DNS
    Health checks every 30 s, or 10 s at extra cost. Failover then waits for client TTLs.
    drain
    One managed balancer waits 300 s by default.
    L4 vs L7
    Go proxies in the lab: L4 forwarded about 61,000 small requests a second, L7 about 25,000. L7 parses every request.
    managed L4
    Millions of requests a second.
    rehash
    Add 1 server to n: consistent hashing moves 1/(n+1) of keys; hash mod n moves n/(n+1).
    tail
    Simulation, steady traffic: p99 24.8 s with round robin, 335 ms with least connections.

    Managed balancer defaults and capacity are from the provider's documentation. Lab: Go 1.26 on an 8-core laptop, 64 clients on the same machine, median of 5 runs.