Load balancing: when, which kind, which algorithm
Add a load balancer when you run two or more copies of a service. Choose L4 or L7 by what it must read, and choose the algorithm by how much the servers differ.
- 1One instance needs no balancer. Two or more stateless copies need one, with health checks and draining.
- 2L4 picks per connection and reads no HTTP. L7 ends TLS and picks per request by path, header or cookie.
- 3Pick by live load, not by turn: least connections or two random choices. Round robin fails when servers differ.
- 4The balancer fails too. Run two or more, and add DNS or anycast for regions.
I add a balancer when I run a second copy, for load or for deploys without downtime. First I make the copies stateless.
| capability | L4 load balancer | L7 proxy |
|---|---|---|
| Balances | Connections | Requests |
| End TLS, hold the certificate | Optional | Yes |
| Route by path or host | No | Yes |
| Route by header or cookie | No | Yes |
| Retry a failed read elsewhere | No | Yes |
| WebSockets | Bytes pass | Upgrade |
| Any TCP or UDP protocol | Yes | HTTP, gRPC |
| Cost per request | Lower | Higher: parses HTTP |
L4 is cheap and protocol-blind. L7 reads the request, so it can route, retry and end TLS.
- Edge: one public entry point. It ends TLS and routes by path.
- Internal: one virtual IP per service. Simple, but every call takes one extra hop.
- In the caller: a client library or a sidecar proxy reads the backend list from service discovery. No extra hop.
- Each caller sees only its own traffic, so it balances with partial information.
At the edge I use a managed L7 balancer. Between services I balance in the caller, so there is no extra hop.
Step 1: Resolve
- DNS returns the address of a healthy region, often the nearest one.
- With anycast, every site announces the same IP. The network delivers to the nearest site.
If it fails
A region fails its health checks: DNS stops returning it. Clients that cached the old answer keep failing until the TTL ends.
DNS picks the region, L4 spreads connections, L7 picks a backend per request, and health checks and draining keep dead copies out.
| tool | capability | what it gives this design | also used for |
|---|---|---|---|
| L4 load balancer | Flow hash on addresses and ports | One pick per connection. It reads no payload, so one managed L4 balancer handles millions of requests a second. | Databases, MQTT, game servers |
| L4 load balancer | Static IP, TLS passthrough | Clients can allow-list one IP. The certificate stays on the backends. | Partner integrations |
| L7 proxy | TLS termination | One place holds certificates. The proxy can then read the request. | HTTP/2 to clients, HTTP/1.1 to backends |
| L7 proxy | Routing by path, host, header, cookie | Send /api and /static to different pools. Send 5% of users to a canary. | Blue-green deploys, A/B tests |
| L7 proxy | Retry an unanswered read on another backend | A crashed backend costs no client error. The lab proxy lost 0 requests when one of 3 backends died. | Hedged requests |
| L7 proxy | Active and passive health checks | Probes remove a bad backend in interval × failures. Errors on live traffic remove it at once. | Outlier detection |
| L7 proxy | Connection draining | A deploy finishes the requests in flight before it stops a copy. | Scale-in, node upgrades |
| L7 proxy | In-flight count per backend | Least connections and two random choices read it. They adapt to slow servers and bursts. | Concurrency limits |
| DNS | Health-checked, weighted or latency records | Send users to a healthy, near region. Failover waits for the TTL in client caches. | Blue-green across regions |
| Anycast | One IP announced from many sites (BGP) | The network delivers to the nearest site. A failed site withdraws its route. | Public DNS resolvers, CDNs |
| Client library | Backend list from service discovery | The caller picks the backend. No extra hop, no central balancer to fail. | gRPC, service mesh sidecars |
| Redis | Shared session store with TTL | Copies stay stateless, so no sticky sessions are necessary. | Carts, rate limits |
The L7 proxy gives me per-request routing, retries, health checks and draining. DNS and anycast give me regions. Redis lets any copy serve any user.
| algorithm | reads | status |
|---|---|---|
| Round robin | Nothing | Equal servers |
| Weighted round robin | Fixed weights | Steady traffic |
| Least connections | In-flight counts | Approved |
| Least response time | Average latency × in flight | Fresh averages |
| Two random choices | In-flight of 2 samples | Approved |
| Consistent hashing | A key: user, cache key | Cache affinity |
| Random | Nothing | Equal servers |
| Source IP hash | Client IP | Not approved |
Source IP hash sends every user behind one corporate NAT to one server. A server that gets slow keeps its full share under any blind algorithm.
- Servers differ in speed, or requests differ in cost: least connections.
- Many balancers, each with its own view: two random choices.
- Servers cache by key: consistent hashing, with bounded loads.
- Equal servers, equal requests, one balancer: round robin is enough.
My default is least connections, or two random choices when there are many balancers. I use consistent hashing only for cache affinity.
least_connections(): // ties: rotate the start
RETURN the up backend with the fewest requests in flight1
two_random_choices():
a, b = two different up backends, at random2
RETURN the one of a, b with fewer requests in flight
weighted_round_robin(): // smooth, as in nginx
FOR EACH up backend: current += weight
best = the backend with the largest current
best.current -= sum of all weights3
RETURN best
consistent_hash(key):
FOR EACH point clockwise4 from hash(key) on the ring:
IF the point's backend is up: RETURN it- 1The count is local to one balancer. With 3 balancers, each one sees only its own third of the traffic.
- 2Random samples stop many balancers from all choosing the same idle server at the same moment.
- 3Weights 5, 1, 1 give a a b a c a a. The heavy server never gets 5 requests in a row.
- 4A backend that is down passes its keys to the next point. Only its own keys move.
Tested source Go: least connections · Go: two random choices · Go: weighted round robin · Go: consistent hashing
// LeastConnections picks the backend with the fewest requests in flight. The scan starts one
// place later on every pick, so ties spread over the backends.
type LeastConnections struct{ start int }
// Pick implements Picker.
func (p *LeastConnections) Pick(v View, _ uint64) (int, bool) {
return p.min(v, func(i int) int64 { return int64(v.InFlight(i)) })
}
func (p *LeastConnections) min(v View, score func(i int) int64) (int, bool) {
n := v.Len()
best, bestScore := -1, int64(0)
for k := range n {
i := (p.start + k) % n
if !v.Up(i) {
continue
}
if s := score(i); best < 0 || s < bestScore {
best, bestScore = i, s
}
}
p.start = (p.start + 1) % n
return best, best >= 0
}// PowerOfTwo samples two different backends at random and sends the request to the one with fewer
// requests in flight.
type PowerOfTwo struct {
rng *rand.Rand
up []int
}
// NewPowerOfTwo returns a power-of-two-choices picker with a seeded random source.
func NewPowerOfTwo(seed uint64) *PowerOfTwo {
return &PowerOfTwo{rng: rand.New(rand.NewPCG(seed, 0x9e3779b97f4a7c15))}
}
// Pick implements Picker.
func (p *PowerOfTwo) Pick(v View, _ uint64) (int, bool) {
p.up = p.up[:0]
for i := range v.Len() {
if v.Up(i) {
p.up = append(p.up, i)
}
}
switch len(p.up) {
case 0:
return 0, false
case 1:
return p.up[0], true
}
a := p.rng.IntN(len(p.up))
b := p.rng.IntN(len(p.up) - 1)
if b >= a {
b++
}
if v.InFlight(p.up[b]) < v.InFlight(p.up[a]) {
return p.up[b], true
}
return p.up[a], true
}// WeightedRoundRobin is the smooth weighted round robin of nginx: over one cycle each backend gets
// requests in proportion to its weight, and the picks of a heavy backend are spread out.
type WeightedRoundRobin struct{ current []int }
// Pick implements Picker.
func (p *WeightedRoundRobin) Pick(v View, _ uint64) (int, bool) {
if len(p.current) != v.Len() {
p.current = make([]int, v.Len())
}
total, best := 0, -1
for i := range v.Len() {
if !v.Up(i) {
continue
}
p.current[i] += v.Weight(i)
total += v.Weight(i)
if best < 0 || p.current[i] > p.current[best] {
best = i
}
}
if best < 0 {
return 0, false
}
p.current[best] -= total
return best, true
}// Pick implements Picker. A backend that is down is skipped; its keys go to the next backend
// clockwise, and only those keys move.
func (c *ConsistentHash) Pick(v View, key uint64) (int, bool) {
h := Mix(key)
at, _ := slices.BinarySearchFunc(c.ring, h, func(p point, h uint64) int {
switch {
case p.hash < h:
return -1
case p.hash > h:
return 1
}
return 0
})
for k := range len(c.ring) {
p := c.ring[(at+k)%len(c.ring)]
if p.backend < v.Len() && v.Up(p.backend) {
return p.backend, true
}
}
return 0, false
}Least connections reads a live count. Two random choices reads two counts. Consistent hashing reads only the key.
Steady: 400 a second, bursts of 800. Speeds 4, 4, 2, 2, 2, 2, 1, 1. A cache hit costs 10 ms at speed 1, a miss 30 ms.
Round robin ignores speed, so slow servers fall behind for good. Least connections follows the live queue. Consistent hashing wins the cache but loses to a hot key.
every interval: // active health check
FOR EACH backend:
ok = GET /healthz returns 200 before the timeout
IF NOT ok 'fall' times in a row1: mark it down
IF ok 'rise' times in a row: mark it up
on a request error, before any response: // passive check
mark the backend down
IF the request is a read with no body2:
send it to the next backend
ELSE: RETURN 502 // never repeat a write3
drain(backend): // before a deploy or a scale-in
send no new requests to it
WAIT until its requests in flight = 04
THEN stop or remove it- 1One lost probe is noise. With a 30 s interval and 2 failures, removal takes about 60 s.
- 2GET, HEAD and OPTIONS change nothing, so a second attempt is safe. The lab proxy lost 0 requests when a backend died.
- 3The first backend may have charged the card before it died. Return the error; the client retries with an idempotency key.
- 4Set the longest wait above the slowest request. One managed balancer waits 300 s by default.
Tested source Go: health check · Go: retry a read on another backend · Go: drain
// HealthCheck probes every backend's /healthz each interval until ctx ends. A backend goes down
// after fall failed probes in a row, and up again after rise good probes in a row.
func (p *Proxy) HealthCheck(ctx context.Context, interval time.Duration, fall, rise int) {
client := &http.Client{Transport: p.transport, Timeout: interval}
bad := make([]int, len(p.backends))
good := make([]int, len(p.backends))
t := time.NewTicker(interval)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
}
for i, b := range p.backends {
if p.probe(ctx, client, b) {
bad[i], good[i] = 0, good[i]+1
if good[i] >= rise && b.down.Swap(false) {
p.log.Info("backend up", "backend", b.URL.Host)
}
continue
}
good[i], bad[i] = 0, bad[i]+1
if bad[i] >= fall && !b.down.Swap(true) {
p.log.Warn("backend down", "backend", b.URL.Host)
}
}
}
}// RoundTrip sends one request to a backend. When the backend cannot be reached and the request is
// safe to repeat, it marks the backend down (a passive health check) and tries the next one.
func (p *Proxy) RoundTrip(req *http.Request) (*http.Response, error) {
v := view{b: p.backends, tried: make([]bool, len(p.backends))}
key := routeKey(req)
for {
p.mu.Lock()
i, ok := p.picker.Pick(v, key)
p.mu.Unlock()
if !ok {
return nil, ErrNoBackend
}
b := p.backends[i]
out := req.Clone(req.Context())
out.URL.Scheme, out.URL.Host, out.Host = b.URL.Scheme, b.URL.Host, b.URL.Host
b.inFlight.Add(1)
start := time.Now()
resp, err := p.transport.RoundTrip(out)
if err != nil {
b.inFlight.Add(-1)
if !idempotent(req) || req.Context().Err() != nil {
return nil, fmt.Errorf("backend %s: %w", b.URL.Host, err)
}
p.log.Warn("backend unreachable; marking it down and retrying", "backend", b.URL.Host, "err", err)
b.down.Store(true)
v.tried[i] = true
continue
}
resp.Body = &trackedBody{ReadCloser: resp.Body, done: func() {
b.inFlight.Add(-1)
b.served.Add(1)
l := time.Since(start).Microseconds()
old := b.ewma.Load()
b.ewma.Store(old + (l-old)/8)
}}
return resp, nil
}
}// Drain stops new requests to backend i and waits until its requests in flight finish. Remove
// the backend, or restart it, after Drain returns.
func (p *Proxy) Drain(ctx context.Context, i int) error {
b := p.backends[i]
b.draining.Store(true)
t := time.NewTicker(5 * time.Millisecond)
defer t.Stop()
for b.inFlight.Load() > 0 {
select {
case <-ctx.Done():
return fmt.Errorf("drain %s: %d requests still in flight: %w", b.URL.Host, b.inFlight.Load(), ctx.Err())
case <-t.C:
}
}
return nil
}Active checks remove a dead backend within interval times failures. Passive checks and a retry of reads cover the gap. Drain before every restart.
- Use it when a server caches by key: sessions, a local cache, a shard of data.
- Each server takes many points (virtual nodes), in proportion to its weight. Shares then stay within about 30% of the weight.
- It ignores load. A hot key overloads its one server. Cap each server's share (bounded loads), or copy the hot key.
Consistent hashing keeps each key on one server, so its cache stays warm. Adding a server moves about one key in n.
| where the session lives | where | status |
|---|---|---|
| In server memory, no stickiness | Service | Not approved |
| In server memory, sticky cookie | L7 proxy | Stopgap |
| Source IP stickiness | L4 load balancer | Not approved |
| Shared store, key from a cookie | Redis | Approved |
| Signed token in the client | Service | Small state |
- A sticky user stays on one server even when that server is busy.
- A failed server loses every session it held.
- Scale-in and deploys break sessions, unless you drain for the full session length.
- WebSockets are sticky by nature: one connection stays on one backend.
I keep sessions out of the server, in Redis or a signed token, so I do not need sticky sessions.
- A managed balancer runs nodes in every zone you enable. A failed node leaves DNS.
- Self-run: two proxies share a floating IP through a heartbeat (for example VRRP).
- Across regions: DNS health checks or anycast route withdrawal.
I run two or more balancers: a floating IP pair, or a managed balancer in each zone, with DNS or anycast in front.
| event | result | why it stays safe | saved by |
|---|---|---|---|
| A backend crashes | New connections to it are refused. | The proxy marks it down and sends reads to the next backend. Clients see no error. | Retry, passive check |
| A backend hangs but accepts connections | Requests to it time out. | Probes time out too. After N failures in a row it leaves the pool. | Active check |
| A backend gets slow | Its queue grows. | Least connections and two random choices send it fewer new requests. | Least connections |
| Every backend fails its check | The pool is empty. | One managed balancer then sends to all targets anyway (fail open). Keep checks shallow. | Fail open |
| A deploy restarts a copy | Its requests in flight would break. | Draining stops new requests and waits for the old ones. | Draining |
| A balancer node dies | Its open connections break. | The standby takes the floating IP, or DNS drops the node. Clients reconnect. | Redundant balancers |
| A region goes down | Its users get errors. | DNS health checks stop returning it. Clients that ignore the TTL keep failing. | DNS |
| One key gets 30% of requests | Under consistent hashing, one server queues. | Bounded loads move the overflow to the next server. Or copy the hot key. | Bounded loads |
| A user's session server dies | With sticky sessions, the user logs in again. | With sessions in Redis, any copy serves the next request. | Redis |
| step | add | it handles | move up when you see |
|---|---|---|---|
| 1 | One instance. DNS points at it. | What one machine can serve. A restart is a short outage. | You need a second copy, for load or for deploys without downtime. |
| 2 | An L7 balancer and 2 or more stateless copies. Sessions in Redis. Health checks and draining. | A small Go proxy forwarded about 25,000 small requests a second in the lab. A managed one scales out by itself. | The balancer is now the single point of failure. |
| 3 | Redundant balancers: a floating IP pair, or a managed balancer across zones. | One balancer node or one zone can fail. | Users far away see high latency, or the target says survive a region outage. |
| 4 | Global routing: DNS with health checks, or anycast, over 2 or more regions. | A whole region can fail. Users reach a near region. | Many services call each other, and the central balancer adds a hop to every call. |
| 5 | Balance in the caller: a client library or a sidecar per service. L4 in front of the L7 tier. | East-west traffic with no central hop. L4 absorbs connection volume. | Top of the ladder. |
Demand example: 100 million requests a day ÷ 86,400 s = 1,157 a second. Lab numbers come from one laptop that also ran the clients, so read them as orders of magnitude. Do not start at step 3. Each step adds a component to run.
I start with one instance and no balancer. I add an L7 balancer when I need a second copy, and redundancy and regions only when the availability target asks for them.
0 of 9 known
You run one instance and want deploys without downtime. Do you need a load balancer?
Round robin over servers of speed 4, 4, 2, 2, 2, 2, 1, 1. What happens at about 475 requests a second?
Weights already match the server speeds. Why does least connections still win at the tail?
You run 10 balancers. Why use two random choices instead of least connections?
gRPC clients keep one HTTP/2 connection open. Why is the load uneven behind an L4 balancer?
A ring has 8 servers. You add a 9th with the same weight. How many keys move? With hash mod n?
Why avoid sticky sessions?
With the defaults of one managed L7 balancer, how long until a dead server leaves the pool?
The health check opens a database connection. The database is down. What happens?
- detect
- One managed L7 balancer: a check every 30 s, 2 failures, so about 60 s. Kubernetes readiness: 10 s × 3 = 30 s.
- DNS
- Health checks every 30 s, or 10 s at extra cost. Failover then waits for client TTLs.
- drain
- One managed balancer waits 300 s by default.
- L4 vs L7
- Go proxies in the lab: L4 forwarded about 61,000 small requests a second, L7 about 25,000. L7 parses every request.
- managed L4
- Millions of requests a second.
- rehash
- Add 1 server to n: consistent hashing moves 1/(n+1) of keys; hash mod n moves n/(n+1).
- tail
- Simulation, steady traffic: p99 24.8 s with round robin, 335 ms with least connections.
Managed balancer defaults and capacity are from the provider's documentation. Lab: Go 1.26 on an 8-core laptop, 64 clients on the same machine, median of 5 runs.