Resilience patterns: survive a slow or failing dependency
A dependency will get slow and then fail. These patterns keep its callers alive, keep it from overload, and let it recover when the cause is gone.
- 1Every call has a timeout, and one deadline travels with the request through every layer.
- 2Retry only idempotent calls, with backoff, full jitter and a budget. Retries without a budget keep a dependency down after its cause is gone.
- 3Breakers, bulkheads and fallbacks stop one slow dependency from taking every thread of its callers.
- 4Under overload the server sheds: bound the queue, serve the newest first, drop work whose caller left.
- A dead dependency fails fast. A slow one holds threads, connections and memory in every caller.
- Retries from the layer above add load at the worst time.
- Each pattern below cuts one link of this chain.
A slow dependency is worse than a dead one: it holds every thread of its callers, so the failure climbs the call graph.
| pattern | protects | status |
|---|---|---|
| Timeout on every call, deadline per request | The caller's threads | Approved |
| Retry with backoff, full jitter and a budget | Users, from short errors | Idempotent only |
| Retry with no backoff, or no budget | Nothing: it multiplies load | Not approved |
| Retry at every layer | Nothing: 3 layers × 3 tries = 27× | Not approved |
| Circuit breaker per dependency | Caller and dependency | Approved |
| Bulkhead: a pool per dependency | Other features of the caller | Approved |
| Bounded queue, LIFO under overload, drop expired | The server's useful work | Approved |
| Fallback: cached or default answer | The user's page | Stale is acceptable |
| Hedged request after the p95 | Tail latency | Reads, spare capacity |
| No timeout: wait for the socket default | Nothing | Not approved |
I set timeouts and a deadline first, then retries with a budget, then breakers and bulkheads. Servers protect themselves with shedding.
Step 1: Deadline
- Set one deadline for the whole request: 2 s.
- Every call below reads the time left from the context.
If it fails
No deadline: each layer waits its own full timeout, and the total can exceed what the user waits.
The request carries a deadline. Each call takes a slot in its own pool, asks the breaker, uses a short timeout, and retries within a budget or falls back.
| tool | capability | what it gives this design | also used for |
|---|---|---|---|
| Go context | Deadline and cancellation on a context | Every call, query and goroutine below the request stops when the deadline passes. | Graceful shutdown |
| gRPC | Deadline sent with each call (grpc-timeout header) | The server sees the time left and stops with the caller. No custom header. | Cancellation across services |
| Service mesh proxy | Retries with a retry budget, per-try timeout | Retry policy in one place for every service, with a cap on retry load. | Canary routing |
| Service mesh proxy | Circuit breaking limits, outlier detection | Caps pending requests per upstream. Ejects one bad host from the pool. | Load balancing |
| Client library | Breaker, bulkhead, backoff in the caller | Per-dependency state with no extra hop. Fallbacks run in the caller. | Rate limits on the client |
| HTTP | 503 or 429 with Retry-After; Idempotency-Key header | The server tells callers when to come back. A retried POST stays safe. | Rate limiting, public APIs |
| Postgres | statement_timeout, lock_timeout | The database stops a slow query itself, even if the caller is gone. | Migrations without long locks |
| Postgres | Unique index on the idempotency key | A retried write finds its key and returns the first result. | Booking, payments |
| Redis | Key with TTL for the last good value | The fallback serves a recent answer while the breaker is open. | Caching |
| Load balancer | Health checks, connection draining | A crashed instance leaves the pool. A restart loses no requests in flight. | Deploys |
Context deadlines and gRPC carry the time left. A mesh or a client library gives me retry budgets, breakers and outlier detection. Postgres enforces its own timeouts and idempotency keys.
handle(request):
deadline = now + 2 s1 // the user's patience, set once at the edge
...
call(dependency, ctx):
left = ctx.deadline - now
IF left <= 0: RETURN error2 // do not start work nobody waits for
wait = min(per_try_timeout, left)3
send header X-Timeout-Ms = left4 // the next service inherits the deadline
RETURN dependency.call(ctx, wait)
on a request with X-Timeout-Ms = left: // the server side
IF left <= 0: RETURN 504 at once5
ctx = deadline now + left // every call it makes inherits the rest- 1One budget for the whole request. It matches how long the user or the client waits.
- 2Work for a caller that has left is wasted. Check before every call.
- 3Per try: the p99.9 of the dependency, so about 0.1% of attempts time out by mistake.
- 4Send time left, not a clock time. Machines disagree on the time; a duration needs no shared clock.
- 5Without the header, the lab backend did all 10 steps of its work. With it, the backend stopped after at most 5.
Tested source Go: outgoing and incoming deadline
// Outgoing copies the time left on ctx into the request, so the next service stops when the
// caller stops waiting.
func Outgoing(ctx context.Context, req *http.Request) error {
dl, ok := ctx.Deadline()
if !ok {
return nil
}
left := time.Until(dl)
if left <= 0 {
return fmt.Errorf("deadline passed before the call: %w", context.DeadlineExceeded)
}
req.Header.Set(TimeoutHeader, strconv.FormatInt(left.Milliseconds(), 10))
return nil
}
// Incoming gives each request a context that ends when the caller's time runs out. A request
// that arrives with no time left gets 504 and costs no work.
func Incoming(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
v := r.Header.Get(TimeoutHeader)
if v == "" {
next.ServeHTTP(w, r)
return
}
ms, err := strconv.ParseInt(v, 10, 64)
if err != nil {
http.Error(w, "bad "+TimeoutHeader, http.StatusBadRequest)
return
}
if ms <= 0 {
http.Error(w, "no time left", http.StatusGatewayTimeout)
return
}
ctx, cancel := context.WithTimeout(r.Context(), time.Duration(ms)*time.Millisecond)
defer cancel()
next.ServeHTTP(w, r.WithContext(ctx))
})
}| choice | status |
|---|---|
| Per-try timeout at the dependency's p99.9 | Approved |
| Library or socket default (often minutes, or none) | Not approved |
| Same fixed timeout at every layer | Not approved |
| Timeout without a server-side stop | Wastes server work |
I set one deadline at the edge and send the time left with every call. Each attempt waits the smaller of its own timeout and the time left.
delay(retry): // retry = 1, 2, 3 ...
ceiling = min(cap, base × 2^(retry - 1))1
RETURN random between 0 and ceiling2 // full jitter
call_with_retries(request):
FOR attempt = 1 TO 43:
result = call(request, timeout)
IF result is ok: RETURN result
IF error is not retryable4: RETURN result // 400, 404, ...
IF budget.try_retry() is false: RETURN result
IF now + delay(attempt) > deadline5: RETURN result
WAIT delay(attempt)
RETURN result- 1Exponential: 200, 400, 800 ms. The cap stops the wait from growing past what the caller can afford.
- 2Full jitter. Clients that failed at the same moment no longer retry at the same moment.
- 3The first try and at most 3 retries, as in the simulation. Google caps a request at 3 attempts.
- 4A 400 or a 404 fails the same way again. Retry timeouts, 503, 429 and connection errors.
- 5No retry that cannot finish before the deadline.
Tested source Go: capped backoff with full jitter
// Backoff is capped exponential backoff.
type Backoff struct {
Base time.Duration
Cap time.Duration
Jitter Jitter
}
// Delay is the wait before retry number `retry` (1 for the first retry).
func (b Backoff) Delay(retry int, r *rand.Rand) time.Duration {
ceiling := b.Base
for i := 1; i < retry && ceiling < b.Cap; i++ {
ceiling *= 2
}
ceiling = min(ceiling, b.Cap)
if b.Jitter == FullJitter {
return time.Duration(r.Int64N(int64(ceiling)))
}
return ceiling
}I wait a random time up to an exponential limit. Without jitter, clients that failed together retry together.
| retries | load on the failing dependency | source |
|---|---|---|
| 3 attempts, 1 layer | up to 3× | derived |
| 3 attempts, 3 layers | up to 27× | derived: 3³ |
| 3 attempts, 5 layers | up to 243× | AWS Builders' Library |
| 3 attempts, 10% budget | about 1.1× | Google SRE book |
| Simulation, no budget | 2.5 to 3.1 attempts per request | measured |
| Simulation, 10% budget | 1.03 attempts per request | measured |
budget = 10 tokens // one budget per caller process1
on every new request:
budget = min(10, budget + 0.12) // each request earns 10% of a retry
before every retry:
IF budget < 1: RETURN no3 // fail fast instead
budget = budget - 1
RETURN yes- 1Local state, no coordination. Each process caps its own retries.
- 2Each request earns a tenth of a retry. Over time, retries stay at or below 10% of requests.
- 3In a long outage the tokens run out, so retry load stops growing. Short errors still get retried.
Tested source Go: retry budget
// RetryBudget caps retries at a share of first attempts. Every request adds Ratio tokens, up to
// Max. A retry spends one token. With no token left, the caller fails fast instead of retrying.
type RetryBudget struct {
mu sync.Mutex
ratio float64
max float64
tokens float64
}
// NewRetryBudget returns a full budget.
func NewRetryBudget(ratio, max float64) *RetryBudget {
return &RetryBudget{ratio: ratio, max: max, tokens: max}
}
// OnRequest is called once for every new request, before its first attempt.
func (b *RetryBudget) OnRequest() {
b.mu.Lock()
defer b.mu.Unlock()
b.tokens = min(b.max, b.tokens+b.ratio)
}
// TryRetry spends a token and reports whether a retry is allowed.
func (b *RetryBudget) TryRetry() bool {
b.mu.Lock()
defer b.mu.Unlock()
if b.tokens < 1 {
return false
}
b.tokens--
return true
}A retry storm can outlast its cause. Every attempt waits in a full queue and times out, so it comes back as more attempts. This is a metastable failure.
Each caller may retry at most 10% of its requests. In an outage, retries then add 10% load, not 3 times the load.
| call | retry it? |
|---|---|
| GET, HEAD: read only | Approved |
| PUT: set the whole value | Approved |
| DELETE by ID | Ignore 404 on repeat |
| POST with an Idempotency-Key header | Approved |
| POST create, no key | Not approved |
| Increment a counter, append to a list | Not approved |
| Charge a card, send an email | With a key only |
- A timeout does not tell you whether the server did the work. Assume it might have.
- The client makes the key once per user action and sends the same key on every retry.
- The server stores the key with the result, under a unique index, in the same transaction as the write.
- The pattern in full: T9, money and exactly once.
I retry only calls that are safe to repeat. For a write, the client sends an idempotency key and the server returns the first result for a repeat.
- One breaker per dependency in each caller process. It needs no coordination.
- Count the last N calls, with a minimum, so 2 failures in 3 calls do not open it.
- Count timeouts and 5xx as failures. A 4xx is the caller's fault: count it as a success.
- In the simulation the breakers opened 75 times. While slow, 20% of requests succeeded, against 3% without them.
- The cost: recovery waited 1.6 s after the restart, for the breakers to close.
allow(now):
IF state = OPEN AND now - opened_at >= open_for:
state = HALF_OPEN1 // test the dependency again
IF state = OPEN: RETURN refuse // fail fast, or serve the fallback2
IF state = HALF_OPEN AND probes_sent >= probes: RETURN refuse
RETURN a ticket for this state
record(ticket, ok):
IF ticket is from an earlier state: IGNORE3
IF state = CLOSED:
add ok to the last N outcomes
IF outcomes >= min_calls AND failures / outcomes >= 50%4: state = OPEN
IF state = HALF_OPEN:
IF NOT ok: state = OPEN5 // still broken
ELSE IF every probe succeeded: state = CLOSED- 1No timer thread: the first call after the wait moves the breaker to half-open.
- 2The caller spends no thread and no timeout on a dependency that is down.
- 3A slow call that started while closed must not reopen a half-open breaker. The lab test checks this.
- 4A rate over the last N calls, not a count of failures since start, which never forgets.
- 5One failed probe opens the breaker for another full wait.
Tested source Go: allow and record
// Allow asks to make one call. It returns ErrOpen when the breaker refuses it.
func (b *Breaker) Allow(now time.Time) (Ticket, error) {
b.mu.Lock()
defer b.mu.Unlock()
if b.state == Open && now.Sub(b.openedAt) >= b.cfg.OpenFor {
b.moveTo(HalfOpen, now)
}
switch b.state {
case Open:
return Ticket{}, ErrOpen
case HalfOpen:
if b.probes >= b.cfg.Probes {
return Ticket{}, ErrOpen
}
b.probes++
}
return Ticket{gen: b.gen}, nil
}
// Record reports the outcome of a call that Allow let through.
func (b *Breaker) Record(t Ticket, now time.Time, ok bool) {
b.mu.Lock()
defer b.mu.Unlock()
if t.gen != b.gen {
return // the call started in an earlier state
}
switch b.state {
case Closed:
b.push(!ok)
if b.n >= b.cfg.MinCalls && float64(b.fails) >= b.cfg.FailureRate*float64(b.n) {
b.moveTo(Open, now)
}
case HalfOpen:
if !ok {
b.moveTo(Open, now)
return
}
b.probesOK++
if b.probesOK >= b.cfg.Probes {
b.moveTo(Closed, now)
}
}
}When half the recent calls fail, the breaker opens and callers fail fast or fall back. After a wait, a few probes test the dependency, and the breaker closes when they succeed.
pools = { payments: 10 slots, catalog: 10 slots }
call(dependency, request):
IF pools[dependency] has no free slot:
RETURN busy at once1 // never wait for a slot
take a slot
result = dependency.call(request)
give the slot back // also on error2
RETURN result- 1Waiting for a slot would hold the caller thread anyway. Fail fast and fall back.
- 2A slot that leaks on error shrinks the pool until it is empty.
Tested source Go: bulkhead
// Bulkhead is a fixed pool of slots for calls to one dependency. When the dependency is slow,
// its calls fill its own pool and fail fast. Calls to other dependencies use other pools, so
// they keep working.
type Bulkhead struct{ slots chan struct{} }
// NewBulkhead returns a pool of n slots.
func NewBulkhead(n int) *Bulkhead { return &Bulkhead{slots: make(chan struct{}, n)} }
// Do runs fn in a free slot, or returns ErrFull at once. It never waits for a slot.
func (b *Bulkhead) Do(ctx context.Context, fn func(context.Context) error) error {
select {
case b.slots <- struct{}{}:
default:
return ErrFull
}
defer func() { <-b.slots }()
return fn(ctx)
}- Size a pool by Little's law: calls a second × p99 latency, plus headroom.
- Separate pools also work at larger scale: a cluster per tier of customer, a cell per group of tenants.
Each dependency gets its own small pool. When payments hang, they fill their own pool and fail fast; browsing keeps its threads.
Same callers (retries with a budget), dependency at 500 requests a second against 1,000 offered. A bound of 100 is 0.2 s of work, below the 300 ms timeout.
admit(request, deadline):
IF queue is full: drop expired entries from the front
IF queue is still full: RETURN 503 + Retry-After1 // backpressure
add request to the queue
next(): // a worker is free
LOOP:
IF queue length > 50: take the NEWEST entry2 // adaptive LIFO
ELSE: take the OLDEST entry
IF its deadline has passed: drop it, CONTINUE3
RETURN it- 1Backpressure: a fast, cheap no. The caller can try another replica or back off.
- 2Adaptive LIFO. The newest caller still waits; the oldest has probably left.
- 3The deadline arrived with the request, so the server knows the caller has gone.
Tested source Go: admission queue
// Push admits an entry, or reports false when the queue is full: the server rejects the
// request at once, and the caller can go elsewhere.
func (q *Queue) Push(e Entry, now time.Time) bool {
if q.policy.MaxLen > 0 && q.Len() >= q.policy.MaxLen {
q.dropExpiredFront(now)
if q.Len() >= q.policy.MaxLen {
return false
}
}
q.items = append(q.items, e)
return true
}
// Pop returns the next entry to serve. Under overload it serves the newest entry, whose caller
// still waits, instead of the oldest, whose caller has probably left.
func (q *Queue) Pop(now time.Time) (Entry, bool) {
for q.Len() > 0 {
var e Entry
if q.policy.LIFOAbove > 0 && q.Len() > q.policy.LIFOAbove {
e = q.items[len(q.items)-1]
q.items = q.items[:len(q.items)-1]
} else {
e = q.items[q.head]
q.head++
}
if q.policy.DropExpired && !now.Before(e.Deadline) {
q.Dropped++
continue
}
q.compact()
return e, true
}
q.compact()
return Entry{}, false
}- Shed by priority too: drop prefetch and batch work before user requests.
- Backpressure inside one system: a bounded channel or queue blocks the producer; TCP and HTTP/2 flow control do the same on the wire.
Under overload I bound the queue, serve the newest request first and drop requests whose deadline passed. A full queue returns 503 with Retry-After.
1,000 requests a second from 10 caller instances. A user waits 2 s. Callers with a timeout wait 300 ms per attempt.
The dependency has 50 workers. Healthy, each request takes 20 ms on average. Slow, 100 ms. Cold after the restart, 40 ms. Retries: 4 attempts at most, backoff from 200 ms. Breaker: opens at 50% of the last 20 calls.
Timeouts protect the caller. Retries without a budget kept the dependency down after it was healthy. A budget let it recover at once, and server shedding kept the most useful work.
| dependency down | fallback | status |
|---|---|---|
| Recommendations | Hide the section, or show best sellers | Approved |
| Prices, stock | Last good value from the cache, with its age | Recheck at checkout |
| Search | Fewer results from a simpler index | Approved |
| Email, notifications | Queue now, send later | Approved |
| Payments | A clear error: "try again in a minute" | No silent fallback |
| Authorization | Allow when the check fails (fail open) | Not approved |
- A fallback runs often only during incidents. Test it, or it fails when you need it.
- Mark degraded answers in the response and in metrics, so you can see how often users get them.
- Never fall back to a wrong answer for money or access.
For each dependency I decide in advance what the page shows without it: a cached value, a default, a hidden section, or a clear error.
Seeded model: a 10 ms median, and 3% of requests wait 50 to 500 ms more for a reason that does not repeat. 200,000 requests per row.
hedge(read, after = p95 latency1):
send read to replica 1
WAIT for an answer or for 'after'
IF no answer yet: send the same read to replica 22
first successful answer wins
cancel the other call3- 1Only the slowest 5% send a copy. In the model the p99 fell from 364.5 ms to 34.9 ms for 5% more requests.
- 2A different replica, so the second copy does not wait behind the first.
- 3Cancel the loser, so it stops using the replica.
Tested source Go: hedge
// Hedge calls fn. If no answer arrives within `after`, it calls fn a second time and returns
// whichever answer succeeds first; the other call is cancelled. Set `after` near the p95
// latency, so only about 5% of requests send a second copy. Use it only for reads and other
// idempotent calls.
func Hedge[T any](ctx context.Context, after time.Duration, fn func(context.Context) (T, error)) (T, error) {
ctx, cancel := context.WithCancel(ctx)
defer cancel() // stops the call that lost
type result struct {
v T
err error
}
results := make(chan result, 2)
call := func() {
v, err := fn(ctx)
results <- result{v, err}
}
go call()
timer := time.NewTimer(after)
defer timer.Stop()
sent, failed := 1, 0
var last error
for {
select {
case <-timer.C:
if sent == 1 {
sent++
go call()
}
case r := <-results:
if r.err == nil {
return r.v, nil
}
failed++
last = r.err
if failed == sent { // a hedge is not a retry: an error before the hedge ends the call
var zero T
return zero, last
}
case <-ctx.Done():
var zero T
return zero, ctx.Err()
}
}
}- It works when slowness is random per request. It fails when the cause is overload, because both copies queue.
- Hedge after the p50 and you send about 50% more requests. Hedge after the p99 and the p99 does not move.
For reads, I send a second copy after the p95 latency and take the first answer. That cuts the p99 for about 5% more requests.
| event | result | why it stays safe | saved by |
|---|---|---|---|
| A dependency gets 5 times slower | Its queue grows past every caller's timeout. | Callers give up at 300 ms and hold at most 336 requests, not 5,099. | Timeout |
| The dependency crashes and restarts cold | Retries from the outage arrive together. | A 10% budget keeps attempts at 1.03 per request, so it recovers at once. | Retry budget |
| Many clients fail at the same moment | They retry in waves. | Full jitter spreads each wave over its whole backoff window. | Full jitter |
| A retried POST reaches the server twice | The second finds the idempotency key. | It returns the first result. No second order, no second charge. | Unique key |
| Payments hang for minutes | Payment calls fill their pool. | New payment calls fail fast. Browsing uses its own pool and keeps working. | Bulkhead |
| A dependency is down | The breaker opens. | Callers serve the last good value with no wait for a timeout. | Fallback |
| Offered load is twice the capacity | The queue grows. | Newest first, expired work dropped: 47% answered in time, against 3% with FIFO. | Admission queue |
| A request reaches a deep service with no time left | The header says 0 ms. | The service returns 504 at once and does no work. | Deadline header |
| One of 10 instances returns errors | A breaker per dependency sees only 10% errors. | It stays closed. The proxy ejects the bad host. | Outlier detection |
| step | add | it handles | move up when you see |
|---|---|---|---|
| 1 | A timeout on every call, from the dependency's p99.9. One deadline per request, sent downstream. | A slow dependency costs only its own calls. Callers keep free threads. | Deploys and failovers cause short bursts of user errors. |
| 2 | Retries with capped backoff and full jitter, on idempotent calls. Idempotency keys on writes. | Short errors disappear for users. | During an incident, load on the failing service grows to several times normal. |
| 3 | A retry budget, a breaker per dependency, and fallbacks. | Outages stay at their own size. The dependency recovers when its cause ends. | One slow dependency still takes threads that other features need. |
| 4 | Bulkheads in callers. Admission control in servers: bounded queues, newest first, drop expired. | Overload costs only the excess. The server keeps doing useful work near capacity. | Many services and teams, each with its own retry and timeout settings. |
| 5 | Policy in a mesh or a shared client library: retry budgets, outlier detection, limits. | One consistent policy for every call, changed in one place. | Top of the ladder. |
27× is 3 × 3 × 3. 243× is from the AWS Builders' Library; 1.1× from the Google SRE book. Do not start at step 5. A mesh adds a proxy hop and its own failure modes.
I start with timeouts and a deadline on every call. I add budgeted retries, then breakers and fallbacks, then bulkheads and shedding, each when an incident shows the gap.
0 of 10 known
A dependency has p99 = 40 ms and p99.9 = 250 ms. Which timeout do you set for one attempt?
Three layers each make up to 3 attempts per call. The bottom dependency fails. How much load reaches it?
In the simulation, full jitter replaced the retry waves with a smooth hump. Why did the dependency still never recover?
Why is a POST that charges a card not safe to retry, and how do you make it safe?
The breaker for payments is open. What does the user see, and when does the breaker test payments again?
Why does a server under overload serve the newest request first?
Dropping expired requests alone answered almost nothing in time. Why?
When does a hedged request make things worse?
With server shedding, adding a breaker lowered success while the dependency was slow. Why?
What is the difference between load shedding and backpressure?
- timeout
- At the dependency's p99.9: about 0.1% false timeouts.
- attempts
- Google: at most 3 attempts per request, and retries under 10% of requests per client.
- layers
- 3 tries at each of 5 layers: up to 243× load.
- storm
- Simulation, no budget: 2.5 to 3.1 attempts per request, 79% of work wasted, no recovery. With a budget: 1.03.
- shedding
- At 2× overload: 47% answered in time with newest first, 3% with FIFO; 31% with a bound alone.
- hedge
- After the p95: about 5% more requests. A 2013 Google benchmark cut the p99.9 from 1,800 ms to 74 ms with 2% more requests.
Timeout and amplification: AWS Builders' Library. Attempts and budget: Google SRE book. Hedge benchmark: "The Tail at Scale", 2013. Simulation: seeded, so the numbers repeat exactly.