Availability: nines, copies and failure domains
Availability is the share of time, or of requests, that succeed. Multiply what a request needs, spread copies so they fail apart, and test the path you will use when they do not.
- 1Each nine cuts allowed downtime by 10: 99.9% is 43 minutes a month, 99.99% is 4.3.
- 2Parts in a row multiply. Copies multiply their failures, but only if they fail independently.
- 3A zone, a region or a shared dependency fails every copy inside it. Spread copies over failure domains.
- 4Your SLO cannot beat a critical dependency. A failover you never tested is not redundancy.
- Both layouts saw the same failures. Only the placement differs.
- Copies help only against failures that hit them one at a time.
Three copies in one zone are as available as the zone. I put one copy in each zone, so a zone failure takes one copy, not all.
| target | a year | a month | a week | a day |
|---|---|---|---|---|
| 90% | 36.5 d | 3 d | 16.8 h | 2.4 h |
| 99% | 3.65 d | 7.2 h | 1.68 h | 14.4 min |
| 99.5% | 1.83 d | 3.6 h | 50.4 min | 7.2 min |
| 99.9% | 8.76 h | 43.2 min | 10.1 min | 1.44 min |
| 99.95% | 4.38 h | 21.6 min | 5.04 min | 43.2 s |
| 99.99% | 52.6 min | 4.32 min | 1.01 min | 8.64 s |
| 99.999% | 5.26 min | 25.9 s | 6.05 s | 864 ms |
Downtime = (1 - target) × period. A month is 30 days; a year is 365.
99.9% is 43 minutes a month. 99.99% is about 4 minutes, which is less than one manual failover takes.
| term | it is | example |
|---|---|---|
| SLI | A measurement of good events | Requests answered without a 5xx in under 300 ms, divided by all requests |
| SLO | Your target for the SLI | 99.9% over a rolling 30 days |
| SLA | A contract with a penalty, set below the SLO | 99.5% a month, or a service credit |
| Error budget | 1 - SLO, spent by every failure | 43.2 min a month, or 0.1% of requests |
| burn rate | alert when | budget lasts | action |
|---|---|---|---|
| 14.4× | 2% gone in 1 h | 2.08 d | page |
| 6× | 5% gone in 6 h | 5 d | page |
| 1× | 10% gone in 3 d | 30 d | ticket |
Burn rates for a 99.9% SLO, from the multiwindow, multi-burn-rate recipe in Google's SRE Workbook. Budget lasts = 30 days / rate.
My SLI is the share of good requests at the load balancer. The SLO is 99.9% over 30 days, so the budget is 43 minutes of full outage.
| layout | survives | status |
|---|---|---|
| One instance | Nothing | Not approved |
| N+1 copies in one zone | A host or a process | Below 99.9% |
| N+1 across 2 or 3 zones | A zone | Approved |
| N+2 across 3 zones | A zone during a deploy | Approved |
| Two regions, active-passive | A region, after the RTO | Tested failover |
| Two regions, active-active | A region, at once | Writes partitioned |
| Cells | A bad deploy, for most users | Approved |
N is the number of copies the peak load needs. Active-active needs each record to have one home region, or a conflict rule.
I run N+1 across three zones for stateless tiers, and a standby database in another zone. A second region comes only when the SLO needs it.
chain(a1 .. an) = a1 × a2 × ... × an // up only if every part is up1
copies(a, n) = 1 - (1 - a)^n // down only if every copy is down2
at_least(k, n, a)3 = sum for j = k..n of C(n, j) · a^j · (1 - a)^(n - j)
zones: // copies in one zone fail together
total = 0
FOR EACH set S of zones that are up4:
p = chance that exactly S is up
FOR EACH tier: p = p × at_least(need, copies in S, a)
total = total + p
availability = region × total × each external dependency5
// 3 copies at 99.5%, 1 zone at 99.9%: 0.999 × (1 - 0.005³) = 99.9%6
// 3 copies at 99.5%, 1 per zone: 1 - (1 - 0.999 × 0.995)³ = 99.99998%- 1Three parts at 99.9% in a row give 99.7%. Every part a request needs lowers the whole.
- 2True only when copies fail independently. A shared zone, deploy or config breaks it.
- 3For a quorum: 2 of 3 nodes must be up. With k = 1 it is the copies formula.
- 4With 3 zones there are 8 sets. Once the zones are fixed, the copies are independent.
- 5A provider outside your zones is in series with everything. It caps the total.
- 6The zone caps three copies in one zone.
Tested source Go: chain, copies, downtime · Go: quorum, zones, regions
// Downtime is the time a service may be down in a period and still meet the target.
func Downtime(availability, periodSeconds float64) float64 {
return (1 - availability) * periodSeconds
}
// Redundant is the availability of n copies when any one copy is enough and copies fail
// independently: the service is down only when every copy is down.
func Redundant(availability float64, n int) float64 {
return 1 - math.Pow(1-availability, float64(n))
}
// Chain is the availability of components that a request needs one after another: the request
// succeeds only when every component is up.
func Chain(availability float64, n int) float64 {
return math.Pow(availability, float64(n))
}
// AtLeast is the chance that at least k of n independent copies are up, each with availability a.
// With k = 1 it is the redundancy formula: down only when every copy is down.
func AtLeast(k, n int, a float64) float64 {
if k <= 0 {
return 1
}
if k > n {
return 0
}
if k == 1 {
return estimate.Redundant(a, n)
}
p := 0.0
for j := k; j <= n; j++ {
p += binomial(n, j) * math.Pow(a, float64(j)) * math.Pow(1-a, float64(n-j))
}
return p
}
// Series is the availability of parts a request needs one after another.
func Series(as ...float64) float64 {
p := 1.0
for _, a := range as {
p *= a
}
return p
}
// inZones counts the instances of a tier that live in the up zones; up is a bit set of zones.
func inZones(t Tier, zones int, up uint) int {
m := 0
for i := range t.Instances {
if up&(1<<(i%zones)) != 0 {
m++
}
}
return m
}
// zoneLayer is one region's internal tiers, with zone failures. It sums over every set of up
// zones: the chance of that set, times the chance that every tier has enough instances up in it.
// A zone is shared by every tier, so the tiers are independent only once the zones are fixed.
func zoneLayer(tiers []Tier, zones int, zoneAvail float64) float64 {
total := 0.0
for up := uint(0); up < 1<<zones; up++ {
p := 1.0
for z := range zones {
if up&(1<<z) != 0 {
p *= zoneAvail
} else {
p *= 1 - zoneAvail
}
}
for _, t := range tiers {
if !t.External {
p *= AtLeast(t.Need, inZones(t, zones, up), t.Avail)
}
}
total += p
}
return total
}
// regions combines one region's availability r over the region mode.
func regions(a Arch, r float64) float64 {
switch a.Mode {
case ActiveActive:
return estimate.Redundant(r, 2)
case ActivePassive:
// While the primary is down, the standby serves if the failover works, once it is done.
lost := math.Min(1, a.RTOMin/a.OutageMin)
return r + (1-r)*a.FailoverOK*r*(1-lost)
default:
return r
}
}
// availability is the whole deployment: regions, then the external dependencies in series.
func availability(a Arch) float64 {
r := a.RegionAvail * zoneLayer(a.Tiers, a.Zones, a.ZoneAvail)
p := regions(a, r)
for _, t := range a.Tiers {
if t.External {
p *= t.Avail
}
}
return p
}
A request path multiplies. Copies multiply their failures. With zones, I add up over which zones are up, because a zone is shared by every tier.
I name the failure domain each copy protects against. Copies on one host protect against a crash, not against a zone.
| tier | each copy | copies | tier alone |
|---|---|---|---|
| app | 3 | 99.9999785% | |
| database | 2 | 99.99641% |
- weakest
- region: perfect, it would remove 52.6 min of downtime a year.
- naive
- 99.9864% if every copy failed on its own.
Assumed: 99.5% per copy (the per-instance figure in a large cloud's compute SLA), 99.9% per zone and 99.99% per region. Active-passive assumes a region outage lasts 4 h on average.
Multi-zone gives about 99.9864%; the region is then the weakest part. A second region, active-active, gives 99.99999815%.
| layout | formula | measured |
|---|---|---|
| 1 copy | 98.9% | 98.9% |
| 3 copies, 1 zone | 99.9% | 99.9028% |
| 3 copies, 3 zones | 99.999867% | 99.999858% |
| 3 app + quorum db, 3 zones | 99.9639% | 99.9637% |
- Each copy fails on average every 99 h and is repaired in 1 h: 99%.
- Each zone fails every 3,996 h and is repaired in 4 h: 99.9%.
- The rates are higher than real ones, so 10,000 years hold enough overlaps to count.
- The quorum layout needs 1 of 3 app copies and 2 of 3 database nodes.
each part: up for Exp(MTBF) hours, then down for Exp(MTTR) hours
take the earliest change from a queue; flip that part
service up = every tier has enough copies up in up zones
measured = 1 - down hours / all hoursTested source Go: event simulation
// Simulate runs l for `years` and measures the share of time every tier had enough instances up
// in zones that were up. Each part alternates between up and down; the next change is always the
// earliest one in a queue. With trace set, it also returns the down intervals of every part and of
// the service in the first traceHours.
func Simulate(l Layout, years float64, seed uint64, traceHours float64) (SimResult, map[string][]Interval) {
ps := parts(l, seed)
q := make(queue, len(ps))
copy(q, ps)
heap.Init(&q)
end := years * 8760
trace := map[string][]Interval{}
mark := func(id string, from, to float64) {
if from < traceHours {
trace[id] = append(trace[id], Interval{from, min(to, traceHours)})
}
}
downSince := map[*part]float64{}
res := SimResult{Years: years}
counts := make([]int, len(l.Tiers))
now, up, serviceDownAt := 0.0, true, 0.0
for q[0].next < end {
p := q[0]
now = p.next
res.Events++
p.up = !p.up
if p.up {
mark(p.id, downSince[p], now)
p.next = now + p.rng.ExpFloat64()*p.fail.MTBF
} else {
downSince[p] = now
p.next = now + p.rng.ExpFloat64()*p.fail.MTTR
}
heap.Fix(&q, 0)
nowUp := serviceUp(l, ps, counts)
switch {
case up && !nowUp:
serviceDownAt = now
res.Outages++
case !up && nowUp:
res.DownHours += now - serviceDownAt
mark("service", serviceDownAt, now)
}
up = nowUp
}
if !up {
res.DownHours += end - serviceDownAt
}
res.Measured = 1 - res.DownHours/end
return res, trace
}
I checked the formula against a simulation of failures. It agrees within 10%, and the naive formula is off by hundreds of times when copies share a zone.
Step 1: Normal
- DNS sends users to region A.
- Region B runs a replica that streams the WAL from A.
- The replica lags by seconds.
If it fails
Nothing has failed yet. Watch the replication lag: it is the data you lose if A dies now.
A zone failure is handled inside the region in seconds. A region failure needs detect, promote, fence and shift traffic, and I practise it.
| tool | capability | what it gives this design | also used for |
|---|---|---|---|
| Postgres | Streaming replication, asynchronous | A standby in another zone or region. RPO = the lag. | Read replicas |
| Postgres | synchronous_standby_names, synchronous_commit | A commit waits for a standby: RPO 0, plus one round trip per commit. | Money, bookings |
| Postgres | pg_promote() | Turn a standby into the primary in one call. | Planned switchover |
| Postgres | Base backup plus WAL archive | Point-in-time recovery to any second the archive covers. | Undo a bad migration |
| Postgres | pg_dump, pg_restore | A logical copy you can restore and verify, as in the lab. | Moving a table, version upgrades |
| Redis | Replicas and Sentinel | Automatic failover for a cache or holds. Asynchronous: recent writes can be lost. | Sessions, rate limits |
| Load balancer | Health checks | Stops sending traffic to a dead copy or zone in seconds. | Rolling deploys |
| DNS | Health-checked records, short TTL, global load balancing | Moves users to another region. | Latency-based routing |
| Kubernetes | Topology spread constraints, pod anti-affinity | The scheduler puts copies in different zones and hosts. | Spreading load |
| Patroni | Leader election through a consensus store | Promotes a Postgres standby and fences the old primary. | Any single-leader database |
| Service | Retries with backoff and an idempotency key | A failover of a few seconds becomes a slow request, not an error. | Every network call |
| Service | Timeouts, circuit breakers, fallbacks | A non-critical dependency fails without failing the request. | Recommendations, search |
Postgres gives me replicas, promotion and point-in-time recovery. Health checks and DNS move traffic. The service hides short failovers with retries.
| method | RPO | RTO |
|---|---|---|
| Nightly logical dump | Up to 24 h | The restore time: hours for large data |
| Base backup + WAL archive | The last archived WAL segment | Restore, then replay the WAL |
| Async replica | The lag: seconds | Detect, promote, shift: minutes |
| Sync replica, other zone | 0 | Detect and promote |
- A replica copies a bad DELETE at once. Only a backup brings the rows back.
- Restore a backup on a schedule, into a scratch database, and check it.
- Time the restore. That time is your RTO for the backup path.
RPO is how much data I can lose; RTO is how long recovery takes. A backup I have never restored gives me neither number.
| step | result |
|---|---|
| Table with an index | 500,000 rows, 48.8 MB |
| Dump, custom format | 5.3 MB in 1,334 ms |
| Restore into an empty database | 6,076 ms: 8 MB/s |
| Row count and row hashes | Equal |
1 TB at 8 MB/s: 1,048,576 MB / 8 = 131,072 s, about 36 h. Parallel restore and physical backups go faster; measure yours.
My restore ran at 8 MB a second, so 1 TB would take about 36 hours on one process. That is the RTO if the replicas are also bad.
| dependency | on the path? | when it fails |
|---|---|---|
| Primary database | Critical | Fail over to the standby. |
| Payment provider | Unless queued | Accept the order, charge later. |
| Auth tokens | Unless cached | Verify signed tokens locally. |
| Recommendations | Off the path | Show a default list. |
| Off the path | Queue and retry. |
Upper bound: SLO ≤ product of the critical dependencies. Two at 99.95% allow at most 99.9%.
With the payment provider at 99.95% in series, multi-zone checkout drops to 99.9364%. I take the provider off the critical path.
- A cell is a complete stack for a fixed set of users or tenants.
- A thin router maps each user to a cell. Keep it simple and highly available.
- Cells do not raise the average much. They cap the worst incident.
I split users into cells, each a full copy of the stack. A bad deploy or a poison request reaches one cell first.
| event | result | how to stay safe | saved by |
|---|---|---|---|
| All copies in one zone, and the zone fails | Every copy goes down together. | One copy per zone. Check placement, not only the count. | Spread constraints |
| A bad deploy or config reaches every copy | A correlated failure: copies give no protection. | Deploy in stages: one cell or zone first, with automatic rollback on errors. | Staged deploy |
| Copies share a hidden dependency | DNS, a config store or one database takes them all down. | List every dependency of the path. Count it in the math. | Dependency map |
| Failover never tested | The standby lacks capacity, config or data when needed. | Fail over on a schedule. Measure the RTO each time. | Game day |
| The old primary returns after a failover | Two primaries accept writes (split brain). | Fence it: revoke its access, or use a lease from a consensus store. | Patroni |
| Two zones left must carry three zones of load | They overload and fail too: a cascade. | Run each zone at most 2/3 full, or shed load. | Headroom |
| Backups were never restored | Corrupt or incomplete backups, found during the disaster. | Restore on a schedule and compare row counts and hashes. | Restore test |
| Clients retry at once after an outage | A retry storm keeps the recovered service down. | Exponential backoff with jitter; a retry budget. | Backoff |
| step | add | it gives | move up when you see |
|---|---|---|---|
| 1 | One instance of the service and the database. | 98.89%: 4.04 d down a year. | Every deploy or host failure is an outage. |
| 2 | Two copies of each tier, in one zone. | 99.885%. The zone is now the weakest part. | The SLO is 99.9% or more: the zone alone uses the budget. |
| 3 | Copies across 3 zones, a standby database in another zone. | 99.9864%. The region is now the weakest part. | The SLO needs more than one region gives, or data must survive a region loss. |
| 4 | A standby region, async replica, tested failover. | 99.99711% if 90% of failovers work in 30 min. | The RTO of a failover is too long, or the standby sits idle at high cost. |
| 5 | Active-active: both regions serve, each record has a home region. | 99.99999815% with independent regions. | Top of the ladder. Add cells to limit the blast radius. |
Each copy at 99.5%, each zone 99.9%, each region 99.99%: assumptions, so read the steps as orders of magnitude. Each step costs more to run and to test.
I start with two copies of each tier across zones. A second region comes only when the SLO needs more than one region can give, or a regulator asks.
0 of 9 known
The SLO is 99.9% over 30 days. How much downtime is that?
Three copies at 99.5%, all in one zone at 99.9%. What is the availability?
What does the simulation show for those two layouts?
Checkout calls a payment provider at 99.95%. Can checkout promise 99.99%?
An alert says the burn rate is 14.4. What does that mean?
You have a standby region but never tested failover. What is it worth?
Async replica in another region: what are the RPO and the RTO?
Why N+2 and not N+1?
What do cells change?
- 99.9%
- 43.2 min a month, 8.76 h a year.
- 99.99%
- 4.32 min a month, 52.6 min a year.
- chain
- 3 parts at 99.9% in a row: 99.7%.
- one zone
- 3 copies in one zone: the zone's availability, not 99.99998%.
- burn
- 14.4× burns a 30-day budget in 50 h.
- restore
- 8 MB/s on one process; 1 TB in about 36 h.
- cloud SLA
- 99.5% per instance; 99.99% for a region across 2 or more zones.
Cloud figures from the Amazon EC2 SLA. Restore measured with Postgres 16 on an 8-core laptop.