System Design
O1

Availability: nines, copies and failure domains

Availability is the share of time, or of requests, that succeed. Multiply what a request needs, spread copies so they fail apart, and test the path you will use when they do not.

Not startedSaved in this browser only.
  1. 1Each nine cuts allowed downtime by 10: 99.9% is 43 minutes a month, 99.99% is 4.3.
  2. 2Parts in a row multiply. Copies multiply their failures, but only if they fail independently.
  3. 3A zone, a region or a shared dependency fails every copy inside it. Spread copies over failure domains.
  4. 4Your SLO cannot beat a critical dependency. A failover you never tested is not redundancy.
O1
    A

    Copies that share a zone

    simulated, same failures
    36 simulated hours · red = down3 copies in 1 zonezone 1copy 1, zone 1copy 2, zone 1copy 3, zone 1servicedown 7.2 h: the zone took every copy3 copies, 1 per zonezone 1copy 1, zone 1copy 2, zone 2copy 3, zone 3serviceup the whole time0 h12 h24 h36 h
    • Both layouts saw the same failures. Only the placement differs.
    • Copies help only against failures that hit them one at a time.

    Three copies in one zone are as available as the zone. I put one copy in each zone, so a zone failure takes one copy, not all.

    B

    The nines

    downtime the target allows
    targeta yeara montha weeka day
    90%36.5 d3 d16.8 h2.4 h
    99%3.65 d7.2 h1.68 h14.4 min
    99.5%1.83 d3.6 h50.4 min7.2 min
    99.9%8.76 h43.2 min10.1 min1.44 min
    99.95%4.38 h21.6 min5.04 min43.2 s
    99.99%52.6 min4.32 min1.01 min8.64 s
    99.999%5.26 min25.9 s6.05 s864 ms

    Downtime = (1 - target) × period. A month is 30 days; a year is 365.

    99.9% is 43 minutes a month. 99.99% is about 4 minutes, which is less than one manual failover takes.

    C

    SLI, SLO, SLA, error budget

    what you measure and promise
    termit isexample
    SLIA measurement of good eventsRequests answered without a 5xx in under 300 ms, divided by all requests
    SLOYour target for the SLI99.9% over a rolling 30 days
    SLAA contract with a penalty, set below the SLO99.5% a month, or a service credit
    Error budget1 - SLO, spent by every failure43.2 min a month, or 0.1% of requests
    burn ratealert whenbudget lastsaction
    14.4×2% gone in 1 h2.08 dpage
    6×5% gone in 6 h5 dpage
    1×10% gone in 3 d30 dticket

    Burn rates for a 99.9% SLO, from the multiwindow, multi-burn-rate recipe in Google's SRE Workbook. Budget lasts = 30 days / rate.

    My SLI is the share of good requests at the load balancer. The SLO is 99.9% over 30 days, so the budget is 43 minutes of full outage.

    D

    Redundancy options

    what survives what
    layoutsurvivesstatus
    One instanceNothingNot approved
    N+1 copies in one zoneA host or a processBelow 99.9%
    N+1 across 2 or 3 zonesA zoneApproved
    N+2 across 3 zonesA zone during a deployApproved
    Two regions, active-passiveA region, after the RTOTested failover
    Two regions, active-activeA region, at onceWrites partitioned
    CellsA bad deploy, for most usersApproved

    N is the number of copies the peak load needs. Active-active needs each record to have one home region, or a conflict rule.

    I run N+1 across three zones for stateless tiers, and a standby database in another zone. A second region comes only when the SLO needs it.

    E

    The math

    pseudo code
    availability of a deploymentpseudo code
    chain(a1 .. an)   = a1 × a2 × ... × an        // up only if every part is up1
    copies(a, n)      = 1 - (1 - a)^n             // down only if every copy is down2
    at_least(k, n, a)3 = sum for j = k..n of C(n, j) · a^j · (1 - a)^(n - j)
    
    zones:                              // copies in one zone fail together
      total = 0
      FOR EACH set S of zones that are up4:
        p = chance that exactly S is up
        FOR EACH tier: p = p × at_least(need, copies in S, a)
        total = total + p
      availability = region × total × each external dependency5
    
    // 3 copies at 99.5%, 1 zone at 99.9%:   0.999 × (1 - 0.005³)       = 99.9%6
    // 3 copies at 99.5%, 1 per zone:        1 - (1 - 0.999 × 0.995)³   = 99.99998%
    1. 1Three parts at 99.9% in a row give 99.7%. Every part a request needs lowers the whole.
    2. 2True only when copies fail independently. A shared zone, deploy or config breaks it.
    3. 3For a quorum: 2 of 3 nodes must be up. With k = 1 it is the copies formula.
    4. 4With 3 zones there are 8 sets. Once the zones are fixed, the copies are independent.
    5. 5A provider outside your zones is in series with everything. It caps the total.
    6. 6The zone caps three copies in one zone.
    Tested source Go: chain, copies, downtime · Go: quorum, zones, regions
    Go: chain, copies, downtimego
    
    // Downtime is the time a service may be down in a period and still meet the target.
    func Downtime(availability, periodSeconds float64) float64 {
      return (1 - availability) * periodSeconds
    }
    
    // Redundant is the availability of n copies when any one copy is enough and copies fail
    // independently: the service is down only when every copy is down.
    func Redundant(availability float64, n int) float64 {
      return 1 - math.Pow(1-availability, float64(n))
    }
    
    // Chain is the availability of components that a request needs one after another: the request
    // succeeds only when every component is up.
    func Chain(availability float64, n int) float64 {
      return math.Pow(availability, float64(n))
    }
    
    Go: quorum, zones, regionsgo
    
    // AtLeast is the chance that at least k of n independent copies are up, each with availability a.
    // With k = 1 it is the redundancy formula: down only when every copy is down.
    func AtLeast(k, n int, a float64) float64 {
      if k <= 0 {
        return 1
      }
      if k > n {
        return 0
      }
      if k == 1 {
        return estimate.Redundant(a, n)
      }
      p := 0.0
      for j := k; j <= n; j++ {
        p += binomial(n, j) * math.Pow(a, float64(j)) * math.Pow(1-a, float64(n-j))
      }
      return p
    }
    
    // Series is the availability of parts a request needs one after another.
    func Series(as ...float64) float64 {
      p := 1.0
      for _, a := range as {
        p *= a
      }
      return p
    }
    
    // inZones counts the instances of a tier that live in the up zones; up is a bit set of zones.
    func inZones(t Tier, zones int, up uint) int {
      m := 0
      for i := range t.Instances {
        if up&(1<<(i%zones)) != 0 {
          m++
        }
      }
      return m
    }
    
    // zoneLayer is one region's internal tiers, with zone failures. It sums over every set of up
    // zones: the chance of that set, times the chance that every tier has enough instances up in it.
    // A zone is shared by every tier, so the tiers are independent only once the zones are fixed.
    func zoneLayer(tiers []Tier, zones int, zoneAvail float64) float64 {
      total := 0.0
      for up := uint(0); up < 1<<zones; up++ {
        p := 1.0
        for z := range zones {
          if up&(1<<z) != 0 {
            p *= zoneAvail
          } else {
            p *= 1 - zoneAvail
          }
        }
        for _, t := range tiers {
          if !t.External {
            p *= AtLeast(t.Need, inZones(t, zones, up), t.Avail)
          }
        }
        total += p
      }
      return total
    }
    
    // regions combines one region's availability r over the region mode.
    func regions(a Arch, r float64) float64 {
      switch a.Mode {
      case ActiveActive:
        return estimate.Redundant(r, 2)
      case ActivePassive:
        // While the primary is down, the standby serves if the failover works, once it is done.
        lost := math.Min(1, a.RTOMin/a.OutageMin)
        return r + (1-r)*a.FailoverOK*r*(1-lost)
      default:
        return r
      }
    }
    
    // availability is the whole deployment: regions, then the external dependencies in series.
    func availability(a Arch) float64 {
      r := a.RegionAvail * zoneLayer(a.Tiers, a.Zones, a.ZoneAvail)
      p := regions(a, r)
      for _, t := range a.Tiers {
        if t.External {
          p *= t.Avail
        }
      }
      return p
    }
    

    A request path multiplies. Copies multiply their failures. With zones, I add up over which zones are up, because a zone is shared by every tier.

    F

    Failure domains

    each fails as one unit
    regioncontrol plane, network, a bad global change✓ a second regionzonepower, cooling, a fibre cut✓ copies in 2 or 3 zonesracktop-of-rack switch, power strip✓ spread hosts over rackshostdisk, memory, kernel✓ a second hostprocesscrash, out of memory, a bad deploy✓ restart, a second copy

    I name the failure domain each copy protects against. Copies on one host protect against a crash, not against a zone.

    G

    Try it: availability calculator

    presets recorded from the lab
    preset
    tiereach copycopiestier alone
    app399.9999785%
    database299.99641%
    zones
    each zone
    each region
    regions
    regionzone 1zone 2zone 3appdatabase
    99.9864%
    1.19 h down a year, 5.88 min a month.
    Meets an SLO of 99.9%: a budget of 43.2 min a month.
    SLO
    weakest
    region: perfect, it would remove 52.6 min of downtime a year.
    naive
    99.9864% if every copy failed on its own.

    Assumed: 99.5% per copy (the per-instance figure in a large cloud's compute SLA), 99.9% per zone and 99.99% per region. Active-passive assumes a region outage lasts 4 h on average.

    Multi-zone gives about 99.9864%; the region is then the weakest part. A second region, active-active, gives 99.99999815%.

    H

    Formula against simulation

    10,000 simulated years per layout
    down minutes a year, log scale · 10,000 simulated years eachformulanaivemeasured0.11101001k10k1 copy896,543 outages5,7765,7765,7563 copies, 1 zone21,827 outages5260.75113 copies, 3 zones341 outages0.70.70.753 app + quorum db, 3 zones59,314 outages190190191
    layoutformulameasured
    1 copy98.9%98.9%
    3 copies, 1 zone99.9%99.9028%
    3 copies, 3 zones99.999867%99.999858%
    3 app + quorum db, 3 zones99.9639%99.9637%
    • Each copy fails on average every 99 h and is repaired in 1 h: 99%.
    • Each zone fails every 3,996 h and is repaired in 4 h: 99.9%.
    • The rates are higher than real ones, so 10,000 years hold enough overlaps to count.
    • The quorum layout needs 1 of 3 app copies and 2 of 3 database nodes.
    the simulationpseudo code
    each part: up for Exp(MTBF) hours, then down for Exp(MTTR) hours
    take the earliest change from a queue; flip that part
    service up = every tier has enough copies up in up zones
    measured = 1 - down hours / all hours
    Tested source Go: event simulation
    Go: event simulationgo
    
    // Simulate runs l for `years` and measures the share of time every tier had enough instances up
    // in zones that were up. Each part alternates between up and down; the next change is always the
    // earliest one in a queue. With trace set, it also returns the down intervals of every part and of
    // the service in the first traceHours.
    func Simulate(l Layout, years float64, seed uint64, traceHours float64) (SimResult, map[string][]Interval) {
      ps := parts(l, seed)
      q := make(queue, len(ps))
      copy(q, ps)
      heap.Init(&q)
    
      end := years * 8760
      trace := map[string][]Interval{}
      mark := func(id string, from, to float64) {
        if from < traceHours {
          trace[id] = append(trace[id], Interval{from, min(to, traceHours)})
        }
      }
      downSince := map[*part]float64{}
    
      res := SimResult{Years: years}
      counts := make([]int, len(l.Tiers))
      now, up, serviceDownAt := 0.0, true, 0.0
      for q[0].next < end {
        p := q[0]
        now = p.next
        res.Events++
        p.up = !p.up
        if p.up {
          mark(p.id, downSince[p], now)
          p.next = now + p.rng.ExpFloat64()*p.fail.MTBF
        } else {
          downSince[p] = now
          p.next = now + p.rng.ExpFloat64()*p.fail.MTTR
        }
        heap.Fix(&q, 0)
    
        nowUp := serviceUp(l, ps, counts)
        switch {
        case up && !nowUp:
          serviceDownAt = now
          res.Outages++
        case !up && nowUp:
          res.DownHours += now - serviceDownAt
          mark("service", serviceDownAt, now)
        }
        up = nowUp
      }
      if !up {
        res.DownHours += end - serviceDownAt
      }
      res.Measured = 1 - res.DownHours/end
      return res, trace
    }
    

    I checked the formula against a simulation of failures. It agrees within 10%, and the naive formula is off by hundreds of times when copies share a zone.

    I

    Regional failover

    click a step; its path lights up
    Usersbrowsers, appsDNS, global LBhealth checkedService, region A3 zonesPostgres primarystandby in zone 2Service, region Bwarm, scaled downPostgres replicastreams WAL from A

    Step 1: Normal

    • DNS sends users to region A.
    • Region B runs a replica that streams the WAL from A.
    • The replica lags by seconds.

    If it fails

    Nothing has failed yet. Watch the replication lag: it is the data you lose if A dies now.

    A zone failure is handled inside the region in seconds. A region failure needs detect, promote, fence and shift traffic, and I practise it.

    J

    Capabilities used

    what each tool gives you
    toolcapabilitywhat it gives this designalso used for
    PostgresStreaming replication, asynchronousA standby in another zone or region. RPO = the lag.Read replicas
    Postgressynchronous_standby_names, synchronous_commitA commit waits for a standby: RPO 0, plus one round trip per commit.Money, bookings
    Postgrespg_promote()Turn a standby into the primary in one call.Planned switchover
    PostgresBase backup plus WAL archivePoint-in-time recovery to any second the archive covers.Undo a bad migration
    Postgrespg_dump, pg_restoreA logical copy you can restore and verify, as in the lab.Moving a table, version upgrades
    RedisReplicas and SentinelAutomatic failover for a cache or holds. Asynchronous: recent writes can be lost.Sessions, rate limits
    Load balancerHealth checksStops sending traffic to a dead copy or zone in seconds.Rolling deploys
    DNSHealth-checked records, short TTL, global load balancingMoves users to another region.Latency-based routing
    KubernetesTopology spread constraints, pod anti-affinityThe scheduler puts copies in different zones and hosts.Spreading load
    PatroniLeader election through a consensus storePromotes a Postgres standby and fences the old primary.Any single-leader database
    ServiceRetries with backoff and an idempotency keyA failover of a few seconds becomes a slow request, not an error.Every network call
    ServiceTimeouts, circuit breakers, fallbacksA non-critical dependency fails without failing the request.Recommendations, search

    Postgres gives me replicas, promotion and point-in-time recovery. Health checks and DNS move traffic. The service hides short failovers with retries.

    K

    RPO and RTO

    what you lose, how long it takes
    methodRPORTO
    Nightly logical dumpUp to 24 hThe restore time: hours for large data
    Base backup + WAL archiveThe last archived WAL segmentRestore, then replay the WAL
    Async replicaThe lag: secondsDetect, promote, shift: minutes
    Sync replica, other zone0Detect and promote
    • A replica copies a bad DELETE at once. Only a backup brings the rows back.
    • Restore a backup on a schedule, into a scratch database, and check it.
    • Time the restore. That time is your RTO for the backup path.

    RPO is how much data I can lose; RTO is how long recovery takes. A backup I have never restored gives me neither number.

    L

    A restore test

    dump, restore, compare
    stepresult
    Table with an index500,000 rows, 48.8 MB
    Dump, custom format5.3 MB in 1,334 ms
    Restore into an empty database6,076 ms: 8 MB/s
    Row count and row hashesEqual

    1 TB at 8 MB/s: 1,048,576 MB / 8 = 131,072 s, about 36 h. Parallel restore and physical backups go faster; measure yours.

    My restore ran at 8 MB a second, so 1 TB would take about 36 hours on one process. That is the RTO if the replicas are also bad.

    M

    Critical dependencies

    your SLO cannot beat them
    dependencyon the path?when it fails
    Primary databaseCriticalFail over to the standby.
    Payment providerUnless queuedAccept the order, charge later.
    Auth tokensUnless cachedVerify signed tokens locally.
    RecommendationsOff the pathShow a default list.
    EmailOff the pathQueue and retry.

    Upper bound: SLO ≤ product of the critical dependencies. Two at 99.95% allow at most 99.9%.

    With the payment provider at 99.95% in series, multi-zone checkout drops to 99.9364%. I take the provider off the critical path.

    N

    Blast radius and cells

    limit who one fault reaches
    the same fault, two layoutsOne shared stackload balancerservice + databasebad deploy here100% of users down4 cellscell router: user id → cellcell 1svc + db25%cell 2svc + db25%cell 3svc + db25%cell 4svc + db25%25% of users downdeploy to 1 cell first; stop on errors
    • A cell is a complete stack for a fixed set of users or tenants.
    • A thin router maps each user to a cell. Keep it simple and highly available.
    • Cells do not raise the average much. They cap the worst incident.

    I split users into cells, each a full copy of the stack. A bad deploy or a poison request reaches one cell first.

    O

    Failure cases

    what breaks, and how to stay safe
    eventresulthow to stay safesaved by
    All copies in one zone, and the zone failsEvery copy goes down together.One copy per zone. Check placement, not only the count.Spread constraints
    A bad deploy or config reaches every copyA correlated failure: copies give no protection.Deploy in stages: one cell or zone first, with automatic rollback on errors.Staged deploy
    Copies share a hidden dependencyDNS, a config store or one database takes them all down.List every dependency of the path. Count it in the math.Dependency map
    Failover never testedThe standby lacks capacity, config or data when needed.Fail over on a schedule. Measure the RTO each time.Game day
    The old primary returns after a failoverTwo primaries accept writes (split brain).Fence it: revoke its access, or use a lease from a consensus store.Patroni
    Two zones left must carry three zones of loadThey overload and fail too: a cascade.Run each zone at most 2/3 full, or shed load.Headroom
    Backups were never restoredCorrupt or incomplete backups, found during the disaster.Restore on a schedule and compare row counts and hashes.Restore test
    Clients retry at once after an outageA retry storm keeps the recovered service down.Exponential backoff with jitter; a retry budget.Backoff
    P

    Scale ladder

    start simple; climb only on a signal
    Each step adds one failure domain1One instance2+ a copy of each3+ zones4+ standby region5active-activemore load →
    Downtime a year, from the calculator presets0.010.11101001k10kOne instance: 4.04 dOne instance4.04 dTwo of each, one zone: 10.1 hTwo of each, one zone10.1 hMulti-zone: 1.19 hMulti-zone1.19 hActive-passive: 15.2 minActive-passive15.2 minActive-active: 585 msActive-active585 msminutes a year, log scale
    stepaddit givesmove up when you see
    1One instance of the service and the database.98.89%: 4.04 d down a year.Every deploy or host failure is an outage.
    2Two copies of each tier, in one zone.99.885%. The zone is now the weakest part.The SLO is 99.9% or more: the zone alone uses the budget.
    3Copies across 3 zones, a standby database in another zone.99.9864%. The region is now the weakest part.The SLO needs more than one region gives, or data must survive a region loss.
    4A standby region, async replica, tested failover.99.99711% if 90% of failovers work in 30 min.The RTO of a failover is too long, or the standby sits idle at high cost.
    5Active-active: both regions serve, each record has a home region.99.99999815% with independent regions.Top of the ladder. Add cells to limit the blast radius.

    Each copy at 99.5%, each zone 99.9%, each region 99.99%: assumptions, so read the steps as orders of magnitude. Each step costs more to run and to test.

    I start with two copies of each tier across zones. A second region comes only when the SLO needs more than one region can give, or a regulator asks.

    Q

    Drill

    predict, then reveal

    0 of 9 known

    1. The SLO is 99.9% over 30 days. How much downtime is that?

    2. Three copies at 99.5%, all in one zone at 99.9%. What is the availability?

    3. What does the simulation show for those two layouts?

    4. Checkout calls a payment provider at 99.95%. Can checkout promise 99.99%?

    5. An alert says the burn rate is 14.4. What does that mean?

    6. You have a standby region but never tested failover. What is it worth?

    7. Async replica in another region: what are the RPO and the RTO?

    8. Why N+2 and not N+1?

    9. What do cells change?

    R

    Numbers to say

    derived, simulated, cited
    99.9%
    43.2 min a month, 8.76 h a year.
    99.99%
    4.32 min a month, 52.6 min a year.
    chain
    3 parts at 99.9% in a row: 99.7%.
    one zone
    3 copies in one zone: the zone's availability, not 99.99998%.
    burn
    14.4× burns a 30-day budget in 50 h.
    restore
    8 MB/s on one process; 1 TB in about 36 h.
    cloud SLA
    99.5% per instance; 99.99% for a region across 2 or more zones.

    Cloud figures from the Amazon EC2 SLA. Restore measured with Postgres 16 on an 8-core laptop.