System Design
R2

Estimation: numbers that pick the design

Turn users into requests a second, bytes and bits. Then compare each number with what one node handles. Five minutes of arithmetic tells you which part of the design needs the most attention.

Not startedSaved in this browser only.
  1. 1A day has 86,400 seconds. Call it 10⁵: 1 million a day is about 10 a second.
  2. 2State a peak factor. Size for the peak, not the average.
  3. 3Storage is writes a day × bytes × days kept, for one copy. Then multiply by the copies.
  4. 4Compare each number with one node. The first number over the line picks the design.
R2
    A

    From users to load

    worked for the booking preset
    requests a second10 Musers a day150 Mrequests a day1.74 k/saverage17.4 k/speak× 15÷ 86,400× 10writes a second150 Mrequests a day1.5 Mwrites a day17.4/saverage174/speak÷ (1 + 99)÷ 86,400× 10storage, one copy1.5 Mwrites a day1.5 GBa day2.74 TBin 5 years× 1 kB× 365 × 5bandwidth out, at peak17.2 k/speak reads138 Mbit/sbits a second× 1 kB × 8
    • Write each step on the board. The interviewer checks the method, not the digits.
    • Split reads from writes early. They scale in different ways.
    • Storage counts one copy. Say the replica count as a separate multiplier.

    Ten million users at 15 requests each is 150 million a day. That is about 1,700 a second on average and 17,000 at a 10× peak.

    B

    Powers of ten

    memorize these
    powercountbytes
    10³thousandKB
    10⁶millionMB
    10⁹billionGB
    10¹²trillionTB
    10¹⁵quadrillionPB
    shortcutarithmetic
    1 day ≈ 10⁵ s24 × 3,600 = 86,400
    1 month ≈ 2.6 M s30 × 86,400 = 2,592,000
    1 year ≈ 31.5 M s365 × 86,400 = 31,536,000
    1 M a day ≈ 12/s10⁶ ÷ 86,400 = 11.6
    1 B a day ≈ 12,000/s10⁹ ÷ 86,400 = 11,574
    2¹⁰ ≈ 10³1,024

    I round to powers of ten and say the rounding out loud.

    C

    Latency, to scale

    orders of magnitude
    nanosecondsmicrosecondsmilliseconds1 ns1 µs1 ms1 sL1 cache read: 0.5 nsL1 cache read0.5 nsMutex lock and unlock: 25 nsMutex lock and unlock25 nsMain memory read: 100 nsMain memory read100 nsSend 2 KB on 1 Gbit/s: 20 µsSend 2 KB on 1 Gbit/s20 µsPostgres key read, local: 56 µsPostgres key read, local56 µsRedis GET, local: 146 µsRedis GET, local146 µsRead 1 MB from memory: 250 µsRead 1 MB from memory250 µsDisk seek: 8 msDisk seek8 msRead 1 MB from disk: 20 msRead 1 MB from disk20 msUS to Europe and back: 150 msUS to Europe and back150 mslog scale; each grid line is 10×

    Grey rows: Peter Norvig's published timing table, quoted as orders of magnitude. Coloured rows: measured in the lab, one client on the same laptop, median of 2,000 calls.

    Memory is 100 nanoseconds, a call on the same network is about 100 microseconds, and a round trip across an ocean is 150 milliseconds.

    D

    What one node handles

    measured, then rounded down
    whereoperationmeasuredplan with
    PostgresPrimary-key read104,800/s100,000/s
    PostgresOne-row insert, committed36,465/s10,000/s
    PostgresBooking confirm, spread (T8)19,500/s10,000/s
    PostgresBooking confirm, one hot row (T8)2,500/s1,000/s
    RedisGET65,900/s10,000/s
    RedisSET57,400/s10,000/s
    Redis3-night hold script (T8)55,000/s10,000/s
    PostgresStorage, one managed instance64 TiB, cited10 TB

    Postgres 16 and Redis 8 on a 16-thread laptop, 32 clients, 3 seconds per run. Other test runs shared the machine. The storage ceiling is from the Amazon RDS documentation.

    One Postgres node does about 10⁵ key reads and 10⁴ committed writes a second. I plan with the lower power of ten.

    E

    Try it: the estimator

    presets from the lab; move any input
    preset

    150 M requests a day÷ 86,400 s1.74 k/s average× 1017.4 k/s peak

    per secondaveragepeak
    requests1.74 k17.4 k
    writes17.4174
    reads1.72 k17.2 k
    storage
    1.5 GB a day, 2.74 TB in 5 yr (one copy)
    bandwidth
    1.39 Mbit/s in, 138 Mbit/s out, at peak
    Peak writes174/s
    log scale, 1 to 1 M one node: 10 k/s
    Peak reads17.2 k/s
    log scale, 1 to 10 M one node: 100 k/s
    Stored bytes2.74 TB
    log scale, 1 to 1 P one node: 10 TB
    1. 1One primaryplus a standby
    2. 2+ replicas, cachereads spread out
    3. 3+ shardswrites and bytes split

    Start at step 1. One primary carries the writes, the reads and the bytes. Add a standby for failover.

    Peak writes, peak reads and bytes each have a line for one node. Whichever crosses its line first decides my first design step.

    F

    The formulas

    pseudo code
    estimatepseudo code
    estimate(users, per_user, reads_per_write, bytes, years, peak):
      per_day  = users × per_user                  // requests a day
      writes   = per_day / (1 + reads_per_write1)
      reads    = per_day - writes
      avg      = per_day / 86,4002                  // seconds in a day
      peak_qps = avg × peak3
      storage  = writes × bytes × 365 × years      // one copy4
      egress   = peak reads a second × bytes × 85   // bits a second
    
      read_nodes = ceil(peak reads / 100,0006)
      shards     = max(ceil(peak writes / 10,000), ceil(storage / 10 TB7))
      IF shards > 1:          step = 3             // shard by key
      ELSE IF read_nodes > 1: step = 2             // replicas and a cache
      ELSE:                   step = 1             // one primary
    1. 1A ratio of 99 reads per write means 1 request in 100 is a write.
    2. 2Seconds in a day. Round it to 10⁵ when you work in your head.
    3. 3The peak factor is an assumption. Say it out loud: 2× for steady traffic, 10× for a flash sale.
    4. 4Replicas, indexes and backups multiply this. Name the multiplier.
    5. 5Links are rated in bits a second. Bytes a second × 8.
    6. 6Measured Postgres key reads, rounded down to a power of ten.
    7. 764 TiB is the largest managed instance. Rounded down, it is 10 TB.
    Tested source Go: the estimator
    Go: the estimatorgo
    
    // Estimate turns daily users into the numbers of the round.
    func Estimate(in Input, lim Limits) Output {
      var o Output
      o.RequestsPerDay = in.DailyUsers * in.ActionsPerUser
      o.WritesPerDay = o.RequestsPerDay / (1 + in.ReadsPerWrite)
      o.ReadsPerDay = o.RequestsPerDay - o.WritesPerDay
    
      o.AvgQPS = o.RequestsPerDay / SecondsPerDay
      o.PeakQPS = o.AvgQPS * in.PeakFactor
      o.AvgWriteQPS = o.WritesPerDay / SecondsPerDay
      o.AvgReadQPS = o.ReadsPerDay / SecondsPerDay
      o.PeakWriteQPS = o.AvgWriteQPS * in.PeakFactor
      o.PeakReadQPS = o.AvgReadQPS * in.PeakFactor
    
      o.BytesPerDay = o.WritesPerDay * in.ObjectBytes
      o.BytesTotal = o.BytesPerDay * DaysPerYear * in.RetentionYears
      o.IngressBitsSec = o.PeakWriteQPS * in.ObjectBytes * 8
      o.EgressBitsSec = o.PeakReadQPS * in.ObjectBytes * 8
    
      o.ReadNodes = atLeastOne(o.PeakReadQPS / lim.ReadsPerNode)
      o.WriteShards = atLeastOne(o.PeakWriteQPS / lim.WritesPerPrimary)
      o.ByteShards = atLeastOne(o.BytesTotal / lim.BytesPerNode)
      o.Shards = max(o.WriteShards, o.ByteShards)
      switch {
      case o.Shards > 1:
        o.Step = 3 // shard the primary
      case o.ReadNodes > 1:
        o.Step = 2 // add read replicas and a cache
      default:
        o.Step = 1 // one primary and a standby
      }
      return o
    }
    
    func atLeastOne(x float64) int {
      return max(1, int(math.Ceil(x)))
    }
    

    Every number on my board comes from one of these lines, so I can show the interviewer where it came from.

    G

    The presets, worked

    recorded from the lab
    presetpeak w/speak r/sbytesstep
    Booking17417.2 k2.74 TB1
    News feed2.86 k286 k1.81 TB2
    URL shortener3.47 k34.7 k183 TB3
    Chat46.3 k46.3 k146 TB3
    presetread nodesby writesby bytes
    Booking111
    News feed311
    URL shortener1119
    Chat1515

    Nodes each number needs. The largest shard count wins. Inputs for each preset:

    booking
    10 M users × 15, 99 : 1 reads, 1 kB, 5 yr, 10× peak
    news feed
    50 M users × 100, 100 : 1 reads, 100 B, 1 yr, 5× peak
    url shortener
    100 M users × 11, 10 : 1 reads, 500 B, 10 yr, 3× peak
    chat
    50 M users × 80, 1 : 1 reads, 200 B, 1 yr, 2× peak

    The URL shortener is small in requests but large in bytes. Ten years of links decide its shards.

    H

    Habits that score

    approved and not approved
    habitstatuswhy
    Round to one significant figure and powers of tenApprovedFast, and errors of 20% do not change the design.
    Say the peak factor and why you chose itApprovedThe interviewer can push back on one number, not on the method.
    Split reads from writesApprovedReads scale with copies and caches. Writes scale only with shards.
    Size the servers for the averageNot approvedThe system fails at the peak, when the most users are watching.
    Long division to three decimalsNot approvedUses minutes and changes no decision.
    Estimate every component before the designNot approvedEstimate only the numbers that pick a component: peak QPS, bytes, bandwidth.
    Storage without replicas or indexesName the multiplierOne copy is the honest base. Then say "× 3 for replicas".
    Skip the numbersOnly if askedSome interviewers say scale is small. Otherwise the numbers are the reason for each box.

    I give the number, the arithmetic behind it, and the decision it changes.

    I

    Scale ladder

    the estimate picks the step
    Each step adds one component1One primary2+ replicas, cache3+ shardsmore load →
    Capacity against demand101001k10k100k1MBooking, peak writes: 174 operations per secondBooking, peak writes174Chat, peak writes: 46,296 operations per secondChat, peak writes46,296Postgres, hot row: 2,500 operations per secondPostgres, hot row2,500Postgres, inserts: 36,465 operations per secondPostgres, inserts36,465Redis, GET: 65,900 operations per secondRedis, GET65,900Postgres, key reads: 104,800 operations per secondPostgres, key reads104,800operations per second, log scale
    stepaddit handlesmove up when you see
    1One Postgres primary and a standby for failover. Stateless services behind a load balancer.Up to about 10,000 writes and 100,000 key reads a second, and 10 TB of data.Peak reads pass one node.
    2Read replicas and a cache. Reads go to copies; writes stay on the primary.Reads grow with copies. Cache hits never reach Postgres.Peak writes pass one primary, or the data passes one node.
    3Shards by key. Each shard is a primary with its own copies.Each shard takes about 10,000 writes a second and 10 TB. Add shards as you grow.One key gets more traffic than one shard takes: a hot key. Split it or queue it.

    Capacity measured on one laptop. Read it as orders of magnitude. A server with durable storage commits slower, and a larger server reads faster.

    My numbers put this design at step 1. I would add replicas when peak reads pass about 100,000 a second, and shard when writes pass 10,000.

    J

    Drill

    predict, then reveal

    0 of 8 known

    1. A service gets 1 million requests a day. About how many is that a second?

    2. The average is 1,000 requests a second and the peak factor is 3. What do you size the servers for?

    3. 100 million new links a day, 500 bytes each, kept 10 years. How much storage?

    4. Reads outnumber writes 100 to 1. Which path do you optimize first?

    5. The chat preset peaks at 46,296 writes a second. One primary takes about 10,000. What do you say?

    6. The lab measured 36,465 committed inserts a second. Why does the estimator plan with 10,000?

    7. A local Redis round trip takes about 150 µs. How many sequential calls fit in a 100 ms budget?

    8. 17,000 reads a second return 1 KB each. What is the bandwidth out?

    K

    Numbers to say

    derived, measured or cited
    a day
    86,400 seconds, about 10⁵.
    1 M a day
    About 12 a second.
    a year
    About 31.5 million seconds.
    memory
    About 100 ns a read.
    local call
    About 50 to 150 µs, measured.
    ocean
    About 150 ms there and back.
    Postgres
    About 10⁵ key reads and 10⁴ committed writes a second, measured.
    Redis
    About 10⁴ to 10⁵ simple commands a second, one round trip each.
    one node
    64 TiB is the largest managed Postgres instance.

    Lab numbers: Postgres 16 and Redis 8 on an 8-core, 16-thread laptop, 32 clients. Use them as orders of magnitude.