System Design
R1

The 60 minutes: run the round

The round is 60 minutes on a virtual whiteboard, and the candidate drives. This sheet is a plan for the hour: what to say, what to draw and when to move on.

Not startedSaved in this browser only.
  1. 1You drive. Say the plan for the hour in the first minute, then follow it.
  2. 2Scope first: 3 to 5 actions, and each quality as a number.
  3. 3By minute 35, one request goes end to end with no unexplained box.
  4. 4Save 10 minutes for operations and a one-minute summary.
R1
    A

    The hour

    minutes by phase, with checkpoints
    min 8 ✓ scope agreedmin 23 ✓ API and data writtenmin 35 ✓ design works end to endmin 50 ✓ two deep dives donemin 60 ✓ summary givenRequirements8 min3 to 5 actions,5 numbersEstimates5 minpeak QPS,bytes, bitsAPI5 min4 to 6callsDatamodel5 mintables,keysHigh-leveldesign12 minboxes and arrows,one request tracedDeepdives15 mintwo parts in depth,options and costsOperationsand wrap-up10 mindashboard, alerts,one-minute summary051015202530354045505560on the board after each phase

    A budget, not a rule. If introductions take 5 minutes, take them from the deep dives, not from the requirements.

    I will spend about 8 minutes on requirements, 5 on numbers, 10 on the API and data, then design and go deep where you want.

    B

    Walk the hour

    click a phase, or start the clock
    00:00 of 60:00Requirements: 08:00 left in this phase

    1. Requirements minutes 0 to 8

    "Before I design, I want to agree on scope. Who uses this, and which three things must they do?"

    Draw

    • Functional: 3 to 5 verbs, numbered.
    • Non-functional: each as a number (latency, availability, consistency).
    • Out of scope: one line.

    Done when

    The interviewer agrees with the list, and you have written it on the board.

    Trap

    Designing before the scope is agreed, or asking 20 questions and leaving no time to design.

    I say what I am doing at each step, so the interviewer always knows where we are.

    C

    Questions to ask

    each one decides something
    areaquestionit decides
    UsersWho uses it? Which 3 actions matter most?The API and the scope
    ScaleDaily users? Requests per user? Growth in 5 years?Peak QPS, shards
    Reads and writesHow many reads per write?Which path to optimize
    DataSize of one object? How long do we keep it?Storage, blob store or rows
    LatencyWhat p99 must the main call meet?Cache, sync or async
    ConsistencyMust a user see their own write at once? Can money or stock be counted twice?Transactions, read path
    AvailabilityWhat does 1 minute of downtime cost? One region or many?Redundancy, failover
    SecurityWho may see what? Personal data?Auth, encryption, audit

    I ask only questions whose answer changes the design, and I state an assumption when the interviewer has no answer.

    D

    Drive the conversation

    moves that keep you in control
    momentmove
    Minute 1Say the plan for the hour, phase by phase.
    No answer to a questionState an assumption with a number, and write it down.
    End of each phaseSummarize in one sentence and ask "shall I move on?"
    A fork in the designName two options, pick one, give the reason.
    The interviewer hintsFollow the hint at once and say it back in your words.
    StuckGo back to the numbers. The largest number shows the problem.
    Minute 35"The design works end to end. I suggest a deep dive on X because of Y."
    Minute 55Summarize: the design, two trade-offs, what you would build next.

    Here are two options. I pick the first because of the read ratio. Tell me if you want me to explore the other.

    E

    The board at minute 35

    five zones, numbered by phase
    1Requirements2Numbers3API4Data model5DiagramFunctional1. search rooms by date2. hold a room, 10 min3. pay and confirmOut of scope: reviewsNon-functionalno double bookingconfirm p99 < 1 s99.9% a month10 M users a day1.74 k/s average, 17.4 k/s peak174 writes/s at peak2.74 TB in 5 years→ one primary is enoughGET /availability?…POST /holdsPOST /bookings Idempotency-KeyGET /bookings/{id}GET /bookings?cursor=inventoryPK hotel, room type, nighttotal, reservedCHECK reserved ≤ totalbookingsPK id · UQ idempotency_keyshard key: hotel_idPostgres: transactions, constraintsClientLoadbalancer ×2Servicestateless ×NCacheRedisPrimaryPostgresStandby, replicasother zonesQueueasync workWorkeremails, search
    • Keep the zones in place. Add to them; do not erase them.
    • Write numbers next to the boxes they size.
    • Colour follows the technology: Service, Redis, Postgres.

    Everything I said is on the board, so the interviewer can point at any part and ask about it.

    F

    What they score

    and where you show it
    scope
    Requirements and clarifying questions, in phase 1.
    completeness
    A design that works end to end, with no gap for an unknown box. Phase 5.
    resilience
    Fault tolerance, high availability and scale: redundancy, and headroom for a lost node. Phases 5 and 6.
    production
    The launch dashboard, and how you debug rising errors or latency. Phase 7.
    trade-offs
    Two options at each fork, with a reason for the choice. Every phase.
    depth
    Real scenarios: fan-out, read and write paths, database choice and sharding. Phase 6.

    I give the trade-off for each choice, not only the choice.

    G

    Fault tolerance in one pass

    redundancy, then headroom
    boxredundancyone fails
    Load balancer2 or more, across zonesDNS or the other node takes the traffic.
    ServiceN stateless copiesThe balancer skips it; others take its share.
    PostgresPrimary, standby in another zoneThe standby is promoted. Writes pause briefly.
    RedisReplica, or a clusterCache misses go to the database. Size it for that.
    QueueReplicated log or brokerProducers retry; consumers resume from their offset.
    equal nodessurvive the loss ofrun each at most at
    2150%
    3167%
    4175%
    5180%
    6183%

    Most load per node = (nodes − lost) ÷ nodes. Running below this line is the over-provisioning that keeps the peak safe after a failure.

    Three zones that must survive the loss of one can each run at two thirds of capacity at most.

    H

    Operations at launch

    the dashboard, then the debug order
    Trafficrequests/sby endpoint, regionErrors5xx and timeouts, %by endpointLatencyp50 and p99by endpointSaturationCPU, memory, poolsqueue depthPostgresp99 query timeconnections, lagRedishit rate, memoryevictionsQueuedepthage of oldestBusinessbookings a minutepayment successdashed line: a deploy · red: a graph that moved after it · shapes are illustrative
    #errors or latency rise: checkthen
    1Scope: one endpoint, host, zone or client version, or all?One host: take it out of the pool.
    2Change: a deploy, a config, a flag, a traffic shift?Roll it back first. Investigate after.
    3Dependencies: which one got slower first?Follow a trace of a slow request.
    4Saturation: CPU, memory, connection pool, queue depth?Scale out, or shed load.
    5Data: one hot key or one large tenant?Rate-limit or split that key.

    The four golden signals (latency, traffic, errors, saturation) are from the Google SRE book.

    If errors or latency rise, I scope the problem, check the last change, then follow the slowest dependency.

    I

    Deep dives they ask for

    options, and what decides
    Push: fan-out on writeNew postFan-outworkerlist 1list 2list 3list NWrite: N list inserts, one per follower.Read: 1 list read. Fast and cheap.✕ An author with 50 M followers: 50 M writes.Followers see the post later: eventual consistency.Pull: fan-out on readposts of 1posts of 2posts of MMergesort by timeReaderWrite: 1 insert.Read: M author reads, then a merge.✕ Slow when a reader follows thousands.✓ Hybrid: push for most, pull for big authors.
    scenariooptionswhat decides
    Timeline fan-outPush on write; pull on read; hybridFollowers per author. Readers accept eventual consistency.
    Read or write pathPrecompute on write; compute on readThe read-to-write ratio from your estimate.
    Database choicePostgres relational; wide-columnTransactions and constraints, or write volume with simple key access.
    Sharding strategyHash of a key; ranges; a directoryHot keys, range scans, and which queries must stay on one shard.
    CachingRedis cache-aside; write-throughHow stale a read may be, and the hit rate.
    Sync or asyncAnswer now; accept, queue, finish laterWhether the user must wait for the result.

    Most users get push fan-out. Authors with millions of followers are merged at read time, and timelines may show a post a little late.

    J

    Common mistakes

    not approved, and the fix
    mistakestatusdo this insteadphase
    Draw boxes in the first minuteNot approvedAgree on scope and write it down first.1
    Ask questions for 15 minutesNot approvedAsk the 8 that decide something; assume the rest aloud.1
    Skip the numbersNot approvedFive minutes of estimates pick the first design step.2
    A box you cannot explainNot approvedDraw only what you can describe: what it stores, how it fails.5
    One of anythingNot approvedTwo or more of every box, across zones, with headroom.5
    "It depends", then silenceNot approvedSay what it depends on, then choose.6
    Wait for the interviewer to leadNot approvedPropose the next step and the reason.all
    Add a cache at onceWhen reads need itAdd it when the read rate or the latency target needs it.5
    No metrics, no alertsNot approvedName the dashboard and two alerts before the summary.7
    Run out of time in a deep diveNot approvedAt minute 55, stop and summarize.7
    K

    Drill

    predict, then reveal

    0 of 8 known

    1. The prompt is "design a news feed". What is your first sentence?

    2. It is minute 30. The interviewer asks about cache eviction, but no request goes end to end yet. What do you do?

    3. Why write an out-of-scope line?

    4. How do you show fault tolerance without a separate deep dive?

    5. What goes on the dashboard at launch?

    6. Errors rose 10 minutes after a deploy. What is your first action?

    7. An author has 50 million followers. Push or pull?

    8. The interviewer says nothing for a long time. What do you do?

    L

    Numbers to say

    the plan in numbers
    hour
    60 minutes, 7 phases.
    scope
    3 to 5 functional requirements; each quality as a number.
    by min 8
    Scope agreed and written.
    by min 23
    API and data model written.
    by min 35
    One request end to end.
    deep dives
    2, chosen from your numbers.
    headroom
    3 zones, lose 1: run at 67% at most.
    signals
    4: latency, traffic, errors, saturation.
    summary
    1 minute, before time runs out.