The 60 minutes: run the round
The round is 60 minutes on a virtual whiteboard, and the candidate drives. This sheet is a plan for the hour: what to say, what to draw and when to move on.
- 1You drive. Say the plan for the hour in the first minute, then follow it.
- 2Scope first: 3 to 5 actions, and each quality as a number.
- 3By minute 35, one request goes end to end with no unexplained box.
- 4Save 10 minutes for operations and a one-minute summary.
A budget, not a rule. If introductions take 5 minutes, take them from the deep dives, not from the requirements.
I will spend about 8 minutes on requirements, 5 on numbers, 10 on the API and data, then design and go deep where you want.
1. Requirements minutes 0 to 8
"Before I design, I want to agree on scope. Who uses this, and which three things must they do?"
Draw
- Functional: 3 to 5 verbs, numbered.
- Non-functional: each as a number (latency, availability, consistency).
- Out of scope: one line.
Done when
The interviewer agrees with the list, and you have written it on the board.
Trap
Designing before the scope is agreed, or asking 20 questions and leaving no time to design.
I say what I am doing at each step, so the interviewer always knows where we are.
| area | question | it decides |
|---|---|---|
| Users | Who uses it? Which 3 actions matter most? | The API and the scope |
| Scale | Daily users? Requests per user? Growth in 5 years? | Peak QPS, shards |
| Reads and writes | How many reads per write? | Which path to optimize |
| Data | Size of one object? How long do we keep it? | Storage, blob store or rows |
| Latency | What p99 must the main call meet? | Cache, sync or async |
| Consistency | Must a user see their own write at once? Can money or stock be counted twice? | Transactions, read path |
| Availability | What does 1 minute of downtime cost? One region or many? | Redundancy, failover |
| Security | Who may see what? Personal data? | Auth, encryption, audit |
I ask only questions whose answer changes the design, and I state an assumption when the interviewer has no answer.
| moment | move |
|---|---|
| Minute 1 | Say the plan for the hour, phase by phase. |
| No answer to a question | State an assumption with a number, and write it down. |
| End of each phase | Summarize in one sentence and ask "shall I move on?" |
| A fork in the design | Name two options, pick one, give the reason. |
| The interviewer hints | Follow the hint at once and say it back in your words. |
| Stuck | Go back to the numbers. The largest number shows the problem. |
| Minute 35 | "The design works end to end. I suggest a deep dive on X because of Y." |
| Minute 55 | Summarize: the design, two trade-offs, what you would build next. |
Here are two options. I pick the first because of the read ratio. Tell me if you want me to explore the other.
- Keep the zones in place. Add to them; do not erase them.
- Write numbers next to the boxes they size.
- Colour follows the technology: Service, Redis, Postgres.
Everything I said is on the board, so the interviewer can point at any part and ask about it.
- scope
- Requirements and clarifying questions, in phase 1.
- completeness
- A design that works end to end, with no gap for an unknown box. Phase 5.
- resilience
- Fault tolerance, high availability and scale: redundancy, and headroom for a lost node. Phases 5 and 6.
- production
- The launch dashboard, and how you debug rising errors or latency. Phase 7.
- trade-offs
- Two options at each fork, with a reason for the choice. Every phase.
- depth
- Real scenarios: fan-out, read and write paths, database choice and sharding. Phase 6.
I give the trade-off for each choice, not only the choice.
| box | redundancy | one fails |
|---|---|---|
| Load balancer | 2 or more, across zones | DNS or the other node takes the traffic. |
| Service | N stateless copies | The balancer skips it; others take its share. |
| Postgres | Primary, standby in another zone | The standby is promoted. Writes pause briefly. |
| Redis | Replica, or a cluster | Cache misses go to the database. Size it for that. |
| Queue | Replicated log or broker | Producers retry; consumers resume from their offset. |
| equal nodes | survive the loss of | run each at most at |
|---|---|---|
| 2 | 1 | 50% |
| 3 | 1 | 67% |
| 4 | 1 | 75% |
| 5 | 1 | 80% |
| 6 | 1 | 83% |
Most load per node = (nodes − lost) ÷ nodes. Running below this line is the over-provisioning that keeps the peak safe after a failure.
Three zones that must survive the loss of one can each run at two thirds of capacity at most.
| # | errors or latency rise: check | then |
|---|---|---|
| 1 | Scope: one endpoint, host, zone or client version, or all? | One host: take it out of the pool. |
| 2 | Change: a deploy, a config, a flag, a traffic shift? | Roll it back first. Investigate after. |
| 3 | Dependencies: which one got slower first? | Follow a trace of a slow request. |
| 4 | Saturation: CPU, memory, connection pool, queue depth? | Scale out, or shed load. |
| 5 | Data: one hot key or one large tenant? | Rate-limit or split that key. |
The four golden signals (latency, traffic, errors, saturation) are from the Google SRE book.
If errors or latency rise, I scope the problem, check the last change, then follow the slowest dependency.
| scenario | options | what decides |
|---|---|---|
| Timeline fan-out | Push on write; pull on read; hybrid | Followers per author. Readers accept eventual consistency. |
| Read or write path | Precompute on write; compute on read | The read-to-write ratio from your estimate. |
| Database choice | Postgres relational; wide-column | Transactions and constraints, or write volume with simple key access. |
| Sharding strategy | Hash of a key; ranges; a directory | Hot keys, range scans, and which queries must stay on one shard. |
| Caching | Redis cache-aside; write-through | How stale a read may be, and the hit rate. |
| Sync or async | Answer now; accept, queue, finish later | Whether the user must wait for the result. |
Most users get push fan-out. Authors with millions of followers are merged at read time, and timelines may show a post a little late.
| mistake | status | do this instead | phase |
|---|---|---|---|
| Draw boxes in the first minute | Not approved | Agree on scope and write it down first. | 1 |
| Ask questions for 15 minutes | Not approved | Ask the 8 that decide something; assume the rest aloud. | 1 |
| Skip the numbers | Not approved | Five minutes of estimates pick the first design step. | 2 |
| A box you cannot explain | Not approved | Draw only what you can describe: what it stores, how it fails. | 5 |
| One of anything | Not approved | Two or more of every box, across zones, with headroom. | 5 |
| "It depends", then silence | Not approved | Say what it depends on, then choose. | 6 |
| Wait for the interviewer to lead | Not approved | Propose the next step and the reason. | all |
| Add a cache at once | When reads need it | Add it when the read rate or the latency target needs it. | 5 |
| No metrics, no alerts | Not approved | Name the dashboard and two alerts before the summary. | 7 |
| Run out of time in a deep dive | Not approved | At minute 55, stop and summarize. | 7 |
0 of 8 known
The prompt is "design a news feed". What is your first sentence?
It is minute 30. The interviewer asks about cache eviction, but no request goes end to end yet. What do you do?
Why write an out-of-scope line?
How do you show fault tolerance without a separate deep dive?
What goes on the dashboard at launch?
Errors rose 10 minutes after a deploy. What is your first action?
An author has 50 million followers. Push or pull?
The interviewer says nothing for a long time. What do you do?
- hour
- 60 minutes, 7 phases.
- scope
- 3 to 5 functional requirements; each quality as a number.
- by min 8
- Scope agreed and written.
- by min 23
- API and data model written.
- by min 35
- One request end to end.
- deep dives
- 2, chosen from your numbers.
- headroom
- 3 zones, lose 1: run at 67% at most.
- signals
- 4: latency, traffic, errors, saturation.
- summary
- 1 minute, before time runs out.