System Design
R3

Requirements to API: verbs, numbers, calls

Write what the system does as verbs and how well it does it as numbers. The verbs become the API; the numbers pick the components of the first diagram.

Not startedSaved in this browser only.
  1. 1A functional requirement is a verb. Each verb becomes one API call.
  2. 2A non-functional requirement is an adjective until you give it a number.
  3. 3Calls that create or pay take an idempotency key. Lists take a cursor.
  4. 4If the user does not need the result now, accept it (202) and finish later.
R3
    A

    Two kinds of requirement

    verbs and adjectives
    Functional: what it doesverbs a user performssearch roomshold a roompay and confirmsee my bookingsAPI callsone per verbNon-functional: how welladjectives, until you add numbersfastavailablecorrectdurablenumbersp99, %, RPOPostgres txnRedis cachecomponents
    • Write 3 to 5 verbs. Agree on them before you draw a box.
    • Write one line of what is out of scope.
    • Turn each adjective into a number. The next panel shows how.

    These four verbs are the functional scope. For each quality, I want a number we can test.

    B

    Adjectives to numbers

    example targets for booking
    adjectiveas a numberit picks
    fastsearch p99 ≤ 300 ms; confirm p99 < 1 sCache for search
    available99.9% of requests succeed in 30 days: 43.2 min down2 or more of every box
    correct0 double bookingsTransaction and a CHECK
    durable0 confirmed bookings lost (RPO 0)Sync standby
    freshsearch at most 60 s staleCache TTL of 60 s
    scalable17.4 k requests/s at peak (R2)Step 1: one primary

    SLI: what you measure (the share of good requests). SLO: the target for it. Measure at the load balancer, where users see the result.

    By available I mean 99.9% of requests succeed over 30 days. That is a budget of about 43 minutes.

    C

    Try it: availability budget

    recorded from the lab
    target
    down per daydown per weekdown per monthdown per year
    1.44 min10.1 min43.2 min8.76 h
    each part
    3 parts at 99.9%, 1 copy each: 99.7003%, or 2.16 h down a month. Misses 99.9%.

    Three parts in a row at 99.9% give only 99.7%. Two copies of each part bring the path back above 99.9%.

    D

    The arithmetic

    pseudo code
    availabilitypseudo code
    downtime(target, period) = (1 - target) × period1
      // 99.9% over 30 days: 0.001 × 2,592,000 s = 2,592 s = 43.2 min
    
    redundant(a, copies) = 1 - (1 - a)^copies   // down only if every copy is down2
    chain(a, parts)      = a^parts               // up only if every part is up3
    
      // 3 parts in a row at 99.9%:  0.999^3        = 99.7%4
      // 2 copies of each part:      (1 - 0.001^2)^3 = 99.9997%5
    1. 1The error budget. Deploys, failovers and incidents all spend it.
    2. 2Assumes copies fail independently. Copies in one zone, or on one bad deploy, fail together.
    3. 3Each part a request must pass lowers the availability of the whole path.
    4. 4Three parts at 99.9% each: the path is down about three times as often.
    5. 5Redundancy repays the chain: each part is now down only when both copies are down.
    Tested source Go: downtime, copies, chains
    Go: downtime, copies, chainsgo
    
    // Downtime is the time a service may be down in a period and still meet the target.
    func Downtime(availability, periodSeconds float64) float64 {
      return (1 - availability) * periodSeconds
    }
    
    // Redundant is the availability of n copies when any one copy is enough and copies fail
    // independently: the service is down only when every copy is down.
    func Redundant(availability float64, n int) float64 {
      return 1 - math.Pow(1-availability, float64(n))
    }
    
    // Chain is the availability of components that a request needs one after another: the request
    // succeeds only when every component is up.
    func Chain(availability float64, n int) float64 {
      return math.Pow(availability, float64(n))
    }
    

    Downtime is one minus the target, times the period. Parts in a row multiply; copies in parallel multiply their failures.

    E

    Resources and calls

    the booking API
    callforretry safe bysuccesserrors
    GET /hotels/{id}/availability?from=&to=1 searcha read200400, 404
    POST /holds {room_type, nights}2 holdthe hold token201 + expires_at409 held, 422 dates
    DELETE /holds/{token}2 releaseDELETE is idempotent204404
    POST /bookings {hold_token} + Idempotency-Key3 confirmthe idempotency key201409 sold out, 422 key reused
    GET /bookings?cursor=&limit=4 lista read200 + next_cursor400 bad cursor
    • Use plural nouns for collections and an opaque id for each item.
    • Put the version in the path (/v1) or a header, and never break a published field.
    • Return the created resource and its Location, so the client needs no second read.

    Paths are nouns and methods are verbs. Every call that creates or pays is safe to retry.

    F

    Idempotency keys

    a retry after a lost answer
    ClientBooking servicePostgresPOST /bookings · key Kclaim K, take nights, commitbooking 981✕201 Createdretry: POST /bookings · key KK exists: read bookingbooking 981201 Created · booking 981timeout: did it book?✓ one booking, one charge, same answer
    caseanswer
    Same key, same body, finishedThe stored result, again
    Same key, first request still running409
    Same key, different body422
    Key missing where required400

    Codes as the IETF HTTP API working group draft for the Idempotency-Key header suggests.

    The client makes one key per user action and sends it on every retry. The server stores the key with the result in the same transaction.

    G

    Pagination

    offset or cursor
    ✕ Offset: LIMIT 3 OFFSET 3✓ Cursor: after id 7, LIMIT 3page 1987654page 2 (after 10)1098765page 1987654page 2 (after 10)109876547 shows twiceno repeat
    offsetcursor
    Stable when items arriveNot approvedApproved
    Cost of a deep pageSkips every rowIndex seek
    Jump to page 50ApprovedNot approved
    Any sort orderApprovedUnique sort key

    Postgres computes and discards the rows an OFFSET skips (PostgreSQL documentation, LIMIT and OFFSET).

    I use a cursor: the last key the client saw. Inserts do not shift it, and each page is an index seek.

    H

    Status codes

    the ones a design uses
    codewhenretry
    200A read, or a retry that returns the stored resultSafe
    201A hold or a booking was createdNot needed
    202Accepted; the work finishes later. Poll the status URL.Not needed
    204Done, nothing to return (a release)Not needed
    400Malformed requestNot approved
    401 / 403Not signed in / not allowedNot approved
    404No such resourceNot approved
    409Conflicts with the state: night held, sold outNot approved
    422Well formed but invalid: check-out before check-inNot approved
    429Rate limitedAfter Retry-After
    500 / 503Server fault or overloadSame key, backoff

    A 4xx means the client must change the request. A 5xx or 429 can be retried with the same key, after a wait.

    I

    Sync or async

    does the user wait for it?
    workanswerwhy
    Hold a roomsync · 201The guest needs the token to pay.
    Confirm and chargesync · 201The guest must know now; p99 < 1 s.
    Confirmation emailQueue asyncNobody waits for it. Retry until sent.
    Search index updateQueue asyncSearch may be 60 s stale.
    Refund after a cancelasync · 202The provider answers later. Give a status URL.
    Export a year of bookingsasync · 202Large; deliver a link when ready.
    • Async work goes through a durable queue, so a crash does not lose it.
    • Write the job in the same transaction as the data (an outbox), so neither is lost alone.
    • 202 returns a status URL. The client polls it, or a webhook calls back.

    The guest waits for the booking, so confirm is synchronous. The email and the search index update go through a queue.

    J

    Worked example: booking

    requirements, then API, then the first diagram
    1 Requirements2 API3 First diagram1 Search rooms for datesGET /hotels/{id}/availability2 Hold a room while payingPOST /holds3 Pay and confirm, oncePOST /bookings · Idempotency-Key4 See my bookingsGET /bookings?cursor=…non-functional, as numbersno double bookingconfirm p99 < 1 ssearch ≤ 60 s stale99.9% a monthBookingservicestateless2+ zonesCache and searchTTL 60 s: stale is fineRedisholds, TTL 10 minPostgresone transaction:no double bookingPayment providersame key, one chargeconfirm p99 < 1 s: one transaction plus one payment call

    Each verb is a call, each call reaches one store, and each number lands on the component that meets it.

    K

    Each call on the first diagram

    click a call; its path lights up
    Guestbrowser or appLoad balancerTLS, routingBooking servicethe API, N copiesCache and search≤ 60 s staleRedisholds, TTL 10 minPostgresbookings, idempotency keysPayment providersame key, one charge

    Step 1: Search

    • Read availability from the cache and search index.
    • Results may be up to 60 s stale. Later steps check again.
    • Answer 200 with a page and a cursor.

    If it fails

    The cache is down: read from a Postgres replica, and shed load if it saturates. Nothing is booked yet.

    Search can be stale, the hold is fast and can be lost, and the confirm transaction decides. The list reads the guest's own bookings by key.

    L

    Drill

    predict, then reveal

    0 of 8 known

    1. The interviewer says "it must be fast". What do you write on the board?

    2. The target is 99.9% over 30 days. How much downtime is that?

    3. A request passes through 3 services in a row, each 99.9%. What can the whole path promise?

    4. POST /bookings times out on the client. What should the client do?

    5. A client reuses an idempotency key with a different body. What does the server answer?

    6. Why use a cursor, not an offset, for "my bookings" or a feed?

    7. Another guest holds the night. Which status code, and what body?

    8. Should the confirm call send the confirmation email before it answers?

    M

    Numbers to say

    derived in the lab
    99%
    7.2 hours down a month.
    99.9%
    43.2 minutes down a month.
    99.99%
    4.32 minutes down a month.
    chain
    3 parts at 99.9% in a row: 99.7003%.
    copies
    The same with 2 copies of each: 99.9997%.
    API
    4 to 6 calls for the first version.
    codes
    409 conflict, 422 invalid, 429 slow down, 202 later.

    A month is 30 days. Copies assume independent failures.