Requirements to API: verbs, numbers, calls
Write what the system does as verbs and how well it does it as numbers. The verbs become the API; the numbers pick the components of the first diagram.
- 1A functional requirement is a verb. Each verb becomes one API call.
- 2A non-functional requirement is an adjective until you give it a number.
- 3Calls that create or pay take an idempotency key. Lists take a cursor.
- 4If the user does not need the result now, accept it (202) and finish later.
- Write 3 to 5 verbs. Agree on them before you draw a box.
- Write one line of what is out of scope.
- Turn each adjective into a number. The next panel shows how.
These four verbs are the functional scope. For each quality, I want a number we can test.
| adjective | as a number | it picks |
|---|---|---|
| fast | search p99 ≤ 300 ms; confirm p99 < 1 s | Cache for search |
| available | 99.9% of requests succeed in 30 days: 43.2 min down | 2 or more of every box |
| correct | 0 double bookings | Transaction and a CHECK |
| durable | 0 confirmed bookings lost (RPO 0) | Sync standby |
| fresh | search at most 60 s stale | Cache TTL of 60 s |
| scalable | 17.4 k requests/s at peak (R2) | Step 1: one primary |
SLI: what you measure (the share of good requests). SLO: the target for it. Measure at the load balancer, where users see the result.
By available I mean 99.9% of requests succeed over 30 days. That is a budget of about 43 minutes.
| down per day | down per week | down per month | down per year |
|---|---|---|---|
| 1.44 min | 10.1 min | 43.2 min | 8.76 h |
Three parts in a row at 99.9% give only 99.7%. Two copies of each part bring the path back above 99.9%.
downtime(target, period) = (1 - target) × period1
// 99.9% over 30 days: 0.001 × 2,592,000 s = 2,592 s = 43.2 min
redundant(a, copies) = 1 - (1 - a)^copies // down only if every copy is down2
chain(a, parts) = a^parts // up only if every part is up3
// 3 parts in a row at 99.9%: 0.999^3 = 99.7%4
// 2 copies of each part: (1 - 0.001^2)^3 = 99.9997%5- 1The error budget. Deploys, failovers and incidents all spend it.
- 2Assumes copies fail independently. Copies in one zone, or on one bad deploy, fail together.
- 3Each part a request must pass lowers the availability of the whole path.
- 4Three parts at 99.9% each: the path is down about three times as often.
- 5Redundancy repays the chain: each part is now down only when both copies are down.
Tested source Go: downtime, copies, chains
// Downtime is the time a service may be down in a period and still meet the target.
func Downtime(availability, periodSeconds float64) float64 {
return (1 - availability) * periodSeconds
}
// Redundant is the availability of n copies when any one copy is enough and copies fail
// independently: the service is down only when every copy is down.
func Redundant(availability float64, n int) float64 {
return 1 - math.Pow(1-availability, float64(n))
}
// Chain is the availability of components that a request needs one after another: the request
// succeeds only when every component is up.
func Chain(availability float64, n int) float64 {
return math.Pow(availability, float64(n))
}
Downtime is one minus the target, times the period. Parts in a row multiply; copies in parallel multiply their failures.
| call | for | retry safe by | success | errors |
|---|---|---|---|---|
| GET /hotels/{id}/availability?from=&to= | 1 search | a read | 200 | 400, 404 |
| POST /holds {room_type, nights} | 2 hold | the hold token | 201 + expires_at | 409 held, 422 dates |
| DELETE /holds/{token} | 2 release | DELETE is idempotent | 204 | 404 |
| POST /bookings {hold_token} + Idempotency-Key | 3 confirm | the idempotency key | 201 | 409 sold out, 422 key reused |
| GET /bookings?cursor=&limit= | 4 list | a read | 200 + next_cursor | 400 bad cursor |
- Use plural nouns for collections and an opaque id for each item.
- Put the version in the path (/v1) or a header, and never break a published field.
- Return the created resource and its Location, so the client needs no second read.
Paths are nouns and methods are verbs. Every call that creates or pays is safe to retry.
| case | answer |
|---|---|
| Same key, same body, finished | The stored result, again |
| Same key, first request still running | 409 |
| Same key, different body | 422 |
| Key missing where required | 400 |
Codes as the IETF HTTP API working group draft for the Idempotency-Key header suggests.
The client makes one key per user action and sends it on every retry. The server stores the key with the result in the same transaction.
| offset | cursor | |
|---|---|---|
| Stable when items arrive | Not approved | Approved |
| Cost of a deep page | Skips every row | Index seek |
| Jump to page 50 | Approved | Not approved |
| Any sort order | Approved | Unique sort key |
Postgres computes and discards the rows an OFFSET skips (PostgreSQL documentation, LIMIT and OFFSET).
I use a cursor: the last key the client saw. Inserts do not shift it, and each page is an index seek.
| code | when | retry |
|---|---|---|
| 200 | A read, or a retry that returns the stored result | Safe |
| 201 | A hold or a booking was created | Not needed |
| 202 | Accepted; the work finishes later. Poll the status URL. | Not needed |
| 204 | Done, nothing to return (a release) | Not needed |
| 400 | Malformed request | Not approved |
| 401 / 403 | Not signed in / not allowed | Not approved |
| 404 | No such resource | Not approved |
| 409 | Conflicts with the state: night held, sold out | Not approved |
| 422 | Well formed but invalid: check-out before check-in | Not approved |
| 429 | Rate limited | After Retry-After |
| 500 / 503 | Server fault or overload | Same key, backoff |
A 4xx means the client must change the request. A 5xx or 429 can be retried with the same key, after a wait.
| work | answer | why |
|---|---|---|
| Hold a room | sync · 201 | The guest needs the token to pay. |
| Confirm and charge | sync · 201 | The guest must know now; p99 < 1 s. |
| Confirmation email | Queue async | Nobody waits for it. Retry until sent. |
| Search index update | Queue async | Search may be 60 s stale. |
| Refund after a cancel | async · 202 | The provider answers later. Give a status URL. |
| Export a year of bookings | async · 202 | Large; deliver a link when ready. |
- Async work goes through a durable queue, so a crash does not lose it.
- Write the job in the same transaction as the data (an outbox), so neither is lost alone.
- 202 returns a status URL. The client polls it, or a webhook calls back.
The guest waits for the booking, so confirm is synchronous. The email and the search index update go through a queue.
Each verb is a call, each call reaches one store, and each number lands on the component that meets it.
Step 1: Search
- Read availability from the cache and search index.
- Results may be up to 60 s stale. Later steps check again.
- Answer 200 with a page and a cursor.
If it fails
The cache is down: read from a Postgres replica, and shed load if it saturates. Nothing is booked yet.
Search can be stale, the hold is fast and can be lost, and the confirm transaction decides. The list reads the guest's own bookings by key.
0 of 8 known
The interviewer says "it must be fast". What do you write on the board?
The target is 99.9% over 30 days. How much downtime is that?
A request passes through 3 services in a row, each 99.9%. What can the whole path promise?
POST /bookings times out on the client. What should the client do?
A client reuses an idempotency key with a different body. What does the server answer?
Why use a cursor, not an offset, for "my bookings" or a feed?
Another guest holds the night. Which status code, and what body?
Should the confirm call send the confirmation email before it answers?
- 99%
- 7.2 hours down a month.
- 99.9%
- 43.2 minutes down a month.
- 99.99%
- 4.32 minutes down a month.
- chain
- 3 parts at 99.9% in a row: 99.7003%.
- copies
- The same with 2 copies of each: 99.9997%.
- API
- 4 to 6 calls for the first version.
- codes
- 409 conflict, 422 invalid, 429 slow down, 202 later.
A month is 30 days. Copies assume independent failures.