System Design
B9

Blob storage: objects, uploads and durability

Store files, images and video outside the database. This sheet covers the object model, direct and resumable uploads, deduplication, and the arithmetic behind durability claims.

Not startedSaved in this browser only.
  1. 1Bytes go to an object store; the database keeps a row with the key. In the lab a 1 MiB bytea row wrote 1.07 MiB of log.
  2. 2The client uploads straight to the store with a presigned URL. The service signs; it never carries the bytes.
  3. 3Erasure coding 6+3 stores 1.5 times the data and survives 3 lost disks. Three replicas store 3 times and survive 2.
  4. 4Content-defined chunks keep deduplication working after an insert: 1 new chunk instead of 33.
B9
    A

    The problem: bytes in the database

    measured in the lab Postgres
    photo1 MiBPostgresbytea, TOASTlog: 1.07 MiBreplica 1+1.07 MiBreplica 2+1.07 MiBbackup, log archive+1.07 MiB✕ 64 files of 1 MiB: 68 MiB of log, copied 3 more timesA metadata row with a key to the object store: 231 bytes of log.
    • 56.5 inserts of 1 MiB a second, against 3,190 metadata rows.
    • Backups, restores and replica builds all grow with the files.

    A 1 MiB file in Postgres writes about 1 MiB of log, and that log goes to every replica and backup. I store the bytes elsewhere and keep a 231-byte row.

    B

    Options

    where the bytes go
    methodwherestatus
    bytea columnPostgresFew, small files
    Files on the app server's diskServiceNot approved
    Object store, upload through the serviceObject storeSmall files
    Object store, presigned direct uploadObject storeApproved
    Multipart upload for large filesObject storeOver 100 MB
    Metadata row with the object keyPostgresApproved
    CDN in front for readsCDNApproved
    Overwrite one key in placeObject storeNot approved

    The app server's disk is lost with the server and is not shared between copies of the service.

    I put bytes in an object store, metadata in Postgres, and let clients upload directly with a presigned URL.

    C

    The object model

    bucket, key, object
    bucket: mediakey → objectu/7/9f1c…/cat.png2.1 MB · image/pngu/7/40ab…/trip.mp41.8 GB · video/mp4u/8/c2e9…/cv.pdf310 kB · application/pdfone objectbytes, immutableContent-Type image/pngContent-Length 2,104,331ETag, checksum sha256x-amz-meta-* your fieldsversion id (if on)✓ GET, PUT, DELETE, LIST by prefix, all by key over HTTP✕ no edit in place, no append, no rename: PUT replaces the object✕ no transaction across two keys; no query on metadata
    consistency
    S3 gives strong read-after-write for PUT and DELETE. A LIST right after a PUT shows the object.
    writers
    Two PUTs to one key at once: the last writer wins. A reader sees one whole object, never a mix.
    keys
    Flat. A slash is a character; LIST by prefix shows folders that do not exist.
    request rate
    S3: at least 3,500 writes and 5,500 reads a second per prefix, and more prefixes add more.
    size
    S3: up to 48.8 TiB an object, uploaded in up to 10,000 parts.
    durability
    S3 Standard is designed for 11 nines a year, across at least 3 Availability Zones.

    Figures are from the AWS S3 documentation. Other stores publish their own; check them before you quote one.

    An object store is a flat map from key to an immutable object over HTTP. It reads its own writes, but it has no transactions across keys and no queries on metadata.

    D

    Schema: metadata in Postgres

    the row points to the object
    filessql
    -- The database keeps what you query: owner, name, size, state.
    -- The bytes live in the object store under object_key.
    CREATE TABLE files (
      id           bigserial   PRIMARY KEY,
      owner_id     bigint      NOT NULL,
      name         text        NOT NULL,
      content_type text        NOT NULL,
      object_key   text        NOT NULL UNIQUE1,
      size_bytes   bigint,
      sha256       text3,
      status       text        NOT NULL DEFAULT 'pending'2
                               CHECK (status IN ('pending', 'ready')),
      created_at   timestamptz NOT NULL DEFAULT now()
    );
    
    -- The sweeper finds abandoned uploads without reading ready files.
    CREATE INDEX files_pending ON files (created_at) WHERE status = 'pending'4;
    1. 1One row per object. The key is new for every upload, so no two uploads write the same object.
    2. 2Pending until the store confirms the bytes. Readers list only ready rows.
    3. 3Taken from the store, not from the client, at finish.
    4. 4A partial index: the sweeper reads only pending rows, however many ready rows exist.
    begin, finish, sweeppseudo code
    begin(owner, name, type):
      INSERT row: status pending, key u/owner/uuid1
      RETURN a presigned PUT for the key, 15 minutes
    the client PUTs the bytes to the object store
    finish(id):
      HEAD key -> size, sha256           // what really arrived2
      UPDATE status = ready WHERE status = pending3
    sweeper, every hour:
      pending for 24 h4: delete the object, then the row
    1. 1A fresh key per upload: no overwrite, no race between two uploads of one file name.
    2. 2The store reports size and hash. A client cannot claim a file it did not send.
    3. 3A retried finish changes no row. The lab calls finish twice.
    4. 4The client never called finish: delete both sides.
    Tested source Go: begin, finish, sweep · SQL: begin and finish
    Go: begin, finish, sweepgo
    // Begin records a pending file and returns a URL the client can PUT the bytes to, until it
    // expires. The application never handles the bytes.
    func (u Uploads) Begin(ctx context.Context, owner int64, name, contentType string) (int64, string, error) {
      var id int64
      var key string
      if err := u.DB.QueryRow(ctx, stmts["begin_upload"], owner, name, contentType).Scan(&id, &key); err != nil {
        return 0, "", fmt.Errorf("begin upload: %w", err)
      }
      url := Presign(u.Store.Secret, http.MethodPut, "/"+key, contentType, u.Store.Now().Add(u.TTL))
      return id, url, nil
    }
    
    // Finish checks that the object exists, then marks the file ready with the size and hash the
    // store reports. Calling it twice is harmless.
    func (u Uploads) Finish(ctx context.Context, id int64) error {
      var key string
      if err := u.DB.QueryRow(ctx, stmts["key_of"], id).Scan(&key); err != nil {
        return fmt.Errorf("file %d: %w", id, err)
      }
      o, ok := u.Store.Head(key)
      if !ok {
        return fmt.Errorf("file %d: %w", id, ErrNotUploaded)
      }
      if _, err := u.DB.Exec(ctx, stmts["finish_upload"], id, len(o.Data), o.SHA256); err != nil {
        return fmt.Errorf("finish file %d: %w", id, err)
      }
      return nil
    }
    
    // Sweep deletes uploads still pending after age: the row, and the object if a client sent one
    // but never called Finish.
    func (u Uploads) Sweep(ctx context.Context, age time.Duration) (int, error) {
      rows, err := u.DB.Query(ctx, stmts["stale_pending"], age.Seconds())
      if err != nil {
        return 0, fmt.Errorf("find stale uploads: %w", err)
      }
      type stale struct {
        id  int64
        key string
      }
      var list []stale
      for rows.Next() {
        var s stale
        if err := rows.Scan(&s.id, &s.key); err != nil {
          rows.Close()
          return 0, fmt.Errorf("scan stale upload: %w", err)
        }
        list = append(list, s)
      }
      rows.Close()
      if err := rows.Err(); err != nil {
        return 0, fmt.Errorf("read stale uploads: %w", err)
      }
      for _, s := range list {
        u.Store.Delete(s.key)
        if _, err := u.DB.Exec(ctx, stmts["delete_file"], s.id); err != nil {
          return 0, fmt.Errorf("delete file %d: %w", s.id, err)
        }
      }
      return len(list), nil
    }
    
    SQL: begin and finishsql
    INSERT INTO files (owner_id, name, content_type, object_key)
    VALUES ($1, $2, $3, 'u/' || $1::bigint || '/' || gen_random_uuid())
    RETURNING id, object_key;
    
    -- Only a pending row moves to ready, so a repeated finish changes nothing.
    UPDATE files SET status = 'ready', size_bytes = $2, sha256 = $3
    WHERE id = $1 AND status = 'pending';

    The row holds owner, name, size, hash and state. A pending row becomes ready only after the store confirms the object, and a sweeper removes abandoned uploads.

    E

    The upload and read path

    click a step; its path lights up
    Clientbrowser or appUpload servicesigns, never proxiesCDNedge cachePostgresfile rowsObject storebytes, lifecycle

    Step 1: Begin

    • Check the user may upload. Insert a pending row with a fresh key.
    • Sign a PUT URL for that key and content type, valid 15 minutes.

    If it fails

    The insert fails: no URL is issued, and nothing is stored.

    The service only signs and records. Bytes go from the client to the store, and reads come through the CDN.

    F

    Capabilities used

    what each tool gives you
    toolcapabilitywhat it gives this designalso used for
    Object storePUT, GET, DELETE, LIST by key over HTTPAny client can read and write without a driver.Static sites, data lakes
    Object storePresigned URLs (S3: up to 7 days)A client uploads or downloads one key, for a while, with no credentials.Private downloads
    Object storeMultipart upload: parts of 5 MiB to 5 GiBParallel, resumable uploads of large files.Server-side copy of big objects
    Object storeChecksums on upload (SHA-256, CRC32C)The store refuses bytes damaged on the way.Integrity audits
    Object storeStrong read-after-write (S3)Finish can HEAD the key right after the PUT.Pipelines that list then read
    Object storeVersioningAn overwrite or delete keeps the old version.Undo, ransomware recovery
    Object storeStorage classes and lifecycle rulesMove cold objects to cheaper classes; expire old ones; abort stale uploads.Log retention
    Object storeReplication across regionsA copy in a second region, asynchronously.Disaster recovery, data residency
    Object storeEvent notificationsA new object can trigger finish or a thumbnail job.Media pipelines
    CDNEdge cache, signed URLs or cookiesRepeat reads never reach the store.Static assets (sheet B5)
    PostgresRows, UNIQUE key, CHECK on statusWho owns what, in which state, queryable.Every listing page
    PostgresPartial index on pending rowsThe sweeper's scan stays small.Job queues
    Postgresbytea and TOASTLimit Works, but every byte goes through the log, replicas and backups.

    The object store gives me presigned URLs, multipart uploads, strong read-after-write, versioning and lifecycle rules. Postgres gives me the rows I query; the CDN serves repeats.

    G

    Presigned URLs

    pseudo code
    sign and verifypseudo code
    presign(method, key, type, ttl):       // on the app server2
      expires = now + ttl
      sig = HMAC-SHA256(secret, method, key, type, expires)1
      RETURN "/" + key + "?expires=" + expires + "&sig=" + sig
    
    the store, on every request:
      IF sig != HMAC(secret, the request's parts): 403
      IF now > expires: 4033
      serve the PUT or GET
    1. 1Change any signed part and the HMAC no longer matches.
    2. 2Signing is local: no call to the store, no bytes through the service.
    3. 3The store checks expiry when the request starts. A long download that started in time finishes.
    Tested source Go: presign and verify
    Go: presign and verifygo
    // Presign returns a URL path with an expiry and a signature. The signature is an HMAC of the
    // method, the path, the content type and the expiry, keyed with a secret only the application
    // and the store know. The client can use the URL; it cannot change any signed part.
    func Presign(secret []byte, method, path, contentType string, expires time.Time) string {
      exp := strconv.FormatInt(expires.Unix(), 10)
      q := url.Values{"X-Expires": {exp}, "X-Signature": {sign(secret, method, path, contentType, exp)}}
      return path + "?" + q.Encode()
    }
    
    func sign(secret []byte, method, path, contentType, exp string) string {
      m := hmac.New(sha256.New, secret)
      fmt.Fprintf(m, "%s\n%s\n%s\n%s", method, path, contentType, exp)
      return hex.EncodeToString(m.Sum(nil))
    }
    
    // Verify checks a presigned request at time now: the signature first, then the expiry.
    func Verify(secret []byte, method, path, contentType string, q url.Values, now time.Time) error {
      exp := q.Get("X-Expires")
      want := sign(secret, method, path, contentType, exp)
      if !hmac.Equal([]byte(want), []byte(q.Get("X-Signature"))) { // constant time
        return ErrBadSignature
      }
      t, err := strconv.ParseInt(exp, 10, 64)
      if err != nil {
        return fmt.Errorf("expiry %q: %w", exp, ErrBadSignature)
      }
      if now.Unix() > t {
        return ErrExpired
      }
      return nil
    }
    
    requeststore answers
    As signed, within 15 minutes, twice200
    After expiry403
    Another key, type or method403
    A later expiry written into the URL403
    Signed with the wrong secret403

    The service signs the method, key, content type and expiry with a secret it shares with the store. The client can use the URL but not change it.

    H

    Multipart and resumable uploads

    pseudo code
    multipart uploadpseudo code
    id = create_upload(key)                // nothing visible yet1
    FOR EACH part n, in parallel:          // 5 MiB to 5 GiB each2
      upload_part(id, n, bytes, sha256)    // bad checksum: refused3
    after a crash:
      have = list_parts(id)                // send only the rest4
    complete(id, [1 .. N])                 // appears all at once
    1. 1Readers see the old object, or none, until Complete.
    2. 2S3 limits. Only the last part may be smaller. Up to 10,000 parts.
    3. 3The lab flips one bit in a part; the store refuses it.
    4. 4A 4 GB upload that drops at 3 GB resends 1 GB, not 4.
    Tested source Go: multipart
    Go: multipartgo
    // CreateUpload starts a multipart upload for key. Nothing is visible under key until Complete.
    func (s *ObjectStore) CreateUpload(key string) string {
      s.mu.Lock()
      defer s.mu.Unlock()
      s.nextID++
      id := "u" + strconv.Itoa(s.nextID)
      s.uploads[id] = &upload{key: key, parts: map[int][]byte{}}
      return id
    }
    
    // UploadPart stores one part. Parts can arrive in any order and in parallel; sending a part again
    // replaces it. The client sends the SHA-256 of the part, and a corrupted part is refused.
    func (s *ObjectStore) UploadPart(id string, n int, data []byte, sum string) error {
      if digest(data) != sum {
        return fmt.Errorf("part %d: %w", n, ErrBadChecksum)
      }
      s.mu.Lock()
      defer s.mu.Unlock()
      u, ok := s.uploads[id]
      if !ok {
        return fmt.Errorf("upload %s: %w", id, ErrNoUpload)
      }
      u.parts[n] = slices.Clone(data)
      return nil
    }
    
    // ListParts returns the part numbers the store holds, so a client can resume after a crash.
    func (s *ObjectStore) ListParts(id string) ([]int, error) {
      s.mu.Lock()
      defer s.mu.Unlock()
      u, ok := s.uploads[id]
      if !ok {
        return nil, fmt.Errorf("upload %s: %w", id, ErrNoUpload)
      }
      nums := make([]int, 0, len(u.parts))
      for n := range u.parts {
        nums = append(nums, n)
      }
      slices.Sort(nums)
      return nums, nil
    }
    
    // Complete joins the listed parts in order into one object. The object appears all at once.
    func (s *ObjectStore) Complete(id string, parts []int) (Object, error) {
      s.mu.Lock()
      defer s.mu.Unlock()
      u, ok := s.uploads[id]
      if !ok {
        return Object{}, fmt.Errorf("upload %s: %w", id, ErrNoUpload)
      }
      var data []byte
      for i, n := range parts {
        p, ok := u.parts[n]
        if !ok {
          return Object{}, fmt.Errorf("part %d: %w", n, ErrMissingPart)
        }
        if i < len(parts)-1 && len(p) < s.MinPart {
          return Object{}, fmt.Errorf("part %d is %d bytes: %w", n, len(p), ErrSmallPart)
        }
        data = append(data, p...)
      }
      o := Object{Data: data, SHA256: digest(data), ContentType: "application/octet-stream"}
      s.objects[u.key] = o
      delete(s.uploads, id)
      return o, nil
    }
    
    • S3 suggests multipart when an object reaches 100 MB.
    • Parts in flight cost storage. A lifecycle rule aborts uploads left incomplete.

    Large files go up in parts, in parallel, each with a checksum. After a drop, I list the parts the store holds and send only the rest.

    I

    Try it: edit a file, see which chunks change

    recorded chunk lists, 256 KiB file
    Fixed, every 8 KiBupload 33 of 33 chunks, 256 KiB (100%)
    beforeafter

    9ef7c5a0 · 8,192 B 1421bb3a · 8,192 B 6e3d6913 · 8,192 B 9db3e2e8 · 8,192 B … 29 more

    Content-defined, 2 to 32 KiBupload 1 of 26 chunks, 4.2 KiB (2%)
    beforeafter

    1c84da2d · 4,265 B

    • hash already stored: send nothing
    • new hash: upload
    • where the edit is

    Seeded text file of 256 KiB. Content-defined chunks: 2 KiB minimum, 8 KiB expected after it, 32 KiB maximum; about 10 KiB on average here.

    Fixed chunks break after an insert because every boundary shifts. Content-defined chunks cut where the content says, so only the chunk around the edit changes.

    J

    Content-defined chunks and dedupe

    pseudo code, then 5 versions
    chunk, hash, store oncepseudo code
    chunks(file):
      FOR EACH byte b:
        h = (h << 1) + GEAR[b]1      // rolling hash, last 64 bytes
        IF size >= MIN AND top 13 bits of h are 02: cut
        IF size == MAX: cut
    store(version):
      FOR EACH chunk: IF sha256(chunk) is new3: upload it
      save the version as its list of hashes
    1. 1A gear hash: each shift pushes an old byte out, so h depends on the last 64 bytes only.
    2. 2True once in 8,192 bytes on average. The cut follows the content, so it moves with an insert.
    3. 3Content addressing: the hash is the chunk name. A chunk already stored costs nothing.
    Tested source Go: content-defined chunks · Go: fixed chunks · Go: dedupe store · Go: the claims checked
    Go: content-defined chunksgo
    // Chunks cuts data where the gear hash of the last 64 bytes has its top 13 bits all zero. That
    // happens once in 8,192 bytes on average. Min and Max bound the chunk size.
    func (c *CDC) Chunks(data []byte) []Chunk {
      var out []Chunk
      start := 0
      for start < len(data) {
        end := min(start+c.Max, len(data))
        var h uint64
        for i := start; i < end; i++ {
          h = h<<1 + c.gear[data[i]] // each shift pushes one old byte out of the 64-bit window
          if i+1-start >= c.Min && h&c.mask == 0 {
            end = i + 1
            break
          }
        }
        out = append(out, newChunk(data, start, end-start))
        start = end
      }
      return out
    }
    
    Go: fixed chunksgo
    // FixedChunks cuts data every size bytes. Insert one byte at the front and every boundary after
    // it moves, so every chunk after the edit gets a new hash.
    func FixedChunks(data []byte, size int) []Chunk {
      var out []Chunk
      for off := 0; off < len(data); off += size {
        out = append(out, newChunk(data, off, min(size, len(data)-off)))
      }
      return out
    }
    
    Go: dedupe storego
    // Put stores the chunks of one file version. A chunk whose hash is already stored costs nothing:
    // the version only records the hash. It returns which chunks were new and the bytes they added.
    func (s *Store) Put(chunks []Chunk) (isNew []bool, added int) {
      isNew = make([]bool, len(chunks))
      for i, c := range chunks {
        if _, ok := s.chunks[c.full]; ok {
          continue
        }
        s.chunks[c.full] = c.Len
        s.Bytes += c.Len
        added += c.Len
        isNew[i] = true
      }
      return isNew, added
    }
    
    Go: the claims checkedgo
    // The claims the sheet makes about the two chunkers.
    size := len(base)
    if r := res["insert"]; r.Fixed.NewBytes < size*9/10 || r.CDC.NewChunks > 2 {
      t.Errorf("insert: fixed adds %d bytes, CDC %d chunks; want 90%% of the file, then at most 2 chunks", r.Fixed.NewBytes, r.CDC.NewChunks)
    }
    if r := res["delete"]; r.Fixed.NewBytes < size*7/10 || r.CDC.NewChunks > 2 {
      t.Errorf("delete: fixed adds %d bytes, CDC %d chunks", r.Fixed.NewBytes, r.CDC.NewChunks)
    }
    for _, id := range []string{"overwrite", "append"} {
      if r := res[id]; r.Fixed.NewChunks > 2 || r.CDC.NewChunks > 2 {
        t.Errorf("%s: fixed %d new chunks, CDC %d; want at most 2 each", id, r.Fixed.NewChunks, r.CDC.NewChunks)
      }
    }
    if out.CDCAll*3 > out.Logical || out.FixedAll < 2*out.CDCAll {
      t.Errorf("history: %d logical, fixed %d, CDC %d; want CDC under a third and fixed twice CDC", out.Logical, out.FixedAll, out.CDCAll)
    }
    versionsizefixed addscontent-defined adds
    v1, the original256 KiB256 KiB256 KiB
    v2, insert256 KiB256 KiB4 KiB
    v3, overwrite256 KiB8 KiB32 KiB
    v4, append260 KiB4 KiB10 KiB
    v5, delete257 KiB201 KiB18 KiB
    stored1,285 KiB725 KiB321 KiB

    Each version applies one edit to the one before. Without dedupe, the store holds every byte of every version.

    • Sync clients and backup tools upload only the new chunks.
    • An overwrite inside one large chunk re-sends that whole chunk: smaller chunks save bytes and cost more hashes.

    Over 5 versions of one file, content-defined chunks stored 25% of the bytes; fixed chunks stored 56%.

    K

    Replication against erasure coding

    bytes to scale, recorded rebuilds
    one object, bytes stored to scale (one copy = the bracket)3 replicas3× · survives 2 lost3 full copies on 3 disksReed-Solomon 6+31.5× · survives 3 lost6 data shards of 1/6 + 3 parity shards, on 9 disksReed-Solomon 10+41.4× · survives 4 lost10 data shards of 1/10 + 4 parity shards, on 14 disks
    Reed-Solomonpseudo code
    encode(object, k, m):
      split into k data shards
      parity[j] = SUM of E[k + j][i] * data[i]  // GF(256) bytes1
      put the k + m shards on k + m disks, racks apart2
    rebuild(any k shards):
      rows = the k matching rows of E
      data = inverse(rows) * shards             // always invertible3
    1. 1Arithmetic on bytes where every non-zero byte has an inverse, so matrices invert.
    2. 2Shards of one object on one rack fail together. Spread them across racks or zones.
    3. 3Built from a Vandermonde matrix, any k rows of E invert. The lab tries every subset.
    Tested source Go: build the code · Go: rebuild
    Go: build the codego
    // NewCode builds the encoding matrix. A Vandermonde matrix has the property that any K of its
    // rows are invertible. Multiplying it by the inverse of its top K rows keeps that property and
    // makes the top K rows the identity, so the data shards are stored as they are.
    func NewCode(k, m int) (*Code, error) {
      n := k + m
      if k < 1 || m < 0 || n > 256 {
        return nil, fmt.Errorf("code %d+%d: need k >= 1 and k + m <= 256", k, m)
      }
      v := newMatrix(n, k)
      for i := range n {
        for j := range k {
          v[i][j] = gfPow(byte(i), j)
        }
      }
      topInv, err := v[:k].invert()
      if err != nil {
        return nil, fmt.Errorf("code %d+%d: %w", k, m, err)
      }
      return &Code{K: k, M: m, enc: v.mul(topInv)}, nil
    }
    
    Go: rebuildgo
    // Reconstruct fills in the missing (nil) shards. Any K surviving shards are K known rows of the
    // encoding matrix times the data, so inverting those K rows recovers the data shards. The lost
    // parity shards are then encoded again.
    func (c *Code) Reconstruct(shards [][]byte) error {
      var rows matrix
      var have [][]byte
      for i, s := range shards {
        if s != nil && len(rows) < c.K {
          rows = append(rows, c.enc[i])
          have = append(have, s)
        }
      }
      if len(rows) < c.K {
        return fmt.Errorf("%d of %d shards left, need %d: %w", len(rows), c.K+c.M, c.K, ErrTooFewShards)
      }
      dec, err := rows.invert()
      if err != nil {
        return fmt.Errorf("decode: %w", err)
      }
      data := make([][]byte, c.K)
      for j := range c.K {
        data[j] = c.combine(dec[j], have)
      }
      for i := range shards {
        if shards[i] == nil {
          shards[i] = c.combine(c.enc[i], data)
        }
      }
      return nil
    }
    

    Lab: 10+4 rebuilt all 1,001 sets of 4 lost shards, and none of the 2,002 sets of 5.

    Reed-Solomon 6+3 cuts an object into 6 data shards and adds 3 parity shards. Any 6 of the 9 rebuild it, at 1.5 times the bytes.

    L

    Try it: lose disks

    recorded rebuilds and durability arithmetic
    layout
    disks failing a year
    time to rebuild a disk
    0 of 9 disks down: the object rebuilds from any 6 survivors. Lab: 1 of 1 such failure sets rebuilt.
    stored per byte
    1.5×
    disks that may fail
    3
    lost a year
    4.14 × 10^-13
    nines
    12.4
    1. one disk fails in one repair windowp = 0.02 × 24 / 8760 = 5.48 × 10^-5
    2. more than 3 of 9 disks fail in one windowC(9, 4) × p^4 = 126 × 9.01 × 10^-18 = 1.14 × 10^-15; with every term: 1.14 × 10^-15
    3. windows a year8760 / 24 = 365
    4. lost a year365 × 1.14 × 10^-15 = 4.14 × 10^-13

    A model: disks fail independently, and each failed disk is rebuilt within the window. Correlated failures (a rack, a bad batch) cost more.

    durability, one objectpseudo code
    p        = AFR * repair_hours / 8760    // one disk, one window1
    lost     = P(more than m of n disks fail in one window)
             ≈ C(n, m + 1) * p^(m + 1)2
    per year = lost * 8760 / repair_hours
    1. 1AFR is the share of disks that fail in a year. Backblaze reported 1.36% over its fleet for 2025.
    2. 2The largest term. p is tiny, so the rest add little; the lab sums them all.
    Tested source Go: durability
    Go: durabilitygo
    // Durable works out the yearly chance to lose an object. The model: disks fail independently at
    // rate afr a year, and a failed disk is rebuilt within repairHours. An object is lost only if more
    // than M of its K+M disks fail inside one repair window.
    func Durable(s Scheme, afr, repairHours float64) Durability {
      n := s.K + s.M
      p := afr * repairHours / 8760 // one disk, one window
      var perWindow float64
      for i := s.M + 1; i <= n; i++ {
        perWindow += binom(n, i) * math.Pow(p, float64(i)) * math.Pow(1-p, float64(n-i))
      }
      windows := 8760 / repairHours
      annual := windows * perWindow
      return Durability{
        Overhead:  float64(n) / float64(s.K),
        Tolerates: s.M,
        P:         p,
        Windows:   windows,
        Lead:      binom(n, s.M+1) * math.Pow(p, float64(s.M+1)),
        PerWindow: perWindow,
        Annual:    annual,
        Nines:     -math.Log10(annual),
      }
    }
    

    At 2% AFR and a 24-hour rebuild, the yearly chance to lose an object is 6.0 × 10^-11 with 3 replicas and 4.1 × 10^-13 with 6+3.

    M

    Failure cases

    what breaks, and the fix
    eventresultwhy it is safe, or the fixsaved by
    The client uploads but never calls finishA pending row and an orphan object.The sweeper deletes both after 24 hours.Pending index
    Finish is retriedThe second call finds the row ready.The update needs status pending, so it changes nothing.Postgres
    An upload URL leaksAnyone can PUT that one key until expiry.Short expiry; key and type are signed; finish checks size and hash.Presigned URL
    The connection drops mid-uploadSome parts arrived.List parts, send the rest; nothing is visible before Complete.Multipart
    A part is damaged on the wayIts checksum does not match.The store refuses the part; the client sends it again.Checksum
    Uploads are abandoned mid-wayParts are stored and billed.A lifecycle rule aborts incomplete uploads after some days.Lifecycle
    More than m disks of one stripe fail in one windowThe object is lost.Fast rebuilds, shards across racks or zones, more parity.Erasure code
    Two uploads write one keyThe last writer wins.A new key per upload; the row names the current key.Unique key
    A row is deleted, the object staysStorage leaks.Delete the object first, or reconcile store listings against rows.Reconcile job
    N

    Scale ladder

    start simple; climb only on a signal
    Each step adds one component1store + row2+ presigned3+ multipart4+ CDN5+ tiers, regionsmore load →
    Capacity against demand101001k10k100kDemand, uploads: 116 requests per secondDemand, uploads116Demand, reads: 11,600 requests per secondDemand, reads11,600Postgres, 1 MiB bytea: 56.5 requests per secondPostgres, 1 MiB bytea56.5Postgres, metadata rows: 3,190 requests per secondPostgres, metadata rows3,190S3 writes, one prefix: ≥ 3,500, citedS3 writes, one prefix≥ 3,500, citedS3 reads, one prefix: ≥ 5,500, citedS3 reads, one prefix≥ 5,500, citedrequests per second, log scale
    stepaddit handlesmove up when you see
    1An object store and a metadata row. Uploads pass through the service.Small files. Postgres takes 3,190 metadata rows a second in the lab.Upload bandwidth or memory on the service; slow uploads hold its connections.
    2Presigned direct uploads.The store takes the bytes; the service signs and records.Files over 100 MB fail on mobile networks and restart from zero.
    3Multipart, resumable uploads.Large files in parallel parts; a drop resends one part.Reads repeat the same objects worldwide; egress and latency grow.
    4A CDN in front.Repeat reads served at the edge (sheet B5).Storage cost grows with old, cold objects.
    5Lifecycle tiers, chunk dedupe, a second region.Cold data on cheaper classes; versions share chunks; a region can fail.Top of the ladder.

    Demand example: 10 million uploads a day is 116 a second; 100 reads per upload is 11,600 a second, mostly from the CDN. S3 rates are per prefix and scale with more prefixes.

    I start with an object store and a metadata row from day one. I add direct uploads, then a CDN, then lifecycle tiers, as traffic and storage grow.

    O

    Drill

    predict, then reveal

    0 of 9 known

    1. Why not keep user photos in a bytea column?

    2. A client has a presigned PUT for u/7/a.png. Can it upload to u/8/b.png instead?

    3. A 4 GB upload drops after 3 of 8 parts. What does the client do?

    4. Insert 100 bytes near the start of a 256 KiB file. How many 8 KiB fixed chunks change?

    5. Why 6+3 erasure coding instead of 3 replicas?

    6. How do you estimate the yearly loss for 3 replicas at 2% AFR and 24 h repair?

    7. Which matters more for durability: a better disk or a faster rebuild?

    8. The upload succeeded but finish never arrived. What is left behind?

    9. Two clients PUT the same key at the same moment. What do readers see?

    P

    Numbers to say

    measured, derived or cited
    bytea
    1 MiB file: 1.07 MiB of log. Metadata row: 231 bytes.
    overhead
    3 replicas 3×; 6+3 1.5×; 10+4 1.4×.
    survives
    2, 3 and 4 lost disks.
    loss a year
    2% AFR, 24 h rebuild: 10, 12 and 15 nines.
    S3
    Designed for 11 nines; strong read-after-write.
    parts
    5 MiB to 5 GiB, up to 10,000; objects to 48.8 TiB.
    rate
    3,500 writes, 5,500 reads a second per prefix.
    URLs
    Presigned for up to 7 days.

    Measured on Postgres 16.14, a 16-thread laptop. Nines are derived from the model on this sheet. S3 figures are cited from its documentation.