Blob storage: objects, uploads and durability
Store files, images and video outside the database. This sheet covers the object model, direct and resumable uploads, deduplication, and the arithmetic behind durability claims.
- 1Bytes go to an object store; the database keeps a row with the key. In the lab a 1 MiB bytea row wrote 1.07 MiB of log.
- 2The client uploads straight to the store with a presigned URL. The service signs; it never carries the bytes.
- 3Erasure coding 6+3 stores 1.5 times the data and survives 3 lost disks. Three replicas store 3 times and survive 2.
- 4Content-defined chunks keep deduplication working after an insert: 1 new chunk instead of 33.
- 56.5 inserts of 1 MiB a second, against 3,190 metadata rows.
- Backups, restores and replica builds all grow with the files.
A 1 MiB file in Postgres writes about 1 MiB of log, and that log goes to every replica and backup. I store the bytes elsewhere and keep a 231-byte row.
| method | where | status |
|---|---|---|
| bytea column | Postgres | Few, small files |
| Files on the app server's disk | Service | Not approved |
| Object store, upload through the service | Object store | Small files |
| Object store, presigned direct upload | Object store | Approved |
| Multipart upload for large files | Object store | Over 100 MB |
| Metadata row with the object key | Postgres | Approved |
| CDN in front for reads | CDN | Approved |
| Overwrite one key in place | Object store | Not approved |
The app server's disk is lost with the server and is not shared between copies of the service.
I put bytes in an object store, metadata in Postgres, and let clients upload directly with a presigned URL.
- consistency
- S3 gives strong read-after-write for PUT and DELETE. A LIST right after a PUT shows the object.
- writers
- Two PUTs to one key at once: the last writer wins. A reader sees one whole object, never a mix.
- keys
- Flat. A slash is a character; LIST by prefix shows folders that do not exist.
- request rate
- S3: at least 3,500 writes and 5,500 reads a second per prefix, and more prefixes add more.
- size
- S3: up to 48.8 TiB an object, uploaded in up to 10,000 parts.
- durability
- S3 Standard is designed for 11 nines a year, across at least 3 Availability Zones.
Figures are from the AWS S3 documentation. Other stores publish their own; check them before you quote one.
An object store is a flat map from key to an immutable object over HTTP. It reads its own writes, but it has no transactions across keys and no queries on metadata.
-- The database keeps what you query: owner, name, size, state.
-- The bytes live in the object store under object_key.
CREATE TABLE files (
id bigserial PRIMARY KEY,
owner_id bigint NOT NULL,
name text NOT NULL,
content_type text NOT NULL,
object_key text NOT NULL UNIQUE1,
size_bytes bigint,
sha256 text3,
status text NOT NULL DEFAULT 'pending'2
CHECK (status IN ('pending', 'ready')),
created_at timestamptz NOT NULL DEFAULT now()
);
-- The sweeper finds abandoned uploads without reading ready files.
CREATE INDEX files_pending ON files (created_at) WHERE status = 'pending'4;- 1One row per object. The key is new for every upload, so no two uploads write the same object.
- 2Pending until the store confirms the bytes. Readers list only ready rows.
- 3Taken from the store, not from the client, at finish.
- 4A partial index: the sweeper reads only pending rows, however many ready rows exist.
begin(owner, name, type):
INSERT row: status pending, key u/owner/uuid1
RETURN a presigned PUT for the key, 15 minutes
the client PUTs the bytes to the object store
finish(id):
HEAD key -> size, sha256 // what really arrived2
UPDATE status = ready WHERE status = pending3
sweeper, every hour:
pending for 24 h4: delete the object, then the row- 1A fresh key per upload: no overwrite, no race between two uploads of one file name.
- 2The store reports size and hash. A client cannot claim a file it did not send.
- 3A retried finish changes no row. The lab calls finish twice.
- 4The client never called finish: delete both sides.
Tested source Go: begin, finish, sweep · SQL: begin and finish
// Begin records a pending file and returns a URL the client can PUT the bytes to, until it
// expires. The application never handles the bytes.
func (u Uploads) Begin(ctx context.Context, owner int64, name, contentType string) (int64, string, error) {
var id int64
var key string
if err := u.DB.QueryRow(ctx, stmts["begin_upload"], owner, name, contentType).Scan(&id, &key); err != nil {
return 0, "", fmt.Errorf("begin upload: %w", err)
}
url := Presign(u.Store.Secret, http.MethodPut, "/"+key, contentType, u.Store.Now().Add(u.TTL))
return id, url, nil
}
// Finish checks that the object exists, then marks the file ready with the size and hash the
// store reports. Calling it twice is harmless.
func (u Uploads) Finish(ctx context.Context, id int64) error {
var key string
if err := u.DB.QueryRow(ctx, stmts["key_of"], id).Scan(&key); err != nil {
return fmt.Errorf("file %d: %w", id, err)
}
o, ok := u.Store.Head(key)
if !ok {
return fmt.Errorf("file %d: %w", id, ErrNotUploaded)
}
if _, err := u.DB.Exec(ctx, stmts["finish_upload"], id, len(o.Data), o.SHA256); err != nil {
return fmt.Errorf("finish file %d: %w", id, err)
}
return nil
}
// Sweep deletes uploads still pending after age: the row, and the object if a client sent one
// but never called Finish.
func (u Uploads) Sweep(ctx context.Context, age time.Duration) (int, error) {
rows, err := u.DB.Query(ctx, stmts["stale_pending"], age.Seconds())
if err != nil {
return 0, fmt.Errorf("find stale uploads: %w", err)
}
type stale struct {
id int64
key string
}
var list []stale
for rows.Next() {
var s stale
if err := rows.Scan(&s.id, &s.key); err != nil {
rows.Close()
return 0, fmt.Errorf("scan stale upload: %w", err)
}
list = append(list, s)
}
rows.Close()
if err := rows.Err(); err != nil {
return 0, fmt.Errorf("read stale uploads: %w", err)
}
for _, s := range list {
u.Store.Delete(s.key)
if _, err := u.DB.Exec(ctx, stmts["delete_file"], s.id); err != nil {
return 0, fmt.Errorf("delete file %d: %w", s.id, err)
}
}
return len(list), nil
}
INSERT INTO files (owner_id, name, content_type, object_key)
VALUES ($1, $2, $3, 'u/' || $1::bigint || '/' || gen_random_uuid())
RETURNING id, object_key;
-- Only a pending row moves to ready, so a repeated finish changes nothing.
UPDATE files SET status = 'ready', size_bytes = $2, sha256 = $3
WHERE id = $1 AND status = 'pending';The row holds owner, name, size, hash and state. A pending row becomes ready only after the store confirms the object, and a sweeper removes abandoned uploads.
Step 1: Begin
- Check the user may upload. Insert a pending row with a fresh key.
- Sign a PUT URL for that key and content type, valid 15 minutes.
If it fails
The insert fails: no URL is issued, and nothing is stored.
The service only signs and records. Bytes go from the client to the store, and reads come through the CDN.
| tool | capability | what it gives this design | also used for |
|---|---|---|---|
| Object store | PUT, GET, DELETE, LIST by key over HTTP | Any client can read and write without a driver. | Static sites, data lakes |
| Object store | Presigned URLs (S3: up to 7 days) | A client uploads or downloads one key, for a while, with no credentials. | Private downloads |
| Object store | Multipart upload: parts of 5 MiB to 5 GiB | Parallel, resumable uploads of large files. | Server-side copy of big objects |
| Object store | Checksums on upload (SHA-256, CRC32C) | The store refuses bytes damaged on the way. | Integrity audits |
| Object store | Strong read-after-write (S3) | Finish can HEAD the key right after the PUT. | Pipelines that list then read |
| Object store | Versioning | An overwrite or delete keeps the old version. | Undo, ransomware recovery |
| Object store | Storage classes and lifecycle rules | Move cold objects to cheaper classes; expire old ones; abort stale uploads. | Log retention |
| Object store | Replication across regions | A copy in a second region, asynchronously. | Disaster recovery, data residency |
| Object store | Event notifications | A new object can trigger finish or a thumbnail job. | Media pipelines |
| CDN | Edge cache, signed URLs or cookies | Repeat reads never reach the store. | Static assets (sheet B5) |
| Postgres | Rows, UNIQUE key, CHECK on status | Who owns what, in which state, queryable. | Every listing page |
| Postgres | Partial index on pending rows | The sweeper's scan stays small. | Job queues |
| Postgres | bytea and TOAST | Limit Works, but every byte goes through the log, replicas and backups. |
The object store gives me presigned URLs, multipart uploads, strong read-after-write, versioning and lifecycle rules. Postgres gives me the rows I query; the CDN serves repeats.
presign(method, key, type, ttl): // on the app server2
expires = now + ttl
sig = HMAC-SHA256(secret, method, key, type, expires)1
RETURN "/" + key + "?expires=" + expires + "&sig=" + sig
the store, on every request:
IF sig != HMAC(secret, the request's parts): 403
IF now > expires: 4033
serve the PUT or GET- 1Change any signed part and the HMAC no longer matches.
- 2Signing is local: no call to the store, no bytes through the service.
- 3The store checks expiry when the request starts. A long download that started in time finishes.
Tested source Go: presign and verify
// Presign returns a URL path with an expiry and a signature. The signature is an HMAC of the
// method, the path, the content type and the expiry, keyed with a secret only the application
// and the store know. The client can use the URL; it cannot change any signed part.
func Presign(secret []byte, method, path, contentType string, expires time.Time) string {
exp := strconv.FormatInt(expires.Unix(), 10)
q := url.Values{"X-Expires": {exp}, "X-Signature": {sign(secret, method, path, contentType, exp)}}
return path + "?" + q.Encode()
}
func sign(secret []byte, method, path, contentType, exp string) string {
m := hmac.New(sha256.New, secret)
fmt.Fprintf(m, "%s\n%s\n%s\n%s", method, path, contentType, exp)
return hex.EncodeToString(m.Sum(nil))
}
// Verify checks a presigned request at time now: the signature first, then the expiry.
func Verify(secret []byte, method, path, contentType string, q url.Values, now time.Time) error {
exp := q.Get("X-Expires")
want := sign(secret, method, path, contentType, exp)
if !hmac.Equal([]byte(want), []byte(q.Get("X-Signature"))) { // constant time
return ErrBadSignature
}
t, err := strconv.ParseInt(exp, 10, 64)
if err != nil {
return fmt.Errorf("expiry %q: %w", exp, ErrBadSignature)
}
if now.Unix() > t {
return ErrExpired
}
return nil
}
| request | store answers |
|---|---|
| As signed, within 15 minutes, twice | 200 |
| After expiry | 403 |
| Another key, type or method | 403 |
| A later expiry written into the URL | 403 |
| Signed with the wrong secret | 403 |
The service signs the method, key, content type and expiry with a secret it shares with the store. The client can use the URL but not change it.
id = create_upload(key) // nothing visible yet1
FOR EACH part n, in parallel: // 5 MiB to 5 GiB each2
upload_part(id, n, bytes, sha256) // bad checksum: refused3
after a crash:
have = list_parts(id) // send only the rest4
complete(id, [1 .. N]) // appears all at once- 1Readers see the old object, or none, until Complete.
- 2S3 limits. Only the last part may be smaller. Up to 10,000 parts.
- 3The lab flips one bit in a part; the store refuses it.
- 4A 4 GB upload that drops at 3 GB resends 1 GB, not 4.
Tested source Go: multipart
// CreateUpload starts a multipart upload for key. Nothing is visible under key until Complete.
func (s *ObjectStore) CreateUpload(key string) string {
s.mu.Lock()
defer s.mu.Unlock()
s.nextID++
id := "u" + strconv.Itoa(s.nextID)
s.uploads[id] = &upload{key: key, parts: map[int][]byte{}}
return id
}
// UploadPart stores one part. Parts can arrive in any order and in parallel; sending a part again
// replaces it. The client sends the SHA-256 of the part, and a corrupted part is refused.
func (s *ObjectStore) UploadPart(id string, n int, data []byte, sum string) error {
if digest(data) != sum {
return fmt.Errorf("part %d: %w", n, ErrBadChecksum)
}
s.mu.Lock()
defer s.mu.Unlock()
u, ok := s.uploads[id]
if !ok {
return fmt.Errorf("upload %s: %w", id, ErrNoUpload)
}
u.parts[n] = slices.Clone(data)
return nil
}
// ListParts returns the part numbers the store holds, so a client can resume after a crash.
func (s *ObjectStore) ListParts(id string) ([]int, error) {
s.mu.Lock()
defer s.mu.Unlock()
u, ok := s.uploads[id]
if !ok {
return nil, fmt.Errorf("upload %s: %w", id, ErrNoUpload)
}
nums := make([]int, 0, len(u.parts))
for n := range u.parts {
nums = append(nums, n)
}
slices.Sort(nums)
return nums, nil
}
// Complete joins the listed parts in order into one object. The object appears all at once.
func (s *ObjectStore) Complete(id string, parts []int) (Object, error) {
s.mu.Lock()
defer s.mu.Unlock()
u, ok := s.uploads[id]
if !ok {
return Object{}, fmt.Errorf("upload %s: %w", id, ErrNoUpload)
}
var data []byte
for i, n := range parts {
p, ok := u.parts[n]
if !ok {
return Object{}, fmt.Errorf("part %d: %w", n, ErrMissingPart)
}
if i < len(parts)-1 && len(p) < s.MinPart {
return Object{}, fmt.Errorf("part %d is %d bytes: %w", n, len(p), ErrSmallPart)
}
data = append(data, p...)
}
o := Object{Data: data, SHA256: digest(data), ContentType: "application/octet-stream"}
s.objects[u.key] = o
delete(s.uploads, id)
return o, nil
}
- S3 suggests multipart when an object reaches 100 MB.
- Parts in flight cost storage. A lifecycle rule aborts uploads left incomplete.
Large files go up in parts, in parallel, each with a checksum. After a drop, I list the parts the store holds and send only the rest.
9ef7c5a0 · 8,192 B 1421bb3a · 8,192 B 6e3d6913 · 8,192 B 9db3e2e8 · 8,192 B … 29 more
1c84da2d · 4,265 B
- hash already stored: send nothing
- new hash: upload
- where the edit is
Seeded text file of 256 KiB. Content-defined chunks: 2 KiB minimum, 8 KiB expected after it, 32 KiB maximum; about 10 KiB on average here.
Fixed chunks break after an insert because every boundary shifts. Content-defined chunks cut where the content says, so only the chunk around the edit changes.
chunks(file):
FOR EACH byte b:
h = (h << 1) + GEAR[b]1 // rolling hash, last 64 bytes
IF size >= MIN AND top 13 bits of h are 02: cut
IF size == MAX: cut
store(version):
FOR EACH chunk: IF sha256(chunk) is new3: upload it
save the version as its list of hashes- 1A gear hash: each shift pushes an old byte out, so h depends on the last 64 bytes only.
- 2True once in 8,192 bytes on average. The cut follows the content, so it moves with an insert.
- 3Content addressing: the hash is the chunk name. A chunk already stored costs nothing.
Tested source Go: content-defined chunks · Go: fixed chunks · Go: dedupe store · Go: the claims checked
// Chunks cuts data where the gear hash of the last 64 bytes has its top 13 bits all zero. That
// happens once in 8,192 bytes on average. Min and Max bound the chunk size.
func (c *CDC) Chunks(data []byte) []Chunk {
var out []Chunk
start := 0
for start < len(data) {
end := min(start+c.Max, len(data))
var h uint64
for i := start; i < end; i++ {
h = h<<1 + c.gear[data[i]] // each shift pushes one old byte out of the 64-bit window
if i+1-start >= c.Min && h&c.mask == 0 {
end = i + 1
break
}
}
out = append(out, newChunk(data, start, end-start))
start = end
}
return out
}
// FixedChunks cuts data every size bytes. Insert one byte at the front and every boundary after
// it moves, so every chunk after the edit gets a new hash.
func FixedChunks(data []byte, size int) []Chunk {
var out []Chunk
for off := 0; off < len(data); off += size {
out = append(out, newChunk(data, off, min(size, len(data)-off)))
}
return out
}
// Put stores the chunks of one file version. A chunk whose hash is already stored costs nothing:
// the version only records the hash. It returns which chunks were new and the bytes they added.
func (s *Store) Put(chunks []Chunk) (isNew []bool, added int) {
isNew = make([]bool, len(chunks))
for i, c := range chunks {
if _, ok := s.chunks[c.full]; ok {
continue
}
s.chunks[c.full] = c.Len
s.Bytes += c.Len
added += c.Len
isNew[i] = true
}
return isNew, added
}
// The claims the sheet makes about the two chunkers.
size := len(base)
if r := res["insert"]; r.Fixed.NewBytes < size*9/10 || r.CDC.NewChunks > 2 {
t.Errorf("insert: fixed adds %d bytes, CDC %d chunks; want 90%% of the file, then at most 2 chunks", r.Fixed.NewBytes, r.CDC.NewChunks)
}
if r := res["delete"]; r.Fixed.NewBytes < size*7/10 || r.CDC.NewChunks > 2 {
t.Errorf("delete: fixed adds %d bytes, CDC %d chunks", r.Fixed.NewBytes, r.CDC.NewChunks)
}
for _, id := range []string{"overwrite", "append"} {
if r := res[id]; r.Fixed.NewChunks > 2 || r.CDC.NewChunks > 2 {
t.Errorf("%s: fixed %d new chunks, CDC %d; want at most 2 each", id, r.Fixed.NewChunks, r.CDC.NewChunks)
}
}
if out.CDCAll*3 > out.Logical || out.FixedAll < 2*out.CDCAll {
t.Errorf("history: %d logical, fixed %d, CDC %d; want CDC under a third and fixed twice CDC", out.Logical, out.FixedAll, out.CDCAll)
}| version | size | fixed adds | content-defined adds |
|---|---|---|---|
| v1, the original | 256 KiB | 256 KiB | 256 KiB |
| v2, insert | 256 KiB | 256 KiB | 4 KiB |
| v3, overwrite | 256 KiB | 8 KiB | 32 KiB |
| v4, append | 260 KiB | 4 KiB | 10 KiB |
| v5, delete | 257 KiB | 201 KiB | 18 KiB |
| stored | 1,285 KiB | 725 KiB | 321 KiB |
Each version applies one edit to the one before. Without dedupe, the store holds every byte of every version.
- Sync clients and backup tools upload only the new chunks.
- An overwrite inside one large chunk re-sends that whole chunk: smaller chunks save bytes and cost more hashes.
Over 5 versions of one file, content-defined chunks stored 25% of the bytes; fixed chunks stored 56%.
encode(object, k, m):
split into k data shards
parity[j] = SUM of E[k + j][i] * data[i] // GF(256) bytes1
put the k + m shards on k + m disks, racks apart2
rebuild(any k shards):
rows = the k matching rows of E
data = inverse(rows) * shards // always invertible3- 1Arithmetic on bytes where every non-zero byte has an inverse, so matrices invert.
- 2Shards of one object on one rack fail together. Spread them across racks or zones.
- 3Built from a Vandermonde matrix, any k rows of E invert. The lab tries every subset.
Tested source Go: build the code · Go: rebuild
// NewCode builds the encoding matrix. A Vandermonde matrix has the property that any K of its
// rows are invertible. Multiplying it by the inverse of its top K rows keeps that property and
// makes the top K rows the identity, so the data shards are stored as they are.
func NewCode(k, m int) (*Code, error) {
n := k + m
if k < 1 || m < 0 || n > 256 {
return nil, fmt.Errorf("code %d+%d: need k >= 1 and k + m <= 256", k, m)
}
v := newMatrix(n, k)
for i := range n {
for j := range k {
v[i][j] = gfPow(byte(i), j)
}
}
topInv, err := v[:k].invert()
if err != nil {
return nil, fmt.Errorf("code %d+%d: %w", k, m, err)
}
return &Code{K: k, M: m, enc: v.mul(topInv)}, nil
}
// Reconstruct fills in the missing (nil) shards. Any K surviving shards are K known rows of the
// encoding matrix times the data, so inverting those K rows recovers the data shards. The lost
// parity shards are then encoded again.
func (c *Code) Reconstruct(shards [][]byte) error {
var rows matrix
var have [][]byte
for i, s := range shards {
if s != nil && len(rows) < c.K {
rows = append(rows, c.enc[i])
have = append(have, s)
}
}
if len(rows) < c.K {
return fmt.Errorf("%d of %d shards left, need %d: %w", len(rows), c.K+c.M, c.K, ErrTooFewShards)
}
dec, err := rows.invert()
if err != nil {
return fmt.Errorf("decode: %w", err)
}
data := make([][]byte, c.K)
for j := range c.K {
data[j] = c.combine(dec[j], have)
}
for i := range shards {
if shards[i] == nil {
shards[i] = c.combine(c.enc[i], data)
}
}
return nil
}
Lab: 10+4 rebuilt all 1,001 sets of 4 lost shards, and none of the 2,002 sets of 5.
Reed-Solomon 6+3 cuts an object into 6 data shards and adds 3 parity shards. Any 6 of the 9 rebuild it, at 1.5 times the bytes.
- stored per byte
- 1.5×
- disks that may fail
- 3
- lost a year
- 4.14 × 10^-13
- nines
- 12.4
- one disk fails in one repair window
p = 0.02 × 24 / 8760 = 5.48 × 10^-5 - more than 3 of 9 disks fail in one window
C(9, 4) × p^4 = 126 × 9.01 × 10^-18 = 1.14 × 10^-15; with every term: 1.14 × 10^-15 - windows a year
8760 / 24 = 365 - lost a year
365 × 1.14 × 10^-15 = 4.14 × 10^-13
A model: disks fail independently, and each failed disk is rebuilt within the window. Correlated failures (a rack, a bad batch) cost more.
p = AFR * repair_hours / 8760 // one disk, one window1
lost = P(more than m of n disks fail in one window)
≈ C(n, m + 1) * p^(m + 1)2
per year = lost * 8760 / repair_hours- 1AFR is the share of disks that fail in a year. Backblaze reported 1.36% over its fleet for 2025.
- 2The largest term. p is tiny, so the rest add little; the lab sums them all.
Tested source Go: durability
// Durable works out the yearly chance to lose an object. The model: disks fail independently at
// rate afr a year, and a failed disk is rebuilt within repairHours. An object is lost only if more
// than M of its K+M disks fail inside one repair window.
func Durable(s Scheme, afr, repairHours float64) Durability {
n := s.K + s.M
p := afr * repairHours / 8760 // one disk, one window
var perWindow float64
for i := s.M + 1; i <= n; i++ {
perWindow += binom(n, i) * math.Pow(p, float64(i)) * math.Pow(1-p, float64(n-i))
}
windows := 8760 / repairHours
annual := windows * perWindow
return Durability{
Overhead: float64(n) / float64(s.K),
Tolerates: s.M,
P: p,
Windows: windows,
Lead: binom(n, s.M+1) * math.Pow(p, float64(s.M+1)),
PerWindow: perWindow,
Annual: annual,
Nines: -math.Log10(annual),
}
}
At 2% AFR and a 24-hour rebuild, the yearly chance to lose an object is 6.0 × 10^-11 with 3 replicas and 4.1 × 10^-13 with 6+3.
| event | result | why it is safe, or the fix | saved by |
|---|---|---|---|
| The client uploads but never calls finish | A pending row and an orphan object. | The sweeper deletes both after 24 hours. | Pending index |
| Finish is retried | The second call finds the row ready. | The update needs status pending, so it changes nothing. | Postgres |
| An upload URL leaks | Anyone can PUT that one key until expiry. | Short expiry; key and type are signed; finish checks size and hash. | Presigned URL |
| The connection drops mid-upload | Some parts arrived. | List parts, send the rest; nothing is visible before Complete. | Multipart |
| A part is damaged on the way | Its checksum does not match. | The store refuses the part; the client sends it again. | Checksum |
| Uploads are abandoned mid-way | Parts are stored and billed. | A lifecycle rule aborts incomplete uploads after some days. | Lifecycle |
| More than m disks of one stripe fail in one window | The object is lost. | Fast rebuilds, shards across racks or zones, more parity. | Erasure code |
| Two uploads write one key | The last writer wins. | A new key per upload; the row names the current key. | Unique key |
| A row is deleted, the object stays | Storage leaks. | Delete the object first, or reconcile store listings against rows. | Reconcile job |
| step | add | it handles | move up when you see |
|---|---|---|---|
| 1 | An object store and a metadata row. Uploads pass through the service. | Small files. Postgres takes 3,190 metadata rows a second in the lab. | Upload bandwidth or memory on the service; slow uploads hold its connections. |
| 2 | Presigned direct uploads. | The store takes the bytes; the service signs and records. | Files over 100 MB fail on mobile networks and restart from zero. |
| 3 | Multipart, resumable uploads. | Large files in parallel parts; a drop resends one part. | Reads repeat the same objects worldwide; egress and latency grow. |
| 4 | A CDN in front. | Repeat reads served at the edge (sheet B5). | Storage cost grows with old, cold objects. |
| 5 | Lifecycle tiers, chunk dedupe, a second region. | Cold data on cheaper classes; versions share chunks; a region can fail. | Top of the ladder. |
Demand example: 10 million uploads a day is 116 a second; 100 reads per upload is 11,600 a second, mostly from the CDN. S3 rates are per prefix and scale with more prefixes.
I start with an object store and a metadata row from day one. I add direct uploads, then a CDN, then lifecycle tiers, as traffic and storage grow.
0 of 9 known
Why not keep user photos in a bytea column?
A client has a presigned PUT for u/7/a.png. Can it upload to u/8/b.png instead?
A 4 GB upload drops after 3 of 8 parts. What does the client do?
Insert 100 bytes near the start of a 256 KiB file. How many 8 KiB fixed chunks change?
Why 6+3 erasure coding instead of 3 replicas?
How do you estimate the yearly loss for 3 replicas at 2% AFR and 24 h repair?
Which matters more for durability: a better disk or a faster rebuild?
The upload succeeded but finish never arrived. What is left behind?
Two clients PUT the same key at the same moment. What do readers see?
- bytea
- 1 MiB file: 1.07 MiB of log. Metadata row: 231 bytes.
- overhead
- 3 replicas 3×; 6+3 1.5×; 10+4 1.4×.
- survives
- 2, 3 and 4 lost disks.
- loss a year
- 2% AFR, 24 h rebuild: 10, 12 and 15 nines.
- S3
- Designed for 11 nines; strong read-after-write.
- parts
- 5 MiB to 5 GiB, up to 10,000; objects to 48.8 TiB.
- rate
- 3,500 writes, 5,500 reads a second per prefix.
- URLs
- Presigned for up to 7 days.
Measured on Postgres 16.14, a 16-thread laptop. Nines are derived from the model on this sheet. S3 figures are cited from its documentation.