Observability: see a problem, find it, explain it
Instrument every service with structured logs, metrics and traces. Alert on what users feel. Then walk a fixed path from the dashboard to one slow request.
- 1Metrics tell you that something is wrong, traces show where, logs say why. One trace ID joins all three.
- 2Measure latency as a histogram. Merge buckets, then take the percentile; never average percentiles.
- 3Alert on SLO burn rate over a long and a short window, so a real outage pages fast and a blip does not.
- 4Debug in order: which signal moved, since when, which dimension isolates it, then one slow trace and its log.
- Metrics aggregate: one series per label set. Cheap to keep and fast to query, but no single request.
- Logs and traces keep each request. Rich, but expensive, so you sample or keep them for days.
- The trace ID is the join key. Without it, you search logs by time and guess.
Each request leaves a log line, metric updates and a trace. The trace ID in the log line and the spans lets me jump from one to the other.
| question | metrics | traces | logs |
|---|---|---|---|
| Is the SLO burning? Page someone? | Approved | Not approved | Not approved |
| Since when, and how much? | Approved | Not approved | Slow to count |
| Which host, route, dependency? | Few values | Approved | With fields |
| Which customer, user, order? | Cardinality | Approved | Approved |
| Where did this request spend its time? | Not approved | Approved | With timings |
| What was the error message and input? | Not approved | Span events | Approved |
| method | for | signals |
|---|---|---|
| RED | Each service | Rate, Errors, Duration |
| USE | Each resource: CPU, disk, pool | Utilization, Saturation, Errors |
| Golden signals | User-facing systems | Latency, traffic, errors, saturation |
Metrics answer 'is it broken and how much'. Traces answer 'where in the call graph'. Logs answer 'what exactly happened to this request'.
Step 1: Emit
- The service writes JSON log lines, updates counters and histograms, and records spans.
- Every log line and span carries the trace ID.
If it fails
Missing instrumentation: you see that the service is slow, not which call is slow.
Services emit, Prometheus scrapes the metrics, and a collector ships traces and logs. Rules alert on burn rate. I debug from the dashboard to a trace to a log line.
| tool | capability | what it gives this design | also used for |
|---|---|---|---|
| Prometheus | Pull scrape of /metrics, instance label added | A target that stops answering is itself a signal. No agent in the service. | Service discovery from Kubernetes |
| Prometheus | Counters, gauges, histograms with buckets | Histograms add across hosts, so a fleet p99 is exact to one bucket. | Queue depth, pool usage |
| Prometheus | PromQL: rate, sum by, histogram_quantile | Break down any SLI by any label, after the fact. | Capacity planning |
| Prometheus | Recording and alerting rules | Burn rates over 5 min to 3 days, evaluated every minute. | Precomputed dashboards |
| Alertmanager | Grouping, deduplication, routing, silences | One page per incident, not one per host. Pages and tickets go to different places. | Maintenance windows |
| OpenTelemetry | Tracing API and SDKs, W3C trace context | Spans from every language join one trace through the traceparent header. | Baggage: tenant across hops |
| OTel collector | Batching, tail sampling, export to many backends | Keeps every slow or failed trace and a small share of the rest. | Redacting fields |
| slog (Go) | Structured JSON logs, a handler per context | Every line has fields and the trace ID, so a log query is exact. | Audit logs |
| Postgres | pg_stat_statements, log_min_duration_statement | Time per query shape, and a log line for each slow statement. | Index tuning |
| Redis | SLOWLOG, LATENCY, INFO | Slow commands and memory, to tell a slow cache from a slow network. | Capacity checks |
Prometheus gives me pull-based metrics, histograms that aggregate and burn-rate rules. OpenTelemetry gives me one API for spans, context propagation and a collector that can tail-sample.
on every log record, with the request context1:
span = span in the context
IF span exists:
add trace_id = span.trace_id2
add span_id = span.span_id
write the record as one JSON line3- 1Log with the context, as slog.InfoContext does, so the handler can find the span.
- 2The join key. Search the log store for it and you get every line of the request, from every service.
- 3Fields, not prose: route=/checkout status=503. A query can then filter and count.
Tested source Go: slog handler
// ContextHandler adds the trace and span IDs from the context to every log record, so a log line
// leads to its trace and a trace leads to its log lines. With a Clock it also stamps records
// with that clock, so a replay logs simulated time.
type ContextHandler struct {
slog.Handler
Clock Clock
}
// Handle adds trace_id and span_id, then passes the record on.
func (h ContextHandler) Handle(ctx context.Context, r slog.Record) error {
if s := SpanFrom(ctx); s != nil {
r.AddAttrs(slog.String("trace_id", s.TraceID.String()), slog.String("span_id", s.SpanID.String()))
}
if h.Clock != nil {
r.Time = h.Clock.Now()
}
if err := h.Handler.Handle(ctx, r); err != nil {
return fmt.Errorf("log: %w", err)
}
return nil
}| level | use it for |
|---|---|
| ERROR | A request failed. Someone may need to act. |
| WARN | A dependency failed; a retry or fallback may hide it. |
| INFO | One line per request or job, with its fields. |
| DEBUG | Off in production; on for one host or one tenant. |
- Never log secrets, tokens or full card numbers. Redact in the handler.
- Sample high-volume INFO lines. Keep every ERROR.
I log one structured line per request, with the route, status, duration, tenant and trace ID. Levels mean something: ERROR is a failed request, not a retry that worked.
handle(request):
span = start span, parent = traceparent header1 (or a new trace)
start = now
FOR EACH dependency of the route:
child = start span under span
call dependency, send traceparent = child
observe dependency_duration_seconds{dep}2 (now - t0)
IF error: log WARN with trace_id; status = 503; STOP
count http_requests_total{route, code, tenant}3 += 1
observe http_request_duration_seconds{route, tenant}4 (now - start)
log one line: route, tenant, status, duration_ms, trace_id
end span- 1Join the caller's trace. A missing or bad header starts a new one; it never fails the request.
- 2RED on the client side: you see a slow dependency from the caller, even if it has no metrics of its own.
- 3A counter. Rate and error ratio come from it at query time.
- 4A histogram, not an average. Buckets can be summed across hosts.
Tested source Go: middleware
// ServeHTTP runs one request: a server span from the caller's trace context, a child span and
// a latency histogram per dependency call, RED metrics for the request, and one log line that
// carries the trace ID.
func (s *Service) ServeHTTP(w http.ResponseWriter, r *http.Request) {
ctx, span := s.Tracer.Start(Extract(r.Context(), r.Header), r.Method+" "+r.URL.Path)
route, tenant := r.URL.Path, r.Header.Get("X-Tenant")
span.Attrs["route"], span.Attrs["tenant"] = route, tenant
start := s.Clock.Now()
deps, ok := s.Routes[route]
status := http.StatusOK
if !ok {
status = http.StatusNotFound
}
if ok && s.Work != nil {
s.Work(ctx)
}
for _, name := range deps {
cctx, child := s.Tracer.Start(ctx, name)
t0 := s.Clock.Now()
err := s.Deps[name].Call(cctx, tenant)
s.Metrics.Observe("dependency_duration_seconds", s.Clock.Now().Sub(t0).Seconds(), Label{"dep", name})
child.Err = err != nil
child.Finish()
if err != nil {
s.Log.WarnContext(cctx, "dependency failed", "dep", name, "tenant", tenant, "err", err.Error())
status = http.StatusServiceUnavailable
break
}
}
w.WriteHeader(status)
took := s.Clock.Now().Sub(start)
code := strconv.Itoa(status)
s.Metrics.Inc("http_requests_total", Label{"route", route}, Label{"code", code}, Label{"tenant", tenant})
s.Metrics.Observe("http_request_duration_seconds", took.Seconds(), Label{"route", route}, Label{"tenant", tenant})
level := slog.LevelInfo
if status >= 500 {
level = slog.LevelError
}
s.Log.Log(ctx, level, "request", "route", route, "tenant", tenant, "status", status, "duration_ms", took.Milliseconds())
span.Err = status >= 500
span.Attrs["status"] = code
span.Finish()
}| type | example | query |
|---|---|---|
| Counter | requests, errors, bytes | rate() |
| Gauge | queue length, pool in use, memory | max, avg |
| Histogram | latency, payload size | histogram_quantile() |
Middleware counts every request by route, code and tenant, and records its duration in a histogram. Each dependency call gets its own histogram.
Seeded run. A 40 ms median, and a slow path of about 1.5 s for 0.2% of requests on healthy hosts and 3% to 5% on the slow one.
p99 over a window, across hosts:
h = sum by (le)1 of bucket counts, end of window - start of window
rank = 0.99 × total count
find the first bucket whose cumulative count >= rank
IF it is the +Inf bucket: RETURN the highest finite bound
RETURN lower + (upper - lower)2 × (rank - count below) / count in bucket
never:
avg(p99 of each host)3 // wrong in both directions- 1Buckets add exactly, host by host. The lab test checks that merged host histograms equal one fleet histogram.
- 2Prometheus assumes values spread evenly in the bucket. The error is at most one bucket width, so put bounds near your SLO threshold.
- 3Canary case: 349 ms against a true 93 ms. Busy-host case: 355 ms against 1,592 ms.
Tested source Go: histogram quantile · Go: sum by a label
// Quantile estimates the q-quantile as Prometheus's histogram_quantile does: find the bucket
// that holds rank q × count, then assume its observations are spread evenly inside it.
func (s Snapshot) Quantile(q float64) float64 {
if s.Count == 0 {
return math.NaN()
}
rank := q * float64(s.Count)
var below uint64
for i, c := range s.Counts {
if float64(below+c) < rank || c == 0 {
below += c
continue
}
if i == len(s.Bounds) { // the +Inf bucket: the best answer is the highest bound
return s.Bounds[len(s.Bounds)-1]
}
lower := 0.0
if i > 0 {
lower = s.Bounds[i-1]
}
return lower + (s.Bounds[i]-lower)*(rank-float64(below))/float64(c)
}
return s.Bounds[len(s.Bounds)-1]
}// HistBy sums one histogram metric by a label, as sum by (label, le) does in PromQL. An empty
// label sums everything into "".
func HistBy(samples []Sample, name, label string) map[string]Snapshot {
out := map[string]Snapshot{}
for _, s := range samples {
if s.Name != name || s.Hist == nil {
continue
}
k := ""
if label != "" {
k = s.Label(label)
}
out[k] = out[k].Add(*s.Hist)
}
return out
}| tenant incident, whole fleet | before | during |
|---|---|---|
| p50 | 37 ms | 37 ms |
| mean | 68 ms | 144 ms |
| p99 | 457 ms | 2,224 ms |
I report p50, p99 and p99.9 from histograms. To get a fleet percentile I sum the buckets across hosts and then take the quantile.
outgoing call:
header traceparent1 = "00-" + trace_id + "-" + span_id + "-01"
incoming request:
parse traceparent: version 00, 32 hex, 16 hex, 2 hex flags
IF valid: new span joins that trace, parent = that span_id
ELSE: start a new trace // never fail the request2
start span(context, name):
parent = span in the context, else the remote parent
span_id = random 8 bytes3- 1The W3C Trace Context header. Every OpenTelemetry SDK reads and writes it, so spans from different languages join one trace.
- 2Telemetry must not take the service down. A bad header costs you one link in the trace, nothing more.
- 3A trace ID is 16 bytes and a span ID 8, random, so no coordination is needed to make them.
Tested source Go: traceparent · Go: start a span, inject, extract
// Traceparent formats the header: version 00, the 32-hex trace ID, the 16-hex parent span ID,
// and flags (01 = sampled). Example: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
func (sc SpanContext) Traceparent() string {
flags := "00"
if sc.Sampled {
flags = "01"
}
return "00-" + sc.TraceID.String() + "-" + sc.SpanID.String() + "-" + flags
}
// ParseTraceparent reads the header. An all-zero trace or span ID is invalid.
func ParseTraceparent(v string) (SpanContext, error) {
parts := strings.Split(strings.TrimSpace(v), "-")
if len(parts) != 4 || parts[0] != "00" || len(parts[1]) != 32 || len(parts[2]) != 16 || len(parts[3]) != 2 {
return SpanContext{}, fmt.Errorf("%w: %q", ErrBadTraceparent, v)
}
var sc SpanContext
if _, err := hex.Decode(sc.TraceID[:], []byte(parts[1])); err != nil {
return SpanContext{}, fmt.Errorf("%w: trace ID: %w", ErrBadTraceparent, err)
}
if _, err := hex.Decode(sc.SpanID[:], []byte(parts[2])); err != nil {
return SpanContext{}, fmt.Errorf("%w: span ID: %w", ErrBadTraceparent, err)
}
flags, err := hex.DecodeString(parts[3])
if err != nil {
return SpanContext{}, fmt.Errorf("%w: flags: %w", ErrBadTraceparent, err)
}
if sc.TraceID == (TraceID{}) || sc.SpanID.IsZero() {
return SpanContext{}, fmt.Errorf("%w: zero ID", ErrBadTraceparent)
}
sc.Sampled = flags[0]&1 == 1
return sc, nil
}// Start begins a span. Its parent is the span already in ctx, or else the remote parent that
// Extract found in the request headers, or else none: a new trace.
func (t *Tracer) Start(ctx context.Context, name string) (context.Context, *Span) {
s := &Span{Name: name, Service: t.Service, Start: t.Clock.Now(), Attrs: map[string]string{}, tracer: t}
t.mu.Lock()
binary.BigEndian.PutUint64(s.SpanID[:], t.rng.Uint64()|1)
if p := SpanFrom(ctx); p != nil {
s.TraceID, s.Parent = p.TraceID, p.SpanID
} else if rc, ok := ctx.Value(remoteKey{}).(SpanContext); ok {
s.TraceID, s.Parent = rc.TraceID, rc.SpanID
} else {
binary.BigEndian.PutUint64(s.TraceID[:8], t.rng.Uint64())
binary.BigEndian.PutUint64(s.TraceID[8:], t.rng.Uint64()|1)
}
t.mu.Unlock()
return context.WithValue(ctx, spanKey{}, s), s
}
// Inject writes the current span into outgoing headers.
func Inject(ctx context.Context, h http.Header) {
if s := SpanFrom(ctx); s != nil {
h.Set(TraceparentHeader, s.Context().Traceparent())
}
}
// Extract reads the caller's span from incoming headers. A missing or bad header starts a new
// trace; it never fails the request.
func Extract(ctx context.Context, h http.Header) context.Context {
sc, err := ParseTraceparent(h.Get(TraceparentHeader))
if err != nil {
return ctx
}
return context.WithValue(ctx, remoteKey{}, sc)
}| sampling | decides | slow or failed kept |
|---|---|---|
| Head, 1% | At the first span | 5 of 497 |
| Tail: all slow or failed, 1% of the rest | After the trace ends, in the collector | 497 of 497 |
Tenant incident replay: 36,000 requests; tail sampling kept 872 traces (2.4%). The collector must hold each trace until it ends.
Each service starts a span per operation and sends the trace context in the traceparent header. A collector keeps every slow or failed trace and samples the rest.
series(name, labels):
IF (name, labels) exists: RETURN it
IF name already has max_series1 series:
count it as dropped2
labels = { overflow = "true" } // one shared series
create the series- 1A cap per metric. In the lab, 10,000 user IDs made 100 real series and one overflow series, and every login was still counted.
- 2Alert on drops: a new label is exploding, usually an ID inside a URL path.
Tested source Go: registry with a series limit
// get returns the series for name and labels. A new label set beyond the metric's limit goes
// into one overflow series instead, and the registry counts it as dropped.
func (r *Registry) get(name string, labels []Label, hist bool) *series {
k := key(name, labels)
if s, ok := r.series[k]; ok {
return s
}
if r.perMetric[name] >= r.maxSeries {
r.dropped++
labels = []Label{OverflowLabel}
k = key(name, labels)
if s, ok := r.series[k]; ok {
return s
}
}
s := &series{name: name, labels: slices.Clone(labels)}
if hist {
s.hist = NewHistogram(r.buckets)
}
r.series[k] = s
r.perMetric[name]++
return s
}- The Prometheus guidance: most metrics need no labels; keep a metric's cardinality below 10, and rethink one that can pass 100.
- Label the route template (/orders/:id), never the raw path.
Every label value multiplies the series count. I keep metric labels to bounded sets like route and status, and put user and order IDs on spans and logs.
| scenario, 99.9% SLO | page 1 h | page 6 h | ticket 3 d | naive 5 min |
|---|---|---|---|---|
| 10% errors for 1 hour | 8 min (2.08%) | 21 min (5.09%) | 38 min (9.03%) | 0 min (0.23%) |
| 0.8% errors for 12 hours | never | 4.5 h (4.98%) | 8.2 h (9.13%) | 0 min (0.02%) |
| 0.2% errors for 4 days | never | never | 34.1 h (9.48%) | 2 min (0.01%) |
| 5% errors for 5 minutes | never | never | never | 0 min (0.12%) |
Time to fire after the errors start, and the share of the 30-day budget spent by then. Thresholds from the SRE Workbook, chapter "Alerting on SLOs".
| term | example |
|---|---|
| SLI: good events ÷ all events | Requests with no 5xx, answered in under 300 ms |
| SLO: the target for the SLI | 99.9% over a rolling 30 days |
| Error budget: 1 - SLO | 0.1% of requests, or 43.2 min a month |
| Burn rate: error ratio ÷ budget ratio | 1% errors on a 99.9% SLO burns at 10 |
allowed = 1 - SLO // 0.001 for 99.9%
fires(rule, minute):
long = error ratio over the last rule.long window
short = error ratio over the last rule.short window
RETURN long >= rule.burn × allowed1
AND short >= rule.burn × allowed2
rules: 1 h and 5 min, burn 14.4 → page // 2% of 30 days
6 h and 30 min, burn 6 → page // 5%
3 d and 6 h, burn 1 → ticket3 // 10%- 1The long window measures how much budget is gone. 14.4 over 1 hour is 2% of 30 days.
- 2The short window, 1/12 of the long one, checks the burn is still happening. The page reset 4 min after the errors stopped, not 51.
- 3A slow leak is not an emergency. It goes to a ticket queue for working hours.
Tested source Go: burn-rate rule
// Firing reports, for each minute, whether the rule fires. The SLO is a target such as 0.999,
// so the allowed error ratio is 1 - SLO.
func (r BurnRule) Firing(ratio []float64, slo float64) []bool {
w := newWindows(ratio)
limit := r.Burn * (1 - slo)
out := make([]bool, len(ratio))
for m := range ratio {
long := w.mean(m, r.Long) >= limit
short := r.Short == 0 || w.mean(m, r.Short) >= limit
out[m] = long && short
}
return out
}My SLI is the share of requests that succeed within 300 ms. I page when the 30-day budget burns 14.4 times too fast over an hour and five minutes, or 6 times over six hours and thirty minutes.
From minute 30, the p99 rose from 457 ms to 2,224 ms. The median moved from 37 to 37 ms and the mean from 68 to 144 ms. Traffic did not change. Next: find who is slow.
Each incident sends 600 requests a minute for 60 minutes through the instrumented service on 6 hosts, with a simulated clock. Charts read the per-minute scrapes, as a query over a 1-minute window would.
The dashboard tells me the p99 rose at minute 30. Breaking down by tenant isolates acme; the slowest trace shows payments.charge; and its log line has the same trace ID.
| step | question | where |
|---|---|---|
| 1 | Which SLI moved, since when, how much? | Golden-signal dashboard |
| 2 | What changed just before? Deploy, flag, config, traffic. | Deploy markers, change log |
| 3 | Is a dependency slow or failing? | Client-side RED per dependency |
| 4 | Is a resource saturated? | USE: CPU, memory, pools, queues |
| 5 | One host, tenant, route, zone? | Break down the SLI by each label |
| 6 | Where does one slow request spend time? | Trace, then its logs by trace ID |
- All hosts rising together means a shared cause. In the tenant incident every host rose about 5 times; only one tenant and one dependency stood out.
- One host rising alone means a host problem. On web-4 the trace spent 3,681 of 3,702 ms outside every child span.
- Mitigate first: roll back, shed the tenant, drain the host. Find the root cause after.
I scope the symptom, check what changed, then break down by dimension until one value stands out, and only then open traces and logs.
| row | panels |
|---|---|
| 1. Users | SLO burn rate, budget left; p50 and p99 latency; error ratio; traffic |
| 2. Business | Sign-ups, orders or payments a minute, against last week |
| 3. Dependencies | Client-side rate, errors, p99 for each database, cache, provider |
| 4. Saturation | CPU, memory, connection pools, queue length and age, disk |
| 5. Changes | Deploy and feature-flag markers on every chart |
- Every latency panel shows percentiles from histograms, never an average alone.
- Load-test before launch, so you know the values of a healthy system.
- Check that the alerts fire: break a staging dependency on purpose.
On launch day I watch the SLO burn and the golden signals first, then each dependency, then saturation, with deploys marked on every chart.
| event | result | fix | saved by |
|---|---|---|---|
| A user ID goes into a metric label | Millions of series; the metrics store runs out of memory. | A series cap per metric, an alert on drops; IDs go on spans and logs. | Series limit |
| An alert on every error spike | Pages for blips nobody can act on. People start to ignore pages. | Burn-rate alerts: the 5-minute blip spent 0.58% of budget and paged nobody. | Burn-rate rules |
| Logs without a trace ID | You find the slow trace but not its error message. | A log handler that adds the trace ID from the context to every line. | Context handler |
| A proxy drops the traceparent header | The trace breaks into two traces at that hop. | Forward trace headers in every proxy and queue message. | OpenTelemetry |
| Head sampling at 1% | It kept 5 of 497 slow or failed traces. | Tail sampling in the collector keeps all of them. | OTel collector |
| A dashboard averages host p99s | It said 355 ms when the fleet p99 was 1,592 ms. | Sum the histogram buckets, then take the quantile. | Histograms |
| The service stops sending metrics | Its graphs go flat, and no alert fires. | Alert on a missing target or an absent series. | Prometheus |
| The monitoring stack shares the outage | You are blind exactly when you need to see. | Run monitoring outside the region it watches, with an external probe. | External probe |
| step | add | it handles | move up when you see |
|---|---|---|---|
| 1 | Structured logs, RED metrics, an external probe. One Prometheus. | One or a few services. You know that it is broken and roughly where. | Pages for blips, or outages found by users first. |
| 2 | SLOs and multiwindow burn-rate alerts, routed by Alertmanager. | Pages that mean real budget loss. Tickets for slow leaks. | A request crosses several services and you cannot tell which is slow. |
| 3 | Tracing with OpenTelemetry, trace IDs in every log line, head sampling. | Where a request spends time, across services. | The slow traces you need were not sampled, or trace storage costs too much. |
| 4 | A collector tier with tail sampling, batching, redaction. | Every slow or failed trace kept, at a few percent of volume. | One Prometheus runs out of memory, or you need a year of history. |
| 5 | Long-term metrics storage, sharded scraping, recording rules. | Many clusters, months of data, one query view. | Top of the ladder. |
14 series per label set: 11 buckets, +Inf, sum and count. The 100,000 example (10,000 nodes) is from the Prometheus instrumentation guide, which calls double-digit millions too much for one server.
I start with structured logs, RED metrics and one external probe. I add SLO alerts before launch, tracing when requests cross several services, and tail sampling when trace volume costs too much.
0 of 9 known
Host A serves 10,000 requests with p99 = 90 ms. A canary serves 500 with p99 = 1.6 s. What is the fleet p99?
During the tenant incident, the median stayed at 37 ms. Why did the p99 still rise about 5 times?
You want latency per customer, and you have 200,000 customers. Do you add a customer label to the histogram?
An SLO of 99.9% over 30 days. 10% of requests fail from now on. When does the fast page fire, and how much budget is gone?
Why does each burn-rate alert need a short window too?
5% of requests fail for 5 minutes, then the service recovers. Should anyone be paged?
You sample 1% of traces at the start of each request. What happens to the slow ones?
Every host got slower at the same time. What does that tell you?
A trace takes 3.7 s, but its child spans add up to 21 ms. Where did the time go?
- burn rules
- Page: 14.4× over 1 h and 5 min; 6× over 6 h and 30 min. Ticket: 1× over 3 d and 6 h.
- budget
- They fire at 2%, 5% and 10% of a 30-day budget.
- detect
- Minutes = long window × threshold ÷ burn. At 10% errors: 60 × 14.4 ÷ 100 = 8.6 min.
- 99.9%
- 43.2 minutes, or 0.1% of requests, a month.
- p99
- Tenant incident: p99 457 → 2,224 ms; median 37 → 37 ms.
- sampling
- Head 1% kept 5 of 497 slow traces; tail kept all, at 2.4% of volume.
- labels
- Keep most metrics under 10 series; rethink one that can pass 100.
Burn rules: SRE Workbook. Labels: Prometheus instrumentation guide. Incident and sampling: replays in the lab with a seeded random source, 36,000 requests each.