System Design
O3

Observability: see a problem, find it, explain it

Instrument every service with structured logs, metrics and traces. Alert on what users feel. Then walk a fixed path from the dashboard to one slow request.

Not startedSaved in this browser only.
  1. 1Metrics tell you that something is wrong, traces show where, logs say why. One trace ID joins all three.
  2. 2Measure latency as a histogram. Merge buckets, then take the percentile; never average percentiles.
  3. 3Alert on SLO burn rate over a long and a short window, so a real outage pages fast and a blip does not.
  4. 4Debug in order: which signal moved, since when, which dimension isolates it, then one slow trace and its log.
O3
    A

    Three signals, one request

    recorded from the lab service
    GET /checkouttenant acme, web-22,007 ms, status 200LOG · one event, all its fields"level":"INFO", "msg":"request", "route":"/checkout","duration_ms":2007, "trace_id":"22bbd4d6…"METRICS · counts, cheap to keep for a yearhttp_requests_total{route, code="200", tenant} +1http_request_duration_seconds_bucket{le="2.5"} +1no trace ID: one series per label set, not per requestTRACE · trace_id 22bbd4d6…, one span per stepGET /checkout2007 mscache.get0 msdb.query6 mspayments.charge1991 ms
    • Metrics aggregate: one series per label set. Cheap to keep and fast to query, but no single request.
    • Logs and traces keep each request. Rich, but expensive, so you sample or keep them for days.
    • The trace ID is the join key. Without it, you search logs by time and guess.

    Each request leaves a log line, metric updates and a trace. The trace ID in the log line and the spans lets me jump from one to the other.

    B

    Which signal answers what

    pick the cheapest one that answers
    questionmetricstraceslogs
    Is the SLO burning? Page someone?ApprovedNot approvedNot approved
    Since when, and how much?ApprovedNot approvedSlow to count
    Which host, route, dependency?Few valuesApprovedWith fields
    Which customer, user, order?CardinalityApprovedApproved
    Where did this request spend its time?Not approvedApprovedWith timings
    What was the error message and input?Not approvedSpan eventsApproved
    methodforsignals
    REDEach serviceRate, Errors, Duration
    USEEach resource: CPU, disk, poolUtilization, Saturation, Errors
    Golden signalsUser-facing systemsLatency, traffic, errors, saturation

    Metrics answer 'is it broken and how much'. Traces answer 'where in the call graph'. Logs answer 'what exactly happened to this request'.

    C

    The telemetry path

    click a step; its path lights up
    Service × Nslog, metrics, spansPrometheusscrapes /metricsOTel collectorbatch, tail sampleAlertmanagergroup, routeTrace storesearch by trace IDLog storeJSON linesOn callpage or ticketDashboardsgolden signals

    Step 1: Emit

    • The service writes JSON log lines, updates counters and histograms, and records spans.
    • Every log line and span carries the trace ID.

    If it fails

    Missing instrumentation: you see that the service is slow, not which call is slow.

    Services emit, Prometheus scrapes the metrics, and a collector ships traces and logs. Rules alert on burn rate. I debug from the dashboard to a trace to a log line.

    D

    Capabilities used

    what each tool gives you
    toolcapabilitywhat it gives this designalso used for
    PrometheusPull scrape of /metrics, instance label addedA target that stops answering is itself a signal. No agent in the service.Service discovery from Kubernetes
    PrometheusCounters, gauges, histograms with bucketsHistograms add across hosts, so a fleet p99 is exact to one bucket.Queue depth, pool usage
    PrometheusPromQL: rate, sum by, histogram_quantileBreak down any SLI by any label, after the fact.Capacity planning
    PrometheusRecording and alerting rulesBurn rates over 5 min to 3 days, evaluated every minute.Precomputed dashboards
    AlertmanagerGrouping, deduplication, routing, silencesOne page per incident, not one per host. Pages and tickets go to different places.Maintenance windows
    OpenTelemetryTracing API and SDKs, W3C trace contextSpans from every language join one trace through the traceparent header.Baggage: tenant across hops
    OTel collectorBatching, tail sampling, export to many backendsKeeps every slow or failed trace and a small share of the rest.Redacting fields
    slog (Go)Structured JSON logs, a handler per contextEvery line has fields and the trace ID, so a log query is exact.Audit logs
    Postgrespg_stat_statements, log_min_duration_statementTime per query shape, and a log line for each slow statement.Index tuning
    RedisSLOWLOG, LATENCY, INFOSlow commands and memory, to tell a slow cache from a slow network.Capacity checks

    Prometheus gives me pull-based metrics, histograms that aggregate and burn-rate rules. OpenTelemetry gives me one API for spans, context propagation and a collector that can tail-sample.

    E

    Logs

    structured, levelled, correlated
    log handlerpseudo code
    on every log record, with the request context1:
      span = span in the context
      IF span exists:
        add trace_id = span.trace_id2
        add span_id  = span.span_id
      write the record as one JSON line3
    1. 1Log with the context, as slog.InfoContext does, so the handler can find the span.
    2. 2The join key. Search the log store for it and you get every line of the request, from every service.
    3. 3Fields, not prose: route=/checkout status=503. A query can then filter and count.
    Tested source Go: slog handler
    Go: slog handlergo
    // ContextHandler adds the trace and span IDs from the context to every log record, so a log line
    // leads to its trace and a trace leads to its log lines. With a Clock it also stamps records
    // with that clock, so a replay logs simulated time.
    type ContextHandler struct {
      slog.Handler
      Clock Clock
    }
    
    // Handle adds trace_id and span_id, then passes the record on.
    func (h ContextHandler) Handle(ctx context.Context, r slog.Record) error {
      if s := SpanFrom(ctx); s != nil {
        r.AddAttrs(slog.String("trace_id", s.TraceID.String()), slog.String("span_id", s.SpanID.String()))
      }
      if h.Clock != nil {
        r.Time = h.Clock.Now()
      }
      if err := h.Handler.Handle(ctx, r); err != nil {
        return fmt.Errorf("log: %w", err)
      }
      return nil
    }
    leveluse it for
    ERRORA request failed. Someone may need to act.
    WARNA dependency failed; a retry or fallback may hide it.
    INFOOne line per request or job, with its fields.
    DEBUGOff in production; on for one host or one tenant.
    • Never log secrets, tokens or full card numbers. Redact in the handler.
    • Sample high-volume INFO lines. Keep every ERROR.

    I log one structured line per request, with the route, status, duration, tenant and trace ID. Levels mean something: ERROR is a failed request, not a retry that worked.

    F

    Metrics and the middleware

    RED for every route
    instrumented handlerpseudo code
    handle(request):
      span = start span, parent = traceparent header1 (or a new trace)
      start = now
      FOR EACH dependency of the route:
        child = start span under span
        call dependency, send traceparent = child
        observe dependency_duration_seconds{dep}2 (now - t0)
        IF error: log WARN with trace_id; status = 503; STOP
      count http_requests_total{route, code, tenant}3 += 1
      observe http_request_duration_seconds{route, tenant}4 (now - start)
      log one line: route, tenant, status, duration_ms, trace_id
      end span
    1. 1Join the caller's trace. A missing or bad header starts a new one; it never fails the request.
    2. 2RED on the client side: you see a slow dependency from the caller, even if it has no metrics of its own.
    3. 3A counter. Rate and error ratio come from it at query time.
    4. 4A histogram, not an average. Buckets can be summed across hosts.
    Tested source Go: middleware
    Go: middlewarego
    // ServeHTTP runs one request: a server span from the caller's trace context, a child span and
    // a latency histogram per dependency call, RED metrics for the request, and one log line that
    // carries the trace ID.
    func (s *Service) ServeHTTP(w http.ResponseWriter, r *http.Request) {
      ctx, span := s.Tracer.Start(Extract(r.Context(), r.Header), r.Method+" "+r.URL.Path)
      route, tenant := r.URL.Path, r.Header.Get("X-Tenant")
      span.Attrs["route"], span.Attrs["tenant"] = route, tenant
      start := s.Clock.Now()
    
      deps, ok := s.Routes[route]
      status := http.StatusOK
      if !ok {
        status = http.StatusNotFound
      }
      if ok && s.Work != nil {
        s.Work(ctx)
      }
      for _, name := range deps {
        cctx, child := s.Tracer.Start(ctx, name)
        t0 := s.Clock.Now()
        err := s.Deps[name].Call(cctx, tenant)
        s.Metrics.Observe("dependency_duration_seconds", s.Clock.Now().Sub(t0).Seconds(), Label{"dep", name})
        child.Err = err != nil
        child.Finish()
        if err != nil {
          s.Log.WarnContext(cctx, "dependency failed", "dep", name, "tenant", tenant, "err", err.Error())
          status = http.StatusServiceUnavailable
          break
        }
      }
      w.WriteHeader(status)
    
      took := s.Clock.Now().Sub(start)
      code := strconv.Itoa(status)
      s.Metrics.Inc("http_requests_total", Label{"route", route}, Label{"code", code}, Label{"tenant", tenant})
      s.Metrics.Observe("http_request_duration_seconds", took.Seconds(), Label{"route", route}, Label{"tenant", tenant})
      level := slog.LevelInfo
      if status >= 500 {
        level = slog.LevelError
      }
      s.Log.Log(ctx, level, "request", "route", route, "tenant", tenant, "status", status, "duration_ms", took.Milliseconds())
      span.Err = status >= 500
      span.Attrs["status"] = code
      span.Finish()
    }
    typeexamplequery
    Counterrequests, errors, bytesrate()
    Gaugequeue length, pool in use, memorymax, avg
    Histogramlatency, payload sizehistogram_quantile()

    Middleware counts every request by route, code and tenant, and records its duration in a histogram. Each dependency call gets its own histogram.

    G

    Percentiles, not averages

    and never an average of percentiles
    A small canary host is slowweb-1 · 10,000 req93 msweb-2 · 10,000 req93 msweb-3 · 10,000 req92 msweb-4 · 10,000 req92 msweb-5 · 10,000 req91 mscanary · 500 req1,633 msaverage of host p99s349 mstrue fleet p9993 msmerged histograms99 msThe busiest host is slowweb-1 · 40,000 req1,653 msweb-2 · 2,000 req96 msweb-3 · 2,000 req95 msweb-4 · 2,000 req97 msweb-5 · 2,000 req90 msweb-6 · 2,000 req96 msaverage of host p99s355 mstrue fleet p991,592 msmerged histograms1,889 ms100 ms1 sp99, log scale

    Seeded run. A 40 ms median, and a slow path of about 1.5 s for 0.2% of requests on healthy hosts and 3% to 5% on the slow one.

    fleet p99 from bucketspseudo code
    p99 over a window, across hosts:
      h = sum by (le)1 of bucket counts, end of window - start of window
      rank = 0.99 × total count
      find the first bucket whose cumulative count >= rank
      IF it is the +Inf bucket: RETURN the highest finite bound
      RETURN lower + (upper - lower)2 × (rank - count below) / count in bucket
    
    never:
      avg(p99 of each host)3                      // wrong in both directions
    1. 1Buckets add exactly, host by host. The lab test checks that merged host histograms equal one fleet histogram.
    2. 2Prometheus assumes values spread evenly in the bucket. The error is at most one bucket width, so put bounds near your SLO threshold.
    3. 3Canary case: 349 ms against a true 93 ms. Busy-host case: 355 ms against 1,592 ms.
    Tested source Go: histogram quantile · Go: sum by a label
    Go: histogram quantilego
    // Quantile estimates the q-quantile as Prometheus's histogram_quantile does: find the bucket
    // that holds rank q × count, then assume its observations are spread evenly inside it.
    func (s Snapshot) Quantile(q float64) float64 {
      if s.Count == 0 {
        return math.NaN()
      }
      rank := q * float64(s.Count)
      var below uint64
      for i, c := range s.Counts {
        if float64(below+c) < rank || c == 0 {
          below += c
          continue
        }
        if i == len(s.Bounds) { // the +Inf bucket: the best answer is the highest bound
          return s.Bounds[len(s.Bounds)-1]
        }
        lower := 0.0
        if i > 0 {
          lower = s.Bounds[i-1]
        }
        return lower + (s.Bounds[i]-lower)*(rank-float64(below))/float64(c)
      }
      return s.Bounds[len(s.Bounds)-1]
    }
    Go: sum by a labelgo
    // HistBy sums one histogram metric by a label, as sum by (label, le) does in PromQL. An empty
    // label sums everything into "".
    func HistBy(samples []Sample, name, label string) map[string]Snapshot {
      out := map[string]Snapshot{}
      for _, s := range samples {
        if s.Name != name || s.Hist == nil {
          continue
        }
        k := ""
        if label != "" {
          k = s.Label(label)
        }
        out[k] = out[k].Add(*s.Hist)
      }
      return out
    }
    tenant incident, whole fleetbeforeduring
    p5037 ms37 ms
    mean68 ms144 ms
    p99457 ms2,224 ms

    I report p50, p99 and p99.9 from histograms. To get a fleet percentile I sum the buckets across hosts and then take the quantile.

    H

    Traces

    spans, propagation, sampling
    context propagationpseudo code
    outgoing call:
      header traceparent1 = "00-" + trace_id + "-" + span_id + "-01"
    
    incoming request:
      parse traceparent: version 00, 32 hex, 16 hex, 2 hex flags
      IF valid: new span joins that trace, parent = that span_id
      ELSE: start a new trace                    // never fail the request2
    
    start span(context, name):
      parent = span in the context, else the remote parent
      span_id = random 8 bytes3
    1. 1The W3C Trace Context header. Every OpenTelemetry SDK reads and writes it, so spans from different languages join one trace.
    2. 2Telemetry must not take the service down. A bad header costs you one link in the trace, nothing more.
    3. 3A trace ID is 16 bytes and a span ID 8, random, so no coordination is needed to make them.
    Tested source Go: traceparent · Go: start a span, inject, extract
    Go: traceparentgo
    // Traceparent formats the header: version 00, the 32-hex trace ID, the 16-hex parent span ID,
    // and flags (01 = sampled). Example: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
    func (sc SpanContext) Traceparent() string {
      flags := "00"
      if sc.Sampled {
        flags = "01"
      }
      return "00-" + sc.TraceID.String() + "-" + sc.SpanID.String() + "-" + flags
    }
    
    // ParseTraceparent reads the header. An all-zero trace or span ID is invalid.
    func ParseTraceparent(v string) (SpanContext, error) {
      parts := strings.Split(strings.TrimSpace(v), "-")
      if len(parts) != 4 || parts[0] != "00" || len(parts[1]) != 32 || len(parts[2]) != 16 || len(parts[3]) != 2 {
        return SpanContext{}, fmt.Errorf("%w: %q", ErrBadTraceparent, v)
      }
      var sc SpanContext
      if _, err := hex.Decode(sc.TraceID[:], []byte(parts[1])); err != nil {
        return SpanContext{}, fmt.Errorf("%w: trace ID: %w", ErrBadTraceparent, err)
      }
      if _, err := hex.Decode(sc.SpanID[:], []byte(parts[2])); err != nil {
        return SpanContext{}, fmt.Errorf("%w: span ID: %w", ErrBadTraceparent, err)
      }
      flags, err := hex.DecodeString(parts[3])
      if err != nil {
        return SpanContext{}, fmt.Errorf("%w: flags: %w", ErrBadTraceparent, err)
      }
      if sc.TraceID == (TraceID{}) || sc.SpanID.IsZero() {
        return SpanContext{}, fmt.Errorf("%w: zero ID", ErrBadTraceparent)
      }
      sc.Sampled = flags[0]&1 == 1
      return sc, nil
    }
    Go: start a span, inject, extractgo
    // Start begins a span. Its parent is the span already in ctx, or else the remote parent that
    // Extract found in the request headers, or else none: a new trace.
    func (t *Tracer) Start(ctx context.Context, name string) (context.Context, *Span) {
      s := &Span{Name: name, Service: t.Service, Start: t.Clock.Now(), Attrs: map[string]string{}, tracer: t}
      t.mu.Lock()
      binary.BigEndian.PutUint64(s.SpanID[:], t.rng.Uint64()|1)
      if p := SpanFrom(ctx); p != nil {
        s.TraceID, s.Parent = p.TraceID, p.SpanID
      } else if rc, ok := ctx.Value(remoteKey{}).(SpanContext); ok {
        s.TraceID, s.Parent = rc.TraceID, rc.SpanID
      } else {
        binary.BigEndian.PutUint64(s.TraceID[:8], t.rng.Uint64())
        binary.BigEndian.PutUint64(s.TraceID[8:], t.rng.Uint64()|1)
      }
      t.mu.Unlock()
      return context.WithValue(ctx, spanKey{}, s), s
    }
    
    // Inject writes the current span into outgoing headers.
    func Inject(ctx context.Context, h http.Header) {
      if s := SpanFrom(ctx); s != nil {
        h.Set(TraceparentHeader, s.Context().Traceparent())
      }
    }
    
    // Extract reads the caller's span from incoming headers. A missing or bad header starts a new
    // trace; it never fails the request.
    func Extract(ctx context.Context, h http.Header) context.Context {
      sc, err := ParseTraceparent(h.Get(TraceparentHeader))
      if err != nil {
        return ctx
      }
      return context.WithValue(ctx, remoteKey{}, sc)
    }
    samplingdecidesslow or failed kept
    Head, 1%At the first span5 of 497
    Tail: all slow or failed, 1% of the restAfter the trace ends, in the collector497 of 497

    Tenant incident replay: 36,000 requests; tail sampling kept 872 traces (2.4%). The collector must hold each trace until it ends.

    Each service starts a span per operation and sends the trace context in the traceparent header. A collector keeps every slow or failed trace and samples the rest.

    I

    Cardinality

    labels multiply
    series for one latency histogram14series per set×20routes×50instances=14,000 ✓× 1,000 tenants=14 million !× 1,000,000 user IDs=14 billion ✕
    series limitpseudo code
    series(name, labels):
      IF (name, labels) exists: RETURN it
      IF name already has max_series1 series:
        count it as dropped2
        labels = { overflow = "true" }           // one shared series
      create the series
    1. 1A cap per metric. In the lab, 10,000 user IDs made 100 real series and one overflow series, and every login was still counted.
    2. 2Alert on drops: a new label is exploding, usually an ID inside a URL path.
    Tested source Go: registry with a series limit
    Go: registry with a series limitgo
    // get returns the series for name and labels. A new label set beyond the metric's limit goes
    // into one overflow series instead, and the registry counts it as dropped.
    func (r *Registry) get(name string, labels []Label, hist bool) *series {
      k := key(name, labels)
      if s, ok := r.series[k]; ok {
        return s
      }
      if r.perMetric[name] >= r.maxSeries {
        r.dropped++
        labels = []Label{OverflowLabel}
        k = key(name, labels)
        if s, ok := r.series[k]; ok {
          return s
        }
      }
      s := &series{name: name, labels: slices.Clone(labels)}
      if hist {
        s.hist = NewHistogram(r.buckets)
      }
      r.series[k] = s
      r.perMetric[name]++
      return s
    }
    • The Prometheus guidance: most metrics need no labels; keep a metric's cardinality below 10, and rethink one that can pass 100.
    • Label the route template (/orders/:id), never the raw path.

    Every label value multiplies the series count. I keep metric labels to bounded sets like route and status, and put user and order IDs on spans and logs.

    J

    SLIs, SLOs and burn-rate alerts

    page on budget, not on blips
    burn rate, log scale · 10% errors for 1 hour0.1114.4100page fires: minute 8, 2.08% of budget1 h window alone: resets at 111resets at 640 min30 min60 min90 min120 min1 h window5 min window- - threshold 14.4
    scenario, 99.9% SLOpage 1 hpage 6 hticket 3 dnaive 5 min
    10% errors for 1 hour8 min (2.08%)21 min (5.09%)38 min (9.03%)0 min (0.23%)
    0.8% errors for 12 hoursnever4.5 h (4.98%)8.2 h (9.13%)0 min (0.02%)
    0.2% errors for 4 daysnevernever34.1 h (9.48%)2 min (0.01%)
    5% errors for 5 minutesnevernevernever0 min (0.12%)

    Time to fire after the errors start, and the share of the 30-day budget spent by then. Thresholds from the SRE Workbook, chapter "Alerting on SLOs".

    termexample
    SLI: good events ÷ all eventsRequests with no 5xx, answered in under 300 ms
    SLO: the target for the SLI99.9% over a rolling 30 days
    Error budget: 1 - SLO0.1% of requests, or 43.2 min a month
    Burn rate: error ratio ÷ budget ratio1% errors on a 99.9% SLO burns at 10
    multiwindow burn-rate rulepseudo code
    allowed = 1 - SLO                    // 0.001 for 99.9%
    
    fires(rule, minute):
      long  = error ratio over the last rule.long window
      short = error ratio over the last rule.short window
      RETURN long  >= rule.burn × allowed1
         AND short >= rule.burn × allowed2
    
    rules:  1 h and 5 min,   burn 14.4  → page      // 2% of 30 days
            6 h and 30 min,  burn 6     → page      // 5%
            3 d and 6 h,     burn 1     → ticket3    // 10%
    1. 1The long window measures how much budget is gone. 14.4 over 1 hour is 2% of 30 days.
    2. 2The short window, 1/12 of the long one, checks the burn is still happening. The page reset 4 min after the errors stopped, not 51.
    3. 3A slow leak is not an emergency. It goes to a ticket queue for working hours.
    Tested source Go: burn-rate rule
    Go: burn-rate rulego
    // Firing reports, for each minute, whether the rule fires. The SLO is a target such as 0.999,
    // so the allowed error ratio is 1 - SLO.
    func (r BurnRule) Firing(ratio []float64, slo float64) []bool {
      w := newWindows(ratio)
      limit := r.Burn * (1 - slo)
      out := make([]bool, len(ratio))
      for m := range ratio {
        long := w.mean(m, r.Long) >= limit
        short := r.Short == 0 || w.mean(m, r.Short) >= limit
        out[m] = long && short
      }
      return out
    }

    My SLI is the share of requests that succeed within 300 ms. I page when the 30-day budget burns 14.4 times too fast over an hour and five minutes, or 6 times over six hours and thirty minutes.

    K

    Try it: the debug path

    recorded from 36,000 requests per incident
    incident
    Traffic: requests a second010200 min15 min30 min45 min60 min
    Errors: 5xx share0%1%2%0 min15 min30 min45 min60 min
    Latency: p99 (violet), mean (dashed), p50 (grey)0 ms1,250 ms2,500 ms0 min15 min30 min45 min60 min
    Saturation: requests in flight, busiest host00.510 min15 min30 min45 min60 min

    From minute 30, the p99 rose from 457 ms to 2,224 ms. The median moved from 37 to 37 ms and the mean from 68 to 144 ms. Traffic did not change. Next: find who is slow.

    Each incident sends 600 requests a minute for 60 minutes through the instrumented service on 6 hosts, with a simulated clock. Charts read the per-minute scrapes, as a query over a 1-minute window would.

    The dashboard tells me the p99 rose at minute 30. Breaking down by tenant isolates acme; the slowest trace shows payments.charge; and its log line has the same trace ID.

    L

    The debug path

    the same order every time
    stepquestionwhere
    1Which SLI moved, since when, how much?Golden-signal dashboard
    2What changed just before? Deploy, flag, config, traffic.Deploy markers, change log
    3Is a dependency slow or failing?Client-side RED per dependency
    4Is a resource saturated?USE: CPU, memory, pools, queues
    5One host, tenant, route, zone?Break down the SLI by each label
    6Where does one slow request spend time?Trace, then its logs by trace ID
    • All hosts rising together means a shared cause. In the tenant incident every host rose about 5 times; only one tenant and one dependency stood out.
    • One host rising alone means a host problem. On web-4 the trace spent 3,681 of 3,702 ms outside every child span.
    • Mitigate first: roll back, shed the tenant, drain the host. Find the root cause after.

    I scope the symptom, check what changed, then break down by dimension until one value stands out, and only then open traces and logs.

    M

    Dashboards for launch day

    one screen, top to bottom
    rowpanels
    1. UsersSLO burn rate, budget left; p50 and p99 latency; error ratio; traffic
    2. BusinessSign-ups, orders or payments a minute, against last week
    3. DependenciesClient-side rate, errors, p99 for each database, cache, provider
    4. SaturationCPU, memory, connection pools, queue length and age, disk
    5. ChangesDeploy and feature-flag markers on every chart
    • Every latency panel shows percentiles from histograms, never an average alone.
    • Load-test before launch, so you know the values of a healthy system.
    • Check that the alerts fire: break a staging dependency on purpose.

    On launch day I watch the SLO burn and the golden signals first, then each dependency, then saturation, with deploys marked on every chart.

    N

    Failure cases

    when observability itself fails
    eventresultfixsaved by
    A user ID goes into a metric labelMillions of series; the metrics store runs out of memory.A series cap per metric, an alert on drops; IDs go on spans and logs.Series limit
    An alert on every error spikePages for blips nobody can act on. People start to ignore pages.Burn-rate alerts: the 5-minute blip spent 0.58% of budget and paged nobody.Burn-rate rules
    Logs without a trace IDYou find the slow trace but not its error message.A log handler that adds the trace ID from the context to every line.Context handler
    A proxy drops the traceparent headerThe trace breaks into two traces at that hop.Forward trace headers in every proxy and queue message.OpenTelemetry
    Head sampling at 1%It kept 5 of 497 slow or failed traces.Tail sampling in the collector keeps all of them.OTel collector
    A dashboard averages host p99sIt said 355 ms when the fleet p99 was 1,592 ms.Sum the histogram buckets, then take the quantile.Histograms
    The service stops sending metricsIts graphs go flat, and no alert fires.Alert on a missing target or an absent series.Prometheus
    The monitoring stack shares the outageYou are blind exactly when you need to see.Run monitoring outside the region it watches, with an external probe.External probe
    O

    Scale ladder

    start simple; climb only on a signal
    Each step adds one component1Logs, RED metrics2+ SLO burn alerts3+ tracing4+ collector, tail5+ long-term storemore load →
    Series for one latency histogram1k10k100k1M10M100M1B10B100B20 routes × 50 hosts: 14,00020 routes × 50 hosts14,000Prometheus example: 100,000: finePrometheus example100,000: fine+ 1,000 tenants: 14 million+ 1,000 tenants14 million+ 1M user IDs: 14 billion+ 1M user IDs14 billiontime series, log scale
    stepaddit handlesmove up when you see
    1Structured logs, RED metrics, an external probe. One Prometheus.One or a few services. You know that it is broken and roughly where.Pages for blips, or outages found by users first.
    2SLOs and multiwindow burn-rate alerts, routed by Alertmanager.Pages that mean real budget loss. Tickets for slow leaks.A request crosses several services and you cannot tell which is slow.
    3Tracing with OpenTelemetry, trace IDs in every log line, head sampling.Where a request spends time, across services.The slow traces you need were not sampled, or trace storage costs too much.
    4A collector tier with tail sampling, batching, redaction.Every slow or failed trace kept, at a few percent of volume.One Prometheus runs out of memory, or you need a year of history.
    5Long-term metrics storage, sharded scraping, recording rules.Many clusters, months of data, one query view.Top of the ladder.

    14 series per label set: 11 buckets, +Inf, sum and count. The 100,000 example (10,000 nodes) is from the Prometheus instrumentation guide, which calls double-digit millions too much for one server.

    I start with structured logs, RED metrics and one external probe. I add SLO alerts before launch, tracing when requests cross several services, and tail sampling when trace volume costs too much.

    P

    Drill

    predict, then reveal

    0 of 9 known

    1. Host A serves 10,000 requests with p99 = 90 ms. A canary serves 500 with p99 = 1.6 s. What is the fleet p99?

    2. During the tenant incident, the median stayed at 37 ms. Why did the p99 still rise about 5 times?

    3. You want latency per customer, and you have 200,000 customers. Do you add a customer label to the histogram?

    4. An SLO of 99.9% over 30 days. 10% of requests fail from now on. When does the fast page fire, and how much budget is gone?

    5. Why does each burn-rate alert need a short window too?

    6. 5% of requests fail for 5 minutes, then the service recovers. Should anyone be paged?

    7. You sample 1% of traces at the start of each request. What happens to the slow ones?

    8. Every host got slower at the same time. What does that tell you?

    9. A trace takes 3.7 s, but its child spans add up to 21 ms. Where did the time go?

    Q

    Numbers to say

    measured, derived, cited
    burn rules
    Page: 14.4× over 1 h and 5 min; 6× over 6 h and 30 min. Ticket: 1× over 3 d and 6 h.
    budget
    They fire at 2%, 5% and 10% of a 30-day budget.
    detect
    Minutes = long window × threshold ÷ burn. At 10% errors: 60 × 14.4 ÷ 100 = 8.6 min.
    99.9%
    43.2 minutes, or 0.1% of requests, a month.
    p99
    Tenant incident: p99 457 → 2,224 ms; median 37 → 37 ms.
    sampling
    Head 1% kept 5 of 497 slow traces; tail kept all, at 2.4% of volume.
    labels
    Keep most metrics under 10 series; rethink one that can pass 100.

    Burn rules: SRE Workbook. Labels: Prometheus instrumentation guide. Incident and sampling: replays in the lab with a seeded random source, 36,000 requests each.