Skip to content

Distributed Tracing#

Problem statement (interviewer prompt)

Design a distributed tracing system (Jaeger / Tempo / Zipkin): collect spans from every microservice, sample intelligently (head + tail), reconstruct end-to-end traces, link traces ↔ metrics ↔ logs, and handle 100k+ trace ingest QPS.

flowchart LR
  S[Service A]
  S2[Service B]
  S3[Service C]
  COL[Collector]
  TR[(Trace store<br/>Jaeger / Tempo)]
  UI[UI]
  S --> S2 --> S3
  S -. span .-> COL
  S2 -. span .-> COL
  S3 -. span .-> COL
  COL --> TR --> UI

    classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
    classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
    classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
    classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
    classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
    classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
    classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
    classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
    classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
    classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
    class S,S2,S3,COL,UI service;
    class TR datastore;
flowchart TB
  subgraph Instrumented[Services]
    A[Service A]
    B[Service B]
    C[Service C]
    SDK([OTel SDK])
    HOOK[Auto-instrumentation]
  end

  subgraph Context[Context propagation]
    HDR[W3C traceparent / tracestate]
    BAG[Baggage]
    QP[[Queue headers]]
  end

  subgraph Collection
    OTELC[OTel Collector]
    SAMP1([Head sampler])
    SAMP2([Tail sampler])
    BATCH[Batching exporter]
  end

  subgraph Store[Storage]
    SS[Span storage<br/>Cassandra / OpenSearch / Tempo S3]
    IDX[(Indexes:<br/>service, operation, ts)]
    BLOOM[Bloom filter for trace IDs]
  end

  subgraph Query
    UI[Jaeger / Tempo UI]
    LINK[Link: metric exemplar -> trace]
    SVCMAP[Service map auto-derived]
  end

  subgraph SLO
    LAT[p99 latency]
    ERR[Error spans]
    SAM[Adaptive sampling tuning]
  end

  Instrumented --> Context --> Collection --> Store --> Query
  SLO --- Query

    classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
    classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
    classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
    classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
    classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
    classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
    classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
    classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
    classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
    classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
    class A,B,C,HOOK,BAG,BATCH,SVCMAP,LAT,ERR,SAM service;
    class IDX datastore;
    class QP queue;
    class SAMP1,SAMP2 compute;
    class SS storage;
    class SDK,HDR,OTELC,BLOOM,UI,LINK obs;

Glossary & fundamentals#

Concepts referenced in this design. Each row links to its canonical page; the tag column shows whether it is a high-level (HLD) or low-level (LLD) concept.

Tag Concept What it is Page
HLD Probabilistic data structures Bloom, HLL, Count-Min, MinHash, t-digest probabilistic-data-structures
HLD Observability metrics, logs, traces, SLOs observability

Quick reference#

Functional#

  • Trace = tree of spans across services.
  • Each span: (trace_id, span_id, parent_id, service, op, ts, duration, attrs).
  • Query by trace ID, service, operation, time range.

Non-functional#

  • Ingestion: 100k+ spans/s.
  • Storage at sampled rate; 1-10% typical head sampling.
  • p99 trace lookup < 1 s.

Trade-offs#

  • Head sampling simpler; tail sampling keeps all errors + slow.
  • OpenTelemetry as standard wins; proprietary SDKs declining.
  • Exemplars link metrics → exact traces.

Refs#

  • OpenTelemetry docs; Jaeger, Tempo, Zipkin papers.
  • Honeycomb / Lightstep blog series.
  • Google Dapper paper (the original).

FAQ#

How does distributed tracing work?#

Each request carries a trace ID and span context across services. Instrumentation libraries emit spans to a collector, which assembles them into end-to-end traces stored for query.

What is the difference between head and tail sampling?#

Head sampling decides at the start of a trace whether to keep it, cheap but blind to outcomes. Tail sampling buffers all spans and selects after seeing errors or latency.

How is trace context propagated?#

Services pass a W3C traceparent header on every outbound call. Instrumentation extracts and injects it automatically for HTTP, gRPC, and message queue protocols.

What is OpenTelemetry?#

OpenTelemetry is a CNCF standard for instrumentation. It defines APIs and SDKs for traces, metrics, and logs, plus a collector that exports to backends like Jaeger or Tempo.

How does distributed tracing complement logs and metrics?#

Metrics tell you something is wrong, logs explain what happened, and traces show where in a distributed call chain the problem occurred and how it propagated.

How are traces stored at scale?#

Backends use columnar stores like ClickHouse or specialized engines like Tempo to keep span data compact. Indexes on trace ID and service make fast lookups practical.