Skip to content

Metrics & Monitoring#

Problem statement (interviewer prompt)

Design a metrics + monitoring platform (Prometheus / Datadog): collect ~10B time-series samples/day with bounded cardinality, support flexible query/aggregation, evaluate alerting + recording rules, and serve dashboards with sub-second latency.

Prometheus logo, the open-source metrics and monitoring system
Prometheus logo, © The Linux Foundation / CNCF, via Wikimedia Commons.
flowchart LR
  APP[Service]
  AG[Exporter / Agent]
  TSDB[(TSDB<br/>Prometheus / Mimir)]
  G[Grafana]
  AL[Alertmanager]
  PD[PagerDuty]
  APP --> AG --> TSDB --> G
  TSDB --> AL --> PD

    classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
    classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
    classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
    classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
    classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
    classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
    classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
    classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
    classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
    classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
    class APP,AG service;
    class TSDB datastore;
    class G,AL,PD obs;
flowchart TB
  subgraph Sources
    APPS[Apps - instrumented]
    NODE[Node exporters]
    CADV[cAdvisor / kubelet]
    BB[Blackbox probes]
  end

  subgraph Collection
    PROM[Prometheus scrapers]
    OTEL[OTel Collector]
    PUSH[Pushgateway]
  end

  subgraph Store[Storage layer]
    LOCAL[Per-Prom WAL + blocks]
    LONG[(Long-term: Mimir / Thanos / Cortex / VictoriaMetrics)]
    S3[(Object storage)]
    DOWN[Downsampled tiers]
  end

  subgraph Query
    PQ[PromQL / Flux / MetricsQL]
    GRAF[Grafana dashboards]
    REC[Recording rules]
    ALERTR[Alert rules]
  end

  subgraph Alert
    AM[Alertmanager / route + dedup]
    SILENCE[Silences / Maintenance]
    PD[PagerDuty / Opsgenie]
    SLACK[Slack]
  end

  subgraph SLO
    SLI[SLI definitions]
    SLOC[SLOs + budgets]
    BURN[Multi-window burn-rate alerts]
  end

  Sources --> Collection --> Store
  Store --> Query --> GRAF
  Query --> Alert --> PD
  SLO --- Query

    classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
    classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
    classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
    classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
    classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
    classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
    classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
    classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
    classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
    classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
    class APPS,NODE,CADV,BB,PUSH,LOCAL,DOWN,REC,SILENCE,SLACK,SLOC service;
    class LONG datastore;
    class S3 storage;
    class PROM,OTEL,PQ,GRAF,ALERTR,AM,PD,SLI,BURN obs;

Glossary & fundamentals#

Concepts referenced in this design. Each row links to its canonical page; the tag column shows whether it is a high-level (HLD) or low-level (LLD) concept.

Tag Concept What it is Page
HLD LSM vs B-Tree engines WAL, memtable, SSTables, compaction storage-engines-lsm-btree
HLD Observability metrics, logs, traces, SLOs observability

Quick reference#

Functional#

  • Pull (Prom) or push (Datadog/StatsD) metric collection.
  • Multi-tenant ingestion, retention, dashboards.
  • Alerting with deduplication and routing.
  • SLOs & error budgets.

Non-functional#

  • 100M active series possible per cluster.
  • p99 dashboard query < 1 s for common ranges.
  • Long-term retention to 1+ year on object storage.

Capacity#

  • ~1 B/sample compressed.
  • 10M active series × 1 sample/15s = ~700k samples/s.

Trade-offs#

  • Pull vs push: pull simpler for service discovery; push for short-lived jobs.
  • Cardinality is the killer - guard label values strictly.
  • Federation vs single big cluster: scaling pattern.

Refs#

  • Prometheus, Cortex, Mimir, Thanos, VictoriaMetrics docs.
  • Google SRE book on SLO/SLI.
  • "Observability Engineering" Charity Majors.

FAQ#

How does a metrics monitoring system work?#

Targets expose metrics endpoints. A collector scrapes them on a schedule, stores samples in a time-series database, and serves queries for dashboards and alert evaluation.

What is the difference between pull-based and push-based metrics?#

Pull-based systems like Prometheus scrape targets, simplifying service discovery. Push-based systems suit short-lived jobs and high-fanout serverless workloads that cannot be scraped.

How is cardinality controlled in metrics systems?#

Operators avoid high-cardinality labels like user_id, drop noisy labels at ingest, and aggregate where possible so the index does not explode and queries stay fast.

How are alerting rules evaluated?#

The system periodically runs PromQL or equivalent queries. If a rule's condition holds for a duration window, it fires an alert routed to Alertmanager or PagerDuty.

What is a recording rule?#

Recording rules precompute expensive aggregations and store the result as a new time series, making dashboards and alerts faster and more reliable.

How is long-term metric storage handled?#

Solutions like Thanos, Cortex, or Mimir back Prometheus with object storage so old samples are downsampled and retained for years without bloating the in-memory database.