Metrics & Monitoring#
Problem statement (interviewer prompt)
Design a metrics + monitoring platform (Prometheus / Datadog): collect ~10B time-series samples/day with bounded cardinality, support flexible query/aggregation, evaluate alerting + recording rules, and serve dashboards with sub-second latency.
flowchart LR
APP[Service]
AG[Exporter / Agent]
TSDB[(TSDB<br/>Prometheus / Mimir)]
G[Grafana]
AL[Alertmanager]
PD[PagerDuty]
APP --> AG --> TSDB --> G
TSDB --> AL --> PD
classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
class APP,AG service;
class TSDB datastore;
class G,AL,PD obs;
flowchart TB
subgraph Sources
APPS[Apps - instrumented]
NODE[Node exporters]
CADV[cAdvisor / kubelet]
BB[Blackbox probes]
end
subgraph Collection
PROM[Prometheus scrapers]
OTEL[OTel Collector]
PUSH[Pushgateway]
end
subgraph Store[Storage layer]
LOCAL[Per-Prom WAL + blocks]
LONG[(Long-term: Mimir / Thanos / Cortex / VictoriaMetrics)]
S3[(Object storage)]
DOWN[Downsampled tiers]
end
subgraph Query
PQ[PromQL / Flux / MetricsQL]
GRAF[Grafana dashboards]
REC[Recording rules]
ALERTR[Alert rules]
end
subgraph Alert
AM[Alertmanager / route + dedup]
SILENCE[Silences / Maintenance]
PD[PagerDuty / Opsgenie]
SLACK[Slack]
end
subgraph SLO
SLI[SLI definitions]
SLOC[SLOs + budgets]
BURN[Multi-window burn-rate alerts]
end
Sources --> Collection --> Store
Store --> Query --> GRAF
Query --> Alert --> PD
SLO --- Query
classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
class APPS,NODE,CADV,BB,PUSH,LOCAL,DOWN,REC,SILENCE,SLACK,SLOC service;
class LONG datastore;
class S3 storage;
class PROM,OTEL,PQ,GRAF,ALERTR,AM,PD,SLI,BURN obs;
Glossary & fundamentals#
Concepts referenced in this design. Each row links to its canonical page; the tag column shows whether it is a high-level (HLD) or low-level (LLD) concept.
| Tag | Concept | What it is | Page |
|---|---|---|---|
HLD |
LSM vs B-Tree engines | WAL, memtable, SSTables, compaction | storage-engines-lsm-btree |
HLD |
Observability | metrics, logs, traces, SLOs | observability |
Quick reference#
Functional#
- Pull (Prom) or push (Datadog/StatsD) metric collection.
- Multi-tenant ingestion, retention, dashboards.
- Alerting with deduplication and routing.
- SLOs & error budgets.
Non-functional#
- 100M active series possible per cluster.
- p99 dashboard query < 1 s for common ranges.
- Long-term retention to 1+ year on object storage.
Capacity#
- ~1 B/sample compressed.
- 10M active series × 1 sample/15s = ~700k samples/s.
Trade-offs#
- Pull vs push: pull simpler for service discovery; push for short-lived jobs.
- Cardinality is the killer - guard label values strictly.
- Federation vs single big cluster: scaling pattern.
Refs#
- Prometheus, Cortex, Mimir, Thanos, VictoriaMetrics docs.
- Google SRE book on SLO/SLI.
- "Observability Engineering" Charity Majors.
FAQ#
How does a metrics monitoring system work?#
Targets expose metrics endpoints. A collector scrapes them on a schedule, stores samples in a time-series database, and serves queries for dashboards and alert evaluation.
What is the difference between pull-based and push-based metrics?#
Pull-based systems like Prometheus scrape targets, simplifying service discovery. Push-based systems suit short-lived jobs and high-fanout serverless workloads that cannot be scraped.
How is cardinality controlled in metrics systems?#
Operators avoid high-cardinality labels like user_id, drop noisy labels at ingest, and aggregate where possible so the index does not explode and queries stay fast.
How are alerting rules evaluated?#
The system periodically runs PromQL or equivalent queries. If a rule's condition holds for a duration window, it fires an alert routed to Alertmanager or PagerDuty.
What is a recording rule?#
Recording rules precompute expensive aggregations and store the result as a new time series, making dashboards and alerts faster and more reliable.
How is long-term metric storage handled?#
Solutions like Thanos, Cortex, or Mimir back Prometheus with object storage so old samples are downsampled and retained for years without bloating the in-memory database.
Related Topics#
- Distributed Tracing: the three pillars of observability
- Distributed Logging: correlate logs with metrics
- Observability: metrics + logs + traces overview
- Time-Series Database: where metrics actually live