Distributed Logging (ELK / EFK)#
Problem statement (interviewer prompt)
Design a centralised logging platform (ELK / EFK): collect structured logs from thousands of services, parse + enrich + redact PII, index for search (last 30 days hot, 1 year cold), and serve dashboards + alerts at 1M+ events/sec.
flowchart LR
APP[Apps]
AG[Agents<br/>Fluent Bit / Filebeat]
BUS[[Kafka buffer]]
PROC([Ingest / parse])
ES[(Elasticsearch / OpenSearch)]
K[Kibana / Grafana]
APP --> AG --> BUS --> PROC --> ES --> K
classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
class APP,AG service;
class ES datastore;
class BUS queue;
class PROC compute;
class K obs;
flowchart TB
subgraph Sources
APPS[Apps / sidecars]
INFRA[Infra logs: nginx, syslog]
K8S[Kubernetes containers]
AUDIT[Security audit]
end
subgraph Agents[Agents per host]
FB[Fluent Bit / Filebeat / Vector]
REDACT[PII redaction]
SAMPLE[Sampling]
PARSE[Light parse]
end
subgraph Bus[Buffer]
KAFKA[[Kafka topics<br/>per environment]]
DLQ[(Parse-failure DLQ)]
end
subgraph Ingest
PIPE([Logstash / Vector aggregator])
GROK[Grok / regex parse]
ENRICH[GeoIP / k8s metadata enrich]
SCHEMA[Schema enforcement]
end
subgraph Storage[Storage]
ES[(Hot index 7d)]
COLD[(Warm 30d - frozen tier)]
S3[(S3 long-term)]
LOKI[Loki - label-index store option]
end
subgraph Query
KIB[Kibana / Grafana / Splunk UI]
ALERT[Alerting rules]
DASH[Dashboards]
end
subgraph Ops
AUTH[AuthN / multi-tenant]
BILL[Per-team quotas + chargeback]
RETN[Retention policy]
end
Sources --> Agents --> Bus --> Ingest --> Storage --> Query
Ingest --> DLQ
Ops --- Storage
classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
class INFRA edge;
class APPS,K8S,AUDIT,FB,REDACT,SAMPLE,PARSE,GROK,ENRICH,SCHEMA,AUTH,BILL,RETN service;
class DLQ,ES,COLD datastore;
class KAFKA queue;
class PIPE compute;
class S3 storage;
class LOKI,KIB,ALERT,DASH obs;
Glossary & fundamentals#
Concepts referenced in this design. Each row links to its canonical page; the tag column shows whether it is a high-level (HLD) or low-level (LLD) concept.
| Tag | Concept | What it is | Page |
|---|---|---|---|
HLD |
Load balancer / GSLB | L4/L7 traffic distribution and failover | load-balancer |
HLD |
Pub/Sub & message brokers | topics, consumer groups, delivery semantics | pub-sub-pattern |
HLD |
Observability | metrics, logs, traces, SLOs | observability |
HLD |
Service mesh | sidecar mesh, mTLS, traffic policy | service-mesh |
Quick reference#
Functional#
- Collect logs from every host / service.
- Parse, enrich, route by tags.
- Index for free-text search.
- Dashboards + alerts.
- Tiered retention (hot / warm / cold).
Non-functional#
- 100k+ events/s for big estates.
- p99 indexing latency < 30 s.
- 99.9% availability for ingest.
Capacity#
- Logs are the most expensive observability pillar; budget by team.
- ES hot tier: ~1 KB/event, 100M events/day = ~100 GB/day per tenant.
Trade-offs#
- ES inverted index = great search, expensive disk.
- Loki labels-only = cheap storage, weaker search (regex over data).
- CDC vs polling at sources: agents always push.
- Structured JSON logs vs free text: enforce JSON + redaction.
Refs#
- ELK / EFK stack docs; Loki paper.
- "Honeycomb: How we built our datastore" blogs.
- Vector + OpenTelemetry Collector docs.
FAQ#
How does a distributed logging system work?#
Agents on each host ship structured logs to a buffered aggregation pipeline that parses, enriches, redacts PII, and indexes them in Elasticsearch or a similar search engine.
What is the difference between ELK and EFK stacks?#
Both index logs in Elasticsearch and visualize with Kibana. ELK uses Logstash as the shipper; EFK uses Fluentd, which is lighter and Kubernetes-friendly.
How are logs retained cost-effectively?#
Pipelines route recent logs to a hot, fast-search tier and roll older indexes to cold object storage. Queries fall through tiers transparently with longer latency.
How is logging different from distributed tracing?#
Logs are unstructured or semi-structured event records. Traces are structured spans linked by trace IDs that show end-to-end request flow across services.
How do log pipelines handle backpressure?#
Agents buffer to local disk, the pipeline uses Kafka for durable queues, and indexers throttle ingest. Sampling and rate limits prevent runaway services from overwhelming storage.
How is PII protected in centralized logs?#
Pipelines apply pattern-based redaction for emails, phone numbers, and tokens at ingest. Access controls and audit logs further restrict who can query sensitive fields.