Service Discovery#
Problem statement (interviewer prompt)
Design a service-discovery system like Consul / Eureka / etcd: services register themselves and their health, clients query the registry to find healthy instances, the registry tolerates partitions, and DNS / xDS / sidecar consumers receive low-latency updates.
flowchart LR
S[Service Instance]
REG[Registry<br/>Consul / etcd / Eureka]
C([Client / Sidecar])
S -. register / heartbeat .-> REG
C -->|lookup| REG
C --> S
classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
class C client;
class S,REG service;
flowchart TB
subgraph Services[Service instances]
A1[Service A pod 1]
A2[Service A pod 2]
B1[Service B pod 1]
end
subgraph Registry
REG[Registry cluster<br/>Consul / etcd / Eureka / ZK]
RAFT[Raft consensus]
KV[(KV store with TTL leases)]
GOSSIP[Gossip / health]
end
subgraph Clients
SDK([Client SDK<br/>watch + cache])
SIDE[Sidecar proxy / Envoy]
DNS[DNS API]
XDS[xDS subscribe]
end
subgraph Health
HC[Active health checks]
PASS[Passive: traffic outcomes]
TTL[Heartbeat TTL leases]
DRAIN[Drain mode]
end
subgraph Mesh[Service mesh integration]
ISTIO[Istio / Linkerd]
SD[xDS service discovery]
MTLS[mTLS identities]
end
Services -. register .-> Registry
Registry --- RAFT
Registry --- KV
Registry --- GOSSIP
Health --> Services
Clients --> Registry
SDK -. cache + watch .-> KV
DNS -. SRV / A .-> KV
XDS --> SIDE
Mesh --- SIDE
classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
class SDK client;
class SIDE,DNS edge;
class A1,A2,B1,REG,RAFT,GOSSIP,XDS,HC,PASS,TTL,DRAIN,ISTIO,SD,MTLS service;
class KV datastore;
Glossary & fundamentals#
Concepts referenced in this design. Each row links to its canonical page; the tag column shows whether it is a high-level (HLD) or low-level (LLD) concept.
| Tag | Concept | What it is | Page |
|---|---|---|---|
HLD |
Load balancer / GSLB | L4/L7 traffic distribution and failover | load-balancer |
HLD |
Raft / Paxos consensus | replicated state machine via majority quorum | consensus-raft-paxos |
HLD |
Service mesh | sidecar mesh, mTLS, traffic policy | service-mesh |
LLD |
Structural patterns | Adapter, Decorator, Facade, Proxy, Composite | structural-patterns |
Quick reference#
Functional#
- Register / deregister service instances.
- Health-checked endpoints.
- Lookup by service name (DNS or API).
- Watch for changes (push update on roster change).
- Tags / metadata for routing decisions.
Non-functional#
- Sub-second propagation typical.
- Registry HA + survives partition.
- Lookups must be fast (cached client-side).
Capacity#
- Tens of thousands of services × instances.
- High read fan-out from clients.
Models#
- Client-side: clients query registry, load balance themselves.
- Server-side: gateway / LB consults registry; clients hit VIP.
- DNS-based: SRV records; simple but slow change propagation.
- xDS (Envoy): explicit subscribe + push.
Trade-offs#
- CP vs AP registry: Consul (CP) vs Eureka (AP) - depends on whether stale routing is acceptable.
- Heartbeat-based TTLs vs active health checks: combine for safety.
- Sidecar mesh simplifies app code but adds operational layer.
Refs#
- Consul, etcd, Eureka, ZooKeeper docs.
- "Service Discovery in a Microservices Architecture" Chris Richardson.
- Envoy xDS protocol spec.
FAQ#
How does service discovery work?#
Services register themselves with a registry on start, send heartbeats, and clients query the registry to find healthy instances. Failed instances are evicted when their lease expires.
What is the difference between client-side and server-side discovery?#
Client-side puts the registry lookup and load balancing in the client. Server-side relies on a load balancer that queries the registry, hiding it from clients.
How does service mesh discovery work?#
Service meshes use a control plane like Istio that pushes endpoint updates via xDS to sidecar proxies. The proxy handles discovery, load balancing, and retries transparently.
Why use Consul vs Eureka vs etcd?#
Consul ships discovery, KV, and service mesh. Eureka is Netflix-style AP-friendly with self-preservation. etcd is a strongly consistent KV often used by Kubernetes for discovery.
How does service discovery survive partitions?#
AP systems like Eureka favor availability, returning possibly stale endpoints. CP systems like etcd refuse writes during partitions but never return stale data to readers.
What is the role of health checks in service discovery?#
Health checks confirm instances are actually serving traffic. Failing checks remove instances from rotation, preventing clients from sending requests to dead or degraded pods.