Skip to content

Service Discovery#

Problem statement (interviewer prompt)

Design a service-discovery system like Consul / Eureka / etcd: services register themselves and their health, clients query the registry to find healthy instances, the registry tolerates partitions, and DNS / xDS / sidecar consumers receive low-latency updates.

flowchart LR
  S[Service Instance]
  REG[Registry<br/>Consul / etcd / Eureka]
  C([Client / Sidecar])
  S -. register / heartbeat .-> REG
  C -->|lookup| REG
  C --> S

    classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
    classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
    classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
    classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
    classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
    classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
    classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
    classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
    classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
    classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
    class C client;
    class S,REG service;
flowchart TB
  subgraph Services[Service instances]
    A1[Service A pod 1]
    A2[Service A pod 2]
    B1[Service B pod 1]
  end

  subgraph Registry
    REG[Registry cluster<br/>Consul / etcd / Eureka / ZK]
    RAFT[Raft consensus]
    KV[(KV store with TTL leases)]
    GOSSIP[Gossip / health]
  end

  subgraph Clients
    SDK([Client SDK<br/>watch + cache])
    SIDE[Sidecar proxy / Envoy]
    DNS[DNS API]
    XDS[xDS subscribe]
  end

  subgraph Health
    HC[Active health checks]
    PASS[Passive: traffic outcomes]
    TTL[Heartbeat TTL leases]
    DRAIN[Drain mode]
  end

  subgraph Mesh[Service mesh integration]
    ISTIO[Istio / Linkerd]
    SD[xDS service discovery]
    MTLS[mTLS identities]
  end

  Services -. register .-> Registry
  Registry --- RAFT
  Registry --- KV
  Registry --- GOSSIP
  Health --> Services
  Clients --> Registry
  SDK -. cache + watch .-> KV
  DNS -. SRV / A .-> KV
  XDS --> SIDE
  Mesh --- SIDE

    classDef client fill:#dbeafe,stroke:#1e40af,stroke-width:1px,color:#0f172a;
    classDef edge fill:#cffafe,stroke:#0e7490,stroke-width:1px,color:#0f172a;
    classDef service fill:#fef3c7,stroke:#92400e,stroke-width:1px,color:#0f172a;
    classDef datastore fill:#fee2e2,stroke:#991b1b,stroke-width:1px,color:#0f172a;
    classDef cache fill:#fed7aa,stroke:#9a3412,stroke-width:1px,color:#0f172a;
    classDef queue fill:#ede9fe,stroke:#5b21b6,stroke-width:1px,color:#0f172a;
    classDef compute fill:#d1fae5,stroke:#065f46,stroke-width:1px,color:#0f172a;
    classDef storage fill:#e5e7eb,stroke:#374151,stroke-width:1px,color:#0f172a;
    classDef external fill:#fce7f3,stroke:#9d174d,stroke-width:1px,color:#0f172a;
    classDef obs fill:#f3e8ff,stroke:#6b21a8,stroke-width:1px,color:#0f172a;
    class SDK client;
    class SIDE,DNS edge;
    class A1,A2,B1,REG,RAFT,GOSSIP,XDS,HC,PASS,TTL,DRAIN,ISTIO,SD,MTLS service;
    class KV datastore;

Glossary & fundamentals#

Concepts referenced in this design. Each row links to its canonical page; the tag column shows whether it is a high-level (HLD) or low-level (LLD) concept.

Tag Concept What it is Page
HLD Load balancer / GSLB L4/L7 traffic distribution and failover load-balancer
HLD Raft / Paxos consensus replicated state machine via majority quorum consensus-raft-paxos
HLD Service mesh sidecar mesh, mTLS, traffic policy service-mesh
LLD Structural patterns Adapter, Decorator, Facade, Proxy, Composite structural-patterns

Quick reference#

Functional#

  • Register / deregister service instances.
  • Health-checked endpoints.
  • Lookup by service name (DNS or API).
  • Watch for changes (push update on roster change).
  • Tags / metadata for routing decisions.

Non-functional#

  • Sub-second propagation typical.
  • Registry HA + survives partition.
  • Lookups must be fast (cached client-side).

Capacity#

  • Tens of thousands of services × instances.
  • High read fan-out from clients.

Models#

  • Client-side: clients query registry, load balance themselves.
  • Server-side: gateway / LB consults registry; clients hit VIP.
  • DNS-based: SRV records; simple but slow change propagation.
  • xDS (Envoy): explicit subscribe + push.

Trade-offs#

  • CP vs AP registry: Consul (CP) vs Eureka (AP) - depends on whether stale routing is acceptable.
  • Heartbeat-based TTLs vs active health checks: combine for safety.
  • Sidecar mesh simplifies app code but adds operational layer.

Refs#

  • Consul, etcd, Eureka, ZooKeeper docs.
  • "Service Discovery in a Microservices Architecture" Chris Richardson.
  • Envoy xDS protocol spec.

FAQ#

How does service discovery work?#

Services register themselves with a registry on start, send heartbeats, and clients query the registry to find healthy instances. Failed instances are evicted when their lease expires.

What is the difference between client-side and server-side discovery?#

Client-side puts the registry lookup and load balancing in the client. Server-side relies on a load balancer that queries the registry, hiding it from clients.

How does service mesh discovery work?#

Service meshes use a control plane like Istio that pushes endpoint updates via xDS to sidecar proxies. The proxy handles discovery, load balancing, and retries transparently.

Why use Consul vs Eureka vs etcd?#

Consul ships discovery, KV, and service mesh. Eureka is Netflix-style AP-friendly with self-preservation. etcd is a strongly consistent KV often used by Kubernetes for discovery.

How does service discovery survive partitions?#

AP systems like Eureka favor availability, returning possibly stale endpoints. CP systems like etcd refuse writes during partitions but never return stale data to readers.

What is the role of health checks in service discovery?#

Health checks confirm instances are actually serving traffic. Failing checks remove instances from rotation, preventing clients from sending requests to dead or degraded pods.