Manikandan — Manikandan
Microservices

Day 22: Client-Side Discovery

ManikandanManikandan
16 min read·Updated Aug 31, 2022

Client-Side Discovery means the *caller* is responsible for finding a service instance.

Intro

Client-Side Discovery means the caller is responsible for finding a service instance. The caller asks the Service Registry “which healthy instances of policy-service exist right now?”, receives a list of addresses, and picks one itself (round-robin, random, least-loaded) inside its own process. There is no load balancer in the middle, so every call is one network hop instead of two. The price is that discovery and load-balancing logic lives in every client, in every language you use.

Why we need this

In a container or cloud world instance IPs and ports are ephemeral: autoscaling adds and removes instances, deployments replace them, and nodes fail. Day 21 (Service Registry) gave us an authoritative list of healthy instances. Someone still has to use that list on every call. Client-Side Discovery is the option where the caller does it.

Business and technical reasons teams choose it:

  • One less hop. In an insurance claims platform, ClaimsApi calls PolicyService on almost every request. Removing an internal load balancer removes roughly one network round trip and one component that must itself be scaled and made highly available.
  • Smarter routing. The client can use knowledge a generic load balancer lacks: prefer instances in the same availability zone, avoid an instance that just failed, use “power of two choices” based on in-flight request counts, or pin by claim ID for cache locality.
  • No central data-path bottleneck. Traffic flows service-to-service; only the (small, cacheable) registry lookup is centralised.
  • Cost. Internal load balancers and their bandwidth are not free at scale.

What problem it solves

Problem from the topic list: routing all internal calls through load balancers adds network hops.

Without it, the typical setup is: ClaimsApi -> internal LB -> PolicyService. Consequences:

  • Extra latency on every call (an LB hop is typically sub-millisecond to a few milliseconds, but it is paid on every call in a chain of five services).
  • The LB becomes a shared dependency and a possible single point of failure or throughput ceiling; it needs its own capacity planning and health-check tuning.
  • The LB is content-agnostic: it cannot easily do zone-aware, failure-aware or load-aware routing based on what the caller knows.
  • Worse, when teams skip discovery entirely and hardcode IPs or DNS names with long TTLs, calls go to instances that no longer exist, causing connection timeouts during every deployment.

When it is needed (and when it is NOT)

Good fit

  • Latency-sensitive, chatty internal call chains (Claims -> Policy -> Fraud Scoring -> Pricing).
  • You control the client stack (mostly .NET with a shared chassis, see Day 40) so one library implements discovery correctly once.
  • gRPC or long-lived HTTP/2 connections, where an L4 load balancer balances connections, not requests, and leaves one instance hot. Client-side balancing spreads requests properly.
  • You need zone-aware or custom routing.

Poor fit / overkill

  • Polyglot estates with many languages and little platform-team capacity: you would have to maintain and patch a discovery client per language. Prefer Server-Side Discovery (Day 23) or a Service Mesh (Day 48).
  • Your platform already gives you a stable virtual name that load-balances (a Kubernetes ClusterIP Service, or an Azure Container Apps app name). Adding your own client-side layer duplicates it.
  • Very small systems (2-3 services, 1-2 instances each): a plain DNS name or a gateway is enough.
  • Third-party or browser clients: an Angular SPA must never talk to the registry. It calls the API Gateway or BFF (Days 19-20).

How to identify the problem (key signals)

  1. Connection timeouts and 502/503 spikes that line up with deployments or scale-in events - callers are still using addresses of instances that are gone.
  2. Hard-coded hostnames, IPs, or ports in appsettings.json that differ per environment and are edited by hand.
  3. One instance at 90% CPU while others idle, especially with gRPC or HTTP/2 behind an L4 balancer (connection-level, not request-level balancing).
  4. Distributed traces show an extra hop (internal-lb) with non-trivial duration in every span chain between services.
  5. Internal load balancer is in the top of your cost report or shows saturation/connection-limit alerts.
  6. DNS-cache complaints: “it works after we restart the caller” - stale DNS or stale endpoint lists inside the client.
  7. Cross-zone data-transfer charges climbing because callers pick instances in other availability zones.

Flow Diagram

The caller queries the registry, caches the list, and picks an instance itself.

flowchart LR
CA["ClaimsApi"] --> RES["Caching resolver - TTL 15s"]
RES -- "refresh" --> REG[("Registry")]
REG -- "instance list" --> RES
RES --> LB{"Round robin, skip ejected"}
LB --> A["PolicyService .11"]
LB --> B["PolicyService .12"]
LB --> C["PolicyService .13"]
B -. "HttpRequestException: eject 30s" .-> RES

Level 1: Beginner

Analogy. You need a taxi. In server-side discovery you call a dispatcher who sends one. In client-side discovery you get the live list of nearby taxis from an app and choose one yourself. Faster, but you must be able to read the app and handle a taxi that turns out to be unavailable.

Minimal working example (C# 13 / .NET 10 console app, top-level statements). The “registry” is an in-memory dictionary just to show the mechanics.

using System.Threading;
// Pretend this came from the Service Registry
var registry = new Dictionary<string, List<Uri>>
{
["policy-service"] = [new("http://10.0.1.11:8080"), new("http://10.0.1.12:8080"), new("http://10.0.1.13:8080")]
};
var counter = 0;
Uri Pick(string service)
{
var instances = registry[service]; // 1. look up instances
var index = (int)((uint)Interlocked.Increment(ref counter) % (uint)instances.Count);
return instances[index]; // 2. client chooses (round-robin)
}
for (var i = 0; i < 6; i++)
Console.WriteLine($"Call {i + 1} -> {Pick("policy-service")}");

Expected output: calls cycle .11, .12, .13, .11, … (the first call starts at index 1 because the counter is incremented before use; that is fine).

Key ideas for beginners: (1) callers use a logical name (policy-service), never an IP; (2) the list of instances comes from the registry; (3) the caller does the picking.

Level 2: Intermediate

In a real .NET + Angular + database system we hide discovery inside HttpClient so business code only sees http://policy-service/....

Angular never does discovery. It calls one stable URL (the gateway/BFF):

// environment.ts - the SPA knows only the edge
export const environment = { apiBaseUrl: 'https://api.claims.example.com' };

.NET: registry client, caching resolver, and a delegating handler.

using System.Collections.Concurrent;
public interface IRegistryClient
{
Task<IReadOnlyList<Uri>> GetHealthyAsync(string service, CancellationToken ct);
}
public interface IServiceResolver
{
ValueTask<Uri> PickAsync(string service, CancellationToken ct);
void ReportFailure(string service, Uri instance);
}
public sealed class CachingResolver(IRegistryClient registry, TimeProvider time) : IServiceResolver
{
private sealed record Entry(IReadOnlyList<Uri> Instances, DateTimeOffset ExpiresAt);
private static readonly TimeSpan CacheTtl = TimeSpan.FromSeconds(15);
private static readonly TimeSpan EjectFor = TimeSpan.FromSeconds(30);
private readonly ConcurrentDictionary<string, Entry> _cache = new();
private readonly ConcurrentDictionary<(string Service, Uri Instance), DateTimeOffset> _ejected = new();
private int _counter;
public async ValueTask<Uri> PickAsync(string service, CancellationToken ct)
{
var now = time.GetUtcNow();
if (!_cache.TryGetValue(service, out var entry) || entry.ExpiresAt <= now)
{
try
{
var fresh = await registry.GetHealthyAsync(service, ct);
entry = new Entry(fresh, now + CacheTtl);
_cache[service] = entry;
}
catch (HttpRequestException) when (entry is not null)
{
// Registry unreachable: keep serving the last known list (stale is better than down).
}
}
if (entry is null || entry.Instances.Count == 0)
throw new InvalidOperationException($"No instances known for '{service}'.");
var candidates = entry.Instances
.Where(i => !_ejected.TryGetValue((service, i), out var until) || until <= now)
.ToList();
if (candidates.Count == 0) candidates = [.. entry.Instances]; // all ejected: try anyway
var index = (int)((uint)Interlocked.Increment(ref _counter) % (uint)candidates.Count);
return candidates[index];
}
public void ReportFailure(string service, Uri instance) =>
_ejected[(service, instance)] = time.GetUtcNow() + EjectFor;
}
public sealed class ClientSideDiscoveryHandler(IServiceResolver resolver) : DelegatingHandler
{
protected override async Task<HttpResponseMessage> SendAsync(HttpRequestMessage request, CancellationToken ct)
{
var logicalName = request.RequestUri!.Host; // "policy-service"
var instance = await resolver.PickAsync(logicalName, ct);
request.RequestUri = new UriBuilder(request.RequestUri)
{
Scheme = instance.Scheme,
Host = instance.Host,
Port = instance.Port
}.Uri;
try
{
return await base.SendAsync(request, ct);
}
catch (HttpRequestException)
{
resolver.ReportFailure(logicalName, instance); // passive health: stop picking it for 30 s
throw;
}
}
}

Wiring in Program.cs (RegistryHttpClient is your implementation of IRegistryClient, for example a thin wrapper over the Consul or Eureka HTTP API):

var builder = WebApplication.CreateBuilder(args);
builder.Services.AddSingleton(TimeProvider.System);
builder.Services.AddHttpClient<IRegistryClient, RegistryHttpClient>(c =>
c.BaseAddress = new Uri(builder.Configuration["Registry:Url"]!));
builder.Services.AddSingleton<IServiceResolver, CachingResolver>();
builder.Services.AddTransient<ClientSideDiscoveryHandler>();
builder.Services.AddHttpClient<IPolicyClient, PolicyClient>(c =>
c.BaseAddress = new Uri("http://policy-service")) // logical name, not a real host
.AddHttpMessageHandler<ClientSideDiscoveryHandler>();
var app = builder.Build();
app.MapGet("/claims/{id:guid}", async (Guid id, IPolicyClient policies, CancellationToken ct) =>
Results.Ok(await policies.GetCoverageAsync(id, ct)));
app.Run();

Ready-made option. Microsoft ships Microsoft.Extensions.ServiceDiscovery (10.x packages line up with .NET 10), used by .NET Aspire. It provides configuration-based and DNS-based endpoint resolution and plugs into HttpClient via AddServiceDiscovery(). Use it before writing your own; the hand-written version above is for understanding and for registries it has no provider for. Check the package documentation for the load-balancing selector options rather than assuming the default is round-robin.

Databases also have client-side discovery. Npgsql (PostgreSQL) accepts several hosts and picks among them:

Host=pg-a.claims.internal,pg-b.claims.internal;Database=claims;Username=app;Target Session Attributes=prefer-standby;Load Balance Hosts=true

For SQL Server availability groups, the client connects to the AG listener name with MultiSubnetFailover=True in the connection string; that is server-side discovery with a client hint.

Level 3: Advanced

Performance and scalability

  • Cache registry results (5-30 s) and refresh in the background; never call the registry on every request or you recreate the bottleneck you removed.
  • Use a shared SocketsHttpHandler with PooledConnectionLifetime (for example 2-5 minutes) so connections to removed instances are eventually dropped. HttpClient created through IHttpClientFactory already rotates handlers (default lifetime 2 minutes).
  • For HTTP/2 or gRPC, balance per request across multiple connections (one channel per instance). In Grpc.Net.Client, the built-in dns:/// resolver with a load-balancing policy does this for headless Kubernetes Services.
  • Prefer “power of two random choices” over plain round-robin when request costs vary widely (a fraud-scoring call can be 100x heavier than a policy lookup).

Security

  • Registry access must be authenticated and authorised; an attacker who can register a fake policy-service receives claim data. Use ACLs/tokens for registration and mTLS between services.
  • TLS with IP addresses: certificates are issued for names. If you rewrite the host to an IP, hostname validation fails. Either keep the logical name in the TLS SNI/host header via a custom SslClientAuthenticationOptions.TargetHost, issue certificates with matching SANs, or terminate mTLS in a mesh.
  • Never trust registry metadata as an authorisation decision.

Failure modes

  • Stale cache sends traffic to dead instances: combine short TTL, passive ejection (as above) and retries (Day 28).
  • Registry outage: keep serving the last known list; do not fail all calls because discovery is down.
  • Thundering herd on registry after a restart of many clients: add jitter to refresh timers.
  • Self-preservation in Eureka: when many instances miss heartbeats, Eureka stops evicting, so the list can contain dead nodes. Clients must tolerate that.
  • Retry storms: retrying against the same bad instance. Retries must re-resolve or exclude the failed instance.

Common mistakes

  1. Resolving DNS once at startup and caching forever (static HttpClient with no PooledConnectionLifetime).
  2. Writing the discovery code in each service instead of a shared chassis (Day 40).
  3. Health checks that only test “process is up”, so the registry advertises instances whose database is down (see Day 36).
  4. Using client-side discovery for browser or partner clients.
  5. No metric for “endpoints known per service”, so nobody notices when the list shrinks to one.

Level 4: Expert and Architect view

Alternatives compared

AspectClient-Side DiscoveryServer-Side Discovery (Day 23)Service Mesh (Day 48)Platform DNS (K8s ClusterIP / Container Apps name)
Extra network hopNoneOne (router/LB)Local sidecar hop (loopback)Depends on kube-proxy/ingress; usually transparent
Load-balancing intelligenceHigh (custom, zone-aware)Medium (LB features)High (Envoy policies)Low to medium
Client complexityHigh (library per language)LowLow (no app code)Very low
Operational componentsRegistry + client libraryRegistry + LB/routerControl plane + sidecarsPlatform only
Polyglot friendlinessPoorExcellentExcellentExcellent
Failure blast radiusBug in client library affects that serviceLB failure affects everyoneControl plane issues affect config, data plane keeps runningPlatform outage
Best forLatency-critical, homogeneous stacksMixed stacks, simple clientsLarge estates with security/telemetry needsMost teams on AKS/Container Apps

Patterns it combines with: Service Registry (Day 21) as the data source; Self-Registration (Day 24) or 3rd Party Registration (Day 25) to fill the registry; Circuit Breaker and Retry (Days 26 and 28) so discovery failures and bad instances are handled; Health Check API (Day 36) as the source of “healthy”; Microservice Chassis (Day 40) to ship the client once.

ADR-style justification

  • Title: ADR-022 - Use client-side discovery for internal calls from ClaimsApi to PolicyService and FraudScoringService.
  • Context: All internal services are .NET 10, deployed on AKS. p95 latency of the claim-submission flow is 480 ms against a 300 ms target; traces show an internal load balancer adds 2-4 ms per hop across 5 hops, and gRPC traffic to FraudScoringService is unevenly distributed across pods.
  • Decision: Adopt Microsoft.Extensions.ServiceDiscovery in the shared chassis for HTTP calls, and headless-Service DNS with client-side round-robin for gRPC. The Angular SPA continues to call only the API Gateway.
  • Consequences: (+) one fewer hop, even gRPC distribution, zone-aware routing possible. (-) discovery behaviour is now part of the chassis and must be versioned and tested; non-.NET services (for example a Python OCR service) still use ClusterIP names; revisit if the estate becomes polyglot (then evaluate a service mesh).
  • Rejected: internal LB per service (extra hop and cost); full mesh now (operational cost not yet justified).

Azure implementation

Services that implement or support it

  • Azure Kubernetes Service (AKS) with CoreDNS: a normal ClusterIP Service gives a stable virtual IP (server-side style, handled by kube-proxy). A headless Service (clusterIP: None) makes DNS return the individual pod IPs, which is the raw material for client-side discovery and per-request gRPC balancing.
  • Azure Container Apps: apps in the same environment reach each other by app name (for example http://policy-service), and the platform load-balances. Use this instead of building your own layer unless you need custom routing.
  • Eureka on Azure: Azure Spring Apps has an announced retirement (see Microsoft’s retirement announcement), and Microsoft documents a managed Eureka Server for Spring as a Java component in Azure Container Apps. Only relevant if you run Spring services next to .NET ones.
  • Consul (HashiCorp) self-hosted on AKS or VMs is the common registry when you want true client-side discovery with health checks.
  • Azure Service Fabric has a built-in naming service used by its reverse proxy and client libraries (relevant only if you already run Service Fabric).
  • Azure App Configuration can hold endpoint lists for the configuration-based provider in small systems.
  • Azure Monitor / Application Insights for the metrics and traces that prove it works.

How to configure (AKS, headless Service for fraud-scoring)

apiVersion: v1
kind: Service
metadata:
name: fraud-scoring
namespace: claims
spec:
clusterIP: None # headless: DNS returns pod IPs
selector:
app: fraud-scoring
ports:
- name: grpc
port: 8080
// gRPC client: resolve all pod IPs via DNS and balance per request
services.AddGrpcClient<FraudScoring.FraudScoringClient>(o =>
o.Address = new Uri("dns:///fraud-scoring.claims.svc.cluster.local:8080"))
.ConfigureChannel(o =>
{
o.Credentials = Grpc.Core.ChannelCredentials.Insecure; // demo only; use mTLS in production
o.ServiceConfig = new Grpc.Net.Client.Configuration.ServiceConfig
{
LoadBalancingConfigs = { new Grpc.Net.Client.Configuration.RoundRobinConfig() }
};
});

Note that the gRPC DNS resolver re-resolves on a refresh interval and when connections fail; tune it (DnsResolverFactory refresh interval) so scale-outs are noticed quickly. Also set ChannelCredentials.Insecure only when the transport is protected another way.

Pricing and tier considerations (check the Azure pricing pages before committing; figures change)

  • Client-side discovery itself costs nothing on Azure: it uses DNS you already have, or a registry you host.
  • AKS: control-plane tiers (Free, Standard, Premium) plus node VM costs; a self-hosted Consul cluster (3-5 small nodes) adds compute and operations cost.
  • Container Apps: consumption vs dedicated workload profiles; internal name-based calls have no extra discovery charge.
  • The saving you are aiming for is the internal load balancer plus its data-processing charges.

Reference architecture (text) Angular SPA -> Azure Front Door/Application Gateway -> API Gateway (Day 19) -> ClaimsApi pods on AKS. ClaimsApi uses the shared chassis with service discovery: HTTP calls to policy-service resolve through a registry (Consul on AKS) or the platform provider; gRPC calls to fraud-scoring resolve through a headless Service. Services register via Kubernetes (3rd Party Registration, Day 25) or Consul agents. Each service writes to its own database (Azure SQL or Azure Database for PostgreSQL, Day 5). Telemetry flows through OpenTelemetry to Application Insights, with an “endpoints per service” metric alerting when a service has fewer than two known endpoints. A network policy restricts who can query and register in the registry.

Teaching guide for my team

2-minute beginner explanation. “Services are like shops that keep moving. A registry is the live phone book. With client-side discovery, you open the phone book, pick a shop from the current list, and walk there yourself. That saves the trip through a receptionist, but you must know how to read the phone book and what to do if the shop is closed.”

5-minute intermediate explanation. Draw ClaimsApi -> [registry lookup, cached 15 s] -> PolicyService instance B. Cover: logical names in code, a delegating handler that swaps the host, local caching and background refresh, passive ejection of failed instances, and why the SPA does not use this. Contrast with Day 23: same registry, but the router does the picking. Mention the trade-off: you now own client code in every language.

Hands-on exercise. Run three copies of a tiny PolicyService (ports 5001-5003) that return their port number. Implement CachingResolver with an in-memory IRegistryClient. (1) Call /claims/{id} 12 times and confirm rotation over 5001, 5002, 5003. (2) Stop the instance on 5002 while calling in a loop. Expected outcome: at most one or two failed calls, then the ejected instance is skipped for 30 seconds and the remaining calls succeed. (3) Restart 5002, wait past the ejection time, confirm it rejoins the rotation.

Interview-style questions

  1. What is the difference between client-side and server-side discovery? In client-side, the caller queries the registry and chooses the instance; in server-side, the caller talks to a router/LB that queries the registry. Client-side saves a hop but puts logic in every client.
  2. Why is a plain L4 load balancer poor for gRPC? gRPC uses long-lived HTTP/2 connections, so the LB balances connections, not requests, leaving some instances overloaded; client-side balancing (or an L7 proxy) balances per request.
  3. What happens if the registry goes down? Clients keep using their cached instance list (stale but usable) and passively eject failed instances; they must not fail every call just because discovery is unavailable.

Mastery checklist

  • I can draw the request flow for client-side and server-side discovery and state the trade-off in one sentence.
  • I can implement a caching resolver with TTL, jitter and a stale-on-error fallback.
  • I can explain why an unhandled DNS or endpoint cache causes errors during deployments.
  • I can explain TLS hostname validation problems when connecting by IP and give two fixes.
  • I can configure a headless Service and a gRPC channel that balances per request on AKS.
  • I can decide between client-side discovery, platform DNS, and a service mesh for a given estate and write the ADR.
  • I know which callers (Angular, partners) must never use client-side discovery.
  • I can name the metrics that prove discovery is healthy (endpoints per service, resolve latency, ejection count).

Key takeaway

Client-side discovery trades an extra network hop for extra client responsibility: cache the registry list, balance per request, eject bad instances, and centralise all of it in one shared library. If your platform already gives you a load-balanced service name, use that instead.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 25: 3rd Party Registration

3rd Party Registration is a service-discovery pattern where the service itself never talks to the service registry.

Manikandan
Manikandan·13 min read
Microservices

Day 24: Self-Registration

Self-Registration is a service-discovery pattern in which each service instance takes responsibility for announcing itself to the service registry when it starts, keeps that registration alive while it is healthy,...

Manikandan
Manikandan·17 min read
Microservices

Day 23: Server-Side Discovery

Server-side discovery means the caller does not look up service instances at all.

Manikandan
Manikandan·17 min read
Microservices

Day 21: Service Registry

A Service Registry is a live "phone book" for your services.

Manikandan
Manikandan·17 min read