Client-Side Discovery means the *caller* is responsible for finding a service instance.
Intro
Client-Side Discovery means the caller is responsible for finding a service instance. The caller asks the Service Registry “which healthy instances of policy-service exist right now?”, receives a list of addresses, and picks one itself (round-robin, random, least-loaded) inside its own process. There is no load balancer in the middle, so every call is one network hop instead of two. The price is that discovery and load-balancing logic lives in every client, in every language you use.
Why we need this
In a container or cloud world instance IPs and ports are ephemeral: autoscaling adds and removes instances, deployments replace them, and nodes fail. Day 21 (Service Registry) gave us an authoritative list of healthy instances. Someone still has to use that list on every call. Client-Side Discovery is the option where the caller does it.
Business and technical reasons teams choose it:
- One less hop. In an insurance claims platform,
ClaimsApicallsPolicyServiceon almost every request. Removing an internal load balancer removes roughly one network round trip and one component that must itself be scaled and made highly available. - Smarter routing. The client can use knowledge a generic load balancer lacks: prefer instances in the same availability zone, avoid an instance that just failed, use “power of two choices” based on in-flight request counts, or pin by claim ID for cache locality.
- No central data-path bottleneck. Traffic flows service-to-service; only the (small, cacheable) registry lookup is centralised.
- Cost. Internal load balancers and their bandwidth are not free at scale.
What problem it solves
Problem from the topic list: routing all internal calls through load balancers adds network hops.
Without it, the typical setup is: ClaimsApi -> internal LB -> PolicyService. Consequences:
- Extra latency on every call (an LB hop is typically sub-millisecond to a few milliseconds, but it is paid on every call in a chain of five services).
- The LB becomes a shared dependency and a possible single point of failure or throughput ceiling; it needs its own capacity planning and health-check tuning.
- The LB is content-agnostic: it cannot easily do zone-aware, failure-aware or load-aware routing based on what the caller knows.
- Worse, when teams skip discovery entirely and hardcode IPs or DNS names with long TTLs, calls go to instances that no longer exist, causing connection timeouts during every deployment.
When it is needed (and when it is NOT)
Good fit
- Latency-sensitive, chatty internal call chains (Claims -> Policy -> Fraud Scoring -> Pricing).
- You control the client stack (mostly .NET with a shared chassis, see Day 40) so one library implements discovery correctly once.
- gRPC or long-lived HTTP/2 connections, where an L4 load balancer balances connections, not requests, and leaves one instance hot. Client-side balancing spreads requests properly.
- You need zone-aware or custom routing.
Poor fit / overkill
- Polyglot estates with many languages and little platform-team capacity: you would have to maintain and patch a discovery client per language. Prefer Server-Side Discovery (Day 23) or a Service Mesh (Day 48).
- Your platform already gives you a stable virtual name that load-balances (a Kubernetes
ClusterIPService, or an Azure Container Apps app name). Adding your own client-side layer duplicates it. - Very small systems (2-3 services, 1-2 instances each): a plain DNS name or a gateway is enough.
- Third-party or browser clients: an Angular SPA must never talk to the registry. It calls the API Gateway or BFF (Days 19-20).
How to identify the problem (key signals)
- Connection timeouts and
502/503spikes that line up with deployments or scale-in events - callers are still using addresses of instances that are gone. - Hard-coded hostnames, IPs, or ports in
appsettings.jsonthat differ per environment and are edited by hand. - One instance at 90% CPU while others idle, especially with gRPC or HTTP/2 behind an L4 balancer (connection-level, not request-level balancing).
- Distributed traces show an extra hop (
internal-lb) with non-trivial duration in every span chain between services. - Internal load balancer is in the top of your cost report or shows saturation/connection-limit alerts.
- DNS-cache complaints: “it works after we restart the caller” - stale DNS or stale endpoint lists inside the client.
- Cross-zone data-transfer charges climbing because callers pick instances in other availability zones.
Flow Diagram
The caller queries the registry, caches the list, and picks an instance itself.
flowchart LR CA["ClaimsApi"] --> RES["Caching resolver - TTL 15s"] RES -- "refresh" --> REG[("Registry")] REG -- "instance list" --> RES RES --> LB{"Round robin, skip ejected"} LB --> A["PolicyService .11"] LB --> B["PolicyService .12"] LB --> C["PolicyService .13"] B -. "HttpRequestException: eject 30s" .-> RESLevel 1: Beginner
Analogy. You need a taxi. In server-side discovery you call a dispatcher who sends one. In client-side discovery you get the live list of nearby taxis from an app and choose one yourself. Faster, but you must be able to read the app and handle a taxi that turns out to be unavailable.
Minimal working example (C# 13 / .NET 10 console app, top-level statements). The “registry” is an in-memory dictionary just to show the mechanics.
using System.Threading;
// Pretend this came from the Service Registryvar registry = new Dictionary<string, List<Uri>>{ ["policy-service"] = [new("http://10.0.1.11:8080"), new("http://10.0.1.12:8080"), new("http://10.0.1.13:8080")]};
var counter = 0;
Uri Pick(string service){ var instances = registry[service]; // 1. look up instances var index = (int)((uint)Interlocked.Increment(ref counter) % (uint)instances.Count); return instances[index]; // 2. client chooses (round-robin)}
for (var i = 0; i < 6; i++) Console.WriteLine($"Call {i + 1} -> {Pick("policy-service")}");Expected output: calls cycle .11, .12, .13, .11, … (the first call starts at index 1 because the counter is incremented before use; that is fine).
Key ideas for beginners: (1) callers use a logical name (policy-service), never an IP; (2) the list of instances comes from the registry; (3) the caller does the picking.
Level 2: Intermediate
In a real .NET + Angular + database system we hide discovery inside HttpClient so business code only sees http://policy-service/....
Angular never does discovery. It calls one stable URL (the gateway/BFF):
// environment.ts - the SPA knows only the edgeexport const environment = { apiBaseUrl: 'https://api.claims.example.com' };.NET: registry client, caching resolver, and a delegating handler.
using System.Collections.Concurrent;
public interface IRegistryClient{ Task<IReadOnlyList<Uri>> GetHealthyAsync(string service, CancellationToken ct);}
public interface IServiceResolver{ ValueTask<Uri> PickAsync(string service, CancellationToken ct); void ReportFailure(string service, Uri instance);}
public sealed class CachingResolver(IRegistryClient registry, TimeProvider time) : IServiceResolver{ private sealed record Entry(IReadOnlyList<Uri> Instances, DateTimeOffset ExpiresAt);
private static readonly TimeSpan CacheTtl = TimeSpan.FromSeconds(15); private static readonly TimeSpan EjectFor = TimeSpan.FromSeconds(30);
private readonly ConcurrentDictionary<string, Entry> _cache = new(); private readonly ConcurrentDictionary<(string Service, Uri Instance), DateTimeOffset> _ejected = new(); private int _counter;
public async ValueTask<Uri> PickAsync(string service, CancellationToken ct) { var now = time.GetUtcNow();
if (!_cache.TryGetValue(service, out var entry) || entry.ExpiresAt <= now) { try { var fresh = await registry.GetHealthyAsync(service, ct); entry = new Entry(fresh, now + CacheTtl); _cache[service] = entry; } catch (HttpRequestException) when (entry is not null) { // Registry unreachable: keep serving the last known list (stale is better than down). } }
if (entry is null || entry.Instances.Count == 0) throw new InvalidOperationException($"No instances known for '{service}'.");
var candidates = entry.Instances .Where(i => !_ejected.TryGetValue((service, i), out var until) || until <= now) .ToList(); if (candidates.Count == 0) candidates = [.. entry.Instances]; // all ejected: try anyway
var index = (int)((uint)Interlocked.Increment(ref _counter) % (uint)candidates.Count); return candidates[index]; }
public void ReportFailure(string service, Uri instance) => _ejected[(service, instance)] = time.GetUtcNow() + EjectFor;}
public sealed class ClientSideDiscoveryHandler(IServiceResolver resolver) : DelegatingHandler{ protected override async Task<HttpResponseMessage> SendAsync(HttpRequestMessage request, CancellationToken ct) { var logicalName = request.RequestUri!.Host; // "policy-service" var instance = await resolver.PickAsync(logicalName, ct);
request.RequestUri = new UriBuilder(request.RequestUri) { Scheme = instance.Scheme, Host = instance.Host, Port = instance.Port }.Uri;
try { return await base.SendAsync(request, ct); } catch (HttpRequestException) { resolver.ReportFailure(logicalName, instance); // passive health: stop picking it for 30 s throw; } }}Wiring in Program.cs (RegistryHttpClient is your implementation of IRegistryClient, for example a thin wrapper over the Consul or Eureka HTTP API):
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddSingleton(TimeProvider.System);builder.Services.AddHttpClient<IRegistryClient, RegistryHttpClient>(c => c.BaseAddress = new Uri(builder.Configuration["Registry:Url"]!));builder.Services.AddSingleton<IServiceResolver, CachingResolver>();builder.Services.AddTransient<ClientSideDiscoveryHandler>();
builder.Services.AddHttpClient<IPolicyClient, PolicyClient>(c => c.BaseAddress = new Uri("http://policy-service")) // logical name, not a real host .AddHttpMessageHandler<ClientSideDiscoveryHandler>();
var app = builder.Build();app.MapGet("/claims/{id:guid}", async (Guid id, IPolicyClient policies, CancellationToken ct) => Results.Ok(await policies.GetCoverageAsync(id, ct)));app.Run();Ready-made option. Microsoft ships Microsoft.Extensions.ServiceDiscovery (10.x packages line up with .NET 10), used by .NET Aspire. It provides configuration-based and DNS-based endpoint resolution and plugs into HttpClient via AddServiceDiscovery(). Use it before writing your own; the hand-written version above is for understanding and for registries it has no provider for. Check the package documentation for the load-balancing selector options rather than assuming the default is round-robin.
Databases also have client-side discovery. Npgsql (PostgreSQL) accepts several hosts and picks among them:
Host=pg-a.claims.internal,pg-b.claims.internal;Database=claims;Username=app;Target Session Attributes=prefer-standby;Load Balance Hosts=trueFor SQL Server availability groups, the client connects to the AG listener name with MultiSubnetFailover=True in the connection string; that is server-side discovery with a client hint.
Level 3: Advanced
Performance and scalability
- Cache registry results (5-30 s) and refresh in the background; never call the registry on every request or you recreate the bottleneck you removed.
- Use a shared
SocketsHttpHandlerwithPooledConnectionLifetime(for example 2-5 minutes) so connections to removed instances are eventually dropped.HttpClientcreated throughIHttpClientFactoryalready rotates handlers (default lifetime 2 minutes). - For HTTP/2 or gRPC, balance per request across multiple connections (one channel per instance). In
Grpc.Net.Client, the built-indns:///resolver with a load-balancing policy does this for headless Kubernetes Services. - Prefer “power of two random choices” over plain round-robin when request costs vary widely (a fraud-scoring call can be 100x heavier than a policy lookup).
Security
- Registry access must be authenticated and authorised; an attacker who can register a fake
policy-servicereceives claim data. Use ACLs/tokens for registration and mTLS between services. - TLS with IP addresses: certificates are issued for names. If you rewrite the host to an IP, hostname validation fails. Either keep the logical name in the TLS SNI/host header via a custom
SslClientAuthenticationOptions.TargetHost, issue certificates with matching SANs, or terminate mTLS in a mesh. - Never trust registry metadata as an authorisation decision.
Failure modes
- Stale cache sends traffic to dead instances: combine short TTL, passive ejection (as above) and retries (Day 28).
- Registry outage: keep serving the last known list; do not fail all calls because discovery is down.
- Thundering herd on registry after a restart of many clients: add jitter to refresh timers.
- Self-preservation in Eureka: when many instances miss heartbeats, Eureka stops evicting, so the list can contain dead nodes. Clients must tolerate that.
- Retry storms: retrying against the same bad instance. Retries must re-resolve or exclude the failed instance.
Common mistakes
- Resolving DNS once at startup and caching forever (
static HttpClientwith noPooledConnectionLifetime). - Writing the discovery code in each service instead of a shared chassis (Day 40).
- Health checks that only test “process is up”, so the registry advertises instances whose database is down (see Day 36).
- Using client-side discovery for browser or partner clients.
- No metric for “endpoints known per service”, so nobody notices when the list shrinks to one.
Level 4: Expert and Architect view
Alternatives compared
| Aspect | Client-Side Discovery | Server-Side Discovery (Day 23) | Service Mesh (Day 48) | Platform DNS (K8s ClusterIP / Container Apps name) |
|---|---|---|---|---|
| Extra network hop | None | One (router/LB) | Local sidecar hop (loopback) | Depends on kube-proxy/ingress; usually transparent |
| Load-balancing intelligence | High (custom, zone-aware) | Medium (LB features) | High (Envoy policies) | Low to medium |
| Client complexity | High (library per language) | Low | Low (no app code) | Very low |
| Operational components | Registry + client library | Registry + LB/router | Control plane + sidecars | Platform only |
| Polyglot friendliness | Poor | Excellent | Excellent | Excellent |
| Failure blast radius | Bug in client library affects that service | LB failure affects everyone | Control plane issues affect config, data plane keeps running | Platform outage |
| Best for | Latency-critical, homogeneous stacks | Mixed stacks, simple clients | Large estates with security/telemetry needs | Most teams on AKS/Container Apps |
Patterns it combines with: Service Registry (Day 21) as the data source; Self-Registration (Day 24) or 3rd Party Registration (Day 25) to fill the registry; Circuit Breaker and Retry (Days 26 and 28) so discovery failures and bad instances are handled; Health Check API (Day 36) as the source of “healthy”; Microservice Chassis (Day 40) to ship the client once.
ADR-style justification
- Title: ADR-022 - Use client-side discovery for internal calls from ClaimsApi to PolicyService and FraudScoringService.
- Context: All internal services are .NET 10, deployed on AKS. p95 latency of the claim-submission flow is 480 ms against a 300 ms target; traces show an internal load balancer adds 2-4 ms per hop across 5 hops, and gRPC traffic to FraudScoringService is unevenly distributed across pods.
- Decision: Adopt
Microsoft.Extensions.ServiceDiscoveryin the shared chassis for HTTP calls, and headless-Service DNS with client-side round-robin for gRPC. The Angular SPA continues to call only the API Gateway. - Consequences: (+) one fewer hop, even gRPC distribution, zone-aware routing possible. (-) discovery behaviour is now part of the chassis and must be versioned and tested; non-.NET services (for example a Python OCR service) still use ClusterIP names; revisit if the estate becomes polyglot (then evaluate a service mesh).
- Rejected: internal LB per service (extra hop and cost); full mesh now (operational cost not yet justified).
Azure implementation
Services that implement or support it
- Azure Kubernetes Service (AKS) with CoreDNS: a normal
ClusterIPService gives a stable virtual IP (server-side style, handled by kube-proxy). A headless Service (clusterIP: None) makes DNS return the individual pod IPs, which is the raw material for client-side discovery and per-request gRPC balancing. - Azure Container Apps: apps in the same environment reach each other by app name (for example
http://policy-service), and the platform load-balances. Use this instead of building your own layer unless you need custom routing. - Eureka on Azure: Azure Spring Apps has an announced retirement (see Microsoft’s retirement announcement), and Microsoft documents a managed Eureka Server for Spring as a Java component in Azure Container Apps. Only relevant if you run Spring services next to .NET ones.
- Consul (HashiCorp) self-hosted on AKS or VMs is the common registry when you want true client-side discovery with health checks.
- Azure Service Fabric has a built-in naming service used by its reverse proxy and client libraries (relevant only if you already run Service Fabric).
- Azure App Configuration can hold endpoint lists for the configuration-based provider in small systems.
- Azure Monitor / Application Insights for the metrics and traces that prove it works.
How to configure (AKS, headless Service for fraud-scoring)
apiVersion: v1kind: Servicemetadata: name: fraud-scoring namespace: claimsspec: clusterIP: None # headless: DNS returns pod IPs selector: app: fraud-scoring ports: - name: grpc port: 8080// gRPC client: resolve all pod IPs via DNS and balance per requestservices.AddGrpcClient<FraudScoring.FraudScoringClient>(o => o.Address = new Uri("dns:///fraud-scoring.claims.svc.cluster.local:8080")) .ConfigureChannel(o => { o.Credentials = Grpc.Core.ChannelCredentials.Insecure; // demo only; use mTLS in production o.ServiceConfig = new Grpc.Net.Client.Configuration.ServiceConfig { LoadBalancingConfigs = { new Grpc.Net.Client.Configuration.RoundRobinConfig() } }; });Note that the gRPC DNS resolver re-resolves on a refresh interval and when connections fail; tune it (DnsResolverFactory refresh interval) so scale-outs are noticed quickly. Also set ChannelCredentials.Insecure only when the transport is protected another way.
Pricing and tier considerations (check the Azure pricing pages before committing; figures change)
- Client-side discovery itself costs nothing on Azure: it uses DNS you already have, or a registry you host.
- AKS: control-plane tiers (Free, Standard, Premium) plus node VM costs; a self-hosted Consul cluster (3-5 small nodes) adds compute and operations cost.
- Container Apps: consumption vs dedicated workload profiles; internal name-based calls have no extra discovery charge.
- The saving you are aiming for is the internal load balancer plus its data-processing charges.
Reference architecture (text) Angular SPA -> Azure Front Door/Application Gateway -> API Gateway (Day 19) -> ClaimsApi pods on AKS. ClaimsApi uses the shared chassis with service discovery: HTTP calls to policy-service resolve through a registry (Consul on AKS) or the platform provider; gRPC calls to fraud-scoring resolve through a headless Service. Services register via Kubernetes (3rd Party Registration, Day 25) or Consul agents. Each service writes to its own database (Azure SQL or Azure Database for PostgreSQL, Day 5). Telemetry flows through OpenTelemetry to Application Insights, with an “endpoints per service” metric alerting when a service has fewer than two known endpoints. A network policy restricts who can query and register in the registry.
Teaching guide for my team
2-minute beginner explanation. “Services are like shops that keep moving. A registry is the live phone book. With client-side discovery, you open the phone book, pick a shop from the current list, and walk there yourself. That saves the trip through a receptionist, but you must know how to read the phone book and what to do if the shop is closed.”
5-minute intermediate explanation. Draw ClaimsApi -> [registry lookup, cached 15 s] -> PolicyService instance B. Cover: logical names in code, a delegating handler that swaps the host, local caching and background refresh, passive ejection of failed instances, and why the SPA does not use this. Contrast with Day 23: same registry, but the router does the picking. Mention the trade-off: you now own client code in every language.
Hands-on exercise. Run three copies of a tiny PolicyService (ports 5001-5003) that return their port number. Implement CachingResolver with an in-memory IRegistryClient. (1) Call /claims/{id} 12 times and confirm rotation over 5001, 5002, 5003. (2) Stop the instance on 5002 while calling in a loop. Expected outcome: at most one or two failed calls, then the ejected instance is skipped for 30 seconds and the remaining calls succeed. (3) Restart 5002, wait past the ejection time, confirm it rejoins the rotation.
Interview-style questions
- What is the difference between client-side and server-side discovery? In client-side, the caller queries the registry and chooses the instance; in server-side, the caller talks to a router/LB that queries the registry. Client-side saves a hop but puts logic in every client.
- Why is a plain L4 load balancer poor for gRPC? gRPC uses long-lived HTTP/2 connections, so the LB balances connections, not requests, leaving some instances overloaded; client-side balancing (or an L7 proxy) balances per request.
- What happens if the registry goes down? Clients keep using their cached instance list (stale but usable) and passively eject failed instances; they must not fail every call just because discovery is unavailable.
Mastery checklist
- I can draw the request flow for client-side and server-side discovery and state the trade-off in one sentence.
- I can implement a caching resolver with TTL, jitter and a stale-on-error fallback.
- I can explain why an unhandled DNS or endpoint cache causes errors during deployments.
- I can explain TLS hostname validation problems when connecting by IP and give two fixes.
- I can configure a headless Service and a gRPC channel that balances per request on AKS.
- I can decide between client-side discovery, platform DNS, and a service mesh for a given estate and write the ADR.
- I know which callers (Angular, partners) must never use client-side discovery.
- I can name the metrics that prove discovery is healthy (endpoints per service, resolve latency, ejection count).
Key takeaway
Client-side discovery trades an extra network hop for extra client responsibility: cache the registry list, balance per request, eject bad instances, and centralise all of it in one shared library. If your platform already gives you a load-balanced service name, use that instead.
