Server-side discovery means the caller does not look up service instances at all.
Intro
Server-side discovery means the caller does not look up service instances at all. It sends every request to one stable address (a router, load balancer, or ingress), and that component asks the service registry which instances are healthy and forwards the request to one of them. The caller stays “dumb” about the network, so every language and framework in the company gets discovery for free. In our insurance claims system, the Angular portal and the Policy service both call claims.internal and never know how many Claims API instances exist or where they run.
Why we need this
In a microservices platform, instances are ephemeral. Kubernetes reschedules pods, Azure Container Apps scales replicas from 1 to 30 and back, and deployments replace every instance. IP addresses and ports change constantly.
Client-side discovery (Day 22) solves this by putting a registry client and a load-balancing library inside every service. That works well when the whole company uses one language and one library. It breaks down when the estate is polyglot (our claims system has .NET services, a Python fraud-scoring model, and a legacy Java policy admin system) because every language needs its own discovery client, and every client must be upgraded when the registry protocol changes.
Server-side discovery moves that logic into infrastructure the platform team owns. The business reasons are:
- Application teams ship business features instead of networking code.
- One place enforces load balancing, health-based routing, TLS, and retries.
- Onboarding a new language costs nothing, because a plain HTTP call is all that is required.
What problem it solves
Problem from the topic list: every language needs custom discovery client code.
Without it, each Claims, Policy, Payment, and Fraud service embeds a registry SDK (Consul, Eureka, etc.), a caching layer, and a load-balancing algorithm. Symptoms:
- The Python fraud service has a home-grown discovery client with a bug: it caches instance lists forever and keeps calling a dead node for minutes.
- The .NET services upgrade the discovery package, but the Java service is two versions behind, so behaviour differs per service.
- Adding a new capability (for example, “prefer instances in the same availability zone”) means changing and redeploying every service.
- Developers put IPs into
appsettings.json“temporarily”, and those become permanent.
With server-side discovery the client code is just http://claims-api/.... The router handles the rest.
When it is needed (and when it is NOT)
Use it when:
- The organisation is polyglot, or many teams with different skill levels build services.
- You run on a platform that already provides it (Kubernetes Services, Azure Container Apps ingress, Azure Load Balancer, Application Gateway). Then you are using it whether you realised it or not.
- You want central control over routing, TLS, rate limits, and observability.
- Clients are external or untrusted (browsers, partners), and must never see internal topology.
Avoid or reconsider it when:
- Every hop must be as fast as possible and you cannot afford the extra network hop (for example, very chatty internal gRPC with microsecond budgets). Client-side discovery or a sidecar (Day 47/48) avoids the hop.
- You need client-specific routing logic such as consistent hashing on a claim ID. A generic load balancer cannot do that unless it supports it.
- You have three services and one team. A static DNS name behind a simple load balancer is enough, and a full registry is overkill.
- The router becomes a single point of failure because it is deployed as one instance. It must itself be redundant.
How to identify the problem (key signals)
- Multiple discovery client libraries or in-house HTTP wrappers exist across repos (search for
Consul,Eureka,ServiceLocatorin every language). - The same outage is handled differently by different services because each has its own retry and failover logic.
- IP addresses, ports, or environment-specific hostnames are hard-coded in config files or pipeline variables.
- Onboarding a new language requires a “port the discovery client” ticket.
- After a scale-in or deploy, callers keep getting connection-refused or timeouts for 30-120 seconds because their cached instance list is stale.
- Registry traffic is unexpectedly high because hundreds of clients poll it directly.
- Security review flags that browsers or partners can see internal service addresses.
Flow Diagram
The caller hits one stable address; the router consults the registry and forwards.
flowchart LR CL["Any client or service"] -- "http://claims-api" --> R["Router / LB / K8s Service"] R -- "query healthy instances" --> REG[("Registry or EndpointSlices")] R --> I1["claims-api replica 1"] R --> I2["claims-api replica 2"] R --> I3["claims-api replica 3"] HC["Readiness probe /health/ready"] --> REGLevel 1: Beginner
Analogy: a hospital reception desk. Patients (clients) do not wander the wards looking for a free doctor. They tell reception what they need; reception checks the duty roster (registry) and sends them to an available doctor. The roster changes all day, but patients always walk to the same desk.
Minimal working example: a YARP reverse proxy in front of two Claims API instances, configured statically. Here the “registry” is the config file. This is the smallest possible server-side discovery.
// Program.cs (.NET 10, package: Yarp.ReverseProxy)var builder = WebApplication.CreateBuilder(args);
builder.Services .AddReverseProxy() .LoadFromConfig(builder.Configuration.GetSection("ReverseProxy"));
var app = builder.Build();app.MapReverseProxy();app.Run();{ "ReverseProxy": { "Routes": { "claims-route": { "ClusterId": "claims-cluster", "Match": { "Path": "/claims/{**catch-all}" } } }, "Clusters": { "claims-cluster": { "LoadBalancingPolicy": "RoundRobin", "Destinations": { "claims-1": { "Address": "http://10.0.1.11:8080/" }, "claims-2": { "Address": "http://10.0.1.12:8080/" } } } } }}A client calls GET https://router/claims/CLM-1001. It never learns 10.0.1.11 or 10.0.1.12. The limitation: someone edits the file by hand when instances change. Levels 2 and 3 remove that.
Level 2: Intermediate
In a real .NET + Angular + database application, three pieces cooperate.
- Claims API exposes health endpoints so the router knows which instances can receive traffic.
- The router gets its destination list from a registry that is updated dynamically.
- Angular calls one stable base URL.
Claims API with readiness and liveness checks (SQL Server dependency in the readiness check):
// Program.cs - Claims.Api (.NET 10)// Packages: AspNetCore.HealthChecks.SqlServer (or Microsoft.Extensions.Diagnostics.HealthChecks)using Microsoft.AspNetCore.Diagnostics.HealthChecks;using Microsoft.Extensions.Diagnostics.HealthChecks;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddHealthChecks() .AddSqlServer( builder.Configuration.GetConnectionString("ClaimsDb")!, name: "claims-db", tags: ["ready"]);
var app = builder.Build();
// Liveness: process is up. No dependencies checked.app.MapHealthChecks("/health/live", new HealthCheckOptions { Predicate = _ => false });
// Readiness: only instances that can reach the database receive traffic.app.MapHealthChecks("/health/ready", new HealthCheckOptions{ Predicate = check => check.Tags.Contains("ready")});
app.MapGet("/claims/{id}", (string id) => Results.Ok(new { id, status = "Open" }));app.Run();A YARP router whose destinations come from a registry. IServiceRegistry is our abstraction over Consul, Kubernetes endpoints, or Azure service metadata. Its implementation is not shown because it depends on your registry.
// Router/Program.cs (.NET 10, package: Yarp.ReverseProxy)using Yarp.ReverseProxy.Configuration;
var routes = new[]{ new RouteConfig { RouteId = "claims-route", ClusterId = "claims-cluster", Match = new RouteMatch { Path = "/claims/{**catch-all}" } }};
var provider = new InMemoryConfigProvider(routes, Array.Empty<ClusterConfig>());
var builder = WebApplication.CreateBuilder(args);builder.Services.AddSingleton(provider);builder.Services.AddSingleton<IProxyConfigProvider>(provider);builder.Services.AddReverseProxy();builder.Services.AddSingleton<IServiceRegistry, MyRegistry>(); // your implementationbuilder.Services.AddHostedService<RegistrySyncService>();
var app = builder.Build();app.MapReverseProxy();app.Run();
public interface IServiceRegistry{ Task<IReadOnlyList<string>> GetHealthyAddressesAsync(string service, CancellationToken ct);}
public sealed class RegistrySyncService( IServiceRegistry registry, InMemoryConfigProvider provider, ILogger<RegistrySyncService> log) : BackgroundService{ private static readonly RouteConfig[] Routes = [ new RouteConfig { RouteId = "claims-route", ClusterId = "claims-cluster", Match = new RouteMatch { Path = "/claims/{**catch-all}" } } ];
protected override async Task ExecuteAsync(CancellationToken ct) { using var timer = new PeriodicTimer(TimeSpan.FromSeconds(5)); do { try { var addresses = await registry.GetHealthyAddressesAsync("claims-api", ct); var destinations = addresses .Select((addr, i) => (Key: $"claims-{i}", Address: addr)) .ToDictionary(x => x.Key, x => new DestinationConfig { Address = x.Address });
var cluster = new ClusterConfig { ClusterId = "claims-cluster", LoadBalancingPolicy = "PowerOfTwoChoices", Destinations = destinations, HealthCheck = new HealthCheckConfig { Active = new ActiveHealthCheckConfig { Enabled = true, Interval = TimeSpan.FromSeconds(10), Timeout = TimeSpan.FromSeconds(3), Policy = "ConsecutiveFailures", Path = "/health/ready" } } };
provider.Update(Routes, [cluster]); } catch (Exception ex) when (ex is not OperationCanceledException) { // Keep the last known-good config; never wipe routes on a registry blip. log.LogWarning(ex, "Registry sync failed; keeping previous destinations"); } } while (await timer.WaitForNextTickAsync(ct)); }}Angular (standalone, functional interceptor). The important point is what is missing: no discovery code.
// api-base-url.interceptor.ts (Angular 21/22, also valid from v17)import { HttpInterceptorFn } from '@angular/common/http';import { environment } from '../environments/environment';
export const apiBaseUrlInterceptor: HttpInterceptorFn = (req, next) => req.url.startsWith('/api/') ? next(req.clone({ url: environment.apiBaseUrl + req.url.substring(4) })) : next(req);
// app.config.tsimport { ApplicationConfig } from '@angular/core';import { provideHttpClient, withInterceptors } from '@angular/common/http';
export const appConfig: ApplicationConfig = { providers: [provideHttpClient(withInterceptors([apiBaseUrlInterceptor]))]};
// claims.service.tsimport { Injectable, inject } from '@angular/core';import { HttpClient } from '@angular/common/http';
@Injectable({ providedIn: 'root' })export class ClaimsService { private http = inject(HttpClient); getClaim(id: string) { return this.http.get<{ id: string; status: string }>(`/api/claims/${id}`); }}environment.apiBaseUrl is one value such as https://api.contoso-claims.com. Whether there are 2 or 40 Claims API instances behind it is invisible to Angular.
On Kubernetes, you already get server-side discovery without writing a router. The Service object is the stable address, and the cluster resolves it to ready pods:
apiVersion: v1kind: Servicemetadata: name: claims-apispec: selector: app: claims-api ports: - port: 80 targetPort: 8080Any pod in the namespace calls http://claims-api/claims/CLM-1001. Kubernetes DNS resolves the name to the Service’s virtual IP, and kube-proxy (or the CNI’s dataplane) picks a ready pod. A pod is only in the endpoint list when its readiness probe passes.
Level 3: Advanced
Performance
- Every request adds one hop. Keep routers close to callers (same region, same zone) and use HTTP/2 or keep-alive connections to the backends so you do not pay a TCP+TLS handshake per request.
- Prefer
PowerOfTwoChoicesorLeastRequestsoverRoundRobinwhen request cost varies widely (for example, a claim PDF export versus a status lookup). Round robin ignores load. - The router should cache the destination list in memory and update it asynchronously (as in Level 2). Never query the registry per request.
Scalability
- Run at least 2-3 router instances across availability zones behind a platform load balancer. Scale them on requests per second and CPU.
- Watch connection limits: L4 load balancers, NAT gateways, and SNAT ports can be exhausted long before CPU is.
Security
- Terminate TLS at the router and re-encrypt to backends (or use mTLS) so traffic inside the network is not plaintext.
- Keep the router’s registry credentials read-only.
- Do not expose the registry’s admin API to application subnets.
- Strip or overwrite
X-Forwarded-*and identity headers from external callers, then add your own.
Failure modes
- Stale registry: instances died but are still listed. Mitigate with short TTLs, active health checks (
/health/ready), and passive health checks on failed requests. - Registry outage: the router must keep serving the last known-good list. Serving from cache during a registry outage is better than failing all traffic.
- Router as single point of failure: solve with redundancy, not hope.
- Retry storms: if the router retries and the caller also retries, one failure multiplies. Decide which layer retries (see Day 28), and only retry idempotent requests.
- Deregistration race during deployments: instance is removed from the registry but still receiving in-flight requests. Use graceful shutdown: fail readiness first, wait for a drain period, then stop.
Common mistakes
- Using the liveness endpoint for routing. A slow database makes the instance “not ready”, but it is still alive, so restarting it just adds load. Route on readiness, restart on liveness.
- Putting business logic (authorisation rules, data mapping) in the router. It turns the router into a hidden monolith.
- Routing by hard-coded IP in “temporary” config.
- Forgetting that
Servicevirtual IPs in Kubernetes balance per connection, not per request. Long-lived gRPC or HTTP/2 connections can pin to one pod, so gRPC needs an L7 proxy or headless-service client balancing.
Level 4: Expert and Architect view
| Aspect | Server-side discovery | Client-side discovery (Day 22) | Service mesh sidecar (Day 48) | Static DNS / config |
|---|---|---|---|---|
| Discovery logic lives in | Router or load balancer | Every client process | Sidecar proxy next to each client | Nowhere (manual) |
| Extra network hop | Yes (one) | No | Yes, but local (loopback) | Depends |
| Language independence | Excellent | Poor (library per language) | Excellent | Excellent |
| Load balancing intelligence | Central, uniform | Per client, can be client-specific | Per client proxy, rich | None |
| Failure blast radius | Router outage affects all callers | One client’s bug affects only it | Sidecar issue affects one pod | Stale config |
| Operational complexity | Medium (run routers) | Medium (upgrade libraries) | High (mesh control plane) | Low, but does not scale |
| Best fit | Polyglot estates, external traffic, platform-owned networking | Homogeneous stack, latency-critical | Large estates needing mTLS and telemetry everywhere | Very small systems |
Patterns it combines with
- Service Registry (Day 21): the source of truth the router reads.
- Third-party registration (Day 25): an orchestrator registers instances, so services carry no registration code. Kubernetes plus server-side discovery is the standard combination.
- API Gateway (Day 19) and BFF (Day 20): the edge router is often the same product, but discovery for east-west traffic is a separate concern from edge concerns like auth and rate limiting.
- Health Check API (Day 36): supplies the readiness signal.
- Circuit Breaker, Retry, Bulkhead (Days 26-28): applied at the router or in the callers, but decide deliberately where.
ADR-style justification
- Title: ADR-023 Use server-side discovery for internal service-to-service calls.
- Context: The claims platform has .NET, Python, and Java services owned by six teams. Instances scale and redeploy several times a day. Client-side discovery libraries diverge per language and caused two incidents through stale instance caches.
- Decision: All internal calls address a stable logical name. The platform (Container Apps ingress or Kubernetes Services plus an ingress/gateway) resolves it to ready instances. Application code contains no registry client.
- Consequences (positive): language-neutral, one place to tune load balancing and TLS, faster onboarding, smaller service code.
- Consequences (negative): extra hop (about a millisecond within a zone), routers must be highly available and monitored, less client-specific routing control.
- Alternatives considered: client-side discovery (rejected: polyglot cost), full service mesh (deferred: operational cost exceeds current need; revisit when mTLS-everywhere becomes a requirement).
- Status: Accepted.
Azure implementation
Azure services that implement or support it
- Azure Container Apps (ACA): built-in ingress backed by Envoy. Each app gets an FQDN, and apps in the same environment can call each other by app name. Traffic is load-balanced across replicas, and you can split traffic between revisions by percentage for canary releases.
- Azure Kubernetes Service (AKS): Kubernetes
Serviceand CoreDNS for east-west discovery; an ingress controller or Gateway API implementation for north-south. Application Gateway for Containers is Azure’s managed layer 7 option for AKS, configured through Kubernetes Ingress or Gateway API resources, with the data plane running outside the cluster. - Azure Load Balancer (layer 4) and Application Gateway v2 (layer 7, WAF) for VM-based or App Service backends using backend pools.
- Azure Front Door for global edge routing to regional backends.
- Azure API Management when you want a managed gateway in front of internal APIs.
How to configure
Container Apps, internal ingress so only apps in the environment can call Claims API:
az containerapp create \ --name claims-api \ --resource-group rg-claims \ --environment cae-claims \ --image myregistry.azurecr.io/claims-api:1.0.0 \ --ingress internal \ --target-port 8080 \ --min-replicas 2 \ --max-replicas 10 \ --scale-rule-name http-rule \ --scale-rule-type http \ --scale-rule-http-concurrency 50Callers in the same environment use http://claims-api. Add a readiness probe in the container app’s probe settings pointing to /health/ready so unready replicas are not sent traffic. Keep min-replicas at 2 or more in production to avoid scale-to-zero cold starts on the claims path.
AKS: define the Service as in Level 2, add readiness probes to the Deployment, and expose external traffic through an Ingress or Gateway API HTTPRoute handled by your chosen controller.
Pricing and tier considerations
- Container Apps has a Consumption plan and a Dedicated plan (workload profiles). Consumption bills on vCPU-seconds, GiB-seconds, and HTTP requests, with a monthly free grant per subscription of 180,000 vCPU-seconds, 360,000 GiB-seconds, and 2 million requests. The ingress itself has no separate charge. Dedicated workload profiles add a plan management fee, as do features like private endpoints and planned maintenance.
- AKS: the control plane has a Free tier (no SLA, for dev/test), a Standard tier with a financially backed uptime SLA, and a Premium tier with long-term support. You pay for node VMs. Pod-to-pod discovery through Kubernetes Services costs nothing extra.
- Application Gateway and Application Gateway for Containers are billed on capacity and usage; check the current rates on the Application Gateway pricing page before sizing, since they vary by region and SKU.
- Cost tip: a router tier that is too small becomes the bottleneck, and one that is too big is wasted spend. Load-test to find requests per second per unit, then set autoscale bounds.
Reference architecture (text)
Internet clients (Angular portal, partner APIs) reach Azure Front Door (WAF, TLS, global routing). Front Door forwards to API Management or Application Gateway in the hub VNet. From there, requests go to the Container Apps environment (or AKS cluster) in a spoke VNet. Inside it, Claims, Policy, Payment, and Fraud services have internal ingress or ClusterIP Services. Services call each other using logical names only; the platform ingress or kube-proxy resolves them to ready replicas. Each service uses Azure SQL Database or PostgreSQL Flexible Server through private endpoints, with credentials in Key Vault via managed identity. Azure Monitor with Application Insights collects request traces, router metrics, and health probe results. Alerts fire on 5xx rate at the router, healthy-instance count dropping below the minimum, and readiness probe failures.
Teaching guide for my team
2-minute beginner explanation
“When the claims service runs on five servers that keep changing, the Angular app cannot remember five addresses. So we give it one address, the router. The router keeps a live list of which servers are healthy and passes each request to one of them. If a server dies, the router just stops using it. The app never notices.”
5-minute intermediate explanation
Cover the difference from Day 22: in client-side discovery each caller queries the registry and picks an instance; in server-side discovery the caller sends to a router that queries the registry. Show the Kubernetes Service YAML and say “this is server-side discovery, you already use it”. Then explain the ingredients: a registry (who is alive), health checks (who is ready), a load-balancing policy (who is next), and a router (who forwards). Finish with trade-offs: one extra hop, router must be redundant, but no discovery code in any language.
Hands-on exercise
Task: run two copies of the Claims API locally on ports 5001 and 5002, each returning its own port in the response. Put the YARP router from Level 1 in front on port 5000 with RoundRobin. Call GET /claims/CLM-1001 ten times, then stop one instance and call it ten more times.
Expected outcome: the first ten responses alternate between 5001 and 5002. After you stop one instance, some requests fail until you enable active health checks in the cluster config (/health/ready); after enabling them, traffic goes to the surviving instance only. Learners should be able to explain why the health check fixed it.
Interview-style questions
- What is the difference between client-side and server-side discovery? Answer: In client-side, the caller queries the registry and load-balances itself. In server-side, the caller sends to a router or load balancer that queries the registry and forwards. Server-side keeps discovery logic out of every service.
- Is a Kubernetes Service server-side discovery? Answer: Yes. The Service gives a stable name and virtual IP, and the cluster routes to ready pods from its endpoint list.
- What is the main risk of server-side discovery and how do you reduce it? Answer: The router is a single point of failure and adds a hop. Run several instances across zones, keep the last known-good config if the registry is down, and monitor healthy-instance counts.
Mastery checklist
- I can explain server-side discovery in two minutes without jargon and draw the request flow.
- I can name three places it already exists in our stack (Kubernetes Service, Container Apps ingress, Application Gateway).
- I can configure a router with dynamic destinations and active health checks.
- I can explain why routing uses the readiness endpoint and restarts use liveness.
- I can describe what happens during a deployment (drain, readiness failure, deregistration) and how to avoid dropped requests.
- I can compare it with client-side discovery and a service mesh and pick one for a given scenario.
- I can identify the router’s failure modes and design its redundancy.
- I can write a short ADR justifying the choice.
Key takeaway
Give callers one stable address and let infrastructure find the healthy instances behind it. You trade one extra hop and a router you must keep highly available for discovery code that no service in any language has to write.
