A Service Registry is a live "phone book" for your services.
Intro
A Service Registry is a live “phone book” for your services. Every running instance of a service (for example, three copies of policy-service) announces where it is (IP and port) and whether it is healthy. Any caller asks the phone book “where is policy-service right now?” instead of using a hardcoded address. Because containers start, stop, scale and move all the time, the phone book is updated automatically and never relies on a human editing a config file.
Why we need this
In a monolith, PolicyService is a class: you call it in-process and it always “exists”. In microservices, PolicyService is a process on a network, and its address changes constantly:
- Autoscaling adds and removes instances several times a day (a storm produces a spike of claims at 09:00).
- Deployments replace instances (rolling update: new pod IPs every release).
- Failures kill instances; the platform restarts them elsewhere with a new IP.
- Environments differ: dev is
localhost:5001, test is a container, prod is a cluster.
Business reasons: teams must be able to deploy and scale their service without asking every consumer to change config. Technical reasons: hardcoded http://10.0.1.11:8080 values are wrong within minutes in any elastic platform, and a load balancer in front of everything becomes a manual-config bottleneck.
What problem it solves
Problem (from the topic list): services run on ephemeral IPs and ports that cannot be hardcoded.
Without a registry in our insurance claims system:
ClaimsApihasPolicyServiceUrl = http://10.0.1.11:8080inappsettings.json.- Ops scales
policy-servicefrom 2 to 5 pods during a storm.ClaimsApistill calls only the original one. The new pods sit idle while the old one is overloaded. - A rolling deployment replaces the pod at
10.0.1.11.ClaimsApinow getsconnection refuseduntil someone edits config and restarts it. - Adjusters see “Cannot verify coverage” errors, and every fix needs a release of the caller, even though the callee is what changed.
With a registry: instances register themselves (or the platform registers them), unhealthy ones drop out, and callers always resolve to a current list of healthy instances.
When it is needed (and when it is NOT)
Needed when:
- Instances are dynamic: autoscaling, containers, rolling updates, multiple environments.
- You have many services calling each other (roughly more than a handful) and manual address management is already painful.
- You run on VMs or a custom platform that does not give you discovery for free.
- You need rich metadata per instance (version, zone, weight, tenant) for routing decisions.
NOT needed (or overkill) when:
- You run on Kubernetes, Azure Container Apps or Service Fabric: the platform already is the registry (Kubernetes Services + CoreDNS, ACA built-in DNS). Do not add Consul or Eureka on top “because the book said so”.
- You have 2-3 services with stable addresses behind a load balancer or App Service hostnames. A DNS name is enough.
- You are still a modular monolith. Discovery solves a distribution problem you do not have yet.
- Databases, Service Bus, Key Vault: these have stable managed DNS names and should be configured, not “discovered”.
How to identify the problem (key signals)
- Service URLs in config files that change with every environment and every release (
appsettings.Production.jsonfull of IPs). - “Connection refused” / “No such host” spikes right after deployments or scale events, which disappear after callers are restarted.
- Uneven load: one instance at 95% CPU while newly scaled instances sit at 2%, because callers pinned to the old address.
- Deployment checklists that include “update the URL in service X and Y” steps.
- Calls routed to dead instances: timeouts of exactly your HTTP timeout value (e.g. 100 s default in
HttpClient) for a fraction of requests, matching the number of dead instances. - Stale DNS behaviour: a caller keeps hitting an old IP for minutes after failover because of OS/JVM/.NET connection or DNS caching.
- A shared “endpoints” wiki page or spreadsheet that people update by hand.
Flow Diagram
Instances register, renew their lease, and are looked up by logical name.
flowchart LR I1["policy-service instance A"] -- "register + heartbeat" --> REG[("Service Registry")] I2["policy-service instance B"] -- "register + heartbeat" --> REG I3["crashed instance C"] -. "lease expired" .-> REG CL["ClaimsApi"] -- "lookup policy-service" --> REG REG -- "healthy: A, B" --> CL CL -- "call" --> I1Level 1: Beginner
Analogy: a hotel reception desk. Guests (instances) check in when they arrive and check out when they leave. If a guest stops answering the room phone (no heartbeat), reception assumes they left. Visitors (callers) never memorise room numbers; they ask reception “where is Mr. Policy?” every time.
The registry has four operations:
| Operation | Who calls it | Meaning |
|---|---|---|
| Register | Instance (or platform) on start | “I am policy-service at 10.0.1.11:8080” |
| Heartbeat / renew | Instance every few seconds | “I am still alive” (a lease renewal) |
| Deregister | Instance on graceful shutdown | “I am leaving” |
| Lookup | Caller | “Give me healthy instances of policy-service” |
Minimal working example (.NET 10, one file, pure in-memory, to show the idea; not for production):
using System.Collections.Concurrent;
var registry = new ServiceRegistry(ttl: TimeSpan.FromSeconds(1));
registry.Register("policy-service", "10.0.1.11:8080");registry.Register("policy-service", "10.0.1.12:8080");
Thread.Sleep(700);registry.Heartbeat("policy-service", "10.0.1.11:8080"); // only instance 11 renews its leaseThread.Sleep(500);
// instance 12 last renewed 1.2 s ago (> 1 s TTL), so it is treated as deadforeach (var address in registry.Lookup("policy-service")){ Console.WriteLine(address); // prints only 10.0.1.11:8080}
public sealed class ServiceRegistry(TimeSpan ttl){ private readonly ConcurrentDictionary<(string Service, string Address), DateTimeOffset> _leases = new();
public void Register(string service, string address) => _leases[(service, address)] = DateTimeOffset.UtcNow;
public void Heartbeat(string service, string address) => Register(service, address);
public void Deregister(string service, string address) => _leases.TryRemove((service, address), out _);
public IReadOnlyList<string> Lookup(string service) { var cutoff = DateTimeOffset.UtcNow - ttl; return _leases .Where(l => l.Key.Service == service && l.Value >= cutoff) .Select(l => l.Key.Address) .Order() .ToList(); }}Key beginner lesson: liveness is proven by a recent heartbeat (a lease with a TTL), not by “the instance registered once”.
Level 2: Intermediate
6.1 The two registry styles you will actually meet
- Platform registry (preferred): Kubernetes (Service + CoreDNS + EndpointSlices) or Azure Container Apps (built-in DNS + Envoy). Your code just uses a name like
http://policy-service. - Standalone registry: HashiCorp Consul, Netflix Eureka (mostly Java/Spring), etcd or ZooKeeper-based. Use when you run outside such a platform (VMs, hybrid, multi-cloud).
6.2 .NET client: Microsoft.Extensions.ServiceDiscovery
On .NET 10 (LTS), the Microsoft.Extensions.ServiceDiscovery NuGet package lets you write logical names in HttpClient base addresses and resolves them at call time. Out of the box it supports a configuration provider (a Services section) and a pass-through provider (returns the name as-is so the platform’s DNS resolves it). DNS and DNS SRV providers are available in the separate Microsoft.Extensions.ServiceDiscovery.Dns package. It works with or without .NET Aspire.
ClaimsApi/Program.cs:
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddServiceDiscovery();
// every HttpClient created via IHttpClientFactory resolves logical namesbuilder.Services.ConfigureHttpClientDefaults(http => http.AddServiceDiscovery());
// "https+http://" = try HTTPS endpoints first, fall back to HTTPvar policyBaseUrl = builder.Configuration["PolicyService:BaseUrl"] ?? "https+http://policy-service";builder.Services.AddHttpClient<PolicyClient>(c => c.BaseAddress = new Uri(policyBaseUrl));
var app = builder.Build();
app.MapGet("/claims/{claimId:guid}/coverage", async (Guid claimId, PolicyClient policies, CancellationToken ct) =>{ var coverage = await policies.GetCoverageAsync(claimId, ct); return coverage is null ? Results.NotFound() : Results.Ok(coverage);});
app.Run();
public sealed record CoverageDto(string PolicyNumber, bool Active, decimal Limit);
public sealed class PolicyClient(HttpClient http){ public Task<CoverageDto?> GetCoverageAsync(Guid claimId, CancellationToken ct) => http.GetFromJsonAsync<CoverageDto>($"policies/by-claim/{claimId}/coverage", ct);}Local development uses the configuration provider (appsettings.Development.json):
{ "Services": { "policy-service": { "https": [ "localhost:7101" ] } }}In production on Kubernetes or Azure Container Apps you add no Services section. The pass-through provider hands policy-service to the platform DNS. On Azure Container Apps, set PolicyService__BaseUrl to http://policy-service (short app-name form) or to the internal FQDN.
The same code, three environments, zero address changes in the caller. That is the point.
6.3 Angular
The browser should not do service discovery. It talks to one stable public address (API Gateway or BFF; see Day 19 and 20). Load that address at runtime, so one build works in every environment (Angular 21/22 standalone style):
import { HttpClient } from '@angular/common/http';import { Injectable, inject } from '@angular/core';import { firstValueFrom } from 'rxjs';
@Injectable({ providedIn: 'root' })export class AppConfigService { private readonly http = inject(HttpClient); apiBaseUrl = '';
async load(): Promise<void> { const cfg = await firstValueFrom(this.http.get<{ apiBaseUrl: string }>('/config.json')); this.apiBaseUrl = cfg.apiBaseUrl; }}import { ApplicationConfig, inject, provideAppInitializer } from '@angular/core';import { provideHttpClient } from '@angular/common/http';import { AppConfigService } from './app-config.service';
export const appConfig: ApplicationConfig = { providers: [ provideHttpClient(), provideAppInitializer(() => inject(AppConfigService).load()), ],};/config.json is a static file replaced per environment at deploy time (for example { "apiBaseUrl": "https://api.claims.contoso.com" }).
6.4 SQL Server / PostgreSQL
Not applicable as a registry concern: Azure SQL Database and Azure Database for PostgreSQL have stable DNS names supplied through configuration or Key Vault. Do not register databases in a service registry. The only database link is if you build a custom registry: keep its state in memory or a purpose-built store (Consul/etcd), never in your business database, because lookups happen on every call path and need heartbeat-grade write rates.
Level 3: Advanced
7.1 Health and leases: what “healthy” means
- Lease/TTL model (Eureka-style): the instance must renew, otherwise it expires. Simple, but a paused process (GC pause, CPU starvation) can be evicted while alive.
- Active health checks (Consul-style): the registry (or agent) probes
/health/ready. Detects “process alive but dependency broken”. - Register readiness, not liveness. An instance whose SQL connection is down must not receive traffic, but it should not be killed either. In ASP.NET Core map separate endpoints (
/health/live,/health/ready) viaMapHealthChecks(see Day 36).
7.2 Graceful shutdown ordering (the most common production bug)
Pseudo-code for a self-registering instance:
on SIGTERM: 1. deregister from registry (stop NEW traffic) 2. wait propagation delay (e.g. 5-10 s) (callers refresh their caches) 3. finish in-flight requests (ASP.NET Core graceful shutdown) 4. exitIf you exit first and deregister last, callers keep sending requests to a dead address for the length of their cache TTL. In Kubernetes the equivalent is a preStop sleep plus a terminationGracePeriodSeconds that exceeds it.
7.3 Caching and staleness
- Callers cache lookups; a longer cache reduces registry load but increases stale routing. Typical: refresh 5-30 s.
HttpClientwithSocketsHttpHandlerkeeps connections alive; new endpoints are not used until connections are recycled. SetPooledConnectionLifetime(for example 2-5 minutes) so DNS/registry changes are honoured.IHttpClientFactoryrotates handlers (default 2 minutes) for the same reason.- Design for the registry being down: callers must keep using their last known list (stale-but-available beats unavailable).
7.4 Consistency vs availability
| Registry | Model | Behaviour in a network partition |
|---|---|---|
| Eureka | AP (available, eventually consistent) | Keeps serving possibly stale data; “self-preservation” stops mass eviction when many heartbeats are missed at once |
| Consul | CP for the catalog (Raft), with agent-side health | Can refuse writes without a quorum; reads can be stale if configured |
| Kubernetes (etcd + EndpointSlices) | CP store, eventually consistent propagation to nodes | Node-local view can lag briefly during changes |
Rule of thumb: for discovery, availability beats consistency. A slightly stale list is survivable; a registry outage that stops all calls is not.
7.5 Security
- Restrict who can register: a rogue process registering as
policy-servicecan intercept claim data. Use ACLs/tokens (Consul ACLs) or platform RBAC. - Use mTLS between services; discovery tells you where, it does not prove who.
- Do not expose the registry UI or API publicly.
7.6 Common mistakes
- Registering
localhostor a container-internal IP that other hosts cannot reach. - Registering on start, before the app is ready (traffic arrives to a warming service). Register on ready.
- Ignoring deregistration on shutdown.
- Building a home-grown registry when the platform already provides one.
- Making the registry a single point of failure (one instance, no cluster, no client-side cache).
- Putting registry lookup on the hot path of every request without caching.
- Confusing discovery with load balancing and resilience: you still need timeouts, retries and circuit breakers (Days 26-28).
Level 4: Expert and Architect view
8.1 Design trade-offs
| Option | Where the registry lives | Pros | Cons | Best for |
|---|---|---|---|---|
| Platform DNS (Kubernetes Service, ACA app name) | Platform control plane | Zero code, zero extra infra, language-neutral | Coarse: no rich metadata, little routing logic | Default choice on AKS/ACA |
| Standalone registry (Consul, Eureka) | Separate cluster you operate | Metadata, multi-DC, works off-platform | You must operate and secure it (HA, ACLs, upgrades) | VMs, hybrid, multi-cloud |
| Config file / env vars | App config | Simplest | Manual, stale, no health awareness | Very small, static systems |
| Hardcoded load balancer DNS | Cloud load balancer | Familiar, health probes built in | One LB per service, cost, manual wiring | Small number of services |
| Service mesh (Istio, Linkerd) | Mesh control plane (reads Kubernetes) | mTLS, traffic shifting, telemetry | High complexity | Large fleets (see Day 48) |
8.2 Patterns it combines with
- Client-Side Discovery (Day 22), Server-Side Discovery (Day 23): how the lookup is done.
- Self-Registration (Day 24), 3rd Party Registration (Day 25): how registration is done.
- Health Check API (Day 36): the source of truth for “healthy”.
- Circuit Breaker / Retry (Days 26, 28): protect against the window between “instance died” and “registry knows”.
- API Gateway (Day 19): the gateway is the biggest consumer of the registry.
8.3 ADR (architecture review ready)
ADR-021: Service discovery for the claims platform
- Status: Proposed
- Context:
ClaimsApi,PolicyService,FraudScoringServiceandDocumentServicescale independently and deploy several times a week. Instance addresses change with every scale or release event. Current config-file URLs caused 3 failed deployments last quarter (illustrative scenario). - Decision: Use the hosting platform’s built-in registry (Azure Container Apps environment DNS today; Kubernetes Services if we move to AKS). Callers use logical names through
Microsoft.Extensions.ServiceDiscovery. Local development uses theServicesconfiguration section. No standalone Consul/Eureka cluster. - Consequences (positive): no additional infrastructure to operate; same caller code across dev, test and prod; scale and deploy without caller changes.
- Consequences (negative): limited instance metadata and routing rules; tied to platform semantics (a move from ACA to AKS changes naming conventions, not code).
- Alternatives rejected: Consul (operational cost with no requirement that the platform cannot meet); hardcoded URLs (current pain).
- Revisit when: we need multi-cloud or on-premise VM workloads, or weighted/metadata-based routing that the platform cannot provide.
Azure implementation
9.1 Which Azure services provide a service registry
| Azure service | How discovery works |
|---|---|
| Azure Container Apps (ACA) | Every app with ingress enabled gets a DNS name inside its environment. Internal calls use http://<app-name> or the FQDN <app-name>.internal.<env-unique-id>.<region>.azurecontainerapps.io (internal ingress). External ingress FQDN: <app-name>.<env-unique-id>.<region>.azurecontainerapps.io. Traffic goes through the environment’s Envoy proxy, which load-balances across replicas and revisions, so apps never talk to replica IPs directly. |
| Azure Kubernetes Service (AKS) | Kubernetes Service objects with CoreDNS: policy-service.claims.svc.cluster.local (or just policy-service inside the same namespace). Headless Services return pod IPs for client-side balancing (for example gRPC). Endpoints are tracked automatically from pod readiness. |
| Azure Service Fabric | Built-in Naming Service and reverse proxy (for existing Service Fabric estates). |
| Azure App Service | No dynamic registry: each app has a stable hostname (policy-service.azurewebsites.net or a custom domain); use configuration plus Front Door or API Management. |
| Dapr on ACA/AKS | Optional: service invocation by app ID via the sidecar (http://localhost:3500/v1.0/invoke/<app-id>/method/<method>), which also brings retries and mTLS. |
| Consul on AKS/VMs | Only when you need a standalone registry (multi-cloud/hybrid). You operate it. |
9.2 How to configure (ACA example)
- Create the Container Apps environment (in a VNet if you need private networking).
- Deploy
policy-servicewith ingress enabled, set to Internal (reachable only by other apps in the same environment). Target port = the port your ASP.NET Core app listens on (for .NET container images, port 8080). - In
ClaimsApi, set the environment variablePolicyService__BaseUrl=http://policy-service. - Configure health probes (liveness, readiness, startup) on
policy-serviceso unhealthy replicas do not receive traffic. - Set replica rules (HTTP concurrency or KEDA scale rules) and a sensible minimum replica count.
- Note: on internal-only environments, requests from a VNet client count as “outside” the environment; use ingress “external” (still not publicly exposed) if VNet clients need access.
Azure CLI sketch:
az containerapp create \ --name policy-service \ --resource-group rg-claims \ --environment cae-claims \ --image myregistry.azurecr.io/claims/policy-service:1.4.0 \ --ingress internal \ --target-port 8080 \ --min-replicas 2 --max-replicas 109.3 Kubernetes (AKS) example
apiVersion: v1kind: Servicemetadata: name: policy-service namespace: claimsspec: selector: app: policy-service ports: - port: 80 targetPort: 8080Readiness probes on the Deployment decide which pods appear as Service endpoints.
9.4 Pricing and tier considerations
- ACA: discovery itself is free; you pay for the app resources. The Consumption plan has monthly free grants per subscription (180,000 vCPU-seconds, 360,000 GiB-seconds, 2 million requests), then per-second vCPU/memory charges; replicas scaled to zero cost nothing, and idle replicas are billed at a reduced idle rate. Dedicated workload profiles add a fixed management fee plus per-instance cost. Check the Azure Container Apps pricing page for current numbers.
- AKS: CoreDNS and Services are included; you pay for node VMs. Choose cluster tier (Free vs Standard vs Premium) based on SLA needs; the Free tier has no financially backed uptime SLA, so use Standard for production. Verify current tier pricing on the AKS pricing page.
- Consul or Eureka: you also pay for their VMs/pods (minimum 3 servers for a Consul quorum) plus the operations time. This is the real cost of a standalone registry.
9.5 Reference architecture (text)
- Angular SPA on Azure Static Web Apps loads
/config.jsonand callshttps://api.claims.contoso.com. - Azure Front Door (WAF) forwards to the API Gateway (YARP or API Management) exposed with external ingress in the ACA environment.
- The gateway resolves
claims-apiby logical name; the ACA environment DNS + Envoy picks a healthy replica. ClaimsApicallspolicy-serviceandfraud-scoring-service(internal ingress) by name viaMicrosoft.Extensions.ServiceDiscovery(pass-through).- Each service uses its own database (Azure SQL / PostgreSQL Flexible Server) referenced by configuration and Key Vault, not discovery.
- Azure Monitor / Application Insights (OpenTelemetry) records which replica served each call, and alerts on
connection refusedspikes.
Teaching guide for my team
10.1 Explain to a beginner in 2 minutes
“When you order a taxi, you do not know which car will arrive, only that the taxi company will send one. Our services are like taxis: they start and stop all day, and each start gives them a new address. So instead of remembering addresses, we ask a directory: ‘give me a working policy-service’. The directory keeps a list, and instances that stop saying ‘I am alive’ are removed. In Azure Container Apps that directory is built in, so we just write http://policy-service.”
10.2 Explain to an intermediate developer in 5 minutes
- Show the failure:
ClaimsApiwith a hardcoded IP; scale/redeploypolicy-service; calls fail. - Introduce register / heartbeat / deregister / lookup and the TTL lease.
- Show the two flavours: platform registry (ACA, Kubernetes) versus standalone (Consul, Eureka), and why we default to the platform.
- Show the
AddServiceDiscovery()+ logicalBaseAddresssnippet and the dev-onlyServicesconfig. - Discuss the traps: stale caches, deregister-before-exit, readiness vs liveness, registry outage.
- Close with: discovery finds an instance; resilience (retry, circuit breaker) handles the instances it gets wrong.
10.3 Hands-on exercise
Task: Extend the Level 1 in-memory registry into a small ASP.NET Core minimal API with three endpoints: POST /register, POST /heartbeat, GET /lookup/{service}. Then start two fake instances of policy-service (two simple minimal APIs on ports 5001 and 5002) that register and heartbeat every 2 seconds. Finally write a console caller that resolves policy-service from the registry before each request and alternates between instances.
Expected outcome:
- Both instances alternate in the caller’s output.
- Stop one instance with Ctrl+C without deregistering: after the TTL (say 6 seconds), the caller stops getting it in lookups, and until then some calls fail (this shows why retries are still needed).
- Repeat with deregistration on shutdown: calls stop failing almost immediately.
10.4 Interview-style questions
- Why is a hardcoded IP or URL a bad idea in microservices? Instances are ephemeral: scaling, redeploys and failures change addresses, so configuration goes stale and callers break or overload old instances.
- What is the difference between client-side and server-side discovery, and which does Azure Container Apps use? Client-side: the caller queries the registry and balances itself. Server-side: the caller uses a router/load balancer that consults the registry. ACA is server-side: apps call a name that resolves to the environment’s Envoy proxy.
- The registry goes down. What should callers do? Keep using their last known (cached) instance list, combined with timeouts, retries and circuit breakers. Availability of stale data beats a total outage.
Mastery checklist
- I can explain register, heartbeat/lease, deregister and lookup, and why TTL is needed.
- I can say why the platform registry (ACA/Kubernetes) is my default and name two cases where I would use Consul.
- I can configure
Microsoft.Extensions.ServiceDiscoveryin a .NET 10 app with theServicesconfig locally and pass-through in Azure. - I can write the correct graceful-shutdown order (deregister, wait, drain, exit).
- I can explain the caching layers (registry client cache,
HttpClientconnection lifetime, DNS TTL) and their effect on stale routing. - I can explain AP versus CP for a registry and choose availability for discovery.
- I can describe how ACA internal FQDNs and AKS Service DNS names are formed.
- I can write a short ADR justifying “use the platform registry” and state when to revisit it.
Key takeaway
A Service Registry replaces hardcoded addresses with a live, health-aware directory of instances; on Azure, let Container Apps or Kubernetes be that directory, and design callers to survive stale data with caching, timeouts and retries.
