Manikandan — Manikandan
Microservices

Day 21: Service Registry

ManikandanManikandan
17 min read·Updated Aug 30, 2022

A Service Registry is a live "phone book" for your services.

Intro

A Service Registry is a live “phone book” for your services. Every running instance of a service (for example, three copies of policy-service) announces where it is (IP and port) and whether it is healthy. Any caller asks the phone book “where is policy-service right now?” instead of using a hardcoded address. Because containers start, stop, scale and move all the time, the phone book is updated automatically and never relies on a human editing a config file.

Why we need this

In a monolith, PolicyService is a class: you call it in-process and it always “exists”. In microservices, PolicyService is a process on a network, and its address changes constantly:

  • Autoscaling adds and removes instances several times a day (a storm produces a spike of claims at 09:00).
  • Deployments replace instances (rolling update: new pod IPs every release).
  • Failures kill instances; the platform restarts them elsewhere with a new IP.
  • Environments differ: dev is localhost:5001, test is a container, prod is a cluster.

Business reasons: teams must be able to deploy and scale their service without asking every consumer to change config. Technical reasons: hardcoded http://10.0.1.11:8080 values are wrong within minutes in any elastic platform, and a load balancer in front of everything becomes a manual-config bottleneck.

What problem it solves

Problem (from the topic list): services run on ephemeral IPs and ports that cannot be hardcoded.

Without a registry in our insurance claims system:

  1. ClaimsApi has PolicyServiceUrl = http://10.0.1.11:8080 in appsettings.json.
  2. Ops scales policy-service from 2 to 5 pods during a storm. ClaimsApi still calls only the original one. The new pods sit idle while the old one is overloaded.
  3. A rolling deployment replaces the pod at 10.0.1.11. ClaimsApi now gets connection refused until someone edits config and restarts it.
  4. Adjusters see “Cannot verify coverage” errors, and every fix needs a release of the caller, even though the callee is what changed.

With a registry: instances register themselves (or the platform registers them), unhealthy ones drop out, and callers always resolve to a current list of healthy instances.

When it is needed (and when it is NOT)

Needed when:

  • Instances are dynamic: autoscaling, containers, rolling updates, multiple environments.
  • You have many services calling each other (roughly more than a handful) and manual address management is already painful.
  • You run on VMs or a custom platform that does not give you discovery for free.
  • You need rich metadata per instance (version, zone, weight, tenant) for routing decisions.

NOT needed (or overkill) when:

  • You run on Kubernetes, Azure Container Apps or Service Fabric: the platform already is the registry (Kubernetes Services + CoreDNS, ACA built-in DNS). Do not add Consul or Eureka on top “because the book said so”.
  • You have 2-3 services with stable addresses behind a load balancer or App Service hostnames. A DNS name is enough.
  • You are still a modular monolith. Discovery solves a distribution problem you do not have yet.
  • Databases, Service Bus, Key Vault: these have stable managed DNS names and should be configured, not “discovered”.

How to identify the problem (key signals)

  1. Service URLs in config files that change with every environment and every release (appsettings.Production.json full of IPs).
  2. “Connection refused” / “No such host” spikes right after deployments or scale events, which disappear after callers are restarted.
  3. Uneven load: one instance at 95% CPU while newly scaled instances sit at 2%, because callers pinned to the old address.
  4. Deployment checklists that include “update the URL in service X and Y” steps.
  5. Calls routed to dead instances: timeouts of exactly your HTTP timeout value (e.g. 100 s default in HttpClient) for a fraction of requests, matching the number of dead instances.
  6. Stale DNS behaviour: a caller keeps hitting an old IP for minutes after failover because of OS/JVM/.NET connection or DNS caching.
  7. A shared “endpoints” wiki page or spreadsheet that people update by hand.

Flow Diagram

Instances register, renew their lease, and are looked up by logical name.

flowchart LR
I1["policy-service instance A"] -- "register + heartbeat" --> REG[("Service Registry")]
I2["policy-service instance B"] -- "register + heartbeat" --> REG
I3["crashed instance C"] -. "lease expired" .-> REG
CL["ClaimsApi"] -- "lookup policy-service" --> REG
REG -- "healthy: A, B" --> CL
CL -- "call" --> I1

Level 1: Beginner

Analogy: a hotel reception desk. Guests (instances) check in when they arrive and check out when they leave. If a guest stops answering the room phone (no heartbeat), reception assumes they left. Visitors (callers) never memorise room numbers; they ask reception “where is Mr. Policy?” every time.

The registry has four operations:

OperationWho calls itMeaning
RegisterInstance (or platform) on start“I am policy-service at 10.0.1.11:8080”
Heartbeat / renewInstance every few seconds“I am still alive” (a lease renewal)
DeregisterInstance on graceful shutdown“I am leaving”
LookupCaller“Give me healthy instances of policy-service”

Minimal working example (.NET 10, one file, pure in-memory, to show the idea; not for production):

using System.Collections.Concurrent;
var registry = new ServiceRegistry(ttl: TimeSpan.FromSeconds(1));
registry.Register("policy-service", "10.0.1.11:8080");
registry.Register("policy-service", "10.0.1.12:8080");
Thread.Sleep(700);
registry.Heartbeat("policy-service", "10.0.1.11:8080"); // only instance 11 renews its lease
Thread.Sleep(500);
// instance 12 last renewed 1.2 s ago (> 1 s TTL), so it is treated as dead
foreach (var address in registry.Lookup("policy-service"))
{
Console.WriteLine(address); // prints only 10.0.1.11:8080
}
public sealed class ServiceRegistry(TimeSpan ttl)
{
private readonly ConcurrentDictionary<(string Service, string Address), DateTimeOffset> _leases = new();
public void Register(string service, string address) =>
_leases[(service, address)] = DateTimeOffset.UtcNow;
public void Heartbeat(string service, string address) => Register(service, address);
public void Deregister(string service, string address) =>
_leases.TryRemove((service, address), out _);
public IReadOnlyList<string> Lookup(string service)
{
var cutoff = DateTimeOffset.UtcNow - ttl;
return _leases
.Where(l => l.Key.Service == service && l.Value >= cutoff)
.Select(l => l.Key.Address)
.Order()
.ToList();
}
}

Key beginner lesson: liveness is proven by a recent heartbeat (a lease with a TTL), not by “the instance registered once”.

Level 2: Intermediate

6.1 The two registry styles you will actually meet

  1. Platform registry (preferred): Kubernetes (Service + CoreDNS + EndpointSlices) or Azure Container Apps (built-in DNS + Envoy). Your code just uses a name like http://policy-service.
  2. Standalone registry: HashiCorp Consul, Netflix Eureka (mostly Java/Spring), etcd or ZooKeeper-based. Use when you run outside such a platform (VMs, hybrid, multi-cloud).

6.2 .NET client: Microsoft.Extensions.ServiceDiscovery

On .NET 10 (LTS), the Microsoft.Extensions.ServiceDiscovery NuGet package lets you write logical names in HttpClient base addresses and resolves them at call time. Out of the box it supports a configuration provider (a Services section) and a pass-through provider (returns the name as-is so the platform’s DNS resolves it). DNS and DNS SRV providers are available in the separate Microsoft.Extensions.ServiceDiscovery.Dns package. It works with or without .NET Aspire.

ClaimsApi/Program.cs:

var builder = WebApplication.CreateBuilder(args);
builder.Services.AddServiceDiscovery();
// every HttpClient created via IHttpClientFactory resolves logical names
builder.Services.ConfigureHttpClientDefaults(http => http.AddServiceDiscovery());
// "https+http://" = try HTTPS endpoints first, fall back to HTTP
var policyBaseUrl = builder.Configuration["PolicyService:BaseUrl"] ?? "https+http://policy-service";
builder.Services.AddHttpClient<PolicyClient>(c => c.BaseAddress = new Uri(policyBaseUrl));
var app = builder.Build();
app.MapGet("/claims/{claimId:guid}/coverage", async (Guid claimId, PolicyClient policies, CancellationToken ct) =>
{
var coverage = await policies.GetCoverageAsync(claimId, ct);
return coverage is null ? Results.NotFound() : Results.Ok(coverage);
});
app.Run();
public sealed record CoverageDto(string PolicyNumber, bool Active, decimal Limit);
public sealed class PolicyClient(HttpClient http)
{
public Task<CoverageDto?> GetCoverageAsync(Guid claimId, CancellationToken ct) =>
http.GetFromJsonAsync<CoverageDto>($"policies/by-claim/{claimId}/coverage", ct);
}

Local development uses the configuration provider (appsettings.Development.json):

{
"Services": {
"policy-service": {
"https": [ "localhost:7101" ]
}
}
}

In production on Kubernetes or Azure Container Apps you add no Services section. The pass-through provider hands policy-service to the platform DNS. On Azure Container Apps, set PolicyService__BaseUrl to http://policy-service (short app-name form) or to the internal FQDN.

The same code, three environments, zero address changes in the caller. That is the point.

6.3 Angular

The browser should not do service discovery. It talks to one stable public address (API Gateway or BFF; see Day 19 and 20). Load that address at runtime, so one build works in every environment (Angular 21/22 standalone style):

app-config.service.ts
import { HttpClient } from '@angular/common/http';
import { Injectable, inject } from '@angular/core';
import { firstValueFrom } from 'rxjs';
@Injectable({ providedIn: 'root' })
export class AppConfigService {
private readonly http = inject(HttpClient);
apiBaseUrl = '';
async load(): Promise<void> {
const cfg = await firstValueFrom(this.http.get<{ apiBaseUrl: string }>('/config.json'));
this.apiBaseUrl = cfg.apiBaseUrl;
}
}
app.config.ts
import { ApplicationConfig, inject, provideAppInitializer } from '@angular/core';
import { provideHttpClient } from '@angular/common/http';
import { AppConfigService } from './app-config.service';
export const appConfig: ApplicationConfig = {
providers: [
provideHttpClient(),
provideAppInitializer(() => inject(AppConfigService).load()),
],
};

/config.json is a static file replaced per environment at deploy time (for example { "apiBaseUrl": "https://api.claims.contoso.com" }).

6.4 SQL Server / PostgreSQL

Not applicable as a registry concern: Azure SQL Database and Azure Database for PostgreSQL have stable DNS names supplied through configuration or Key Vault. Do not register databases in a service registry. The only database link is if you build a custom registry: keep its state in memory or a purpose-built store (Consul/etcd), never in your business database, because lookups happen on every call path and need heartbeat-grade write rates.

Level 3: Advanced

7.1 Health and leases: what “healthy” means

  • Lease/TTL model (Eureka-style): the instance must renew, otherwise it expires. Simple, but a paused process (GC pause, CPU starvation) can be evicted while alive.
  • Active health checks (Consul-style): the registry (or agent) probes /health/ready. Detects “process alive but dependency broken”.
  • Register readiness, not liveness. An instance whose SQL connection is down must not receive traffic, but it should not be killed either. In ASP.NET Core map separate endpoints (/health/live, /health/ready) via MapHealthChecks (see Day 36).

7.2 Graceful shutdown ordering (the most common production bug)

Pseudo-code for a self-registering instance:

on SIGTERM:
1. deregister from registry (stop NEW traffic)
2. wait propagation delay (e.g. 5-10 s) (callers refresh their caches)
3. finish in-flight requests (ASP.NET Core graceful shutdown)
4. exit

If you exit first and deregister last, callers keep sending requests to a dead address for the length of their cache TTL. In Kubernetes the equivalent is a preStop sleep plus a terminationGracePeriodSeconds that exceeds it.

7.3 Caching and staleness

  • Callers cache lookups; a longer cache reduces registry load but increases stale routing. Typical: refresh 5-30 s.
  • HttpClient with SocketsHttpHandler keeps connections alive; new endpoints are not used until connections are recycled. Set PooledConnectionLifetime (for example 2-5 minutes) so DNS/registry changes are honoured. IHttpClientFactory rotates handlers (default 2 minutes) for the same reason.
  • Design for the registry being down: callers must keep using their last known list (stale-but-available beats unavailable).

7.4 Consistency vs availability

RegistryModelBehaviour in a network partition
EurekaAP (available, eventually consistent)Keeps serving possibly stale data; “self-preservation” stops mass eviction when many heartbeats are missed at once
ConsulCP for the catalog (Raft), with agent-side healthCan refuse writes without a quorum; reads can be stale if configured
Kubernetes (etcd + EndpointSlices)CP store, eventually consistent propagation to nodesNode-local view can lag briefly during changes

Rule of thumb: for discovery, availability beats consistency. A slightly stale list is survivable; a registry outage that stops all calls is not.

7.5 Security

  • Restrict who can register: a rogue process registering as policy-service can intercept claim data. Use ACLs/tokens (Consul ACLs) or platform RBAC.
  • Use mTLS between services; discovery tells you where, it does not prove who.
  • Do not expose the registry UI or API publicly.

7.6 Common mistakes

  1. Registering localhost or a container-internal IP that other hosts cannot reach.
  2. Registering on start, before the app is ready (traffic arrives to a warming service). Register on ready.
  3. Ignoring deregistration on shutdown.
  4. Building a home-grown registry when the platform already provides one.
  5. Making the registry a single point of failure (one instance, no cluster, no client-side cache).
  6. Putting registry lookup on the hot path of every request without caching.
  7. Confusing discovery with load balancing and resilience: you still need timeouts, retries and circuit breakers (Days 26-28).

Level 4: Expert and Architect view

8.1 Design trade-offs

OptionWhere the registry livesProsConsBest for
Platform DNS (Kubernetes Service, ACA app name)Platform control planeZero code, zero extra infra, language-neutralCoarse: no rich metadata, little routing logicDefault choice on AKS/ACA
Standalone registry (Consul, Eureka)Separate cluster you operateMetadata, multi-DC, works off-platformYou must operate and secure it (HA, ACLs, upgrades)VMs, hybrid, multi-cloud
Config file / env varsApp configSimplestManual, stale, no health awarenessVery small, static systems
Hardcoded load balancer DNSCloud load balancerFamiliar, health probes built inOne LB per service, cost, manual wiringSmall number of services
Service mesh (Istio, Linkerd)Mesh control plane (reads Kubernetes)mTLS, traffic shifting, telemetryHigh complexityLarge fleets (see Day 48)

8.2 Patterns it combines with

  • Client-Side Discovery (Day 22), Server-Side Discovery (Day 23): how the lookup is done.
  • Self-Registration (Day 24), 3rd Party Registration (Day 25): how registration is done.
  • Health Check API (Day 36): the source of truth for “healthy”.
  • Circuit Breaker / Retry (Days 26, 28): protect against the window between “instance died” and “registry knows”.
  • API Gateway (Day 19): the gateway is the biggest consumer of the registry.

8.3 ADR (architecture review ready)

ADR-021: Service discovery for the claims platform

  • Status: Proposed
  • Context: ClaimsApi, PolicyService, FraudScoringService and DocumentService scale independently and deploy several times a week. Instance addresses change with every scale or release event. Current config-file URLs caused 3 failed deployments last quarter (illustrative scenario).
  • Decision: Use the hosting platform’s built-in registry (Azure Container Apps environment DNS today; Kubernetes Services if we move to AKS). Callers use logical names through Microsoft.Extensions.ServiceDiscovery. Local development uses the Services configuration section. No standalone Consul/Eureka cluster.
  • Consequences (positive): no additional infrastructure to operate; same caller code across dev, test and prod; scale and deploy without caller changes.
  • Consequences (negative): limited instance metadata and routing rules; tied to platform semantics (a move from ACA to AKS changes naming conventions, not code).
  • Alternatives rejected: Consul (operational cost with no requirement that the platform cannot meet); hardcoded URLs (current pain).
  • Revisit when: we need multi-cloud or on-premise VM workloads, or weighted/metadata-based routing that the platform cannot provide.

Azure implementation

9.1 Which Azure services provide a service registry

Azure serviceHow discovery works
Azure Container Apps (ACA)Every app with ingress enabled gets a DNS name inside its environment. Internal calls use http://<app-name> or the FQDN <app-name>.internal.<env-unique-id>.<region>.azurecontainerapps.io (internal ingress). External ingress FQDN: <app-name>.<env-unique-id>.<region>.azurecontainerapps.io. Traffic goes through the environment’s Envoy proxy, which load-balances across replicas and revisions, so apps never talk to replica IPs directly.
Azure Kubernetes Service (AKS)Kubernetes Service objects with CoreDNS: policy-service.claims.svc.cluster.local (or just policy-service inside the same namespace). Headless Services return pod IPs for client-side balancing (for example gRPC). Endpoints are tracked automatically from pod readiness.
Azure Service FabricBuilt-in Naming Service and reverse proxy (for existing Service Fabric estates).
Azure App ServiceNo dynamic registry: each app has a stable hostname (policy-service.azurewebsites.net or a custom domain); use configuration plus Front Door or API Management.
Dapr on ACA/AKSOptional: service invocation by app ID via the sidecar (http://localhost:3500/v1.0/invoke/<app-id>/method/<method>), which also brings retries and mTLS.
Consul on AKS/VMsOnly when you need a standalone registry (multi-cloud/hybrid). You operate it.

9.2 How to configure (ACA example)

  1. Create the Container Apps environment (in a VNet if you need private networking).
  2. Deploy policy-service with ingress enabled, set to Internal (reachable only by other apps in the same environment). Target port = the port your ASP.NET Core app listens on (for .NET container images, port 8080).
  3. In ClaimsApi, set the environment variable PolicyService__BaseUrl=http://policy-service.
  4. Configure health probes (liveness, readiness, startup) on policy-service so unhealthy replicas do not receive traffic.
  5. Set replica rules (HTTP concurrency or KEDA scale rules) and a sensible minimum replica count.
  6. Note: on internal-only environments, requests from a VNet client count as “outside” the environment; use ingress “external” (still not publicly exposed) if VNet clients need access.

Azure CLI sketch:

Terminal window
az containerapp create \
--name policy-service \
--resource-group rg-claims \
--environment cae-claims \
--image myregistry.azurecr.io/claims/policy-service:1.4.0 \
--ingress internal \
--target-port 8080 \
--min-replicas 2 --max-replicas 10

9.3 Kubernetes (AKS) example

apiVersion: v1
kind: Service
metadata:
name: policy-service
namespace: claims
spec:
selector:
app: policy-service
ports:
- port: 80
targetPort: 8080

Readiness probes on the Deployment decide which pods appear as Service endpoints.

9.4 Pricing and tier considerations

  • ACA: discovery itself is free; you pay for the app resources. The Consumption plan has monthly free grants per subscription (180,000 vCPU-seconds, 360,000 GiB-seconds, 2 million requests), then per-second vCPU/memory charges; replicas scaled to zero cost nothing, and idle replicas are billed at a reduced idle rate. Dedicated workload profiles add a fixed management fee plus per-instance cost. Check the Azure Container Apps pricing page for current numbers.
  • AKS: CoreDNS and Services are included; you pay for node VMs. Choose cluster tier (Free vs Standard vs Premium) based on SLA needs; the Free tier has no financially backed uptime SLA, so use Standard for production. Verify current tier pricing on the AKS pricing page.
  • Consul or Eureka: you also pay for their VMs/pods (minimum 3 servers for a Consul quorum) plus the operations time. This is the real cost of a standalone registry.

9.5 Reference architecture (text)

  1. Angular SPA on Azure Static Web Apps loads /config.json and calls https://api.claims.contoso.com.
  2. Azure Front Door (WAF) forwards to the API Gateway (YARP or API Management) exposed with external ingress in the ACA environment.
  3. The gateway resolves claims-api by logical name; the ACA environment DNS + Envoy picks a healthy replica.
  4. ClaimsApi calls policy-service and fraud-scoring-service (internal ingress) by name via Microsoft.Extensions.ServiceDiscovery (pass-through).
  5. Each service uses its own database (Azure SQL / PostgreSQL Flexible Server) referenced by configuration and Key Vault, not discovery.
  6. Azure Monitor / Application Insights (OpenTelemetry) records which replica served each call, and alerts on connection refused spikes.

Teaching guide for my team

10.1 Explain to a beginner in 2 minutes

“When you order a taxi, you do not know which car will arrive, only that the taxi company will send one. Our services are like taxis: they start and stop all day, and each start gives them a new address. So instead of remembering addresses, we ask a directory: ‘give me a working policy-service’. The directory keeps a list, and instances that stop saying ‘I am alive’ are removed. In Azure Container Apps that directory is built in, so we just write http://policy-service.”

10.2 Explain to an intermediate developer in 5 minutes

  1. Show the failure: ClaimsApi with a hardcoded IP; scale/redeploy policy-service; calls fail.
  2. Introduce register / heartbeat / deregister / lookup and the TTL lease.
  3. Show the two flavours: platform registry (ACA, Kubernetes) versus standalone (Consul, Eureka), and why we default to the platform.
  4. Show the AddServiceDiscovery() + logical BaseAddress snippet and the dev-only Services config.
  5. Discuss the traps: stale caches, deregister-before-exit, readiness vs liveness, registry outage.
  6. Close with: discovery finds an instance; resilience (retry, circuit breaker) handles the instances it gets wrong.

10.3 Hands-on exercise

Task: Extend the Level 1 in-memory registry into a small ASP.NET Core minimal API with three endpoints: POST /register, POST /heartbeat, GET /lookup/{service}. Then start two fake instances of policy-service (two simple minimal APIs on ports 5001 and 5002) that register and heartbeat every 2 seconds. Finally write a console caller that resolves policy-service from the registry before each request and alternates between instances.

Expected outcome:

  • Both instances alternate in the caller’s output.
  • Stop one instance with Ctrl+C without deregistering: after the TTL (say 6 seconds), the caller stops getting it in lookups, and until then some calls fail (this shows why retries are still needed).
  • Repeat with deregistration on shutdown: calls stop failing almost immediately.

10.4 Interview-style questions

  1. Why is a hardcoded IP or URL a bad idea in microservices? Instances are ephemeral: scaling, redeploys and failures change addresses, so configuration goes stale and callers break or overload old instances.
  2. What is the difference between client-side and server-side discovery, and which does Azure Container Apps use? Client-side: the caller queries the registry and balances itself. Server-side: the caller uses a router/load balancer that consults the registry. ACA is server-side: apps call a name that resolves to the environment’s Envoy proxy.
  3. The registry goes down. What should callers do? Keep using their last known (cached) instance list, combined with timeouts, retries and circuit breakers. Availability of stale data beats a total outage.

Mastery checklist

  • I can explain register, heartbeat/lease, deregister and lookup, and why TTL is needed.
  • I can say why the platform registry (ACA/Kubernetes) is my default and name two cases where I would use Consul.
  • I can configure Microsoft.Extensions.ServiceDiscovery in a .NET 10 app with the Services config locally and pass-through in Azure.
  • I can write the correct graceful-shutdown order (deregister, wait, drain, exit).
  • I can explain the caching layers (registry client cache, HttpClient connection lifetime, DNS TTL) and their effect on stale routing.
  • I can explain AP versus CP for a registry and choose availability for discovery.
  • I can describe how ACA internal FQDNs and AKS Service DNS names are formed.
  • I can write a short ADR justifying “use the platform registry” and state when to revisit it.

Key takeaway

A Service Registry replaces hardcoded addresses with a live, health-aware directory of instances; on Azure, let Container Apps or Kubernetes be that directory, and design callers to survive stale data with caching, timeouts and retries.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 25: 3rd Party Registration

3rd Party Registration is a service-discovery pattern where the service itself never talks to the service registry.

Manikandan
Manikandan·13 min read
Microservices

Day 24: Self-Registration

Self-Registration is a service-discovery pattern in which each service instance takes responsibility for announcing itself to the service registry when it starts, keeps that registration alive while it is healthy,...

Manikandan
Manikandan·17 min read
Microservices

Day 23: Server-Side Discovery

Server-side discovery means the caller does not look up service instances at all.

Manikandan
Manikandan·17 min read
Microservices

Day 22: Client-Side Discovery

Client-Side Discovery means the *caller* is responsible for finding a service instance.

Manikandan
Manikandan·16 min read