Self-Registration is a service-discovery pattern in which each service instance takes responsibility for announcing itself to the service registry when it starts, keeps that registration alive while it is healthy,...
Intro
Self-Registration is a service-discovery pattern in which each service instance takes responsibility for announcing itself to the service registry when it starts, keeps that registration alive while it is healthy, and removes it when it shuts down. Callers never need hard-coded IPs or ports. They ask the registry “where is claims-service right now?” and get back the live instances. In our insurance claims system, every new copy of ClaimsService that the platform spins up during a storm-season spike registers itself, and every copy that is scaled in or crashes disappears from the registry.
Versions referenced in this lesson (verified September 2026): .NET 10 (LTS, released 11 Nov 2025, supported until 14 Nov 2028), Angular 22 (released 3 Jun 2026; Angular 21 is in LTS), Consul .NET client (Consul NuGet package).
Why we need this
Modern services run on ephemeral infrastructure. Containers get new IPs on every restart, autoscalers add and remove instances, and blue/green rollouts create instances that live for hours. A static configuration file (ClaimsServiceUrl=http://10.0.1.15:8080) is wrong the moment it is committed.
- Business reason: claim volumes are spiky (hailstorm, flood, end of policy year). The system must add capacity without an operator editing configuration files or redeploying every caller.
- Technical reason: something must maintain the mapping of logical service name to live network locations. Self-Registration is the simplest way to populate that mapping: the instance itself knows its own address, port, version, and readiness better than anyone else.
- Operational reason: no separate agent or orchestrator integration is needed. Any environment that can run your code (VMs, bare metal, containers without an orchestrator) can participate in discovery.
What problem it solves
Problem from the topic list: manually tracking scaling instances is unmanageable.
Without the pattern, someone (or some script) must edit a config file, load-balancer pool, or DNS record each time an instance starts or stops. Typical failures:
- A claims-intake instance is scaled out from 2 to 6 but callers still only know about the original 2, so the new capacity is idle while the old ones are overloaded.
- An instance crashes at 02:00. Its IP stays in the load-balancer pool, so roughly one in N requests fails until someone removes it manually.
- A rolling deployment replaces all instances; callers keep sending traffic to IPs that no longer exist.
- Environments drift: dev, test, and prod each have hand-maintained lists that silently diverge.
With Self-Registration, the lifecycle of the registry entry is tied to the lifecycle of the process: register on start, heartbeat while alive, deregister on graceful stop, and expire on crash.
When it is needed (and when it is NOT)
Good fit
- You run your own registry (Consul, Eureka, etcd, ZooKeeper) and services run on VMs or on containers without an orchestrator that handles discovery.
- You are migrating a monolith to services on VMs and need discovery quickly, without adopting Kubernetes.
- You need rich metadata per instance (version, zone, weight, feature tags) that only the instance knows.
- Most of your services are in one language or framework, so a shared registration library is affordable (see Day 40, Microservice Chassis).
Poor fit or overkill
- You deploy on Kubernetes, Azure Container Apps, or similar. The platform already registers instances for you (this is the 3rd Party Registration pattern, Day 25). Adding Consul registration inside the app is redundant and creates two sources of truth.
- You have two or three services with stable addresses behind a load balancer. A DNS name is enough.
- You run many languages and cannot maintain a registration client for each. Prefer platform-managed registration.
- Your service is short-lived (a function or a batch job). Registration overhead outweighs the benefit; use serverless triggers instead.
How to identify the problem (key signals)
- Deployments require editing URLs, IP lists, or
appsettings.Production.jsonin other services. - After a scale-out, the new instances show near-zero CPU and request counts while old instances are saturated.
- Intermittent
502/503/connection refusederrors that correlate with restarts or scale-in events, and disappear after a manual pool cleanup. - Runbooks contain steps like “remove the node from the load balancer before patching”.
- Incident reviews mention “stale endpoint” or “wrong IP” as a contributing factor.
- Config repos contain many hard-coded
http://10.x.x.xvalues, or environment-specific URL lists that differ between environments. - Autoscaling is disabled or limited “because callers wouldn’t find the new instances”.
Flow Diagram
The instance itself manages its registry entry for its whole lifecycle.
flowchart TD ST(["Process starts"]) --> REG["Register: unique ID, address, port"] REG --> CHK["Registry runs health check /health/ready"] CHK --> SERVE["Serve traffic"] SERVE --> STOP{"How does it stop?"} STOP -- "graceful" --> DEREG["Deregister on ApplicationStopping"] STOP -- "kill -9 / crash" --> CRIT["Check fails: marked critical"] CRIT --> EXP["Removed after DeregisterCriticalServiceAfter"]Level 1: Beginner
Analogy: a food truck that posts its location on a shared map every morning (“I’m at Pier 9 today”), keeps updating “still open” during the day, and removes the pin when it leaves. Customers never call the owner; they look at the map. The truck (the instance) updates its own pin (registry entry).
Lifecycle in four steps
- Start: instance boots and registers
{name, id, address, port}with the registry. - Live: the registry keeps checking health (registry polls a health URL, or the instance sends TTL heartbeats).
- Stop: on graceful shutdown the instance deregisters itself.
- Crash: no heartbeat or failing health check means the registry marks the instance critical and eventually removes it.
Minimal example (ASP.NET Core, .NET 10, Consul registry)
dotnet new web -n ClaimsServicecd ClaimsServicedotnet add package Consul// Program.cs - minimal self-registration for ClaimsServiceusing Consul;
var builder = WebApplication.CreateBuilder(args);var app = builder.Build();
app.MapGet("/health", () => Results.Ok("healthy"));app.MapGet("/claims/{id}", (string id) => Results.Ok(new { id, status = "Open" }));
var consul = new ConsulClient(c => c.Address = new Uri("http://localhost:8500"));var registration = new AgentServiceRegistration{ ID = $"claims-service-{Environment.MachineName}-5000", Name = "claims-service", Address = "localhost", Port = 5000, Check = new AgentServiceCheck { HTTP = "http://localhost:5000/health", Interval = TimeSpan.FromSeconds(10), Timeout = TimeSpan.FromSeconds(3), DeregisterCriticalServiceAfter = TimeSpan.FromMinutes(1) }};
app.Lifetime.ApplicationStarted.Register(() => consul.Agent.ServiceRegister(registration).GetAwaiter().GetResult());app.Lifetime.ApplicationStopping.Register(() => consul.Agent.ServiceDeregister(registration.ID).GetAwaiter().GetResult());
app.Run("http://localhost:5000");Run a local Consul in dev mode with consul agent -dev, start the service, then open http://localhost:8500/ui and you will see claims-service listed as passing. Stop the service with Ctrl+C and it disappears.
Level 2: Intermediate
In a real system, registration is a reusable component (a BackgroundService/hosted service), configured per environment, and callers resolve names through the registry rather than URLs.
6.1 Registration as a hosted service (.NET 10)
using Consul;using Microsoft.Extensions.Options;
public sealed class ConsulOptions{ public string ConsulAddress { get; set; } = "http://localhost:8500"; public string ServiceName { get; set; } = "claims-service"; public string ServiceAddress { get; set; } = ""; // set by platform/env var public int ServicePort { get; set; } = 8080; public string Version { get; set; } = "1.0.0";}
public sealed class ConsulRegistrationService( IConsulClient consul, IOptions<ConsulOptions> options, ILogger<ConsulRegistrationService> logger) : IHostedService{ private readonly ConsulOptions _o = options.Value; private string? _serviceId;
public async Task StartAsync(CancellationToken ct) { _serviceId = $"{_o.ServiceName}-{Guid.NewGuid():N}"; // unique per instance
var registration = new AgentServiceRegistration { ID = _serviceId, Name = _o.ServiceName, Address = _o.ServiceAddress, Port = _o.ServicePort, Tags = ["api", $"v{_o.Version}"], Meta = new Dictionary<string, string> { ["version"] = _o.Version }, Check = new AgentServiceCheck { HTTP = $"http://{_o.ServiceAddress}:{_o.ServicePort}/health/ready", Interval = TimeSpan.FromSeconds(10), Timeout = TimeSpan.FromSeconds(3), // remove instances that stay critical, e.g. after a hard crash DeregisterCriticalServiceAfter = TimeSpan.FromMinutes(2) } };
await consul.Agent.ServiceRegister(registration, ct); logger.LogInformation("Registered {ServiceId} at {Address}:{Port}", _serviceId, _o.ServiceAddress, _o.ServicePort); }
public async Task StopAsync(CancellationToken ct) { if (_serviceId is null) return; await consul.Agent.ServiceDeregister(_serviceId, ct); logger.LogInformation("Deregistered {ServiceId}", _serviceId); }}using Consul;
var builder = WebApplication.CreateBuilder(args);
builder.Services.Configure<ConsulOptions>(builder.Configuration.GetSection("Consul"));builder.Services.AddSingleton<IConsulClient>(sp =>{ var o = sp.GetRequiredService<Microsoft.Extensions.Options.IOptions<ConsulOptions>>().Value; return new ConsulClient(c => c.Address = new Uri(o.ConsulAddress));});builder.Services.AddHostedService<ConsulRegistrationService>();builder.Services.AddHealthChecks();
var app = builder.Build();app.MapHealthChecks("/health/ready");app.MapGet("/claims/{id}", (string id) => Results.Ok(new { id, status = "Open" }));app.Run();Notes: hosted services start before the web server begins listening in the default host ordering, so the registry could route traffic to an instance that is not yet accepting connections. Two mitigations: (1) the registry health check keeps the instance out of rotation until /health/ready passes (used above), and (2) register from IHostApplicationLifetime.ApplicationStarted when you must guarantee the server is listening.
6.2 Caller side: resolve by name
// ClaimsClient.cs - PolicyService calling ClaimsService by logical nameusing Consul;
public sealed class ClaimsClient(IConsulClient consul, HttpClient http){ public async Task<string> GetClaimAsync(string id, CancellationToken ct) { // passingOnly: true returns only instances whose health checks pass var result = await consul.Health.Service("claims-service", tag: null, passingOnly: true, ct); var instances = result.Response; if (instances.Length == 0) throw new InvalidOperationException("No healthy claims-service instance.");
var pick = instances[Random.Shared.Next(instances.Length)].Service; // simple client-side LB return await http.GetStringAsync($"http://{pick.Address}:{pick.Port}/claims/{id}", ct); }}This is Client-Side Discovery (Day 22) layered on top of Self-Registration. Self-Registration answers “how does the registry learn about instances”, while Client-Side or Server-Side Discovery answers “how does the caller use the registry”.
6.3 Angular and the database
- Angular never talks to the registry. Browsers should call a stable gateway (Day 19 API Gateway or Day 20 BFF) whose URL is in environment config; the gateway resolves service names. Registering browser clients would be a design error.
// claims.service.ts (Angular 22, standalone, uses HttpClient)import { inject, Injectable } from '@angular/core';import { HttpClient } from '@angular/common/http';import { environment } from '../environments/environment';
export interface Claim { id: string; status: string; }
@Injectable({ providedIn: 'root' })export class ClaimsApi { private readonly http = inject(HttpClient); // Always the gateway, never an individual service instance private readonly base = `${environment.apiBaseUrl}/claims`;
get(id: string) { return this.http.get<Claim>(`${this.base}/${id}`); }}- Database: registration state lives in the registry, not in SQL Server or PostgreSQL. A database is a dependency your readiness check may probe (
AddNpgSqlorAddSqlServerhealth checks from theAspNetCore.HealthChecks.*packages), so an instance that cannot reach PostgreSQL reports not-ready and is excluded from discovery results.
Level 3: Advanced
Failure modes and how to handle them
| Failure | What happens | Mitigation |
|---|---|---|
| Process crashes (no deregister) | Stale entry remains | Health checks plus DeregisterCriticalServiceAfter, or TTL checks that expire |
| Instance registers before it is ready | Requests hit a cold or unmigrated instance | Register as ready only after warm-up; use a readiness endpoint that checks dependencies |
| Registry outage at startup | Service cannot register | Retry with backoff (Day 28) instead of crashing; decide explicitly whether the service starts unregistered |
| Registry outage while running | Callers cannot resolve | Callers cache last-known instance list; existing registrations survive on the local agent |
| Split-brain or network partition | Healthy instances look dead | Use a local registry agent per node, and tune check intervals |
| Instance ID collision | One instance overwrites another | Use a unique ID per instance (name-guid), never just the name |
| Wrong advertised address | Registered localhost or a container-internal IP unreachable by callers | Inject the routable address from environment (host IP, pod IP) and never guess |
Performance and scalability
- Every instance issues a write on start and stop, and health checks add steady load. A Consul cluster handles thousands of instances, but poll intervals of 1 to 2 seconds across hundreds of services create needless load. 10 seconds is a common default.
- Caller-side lookups must be cached or use blocking queries (
WaitTime/WaitIndex) instead of hitting the registry on every request.
Security
- Protect the registry API with ACL tokens, so a compromised service can only register its own name (Consul ACL policies with
service "claims-service" { policy = "write" }). - Use TLS between services and the registry; do not expose port 8500 to a network segment where untrusted workloads run.
- Do not put secrets or connection strings in registry metadata (
Meta,Tags). They are readable by anyone who can query the catalog.
Common mistakes
- Using the service name as the instance ID, so the second instance replaces the first.
- Registering
localhostor127.0.0.1in a container. Other hosts cannot reach that address. - Skipping deregistration on shutdown and having no crash timeout, which leaves ghost instances.
- A
/healthendpoint that always returns 200 even when the database is down. - Coupling business code to the registry client instead of isolating it in a hosted service or shared library.
- Blocking startup forever when the registry is temporarily down.
Level 4: Expert and Architect view
8.1 Self-Registration vs. alternatives
| Aspect | Self-Registration | 3rd Party Registration (Day 25) | DNS-only (static names) | Platform-native (Kubernetes Service, Container Apps) |
|---|---|---|---|---|
| Who registers | The instance | Orchestrator or agent | Operator or IaC | Platform automatically |
| Coupling to registry | Service code depends on registry client | Service unaware | None | None |
| Language coverage | Client library per language | Any language | Any | Any |
| Metadata richness | High, since the instance knows itself | Medium, limited to what the agent can see | Low | Medium (labels/annotations) |
| Failure knowledge | Instance may die without deregistering, so needs TTL/health checks | Agent observes container state, more reliable | None | Platform probes readiness |
| Operational effort | Registry cluster plus client libraries | Registry plus registrar | Lowest | Lowest |
| Best fit | VMs, mixed hosts, non-orchestrated | Orchestrated environments with a registry | Small, stable systems | Kubernetes and PaaS |
8.2 Patterns it combines with
- Service Registry (Day 21): the datastore that Self-Registration writes to.
- Client-Side Discovery (Day 22) or Server-Side Discovery (Day 23): how callers read from it.
- Health Check API (Day 36): the liveness/readiness endpoints the registry uses to trust an entry.
- Microservice Chassis (Day 40): the shared library that packages the registration code so teams do not reinvent it.
- Circuit Breaker and Retry (Days 26 and 28): required because the registry can briefly return dead instances.
- Externalized Configuration (Day 39): registry address, advertised address, and tokens come from config or Key Vault.
8.3 ADR (architecture review format)
ADR-024: Service instances self-register with Consul for claims-platform services on VMs
- Status: Proposed
- Context: The claims platform (ClaimsService, PolicyService, PaymentsService, DocumentService) runs on Azure VM Scale Sets during the migration from the monolith. Instance counts change with load and with patch rollouts. There is no orchestrator that can register instances, and all services are .NET, so a shared library is feasible.
- Decision: Each service registers itself with the local Consul agent through a shared
Platform.DiscoveryNuGet package (hosted service). Registration uses a unique instance ID, an HTTP readiness check against/health/ready, andDeregisterCriticalServiceAfterof 2 minutes. Callers use passing-only health queries with a cached, in-process load balancer. - Consequences (positive): automatic scale-out and scale-in visibility, no manual pool edits, instance metadata (version) available for canary routing.
- Consequences (negative): a dependency on Consul availability and operations; each service embeds registry client code; we must maintain the library across .NET versions; ghost instances are possible if checks are misconfigured.
- Alternatives considered: platform-managed registration (rejected for now, since no orchestrator exists), DNS-only (rejected, no health awareness), Azure Load Balancer only (viable for a small number of stable services but lacks per-service metadata).
- Revisit when: services move to Azure Container Apps or AKS. At that point platform-native discovery replaces this ADR and the library is retired.
Azure implementation
Azure does not offer a managed Consul-style registry as a first-party service. Self-Registration on Azure therefore appears in three forms.
Option A: Consul (self-managed or HashiCorp-managed) on Azure
- Run a 3 or 5 server Consul cluster on VMs, or on AKS via the official Helm chart. Place clients on each VM Scale Set node or as a DaemonSet. HashiCorp also offers HCP Consul, so verify current availability and regions before choosing it.
- Configure: inject
Consul__ConsulAddress,Consul__ServiceAddress(VM private IP via Azure Instance Metadata Service athttp://169.254.169.254/metadata/instance), and the ACL token from Azure Key Vault using Managed Identity. Put the cluster in the same VNet and restrict port 8500/8501 with NSGs. - Cost: no license cost for open-source Consul; you pay for VMs (for example three small VMs for servers), disks, and monitoring.
Option B: Azure Container Apps (registration done for you)
- Apps in the same Container Apps environment discover each other by app name (
http://claims-service) or by internal FQDN of the form<app>.internal.<env-unique-id>.<region>.azurecontainerapps.io. The environment DNS resolves to the built-in Envoy proxy, which routes to healthy replicas. You write no registration code. This is the platform performing the registration (Day 25), not Self-Registration. - Configure: set ingress to internal for private services; set min/max replicas and HTTP scale rules; define liveness and readiness probes so the platform only routes to ready replicas.
- Pricing: Consumption plan bills per second for active and idle vCPU and memory plus requests, with a monthly free grant of 180,000 vCPU-seconds, 360,000 GiB-seconds and 2 million requests, and no charge when scaled to zero. Check the pricing page for current per-second rates in your region. Dedicated workload profiles and savings plans exist for steady workloads.
Option C: AKS
- Kubernetes Services and EndpointSlices track ready pods, and CoreDNS gives names such as
claims-service.claims.svc.cluster.local. Again platform-managed, so do not add Consul registration inside pods unless you need multi-cluster or hybrid discovery. - Pricing: control plane has Free, Standard, and Premium tiers (Free has no control-plane fee but no uptime SLA; Standard and Premium carry an hourly per-cluster charge). Node VMs are billed separately. Verify current rates.
Reference architecture (text) for the VM migration phase
Azure Front Door or Application Gateway terminates internet traffic and forwards to an internal API Gateway (YARP on .NET 10, or Azure API Management). The gateway queries the Consul cluster (3 servers across availability zones in one VNet) for claims-service. ClaimsService instances run in a VM Scale Set. Each VM runs a Consul client agent, and the app registers with the local agent at startup using an ID that includes the VM instance ID. Instances read secrets from Key Vault through Managed Identity. Health probes (/health/ready) verify SQL Server or Azure Database for PostgreSQL connectivity. Application Insights and Azure Monitor receive logs, traces, and metrics, with an alert on “registered instance count below N” and “critical checks greater than 0 for 5 minutes”. When the scale set autoscale rule adds VMs, they register themselves within seconds of becoming ready. On scale-in, the graceful shutdown hook deregisters (the scale set’s terminate notification gives the app time to do this).
Teaching guide for my team
Explain to a beginner in 2 minutes
“Services are like food trucks that move around. Instead of us writing down where each truck parks, each truck posts its own location on a shared board when it opens, keeps saying ‘still open’, and removes itself when it closes. The board is the service registry. Other services just read the board to find a truck. If a truck breaks down and stops saying ‘still open’, the board removes it after a short while.”
Explain to an intermediate developer in 5 minutes
Cover: (1) ephemeral IPs make static config impossible; (2) the instance registers {unique ID, name, address, port, health check} at startup through a hosted service; (3) the registry runs health checks or expects TTL heartbeats, and marks failing instances critical; (4) graceful shutdown deregisters, and crash cases are handled by DeregisterCriticalServiceAfter; (5) callers query passing-only instances and cache; (6) the downside: registry client code in every service and a registry to operate, so on Kubernetes or Container Apps you use the platform instead (3rd Party Registration). Draw the sequence: start, register, check, serve, stop, deregister.
Hands-on exercise
- Run
consul agent -devlocally. - Build
ClaimsServicefrom section 6.1 and start three copies on ports 5001, 5002, 5003 using differentConsul__ServicePortvalues. - Build a small
PolicyServiceendpoint using theClaimsClientfrom section 6.2 and call it repeatedly. - Kill one ClaimsService process with
kill -9(no graceful shutdown). Watch the Consul UI.
Expected outcome: three instances appear as passing. Calls are spread across all three. After kill -9, the killed instance turns critical within about one check interval, callers using passingOnly: true stop receiving it, and after DeregisterCriticalServiceAfter it disappears from the catalog. A graceful Ctrl+C on another instance removes it immediately.
Interview-style questions
- Why must the service instance ID be unique, and what breaks if it isn’t? Registration is keyed by ID. If two instances share an ID, the second overwrites the first, so one instance is invisible and its capacity is wasted.
- What happens if a self-registered service is killed with SIGKILL? It cannot deregister. The registry relies on health checks or TTL expiry to mark it critical, and on a deregister-after-critical timeout to remove it. Callers must also use retries and circuit breakers for the window in between.
- When would you not use Self-Registration? On Kubernetes, Azure Container Apps, or other platforms that register instances automatically, or in polyglot fleets where maintaining registration libraries is impractical. Use 3rd Party Registration or platform-native discovery.
Mastery checklist
- I can explain the register, heartbeat/health-check, and deregister lifecycle and what handles the crash case.
- I can implement registration as a .NET hosted service with unique IDs, an advertised routable address, and a readiness check.
- I can explain why registering
localhostor reusing the service name as ID are bugs. - I can distinguish Self-Registration from 3rd Party Registration and from Client-Side and Server-Side Discovery.
- I can decide, for a given hosting platform (VMs, AKS, Container Apps), whether the app should self-register at all.
- I can describe how to secure registry access (ACLs, TLS, no secrets in metadata).
- I can test failure paths: graceful stop, hard kill, and registry outage at startup.
- I can write an ADR that names the trade-offs and the trigger to retire the pattern.
Key takeaway
In Self-Registration, each instance announces itself to the registry on start and withdraws on stop, which makes discovery easy on unmanaged hosts but pushes registry code and crash-cleanup risk into every service. On orchestrated platforms such as Kubernetes and Azure Container Apps, let the platform do it instead.
