3rd Party Registration is a service-discovery pattern where the service itself never talks to the service registry.
Intro
3rd Party Registration is a service-discovery pattern where the service itself never talks to the service registry. A separate component, the registrar (a Kubernetes controller, a platform agent, or a sidecar), watches what is running, registers each healthy instance in the registry, and removes it when it stops or fails. The application code stays free of registry client libraries. In Kubernetes and Azure Container Apps this is the default behaviour, so most teams use it without naming it.
Why we need this
In Day 24 (Self-Registration) every service embeds a registry client: it registers on startup, sends heartbeats, and deregisters on shutdown. That works, but it has costs:
- Dependency debt. Every service, in every language, must carry a registry client library (Consul, Eureka, etc.) and keep it upgraded.
- Lifecycle coupling. If the process is killed with
SIGKILL, OOM-killed, or the node dies, it never deregisters. The registry keeps a dead entry until a TTL expires. - Business code knows about infrastructure. Startup code decides when the service is “ready”, mixes it with registry calls, and has to handle registry outages.
- Inconsistent behaviour. Team A registers before warm-up finishes, team B after; team C forgets heartbeats.
3rd Party Registration moves this responsibility to the platform. The platform already knows when a container starts, passes its health probes, is scaled, or is terminated, so it is the best-placed component to keep the registry correct.
What problem it solves
Problem from the topic list: registry client libraries in business code add dependency debt.
Example in our insurance claims system: ClaimsService (.NET), PolicyService (.NET), FraudScoringService (Python), and DocumentOcrService (Node.js) all need to be discoverable. With self-registration, four teams write four registration clients. During a rolling deployment, FraudScoringService is OOM-killed under load and never deregisters. ClaimsService keeps routing to the dead IP for the TTL window (say 30 seconds) and claim submissions time out.
Without the pattern you get: stale registry entries, registry client upgrades that require redeploying every service, services that cannot start when the registry is briefly unavailable, and a different discovery story per language.
With it, an external registrar owns the registry entries. Instances appear only when they are actually ready, and disappear when the platform removes them, regardless of how the process died.
When it is needed (and when it is NOT)
Good fit
- You run on an orchestrator (Kubernetes/AKS, Azure Container Apps, Nomad, ECS) that already tracks instance lifecycle.
- You have polyglot services and do not want a discovery library per language.
- You want registration behaviour to be uniform and centrally governed.
- Instances are ephemeral and scale frequently.
Poor fit or overkill
- A small system of two or three services on a couple of App Service instances with a fixed URL each. A DNS name is enough; a registry is overkill.
- Bare VMs with no orchestrator or agent. Then there is no third party to do the registering, and you would have to build one (or use Consul agents / self-registration).
- You need registration metadata that only the application knows at runtime (for example, dynamic capability tags computed after startup). A registrar can only see what the platform exposes (labels, annotations, ports), so self-registration may fit better.
- The registrar becomes a single point of failure you do not operate well. It must be highly available.
How to identify the problem (key signals)
- Each service repository has a different registry client package, and Dependabot PRs for it appear in every repo.
- Callers get connection refused or timeouts to IPs that no longer exist, especially minutes after a crash or node loss.
- Registry has entries with old heartbeat timestamps, or
critical/unknown health states nobody cleans up. - Service startup code contains retry loops around registry calls, and services fail to start when the registry is down.
- Non-.NET teams say “there is no good client library for our language”.
- Deployments show a window of failed requests because a terminating instance is still listed as healthy.
- Code reviews repeatedly contain “did you remember to deregister on shutdown?”
Flow Diagram
The platform, not the app, keeps the registry accurate.
flowchart LR POD["Pod starts"] --> SP["startupProbe passes"] SP --> RP{"readinessProbe passes?"} RP -- "yes" --> EC["Endpoint controller adds IP to EndpointSlice"] RP -- "no" --> OUT["Kept out of rotation"] EC --> DNS["CoreDNS: claims-service"] DNS --> CALLER["Callers"] TERM["Pod terminating"] --> RM["Removed from endpoints + preStop drain"]Level 1: Beginner
Analogy. A hotel guest does not write their own name on the front-desk list. The receptionist (registrar) checks them in when they arrive with a working room key, and removes them at checkout, or when housekeeping finds the room empty. Guests just live in the room.
Minimal working example. A service that knows nothing about discovery, with a health endpoint. The registrar (Kubernetes) does the rest.
// ClaimsService/Program.cs (.NET 10 LTS, minimal API)var builder = WebApplication.CreateBuilder(args);builder.Services.AddHealthChecks();
var app = builder.Build();
app.MapHealthChecks("/health/ready");app.MapGet("/claims/{id:guid}", (Guid id) => Results.Ok(new { id, status = "Submitted" }));
app.Run();# claims-service.yaml - the Service selects pods; Kubernetes' endpoint controller# is the "3rd party" that registers ready pods and removes unready ones.apiVersion: apps/v1kind: Deploymentmetadata: name: claims-servicespec: replicas: 3 selector: matchLabels: { app: claims-service } template: metadata: labels: { app: claims-service } spec: containers: - name: api image: myregistry.azurecr.io/claims-service:1.0.0 ports: [{ containerPort: 8080 }] readinessProbe: httpGet: { path: /health/ready, port: 8080 } periodSeconds: 5---apiVersion: v1kind: Servicemetadata: name: claims-servicespec: selector: { app: claims-service } ports: [{ port: 80, targetPort: 8080 }]Other services call http://claims-service/claims/.... There is no registry code in ClaimsService. Kubernetes maintains EndpointSlice objects listing the IPs of pods that pass the readiness probe, and cluster DNS resolves claims-service to the Service.
Level 2: Intermediate
Real .NET + Angular + database flow. Angular calls an API gateway; the gateway (YARP) calls ClaimsService; ClaimsService uses SQL Server or PostgreSQL. The registration is external, so the .NET code uses only a normal named HttpClient.
// ApiGateway/Program.cs - call by stable Kubernetes DNS name, no registry client.builder.Services.AddHttpClient("claims", c =>{ c.BaseAddress = new Uri("http://claims-service"); // resolved by cluster DNS c.Timeout = TimeSpan.FromSeconds(5);});
// YARP config can also point at the same DNS name.// appsettings.json (gateway){ "ReverseProxy": { "Routes": { "claims": { "ClusterId": "claims", "Match": { "Path": "/api/claims/{**rest}" } } }, "Clusters": { "claims": { "Destinations": { "d1": { "Address": "http://claims-service/" } } } } }}Readiness should reflect real dependencies, otherwise the registrar will register instances that cannot serve. Use liveness and readiness separately:
// ClaimsService: readiness checks the database; liveness does not (avoid restart storms).builder.Services.AddHealthChecks() .AddCheck("self", () => HealthCheckResult.Healthy(), tags: new[] { "live" }) .AddNpgSql(builder.Configuration.GetConnectionString("Claims")!, tags: new[] { "ready" }); // package: AspNetCore.HealthChecks.NpgSql
app.MapHealthChecks("/health/live", new HealthCheckOptions { Predicate = r => r.Tags.Contains("live") });app.MapHealthChecks("/health/ready", new HealthCheckOptions { Predicate = r => r.Tags.Contains("ready") });Angular is unaffected. It calls /api/claims on the gateway. The key point for the team: discovery is a platform concern, and the frontend and business code stay unchanged.
Graceful shutdown. The registrar removes a terminating pod from endpoints in parallel with sending SIGTERM, so there is a short race. Handle it in code and config:
lifecycle: preStop: exec: { command: ["sleep", "10"] } # let endpoint removal propagate before shutdownterminationGracePeriodSeconds: 45Level 3: Advanced
Performance and scalability
- The registrar works on watch streams (Kubernetes EndpointSlices), not polling, so it scales to large clusters. Endpoint updates propagate to kube-proxy or the dataplane on every node, so very high pod churn causes control-plane load.
- DNS caching matters. A client that caches DNS forever (or a long-lived
HttpClientconnection pool) will keep using old IPs. SetPooledConnectionLifetimeonSocketsHttpHandler.
builder.Services.AddHttpClient("claims") .ConfigurePrimaryHttpMessageHandler(() => new SocketsHttpHandler { PooledConnectionLifetime = TimeSpan.FromMinutes(2) // re-resolve DNS periodically });Security
- Restrict who can write to the registry. Only the registrar’s identity should have write access; services get read-only access.
- Registry entries can be spoofed if any workload can create Services or labels. Use RBAC and admission policies, and NetworkPolicy to restrict who can call whom.
- Registration does not equal authorization. Combine with mTLS or token validation.
Failure modes
- Registrar down. Existing entries stay (stale but usually valid); new instances are not registered and dead ones not removed. Run it highly available (in Kubernetes this is the managed control plane).
- Bad readiness probe. Always-true probe registers broken instances; a probe that fails on a shared dependency outage removes every instance at once, producing a total outage instead of degraded service.
- Race on termination (see preStop above).
- Split brain between registries in multi-cluster setups.
Common mistakes
- Using the liveness probe for registration decisions. Readiness governs registration; liveness governs restarts.
- Forgetting
startupProbefor slow-starting services, so they get killed before being registered. - Mixing patterns: leaving a self-registration client in the app while the platform also registers it, causing duplicate or conflicting entries.
Level 4: Expert and Architect view
| Aspect | 3rd Party Registration | Self-Registration | Platform-native DNS only (no separate registry) | Service mesh discovery |
|---|---|---|---|---|
| Who registers | Platform agent or controller | The service | Platform (implicit) | Control plane |
| Client library in code | None | Required | None | None (sidecar) |
| Handles crashes | Yes, registrar observes state | Only after TTL | Yes | Yes |
| Language independence | High | Low to medium | High | High |
| Custom runtime metadata | Limited to labels/annotations | Full | Limited | Limited |
| Operational cost | Depends on the platform | Low infra, high code cost | Lowest | Highest |
| Typical tech | Kubernetes controllers, Consul K8s sync, Registrator | Eureka client, Consul SDK | Kubernetes Services + CoreDNS | Istio, Linkerd |
Combines with: Client-Side Discovery or Server-Side Discovery (registration is only half of discovery; something still has to query the registry), Health Check API (readiness drives registration), Sidecar and Service Mesh (the sidecar or mesh control plane can be the registrar), Service Registry (the store being written to), Externalized Configuration.
ADR-style justification
ADR-025: Use platform-managed (3rd Party) registration for service discovery Status: Proposed Context: The claims platform has .NET, Python, and Node.js services on AKS. Self-registration required a registry client per language and left stale entries after crashes. Decision: Services do not register themselves. Kubernetes (endpoint controller with readiness probes) registers instances; callers use cluster DNS names. Every service must expose
/health/liveand/health/ready. Consequences: (+) no registry SDK in application code; (+) accurate entries after crashes; (+) uniform behaviour. (-) registration metadata limited to what Kubernetes exposes; (-) readiness probe quality now determines availability; (-) tied to the orchestrator, so a move off Kubernetes needs a new registrar. Alternatives considered: Self-Registration with Consul SDK (rejected: dependency debt); Service Mesh (deferred: operational cost, revisit if mTLS and traffic shaping are needed).
Azure implementation
Services
- Azure Kubernetes Service (AKS). The Kubernetes endpoint controller is the registrar; CoreDNS provides discovery. Deployments and Services as shown above.
- Azure Container Apps. The platform registers replicas and gives every app an internal DNS name within the environment. You do not run a registry at all. Ingress can be internal-only.
- Azure Service Fabric (existing customers): its Naming Service is registered by the platform when services are deployed.
- Azure App Service / Azure Functions: platform-managed hostnames; instances are load-balanced by the front end, so there is nothing to register.
- Azure Application Gateway / Azure Front Door / API Management can sit at the edge, pointing at stable internal names.
- Consul on AKS (self-managed or HashiCorp-managed) with its Kubernetes sync is an option when you need a registry that spans clusters or VMs.
Configuration checklist
- Create the AKS cluster; deploy each service with a
Deployment,Service, and readiness, liveness, and startup probes. - Set
terminationGracePeriodSecondsand apreStopdelay. - Use Container Apps? Enable internal ingress and call
http://<app-name>inside the environment. - Enable Azure Monitor / Container insights and alert on pods that are NotReady and on endpoint counts dropping to zero.
- Apply Kubernetes NetworkPolicies and Azure RBAC so only the platform can modify Services.
Pricing and tiers (verify current numbers on the Azure pricing pages before budgeting)
- AKS has Free, Standard, and Premium cluster management tiers. Free has no financially backed uptime SLA and is suited to dev/test; Standard adds the financially backed SLA and larger control-plane limits; Premium adds long-term support options. You pay per cluster-hour for Standard and Premium, plus the VMs, disks, and networking of node pools.
- Container Apps is consumption-based (vCPU-seconds, GiB-seconds, and requests), with a free monthly grant, and also offers dedicated workload profiles.
- Registration itself has no separate charge in either service. The cost of the pattern is the health-probe traffic (negligible) and the platform tier you choose.
Reference architecture (text)
Angular SPA on Azure Static Web Apps → Azure Front Door / Application Gateway → YARP API gateway pod in AKS → internal Kubernetes Services (claims-service, policy-service, fraud-scoring-service) → Azure Database for PostgreSQL Flexible Server and Azure SQL. Kubernetes’ endpoint controller registers only pods passing readiness; CoreDNS resolves names; Azure Monitor and Application Insights (OpenTelemetry) collect health and traces; Key Vault holds secrets via the CSI driver; Azure Container Registry supplies images.
Teaching guide for my team
Beginner, 2 minutes. “A service needs to be findable by others. Instead of each service shouting its address to a phonebook, we let the platform do it. Kubernetes watches which containers are healthy and keeps the phonebook up to date. Your job is only to provide a health endpoint that tells the truth. Other services call http://claims-service and never care which machine it is on.”
Intermediate, 5 minutes. Walk through the flow: pod starts → startup probe passes → readiness probe passes → endpoint controller adds the pod to the EndpointSlice → kube-proxy/dataplane and DNS make it reachable. On termination: pod marked terminating → removed from endpoints → preStop gives time to drain → SIGTERM. Then compare with Self-Registration (Day 24): what does each pattern do when the process is kill -9ed? Discuss readiness vs liveness, DNS caching in HttpClient, and why a database check in readiness is a trade-off.
Hands-on exercise. Deploy claims-service with 3 replicas on a local cluster (kind, minikube) or AKS. Add a /health/ready endpoint that returns 503 when a file /tmp/unready exists. Run kubectl get endpointslices -l kubernetes.io/service-name=claims-service -w, then kubectl exec into one pod and create the file. Expected outcome: that pod’s IP disappears from the EndpointSlice within a few seconds while the pod keeps running; removing the file brings it back. Then kubectl delete pod and observe the entry removed without any code in the app.
Interview questions
- What is the difference between Self-Registration and 3rd Party Registration? In self-registration the service calls the registry itself; in 3rd party, an external registrar (platform/agent) does it, so the app has no registry code.
- Which probe should drive registration, and why? Readiness. It says “can serve traffic now”. Liveness triggers restarts and should not check external dependencies.
- What are the drawbacks of 3rd Party Registration? Coupling to the platform, limited metadata, dependency on the registrar’s availability, and probe quality directly affecting availability.
Mastery checklist
- Can explain how 3rd Party Registration differs from Self-Registration, and give one case where self-registration is still better.
- Can name the registrar in Kubernetes and in Azure Container Apps.
- Can write correct readiness, liveness, and startup probes for a .NET service and justify what each checks.
- Can demonstrate the endpoint list changing when a pod becomes unready or is deleted.
- Can explain the termination race and configure
preStopand grace period. - Can explain why
HttpClientneedsPooledConnectionLifetimewhen instances change. - Can list failure modes (registrar down, bad probes) and mitigations.
- Can write an ADR justifying the pattern for a team.
Key takeaway
Let the platform, not your application code, keep the registry accurate: expose honest readiness endpoints and let a third-party registrar add and remove instances.
