Manikandan — Manikandan
Microservices

Day 36: Health Check API

ManikandanManikandan
19 min read·Updated Sep 14, 2022

A Health Check API is a small set of HTTP endpoints that every service exposes so that machines (Kubernetes, Azure Container Apps, App Service, load balancers, monitoring) can ask "are you alive?", "are you ready to...

Intro

A Health Check API is a small set of HTTP endpoints that every service exposes so that machines (Kubernetes, Azure Container Apps, App Service, load balancers, monitoring) can ask “are you alive?”, “are you ready to take traffic?” and “have you finished starting?”. Instead of guessing from the outside whether a container is deadlocked, still warming up, or cannot reach its database, the service reports its own state and the platform reacts: restart it, stop routing traffic to it, or wait for it. In our insurance claims system, this is what stops the Claims service from receiving customer claim submissions while its SQL connection pool is exhausted, and what restarts a Policy service pod whose thread pool is deadlocked.

Versions used in this lesson: .NET 10 (current LTS, Microsoft.Extensions.Diagnostics.HealthChecks in the shared framework) and Angular 21/22 style code (standalone components, provideHttpClient, signals). The .NET health check APIs shown here have been stable since ASP.NET Core 2.2 and are unchanged in .NET 10.

Why we need this

Business reasons. A claims customer uploading photos of a damaged car at 11pm does not care that one of six pods is broken; they care that their request succeeds. Health checks let the platform hide individual instance failures from users, which directly protects availability targets (SLOs) and support-ticket volume.

Technical reasons.

  • A process can be running (the OS says PID exists) yet be useless: deadlocked, out of threads, out of DB connections, stuck in an infinite loop, or missing a required secret. “Container is up” is not the same as “service works”.
  • Startup takes time (EF Core migrations, cache warm-up, loading reference data). Without a readiness signal, orchestrators send traffic to instances that would return 500s.
  • Rolling deployments need a definition of “the new version is good” before the old version is removed. That definition is the readiness probe.
  • Dependencies fail independently. You need a place to encode “I can serve requests only if SQL Server and Service Bus are reachable”.
  • Operations teams need a cheap, uniform, scriptable way to check any service without knowing its internals.

What problem it solves

The problem (from the topic list): orchestrators cannot tell whether a container is deadlocked or ready.

What goes wrong without it:

  • Kubernetes/Container Apps only knows the process is alive, so a deadlocked Claims API keeps receiving traffic; every request that lands on it times out after 30 seconds while the pod shows Running.
  • During a rollout, the new Claims API pod is added to the load balancer the instant the container starts, while it is still running migrations. Users get 502/503 for 20-40 seconds on every deploy.
  • A pod that lost its database connection keeps accepting requests and failing them, instead of being taken out of rotation.
  • The opposite failure: an over-eager health check that depends on every downstream service makes all pods report unhealthy when a non-critical dependency (say, the SMS gateway) is down, and the orchestrator restarts healthy pods in a loop, turning a small incident into a full outage.

When it is needed (and when it is NOT)

Needed when:

  • The service runs behind a load balancer, gateway, or orchestrator with more than one instance.
  • You do rolling, blue/green, or canary deployments and want automatic “is the new version healthy” gating.
  • The service has a meaningful startup phase or depends on external resources (DB, broker, cache, secrets).
  • You want external uptime monitoring or auto-healing (App Service Health Check, Kubernetes restarts).

Not needed, or the wrong tool, when:

  • A short-lived batch job or one-off console tool that runs to completion (use exit codes and job success/failure instead). A long-running worker service still benefits from a probe.
  • A single-instance internal tool where a human notices failure anyway. A trivial /healthz is still cheap, but do not invest in dependency checks.
  • You are tempted to use the health endpoint as a business dashboard or metrics source. Health answers a yes/no question; latency, error rates, and throughput belong in metrics (Day 33).
  • You are tempted to put deep end-to-end business tests in it (for example, “submit a fake claim”). That belongs in synthetic monitoring, not in a probe hit every few seconds.

How to identify the problem (key signals)

  1. Users report intermittent failures that stop when they retry, and dashboards show only one instance producing 5xx or timeouts.
  2. Every deployment produces a short burst of 502/503/504 errors right after new pods start.
  3. kubectl get pods (or the Container Apps replica list) shows pods Running/1/1 Ready while their logs show repeated database or broker connection failures.
  4. A pod has to be restarted manually to recover from hangs; the incident runbook literally says “restart the pod”.
  5. Load balancer metrics show uneven latency: one backend has p99 of 30s while the others have p99 of 200ms.
  6. Restart loops (CrashLoopBackOff, high RestartCount) that correlate with a downstream outage, a sign your liveness check is too deep.
  7. No /health-style endpoint exists, or every team invented a different one (/ping, /status, /api/health, returning “OK” without checking anything).

Flow Diagram

Liveness, readiness and startup probes drive restart and traffic decisions.

flowchart TD
S["startupProbe /health/live"] -->|passes| L["livenessProbe /health/live"]
S -->|never passes| RS1["Restart"]
L -->|fails x3| RS2["Restart container"]
L -->|ok| R["readinessProbe /health/ready"]
R -->|DB reachable| IN["In load balancer rotation"]
R -->|DB down| OUT["Removed from rotation, not restarted"]

Level 1: Beginner

Concept and analogy. Think of a restaurant kitchen. The manager sometimes shouts “Are you alive?” (liveness: is the kitchen on fire or frozen?) and sometimes “Are you ready for orders?” (readiness: is the stove hot, is the cook here?). A kitchen that is still setting up is alive but not ready. A kitchen where the cook has fainted is not alive and needs to be replaced.

Three questions, three probe types:

ProbeQuestionFailure meansPlatform action
StartupHas the app finished booting?Still starting or stuck startingWait; restart if it never finishes
LivenessIs the process able to make progress at all?Deadlocked or broken beyond self-repairRestart the instance
ReadinessCan it serve traffic right now?Dependency down, warming up, overloadedStop sending traffic; do NOT restart

Minimal working example (.NET 10 minimal API).

var builder = WebApplication.CreateBuilder(args);
builder.Services.AddHealthChecks(); // registers the health check service, no checks yet
var app = builder.Build();
app.MapHealthChecks("/health"); // returns 200 "Healthy" or 503 "Unhealthy"
app.MapGet("/claims/ping", () => "Claims API is running");
app.Run();

Run it and call GET /health: you get HTTP 200 with body Healthy. With no checks registered, this only proves the process can serve HTTP, which is a valid (if shallow) liveness signal.

Level 2: Intermediate

In a real .NET + Angular + SQL Server application you separate liveness (cheap, no dependencies) from readiness (checks the dependencies this instance needs), using check tags.

Step 1: Register checks with tags (Claims API, .NET 10, EF Core + SQL Server).

Packages: Microsoft.Extensions.Diagnostics.HealthChecks.EntityFrameworkCore (for AddDbContextCheck). For Service Bus, AspNetCore.HealthChecks.AzureServiceBus from the community Xabaril project is common.

using Microsoft.AspNetCore.Diagnostics.HealthChecks;
using Microsoft.EntityFrameworkCore;
using Microsoft.Extensions.Diagnostics.HealthChecks;
using System.Text.Json;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddDbContext<ClaimsDbContext>(o =>
o.UseSqlServer(builder.Configuration.GetConnectionString("ClaimsDb")));
builder.Services.AddHealthChecks()
// Liveness: no external dependencies, always cheap
.AddCheck("self", () => HealthCheckResult.Healthy(), tags: new[] { "live" })
// Readiness: can this instance actually serve claim requests?
.AddDbContextCheck<ClaimsDbContext>("claims-db", tags: new[] { "ready" })
.AddCheck<PolicyServiceCheck>("policy-service",
failureStatus: HealthStatus.Degraded, // non-critical: degrade, do not fail
tags: new[] { "ready" });
builder.Services.AddHttpClient("policy", c =>
c.BaseAddress = new Uri(builder.Configuration["Services:Policy"]!));
var app = builder.Build();
// Liveness: only checks tagged "live"
app.MapHealthChecks("/health/live", new HealthCheckOptions
{
Predicate = r => r.Tags.Contains("live")
});
// Readiness: checks tagged "ready"; Degraded still returns 200 by default
app.MapHealthChecks("/health/ready", new HealthCheckOptions
{
Predicate = r => r.Tags.Contains("ready"),
ResponseWriter = WriteJson
});
app.Run();
static Task WriteJson(HttpContext ctx, HealthReport report)
{
ctx.Response.ContentType = "application/json";
var body = new
{
status = report.Status.ToString(),
totalMs = report.TotalDuration.TotalMilliseconds,
checks = report.Entries.Select(e => new
{
name = e.Key,
status = e.Value.Status.ToString(),
ms = e.Value.Duration.TotalMilliseconds
// Do NOT include e.Value.Exception or connection strings here.
})
};
return ctx.Response.WriteAsync(JsonSerializer.Serialize(body));
}
public class ClaimsDbContext(DbContextOptions<ClaimsDbContext> options) : DbContext(options) { }
public class PolicyServiceCheck(IHttpClientFactory factory) : IHealthCheck
{
public async Task<HealthCheckResult> CheckHealthAsync(
HealthCheckContext context, CancellationToken ct = default)
{
try
{
var client = factory.CreateClient("policy");
using var cts = CancellationTokenSource.CreateLinkedTokenSource(ct);
cts.CancelAfter(TimeSpan.FromSeconds(2)); // a probe must never hang
var resp = await client.GetAsync("/health/live", cts.Token);
return resp.IsSuccessStatusCode
? HealthCheckResult.Healthy()
: HealthCheckResult.Degraded($"Policy returned {(int)resp.StatusCode}");
}
catch (Exception ex)
{
return new HealthCheckResult(context.Registration.FailureStatus,
"Policy service unreachable", ex);
}
}
}

Key points:

  • HealthStatus.Healthy and Degraded map to HTTP 200, Unhealthy maps to 503 by default (ResultStatusCodes can change this).
  • AddDbContextCheck runs CanConnectAsync by default. It proves connectivity, not that migrations are applied; pass a custom query via the customTestQuery parameter if you need more.
  • The Policy check is registered as Degraded on failure because the Claims API can still accept a claim without a live policy lookup if it has a fallback (Day 29).

Step 2: Protect the endpoints and hold back detail. Internal probes are usually called by the platform on the container network. If you expose them publicly through a gateway, expose only /health/live or nothing, and never return exception messages.

Step 3: Angular consumption. Angular should not call /health/* for its own business logic. Two legitimate uses are an internal ops page and an app-level “service degraded” banner fed by your BFF/gateway.

// ops-status.service.ts (Angular 21/22, standalone, signals)
import { Injectable, inject, signal } from '@angular/core';
import { HttpClient } from '@angular/common/http';
import { timer, switchMap, catchError, of } from 'rxjs';
import { takeUntilDestroyed } from '@angular/core/rxjs-interop';
export interface HealthReport {
status: 'Healthy' | 'Degraded' | 'Unhealthy';
checks: { name: string; status: string; ms: number }[];
}
@Injectable({ providedIn: 'root' })
export class OpsStatusService {
private http = inject(HttpClient);
readonly report = signal<HealthReport | null>(null);
constructor() {
timer(0, 30_000)
.pipe(
switchMap(() =>
this.http.get<HealthReport>('/api/claims/health/ready').pipe(
catchError(() => of({ status: 'Unhealthy', checks: [] } as HealthReport))
)
),
takeUntilDestroyed()
)
.subscribe(r => this.report.set(r));
}
}

Note that the 503 response body still comes back as an HttpErrorResponse; the catchError above collapses it to “Unhealthy” for simplicity.

Step 4: Kubernetes wiring (works on AKS, see section 9).

containers:
- name: claims-api
image: acrclaims.azurecr.io/claims-api:1.4.2
ports: [{ containerPort: 8080 }]
startupProbe:
httpGet: { path: /health/live, port: 8080 }
periodSeconds: 3
failureThreshold: 30 # up to 90s to boot
livenessProbe:
httpGet: { path: /health/live, port: 8080 }
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
readinessProbe:
httpGet: { path: /health/ready, port: 8080 }
periodSeconds: 5
timeoutSeconds: 3
failureThreshold: 2

Note: .NET 8 and later container images listen on port 8080 by default (non-root), not 80, which is why the port is 8080 here.

Level 3: Advanced

Performance.

  • Probes run continuously on every instance. If a readiness check runs a heavy SQL query every 5 seconds across 20 pods, you have built a self-inflicted load test. Keep checks cheap (SELECT 1-level) and set a timeout on every dependency call.
  • Use HealthCheckPublisher or caching to decouple probe frequency from dependency load: run checks on a timer (for example every 10 seconds) and let the endpoint return the last cached result.
builder.Services.Configure<HealthCheckPublisherOptions>(o =>
{
o.Delay = TimeSpan.FromSeconds(5);
o.Period = TimeSpan.FromSeconds(15);
o.Predicate = r => r.Tags.Contains("ready");
});
builder.Services.AddSingleton<IHealthCheckPublisher, CachedReadinessPublisher>();
// CachedReadinessPublisher stores the latest HealthReport (pseudo-code: keep it in a
// static/volatile field) and a minimal endpoint returns 200/503 from that cache.
  • Health check registrations also accept a per-check timeout argument; use it so one slow dependency cannot make the whole probe exceed the platform’s timeoutSeconds.

Scalability. Readiness is your back-pressure valve. A service can deliberately report Unhealthy on readiness when overloaded (for example, when the request queue is full) so the load balancer sheds traffic away and lets it recover. Be careful: if all instances do this simultaneously you remove all capacity; prefer rate limiting (Day 30) for overload and reserve readiness for “cannot serve correctly”.

Security.

  • Do not leak internals: no stack traces, connection strings, server names, or version numbers in the public response body. The default Healthy/Unhealthy plain text is safe; a detailed JSON writer should sit on an internal-only route.
  • Restrict access: use .RequireHost("*:8080") with a separate internal port, network policies, or .RequireAuthorization("ops") for the detailed endpoint. Public API gateways should not route /health/ready to the internet.
  • Health endpoints bypass your normal auth. Make sure they do not perform side effects (no writes, no message publishing) and cannot be abused for amplification (each hit triggers downstream calls).

Failure modes and common mistakes.

MistakeConsequenceFix
Liveness probe checks the databaseDB outage causes every pod to be restarted in a loop, prolonging recoveryLiveness must be dependency-free; DB belongs in readiness
Readiness includes every downstream serviceOne non-critical outage removes all pods from rotation, total outageOnly include critical dependencies; mark others Degraded
No startupProbe and short initialDelaySecondsSlow-starting pod is killed before it finishes booting, then restarts foreverUse startup probe with generous failureThreshold * periodSeconds
Probe timeout shorter than a GC pause or cold startRandom restarts under loadtimeoutSeconds of 2-5s, failureThreshold of 3 for liveness
Health check swallows exceptions and always returns HealthyEndpoint is decoration, not signalTest by breaking the dependency in a staging environment
Health check opens a new connection per call without disposingConnection pool exhaustion caused by the probe itselfReuse pooled connections/DbContext from DI and dispose correctly
Returning 200 with "status":"Unhealthy" in the bodyPlatforms judge by HTTP status code, not bodyEnsure mapping to 503, or use the built-in middleware
Same endpoint for liveness and readinessCannot restart without also draining, or vice versaSeparate endpoints and tags
Health status cached foreverStale “healthy” after failureCache for seconds, not minutes

Level 4: Expert and Architect view

Design trade-offs and alternatives.

ApproachStrengthsWeaknessesUse when
Shallow liveness + dependency-aware readiness (recommended default)Avoids restart storms; drains traffic on dependency lossNeeds discipline on what counts as criticalAlmost every HTTP service
Single deep /health endpointSimple to buildConflates restart and drain; restart stormsVery small internal tools only
TCP-only probeZero code, works for any protocolProves the port is open, not that the app worksLegacy or non-HTTP workloads
gRPC health checking protocol (grpc.health.v1)Standard for gRPC, native Kubernetes gRPC probesExtra package and wiringgRPC services
Push-based heartbeat (service reports to a monitor)Works for workers with no inbound portMonitor must detect missing heartbeatsBackground workers, scheduled jobs
Synthetic transactions (external monitoring)Tests the real user path end to endSlow, costly, not for orchestrator probesComplement, not replacement
Health via metrics only (Prometheus alerts)Rich signal, trendsOrchestrator cannot act on it directlyAlerting alongside probes

Patterns it combines with.

  • Circuit Breaker (Day 26) and Fallback (Day 29): a service whose circuit to Policy is open can report Degraded rather than Unhealthy.
  • Service Registry / Discovery (Day 21-25): registries use health status to add and remove instances; Kubernetes readiness does this natively for Services.
  • Application Metrics (Day 33): export health check status as a metric (aspnetcore_healthcheck_status) so you can alert on flapping, not just current state.
  • Log Deployments & Changes (Day 37): readiness gating during rollout plus deploy markers makes “did this release break things” answerable.
  • Sidecar / Service Mesh (Day 47-48): meshes have their own outlier detection; be aware that probes may need to bypass mTLS.
  • Externalized Configuration (Day 39): a readiness check that verifies required config/secrets loaded catches bad deployments early.

ADR-style justification.

ADR-036: Standard health endpoints for all Claims platform services

Status: Accepted

Context: We run 14 .NET services on Azure Container Apps and AKS. Deployments cause 20-40s of 5xx errors because new replicas receive traffic before migrations and warm-up complete. Two incidents in the last quarter required manual pod restarts for deadlocked instances, and each team has its own /status implementation.

Decision: Every service will expose /health/live (dependency-free, tag live) and /health/ready (critical dependencies only, tag ready), built with Microsoft.Extensions.Diagnostics.HealthChecks. Platform probes: startup and liveness hit /health/live; readiness hits /health/ready. Non-critical dependencies report Degraded, not Unhealthy. Detailed JSON is available only on an internal port. The shared service chassis (Day 40) will supply the registration so teams do not reimplement it.

Consequences: (+) Zero-downtime rollouts, automatic recovery from hangs, uniform ops experience. (-) Teams must classify dependencies as critical or not; misclassification can still cause outages. (-) Probe traffic adds small constant load; mitigated by cached checks. Review after two quarters using restart counts and deploy error rates.

Alternatives rejected: single deep endpoint (restart storms); TCP probes only (no application-level signal); relying on external uptime monitoring alone (cannot drain or restart instances).

Azure implementation

Which Azure services use or implement this topic.

  • Azure Kubernetes Service (AKS): native startupProbe, livenessProbe, readinessProbe (HTTP, TCP, exec, gRPC). Standard Kubernetes behaviour.
  • Azure Container Apps: supports startup, liveness, and readiness probes (TCP or HTTP/S) configured per container. If you configure none, default probes are applied (TCP-based); check the current defaults in the Container Apps health probes documentation for the exact values, because they differ per probe type.
  • Azure App Service: the Health Check feature pings a path you configure (about once a minute); after a configurable number of consecutive failures the instance is removed from load-balancer rotation, and if it stays unhealthy it can be replaced. It only acts when the App Service plan runs two or more instances. Configure the path in the portal under Monitoring > Health check, or by IaC (healthCheckPath site property) and set the WEBSITE_HEALTHCHECK_MAXPINGFAILURES app setting to tune the failure count.
  • Azure Application Gateway and Azure Front Door: custom health probes decide which backends receive traffic. Point them at a lightweight endpoint. Front Door probes come from many edge locations, so a heavy probe endpoint multiplies load.
  • Azure Load Balancer: health probes (TCP/HTTP/HTTPS) for VM or VMSS backends.
  • Azure Monitor / Application Insights: availability tests (Standard tests) hit a public URL from multiple regions and alert on failure; this is external, user-perspective monitoring. Use an endpoint that is safe to expose publicly. Verify the current availability-test types, since the older URL ping tests are being retired in favour of Standard tests.
  • Azure Functions: no separate liveness/readiness concept; App Service Health Check applies on some hosting plans, and Application Insights covers monitoring.

How to configure (Container Apps, Bicep excerpt).

resource claimsApi 'Microsoft.App/containerApps@2024-03-01' = {
name: 'ca-claims-api'
location: location
properties: {
managedEnvironmentId: env.id
configuration: {
ingress: { external: false, targetPort: 8080 }
}
template: {
containers: [
{
name: 'claims-api'
image: '${acr.properties.loginServer}/claims-api:1.4.2'
probes: [
{
type: 'Startup'
httpGet: { path: '/health/live', port: 8080 }
periodSeconds: 3
failureThreshold: 30
}
{
type: 'Liveness'
httpGet: { path: '/health/live', port: 8080 }
periodSeconds: 10
failureThreshold: 3
timeoutSeconds: 2
}
{
type: 'Readiness'
httpGet: { path: '/health/ready', port: 8080 }
periodSeconds: 5
failureThreshold: 3
timeoutSeconds: 3
}
]
}
]
scale: { minReplicas: 2, maxReplicas: 10 }
}
}
}

The apiVersion above is illustrative; use the current stable API version from the Azure resource reference when you deploy.

App Service via Azure CLI:

Terminal window
az webapp config set -g rg-claims -n app-claims-api --generic-configurations '{"healthCheckPath": "/health/ready"}'
az webapp config appsettings set -g rg-claims -n app-claims-api --settings WEBSITE_HEALTHCHECK_MAXPINGFAILURES=5

For App Service, pointing at the readiness endpoint is a reasonable choice; unlike Kubernetes there is no separate liveness concept, and a failing instance is removed from rotation.

Pricing and tier considerations.

  • The probes themselves (AKS, Container Apps, App Service Health Check, Application Gateway/Front Door/Load Balancer probes) carry no separate charge; you pay for the underlying compute and networking.
  • App Service Health Check needs at least two instances, so it implies a plan tier and instance count that supports scale-out (Basic and above for manual scale-out). Single-instance Free/Shared plans cannot benefit.
  • Application Insights availability tests are billed per test execution/per test per month depending on the type; confirm current rates on the Azure Monitor pricing page before creating many tests across regions.
  • Front Door and Application Gateway have tiered pricing (Standard/Premium; Application Gateway v2 SKUs); health probing is included, but the gateway SKU choice is the cost driver, not the probe.
  • Verify all SKUs and rates against the pricing pages at decision time; Azure tiers and names change.

Reference architecture (text).

Customers use the Angular app served from Azure Static Web Apps or Blob storage behind Azure Front Door. Front Door routes /api/* to an internal Application Gateway or directly to API Management, which forwards to the Claims, Policy, and Payments services running in an Azure Container Apps environment (or AKS). Each service exposes /health/live and /health/ready. Container Apps/AKS use startup, liveness, and readiness probes to gate rollouts and restart deadlocked replicas. Front Door and the gateway probe a shallow public-safe endpoint to detect regional or backend loss and fail over. Application Insights receives request telemetry and availability test results, and an alert rule fires when readiness failures or restart counts exceed a threshold, notifying the on-call team via an Action Group. Detailed health JSON stays on an internal port reachable only inside the virtual network.

Teaching guide for my team

Explain to a beginner in 2 minutes. “Your service needs a way to answer three questions when the platform asks. Am I still running properly (liveness)? Am I ready to take customer requests (readiness)? Have I finished starting up (startup)? We add two URLs, /health/live and /health/ready. If live fails, the platform restarts us. If ready fails, the platform stops sending customers to us until we recover. Liveness never checks the database. Readiness does.” Draw a kitchen: on fire versus not ready for orders.

Explain to an intermediate developer in 5 minutes. Walk through: tags and predicates in AddHealthChecks/MapHealthChecks; the Healthy/Degraded/Unhealthy mapping to HTTP 200/200/503; why the DB goes in readiness only; what a restart storm is and how a deep liveness probe causes one; startup probes for slow boots; timeouts on every check; keeping detail off public routes. Show the AKS/Container Apps probe YAML next to the code so they see the two halves of the contract. Finish with the Degraded versus Unhealthy decision for the Policy dependency.

Hands-on exercise.

  1. Create a .NET 10 minimal API “Claims.Api” with /health/live and /health/ready as in section 6, with a SQL Server container as the database (docker compose).
  2. Run both in Docker Compose or a local Kubernetes (kind/minikube) with the probes from section 6.
  3. Stop the SQL container.

Expected outcome: /health/live stays 200 and the pod is not restarted; /health/ready returns 503, the pod shows 0/1 Ready, and the Service has no endpoints (requests get connection errors from the load balancer). Restart SQL: readiness recovers within a few probe periods and traffic resumes. Bonus: temporarily add the DB check to liveness and observe the restart loop, then discuss why that is worse.

Interview-style questions.

  1. What is the difference between liveness and readiness? Liveness asks whether the process should be restarted; readiness asks whether it should receive traffic right now. Failing liveness restarts the instance; failing readiness only removes it from load balancing.
  2. Why should a liveness probe not check the database? If the DB is down, every instance fails liveness and is restarted repeatedly, adding load and prolonging recovery, while a restart cannot fix the DB. Dependencies belong in readiness.
  3. What is a startup probe for? It gives slow-starting apps time to boot (migrations, warm-up) before liveness and readiness begin, so the platform does not kill an instance that is merely still starting.

Mastery checklist

  • I can explain liveness, readiness, and startup with one example each from the claims system.
  • I can implement tagged health checks in .NET 10 and map them to separate endpoints.
  • I can predict what happens (restart, drain, or nothing) for each failure: deadlock, DB down, SMS gateway down, slow boot.
  • I can configure probes correctly on AKS and Azure Container Apps, including timeouts and thresholds.
  • I can explain a restart storm and show how to avoid it.
  • I can secure health endpoints: no sensitive detail, restricted routes, no side effects.
  • I can decide Degraded versus Unhealthy for a given dependency and justify it.
  • I can write a short ADR standardizing health endpoints across services.

Key takeaway

Liveness answers “restart me?” and must never depend on other systems; readiness answers “send me traffic?” and should check only what this instance truly needs. Get that split right and the platform heals your services instead of amplifying their failures.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 37: Log Deployments & Changes

Log Deployments & Changes means every release, configuration change, feature-flag flip, and infrastructure change is recorded as a timestamped event and drawn as a marker on the same dashboards where you watch errors...

Manikandan
Manikandan·24 min read
Microservices

Day 35: Exception Tracking

Exception Tracking means capturing every unhandled (and important handled) error from your services and your browser app, attaching context to it (release, user, request, breadcrumbs), grouping identical errors into...

Manikandan
Manikandan·18 min read
Microservices

Day 34: Audit Logging

Audit Logging is the practice of writing a structured, tamper-resistant record of *who* did *what*, *to which thing*, *when*, *from where*, and *with what result* every time a security-relevant or business-relevant...

Manikandan
Manikandan·24 min read
Microservices

Day 33: Application Metrics

Application Metrics means every service continuously exposes small, cheap, numeric measurements (request latency, error counts, throughput, queue depth, memory, business counters like "claims submitted") to a...

Manikandan
Manikandan·18 min read