A Health Check API is a small set of HTTP endpoints that every service exposes so that machines (Kubernetes, Azure Container Apps, App Service, load balancers, monitoring) can ask "are you alive?", "are you ready to...
Intro
A Health Check API is a small set of HTTP endpoints that every service exposes so that machines (Kubernetes, Azure Container Apps, App Service, load balancers, monitoring) can ask “are you alive?”, “are you ready to take traffic?” and “have you finished starting?”. Instead of guessing from the outside whether a container is deadlocked, still warming up, or cannot reach its database, the service reports its own state and the platform reacts: restart it, stop routing traffic to it, or wait for it. In our insurance claims system, this is what stops the Claims service from receiving customer claim submissions while its SQL connection pool is exhausted, and what restarts a Policy service pod whose thread pool is deadlocked.
Versions used in this lesson: .NET 10 (current LTS, Microsoft.Extensions.Diagnostics.HealthChecks in the shared framework) and Angular 21/22 style code (standalone components, provideHttpClient, signals). The .NET health check APIs shown here have been stable since ASP.NET Core 2.2 and are unchanged in .NET 10.
Why we need this
Business reasons. A claims customer uploading photos of a damaged car at 11pm does not care that one of six pods is broken; they care that their request succeeds. Health checks let the platform hide individual instance failures from users, which directly protects availability targets (SLOs) and support-ticket volume.
Technical reasons.
- A process can be running (the OS says PID exists) yet be useless: deadlocked, out of threads, out of DB connections, stuck in an infinite loop, or missing a required secret. “Container is up” is not the same as “service works”.
- Startup takes time (EF Core migrations, cache warm-up, loading reference data). Without a readiness signal, orchestrators send traffic to instances that would return 500s.
- Rolling deployments need a definition of “the new version is good” before the old version is removed. That definition is the readiness probe.
- Dependencies fail independently. You need a place to encode “I can serve requests only if SQL Server and Service Bus are reachable”.
- Operations teams need a cheap, uniform, scriptable way to check any service without knowing its internals.
What problem it solves
The problem (from the topic list): orchestrators cannot tell whether a container is deadlocked or ready.
What goes wrong without it:
- Kubernetes/Container Apps only knows the process is alive, so a deadlocked Claims API keeps receiving traffic; every request that lands on it times out after 30 seconds while the pod shows
Running. - During a rollout, the new Claims API pod is added to the load balancer the instant the container starts, while it is still running migrations. Users get 502/503 for 20-40 seconds on every deploy.
- A pod that lost its database connection keeps accepting requests and failing them, instead of being taken out of rotation.
- The opposite failure: an over-eager health check that depends on every downstream service makes all pods report unhealthy when a non-critical dependency (say, the SMS gateway) is down, and the orchestrator restarts healthy pods in a loop, turning a small incident into a full outage.
When it is needed (and when it is NOT)
Needed when:
- The service runs behind a load balancer, gateway, or orchestrator with more than one instance.
- You do rolling, blue/green, or canary deployments and want automatic “is the new version healthy” gating.
- The service has a meaningful startup phase or depends on external resources (DB, broker, cache, secrets).
- You want external uptime monitoring or auto-healing (App Service Health Check, Kubernetes restarts).
Not needed, or the wrong tool, when:
- A short-lived batch job or one-off console tool that runs to completion (use exit codes and job success/failure instead). A long-running worker service still benefits from a probe.
- A single-instance internal tool where a human notices failure anyway. A trivial
/healthzis still cheap, but do not invest in dependency checks. - You are tempted to use the health endpoint as a business dashboard or metrics source. Health answers a yes/no question; latency, error rates, and throughput belong in metrics (Day 33).
- You are tempted to put deep end-to-end business tests in it (for example, “submit a fake claim”). That belongs in synthetic monitoring, not in a probe hit every few seconds.
How to identify the problem (key signals)
- Users report intermittent failures that stop when they retry, and dashboards show only one instance producing 5xx or timeouts.
- Every deployment produces a short burst of 502/503/504 errors right after new pods start.
kubectl get pods(or the Container Apps replica list) shows podsRunning/1/1 Readywhile their logs show repeated database or broker connection failures.- A pod has to be restarted manually to recover from hangs; the incident runbook literally says “restart the pod”.
- Load balancer metrics show uneven latency: one backend has p99 of 30s while the others have p99 of 200ms.
- Restart loops (
CrashLoopBackOff, highRestartCount) that correlate with a downstream outage, a sign your liveness check is too deep. - No
/health-style endpoint exists, or every team invented a different one (/ping,/status,/api/health, returning “OK” without checking anything).
Flow Diagram
Liveness, readiness and startup probes drive restart and traffic decisions.
flowchart TD S["startupProbe /health/live"] -->|passes| L["livenessProbe /health/live"] S -->|never passes| RS1["Restart"] L -->|fails x3| RS2["Restart container"] L -->|ok| R["readinessProbe /health/ready"] R -->|DB reachable| IN["In load balancer rotation"] R -->|DB down| OUT["Removed from rotation, not restarted"]Level 1: Beginner
Concept and analogy. Think of a restaurant kitchen. The manager sometimes shouts “Are you alive?” (liveness: is the kitchen on fire or frozen?) and sometimes “Are you ready for orders?” (readiness: is the stove hot, is the cook here?). A kitchen that is still setting up is alive but not ready. A kitchen where the cook has fainted is not alive and needs to be replaced.
Three questions, three probe types:
| Probe | Question | Failure means | Platform action |
|---|---|---|---|
| Startup | Has the app finished booting? | Still starting or stuck starting | Wait; restart if it never finishes |
| Liveness | Is the process able to make progress at all? | Deadlocked or broken beyond self-repair | Restart the instance |
| Readiness | Can it serve traffic right now? | Dependency down, warming up, overloaded | Stop sending traffic; do NOT restart |
Minimal working example (.NET 10 minimal API).
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddHealthChecks(); // registers the health check service, no checks yet
var app = builder.Build();
app.MapHealthChecks("/health"); // returns 200 "Healthy" or 503 "Unhealthy"
app.MapGet("/claims/ping", () => "Claims API is running");
app.Run();Run it and call GET /health: you get HTTP 200 with body Healthy. With no checks registered, this only proves the process can serve HTTP, which is a valid (if shallow) liveness signal.
Level 2: Intermediate
In a real .NET + Angular + SQL Server application you separate liveness (cheap, no dependencies) from readiness (checks the dependencies this instance needs), using check tags.
Step 1: Register checks with tags (Claims API, .NET 10, EF Core + SQL Server).
Packages: Microsoft.Extensions.Diagnostics.HealthChecks.EntityFrameworkCore (for AddDbContextCheck). For Service Bus, AspNetCore.HealthChecks.AzureServiceBus from the community Xabaril project is common.
using Microsoft.AspNetCore.Diagnostics.HealthChecks;using Microsoft.EntityFrameworkCore;using Microsoft.Extensions.Diagnostics.HealthChecks;using System.Text.Json;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddDbContext<ClaimsDbContext>(o => o.UseSqlServer(builder.Configuration.GetConnectionString("ClaimsDb")));
builder.Services.AddHealthChecks() // Liveness: no external dependencies, always cheap .AddCheck("self", () => HealthCheckResult.Healthy(), tags: new[] { "live" }) // Readiness: can this instance actually serve claim requests? .AddDbContextCheck<ClaimsDbContext>("claims-db", tags: new[] { "ready" }) .AddCheck<PolicyServiceCheck>("policy-service", failureStatus: HealthStatus.Degraded, // non-critical: degrade, do not fail tags: new[] { "ready" });
builder.Services.AddHttpClient("policy", c => c.BaseAddress = new Uri(builder.Configuration["Services:Policy"]!));
var app = builder.Build();
// Liveness: only checks tagged "live"app.MapHealthChecks("/health/live", new HealthCheckOptions{ Predicate = r => r.Tags.Contains("live")});
// Readiness: checks tagged "ready"; Degraded still returns 200 by defaultapp.MapHealthChecks("/health/ready", new HealthCheckOptions{ Predicate = r => r.Tags.Contains("ready"), ResponseWriter = WriteJson});
app.Run();
static Task WriteJson(HttpContext ctx, HealthReport report){ ctx.Response.ContentType = "application/json"; var body = new { status = report.Status.ToString(), totalMs = report.TotalDuration.TotalMilliseconds, checks = report.Entries.Select(e => new { name = e.Key, status = e.Value.Status.ToString(), ms = e.Value.Duration.TotalMilliseconds // Do NOT include e.Value.Exception or connection strings here. }) }; return ctx.Response.WriteAsync(JsonSerializer.Serialize(body));}
public class ClaimsDbContext(DbContextOptions<ClaimsDbContext> options) : DbContext(options) { }
public class PolicyServiceCheck(IHttpClientFactory factory) : IHealthCheck{ public async Task<HealthCheckResult> CheckHealthAsync( HealthCheckContext context, CancellationToken ct = default) { try { var client = factory.CreateClient("policy"); using var cts = CancellationTokenSource.CreateLinkedTokenSource(ct); cts.CancelAfter(TimeSpan.FromSeconds(2)); // a probe must never hang var resp = await client.GetAsync("/health/live", cts.Token); return resp.IsSuccessStatusCode ? HealthCheckResult.Healthy() : HealthCheckResult.Degraded($"Policy returned {(int)resp.StatusCode}"); } catch (Exception ex) { return new HealthCheckResult(context.Registration.FailureStatus, "Policy service unreachable", ex); } }}Key points:
HealthStatus.HealthyandDegradedmap to HTTP 200,Unhealthymaps to 503 by default (ResultStatusCodescan change this).AddDbContextCheckrunsCanConnectAsyncby default. It proves connectivity, not that migrations are applied; pass a custom query via thecustomTestQueryparameter if you need more.- The Policy check is registered as
Degradedon failure because the Claims API can still accept a claim without a live policy lookup if it has a fallback (Day 29).
Step 2: Protect the endpoints and hold back detail. Internal probes are usually called by the platform on the container network. If you expose them publicly through a gateway, expose only /health/live or nothing, and never return exception messages.
Step 3: Angular consumption. Angular should not call /health/* for its own business logic. Two legitimate uses are an internal ops page and an app-level “service degraded” banner fed by your BFF/gateway.
// ops-status.service.ts (Angular 21/22, standalone, signals)import { Injectable, inject, signal } from '@angular/core';import { HttpClient } from '@angular/common/http';import { timer, switchMap, catchError, of } from 'rxjs';import { takeUntilDestroyed } from '@angular/core/rxjs-interop';
export interface HealthReport { status: 'Healthy' | 'Degraded' | 'Unhealthy'; checks: { name: string; status: string; ms: number }[];}
@Injectable({ providedIn: 'root' })export class OpsStatusService { private http = inject(HttpClient); readonly report = signal<HealthReport | null>(null);
constructor() { timer(0, 30_000) .pipe( switchMap(() => this.http.get<HealthReport>('/api/claims/health/ready').pipe( catchError(() => of({ status: 'Unhealthy', checks: [] } as HealthReport)) ) ), takeUntilDestroyed() ) .subscribe(r => this.report.set(r)); }}Note that the 503 response body still comes back as an HttpErrorResponse; the catchError above collapses it to “Unhealthy” for simplicity.
Step 4: Kubernetes wiring (works on AKS, see section 9).
containers: - name: claims-api image: acrclaims.azurecr.io/claims-api:1.4.2 ports: [{ containerPort: 8080 }] startupProbe: httpGet: { path: /health/live, port: 8080 } periodSeconds: 3 failureThreshold: 30 # up to 90s to boot livenessProbe: httpGet: { path: /health/live, port: 8080 } periodSeconds: 10 timeoutSeconds: 2 failureThreshold: 3 readinessProbe: httpGet: { path: /health/ready, port: 8080 } periodSeconds: 5 timeoutSeconds: 3 failureThreshold: 2Note: .NET 8 and later container images listen on port 8080 by default (non-root), not 80, which is why the port is 8080 here.
Level 3: Advanced
Performance.
- Probes run continuously on every instance. If a readiness check runs a heavy SQL query every 5 seconds across 20 pods, you have built a self-inflicted load test. Keep checks cheap (
SELECT 1-level) and set a timeout on every dependency call. - Use
HealthCheckPublisheror caching to decouple probe frequency from dependency load: run checks on a timer (for example every 10 seconds) and let the endpoint return the last cached result.
builder.Services.Configure<HealthCheckPublisherOptions>(o =>{ o.Delay = TimeSpan.FromSeconds(5); o.Period = TimeSpan.FromSeconds(15); o.Predicate = r => r.Tags.Contains("ready");});builder.Services.AddSingleton<IHealthCheckPublisher, CachedReadinessPublisher>();// CachedReadinessPublisher stores the latest HealthReport (pseudo-code: keep it in a// static/volatile field) and a minimal endpoint returns 200/503 from that cache.- Health check registrations also accept a per-check
timeoutargument; use it so one slow dependency cannot make the whole probe exceed the platform’stimeoutSeconds.
Scalability. Readiness is your back-pressure valve. A service can deliberately report Unhealthy on readiness when overloaded (for example, when the request queue is full) so the load balancer sheds traffic away and lets it recover. Be careful: if all instances do this simultaneously you remove all capacity; prefer rate limiting (Day 30) for overload and reserve readiness for “cannot serve correctly”.
Security.
- Do not leak internals: no stack traces, connection strings, server names, or version numbers in the public response body. The default
Healthy/Unhealthyplain text is safe; a detailed JSON writer should sit on an internal-only route. - Restrict access: use
.RequireHost("*:8080")with a separate internal port, network policies, or.RequireAuthorization("ops")for the detailed endpoint. Public API gateways should not route/health/readyto the internet. - Health endpoints bypass your normal auth. Make sure they do not perform side effects (no writes, no message publishing) and cannot be abused for amplification (each hit triggers downstream calls).
Failure modes and common mistakes.
| Mistake | Consequence | Fix |
|---|---|---|
| Liveness probe checks the database | DB outage causes every pod to be restarted in a loop, prolonging recovery | Liveness must be dependency-free; DB belongs in readiness |
| Readiness includes every downstream service | One non-critical outage removes all pods from rotation, total outage | Only include critical dependencies; mark others Degraded |
No startupProbe and short initialDelaySeconds | Slow-starting pod is killed before it finishes booting, then restarts forever | Use startup probe with generous failureThreshold * periodSeconds |
| Probe timeout shorter than a GC pause or cold start | Random restarts under load | timeoutSeconds of 2-5s, failureThreshold of 3 for liveness |
| Health check swallows exceptions and always returns Healthy | Endpoint is decoration, not signal | Test by breaking the dependency in a staging environment |
| Health check opens a new connection per call without disposing | Connection pool exhaustion caused by the probe itself | Reuse pooled connections/DbContext from DI and dispose correctly |
Returning 200 with "status":"Unhealthy" in the body | Platforms judge by HTTP status code, not body | Ensure mapping to 503, or use the built-in middleware |
| Same endpoint for liveness and readiness | Cannot restart without also draining, or vice versa | Separate endpoints and tags |
| Health status cached forever | Stale “healthy” after failure | Cache for seconds, not minutes |
Level 4: Expert and Architect view
Design trade-offs and alternatives.
| Approach | Strengths | Weaknesses | Use when |
|---|---|---|---|
| Shallow liveness + dependency-aware readiness (recommended default) | Avoids restart storms; drains traffic on dependency loss | Needs discipline on what counts as critical | Almost every HTTP service |
Single deep /health endpoint | Simple to build | Conflates restart and drain; restart storms | Very small internal tools only |
| TCP-only probe | Zero code, works for any protocol | Proves the port is open, not that the app works | Legacy or non-HTTP workloads |
gRPC health checking protocol (grpc.health.v1) | Standard for gRPC, native Kubernetes gRPC probes | Extra package and wiring | gRPC services |
| Push-based heartbeat (service reports to a monitor) | Works for workers with no inbound port | Monitor must detect missing heartbeats | Background workers, scheduled jobs |
| Synthetic transactions (external monitoring) | Tests the real user path end to end | Slow, costly, not for orchestrator probes | Complement, not replacement |
| Health via metrics only (Prometheus alerts) | Rich signal, trends | Orchestrator cannot act on it directly | Alerting alongside probes |
Patterns it combines with.
- Circuit Breaker (Day 26) and Fallback (Day 29): a service whose circuit to Policy is open can report
Degradedrather thanUnhealthy. - Service Registry / Discovery (Day 21-25): registries use health status to add and remove instances; Kubernetes readiness does this natively for Services.
- Application Metrics (Day 33): export health check status as a metric (
aspnetcore_healthcheck_status) so you can alert on flapping, not just current state. - Log Deployments & Changes (Day 37): readiness gating during rollout plus deploy markers makes “did this release break things” answerable.
- Sidecar / Service Mesh (Day 47-48): meshes have their own outlier detection; be aware that probes may need to bypass mTLS.
- Externalized Configuration (Day 39): a readiness check that verifies required config/secrets loaded catches bad deployments early.
ADR-style justification.
ADR-036: Standard health endpoints for all Claims platform services
Status: Accepted
Context: We run 14 .NET services on Azure Container Apps and AKS. Deployments cause 20-40s of 5xx errors because new replicas receive traffic before migrations and warm-up complete. Two incidents in the last quarter required manual pod restarts for deadlocked instances, and each team has its own
/statusimplementation.Decision: Every service will expose
/health/live(dependency-free, taglive) and/health/ready(critical dependencies only, tagready), built withMicrosoft.Extensions.Diagnostics.HealthChecks. Platform probes: startup and liveness hit/health/live; readiness hits/health/ready. Non-critical dependencies reportDegraded, notUnhealthy. Detailed JSON is available only on an internal port. The shared service chassis (Day 40) will supply the registration so teams do not reimplement it.Consequences: (+) Zero-downtime rollouts, automatic recovery from hangs, uniform ops experience. (-) Teams must classify dependencies as critical or not; misclassification can still cause outages. (-) Probe traffic adds small constant load; mitigated by cached checks. Review after two quarters using restart counts and deploy error rates.
Alternatives rejected: single deep endpoint (restart storms); TCP probes only (no application-level signal); relying on external uptime monitoring alone (cannot drain or restart instances).
Azure implementation
Which Azure services use or implement this topic.
- Azure Kubernetes Service (AKS): native
startupProbe,livenessProbe,readinessProbe(HTTP, TCP, exec, gRPC). Standard Kubernetes behaviour. - Azure Container Apps: supports startup, liveness, and readiness probes (TCP or HTTP/S) configured per container. If you configure none, default probes are applied (TCP-based); check the current defaults in the Container Apps health probes documentation for the exact values, because they differ per probe type.
- Azure App Service: the Health Check feature pings a path you configure (about once a minute); after a configurable number of consecutive failures the instance is removed from load-balancer rotation, and if it stays unhealthy it can be replaced. It only acts when the App Service plan runs two or more instances. Configure the path in the portal under Monitoring > Health check, or by IaC (
healthCheckPathsite property) and set theWEBSITE_HEALTHCHECK_MAXPINGFAILURESapp setting to tune the failure count. - Azure Application Gateway and Azure Front Door: custom health probes decide which backends receive traffic. Point them at a lightweight endpoint. Front Door probes come from many edge locations, so a heavy probe endpoint multiplies load.
- Azure Load Balancer: health probes (TCP/HTTP/HTTPS) for VM or VMSS backends.
- Azure Monitor / Application Insights: availability tests (Standard tests) hit a public URL from multiple regions and alert on failure; this is external, user-perspective monitoring. Use an endpoint that is safe to expose publicly. Verify the current availability-test types, since the older URL ping tests are being retired in favour of Standard tests.
- Azure Functions: no separate liveness/readiness concept; App Service Health Check applies on some hosting plans, and Application Insights covers monitoring.
How to configure (Container Apps, Bicep excerpt).
resource claimsApi 'Microsoft.App/containerApps@2024-03-01' = { name: 'ca-claims-api' location: location properties: { managedEnvironmentId: env.id configuration: { ingress: { external: false, targetPort: 8080 } } template: { containers: [ { name: 'claims-api' image: '${acr.properties.loginServer}/claims-api:1.4.2' probes: [ { type: 'Startup' httpGet: { path: '/health/live', port: 8080 } periodSeconds: 3 failureThreshold: 30 } { type: 'Liveness' httpGet: { path: '/health/live', port: 8080 } periodSeconds: 10 failureThreshold: 3 timeoutSeconds: 2 } { type: 'Readiness' httpGet: { path: '/health/ready', port: 8080 } periodSeconds: 5 failureThreshold: 3 timeoutSeconds: 3 } ] } ] scale: { minReplicas: 2, maxReplicas: 10 } } }}The apiVersion above is illustrative; use the current stable API version from the Azure resource reference when you deploy.
App Service via Azure CLI:
az webapp config set -g rg-claims -n app-claims-api --generic-configurations '{"healthCheckPath": "/health/ready"}'az webapp config appsettings set -g rg-claims -n app-claims-api --settings WEBSITE_HEALTHCHECK_MAXPINGFAILURES=5For App Service, pointing at the readiness endpoint is a reasonable choice; unlike Kubernetes there is no separate liveness concept, and a failing instance is removed from rotation.
Pricing and tier considerations.
- The probes themselves (AKS, Container Apps, App Service Health Check, Application Gateway/Front Door/Load Balancer probes) carry no separate charge; you pay for the underlying compute and networking.
- App Service Health Check needs at least two instances, so it implies a plan tier and instance count that supports scale-out (Basic and above for manual scale-out). Single-instance Free/Shared plans cannot benefit.
- Application Insights availability tests are billed per test execution/per test per month depending on the type; confirm current rates on the Azure Monitor pricing page before creating many tests across regions.
- Front Door and Application Gateway have tiered pricing (Standard/Premium; Application Gateway v2 SKUs); health probing is included, but the gateway SKU choice is the cost driver, not the probe.
- Verify all SKUs and rates against the pricing pages at decision time; Azure tiers and names change.
Reference architecture (text).
Customers use the Angular app served from Azure Static Web Apps or Blob storage behind Azure Front Door. Front Door routes /api/* to an internal Application Gateway or directly to API Management, which forwards to the Claims, Policy, and Payments services running in an Azure Container Apps environment (or AKS). Each service exposes /health/live and /health/ready. Container Apps/AKS use startup, liveness, and readiness probes to gate rollouts and restart deadlocked replicas. Front Door and the gateway probe a shallow public-safe endpoint to detect regional or backend loss and fail over. Application Insights receives request telemetry and availability test results, and an alert rule fires when readiness failures or restart counts exceed a threshold, notifying the on-call team via an Action Group. Detailed health JSON stays on an internal port reachable only inside the virtual network.
Teaching guide for my team
Explain to a beginner in 2 minutes. “Your service needs a way to answer three questions when the platform asks. Am I still running properly (liveness)? Am I ready to take customer requests (readiness)? Have I finished starting up (startup)? We add two URLs, /health/live and /health/ready. If live fails, the platform restarts us. If ready fails, the platform stops sending customers to us until we recover. Liveness never checks the database. Readiness does.” Draw a kitchen: on fire versus not ready for orders.
Explain to an intermediate developer in 5 minutes. Walk through: tags and predicates in AddHealthChecks/MapHealthChecks; the Healthy/Degraded/Unhealthy mapping to HTTP 200/200/503; why the DB goes in readiness only; what a restart storm is and how a deep liveness probe causes one; startup probes for slow boots; timeouts on every check; keeping detail off public routes. Show the AKS/Container Apps probe YAML next to the code so they see the two halves of the contract. Finish with the Degraded versus Unhealthy decision for the Policy dependency.
Hands-on exercise.
- Create a .NET 10 minimal API “Claims.Api” with
/health/liveand/health/readyas in section 6, with a SQL Server container as the database (docker compose). - Run both in Docker Compose or a local Kubernetes (kind/minikube) with the probes from section 6.
- Stop the SQL container.
Expected outcome: /health/live stays 200 and the pod is not restarted; /health/ready returns 503, the pod shows 0/1 Ready, and the Service has no endpoints (requests get connection errors from the load balancer). Restart SQL: readiness recovers within a few probe periods and traffic resumes. Bonus: temporarily add the DB check to liveness and observe the restart loop, then discuss why that is worse.
Interview-style questions.
- What is the difference between liveness and readiness? Liveness asks whether the process should be restarted; readiness asks whether it should receive traffic right now. Failing liveness restarts the instance; failing readiness only removes it from load balancing.
- Why should a liveness probe not check the database? If the DB is down, every instance fails liveness and is restarted repeatedly, adding load and prolonging recovery, while a restart cannot fix the DB. Dependencies belong in readiness.
- What is a startup probe for? It gives slow-starting apps time to boot (migrations, warm-up) before liveness and readiness begin, so the platform does not kill an instance that is merely still starting.
Mastery checklist
- I can explain liveness, readiness, and startup with one example each from the claims system.
- I can implement tagged health checks in .NET 10 and map them to separate endpoints.
- I can predict what happens (restart, drain, or nothing) for each failure: deadlock, DB down, SMS gateway down, slow boot.
- I can configure probes correctly on AKS and Azure Container Apps, including timeouts and thresholds.
- I can explain a restart storm and show how to avoid it.
- I can secure health endpoints: no sensitive detail, restricted routes, no side effects.
- I can decide
DegradedversusUnhealthyfor a given dependency and justify it. - I can write a short ADR standardizing health endpoints across services.
Key takeaway
Liveness answers “restart me?” and must never depend on other systems; readiness answers “send me traffic?” and should check only what this instance truly needs. Get that split right and the platform heals your services instead of amplifying their failures.
