A Service Deployment Platform is an orchestrator (in practice, usually Kubernetes) that takes a declared desired state, such as "run 3 copies of the Claims API at version 2.4", and continuously makes reality match it.
Intro
A Service Deployment Platform is an orchestrator (in practice, usually Kubernetes) that takes a declared desired state, such as “run 3 copies of the Claims API at version 2.4”, and continuously makes reality match it. It schedules containers onto machines, scales them, rolls out new versions gradually, and restarts or replaces anything that fails. Teams stop deploying to servers and start deploying to a platform.
Why we need this
- Once you have tens of services and several instances of each, “which server runs what” becomes a scheduling problem no human can solve reliably by hand.
- Business demand: insurers see spikes (storm events cause thousands of claims in hours). Capacity must grow and shrink automatically.
- Technical demand: independent releases per team (Day 4, Service per Team) only work if deploying is cheap, safe and repeatable.
- A platform gives one consistent contract (a manifest) for deploy, scale, config, secrets, networking and health, regardless of the language of the service.
What problem it solves
Problem (from the topic list): manual scheduling, rollouts and healing do not scale.
Without a platform in our claims system:
- An ops engineer RDPs/SSHs to VMs, pulls a new image, restarts the
ClaimsApiprocess, and hopes. One typo takes the service down. - A crashed
PaymentsServicestays down until someone notices a customer complaint. - Capacity is sized for the worst day of the year, so 90% of the time you pay for idle machines.
- Rollbacks mean finding the previous build and repeating the manual steps under pressure.
- Each service is deployed differently (scripts, Word documents, tribal knowledge).
When it is needed (and when it is NOT)
Fits when:
- 10+ services, or fewer services with several replicas and frequent releases.
- You need zero-downtime rollouts, autoscaling, self-healing, and multi-team self-service.
- You already package services as containers (Day 42).
Overkill or wrong when:
- One or two services with steady low traffic: Azure App Service or Azure Container Apps is cheaper to run and to learn.
- Purely event-driven, spiky, short jobs: serverless (Day 45, Azure Functions) may fit better.
- The team has no capacity to learn and operate Kubernetes (networking, RBAC, upgrades). A managed higher-level option such as Container Apps or AKS Automatic reduces this cost but does not remove it.
- Hard hypervisor isolation requirements per service (Day 43).
How to identify the problem (key signals)
- Deployments are done by a named person, out of hours, from a checklist.
- Mean time to recover from a crashed instance is measured in minutes or hours because a human must notice it first.
- Production and staging differ (“works on staging”) because servers were configured by hand.
- Utilisation graphs show servers at 10-20% CPU most of the day, yet you still hit capacity during peaks.
- Rollbacks are feared and take longer than the original release.
- Different services have different deployment scripts, and new joiners cannot deploy on day one.
- Release frequency is throttled (“we only deploy on Thursdays”) because deploys are risky.
Flow Diagram
You declare desired state; the platform reconciles, scales and heals.
flowchart LR GIT["Git manifests"] -- "GitOps sync" --> API["Kubernetes API"] API --> LOOP{"Desired equals actual?"} LOOP -- "no" --> SCH["Scheduler places pods"] SCH --> NODES["Node pools across zones"] NODES --> LOOP LOOP -- "yes" --> WATCH["Watch for changes"] HPA["HPA / KEDA"] --> API CRASH["Pod crashes"] --> LOOPLevel 1: Beginner
Analogy: an airport operations centre. Airlines (teams) say what flights they want. The centre decides which gate, which crew and which runway is free, replaces a broken aircraft, and adds flights when demand rises. Nobody phones each pilot.
Core ideas: you describe desired state in YAML; the platform’s control loop compares it with actual state and fixes the difference.
Minimal Kubernetes example for the Claims API:
apiVersion: apps/v1kind: Deploymentmetadata: name: claims-apispec: replicas: 3 selector: matchLabels: app: claims-api template: metadata: labels: app: claims-api spec: containers: - name: claims-api image: contosoinsurance.azurecr.io/claims-api:2.4.0 ports: - containerPort: 8080---apiVersion: v1kind: Servicemetadata: name: claims-apispec: selector: app: claims-api ports: - port: 80 targetPort: 8080Apply with kubectl apply -f claims-api.yaml. Delete one pod and watch Kubernetes create a replacement: that is self-healing.
Level 2: Intermediate
A real .NET claims API needs health probes, resource requests/limits, config from outside the image, and a rolling update strategy. Note: recent official .NET container images (.NET 8 and later) listen on port 8080 and run as a non-root user by default.
apiVersion: apps/v1kind: Deploymentmetadata: name: claims-api labels: app: claims-apispec: replicas: 3 strategy: type: RollingUpdate rollingUpdate: maxUnavailable: 0 maxSurge: 1 selector: matchLabels: app: claims-api template: metadata: labels: app: claims-api spec: containers: - name: claims-api image: contosoinsurance.azurecr.io/claims-api:2.4.0 ports: - containerPort: 8080 env: - name: ConnectionStrings__ClaimsDb valueFrom: secretKeyRef: name: claims-db key: connection-string resources: requests: { cpu: "250m", memory: "256Mi" } limits: { memory: "512Mi" } readinessProbe: httpGet: { path: /health/ready, port: 8080 } periodSeconds: 5 livenessProbe: httpGet: { path: /health/live, port: 8080 } initialDelaySeconds: 10 periodSeconds: 10Matching ASP.NET Core health endpoints (Day 36):
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddHealthChecks() .AddCheck("self", () => HealthCheckResult.Healthy(), tags: new[] { "live" }) .AddSqlServer( builder.Configuration.GetConnectionString("ClaimsDb")!, name: "sql", tags: new[] { "ready" }); // needs AspNetCore.HealthChecks.SqlServer package
var app = builder.Build();
app.MapHealthChecks("/health/live", new HealthCheckOptions { Predicate = r => r.Tags.Contains("live") });app.MapHealthChecks("/health/ready", new HealthCheckOptions { Predicate = r => r.Tags.Contains("ready") });
app.Run();Autoscaling on CPU:
apiVersion: autoscaling/v2kind: HorizontalPodAutoscalermetadata: name: claims-apispec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: claims-api minReplicas: 3 maxReplicas: 20 metrics: - type: Resource resource: name: cpu target: type: Utilization averageUtilization: 65Angular side: the Angular claims portal is built to static files and served by an nginx container as its own Deployment, or from Azure Static Web Apps / Blob storage behind a CDN. The Angular app calls the API through an Ingress or API gateway (Day 19) using a relative path such as /api/claims, so no environment-specific URL is baked into the bundle.
Database: SQL Server / PostgreSQL usually stay outside the cluster as managed services (Azure SQL, Azure Database for PostgreSQL). Pods reach them via private endpoints. Avoid running production databases as ordinary Deployments.
Level 3: Advanced
Performance and scalability
- Set memory and CPU requests accurately: the scheduler bin-packs on requests. Over-requesting wastes money; under-requesting causes noisy neighbours.
- Prefer a memory limit and consider leaving CPU limit unset to avoid throttling of latency-sensitive .NET services. Remember .NET sizes its thread pool and GC heaps from container limits, so set them deliberately.
- Use HPA for pods and the cluster autoscaler (or Karpenter-based node autoprovisioning on AKS) for nodes. KEDA scales on queue length, e.g. Service Bus messages waiting for the claims-processing worker.
- Use PodDisruptionBudgets and topology spread constraints so node upgrades and zone failures do not remove all replicas at once.
Security
- Non-root containers, read-only root filesystem, dropped capabilities.
- Namespaces + RBAC per team; NetworkPolicies default-deny between services.
- Secrets from Azure Key Vault via the Secrets Store CSI driver, and Microsoft Entra Workload Identity instead of stored credentials (Day 39).
- Admission policy (Azure Policy for Kubernetes / Pod Security Standards) to block privileged pods and untrusted registries.
Failure modes
- Wrong readiness probe: pod receives traffic before the app is ready, or a slow dependency outage makes every pod unready and takes the whole service offline. Keep readiness about “can I serve”, not deep dependency checks of everything.
- Liveness probe that checks the database: a DB blip restarts every pod (restart storm). Liveness should check only the process itself.
- Missing graceful shutdown: on SIGTERM ASP.NET Core drains requests for
HostOptions.ShutdownTimeout(default 30 s in .NET 8+). Align this withterminationGracePeriodSeconds. :latesttags and no resource requests lead to unpredictable rollouts and evictions.- Stateful services with local disk written into pods that get rescheduled.
Common mistakes
- Adopting Kubernetes before having CI/CD, observability and container basics.
- One giant cluster with no namespace or quota boundaries.
- Treating YAML as a copy-paste artefact instead of versioned, reviewed code (GitOps).
- Ignoring the upgrade cadence: Kubernetes minor versions have a limited support window.
Level 4: Expert and Architect view
| Option | Ops effort | Control | Scale-to-zero | Best fit | Main drawback |
|---|---|---|---|---|---|
| Kubernetes (AKS) | High (Medium with AKS Automatic) | Very high | With KEDA / add-ons | Many services, platform team, custom networking | Complexity, needs platform skills |
| Azure Container Apps | Low | Medium | Yes | Microservices without Kubernetes admin, event-driven | Less low-level control, fewer Kubernetes features exposed |
| Azure App Service | Low | Low-medium | No | A handful of web apps and APIs | Coarse scaling, not container-scheduler semantics |
| Azure Functions (Day 45) | Very low | Low | Yes | Event-driven, short tasks | Execution limits, cold starts on some plans |
| Service per VM (Day 43) | High | High isolation | No | Compliance / hypervisor isolation | Slow, costly, manual scaling |
| Docker Swarm / Nomad | Medium | Medium | No | Smaller estates | Smaller ecosystem, less Azure integration |
Combines with: Service per Container (Day 42), Sidecar (Day 47), Service Mesh (Day 48), Health Check API (Day 36), Externalized Configuration (Day 39), Service Registry patterns (Days 21-25; Kubernetes Services and DNS act as built-in server-side discovery), Circuit Breaker and Retry (Days 26, 28), Log Deployments & Changes (Day 37).
ADR (architecture review)
- Title: ADR-046 Adopt AKS as the service deployment platform for the claims system
- Status: Proposed
- Context: 14 services across 5 teams, weekly-to-daily releases, seasonal claim spikes (10x), 99.9% availability target, existing .NET 8+/Angular skills, Azure-first strategy.
- Decision: Run all containerised services on AKS (Standard tier, three availability zones), with a platform team owning cluster operations, GitOps (Flux or Argo CD) for deployment, and namespaces per team. Evaluate Azure Container Apps for low-traffic edge services.
- Consequences (positive): self-healing, uniform deployment contract, autoscaling, portable across clouds at the manifest level.
- Consequences (negative): platform team headcount, Kubernetes learning curve, upgrade duty every few months, cost of always-on system node pool.
- Alternatives considered: Container Apps (rejected for now: need custom network policy and mesh options), App Service (rejected: 14 services with per-service scaling profiles and sidecars).
- Review trigger: revisit if service count drops below ~6 or no platform team can be staffed.
Azure implementation
Services
- Azure Kubernetes Service (AKS): managed control plane. Options include standard AKS and AKS Automatic (opinionated, more automated node and add-on management, Standard tier by default).
- Azure Container Registry (ACR): private image storage, attached to AKS with managed identity (
--attach-acr). - Azure Container Apps: a simpler managed platform (built on Kubernetes, without exposing its API) for teams who want less operational surface.
- Supporting: Azure Key Vault, Azure Monitor / Container Insights + Managed Prometheus + Managed Grafana, Application Insights (OpenTelemetry), Microsoft Defender for Containers, Azure Policy, Azure Application Gateway for Containers or NGINX / Gateway API ingress, Azure SQL / PostgreSQL Flexible Server via private endpoints.
Configuration (example)
az group create -n rg-claims-prod -l westeurope
az acr create -g rg-claims-prod -n contosoinsurance --sku Standard
az aks create \ -g rg-claims-prod -n aks-claims-prod \ --tier standard \ --node-count 3 --zones 1 2 3 \ --enable-managed-identity \ --enable-oidc-issuer --enable-workload-identity \ --enable-cluster-autoscaler --min-count 3 --max-count 10 \ --attach-acr contosoinsurance \ --network-plugin azure --network-plugin-mode overlay
az aks get-credentials -g rg-claims-prod -n aks-claims-prodPricing and tier considerations (verified against Microsoft Learn; check the Azure pricing page for current per-hour rates in your region)
- Free tier: no management charge and no financially backed SLA; up to 1,000 nodes. Use for dev/test and learning.
- Standard tier: financially backed uptime SLA of 99.95% with availability zones (99.9% without); up to 5,000 nodes. Use for production. This is a per-cluster hourly management fee on top of node costs.
- Premium tier: Standard features plus 24-month Long Term Support for Kubernetes versions (
--k8s-support-plan AKSLongTermSupport), for regulated estates that cannot upgrade yearly. - You always pay for the worker VMs, disks, load balancers, public IPs, egress and monitoring ingestion. Node VMs are typically the dominant cost: use right-sized SKUs, spot node pools for non-critical batch workers, reserved instances / savings plans for steady baseline, and cluster start/stop for non-production.
- Log Analytics ingestion from Container Insights can become a surprise cost; tune collection and use Managed Prometheus for metrics.
Reference architecture (text) Users reach Azure Front Door (WAF, CDN). Front Door routes to an ingress (Application Gateway for Containers) in front of an AKS Standard cluster spread over three zones. A system node pool runs cluster components; user node pools run the Angular portal (nginx), Claims API, Payments, Notifications and worker services in team namespaces. Images come from ACR. Pods use Workload Identity to fetch secrets from Key Vault and to reach Azure SQL and PostgreSQL over private endpoints, and to send/receive messages via Azure Service Bus. KEDA scales workers on Service Bus queue depth. OpenTelemetry sends traces to Application Insights, metrics go to Managed Prometheus/Grafana, and Flux syncs manifests from a Git repository. Defender for Containers and Azure Policy enforce guardrails.
Teaching guide for my team
2-minute beginner explanation: “You used to log in to a server and start your app. Now you write a small file that says: run 3 copies of this container, keep them healthy, and use this much memory. Kubernetes reads the file and does the work. If one copy crashes, it starts another. If we ship a new version, it swaps copies one by one so users see no downtime. We describe what we want; the platform figures out how.”
5-minute intermediate explanation: Walk through Deployment, ReplicaSet, Pod, Service, Ingress, ConfigMap/Secret. Explain the reconcile loop, requests vs limits, readiness vs liveness, rolling updates with maxSurge/maxUnavailable, and HPA. Show the claims-api YAML from Level 2 and point out how each field prevents one failure from Section 4.
Hands-on exercise
- Create a local cluster (kind or minikube) and deploy the
claims-apimanifest with 3 replicas. - Delete one pod:
kubectl delete pod <name>. - Change the image tag to a new version and watch
kubectl rollout status deployment/claims-api. - Deploy a version whose
/health/readyreturns 503 and observe that traffic never reaches it and the rollout stalls, then runkubectl rollout undo.
Expected outcome: the replacement pod appears within seconds; the good rollout completes with no failed requests; the broken rollout stays halted with old pods still serving, and undo restores the previous version.
Interview questions
- What is the difference between a liveness and a readiness probe? Readiness controls whether a pod receives traffic; liveness controls whether the container is restarted.
- Why set resource requests? The scheduler uses requests to place pods and autoscalers use them for utilisation; without them, placement and scaling are unpredictable.
- When would you choose Container Apps over AKS? When you want microservices, autoscaling and revisions without operating Kubernetes, and do not need direct Kubernetes API or custom cluster-level networking.
Mastery checklist
- I can explain desired state vs actual state and the reconcile loop.
- I can write a Deployment, Service, Ingress, and HPA from scratch and explain every field.
- I can design readiness and liveness probes that avoid restart storms.
- I can perform, pause and roll back a rolling update with zero downtime.
- I can explain the AKS Free, Standard and Premium tiers and choose one for a workload.
- I can secure a workload with Workload Identity, Key Vault, RBAC and NetworkPolicy.
- I can justify AKS vs Container Apps vs App Service in an ADR.
- I can diagnose a CrashLoopBackOff or Pending pod using
kubectl describeand logs.
Key takeaway
A service deployment platform turns deployment from a manual procedure into a declared, self-correcting system: you state the desired result, and the orchestrator schedules, scales, rolls out and heals. Adopt it when service count and release pace justify its operating cost, and pick the lowest-effort Azure option that meets your needs.
