Manikandan — Manikandan
Microservices

Day 46: Service Deployment Platform

ManikandanManikandan
12 min read·Updated Sep 24, 2022

A Service Deployment Platform is an orchestrator (in practice, usually Kubernetes) that takes a declared desired state, such as "run 3 copies of the Claims API at version 2.4", and continuously makes reality match it.

Intro

A Service Deployment Platform is an orchestrator (in practice, usually Kubernetes) that takes a declared desired state, such as “run 3 copies of the Claims API at version 2.4”, and continuously makes reality match it. It schedules containers onto machines, scales them, rolls out new versions gradually, and restarts or replaces anything that fails. Teams stop deploying to servers and start deploying to a platform.

Why we need this

  • Once you have tens of services and several instances of each, “which server runs what” becomes a scheduling problem no human can solve reliably by hand.
  • Business demand: insurers see spikes (storm events cause thousands of claims in hours). Capacity must grow and shrink automatically.
  • Technical demand: independent releases per team (Day 4, Service per Team) only work if deploying is cheap, safe and repeatable.
  • A platform gives one consistent contract (a manifest) for deploy, scale, config, secrets, networking and health, regardless of the language of the service.

What problem it solves

Problem (from the topic list): manual scheduling, rollouts and healing do not scale.

Without a platform in our claims system:

  • An ops engineer RDPs/SSHs to VMs, pulls a new image, restarts the ClaimsApi process, and hopes. One typo takes the service down.
  • A crashed PaymentsService stays down until someone notices a customer complaint.
  • Capacity is sized for the worst day of the year, so 90% of the time you pay for idle machines.
  • Rollbacks mean finding the previous build and repeating the manual steps under pressure.
  • Each service is deployed differently (scripts, Word documents, tribal knowledge).

When it is needed (and when it is NOT)

Fits when:

  • 10+ services, or fewer services with several replicas and frequent releases.
  • You need zero-downtime rollouts, autoscaling, self-healing, and multi-team self-service.
  • You already package services as containers (Day 42).

Overkill or wrong when:

  • One or two services with steady low traffic: Azure App Service or Azure Container Apps is cheaper to run and to learn.
  • Purely event-driven, spiky, short jobs: serverless (Day 45, Azure Functions) may fit better.
  • The team has no capacity to learn and operate Kubernetes (networking, RBAC, upgrades). A managed higher-level option such as Container Apps or AKS Automatic reduces this cost but does not remove it.
  • Hard hypervisor isolation requirements per service (Day 43).

How to identify the problem (key signals)

  1. Deployments are done by a named person, out of hours, from a checklist.
  2. Mean time to recover from a crashed instance is measured in minutes or hours because a human must notice it first.
  3. Production and staging differ (“works on staging”) because servers were configured by hand.
  4. Utilisation graphs show servers at 10-20% CPU most of the day, yet you still hit capacity during peaks.
  5. Rollbacks are feared and take longer than the original release.
  6. Different services have different deployment scripts, and new joiners cannot deploy on day one.
  7. Release frequency is throttled (“we only deploy on Thursdays”) because deploys are risky.

Flow Diagram

You declare desired state; the platform reconciles, scales and heals.

flowchart LR
GIT["Git manifests"] -- "GitOps sync" --> API["Kubernetes API"]
API --> LOOP{"Desired equals actual?"}
LOOP -- "no" --> SCH["Scheduler places pods"]
SCH --> NODES["Node pools across zones"]
NODES --> LOOP
LOOP -- "yes" --> WATCH["Watch for changes"]
HPA["HPA / KEDA"] --> API
CRASH["Pod crashes"] --> LOOP

Level 1: Beginner

Analogy: an airport operations centre. Airlines (teams) say what flights they want. The centre decides which gate, which crew and which runway is free, replaces a broken aircraft, and adds flights when demand rises. Nobody phones each pilot.

Core ideas: you describe desired state in YAML; the platform’s control loop compares it with actual state and fixes the difference.

Minimal Kubernetes example for the Claims API:

apiVersion: apps/v1
kind: Deployment
metadata:
name: claims-api
spec:
replicas: 3
selector:
matchLabels:
app: claims-api
template:
metadata:
labels:
app: claims-api
spec:
containers:
- name: claims-api
image: contosoinsurance.azurecr.io/claims-api:2.4.0
ports:
- containerPort: 8080
---
apiVersion: v1
kind: Service
metadata:
name: claims-api
spec:
selector:
app: claims-api
ports:
- port: 80
targetPort: 8080

Apply with kubectl apply -f claims-api.yaml. Delete one pod and watch Kubernetes create a replacement: that is self-healing.

Level 2: Intermediate

A real .NET claims API needs health probes, resource requests/limits, config from outside the image, and a rolling update strategy. Note: recent official .NET container images (.NET 8 and later) listen on port 8080 and run as a non-root user by default.

apiVersion: apps/v1
kind: Deployment
metadata:
name: claims-api
labels:
app: claims-api
spec:
replicas: 3
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 1
selector:
matchLabels:
app: claims-api
template:
metadata:
labels:
app: claims-api
spec:
containers:
- name: claims-api
image: contosoinsurance.azurecr.io/claims-api:2.4.0
ports:
- containerPort: 8080
env:
- name: ConnectionStrings__ClaimsDb
valueFrom:
secretKeyRef:
name: claims-db
key: connection-string
resources:
requests: { cpu: "250m", memory: "256Mi" }
limits: { memory: "512Mi" }
readinessProbe:
httpGet: { path: /health/ready, port: 8080 }
periodSeconds: 5
livenessProbe:
httpGet: { path: /health/live, port: 8080 }
initialDelaySeconds: 10
periodSeconds: 10

Matching ASP.NET Core health endpoints (Day 36):

var builder = WebApplication.CreateBuilder(args);
builder.Services.AddHealthChecks()
.AddCheck("self", () => HealthCheckResult.Healthy(), tags: new[] { "live" })
.AddSqlServer(
builder.Configuration.GetConnectionString("ClaimsDb")!,
name: "sql",
tags: new[] { "ready" }); // needs AspNetCore.HealthChecks.SqlServer package
var app = builder.Build();
app.MapHealthChecks("/health/live", new HealthCheckOptions { Predicate = r => r.Tags.Contains("live") });
app.MapHealthChecks("/health/ready", new HealthCheckOptions { Predicate = r => r.Tags.Contains("ready") });
app.Run();

Autoscaling on CPU:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: claims-api
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: claims-api
minReplicas: 3
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 65

Angular side: the Angular claims portal is built to static files and served by an nginx container as its own Deployment, or from Azure Static Web Apps / Blob storage behind a CDN. The Angular app calls the API through an Ingress or API gateway (Day 19) using a relative path such as /api/claims, so no environment-specific URL is baked into the bundle.

Database: SQL Server / PostgreSQL usually stay outside the cluster as managed services (Azure SQL, Azure Database for PostgreSQL). Pods reach them via private endpoints. Avoid running production databases as ordinary Deployments.

Level 3: Advanced

Performance and scalability

  • Set memory and CPU requests accurately: the scheduler bin-packs on requests. Over-requesting wastes money; under-requesting causes noisy neighbours.
  • Prefer a memory limit and consider leaving CPU limit unset to avoid throttling of latency-sensitive .NET services. Remember .NET sizes its thread pool and GC heaps from container limits, so set them deliberately.
  • Use HPA for pods and the cluster autoscaler (or Karpenter-based node autoprovisioning on AKS) for nodes. KEDA scales on queue length, e.g. Service Bus messages waiting for the claims-processing worker.
  • Use PodDisruptionBudgets and topology spread constraints so node upgrades and zone failures do not remove all replicas at once.

Security

  • Non-root containers, read-only root filesystem, dropped capabilities.
  • Namespaces + RBAC per team; NetworkPolicies default-deny between services.
  • Secrets from Azure Key Vault via the Secrets Store CSI driver, and Microsoft Entra Workload Identity instead of stored credentials (Day 39).
  • Admission policy (Azure Policy for Kubernetes / Pod Security Standards) to block privileged pods and untrusted registries.

Failure modes

  • Wrong readiness probe: pod receives traffic before the app is ready, or a slow dependency outage makes every pod unready and takes the whole service offline. Keep readiness about “can I serve”, not deep dependency checks of everything.
  • Liveness probe that checks the database: a DB blip restarts every pod (restart storm). Liveness should check only the process itself.
  • Missing graceful shutdown: on SIGTERM ASP.NET Core drains requests for HostOptions.ShutdownTimeout (default 30 s in .NET 8+). Align this with terminationGracePeriodSeconds.
  • :latest tags and no resource requests lead to unpredictable rollouts and evictions.
  • Stateful services with local disk written into pods that get rescheduled.

Common mistakes

  • Adopting Kubernetes before having CI/CD, observability and container basics.
  • One giant cluster with no namespace or quota boundaries.
  • Treating YAML as a copy-paste artefact instead of versioned, reviewed code (GitOps).
  • Ignoring the upgrade cadence: Kubernetes minor versions have a limited support window.

Level 4: Expert and Architect view

OptionOps effortControlScale-to-zeroBest fitMain drawback
Kubernetes (AKS)High (Medium with AKS Automatic)Very highWith KEDA / add-onsMany services, platform team, custom networkingComplexity, needs platform skills
Azure Container AppsLowMediumYesMicroservices without Kubernetes admin, event-drivenLess low-level control, fewer Kubernetes features exposed
Azure App ServiceLowLow-mediumNoA handful of web apps and APIsCoarse scaling, not container-scheduler semantics
Azure Functions (Day 45)Very lowLowYesEvent-driven, short tasksExecution limits, cold starts on some plans
Service per VM (Day 43)HighHigh isolationNoCompliance / hypervisor isolationSlow, costly, manual scaling
Docker Swarm / NomadMediumMediumNoSmaller estatesSmaller ecosystem, less Azure integration

Combines with: Service per Container (Day 42), Sidecar (Day 47), Service Mesh (Day 48), Health Check API (Day 36), Externalized Configuration (Day 39), Service Registry patterns (Days 21-25; Kubernetes Services and DNS act as built-in server-side discovery), Circuit Breaker and Retry (Days 26, 28), Log Deployments & Changes (Day 37).

ADR (architecture review)

  • Title: ADR-046 Adopt AKS as the service deployment platform for the claims system
  • Status: Proposed
  • Context: 14 services across 5 teams, weekly-to-daily releases, seasonal claim spikes (10x), 99.9% availability target, existing .NET 8+/Angular skills, Azure-first strategy.
  • Decision: Run all containerised services on AKS (Standard tier, three availability zones), with a platform team owning cluster operations, GitOps (Flux or Argo CD) for deployment, and namespaces per team. Evaluate Azure Container Apps for low-traffic edge services.
  • Consequences (positive): self-healing, uniform deployment contract, autoscaling, portable across clouds at the manifest level.
  • Consequences (negative): platform team headcount, Kubernetes learning curve, upgrade duty every few months, cost of always-on system node pool.
  • Alternatives considered: Container Apps (rejected for now: need custom network policy and mesh options), App Service (rejected: 14 services with per-service scaling profiles and sidecars).
  • Review trigger: revisit if service count drops below ~6 or no platform team can be staffed.

Azure implementation

Services

  • Azure Kubernetes Service (AKS): managed control plane. Options include standard AKS and AKS Automatic (opinionated, more automated node and add-on management, Standard tier by default).
  • Azure Container Registry (ACR): private image storage, attached to AKS with managed identity (--attach-acr).
  • Azure Container Apps: a simpler managed platform (built on Kubernetes, without exposing its API) for teams who want less operational surface.
  • Supporting: Azure Key Vault, Azure Monitor / Container Insights + Managed Prometheus + Managed Grafana, Application Insights (OpenTelemetry), Microsoft Defender for Containers, Azure Policy, Azure Application Gateway for Containers or NGINX / Gateway API ingress, Azure SQL / PostgreSQL Flexible Server via private endpoints.

Configuration (example)

Terminal window
az group create -n rg-claims-prod -l westeurope
az acr create -g rg-claims-prod -n contosoinsurance --sku Standard
az aks create \
-g rg-claims-prod -n aks-claims-prod \
--tier standard \
--node-count 3 --zones 1 2 3 \
--enable-managed-identity \
--enable-oidc-issuer --enable-workload-identity \
--enable-cluster-autoscaler --min-count 3 --max-count 10 \
--attach-acr contosoinsurance \
--network-plugin azure --network-plugin-mode overlay
az aks get-credentials -g rg-claims-prod -n aks-claims-prod

Pricing and tier considerations (verified against Microsoft Learn; check the Azure pricing page for current per-hour rates in your region)

  • Free tier: no management charge and no financially backed SLA; up to 1,000 nodes. Use for dev/test and learning.
  • Standard tier: financially backed uptime SLA of 99.95% with availability zones (99.9% without); up to 5,000 nodes. Use for production. This is a per-cluster hourly management fee on top of node costs.
  • Premium tier: Standard features plus 24-month Long Term Support for Kubernetes versions (--k8s-support-plan AKSLongTermSupport), for regulated estates that cannot upgrade yearly.
  • You always pay for the worker VMs, disks, load balancers, public IPs, egress and monitoring ingestion. Node VMs are typically the dominant cost: use right-sized SKUs, spot node pools for non-critical batch workers, reserved instances / savings plans for steady baseline, and cluster start/stop for non-production.
  • Log Analytics ingestion from Container Insights can become a surprise cost; tune collection and use Managed Prometheus for metrics.

Reference architecture (text) Users reach Azure Front Door (WAF, CDN). Front Door routes to an ingress (Application Gateway for Containers) in front of an AKS Standard cluster spread over three zones. A system node pool runs cluster components; user node pools run the Angular portal (nginx), Claims API, Payments, Notifications and worker services in team namespaces. Images come from ACR. Pods use Workload Identity to fetch secrets from Key Vault and to reach Azure SQL and PostgreSQL over private endpoints, and to send/receive messages via Azure Service Bus. KEDA scales workers on Service Bus queue depth. OpenTelemetry sends traces to Application Insights, metrics go to Managed Prometheus/Grafana, and Flux syncs manifests from a Git repository. Defender for Containers and Azure Policy enforce guardrails.

Teaching guide for my team

2-minute beginner explanation: “You used to log in to a server and start your app. Now you write a small file that says: run 3 copies of this container, keep them healthy, and use this much memory. Kubernetes reads the file and does the work. If one copy crashes, it starts another. If we ship a new version, it swaps copies one by one so users see no downtime. We describe what we want; the platform figures out how.”

5-minute intermediate explanation: Walk through Deployment, ReplicaSet, Pod, Service, Ingress, ConfigMap/Secret. Explain the reconcile loop, requests vs limits, readiness vs liveness, rolling updates with maxSurge/maxUnavailable, and HPA. Show the claims-api YAML from Level 2 and point out how each field prevents one failure from Section 4.

Hands-on exercise

  1. Create a local cluster (kind or minikube) and deploy the claims-api manifest with 3 replicas.
  2. Delete one pod: kubectl delete pod <name>.
  3. Change the image tag to a new version and watch kubectl rollout status deployment/claims-api.
  4. Deploy a version whose /health/ready returns 503 and observe that traffic never reaches it and the rollout stalls, then run kubectl rollout undo.

Expected outcome: the replacement pod appears within seconds; the good rollout completes with no failed requests; the broken rollout stays halted with old pods still serving, and undo restores the previous version.

Interview questions

  1. What is the difference between a liveness and a readiness probe? Readiness controls whether a pod receives traffic; liveness controls whether the container is restarted.
  2. Why set resource requests? The scheduler uses requests to place pods and autoscalers use them for utilisation; without them, placement and scaling are unpredictable.
  3. When would you choose Container Apps over AKS? When you want microservices, autoscaling and revisions without operating Kubernetes, and do not need direct Kubernetes API or custom cluster-level networking.

Mastery checklist

  • I can explain desired state vs actual state and the reconcile loop.
  • I can write a Deployment, Service, Ingress, and HPA from scratch and explain every field.
  • I can design readiness and liveness probes that avoid restart storms.
  • I can perform, pause and roll back a rolling update with zero downtime.
  • I can explain the AKS Free, Standard and Premium tiers and choose one for a workload.
  • I can secure a workload with Workload Identity, Key Vault, RBAC and NetworkPolicy.
  • I can justify AKS vs Container Apps vs App Service in an ADR.
  • I can diagnose a CrashLoopBackOff or Pending pod using kubectl describe and logs.

Key takeaway

A service deployment platform turns deployment from a manual procedure into a declared, self-correcting system: you state the desired result, and the orchestrator schedules, scales, rolls out and heals. Adopt it when service count and release pace justify its operating cost, and pick the lowest-effort Azure option that meets your needs.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 48: Service Mesh

A service mesh is an infrastructure layer that moves service-to-service networking concerns (mutual TLS, retries, timeouts, traffic routing, telemetry, authorization) out of your application code and into a fleet of...

Manikandan
Manikandan·20 min read
Microservices

Day 47: Sidecar

A sidecar is a helper container that is deployed and scaled together with your application container, in the same pod (Kubernetes) or the same replica (Azure Container Apps).

Manikandan
Manikandan·14 min read
Microservices

Day 45: Serverless Deployment

Serverless deployment means you ship only your code and its triggers (an HTTP call, a queue message, a timer) and let the cloud platform decide where it runs, how many copies run, and when they are switched off.

Manikandan
Manikandan·17 min read
Microservices

Day 44: Multiple Services per Host

Multiple Services per Host means running several independently built and versioned services on the same server (VM, App Service plan, or physical machine) instead of giving each its own VM or container host.

Manikandan
Manikandan·14 min read