Service per Container means every microservice is packaged as its own OCI (Docker) image and runs as one container (or a set of identical container replicas) with its own process, filesystem, and resource limits.
Intro
Service per Container means every microservice is packaged as its own OCI (Docker) image and runs as one container (or a set of identical container replicas) with its own process, filesystem, and resource limits. Instead of installing a full operating system per service, you ship only the app plus its runtime and dependencies, and the container runtime shares the host kernel. Containers start in seconds, are identical from laptop to production, and can be scaled, replaced, and rolled back independently.
Versions used in this lesson: .NET 10 (current LTS, released November 2025), Angular 22 (current major, released June 2026), Docker/OCI images, Azure Container Registry, Azure Container Apps, and Azure Kubernetes Service.
Why we need this
Business reasons
- Claims volume is spiky (storms, floods, month-end). The business wants to add capacity for ClaimsIntake in minutes without touching Payments.
- Releases must be frequent and low-risk. A fraud-scoring fix should not require a full-system release window.
- Infrastructure cost must follow load. Paying for a dedicated VM per small service is hard to justify.
Technical reasons
- A VM carries a full guest OS: boot takes minutes, patching is per-VM, and the disk image is many GB. A container image for an ASP.NET Core service is typically tens to a few hundred MB and starts in seconds.
- “Works on my machine” disappears when the exact runtime, libraries, and config layout ship inside the image.
- Each service can use its own runtime version (ClaimsIntake on .NET 10 while a legacy adjuster-report service still runs on .NET 8) without conflicts on a shared host.
- Containers give a uniform unit for orchestrators (Kubernetes, Azure Container Apps) to schedule, health-check, restart, and scale.
What problem it solves
Problem (from the topic list): VMs are slow to start and carry heavy OS overhead.
What goes wrong without it (claims system example):
- ClaimsIntake, PolicyLookup, FraudScoring, and Payments each get a VM. During a hailstorm, intake needs 10 more instances. Each new VM takes 3 to 8 minutes to boot, configure, and pass health checks; the queue of unprocessed claims grows meanwhile.
- The four VMs are each 10-20% utilised, but you pay for 100% of them.
- The Payments VM has .NET 8 and a specific OpenSSL configuration; the FraudScoring team installs a different runtime patch on their VM and “configuration drift” begins. A bug reproduces only in production.
- Deployments are scripts that RDP/SSH into machines and copy files. Rollback means re-copying the old files and hoping nothing else changed.
With Service per Container, each service is an immutable, versioned image. Deploy = start new containers from a new tag. Rollback = start containers from the old tag.
When it is needed (and when it is NOT)
Good fit
- You already have, or are moving to, more than a handful of independently deployed services.
- You need fast, automated scale-out and scale-in.
- You want the same artifact in dev, test, and production.
- Teams use different runtimes or dependency versions.
- You plan to use an orchestrator (Kubernetes, Azure Container Apps, Azure App Service for Containers).
Not a good fit / overkill
- A single monolith with one team and low traffic: an Azure App Service (code deployment) is simpler and cheaper than building container pipelines you do not need.
- Workloads needing hypervisor-level isolation for hostile multi-tenant code: containers share the host kernel. Consider VM-level isolation (Day 43) or hardened sandboxes (for example, Azure Container Apps dynamic sessions or Kata-based isolation options in AKS).
- Software that needs a GUI, kernel modules, or licensing tied to VM hardware IDs.
- Stateful databases run naively in containers without persistent volume design. Prefer managed services (Azure SQL, Azure Database for PostgreSQL) in the claims system.
- Very small teams without container skills where an extra layer (Dockerfiles, registries, image scanning) adds more operational load than value.
How to identify the problem (key signals)
- Scale-out takes minutes. Autoscale rules fire, but new capacity arrives after the spike is over.
- VM CPU averages under 20% while the bill shows one VM per service.
- “Works in test, fails in prod” bugs caused by different runtime versions, OS patches, or environment variables.
- Deployment runbooks with manual steps (copy files, edit config, restart IIS/systemd) and a long “release window.”
- Two services on the same VM fight over ports, memory, or a shared runtime, so someone proposes a “bigger VM.”
- Onboarding a new developer takes days of installing SQL, SDKs, and tools instead of
docker compose up. - Patching the OS forces coordinated downtime of all services on that VM.
Flow Diagram
A multi-stage build produces a small immutable image that runs as one container per service.
flowchart LR SRC["Source"] --> SDK["Stage 1: sdk:10.0 build + publish"] SDK --> RT["Stage 2: aspnet:10.0 runtime, non-root"] RT --> IMG["claims-intake:1.0.0"] IMG --> ACR["Azure Container Registry"] ACR --> ACA["Container Apps revision"] ACA --> R1["Replica"] ACA --> R2["Replica"] ACA -- "rollback = old tag" --> ACRLevel 1: Beginner
Analogy: A VM is like renting a whole apartment (kitchen, plumbing, meters) for each tenant. A container is a furnished room in a shared building: the building’s plumbing and electricity (the host kernel) are shared, but each tenant has a locked door, their own furniture (files and libraries), and a usage limit.
Key terms
- Image: immutable, layered template (read-only). Built from a Dockerfile.
- Container: a running instance of an image.
- Registry: storage for images (Azure Container Registry).
- Tag / digest: version label / content hash of an image.
Minimal working example. A tiny ClaimsIntake API.
ClaimsIntake.Api.csproj
<Project Sdk="Microsoft.NET.Sdk.Web"> <PropertyGroup> <TargetFramework>net10.0</TargetFramework> <Nullable>enable</Nullable> <ImplicitUsings>enable</ImplicitUsings> </PropertyGroup></Project>Program.cs
var builder = WebApplication.CreateBuilder(args);var app = builder.Build();
app.MapGet("/", () => "ClaimsIntake is running");
app.MapPost("/claims", (NewClaim claim) => Results.Accepted($"/claims/{Guid.NewGuid()}", new { claim.PolicyNumber, Status = "Received" }));
app.Run();
record NewClaim(string PolicyNumber, decimal Amount, string Description);Dockerfile (multi-stage)
# Stage 1: build with the SDK imageFROM mcr.microsoft.com/dotnet/sdk:10.0 AS buildWORKDIR /srcCOPY ClaimsIntake.Api.csproj .RUN dotnet restoreCOPY . .RUN dotnet publish -c Release -o /app/publish --no-restore
# Stage 2: run with the small ASP.NET runtime imageFROM mcr.microsoft.com/dotnet/aspnet:10.0 AS finalWORKDIR /appCOPY --from=build /app/publish .# .NET 8+ images listen on 8080 and include a non-root user named "app"USER $APP_UIDENTRYPOINT ["dotnet", "ClaimsIntake.Api.dll"]Run it:
docker build -t claims-intake:1.0.0 .docker run --rm -p 8080:8080 claims-intake:1.0.0curl -X POST http://localhost:8080/claims -H "Content-Type: application/json" \ -d '{"policyNumber":"P-1001","amount":2500,"description":"Hail damage"}'Note: since .NET 8 the official images listen on port 8080 (not 80) and run non-root friendly. The .NET 10 aspnet:10.0 tag is based on Ubuntu 24.04 (Debian images are not published for .NET 10).
Level 2: Intermediate
One container per service in a real stack
The claims system has ClaimsIntake (.NET), PolicyLookup (.NET), an Angular adjuster portal, and PostgreSQL for local development. Each app is one container; the database is a container only locally (use managed Azure Database for PostgreSQL in production).
docker-compose.yml
services: claims-intake: build: ./ClaimsIntake ports: ["8081:8080"] environment: ConnectionStrings__Claims: "Host=claims-db;Database=claims;Username=claims;Password=dev-only-password" PolicyLookup__BaseUrl: "http://policy-lookup:8080" depends_on: claims-db: condition: service_healthy deploy: resources: limits: cpus: "1.0" memory: 512M
policy-lookup: build: ./PolicyLookup ports: ["8082:8080"]
adjuster-portal: build: ./AdjusterPortal ports: ["4200:8080"]
claims-db: image: postgres:17 environment: POSTGRES_DB: claims POSTGRES_USER: claims POSTGRES_PASSWORD: dev-only-password volumes: - claims-data:/var/lib/postgresql/data healthcheck: test: ["CMD-SHELL", "pg_isready -U claims -d claims"] interval: 5s timeout: 3s retries: 10
volumes: claims-data:The password above is a local development placeholder only. In shared environments use Key Vault (Day 39).
Configuration and health inside the container
Configuration comes from environment variables (12-factor). ASP.NET Core maps ConnectionStrings__Claims to ConnectionStrings:Claims. Add health endpoints so the orchestrator can check the container (Day 36).
using Microsoft.EntityFrameworkCore;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddDbContext<ClaimsDbContext>(o => o.UseNpgsql(builder.Configuration.GetConnectionString("Claims")));
builder.Services.AddHealthChecks() .AddDbContextCheck<ClaimsDbContext>("db", tags: new[] { "ready" });
var app = builder.Build();
// Liveness: process is up. No dependency checks, so a DB outage does not restart every pod.app.MapHealthChecks("/health/live", new() { Predicate = _ => false });// Readiness: safe to receive traffic.app.MapHealthChecks("/health/ready", new() { Predicate = r => r.Tags.Contains("ready") });
app.MapPost("/claims", async (NewClaim claim, ClaimsDbContext db) =>{ var entity = new Claim { PolicyNumber = claim.PolicyNumber, Amount = claim.Amount, Description = claim.Description }; db.Claims.Add(entity); await db.SaveChangesAsync(); return Results.Created($"/claims/{entity.Id}", entity.Id);});
app.Run();
record NewClaim(string PolicyNumber, decimal Amount, string Description);
class Claim{ public Guid Id { get; set; } = Guid.NewGuid(); public string PolicyNumber { get; set; } = ""; public decimal Amount { get; set; } public string Description { get; set; } = "";}
class ClaimsDbContext(DbContextOptions<ClaimsDbContext> options) : DbContext(options){ public DbSet<Claim> Claims => Set<Claim>();}This needs the packages Microsoft.EntityFrameworkCore, Npgsql.EntityFrameworkCore.PostgreSQL, and Microsoft.Extensions.Diagnostics.HealthChecks.EntityFrameworkCore.
Angular in its own container
The Angular portal is static files. Build with Node, serve with a non-root nginx image on port 8080.
FROM node:24-alpine AS buildWORKDIR /appCOPY package*.json ./RUN npm ciCOPY . .RUN npm run build -- --configuration production
FROM nginxinc/nginx-unprivileged:stable-alpine# Angular's application builder outputs to dist/<project-name>/browserCOPY --from=build /app/dist/adjuster-portal/browser /usr/share/nginx/htmlCOPY nginx.conf /etc/nginx/conf.d/default.confEXPOSE 8080nginx.conf (SPA fallback so deep links like /claims/123 work)
server { listen 8080; root /usr/share/nginx/html; location / { try_files $uri $uri/ /index.html; }}Use a Node version supported by your Angular version (check the Angular version compatibility table); the example uses Node 24, an LTS line. Runtime configuration such as the API base URL should not be baked into the image; load it from a config.json fetched at startup, or route via a gateway (Day 19), so one image works in every environment.
Building without a Dockerfile (option)
The .NET SDK can produce an image directly:
dotnet publish ./ClaimsIntake.Api -c Release /t:PublishContainer \ -p:ContainerRepository=claims-intake -p:ContainerImageTags='"1.0.0;latest"'This is convenient for simple services. Keep a Dockerfile when you need OS packages, custom users, or multi-step builds.
Level 3: Advanced
Image size, security, and startup
| Concern | Practice |
|---|---|
| Size | Multi-stage builds; runtime image only in final stage; consider chiseled Ubuntu images (10.0-noble-chiseled), which are distroless with no shell or package manager |
| Attack surface | Run as non-root (USER $APP_UID), read-only root filesystem where possible, drop Linux capabilities, no secrets in image layers |
| Supply chain | Pin base images by digest, scan images (Microsoft Defender for Containers, Trivy), sign images, rebuild regularly to pick up base image patches |
| Startup | Keep startup work small; ReadyToRun or Native AOT can cut cold start for scale-to-zero services (AOT restricts reflection-heavy libraries) |
| Reproducibility | Never rely on latest in production; deploy immutable version tags or digests |
Chiseled example (note: no shell, so no curl-based Docker HEALTHCHECK; use orchestrator probes):
FROM mcr.microsoft.com/dotnet/aspnet:10.0-noble-chiseled AS finalWORKDIR /appCOPY --from=build /app/publish .USER $APP_UIDENTRYPOINT ["dotnet", "ClaimsIntake.Api.dll"]Resource limits and .NET
- Set CPU and memory limits. .NET reads container limits: the GC heap hard limit defaults to 75% of the container memory limit, and
Environment.ProcessorCountreflects the CPU limit (rounded up). A very low CPU limit (for example 0.25) can cause thread pool starvation and slow startup; measure before choosing. - Fractional CPU limits with a synchronous, blocking code path are a common cause of “slow only in the container” complaints.
Failure modes
- OOMKilled (exit code 137): memory limit too low or a leak. Watch working set metrics and container restart count.
- Crash loops: bad config, or a liveness probe that depends on the database so a DB outage restarts all replicas.
- Slow shutdown: container receives SIGTERM; if the app ignores it, the orchestrator SIGKILLs after the grace period (30s default in Kubernetes) and in-flight claims requests fail. ASP.NET Core handles SIGTERM via the host; ensure
HostOptions.ShutdownTimeoutfits inside the grace period. - Ephemeral filesystem: anything written to the container’s writable layer is lost on restart. Uploaded claim photos belong in Azure Blob Storage, not the container disk.
- Noisy neighbour on the node: no limits means one container can starve others.
Common mistakes
- Multiple unrelated processes in one container (an anti-pattern; use sidecars, Day 47).
- Putting secrets in
ENVin the Dockerfile or in image layers. - Using
latest, so two replicas run different code. - Running as root.
- Giant images that include the SDK and build tools.
- Storing state on local disk.
- Not adding a
.dockerignore, which sendsbin/,obj/, and.gitinto the build context.
Level 4: Expert and Architect view
Alternatives compared
| Option | Isolation | Start time | Density / cost | Ops effort | Best for |
|---|---|---|---|---|---|
| Service per Container (this pattern) | Process + namespaces, shared kernel | Seconds | High | Medium (needs registry, orchestrator) | Most microservices |
| Service per VM (Day 43) | Hypervisor | Minutes | Low | Medium to high (patching per VM) | Strong isolation, licensing tied to VMs, legacy apps |
| Multiple Services per Host (Day 44) | Shared process space / OS | Fast | Highest, but blast radius large | Low initially | Few small services, tight budgets |
| Serverless (Day 45) | Platform managed | Cold start possible | Pay per execution, scale to zero | Low | Event-driven, spiky workloads |
| PaaS code deploy (App Service, no container) | Platform managed | Seconds to minutes | Medium | Low | Simple web apps and APIs |
Patterns it combines with
- Service Deployment Platform (Day 46): containers need an orchestrator for scheduling, rollout, and healing.
- Health Check API (Day 36): probes are how the orchestrator judges each container.
- Externalized Configuration (Day 39): one image, many environments.
- Sidecar (Day 47) and Service Mesh (Day 48): cross-cutting concerns without stuffing them into the app image.
- Database per Service (Day 5): each container owns its data store; the container itself stays stateless.
- Service Template (Day 41) and Microservice Chassis (Day 40): the Dockerfile, CI pipeline, and base image policy live in the template.
ADR (architecture review)
ADR-042: Package every claims microservice as one container image
- Status: Proposed
- Context: The claims platform has 6 services with different scaling profiles. Intake must scale 10x during catastrophe events; Payments must remain stable and auditable. Current VM-based hosting takes 5+ minutes to add capacity and averages under 20% CPU utilisation.
- Decision: Package each service (and the Angular portal) as its own OCI image, one main process per container, built by CI from a shared service template, stored in Azure Container Registry, and deployed to Azure Container Apps. Databases stay on managed Azure services.
- Consequences (positive): faster scale-out, consistent environments, independent rollback, higher density.
- Consequences (negative): new skills (image hygiene, registry, probes), a shared-kernel isolation model, image patching becomes our responsibility (rebuild on base image updates), and distributed troubleshooting needs tracing and central logs (Days 31, 32).
- Alternatives considered: Service per VM (rejected: slow scale, cost), Multiple Services per Host (rejected: blast radius on Payments), App Service without containers (acceptable for the portal, but two deployment models add cost).
- Revisit when: more than about 20 services or a need for custom networking/operators, at which point evaluate AKS.
Azure implementation
Services
| Need | Azure service |
|---|---|
| Image storage | Azure Container Registry (ACR) |
| Run containers, simplest path | Azure Container Apps (ACA) |
| Run containers, full Kubernetes control | Azure Kubernetes Service (AKS) |
| Single web container, PaaS style | Azure App Service (Web App for Containers) |
| Short jobs / burst | Azure Container Instances, ACA Jobs |
| Image scanning / runtime protection | Microsoft Defender for Containers |
| Logs, metrics, traces | Azure Monitor, Log Analytics, Application Insights |
| Secrets | Azure Key Vault |
| Build | Azure DevOps or GitHub Actions; ACR Tasks for in-registry builds |
Configure (Azure CLI)
RG=rg-claims-prodLOC=westeuropeACR=claimsacr$RANDOMENV=cae-claims-prod
az group create -n $RG -l $LOC
# Registry (Standard is a common production starting tier)az acr create -g $RG -n $ACR --sku Standard
# Build and push the image inside ACR (no local Docker needed)az acr build -r $ACR -t claims-intake:1.0.0 ./ClaimsIntake
# Container Apps environment and appaz containerapp env create -g $RG -n $ENV -l $LOC
az containerapp create -g $RG -n claims-intake --environment $ENV \ --image $ACR.azurecr.io/claims-intake:1.0.0 \ --registry-server $ACR.azurecr.io --registry-identity system \ --target-port 8080 --ingress external \ --cpu 0.5 --memory 1.0Gi \ --min-replicas 1 --max-replicas 20 \ --scale-rule-name http-load --scale-rule-type http --scale-rule-http-concurrency 50Key configuration points:
- Use a managed identity with the
AcrPullrole for pulling images, not admin credentials. - Configure probes (
/health/live,/health/ready) in the container app YAML. - Use revisions for safe rollout: keep multiple revisions active and split traffic (for example 10% to the new revision) before full cutover.
- Pull secrets from Key Vault references rather than plain environment variables.
- Scale rules: HTTP concurrency, CPU/memory, or event sources via KEDA (for example Azure Service Bus queue length for a claims-processing worker).
Pricing and tier considerations
Exact unit prices vary by region and change over time, so verify in the Azure pricing calculator before committing. Structure that is stable and worth knowing:
- ACR charges a flat daily fee per registry by tier, with included storage: Basic 10 GiB, Standard 100 GiB, Premium 500 GiB, plus a daily charge for storage above that. Premium adds geo-replication, private link, content trust, customer-managed keys, retention policies, and artifact streaming. Choose Premium when you need private endpoints or multi-region replication.
- Azure Container Apps Consumption plan bills per second for vCPU and memory, plus requests. Each subscription gets a monthly free grant of 180,000 vCPU-seconds, 360,000 GiB-seconds, and 2 million requests. Replicas scaled to zero cost nothing; replicas kept at a minimum count while idle are billed at a lower idle rate. Health probe requests are not billable.
- Azure Container Apps Dedicated plan (workload profiles) bills per workload profile instance rather than per app, plus a management fee when any dedicated profile exists. Use it for steady, predictable load, GPU, or larger instance sizes.
- AKS: the control plane has a free tier (no uptime SLA) and a paid tier with an SLA; you pay for the node VMs. Higher operational effort, best when you need Kubernetes-native tooling and many services.
- Reserve Azure Savings Plan or reservations for steady baseline compute.
Reference architecture (text)
- Developers push to Git; a GitHub Actions or Azure DevOps pipeline restores, tests, and runs
az acr build(or docker build + push) to produceclaims-intake:<git-sha>in ACR. Defender for Containers scans the image. - The pipeline deploys a new revision to the Azure Container Apps environment inside a virtual network.
- Azure Front Door (with WAF) or Application Gateway receives internet traffic and forwards to the public-facing apps: the Angular portal container and the API gateway container (Day 19). Internal services (PolicyLookup, FraudScoring, Payments) use internal ingress only.
- Each service reads secrets via managed identity from Key Vault and connects to its own Azure Database for PostgreSQL or Azure SQL (Day 5). Files go to Blob Storage.
- Asynchronous work uses Azure Service Bus; KEDA scales the worker containers on queue length.
- All containers emit OpenTelemetry traces and logs to Application Insights and Log Analytics; alerts watch restart count, 5xx rate, and p95 latency.
Teaching guide for my team
Explain to a beginner in 2 minutes
“A container is a sealed box that has our app plus exactly the libraries it needs. We build the box once, give it a version number, and run the same box on your laptop, in test, and in production. It starts in seconds because it does not boot a whole operating system; it shares the host’s kernel. One service per box means we can start ten more ClaimsIntake boxes during a storm without touching Payments. If a new version is bad, we start the old box again.”
Explain to an intermediate developer in 5 minutes
Cover these in order: (1) image vs container vs registry; (2) multi-stage Dockerfile and why the runtime image is separate from the SDK image; (3) one main process per container, config through environment variables, state outside the container; (4) resource limits and how .NET reads them; (5) liveness vs readiness probes and why liveness must not call the database; (6) immutable tags and digests, revisions, and rollback; (7) the security baseline: non-root, no secrets in layers, scan and rebuild images; (8) where this stops being enough and you need an orchestrator (Day 46).
Hands-on exercise
Task: Containerise the ClaimsIntake API and prove the container behaves correctly under limits.
- Build the Level 1 API and its Dockerfile. Confirm the image runs as non-root:
docker run --rm --entrypoint id claims-intake:1.0.0(use the non-chiseled image, because chiseled has no shell tools). - Add
/health/liveand/health/ready. - Run with
--memory=256m --cpus=0.5and hit/claims200 times with a loop. - Change the image tag to
1.0.1with a small change, run both versions on different ports, and switch traffic by changing which port the client calls (simulated rollback). - Compare image sizes for
aspnet:10.0andaspnet:10.0-noble-chiseledwithdocker images.
Expected outcome: the container runs as a non-root user, both health endpoints return 200, the app stays within its memory limit, 1.0.0 remains runnable after 1.0.1 exists (rollback works), and the chiseled variant is noticeably smaller than the standard runtime image.
Interview-style questions
- Why is the liveness probe usually not allowed to check the database? If the database is down, every replica would fail liveness and be restarted repeatedly, making the outage worse. Liveness answers “is the process stuck”; readiness answers “can it serve traffic now”.
- What is the difference between a VM and a container in terms of isolation? A VM has its own guest OS kernel under a hypervisor, so isolation is strong. A container shares the host kernel and is isolated with namespaces and cgroups, so it starts faster and is lighter but has a larger shared attack surface.
- A container keeps restarting with exit code 137. What do you check? 137 is SIGKILL, most often OOMKilled. Check the memory limit against the working set, look for leaks, and check whether the GC heap limit and the container limit are consistent. Confirm through the orchestrator’s events and Azure Monitor metrics.
Mastery checklist
- I can write a multi-stage Dockerfile for a .NET 10 service and explain each stage.
- I can run the container as non-root and explain why it matters.
- I know how configuration and secrets reach a container without being baked into the image.
- I can explain liveness vs readiness and design correct probes for a service with a database dependency.
- I can set CPU and memory limits and predict how .NET reacts to them.
- I can push an image to ACR and deploy it to Azure Container Apps using managed identity for pulls.
- I can roll back by redeploying a previous immutable tag or revision.
- I can justify, in an ADR, when containers are the right choice versus VMs, App Service, or serverless.
Key takeaway
Package each service as its own small, immutable, non-root container image so it starts in seconds, scales on its own, and runs identically everywhere; keep state, secrets, and cross-service concerns outside the container.
