Manikandan — Manikandan
Microservices

Day 43: Service per VM

ManikandanManikandan
17 min read·Updated Sep 21, 2022

Summary: "Service per VM" means every microservice runs on its own dedicated virtual machine (or its own small group of identical VMs behind a load balancer).

Intro

Summary: “Service per VM” means every microservice runs on its own dedicated virtual machine (or its own small group of identical VMs behind a load balancer). Nothing else shares that machine’s operating system, CPU, memory, disk or network identity. It gives the strongest isolation and the most predictable performance of the common hosting styles, at the price of slower start-up, heavier patching work and a higher bill. In our insurance claims system, the Payments service and the Document-OCR service would each get their own VMs so that a memory-hungry OCR job can never starve payment processing.


Why we need this

  • Hard isolation. A hypervisor boundary is a much stronger wall than a container boundary. Containers share the host kernel; VMs do not. When services handle regulated data (payment card data, health-related claim attachments) auditors often prefer, or require, separate OS instances.
  • Guaranteed resources. A VM with 4 vCPUs and 16 GB RAM has exactly that, with no noisy neighbour from another service on the same OS. Capacity planning becomes simple arithmetic.
  • Independent OS-level needs. One service needs Windows Server with a legacy COM component or a specific .NET Framework 4.8 runtime; another needs Linux with a native library like Tesseract. Per-service VMs let each pick its own OS, patch level, kernel settings and installed agents.
  • Simple failure and blast-radius model. If the Claims-Intake VM crashes or is compromised, the Payments VM is a different machine, disk and network interface.
  • Familiar operations. Many enterprise teams already have VM tooling: image pipelines, patching, backup, monitoring agents, firewall rules. Reusing it is often the fastest path from a monolith to services.

What problem it solves

Problem (from the topic list): you need hypervisor-level isolation and guaranteed resources for a service.

Without it, in a “multiple services per host” setup:

  • A memory leak in the OCR service triggers the OS out-of-memory killer, which may kill the Payments process on the same host.
  • One service’s CPU spike raises latency for every other service on that OS.
  • Two services need different versions of the same runtime or a conflicting native library, so deployments break each other.
  • A vulnerability in one service gives an attacker a foothold on the same OS as the sensitive service.
  • A single reboot for patching takes down unrelated business functions.

When it is needed (and when it is NOT)

Good fit

  • Strict compliance or tenancy isolation (PCI-DSS scope reduction: only the Payments VM is in scope, for example).
  • Services with very different or large resource profiles (OCR, fraud scoring, reporting) that must not compete.
  • Legacy or stateful workloads that cannot be containerised easily (Windows services, COM, licensed software bound to a machine, specialised drivers, GPU or large-memory machines).
  • Small number of services (roughly 3 to 15) where the VM count is manageable.
  • Teams with strong VM operations skills and no Kubernetes skills yet.

Poor fit / overkill

  • Dozens or hundreds of small services: VM count, patching effort and cost explode.
  • Services that are idle most of the day: a whole VM is billed even at 3% CPU.
  • Need for fast scale-out in seconds: a VM takes minutes to boot and configure; containers take seconds.
  • Teams that want rapid, frequent deployments with minimal infrastructure overhead: consider containers on Azure Container Apps or AKS, or App Service.
  • Workloads whose isolation needs are met by containers (most business web APIs).

How to identify the problem (key signals)

  1. Noisy-neighbour incidents: latency graphs of two unrelated services rise together when only one received extra traffic (shared host CPU or memory).
  2. OOM kills in logs: Out of memory: Killed process ... in journalctl or Windows Event ID 2004 (resource exhaustion) naming a process that is not the one that leaked.
  3. Deployment collisions: “we can’t upgrade the runtime for service A because service B on the same box needs the old one.”
  4. Audit findings: the auditor asks why a low-trust service shares an OS with the service that handles card or bank data.
  5. Blast radius surprises: one host reboot or one compromised account takes down several business capabilities.
  6. Capacity guesswork: nobody can say how much CPU or RAM a given service really uses because the host graph mixes everyone.
  7. Inverse signals (you have over-applied the pattern): 20+ VMs each under 10% average CPU, patch Tuesday takes days, and deployment of a one-line fix takes 40 minutes of VM work.

Flow Diagram

Each service runs on its own immutable, zone-redundant VM scale set in its own subnet.

flowchart LR
FD["Front Door + WAF"] --> AG["Application Gateway"]
AG --> LB1["ILB intake"] --> V1["VMSS claims-intake - 3 zones"]
AG --> LB2["ILB payments"] --> V2["VMSS payments - PCI subnet"]
AG --> LB3["ILB OCR"] --> V3["VMSS OCR - memory optimised"]
GAL["Compute Gallery image"] -. "rolling upgrade" .-> V1
GAL -.-> V2
GAL -.-> V3

Level 1: Beginner

Analogy: an apartment block versus separate houses. In “multiple services per host” everyone lives in one apartment and shares the kitchen and bathroom. In “service per VM” each family has its own house: their own walls, kitchen, meter and front door. Noise and mess stay inside. The cost is that you now maintain many houses.

Minimal working example. A tiny ASP.NET Core API (target the current LTS, .NET 10) that will be the only thing on its VM. Package needed for Linux systemd integration: Microsoft.Extensions.Hosting.Systemd.

// Program.cs (ClaimsIntake.Api)
var builder = WebApplication.CreateBuilder(args);
builder.Host.UseSystemd(); // lets systemd track start/stop; no-op elsewhere
builder.Services.AddHealthChecks();
var app = builder.Build();
app.MapHealthChecks("/health");
app.MapPost("/claims", (NewClaim claim) =>
{
var id = Guid.NewGuid();
return Results.Created($"/claims/{id}", new { id, claim.PolicyNumber, status = "Received" });
});
app.Run();
public record NewClaim(string PolicyNumber, decimal Amount);

Publish it and run it as the VM’s one service. A systemd unit on Linux:

/etc/systemd/system/claims-intake.service
[Unit]
Description=Claims Intake API
After=network.target
[Service]
Type=notify
WorkingDirectory=/opt/claims-intake
ExecStart=/opt/claims-intake/ClaimsIntake.Api
Restart=always
RestartSec=5
User=claims
Environment=ASPNETCORE_URLS=http://0.0.0.0:8080
Environment=DOTNET_ENVIRONMENT=Production
[Install]
WantedBy=multi-user.target

Key idea: one VM = one service = one deployable unit. The VM is named after the service (vm-claims-intake-prod-01) so anyone can tell what is on it.

Level 2: Intermediate

In a real .NET + Angular + SQL Server/PostgreSQL system the shape is:

Browser (Angular SPA on Azure Static Web Apps or Storage static site)
-> Application Gateway / Front Door
-> Load Balancer -> VMs: claims-intake (2 instances, different zones)
-> Load Balancer -> VMs: claims-payments (2 instances, PCI subnet)
-> Load Balancer -> VMs: claims-ocr (memory-optimised)
each service -> its own database (Azure SQL / PostgreSQL Flexible Server)

Rules that make it work

  • Each service tier has its own subnet and NSG (network security group). Only the gateway subnet may reach the service port; only the Payments subnet may reach the payment database.
  • Services call each other by an internal load balancer DNS name, not by VM IP.
  • Configuration and secrets come from Key Vault via managed identity, never from files baked in the image.
  • Deployment is immutable: build a new VM image (or new scale set version) and replace instances, instead of editing servers by hand.

Reading a secret in .NET with the VM’s managed identity (packages: Azure.Identity, Azure.Extensions.AspNetCore.Configuration.Secrets):

using Azure.Identity;
var builder = WebApplication.CreateBuilder(args);
var vaultUri = builder.Configuration["KeyVault:Uri"]
?? throw new InvalidOperationException("KeyVault:Uri is not configured");
builder.Configuration.AddAzureKeyVault(new Uri(vaultUri), new DefaultAzureCredential());
builder.Services.AddHealthChecks();
var app = builder.Build();
app.MapHealthChecks("/health");
app.Run();

DefaultAzureCredential picks up the VM’s managed identity automatically. No password is stored on the machine.

Angular side. The Angular app never knows which VM serves it; it only calls the public gateway URL. Environment-specific base URLs belong in configuration, not in code:

// claims.service.ts (standalone-component era Angular, using inject())
import { Injectable, inject } from '@angular/core';
import { HttpClient } from '@angular/common/http';
import { Observable } from 'rxjs';
export interface NewClaim { policyNumber: string; amount: number; }
@Injectable({ providedIn: 'root' })
export class ClaimsService {
private readonly http = inject(HttpClient);
private readonly baseUrl = '/api/claims'; // same origin, routed by the gateway
submit(claim: NewClaim): Observable<{ id: string }> {
return this.http.post<{ id: string }>(this.baseUrl, claim);
}
}

Deployment flow (CI/CD)

  1. Build and test the .NET app.
  2. Bake an image: base OS + .NET runtime + the app + monitoring agent (Packer, or Azure VM Image Builder).
  3. Publish the image version to an Azure Compute Gallery.
  4. Update the scale set to the new image version with a rolling upgrade; health probes decide whether each instance is healthy before the next is replaced.

Level 3: Advanced

Performance and capacity

  • Right-size per service from real metrics (CPU, memory, disk IOPS, network). Start with measured 95th-percentile usage plus headroom, then adjust. A common waste is copying one VM size to every service.
  • Choose the VM family for the workload: general purpose for APIs, memory-optimised for OCR or in-memory rating, compute-optimised for scoring. Burstable (B-series) sizes suit low, spiky services such as internal admin tools; they use CPU credits and throttle when credits run out, so avoid them for steady production load.
  • Boot time is minutes. Do all installation at image build time, not at first boot, or scale-out will be slow and fragile.

Scalability

  • Scale a service by adding identical instances (horizontal). Use a Virtual Machine Scale Set (VMSS) rather than hand-built VMs. Prefer Flexible orchestration mode, which is the recommended mode for new scale sets.
  • Autoscale on a metric (CPU, queue length, requests) with a cooldown; scale-in policies must drain connections first.
  • The service must be stateless (session and files in Redis, Blob storage or the database), or scale-out and instance replacement lose data.

Security

  • Enable Trusted Launch (secure boot, vTPM) on supported generation 2 images.
  • Turn on encryption at host or disk encryption; keep secrets in Key Vault.
  • No public IPs on service VMs. Admin access through Azure Bastion or just-in-time access; disable password SSH.
  • Separate NSGs per service tier; deny by default.
  • Patch through Azure Update Manager or by replacing instances with a newly patched image (preferred: immutable).
  • Least-privilege managed identity per service, so the OCR VM cannot read the Payments Key Vault.

Failure modes

FailureEffectMitigation
Single VM per serviceHost failure = outage2+ instances across availability zones behind a load balancer
Config drift from manual fixes“Works on VM 1, not VM 2”Immutable images, no SSH fixes, drift detection
Slow autoscaleTraffic spike before new VM is readyPre-baked images, scale on leading indicators, keep minimum capacity
Full disk from logsService stopsLog rotation, ship logs off-box, disk alerts
Patch reboot of all instances at onceOutageRolling upgrade with health probes and max-unhealthy limits
Stateful data on local diskData lost on replaceExternalise state to managed services

Common mistakes

  • Putting two services on one VM “just for now” and never separating them, which defeats the pattern.
  • Treating VMs as pets: manual patching, hand edits, unique hostnames nobody can recreate.
  • One VM, no redundancy, because “it’s a small service”.
  • Sizing every VM the same.
  • Baking secrets or connection strings into the image.

Level 4: Expert and Architect view

Comparison of deployment styles

AspectService per VMMultiple services per hostService per containerServerless (Functions)Orchestrated (AKS / Container Apps)
IsolationStrong (hypervisor)Weak (shared OS)Medium (shared kernel)Platform-managedMedium; stronger with sandboxed pods or confidential options
Start timeMinutesSeconds (process)SecondsMilliseconds to seconds (cold start)Seconds
Resource efficiencyLow to mediumHighHighVery high (scale to zero)High
Ops effort per serviceHigh (OS, patching, images)MediumLow to mediumLowMedium (platform team needed)
Cost at low utilisationHighLowLowLowestLow to medium
Fits legacy Windows / COMYesYesLimited (Windows containers)NoLimited
Blast radiusSmallLargeSmall to mediumSmallSmall to medium
Best forCompliance, legacy, big/steady workloadsTiny systems, cost-firstGeneral microservicesEvent-driven, spikyMany services at scale

Combines well with

  • Database per Service: each VM-hosted service owns its database, completing the isolation.
  • API Gateway / BFF: the gateway is the only public entry point; VMs stay private.
  • Health Check API: load balancer probes and scale-set health extensions rely on /health.
  • Externalized Configuration: Key Vault plus managed identity.
  • Service Registry / Server-Side Discovery: an internal load balancer plus private DNS acts as the “router”.
  • Circuit Breaker and Retry: still required; separate VMs do not remove network failure.
  • Strangler Fig: lift a legacy module onto its own VM first, then modernise it later.

ADR (architecture decision record)

ADR-043: Host the Payments and Document-OCR services on dedicated VM scale sets Status: Proposed Context: Payments is in PCI-DSS scope and must be isolated from lower-trust workloads. Document-OCR uses a native library and spikes memory to 12 GB during batch runs. The team has strong VM operations skills and no Kubernetes production experience. The system has 8 services today. Decision: Run Payments and OCR each on their own Flexible-orchestration VM scale sets across three availability zones, in separate subnets with dedicated NSGs, built from immutable images published through an Azure Compute Gallery. Run the remaining low-risk services on Azure Container Apps. Consequences (positive): Clear PCI boundary that shrinks audit scope; OCR memory spikes cannot affect Payments; predictable capacity. Consequences (negative): Higher cost at low utilisation; image pipeline and patching to maintain; slower scale-out (minutes) than containers. Alternatives considered: All services on AKS (rejected for now: platform skills gap, added operational surface); all on shared VMs (rejected: fails isolation requirement). Review trigger: Revisit when service count exceeds 15 or when the team has an AKS platform capability.

Azure implementation

Services that implement or support the pattern

NeedAzure service
The VM itselfAzure Virtual Machines
Multiple identical instances, autoscale, rolling upgradeVirtual Machine Scale Sets (Flexible orchestration)
Image storage and replicationAzure Compute Gallery
Image build automationAzure VM Image Builder or HashiCorp Packer
Load balancingAzure Load Balancer (layer 4, internal or public), Application Gateway (layer 7, WAF)
Global entry / WAF / CDNAzure Front Door
Network isolationVirtual Network, subnets, NSGs, Azure Firewall, Bastion
SecretsAzure Key Vault + managed identity
MonitoringAzure Monitor Agent, Log Analytics, Application Insights, alerts
PatchingAzure Update Manager
BackupAzure Backup (for stateful VMs)
DatabasesAzure SQL Database, Azure Database for PostgreSQL Flexible Server, SQL Server on VM if truly required

How to configure (Azure CLI sketch; confirm current flags with az vmss create --help)

Terminal window
az group create -n rg-claims-payments-prod -l westeurope
az vmss create \
--resource-group rg-claims-payments-prod \
--name vmss-claims-payments \
--orchestration-mode Flexible \
--image "/subscriptions/<sub>/resourceGroups/rg-images/providers/Microsoft.Compute/galleries/gal_claims/images/payments-linux/versions/1.4.0" \
--vm-sku Standard_D2as_v6 \
--instance-count 2 \
--zones 1 2 3 \
--platform-fault-domain-count 1 \
--security-type TrustedLaunch \
--enable-secure-boot true --enable-vtpm true \
--assign-identity '[system]' \
--vnet-name vnet-claims --subnet snet-payments \
--lb lb-payments-internal \
--public-ip-address ""

Then: attach an autoscale rule (for example, scale out at average CPU above 65% over 5 minutes, scale in below 30% over 10 minutes, with cooldown), a load-balancer health probe on /health, and Azure Monitor Agent through a data collection rule.

Pricing and tier considerations (always confirm current numbers in the Azure pricing calculator; prices differ by region and OS)

  • Billing model: you pay for VM compute per unit time while it is allocated, plus the OS disk, data disks, public IPs and egress. Stopping (deallocating) a VM stops compute charges but disks still cost money.
  • Windows and SQL licensing: Windows Server VMs include a licence cost in the hourly price; Azure Hybrid Benefit lets you bring eligible on-premises licences and reduce it. Linux images avoid the Windows licence charge.
  • Commitment discounts: Reserved VM Instances (1 or 3 year, tied to a size and region) and Azure savings plan for compute (flexible hourly commitment) can cut steady-state cost substantially versus pay-as-you-go; Microsoft advertises savings of up to roughly 70% for long commitments, and the actual figure depends on size, region and term. Use them for always-on production VMs.
  • Spot VMs: deeply discounted but evictable. Suitable only for interruptible work such as batch OCR, never Payments.
  • Sizing families: general purpose (D-series) for APIs; memory-optimised (E-series) for OCR or rating; burstable (B-series) for low-traffic internal tools; compute-optimised (F-series) for CPU-heavy scoring. Newer generations (for example v6) usually give better price-performance than older ones; check availability in your region.
  • Cost rule of thumb: with two zone-spread instances per service as the minimum production footprint, 8 services means at least 16 VMs before any scale-out. Compare this against Container Apps or AKS before committing.
  • Scale set itself: there is no extra charge for the scale set; you pay for the VMs and related resources.

Reference architecture (text)

Front Door (with WAF) receives all internet traffic and forwards to an Application Gateway in the “edge” subnet. The gateway routes /api/claims/* to the internal load balancer of the Claims-Intake scale set (subnet snet-intake, three zones, two to six instances), /api/payments/* to the Payments scale set (subnet snet-payments, its own NSG, three zones, two to four instances, no route from the internet), and /api/documents/* to the OCR scale set (memory-optimised E-series). Each scale set uses a system-assigned managed identity to read its own Key Vault and to connect to its own database through a private endpoint. Azure Monitor Agent ships logs and metrics to a Log Analytics workspace; Application Insights collects traces via OpenTelemetry; alerts page the on-call engineer. The Angular SPA is hosted on Azure Static Web Apps and calls the same-origin /api path. Images come from an Azure Compute Gallery populated by a Packer/Image Builder pipeline in Azure DevOps or GitHub Actions; each release creates a new image version and triggers a rolling upgrade. Azure Update Manager reports patch compliance for any long-lived VMs; Bastion provides break-glass admin access.

Teaching guide for my team

2-minute beginner explanation

“Imagine every microservice gets its own computer. Payments has one, Claims Intake has one, Document OCR has one. If OCR uses all its memory, only OCR suffers; Payments carries on. It’s safe and simple to reason about, but computers cost money, need updates, and take a few minutes to start. So we use it when a service needs strong separation or special software, not for everything.”

5-minute intermediate explanation

“Service per VM gives hypervisor-level isolation and reserved CPU/memory per service. In Azure we don’t hand-craft VMs; we bake an image containing the OS, runtime and app, put it in a Compute Gallery, and run it as a Virtual Machine Scale Set across availability zones behind a load balancer. The load balancer probes /health, autoscale adds instances on CPU or queue length, and a deployment is a rolling replacement with a new image version. Secrets come from Key Vault using managed identity; each service sits in its own subnet with its own NSG; each has its own database. The trade-offs are cost at low utilisation, minutes-long scale-out, and patching effort. Compared with containers, we give up density and speed to gain isolation and OS control. Compared with shared hosts, we give up cost efficiency to remove noisy-neighbour and blast-radius problems.”

Hands-on exercise: “Isolate the noisy neighbour” (about 90 minutes)

  1. Start with two small APIs (ClaimsIntake.Api and Ocr.Api) on ONE Linux VM (or two local processes limited to one machine).
  2. Add an endpoint to Ocr.Api that allocates memory in a loop (for example, POST /ocr/stress allocating byte arrays and holding them).
  3. Load-test ClaimsIntake.Api /health with hey or k6 while calling /ocr/stress. Record latency and check journalctl -k | grep -i oom.
  4. Create a second VM. Move Ocr.Api to it using a systemd unit (as in Level 1). Repeat the test.
  5. Add an Azure Load Balancer health probe on /health for the Intake VM.

Expected outcome: in step 3, Intake latency rises and possibly a process is OOM-killed; in step 4, Intake latency stays flat while OCR alone degrades. The team can explain why (separate kernels, separate memory) and name the cost of the extra VM.

Interview-style questions

  1. Why choose service-per-VM over containers? When you need hypervisor-level isolation (compliance, multi-tenant safety), guaranteed resources, a specific OS or kernel, or when the workload cannot be containerised (legacy Windows, drivers, licences).
  2. Why must services on VM scale sets be stateless and images immutable? Instances are replaced during scale-in, upgrades and failures; local state would be lost, and hand-edited servers drift so that instances behave differently.
  3. What are the main downsides, and how do you limit them? Cost at low utilisation, slow boot, and patching effort. Limit them with right-sizing, reservations or savings plans, pre-baked images, autoscale with sensible minimums, and automated image pipelines.

Mastery checklist

  • I can state at least three situations where service-per-VM is the right choice and three where it is not.
  • I can explain the isolation difference between a VM and a container in one sentence.
  • I have built an immutable image and deployed it to a scale set with a health probe.
  • I can configure multi-zone redundancy, autoscale rules and rolling upgrades, and explain their failure modes.
  • I can apply network isolation: private subnets, NSGs, no public IPs, Bastion for admin access.
  • I use managed identity and Key Vault, and no secrets exist in images or repositories.
  • I can estimate monthly cost for a service (VM size, instance count, disks, licences, commitment discount) and compare it with a container alternative.
  • I can write an ADR that justifies (or rejects) the pattern for a given service.

Key takeaway

Give a service its own VM when isolation, guaranteed resources or OS control are worth more than density and speed; then treat those VMs as disposable, immutable, multi-zone instances built from images, never as hand-tended servers.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 48: Service Mesh

A service mesh is an infrastructure layer that moves service-to-service networking concerns (mutual TLS, retries, timeouts, traffic routing, telemetry, authorization) out of your application code and into a fleet of...

Manikandan
Manikandan·20 min read
Microservices

Day 47: Sidecar

A sidecar is a helper container that is deployed and scaled together with your application container, in the same pod (Kubernetes) or the same replica (Azure Container Apps).

Manikandan
Manikandan·14 min read
Microservices

Day 46: Service Deployment Platform

A Service Deployment Platform is an orchestrator (in practice, usually Kubernetes) that takes a declared desired state, such as "run 3 copies of the Claims API at version 2.4", and continuously makes reality match it.

Manikandan
Manikandan·12 min read
Microservices

Day 45: Serverless Deployment

Serverless deployment means you ship only your code and its triggers (an HTTP call, a queue message, a timer) and let the cloud platform decide where it runs, how many copies run, and when they are switched off.

Manikandan
Manikandan·17 min read