Manikandan — Manikandan
Microservices

Day 48: Service Mesh

ManikandanManikandan
20 min read·Updated Sep 26, 2022

A service mesh is an infrastructure layer that moves service-to-service networking concerns (mutual TLS, retries, timeouts, traffic routing, telemetry, authorization) out of your application code and into a fleet of...

Intro

A service mesh is an infrastructure layer that moves service-to-service networking concerns (mutual TLS, retries, timeouts, traffic routing, telemetry, authorization) out of your application code and into a fleet of proxies (Envoy in Istio) that sit next to each service. A central control plane (istiod) pushes policy and certificates to those proxies. Your .NET, Java, or Node services keep making plain HTTP/gRPC calls; the mesh makes those calls encrypted, observable, and policy-controlled, uniformly, in every language.

Why we need this

In an insurance claims system you might have ClaimsApi, PolicyService, FraudScoringService, PaymentService, DocumentService (Python OCR), and NotificationService. Each needs the same cross-cutting network behaviour:

  • Encrypted, authenticated service-to-service traffic (claims data is personal and financial; regulators expect encryption in transit inside the cluster, not only at the edge).
  • Consistent timeouts, retries and circuit breaking.
  • Traffic-level telemetry (request rate, errors, latency per service pair) and trace propagation.
  • Fine-grained “who may call whom” rules (only ClaimsApi may call PaymentService).
  • Safe releases: send 5% of traffic to FraudScoringService v2.

Without a mesh, each team implements these in a shared library (the Microservice Chassis from Day 40). That works while everything is one language and one team. It breaks down when services are polyglot, when library upgrades must be rolled out across 40 repos to fix a TLS bug, or when security wants to enforce a policy without asking every team to redeploy. The mesh makes networking behaviour an operations-owned platform capability instead of a library each team must remember to upgrade.

What problem it solves

Problem (from the topic list): networking logic (mTLS, telemetry, routing) is hardcoded across languages.

What goes wrong without it:

  • The .NET services use one Polly retry config, the Python OCR service uses tenacity with different numbers, and the Node BFF has no retries at all. Behaviour during an incident is unpredictable.
  • Certificates for internal TLS are managed by hand or not at all; traffic inside the cluster is plaintext, so anyone with network access on a node can read claim payloads.
  • A CVE in the TLS stack of your shared library means a coordinated redeploy of every service.
  • Nobody can answer “which services call PaymentService, and with what error rate?” without reading code.
  • Canary releases need custom code or a separate deployment tool per team.

The mesh solves this by intercepting all traffic in and out of each pod through a proxy that is configured centrally and identically for every language.

When it is needed (and when it is NOT)

Good fit:

  • 15-20+ services, several languages, several teams, and a platform/SRE team that can own the mesh.
  • A compliance requirement for mTLS between all workloads (zero-trust), with workload identity rather than IP-based trust.
  • You need uniform L7 traffic control: weighted routing, fault injection, per-route timeouts and retries, without touching application code.
  • You already run Kubernetes (AKS) and have observability tooling (Prometheus, Grafana, tracing backend) to consume mesh telemetry.

Overkill or wrong:

  • Fewer than ~10 services, one language, one team: a chassis library plus Kubernetes NetworkPolicy is simpler and cheaper.
  • Nobody owns platform operations. A mesh adds a control plane, CRDs, sidecar upgrades, and a new class of failures to debug.
  • Latency-critical paths where every extra hop matters and you cannot afford the proxy overhead (each sidecar hop adds a small but non-zero latency and CPU/memory cost per pod).
  • Serverless-only or Azure Container Apps/App Service-only estates: there is no pod to inject a sidecar into (Container Apps has its own built-in Envoy-based ingress and Dapr; you do not run Istio there).
  • Windows containers or virtual nodes on AKS: not supported by the AKS Istio add-on (see section 9).
  • Teams that expect the mesh to fix bad service design. It will not fix chatty synchronous chains; it will only make them observable.

How to identify the problem (key signals)

  1. Copy-pasted resilience code with different timeouts and retry counts in each service repo (Polly in .NET, tenacity in Python, hand-written loops elsewhere).
  2. Security audit finding: “east-west traffic inside the cluster is unencrypted” or “we cannot prove which service called which”.
  3. A library CVE or TLS change requires touching dozens of repos and coordinated deployments.
  4. During incidents nobody can see per-hop metrics; the only signal is the user-facing 500 and you grep logs across services.
  5. Release friction: canary and blue/green are done with ad-hoc scripts, or teams avoid them.
  6. Client libraries for service discovery/auth exist in some languages but not others, so newer-language services are second-class citizens.
  7. Certificate expiry outages for internal services because rotation is manual.

Flow Diagram

Envoy sidecars carry mTLS traffic; istiod pushes config and certificates.

flowchart LR
subgraph P1["Pod: claims-api"]
A1["app"] --> E1["envoy"]
end
subgraph P2["Pod: policy-service"]
E2["envoy"] --> A2["app"]
end
E1 == "mTLS + retries + telemetry" ==> E2
IST["istiod control plane"] -. "xDS config + certs" .-> E1
IST -.-> E2
IGW["Ingress gateway"] --> E1

Level 1: Beginner

Analogy: Think of every service as an office. Instead of each office training its own receptionist to check ID badges, encrypt mail and log visitors, the company puts an identical, professionally trained security desk (the sidecar proxy) at every office door. The head of security (the control plane) hands each desk the same rulebook and issues fresh badges (certificates) automatically. Employees (your code) just walk out and say “send this to PolicyService”; the desk does the rest.

Architecture in one picture (text):

Pod: claims-api Pod: policy-service
+-----------------------------+ +-----------------------------+
| app container --> envoy | ==mTLS=> | envoy --> app container |
| (plain HTTP) (sidecar) | | (sidecar) (plain HTTP) |
+-----------------------------+ +-----------------------------+
^ config + certs (xDS) ^
+------------ istiod (control plane) --+
  • Data plane: the Envoy proxies. They carry the traffic.
  • Control plane: istiod. It watches Kubernetes, converts your Istio YAML into proxy config, and acts as the certificate authority for workload identities (SPIFFE IDs).

Minimal working example. The app code stays plain. This is a complete .NET minimal API caller; note there is no TLS or retry code:

// Program.cs (ClaimsApi) - .NET 10 LTS
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddHttpClient("policy", c =>
c.BaseAddress = new Uri("http://policy-service")); // plain HTTP inside the mesh
var app = builder.Build();
app.MapGet("/claims/{id:guid}/coverage", async (Guid id, IHttpClientFactory f) =>
{
var http = f.CreateClient("policy");
var resp = await http.GetAsync($"/policies/by-claim/{id}");
return resp.IsSuccessStatusCode
? Results.Content(await resp.Content.ReadAsStringAsync(), "application/json")
: Results.StatusCode((int)resp.StatusCode);
});
app.Run();

Turn the mesh on for a namespace and enforce mTLS:

apiVersion: v1
kind: Namespace
metadata:
name: claims
labels:
istio-injection: enabled # upstream Istio; on the AKS add-on use istio.io/rev=<revision> instead
---
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata:
name: default
namespace: claims
spec:
mtls:
mode: STRICT # reject any plaintext traffic to pods in this namespace

After you redeploy the pods, each one has two containers (app + istio-proxy), and ClaimsApi -> PolicyService traffic is now mutually authenticated and encrypted with no code change.

Level 2: Intermediate

6.1 Traffic policy: timeouts, retries, outlier ejection

Put resilience for one specific route in the mesh, expressed declaratively:

apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
name: policy-service
namespace: claims
spec:
hosts: [policy-service]
http:
- route:
- destination: { host: policy-service }
timeout: 3s
retries:
attempts: 2
perTryTimeout: 1s
retryOn: connect-failure,refused-stream,unavailable,503
---
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata:
name: policy-service
namespace: claims
spec:
host: policy-service
trafficPolicy:
connectionPool:
http: { http1MaxPendingRequests: 100, maxRequestsPerConnection: 50 }
outlierDetection: # mesh-level "circuit breaker": eject bad endpoints
consecutive5xxErrors: 5
interval: 10s
baseEjectionTime: 30s
maxEjectionPercent: 50

Decide once where retries live. If the mesh retries 2 times and your HttpClient also has a Polly retry of 3, one user request can become 9 calls to PolicyService during an incident. Rule of thumb: keep retries in the mesh for connection-level failures on idempotent calls, and keep business-aware policy (fallbacks, compensation) in code.

6.2 Authorization: who may call whom

apiVersion: security.istio.io/v1
kind: AuthorizationPolicy
metadata:
name: payment-allow-claims-api-only
namespace: claims
spec:
selector:
matchLabels: { app: payment-service }
action: ALLOW
rules:
- from:
- source:
principals: ["cluster.local/ns/claims/sa/claims-api"] # identity = Kubernetes ServiceAccount
to:
- operation:
methods: ["POST"]
paths: ["/payments/*"]

The identity comes from the client certificate that istiod issued to the claims-api service account, not from an IP address, so it survives pod rescheduling. Each Deployment must therefore use its own ServiceAccount.

6.3 Canary release

http:
- route:
- destination: { host: fraud-scoring, subset: v1 }
weight: 95
- destination: { host: fraud-scoring, subset: v2 }
weight: 5

subset values are defined in a DestinationRule by pod labels (version: v1, version: v2).

6.4 .NET and Angular specifics

  • .NET: HttpClient and ASP.NET Core propagate the W3C traceparent header automatically through System.Diagnostics.Activity. Envoy creates spans for each hop but cannot link inbound and outbound calls inside your service unless your app forwards the trace headers, which the built-in instrumentation does. If you use raw sockets or a non-instrumented client, you lose the trace.
  • Health probes: Kubernetes probes go to the pod, and with strict mTLS the probe traffic would be rejected; Istio rewrites probe requests through the sidecar agent by default, so keep your /health/live and /health/ready endpoints (Day 36) and do not disable probe rewriting without reason.
  • Angular: the mesh is invisible to the browser. The Angular app calls the public API through the ingress gateway (or your API Gateway from Day 19), which terminates TLS and can enforce JWT validation (RequestAuthentication) and CORS. Mesh mTLS applies only between pods. Keep an HTTP interceptor that adds a correlation ID header so a user complaint can be traced to a mesh-level trace.
  • Database: SQL Server and PostgreSQL usually sit outside the mesh (Azure SQL / Azure Database for PostgreSQL). Traffic to them is TCP through the sidecar, and mTLS is not applied to them. Use a ServiceEntry if you run REGISTRY_ONLY outbound policy, and use the database’s own TLS.

Level 3: Advanced

Performance and cost

  • Each sidecar consumes CPU and memory and adds a proxy hop on both sides of every call (two extra hops per request: client sidecar and server sidecar). For typical business APIs the added latency is in the low single-digit milliseconds, but measure on your own workload, do not trust a generic number.
  • Set explicit resource requests/limits on istio-proxy; a fleet of 300 pods each reserving 100m CPU / 128Mi for a sidecar is 30 vCPU / ~37 GiB reserved before your apps do any work.
  • Limit what each sidecar knows using the Sidecar resource (egress.hosts) so proxies do not receive configuration for every service in a large cluster. This is the single most effective control-plane scalability setting.
  • Ambient mode (upstream Istio, GA since Istio 1.24) removes per-pod sidecars: a per-node ztunnel provides L4 mTLS and optional per-namespace/service waypoint proxies handle L7. It cuts resource overhead and removes the pod-restart-on-upgrade problem. The AKS Istio add-on documentation lists ambient mode as not yet supported, so on AKS today you choose between the sidecar add-on and self-managing upstream Istio (which the add-on cannot coexist with). Re-check the AKS docs before deciding.

Security

  • Use STRICT mTLS after a PERMISSIVE migration period. Watch for clients that are not in the mesh (jobs, legacy VMs, the ingress from a non-mesh load balancer).
  • Add a default-deny AuthorizationPolicy per namespace, then allow explicit callers.
  • Set outboundTrafficPolicy to REGISTRY_ONLY if you want to block calls to unknown external hosts; you must then declare ServiceEntry objects for every external dependency (payment gateway, fraud vendor). This is powerful and a common source of “everything broke” incidents.

Failure modes

FailureSymptomMitigation
App starts before its sidecar is readyFirst outbound calls fail with connection refused after each deployEnable holdApplicationUntilProxyStarts (mesh config or pod annotation)
Job/CronJob pod never completesPod stays Running because the sidecar keeps runningUse native sidecar containers (Kubernetes 1.29+ feature, supported by newer Istio versions) or explicitly quit the proxy at job end
Stale sidecars after control plane upgradeOld proxy version keeps running until the pod restartsPlan rolling restarts (kubectl rollout restart) after each upgrade; the AKS support policy also states patch upgrades of istio-proxy need workload restarts
Retry stormLatency and error spike; upstream saturatedOne retry layer only; retry budget via attempts small; outlierDetection
Control plane (istiod) downExisting traffic keeps flowing with cached config; new pods cannot get certs/configRun 2+ replicas with PodDisruptionBudget, autoscale istiod
Certificate or clock issuesTLS handshake failures between servicesMonitor cert expiry and node time sync; rely on istiod auto-rotation
Protocol sniffing surprisesgRPC or TCP traffic misidentifiedName service ports with a protocol prefix (http, grpc, tcp) or set appProtocol

Common mistakes

  1. Installing the mesh for one feature (e.g., just mTLS) and then paying the operational cost of all of it. Consider whether NetworkPolicy plus app-level TLS is enough.
  2. Debugging by editing YAML in production instead of using istioctl analyze, istioctl proxy-config, and the Envoy access logs.
  3. Forgetting that sidecar injection only happens at pod creation; labelling the namespace does nothing for running pods.
  4. Using one shared ServiceAccount (default) for all workloads, which makes identity-based authorization meaningless.
  5. Expecting the mesh to secure the browser-to-cluster path. It does not; that is the ingress gateway’s job.

Level 4: Expert and Architect view

Where the networking concerns can live

OptionStrengthsWeaknessesChoose when
Chassis library (Day 40) + NetworkPolicySimple, no new infra, in-process, easy to debugPer-language, upgrades need redeploys, no uniform identityFew services, one stack
Istio, sidecar modelRichest L7 feature set, huge ecosystem, mTLS + authz + traffic mgmtHighest complexity and per-pod overhead, restarts needed to upgrade proxiesMany services, polyglot, strong platform team
Istio, ambient modeLower overhead, no sidecar injection or pod restarts for upgradesNewer operational model; not available in the AKS managed add-on at the time of writingSelf-managed clusters wanting lower overhead
LinkerdLightweight, simpler operations, own Rust proxySmaller feature set than Istio; check current licensing/distribution terms of stable releases before adoptingWant mTLS + basic traffic features with minimal ops
Cilium (eBPF) service mesh / network policyNo sidecars, strong L3/L4 and identity-aware policyL7 features less complete than IstioAlready using Cilium as CNI
Managed platform ingress (Azure Container Apps, Dapr)No cluster to runLess control; not a general meshContainers without Kubernetes ops
Open Service Mesh (OSM)(historical) simple SMI meshUpstream project retired by the CNCF; the AKS OSM add-on is scheduled to reach end of support on 30 September 2027Do not choose for new work; migrate to the Istio add-on

Combines with: Sidecar (Day 47, the mesh’s data-plane mechanism), Service Deployment Platform (Day 46, Kubernetes is the prerequisite), API Gateway (Day 19, north-south vs the mesh’s east-west traffic; Istio ingress gateway can play the gateway role), Circuit Breaker / Retry / Bulkhead (Days 26-28, partly expressible as mesh policy), Distributed Tracing and Metrics (Days 31, 33), Access Token (Day 38, RequestAuthentication validates JWTs at the edge), Health Check API (Day 36), Service Registry (Day 21, Kubernetes DNS + the mesh’s endpoint discovery replace a separate registry).

ADR (suitable for architecture review)

  • Title: ADR-048 Adopt the AKS Istio-based service mesh add-on for the Claims platform east-west traffic.
  • Status: Proposed.
  • Context: 22 services across .NET, Python and Node on AKS, owned by five teams. Audit requires encryption in transit and service-level access control between claims, policy and payment services. Resilience configuration is inconsistent across services. A platform team of three engineers exists.
  • Decision: Enable the Istio-based service mesh add-on on the production AKS cluster, starting with the claims and payments namespaces in PERMISSIVE mTLS, moving to STRICT after telemetry shows no plaintext callers. Use AuthorizationPolicy default-deny plus explicit allows. Keep application-level retries only for non-idempotent-safe, business-specific cases; standard retries and outlier ejection go in mesh policy. Use the Istio ingress gateway for internal traffic only; public traffic continues through Azure Application Gateway / API Management.
  • Consequences (positive): uniform mTLS and identity, per-hop telemetry with no code change, central traffic policy and canary releases, Microsoft-managed control-plane lifecycle.
  • Consequences (negative): sidecar CPU/memory overhead and extra latency; new failure modes (injection, startup ordering, stale proxies); staff must learn Istio CRDs; add-on limits (no ambient mode, no multi-cluster, no virtual nodes or Windows containers per current docs); revisions have a defined support window, so upgrades become a scheduled task.
  • Alternatives considered: chassis library only (rejected: polyglot fleet), Linkerd self-managed (rejected: not managed by Azure, fewer L7 features), upstream self-managed Istio (rejected: add-on removes control-plane operations; cannot coexist with the add-on).
  • Revisit when: the add-on supports ambient mode or multi-cluster, or if the fleet shrinks below the point where the platform cost is justified.

Azure implementation

Services that implement or support this topic

  • Azure Kubernetes Service (AKS) Istio-based service mesh add-on: Microsoft-managed Istio control plane (istiod), sidecar injection, and optional ingress gateways, integrated with Azure Monitor managed service for Prometheus and Azure Managed Grafana. Generally available.
  • Azure Monitor managed Prometheus + Managed Grafana + Application Insights/OpenTelemetry: consume mesh metrics and traces.
  • Azure Key Vault: can supply a custom root/intermediate CA for the mesh (plug-in CA) through the Key Vault Secrets Provider; the default is a self-signed root generated by istiod.
  • Azure Application Gateway / Azure Front Door / API Management: public north-south entry in front of the mesh.
  • Azure Container Apps: not a mesh target; it has its own managed Envoy ingress and Dapr, and you do not run Istio there.

How to configure (Azure CLI)

Terminal window
# 1. See which Istio revisions your region and AKS version support (do not hardcode a revision)
az aks mesh get-revisions --location <region> -o table
# 2. Enable the add-on on an existing cluster (optionally pass --revision <asm-1-XX>)
az aks mesh enable --resource-group rg-claims-prod --name aks-claims-prod
# 3. Optional: internal ingress gateway for east-west or private entry
az aks mesh enable-ingress-gateway --resource-group rg-claims-prod \
--name aks-claims-prod --ingress-gateway-type internal
# 4. Label namespace for injection using the REVISION label, not istio-injection
kubectl label namespace claims istio.io/rev=<revision-shown-by-get-revisions>
# 5. Restart workloads so sidecars are injected
kubectl rollout restart deployment -n claims

Notes verified against Microsoft Learn documentation at the time of writing:

  • On the add-on, injection uses the istio.io/rev=<revision> label; the upstream istio-injection=enabled label is not the mechanism.
  • Revisions follow an N-2 style support policy: at least two revisions are supported, older ones are retired roughly six weeks after the newest starts rolling out, and deprecated revisions cannot be newly installed. Always check az aks mesh get-revisions and the AKS release notes. Any revision named in this lesson is illustrative.
  • Patches to istiod and ingress gateways roll out with AKS releases; workloads must be restarted to pick up a new istio-proxy.
  • The Istio add-on cannot coexist with the OSM add-on or with a self-managed Istio install, and does not support Windows containers, virtual nodes, multi-cluster, or (currently) ambient mode.
  • Some Istio resources (e.g., EnvoyFilter, WasmPlugin, IstioOperator) are blocked or unsupported depending on the support tier; consult the add-on’s customization matrix before relying on them.

Pricing and tier considerations

  • AKS has Free, Standard and Premium cluster tiers. Free has no uptime SLA and no management fee; Standard adds a financially backed API server SLA and an hourly per-cluster management fee; Premium adds long-term support. Use Standard or above for production. Check the AKS pricing page for the current hourly rate in your region.
  • Verify on the AKS pricing page whether any add-on-specific charge applies. The pages reviewed describe compute as the metered cost; in practice the real bill impact is the node capacity consumed by istiod, ingress gateways, and one istio-proxy per pod, plus the added Prometheus/Log Analytics ingestion if you scrape mesh metrics.
  • Plan node sizing with sidecar requests included. Right-size istio-proxy requests and use the Sidecar resource to reduce memory.
  • Log Analytics and managed Prometheus ingestion charges grow with high-cardinality mesh metrics; drop unused labels.

Reference architecture (text)

  1. Users hit the Angular SPA (Azure Static Web Apps or Storage + Front Door). API calls go to Azure Front Door -> Application Gateway (WAF) -> API Management.
  2. API Management forwards to an internal Istio ingress gateway (private load balancer) in the AKS cluster in a spoke VNet.
  3. Inside AKS, namespaces claims, payments, documents are mesh-enabled. Pods run app + istio-proxy; all pod-to-pod traffic is mTLS. Each Deployment has its own ServiceAccount, and AuthorizationPolicy objects encode allowed callers.
  4. istiod (add-on managed) issues workload certificates; optionally chained to an intermediate CA stored in Azure Key Vault.
  5. Data tier: Azure SQL Database and Azure Database for PostgreSQL via Private Endpoints, outside the mesh, with database-level TLS; declared as ServiceEntry if outbound is REGISTRY_ONLY.
  6. Observability: mesh metrics -> Azure Monitor managed Prometheus -> Managed Grafana dashboards; traces via OpenTelemetry to Application Insights; Envoy access logs to Log Analytics with sampling.
  7. Delivery: GitOps or Azure DevOps/GitHub Actions applies Istio YAML (VirtualService, DestinationRule, AuthorizationPolicy) alongside the app manifests; upgrades of the mesh revision are canaried namespace by namespace.

Teaching guide for my team

2-minute beginner explanation

“Every service needs to talk to other services safely: encrypted, with timeouts, and with a record of what happened. Instead of every team coding that, we place a small proxy next to every service. Your code sends a normal request; the proxy encrypts it, retries it if the network hiccups, and reports metrics. A central brain called the control plane gives every proxy the same rules and the certificates. So we get the same security and visibility in .NET, Python and Node, without changing the code.”

5-minute intermediate explanation

Cover: (1) data plane vs control plane; (2) sidecar injection at pod creation and why restarts matter; (3) mTLS with SPIFFE identity from the ServiceAccount; (4) VirtualService / DestinationRule for timeouts, retries, weighted routing, outlier detection; (5) AuthorizationPolicy default-deny; (6) the cost: resources per pod, upgrade restarts, new failure modes; (7) the rule “one retry layer only”; (8) north-south (ingress/API gateway) vs east-west (mesh).

Hands-on exercise (local, ~60 minutes)

  1. Create a local cluster (kind or minikube) and install Istio with istioctl install --set profile=demo (upstream Istio locally; on AKS you would use the add-on).
  2. Deploy claims-api and policy-service (any two small ASP.NET Core APIs) into namespace claims with istio-injection=enabled.
  3. Verify each pod shows 2/2 containers.
  4. Apply PeerAuthentication STRICT. From a pod in a non-injected namespace, curl policy-service.claims and observe it fails; from claims-api it succeeds.
  5. Make policy-service return 503 randomly (e.g., 30%). Apply the retry VirtualService; observe fewer client-visible errors. Then also add a Polly retry in claims-api and count requests hitting policy-service to see the multiplication.
  6. Apply a default-deny AuthorizationPolicy and an allow rule for claims-api’s ServiceAccount only.

Expected outcome: plaintext callers are rejected, retries reduce visible errors, doubling retries multiplies backend load, and only the allowed identity can call policy-service.

Interview-style questions

  1. What is the difference between an API Gateway and a service mesh? The gateway handles north-south traffic (external clients to your system: auth, rate limiting, aggregation). The mesh handles east-west traffic between services (mTLS, retries, routing, telemetry). An ingress gateway blurs the line but the concerns are different.
  2. Why did enabling STRICT mTLS break our nightly job? The job pod is either not injected (plaintext client rejected) or its sidecar was not ready when the job started. Fix by injecting it, using PERMISSIVE during migration, and holdApplicationUntilProxyStarts.
  3. When would you choose not to adopt a mesh? Few services, one language, no platform team, latency-critical paths, or when NetworkPolicy plus a shared chassis meets the requirements; the mesh’s operational cost would exceed its benefit.

Mastery checklist

  • I can explain data plane vs control plane and draw the request path through two sidecars.
  • I can explain how mTLS identity is derived (ServiceAccount -> SPIFFE ID -> certificate) and why a shared default ServiceAccount defeats authorization.
  • I can write a VirtualService, DestinationRule, PeerAuthentication, and AuthorizationPolicy from scratch and predict their effect.
  • I can migrate a namespace from no mesh to STRICT mTLS safely (PERMISSIVE first, verify with telemetry).
  • I can identify retry amplification between mesh and application code and decide where retries belong.
  • I can debug with istioctl analyze, istioctl proxy-status, istioctl proxy-config, and Envoy access logs.
  • I can enable the AKS Istio add-on, choose a supported revision with az aks mesh get-revisions, and plan a revision upgrade including workload restarts.
  • I can argue, in an ADR, when a mesh is and is not worth its cost.

Key takeaway

A service mesh turns mTLS, resilience policy and telemetry into a uniform, centrally managed platform capability at the price of a proxy per pod and real operational complexity; adopt it when many services in many languages need consistent security and traffic control, not before.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 47: Sidecar

A sidecar is a helper container that is deployed and scaled together with your application container, in the same pod (Kubernetes) or the same replica (Azure Container Apps).

Manikandan
Manikandan·14 min read
Microservices

Day 46: Service Deployment Platform

A Service Deployment Platform is an orchestrator (in practice, usually Kubernetes) that takes a declared desired state, such as "run 3 copies of the Claims API at version 2.4", and continuously makes reality match it.

Manikandan
Manikandan·12 min read
Microservices

Day 45: Serverless Deployment

Serverless deployment means you ship only your code and its triggers (an HTTP call, a queue message, a timer) and let the cloud platform decide where it runs, how many copies run, and when they are switched off.

Manikandan
Manikandan·17 min read
Microservices

Day 44: Multiple Services per Host

Multiple Services per Host means running several independently built and versioned services on the same server (VM, App Service plan, or physical machine) instead of giving each its own VM or container host.

Manikandan
Manikandan·14 min read