A service mesh is an infrastructure layer that moves service-to-service networking concerns (mutual TLS, retries, timeouts, traffic routing, telemetry, authorization) out of your application code and into a fleet of...
Intro
A service mesh is an infrastructure layer that moves service-to-service networking concerns (mutual TLS, retries, timeouts, traffic routing, telemetry, authorization) out of your application code and into a fleet of proxies (Envoy in Istio) that sit next to each service. A central control plane (istiod) pushes policy and certificates to those proxies. Your .NET, Java, or Node services keep making plain HTTP/gRPC calls; the mesh makes those calls encrypted, observable, and policy-controlled, uniformly, in every language.
Why we need this
In an insurance claims system you might have ClaimsApi, PolicyService, FraudScoringService, PaymentService, DocumentService (Python OCR), and NotificationService. Each needs the same cross-cutting network behaviour:
- Encrypted, authenticated service-to-service traffic (claims data is personal and financial; regulators expect encryption in transit inside the cluster, not only at the edge).
- Consistent timeouts, retries and circuit breaking.
- Traffic-level telemetry (request rate, errors, latency per service pair) and trace propagation.
- Fine-grained “who may call whom” rules (only
ClaimsApimay callPaymentService). - Safe releases: send 5% of traffic to
FraudScoringServicev2.
Without a mesh, each team implements these in a shared library (the Microservice Chassis from Day 40). That works while everything is one language and one team. It breaks down when services are polyglot, when library upgrades must be rolled out across 40 repos to fix a TLS bug, or when security wants to enforce a policy without asking every team to redeploy. The mesh makes networking behaviour an operations-owned platform capability instead of a library each team must remember to upgrade.
What problem it solves
Problem (from the topic list): networking logic (mTLS, telemetry, routing) is hardcoded across languages.
What goes wrong without it:
- The .NET services use one Polly retry config, the Python OCR service uses
tenacitywith different numbers, and the Node BFF has no retries at all. Behaviour during an incident is unpredictable. - Certificates for internal TLS are managed by hand or not at all; traffic inside the cluster is plaintext, so anyone with network access on a node can read claim payloads.
- A CVE in the TLS stack of your shared library means a coordinated redeploy of every service.
- Nobody can answer “which services call
PaymentService, and with what error rate?” without reading code. - Canary releases need custom code or a separate deployment tool per team.
The mesh solves this by intercepting all traffic in and out of each pod through a proxy that is configured centrally and identically for every language.
When it is needed (and when it is NOT)
Good fit:
- 15-20+ services, several languages, several teams, and a platform/SRE team that can own the mesh.
- A compliance requirement for mTLS between all workloads (zero-trust), with workload identity rather than IP-based trust.
- You need uniform L7 traffic control: weighted routing, fault injection, per-route timeouts and retries, without touching application code.
- You already run Kubernetes (AKS) and have observability tooling (Prometheus, Grafana, tracing backend) to consume mesh telemetry.
Overkill or wrong:
- Fewer than ~10 services, one language, one team: a chassis library plus Kubernetes NetworkPolicy is simpler and cheaper.
- Nobody owns platform operations. A mesh adds a control plane, CRDs, sidecar upgrades, and a new class of failures to debug.
- Latency-critical paths where every extra hop matters and you cannot afford the proxy overhead (each sidecar hop adds a small but non-zero latency and CPU/memory cost per pod).
- Serverless-only or Azure Container Apps/App Service-only estates: there is no pod to inject a sidecar into (Container Apps has its own built-in Envoy-based ingress and Dapr; you do not run Istio there).
- Windows containers or virtual nodes on AKS: not supported by the AKS Istio add-on (see section 9).
- Teams that expect the mesh to fix bad service design. It will not fix chatty synchronous chains; it will only make them observable.
How to identify the problem (key signals)
- Copy-pasted resilience code with different timeouts and retry counts in each service repo (Polly in .NET, tenacity in Python, hand-written loops elsewhere).
- Security audit finding: “east-west traffic inside the cluster is unencrypted” or “we cannot prove which service called which”.
- A library CVE or TLS change requires touching dozens of repos and coordinated deployments.
- During incidents nobody can see per-hop metrics; the only signal is the user-facing 500 and you grep logs across services.
- Release friction: canary and blue/green are done with ad-hoc scripts, or teams avoid them.
- Client libraries for service discovery/auth exist in some languages but not others, so newer-language services are second-class citizens.
- Certificate expiry outages for internal services because rotation is manual.
Flow Diagram
Envoy sidecars carry mTLS traffic; istiod pushes config and certificates.
flowchart LR subgraph P1["Pod: claims-api"] A1["app"] --> E1["envoy"] end subgraph P2["Pod: policy-service"] E2["envoy"] --> A2["app"] end E1 == "mTLS + retries + telemetry" ==> E2 IST["istiod control plane"] -. "xDS config + certs" .-> E1 IST -.-> E2 IGW["Ingress gateway"] --> E1Level 1: Beginner
Analogy: Think of every service as an office. Instead of each office training its own receptionist to check ID badges, encrypt mail and log visitors, the company puts an identical, professionally trained security desk (the sidecar proxy) at every office door. The head of security (the control plane) hands each desk the same rulebook and issues fresh badges (certificates) automatically. Employees (your code) just walk out and say “send this to PolicyService”; the desk does the rest.
Architecture in one picture (text):
Pod: claims-api Pod: policy-service +-----------------------------+ +-----------------------------+ | app container --> envoy | ==mTLS=> | envoy --> app container | | (plain HTTP) (sidecar) | | (sidecar) (plain HTTP) | +-----------------------------+ +-----------------------------+ ^ config + certs (xDS) ^ +------------ istiod (control plane) --+- Data plane: the Envoy proxies. They carry the traffic.
- Control plane:
istiod. It watches Kubernetes, converts your Istio YAML into proxy config, and acts as the certificate authority for workload identities (SPIFFE IDs).
Minimal working example. The app code stays plain. This is a complete .NET minimal API caller; note there is no TLS or retry code:
// Program.cs (ClaimsApi) - .NET 10 LTSvar builder = WebApplication.CreateBuilder(args);builder.Services.AddHttpClient("policy", c => c.BaseAddress = new Uri("http://policy-service")); // plain HTTP inside the mesh
var app = builder.Build();
app.MapGet("/claims/{id:guid}/coverage", async (Guid id, IHttpClientFactory f) =>{ var http = f.CreateClient("policy"); var resp = await http.GetAsync($"/policies/by-claim/{id}"); return resp.IsSuccessStatusCode ? Results.Content(await resp.Content.ReadAsStringAsync(), "application/json") : Results.StatusCode((int)resp.StatusCode);});
app.Run();Turn the mesh on for a namespace and enforce mTLS:
apiVersion: v1kind: Namespacemetadata: name: claims labels: istio-injection: enabled # upstream Istio; on the AKS add-on use istio.io/rev=<revision> instead---apiVersion: security.istio.io/v1kind: PeerAuthenticationmetadata: name: default namespace: claimsspec: mtls: mode: STRICT # reject any plaintext traffic to pods in this namespaceAfter you redeploy the pods, each one has two containers (app + istio-proxy), and ClaimsApi -> PolicyService traffic is now mutually authenticated and encrypted with no code change.
Level 2: Intermediate
6.1 Traffic policy: timeouts, retries, outlier ejection
Put resilience for one specific route in the mesh, expressed declaratively:
apiVersion: networking.istio.io/v1kind: VirtualServicemetadata: name: policy-service namespace: claimsspec: hosts: [policy-service] http: - route: - destination: { host: policy-service } timeout: 3s retries: attempts: 2 perTryTimeout: 1s retryOn: connect-failure,refused-stream,unavailable,503---apiVersion: networking.istio.io/v1kind: DestinationRulemetadata: name: policy-service namespace: claimsspec: host: policy-service trafficPolicy: connectionPool: http: { http1MaxPendingRequests: 100, maxRequestsPerConnection: 50 } outlierDetection: # mesh-level "circuit breaker": eject bad endpoints consecutive5xxErrors: 5 interval: 10s baseEjectionTime: 30s maxEjectionPercent: 50Decide once where retries live. If the mesh retries 2 times and your HttpClient also has a Polly retry of 3, one user request can become 9 calls to PolicyService during an incident. Rule of thumb: keep retries in the mesh for connection-level failures on idempotent calls, and keep business-aware policy (fallbacks, compensation) in code.
6.2 Authorization: who may call whom
apiVersion: security.istio.io/v1kind: AuthorizationPolicymetadata: name: payment-allow-claims-api-only namespace: claimsspec: selector: matchLabels: { app: payment-service } action: ALLOW rules: - from: - source: principals: ["cluster.local/ns/claims/sa/claims-api"] # identity = Kubernetes ServiceAccount to: - operation: methods: ["POST"] paths: ["/payments/*"]The identity comes from the client certificate that istiod issued to the claims-api service account, not from an IP address, so it survives pod rescheduling. Each Deployment must therefore use its own ServiceAccount.
6.3 Canary release
http:- route: - destination: { host: fraud-scoring, subset: v1 } weight: 95 - destination: { host: fraud-scoring, subset: v2 } weight: 5subset values are defined in a DestinationRule by pod labels (version: v1, version: v2).
6.4 .NET and Angular specifics
- .NET:
HttpClientand ASP.NET Core propagate the W3Ctraceparentheader automatically throughSystem.Diagnostics.Activity. Envoy creates spans for each hop but cannot link inbound and outbound calls inside your service unless your app forwards the trace headers, which the built-in instrumentation does. If you use raw sockets or a non-instrumented client, you lose the trace. - Health probes: Kubernetes probes go to the pod, and with strict mTLS the probe traffic would be rejected; Istio rewrites probe requests through the sidecar agent by default, so keep your
/health/liveand/health/readyendpoints (Day 36) and do not disable probe rewriting without reason. - Angular: the mesh is invisible to the browser. The Angular app calls the public API through the ingress gateway (or your API Gateway from Day 19), which terminates TLS and can enforce JWT validation (
RequestAuthentication) and CORS. Mesh mTLS applies only between pods. Keep an HTTP interceptor that adds a correlation ID header so a user complaint can be traced to a mesh-level trace. - Database: SQL Server and PostgreSQL usually sit outside the mesh (Azure SQL / Azure Database for PostgreSQL). Traffic to them is TCP through the sidecar, and mTLS is not applied to them. Use a
ServiceEntryif you runREGISTRY_ONLYoutbound policy, and use the database’s own TLS.
Level 3: Advanced
Performance and cost
- Each sidecar consumes CPU and memory and adds a proxy hop on both sides of every call (two extra hops per request: client sidecar and server sidecar). For typical business APIs the added latency is in the low single-digit milliseconds, but measure on your own workload, do not trust a generic number.
- Set explicit resource requests/limits on
istio-proxy; a fleet of 300 pods each reserving 100m CPU / 128Mi for a sidecar is 30 vCPU / ~37 GiB reserved before your apps do any work. - Limit what each sidecar knows using the
Sidecarresource (egress.hosts) so proxies do not receive configuration for every service in a large cluster. This is the single most effective control-plane scalability setting. - Ambient mode (upstream Istio, GA since Istio 1.24) removes per-pod sidecars: a per-node
ztunnelprovides L4 mTLS and optional per-namespace/servicewaypointproxies handle L7. It cuts resource overhead and removes the pod-restart-on-upgrade problem. The AKS Istio add-on documentation lists ambient mode as not yet supported, so on AKS today you choose between the sidecar add-on and self-managing upstream Istio (which the add-on cannot coexist with). Re-check the AKS docs before deciding.
Security
- Use
STRICTmTLS after aPERMISSIVEmigration period. Watch for clients that are not in the mesh (jobs, legacy VMs, the ingress from a non-mesh load balancer). - Add a default-deny
AuthorizationPolicyper namespace, then allow explicit callers. - Set
outboundTrafficPolicytoREGISTRY_ONLYif you want to block calls to unknown external hosts; you must then declareServiceEntryobjects for every external dependency (payment gateway, fraud vendor). This is powerful and a common source of “everything broke” incidents.
Failure modes
| Failure | Symptom | Mitigation |
|---|---|---|
| App starts before its sidecar is ready | First outbound calls fail with connection refused after each deploy | Enable holdApplicationUntilProxyStarts (mesh config or pod annotation) |
| Job/CronJob pod never completes | Pod stays Running because the sidecar keeps running | Use native sidecar containers (Kubernetes 1.29+ feature, supported by newer Istio versions) or explicitly quit the proxy at job end |
| Stale sidecars after control plane upgrade | Old proxy version keeps running until the pod restarts | Plan rolling restarts (kubectl rollout restart) after each upgrade; the AKS support policy also states patch upgrades of istio-proxy need workload restarts |
| Retry storm | Latency and error spike; upstream saturated | One retry layer only; retry budget via attempts small; outlierDetection |
| Control plane (istiod) down | Existing traffic keeps flowing with cached config; new pods cannot get certs/config | Run 2+ replicas with PodDisruptionBudget, autoscale istiod |
| Certificate or clock issues | TLS handshake failures between services | Monitor cert expiry and node time sync; rely on istiod auto-rotation |
| Protocol sniffing surprises | gRPC or TCP traffic misidentified | Name service ports with a protocol prefix (http, grpc, tcp) or set appProtocol |
Common mistakes
- Installing the mesh for one feature (e.g., just mTLS) and then paying the operational cost of all of it. Consider whether NetworkPolicy plus app-level TLS is enough.
- Debugging by editing YAML in production instead of using
istioctl analyze,istioctl proxy-config, and the Envoy access logs. - Forgetting that sidecar injection only happens at pod creation; labelling the namespace does nothing for running pods.
- Using one shared ServiceAccount (
default) for all workloads, which makes identity-based authorization meaningless. - Expecting the mesh to secure the browser-to-cluster path. It does not; that is the ingress gateway’s job.
Level 4: Expert and Architect view
Where the networking concerns can live
| Option | Strengths | Weaknesses | Choose when |
|---|---|---|---|
| Chassis library (Day 40) + NetworkPolicy | Simple, no new infra, in-process, easy to debug | Per-language, upgrades need redeploys, no uniform identity | Few services, one stack |
| Istio, sidecar model | Richest L7 feature set, huge ecosystem, mTLS + authz + traffic mgmt | Highest complexity and per-pod overhead, restarts needed to upgrade proxies | Many services, polyglot, strong platform team |
| Istio, ambient mode | Lower overhead, no sidecar injection or pod restarts for upgrades | Newer operational model; not available in the AKS managed add-on at the time of writing | Self-managed clusters wanting lower overhead |
| Linkerd | Lightweight, simpler operations, own Rust proxy | Smaller feature set than Istio; check current licensing/distribution terms of stable releases before adopting | Want mTLS + basic traffic features with minimal ops |
| Cilium (eBPF) service mesh / network policy | No sidecars, strong L3/L4 and identity-aware policy | L7 features less complete than Istio | Already using Cilium as CNI |
| Managed platform ingress (Azure Container Apps, Dapr) | No cluster to run | Less control; not a general mesh | Containers without Kubernetes ops |
| Open Service Mesh (OSM) | (historical) simple SMI mesh | Upstream project retired by the CNCF; the AKS OSM add-on is scheduled to reach end of support on 30 September 2027 | Do not choose for new work; migrate to the Istio add-on |
Combines with: Sidecar (Day 47, the mesh’s data-plane mechanism), Service Deployment Platform (Day 46, Kubernetes is the prerequisite), API Gateway (Day 19, north-south vs the mesh’s east-west traffic; Istio ingress gateway can play the gateway role), Circuit Breaker / Retry / Bulkhead (Days 26-28, partly expressible as mesh policy), Distributed Tracing and Metrics (Days 31, 33), Access Token (Day 38, RequestAuthentication validates JWTs at the edge), Health Check API (Day 36), Service Registry (Day 21, Kubernetes DNS + the mesh’s endpoint discovery replace a separate registry).
ADR (suitable for architecture review)
- Title: ADR-048 Adopt the AKS Istio-based service mesh add-on for the Claims platform east-west traffic.
- Status: Proposed.
- Context: 22 services across .NET, Python and Node on AKS, owned by five teams. Audit requires encryption in transit and service-level access control between claims, policy and payment services. Resilience configuration is inconsistent across services. A platform team of three engineers exists.
- Decision: Enable the Istio-based service mesh add-on on the production AKS cluster, starting with the
claimsandpaymentsnamespaces inPERMISSIVEmTLS, moving toSTRICTafter telemetry shows no plaintext callers. UseAuthorizationPolicydefault-deny plus explicit allows. Keep application-level retries only for non-idempotent-safe, business-specific cases; standard retries and outlier ejection go in mesh policy. Use the Istio ingress gateway for internal traffic only; public traffic continues through Azure Application Gateway / API Management. - Consequences (positive): uniform mTLS and identity, per-hop telemetry with no code change, central traffic policy and canary releases, Microsoft-managed control-plane lifecycle.
- Consequences (negative): sidecar CPU/memory overhead and extra latency; new failure modes (injection, startup ordering, stale proxies); staff must learn Istio CRDs; add-on limits (no ambient mode, no multi-cluster, no virtual nodes or Windows containers per current docs); revisions have a defined support window, so upgrades become a scheduled task.
- Alternatives considered: chassis library only (rejected: polyglot fleet), Linkerd self-managed (rejected: not managed by Azure, fewer L7 features), upstream self-managed Istio (rejected: add-on removes control-plane operations; cannot coexist with the add-on).
- Revisit when: the add-on supports ambient mode or multi-cluster, or if the fleet shrinks below the point where the platform cost is justified.
Azure implementation
Services that implement or support this topic
- Azure Kubernetes Service (AKS) Istio-based service mesh add-on: Microsoft-managed Istio control plane (istiod), sidecar injection, and optional ingress gateways, integrated with Azure Monitor managed service for Prometheus and Azure Managed Grafana. Generally available.
- Azure Monitor managed Prometheus + Managed Grafana + Application Insights/OpenTelemetry: consume mesh metrics and traces.
- Azure Key Vault: can supply a custom root/intermediate CA for the mesh (plug-in CA) through the Key Vault Secrets Provider; the default is a self-signed root generated by istiod.
- Azure Application Gateway / Azure Front Door / API Management: public north-south entry in front of the mesh.
- Azure Container Apps: not a mesh target; it has its own managed Envoy ingress and Dapr, and you do not run Istio there.
How to configure (Azure CLI)
# 1. See which Istio revisions your region and AKS version support (do not hardcode a revision)az aks mesh get-revisions --location <region> -o table
# 2. Enable the add-on on an existing cluster (optionally pass --revision <asm-1-XX>)az aks mesh enable --resource-group rg-claims-prod --name aks-claims-prod
# 3. Optional: internal ingress gateway for east-west or private entryaz aks mesh enable-ingress-gateway --resource-group rg-claims-prod \ --name aks-claims-prod --ingress-gateway-type internal
# 4. Label namespace for injection using the REVISION label, not istio-injectionkubectl label namespace claims istio.io/rev=<revision-shown-by-get-revisions>
# 5. Restart workloads so sidecars are injectedkubectl rollout restart deployment -n claimsNotes verified against Microsoft Learn documentation at the time of writing:
- On the add-on, injection uses the
istio.io/rev=<revision>label; the upstreamistio-injection=enabledlabel is not the mechanism. - Revisions follow an N-2 style support policy: at least two revisions are supported, older ones are retired roughly six weeks after the newest starts rolling out, and deprecated revisions cannot be newly installed. Always check
az aks mesh get-revisionsand the AKS release notes. Any revision named in this lesson is illustrative. - Patches to istiod and ingress gateways roll out with AKS releases; workloads must be restarted to pick up a new
istio-proxy. - The Istio add-on cannot coexist with the OSM add-on or with a self-managed Istio install, and does not support Windows containers, virtual nodes, multi-cluster, or (currently) ambient mode.
- Some Istio resources (e.g.,
EnvoyFilter,WasmPlugin,IstioOperator) are blocked or unsupported depending on the support tier; consult the add-on’s customization matrix before relying on them.
Pricing and tier considerations
- AKS has Free, Standard and Premium cluster tiers. Free has no uptime SLA and no management fee; Standard adds a financially backed API server SLA and an hourly per-cluster management fee; Premium adds long-term support. Use Standard or above for production. Check the AKS pricing page for the current hourly rate in your region.
- Verify on the AKS pricing page whether any add-on-specific charge applies. The pages reviewed describe compute as the metered cost; in practice the real bill impact is the node capacity consumed by
istiod, ingress gateways, and oneistio-proxyper pod, plus the added Prometheus/Log Analytics ingestion if you scrape mesh metrics. - Plan node sizing with sidecar requests included. Right-size
istio-proxyrequests and use theSidecarresource to reduce memory. - Log Analytics and managed Prometheus ingestion charges grow with high-cardinality mesh metrics; drop unused labels.
Reference architecture (text)
- Users hit the Angular SPA (Azure Static Web Apps or Storage + Front Door). API calls go to Azure Front Door -> Application Gateway (WAF) -> API Management.
- API Management forwards to an internal Istio ingress gateway (private load balancer) in the AKS cluster in a spoke VNet.
- Inside AKS, namespaces
claims,payments,documentsare mesh-enabled. Pods run app +istio-proxy; all pod-to-pod traffic is mTLS. Each Deployment has its own ServiceAccount, andAuthorizationPolicyobjects encode allowed callers. istiod(add-on managed) issues workload certificates; optionally chained to an intermediate CA stored in Azure Key Vault.- Data tier: Azure SQL Database and Azure Database for PostgreSQL via Private Endpoints, outside the mesh, with database-level TLS; declared as
ServiceEntryif outbound isREGISTRY_ONLY. - Observability: mesh metrics -> Azure Monitor managed Prometheus -> Managed Grafana dashboards; traces via OpenTelemetry to Application Insights; Envoy access logs to Log Analytics with sampling.
- Delivery: GitOps or Azure DevOps/GitHub Actions applies Istio YAML (
VirtualService,DestinationRule,AuthorizationPolicy) alongside the app manifests; upgrades of the mesh revision are canaried namespace by namespace.
Teaching guide for my team
2-minute beginner explanation
“Every service needs to talk to other services safely: encrypted, with timeouts, and with a record of what happened. Instead of every team coding that, we place a small proxy next to every service. Your code sends a normal request; the proxy encrypts it, retries it if the network hiccups, and reports metrics. A central brain called the control plane gives every proxy the same rules and the certificates. So we get the same security and visibility in .NET, Python and Node, without changing the code.”
5-minute intermediate explanation
Cover: (1) data plane vs control plane; (2) sidecar injection at pod creation and why restarts matter; (3) mTLS with SPIFFE identity from the ServiceAccount; (4) VirtualService / DestinationRule for timeouts, retries, weighted routing, outlier detection; (5) AuthorizationPolicy default-deny; (6) the cost: resources per pod, upgrade restarts, new failure modes; (7) the rule “one retry layer only”; (8) north-south (ingress/API gateway) vs east-west (mesh).
Hands-on exercise (local, ~60 minutes)
- Create a local cluster (kind or minikube) and install Istio with
istioctl install --set profile=demo(upstream Istio locally; on AKS you would use the add-on). - Deploy
claims-apiandpolicy-service(any two small ASP.NET Core APIs) into namespaceclaimswithistio-injection=enabled. - Verify each pod shows 2/2 containers.
- Apply
PeerAuthenticationSTRICT. From a pod in a non-injected namespace,curl policy-service.claimsand observe it fails; fromclaims-apiit succeeds. - Make
policy-servicereturn 503 randomly (e.g., 30%). Apply the retryVirtualService; observe fewer client-visible errors. Then also add a Polly retry inclaims-apiand count requests hittingpolicy-serviceto see the multiplication. - Apply a default-deny
AuthorizationPolicyand an allow rule forclaims-api’s ServiceAccount only.
Expected outcome: plaintext callers are rejected, retries reduce visible errors, doubling retries multiplies backend load, and only the allowed identity can call policy-service.
Interview-style questions
- What is the difference between an API Gateway and a service mesh? The gateway handles north-south traffic (external clients to your system: auth, rate limiting, aggregation). The mesh handles east-west traffic between services (mTLS, retries, routing, telemetry). An ingress gateway blurs the line but the concerns are different.
- Why did enabling STRICT mTLS break our nightly job? The job pod is either not injected (plaintext client rejected) or its sidecar was not ready when the job started. Fix by injecting it, using
PERMISSIVEduring migration, andholdApplicationUntilProxyStarts. - When would you choose not to adopt a mesh? Few services, one language, no platform team, latency-critical paths, or when NetworkPolicy plus a shared chassis meets the requirements; the mesh’s operational cost would exceed its benefit.
Mastery checklist
- I can explain data plane vs control plane and draw the request path through two sidecars.
- I can explain how mTLS identity is derived (ServiceAccount -> SPIFFE ID -> certificate) and why a shared
defaultServiceAccount defeats authorization. - I can write a
VirtualService,DestinationRule,PeerAuthentication, andAuthorizationPolicyfrom scratch and predict their effect. - I can migrate a namespace from no mesh to STRICT mTLS safely (PERMISSIVE first, verify with telemetry).
- I can identify retry amplification between mesh and application code and decide where retries belong.
- I can debug with
istioctl analyze,istioctl proxy-status,istioctl proxy-config, and Envoy access logs. - I can enable the AKS Istio add-on, choose a supported revision with
az aks mesh get-revisions, and plan a revision upgrade including workload restarts. - I can argue, in an ADR, when a mesh is and is not worth its cost.
Key takeaway
A service mesh turns mTLS, resilience policy and telemetry into a uniform, centrally managed platform capability at the price of a proxy per pod and real operational complexity; adopt it when many services in many languages need consistent security and traffic control, not before.
