A sidecar is a helper container that is deployed and scaled together with your application container, in the same pod (Kubernetes) or the same replica (Azure Container Apps).
Intro
A sidecar is a helper container that is deployed and scaled together with your application container, in the same pod (Kubernetes) or the same replica (Azure Container Apps). It shares the application’s network namespace and can share volumes, so it can take over platform chores such as log shipping, telemetry export, TLS termination, authentication proxying, secret refresh, or talking to a message broker. The application stays focused on business logic (for us: insurance claims) and the platform tooling lives in a separate, independently versioned image, in any language.
Why we need this
Every microservice needs the same non-business capabilities: ship logs, export traces and metrics, terminate mTLS, refresh secrets, authenticate callers, publish to a broker. There are three ways to provide them:
- Library inside the app (chassis, Day 40). Fast and simple, but it ties you to one language and one release train. Ten services in .NET, Python and Java means three implementations to keep in sync.
- Node-level agent (DaemonSet). One agent per node is cheap, but it cannot be tailored per service, cannot share the pod’s localhost, and breaks isolation between tenants and teams.
- Sidecar. One helper per app instance. Language-neutral, versioned and configured per service, and it can be upgraded without recompiling the application.
Business reasons: platform teams can roll out a security fix (say a new TLS policy) by changing a container image instead of asking 40 teams to rebuild; product teams ship faster because they stop re-implementing plumbing.
What problem it solves
Problem (from the topic list): platform tooling is duplicated inside application code.
In our ClaimsApi (.NET), PolicyService (.NET) and FraudScoring (Python) services, each team wrote its own log shipper, its own token-refresh logic and its own retry-to-broker code. Without a sidecar:
- A CVE in the logging library means three separate patches in three languages.
- Behaviour drifts: one service exports traces with sampling, another does not.
- App images become large and slow to build because they bundle agents, CLIs and certificates.
- Business code is cluttered with infrastructure code, and unit tests need to mock it all.
- Non-.NET services never get the feature at all, because the shared library exists only for .NET.
When it is needed (and when it is NOT)
Good fit
- Polyglot estates where a shared library is impractical.
- Cross-cutting capability that must be identical everywhere: mTLS, telemetry export, egress proxy, auth proxy.
- Legacy or third-party apps you cannot modify (wrap them with a sidecar for TLS, logging or auth).
- Runtime-adapter needs: Dapr for pub/sub and state, a database connection proxy for Azure SQL / PostgreSQL, a config/secret reloader.
- Per-instance lifecycle coupling: the helper must start, stop and scale exactly with the app.
Not a fit
- A single-language, small system where a NuGet package does the job with less operational cost.
- Ultra-latency-sensitive hot paths where the extra localhost hop is unacceptable (it is typically sub-millisecond, but not free).
- Very high replica counts with tiny apps: a 100 MB sidecar on every 50 MB pod multiplies memory cost. Consider a node-level agent or a sidecarless mesh (Istio ambient mode).
- Batch jobs: a sidecar that never exits keeps a Job from completing unless you use native sidecar containers (Section 5) or handle shutdown explicitly.
- When the team lacks Kubernetes/container maturity to debug two containers per pod.
How to identify the problem (key signals)
- The same logging/telemetry/auth code exists in several repos, in different languages, at different versions.
- A platform-wide change (new collector endpoint, TLS min version) needs coordinated releases of many services.
- Pull requests to business services contain infrastructure changes such as exporter config or certificate loading.
- Dashboards show services with missing or inconsistent traces, because some services never got the instrumentation.
- Docker images contain unrelated tools (fluentd, az CLI, openssl scripts) and take minutes to build.
- Teams ask “can our Python service also get the .NET chassis features?”
- Legacy or vendor apps cannot speak your auth or telemetry standard and have no source access.
Flow Diagram
Helper containers share the app’s network and lifecycle inside one pod or replica.
flowchart LR subgraph Pod["Pod / replica"] APP["claims-api container"] DAPR["daprd sidecar :3500"] OTEL["OTel collector sidecar :4317"] APP -- "localhost publish" --> DAPR APP -- "OTLP" --> OTEL end DAPR --> SB{{"Service Bus"}} DAPR --> KV["Key Vault"] OTEL --> AI["Application Insights"]Level 1: Beginner
Analogy: a motorcycle with a sidecar. The rider (application) drives and does the main job; the sidecar carries extra equipment. It is attached, travels with the bike, and is detached when the bike is retired. It is not a second vehicle on a different road.
Key facts
- Containers in one pod share
localhostand can share volumes. - They start and stop together and are scheduled on the same node.
- They are scaled together: 5 app replicas means 5 sidecars.
Minimal example: a log-shipping sidecar with a shared volume. The app writes files to /var/log/claims; the sidecar tails and ships them.
apiVersion: v1kind: Podmetadata: name: claims-apispec: volumes: - name: logs emptyDir: {} containers: - name: claims-api image: myacr.azurecr.io/claims-api:1.4.0 volumeMounts: - name: logs mountPath: /var/log/claims - name: log-shipper image: busybox:1.36 command: ["sh", "-c", "tail -F /var/log/claims/app.log"] volumeMounts: - name: logs mountPath: /var/log/claims readOnly: trueNative sidecar containers. Since Kubernetes 1.28 (beta from 1.29, stable in 1.33) a sidecar is declared under initContainers with restartPolicy: Always. It starts before the app, runs for the pod’s whole life, and is shut down after the app. This fixes the old problems of startup ordering and of Jobs that never finish.
spec: initContainers: - name: log-shipper image: fluent/fluent-bit:3.2 restartPolicy: Always # this line makes it a native sidecar containers: - name: claims-api image: myacr.azurecr.io/claims-api:1.4.0Level 2: Intermediate
Real scenario: ClaimsApi (ASP.NET Core) must publish a ClaimSubmitted event, but the team does not want broker SDKs in the code. We use the Dapr sidecar (daprd): the app makes a plain HTTP or gRPC call to localhost:3500, and Dapr talks to Azure Service Bus.
.NET (target: current .NET LTS, NuGet Dapr.AspNetCore)
using Dapr.Client;
var builder = WebApplication.CreateBuilder(args);builder.Services.AddDaprClient(); // talks to the sidecar on localhostvar app = builder.Build();
app.MapPost("/claims", async (ClaimSubmitted claim, DaprClient dapr) =>{ // The app does not know about Service Bus; the sidecar does. await dapr.PublishEventAsync("pubsub", "claim-submitted", claim); return Results.Accepted($"/claims/{claim.ClaimId}");});
app.Run();
public record ClaimSubmitted(Guid ClaimId, string PolicyNumber, decimal Amount);Dapr component (swap brokers without code change)
apiVersion: dapr.io/v1alpha1kind: Componentmetadata: name: pubsubspec: type: pubsub.azure.servicebus.topics version: v1 metadata: - name: namespaceName value: "claims-sb.servicebus.windows.net" # uses managed identity, no secretKubernetes annotations that inject the Dapr sidecar
metadata: annotations: dapr.io/enabled: "true" dapr.io/app-id: "claims-api" dapr.io/app-port: "8080"Angular side. The Angular SPA never sees the sidecar. It calls the gateway or BFF over HTTPS as usual. That is a benefit: moving from direct broker access to Dapr or a mesh does not change the front end.
Database side. A sidecar is a good place for a connection proxy. For example, a small proxy container that holds the Azure SQL or PostgreSQL connection and refreshes Microsoft Entra ID tokens, so the app connects to localhost:1433 or localhost:5432 without handling token rotation.
Where the sidecar fits with earlier days
- Publishing through Dapr is still a dual write with your database. Combine it with the Transactional Outbox (Day 12) or Dapr’s outbox support.
- Consumers still need to be idempotent (Day 17).
Level 3: Advanced
Performance and cost
- Every sidecar requests CPU and memory. Set explicit
requestsandlimitsfor the sidecar; unbounded sidecars are a classic cause of OOMKilled pods and noisy nodes. - Localhost hop cost is small, but a proxy sidecar adds two extra hops per call (in and out). Measure p99, not averages.
- At scale, sidecar memory multiplies by replica count. 500 pods x 100 MiB = about 50 GiB just for helpers.
Scalability
- HPA/KEDA scale the pod as a unit, and the sidecar’s CPU counts toward the pod’s utilisation. A busy log shipper can trigger scale-out of the application even though the app is idle. Use per-container metrics or custom metrics to avoid this.
Security
- Containers in a pod share the network namespace: a compromised app can reach a sidecar’s admin port on localhost. Bind admin and metrics ports carefully and require authentication or use a Unix socket.
- Give the sidecar its own, least-privilege identity where possible (workload identity), and drop Linux capabilities.
- Use
readOnlyRootFilesystemand mount shared volumes read-only on the consuming side. - Sidecar injection through mutating webhooks makes the webhook a critical component. If it is down, pods may start without their sidecar (fail-open) or not start at all (fail-closed). Decide deliberately.
Failure modes
| Failure | Effect | Mitigation |
|---|---|---|
| Sidecar not ready when app starts | First requests fail (app tries to publish, mesh not ready) | Native sidecars (start first, startup probe), or app waits on /v1.0/healthz |
| Sidecar crashes | Pod is unhealthy or telemetry silently lost | Probes on both containers; alert on sidecar restarts |
| Sidecar outlives the app in a Job | Job never completes | Native sidecars (stopped automatically) |
| Sidecar shuts down first | In-flight requests fail on termination | Native sidecars terminate after the app; add preStop sleep for proxies |
| Version skew (mesh proxy vs control plane) | Odd behaviour | Pin versions, upgrade via rollout, track in CI |
Common mistakes
- Using a sidecar for something a 20-line library does well.
- No resource limits on the sidecar.
- Putting business logic in the sidecar (it then becomes a hidden, untested part of the service).
- Treating sidecar logs as the app’s logs and losing them on crash; ensure both containers ship logs.
- Ignoring shutdown order for old-style sidecars (before Kubernetes 1.28, no ordering guarantee).
- Tagging images
latest, so a pod restart silently upgrades the sidecar.
Level 4: Expert and Architect view
Alternatives compared
| Option | Language neutral | Per-service tailoring | Extra cost | Isolation | Upgrade path | Best for |
|---|---|---|---|---|---|---|
| Sidecar | Yes | High | One container per pod | Per pod | Roll pods | Polyglot, per-app policy, Dapr, mesh |
| Library / chassis (Day 40) | No | High | None at runtime | In-process | Rebuild every service | Single-language stacks |
| Node agent (DaemonSet) | Yes | Low | One per node | Shared per node | Roll nodes | Logs, node metrics |
| Sidecarless mesh (ambient / eBPF) | Yes | Medium | Per-node proxy + optional waypoint | Per node | Roll node proxy | Large clusters, mesh-only needs |
| Gateway / shared proxy | Yes | Low | Central component | Central | Central | North-south traffic |
Combines with: Service Mesh (Day 48, sidecars are its data plane), Service per Container (Day 42), Externalized Configuration (Day 39), Health Check API (Day 36), Distributed Tracing (Day 31), Log Aggregation (Day 32), Transactional Outbox (Day 12).
ADR (for the architecture review)
- Title: ADR-047 Use sidecars for telemetry export and broker access in the claims platform.
- Status: Proposed.
- Context:
ClaimsApiandPolicyServiceare .NET,FraudScoringis Python. Logging, tracing, and broker code is duplicated and drifts. Platform team owns AKS/Container Apps. - Decision: Run an OpenTelemetry Collector sidecar (or agent) for telemetry export, and adopt Dapr sidecar for pub/sub. Use native sidecar containers on AKS (Kubernetes 1.29+ is required; 1.33+ recommended for stable behaviour). Keep business rules in the app.
- Consequences: (+) one place to change exporters and brokers, language neutral, faster onboarding. (-) higher per-pod memory and CPU, more moving parts to debug, need for sidecar resource governance and version pinning.
- Alternatives rejected: shared NuGet chassis only (does not help Python); node-level agents only (cannot tailor per service).
- Review trigger: revisit if sidecar overhead exceeds 15% of cluster memory or if a sidecarless mesh becomes standard on our platform.
Azure implementation
Services
- Azure Kubernetes Service (AKS): native sidecar containers (Kubernetes version dependent, verify your cluster version), Istio-based service mesh add-on (sidecar proxies), Azure Monitor Container Insights, Dapr as an AKS extension.
- Azure Container Apps (ACA): supports multiple containers per app; additional containers run as sidecars in the same replica and share network and volumes (init containers are also supported). Managed Dapr is enabled per app, and the platform injects the Dapr sidecar.
- Azure Monitor / Application Insights: the OpenTelemetry Collector sidecar exports to it. ACA also has a managed OpenTelemetry agent option.
- Azure Service Bus, Key Vault, Azure SQL / PostgreSQL Flexible Server: the back ends the sidecars talk to, ideally with managed identity.
- Azure Container Registry: stores app and sidecar images; pin by digest.
Configuration: enable Dapr on Azure Container Apps
az containerapp create \ --name claims-api --resource-group rg-claims \ --environment cae-claims \ --image myacr.azurecr.io/claims-api:1.4.0 \ --target-port 8080 --ingress internal \ --enable-dapr --dapr-app-id claims-api --dapr-app-port 8080Configuration: an extra sidecar container in ACA (YAML template excerpt for az containerapp update --yaml)
properties: template: containers: - name: claims-api image: myacr.azurecr.io/claims-api:1.4.0 resources: { cpu: 0.5, memory: 1Gi } - name: otel-collector image: otel/opentelemetry-collector-contrib:0.110.0 resources: { cpu: 0.25, memory: 0.5Gi }Pricing and tier considerations
- On AKS you pay for the nodes (VM size and count); sidecars cost you node capacity, not a separate line item. Sidecar requests reduce pod density, so plan node size with the sidecar’s requests included. The Standard tier adds the uptime SLA (verify current AKS tier pricing on the Azure pricing page).
- On ACA, billing depends on the plan. On the Consumption plan you pay for vCPU-seconds and GiB-seconds of the resources allocated to the replica plus requests, and there is a monthly free grant. All containers in the replica, including sidecars, contribute to the replica’s resource allocation, so size them together. Dedicated (workload profiles) plans bill for the profile instead. Check the current ACA billing page before quoting numbers, because rates and grants change.
- Dapr on ACA is a platform feature; you still pay for the resources your replica uses and for backing services such as Service Bus (Standard is required for topics; Premium adds isolation and larger messages).
- Application Insights / Log Analytics cost scales with ingested GB. A chatty log-shipper sidecar can become your largest Azure Monitor bill: filter and sample at the collector.
Reference architecture (text)
Clients (Angular SPA) call Azure Front Door, then API Management or the BFF. Traffic reaches the ClaimsApi pod or replica. Inside it, the claims-api container serves HTTP; a Dapr daprd sidecar handles pub/sub to Azure Service Bus and secrets from Key Vault; an OpenTelemetry Collector sidecar receives OTLP on localhost:4317 and exports to Application Insights. Everything authenticates with managed identity (workload identity on AKS). Azure SQL is accessed by the app directly with Entra ID tokens. Images come from ACR and are deployed by GitHub Actions or Azure DevOps with the same pipeline template for app and sidecar versions.
Teaching guide for my team
2-minute beginner explanation
“Your app is the driver. Some jobs, like sending logs or securing traffic, are not the driver’s job, so we bolt a helper container next to it. It lives in the same pod, so it can talk over localhost and share files. It starts and stops with the app, and if we run 5 copies of the app, we get 5 copies of the helper. The point: platform people change the helper, and you do not rebuild your code.”
5-minute intermediate explanation
Cover: (1) why libraries do not work across languages; (2) shared network and volumes in a pod; (3) native sidecars with restartPolicy: Always and what they fix; (4) show the Dapr PublishEventAsync example and how the component file swaps the broker; (5) costs: memory per pod, extra hop, more to debug; (6) when not to use it (single-language system, tiny pods).
Hands-on exercise
Deploy ClaimsApi to a local cluster (kind or minikube, Kubernetes 1.29+). Add a native sidecar that tails /var/log/claims/app.log from a shared emptyDir.
Expected outcome: kubectl get pod shows 2/2 containers running; kubectl logs claims-api -c log-shipper shows the app’s log lines; when you delete the pod, the sidecar stops after the app (check with kubectl get events). Stretch: set a 16Mi memory limit on the sidecar and observe what happens when it is exceeded.
Interview-style questions
- What do containers in the same pod share? The network namespace (one IP,
localhost), IPC and, if declared, volumes. They do not share the filesystem by default. - What problem do native sidecar containers solve? Startup ordering (sidecar starts before the app), shutdown ordering (stops after the app), and Jobs that never complete because a regular sidecar never exits.
- When would you choose a library over a sidecar? When the whole estate is one language, the capability is small, and the extra resource cost and operational complexity of another container are not justified.
Mastery checklist
- I can explain why sidecar, library, and node agent are different trade-offs.
- I can write a pod spec with a native sidecar (
initContainers+restartPolicy: Always) and explain the ordering guarantees. - I can set correct resource requests and limits for a sidecar and calculate its cluster-wide cost.
- I can call a Dapr sidecar from .NET and explain why publish still needs an outbox.
- I know how sidecar injection works (mutating webhook or platform feature) and what happens if it fails.
- I can enable Dapr and add a second container in Azure Container Apps.
- I can name three situations where a sidecar is the wrong choice.
- I can write an ADR justifying sidecars for a real system.
Key takeaway
A sidecar moves shared platform plumbing out of your application into a helper container that lives and scales with it. Use it when many services or languages need the same capability, and pay attention to its resource cost, startup and shutdown order, and security boundary.
