Manikandan — Manikandan
Microservices

Day 32: Log Aggregation

ManikandanManikandan
18 min read·Updated Sep 10, 2022

Log Aggregation means every service writes its logs as structured events to stdout or an agent, and a pipeline ships them to one central, searchable, access-controlled store.

Intro

Log Aggregation means every service writes its logs as structured events to stdout or an agent, and a pipeline ships them to one central, searchable, access-controlled store. In our insurance claims system, when a policyholder says “my claim submission failed at 10:42”, you type one query and see the Angular request, the API log lines, the Claims service log lines and the payment-service error together, instead of SSH-ing or kubectl logs-ing into a dozen short-lived containers that may already be gone.


Why we need this

Business reasons

  • Support and operations must answer “what happened to claim CLM-10482?” in minutes, not hours. Slow answers cost customer trust and, in insurance, can breach regulatory response times.
  • Compliance and dispute handling often require proof of what the system did and when. Logs that vanish with a container are useless as evidence (see Day 34, Audit Logging, for the non-repudiable variant).
  • Incident cost is dominated by time-to-diagnose. Central logs cut that time.

Technical reasons

  • Containers, App Service instances and Functions are ephemeral. When an instance is recycled or scaled in, its local log files disappear with it.
  • One user request crosses the Angular app, the API Gateway, several services and a database. No single machine holds the full story.
  • Services scale to N instances. Without aggregation you must know which instance handled the request.
  • Different teams and languages log differently. A central schema (timestamp, level, service, trace id, message) makes cross-service queries possible.

What problem it solves

Problem (from the topic list): logs are scattered across transient containers.

What goes wrong without it

  • The claim payment failed on claims-api replica 3, which was scaled in 20 minutes ago. Its logs are gone.
  • An engineer opens 4 terminals, runs kubectl logs on each pod and greps by hand, then still cannot correlate a request across services.
  • A crash-looping container loses the exact lines that explain the crash because the file system is discarded on restart.
  • Logs are found only after the customer complains, because nobody can alert on text spread across machines.
  • Developers get production shell access just to read logs, which is a security and audit problem.

What it gives you: one query surface, retention you control, alerts on log patterns, access control on who can read what, and a base on which distributed tracing (Day 31) and metrics (Day 33) are joined.


When it is needed (and when it is NOT)

Needed when

  • You run more than one service or more than one instance of a service.
  • You run on containers, App Service, Functions or Kubernetes, where local disk is not durable.
  • You have on-call responsibility, SLAs or regulatory duties to reconstruct events.
  • More than one team must read the same logs.

Not needed, or overkill, when

  • A single-instance internal tool on one VM with a handful of users. Rolling file logs plus a weekly look may be enough.
  • A prototype or spike that will be thrown away.
  • You are tempted to build your own Elasticsearch/OpenSearch cluster for 3 small services. A managed option (Azure Monitor Logs) is usually cheaper in people-time than operating a search cluster.

Wrong choices to avoid

  • Using the log store as your system of record for business events. Logs are diagnostic and may be sampled, dropped or expire. Use the database or an event store.
  • Logging everything at Debug in production “just in case”. You pay to ingest noise and bury the signal.

How to identify the problem (key signals)

  1. Engineers ask for kubectl exec/SSH access to production only to read logs.
  2. “Which instance handled that request?” is a routine question during incidents.
  3. Post-incident notes say “logs were lost when the pod restarted”.
  4. Mean time to diagnose is dominated by finding the right log, not understanding it.
  5. The same failure is investigated by grepping different formats in different services (plain text here, JSON there, no shared request id).
  6. Alerts exist for CPU and HTTP 5xx but nothing alerts on specific business errors such as “payment gateway rejected” spikes.
  7. A support ticket needs a claim’s history and the answer is “we cannot see it anymore”.

Flow Diagram

Structured logs leave ephemeral instances immediately and land in one searchable store.

flowchart LR
A["Claims replicas"] -- "ILogger + OpenTelemetry" --> W[("Log Analytics workspace")]
B["Policy replicas"] -- "stdout" --> CAE["Container Apps env logs"] --> W
C["APIM, Key Vault, SQL"] -- "diagnostic settings" --> W
DCR["Data Collection Rule - drop noise"] --> W
W --> KQL["KQL queries + workbooks"]
W --> AL["Log search alerts"]
W --> AR["Archive to Storage"]

Level 1: Beginner

Analogy: Imagine 20 shops each keeping a paper visitors’ book by the door. If a thief visited three shops, you would have to visit all three to compare notes. Log aggregation is a single security desk that receives a copy of every shop’s book, in one format, in one room.

Core ideas

  • Log event: one record: timestamp, level, message, and properties.
  • Structured logging: properties are named fields, not text glued into a sentence.
  • Collector or shipper: gets logs off the instance (agent, sidecar, platform feature or SDK).
  • Central store: indexed and searchable (Azure Log Analytics, Elasticsearch, OpenSearch).

Bad vs good (C#, .NET 10, ASP.NET Core minimal API)

// Program.cs (top-level statements)
var builder = WebApplication.CreateBuilder(args);
var app = builder.Build();
app.MapPost("/claims/{claimId}/submit", (string claimId, ILogger<Program> logger) =>
{
// BAD: string interpolation. Every message is unique text, hard to query,
// and the id is baked into the sentence.
logger.LogInformation($"Claim {claimId} submitted");
// GOOD: message template. The template stays constant; ClaimId is a searchable property.
logger.LogInformation("Claim {ClaimId} submitted", claimId);
return Results.Accepted();
});
app.Run();

Why it matters: with the template form, a backend can group by the template and filter by ClaimId = "CLM-10482". With interpolation, every line is a different string.


Level 2: Intermediate

Real setup: Angular + ASP.NET Core + SQL Server/PostgreSQL, logs to Azure Monitor

The goals are: structured logs, one correlation id from browser to database call, no PII in logs, and a single export path.

6.1 ASP.NET Core: structured logging with the source generator

using Microsoft.Extensions.Logging;
public static partial class ClaimLog
{
[LoggerMessage(EventId = 1001, Level = LogLevel.Information,
Message = "Claim {ClaimId} submitted by policy {PolicyNumber}")]
public static partial void Submitted(this ILogger logger, string claimId, string policyNumber);
[LoggerMessage(EventId = 2001, Level = LogLevel.Error,
Message = "Payment gateway rejected claim {ClaimId} with code {GatewayCode}")]
public static partial void PaymentRejected(this ILogger logger, string claimId, string gatewayCode, Exception ex);
}

LoggerMessage generates allocation-free logging code at compile time and gives every message a stable EventId, which you can alert on.

6.2 Sending logs to Azure Monitor (Application Insights, workspace-based)

Package: Azure.Monitor.OpenTelemetry.AspNetCore (the Azure Monitor OpenTelemetry Distro). It collects logs from Microsoft.Extensions.Logging, ASP.NET Core and HttpClient traces, SQL client traces, and standard metrics.

Program.cs
using Azure.Monitor.OpenTelemetry.AspNetCore;
var builder = WebApplication.CreateBuilder(args);
// Reads APPLICATIONINSIGHTS_CONNECTION_STRING from the environment.
builder.Services.AddOpenTelemetry().UseAzureMonitor();
var app = builder.Build();
app.MapPost("/claims/{claimId}/submit", (string claimId, ILogger<Program> logger) =>
{
logger.Submitted(claimId, "POL-77120");
return Results.Accepted();
});
app.Run();

Set the connection string as an environment variable (Container Apps secret, App Service setting, Key Vault reference). Do not hard-code it.

6.3 Correlation from Angular to the API

The logs are useful only if lines from the same request share an id. In ASP.NET Core, the request’s Activity trace id is automatically attached to log records by the OpenTelemetry logging provider, and by ILogger scopes when IncludeScopes is enabled. The browser can pass a W3C traceparent header, or you can add your own id:

// correlation.interceptor.ts (Angular, standalone, functional interceptor)
import { HttpInterceptorFn } from '@angular/common/http';
export const correlationInterceptor: HttpInterceptorFn = (req, next) => {
const correlationId = crypto.randomUUID();
return next(req.clone({ setHeaders: { 'X-Correlation-Id': correlationId } }));
};
app.config.ts
import { ApplicationConfig } from '@angular/core';
import { provideHttpClient, withInterceptors } from '@angular/common/http';
import { correlationInterceptor } from './correlation.interceptor';
export const appConfig: ApplicationConfig = {
providers: [provideHttpClient(withInterceptors([correlationInterceptor]))]
};
// Middleware: put the header into the logging scope so every log line in the request carries it.
app.Use(async (context, next) =>
{
var logger = context.RequestServices.GetRequiredService<ILogger<Program>>();
var correlationId = context.Request.Headers["X-Correlation-Id"].FirstOrDefault()
?? Guid.NewGuid().ToString();
context.Response.Headers["X-Correlation-Id"] = correlationId;
using (logger.BeginScope(new Dictionary<string, object> { ["CorrelationId"] = correlationId }))
{
await next();
}
});

Your own X-Correlation-Id is useful for support (the user can quote it), while the trace id from tracing is used for machine correlation. Prefer W3C traceparent between services (Day 31) and keep the custom header only for the user-visible id.

6.4 Querying (KQL) in Log Analytics

In a workspace-based Application Insights resource, ILogger output lands in the AppTraces table and exceptions in AppExceptions.

AppTraces
| where TimeGenerated > ago(1h)
| where Properties["ClaimId"] == "CLM-10482"
| order by TimeGenerated asc
| project TimeGenerated, AppRoleName, SeverityLevel, Message, OperationId

Then follow the whole request:

let op = toscalar(AppTraces
| where Properties["ClaimId"] == "CLM-10482"
| top 1 by TimeGenerated desc
| project OperationId);
union AppRequests, AppDependencies, AppTraces, AppExceptions
| where OperationId == op
| order by TimeGenerated asc

6.5 Do not log PII

Claims contain names, national ids, medical details and bank data. Log identifiers (ClaimId, PolicyNumber), never the payload. Mask at the source; a central store multiplies who can see the data.


Level 3: Advanced

Performance and cost

  • Ingestion volume is the bill. Azure Monitor Logs bills primarily on GB ingested. Cost control is a design activity: set sensible minimum levels per namespace, drop health-check noise, avoid logging whole request/response bodies.
  • Use the logging levels properly. In appsettings.Production.json set "Logging": { "LogLevel": { "Default": "Information", "Microsoft.AspNetCore": "Warning" } } and raise a single namespace to Debug temporarily during an investigation.
  • Async and non-blocking. Exporters buffer and send in the background. Never write logs synchronously to a remote store on the request path.
  • Sampling. Trace sampling in the Azure Monitor distro is rate-limited by default (5 traces per second per the distro documentation). Log volume is a separate concern; do not assume log records are sampled the same way. Verify behaviour for your distro version.

Scalability

  • Prefer stdout/stderr in containers and let the platform ship them (Container Apps log options, AKS Container Insights). It removes agents from your images and keeps the app simple (12-factor).
  • Choose table plans deliberately. Log Analytics has Analytics, Basic and Auxiliary plans. Analytics supports rich querying and is the only plan eligible for commitment tiers. Basic and Auxiliary cost less per GB ingested but restrict query capability and bill queries by data scanned. Use them for high-volume, rarely-queried logs, checking which tables support them.
  • Data Collection Rules (DCR) transformations can filter or trim data before it is stored, which is the cleanest place to remove noisy or sensitive fields.

Security

  • Logs are a data leak surface. Apply Azure RBAC on the workspace (or per-table RBAC), separate “read logs” from “manage workspace”, and use resource-context access so a team sees only its own resources’ logs.
  • Use managed identity / Microsoft Entra authentication for ingestion where the SDK supports it, instead of only relying on the connection string.
  • Protect against log injection: never write raw user input into a message template’s text; pass it as a property, and keep structured output so newlines cannot forge entries.
  • Retention must match your policy and law. Deleting data by purge does not lower retention cost; only reducing retention does.

Failure modes

FailureEffectMitigation
Collector or exporter downGap in logs during the incident you need them forBuffer locally in the agent, alert on “no logs from service X for 10 minutes”
Log flood (retry storm logging errors)Cost spike and slow queriesRate-limit repeated errors, alert on ingestion volume, daily cap as a safety net
Clock skew between hostsEvents appear out of orderUse platform time sync, order by ingestion and trace, not only wall clock
Schema driftQueries breakShared log contract (Day 40 chassis), stable property names
PII leaked into logsCompliance breachCode review rule, destructuring policies, DCR transformation, scanning

Common mistakes

  1. Logging with string interpolation instead of message templates.
  2. Logging exceptions with ex.Message only, dropping the stack trace. Pass the exception object.
  3. Different property names per team (claimId, ClaimID, claim_id).
  4. No correlation id.
  5. Treating logs as an audit trail.
  6. Turning Debug on in production permanently.
  7. Running two competing logging pipelines in one process (for example Serilog and the Azure Monitor distro both trying to own ILogger) without deciding which one exports. Pick one export path and test that logs actually arrive.

Level 4: Expert and Architect view

Options compared

OptionStrengthsWeaknessesBest fit
Azure Monitor Logs + Application InsightsManaged, KQL, native integration with Azure services, Sentinel, alerts, RBACCost grows with ingestion, KQL learning curve, Azure-centricAzure-first teams, .NET workloads
Elastic Stack (Elasticsearch + Kibana, or Elastic Cloud on Azure)Powerful full-text search, mature ecosystem, portableCluster operations or SaaS bill, sizing and mapping tuningMulti-cloud, search-heavy use cases, existing Elastic skills
OpenSearch (self-managed or managed elsewhere)Open-source, Elasticsearch-compatible API lineageYou operate it on Azure (no first-party managed service), upgrade and shard managementPortability requirements, cost of licence concerns
Grafana Loki (with Azure Managed Grafana or self-hosted)Cheap storage, indexes labels only, pairs with Prometheus and GrafanaWeaker ad-hoc full-text search, label design mattersKubernetes-heavy teams already on Grafana
Files on host + manual accessZero costEverything in section 2Single-VM hobby tools only

Patterns it combines with

  • Distributed Tracing (Day 31): logs carry trace_id/span_id, so you jump from a slow trace to its log lines.
  • Application Metrics (Day 33): alerts on log-derived counts, and metrics tell you where to look in the logs.
  • Audit Logging (Day 34): separate, immutable, longer-retention pipeline; do not mix with diagnostic logs.
  • Exception Tracking (Day 35): deduplicates and groups the errors that aggregation stores raw.
  • Health Check API (Day 36): health probe noise should be filtered out of logs.
  • Microservice Chassis (Day 40): enforces the log format, correlation, and PII rules once for all services.
  • Sidecar (Day 47): a log-shipping agent can run as a sidecar where the platform does not do it.

ADR (suitable for an architecture review)

ADR-032: Centralised structured log aggregation with Azure Monitor Logs

  • Status: Proposed
  • Context: The claims platform has 9 services on Azure Container Apps, autoscaling between 1 and 10 replicas. Diagnosing production issues currently requires reading per-replica console output, which disappears on scale-in. Support needs claim-level history within minutes. The team is .NET/Angular with Azure experience and no dedicated platform engineers to run a search cluster.
  • Decision: All services emit structured logs (message templates, LoggerMessage) to stdout/OpenTelemetry. Application Insights (workspace-based) and a shared Log Analytics workspace per environment are the central store. Every request carries a W3C trace context and a user-visible correlation id. No PII in logs. A DCR transformation drops health-check noise. Audit records go to a separate immutable store (ADR-034).
  • Alternatives considered: Elastic Cloud on Azure (better full-text search, extra vendor and cost, another identity/RBAC model); self-managed OpenSearch (operational burden we cannot staff); Loki (team has no Grafana/Prometheus practice).
  • Consequences: (+) No infrastructure to run, native alerts and RBAC, one query language. (-) Vendor lock-in to KQL and Azure Monitor, ingestion cost must be governed (daily cap, level policy, Basic/Auxiliary plans for bulky tables), KQL training needed. (-) Log store is diagnostic only; business facts stay in the database.
  • Review trigger: Monthly ingestion exceeds the agreed budget, or a multi-cloud requirement appears.

Azure implementation

Services that implement or support this topic

  • Log Analytics workspace: the central store and query engine (KQL).
  • Application Insights (workspace-based): application telemetry (requests, dependencies, traces, exceptions) stored in the workspace tables (AppRequests, AppTraces, and so on).
  • Azure Monitor OpenTelemetry Distro (Azure.Monitor.OpenTelemetry.AspNetCore): in-process exporter for .NET.
  • Azure Container Apps logging options: console and system logs can go to Log Analytics, Azure Monitor (via diagnostic settings), or be disabled.
  • AKS Container Insights: collects container stdout/stderr into the workspace (the ContainerLogV2 schema is the current one).
  • Diagnostic settings: send platform logs from Azure resources (Key Vault, Azure SQL, Application Gateway, API Management) to the workspace, Storage, or Event Hubs.
  • Data Collection Rules: define what is collected and allow transformations before storage.
  • Azure Monitor alerts (log search alerts) and Workbooks: alerting and dashboards on KQL queries.
  • Event Hubs / Storage (data export): stream or archive logs to external tools such as Elastic, or for cheap long-term archive.
  • Microsoft Sentinel: optional SIEM layered on the same workspace.

How to configure (outline)

  1. Create one Log Analytics workspace per environment (dev, test, prod), in the region of the workloads.
  2. Create a workspace-based Application Insights resource linked to it.
  3. Store APPLICATIONINSIGHTS_CONNECTION_STRING as a secret (Key Vault reference) and inject as an environment variable.
  4. Enable UseAzureMonitor() in each service (see section 6.2).
  5. For Container Apps, set the environment’s logs destination to Log Analytics.
  6. Add diagnostic settings on shared resources (API Management, Key Vault, SQL).
  7. Set retention per workspace and per table. Set a daily cap in non-production. In production treat a daily cap as an emergency brake, since hitting it stops ingestion of new data.
  8. Grant access with Azure RBAC (Log Analytics Reader for engineers, table-level or resource-context access where required).
  9. Create log search alerts, for example: more than N PaymentRejected events in 5 minutes.

Pricing and tier considerations

Confirm current numbers on the Azure Monitor pricing page before committing; I have deliberately not quoted per-GB prices because they vary by region and change.

  • Model: billed mainly on data ingested (per GB), per workspace.
  • Analytics Logs plan: includes 31 days of interactive (analytics) retention at no extra cost; interactive queries are not charged per data scanned; longer retention is charged per GB per month.
  • Basic and Auxiliary plans: lower ingestion price, per-table configuration, queries billed by data scanned, restricted query features. Use for verbose, rarely queried logs.
  • Commitment tiers: Analytics Logs only, starting at 100 GB/day, with savings of up to about 30% versus pay-as-you-go, with a 31-day commitment period. Not worth it below that daily volume.
  • Free ingestion: a few tables such as Heartbeat, Usage, Operation and AzureActivity are not billed for ingestion.
  • Levers you control: log levels, filtering health checks, DCR transformations, per-table plan, retention, sampling of traces, and avoiding body logging.
  • Cost tip: track the Usage table (Quantity by DataType) to find your top ingest contributors weekly.

Reference architecture (text)

The Angular 22 SPA is served from Azure Static Web Apps or a container, and calls the API through Azure API Management. Each request gets a W3C trace context and an X-Correlation-Id. The services (Claims, Policy, Payments, Documents) run on Azure Container Apps and use ASP.NET Core on .NET 10 (the current LTS, supported until November 2028). Each service writes structured logs via ILogger, exported by the Azure Monitor OpenTelemetry Distro to the workspace-based Application Insights resource, which stores them in the shared Log Analytics workspace. Container Apps environment console and system logs go to the same workspace. API Management, Key Vault and Azure SQL/PostgreSQL send diagnostic logs to it via diagnostic settings. A data collection rule filters noise. Engineers query with KQL and Workbooks; log search alerts notify an on-call action group (Teams, email, PagerDuty webhook). A data export rule archives selected tables to a Storage account for long-term retention. A separate immutable audit store (Day 34) is kept outside this pipeline. Optionally, Event Hubs feeds the SIEM.

Version note: Angular 22 is the current major release (June 2026) with LTS through mid-2028. Angular 20 security support ends in November 2026, so teams still on it should plan an upgrade. .NET 8 LTS support ends in November 2026, so new services should target .NET 10.


Teaching guide for my team

Explain to a beginner in 2 minutes

“When our API runs in the cloud it may be five copies today and two tomorrow, and a copy can disappear at any time, taking its notes with it. So we do not keep notes on the copy. Every copy sends its notes, in a fixed format, to one shared notebook that everyone can search. When a customer says their claim failed, we search that notebook for the claim number and see what every service said. Two habits make it work: write notes with named fields (ClaimId), not sentences with the number stuck inside, and never write private customer data into the notebook.”

Explain to an intermediate developer in 5 minutes

  1. The pipeline: app to stdout/OpenTelemetry exporter, then agent or platform, then store (Log Analytics), then query/alert.
  2. Structured logging: message templates and LoggerMessage; properties are indexed fields.
  3. Correlation: W3C trace context plus X-Correlation-Id; scope so every line carries it.
  4. Cost: ingestion is the bill, so levels, noise filtering, plans, retention.
  5. Safety: no PII, RBAC, no logs as audit.
  6. Demo: run the KQL from section 6.4 and follow one request across services.

Hands-on exercise

Task: In a small ASP.NET Core .NET 10 API for “submit claim”:

  1. Add a LoggerMessage-based log for submission and for a simulated payment rejection.
  2. Add the correlation middleware from section 6.3 and an Angular functional interceptor.
  3. Wire UseAzureMonitor() with a dev Application Insights connection string (environment variable).
  4. Submit 20 claims, 3 of them forced to fail.
  5. In Log Analytics, write a KQL query that lists the failed claims with their CorrelationId, then a second query that shows all telemetry for one failed request.

Expected outcome: the first query returns exactly 3 rows with claim ids and correlation ids; the second returns the request, the dependency call and the log lines in time order under one OperationId. The student can also show that changing the message from a template to interpolation breaks the “group by message” query.

Interview-style questions

  1. Why is logger.LogInformation($"Claim {id}") worse than logger.LogInformation("Claim {ClaimId}", id)? The template keeps a constant message and captures ClaimId as a searchable property; interpolation creates a unique string per event, so you cannot group or filter by field and you lose structured data.
  2. Containers are ephemeral. How do you avoid losing logs? Write to stdout or an exporter and have the platform or agent ship them off the instance continuously to a central store; never rely on the local disk.
  3. Is the log store a good place for audit records? Why or why not? No. Logs are diagnostic, can be filtered, sampled, expired and are broadly readable. Audit records need immutability, defined retention and restricted access, so use a dedicated store.

Mastery checklist

  • I can explain why local logs fail in an ephemeral, multi-instance environment.
  • I write logs with message templates or LoggerMessage, never interpolation.
  • I can make one request traceable from Angular through every service using trace context and a correlation id.
  • I can write KQL to find a claim’s history and to follow one operation across tables.
  • I can estimate and control ingestion cost (levels, filtering, table plans, retention, commitment tier break-even).
  • I can keep PII out of logs and explain how RBAC limits who can read them.
  • I can compare Azure Monitor Logs, Elastic/OpenSearch and Loki and justify a choice in an ADR.
  • I can name the failure modes (collector down, log flood, schema drift) and their mitigations.

Key takeaway

Get logs off ephemeral instances immediately, in a structured, correlated, PII-free format, into one searchable store, and treat ingestion volume as a cost you design, not a side effect. Logs explain what happened; they are not your audit trail or your system of record.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 37: Log Deployments & Changes

Log Deployments & Changes means every release, configuration change, feature-flag flip, and infrastructure change is recorded as a timestamped event and drawn as a marker on the same dashboards where you watch errors...

Manikandan
Manikandan·24 min read
Microservices

Day 36: Health Check API

A Health Check API is a small set of HTTP endpoints that every service exposes so that machines (Kubernetes, Azure Container Apps, App Service, load balancers, monitoring) can ask "are you alive?", "are you ready to...

Manikandan
Manikandan·19 min read
Microservices

Day 35: Exception Tracking

Exception Tracking means capturing every unhandled (and important handled) error from your services and your browser app, attaching context to it (release, user, request, breadcrumbs), grouping identical errors into...

Manikandan
Manikandan·18 min read
Microservices

Day 34: Audit Logging

Audit Logging is the practice of writing a structured, tamper-resistant record of *who* did *what*, *to which thing*, *when*, *from where*, and *with what result* every time a security-relevant or business-relevant...

Manikandan
Manikandan·24 min read