Manikandan — Manikandan
Microservices

Day 31: Distributed Tracing

ManikandanManikandan
25 min read·Updated Sep 9, 2022

Distributed tracing gives every request a single identity (a *trace ID*) that travels with it through every service, queue and database call it touches.

Intro

Distributed tracing gives every request a single identity (a trace ID) that travels with it through every service, queue and database call it touches. Each unit of work along the way is recorded as a span with a start time, duration, status and a pointer to its parent. Put together, the spans form a timeline and a call tree for one request, so you can answer “where did this request spend its time, and where did it fail?” without grepping logs on ten machines. In this lesson we use an insurance claims system: an Angular portal, an API gateway, a Claims API on .NET, a Policy API, SQL Server, and an asynchronous fraud-scoring worker fed by Azure Service Bus.

Version baseline used in this lesson: .NET 10 (LTS), Angular 22 (current major; Angular 21 is in LTS), OpenTelemetry for .NET, and the Azure.Monitor.OpenTelemetry.AspNetCore distro package for Azure Monitor / Application Insights. Version-specific notes are called out where they matter.


Why we need this

Business reasons

  • A customer calls: “I submitted my claim and the page spun for 40 seconds.” Support needs to find that exact request in minutes, not hours.
  • Claims have SLAs (for example “acknowledge within 5 seconds”, “fraud score within 2 minutes”). You can only prove or fix SLA breaches when you can measure each hop.
  • Incidents are expensive. The longer it takes to find the failing service, the longer claims processing is degraded.

Technical reasons

  • In a monolith, one stack trace and one profiler session explain a slow request. In microservices the same request crosses processes, networks, brokers and databases. No single log file or debugger holds the whole story.
  • Metrics tell you that p95 latency is up. Logs tell you what one service did. Only a trace tells you which hop in which request caused it, and how the hops relate.
  • Asynchronous flows (outbox, Service Bus, background workers) break the “one thread = one request” assumption. Trace context must be carried explicitly.
  • Standards now exist (W3C Trace Context, OpenTelemetry), so this is a commodity capability, not a custom build.

What problem it solves

The problem (from the topic list): multi-hop, often asynchronous requests cannot be traced end to end.

What goes wrong without it

Consider POST /claims in our system. The call path is:

  1. Angular portal calls the API Gateway.
  2. Gateway calls Claims API.
  3. Claims API calls Policy API (HTTP) to validate cover, then writes to SQL Server, and writes an outbox row.
  4. An outbox relay publishes ClaimSubmitted to Azure Service Bus.
  5. Fraud Scoring worker consumes the message, calls an external scoring vendor, and publishes ClaimScored.

Without tracing:

  • The portal reports 12 s. The gateway log says 11.8 s. The Claims API log says 11.7 s. Nobody can tell whether the time went into the Policy API, a slow SQL query, or a retry loop.
  • Logs from the five services are timestamped on different machines. Engineers correlate them by guessing timestamps and claim numbers, and miss the requests where the claim number was never logged.
  • The async half (“claim stuck in Scoring for 3 hours”) is invisible: nothing links the message to the HTTP request that caused it.
  • Teams blame each other. Each service’s dashboard looks “green” in isolation.

When it is needed (and when it is NOT)

Needed when

  • A single user action fans out across three or more services or queues.
  • You have asynchronous messaging and need to follow a business flow across it.
  • Latency SLOs exist and you must attribute latency to a hop.
  • Multiple teams own different services and need a shared, neutral view of a request.
  • You run retries, timeouts and circuit breakers and need to see how they compound (a retry storm is obvious in a trace, invisible in one service’s logs).

Not needed, or overkill, when

  • A single deployable (a modular monolith) with one database. Structured logs with a correlation ID and in-process Activity spans are enough. Do not build a tracing backend for one process.
  • Prototype or low-traffic internal tools where “grep the log” is genuinely faster than operating a tracing stack.
  • As a replacement for metrics or audit logs. Traces are sampled, short-lived and not a system of record. Use metrics for alerting on rates and SLOs, and audit logging (Day 34) for compliance.
  • Tracing every single request at 100% in a very high-volume system without a cost plan. Use sampling.

How to identify the problem (key signals)

  1. “Which service is slow?” takes more than 15 minutes to answer during an incident, and the answer comes from guessing.
  2. Engineers correlate logs by timestamp or by pasting a claim number into five different log queries. This is the classic smell.
  3. End-to-end latency (portal or gateway) is far higher than the sum of what any single service reports, and nobody can explain the gap (queueing, DNS, TLS, connection pool waits, retries).
  4. Support tickets say “it hung” with no request identifier, and the error page shows no reference number the user can quote.
  5. Async work goes missing: claims stuck in a status, dead-letter queue messages with no way to find the originating HTTP request.
  6. Post-mortems contain sentences like “we think the retry in Policy API caused it”: hypotheses rather than evidence.
  7. Logs show the same request handled twice (retry or redelivery) and you cannot tell whether it was one trace with two attempts or two separate requests.
  8. Dependency maps are hand-drawn and out of date; nobody knows all callers of the Policy API.

Flow Diagram

One trace ID follows the request across HTTP, the outbox and Service Bus.

flowchart LR
NG["Angular - traceparent"] --> GW["API Gateway"]
GW --> CA["Claims API span"]
CA --> PA["Policy API span"]
CA --> SQL[("SQL span")]
CA --> OB[("Outbox row stores traceparent")]
OB --> RL["Outbox relay - resumes trace"]
RL --> SB{{"Service Bus - Diagnostic-Id"}}
SB --> FW["Fraud worker span"]
CA -.-> AI["Application Insights"]
FW -.-> AI

Level 1: Beginner

Analogy. Think of a courier parcel with a tracking number. At every depot it is scanned: “arrived 10:02, left 10:15”. The tracking number is the trace ID; each depot scan is a span. If the parcel is late you open the tracking page and see exactly which depot held it. Without the tracking number, you phone every depot and ask “have you seen a brown box?”.

Vocabulary

TermMeaning.NET name
TraceThe whole journey of one requestActivity.TraceId
SpanOne timed operation inside the traceActivity
Parent spanThe span that caused this oneActivity.ParentSpanId
Tags / attributesKey-value facts on a spanactivity.SetTag(...)
Context propagationPassing the trace ID to the next serviceW3C traceparent HTTP header

In .NET, tracing is built in: System.Diagnostics.Activity is the span, and ActivitySource creates them. OpenTelemetry simply listens to these and exports them.

Minimal working example (console app, no packages needed). It creates a parent span and two children and prints them, so you can see the parent/child relationship:

using System.Diagnostics;
var source = new ActivitySource("Claims.Demo");
using var listener = new ActivityListener
{
ShouldListenTo = s => s.Name == "Claims.Demo",
Sample = (ref ActivityCreationOptions<ActivityContext> _) =>
ActivitySamplingResult.AllDataAndRecorded,
ActivityStarted = a =>
Console.WriteLine($"START {a.OperationName,-16} trace={a.TraceId} span={a.SpanId} parent={a.ParentSpanId}"),
ActivityStopped = a =>
Console.WriteLine($"STOP {a.OperationName,-16} {a.Duration.TotalMilliseconds:F0} ms")
};
ActivitySource.AddActivityListener(listener);
using (var root = source.StartActivity("SubmitClaim"))
{
using (var validate = source.StartActivity("ValidatePolicy"))
{
await Task.Delay(30);
}
using (var score = source.StartActivity("ScoreFraud"))
{
await Task.Delay(80);
}
}

Expected output (IDs will differ): all three lines share the same trace= value; ValidatePolicy and ScoreFraud show parent= equal to the span= of SubmitClaim. SubmitClaim reports about 110 ms, made up of 30 + 80. That is a trace: same trace ID, parent links, durations.

The one rule to remember: if the trace ID is not passed across a boundary, the trace breaks into two unrelated traces. Most of the work in distributed tracing is making sure every boundary (HTTP, queue, outbox, background job) passes it on.


Level 2: Intermediate

6.1 The .NET service: Claims API

Install (all versions are the current stable releases at the time you read this; pin them in your repo):

Terminal window
dotnet add package Azure.Monitor.OpenTelemetry.AspNetCore
dotnet add package Azure.Messaging.ServiceBus

The Azure Monitor distro already wires up ASP.NET Core, HttpClient and SQL client instrumentation, and exports to Application Insights. You only add your own sources.

// Program.cs (Claims API, .NET 10)
using System.Diagnostics;
using Azure.Messaging.ServiceBus;
using Azure.Monitor.OpenTelemetry.AspNetCore;
using OpenTelemetry.Trace;
// Service Bus SDK tracing: on many SDK versions this switch is required.
// Check the Azure.Messaging.ServiceBus release notes for your version.
AppContext.SetSwitch("Azure.Experimental.EnableActivitySource", true);
var builder = WebApplication.CreateBuilder(args);
builder.Services
.AddOpenTelemetry()
.UseAzureMonitor() // reads APPLICATIONINSIGHTS_CONNECTION_STRING
.WithTracing(t => t
.AddSource("Claims.Api") // our own spans
.AddSource("Azure.Messaging.ServiceBus.*")); // SDK spans for send/process
builder.Services.AddDbContext<ClaimsDb>(o =>
o.UseSqlServer(builder.Configuration.GetConnectionString("Claims")));
builder.Services.AddHttpClient<IPolicyClient, PolicyClient>(c =>
c.BaseAddress = new Uri(builder.Configuration["Services:Policy"]!));
builder.Services.AddProblemDetails(); // adds traceId to error responses
var app = builder.Build();
app.UseExceptionHandler();
app.MapPost("/claims", async (SubmitClaim cmd, ClaimsDb db, IPolicyClient policies) =>
{
using var activity = Telemetry.Source.StartActivity("claims.submit");
activity?.SetTag("claims.policy_id", cmd.PolicyId); // ID, never names or health data
var cover = await policies.CheckCoverAsync(cmd.PolicyId); // HttpClient span is automatic
if (!cover.IsActive)
{
activity?.SetStatus(ActivityStatusCode.Error, "policy inactive");
return Results.Problem("Policy is not active", statusCode: 422);
}
var claim = Claim.Create(cmd);
db.Claims.Add(claim);
db.Outbox.Add(OutboxMessage.For(claim, Activity.Current?.Id)); // carry the trace context (see 6.2)
await db.SaveChangesAsync(); // SQL span is automatic
activity?.SetTag("claims.claim_id", claim.Id);
return Results.Accepted($"/claims/{claim.Id}", new { claim.Id });
});
app.Run();
public static class Telemetry
{
public static readonly ActivitySource Source = new("Claims.Api");
}

Notes:

  • builder.Services.AddProblemDetails() makes error responses include a traceId. Show that ID in the Angular error screen so support can quote it.
  • Incoming traceparent headers are read by ASP.NET Core automatically, and outgoing HttpClient calls add it. That is why the Policy API span appears under the Claims API span with no code.

6.2 Async hop 1: the outbox (context must be stored)

A database row has no HTTP headers, so the trace context has to be saved in the row, then restored by the relay. This is the most common place traces break.

// OutboxMessage.cs (excerpt)
public class OutboxMessage
{
public Guid Id { get; set; } = Guid.NewGuid();
public string Type { get; set; } = "";
public string Payload { get; set; } = "";
public string? Traceparent { get; set; } // W3C format: 00-<traceId>-<spanId>-<flags>
public DateTime? PublishedUtc { get; set; }
public static OutboxMessage For(Claim c, string? traceparent) => new()
{
Type = "ClaimSubmitted",
Payload = System.Text.Json.JsonSerializer.Serialize(new { c.Id, c.PolicyId }),
Traceparent = traceparent
};
}
// OutboxRelay.cs (BackgroundService excerpt)
foreach (var msg in pending)
{
ActivityContext.TryParse(msg.Traceparent, null, out var parent);
// Resume the original trace. Producer kind marks this as "sending a message".
using var span = Telemetry.Source.StartActivity(
"outbox.publish", ActivityKind.Producer, parent);
span?.SetTag("messaging.destination.name", "claims-submitted");
await sender.SendMessageAsync(new ServiceBusMessage(msg.Payload)
{
Subject = msg.Type,
MessageId = msg.Id.ToString() // also gives dedup on the consumer
});
msg.PublishedUtc = DateTime.UtcNow;
}
await db.SaveChangesAsync(stoppingToken);

ActivityContext.TryParse takes the stored traceparent; the relay span becomes a child of the original request span, even if it runs minutes later on another pod. Because Activity.Current is now the relay span, the Service Bus SDK stamps the outgoing message with that context.

6.3 Async hop 2: the consumer

With the Service Bus SDK’s ServiceBusProcessor, the SDK creates a processing span and links it to the send span through the message’s Diagnostic-Id property (Microsoft documents this mechanism for Azure.Messaging.ServiceBus). Your handler code runs inside that span:

processor.ProcessMessageAsync += async args =>
{
using var span = Telemetry.Source.StartActivity("fraud.score"); // child of the SDK processing span
var claimId = System.Text.Json.JsonDocument.Parse(args.Message.Body)
.RootElement.GetProperty("Id").GetGuid();
span?.SetTag("claims.claim_id", claimId);
var score = await vendor.ScoreAsync(claimId); // HttpClient span, if vendor client is created via IHttpClientFactory
span?.SetTag("fraud.score", score);
await args.CompleteMessageAsync(args.Message);
};

6.4 The Angular side

Angular’s HttpClient does not create traces by itself. The simplest correct option is the Application Insights JavaScript SDK (it creates a browser span and sends traceparent). If you want a dependency-free start, a functional interceptor works in Angular 22 (and 21, 20):

trace.interceptor.ts
import { HttpInterceptorFn } from '@angular/common/http';
function randomHex(bytes: number): string {
const a = new Uint8Array(bytes);
crypto.getRandomValues(a);
return Array.from(a, b => b.toString(16).padStart(2, '0')).join('');
}
export const traceInterceptor: HttpInterceptorFn = (req, next) => {
if (!req.url.startsWith('/api')) return next(req); // only our own API
const traceparent = `00-${randomHex(16)}-${randomHex(8)}-01`;
return next(req.clone({ setHeaders: { traceparent } }));
};
app.config.ts
import { ApplicationConfig } from '@angular/core';
import { provideHttpClient, withInterceptors } from '@angular/common/http';
import { traceInterceptor } from './trace.interceptor';
export const appConfig: ApplicationConfig = {
providers: [provideHttpClient(withInterceptors([traceInterceptor]))]
};

Two caveats: the trailing -01 flag says “sampled”, so a hand-rolled header forces recording for every browser request; and if the API is on another origin, the API’s CORS policy must allow the traceparent header (WithHeaders("traceparent", "tracestate", ...)). On errors, show the traceId from the ProblemDetails body:

this.http.post('/api/claims', body).subscribe({
error: (e) => this.error.set(`Submission failed. Reference: ${e.error?.traceId ?? 'n/a'}`)
});

6.5 Reading a trace: KQL in Application Insights / Log Analytics

Given a trace ID from a support ticket:

union requests, dependencies, exceptions, traces
| where operation_Id == "4bf92f3577b34da6a3ce929d0e0e4736"
| project timestamp, itemType, cloud_RoleName, name, message, duration, success
| order by timestamp asc

operation_Id is the trace ID; cloud_RoleName tells you which service emitted each item. To find the slowest hops for the claim endpoint over the last hour:

dependencies
| where timestamp > ago(1h) and operation_Name == "POST /claims"
| summarize p95 = percentile(duration, 95), calls = count() by target, cloud_RoleName
| order by p95 desc

Also use the portal’s Transaction search → End-to-end transaction details and Application map views for the visual waterfall.


Level 3: Advanced

7.1 Sampling: the cost/visibility dial

  • Head sampling decides at the start of the trace (for example keep 10%). Cheap and simple, but it decides before knowing whether the request will fail or be slow, so it can discard exactly the traces you want.
  • Tail sampling decides after the trace completes (keep all errors and slow traces, sample the rest). It needs a component that sees all spans of a trace, typically an OpenTelemetry Collector with the tail_sampling processor, and it costs memory.
  • The Azure Monitor distro supports configurable sampling (a fixed ratio and, in recent versions, a rate-limited mode measured in traces per second). Check the docs for your package version, and note how the two settings interact before you rely on either.
  • Sampling must be consistent across services. If the Gateway keeps a trace but the Claims API drops it, you get holes. Use the parent’s sampled flag (ParentBased sampler; the -01 flag in traceparent) so children follow the root’s decision.
  • Do not sample aggressively on low-volume, high-value flows (claim payouts). Sample the chatty ones (health checks, reads).

7.2 Performance and cost

  • Tracing overhead is small (microseconds per span) when spans are coarse. It becomes real when you create spans in tight loops (per row, per property). Rule of thumb: one span per I/O call or meaningful business step.
  • Telemetry volume is the real cost. Exporters batch and are asynchronous, but ingestion is billed by the GB. A chatty service with 200 spans per request at 500 requests/s adds up fast.
  • Exclude noise: health-check endpoints, static files, and Service Bus lock-renewal chatter.

7.3 Security and privacy

  • Never put PII or claim details in span tags or baggage. Names, addresses, medical information and card numbers must not be tags. Use opaque IDs (claim_id) and look details up in the system of record. Traces are widely readable by engineers and are often retained for weeks.
  • SQL text: instrumentation can capture db.statement. Parameterised text is fine; make sure you never build SQL with literal values, or those values end up in telemetry.
  • Baggage (key-value pairs propagated to every downstream call, including third parties) is dangerous: it leaks in outbound calls. Keep it tiny and non-sensitive, or don’t use it.
  • Do not trust inbound traceparent from the public internet blindly. A malicious client can supply an arbitrary trace ID or force the sampled flag on. At the edge, start a fresh trace (or restrict who can influence sampling) rather than honouring untrusted headers.
  • Protect access to the telemetry store with RBAC; it contains request paths, IDs and error messages.

7.4 Failure modes and how they look

FailureSymptom in the trace UITypical cause
Broken traceTwo unrelated traces for one user actionA hop dropped traceparent: custom HttpClient without the factory, a proxy stripping headers, an outbox with no stored context
Orphan spansSpans whose parent is missingParent service not instrumented, or sampled out
Missing async continuationTrace ends at “publish”Consumer not listening to the Azure.Messaging.ServiceBus.* source, or SDK tracing switch off
Giant traceThousands of spans in one traceLong-lived background loop reuses one root activity; batch job creating a span per item
Wrong durationsChild longer than parent, negative gapsClock skew between hosts; fire-and-forget child outliving the parent
Lost context in background workActivity.Current is null in Task.Run or a Timer callbackContext flows with async/await but not through arbitrary queues or long-lived singletons; capture and pass it explicitly

7.5 Common mistakes

  1. Instrumenting only the API and forgetting workers, the outbox relay and scheduled jobs.
  2. Creating a new HttpClient() per call (bypasses IHttpClientFactory instrumentation and causes socket exhaustion).
  3. Using one trace for a batch consumer: a consumer that processes 100 messages should create one span with links to 100 producer contexts, not make all 100 traces children of one.
  4. High-cardinality data as span names (GET /claims/8f3a...). Names must be low-cardinality (GET /claims/{id}); put IDs in tags.
  5. Treating traces as logs: writing large payloads into span events.
  6. Not returning the trace ID to the caller, so support can never find the trace.
  7. Never testing propagation: it silently breaks on refactors. Add an integration test that asserts the child span shares the parent’s TraceId.

Level 4: Expert and Architect view

8.1 Trade-offs and alternatives

OptionStrengthsWeaknessesBest fit
OpenTelemetry SDK + Azure Monitor / Application InsightsVendor-neutral instrumentation, managed backend, Application Map, tight Azure integrationBackend is Azure-specific; cost grows with volume; some OTel features arrive later than in OSS backendsAzure-centred estates (our case)
OpenTelemetry + OTel Collector + Jaeger or Grafana TempoFull control, tail sampling, no per-GB vendor costYou operate storage, scaling and upgradesMulti-cloud or strict data residency, platform team available
Commercial APM (Datadog, Dynatrace, New Relic)Rich UI, AI analysis, one pane for logs, metrics and tracesLicence cost, agent lock-in riskLarge orgs standardised on one vendor
ZipkinSimple, lightweightSmaller ecosystem than OTel-native backendsSmall or legacy Java estates
Correlation IDs in logs onlyTrivial to build, no new infrastructureNo timing tree, no parent/child, no dependency map; manual joinsStop-gap for a 2-3 service system
Application Insights classic SDK (Microsoft.ApplicationInsights.AspNetCore)Familiar, feature-richMicrosoft’s direction is OpenTelemetry-based instrumentation; classic SDK is not the recommended path for new workExisting apps not yet migrated

Patterns it combines with

  • Log Aggregation (Day 32): logs enriched with TraceId/SpanId so you jump from a trace span to its logs and back.
  • Application Metrics (Day 33): metrics for detection and alerting; exemplars link a slow metric bucket to a concrete trace.
  • Exception Tracking (Day 35) and Audit Logging (Day 34): attach the trace ID to each record.
  • Transactional Outbox (Day 12), Messaging (Day 16), Saga (Day 7): each needs explicit context propagation; a saga is a long-running trace or a set of linked traces.
  • API Gateway (Day 19) / BFF (Day 20): the natural place to start or validate the trace and to add the trace ID to responses.
  • Circuit Breaker / Retry (Days 26, 28): make each retry attempt its own span so retry storms are visible.
  • Health Check API and Log Deployments & Changes: overlay deployment markers on latency charts to tie regressions to releases.

8.2 ADR (architecture review)

ADR-031: Adopt OpenTelemetry-based distributed tracing with Azure Monitor as the backend

Status: Proposed

Context. The claims platform has grown to seven services plus asynchronous processing via Service Bus. Incident diagnosis relies on timestamp-based log correlation, taking 1-3 hours for cross-service latency issues. SLAs on claim acknowledgement (5 s) and fraud scoring (2 min) cannot be verified per hop. The platform runs on Azure with .NET 10 and Angular.

Decision. Instrument all services with OpenTelemetry (ActivitySource/Activity in .NET), propagate W3C Trace Context on HTTP and through Service Bus and the outbox, and export to Application Insights (workspace-based) through the Azure Monitor OpenTelemetry distro. Return the trace ID in every error response. Sample with a rate limit by default and keep 100% of failed and slow requests where the platform allows. No PII in span attributes or baggage.

Alternatives considered. Self-hosted Jaeger/Tempo (rejected for now: operational burden with no platform team capacity); commercial APM (rejected: cost and duplication of existing Azure Monitor licences); log correlation IDs only (rejected: no timing tree).

Consequences. (+) Per-hop latency evidence, dependency map, shorter incident triage. (+) Instrumentation stays vendor-neutral, so the backend can change later. (-) Telemetry ingestion cost that must be governed by sampling and daily caps. (-) Every new service and worker must follow the propagation checklist; a propagation test is added to the service template (Day 41). (-) Sampling means some requests will have no trace.

Review trigger. Revisit if monthly telemetry ingestion exceeds the agreed budget, or if tail-sampling requirements outgrow the managed sampler.


Azure implementation

9.1 Which Azure services

  • Application Insights (workspace-based) with a Log Analytics workspace: stores and queries traces (requests, dependencies, exceptions, traces tables). Features: Application Map, Transaction search and end-to-end transaction view, Live Metrics, availability tests.
  • Azure Monitor OpenTelemetry Distro (Azure.Monitor.OpenTelemetry.AspNetCore): the recommended way to send OpenTelemetry data from ASP.NET Core to Application Insights. For worker services and console apps, the exporter package Azure.Monitor.OpenTelemetry.Exporter is used with the OpenTelemetry SDK.
  • Azure Service Bus: the .NET SDK emits Activity spans for send and process and propagates context via the Diagnostic-Id property (see Level 2).
  • Azure API Management: can forward the traceparent header and send its own diagnostics to Application Insights; verify the current APIM diagnostics and trace-context options for your tier.
  • Azure Functions (isolated worker, .NET): supports OpenTelemetry mode; check the current Functions documentation for the host setting and package versions for your runtime.
  • Azure Container Apps / AKS / App Service: hosts for your services. All simply need the connection string in an environment variable. Container Apps and AKS can also run an OpenTelemetry Collector (as a separate app or a DaemonSet/sidecar) if you need tail sampling or multi-backend export.
  • Azure Monitor alerts, Workbooks, Azure Managed Grafana: alerting and dashboards on top of the same data.

9.2 Configuration

  1. Create a Log Analytics workspace and a workspace-based Application Insights resource (one per environment; one shared workspace per environment is common so cross-service queries work).
  2. Give each service a distinct cloud role name so the Application Map and KQL show claims-api, policy-api, fraud-worker:
builder.Services.AddOpenTelemetry()
.ConfigureResource(r => r.AddService(serviceName: "claims-api", serviceVersion: "1.4.2"))
.UseAzureMonitor();

(AddService needs using OpenTelemetry.Resources;. The distro maps service.name to the cloud role name.)

  1. Provide the connection string as an app setting or environment variable, ideally from Key Vault:
Terminal window
az containerapp update -n claims-api -g rg-claims-prod \
--set-env-vars APPLICATIONINSIGHTS_CONNECTION_STRING=secretref:appinsights-conn
  1. Set a daily cap on the workspace (or Application Insights) as a safety net, but remember a cap causes data loss when hit, so alert at 80% of the cap.
  2. Configure sampling in the app (ratio or rate-limited) rather than relying on the daily cap.
  3. Restrict access with Azure RBAC (Monitoring Reader for engineers; Contributor only for platform team). Consider Microsoft Entra authentication for ingestion instead of only the instrumentation key.
  4. Add an alert rule on failed dependencies (dependencies | where success == false) and on p95 request duration.

9.3 Pricing and tier considerations

Prices change and vary by region, so always confirm on the Azure Monitor pricing page and with the Azure pricing calculator. What is stable in structure:

  • Workspace-based Application Insights data is billed as Log Analytics ingestion, per GB, on the default pay-as-you-go model. There is no per-trace or per-host charge.
  • Commitment tiers (starting at 100 GB/day per Microsoft’s cost documentation) can give a discount of up to about 30% versus pay-as-you-go. They only pay off at sustained high volume, and you cannot drop to a lower tier during the 31-day commitment.
  • Retention: Analytics-plan tables include 31 days of retention in the ingestion price and longer retention is charged per GB per month. Application Insights tables have their own included-retention rules; verify the current included period for your resource before planning retention.
  • Billed size differs from incoming size (Microsoft states billed size is on average about 25% smaller than the raw JSON).
  • Some Azure Monitor features include a small monthly free ingestion allowance; check the current amount rather than assuming.
  • Cost levers, in order of impact: sampling; excluding health-check and other noisy spans; not putting large payloads in span events; capping cardinality; choosing a commitment tier only once volume is proven.
  • Self-hosting Jaeger/Tempo on AKS trades the per-GB charge for compute, storage and engineering time.

9.4 Reference architecture (text)

Angular 22 portal (App Insights JS SDK or traceInterceptor)
| traceparent
v
Azure API Management / Application Gateway -- forwards traceparent, returns trace id
|
v
Azure Container Apps: claims-api (.NET 10, Azure Monitor distro)
|-- HTTP --> policy-api (Container Apps) -> Azure SQL / SQL Server
|-- EF Core --> Azure SQL (Claims DB) + Outbox table (stores traceparent)
v
outbox-relay (hosted service) -- resumes trace from stored traceparent
v
Azure Service Bus topic "claims-submitted" (Diagnostic-Id on message)
v
fraud-worker (Container Apps, .NET 10) --> external scoring API
All services --OTLP/Azure Monitor exporter--> Application Insights (workspace-based)
--> Log Analytics workspace
--> Workbooks / alerts / Managed Grafana
Secrets (connection string) from Azure Key Vault via managed identity.
Optional: OpenTelemetry Collector app between services and Azure Monitor for tail sampling.

Teaching guide for my team

10.1 Beginner: 2-minute explanation

“When you send a parcel, it gets a tracking number and every depot scans it. A trace ID is the tracking number for a request. Every service that touches the request records a span: ‘I started at this time, I took this long, I succeeded or failed.’ All spans with the same trace ID make one timeline. The only rule: every service must pass the trace ID on to the next one, in the traceparent header for HTTP or inside the message for queues. In .NET this is built in: Activity is a span, and libraries like HttpClient and ASP.NET Core do the passing for us. Our jobs are to keep it unbroken across queues and the outbox, and to not put personal data in it.”

10.2 Intermediate: 5-minute explanation

  1. Show the three signals: metrics say “p95 latency is up”, logs say “what one service did”, traces say “which hop of which request”. Open a real trace in Application Insights and walk the waterfall.
  2. Explain propagation: W3C traceparent = version, trace ID, parent span ID, flags. ASP.NET Core reads it, HttpClient writes it, Service Bus SDK puts it in Diagnostic-Id. The outbox and background jobs are on us.
  3. Explain instrumentation: ActivitySource in our code, AddSource to register it, distro exports it. Show the claims.submit span and the two automatic children.
  4. Explain sampling: we cannot afford 100%, so we sample, and the decision must be consistent through the chain.
  5. Explain safe use: no PII in tags; return the trace ID on errors; low-cardinality span names; use links for batches.
  6. Show the query: paste a trace ID in the KQL from section 6.5.

10.3 Hands-on exercise

Goal: make a trace that crosses HTTP and a queue, then find and fix a deliberately broken hop.

  1. Run Claims API, Policy API and Fraud Worker locally with the console exporter (AddConsoleExporter()), or use the Azure Monitor distro with a dev Application Insights resource.
  2. Send POST /claims and copy the traceId from the response (ProblemDetails on failure, or a response header you add).
  3. Confirm you see spans from claims-api, policy-api, and fraud-worker sharing one trace ID.
  4. Break it: in the outbox relay, remove the line that parses Traceparent and starts the span with a parent. Resend the request.
  5. Observe: the fraud worker spans now sit in a different trace, and the Application Map shows no edge from the relay to the worker.
  6. Fix it: restore ActivityContext.TryParse(...) and the parent argument.
  7. Add a test: an integration test that publishes through the relay and asserts the consumer activity’s TraceId equals the originating request’s TraceId.

Expected outcome: after step 3 there is one trace with roughly 8-12 spans; after step 4 there are two traces; after step 6 one trace again, and the test in step 7 fails on the broken version and passes on the fixed one.

10.4 Interview-style questions

  1. What is the difference between a trace and a span? A trace is the end-to-end record of one request, identified by a trace ID. A span is one timed operation in that trace, with its own span ID and a parent span ID. A trace is a tree of spans.

  2. Why do traces often “break” at message queues, and how do you fix it? HTTP libraries propagate traceparent automatically; queues and database rows do not. You must put the trace context in the message (or outbox row) when producing, and restore it as the parent (or a link) when consuming. The Azure Service Bus SDK does this for you via Diagnostic-Id; the outbox and custom transports need explicit code.

  3. Your telemetry bill doubled. What do you change first, and what is the risk? Introduce or tighten sampling and drop noisy spans such as health checks. The risk is losing rare failure traces with head sampling, so keep errors and slow requests through tail sampling or a rate-limited/adaptive approach, and keep sampling consistent across services.


Mastery checklist

  • I can explain trace, span, parent span, tags, links, baggage and traceparent without notes.
  • I can add a custom ActivitySource span in a .NET service, register it with AddSource, and see it in Application Insights.
  • I can trace a request across HTTP, the outbox and Service Bus and prove it is a single trace.
  • I can diagnose a broken trace (missing propagation, unregistered source, sampled-out parent) from what the UI shows.
  • I can write a KQL query that reconstructs one request from its operation_Id and finds the slowest dependencies.
  • I can choose head vs tail sampling for a given system and justify it with cost and visibility trade-offs.
  • I can list what must never go into spans or baggage, and I return the trace ID to callers on errors.
  • I can write an ADR comparing Azure Monitor, self-hosted OpenTelemetry backends and commercial APM.

Key takeaway

A distributed trace is one request’s tracking number plus a timed record from every hop; it only works if the context is passed across every boundary, especially queues and the outbox. Instrument with OpenTelemetry, sample deliberately, keep PII out, and always hand the trace ID back to the caller.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 37: Log Deployments & Changes

Log Deployments & Changes means every release, configuration change, feature-flag flip, and infrastructure change is recorded as a timestamped event and drawn as a marker on the same dashboards where you watch errors...

Manikandan
Manikandan·24 min read
Microservices

Day 36: Health Check API

A Health Check API is a small set of HTTP endpoints that every service exposes so that machines (Kubernetes, Azure Container Apps, App Service, load balancers, monitoring) can ask "are you alive?", "are you ready to...

Manikandan
Manikandan·19 min read
Microservices

Day 35: Exception Tracking

Exception Tracking means capturing every unhandled (and important handled) error from your services and your browser app, attaching context to it (release, user, request, breadcrumbs), grouping identical errors into...

Manikandan
Manikandan·18 min read
Microservices

Day 34: Audit Logging

Audit Logging is the practice of writing a structured, tamper-resistant record of *who* did *what*, *to which thing*, *when*, *from where*, and *with what result* every time a security-relevant or business-relevant...

Manikandan
Manikandan·24 min read