Manikandan — Manikandan
Microservices

Day 35: Exception Tracking

ManikandanManikandan
18 min read·Updated Sep 13, 2022

Exception Tracking means capturing every unhandled (and important handled) error from your services and your browser app, attaching context to it (release, user, request, breadcrumbs), grouping identical errors into...

Intro

Exception Tracking means capturing every unhandled (and important handled) error from your services and your browser app, attaching context to it (release, user, request, breadcrumbs), grouping identical errors into one “issue”, counting them, and alerting the right person when a new or spiking issue appears. Instead of grepping raw logs for stack traces, you get one deduplicated list: “this exception, first seen in release 1.42, 3,200 events, 41 users affected”. This lesson uses an insurance claims system (Claims API in .NET, Claims Web in Angular, SQL Server/PostgreSQL) and covers both Sentry and Azure Application Insights.

Why we need this

  • Logs record events; they do not manage problems. A log line saying NullReferenceException appears 8,000 times as 8,000 lines. Nobody knows if that is one bug or eight, whether it is new, or who it affects.
  • Users do not report most errors. A claims adjuster whose “Approve” button silently fails usually retries or phones a colleague. Browser-side errors never reach server logs at all.
  • Release confidence. After a deployment you need an answer to “did this release introduce new errors?” within minutes, not after the next customer escalation.
  • Triage and ownership. Issues can be assigned, resolved, ignored, and reopened automatically if the error returns in a later release (a regression).
  • Cost of debugging. An error event carries the stack trace, request URL, user id, release, and the breadcrumbs (log lines, SQL calls, HTTP calls) leading up to it, so the developer often needs no reproduction.

What problem it solves

Problem (from the topic list): unhandled errors hide in raw logs. The tracker captures, deduplicates, and alerts on stack traces (e.g., Sentry).

Without it, in the claims system:

  1. POST /claims/{id}/approve starts throwing DbUpdateConcurrencyException after a Friday deployment. It is one line among 2 million log lines per day. Nobody looks until Monday.
  2. An Angular component throws Cannot read properties of undefined (reading 'policyNumber') for users on one browser version. Backend logs show nothing; the only evidence is a phone call.
  3. The same root cause produces different-looking log messages (different claim ids in the text), so counting by message text gives 5,000 “unique” errors.
  4. The on-call engineer is paged by a generic “5xx rate high” alert and then spends 40 minutes finding which exception caused it.

When it is needed (and when it is NOT)

Needed when:

  • You run more than one service or more than one deployment per week, so regressions are likely.
  • You have a browser or mobile client whose errors never reach your server.
  • Several teams own code paths and need error ownership and routing.
  • Support asks “can you find what went wrong for claim C-10293?” and you need to search by user, claim, or trace id.

NOT needed (or overkill) when:

  • A single internal script or batch job that emails on failure is enough.
  • You already have Application Insights and only need basic server-side failure counts; adding a second tool creates two places to look. Choose one primary tracker per system.
  • As a replacement for metrics or tracing. Exception tracking answers “what broke”; metrics (Day 33) answer “how much and how fast”; traces (Day 31) answer “where in the call chain”.
  • As an audit trail (Day 34). Error tools sample, drop, and expire data by design, so they must never be your record of who did what.

How to identify the problem (key signals)

  1. Engineers answer “is this new?” by scrolling or grep -c in a log viewer.
  2. The team learns about production errors from users, support tickets, or the CEO, before any alert.
  3. Log search for Exception returns thousands of hits, and nobody can say how many distinct bugs that is.
  4. Frontend bugs are only diagnosed from screenshots and “it happens in Chrome” messages.
  5. After each release someone manually watches logs for 30 minutes (“release babysitting”).
  6. The same bug is fixed twice by two people because there is no shared issue with an owner.
  7. Catch blocks that swallow exceptions (catch { }) or log-and-continue exist in many places, hiding failures that never surface anywhere.

Flow Diagram

Exceptions are captured once, grouped by fingerprint, and alerted on when new or spiking.

flowchart LR
BE["Claims API global handler"] --> EV["Error event + release + trace_id"]
FE["Angular ErrorHandler"] --> EV
EV --> FIL{"Expected outcome? 404/409/cancel"}
FIL -- "yes" --> DROP["Drop"]
FIL -- "no" --> SCR["Scrub PII"]
SCR --> FP["Fingerprint: type + in-app frame"]
FP --> IS["Issue: count, users, first seen"]
IS --> AL["Alert on new / regression / spike"]

Level 1: Beginner

Analogy. A hospital triage desk. Patients (exceptions) arrive constantly. The desk does not treat them; it records who came, groups similar cases (“30 people with the same food poisoning”), and calls the right doctor when a new or serious case arrives.

Core ideas:

  • Event: one occurrence of an exception, with stack trace and context.
  • Issue (group): many events that share a root cause, identified by a fingerprint.
  • Fingerprint: a hash of things that identify the root cause, typically the exception type plus the top in-app stack frames, and not the message text (which contains variable data like claim ids).
  • Release: the version that produced the event, which lets you see regressions.
  • Breadcrumbs: the trail of actions before the crash.

A minimal, working illustration of grouping (this is a teaching tool, not a replacement for a real tracker):

using System.Collections.Concurrent;
using System.Diagnostics;
using System.Security.Cryptography;
using System.Text;
public sealed record ErrorGroup(
string Fingerprint, string Title, int Count,
DateTimeOffset FirstSeen, DateTimeOffset LastSeen);
public sealed class MiniErrorTracker
{
private readonly ConcurrentDictionary<string, ErrorGroup> _groups = new();
public ErrorGroup Capture(Exception ex)
{
var fingerprint = Fingerprint(ex);
var now = DateTimeOffset.UtcNow;
return _groups.AddOrUpdate(
fingerprint,
_ => new ErrorGroup(fingerprint, $"{ex.GetType().Name}: {ex.Message}", 1, now, now),
(_, g) => g with { Count = g.Count + 1, LastSeen = now });
}
// Group by exception type + first frame inside OUR code, ignoring the message text.
private static string Fingerprint(Exception ex)
{
var frame = new StackTrace(ex, false).GetFrames()
.Select(f => f.GetMethod())
.FirstOrDefault(m => m?.DeclaringType?.Namespace?.StartsWith("Contoso.Claims") == true);
var key = $"{ex.GetType().FullName}|{frame?.DeclaringType?.FullName}.{frame?.Name}";
return Convert.ToHexString(SHA256.HashData(Encoding.UTF8.GetBytes(key)))[..12];
}
}

Throw new InvalidOperationException("Claim C-1 already approved") and new InvalidOperationException("Claim C-2 already approved") from the same method: you get one group with Count = 2, even though the messages differ. That is the whole idea of deduplication.

Level 2: Intermediate

6.1 ASP.NET Core (.NET 10 LTS) with Sentry

Package: Sentry.AspNetCore. Setup in Program.cs, with a global exception handler that returns a safe ProblemDetails response carrying a trace id support can search for:

using System.Diagnostics;
using System.Text.RegularExpressions;
using Microsoft.AspNetCore.Diagnostics;
using Microsoft.AspNetCore.Mvc;
using Sentry;
var policyNo = new Regex(@"POL-\d{8}", RegexOptions.Compiled);
var builder = WebApplication.CreateBuilder(args);
builder.WebHost.UseSentry(o =>
{
o.Dsn = builder.Configuration["Sentry:Dsn"];
o.Environment = builder.Environment.EnvironmentName;
o.Release = builder.Configuration["APP_RELEASE"]; // e.g. claims-api@1.42.0, set by CI
o.SendDefaultPii = false; // no IPs/cookies/user data by default
o.TracesSampleRate = 0.05;
// Expected outcomes are not bugs: keep them out of the error inbox.
o.AddExceptionFilterForType<OperationCanceledException>();
o.AddExceptionFilterForType<ClaimNotFoundException>();
o.SetBeforeSend((ev, hint) =>
{
// Scrub business identifiers from exception messages before they leave the process.
foreach (var se in ev.SentryExceptions ?? [])
se.Value = se.Value is null ? null : policyNo.Replace(se.Value, "POL-********");
// Group all SQL Server deadlocks (error 1205) into one issue, whatever the query.
if (ev.Exception is Microsoft.Data.SqlClient.SqlException { Number: 1205 })
ev.SetFingerprint(["sql-deadlock"]);
return ev;
});
});
builder.Services.AddExceptionHandler<GlobalExceptionHandler>();
builder.Services.AddProblemDetails();
var app = builder.Build();
app.UseExceptionHandler();
app.MapGet("/claims/{id}", (string id) =>
id == "missing" ? throw new ClaimNotFoundException(id) : Results.Ok(new { id }));
app.Run();
public sealed class ClaimNotFoundException(string claimId)
: Exception($"Claim {claimId} not found");
internal sealed class GlobalExceptionHandler(ILogger<GlobalExceptionHandler> logger) : IExceptionHandler
{
public async ValueTask<bool> TryHandleAsync(HttpContext ctx, Exception ex, CancellationToken ct)
{
var traceId = Activity.Current?.TraceId.ToString() ?? ctx.TraceIdentifier;
if (ex is ClaimNotFoundException)
{
logger.LogInformation("Claim not found. TraceId {TraceId}", traceId);
ctx.Response.StatusCode = StatusCodes.Status404NotFound;
await ctx.Response.WriteAsJsonAsync(
new ProblemDetails { Status = 404, Title = "Claim not found" }, ct);
return true;
}
SentrySdk.ConfigureScope(s => s.SetTag("trace_id", traceId));
// Sentry's ILogger integration turns Error-level logs with an exception into events.
logger.LogError(ex, "Unhandled exception. TraceId {TraceId}", traceId);
ctx.Response.StatusCode = StatusCodes.Status500InternalServerError;
await ctx.Response.WriteAsJsonAsync(new ProblemDetails
{
Status = 500,
Title = "Unexpected error",
Extensions = { ["traceId"] = traceId }
}, ct);
return true;
}
}

Notes: the call to AddExceptionFilterForType and the Sentry:Dsn key come from the Sentry .NET SDK options; the required Microsoft.Data.SqlClient reference is already present in SQL Server projects. For PostgreSQL, the equivalent check is Npgsql.PostgresException { SqlState: "40P01" } (deadlock) or "40001" (serialization failure). Ship portable PDB files (or <DebugType>embedded</DebugType>) with the deployment so stack traces carry line numbers.

6.2 Angular (current major: 22) with Sentry

Angular 22 is the current major (released June 2026). The Sentry Angular SDK supports Angular 17 and later via its setup wizard (npx @sentry/wizard@latest -i angular).

main.ts:

import { bootstrapApplication } from '@angular/platform-browser';
import { HttpErrorResponse } from '@angular/common/http';
import * as Sentry from '@sentry/angular';
import { appConfig } from './app/app.config';
import { App } from './app/app';
Sentry.init({
dsn: 'https://<key>@o<orgId>.ingest.sentry.io/<projectId>',
environment: 'production',
release: 'claims-web@1.42.0', // injected by the build pipeline
sendDefaultPii: false,
integrations: [Sentry.browserTracingIntegration()],
tracesSampleRate: 0.05,
// Send trace headers only to our own API so backend and frontend errors link up.
tracePropagationTargets: [/^https:\/\/api\.claims\.contoso\.com/],
ignoreErrors: ['ResizeObserver loop completed with undelivered notifications.'],
beforeSend(event, hint) {
const err = hint?.originalException;
// A 4xx from our API is a business outcome (validation, not found), not a client bug.
if (err instanceof HttpErrorResponse && err.status >= 400 && err.status < 500) {
return null;
}
return event;
},
});
bootstrapApplication(App, appConfig).catch((e) => console.error(e));

app.config.ts:

import { ApplicationConfig, ErrorHandler } from '@angular/core';
import { provideRouter } from '@angular/router';
import { provideHttpClient } from '@angular/common/http';
import * as Sentry from '@sentry/angular';
import { routes } from './app.routes';
export const appConfig: ApplicationConfig = {
providers: [
provideRouter(routes),
provideHttpClient(),
// Route Angular's unhandled errors to Sentry instead of only console.error.
{ provide: ErrorHandler, useValue: Sentry.createErrorHandler({ showDialog: false }) },
],
};

After login, attach an opaque user id (never the name or email) so “users affected” is meaningful: Sentry.setUser({ id: hashedAdjusterId }). Check in staging which exception object your SDK version passes to beforeSend for HTTP failures, because Angular wraps some errors.

Upload source maps in CI, otherwise minified stack traces are unreadable:

Terminal window
npx sentry-cli sourcemaps inject ./dist/claims-web/browser
npx sentry-cli sourcemaps upload ./dist/claims-web/browser --release "claims-web@1.42.0"

6.3 Database angle (SQL Server / PostgreSQL)

Database exceptions are the most common repeat offenders: deadlocks (SQL Server 1205, PostgreSQL 40P01), timeouts, unique-key violations, and DbUpdateConcurrencyException from EF Core row-version checks. Treat them deliberately: a concurrency conflict on claim approval is an expected 409 that you handle and do not report; a deadlock is a reportable signal that you fingerprint into one issue and watch the count. Keep SendDefaultPii = false and use parameterised queries so SQL breadcrumbs do not carry personal data.

Level 3: Advanced

Performance and cost.

  • Error volume, not user count, drives price. One tight retry loop hitting a dead database can emit millions of events in an hour and burn a monthly quota. Use client-side rate limiting, per-issue spike protection (server side), and filters for known-noisy exceptions.
  • Do not capture the same exception at every layer (catch, log, rethrow at three levels). Capture once, at the boundary (the global handler). Duplicates inflate counts and cost.
  • Keep the SDK asynchronous. Sentry and App Insights both send in the background; make sure a shutdown flush is allowed (SentrySdk.FlushAsync in short-lived jobs and Azure Functions), otherwise the last events are lost when the process exits.

Security and privacy.

  • Stack traces, request bodies, and SQL text can carry personal or claim data. Default to no PII, scrub in BeforeSend, and never attach request bodies from claim endpoints.
  • The DSN is a write-only ingest key and is public in browser code; the auth token used for source-map upload is a secret and belongs in the pipeline secret store or Key Vault.
  • Define retention and access. Error tools are not compliant record stores; restrict who can see events from claims containing medical or financial detail.

Failure modes.

  • The tracker itself is down or blocked (ad blockers block browser SDK traffic). Never make the app depend on it, and keep logs as a second source.
  • Sampling that drops error events by accident. Sample traces aggressively, but do not sample errors unless volume forces it.
  • Missing release or environment tags, so production and staging errors mix and regressions cannot be detected.
  • Alert fatigue: alerting on every event. Alert on new issues, regressions, and spikes.

Common mistakes.

  1. Grouping by message text, so every claim id makes a new issue.
  2. Reporting expected business outcomes (404, 409, validation) as errors.
  3. catch (Exception) { } with no capture, which is worse than having no tracker.
  4. Forgetting source maps or PDBs, so traces are unreadable.
  5. No owner: issues pile up, and the inbox is ignored within weeks.

Level 4: Expert and Architect view

Options compared

OptionStrengthsWeaknessesChoose when
Sentry (SaaS)Best-in-class grouping, breadcrumbs, release tracking, regression detection, strong Angular and .NET SDKs, issue workflowSeparate vendor and data residency review; per-event pricing; not Azure-nativeYou want the best developer triage experience across .NET and Angular
Azure Application InsightsAzure-native, in your subscription and region, one tool for traces, metrics, and exceptions, KQL over everything, Azure RBAC and alertsGrouping and issue workflow are more basic; browser source-map handling is more manual; ingestion cost is per GBThe organisation is Azure-first, standardised on Azure Monitor, or has data-residency constraints
Self-hosted Sentry-compatible (Sentry self-hosted, GlitchTip)Data stays with you; Sentry SDKs work unchangedYou operate a multi-component stack (databases, queues) yourself; upgrades and scaling are your jobStrict data-control rules and a platform team able to run it
Logs only (Day 32) plus queriesNo new toolNo dedupe, no ownership, no regression detectionVery small systems, temporarily

Combines with

  • Distributed Tracing (Day 31): attach trace_id to each error so you jump from an issue to the full request path.
  • Log Aggregation (Day 32): the tracker points at the bug; logs give the surrounding narrative.
  • Application Metrics (Day 33): alert on error rate; use the tracker to find the cause.
  • Log Deployments & Changes (Day 37): release markers let you tie an issue’s first-seen time to a deployment.
  • Health Check API (Day 36), Circuit Breaker, Retry: exceptions from open circuits or exhausted retries are high-value issues.

ADR (suitable for an architecture review)

Title: ADR-035 Adopt centralised exception tracking for Claims Platform

Status: Proposed

Context: Claims API (.NET 10), Claims Web (Angular 22), and three supporting services log to a central store, but production errors are found by users and support, not engineering. Frontend errors are invisible. Post-release verification is manual.

Decision: Use one exception tracker for all services and the web client. Capture at service boundaries only, tag every event with release, environment, and trace_id, do not send PII, do not report expected 4xx outcomes, and alert on new issues, regressions, and spikes to the owning team’s channel. Tool choice: Sentry if developer triage is the priority and data-residency review passes; Application Insights if the organisation mandates Azure-native tooling.

Consequences: (+) Faster detection and triage, release health visible in minutes, frontend visibility. (-) An additional dependency and cost that must be budgeted by event volume, a data-privacy review, and the discipline of assigning issue owners. Rejected alternative: logs plus manual search, because it gives no deduplication or ownership.

Azure implementation

Services:

  • Azure Monitor Application Insights (workspace-based), backed by a Log Analytics workspace. This is the Azure-native exception tracker: server exceptions, browser exceptions, the Failures view, and Smart Detection (Failure Anomalies).
  • Azure Monitor alerts (log search alert rules and action groups) to page the team.
  • Azure Key Vault for secrets (source-map upload tokens, and DSN or connection string if you prefer not to place them in config).
  • Azure DevOps / GitHub Actions to set the release and upload source maps.
  • If using Sentry: it is SaaS outside Azure, so the app only needs outbound HTTPS. Self-hosting on AKS is possible but is a substantial workload to run.

Configuration (.NET, Application Insights via OpenTelemetry distro): package Azure.Monitor.OpenTelemetry.AspNetCore.

using Azure.Monitor.OpenTelemetry.AspNetCore;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddOpenTelemetry().UseAzureMonitor(); // reads APPLICATIONINSIGHTS_CONNECTION_STRING
var app = builder.Build();
app.Run();

Set APPLICATIONINSIGHTS_CONNECTION_STRING as an app setting (App Service, Container Apps, AKS secret). Exceptions logged through ILogger with an exception, and unhandled request exceptions, appear in the exceptions table.

Configuration (Angular): package @microsoft/applicationinsights-web; call appInsights.trackException({ exception: error }) from a custom ErrorHandler, the same place you would use Sentry’s handler.

Useful KQL (Application Insights Logs):

exceptions
| where timestamp > ago(24h)
| summarize events = count(), firstSeen = min(timestamp), lastSeen = max(timestamp)
by problemId, type, outerMessage, cloud_RoleName
| order by events desc

Create a log search alert on new problemId values (compare the last hour with the previous 7 days) and route it to an action group (Teams, email, on-call).

Pricing and tier considerations (confirm on the Azure Monitor and Sentry pricing pages, since both change):

  • Application Insights bills by data ingested through the Log Analytics workspace, pay-as-you-go at roughly 2.30 USD per GB in many regions, with a monthly free allowance (5 GB per billing account) and commitment tiers that lower the per-GB price at high volume. Exceptions are usually small compared with traces and dependencies, so sampling and filtering traces controls most of the cost.
  • Sentry plans: Developer (free, 5,000 errors per month, one user), Team (about 26 USD per month, 50,000 errors), Business (about 80 USD per month), and Enterprise (custom), with pay-as-you-go overage. Cost scales with event volume.
  • Retention is a separate cost lever in both tools; set it to your operational need, not the maximum.

Reference architecture (text): Angular 22 app served from Azure Static Web Apps or Front Door reports browser errors (with release and trace headers) to the tracker. It calls the Claims API on Azure Container Apps (or AKS), fronted by API Management. The API reports unhandled exceptions from its global handler, tagged with release, environment, and trace id. SQL Server on Azure SQL Database and PostgreSQL Flexible Server errors surface through the API’s exception path. Application Insights and its Log Analytics workspace store telemetry; alert rules fire to an Azure Monitor action group that posts to the owning team’s Teams channel. The CI/CD pipeline stamps the release, uploads source maps or symbols using a token from Key Vault, and records a deployment marker. Access to the workspace is restricted by Azure RBAC.

Teaching guide for my team

2-minute explanation for a beginner. “When our code crashes in production, we do not want to find out from a customer. An exception tracker catches every crash, adds details like who, where, and which version, and groups the same crash together so 5,000 crashes show up as one problem with a counter. It then tells the right person. Logs tell you what happened; the tracker tells you what is broken right now and whether it is new.”

5-minute explanation for an intermediate developer. Cover: capture at the boundary once; fingerprint by exception type plus in-app frame, not message; release and environment tags for regression detection; no PII and scrub in BeforeSend; filter expected outcomes (404, 409, cancellations); link the error to a trace id; source maps and PDBs; alert on new issue, regression, and spike, not on every event. Demo: throw the same exception with two different claim ids and show it grouped as one issue.

Hands-on exercise (45 minutes).

  1. Add Sentry (or Application Insights) to a sample Claims API and Angular app using the code in Level 2.
  2. Add an endpoint that throws InvalidOperationException($"Claim {id} already approved") and call it with five different ids.
  3. Add a button in Angular that reads claim.policy.number when policy is undefined.
  4. Add a ClaimNotFoundException path and confirm it returns 404.

Expected outcome: the tracker shows exactly one issue for the five API calls (grouped), one browser issue with a readable stack trace (source maps working), no issue for the 404 path, both events tagged with release and environment, and the API error carries a trace id that matches the traceId in the response body.

Interview-style questions.

  1. Why not just search the logs for “Exception”? Logs are unstructured occurrences with no grouping, ownership, release awareness, or regression detection; a tracker turns thousands of lines into a small set of actionable issues.
  2. How does an exception tracker decide two errors are the same? By a fingerprint, usually exception type plus the top in-app stack frames, not the message, because the message contains variable data. You can override it, for example grouping every SQL deadlock together.
  3. Which exceptions should you not report? Expected business outcomes such as not-found, validation, and optimistic-concurrency conflicts, and cancellations. Reporting them buries the real bugs and raises cost.

Mastery checklist

  • I can explain the difference between an event, an issue, and a fingerprint.
  • I can capture exceptions once at the boundary in ASP.NET Core with a global handler and return a ProblemDetails carrying a trace id.
  • I have wired the Angular ErrorHandler and verified source maps produce readable stack traces.
  • I can tag events with release and environment and use them to spot a regression after a deployment.
  • I filter expected outcomes and scrub personal and claim data before events leave the process.
  • I can write a custom fingerprint for a known noisy error such as a database deadlock.
  • I can query exceptions in Application Insights with KQL and build an alert on new problem ids.
  • I can defend the choice between Sentry and Application Insights in an ADR, including cost drivers.

Key takeaway

Logs tell you what happened; exception tracking tells you what is broken, whether it is new, how many people it hurts, and who owns it. Capture once at the boundary, group by root cause, keep personal data out, and alert only on new issues, regressions, and spikes.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 37: Log Deployments & Changes

Log Deployments & Changes means every release, configuration change, feature-flag flip, and infrastructure change is recorded as a timestamped event and drawn as a marker on the same dashboards where you watch errors...

Manikandan
Manikandan·24 min read
Microservices

Day 36: Health Check API

A Health Check API is a small set of HTTP endpoints that every service exposes so that machines (Kubernetes, Azure Container Apps, App Service, load balancers, monitoring) can ask "are you alive?", "are you ready to...

Manikandan
Manikandan·19 min read
Microservices

Day 34: Audit Logging

Audit Logging is the practice of writing a structured, tamper-resistant record of *who* did *what*, *to which thing*, *when*, *from where*, and *with what result* every time a security-relevant or business-relevant...

Manikandan
Manikandan·24 min read
Microservices

Day 33: Application Metrics

Application Metrics means every service continuously exposes small, cheap, numeric measurements (request latency, error counts, throughput, queue depth, memory, business counters like "claims submitted") to a...

Manikandan
Manikandan·18 min read