Manikandan — Manikandan
Microservices

Day 37: Log Deployments & Changes

ManikandanManikandan
24 min read·Updated Sep 15, 2022

Log Deployments & Changes means every release, configuration change, feature-flag flip, and infrastructure change is recorded as a timestamped event and drawn as a marker on the same dashboards where you watch errors...

Intro

Log Deployments & Changes means every release, configuration change, feature-flag flip, and infrastructure change is recorded as a timestamped event and drawn as a marker on the same dashboards where you watch errors and latency. When the error rate jumps at 14:07, the on-call engineer sees a “claims-api v2.14.3 deployed 14:05” line right there and starts with the most likely suspect, instead of spending forty minutes asking in chat “did anyone deploy anything?”. This lesson targets the current .NET LTS (.NET 10) and Angular 22 (the current major at the time of writing; Angular 21 is also still in security support), with Azure Monitor / Application Insights as the main implementation.


Why we need this

Most production incidents are caused by change. Deployments, config edits, certificate rotations, schema migrations, feature-flag toggles and scaling changes are the usual triggers; random hardware failure is the minority. If your observability tools show only symptoms (errors, latency, saturation) and not causes (changes), the first and most expensive question in every incident, “what changed?”, has to be answered by humans asking each other.

Business reasons:

  • Lower mean time to recovery (MTTR). If a marker shows the release directly before the spike, rollback is a five-minute decision instead of a one-hour investigation.
  • Blameless, factual post-mortems. The timeline is data, not memory.
  • Change-failure-rate measurement. DORA’s “change failure rate” and “time to restore” metrics need a reliable list of deployments and which ones were followed by incidents or rollbacks.
  • Audit and compliance. Regulated domains such as insurance need to show who released what, when, and with which approval. (This overlaps with Audit Logging from Day 34, but the audience differs: audit logging records user actions inside the business system; this pattern records engineering changes to the system itself.)

Technical reasons: in a microservices estate with 30 services deployed independently and several times a day, no one person knows what is running. The number of possible “recent changes” grows with the number of services, so it must be automated.

What problem it solves

Problem (from the topic list): teams cannot tell whether an incident was caused by a release.

Without the pattern, this is what happens in our insurance claims system:

  1. At 14:07 the claims-api p95 latency goes from 180 ms to 2.4 s and 5xx errors reach 8%.
  2. The on-call engineer opens the dashboard and sees the graph go bad, but nothing on it says why.
  3. She asks in the team chat: “anyone deploy anything?” Three services are owned by three teams in two time zones. Someone answers 20 minutes later: “Payments team released at 14:05, but that shouldn’t affect you.”
  4. Meanwhile, a fourth change is invisible to everyone: an Azure App Configuration value for the fraud-scoring timeout was edited at 14:02 by an engineer who did not consider it a “deployment”.
  5. The team spends 45 minutes looking at database locks and CPU before finding the config edit.

What goes wrong: slow diagnosis, wrong hypotheses, rollbacks of the wrong thing, and no data to improve on next time. With markers, the graph itself shows a vertical line at 14:02 labelled “App Config: FraudScoring:TimeoutMs 800 -> 8000 by t.mani” and one at 14:05 for the Payments release, and the investigation starts at the right place.

When it is needed (and when it is NOT)

Needed when:

  • More than one team or more than a handful of services release independently.
  • You deploy more often than about once a week (frequent change means frequent suspicion).
  • You have an on-call rotation and incident reviews.
  • You run production configuration, feature flags, or infrastructure-as-code that people can change outside the normal release pipeline.
  • You need DORA metrics or change-audit evidence for compliance.

Not needed, or overkill, when:

  • A single small internal tool is deployed by one person a few times a year; a line in a changelog is enough.
  • You are building a full “change management” product with approval workflows just to get markers. The pattern is about recording and displaying events; ticketing and CAB approvals are a separate concern.
  • You would only log deployments but skip config and flag changes. That gives false confidence: “no deploy happened, so it isn’t a change”. Cover all change types or clearly state what is covered.

Wrong choice: using this as a substitute for observability of the system itself. Markers explain when behaviour changed; you still need traces, metrics and logs (Days 31 to 33) to understand why.

How to identify the problem (key signals)

  1. “Did anyone deploy?” appears in the incident channel in the first ten minutes of most incidents.
  2. Post-mortem timelines are reconstructed from Slack messages and pipeline screenshots, and have gaps of tens of minutes.
  3. Dashboards show a step change with no explanation. Latency or error graphs have a clean vertical step, typical of a release or config change rather than organic load.
  4. Nobody can say which version is running in production for a given service right now, or two instances run different versions during a rollout and nobody notices.
  5. Rollbacks are done “just in case” because the team cannot tell whether the latest release is involved.
  6. Config, secret, or flag edits are made in the portal and leave no trace outside the platform’s own hidden logs.
  7. DORA metrics are guessed. Nobody can produce the number of deployments last month or the share that caused an incident.

Flow Diagram

Every change becomes a timestamped event drawn on the same charts as errors and latency.

flowchart LR
GH["Pipeline deploys claims-api v2.14.3"] --> ANN["App Insights release annotation"]
GH --> CLOG[("Change log API / table")]
APPC["App Configuration edit"] -- "Event Grid" --> FN["Azure Function"] --> CLOG
ACT["Azure Activity Log"] --> LA[("Log Analytics")]
CLOG --> WB["Operations workbook"]
ANN --> WB
LA --> WB
TEL["Telemetry with service.version"] --> WB
WB --> ONC["On-call: what changed?"]

Level 1: Beginner

Core concept and analogy

Think of a hospital patient’s chart. The heart-rate line is on the monitor, and the nurse writes on the same chart “14:05 - new medication given”. If the heart rate changes at 14:10, everybody looks at the medication note first. A deployment marker is the note on the chart.

The minimum viable version has three parts:

  1. Stamp the version into the running app (so the app itself can say what it is).
  2. Record an event when something changes (who, what, when, where).
  3. Show it next to your graphs.

Minimal working example (ASP.NET Core, .NET 10)

The service exposes its version, which is set at build time, so anyone (and any dashboard) can ask “what is running?”.

// Program.cs - Claims.Api (net10.0)
using System.Reflection;
var builder = WebApplication.CreateBuilder(args);
var app = builder.Build();
// The informational version is set in CI, e.g. dotnet publish -p:Version=2.14.3 -p:SourceRevisionId=$(git rev-parse --short HEAD)
var info = Assembly.GetExecutingAssembly()
.GetCustomAttribute<AssemblyInformationalVersionAttribute>()?.InformationalVersion ?? "unknown";
app.MapGet("/api/version", () => Results.Ok(new
{
service = "claims-api",
version = info,
environment = app.Environment.EnvironmentName
}));
app.Run();

In the .csproj, IncludeSourceRevisionInInformationalVersion (on by default in the modern SDK) appends the commit hash to the informational version, giving values like 2.14.3+a1b2c3d.

Beginner “poor man’s” marker

A CI step that writes one row per deployment. It is not fancy, but it already answers “what changed?”.

-- SQL Server (T-SQL)
CREATE TABLE dbo.DeploymentLog
(
Id BIGINT IDENTITY PRIMARY KEY,
ServiceName NVARCHAR(100) NOT NULL,
Version NVARCHAR(50) NOT NULL,
Environment NVARCHAR(30) NOT NULL,
ChangeKind NVARCHAR(30) NOT NULL, -- Deployment | Config | FeatureFlag | Infra | Rollback
Summary NVARCHAR(400) NOT NULL,
ChangedBy NVARCHAR(100) NOT NULL,
GitCommit NVARCHAR(40) NULL,
OccurredAtUtc DATETIME2(0) NOT NULL DEFAULT SYSUTCDATETIME()
);
CREATE INDEX IX_DeploymentLog_Env_Time ON dbo.DeploymentLog (Environment, OccurredAtUtc DESC);

Level 2: Intermediate

In a real .NET + Angular + database application, you want four things: (a) a consistent version identity flowing into all telemetry, (b) an automatic marker on the monitoring platform at every release, (c) a queryable change log that also covers config and flags, and (d) the version visible in the UI for support staff.

6.1 Version identity in all telemetry (OpenTelemetry + Azure Monitor)

If every span, log and metric carries service.name and service.version, you can split any graph by version and see the new release’s behaviour next to the old one during a rollout.

Azure.Monitor.OpenTelemetry.AspNetCore
using Azure.Monitor.OpenTelemetry.AspNetCore;
using OpenTelemetry.Resources;
var builder = WebApplication.CreateBuilder(args);
var version = typeof(Program).Assembly
.GetCustomAttributes(typeof(System.Reflection.AssemblyInformationalVersionAttribute), false)
.OfType<System.Reflection.AssemblyInformationalVersionAttribute>()
.FirstOrDefault()?.InformationalVersion ?? "unknown";
builder.Services.AddOpenTelemetry()
.UseAzureMonitor() // reads APPLICATIONINSIGHTS_CONNECTION_STRING
.ConfigureResource(r => r
.AddService(serviceName: "claims-api", serviceVersion: version)
.AddAttributes(new Dictionary<string, object>
{
["deployment.environment.name"] = builder.Environment.EnvironmentName
}));
builder.Services.AddControllers();
var app = builder.Build();
app.MapControllers();
app.Run();

Note: OpenTelemetry semantic conventions renamed deployment.environment to deployment.environment.name in recent releases. If your collector or dashboards still expect the old key, emit both during migration.

In Application Insights, service.version shows up as application_Version on request, dependency and trace records (verify in your own workspace, since mapping depends on the exporter version). A query then splits errors by version:

// Failed requests per 5 minutes, split by running version
requests
| where timestamp > ago(3h) and cloud_RoleName == "claims-api"
| summarize failed = countif(success == false), total = count() by application_Version, bin(timestamp, 5m)
| extend failureRate = todouble(failed) / total
| render timechart

6.2 Automatic marker on release (Application Insights annotation)

Application Insights supports release annotations: vertical markers shown on Performance, Failures, Usage and Workbooks charts (not the Metrics pane). Key facts to remember:

  • The old Application Insights annotation deployment task for Azure DevOps is deprecated; Microsoft says to delete it if you still use it.
  • Annotations are created automatically by the current Azure Pipelines tasks for App Service and Azure Functions (for example AzureWebApp, AzureRmWebAppDeployment V3+, AzureFunctionApp, AzureWebAppContainer) when the target is linked to an Application Insights resource in the same subscription through APPLICATIONINSIGHTS_CONNECTION_STRING.
  • For everything else (AKS, Container Apps, GitHub Actions, scripts), create the annotation yourself through the REST API or the CreateReleaseAnnotation.ps1 script. The annotation’s Category must be Deployment or the portal will not show it.

A small .NET tool that a pipeline can run after a successful deploy:

Azure.Identity
using System.Net.Http.Headers;
using System.Text;
using System.Text.Json;
using Azure.Core;
using Azure.Identity;
// args: <appInsightsResourceId> <releaseName> <version> <commit> <deployedBy>
var (aiResourceId, releaseName, version, commit, by) = (args[0], args[1], args[2], args[3], args[4]);
var credential = new DefaultAzureCredential(); // pipeline identity via workload identity / OIDC
var token = await credential.GetTokenAsync(
new TokenRequestContext(new[] { "https://management.azure.com/.default" }));
var annotation = new
{
Id = Guid.NewGuid().ToString(),
AnnotationName = releaseName,
EventTime = DateTime.UtcNow.ToString("yyyy-MM-ddTHH:mm:ss.fffZ"),
Category = "Deployment",
// Properties is a JSON *string*, not a nested object
Properties = JsonSerializer.Serialize(new
{
Version = version,
Commit = commit,
TriggerBy = by,
Environment = "Production"
})
};
using var http = new HttpClient();
http.DefaultRequestHeaders.Authorization = new AuthenticationHeaderValue("Bearer", token.Token);
var url = $"https://management.azure.com{aiResourceId}/Annotations?api-version=2015-05-01";
var response = await http.PutAsync(url,
new StringContent(JsonSerializer.Serialize(annotation), Encoding.UTF8, "application/json"));
response.EnsureSuccessStatusCode();
Console.WriteLine($"Marker created for {releaseName}");

The pipeline identity needs permission to write to the Application Insights component (for example Contributor on that resource, or a narrower custom role). Test this once in a non-production resource before rolling out.

6.3 A change log that also covers config and feature flags

Deployment annotations do not cover “someone edited a setting”. Add a small change-log API that any automation can call, and store it in the database. Example in PostgreSQL for a shared platform database:

-- PostgreSQL
CREATE TABLE change_log (
id BIGINT GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
service_name TEXT NOT NULL,
environment TEXT NOT NULL,
change_kind TEXT NOT NULL CHECK (change_kind IN ('deployment','config','feature_flag','infra','rollback','migration')),
version TEXT,
summary TEXT NOT NULL,
changed_by TEXT NOT NULL,
git_commit TEXT,
correlation_id TEXT,
occurred_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX ix_change_log_env_time ON change_log (environment, occurred_at DESC);
// Claims.Platform.ChangeLog - minimal endpoint (EF Core 10 + Npgsql)
public sealed record ChangeEvent(
string ServiceName, string Environment, string ChangeKind,
string? Version, string Summary, string ChangedBy, string? GitCommit);
app.MapPost("/api/changes", async (ChangeEvent e, ChangeDb db, CancellationToken ct) =>
{
db.Changes.Add(new ChangeRow
{
ServiceName = e.ServiceName, Environment = e.Environment, ChangeKind = e.ChangeKind,
Version = e.Version, Summary = e.Summary, ChangedBy = e.ChangedBy,
GitCommit = e.GitCommit, OccurredAt = DateTimeOffset.UtcNow
});
await db.SaveChangesAsync(ct);
return Results.Accepted();
}).RequireAuthorization("ChangeWriter"); // pipelines only, via workload identity
app.MapGet("/api/changes", async (string env, int hours, ChangeDb db, CancellationToken ct) =>
await db.Changes
.Where(c => c.Environment == env && c.OccurredAt > DateTimeOffset.UtcNow.AddHours(-hours))
.OrderByDescending(c => c.OccurredAt)
.Take(200)
.ToListAsync(ct));

Call it from a GitHub Actions or Azure Pipelines step, and also from anything that changes configuration through automation (for example the script that updates Azure App Configuration keys or toggles a feature flag).

6.4 Angular: show the running version and recent changes

Support staff should be able to read the UI version without opening a terminal. This standalone component (Angular 22, signals) shows the API version and the last few changes.

version-badge.component.ts
import { ChangeDetectionStrategy, Component, inject, signal } from '@angular/core';
import { HttpClient } from '@angular/common/http';
interface VersionInfo { service: string; version: string; environment: string; }
@Component({
selector: 'app-version-badge',
changeDetection: ChangeDetectionStrategy.OnPush,
template: `
@if (info(); as i) {
<small class="version-badge" [title]="i.service + ' ' + i.environment">
API {{ i.version }}
</small>
}
`,
})
export class VersionBadgeComponent {
private readonly http = inject(HttpClient);
protected readonly info = signal<VersionInfo | null>(null);
constructor() {
this.http.get<VersionInfo>('/api/version').subscribe({
next: v => this.info.set(v),
error: () => this.info.set(null), // never break the page for a footer badge
});
}
}

Make sure the SPA build version is also stamped (for example through the build pipeline writing a version.json file into the dist folder). Otherwise you know the API version but not the front-end version, which is a classic gap.

Level 3: Advanced

Performance and scalability

  • Markers are low-volume events. Even at 200 deployments a day the volume is trivial. The real scalability issue is cardinality of labels if you push versions into metrics. Adding a version label to a Prometheus/OTel metric multiplies series by the number of versions alive; keep only a bounded set (current plus previous) or use a dedicated “build info” gauge rather than putting version on every metric.
  • Keep marker calls off the critical path of deployment. A monitoring outage must never block a release. Mark the step continueOnError, retry briefly, and alert if markers repeatedly fail.
  • Query cost. Joining a change table with telemetry on wide time windows is expensive. Query a bounded window (for example +/- 2 hours around the incident).

Security

  • Only pipeline identities (workload identity federation or managed identity, not personal tokens or long-lived secrets) can write to the change API and to Application Insights annotations.
  • Markers are integrity-sensitive: an attacker who can forge or delete them can hide a malicious change. Store the authoritative record in an append-only place (see Day 34) and treat dashboard markers as a convenience view.
  • Do not put secrets, connection strings, or personal data in marker properties. A config-change event should say “FraudScoring:TimeoutMs changed”, and for secrets only “secret X rotated”, never the value.
  • Restrict who can read the change API; commit messages sometimes contain ticket data.

Failure modes

  • Clock skew and time zones. Use UTC everywhere; use the time of deployment completion, not pipeline start, or note both (started/finished) for long rollouts.
  • Staged rollouts. With canary or blue/green deployments one “deployment” is really a sequence: 5% at 14:05, 50% at 14:20, 100% at 14:35. Record each stage, otherwise the marker at 14:05 will not explain the problem at 14:22.
  • Out-of-band changes. Portal edits, kubectl edit, manual SQL. Cover them by collecting platform-native change records (Azure Activity Log, Azure Resource Graph change data, Kubernetes events) instead of relying on people calling your API.
  • Marker noise. If every trivial change shows up, people ignore them. Tag by kind and severity and let dashboards filter.
  • Wrong environment. A marker created against the staging Application Insights resource never shows on production charts. Put the resource ID in one variable per environment and validate it.
  • Rollback not recorded. A rollback is a change too. Record it with ChangeKind = Rollback and a reference to the version it replaced.

Common mistakes

  1. Only marking the “successful deploy” event and forgetting migrations, config and flags.
  2. Using a version like 1.0.0 for every build, so markers carry no useful identity. Use SemVer plus commit hash.
  3. Building the annotation with the deprecated classic task and never noticing it stopped working.
  4. Putting the version only in the pipeline logs, not in the running process, so you cannot verify what actually runs.
  5. Treating the marker as proof of causation. It is a lead, not a verdict; always confirm with telemetry.

Level 4: Expert and Architect view

Design trade-offs and alternatives

ApproachWhat it givesStrengthsWeaknessesBest for
App Insights release annotationsVertical markers on Performance/Failures/Usage/WorkbooksNative, zero infrastructure, shows in the tools engineers already useDeployments mainly; not shown on Metrics pane; limited to App Insights resource scopeTeams already on Azure Monitor
Custom change-log table + workbook/dashboardAny change kind, queryable historyFull control, covers config/flags/infra, drives DORA reportsYou build and maintain itPlatform teams with multiple change sources
Grafana annotations (Azure Managed Grafana or self-hosted)Markers over any data sourceOne place for metrics, logs, traces; API-driven annotationsAnother system to operate; annotation quality depends on your automationMulti-cloud or Prometheus-heavy estates
Platform-native change records (Azure Activity Log, Resource Graph change data, Kubernetes events)Captures out-of-band changes automaticallyIncludes portal/CLI changes nobody reportedInfra-level detail only, no business meaning; retention limitsDetecting “shadow” changes
Deployment events in the CI/CD system (GitHub Deployments API, Azure Pipelines environments)Authoritative deployment historyAlready exists; approvals attachedNot visible on runtime dashboards without integrationDORA metrics, audit
Version as telemetry attribute (service.version)Per-version comparison in any queryCheap, works everywhere, enables canary analysisShows what ran, not who changed configAlways, as a baseline

Patterns it combines with

  • Distributed Tracing, Application Metrics, Log Aggregation (Days 31 to 33): markers give the “when”, telemetry gives the “why”.
  • Audit Logging (Day 34): authoritative, tamper-resistant record of who did what; deployment markers can be derived from it.
  • Health Check API (Day 36): during rollouts, readiness results explain why a stage paused.
  • Externalized Configuration (Day 39): every configuration change should emit a change event.
  • Service Template / Chassis (Days 40, 41): bake version stamping and the marker step into the template so no team can forget it.
  • Strangler Fig (Day 49), canary and blue/green releases: stage-level markers are essential for judging a gradual cutover.

ADR-style justification

Title: ADR-037 Record all production changes as events and display them on operational dashboards

Status: Proposed

Context: The claims platform has 14 services owned by 5 teams, releasing independently up to 20 times per day. In the last quarter, 9 of 12 Sev-2 incidents began within 30 minutes of a change, but median time to identify the responsible change was 38 minutes because deployment, configuration and feature-flag changes were tracked in different places. Post-mortem timelines are compiled manually.

Decision: (1) All services publish service.name, service.version and deployment.environment.name in telemetry through OpenTelemetry using the Azure Monitor distro. (2) Every production deployment creates an Application Insights release annotation from the pipeline using workload identity. (3) A central change-log API records deployments, configuration edits, feature-flag changes, migrations and rollbacks; pipelines and config tooling call it. (4) Azure Activity Log and Resource Graph change data are queried as a second source for out-of-band infrastructure changes. (5) The Azure Monitor Workbook “Operations Overview” overlays change events on latency and error charts.

Consequences: Positive: faster diagnosis, data for DORA metrics, consistent post-mortems. Negative: pipeline steps and a small service to maintain; teams must route config changes through automation or accept that only the platform-native record will exist. Risk: marker failure must not block deployments (mitigated by non-blocking steps and an alert on missing markers).

Alternatives considered: annotations only (rejected: no config/flag coverage); Grafana-only (deferred: adds a new platform while the team standardizes on Azure Monitor); manual change calendar (rejected: not reliable).

Azure implementation

Services that implement or support this topic

  • Application Insights (workspace-based) with Azure Monitor OpenTelemetry Distro: telemetry with service.version, plus release annotations.
  • Azure Monitor Workbooks: dashboards that show annotations and can also plot change events from your own table or Azure Resource Graph.
  • Log Analytics workspace: stores telemetry, AzureActivity (control-plane operations) and optionally your custom change table.
  • Azure Activity Log: records control-plane changes (resource updates, deployments, role assignments) automatically. Route it to Log Analytics through a diagnostic setting for long retention and KQL.
  • Azure Resource Graph change data: records property-level changes to Azure resources (for example an App Service setting or a Container App revision setting). Retention is limited; check the current documented period before depending on it for audits.
  • Azure Pipelines / GitHub Actions: the source of the deployment event. Both can call the annotation REST API and your change-log API.
  • Azure App Configuration (feature flags and settings): supports Event Grid events on key-value changes, which you can forward to the change-log API. This is how config and flag edits become events without depending on people to report them.
  • Azure Container Apps / AKS: revisions and Kubernetes events already carry deployment history; include revision name or image tag as service.version.
  • Azure Managed Grafana (optional): annotation API for teams standardized on Grafana.
  • Azure Alerts: an alert rule can fire when error rate rises within N minutes after a change event, which turns markers into automated “suspect release” warnings.

How to configure

  1. Create a workspace-based Application Insights resource per environment (dev, test, prod), all linked to the right Log Analytics workspace.
  2. Set APPLICATIONINSIGHTS_CONNECTION_STRING on each service (App Service setting, Container Apps environment variable/secret reference, or Kubernetes secret). Use Key Vault references, not literal strings, in shared templates.
  3. Add the Azure Monitor OpenTelemetry distro and set service.name, service.version and environment as in Level 2.
  4. In the pipeline: for App Service or Azure Functions, use the current deploy tasks and confirm the annotation appears automatically. For Container Apps or AKS, add the ReleaseMarker step after a successful rollout.
  5. Give the pipeline’s federated identity the minimum role required to write annotations on that Application Insights resource.
  6. Create an Event Grid subscription on the App Configuration store (key-value modified/deleted events) targeting an Azure Function that calls /api/changes.
  7. Enable a diagnostic setting sending the Activity Log to Log Analytics.
  8. Build a Workbook with: latency/error charts by application_Version, a table of the last 24 hours of changes from your change_log, and an AzureActivity panel filtered to the resource group.

Example KQL for the Activity Log panel:

AzureActivity
| where TimeGenerated > ago(24h)
| where ResourceGroup =~ "rg-claims-prod"
| where CategoryValue == "Administrative" and ActivityStatusValue == "Success"
| where OperationNameValue has_any ("Microsoft.App/containerApps/write", "Microsoft.Web/sites/config/write", "Microsoft.AppConfiguration/configurationStores/write")
| project TimeGenerated, Caller, OperationNameValue, Resource
| order by TimeGenerated desc

Pricing and tier considerations

  • The annotation and Workbook features have no separate charge; cost comes from telemetry ingestion and retention in the Log Analytics workspace (per GB). A monthly free ingestion allowance exists; check the current Azure Monitor pricing page for the exact figure and regional prices before budgeting.
  • Activity Log collection into a workspace via diagnostic setting is billed as log ingestion/retention; volume is usually small compared with application telemetry.
  • Event Grid and Azure Functions (Consumption or Flex Consumption) costs for the config-change forwarder are negligible at this volume.
  • Azure Managed Grafana is billed per instance and per active user depending on tier; verify current tiers (for example Essential vs Standard) and their SLA and feature differences before choosing.
  • To save cost, do not add version to high-cardinality custom metrics; use the resource attribute and one build-info metric instead.

Reference architecture (text)

  1. A developer merges to main. GitHub Actions builds the claims-api image tagged 2.14.3-a1b2c3d, pushes to Azure Container Registry, and deploys a new revision to Azure Container Apps using OIDC federation to Entra ID.
  2. After the revision is healthy (readiness probe from Day 36), the pipeline runs the ReleaseMarker tool: it creates an Application Insights annotation (Category Deployment) and POSTs a deployment event to the change-log API.
  3. The running container reports telemetry through the Azure Monitor OpenTelemetry distro with service.version = 2.14.3+a1b2c3d.
  4. An operator edits a value in Azure App Configuration. Event Grid triggers an Azure Function that posts a config event to the change-log API (key name only, value redacted for secrets).
  5. The Activity Log diagnostic setting streams control-plane operations into Log Analytics for out-of-band changes.
  6. The on-call engineer opens the “Operations Overview” Workbook: latency and error charts split by version, with annotations as vertical lines, plus a table of all change events for the last 24 hours.
  7. An Azure Monitor alert rule compares error rate in the 15 minutes after any change event with the previous hour and notifies the owning team’s channel with the change details.

Teaching guide for my team

Explain to a beginner in 2 minutes

“When something breaks in production, the first question is always ‘what changed?’. Right now we answer that by asking around. This pattern makes the system answer for us. Every time we deploy, change a setting, or flip a feature flag, we write a small note with the time, what changed, and who did it. Then we draw that note as a vertical line on our dashboards. When the graph goes wrong, we look at the line just before it. Also, every service tells us its own version, so we always know what is actually running. Most problems come from the last thing that changed, so this saves us a lot of guessing.”

Explain to an intermediate developer in 5 minutes

Cover these points in order:

  1. Identity: stamp service.name and service.version (SemVer plus commit) into the assembly at build time and into all telemetry via OpenTelemetry resource attributes; show the API and SPA version in the UI.
  2. Events: every change type (deployment, config, flag, migration, infra, rollback) is an event with time in UTC, actor, target, version, and a correlation id to the pipeline run.
  3. Display: use Application Insights release annotations for deployments (automatic for App Service and Functions tasks, REST API or script for Container Apps/AKS/GitHub Actions) and a Workbook that also lists events from a change table.
  4. Coverage: capture out-of-band changes through Activity Log, Resource Graph changes and App Configuration events.
  5. Caveats: markers are leads, not proof; they must never block a release; never include secret values; record each stage of a canary rollout; treat integrity of the record seriously.

Hands-on exercise

Goal: make a deployment visible and prove it helps find a bad release.

  1. Take the sample claims-api (Level 1). Add the /api/version endpoint and OpenTelemetry with service.version from the assembly (Level 2, 6.1). Deploy version 1.0.0 to a dev App Service or Container App linked to a dev Application Insights resource.
  2. Generate steady traffic against GET /api/claims/{id} for 5 minutes (a simple while loop with curl, or a load-testing tool).
  3. Build and deploy 1.1.0 containing a deliberate regression: add await Task.Delay(1500) to the claims lookup and return HTTP 500 for 10% of requests.
  4. Run the ReleaseMarker tool (or use the automatic annotation if using App Service tasks) at the moment of deployment. Also POST a config event for a fake setting change 2 minutes later.
  5. Run the KQL query from 6.1 to split failures by application_Version, and open the Failures blade to view the annotation.

Expected outcome: the failure rate and latency step up right at the annotation; the query shows the 500s belong only to version 1.1.0; the config event is visible in the change list but, by comparing timestamps, the learner concludes the release (not the config edit) is the cause. Roll back to 1.0.0, record a rollback event, and confirm the metrics recover.

Three interview-style questions

  1. Why is a deployment annotation not enough for change tracking? Because many incidents come from non-deployment changes such as configuration edits, feature flags, migrations, secret rotations and portal changes. You need to record every change kind and also collect platform-native records for out-of-band edits.

  2. How would you avoid putting a version label on every metric, and why? Version labels multiply time-series by the number of live versions and can blow up cardinality and cost. Use a resource attribute (service.version) and one build-info metric, and split queries on the trace and log data where version is already attached.

  3. A canary rollout goes 5% then 50% then 100% over an hour. How do you record it? Emit an event per stage (with percentage and revision) plus start and finish times, so the marker precedes the exact point where behaviour changed. Compare the canary’s service.version against the stable version to judge it.

Mastery checklist

  • I can explain the difference between a deployment marker, an audit log entry, and a trace, and when each is used.
  • Every service I own reports service.name, service.version and environment in telemetry, and the version includes the commit.
  • A production deployment of my service creates a visible marker without anyone doing it manually, and I have verified it appears on the correct environment’s charts.
  • Config changes, feature-flag changes, migrations and rollbacks are recorded as change events, not just deployments.
  • I can write a KQL query that compares error rate and latency across two versions and correlates a spike with a change event.
  • I know which markers must not block a release, and what secrets must never appear in event properties.
  • I can describe how out-of-band changes are detected (Activity Log, Resource Graph change data, App Configuration events) and their retention limits.
  • I can run a mock incident where the team uses markers to identify a bad release and record the rollback.

Key takeaway

Most incidents are caused by change, so record every change (deployments, config, flags, infrastructure) as a timestamped event and draw it on the same dashboards as your errors and latency. Then “what changed?” becomes a glance at the graph instead of a chat thread.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 36: Health Check API

A Health Check API is a small set of HTTP endpoints that every service exposes so that machines (Kubernetes, Azure Container Apps, App Service, load balancers, monitoring) can ask "are you alive?", "are you ready to...

Manikandan
Manikandan·19 min read
Microservices

Day 35: Exception Tracking

Exception Tracking means capturing every unhandled (and important handled) error from your services and your browser app, attaching context to it (release, user, request, breadcrumbs), grouping identical errors into...

Manikandan
Manikandan·18 min read
Microservices

Day 34: Audit Logging

Audit Logging is the practice of writing a structured, tamper-resistant record of *who* did *what*, *to which thing*, *when*, *from where*, and *with what result* every time a security-relevant or business-relevant...

Manikandan
Manikandan·24 min read
Microservices

Day 33: Application Metrics

Application Metrics means every service continuously exposes small, cheap, numeric measurements (request latency, error counts, throughput, queue depth, memory, business counters like "claims submitted") to a...

Manikandan
Manikandan·18 min read