Application Metrics means every service continuously exposes small, cheap, numeric measurements (request latency, error counts, throughput, queue depth, memory, business counters like "claims submitted") to a...
Intro
Application Metrics means every service continuously exposes small, cheap, numeric measurements (request latency, error counts, throughput, queue depth, memory, business counters like “claims submitted”) to a collector such as Prometheus or Azure Monitor. Because numbers are aggregated in the process and only a few data points leave the service, you can store months of history, draw dashboards, and fire alerts long before a customer calls to say the system is down. In our insurance claims system, metrics tell you that the Claims API p95 latency doubled at 10:05 while logs and traces tell you why.
Why we need this
- Microservices multiply failure points. One claim submission touches the Angular UI, API gateway, Claims API, Policy service, Fraud service, a message broker and two databases. Without numbers per hop, nobody knows which hop is slow.
- Logs are too expensive and too slow to answer “how many, how fast, how often”. Counting 500 errors by grepping millions of log lines is slow and costly. A counter increment is nanoseconds and a time series is a few bytes per scrape.
- Alerting needs a signal that is cheap, continuous, and aggregatable. Alerts such as “error rate above 2% for 5 minutes” or “queue depth rising for 10 minutes” are metric queries.
- Capacity planning and autoscaling need numbers. Kubernetes HPA, KEDA and Azure App Service autoscale all scale on metrics (CPU, requests per second, queue length).
- Business and engineering share one truth. Claims submitted per minute and claims stuck in “Pending Fraud Check” are as important as CPU.
- SLOs are defined on metrics. “99.5% of claim submissions succeed within 2 seconds” can only be measured with a latency histogram plus a success counter.
What problem it solves
Problem (from the topic list): bottlenecks, error spikes, and memory leaks stay invisible until an outage.
What goes wrong without it:
- A memory leak in the Claims API grows 20 MB per hour. Nobody sees it. Every Friday at 3 a.m. the container is OOM-killed and Kubernetes restarts it. Customers see intermittent 502 errors and the team blames the network.
- Payment retries slowly increase database load for two weeks. The first sign is a full outage on the busiest day.
- A deployment adds an N+1 query. The average latency looks fine, but the p99 goes from 400 ms to 6 s. Support tickets are the only signal.
- Autoscaling is configured by guessing (“scale at 70% CPU”) because nobody has ever measured what the real bottleneck (thread pool, DB connections, queue depth) is.
When it is needed (and when it is NOT)
Needed
- Any service in production with more than one instance or more than one dependency.
- Any service with an SLA/SLO, or which is on a revenue or regulatory path (claims, payments, policy issue).
- Any asynchronous flow (queues, sagas, outbox publishers), where “is it keeping up?” is only answerable with queue-depth and processing-lag metrics.
- Anything that autoscales.
Not needed / overkill
- A throwaway prototype or a one-off batch script: a few log lines and the platform’s built-in CPU/memory graphs are enough.
- Instrumenting every method with its own histogram. Metrics should describe service behaviour at boundaries (incoming request, outgoing call, queue, key business event), not every function.
- Using metrics to debug one specific request. That is the job of distributed tracing (Day 31) and logs (Day 32). Metrics tell you that and where; traces and logs tell you why.
- High-cardinality data (claim ID, customer ID, raw URL) as metric labels. That belongs in traces or logs.
How to identify the problem (key signals)
- “Is it slow for everyone or just me?” cannot be answered. Nobody can state the current p95 latency of the Claims API.
- You learn about outages from customers or the call centre, not from an alert.
- Recurring restarts (
OOMKilled, container restart count climbing, App Service recycles) with no memory or GC graph to explain them. - Averages hide pain. Dashboards, if any, show average latency only. Users complain while the “average is fine”.
- Scaling by superstition. “We added two more instances and it did not help,” because the real bottleneck (DB connection pool, thread pool starvation, downstream rate limit) was never measured.
- Queue backlog discovered late. Messages in the Service Bus queue are found to be hours old during month-end.
- Post-incident reviews contain the sentence “we had no data for that.”
- Every team invents different metric names (
reqs,RequestCount,http_requests) so cross-service dashboards are impossible.
Flow Diagram
Services emit RED metrics and business counters; the backend drives dashboards, alerts and autoscaling.
flowchart LR SVC["Claims API"] --> M1["Counter: claims.submitted"] SVC --> M2["Histogram: processing.duration"] SVC --> M3["Gauge: workers.active"] M1 --> OT["OpenTelemetry SDK"] M2 --> OT M3 --> OT OT --> PROM[("Managed Prometheus / App Insights")] PROM --> GR["Grafana dashboards"] PROM --> ALR["SLO burn-rate alerts"] PROM --> HPA["Autoscaling"]Level 1: Beginner
Core concept and analogy
Think of a car dashboard. You do not open the engine to know if you are in trouble: the speedometer (rate), fuel gauge (saturation), warning lights (errors) and trip counter (totals) tell you continuously. Application metrics are the dashboard of your service.
Three instrument types cover most needs:
| Instrument | Meaning | Claims example |
|---|---|---|
| Counter | Only goes up (resets on restart) | claims.submitted, claims.failed |
| Gauge | A value that goes up and down, sampled | claims.queue.depth, memory in use |
| Histogram | Distribution of values in buckets | claims.processing.duration (gives p50/p95/p99) |
Two quick rules of thumb: RED for services (Rate, Errors, Duration) and USE for resources (Utilisation, Saturation, Errors).
Minimal working example (.NET, compiles as a console app)
System.Diagnostics.Metrics is built into .NET. No NuGet package is needed to create and observe metrics.
using System.Diagnostics.Metrics;
var meter = new Meter("Contoso.Claims");var submitted = meter.CreateCounter<long>("claims.submitted", unit: "{claim}", description: "Number of claims submitted");
// A listener plays the role a collector (OpenTelemetry, Prometheus) plays in production.using var listener = new MeterListener();listener.InstrumentPublished = (instrument, l) =>{ if (instrument.Meter.Name == "Contoso.Claims") l.EnableMeasurementEvents(instrument);};listener.SetMeasurementEventCallback<long>((instrument, value, tags, state) => Console.WriteLine($"{instrument.Name} += {value}"));listener.Start();
submitted.Add(1);submitted.Add(1);Expected output: two lines claims.submitted += 1. In a running service you can also watch live values with the dotnet-counters tool: dotnet-counters monitor --name ClaimsApi Contoso.Claims.
Level 2: Intermediate
6.1 A metrics class for the Claims API (ASP.NET Core, current .NET LTS = .NET 10)
Use IMeterFactory (available since .NET 8). It ties the meter lifetime to dependency injection and keeps unit tests isolated.
using System.Diagnostics;using System.Diagnostics.Metrics;
public sealed class ClaimsMetrics{ public const string MeterName = "Contoso.Claims";
private readonly Counter<long> _submitted; private readonly Histogram<double> _processingSeconds;
public ClaimsMetrics(IMeterFactory meterFactory) { var meter = meterFactory.Create(MeterName);
_submitted = meter.CreateCounter<long>( "claims.submitted", unit: "{claim}", description: "Claims accepted by the API");
_processingSeconds = meter.CreateHistogram<double>( "claims.processing.duration", unit: "s", description: "Time to validate and persist a claim");
// Observable gauge: the value is read when the collector scrapes. meter.CreateObservableGauge( "claims.workers.active", () => ActiveWorkers.Current, unit: "{worker}"); }
// Tags must be low-cardinality: a handful of known values. public void ClaimSubmitted(string claimType) => _submitted.Add(1, new KeyValuePair<string, object?>("claim.type", claimType));
public void ClaimProcessed(double seconds, string outcome) => _processingSeconds.Record(seconds, new TagList { { "outcome", outcome } });}
public static class ActiveWorkers{ private static int _current; public static int Current => Volatile.Read(ref _current); public static void Increment() => Interlocked.Increment(ref _current); public static void Decrement() => Interlocked.Decrement(ref _current);}6.2 Wiring OpenTelemetry and a Minimal API endpoint
NuGet packages: OpenTelemetry.Extensions.Hosting, OpenTelemetry.Instrumentation.AspNetCore, OpenTelemetry.Instrumentation.Http, OpenTelemetry.Instrumentation.Runtime, OpenTelemetry.Exporter.OpenTelemetryProtocol.
using System.Diagnostics;using OpenTelemetry.Metrics;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddSingleton<ClaimsMetrics>();
builder.Services.AddOpenTelemetry() .WithMetrics(metrics => metrics .AddAspNetCoreInstrumentation() // http.server.request.duration, active requests .AddHttpClientInstrumentation() // http.client.request.duration .AddRuntimeInstrumentation() // GC, heap, thread pool, exceptions .AddMeter(ClaimsMetrics.MeterName) .AddOtlpExporter()); // endpoint from OTEL_EXPORTER_OTLP_ENDPOINT
var app = builder.Build();
app.MapPost("/claims", async (ClaimRequest request, ClaimsMetrics metrics) =>{ var started = Stopwatch.GetTimestamp(); try { await Task.Delay(50); // stand-in for validation + DB save metrics.ClaimSubmitted(request.ClaimType); metrics.ClaimProcessed(Stopwatch.GetElapsedTime(started).TotalSeconds, "success"); return Results.Accepted(); } catch { metrics.ClaimProcessed(Stopwatch.GetElapsedTime(started).TotalSeconds, "error"); throw; }});
app.Run();
public record ClaimRequest(string ClaimType, decimal Amount);The ASP.NET Core and System.Net.Http instrumentations use the built-in meters Microsoft.AspNetCore.Hosting and System.Net.Http. Since .NET 8 these emit OpenTelemetry semantic-convention metrics such as http.server.request.duration, so you do not have to hand-write request metrics.
6.3 Angular side (browser to API latency)
Backend metrics do not show what the user feels. This functional interceptor (works with any current standalone-based Angular version; register with provideHttpClient(withInterceptors([timingInterceptor]))) measures the round trip and reports it to the backend, which records it in a histogram.
import { HttpInterceptorFn } from '@angular/common/http';import { tap } from 'rxjs';
export const timingInterceptor: HttpInterceptorFn = (req, next) => { const started = performance.now(); return next(req).pipe( tap({ next: (event) => { if ((event as { type?: number }).type === 4) { // HttpEventType.Response report(req.method, (event as { status: number }).status, started); } }, error: (err) => report(req.method, err.status ?? 0, started), }) );};
function report(method: string, status: number, started: number): void { const body = JSON.stringify({ // Send a route template, never the raw URL with IDs (cardinality!). route: 'claims', method, statusClass: `${Math.floor(status / 100)}xx`, durationMs: Math.round(performance.now() - started), }); navigator.sendBeacon('/api/telemetry/ui-timing', new Blob([body], { type: 'application/json' }));}For production, Azure Monitor’s browser SDK (Application Insights JavaScript SDK) already collects page-load and AJAX timings, so write this only when you need custom, business-specific UI timings.
6.4 Database metrics
Most database numbers come from the platform, not your code:
- SQL Server / Azure SQL: platform metrics
cpu_percent,dtu_consumption_percent,connection_failed,deadlock; query-level data from Query Store. - PostgreSQL (Azure Database for PostgreSQL flexible server): platform metrics
cpu_percent,active_connections,storage_percent; query-level data frompg_stat_statements. - In your app: connection-pool usage and command duration. The Npgsql and
Microsoft.Data.SqlClientdrivers publish their own metrics/counters, which you can add to the OpenTelemetry pipeline (check the driver version documentation for the exact meter name before enabling).
Level 3: Advanced
Performance
- Recording a counter or histogram measurement is allocation-free and takes tens of nanoseconds when a listener is attached. Prefer
TagList(a struct) for two or more tags to avoid array allocations. - Aggregation happens in-process. The export interval (default 60 s in the OpenTelemetry .NET SDK; Prometheus typically scrapes every 15-60 s) decides the freshness of your alerts. Do not set a 1-second interval “for better data”.
Cardinality is the number one production risk
Every distinct combination of tag values creates a new time series. claim.type (6 values) x outcome (3 values) = 18 series, which is fine. Adding customer.id with 500,000 values yields millions of series, an exploding bill, slow queries and, in Prometheus, memory exhaustion. Rules:
- Tags must have a small, bounded, known set of values.
- Never use IDs, e-mail addresses, raw URLs, exception messages or free text as tags.
- Use route templates (
/claims/{id}), which ASP.NET Core instrumentation does by default viahttp.route. - Put per-claim detail in traces (exemplars) or structured logs.
Histograms and percentiles
- Never alert on or report the average latency. Use p95/p99 via
histogram_quantile(Prometheus) or the percentile aggregation in Azure Monitor. - Bucket boundaries decide accuracy. The default OpenTelemetry boundaries are generic. For a metric measured in seconds, set boundaries around your SLO (for example 2 s):
// In the WithMetrics(...) builder.AddView("claims.processing.duration", new ExplicitBucketHistogramConfiguration { Boundaries = new double[] { 0.05, 0.1, 0.25, 0.5, 1, 2, 5, 10 } })On .NET 9 and later you can alternatively suggest boundaries at instrument creation using InstrumentAdvice<T> (HistogramBucketBoundaries).
- Percentiles cannot be averaged across instances. Aggregate the buckets (
sum by (le)) first, then compute the quantile.
Example PromQL (Prometheus / Azure Managed Prometheus)
# p95 claim processing time over 5 minutes, across all instanceshistogram_quantile(0.95, sum by (le) (rate(claims_processing_duration_seconds_bucket[5m])))
# Error ratio of claim processingsum(rate(claims_processing_duration_seconds_count{outcome="error"}[5m]))/sum(rate(claims_processing_duration_seconds_count[5m]))OpenTelemetry names are converted by the Prometheus exporters: dots become underscores, the unit s becomes the _seconds suffix, and counters get _total.
Security
- Do not expose
/metricspublicly. Bind it to an internal port or a separate management listener, restrict with network policy, or use OTLP push to an internal collector. - Metrics leak business volume (claims per minute), so treat them as internal data and never put PII in labels.
Failure modes and common mistakes
| Mistake | Consequence | Fix |
|---|---|---|
| High-cardinality tags | Cost explosion, slow dashboards | Bound tag values, review new tags in PRs |
| Alerting on averages or CPU only | Misses user pain | Alert on SLO burn rate, p95/p99, error ratio |
| Counters read as instantaneous values | Wrong graphs | Always use rate() / increase() on counters |
| Metrics lost on scale-in (push pipeline flush) | Missing last data points | Let the OpenTelemetry SDK flush on host shutdown, increase graceful shutdown time |
| Different names per team | Cannot compare services | Naming convention plus shared library (Day 40 Chassis) |
| Only technical metrics | Business failures invisible | Add business counters (claims submitted/rejected) |
Creating a Meter/instrument per request | Memory leak | Create instruments once, at startup (singleton) |
| Ignoring the alert noise | Alert fatigue, pager ignored | Alert on symptoms with a duration (for: 5m), route by severity |
Level 4: Expert and Architect view
Pull vs push and tool choices
| Option | Model | Strengths | Trade-offs | Fit for us |
|---|---|---|---|---|
| Prometheus scrape + Grafana | Pull | Open standard, PromQL, rich ecosystem, great with Kubernetes | You run/scale storage (or use managed) | AKS workloads |
| OpenTelemetry OTLP to a Collector | Push | Vendor-neutral, same pipeline for traces and logs, easy to switch backends | Extra component (Collector) to operate | Default choice for new services |
| Azure Monitor / Application Insights (OTel distro) | Push to Azure | Zero infrastructure, integrated alerting and portal | Vendor coupling, log-based billing for custom metrics | App Service / Container Apps / Functions |
| StatsD / Graphite | Push (UDP) | Simple | Weak dimensional model, legacy | Not for new work |
| EventCounters / PerfMon | Local | Built in | No dimensions, no central store | Diagnostics only |
Combines with
- Distributed Tracing (Day 31) via exemplars: click a slow histogram bucket and jump to a representative trace.
- Health Check API (Day 36): health answers “should I get traffic?”, metrics answer “how well am I doing?”.
- Circuit Breaker, Bulkhead, Retry (Days 26-28): expose breaker state, bulkhead queue length and retry counts as metrics, otherwise these patterns hide problems.
- Rate Limiter (Day 30): rejected-request counters.
- Log Deployments & Changes (Day 37): deployment markers on the metric dashboards.
- Microservice Chassis (Day 40): one shared library that wires the meters, naming and exporters for every service.
- Autoscaling (KEDA/HPA): metrics drive scaling.
Design guidance
- Standardise on RED for every service and USE for every resource; add 2-5 business metrics per service.
- Define SLIs as metric queries, then SLOs, then alert on error-budget burn rate (multi-window) rather than raw thresholds.
- Standard metric naming: OpenTelemetry semantic conventions for HTTP/DB/messaging,
claims.*namespace for domain metrics, units always declared.
ADR-style justification
ADR-033: Standardise application metrics on OpenTelemetry
- Status: Proposed
- Context: The Claims platform has 12 services on Azure (AKS and App Service) with .NET back ends and an Angular front end. Incidents in the last quarter were detected by customers because only CPU and memory graphs exist. Each team names metrics differently.
- Decision: All .NET services emit metrics with
System.Diagnostics.Metrics, exported via OpenTelemetry (OTLP). AKS workloads send to an OpenTelemetry Collector, then to Azure Monitor managed Prometheus, and are visualised in Azure Managed Grafana. App Service workloads use the Azure Monitor OpenTelemetry distro. RED metrics and OpenTelemetry semantic conventions are mandatory; aClaimsMetrics-style class per bounded context holds business metrics. Tags must be low-cardinality and are reviewed in code review. - Consequences (+): vendor-neutral instrumentation, consistent dashboards, SLO-based alerts, ability to change backend without touching code. (-): Collector to operate, cost governance for cardinality, team training needed.
- Alternatives rejected: Application Insights classic SDK only (locks instrumentation to one vendor and is superseded by the OpenTelemetry distro for new work); self-hosted Prometheus/Grafana (operational burden for our team size).
Azure implementation
Services
| Azure service | Role for metrics |
|---|---|
| Azure Monitor Metrics | Free platform metrics for every resource (App Service, AKS, Service Bus, Azure SQL, PostgreSQL) with alert rules and autoscale |
| Application Insights (workspace-based) + Azure Monitor OpenTelemetry distro | Application metrics, traces and logs for App Service, Container Apps, Functions and VMs |
| Azure Monitor managed service for Prometheus (Azure Monitor workspace) | Prometheus-compatible storage, PromQL and Prometheus rule/alert groups; the metrics add-on for AKS scrapes pods |
| Azure Managed Grafana | Dashboards over Prometheus, Azure Monitor and Log Analytics |
| Azure Monitor alerts + Action Groups | Notify Teams, e-mail, webhook, ITSM |
| Container Insights | Node/pod CPU, memory, restarts on AKS |
Configuration
- ASP.NET Core on App Service / Container Apps: add the
Azure.Monitor.OpenTelemetry.AspNetCorepackage and callbuilder.Services.AddOpenTelemetry().UseAzureMonitor();. Set the connection string through theAPPLICATIONINSIGHTS_CONNECTION_STRINGapp setting (do not hard-code it). Register custom meters with.WithMetrics(m => m.AddMeter(ClaimsMetrics.MeterName))on the same OpenTelemetry builder. - AKS: create an Azure Monitor workspace, enable managed Prometheus on the cluster (
az aks update --enable-azure-monitor-metrics ...with the workspace ID), and expose your pods’ metrics through aPodMonitor/ServiceMonitorcustom resource or the pod annotation-based scrape settings. Alternatively send OTLP to an OpenTelemetry Collector that remote-writes to the workspace. - Grafana: create an Azure Managed Grafana instance, link it to the Azure Monitor workspace (its managed identity needs Monitoring Reader) and import the standard dashboards.
- Alerts: create Prometheus rule groups for SLO-style rules (error ratio, p95) and Azure Monitor metric alert rules for platform metrics (Service Bus active messages, SQL CPU). Send to an Action Group.
- Infrastructure as code: define all of the above in Bicep or Terraform so alerts are versioned with the service.
Pricing and tier considerations
Always confirm current numbers in the Azure pricing calculator, as they vary by region and currency.
- Platform metrics are collected and stored free for Azure resources.
- Application Insights / Log Analytics is billed by data ingested (Analytics logs, pay-as-you-go or commitment tiers). The first 5 GB per month per billing account on the pay-as-you-go Analytics tier is free. With workspace-based Application Insights, custom metrics sent through OpenTelemetry count as ingested data, so cardinality and export frequency drive cost. Use sampling for traces (not for metrics, which are already aggregated) and daily caps as a safety net.
- Managed Prometheus is billed on the number of samples ingested and on samples processed by queries; there is no per-instance charge. Cardinality control (drop unused metrics with
metric_relabel_configs/keep lists) is the main cost lever. - Azure Managed Grafana: the Standard SKU (instance sizes X1 default and X2) is the supported tier. Microsoft documents the Essential tier as deprecated, with new Essential workspaces blocked and full deprecation on March 31, 2027, so do not start new work on Essential. Microsoft also offers “Azure Monitor dashboards with Grafana” (preview) as a free, Azure-native alternative for basic needs.
- Alert rules are billed per rule / signal type; check the Azure Monitor pricing page.
Reference architecture (text)
Angular SPA (Application Insights JS SDK for page and AJAX timings) calls Azure Front Door and API Management, which route to the Claims API, Policy API and Fraud service running on AKS (and the Notification function on Azure Functions). Every .NET pod emits System.Diagnostics.Metrics data (ASP.NET Core, HttpClient, runtime and Contoso.Claims meters). On AKS, the managed Prometheus add-on scrapes the pods and writes to an Azure Monitor workspace; on App Service and Functions the Azure Monitor OpenTelemetry distro pushes to workspace-based Application Insights. Azure SQL, PostgreSQL and Service Bus platform metrics flow into Azure Monitor automatically. Azure Managed Grafana shows one RED dashboard per service, with a claims-business dashboard on top. Prometheus rule groups and Azure Monitor alert rules notify the on-call channel through an Action Group; deployment markers from the pipeline are annotated on the dashboards.
Teaching guide for my team
Explain to a beginner in 2 minutes
“Your car has a dashboard: speed, fuel, warning lights. Your service needs one too. A metric is a number your code updates as it works: a counter of claims submitted, a timer of how long each claim took, a gauge of how many are waiting. A collector reads these numbers every few seconds and draws graphs and sends alerts. Logs tell a story about one event; metrics tell you the health of the whole system right now. Remember three words: Rate (how many), Errors (how many failed), Duration (how long).”
Explain to an intermediate developer in 5 minutes
Cover: (1) the three instruments and when to choose each; (2) Meter created via IMeterFactory once, instruments as singletons; (3) OpenTelemetry pipeline: AddMeter, instrumentation packages, exporter; (4) tags and cardinality with the 18-series versus millions-of-series example; (5) why histograms and percentiles beat averages; (6) how metrics feed alerts and autoscaling; (7) metrics vs traces vs logs: metrics say something is wrong, traces say where, logs say why.
Hands-on exercise
Add metrics to a sample Claims API:
- Create
ClaimsMetricswith a counterclaims.submitted(tagclaim.type) and a histogramclaims.processing.duration(tagoutcome). - Wire OpenTelemetry with ASP.NET Core, HttpClient and runtime instrumentation and the console exporter (
OpenTelemetry.Exporter.Console) or a local OTLP collector. - Call the endpoint 200 times with a random 50-2000 ms delay and 10% forced failures.
- Write the PromQL (or view in the console exporter) for p95 latency and error ratio.
Expected outcome: the counter shows 200 total split by claim type; the histogram shows the p95 near the upper delay range; the error ratio is close to 10%. As a stretch, a student who adds claimId as a tag should observe the series count exploding and explain why that is wrong.
Interview-style questions
- Why use a histogram instead of an average latency? Averages hide the slow tail. A histogram lets you compute p95/p99 and aggregate across instances by summing buckets.
- What is metric cardinality and why does it matter? It is the number of unique tag-value combinations, each a separate time series. High cardinality (IDs, raw URLs) explodes cost and memory and slows queries.
- When would you choose metrics over logs or traces? For continuous health, alerting, trends, SLOs and autoscaling. Use traces to follow one request across services and logs for the detailed reason for a failure.
Mastery checklist
- I can choose between a counter, gauge and histogram for a given claims scenario.
- I can create a
MeterwithIMeterFactoryand register it with OpenTelemetry usingAddMeter. - I can explain RED and USE and list the metrics each Claims service must expose.
- I can identify and fix a high-cardinality tag before it reaches production.
- I can write a PromQL query for p95 latency and an error ratio and explain why we sum buckets first.
- I can design an alert on symptoms (error ratio, burn rate) with a duration, not on raw CPU.
- I can describe how metrics reach Azure Monitor / managed Prometheus / Managed Grafana and what drives their cost.
- I can explain when metrics are the wrong tool and traces or logs should be used.
Key takeaway
Measure at the boundaries with a few low-cardinality Rate, Error and Duration metrics plus a handful of business counters, and alert on symptoms users feel rather than on averages or raw CPU. Metrics tell you that and where something is wrong; traces and logs tell you why.
