A Microservice Chassis is a shared, versioned starter library (in .NET: a NuGet package, or a small set of them) that every service references so that logging, tracing, metrics, health checks, authentication, error...
Intro
A Microservice Chassis is a shared, versioned starter library (in .NET: a NuGet package, or a small set of them) that every service references so that logging, tracing, metrics, health checks, authentication, error handling and resilient HTTP clients are configured the same way, with one line of code in Program.cs. Instead of every team re-implementing these concerns slightly differently, the platform team ships them once, and service teams write only business logic. In our insurance claims system, Claims, Policy, Payments and Fraud services all call builder.AddClaimsChassis() and get identical behaviour on day one.
Why we need this
- Every service needs the same 80% of plumbing. Structured logging, correlation IDs, OpenTelemetry,
/health/liveand/health/ready, JWT validation, ProblemDetails error responses, and resilient outbound HTTP are needed by Claims, Policy, Payments and Fraud alike. - Copy-paste drifts. When each team copies the setup from the last service, versions and settings diverge. One service logs JSON, another logs plain text; one propagates
traceparent, another does not. - Security fixes must roll out fast. If a JWT validation bug is fixed in one place (the chassis) and consumed as a version bump, you patch 40 services with a Dependabot/Renovate PR, not 40 manual code reviews.
- New service = productive in hours. A developer creating the “Subrogation” service should not spend a week on plumbing.
- Operations gets uniformity. On-call engineers can rely on every service exposing the same health endpoints, metric names and log fields.
What problem it solves
Problem (from the topic list): Teams reinvent logging, health checks and auth in every service.
Without a chassis:
- Claims logs
claimId, Payments logsClaimID, Fraud logsclaim_id. A single Kibana/Log Analytics query cannot follow one claim across services. - The Payments service forgets to validate the token
audience, and any valid token from another API is accepted. Nobody notices for months. - Two services return raw exception messages to clients (information leak), while others return RFC 7807 problem details.
- Kubernetes/Container Apps cannot tell if a service is ready because half the services have no readiness probe.
- A CVE in a logging package requires 40 separate investigations.
The chassis moves these decisions to one owned, tested, versioned place.
When it is needed (and when it is NOT)
Use it when:
- You have (or will soon have) roughly five or more services on the same stack, owned by more than one team.
- Cross-cutting behaviour must be provably consistent (security, audit, compliance).
- A platform or enablement team exists that can own and version it.
Do not use it (or keep it very thin) when:
- You have 1-3 services and one team. A shared library is overhead; copy a
ServiceDefaults-style project instead. - Services are polyglot with no dominant stack. A library cannot cross languages; use a Service Mesh or Sidecar for network concerns (Days 47-48) and a Service Template (Day 41) for repo layout.
- The “chassis” would contain business logic or shared domain models. That creates a distributed monolith: every release forces a coordinated upgrade.
- Teams cannot upgrade regularly. A chassis that lags 2 majors behind becomes a liability.
How to identify the problem (key signals)
- Diffing
Program.csacross two services shows 100+ lines of near-identical, slightly different setup. - Log queries need per-service field names (
claimIdvsClaimID), or trace IDs break at some hops. - A security review finds inconsistent JWT settings (
ValidateAudience = falsein some services). - Some services lack
/health/ready; deployments route traffic to pods that are not ready. - Error responses differ: HTML error pages, plain strings, ProblemDetails, and stack traces, all from the same API surface.
- New-service onboarding takes days, mostly copying and fixing config from another repo.
- A dependency upgrade (for example OpenTelemetry) needs a separate PR in every repo, and half are stuck on old versions.
Flow Diagram
A versioned chassis package gives every service the same cross-cutting baseline.
flowchart TB PLAT["Platform team"] --> PKG["Claims.Chassis NuGet - SemVer"] PKG --> F1["OpenTelemetry"] PKG --> F2["Health live/ready"] PKG --> F3["JWT defaults"] PKG --> F4["ProblemDetails"] PKG --> F5["Resilient HttpClient"] PKG -- "AddClaimsChassis()" --> S1["Claims service"] PKG -- "AddClaimsChassis()" --> S2["Policy service"] PKG -- "AddClaimsChassis()" --> S3["Payments service"] REN["Renovate / Dependabot"] -. "upgrade PRs" .-> S1Level 1: Beginner
Analogy: A car chassis. Every model (sedan, SUV, van) is built on the same frame with the same brakes, steering and wiring. Designers focus on the body. Here, the “body” is claim-handling logic, and the chassis provides brakes (resilience), dashboard (metrics/health), and locks (auth).
Minimal example: a class library Claims.Chassis with a single extension method.
using Microsoft.AspNetCore.Builder;using Microsoft.Extensions.DependencyInjection;
namespace Claims.Chassis;
public static class ChassisExtensions{ public static WebApplicationBuilder AddClaimsChassis(this WebApplicationBuilder builder) { builder.Services.AddProblemDetails(); builder.Services.AddHealthChecks(); return builder; }
public static WebApplication UseClaimsChassis(this WebApplication app) { app.UseExceptionHandler(); app.UseStatusCodePages(); app.MapHealthChecks("/health/live"); return app; }}A service consumes it:
using Claims.Chassis;
var builder = WebApplication.CreateBuilder(args);builder.AddClaimsChassis();
var app = builder.Build();app.UseClaimsChassis();app.MapGet("/claims/{id:guid}", (Guid id) => Results.Ok(new { id, status = "Open" }));app.Run();Result: two lines give the service uniform error responses and a health endpoint.
Level 2: Intermediate
A realistic chassis for .NET 10 (current LTS) bundles: OpenTelemetry (traces, metrics, logs), health checks split into live/ready, JWT bearer auth, resilient HttpClient defaults, and a correlation/log enricher. Packages used: OpenTelemetry.Extensions.Hosting, OpenTelemetry.Instrumentation.AspNetCore, OpenTelemetry.Instrumentation.Http, OpenTelemetry.Exporter.OpenTelemetryProtocol, Microsoft.Extensions.Http.Resilience, Microsoft.AspNetCore.Authentication.JwtBearer.
using Microsoft.AspNetCore.Authentication.JwtBearer;using Microsoft.AspNetCore.Builder;using Microsoft.AspNetCore.Diagnostics.HealthChecks;using Microsoft.Extensions.Configuration;using Microsoft.Extensions.DependencyInjection;using Microsoft.Extensions.Diagnostics.HealthChecks;using Microsoft.Extensions.Logging;using OpenTelemetry.Metrics;using OpenTelemetry.Trace;
namespace Claims.Chassis;
public sealed class ChassisOptions{ public string ServiceName { get; set; } = "unknown"; public string? Authority { get; set; } public string? Audience { get; set; }}
public static class ChassisExtensions{ public static WebApplicationBuilder AddClaimsChassis( this WebApplicationBuilder builder, Action<ChassisOptions>? configure = null) { var options = new ChassisOptions { ServiceName = builder.Environment.ApplicationName }; builder.Configuration.GetSection("Chassis").Bind(options); configure?.Invoke(options);
// 1. Structured logging with OpenTelemetry builder.Logging.AddOpenTelemetry(o => { o.IncludeFormattedMessage = true; o.IncludeScopes = true; });
// 2. Traces + metrics builder.Services.AddOpenTelemetry() .ConfigureResource(r => r.AddService(options.ServiceName)) .WithTracing(t => t.AddAspNetCoreInstrumentation().AddHttpClientInstrumentation()) .WithMetrics(m => m.AddAspNetCoreInstrumentation() .AddHttpClientInstrumentation() .AddRuntimeInstrumentation());
// 3. Errors: RFC 7807 problem details, no stack traces to clients builder.Services.AddProblemDetails();
// 4. Health: "live" = process is up; "ready" = dependencies OK builder.Services.AddHealthChecks() .AddCheck("self", () => HealthCheckResult.Healthy(), tags: new[] { "live" });
// 5. Auth: validate issuer, audience and lifetime by default if (options.Authority is not null) { builder.Services.AddAuthentication(JwtBearerDefaults.AuthenticationScheme) .AddJwtBearer(o => { o.Authority = options.Authority; o.Audience = options.Audience; o.TokenValidationParameters.ValidateAudience = true; o.TokenValidationParameters.ValidateLifetime = true; }); builder.Services.AddAuthorization(); }
// 6. Resilient outbound HTTP for every HttpClient builder.Services.ConfigureHttpClientDefaults(http => http.AddStandardResilienceHandler());
return builder; }
public static WebApplication UseClaimsChassis(this WebApplication app) { app.UseExceptionHandler(); app.UseStatusCodePages();
if (app.Configuration.GetSection("Chassis:Authority").Exists()) { app.UseAuthentication(); app.UseAuthorization(); }
app.MapHealthChecks("/health/live", new HealthCheckOptions { Predicate = r => r.Tags.Contains("live") }); app.MapHealthChecks("/health/ready"); return app; }}Usage in the Claims service, with a database readiness check added by the service itself (only the service knows its dependencies):
var builder = WebApplication.CreateBuilder(args);builder.AddClaimsChassis(o => o.ServiceName = "claims-api");builder.Services.AddHealthChecks() .AddNpgSql(builder.Configuration.GetConnectionString("Claims")!, name: "postgres");
var app = builder.Build();app.UseClaimsChassis();app.MapGet("/claims/{id:guid}", (Guid id) => Results.Ok(new { id })).RequireAuthorization();app.Run();(AddNpgSql comes from AspNetCore.HealthChecks.NpgSql.)
Angular side: the chassis idea also applies to the front end as a shared npm package (for example @claims/http-core) exposing an HTTP interceptor that attaches the bearer token and the traceparent/correlation header, and a global error handler. Keep it small and versioned the same way.
import { HttpInterceptorFn } from '@angular/common/http';
export const correlationInterceptor: HttpInterceptorFn = (req, next) => { const id = crypto.randomUUID(); return next(req.clone({ setHeaders: { 'X-Correlation-Id': id } }));};
// app.config.ts// provideHttpClient(withInterceptors([correlationInterceptor]))Level 3: Advanced
Performance: The chassis runs in every request path. Keep middleware light, avoid synchronous I/O at startup, and never add expensive global filters. Measure startup time and per-request overhead in the chassis’s own benchmark project.
Scalability (of the organisation): Publish to a private NuGet feed with semantic versioning. Use automated dependency PRs (Renovate/Dependabot) so services stay within one or two minor versions of the latest.
Security: Default to secure: audience and issuer validation on, HTTPS metadata required, no detailed errors in production. Make insecure choices explicit and reviewable, not the default. Never let the chassis read secrets from source control; consume Key Vault (Day 39).
Failure modes and common mistakes:
- Distributed monolith by dependency. A breaking chassis release forces every service to upgrade at once. Mitigation: SemVer, deprecation windows, additive changes, and multi-targeting where needed.
- Chassis grows into a “framework”. Everything ends up in it, including domain code. Rule: only cross-cutting, business-agnostic concerns.
- Transitive dependency conflicts. The chassis pins a package version that clashes with a service’s own. Mitigation: use
Directory.Packages.props(central package management) and keep chassis dependencies minimal. - Hidden magic. Developers cannot tell what
AddClaimsChassis()registers. Mitigation: document each feature, make each one opt-out through options, and log at startup what was enabled. - No escape hatch. A service with special needs (for example, a file-upload API with a different timeout) forks the chassis. Mitigation: options and extension points.
- Coupling to one team’s release cadence. Give the chassis its own CI, tests and changelog.
Level 4: Expert and Architect view
Alternatives compared
| Approach | How it works | Strengths | Weaknesses | Best fit |
|---|---|---|---|---|
| Microservice Chassis (shared library) | NuGet/npm package referenced by each service | In-process, full access to app code (logging enrichers, auth policies); easy to debug | Language-specific; version-upgrade effort; coupling risk | One dominant stack (.NET), many teams |
| Service Template (Day 41) | Skeleton repo copied at creation | Full flexibility afterwards; gives repo layout, CI/CD | Copies drift; no automatic updates | Consistent scaffolding, combined with a chassis |
| Sidecar (Day 47) | Helper container in the same pod/replica | Language-agnostic; independent lifecycle | Extra resources; cannot touch in-process concerns | Network-level concerns, polyglot fleets |
| Service Mesh (Day 48) | Proxy layer (Istio, Envoy) handles mTLS, retries, telemetry | No code changes; uniform network policy | Operational complexity; not app-level (no business logging, no auth claims logic) | Large Kubernetes estates |
Aspire-style ServiceDefaults project | A shared project inside one solution | Very light; ideal for small estates | Not versioned independently across repos | Few services, one repo |
| Copy-paste | Each team writes its own | No coupling | Drift, inconsistent security | Prototypes only |
Combines with: Externalized Configuration (Day 39: the chassis reads config and secrets), Health Check API (Day 36), Distributed Tracing (Day 31), Access Token (Day 38), Circuit Breaker and Retry (Days 26 and 28, provided by the resilience handler), Service Template (Day 41), and Sidecar/Service Mesh (Days 47-48) for the concerns a library should not own.
ADR-style justification
- Title: ADR-040: Adopt a shared .NET chassis (
Claims.Chassis) for cross-cutting concerns. - Status: Proposed.
- Context: Nine .NET services owned by four teams. Audit found inconsistent JWT validation, three logging formats, and missing readiness probes in four services.
- Decision: Create
Claims.Chassis(NuGet, private feed) owned by the platform team. It provides OpenTelemetry, health checks, JWT defaults, ProblemDetails and standard resilience. It contains no domain code. SemVer is mandatory; breaking changes need one minor version of deprecation notice. - Alternatives considered: Service mesh only (rejected: does not cover in-process logging and claims-based auth); copy-paste template only (rejected: cannot deliver security fixes centrally).
- Consequences: Positive: uniform telemetry and security, faster onboarding. Negative: platform team becomes a dependency, and upgrade discipline is required. Mitigations: automated update PRs, a compatibility test suite, and options-based opt-outs.
Azure implementation
A chassis is a code library, so Azure does not “implement” it; Azure hosts, distributes and observes what it configures.
Services that support it
- Azure Artifacts (Azure DevOps): private NuGet/npm feed for the chassis packages. Feed storage is free up to a per-organisation quota; beyond that it is billed per GiB (check current Azure DevOps pricing). GitHub Packages is an alternative.
- Azure Monitor / Application Insights: the chassis can wire telemetry with the
Azure.Monitor.OpenTelemetry.AspNetCorepackage (builder.Services.AddOpenTelemetry().UseAzureMonitor()), with the connection string coming from configuration. Billing is by data ingested (GB), so sampling defaults belong in the chassis. - Azure Key Vault and Azure App Configuration: the chassis registers them as configuration sources using
DefaultAzureCredential/ managed identity, so no secrets live in config (ties to Day 39). - Microsoft Entra ID: the
AuthorityandAudiencethe chassis validates against (https://login.microsoftonline.com/{tenant}/v2.0). UseMicrosoft.Identity.Webin the chassis if you prefer its helpers. - Azure Container Apps / AKS / App Service: map
/health/liveand/health/readyto liveness and readiness probes. Container Apps supports probes in the container configuration; hosting cost depends on the plan (Container Apps consumption is billed per vCPU-second and GiB-second with a monthly free grant; verify current numbers). - Azure Container Registry: stores the images built from the template that consume the chassis.
How to configure (example)
// inside AddClaimsChassis, when APPLICATIONINSIGHTS_CONNECTION_STRING is presentif (!string.IsNullOrEmpty(builder.Configuration["APPLICATIONINSIGHTS_CONNECTION_STRING"])){ builder.Services.AddOpenTelemetry().UseAzureMonitor();}Build pipeline: on tag v*, run chassis tests, dotnet pack, and push to the Azure Artifacts feed. Consumers use Directory.Packages.props to pin the version, and Renovate raises upgrade PRs.
Reference architecture (text): The platform team’s repo builds Claims.Chassis in Azure Pipelines, publishing to an Azure Artifacts feed. Service repos (Claims, Policy, Payments, Fraud) restore from that feed, build container images pushed to Azure Container Registry, and deploy to Azure Container Apps. At runtime each service authenticates callers through Entra ID tokens, reads secrets from Key Vault via managed identity, exports OpenTelemetry to Application Insights, and exposes live/ready probes used by the platform. Log Analytics queries and dashboards work across all services because field names are identical.
Teaching guide for my team
2-minute explanation for a beginner “Every service we build needs the same boring parts: logs, health checks, login checks, and safe error messages. Instead of rewriting them in each service, we put them in one package. You add the package, call one method, and your service already behaves like all the others. You only write the claim logic.”
5-minute explanation for an intermediate developer Walk through AddClaimsChassis(): what it registers (OpenTelemetry, health checks, JWT, ProblemDetails, resilience), how to override via ChassisOptions, and why the Claims service adds its own Postgres readiness check (the chassis cannot know your dependencies). Explain versioning: the chassis follows SemVer; upgrade through the automated PR; never fork it. Explain the boundary rule: cross-cutting and business-agnostic only, no domain models, no shared DTOs.
Hands-on exercise Create Chassis.Demo (class library) and Demo.Api (minimal API). Implement AddChassis() that registers ProblemDetails and health checks (/health/live, /health/ready). Add a ChassisOptions.ServiceName and log it at startup. Expected outcome: GET /health/live returns 200 Healthy; an endpoint that throws returns application/problem+json with status 500 and no stack trace; a second API project can adopt the same behaviour by adding the reference and one method call.
Interview-style questions
- How is a chassis different from a service template? A chassis is a versioned library that provides runtime behaviour and can be upgraded centrally; a template is a one-time copy of repo structure and pipelines that then drifts.
- What is the main risk of a shared chassis? Coupling: breaking changes force coordinated upgrades and can create a distributed monolith. Mitigate with SemVer, deprecation windows, and keeping it free of business logic.
- Why not use a service mesh instead? A mesh handles network concerns (mTLS, retries, routing) without code changes, but it cannot do in-process work such as structured log enrichment, claims-based authorization or app-level health checks. They complement each other.
Mastery checklist
- I can list what belongs in a chassis (logging, tracing, metrics, health, auth defaults, error handling, resilience) and what does not (domain logic, shared DTOs).
- I can build a working
AddXxxChassis()extension for .NET 10 and consume it from a new service. - I can explain how services extend the chassis (extra health checks, options) without forking it.
- I can describe a SemVer and deprecation policy that avoids forced big-bang upgrades.
- I can distinguish chassis, template, sidecar and mesh, and choose a combination for a given estate.
- I can publish the package to a private feed (Azure Artifacts) and automate consumer upgrades.
- I can write an ADR that justifies (or rejects) a chassis for a given team size.
- I can wire live/ready endpoints to Azure Container Apps or AKS probes.
Key takeaway
A chassis gives every service the same secure, observable baseline through one versioned package, but only if it stays small, business-agnostic and easy to upgrade. Once it holds domain logic, it becomes a distributed monolith.
