A self-contained service answers a client request using only its own code and its own local data.
Intro
A self-contained service answers a client request using only its own code and its own local data. It never has to wait on a synchronous call to another service in the middle of that request. Data it needs from other services arrives earlier, through events or scheduled replication, and is stored locally as a read-only copy. The result is a service whose availability and latency depend on itself, not on a chain of neighbours.
Why we need this
In a microservices system, availability multiplies. If a request touches five services in a chain and each is 99.9% available, the request is about 99.5% available (0.999^5), which is roughly 44 hours of downtime a year instead of 8.8. Latency adds up the same way: five hops of 80 ms is 400 ms before any business logic runs.
Teams also split a monolith into services and then rebuild the monolith’s coupling over the network. Every “get the customer, then get the policy, then get the coverage” call becomes a distributed dependency that must be up, fast, and version-compatible at the same instant.
The self-contained service pattern says: design service boundaries and data ownership so that the common request path is answerable locally. Cross-service data is copied in ahead of time, not fetched on demand.
What problem it solves
Problem: synchronous RPC dependency chains create cascading latency and low availability.
Insurance example: the Claims API POST /claims needs to confirm the policy is active, read the coverage limits, and check the claimant’s status. In a naive design, ClaimsService calls PolicyService (which calls CustomerService, which calls BillingService to check arrears). Failure modes without the pattern:
BillingServiceslows down under month-end load, so claim submission times out even though claims logic is healthy.- A deploy of
CustomerServicebreaks a response field and claims fail at runtime. - Thread pools in
ClaimsServicefill with waiting requests and the service stops answering health checks. - Nobody can say which team owns the outage, because one user action spans four services.
With the pattern, ClaimsService holds a local PolicySnapshot table (policy number, status, coverage limits, effective dates) kept current by PolicyActivated, PolicyLapsed, and CoverageChanged events. Submitting a claim is one local database transaction.
When it is needed (and when it is NOT)
Use it when:
- The operation is on a hot, user-facing path (claim submission, quote, checkout) where latency and availability targets are strict.
- The needed foreign data changes slowly compared with how often it is read (policy status changes daily; claims are read hundreds of times).
- Teams own separate services and you want each to keep working when another is deploying or down.
- You already have, or can add, reliable event publishing (see Days 11 to 14 on domain events and the outbox).
Do not use it when:
- The data must be strictly current at the moment of use and a stale copy is unacceptable (for example, checking a fraud block list that must reflect changes within seconds, with legal consequences). A synchronous call, with a circuit breaker, may be right.
- The foreign data is huge or changes constantly; replicating it costs more than calling for it.
- The system is small, one team owns everything, and a modular monolith would be simpler. Copying data between services you could have kept in one database is accidental complexity.
- You cannot tolerate the operational load of maintaining replicas, handling event replay, and reconciling drift.
How to identify the problem (key signals)
- Distributed traces show a request fanning out serially across three or more services before returning.
- p99 latency of one service closely tracks the p99 latency of a downstream service.
- Incident timelines read “Service A was down because Service B was slow”, and A’s own code was not at fault.
- HTTP client timeouts,
TaskCanceledException,HttpRequestException, andSocketExceptionare the top exceptions in a service’s logs. - Thread pool starvation or a rising count of in-flight requests when a dependency degrades.
- Teams cannot deploy independently because integration environments must have every dependency running.
- The composite SLO is worse than any individual service’s SLO, and the gap is explained by call depth.
Flow Diagram
Policy data is replicated ahead of time, so the submit-claim request never blocks on another service.
flowchart LR subgraph Background["Ahead of time - asynchronous"] PS["PolicyService"] --> OB[("Outbox")] OB --> SB{{"Service Bus: policy-events"}} SB --> C["PolicyChanged consumer"] C -- "upsert if Version is newer" --> SNAP[("PolicySnapshots")] end subgraph Hot["Request path - local only"] U["Client"] -- "POST /claims" --> API["ClaimsService API"] API -- "read" --> SNAP API -- "write" --> CL[("Claims")] API -- "201 Created" --> U endLevel 1: Beginner
Analogy: a chef who keeps a prepped station (chopped onions, stock, sauces) beside the stove. Cooking an order does not require walking to the pantry, the fridge, and the supplier each time. A runner restocks the station in the background. If the supplier is closed for an hour, service continues.
Minimal example (C#, .NET 10, minimal API). The service stores a local copy of policy data and uses it on the request path. Assume a top-level Program.cs.
using Microsoft.EntityFrameworkCore;
var builder = WebApplication.CreateBuilder(args);builder.Services.AddDbContext<ClaimsDb>(o => o.UseSqlServer(builder.Configuration.GetConnectionString("ClaimsDb")));var app = builder.Build();
app.MapPost("/claims", async (SubmitClaim cmd, ClaimsDb db) =>{ // Local read only. No HTTP call to PolicyService. var policy = await db.PolicySnapshots .AsNoTracking() .SingleOrDefaultAsync(p => p.PolicyNumber == cmd.PolicyNumber);
if (policy is null || policy.Status != "Active") return Results.UnprocessableEntity("Policy is not active.");
if (cmd.LossDate < policy.EffectiveFrom || cmd.LossDate > policy.EffectiveTo) return Results.UnprocessableEntity("Loss date is outside the policy period.");
if (cmd.Amount > policy.CoverageLimit) return Results.UnprocessableEntity("Amount exceeds coverage limit.");
var claim = new Claim { Id = Guid.NewGuid(), PolicyNumber = cmd.PolicyNumber, LossDate = cmd.LossDate, Amount = cmd.Amount, Status = "Submitted" }; db.Claims.Add(claim); await db.SaveChangesAsync(); return Results.Created($"/claims/{claim.Id}", claim);});
app.Run();
public record SubmitClaim(string PolicyNumber, DateOnly LossDate, decimal Amount);
public class Claim{ public Guid Id { get; set; } public string PolicyNumber { get; set; } = ""; public DateOnly LossDate { get; set; } public decimal Amount { get; set; } public string Status { get; set; } = "";}
public class PolicySnapshot{ public string PolicyNumber { get; set; } = ""; // key public string Status { get; set; } = ""; public DateOnly EffectiveFrom { get; set; } public DateOnly EffectiveTo { get; set; } public decimal CoverageLimit { get; set; } public DateTime LastUpdatedUtc { get; set; }}
public class ClaimsDb(DbContextOptions<ClaimsDb> options) : DbContext(options){ public DbSet<Claim> Claims => Set<Claim>(); public DbSet<PolicySnapshot> PolicySnapshots => Set<PolicySnapshot>(); protected override void OnModelCreating(ModelBuilder b) => b.Entity<PolicySnapshot>().HasKey(p => p.PolicyNumber);}Something else must fill PolicySnapshots. That is Level 2.
Level 2: Intermediate
The snapshot is maintained by an event consumer. PolicyService publishes events (via a transactional outbox, Day 12) to Azure Service Bus; ClaimsService consumes them and upserts its local copy.
Event contract (versioned, additive changes only):
public record PolicyChanged( string PolicyNumber, string Status, DateOnly EffectiveFrom, DateOnly EffectiveTo, decimal CoverageLimit, long Version, // monotonically increasing per policy DateTime OccurredUtc);Consumer using Azure.Messaging.ServiceBus. It guards against out-of-order delivery with the Version field and is idempotent:
using Azure.Messaging.ServiceBus;using System.Text.Json;using Microsoft.EntityFrameworkCore;
public class PolicyChangedConsumer( ServiceBusClient client, IServiceScopeFactory scopes, ILogger<PolicyChangedConsumer> log) : BackgroundService{ protected override async Task ExecuteAsync(CancellationToken ct) { var processor = client.CreateProcessor("policy-events", "claims-service", new ServiceBusProcessorOptions { MaxConcurrentCalls = 4, AutoCompleteMessages = false });
processor.ProcessMessageAsync += async args => { var evt = JsonSerializer.Deserialize<PolicyChanged>(args.Message.Body)!; using var scope = scopes.CreateScope(); var db = scope.ServiceProvider.GetRequiredService<ClaimsDb>();
var row = await db.PolicySnapshots.FindAsync([evt.PolicyNumber], args.CancellationToken); if (row is null) { db.PolicySnapshots.Add(Map(new PolicySnapshot(), evt)); } else if (evt.Version > row.Version) // ignore stale or duplicate events { Map(row, evt); } await db.SaveChangesAsync(args.CancellationToken); await args.CompleteMessageAsync(args.Message, args.CancellationToken); };
processor.ProcessErrorAsync += args => { log.LogError(args.Exception, "Service Bus error from {Source}", args.ErrorSource); return Task.CompletedTask; };
await processor.StartProcessingAsync(ct); try { await Task.Delay(Timeout.Infinite, ct); } catch (OperationCanceledException) { } await processor.StopProcessingAsync(CancellationToken.None); }
static PolicySnapshot Map(PolicySnapshot s, PolicyChanged e) { s.PolicyNumber = e.PolicyNumber; s.Status = e.Status; s.EffectiveFrom = e.EffectiveFrom; s.EffectiveTo = e.EffectiveTo; s.CoverageLimit = e.CoverageLimit; s.Version = e.Version; s.LastUpdatedUtc = e.OccurredUtc; return s; }}This requires adding public long Version { get; set; } to PolicySnapshot. Register with builder.Services.AddSingleton(new ServiceBusClient(fullyQualifiedNamespace, new DefaultAzureCredential())) and AddHostedService<PolicyChangedConsumer>().
Angular side: the frontend talks only to the Claims API, so the form never waits on other services. Using the current Angular major (standalone components, signals; confirm the exact version on angular.dev before teaching):
import { Component, inject, signal } from '@angular/core';import { HttpClient } from '@angular/common/http';import { FormBuilder, ReactiveFormsModule, Validators } from '@angular/forms';
@Component({ selector: 'app-submit-claim', imports: [ReactiveFormsModule], template: ` <form [formGroup]="form" (ngSubmit)="submit()"> <input formControlName="policyNumber" placeholder="Policy number" /> <input formControlName="lossDate" type="date" /> <input formControlName="amount" type="number" /> <button [disabled]="form.invalid || busy()">Submit claim</button> </form> @if (error()) { <p class="error">{{ error() }}</p> } `,})export class SubmitClaimComponent { private http = inject(HttpClient); private fb = inject(FormBuilder); busy = signal(false); error = signal<string | null>(null);
form = this.fb.nonNullable.group({ policyNumber: ['', Validators.required], lossDate: ['', Validators.required], amount: [0, [Validators.required, Validators.min(1)]], });
submit() { this.busy.set(true); this.http.post('/api/claims', this.form.getRawValue()).subscribe({ next: () => this.busy.set(false), error: e => { this.error.set(e.error ?? 'Submission failed'); this.busy.set(false); }, }); }}Database notes: the snapshot table is owned by ClaimsService and is never written by anyone else. Keep it narrow (only fields Claims needs), index on the lookup key, and store LastUpdatedUtc so you can measure staleness.
Level 3: Advanced
Performance and scalability
- Reads are a primary-key lookup on local storage, so latency is stable and independent of other services. Scale the service and its database as one unit.
- Snapshot tables stay small if you copy only required attributes. Copying entire aggregates is the most common way this pattern gets expensive.
- Replay throughput matters: when you add a new consumer or rebuild a snapshot, you need to reprocess history. Design consumers to handle bulk replay without hammering the database (batching,
MaxConcurrentCallstuning).
Consistency and staleness
- The local copy is eventually consistent. Decide, per rule, how stale is acceptable and write it down. Example: coverage limit lag of a few seconds is acceptable; a policy cancelled for fraud may need a stricter path.
- Track lag as a metric:
now - LastUpdatedUtcof the latest event applied, and Service Bus subscription active message count and oldest message age. - For decisions that are costly to get wrong, use a hybrid: decide locally, then re-verify asynchronously and compensate (a saga, Day 7) if the source of truth disagrees.
Security
- Events carry data across boundaries. Do not put personal data in events unless the consumer needs it, and classify what you replicate. A copy of customer data is a second place a breach or a GDPR erasure request must reach; plan for “forgotten” events.
- Use managed identity with Service Bus roles (
Azure Service Bus Data Receiverfor the consumer,Data Senderfor the publisher) rather than connection strings.
Failure modes
- Out-of-order or duplicate events: solved with a per-entity version and idempotent upserts (see Day 17).
- Poison messages: configure
MaxDeliveryCountand monitor the dead-letter queue; a malformed event must not block the subscription. - Missed events after an outage or a bug: provide a reconciliation job or an admin endpoint that rebuilds a snapshot from the owner’s API, used rarely and off the hot path.
- Cold start: a brand-new service has an empty snapshot. Seed it with a one-off bulk load before it takes traffic.
- Schema drift: version events, add fields only, and tolerate unknown fields.
Common mistakes
- Turning the local copy into a second source of truth by letting the consumer service edit it.
- Copying data “just in case”, then never using it.
- Hiding a synchronous call behind a cache with a short TTL and calling that self-contained. A cache miss still needs the remote service.
- Ignoring the staleness requirement until an audit asks about it.
Level 4: Expert and Architect view
Trade-offs of the alternatives:
| Approach | Availability of the request path | Latency | Data freshness | Complexity | Best fit |
|---|---|---|---|---|---|
| Synchronous call chain (RPI, Day 15) | Product of all dependencies | Sum of hops | Always current | Low at first, rises with depth | Low-volume, non-critical, must-be-current reads |
| Sync call plus Circuit Breaker, Retry, Fallback (Days 26 to 29) | Better, still coupled | Sum of hops | Current or degraded | Medium | Dependencies you cannot copy |
| API Composition (Day 10) | Depends on all composed services | Slowest branch (parallel) | Current | Medium | Read-only aggregate views |
| Self-contained service (local replica via events) | Own availability only | Local read | Eventually consistent | Medium-high (events, replay, reconciliation) | Hot paths, strict SLOs |
| CQRS read model (Day 8) | Own availability only | Local read | Eventually consistent | High | Rich query needs across services |
| Modular monolith | One process | In-process | Immediate | Low | Small teams, early stage |
Patterns it combines with:
- Database per Service (Day 5): the replica lives in the service’s own store.
- Domain Event (Day 11) and Transactional Outbox (Day 12): reliable way to feed the replica.
- Idempotent Consumer (Day 17): makes replay and duplicates safe.
- Saga (Day 7): compensates when a decision made on stale data proves wrong.
- CQRS (Day 8): a self-contained read side is CQRS applied to cross-service data.
ADR-style justification:
ADR-003: Claims submission will use a locally replicated policy snapshot
Status: Proposed
Context: Claim submission currently makes three synchronous calls (Policy, Customer, Billing). Measured p99 is 1.4 s, submission availability is 99.2% against a 99.9% target, and Billing incidents have caused claim outages.
Decision:
ClaimsServicewill keep a read-onlyPolicySnapshotpopulated fromPolicyChangedevents over Azure Service Bus (via the outbox inPolicyService). The submit-claim path will make no synchronous calls to other services.Consequences: (+) Submission availability and latency depend only on Claims. (+) Teams deploy independently. (-) Snapshot can lag; we accept up to 60 seconds and alert above 30. (-) We own replay, reconciliation, and dead-letter handling. (-) Policy data is duplicated and must be covered by retention and erasure processes.
Alternatives rejected: Circuit breaker only (still coupled, degrades to errors); API composition (does not help writes); merging services (loses team autonomy).
Azure implementation
Services that support the pattern:
- Azure Service Bus (topics and subscriptions) to carry domain events. Standard tier supports topics and is pay-per-use with a base charge; Premium gives dedicated capacity, higher message size limits, VNet integration and private endpoints. Choose Standard to start and Premium when you need isolation, predictable latency, or private networking. Confirm current pricing on the Azure pricing page before quoting numbers.
- Azure SQL Database or Azure Database for PostgreSQL (Flexible Server) as the service’s own datastore holding the snapshot. Azure Cosmos DB is an option when the snapshot store needs global distribution; serverless mode suits spiky or low traffic, provisioned throughput suits steady load.
- Azure Container Apps or AKS to host the API and the consumer. Container Apps can scale the consumer on Service Bus queue depth using a KEDA
azure-servicebusscaler, and can scale to zero on the consumption plan. - Azure Event Grid is an alternative for lightweight notification-style events, but Service Bus is the better default for business events that need ordering options, sessions, dead-lettering, and duplicate detection.
- Azure Monitor / Application Insights for traces, dependency maps, and lag metrics; Azure Key Vault for secrets; Microsoft Entra ID managed identities for access.
Configuration highlights:
- Create a Service Bus namespace, a topic
policy-events, and a subscriptionclaims-servicewithMaxDeliveryCount(for example 10), dead-lettering on message expiration, and a lock duration that exceeds your worst-case processing time. - Enable duplicate detection on the topic if the publisher sets a stable
MessageId(use the outbox row ID). - Assign
Azure Service Bus Data Senderto the Policy service identity andAzure Service Bus Data Receiverto the Claims service identity. - Set alerts on subscription active message count, dead-letter count, and oldest message age.
- Add a custom metric or log for snapshot age and chart it in a workbook.
Cost considerations: the main cost drivers are Service Bus tier and operations, database size and tier for each service’s own store, and container compute. Duplicating a small snapshot is cheap; replicating large datasets is where cost and operational burden grow.
Reference architecture (text): Client (Angular SPA) calls the API Gateway (Azure API Management or Application Gateway) which routes to ClaimsService running on Azure Container Apps. ClaimsService reads and writes only its Azure SQL database (tables: Claims, PolicySnapshots). PolicyService writes to its own database and an outbox table in one transaction; a publisher sends outbox rows to the Service Bus topic policy-events. A consumer hosted in the Claims app (scaled by KEDA on subscription depth) upserts PolicySnapshots. Application Insights receives traces from all components; alerts fire on lag and dead-letters. No arrow exists from ClaimsService to PolicyService on the submit path.
Teaching guide for my team
Two-minute explanation for a beginner: “When you order at a restaurant, the cook shouldn’t walk to the supplier to get onions for your dish. They keep onions at the station. A self-contained service does the same: it keeps a copy of the small bits of data it needs, updated by messages, so it can answer you without calling anyone else. If the other services are slow or down, this one still works. The price is that the copy can be a few seconds old.”
Five-minute explanation for an intermediate developer: Start with availability math (five services at 99.9% give about 99.5%). Show the trace of a chained call. Then show the alternative: the owner publishes PolicyChanged through an outbox; the consumer upserts a PolicySnapshot guarded by Version. Walk through the three questions every replica needs answered: how stale can it be, how do we rebuild it, and how do we handle duplicates and ordering. Finish with when not to use it: data that must be strictly current, or a system small enough to be a modular monolith.
Hands-on exercise: Build two small .NET 10 services, PolicyService and ClaimsService, with a Service Bus topic (or the local emulator, or RabbitMQ as a substitute). Implement POST /claims reading only from PolicySnapshots. Steps: (1) publish PolicyChanged for policy P-100 with version 1 and status Active; (2) submit a claim and confirm success; (3) stop PolicyService and submit another claim; (4) publish version 2 with status Lapsed, then version 1 again out of order. Expected outcome: claims succeed while PolicyService is down; after version 2 a claim is rejected; the stale version 1 replay does not revert the status; the duplicate delivery causes no error.
Interview questions:
- What does “self-contained” mean here? The service can complete its main request from its own code and data without blocking on a synchronous call to another service.
- How do you keep the local copy correct if events arrive twice or out of order? Use idempotent upserts and a per-entity version or timestamp; ignore events whose version is not newer.
- When would you not use it? When data must be strictly current at decision time, when the data is large or changes constantly, or when the system is small enough that a modular monolith is simpler.
Mastery checklist
- Can compute composite availability for a call chain and explain why depth hurts.
- Can identify a chained-call hot path in a distributed trace.
- Can define which data to replicate, and why only the minimum fields.
- Can implement an idempotent, version-guarded event consumer.
- Can state and defend an acceptable staleness bound for a business rule.
- Can design seeding, replay, and reconciliation for the replica.
- Can explain the privacy impact of replicating personal data and how erasure is handled.
- Can compare the pattern with circuit breaker, API composition, CQRS, and a modular monolith and choose with a written ADR.
Key takeaway
Move data to the request instead of moving the request to the data: replicate what you need through events so the hot path depends only on your own service, and pay for that with managed eventual consistency.
