A saga is a way to complete one business process that spans several services, each with its own database, without a distributed lock or a two-phase commit.
Intro
A saga is a way to complete one business process that spans several services, each with its own database, without a distributed lock or a two-phase commit. The process is broken into a sequence of small local transactions. Each local transaction commits on its own and then triggers the next one. If a later step fails, the saga runs compensating transactions that semantically undo the earlier steps. The system is never atomic across services, but it is eventually consistent and always ends in a known, valid state (completed or compensated).
Running example: an insurance claims system where “Submit Claim” touches Claims, Policy, Fraud, and Payment services.
Why we need this
After you adopt Database per Service (Day 5), a single business action no longer fits inside one database transaction. “Approve and pay a claim” now means: update the claim (Claims DB), reserve the coverage limit (Policy DB), record a fraud score (Fraud DB), and issue a payment (Payment DB).
- Business reason: Customers and adjusters expect “the claim was either fully processed or cleanly rolled back”. Half-processed claims mean money paid without a reserve, or a reserve held with no claim.
- Technical reason: Classic distributed transactions (2PC / XA / MSDTC) need every participant to support the protocol, hold locks across the network while waiting for the slowest participant, and stop working if the coordinator dies mid-commit. Most modern brokers, REST APIs, and cloud databases (Azure Service Bus, Cosmos DB, third-party payment gateways) simply do not participate in 2PC.
- Availability reason: With 2PC the availability of the whole operation is the product of every participant’s availability. Sagas let each service commit locally and move on.
What problem it solves
Problem (from the topic list): 2PC fails to scale across networks and locks distributed resources.
Without a saga, teams usually fall into one of these:
- Chained synchronous calls with no undo.
ClaimsServicecalls Policy, then Fraud, then Payment. Payment fails, and the Policy reserve is left dangling forever. - Dual writes. Update local DB and call the next service in the same method. Crash between the two and the data diverges.
- Shared database + big transaction. This re-couples services (undoing Day 5) and creates lock contention.
- Manual clean-up by support staff. A nightly SQL script “fixes” stuck claims. Nobody trusts the numbers.
A saga replaces all of these with an explicit, observable, recoverable workflow whose failure paths are designed up front.
When it is needed (and when it is NOT)
Needed when:
- One business transaction updates data owned by two or more services with separate datastores.
- Some steps are slow or external (payment gateway, third-party fraud API, human approval).
- You need “all or compensated” semantics but can tolerate temporary intermediate states (e.g., claim shows
PendingPayment). - Steps can be undone or neutralised in a business-meaningful way (release a reserve, void a payment authorisation, issue a refund).
NOT needed / wrong choice when:
- All the data lives in one service and one database. Use a normal local ACID transaction. A saga there is pure overhead.
- The business truly requires strict, immediate, all-or-nothing consistency across services (rare; e.g., some ledger postings). Then reconsider the service boundary, since it is probably drawn wrongly (see Day 1 and Day 2).
- A step cannot be compensated and cannot be moved to the end (e.g., an irreversible external e-mail or a wire transfer already settled). Design it as the last step (the “pivot”/“point of no return”) or add a human process.
- Only one remote call is involved and it is idempotent. A simple retry (Day 28) is enough.
- The team has no messaging, observability, or idempotency discipline yet. Sagas amplify those gaps.
How to identify the problem (key signals)
- Orphaned records: claims stuck in
Pendingfor days; reserves in Policy with no matching claim. - Reconciliation mismatches: Finance report shows payments that Claims does not know about (or the opposite).
- “Fix it in prod” scripts that manually update several databases to repair one business case.
- Try/catch blocks that call
Undo...methods inside controllers or handlers, with no persistence of what has already been done. - Dual-write code smell:
SaveChanges()followed byhttpClient.PostAsync(...)orbus.Publish(...)in the same method. - Latency and lock symptoms: long-held locks or
MSDTCtimeouts in SQL Server; deadlocks correlated with cross-service operations. - Support tickets like “my claim says approved but I was not paid” or “I was paid twice”.
Flow Diagram
Orchestrated claim-approval saga with compensations on failure.
flowchart TD S(["ClaimApprovalRequested"]) --> R["Policy: ReserveLimit"] R -->|LimitReserved| P["Payment: AuthorizePayment"] R -->|LimitReservationFailed| RJ["Claims: MarkClaimRejected"] P -->|PaymentAuthorized| PAID["Claims: MarkClaimPaid"] P -->|PaymentFailed| REL["Compensate: ReleaseLimit"] REL --> RJ PAID --> DONE(["Completed"]) RJ --> COMP(["Compensated"])Level 1: Beginner
Analogy: Booking a holiday: flight, hotel, car. You book them one by one. If the car is unavailable, you cancel the hotel and the flight (perhaps paying a small fee). Each booking is its own transaction, and each cancellation is a compensating action. You cannot “roll back” the airline’s database; you can only send a cancel request.
Key vocabulary:
| Term | Meaning |
|---|---|
| Local transaction | An ordinary ACID transaction inside one service |
| Compensating transaction | A new transaction that semantically reverses an earlier committed one |
| Choreography | Services react to each other’s events; no central controller |
| Orchestration | A central saga orchestrator tells each service what to do next |
| Pivot transaction | The step after which the saga can only go forward (no compensation) |
Minimal example (pseudo-code, orchestration style):
// PSEUDO-CODEApproveClaimSaga(claimId): step 1: Claims.MarkPendingPayment(claimId) compensate: Claims.Reopen(claimId) step 2: Policy.ReserveLimit(claimId, amount) compensate: Policy.ReleaseLimit(claimId) step 3: Payment.Authorize(claimId, amount) compensate: Payment.VoidAuthorization(claimId) step 4: Claims.MarkPaid(claimId) (no compensation needed)
run steps in orderif step N fails: run compensations for steps N-1 ... 1 in reverse order mark saga FailedMinimal working C# (in-process, to teach the idea, not for production):
public interface ISagaStep{ string Name { get; } Task ExecuteAsync(CancellationToken ct); Task CompensateAsync(CancellationToken ct);}
public sealed class SagaRunner{ public async Task<bool> RunAsync(IReadOnlyList<ISagaStep> steps, CancellationToken ct) { var done = new Stack<ISagaStep>(); foreach (var step in steps) { try { await step.ExecuteAsync(ct); done.Push(step); } catch (Exception ex) { Console.WriteLine($"Step {step.Name} failed: {ex.Message}. Compensating..."); while (done.Count > 0) await done.Pop().CompensateAsync(ct); return false; } } return true; }}Teaching note: this in-memory version loses everything if the process crashes. Levels 2 and 3 fix that by persisting saga state.
Level 2: Intermediate
Choreography vs orchestration
| Choreography | Orchestration | |
|---|---|---|
| Control | Distributed: each service reacts to events | Central: orchestrator sends commands |
| Coupling | Services know which events to listen to | Services only know their own commands; orchestrator knows the flow |
| Visibility | Hard: flow is implicit across services | Easy: state machine shows current step |
| Good for | 2 to 3 steps, simple flows | 4+ steps, branching, timeouts, compensation |
| Risk | Cyclic dependencies, “event spaghetti” | Orchestrator can become a god service if it holds business logic |
For the claims example (4 steps, timeouts, compensation) we use orchestration.
Orchestrated saga in .NET with MassTransit (state machine, EF Core, SQL Server)
Version note: MassTransit v8 is open source (Apache 2.0). MassTransit v9 moves to a commercial licence model. Check the current licence terms before choosing a version for a new project. Alternatives with saga support: NServiceBus (commercial), Wolverine (open source), Rebus (open source), or Azure Durable Functions.
State (persisted):
using MassTransit;
public class ApproveClaimState : SagaStateMachineInstance{ public Guid CorrelationId { get; set; } // = ClaimId public string CurrentState { get; set; } = default!; public decimal Amount { get; set; } public string? FailureReason { get; set; } public byte[] RowVersion { get; set; } = default!; // optimistic concurrency}Messages (records):
public record ClaimApprovalRequested(Guid ClaimId, decimal Amount);
public record ReserveLimit(Guid ClaimId, decimal Amount);public record LimitReserved(Guid ClaimId);public record LimitReservationFailed(Guid ClaimId, string Reason);public record ReleaseLimit(Guid ClaimId);
public record AuthorizePayment(Guid ClaimId, decimal Amount);public record PaymentAuthorized(Guid ClaimId);public record PaymentFailed(Guid ClaimId, string Reason);
public record MarkClaimPaid(Guid ClaimId);public record MarkClaimRejected(Guid ClaimId, string Reason);State machine:
public class ApproveClaimStateMachine : MassTransitStateMachine<ApproveClaimState>{ public State ReservingLimit { get; private set; } = default!; public State AuthorizingPayment { get; private set; } = default!; public State Completed { get; private set; } = default!; public State Compensated { get; private set; } = default!;
public Event<ClaimApprovalRequested> Requested { get; private set; } = default!; public Event<LimitReserved> LimitOk { get; private set; } = default!; public Event<LimitReservationFailed> LimitFailed { get; private set; } = default!; public Event<PaymentAuthorized> PaymentOk { get; private set; } = default!; public Event<PaymentFailed> PaymentNotOk { get; private set; } = default!;
public ApproveClaimStateMachine() { InstanceState(x => x.CurrentState);
Event(() => Requested, e => e.CorrelateById(m => m.Message.ClaimId)); Event(() => LimitOk, e => e.CorrelateById(m => m.Message.ClaimId)); Event(() => LimitFailed, e => e.CorrelateById(m => m.Message.ClaimId)); Event(() => PaymentOk, e => e.CorrelateById(m => m.Message.ClaimId)); Event(() => PaymentNotOk, e => e.CorrelateById(m => m.Message.ClaimId));
Initially( When(Requested) .Then(ctx => ctx.Saga.Amount = ctx.Message.Amount) .Send(new Uri("queue:policy-commands"), ctx => new ReserveLimit(ctx.Saga.CorrelationId, ctx.Saga.Amount)) .TransitionTo(ReservingLimit));
During(ReservingLimit, When(LimitOk) .Send(new Uri("queue:payment-commands"), ctx => new AuthorizePayment(ctx.Saga.CorrelationId, ctx.Saga.Amount)) .TransitionTo(AuthorizingPayment), When(LimitFailed) .Then(ctx => ctx.Saga.FailureReason = ctx.Message.Reason) .Send(new Uri("queue:claims-commands"), ctx => new MarkClaimRejected(ctx.Saga.CorrelationId, ctx.Message.Reason)) .TransitionTo(Compensated));
During(AuthorizingPayment, When(PaymentOk) .Send(new Uri("queue:claims-commands"), ctx => new MarkClaimPaid(ctx.Saga.CorrelationId)) .TransitionTo(Completed), When(PaymentNotOk) .Then(ctx => ctx.Saga.FailureReason = ctx.Message.Reason) // compensate step 1: release the reserved limit .Send(new Uri("queue:policy-commands"), ctx => new ReleaseLimit(ctx.Saga.CorrelationId)) .Send(new Uri("queue:claims-commands"), ctx => new MarkClaimRejected(ctx.Saga.CorrelationId, ctx.Message.Reason)) .TransitionTo(Compensated));
SetCompletedWhenFinalized(); }}Registration (Program.cs, .NET 10 minimal hosting):
builder.Services.AddDbContext<SagaDbContext>(o => o.UseSqlServer(builder.Configuration.GetConnectionString("SagaDb")));
builder.Services.AddMassTransit(x =>{ x.AddSagaStateMachine<ApproveClaimStateMachine, ApproveClaimState>() .EntityFrameworkRepository(r => { r.ConcurrencyMode = ConcurrencyMode.Optimistic; r.ExistingDbContext<SagaDbContext>(); });
x.UsingAzureServiceBus((ctx, cfg) => { cfg.Host(builder.Configuration["ServiceBus:Namespace"]); // uses DefaultAzureCredential / managed identity cfg.ConfigureEndpoints(ctx); });});Notes: (a) the SagaDbContext needs the saga map registered (ApproveClaimStateMap : SagaClassMap<ApproveClaimState>); (b) production setups should also enable the MassTransit EF Core outbox so the state change and the outgoing command commit atomically (see Day 12, Transactional Outbox).
Angular side
The saga is asynchronous, so the UI must not pretend the result is immediate. The API returns 202 Accepted with a status URL; Angular polls (or receives SignalR pushes).
// claim-status.service.ts (Angular standalone, signals + HttpClient)import { Injectable, inject, signal } from '@angular/core';import { HttpClient } from '@angular/common/http';import { timer, switchMap, takeWhile } from 'rxjs';
export type ClaimStatus = 'PendingPayment' | 'Paid' | 'Rejected';
@Injectable({ providedIn: 'root' })export class ClaimStatusService { private http = inject(HttpClient); status = signal<ClaimStatus>('PendingPayment');
approve(claimId: string, amount: number) { return this.http.post<{ statusUrl: string }>(`/api/claims/${claimId}/approve`, { amount }); }
watch(claimId: string) { timer(0, 2000) .pipe( switchMap(() => this.http.get<{ status: ClaimStatus }>(`/api/claims/${claimId}/status`)), takeWhile(r => r.status === 'PendingPayment', true) ) .subscribe(r => this.status.set(r.status)); }}API endpoint that starts the saga:
app.MapPost("/api/claims/{id:guid}/approve", async (Guid id, ApproveRequest req, IPublishEndpoint bus) =>{ await bus.Publish(new ClaimApprovalRequested(id, req.Amount)); return Results.Accepted($"/api/claims/{id}/status");});
public record ApproveRequest(decimal Amount);Compensations must be idempotent and business-meaningful
ReleaseLimit may be delivered twice. Handle it so the second call does nothing (see Day 17, Idempotent Consumer):
public class ReleaseLimitConsumer(PolicyDbContext db) : IConsumer<ReleaseLimit>{ public async Task Consume(ConsumeContext<ReleaseLimit> ctx) { var reserve = await db.Reserves .SingleOrDefaultAsync(r => r.ClaimId == ctx.Message.ClaimId && r.Status == ReserveStatus.Active);
if (reserve is null) return; // already released or never created: safe no-op
reserve.Status = ReserveStatus.Released; await db.SaveChangesAsync(); }}Level 3: Advanced
Isolation anomalies (the “I” in ACID is gone)
Sagas are ACD, not ACID. Between steps, other transactions can see intermediate data. Typical anomalies and countermeasures (from the “semantic lock” family of techniques):
| Anomaly | Example | Countermeasure |
|---|---|---|
| Lost update | Another process edits the claim while the saga is mid-flight | Semantic lock: set Status = PendingPayment so others reject or queue changes |
| Dirty read | Dashboard shows the reserve before the payment succeeded | Commutative updates; show “pending” states; read from saga-aware views |
| Non-repeatable read | Two reads of the same claim see different states | Versioning (RowVersion), reread by value |
Failure modes and how to design for them
- Duplicate messages (at-least-once delivery): every command handler and every saga event must be idempotent.
- Lost/stuck sagas: add timeouts. In MassTransit use
Schedule; in Durable Functions use durable timers. On timeout, either retry, compensate, or escalate to a human queue. - Compensation fails: retry with backoff; after N attempts move to a dead-letter queue and raise an alert. A failed compensation is an operational incident, not something to swallow.
- Out-of-order events: correlate on ClaimId and guard on state. Ignore events that are not valid in the current state (or log them).
- Concurrent events for the same saga: use optimistic concurrency (
RowVersion) or pessimistic locking; retry onDbUpdateConcurrencyException. - Orchestrator crash: because state is persisted after each transition, another instance resumes. This requires the outbox pattern so “state saved” and “command sent” are atomic.
- Non-compensable steps: order steps as compensable steps first, pivot next, retriable steps last. Example: reserve limit (compensable), authorise payment (pivot, can be voided until capture), capture/settle and notify (retriable, must eventually succeed).
Performance and scalability
- A saga instance is a small row; the cost is I/O per transition. Keep saga state minimal (IDs, amounts, current state), not whole documents.
- Partition by correlation ID (Service Bus sessions or partitioned queues) so one saga’s messages are processed in order and different sagas run in parallel.
- Index the saga table on
CorrelationId(primary key) and onCurrentStateif you query “all stuck sagas”. - Avoid orchestrator calls into other services synchronously; send commands and wait for events.
Security
- Authorise the initiator at the edge (API gateway / Access Token, Day 38) and carry the caller identity in message headers for audit (Day 34).
- Sign or authenticate messages: use managed identity and RBAC on the broker so only known services can send commands to
payment-commands. - Do not put sensitive data (card numbers, health details) in saga state or message bodies that are logged. Store references.
Common mistakes
- Using a saga where a local transaction would do.
- Putting business rules inside the orchestrator so it becomes a god service.
- Compensations that are not idempotent, or that assume the forward step fully succeeded.
- No timeouts, so sagas hang forever.
- No correlation ID in logs/traces, so nobody can debug a stuck saga (see Day 31, Distributed Tracing).
- Pretending compensation is a database rollback. It is a new business action (e.g., “void authorisation”, not “delete the row”).
Level 4: Expert and Architect view
Alternatives compared
| Approach | Consistency | Coupling | Complexity | Fits claims example? |
|---|---|---|---|---|
| Local ACID transaction | Strong | N/A (single service) | Lowest | Only if data is in one service |
| 2PC / MSDTC / XA | Strong (blocking) | High (all participants must support it) | Medium, fragile at scale | No: payment gateway and Service Bus cannot join |
| Saga, choreography | Eventual | Low to medium (event contracts) | Grows quickly with steps | OK for 2 to 3 simple steps |
| Saga, orchestration | Eventual | Low (commands), central flow | Medium, visible | Yes (chosen) |
| Try-Confirm/Cancel (TCC) | Eventual, reservation-based | Medium (every service exposes try/confirm/cancel) | High | Possible for limit reservation |
| Event-driven + reconciliation only | Eventually consistent, repaired later | Low | Low up front, high ops cost | Only for low-value, high-volume flows |
Patterns it combines with
- Database per Service (Day 5): the reason sagas exist.
- Transactional Outbox (Day 12): makes “save state + send message” atomic.
- Idempotent Consumer (Day 17): required for safe redelivery.
- Messaging (Day 16): the transport.
- Retry & Backoff (Day 28), Circuit Breaker (Day 26): protect saga steps that call flaky externals.
- Domain Event (Day 11): services emit events that choreographed sagas react to.
- Distributed Tracing (Day 31) and Audit Logging (Day 34): make sagas debuggable and defensible.
ADR (architecture review format)
ADR-007: Use an orchestrated saga for claim approval and payment
- Status: Proposed
- Context: Claim approval spans Claims, Policy, and Payment services, each with its own database. Payment goes through an external gateway. Distributed transactions are not available. Approval must end either as Paid or Rejected with the coverage reserve released.
- Decision: Implement
ApproveClaimSagaas an orchestrated state machine with persisted state in SQL Server, using commands and events over Azure Service Bus. Use the transactional outbox on the orchestrator and all participants. All handlers are idempotent. Steps are ordered: reserve limit (compensable), authorise payment (pivot), mark paid and notify (retriable). - Consequences (positive): no cross-service locks; clear visibility of each claim’s progress; independent service deployment; explicit failure handling.
- Consequences (negative): eventual consistency and intermediate states visible to users; more infrastructure (broker, saga store); need for timeouts, dead-letter handling, and monitoring; team must learn state-machine design.
- Alternatives rejected: 2PC (not supported by Payment gateway or Service Bus); choreography (too many steps and branches to follow implicitly); shared database (violates Day 5 goals).
- Review trigger: revisit if the saga grows beyond about 8 steps or if more than two teams must edit the same flow. Consider splitting into sub-sagas.
Azure implementation
Services that implement or support sagas
| Need | Azure service | Notes |
|---|---|---|
| Message transport (commands/events) | Azure Service Bus (Standard or Premium) | Queues and topics, sessions for ordered per-saga processing, dead-letter queues, duplicate detection, scheduled messages for timeouts |
| Orchestration as code | Azure Durable Functions (.NET isolated worker) | Orchestrator function with try/catch and compensation calls; built-in durable timers and checkpointing |
| Managed durable execution backend | Durable Task Scheduler | Managed backend for Durable Functions and the Durable Task SDKs; check current tiers and regions before choosing |
| Saga state store (state-machine libraries) | Azure SQL Database or Azure Database for PostgreSQL (flexible server) | Table per saga type, optimistic concurrency |
| Event fan-out (choreography, notifications) | Azure Event Grid | Good for reactive integration, less so for ordered command flows |
| Hosting participants | Azure Container Apps or AKS or App Service | Container Apps scale on Service Bus queue length via KEDA rules |
| Secrets/identity | Managed Identity + Azure Key Vault | No connection strings in code |
| Observability | Application Insights / Azure Monitor (OpenTelemetry) | Correlate by ClaimId/operation ID |
How to configure (key points)
- Service Bus namespace: create queues
policy-commands,payment-commands,claims-commands, and a topic for saga events. Enable dead-lettering on message expiration, setMaxDeliveryCount(for example 5 to 10), and turn on duplicate detection with a window matching your retry horizon. Use sessions if you need strict ordering per ClaimId. - Access: disable shared-key auth where possible. Grant
Azure Service Bus Data Sender/Data Receiverroles to each service’s managed identity. - Saga state DB: one schema owned by the orchestrator service only. Enable geo-redundant backup as appropriate. Use
rowversionfor concurrency. - Timeouts: use Service Bus scheduled messages (or MassTransit
Schedule, or Durable Functions timers) to fire aClaimApprovalTimeoutevent. - Autoscaling: in Container Apps add a KEDA
azure-servicebusscale rule on queue message count, with a cap that respects downstream limits. - Alerts: dead-letter queue length > 0, saga age > SLA, compensation failure count, and queue oldest-message age.
Durable Functions alternative (orchestration as code)
using Microsoft.Azure.Functions.Worker;using Microsoft.DurableTask;
public static class ApproveClaimOrchestrator{ [Function(nameof(ApproveClaimOrchestrator))] public static async Task<string> Run([OrchestrationTrigger] TaskOrchestrationContext ctx) { var input = ctx.GetInput<ApproveInput>()!; var compensations = new Stack<Func<Task>>();
try { await ctx.CallActivityAsync(nameof(Activities.ReserveLimit), input); compensations.Push(() => ctx.CallActivityAsync(nameof(Activities.ReleaseLimit), input));
await ctx.CallActivityAsync(nameof(Activities.AuthorizePayment), input); compensations.Push(() => ctx.CallActivityAsync(nameof(Activities.VoidPayment), input));
await ctx.CallActivityAsync(nameof(Activities.MarkClaimPaid), input); return "Paid"; } catch (TaskFailedException) { while (compensations.Count > 0) await compensations.Pop().Invoke();
await ctx.CallActivityAsync(nameof(Activities.MarkClaimRejected), input); return "Rejected"; } }}
public record ApproveInput(Guid ClaimId, decimal Amount);Orchestrator code must be deterministic: no DateTime.Now, random values, direct I/O, or Task.Delay. Use ctx.CurrentUtcDateTime, activities, and ctx.CreateTimer.
Pricing and tier considerations
Prices change and vary by region, so confirm on the Azure pricing pages before committing. The structure to remember:
- Service Bus Standard: pay-per-use (a small monthly base charge plus per-million-operations charges). Shared infrastructure, so throughput can vary. Good for dev/test and moderate production loads.
- Service Bus Premium: billed per messaging unit per hour with dedicated resources, predictable latency, larger message sizes, and VNet/private endpoint support. Choose it for regulated or high-throughput saga traffic, or when you need private networking.
- Durable Functions: you pay for the hosting plan (Consumption/Flex Consumption/Premium/Dedicated) plus storage or Durable Task Scheduler charges depending on the backend you choose. Lots of activity calls means lots of storage or scheduler operations, so model the cost per saga.
- Saga state DB: small footprint; a low-tier Azure SQL or PostgreSQL flexible server is usually enough. Watch IOPS if you have millions of daily sagas.
- Cost driver to watch: each saga step is at least two broker operations (command + event) plus a DB write. Multiply by steps and volume.
Reference architecture (text)
- The Angular app calls the API Gateway (Azure API Management or YARP on Container Apps) with an Entra ID token.
Claims API(Container App, .NET 10) validates the request, stores the claim asPendingPaymentvia the outbox, and publishesClaimApprovalRequestedto Service Bus. It returns202 Acceptedwith a status URL.Approval Orchestrator(Container App) hosts the MassTransit state machine, persisting state in Azure SQL. It sendsReserveLimittopolicy-commands.Policy Servicereserves the limit in its own Azure SQL/PostgreSQL DB and emitsLimitReservedorLimitReservationFailed.- The orchestrator sends
AuthorizePaymenttopayment-commands.Payment Servicecalls the external gateway (with Retry and Circuit Breaker) and emitsPaymentAuthorizedorPaymentFailed. - On success the orchestrator sends
MarkClaimPaid. On failure it sendsReleaseLimitandMarkClaimRejected. - All services send OpenTelemetry traces and logs to Application Insights, keyed by ClaimId. Alerts fire on dead-letter messages and stuck sagas.
- Managed identities and Key Vault secure access; private endpoints protect Service Bus and databases in the Premium setup.
Teaching guide for my team
Explain to a beginner in 2 minutes
“When you book a trip, you book the flight, then the hotel, then the car. If the car fails, you cancel the hotel and the flight. Nobody can press one big undo button because each company has its own system. A saga is exactly that: a list of small steps, and for each step, a defined ‘cancel’ action. If something goes wrong halfway, we run the cancels in reverse. We also write down which step we are on, so if our app crashes we can carry on where we left off.”
Explain to an intermediate developer in 5 minutes
- Database per Service means no single transaction across services, and 2PC does not scale or work with most cloud services.
- So we split the business process into local transactions and connect them with messages.
- Two styles: choreography (services react to events) and orchestration (a state machine sends commands). We prefer orchestration when there are more than about three steps.
- Every step has a compensation. Compensations are business actions (release reserve, void payment), not database rollbacks, and must be idempotent.
- The saga’s state is persisted, and state changes and outgoing messages must be atomic (outbox).
- Isolation is lost: other users may see intermediate states, so we model them explicitly (
PendingPayment) and use semantic locks. - We add timeouts, dead-letter handling, and correlation IDs so we can operate it.
Hands-on exercise
Task: Extend the in-memory SagaRunner from Level 1 into a persisted mini-saga for “Approve Claim” with three steps (ReserveLimit, AuthorizePayment, MarkClaimPaid).
- Create a
SagaInstancetable (Id,CurrentStep,Status,FailureReason,RowVersion) with EF Core. - Persist the state after each step.
- Simulate a failure in
AuthorizePayment(throw for amounts above 10,000). - Simulate a crash by killing the process after step 1, restart it, and have it resume from the stored
CurrentStep. - Make
ReleaseLimitidempotent and call it twice in a test.
Expected outcome: Amount 5,000 finishes as Completed. Amount 20,000 ends as Compensated with the reserve released exactly once. Killing the process mid-saga and restarting completes or compensates the saga instead of leaving it stuck. Calling ReleaseLimit twice leaves the data unchanged the second time.
Interview-style questions
- Why can’t we just use a distributed transaction (2PC) across microservices? It blocks and holds locks across the network, needs every participant (databases, brokers, external APIs) to support the protocol, and reduces availability because a coordinator or slow participant stalls everyone. Many cloud services and third-party APIs do not support it at all.
- What is the difference between choreography and orchestration, and when do you pick each? Choreography has services react to each other’s events with no central controller, which suits short flows of two to three steps. Orchestration uses a central state machine that sends commands and tracks progress, which suits longer flows with branching, timeouts, and compensation because the flow is visible in one place.
- What is a compensating transaction and what makes a good one? It is a new local transaction that semantically undoes a previously committed step (for example, releasing a reserve or voiding a payment authorisation). A good one is idempotent, retriable, tolerant of the forward step having partially completed, and expressed in business terms rather than as a data rollback.
Mastery checklist
- I can explain why 2PC is a poor fit for microservices and cloud services.
- I can draw the saga for a real workflow with steps, compensations, and the pivot transaction.
- I can choose between choreography and orchestration and justify it.
- I can implement an idempotent compensating handler and prove it with a test.
- I can explain the isolation anomalies of sagas and apply at least one countermeasure (semantic lock).
- I can make “save state + send message” atomic using the outbox pattern.
- I can design timeouts, dead-letter handling, and alerts for stuck sagas.
- I can produce an ADR comparing saga, 2PC, and other alternatives for a given scenario.
Key takeaway
A saga replaces one impossible distributed transaction with a sequence of local transactions plus designed-in compensations. It buys availability and independence at the cost of eventual consistency, so idempotency, persisted state, and observability are not optional.
