Manikandan — Manikandan
Microservices

Day 7: Saga Pattern

ManikandanManikandan
20 min read·Updated Aug 16, 2022

A saga is a way to complete one business process that spans several services, each with its own database, without a distributed lock or a two-phase commit.

Intro

A saga is a way to complete one business process that spans several services, each with its own database, without a distributed lock or a two-phase commit. The process is broken into a sequence of small local transactions. Each local transaction commits on its own and then triggers the next one. If a later step fails, the saga runs compensating transactions that semantically undo the earlier steps. The system is never atomic across services, but it is eventually consistent and always ends in a known, valid state (completed or compensated).

Running example: an insurance claims system where “Submit Claim” touches Claims, Policy, Fraud, and Payment services.


Why we need this

After you adopt Database per Service (Day 5), a single business action no longer fits inside one database transaction. “Approve and pay a claim” now means: update the claim (Claims DB), reserve the coverage limit (Policy DB), record a fraud score (Fraud DB), and issue a payment (Payment DB).

  • Business reason: Customers and adjusters expect “the claim was either fully processed or cleanly rolled back”. Half-processed claims mean money paid without a reserve, or a reserve held with no claim.
  • Technical reason: Classic distributed transactions (2PC / XA / MSDTC) need every participant to support the protocol, hold locks across the network while waiting for the slowest participant, and stop working if the coordinator dies mid-commit. Most modern brokers, REST APIs, and cloud databases (Azure Service Bus, Cosmos DB, third-party payment gateways) simply do not participate in 2PC.
  • Availability reason: With 2PC the availability of the whole operation is the product of every participant’s availability. Sagas let each service commit locally and move on.

What problem it solves

Problem (from the topic list): 2PC fails to scale across networks and locks distributed resources.

Without a saga, teams usually fall into one of these:

  1. Chained synchronous calls with no undo. ClaimsService calls Policy, then Fraud, then Payment. Payment fails, and the Policy reserve is left dangling forever.
  2. Dual writes. Update local DB and call the next service in the same method. Crash between the two and the data diverges.
  3. Shared database + big transaction. This re-couples services (undoing Day 5) and creates lock contention.
  4. Manual clean-up by support staff. A nightly SQL script “fixes” stuck claims. Nobody trusts the numbers.

A saga replaces all of these with an explicit, observable, recoverable workflow whose failure paths are designed up front.

When it is needed (and when it is NOT)

Needed when:

  • One business transaction updates data owned by two or more services with separate datastores.
  • Some steps are slow or external (payment gateway, third-party fraud API, human approval).
  • You need “all or compensated” semantics but can tolerate temporary intermediate states (e.g., claim shows PendingPayment).
  • Steps can be undone or neutralised in a business-meaningful way (release a reserve, void a payment authorisation, issue a refund).

NOT needed / wrong choice when:

  • All the data lives in one service and one database. Use a normal local ACID transaction. A saga there is pure overhead.
  • The business truly requires strict, immediate, all-or-nothing consistency across services (rare; e.g., some ledger postings). Then reconsider the service boundary, since it is probably drawn wrongly (see Day 1 and Day 2).
  • A step cannot be compensated and cannot be moved to the end (e.g., an irreversible external e-mail or a wire transfer already settled). Design it as the last step (the “pivot”/“point of no return”) or add a human process.
  • Only one remote call is involved and it is idempotent. A simple retry (Day 28) is enough.
  • The team has no messaging, observability, or idempotency discipline yet. Sagas amplify those gaps.

How to identify the problem (key signals)

  1. Orphaned records: claims stuck in Pending for days; reserves in Policy with no matching claim.
  2. Reconciliation mismatches: Finance report shows payments that Claims does not know about (or the opposite).
  3. “Fix it in prod” scripts that manually update several databases to repair one business case.
  4. Try/catch blocks that call Undo... methods inside controllers or handlers, with no persistence of what has already been done.
  5. Dual-write code smell: SaveChanges() followed by httpClient.PostAsync(...) or bus.Publish(...) in the same method.
  6. Latency and lock symptoms: long-held locks or MSDTC timeouts in SQL Server; deadlocks correlated with cross-service operations.
  7. Support tickets like “my claim says approved but I was not paid” or “I was paid twice”.

Flow Diagram

Orchestrated claim-approval saga with compensations on failure.

flowchart TD
S(["ClaimApprovalRequested"]) --> R["Policy: ReserveLimit"]
R -->|LimitReserved| P["Payment: AuthorizePayment"]
R -->|LimitReservationFailed| RJ["Claims: MarkClaimRejected"]
P -->|PaymentAuthorized| PAID["Claims: MarkClaimPaid"]
P -->|PaymentFailed| REL["Compensate: ReleaseLimit"]
REL --> RJ
PAID --> DONE(["Completed"])
RJ --> COMP(["Compensated"])

Level 1: Beginner

Analogy: Booking a holiday: flight, hotel, car. You book them one by one. If the car is unavailable, you cancel the hotel and the flight (perhaps paying a small fee). Each booking is its own transaction, and each cancellation is a compensating action. You cannot “roll back” the airline’s database; you can only send a cancel request.

Key vocabulary:

TermMeaning
Local transactionAn ordinary ACID transaction inside one service
Compensating transactionA new transaction that semantically reverses an earlier committed one
ChoreographyServices react to each other’s events; no central controller
OrchestrationA central saga orchestrator tells each service what to do next
Pivot transactionThe step after which the saga can only go forward (no compensation)

Minimal example (pseudo-code, orchestration style):

// PSEUDO-CODE
ApproveClaimSaga(claimId):
step 1: Claims.MarkPendingPayment(claimId) compensate: Claims.Reopen(claimId)
step 2: Policy.ReserveLimit(claimId, amount) compensate: Policy.ReleaseLimit(claimId)
step 3: Payment.Authorize(claimId, amount) compensate: Payment.VoidAuthorization(claimId)
step 4: Claims.MarkPaid(claimId) (no compensation needed)
run steps in order
if step N fails:
run compensations for steps N-1 ... 1 in reverse order
mark saga Failed

Minimal working C# (in-process, to teach the idea, not for production):

public interface ISagaStep
{
string Name { get; }
Task ExecuteAsync(CancellationToken ct);
Task CompensateAsync(CancellationToken ct);
}
public sealed class SagaRunner
{
public async Task<bool> RunAsync(IReadOnlyList<ISagaStep> steps, CancellationToken ct)
{
var done = new Stack<ISagaStep>();
foreach (var step in steps)
{
try
{
await step.ExecuteAsync(ct);
done.Push(step);
}
catch (Exception ex)
{
Console.WriteLine($"Step {step.Name} failed: {ex.Message}. Compensating...");
while (done.Count > 0)
await done.Pop().CompensateAsync(ct);
return false;
}
}
return true;
}
}

Teaching note: this in-memory version loses everything if the process crashes. Levels 2 and 3 fix that by persisting saga state.

Level 2: Intermediate

Choreography vs orchestration

ChoreographyOrchestration
ControlDistributed: each service reacts to eventsCentral: orchestrator sends commands
CouplingServices know which events to listen toServices only know their own commands; orchestrator knows the flow
VisibilityHard: flow is implicit across servicesEasy: state machine shows current step
Good for2 to 3 steps, simple flows4+ steps, branching, timeouts, compensation
RiskCyclic dependencies, “event spaghetti”Orchestrator can become a god service if it holds business logic

For the claims example (4 steps, timeouts, compensation) we use orchestration.

Orchestrated saga in .NET with MassTransit (state machine, EF Core, SQL Server)

Version note: MassTransit v8 is open source (Apache 2.0). MassTransit v9 moves to a commercial licence model. Check the current licence terms before choosing a version for a new project. Alternatives with saga support: NServiceBus (commercial), Wolverine (open source), Rebus (open source), or Azure Durable Functions.

State (persisted):

using MassTransit;
public class ApproveClaimState : SagaStateMachineInstance
{
public Guid CorrelationId { get; set; } // = ClaimId
public string CurrentState { get; set; } = default!;
public decimal Amount { get; set; }
public string? FailureReason { get; set; }
public byte[] RowVersion { get; set; } = default!; // optimistic concurrency
}

Messages (records):

public record ClaimApprovalRequested(Guid ClaimId, decimal Amount);
public record ReserveLimit(Guid ClaimId, decimal Amount);
public record LimitReserved(Guid ClaimId);
public record LimitReservationFailed(Guid ClaimId, string Reason);
public record ReleaseLimit(Guid ClaimId);
public record AuthorizePayment(Guid ClaimId, decimal Amount);
public record PaymentAuthorized(Guid ClaimId);
public record PaymentFailed(Guid ClaimId, string Reason);
public record MarkClaimPaid(Guid ClaimId);
public record MarkClaimRejected(Guid ClaimId, string Reason);

State machine:

public class ApproveClaimStateMachine : MassTransitStateMachine<ApproveClaimState>
{
public State ReservingLimit { get; private set; } = default!;
public State AuthorizingPayment { get; private set; } = default!;
public State Completed { get; private set; } = default!;
public State Compensated { get; private set; } = default!;
public Event<ClaimApprovalRequested> Requested { get; private set; } = default!;
public Event<LimitReserved> LimitOk { get; private set; } = default!;
public Event<LimitReservationFailed> LimitFailed { get; private set; } = default!;
public Event<PaymentAuthorized> PaymentOk { get; private set; } = default!;
public Event<PaymentFailed> PaymentNotOk { get; private set; } = default!;
public ApproveClaimStateMachine()
{
InstanceState(x => x.CurrentState);
Event(() => Requested, e => e.CorrelateById(m => m.Message.ClaimId));
Event(() => LimitOk, e => e.CorrelateById(m => m.Message.ClaimId));
Event(() => LimitFailed, e => e.CorrelateById(m => m.Message.ClaimId));
Event(() => PaymentOk, e => e.CorrelateById(m => m.Message.ClaimId));
Event(() => PaymentNotOk, e => e.CorrelateById(m => m.Message.ClaimId));
Initially(
When(Requested)
.Then(ctx => ctx.Saga.Amount = ctx.Message.Amount)
.Send(new Uri("queue:policy-commands"),
ctx => new ReserveLimit(ctx.Saga.CorrelationId, ctx.Saga.Amount))
.TransitionTo(ReservingLimit));
During(ReservingLimit,
When(LimitOk)
.Send(new Uri("queue:payment-commands"),
ctx => new AuthorizePayment(ctx.Saga.CorrelationId, ctx.Saga.Amount))
.TransitionTo(AuthorizingPayment),
When(LimitFailed)
.Then(ctx => ctx.Saga.FailureReason = ctx.Message.Reason)
.Send(new Uri("queue:claims-commands"),
ctx => new MarkClaimRejected(ctx.Saga.CorrelationId, ctx.Message.Reason))
.TransitionTo(Compensated));
During(AuthorizingPayment,
When(PaymentOk)
.Send(new Uri("queue:claims-commands"),
ctx => new MarkClaimPaid(ctx.Saga.CorrelationId))
.TransitionTo(Completed),
When(PaymentNotOk)
.Then(ctx => ctx.Saga.FailureReason = ctx.Message.Reason)
// compensate step 1: release the reserved limit
.Send(new Uri("queue:policy-commands"),
ctx => new ReleaseLimit(ctx.Saga.CorrelationId))
.Send(new Uri("queue:claims-commands"),
ctx => new MarkClaimRejected(ctx.Saga.CorrelationId, ctx.Message.Reason))
.TransitionTo(Compensated));
SetCompletedWhenFinalized();
}
}

Registration (Program.cs, .NET 10 minimal hosting):

builder.Services.AddDbContext<SagaDbContext>(o =>
o.UseSqlServer(builder.Configuration.GetConnectionString("SagaDb")));
builder.Services.AddMassTransit(x =>
{
x.AddSagaStateMachine<ApproveClaimStateMachine, ApproveClaimState>()
.EntityFrameworkRepository(r =>
{
r.ConcurrencyMode = ConcurrencyMode.Optimistic;
r.ExistingDbContext<SagaDbContext>();
});
x.UsingAzureServiceBus((ctx, cfg) =>
{
cfg.Host(builder.Configuration["ServiceBus:Namespace"]); // uses DefaultAzureCredential / managed identity
cfg.ConfigureEndpoints(ctx);
});
});

Notes: (a) the SagaDbContext needs the saga map registered (ApproveClaimStateMap : SagaClassMap<ApproveClaimState>); (b) production setups should also enable the MassTransit EF Core outbox so the state change and the outgoing command commit atomically (see Day 12, Transactional Outbox).

Angular side

The saga is asynchronous, so the UI must not pretend the result is immediate. The API returns 202 Accepted with a status URL; Angular polls (or receives SignalR pushes).

// claim-status.service.ts (Angular standalone, signals + HttpClient)
import { Injectable, inject, signal } from '@angular/core';
import { HttpClient } from '@angular/common/http';
import { timer, switchMap, takeWhile } from 'rxjs';
export type ClaimStatus = 'PendingPayment' | 'Paid' | 'Rejected';
@Injectable({ providedIn: 'root' })
export class ClaimStatusService {
private http = inject(HttpClient);
status = signal<ClaimStatus>('PendingPayment');
approve(claimId: string, amount: number) {
return this.http.post<{ statusUrl: string }>(`/api/claims/${claimId}/approve`, { amount });
}
watch(claimId: string) {
timer(0, 2000)
.pipe(
switchMap(() => this.http.get<{ status: ClaimStatus }>(`/api/claims/${claimId}/status`)),
takeWhile(r => r.status === 'PendingPayment', true)
)
.subscribe(r => this.status.set(r.status));
}
}

API endpoint that starts the saga:

app.MapPost("/api/claims/{id:guid}/approve",
async (Guid id, ApproveRequest req, IPublishEndpoint bus) =>
{
await bus.Publish(new ClaimApprovalRequested(id, req.Amount));
return Results.Accepted($"/api/claims/{id}/status");
});
public record ApproveRequest(decimal Amount);

Compensations must be idempotent and business-meaningful

ReleaseLimit may be delivered twice. Handle it so the second call does nothing (see Day 17, Idempotent Consumer):

public class ReleaseLimitConsumer(PolicyDbContext db) : IConsumer<ReleaseLimit>
{
public async Task Consume(ConsumeContext<ReleaseLimit> ctx)
{
var reserve = await db.Reserves
.SingleOrDefaultAsync(r => r.ClaimId == ctx.Message.ClaimId && r.Status == ReserveStatus.Active);
if (reserve is null) return; // already released or never created: safe no-op
reserve.Status = ReserveStatus.Released;
await db.SaveChangesAsync();
}
}

Level 3: Advanced

Isolation anomalies (the “I” in ACID is gone)

Sagas are ACD, not ACID. Between steps, other transactions can see intermediate data. Typical anomalies and countermeasures (from the “semantic lock” family of techniques):

AnomalyExampleCountermeasure
Lost updateAnother process edits the claim while the saga is mid-flightSemantic lock: set Status = PendingPayment so others reject or queue changes
Dirty readDashboard shows the reserve before the payment succeededCommutative updates; show “pending” states; read from saga-aware views
Non-repeatable readTwo reads of the same claim see different statesVersioning (RowVersion), reread by value

Failure modes and how to design for them

  1. Duplicate messages (at-least-once delivery): every command handler and every saga event must be idempotent.
  2. Lost/stuck sagas: add timeouts. In MassTransit use Schedule; in Durable Functions use durable timers. On timeout, either retry, compensate, or escalate to a human queue.
  3. Compensation fails: retry with backoff; after N attempts move to a dead-letter queue and raise an alert. A failed compensation is an operational incident, not something to swallow.
  4. Out-of-order events: correlate on ClaimId and guard on state. Ignore events that are not valid in the current state (or log them).
  5. Concurrent events for the same saga: use optimistic concurrency (RowVersion) or pessimistic locking; retry on DbUpdateConcurrencyException.
  6. Orchestrator crash: because state is persisted after each transition, another instance resumes. This requires the outbox pattern so “state saved” and “command sent” are atomic.
  7. Non-compensable steps: order steps as compensable steps first, pivot next, retriable steps last. Example: reserve limit (compensable), authorise payment (pivot, can be voided until capture), capture/settle and notify (retriable, must eventually succeed).

Performance and scalability

  • A saga instance is a small row; the cost is I/O per transition. Keep saga state minimal (IDs, amounts, current state), not whole documents.
  • Partition by correlation ID (Service Bus sessions or partitioned queues) so one saga’s messages are processed in order and different sagas run in parallel.
  • Index the saga table on CorrelationId (primary key) and on CurrentState if you query “all stuck sagas”.
  • Avoid orchestrator calls into other services synchronously; send commands and wait for events.

Security

  • Authorise the initiator at the edge (API gateway / Access Token, Day 38) and carry the caller identity in message headers for audit (Day 34).
  • Sign or authenticate messages: use managed identity and RBAC on the broker so only known services can send commands to payment-commands.
  • Do not put sensitive data (card numbers, health details) in saga state or message bodies that are logged. Store references.

Common mistakes

  • Using a saga where a local transaction would do.
  • Putting business rules inside the orchestrator so it becomes a god service.
  • Compensations that are not idempotent, or that assume the forward step fully succeeded.
  • No timeouts, so sagas hang forever.
  • No correlation ID in logs/traces, so nobody can debug a stuck saga (see Day 31, Distributed Tracing).
  • Pretending compensation is a database rollback. It is a new business action (e.g., “void authorisation”, not “delete the row”).

Level 4: Expert and Architect view

Alternatives compared

ApproachConsistencyCouplingComplexityFits claims example?
Local ACID transactionStrongN/A (single service)LowestOnly if data is in one service
2PC / MSDTC / XAStrong (blocking)High (all participants must support it)Medium, fragile at scaleNo: payment gateway and Service Bus cannot join
Saga, choreographyEventualLow to medium (event contracts)Grows quickly with stepsOK for 2 to 3 simple steps
Saga, orchestrationEventualLow (commands), central flowMedium, visibleYes (chosen)
Try-Confirm/Cancel (TCC)Eventual, reservation-basedMedium (every service exposes try/confirm/cancel)HighPossible for limit reservation
Event-driven + reconciliation onlyEventually consistent, repaired laterLowLow up front, high ops costOnly for low-value, high-volume flows

Patterns it combines with

  • Database per Service (Day 5): the reason sagas exist.
  • Transactional Outbox (Day 12): makes “save state + send message” atomic.
  • Idempotent Consumer (Day 17): required for safe redelivery.
  • Messaging (Day 16): the transport.
  • Retry & Backoff (Day 28), Circuit Breaker (Day 26): protect saga steps that call flaky externals.
  • Domain Event (Day 11): services emit events that choreographed sagas react to.
  • Distributed Tracing (Day 31) and Audit Logging (Day 34): make sagas debuggable and defensible.

ADR (architecture review format)

ADR-007: Use an orchestrated saga for claim approval and payment

  • Status: Proposed
  • Context: Claim approval spans Claims, Policy, and Payment services, each with its own database. Payment goes through an external gateway. Distributed transactions are not available. Approval must end either as Paid or Rejected with the coverage reserve released.
  • Decision: Implement ApproveClaimSaga as an orchestrated state machine with persisted state in SQL Server, using commands and events over Azure Service Bus. Use the transactional outbox on the orchestrator and all participants. All handlers are idempotent. Steps are ordered: reserve limit (compensable), authorise payment (pivot), mark paid and notify (retriable).
  • Consequences (positive): no cross-service locks; clear visibility of each claim’s progress; independent service deployment; explicit failure handling.
  • Consequences (negative): eventual consistency and intermediate states visible to users; more infrastructure (broker, saga store); need for timeouts, dead-letter handling, and monitoring; team must learn state-machine design.
  • Alternatives rejected: 2PC (not supported by Payment gateway or Service Bus); choreography (too many steps and branches to follow implicitly); shared database (violates Day 5 goals).
  • Review trigger: revisit if the saga grows beyond about 8 steps or if more than two teams must edit the same flow. Consider splitting into sub-sagas.

Azure implementation

Services that implement or support sagas

NeedAzure serviceNotes
Message transport (commands/events)Azure Service Bus (Standard or Premium)Queues and topics, sessions for ordered per-saga processing, dead-letter queues, duplicate detection, scheduled messages for timeouts
Orchestration as codeAzure Durable Functions (.NET isolated worker)Orchestrator function with try/catch and compensation calls; built-in durable timers and checkpointing
Managed durable execution backendDurable Task SchedulerManaged backend for Durable Functions and the Durable Task SDKs; check current tiers and regions before choosing
Saga state store (state-machine libraries)Azure SQL Database or Azure Database for PostgreSQL (flexible server)Table per saga type, optimistic concurrency
Event fan-out (choreography, notifications)Azure Event GridGood for reactive integration, less so for ordered command flows
Hosting participantsAzure Container Apps or AKS or App ServiceContainer Apps scale on Service Bus queue length via KEDA rules
Secrets/identityManaged Identity + Azure Key VaultNo connection strings in code
ObservabilityApplication Insights / Azure Monitor (OpenTelemetry)Correlate by ClaimId/operation ID

How to configure (key points)

  1. Service Bus namespace: create queues policy-commands, payment-commands, claims-commands, and a topic for saga events. Enable dead-lettering on message expiration, set MaxDeliveryCount (for example 5 to 10), and turn on duplicate detection with a window matching your retry horizon. Use sessions if you need strict ordering per ClaimId.
  2. Access: disable shared-key auth where possible. Grant Azure Service Bus Data Sender / Data Receiver roles to each service’s managed identity.
  3. Saga state DB: one schema owned by the orchestrator service only. Enable geo-redundant backup as appropriate. Use rowversion for concurrency.
  4. Timeouts: use Service Bus scheduled messages (or MassTransit Schedule, or Durable Functions timers) to fire a ClaimApprovalTimeout event.
  5. Autoscaling: in Container Apps add a KEDA azure-servicebus scale rule on queue message count, with a cap that respects downstream limits.
  6. Alerts: dead-letter queue length > 0, saga age > SLA, compensation failure count, and queue oldest-message age.

Durable Functions alternative (orchestration as code)

using Microsoft.Azure.Functions.Worker;
using Microsoft.DurableTask;
public static class ApproveClaimOrchestrator
{
[Function(nameof(ApproveClaimOrchestrator))]
public static async Task<string> Run([OrchestrationTrigger] TaskOrchestrationContext ctx)
{
var input = ctx.GetInput<ApproveInput>()!;
var compensations = new Stack<Func<Task>>();
try
{
await ctx.CallActivityAsync(nameof(Activities.ReserveLimit), input);
compensations.Push(() => ctx.CallActivityAsync(nameof(Activities.ReleaseLimit), input));
await ctx.CallActivityAsync(nameof(Activities.AuthorizePayment), input);
compensations.Push(() => ctx.CallActivityAsync(nameof(Activities.VoidPayment), input));
await ctx.CallActivityAsync(nameof(Activities.MarkClaimPaid), input);
return "Paid";
}
catch (TaskFailedException)
{
while (compensations.Count > 0)
await compensations.Pop().Invoke();
await ctx.CallActivityAsync(nameof(Activities.MarkClaimRejected), input);
return "Rejected";
}
}
}
public record ApproveInput(Guid ClaimId, decimal Amount);

Orchestrator code must be deterministic: no DateTime.Now, random values, direct I/O, or Task.Delay. Use ctx.CurrentUtcDateTime, activities, and ctx.CreateTimer.

Pricing and tier considerations

Prices change and vary by region, so confirm on the Azure pricing pages before committing. The structure to remember:

  • Service Bus Standard: pay-per-use (a small monthly base charge plus per-million-operations charges). Shared infrastructure, so throughput can vary. Good for dev/test and moderate production loads.
  • Service Bus Premium: billed per messaging unit per hour with dedicated resources, predictable latency, larger message sizes, and VNet/private endpoint support. Choose it for regulated or high-throughput saga traffic, or when you need private networking.
  • Durable Functions: you pay for the hosting plan (Consumption/Flex Consumption/Premium/Dedicated) plus storage or Durable Task Scheduler charges depending on the backend you choose. Lots of activity calls means lots of storage or scheduler operations, so model the cost per saga.
  • Saga state DB: small footprint; a low-tier Azure SQL or PostgreSQL flexible server is usually enough. Watch IOPS if you have millions of daily sagas.
  • Cost driver to watch: each saga step is at least two broker operations (command + event) plus a DB write. Multiply by steps and volume.

Reference architecture (text)

  1. The Angular app calls the API Gateway (Azure API Management or YARP on Container Apps) with an Entra ID token.
  2. Claims API (Container App, .NET 10) validates the request, stores the claim as PendingPayment via the outbox, and publishes ClaimApprovalRequested to Service Bus. It returns 202 Accepted with a status URL.
  3. Approval Orchestrator (Container App) hosts the MassTransit state machine, persisting state in Azure SQL. It sends ReserveLimit to policy-commands.
  4. Policy Service reserves the limit in its own Azure SQL/PostgreSQL DB and emits LimitReserved or LimitReservationFailed.
  5. The orchestrator sends AuthorizePayment to payment-commands. Payment Service calls the external gateway (with Retry and Circuit Breaker) and emits PaymentAuthorized or PaymentFailed.
  6. On success the orchestrator sends MarkClaimPaid. On failure it sends ReleaseLimit and MarkClaimRejected.
  7. All services send OpenTelemetry traces and logs to Application Insights, keyed by ClaimId. Alerts fire on dead-letter messages and stuck sagas.
  8. Managed identities and Key Vault secure access; private endpoints protect Service Bus and databases in the Premium setup.

Teaching guide for my team

Explain to a beginner in 2 minutes

“When you book a trip, you book the flight, then the hotel, then the car. If the car fails, you cancel the hotel and the flight. Nobody can press one big undo button because each company has its own system. A saga is exactly that: a list of small steps, and for each step, a defined ‘cancel’ action. If something goes wrong halfway, we run the cancels in reverse. We also write down which step we are on, so if our app crashes we can carry on where we left off.”

Explain to an intermediate developer in 5 minutes

  1. Database per Service means no single transaction across services, and 2PC does not scale or work with most cloud services.
  2. So we split the business process into local transactions and connect them with messages.
  3. Two styles: choreography (services react to events) and orchestration (a state machine sends commands). We prefer orchestration when there are more than about three steps.
  4. Every step has a compensation. Compensations are business actions (release reserve, void payment), not database rollbacks, and must be idempotent.
  5. The saga’s state is persisted, and state changes and outgoing messages must be atomic (outbox).
  6. Isolation is lost: other users may see intermediate states, so we model them explicitly (PendingPayment) and use semantic locks.
  7. We add timeouts, dead-letter handling, and correlation IDs so we can operate it.

Hands-on exercise

Task: Extend the in-memory SagaRunner from Level 1 into a persisted mini-saga for “Approve Claim” with three steps (ReserveLimit, AuthorizePayment, MarkClaimPaid).

  1. Create a SagaInstance table (Id, CurrentStep, Status, FailureReason, RowVersion) with EF Core.
  2. Persist the state after each step.
  3. Simulate a failure in AuthorizePayment (throw for amounts above 10,000).
  4. Simulate a crash by killing the process after step 1, restart it, and have it resume from the stored CurrentStep.
  5. Make ReleaseLimit idempotent and call it twice in a test.

Expected outcome: Amount 5,000 finishes as Completed. Amount 20,000 ends as Compensated with the reserve released exactly once. Killing the process mid-saga and restarting completes or compensates the saga instead of leaving it stuck. Calling ReleaseLimit twice leaves the data unchanged the second time.

Interview-style questions

  1. Why can’t we just use a distributed transaction (2PC) across microservices? It blocks and holds locks across the network, needs every participant (databases, brokers, external APIs) to support the protocol, and reduces availability because a coordinator or slow participant stalls everyone. Many cloud services and third-party APIs do not support it at all.
  2. What is the difference between choreography and orchestration, and when do you pick each? Choreography has services react to each other’s events with no central controller, which suits short flows of two to three steps. Orchestration uses a central state machine that sends commands and tracks progress, which suits longer flows with branching, timeouts, and compensation because the flow is visible in one place.
  3. What is a compensating transaction and what makes a good one? It is a new local transaction that semantically undoes a previously committed step (for example, releasing a reserve or voiding a payment authorisation). A good one is idempotent, retriable, tolerant of the forward step having partially completed, and expressed in business terms rather than as a data rollback.

Mastery checklist

  • I can explain why 2PC is a poor fit for microservices and cloud services.
  • I can draw the saga for a real workflow with steps, compensations, and the pivot transaction.
  • I can choose between choreography and orchestration and justify it.
  • I can implement an idempotent compensating handler and prove it with a test.
  • I can explain the isolation anomalies of sagas and apply at least one countermeasure (semantic lock).
  • I can make “save state + send message” atomic using the outbox pattern.
  • I can design timeouts, dead-letter handling, and alerts for stuck sagas.
  • I can produce an ADR comparing saga, 2PC, and other alternatives for a given scenario.

Key takeaway

A saga replaces one impossible distributed transaction with a sequence of local transactions plus designed-in compensations. It buys availability and independence at the cost of eventual consistency, so idempotency, persisted state, and observability are not optional.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 11: Domain Event

A Domain Event is an immutable record that something meaningful happened in the business domain, named in the past tense (for example ClaimApproved).

Manikandan
Manikandan·16 min read
Microservices

Day 10: API Composition

API Composition solves the "no more SQL JOIN" problem that appears once each microservice owns its own database.

Manikandan
Manikandan·13 min read
Microservices

Day 9: Event Sourcing

Event Sourcing stores every change to a business object as an immutable event in an append-only log, instead of overwriting the current state in a row.

Manikandan
Manikandan·18 min read
Microservices

Day 8: CQRS

CQRS (Command Query Responsibility Segregation) splits an application's model into two sides.

Manikandan
Manikandan·18 min read