Manikandan — Manikandan
Microservices

Day 8: CQRS

ManikandanManikandan
18 min read·Updated Aug 17, 2022

CQRS (Command Query Responsibility Segregation) splits an application's model into two sides.

Intro

CQRS (Command Query Responsibility Segregation) splits an application’s model into two sides. The write side (commands) enforces business rules and changes state. The read side (queries) returns data shaped exactly for the screens and reports that need it. The two sides can use different classes, different schemas, or even different databases, so each can be optimised and scaled on its own. In our insurance claims system, “submit a claim” and “show the adjuster’s claim dashboard” stop fighting over one table and one model.

Why we need this

Most business applications read far more than they write, and the two workloads want opposite things.

  • Writes need a normalised, consistent model with strict invariants (a claim cannot be approved above the policy limit, a closed claim cannot be edited). They need short transactions and few indexes, because every index slows inserts and updates.
  • Reads need denormalised, pre-joined, index-heavy data shaped like the UI (claim + policyholder name + latest status + total paid). They tolerate slight staleness but must be very fast.

A single model serves both badly. The domain entity gets polluted with display-only properties, queries drag in navigation properties they do not need, and every new dashboard adds an index that slows claim submission. CQRS exists so that each side can be designed for its own job. It is also a natural companion to Event Sourcing (Day 9), Domain Events (Day 11) and the Transactional Outbox (Day 12).

What problem it solves

Problem (from the topic list): A single model suffers lock contention and poor indexing under mixed read/write load.

Without CQRS in our claims system:

  • The adjuster dashboard runs a 6-table join with ORDER BY and paging while intake staff are inserting claims. Reads take shared locks (or, under default READ COMMITTED, block on writers’ locks) and writers wait behind long reads.
  • To make the dashboard fast, the team adds 8 indexes on the Claims table. Claim inserts and status updates get slower, and index maintenance causes page splits.
  • The same Claim entity is used for the write rules and for the JSON returned to Angular, so the API leaks internal fields, or the entity grows [NotMapped] display properties.
  • Scaling reads means scaling the one database that also handles writes, which is the most expensive way to scale.

CQRS resolves this by separating the paths. The write path stays small and rule-focused. The read path uses purpose-built projections (read models) that can live in the same database, a read replica, or a different store.

When it is needed (and when it is NOT)

Good fit

  • Read/write ratio is heavily skewed and the two sides have different scaling or latency needs (dashboards, search, reporting over transactional data).
  • The write model has real business rules (aggregates, invariants), while reads are mostly display shaping.
  • Multiple, very different read views exist for the same data (adjuster queue, customer portal, fraud analytics).
  • You already publish domain events and can build projections from them.
  • You need to use a different store for reads (Azure AI Search, Cosmos DB, a reporting schema) without touching the write model.

Poor fit or overkill

  • Simple CRUD screens with little business logic. A single EF Core model is faster to build and easier to teach.
  • Small teams with low load. The extra moving parts (projections, sync, lag) cost more than they save.
  • Screens that must always show the result of the user’s own write instantly and you cannot tolerate or design around eventual consistency.
  • Applying CQRS to the whole system by default. Apply it per bounded context or per aggregate where the pain exists (for example, only to Claims, not to Reference Data).

Important clarification for your team: CQRS has levels. Separating command and query code paths on one database is CQRS. Separate databases is an optional extra step. Many teams need only the first level.

How to identify the problem (key signals)

  1. Lock waits and deadlocks in SQL Server (sys.dm_tran_locks, deadlock graphs in Extended Events) where a reporting query and an update touch the same rows.
  2. Index bloat on a hot table: more than a handful of nonclustered indexes on a table that receives frequent inserts, added purely for read screens.
  3. Write latency degrades when a report runs, and p95 for POST /claims correlates with dashboard usage in Application Insights.
  4. Entity classes with display-only properties, DTO mapping full of Include(...).ThenInclude(...) chains, or “god” entities serving ten different screens.
  5. Query handlers that ignore the domain rules but still load full aggregates, causing over-fetching (SELECT * on wide tables, large change-tracker memory).
  6. Conflicting change requests: the reporting team wants denormalisation, the domain team wants normalisation, and every schema change needs both to agree.
  7. Reads cannot be scaled independently: the only way to handle more dashboard traffic is a bigger primary database tier.

Flow Diagram

Commands and queries take different paths; projections keep the read model current.

flowchart LR
UI["Angular UI"] -- "POST /api/claims" --> CMD["Command handler"]
CMD --> AGG["Claim aggregate - rules"]
AGG --> WDB[("Write DB - normalised")]
AGG -. "ClaimSubmitted" .-> PROJ["Projection handler"]
PROJ --> RDB[("ClaimListView - read model")]
UI -- "GET /api/claims" --> Q["Query handler"]
Q --> RDB

Level 1: Beginner

Analogy: A restaurant. Waiters take orders (commands): they must be accurate, and the kitchen validates them. The menu board and the “today’s specials” screen (queries) are read-only, prepared in advance, and can be looked at by hundreds of people without disturbing the kitchen.

Core rules

  • A command expresses intent and changes state: SubmitClaim, ApproveClaim. It returns little or nothing (an ID or success/failure).
  • A query returns data and never changes state: GetClaimSummary, ListOpenClaims.
  • Commands go through the domain model. Queries do not need to.

Minimal example (single database, separate code paths), .NET 10 / C# 14

Commands/SubmitClaim.cs
public sealed record SubmitClaimCommand(Guid PolicyId, decimal Amount, string Description);
public sealed class SubmitClaimHandler(ClaimsDbContext db)
{
public async Task<Guid> HandleAsync(SubmitClaimCommand cmd, CancellationToken ct)
{
var policy = await db.Policies.FindAsync([cmd.PolicyId], ct)
?? throw new InvalidOperationException("Policy not found");
var claim = Claim.Submit(policy, cmd.Amount, cmd.Description); // domain rules live here
db.Claims.Add(claim);
await db.SaveChangesAsync(ct);
return claim.Id;
}
}
Queries/GetClaimSummary.cs
public sealed record ClaimSummaryDto(Guid Id, string Status, decimal Amount, string PolicyHolder);
public sealed class GetClaimSummaryHandler(ClaimsDbContext db)
{
public Task<ClaimSummaryDto?> HandleAsync(Guid id, CancellationToken ct) =>
db.Claims
.AsNoTracking() // read side: no change tracking
.Where(c => c.Id == id)
.Select(c => new ClaimSummaryDto(c.Id, c.Status.ToString(), c.Amount, c.Policy.HolderName))
.SingleOrDefaultAsync(ct); // projects only needed columns
}

Even this small step already gives value: queries do not load aggregates, and commands do not care about display shapes. Note that we have not added a message bus or a second database.

Level 2: Intermediate

Here we use CQRS in a real .NET + Angular + SQL application, with a separate read model kept up to date by domain events.

Architecture for the Claims service

  • Write side: ASP.NET Core Minimal API, EF Core 10, SQL Server (normalised: Claims, Policies, ClaimLines).
  • Read side: a denormalised table ClaimListView (one row per claim, already joined) in the same SQL Server database or a replica.
  • Sync: the command handler raises a domain event, and a projection handler updates the read table. For cross-process sync, publish through the Transactional Outbox (Day 12).

Write side: aggregate raises an event

public sealed record ClaimSubmitted(Guid ClaimId, Guid PolicyId, string HolderName, decimal Amount, DateTime SubmittedUtc);
public sealed class Claim
{
private readonly List<object> _events = [];
public IReadOnlyList<object> Events => _events;
public Guid Id { get; private set; } = Guid.NewGuid();
public Guid PolicyId { get; private set; }
public decimal Amount { get; private set; }
public ClaimStatus Status { get; private set; }
public static Claim Submit(Policy policy, decimal amount, string description)
{
if (amount <= 0) throw new DomainException("Amount must be positive");
if (amount > policy.CoverageLimit) throw new DomainException("Exceeds coverage limit");
var claim = new Claim { PolicyId = policy.Id, Amount = amount, Status = ClaimStatus.Submitted };
claim._events.Add(new ClaimSubmitted(claim.Id, policy.Id, policy.HolderName, amount, DateTime.UtcNow));
return claim;
}
}
public enum ClaimStatus { Submitted, UnderReview, Approved, Rejected, Closed }

Read model and projection (same database, updated in the same transaction as the simplest option)

public sealed class ClaimListView
{
public Guid ClaimId { get; set; }
public string HolderName { get; set; } = "";
public decimal Amount { get; set; }
public string Status { get; set; } = "";
public DateTime SubmittedUtc { get; set; }
}
public sealed class ClaimListProjection(ClaimsDbContext db)
{
public async Task On(ClaimSubmitted e, CancellationToken ct)
{
db.ClaimListViews.Add(new ClaimListView
{
ClaimId = e.ClaimId, HolderName = e.HolderName, Amount = e.Amount,
Status = "Submitted", SubmittedUtc = e.SubmittedUtc
});
await db.SaveChangesAsync(ct);
}
}

Query endpoint (fast path, no joins)

app.MapGet("/api/claims", async (ClaimsDbContext db, string? status, int page = 1, int size = 20, CancellationToken ct) =>
{
var q = db.ClaimListViews.AsNoTracking();
if (!string.IsNullOrEmpty(status)) q = q.Where(x => x.Status == status);
var items = await q.OrderByDescending(x => x.SubmittedUtc)
.Skip((page - 1) * size).Take(size)
.ToListAsync(ct);
return Results.Ok(items);
});
app.MapPost("/api/claims", async (SubmitClaimCommand cmd, SubmitClaimHandler h, CancellationToken ct) =>
{
var id = await h.HandleAsync(cmd, ct);
return Results.Accepted($"/api/claims/{id}", new { id }); // 202: read model may lag slightly
});

Index for the read model (SQL Server)

CREATE INDEX IX_ClaimListView_Status_Submitted
ON dbo.ClaimListView (Status, SubmittedUtc DESC)
INCLUDE (ClaimId, HolderName, Amount);

Angular (standalone component, signals) that handles eventual consistency

import { Component, inject, signal } from '@angular/core';
import { HttpClient } from '@angular/common/http';
interface ClaimRow { claimId: string; holderName: string; amount: number; status: string; }
@Component({
selector: 'app-claim-list',
template: `
<button (click)="submit()">Submit test claim</button>
@if (pendingId()) { <p>Claim {{ pendingId() }} accepted, it will appear shortly.</p> }
<ul>
@for (c of claims(); track c.claimId) {
<li>{{ c.holderName }} - {{ c.amount | currency }} - {{ c.status }}</li>
}
</ul>`
})
export class ClaimListComponent {
private http = inject(HttpClient);
claims = signal<ClaimRow[]>([]);
pendingId = signal<string | null>(null);
constructor() { this.load(); }
load() { this.http.get<ClaimRow[]>('/api/claims').subscribe(r => this.claims.set(r)); }
submit() {
this.http.post<{ id: string }>('/api/claims',
{ policyId: '...', amount: 1200, description: 'Windscreen' })
.subscribe(r => {
this.pendingId.set(r.id); // show optimistic message, do not assume the list is updated
setTimeout(() => this.load(), 1000); // simple refresh; production: SignalR push or polling by id
});
}
}

The key UI lesson is that after a command returns, the read model may not have caught up. Design the UX for that (optimistic message, refresh, push notification), rather than pretending it is instant.

Level 3: Advanced

Performance and scalability

  • Read models are cheap to scale: add replicas, add caching, or move to a store built for reads. Writes remain on a small, well-tuned transactional store.
  • Keep read models rebuildable. If a projection is wrong, you should be able to drop and rebuild it from the source of truth (the write DB or the event stream).
  • Use keyset paging for very large lists instead of Skip/Take with large offsets.

Consistency and failure modes

  • Projection lag: measure it (time between SubmittedUtc and the read row’s ProjectedUtc) and alert on it.
  • Lost or duplicated events: publishing an event and saving state in two steps is a dual-write. Use the Transactional Outbox (Day 12) and make projections idempotent (Day 17), for example by storing the last processed event version per row.
  • Out-of-order events: carry a version or sequence number per aggregate, and ignore stale updates.
  • Read-your-own-writes: options include returning the new state from the command, routing that user’s next read to the primary, or waiting until the projection reaches a given version.

Security

  • Authorise both sides separately. Queries must apply the same row-level rules (a customer sees only their own claims). Do not assume that “reads are safe”.
  • Read models can contain denormalised PII. Apply the same encryption, masking, and retention rules as the source.

Common mistakes

  1. Using CQRS everywhere, including trivial CRUD contexts.
  2. Adding a second database on day one when a separate query path on one database would have solved the problem.
  3. Letting query handlers call command handlers, or commands returning large read DTOs.
  4. No monitoring of projection lag, so staleness is discovered by users.
  5. Read models that cannot be rebuilt.
  6. Putting business rules in projections. Projections should only reshape data.
  7. Confusing CQRS with a mediator library. A mediator is optional plumbing; CQRS is a modelling decision. Note that some popular mediator libraries have moved to commercial licensing in recent major versions, so check licences before standardising, or use plain handler classes as shown above.

Level 4: Expert and Architect view

Levels of CQRS compared

OptionDescriptionStrengthsWeaknessesUse when
Single model, no CQRSOne entity/model for reads and writesSimplest, fastest to buildRead/write contention, bloated entitiesSmall CRUD apps
CQRS level 1: separate code paths, one DBCommand handlers use aggregates, query handlers project directlyLow cost, no lag, clean codeSame DB scaling limitsFirst step for most teams
CQRS level 2: separate read tables/views, same DBDenormalised read tables, updated by events or triggersFast queries, custom indexesSync logic, small lagHeavy read screens
CQRS level 3: separate read storeRead replica, Cosmos DB, or search index fed by eventsIndependent scale, polyglot storageEventual consistency, operational costVery different read/write scale or search needs
CQRS + Event SourcingWrite side is an event log, projections build readsFull audit, rebuildable viewsHighest complexity, steep learning curveAudit-heavy domains
API Composition (Day 10)Join at query time across servicesNo sync neededSlow for large joins, availability couplingSmall cross-service reads

Patterns it combines with

  • Domain Event (Day 11) and Transactional Outbox (Day 12) to feed projections reliably.
  • Idempotent Consumer (Day 17) so projections tolerate redelivery.
  • Event Sourcing (Day 9) as an optional write-side storage choice.
  • Database per Service (Day 5): CQRS builds read models locally so services avoid cross-service joins.
  • API Gateway / BFF (Days 19-20) to route queries and commands to the right endpoints.

ADR (architecture review style)

ADR-008: Adopt CQRS (level 2) for the Claims bounded context

Status: Proposed

Context: The Claims database serves intake (write-heavy bursts) and the adjuster dashboard (read-heavy, multi-join). Dashboard queries cause lock waits during peak intake, and 9 read-oriented indexes slow claim inserts. p95 for POST /claims rises by roughly 2x when reports run (to be validated with Application Insights data).

Decision: Separate the write model (EF Core aggregates, normalised tables) from a denormalised ClaimListView read model, updated by domain events published via a transactional outbox. Queries bypass aggregates and read the view. Other contexts (Reference Data, Policy Admin) stay as simple CRUD.

Consequences: (+) Faster dashboards, fewer indexes on write tables, independent evolution of screens. (-) Eventual consistency (target lag under 2 seconds at p95), extra code for projections, need for lag monitoring and a rebuild tool.

Alternatives considered: Add more indexes and READ_COMMITTED_SNAPSHOT (cheaper, tried first, a valid step if it is enough); read replica only (helps load, not the model problem); full Event Sourcing (rejected, audit needs are met by Day 34 audit logging).

Revisit when: projection lag breaches its SLO repeatedly, or reads need a different store.

Azure implementation

Services that support CQRS

NeedAzure serviceNotes
Write storeAzure SQL Database (or Azure Database for PostgreSQL flexible server)Normalised transactional store
Read scale-outAzure SQL read scale-out replica (available on Premium and Business Critical tiers, and on Hyperscale via named/HA replicas)Route read queries with ApplicationIntent=ReadOnly in the connection string
Read store (documents)Azure Cosmos DB for NoSQLDenormalised documents, feed from the change feed
Read store (search)Azure AI SearchFull-text and faceted claim search
Event transportAzure Service Bus (topics/subscriptions) or Azure Event HubsService Bus for reliable business events with sessions, dead-lettering, duplicate detection
Projection hostAzure Container Apps, Azure Functions, or AKSConsumers that update read models
CacheAzure Managed Redis (Azure Cache for Redis is being succeeded by Managed Redis, so verify the current SKU names in your region)Hot read paths
MonitoringApplication Insights + Azure MonitorTrack projection lag and DLQ depth

How to configure the main pieces

  1. Azure SQL read replica routing: add ApplicationIntent=ReadOnly to the connection string used by the query side. Register a second DbContext (for example ClaimsReadDbContext) with that connection string and UseQueryTrackingBehavior(QueryTrackingBehavior.NoTracking).
  2. Service Bus: create a topic claims-events and a subscription per projection (claims-list-projection, claims-search-indexer). Enable dead-lettering, set MaxDeliveryCount (for example 10), and enable duplicate detection when the publisher sets MessageId to the event ID.
  3. Cosmos DB read store (optional): partition by the query access pattern (for example /adjusterId). A change feed processor or an Azure Function with a Cosmos trigger can keep downstream stores in sync.
  4. Identity: use managed identities for Service Bus and Cosmos DB access, with Azure RBAC roles instead of connection-string secrets. Keep remaining secrets in Key Vault (Day 39).

Pricing and tier considerations (always confirm current numbers on the Azure pricing pages, since prices vary by region and change)

  • Azure SQL read scale-out is included at no extra cost on Premium/Business Critical, but those tiers cost more than General Purpose. On General Purpose you do not get a built-in readable secondary, so consider Hyperscale replicas or a read model in the same database.
  • Service Bus: Standard tier (pay per operation plus a base charge) supports topics and duplicate detection. Premium gives dedicated capacity (messaging units), predictable latency, and larger messages, and is suited to production workloads with strict performance needs. Basic has no topics, so it does not fit CQRS fan-out.
  • Cosmos DB: choose provisioned throughput (manual or autoscale) or serverless for spiky, low-volume read models. Request Units (RU) per query drive cost, so model the read documents to answer queries with a single partition read.
  • Azure AI Search: billed per search unit by tier; start with the smallest tier that fits your index size.
  • Cost principle: a second read store is only worth it when it removes more cost or risk than it adds. Start with level 1 or 2 on the same SQL database.

Reference architecture (text)

  1. An Angular 21+ app (hosted on Azure Static Web Apps) calls the API through Azure API Management or Application Gateway.
  2. The Claims API on Azure Container Apps exposes command endpoints (POST/PUT) and query endpoints (GET).
  3. Command handlers write to Azure SQL (primary) and, in the same transaction, insert an event row into the outbox table.
  4. An outbox publisher (Day 12/14) sends events to the Service Bus topic claims-events.
  5. A projection worker (Container Apps job or Azure Function with a Service Bus trigger) consumes events idempotently and updates ClaimListView and/or Cosmos DB and the AI Search index.
  6. Query endpoints read from the read replica (ApplicationIntent=ReadOnly), Cosmos DB, or AI Search, with optional Redis caching.
  7. Application Insights collects traces (Day 31) with a custom metric for projection lag; alerts fire on lag and dead-letter growth.

Teaching guide for my team

Explain to a beginner in 2 minutes

“Think of a bank. When you deposit money, a teller carefully records it and checks the rules. When you look at your balance on the app, you see a ready-made summary. CQRS means we build our software the same way: one path for changing data, which is careful and strict, and another for reading data, which is fast and shaped for the screen. They can be different code, different tables, even different databases. The price is that the summary may be a moment behind, so the screen must be designed for that.”

Explain to an intermediate developer in 5 minutes

  1. Show the pain: a dashboard query with 6 joins blocking claim inserts, plus a list of indexes on Claims.
  2. Introduce levels: (a) separate handlers on one DB, (b) read tables, (c) separate read store.
  3. Walk through the flow: command, aggregate rules, save state and outbox event, projection updates the read model, query reads it.
  4. Discuss consistency: lag, idempotent projections, versioning, read-your-own-writes.
  5. Close with when not to use it: simple CRUD, small team, no real pain yet.

Hands-on exercise

Task: Starting from a CRUD Claims API, (1) split it into SubmitClaimHandler and GetClaimSummaryHandler; (2) add a ClaimListView table with an index on (Status, SubmittedUtc DESC); (3) write a projection that updates it when ClaimSubmitted is raised; (4) add a ProjectedUtc column and an endpoint that reports current lag; (5) drop and rebuild the view from the Claims table with a script.

Expected outcome: GET /api/claims reads only from ClaimListView with no joins. Submitting a claim makes it appear after the projection runs. The lag endpoint shows a small number. The rebuild script recreates the same rows, proving the read model is disposable. The team can explain which parts are the command side and which are the query side.

Interview-style questions

  1. What is the difference between CQRS and Event Sourcing? CQRS separates read and write models. Event Sourcing stores state as a sequence of events. They are independent, and CQRS works fine with a normal database.
  2. How do you handle a user who submits a claim and immediately refreshes the list? Return the created ID and state from the command, show it optimistically in the UI, or route the next read to the primary or wait for the projection version. Accept eventual consistency where the business allows.
  3. When would you avoid CQRS? For simple CRUD contexts, small teams with no measurable contention, or where strong immediate consistency is required on every read and the load does not justify the complexity.

Mastery checklist

  • I can explain commands vs queries and why queries must not change state.
  • I can describe the three levels of CQRS and choose the cheapest one that solves a given problem.
  • I can implement a query handler that projects to a DTO with AsNoTracking and no aggregate loading.
  • I can build and index a denormalised read model and keep it in sync using domain events.
  • I can make projections idempotent, order-aware, and rebuildable.
  • I can design the UI and API (202 Accepted, optimistic updates) around eventual consistency.
  • I can measure and alert on projection lag and dead-letter queues in Azure Monitor.
  • I can write an ADR that justifies CQRS for one bounded context and states when to revisit it.

Key takeaway

CQRS lets the write side protect business rules while the read side is shaped purely for speed, at the price of eventual consistency and extra plumbing. Apply it only where mixed read/write load is a measured problem, and start with the simplest level.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 11: Domain Event

A Domain Event is an immutable record that something meaningful happened in the business domain, named in the past tense (for example ClaimApproved).

Manikandan
Manikandan·16 min read
Microservices

Day 10: API Composition

API Composition solves the "no more SQL JOIN" problem that appears once each microservice owns its own database.

Manikandan
Manikandan·13 min read
Microservices

Day 9: Event Sourcing

Event Sourcing stores every change to a business object as an immutable event in an append-only log, instead of overwriting the current state in a row.

Manikandan
Manikandan·18 min read
Microservices

Day 7: Saga Pattern

A saga is a way to complete one business process that spans several services, each with its own database, without a distributed lock or a two-phase commit.

Manikandan
Manikandan·20 min read