The Strangler Fig pattern replaces a legacy monolith gradually instead of rewriting it in one go.
Intro
The Strangler Fig pattern replaces a legacy monolith gradually instead of rewriting it in one go. You place a routing layer (a facade) in front of the monolith, build new functionality as separate services, and move traffic route by route from old to new. When the last route has moved, the monolith is switched off. Like the fig tree that grows around a host tree until the host is gone, the new system grows around the old one while production keeps running.
Running example: an insurance claims system. A ten-year-old ASP.NET monolith (ClaimsCore) handles claim intake, adjudication, payments, and document handling. We migrate it to services one capability at a time.
Why we need this
- Business reason: The claims system earns money every minute. A “stop the world and rewrite” project produces no customer value for 12-24 months and often ends with a system that misses years of accumulated business rules.
- Technical reason: Legacy monoliths hold undocumented behaviour (rounding rules, regulatory edge cases). Only production traffic reveals all of it. Incremental migration lets you compare old and new behaviour on real traffic.
- Risk reason: Each slice is small and reversible. If the new
Claim Intakeservice misbehaves, you flip the route back to the monolith in minutes. - Funding reason: Value ships every few weeks, so the migration survives budget reviews and staff changes.
What problem it solves
Problem: Big-bang monolith rewrites carry extreme risk.
What goes wrong without it:
- The rewrite runs in parallel with a monolith that keeps changing. The target moves, and the new system is never “feature complete”.
- Cut-over happens on one weekend with all data and all users. Any defect affects everything and rollback means restoring database backups.
- Domain knowledge hidden in old code is lost, and defects surface in production as wrong claim payouts.
- Teams burn out and the project is cancelled at 70%, leaving two half-systems to maintain.
The Strangler Fig removes the single cut-over. Every request is served by exactly one system at any time, and the decision is a routing rule that can be changed without deployment of either system.
When it is needed (and when it is NOT)
Use it when:
- The monolith must stay live during migration (24/7 claims intake, regulatory SLAs).
- You can put a routing layer in front of the system (HTTP/API, or a message channel you can intercept).
- The monolith can be split along seams: URLs, API operations, or business capabilities.
- You want to migrate incrementally by capability (see Day 1) or subdomain (see Day 2).
Do NOT use it when:
- The system is small and a rewrite takes a few weeks. The routing layer is overhead.
- You cannot intercept the traffic (for example a thick desktop client with hard-coded database connections and no server API). Consider a different approach first, such as introducing an API.
- The monolith is being retired with no replacement (just decommission it).
- The legacy codebase is so unstable that even the facade cannot be added safely. Stabilise first.
- The data model is so entangled that no slice can be moved without moving most tables. Do the data-untangling work first (see Day 5 and Day 6), otherwise you end up with a distributed monolith.
How to identify the problem (key signals)
- Release cadence is measured in quarters, and every release needs a weekend outage window.
- A small change (for example a new claim status) touches dozens of projects and needs a full regression suite that takes days.
- Nobody can explain how a business rule works; “ask the one developer who knows”.
- The runtime or framework is out of support (.NET Framework 4.x on Windows Server only, ASP.NET Web Forms), blocking security patches and cloud adoption.
- Scaling one hot feature (claim document upload) means scaling the whole application.
- A previous rewrite attempt was cancelled or keeps slipping (“v2 is nearly done” for two years).
- Production incidents in one module take down unrelated modules because they share memory, threads, and a database.
Flow Diagram
A facade routes migrated slices to new services and everything else to the monolith.
flowchart LR U["Users / Angular"] --> FAC["YARP facade"] FAC -- "POST /api/claims - flag on" --> NEW["Claim Intake service"] FAC -- "everything else" --> MON["ClaimsCore monolith"] NEW --> PG[("PostgreSQL")] MON --> SQL[("SQL Server")] NEW -. "ClaimSubmitted via outbox" .-> SB{{"Service Bus"}} SB -.-> ACL["Anti-corruption layer"] --> SQLLevel 1: Beginner
Analogy: Renovating a restaurant while it stays open. You do not close for a year. You rebuild one section at a time, put up a temporary wall, and route customers to the finished section. When all sections are done, the old kitchen is removed.
Three parts:
- Facade (proxy/router): receives all requests and decides where each goes.
- Legacy system: keeps serving everything not yet migrated.
- New service(s): serve the migrated routes.
Three steps (repeat per slice): Transform (build the new service), Coexist (route some or all traffic to it while the old code remains), Eliminate (delete the old code once traffic is fully moved).
Minimal working example (C#, .NET 10, ASP.NET Core minimal API acting as a hand-rolled facade):
// Program.cs - a tiny facade using HttpClient. For real use, prefer YARP (Level 2).var builder = WebApplication.CreateBuilder(args);
builder.Services.AddHttpClient("legacy", c => c.BaseAddress = new Uri("http://legacy-claims:8080"));builder.Services.AddHttpClient("intake", c => c.BaseAddress = new Uri("http://claim-intake-svc:8080"));
var app = builder.Build();
// Migrated route: goes to the new serviceapp.MapPost("/api/claims", async (HttpContext ctx, IHttpClientFactory f) =>{ var client = f.CreateClient("intake"); using var req = new HttpRequestMessage(HttpMethod.Post, "/api/claims") { Content = new StreamContent(ctx.Request.Body) }; req.Content.Headers.ContentType = new System.Net.Http.Headers.MediaTypeHeaderValue("application/json"); var resp = await client.SendAsync(req, ctx.RequestAborted); ctx.Response.StatusCode = (int)resp.StatusCode; await resp.Content.CopyToAsync(ctx.Response.Body);});
// Everything else still goes to the monolithapp.MapFallback(async (HttpContext ctx, IHttpClientFactory f) =>{ var client = f.CreateClient("legacy"); using var req = new HttpRequestMessage(new HttpMethod(ctx.Request.Method), ctx.Request.Path + ctx.Request.QueryString); if (ctx.Request.ContentLength > 0) req.Content = new StreamContent(ctx.Request.Body); var resp = await client.SendAsync(req, ctx.RequestAborted); ctx.Response.StatusCode = (int)resp.StatusCode; await resp.Content.CopyToAsync(ctx.Response.Body);});
app.Run();This hand-rolled version drops headers and does not stream well. It shows the idea only. Use YARP in real projects.
Level 2: Intermediate
Real-world setup: Angular SPA + .NET facade (YARP) + legacy ASP.NET monolith + new .NET 10 service + SQL Server (legacy) and PostgreSQL (new service).
6.1 The facade with YARP
YARP (Yarp.ReverseProxy, currently the 2.x line) is Microsoft’s reverse proxy toolkit for ASP.NET Core. Routes and clusters are configuration, so shifting traffic is a config change.
var builder = WebApplication.CreateBuilder(args);
builder.Services .AddReverseProxy() .LoadFromConfig(builder.Configuration.GetSection("ReverseProxy"));
var app = builder.Build();app.MapReverseProxy();app.Run();{ "ReverseProxy": { "Routes": { "claims-intake": { "ClusterId": "intake-cluster", "Order": 1, "Match": { "Path": "/api/claims/{**catch-all}", "Methods": [ "POST" ] } }, "legacy-catch-all": { "ClusterId": "legacy-cluster", "Order": 1000, "Match": { "Path": "{**catch-all}" } } }, "Clusters": { "intake-cluster": { "Destinations": { "d1": { "Address": "http://claim-intake-svc:8080/" } } }, "legacy-cluster": { "Destinations": { "d1": { "Address": "http://legacy-claims:8080/" } } } } }}Lower Order values are evaluated first, so specific migrated routes win over the catch-all. Slice migration = add a route with a lower order.
6.2 Percentage-based (canary) routing
To send 10% of intake traffic to the new service, add a custom middleware or transform that picks the cluster. A simple approach with an in-process feature flag:
using Microsoft.FeatureManagement;using Yarp.ReverseProxy.Model;
builder.Services.AddFeatureManagement();
// Program.cs after builder.Build()app.MapReverseProxy(proxy =>{ proxy.Use(async (ctx, next) => { var features = ctx.RequestServices.GetRequiredService<IFeatureManager>(); var route = ctx.GetRouteModel(); if (route.Config.RouteId == "claims-intake" && !await features.IsEnabledAsync("NewIntake")) { // Flag off: send this request to the monolith instead var legacy = ctx.RequestServices .GetRequiredService<IProxyStateLookup>(); if (legacy.TryGetCluster("legacy-cluster", out var cluster)) ctx.ReassignProxyRequest(cluster); } await next(); });});Microsoft.FeatureManagement supports percentage filters, so NewIntake can be enabled for 10%, then 50%, then 100%. ReassignProxyRequest is a YARP extension for changing the target cluster in the proxy pipeline.
6.3 Angular does not need to change
The Angular app keeps calling the same base URL (/api/claims). Only the facade knows the backend moved.
// claims.service.ts (Angular, standalone, unchanged during migration)import { Injectable, inject } from '@angular/core';import { HttpClient } from '@angular/common/http';
@Injectable({ providedIn: 'root' })export class ClaimsService { private http = inject(HttpClient);
submit(claim: NewClaim) { return this.http.post<ClaimReceipt>('/api/claims', claim); }}6.4 Data: the hard part
The new Intake service should own its data (see Day 5). During coexistence, the monolith still needs claim records that Intake creates. Two proven options:
- Sync via events: Intake writes to its PostgreSQL database and publishes
ClaimSubmittedusing the Transactional Outbox (Day 12). An anti-corruption layer (ACL) inside the monolith consumes the event and inserts into the legacy SQL Server tables. - Legacy remains the source of truth for a period: Intake calls the monolith for reads (API Composition, Day 10) until reporting moves too.
Anti-corruption layer sketch (pseudo-code for the mapping idea):
// Inside a legacy-side worker (compiles as a plain class; wiring omitted)public sealed class ClaimSubmittedHandler(LegacyDbContext db){ public async Task HandleAsync(ClaimSubmitted e, CancellationToken ct) { // Map new model -> legacy schema; keep legacy naming out of the new service db.LegacyClaims.Add(new LegacyClaim { CLM_NO = e.ClaimNumber, POL_NO = e.PolicyNumber, LOSS_DT = e.LossDate, CLM_STAT = "NEW" }); await db.SaveChangesAsync(ct); }}Level 3: Advanced
7.1 Performance
- The facade adds one network hop. Keep it stateless, co-located with the backends, and use HTTP keep-alive (YARP uses
SocketsHttpHandlerconnection pooling). - Do not buffer large uploads (claim photos, PDFs). Stream them through.
- Measure added latency per route before and after; a few milliseconds is normal, tens of milliseconds means a problem.
7.2 Scalability and availability
- The facade is now a single point of failure. Run at least two instances across availability zones behind a load balancer, and keep it free of business logic.
- Keep configuration in a central store so all instances switch routes together.
7.3 Security
- Terminate TLS and validate tokens once at the edge (see Day 38, Day 19), then forward identity to both old and new systems. The monolith may use cookie/forms auth: use the facade to translate JWT to the legacy identity header, or run the new service and the monolith under a shared identity provider.
- Do not expose the monolith directly once the facade exists. Restrict the network path (private endpoints, NSGs) so all traffic goes through the facade, otherwise routing rules can be bypassed.
7.4 Failure modes and how to handle them
| Failure mode | Effect | Mitigation |
|---|---|---|
| New service down | Migrated routes return errors | Circuit breaker + automatic fallback route to monolith while legacy code still exists |
| Data drift between old and new stores | Reports disagree, wrong payouts | Reconciliation job, compare counts and checksums daily, dual-run “shadow” comparison |
| Facade misconfiguration | Wrong route order sends traffic to the wrong system | Route config in Git, reviewed, tested by automated routing tests |
| Session/state split | User logs in on legacy, new service does not know the session | Shared token-based auth, no server-side session dependence |
| Distributed transaction gap | Claim created in new system but not in legacy | Outbox + idempotent consumer (Day 12, Day 17) |
7.5 Shadow traffic (dark launch)
Before switching, send a copy of real requests to the new service and compare its responses to the monolith’s without returning the new result to users. This exposes behavioural differences safely. Only do this for read-only or idempotent calls, never for payments.
7.6 Common mistakes
- Never eliminating: the monolith code stays “just in case”. The third step is mandatory, otherwise you maintain two systems forever.
- Slicing by technical layer (move the “data access layer” first) instead of by business capability.
- Sharing the database indefinitely so the new service is still coupled to the legacy schema.
- Putting business logic in the facade, turning it into a new monolith.
- Migrating the easiest pieces only, leaving the hardest core untouched with no plan.
- No rollback path or no monitoring per route.
Level 4: Expert and Architect view
8.1 Alternatives compared
| Approach | Risk | Time to first value | Rollback | Best when |
|---|---|---|---|---|
| Big-bang rewrite | Very high | Long (12-24 months) | Hard (data restore) | Small system, well-understood, downtime acceptable |
| Strangler Fig | Low per slice | Weeks | Easy (route flip) | Live system, traffic interceptable |
| Branch by Abstraction | Low | Weeks | Easy (flag) | Change is inside the codebase, no network seam |
| Parallel run (dual run) | Medium | Medium | Easy | Need proof of equivalence (financial calcs) |
| Lift and shift only | Low | Fast | Easy | Goal is hosting change, not architecture change |
| Modular monolith first | Low-medium | Medium | Easy | Boundaries unclear; not ready for services |
8.2 Patterns it combines with
- API Gateway (Day 19): the facade is often the gateway itself.
- Decompose by Business Capability / Subdomain (Day 1, Day 2): decides the slices.
- Anti-Corruption Layer (DDD): protects the new model from legacy naming and rules.
- Transactional Outbox + Idempotent Consumer (Day 12, Day 17): reliable data sync between old and new.
- Circuit Breaker and Fallback (Day 26, Day 29): safe routing while both exist.
- Feature flags / canary release: gradual traffic shift.
- Database per Service (Day 5) and Shared Database (Interim) (Day 6): the data strategy during coexistence.
8.3 ADR
ADR-049: Migrate ClaimsCore to services using the Strangler Fig pattern
- Status: Proposed
- Context: ClaimsCore is a .NET Framework 4.8 monolith on Windows VMs. .NET Framework 4.8 is still supported as a Windows component, but it cannot use modern .NET features or run on Linux containers. Releases happen quarterly with outages. Claims intake must stay available 24/7. A previous rewrite attempt was cancelled.
- Decision: Introduce a YARP-based facade in front of ClaimsCore. Migrate by business capability in this order: Claim Intake, Document Handling, Payments, Adjudication. Each slice gets its own service and datastore, synchronised with the monolith through outbox events and an ACL until the slice is fully cut over. The monolith is decommissioned when the last route is migrated.
- Consequences (positive): value shipped every 4-6 weeks; per-slice rollback; no big cut-over.
- Consequences (negative): temporary complexity (two systems, data sync); facade becomes critical infrastructure; team must maintain both stacks during migration.
- Alternatives rejected: big-bang rewrite (risk), lift-and-shift only (does not fix coupling).
- Exit criteria per slice: 100% traffic on new service for 2 weeks with error rate and latency within SLO; legacy code for the slice deleted.
Azure implementation
9.1 Azure services
| Role | Service | Notes |
|---|---|---|
| Facade / router | YARP on Azure Container Apps or AKS or App Service | Full control, config-driven, .NET-native |
| Global edge routing | Azure Front Door (Standard/Premium) | Path-based routing to origins, weighted origins for canary, WAF on Premium |
| Regional L7 routing | Azure Application Gateway v2 | URL path-based routing rules, WAF option |
| API layer | Azure API Management | Policies, versioning, backend routing per operation |
| Legacy hosting | Azure VMs, App Service, or Azure Migrate rehost | Move the monolith to Azure first if it is on-premises |
| New services | Azure Container Apps, AKS, App Service | Linux containers on .NET 10 |
| Messaging | Azure Service Bus | Reliable events between old and new |
| Data | Azure SQL / SQL Server on VM (legacy), Azure Database for PostgreSQL Flexible Server (new) | Data per service |
| Config / flags | Azure App Configuration | Feature flags and percentage rollout with Microsoft.FeatureManagement |
| Observability | Application Insights / Azure Monitor | Per-route metrics, distributed traces |
Microsoft’s Azure Architecture Center documents this pattern as “Strangler Fig”, and it lists the facade, gradual migration, and decommissioning steps described above.
9.2 Configuration outline
- Host the monolith unchanged on Azure (VM or App Service) behind a private endpoint or internal load balancer.
- Deploy the facade (YARP container) on Azure Container Apps with min replicas of 2, external ingress, and the monolith as the default cluster.
- Store routing flags in Azure App Configuration:
NewIntakewith a percentage filter (10, 50, 100). The facade loads them with theMicrosoft.Extensions.Configuration.AzureAppConfigurationprovider and refreshes on a sentinel key. - Add Front Door in front if you need global reach, TLS, and WAF. Use it for coarse routing (for example
/api/*to the facade,/*to the Angular static site); keep fine-grained slice routing in the facade so it is testable in code and config. - Instrument each route with Application Insights and add a workbook comparing error rate and latency of legacy vs new per route.
- Secure: put the monolith in a VNet with no public IP, allow inbound only from the facade subnet.
9.3 Pricing and tier considerations
Always check the current Azure pricing pages before budgeting, since prices vary by region and change over time.
- Azure Container Apps: consumption plan bills per vCPU-second, GiB-second, and requests, with a monthly free grant; a dedicated plan is available for predictable workloads. Good for a small facade.
- Azure Front Door: Standard and Premium tiers (Classic is being retired). Premium adds WAF managed rulesets and Private Link to origins; both charge a base fee per profile plus request and data transfer.
- Application Gateway: v2 SKUs (Standard_v2 / WAF_v2) bill per gateway hour plus capacity units. v1 is retired, so use v2 only.
- API Management: tiers include Consumption, Developer (non-production), Basic, Standard, Premium, plus v2 tiers. Use Developer for learning; do not use it in production because it has no SLA.
- Azure App Configuration: a free tier exists with limits; Standard adds higher limits and an SLA.
- Cost trap: running old and new stacks in parallel doubles hosting cost for the migration period. Budget for it and set an end date.
9.4 Reference architecture (text)
Users (Angular SPA on Azure Static Web Apps / Storage + CDN) | Azure Front Door (TLS, WAF, global routing) | YARP Facade (Azure Container Apps, 2+ replicas, config from App Configuration) |------------------------------|-----------------------------| /api/claims (POST) /api/documents/* everything else | | | Claim Intake Service Document Service ClaimsCore monolith (Container Apps, .NET 10) (Container Apps) (VM/App Service, private) | | | PostgreSQL Flexible Server Blob Storage Azure SQL / SQL Server \____ Service Bus (ClaimSubmitted events, outbox) ____/ | Application Insights + Log Analytics (per-route dashboards)Teaching guide for my team
10.1 Two-minute beginner explanation
“We cannot switch off the claims system to rebuild it. So we put a traffic director in front. Today it sends everything to the old system. Then we build a new claim-submission service and tell the director: submissions go to the new one, everything else stays. We repeat for each feature. When the old system gets no traffic, we delete it. If the new piece breaks, we point the director back to the old one in minutes.”
10.2 Five-minute intermediate explanation
Cover: the facade (YARP) and route ordering; the three steps Transform, Coexist, Eliminate; choosing slices by business capability; canary with feature flags; data sync using outbox events and an anti-corruption layer; the risk that the facade becomes a new monolith; and the rule that every slice ends with deleting legacy code. Show the YARP JSON and the Order property, and demo flipping a flag from 0% to 10%.
10.3 Hands-on exercise
Goal: Build a facade that routes one endpoint to a new service and everything else to a legacy stub.
- Create three projects:
Legacy(returns “legacy” for any path),NewIntake(returns “new intake” forPOST /api/claims), andFacade(YARP). - Configure the two clusters and two routes as in section 6.1.
- Call
POST /api/claimsandGET /api/policies/1through the facade. - Add a
NewIntakefeature flag and confirm that turning it off sendsPOST /api/claimsto the legacy stub.
Expected outcome: POST /api/claims returns “new intake”; GET /api/policies/1 returns “legacy”; with the flag off, both return “legacy”. Trainees see that no client changed.
10.4 Interview-style questions
- What are the three steps of the Strangler Fig pattern? Transform (build new functionality), Coexist (route traffic between old and new), Eliminate (remove the old code once unused).
- Why is the facade a risk and how do you limit it? It is a single point of failure and can grow into a new monolith. Keep it stateless, run multiple instances, and keep business logic out of it.
- How do you keep data consistent between old and new during coexistence? Give the new service its own data, publish changes with a transactional outbox, consume them idempotently through an anti-corruption layer in the legacy side, and reconcile regularly. Alternatively keep the legacy database as the source of truth temporarily.
Mastery checklist
- I can explain why a big-bang rewrite is riskier than incremental replacement, with a concrete example.
- I can configure YARP routes with correct ordering so a specific path overrides a catch-all.
- I can choose migration slices by business capability and justify the order (risk, value, coupling).
- I can design the data sync strategy for a slice, including the anti-corruption layer and reconciliation.
- I can roll out a slice with a feature flag and percentage canary, and define rollback triggers.
- I can identify when the pattern does not apply (no interception point, tiny system).
- I can define exit criteria and make sure the legacy code for each slice is actually deleted.
- I can map the pattern to Azure services and estimate the cost of running both systems in parallel.
Key takeaway
Put a router in front of the monolith and move one slice of traffic at a time to new services, so every step is small, reversible, and ends with deleting legacy code.
