Manikandan — Manikandan
Microservices

Day 29: Fallback

ManikandanManikandan
21 min read·Updated Sep 7, 2022

A Fallback is the "plan B" your service runs when a dependency call fails or is rejected (timeout, open circuit, bulkhead full, 5xx).

Intro

A Fallback is the “plan B” your service runs when a dependency call fails or is rejected (timeout, open circuit, bulkhead full, 5xx). Instead of returning an HTTP 500 to the user, the service returns something acceptable: a safe default, a cached or last-known value, or a reduced version of the feature. The skill is not in writing the catch block; it is in deciding, per dependency, which failures may be hidden, what the substitute is, and how to make the degradation visible to monitoring so you never quietly serve wrong data.

Versions used in this lesson: .NET 10 (LTS, released Nov 2025), Angular 22 (current major), Polly v8 with Microsoft.Extensions.Http.Resilience / Microsoft.Extensions.Resilience, SQL Server / PostgreSQL. Running domain: an insurance claims system.


Why we need this

  • Availability is a product of dependencies. If a “Submit Claim” request touches 5 downstream services, each 99.9% available, the combined ceiling is about 99.5% (0.999^5). Every non-essential dependency you can survive without raises the ceiling.
  • Not every dependency is equally important. In a claims system, the policy coverage check is essential (without it you may pay an invalid claim). The repair-shop rating, fraud-score enrichment, document thumbnail and weather-event tag are useful but not essential. Failing the whole request because a thumbnail service is down is a bad trade.
  • Users judge the product, not the architecture. A claims portal that shows “Rating temporarily unavailable” is fine; one that shows a blank page with a stack trace is a support ticket and a lost customer.
  • It completes the resilience toolkit. Retry handles blips, Circuit Breaker stops hammering, Bulkhead isolates, Rate Limiter protects. Fallback is the last layer that decides what the caller gets once the others have given up.

What problem it solves

Problem (from the topic list): users see 500s when non-critical dependencies fail.

Without it: A GET /claims/{id} handler calls Claims DB (essential), Repair-Shop Rating API (nice-to-have) and Document Preview API (nice-to-have). The rating API deploys a bad build and throws for 40 minutes. Every claim-detail page returns 500, although 100% of the essential data was available. Adjusters cannot work; the incident is a Sev-1 caused by a Sev-3 feature.

With it: the handler still returns claim data, with repairShopRating: null and degraded: ["repairShopRating"]. The UI shows “Rating unavailable”. A metric fallback_invocations_total{dependency="repair-rating"} increases and an alert fires at a threshold. Blast radius: one widget.

When it is needed (and when it is NOT)

Use Fallback when:

  • The dependency is non-critical or enrichment-only (ratings, recommendations, thumbnails, exchange-rate hints, geocoding).
  • A safe substitute exists: last-known value from cache, a static default, a simpler local computation, a “pending” marker, or a queue-for-later.
  • Staleness is acceptable and bounded (e.g., a repair-shop rating that is 1 day old is fine).
  • You have already applied timeout + retry + circuit breaker, and still need a defined outcome.

Do NOT use (or use very carefully) when:

  • The data drives a money or legal decision: coverage verification, payout amount, sanctions/fraud blocking checks. Returning a default “approved” or “score 0” here is a silent defect. The right answer is to fail (or park the claim for manual review), not to guess.
  • The failure is a client error (4xx / validation). Falling back on a 400 or 404 hides bugs. Handle only transient/infrastructure failures.
  • The operation is a write with side effects and you cannot make the substitute idempotent and honest (“we saved it” when you didn’t).
  • A fallback would mask a persistent outage with no alerting. Fallback without telemetry converts an outage into invisible data rot.
  • Fallback code paths are never tested. An untested fallback often fails at the worst moment (e.g., cache is also down).

How to identify the problem (key signals)

  1. Error rate of a page/endpoint tracks the error rate of a minor dependency (in traces or dashboards, the 5xx on /claims/{id} mirrors 5xx on repair-rating).
  2. Stack traces in logs whose root cause is an HttpRequestException / TimeoutRejectedException / BrokenCircuitException from a dependency that only supplies decoration data.
  3. Support tickets like “the whole claim screen is blank” while the primary database is healthy.
  4. Code smell: a controller/service method with 6 sequential await client.GetAsync(...) calls and no try/catch or resilience pipeline around the optional ones.
  5. Availability SLO breaches caused by a dependency with a lower tier (your SLO 99.9%, the vendor’s SLA 99.5%).
  6. Feature flags being flipped manually during incidents (“comment out the rating call and redeploy”) is a sign a fallback should have been designed in.
  7. Opposite smell: a catch (Exception) { return default; } scattered around with no logging or metric. That is not a fallback; it is error swallowing, and it makes the system look healthy while data is wrong.

Flow Diagram

Fallback is the outermost layer and only fires on infrastructure failures of optional data.

flowchart LR
REQ["GET /claims/id"] --> DB["Claims DB - essential, no fallback"]
REQ --> PIPE["Fallback > Retry > Circuit Breaker > Timeout"]
PIPE --> RT["Rating API"]
RT -- "ok" --> LIVE["Live rating"]
RT -- "5xx / timeout / circuit open" --> FB["Fallback"]
FB --> CACHE["Redis / snapshot, max age 24h"]
FB --> NUL["null + degraded flag"]
LIVE --> RESP["Response"]
CACHE --> RESP
NUL --> RESP
DB --> RESP

Level 1: Beginner

Analogy: A restaurant runs out of the fresh tomato sauce. A good waiter does not tell you “the kitchen is broken”. They say “we can serve it with the pesto instead”. The meal (essential) arrives; the sauce (non-essential) is swapped. A bad restaurant would send you home hungry, and a dishonest one would serve you pesto and call it tomato.

Core rule: try the real thing, and if it fails for a transient reason, return a safe, clearly-marked substitute.

Minimal working example (plain C#, no libraries):

public sealed record RepairShopRating(decimal? Stars, bool IsFallback);
public sealed class RatingService(HttpClient http, ILogger<RatingService> log)
{
public async Task<RepairShopRating> GetRatingAsync(int shopId, CancellationToken ct)
{
try
{
var stars = await http.GetFromJsonAsync<decimal>($"ratings/{shopId}", ct);
return new RepairShopRating(stars, IsFallback: false);
}
catch (Exception ex) when (ex is HttpRequestException or TaskCanceledException)
{
log.LogWarning(ex, "Rating service unavailable for shop {ShopId}; using fallback", shopId);
return new RepairShopRating(Stars: null, IsFallback: true);
}
}
}

What to notice: (a) only infrastructure exceptions are caught, (b) the fallback is explicit (IsFallback = true) so the UI can say “unavailable”, (c) it is logged. Using HttpClient without a timeout would make this useless, because a hung call never throws; always pair Fallback with a timeout.

Level 2: Intermediate

6.1 Where fallback sits in the pipeline

Order matters. Strategies added first are the outermost. The standard order for a dependency call:

Fallback → Retry → Circuit Breaker → Timeout (per attempt) → HTTP call
(outermost: runs only after everything inside gave up)

So Fallback sees the final outcome (including BrokenCircuitException when the breaker is open and TimeoutRejectedException), which is exactly what you want.

6.2 .NET 10 + Polly v8: a typed result with a fallback

Packages: Microsoft.Extensions.Http.Resilience (brings Polly v8), Polly.Extensions.

// Program.cs (.NET 10, minimal hosting)
using System.Net.Http.Json;
using Polly;
using Polly.CircuitBreaker;
using Polly.Fallback;
using Polly.Retry;
using Polly.Timeout;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddHttpClient<IRatingClient, RatingClient>(c =>
c.BaseAddress = new Uri(builder.Configuration["RatingApi:BaseUrl"]!));
// Register a named pipeline that returns RepairShopRating
builder.Services.AddScoped<IClaimReadService, ClaimReadService>();
builder.Services.AddResiliencePipeline<string, RepairShopRating>("rating", (pipeline, ctx) =>
{
var logger = ctx.ServiceProvider.GetRequiredService<ILoggerFactory>().CreateLogger("Resilience.Rating");
pipeline
.AddFallback(new FallbackStrategyOptions<RepairShopRating>
{
// Only infrastructure failures, never 4xx/validation errors
ShouldHandle = new PredicateBuilder<RepairShopRating>()
.Handle<HttpRequestException>()
.Handle<TimeoutRejectedException>()
.Handle<BrokenCircuitException>(),
FallbackAction = _ => Outcome.FromResultAsValueTask(new RepairShopRating(null, IsFallback: true)),
OnFallback = args =>
{
logger.LogWarning(args.Outcome.Exception, "Fallback used for rating");
FallbackMetrics.Rating.Add(1); // custom counter, see 6.4
return default;
}
})
.AddRetry(new RetryStrategyOptions<RepairShopRating>
{
MaxRetryAttempts = 2,
Delay = TimeSpan.FromMilliseconds(200),
BackoffType = DelayBackoffType.Exponential,
UseJitter = true,
ShouldHandle = new PredicateBuilder<RepairShopRating>()
.Handle<HttpRequestException>()
.Handle<TimeoutRejectedException>()
})
.AddCircuitBreaker(new CircuitBreakerStrategyOptions<RepairShopRating>
{
FailureRatio = 0.5,
MinimumThroughput = 20,
SamplingDuration = TimeSpan.FromSeconds(30),
BreakDuration = TimeSpan.FromSeconds(15),
ShouldHandle = new PredicateBuilder<RepairShopRating>()
.Handle<HttpRequestException>()
.Handle<TimeoutRejectedException>()
})
.AddTimeout(TimeSpan.FromMilliseconds(800)); // per attempt, innermost
});
var app = builder.Build();
app.MapGet("/api/claims/{id:int}", async (int id, IClaimReadService svc, CancellationToken ct) =>
Results.Ok(await svc.GetAsync(id, ct)));
app.Run();
public sealed record RepairShopRating(decimal? Stars, bool IsFallback);
public interface IRatingClient { Task<RepairShopRating> GetAsync(int shopId, CancellationToken ct); }
public sealed class RatingClient(HttpClient http, ResiliencePipelineProvider<string> pipelines) : IRatingClient
{
public async Task<RepairShopRating> GetAsync(int shopId, CancellationToken ct)
{
var pipeline = pipelines.GetPipeline<RepairShopRating>("rating");
return await pipeline.ExecuteAsync(async token =>
{
var stars = await http.GetFromJsonAsync<decimal>($"ratings/{shopId}", token);
return new RepairShopRating(stars, IsFallback: false);
}, ct);
}
}

The claim read service composes essential and optional data:

public sealed record ClaimDetailDto(
int Id, string Status, decimal ReserveAmount,
RepairShopRating? RepairShop, string[] Degraded);
public interface IClaimReadService { Task<ClaimDetailDto> GetAsync(int id, CancellationToken ct); }
public sealed class ClaimReadService(ClaimsDbContext db, IRatingClient rating) : IClaimReadService
{
public async Task<ClaimDetailDto> GetAsync(int id, CancellationToken ct)
{
// ESSENTIAL: no fallback. If the DB is down, the request must fail.
var claim = await db.Claims.AsNoTracking()
.Where(c => c.Id == id)
.Select(c => new { c.Id, c.Status, c.ReserveAmount, c.RepairShopId })
.SingleAsync(ct);
// OPTIONAL: fallback is inside the pipeline
var shop = await rating.GetAsync(claim.RepairShopId, ct);
var degraded = shop.IsFallback ? new[] { "repairShopRating" } : Array.Empty<string>();
return new ClaimDetailDto(claim.Id, claim.Status, claim.ReserveAmount, shop, degraded);
}
}

(ClaimsDbContext and the Claim entity are assumed to exist in your project.)

6.3 Database: “last known good” table

For values worth keeping across restarts, keep a snapshot next to the source of truth.

-- SQL Server
CREATE TABLE dbo.RepairShopRatingSnapshot (
RepairShopId INT NOT NULL PRIMARY KEY,
Stars DECIMAL(3,2) NOT NULL,
CapturedAtUtc DATETIME2(0) NOT NULL
);
-- Upsert after every successful live call
MERGE dbo.RepairShopRatingSnapshot AS t
USING (VALUES (@RepairShopId, @Stars, SYSUTCDATETIME())) AS s(RepairShopId, Stars, CapturedAtUtc)
ON t.RepairShopId = s.RepairShopId
WHEN MATCHED THEN UPDATE SET Stars = s.Stars, CapturedAtUtc = s.CapturedAtUtc
WHEN NOT MATCHED THEN INSERT (RepairShopId, Stars, CapturedAtUtc)
VALUES (s.RepairShopId, s.Stars, s.CapturedAtUtc);
-- PostgreSQL equivalent
INSERT INTO repair_shop_rating_snapshot (repair_shop_id, stars, captured_at_utc)
VALUES (@repair_shop_id, @stars, now() at time zone 'utc')
ON CONFLICT (repair_shop_id)
DO UPDATE SET stars = EXCLUDED.stars, captured_at_utc = EXCLUDED.captured_at_utc;

The fallback action reads this snapshot and returns it with IsFallback = true and its age. Cap the acceptable age (for example 24 hours); older than that → return null, not stale data.

6.4 Make it visible: a metric

using System.Diagnostics.Metrics;
public static class FallbackMetrics
{
private static readonly Meter Meter = new("Claims.Resilience");
public static readonly Counter<long> Rating =
Meter.CreateCounter<long>("fallback_invocations_total", description: "Times a fallback replaced a live result");
}

(Tag it: Rating.Add(1, new KeyValuePair<string, object?>("dependency", "repair-rating")). Export via OpenTelemetry, see section 9.)

6.5 Angular 22: show degraded data honestly

The API tells the client what degraded; the UI must not pretend.

// claim-detail.component.ts (standalone component, Angular 22)
import { Component, inject, input, computed } from '@angular/core';
import { HttpClient } from '@angular/common/http';
import { toSignal, toObservable } from '@angular/core/rxjs-interop';
import { switchMap, catchError, of } from 'rxjs';
interface ClaimDetail {
id: number; status: string; reserveAmount: number;
repairShop: { stars: number | null; isFallback: boolean } | null;
degraded: string[];
}
@Component({
selector: 'app-claim-detail',
template: `
@if (claim(); as c) {
<h2>Claim {{ c.id }} – {{ c.status }}</h2>
@if (c.degraded.length) {
<p class="banner" role="status">
Some information is temporarily unavailable ({{ c.degraded.join(', ') }}). Core claim data is current.
</p>
}
<p>Repair shop rating:
@if (c.repairShop?.stars != null) { {{ c.repairShop!.stars }} / 5 }
@else { <em>unavailable right now</em> }
</p>
} @else if (failed()) {
<p role="alert">We could not load this claim. Please retry.</p>
}
`,
})
export class ClaimDetailComponent {
private http = inject(HttpClient);
id = input.required<number>();
private result = toSignal(
toObservable(this.id).pipe(
switchMap(id => this.http.get<ClaimDetail>(`/api/claims/${id}`).pipe(
// Client-side last resort: only for essential data failure
catchError(() => of(null))
))
),
{ initialValue: undefined }
);
claim = computed(() => this.result() ?? null);
failed = computed(() => this.result() === null);
}

Key design point: the backend decides what may degrade; the frontend renders it truthfully. The Angular side has its own fallback too (placeholder image via (error) handler, skeleton state), but it never invents business data.

Level 3: Advanced

7.1 Kinds of fallback (choose deliberately)

KindExample in claims domainRisk
Static defaultRating = null, “Unavailable”Lowest; UI must handle null
Cached / last-known-goodYesterday’s rating from RepairShopRatingSnapshotStaleness; needs max-age
Degraded computationLocal rule-based fraud hint instead of ML score, marked “provisional”Lower accuracy; must not auto-approve
Alternative providerSecondary geocoding vendorCost, data-format differences
Deferred completionAccept the claim, queue enrichment, fill laterNeeds async pipeline and idempotency
Fail-safe manual routeFraud service down → route claim to manual review queueSlower, but safe. Usually the right choice for risk checks

7.2 Performance and scalability

  • Fallback must be cheaper and more reliable than the primary. A fallback that calls another remote service simply moves the failure. Prefer in-memory / local data.
  • Fallback amplifies load on the fallback source. When the breaker opens, every request now hits your cache or DB snapshot. Size it for 100% of traffic, not the normal miss rate.
  • Avoid cache stampede on warm-up after an outage: use HybridCache (Microsoft.Extensions.Caching.Hybrid, generally available) or per-key locks so one caller refreshes while others get the stale value.
  • Total time budget: timeout × (retries + 1) + backoff must still fit inside the caller’s own timeout, otherwise the fallback fires after the client has given up.

7.3 Security

  • A fail-open fallback on an authorization or fraud check is a vulnerability. Attackers can deliberately overload a dependency to trigger the fallback. Default rule: security-relevant checks fail closed.
  • Do not log full PII/claim payloads in OnFallback; log IDs and error class.
  • Cached fallback data must respect tenant boundaries (key includes tenant ID).

7.4 Failure modes

  • Fallback also fails. Wrap carefully: the fallback action should not throw. If it can (cache/DB read), define what happens next (typically null + degraded flag).
  • Silent degradation for days. Solved by metrics + alert on fallback rate > X% for Y minutes, and a dashboard showing ”% of responses degraded”.
  • Poisoned cache: never write fallback values back into the cache as if they were real data.
  • Retry storms behind fallback: Fallback hides errors from users, but the dependency may still be under attack from your retries. Keep the circuit breaker.
  • Fallback masks a contract change: a JsonException after an API change gets “handled”, and you serve defaults forever. Do not include deserialization errors in ShouldHandle for long; alert on them separately.

7.5 Common mistakes

  1. Catching Exception and returning default (hides bugs and cancellation). Never handle OperationCanceledException caused by the caller’s token.
  2. Fallback without a timeout (hung calls never trigger it).
  3. Putting Fallback inside Retry (fallback would trigger on the first failure and retry would never run).
  4. Returning fallback data with HTTP 200 and no marker. Consumers cannot distinguish real from substituted.
  5. Applying a fallback to a write/payment call.
  6. Never testing it (see exercise in section 10).

Level 4: Expert and Architect view

8.1 Trade-offs and alternatives

ApproachStrengthWeaknessChoose when
Fallback (default/cached)Simple, immediate, keeps UX aliveData may be stale/absent; can hide outagesOptional enrichment, read paths
Fail fast (no fallback)Honest, simple, no wrong dataUser-visible errorPayments, coverage, authz
Fail-safe manual routeCorrect + safe under uncertaintyHuman cost, latencyFraud/sanctions checks, high-value claims
Graceful degradation via feature flagsOperator control, can turn off featuresManual, slower than automaticKnown long outages, planned maintenance
Async decoupling (events/outbox)Removes sync dependency entirelyEventual consistency, more infrastructureThe dependency is inherently slow or unreliable
Data replication (own read model / CQRS)No runtime call at allSync complexity, storageFrequently-read reference data
Multi-region / active-activeSurvives infra failureCost, data consistencyWhole-service availability needs

8.2 Patterns it combines with

  • Timeout + Retry + Circuit Breaker + Bulkhead (Days 26–28): the standard resilience pipeline; Fallback is the outermost layer.
  • Self-contained Service (Day 3) / CQRS (Day 8): the strongest “fallback” is not needing the remote call because a local read model already has the data.
  • API Composition / BFF (Days 10, 20): composition endpoints are the natural place to define per-widget degradation.
  • Health Check API (Day 36): expose degraded status (not just up/down) for dependencies using fallbacks.
  • Saga (Day 7): for writes, the equivalent of fallback is a compensating action, not a default value.

8.3 ADR (architecture review ready)

ADR-029: Use Fallback for non-critical enrichment dependencies in Claims read APIs

  • Status: Proposed
  • Context (illustrative): GET /claims/{id} aggregates the Claims DB with Repair-Shop Rating and Document Preview services. In the last quarter, two incidents took the claim screen down when only these optional services failed. Our availability SLO is 99.9%; the optional services offer 99.5%.
  • Decision: Wrap optional dependencies in a Polly v8 pipeline (Fallback → Retry → Circuit Breaker → Timeout). Fallback returns a null/snapshot value marked degraded, with max snapshot age 24 h. Essential dependencies (Claims DB, Coverage service) have no fallback. Fraud-scoring failures route the claim to manual review instead of using default scores.
  • Consequences: (+) Optional-service outages no longer affect claim viewing; (+) uniform pattern via a shared chassis; (−) users may see stale/absent data, mitigated by UI markers; (−) additional code paths to test; (−) requires dashboards/alerts on fallback_invocations_total.
  • Alternatives rejected: Fail-fast everywhere (poor UX); copying rating data into the Claims DB (adds sync ownership we do not want yet).
  • Review trigger: Fallback rate > 5% of requests for 30 minutes, or any fallback found on a financial decision path.

Azure implementation

Fallback is mostly an application-level pattern, but Azure provides the infrastructure that makes it possible (cache, snapshot storage, failover routing) and observable.

9.1 Services and roles

ConcernAzure serviceHow it supports Fallback
Host the APIAzure App Service or Azure Container Apps (or AKS)Runs the .NET 10 service with the resilience pipeline
Last-known-good cacheAzure Managed Redis (successor to Azure Cache for Redis, which Microsoft has announced for retirement; check the current retirement FAQ for dates before choosing a tier)Low-latency cache for fallback values; use HybridCache with a Redis L2
Durable snapshotAzure SQL Database or Azure Database for PostgreSQL – Flexible ServerRepairShopRatingSnapshot table (section 6.3)
Edge / gateway fallbackAzure API Managementretry policy plus return-response policy in on-error to send a static/degraded response when the backend fails
Regional failoverAzure Front Door (Standard/Premium)Origin groups with health probes route to a healthy origin; supports priority-based active/passive failover
Static fallback contentAzure Blob Storage static website / Front DoorServe a maintenance/degraded page
ObservabilityAzure Monitor + Application Insights (OpenTelemetry distro)Track fallback counter, dependency failures, alert rules
Config / feature flagsAzure App ConfigurationToggle fallbacks, TTLs, kill-switches without redeploy
Async deferralAzure Service BusQueue enrichment for later when a dependency is down

9.2 How to configure (key points)

  1. Application Insights via OpenTelemetry (.NET 10): add Azure.Monitor.OpenTelemetry.AspNetCore, call builder.Services.AddOpenTelemetry().UseAzureMonitor(), and register the meter: .WithMetrics(m => m.AddMeter("Claims.Resilience")). Set the connection string via App Service/Container Apps configuration (APPLICATIONINSIGHTS_CONNECTION_STRING), stored as a Key Vault reference.
  2. Alert rule: on the custom metric fallback_invocations_total (split by dependency), or a KQL log alert on the OnFallback warning; threshold example: > 50 in 5 minutes, severity 2.
  3. API Management fallback (edge-level): in the API’s outbound/on-error section use retry (condition on backend 5xx / timeouts) and then return-response with a small JSON body and an X-Degraded: true header. Note: policy behaviour and availability differ by tier; verify in the API Management policy reference for your tier.
  4. Front Door: create an origin group with a health probe path (for example /health/ready), set priorities (primary = 1, secondary region = 2) or use latency-based routing; configure sample size / successful samples required so a flapping origin does not thrash.
  5. Redis / HybridCache: set a long fallback TTL separate from the normal TTL (e.g., normal 5 min, stale-allowed 24 h), with tenant-scoped keys.
  6. App Configuration: store Fallback:Rating:MaxStaleHours and a Fallback:Rating:Enabled flag to disable in case the fallback itself misbehaves.

9.3 Pricing and tier considerations

Prices change and vary by region; confirm on the Azure pricing pages before budgeting.

  • App Service / Container Apps: no extra cost for the pattern itself; you pay for compute you already run. Container Apps has a consumption plan and dedicated options.
  • Azure Managed Redis: priced by memory/performance tier. For fallback caches, start with the smallest tier that holds the “hot set” of last-known-good values, and size for full read traffic when breakers are open.
  • Azure SQL / PostgreSQL snapshot table: negligible storage; runs on the DB you already have.
  • API Management: Consumption tier is pay-per-call; Developer/Basic/Standard/Premium and the v2 tiers are capacity-based. Choose based on VNet, SLA and multi-region needs, not on this pattern.
  • Front Door: Standard and Premium have a monthly base fee plus request/data charges; Premium adds Private Link origins and advanced WAF. Multi-region failover means paying for the second region’s compute (or scale it down in a warm-standby).
  • Application Insights / Log Analytics: billed by data ingested; keep the OnFallback log at Warning with small payloads and use metrics for alerting to control cost.

9.4 Reference architecture (text)

  1. The Angular 22 SPA is served from Azure Static Web Apps or Blob + Front Door, calling https://api.claims.example through Azure Front Door.
  2. Front Door routes to the primary region’s API Management (or directly to the Container App), with a secondary region as lower-priority origin.
  3. The Claims Read API (.NET 10, Container Apps) executes the request: Claims DB read (no fallback) + optional calls to Rating and Document Preview services via Polly pipelines (Fallback → Retry → Circuit Breaker → Timeout).
  4. On success, the API writes the live value to Azure Managed Redis and to the SQL snapshot table. On failure, the fallback reads Redis first, then the snapshot, else returns null with degraded.
  5. Fraud scoring runs via Service Bus: if the Fraud service is unavailable, the message is parked and the claim gets a “manual review” state.
  6. OpenTelemetry → Application Insights collect traces, the fallback counter and dependency failures; Azure Monitor alerts page the on-call when fallback rate crosses the threshold; App Configuration holds kill-switches; secrets live in Key Vault.

Teaching guide for my team

10.1 Explain to a beginner in 2 minutes

“When our app calls another service and it fails, we have two options: crash the whole page, or show what we can. Fallback means we decide in advance what to show instead: a default, the last value we saw, or a message saying ‘not available right now’. The rule of thumb: if the missing piece is decoration (ratings, images), use a fallback; if it’s about money or safety (coverage, fraud), don’t guess, fail or send it to a human. And always leave a note (a log and a metric), otherwise nobody knows the fallback is running.”

10.2 Explain to an intermediate developer in 5 minutes

  1. Draw the call chain: API → Claims DB (essential) and API → Rating, API → Preview (optional).
  2. Show the pipeline order: Fallback → Retry → Circuit Breaker → Timeout. Explain why Fallback is outermost: it should only fire after retry and breaker finished.
  3. Show ShouldHandle: only HttpRequestException, TimeoutRejectedException, BrokenCircuitException; never 4xx, never caller cancellation.
  4. Show the response contract: degraded: ["repairShopRating"] so the UI is honest.
  5. Show the three safety nets: OnFallback logging, fallback_invocations_total metric, alert threshold.
  6. Close with the decision table: static default vs cached vs alternate provider vs manual route, and the “fail closed for security/money” rule.

10.3 Hands-on exercise

Task: In a sample .NET 10 API, implement GET /claims/{id} that reads a claim from an in-memory list and calls a fake RatingClient (a small second minimal API you run locally).

  1. Add the Polly pipeline from section 6.2 with a fallback returning RepairShopRating(null, true).
  2. Add degraded to the response.
  3. Stop the fake Rating API, then call /claims/1.
  4. Add a unit test that injects an HttpMessageHandler that throws HttpRequestException, and assert IsFallback == true and that the counter increments.
  5. Make the fake Rating API return 400; assert that the fallback is not used (the error surfaces).

Expected outcome: Rating API up → real stars, degraded: []. Rating API down → HTTP 200 with repairShop.stars = null, degraded: ["repairShopRating"], one warning log, counter > 0. Rating API returns 400 → request fails (bug surfaced), proving ShouldHandle is correctly scoped.

10.4 Interview-style questions

  1. Why must Fallback be the outermost strategy in a Polly pipeline? Because it should run only on the final outcome after retries, circuit breaker and timeout have finished. If it is inside Retry, it would “succeed” on the first failure and retries never happen; if the breaker is outside it, the fallback wouldn’t see BrokenCircuitException consistently.
  2. Give an example where a fallback is dangerous and what to do instead. Returning “fraud score = 0” when the fraud service is down would let risky claims through automatically. Instead, fail closed by routing the claim to manual review (or reject/park it) and alert.
  3. How do you make sure a fallback isn’t silently hiding a permanent outage? Emit a metric and a warning log on every fallback invocation, mark responses as degraded, alert when the fallback rate exceeds a threshold for a period, and cap the age of cached values.

Mastery checklist

  • I can classify each dependency of a service as essential vs optional and justify it.
  • I can name at least four fallback types and pick the right one for a given dependency.
  • I can build a Polly v8 pipeline (Fallback → Retry → Circuit Breaker → Timeout) and explain the ordering.
  • I can restrict ShouldHandle to transient/infrastructure failures and exclude 4xx and caller cancellation.
  • I can design the API contract so degraded responses are explicit and the Angular UI renders them honestly.
  • I can add a metric, log and alert so fallbacks are observable.
  • I can explain why security- and money-related checks fail closed, with an example.
  • I have tested the fallback path (unit test and a real dependency-down test) and know the fallback itself can fail.

Key takeaway

Fallback turns “a minor dependency failed” into “one widget is degraded”, but only when you choose the substitute deliberately, keep it honest and observable, and never use it to guess on decisions involving money or safety.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 30: Rate Limiter / Throttling

A rate limiter decides, per caller, how many requests the system will accept in a given time (or how many it will process at once) and rejects or delays the rest, usually with HTTP 429 and a Retry-After header.

Manikandan
Manikandan·20 min read
Microservices

Day 28: Retry & Backoff

Retry & Backoff means that when a call to another service or resource fails with a *transient* error (a network blip, a 503, a throttling response, a deadlock victim), the caller waits a short, growing, randomised...

Manikandan
Manikandan·18 min read
Microservices

Day 27: Bulkhead

The Bulkhead pattern splits a service's shared resources (threads, connections, memory, queue slots, CPU) into isolated compartments, so that when one dependency or one traffic flow becomes slow or overloaded, only...

Manikandan
Manikandan·19 min read
Microservices

Day 26: Circuit Breaker

A circuit breaker sits between your service and a dependency it calls (a fraud-scoring API, a payment gateway, another microservice).

Manikandan
Manikandan·20 min read