Manikandan — Manikandan
Microservices

Day 26: Circuit Breaker

ManikandanManikandan
20 min read·Updated Sep 4, 2022

A circuit breaker sits between your service and a dependency it calls (a fraud-scoring API, a payment gateway, another microservice).

Intro

A circuit breaker sits between your service and a dependency it calls (a fraud-scoring API, a payment gateway, another microservice). It watches the calls. When too many fail, it “trips” and immediately rejects further calls for a short cool-down instead of letting every request wait and fail. After the cool-down it lets a few probe calls through; if they succeed the circuit closes and traffic resumes. The result: a sick dependency cannot drag your healthy service down with it, and it gets breathing room to recover.

Versions used in this lesson: .NET 10 (current LTS), Polly 8.x with Microsoft.Extensions.Http.Resilience, Angular 22 (current major at the time of writing), SQL Server / PostgreSQL. Running example: an insurance claims system (Claims API calling Fraud Scoring, Policy, and Payment services).

Why we need this

In a microservice system every synchronous call is a place where failure can spread. A dependency rarely fails cleanly; it usually gets slow first. A slow dependency is more dangerous than a dead one, because a dead one fails fast (connection refused) while a slow one holds your threads, connections, and memory until you run out.

Business reasons: adjusters keep working even when a third-party fraud vendor has an outage; a payment gateway incident does not take down claim intake; you avoid hammering a struggling system with retries, which is how a 5-minute blip becomes a 2-hour outage.

Technical reasons: bounded resource usage (threads, sockets, HttpClient connections), fast failure with a clear error instead of a 30-second hang, automatic recovery without a human restarting anything, and a single place to emit “dependency X is unhealthy” telemetry.

What problem it solves

Problem statement: cascading timeouts exhaust upstream threads and memory. The circuit breaker trips open past a failure threshold to fail fast during recovery.

Without it, consider the Claims API calling Fraud Scoring on every claim submission. Fraud Scoring starts taking 30 seconds per call because its database is saturated. Claims API requests pile up waiting; the ASP.NET Core thread pool and the outbound connection pool fill; the Claims API stops answering even for endpoints that never call Fraud Scoring (health checks, claim lookups). The API gateway sees timeouts, clients retry, load multiplies, and now Claims API is down too. One slow service took out two.

With a breaker: after, say, 50% failures over 30 seconds, the circuit opens. For the next 20 seconds every fraud call fails in microseconds with BrokenCircuitException. The Claims API applies its fallback (queue the claim for later scoring), stays responsive, and Fraud Scoring sees zero traffic while it recovers.

When it is needed (and when it is NOT)

Needed when:

  • A call crosses a network boundary to something you do not control or cannot scale instantly (third-party APIs, legacy mainframe adapters, shared services).
  • The caller has a sensible degraded behaviour: queue it, use a cached value, show partial data, or return a fast, clear error.
  • Traffic is high enough that failures accumulate quickly (tens of calls per sampling window or more).
  • The dependency has known bad days: rate limiting (HTTP 429), nightly batch windows, deployments.

Not needed or wrong when:

  • In-process calls or calls to your own database via a normal connection pool. Use timeouts, pool limits, and retries for transient errors; a breaker around every SQL call is noise.
  • Very low traffic (a nightly job making 10 calls). The breaker never gathers enough samples to decide; use a plain retry with a timeout.
  • Failures caused by the request itself (HTTP 400, 404, 422). These say nothing about dependency health and must not count toward tripping.
  • There is no acceptable fallback and failing fast is no better than waiting. Then a breaker only changes the error message; still useful for protecting resources, but do not expect availability gains.
  • A message-based flow already provides buffering: a consumer reading from Azure Service Bus at its own pace does not need a breaker on the queue, only on the calls the handler makes downstream.

How to identify the problem (key signals)

  1. Thread-pool starvation or HttpClient port exhaustion in the caller when a downstream service is slow (dotnet-counters shows growing ThreadPool Queue Length; SocketException / “An attempt was made to access a socket…” errors).
  2. p99 latency of an unrelated endpoint rises whenever one downstream service degrades.
  3. Logs show the same timeout stack trace repeated hundreds of times per minute against one host, all with identical durations (e.g., exactly 30,000 ms).
  4. Incident timelines where the caller’s error rate follows the dependency’s error rate with a 1-2 minute lag, and recovery takes far longer than the dependency’s own recovery (“retry storm”).
  5. Application Insights dependency telemetry shows a dependency with success rate below ~90% while request rate to it stays flat: nobody is backing off.
  6. The dependency’s team asks you to “stop calling us while we recover” (or their rate limiter is returning 429 with Retry-After and you keep calling).
  7. Code smell: try/catch around an HTTP call that just logs and retries in a while loop, with no state shared between calls.

Flow Diagram

Closed, open and half-open states of a circuit breaker.

stateDiagram-v2
[*] --> Closed
Closed --> Open: failure ratio over threshold
Open --> HalfOpen: break duration elapsed
HalfOpen --> Closed: probe succeeds
HalfOpen --> Open: probe fails
Closed --> Closed: calls succeed

Level 1: Beginner

Analogy: the electrical breaker in your home. When a faulty appliance draws too much current, the breaker flips off so the wiring does not burn. You fix the appliance, then flip it back. A software breaker does the same, but flips itself back after a cool-down.

Three states:

  • Closed: normal. Calls go through, failures are counted.
  • Open: tripped. Calls are rejected immediately without touching the dependency.
  • Half-open: after the break duration, a trial call is allowed. Success closes the circuit; failure re-opens it.

Minimal working example (console app, dotnet add package Polly.Core). It simulates a fraud API that is down for the first 8 seconds:

using Polly;
using Polly.CircuitBreaker;
var outageEndsAt = DateTime.UtcNow.AddSeconds(8);
var pipeline = new ResiliencePipelineBuilder()
.AddCircuitBreaker(new CircuitBreakerStrategyOptions
{
FailureRatio = 0.5, // trip at 50% failures...
MinimumThroughput = 4, // ...but only after at least 4 calls
SamplingDuration = TimeSpan.FromSeconds(10), // measured over 10 seconds
BreakDuration = TimeSpan.FromSeconds(5), // stay open for 5 seconds
ShouldHandle = new PredicateBuilder().Handle<HttpRequestException>(),
OnOpened = args => { Console.WriteLine($" >> OPEN for {args.BreakDuration}"); return default; },
OnHalfOpened = _ => { Console.WriteLine(" >> HALF-OPEN: sending a probe"); return default; },
OnClosed = _ => { Console.WriteLine(" >> CLOSED: dependency healthy"); return default; }
})
.Build();
for (var i = 1; i <= 15; i++)
{
try
{
var score = await pipeline.ExecuteAsync(async ct =>
{
await Task.Delay(20, ct);
if (DateTime.UtcNow < outageEndsAt)
throw new HttpRequestException("Fraud API returned 503");
return 0.12m;
});
Console.WriteLine($"Call {i}: fraud score {score}");
}
catch (BrokenCircuitException)
{
Console.WriteLine($"Call {i}: rejected instantly (circuit open)");
}
catch (HttpRequestException ex)
{
Console.WriteLine($"Call {i}: dependency failed ({ex.Message})");
}
await Task.Delay(1000);
}

Expected output shape: four failures, then OPEN, several instant rejections, then HALF-OPEN, and once the simulated outage is over the probe succeeds and CLOSED appears. Exact call numbers vary with timing.

Level 2: Intermediate

In a real .NET service you do not hand-build pipelines around every call. You attach the breaker to a typed HttpClient using Microsoft.Extensions.Http.Resilience, so every call through that client is protected.

Program.cs of the Claims API (excerpt):

using Microsoft.Extensions.Http.Resilience;
using Polly;
using Polly.CircuitBreaker;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddHttpClient<IFraudScoreClient, FraudScoreClient>(c =>
c.BaseAddress = new Uri(builder.Configuration["Fraud:BaseUrl"]!))
.AddResilienceHandler("fraud-scoring", pipeline =>
{
// Order matters: outermost first.
pipeline.AddRetry(new HttpRetryStrategyOptions
{
MaxRetryAttempts = 2,
BackoffType = DelayBackoffType.Exponential,
UseJitter = true,
Delay = TimeSpan.FromMilliseconds(300)
});
pipeline.AddCircuitBreaker(new HttpCircuitBreakerStrategyOptions
{
FailureRatio = 0.5,
MinimumThroughput = 20,
SamplingDuration = TimeSpan.FromSeconds(30),
BreakDuration = TimeSpan.FromSeconds(20)
});
pipeline.AddTimeout(TimeSpan.FromSeconds(3)); // per attempt, innermost
});
builder.Services.AddDbContext<ClaimsDb>(o =>
o.UseSqlServer(builder.Configuration.GetConnectionString("Claims"))); // or UseNpgsql(...)
var app = builder.Build();
app.MapPost("/api/claims/{id:guid}/submit", async (
Guid id, IFraudScoreClient fraud, ClaimsDb db, CancellationToken ct) =>
{
var claim = await db.Claims.FindAsync([id], ct);
if (claim is null) return Results.NotFound();
try
{
var score = await fraud.GetScoreAsync(id, ct);
claim.FraudBand = score.Band;
claim.Status = score.Band == "High" ? "ManualReview" : "Approved";
}
catch (BrokenCircuitException)
{
// Fallback: accept the claim, score it later. Business stays open.
claim.FraudBand = "Unknown";
claim.Status = "PendingFraudCheck";
}
await db.SaveChangesAsync(ct);
return Results.Ok(new { claim.Id, claim.Status });
});
app.Run();
public sealed record FraudScore(decimal Score, string Band);
public interface IFraudScoreClient
{
Task<FraudScore> GetScoreAsync(Guid claimId, CancellationToken ct);
}
public sealed class FraudScoreClient(HttpClient http) : IFraudScoreClient
{
public async Task<FraudScore> GetScoreAsync(Guid claimId, CancellationToken ct) =>
await http.GetFromJsonAsync<FraudScore>($"scores/{claimId}", ct)
?? throw new InvalidOperationException("Empty fraud score response");
}
public sealed class Claim
{
public Guid Id { get; set; }
public string Status { get; set; } = "Submitted";
public string? FraudBand { get; set; }
public DateTime SubmittedAtUtc { get; set; } = DateTime.UtcNow;
}
public sealed class ClaimsDb(DbContextOptions<ClaimsDb> options) : DbContext(options)
{
public DbSet<Claim> Claims => Set<Claim>();
}

A background worker later re-scores claims parked in PendingFraudCheck. Index it so the poll is cheap (works on SQL Server and PostgreSQL):

CREATE INDEX IX_Claims_PendingFraud
ON Claims (SubmittedAtUtc)
WHERE Status = 'PendingFraudCheck';
-- SQL Server
SELECT TOP (100) Id FROM Claims WHERE Status = 'PendingFraudCheck' ORDER BY SubmittedAtUtc;
-- PostgreSQL
SELECT "Id" FROM "Claims" WHERE "Status" = 'PendingFraudCheck' ORDER BY "SubmittedAtUtc" LIMIT 100;

For a dependency with no acceptable fallback (payments), return a fast, honest 503 with Retry-After so the UI can react. BrokenCircuitException.RetryAfter tells you how long the circuit stays open:

// Excerpt from a minimal API handler that receives HttpContext ctx and calls IPaymentClient.
catch (BrokenCircuitException ex)
{
var seconds = (int)Math.Ceiling((ex.RetryAfter ?? TimeSpan.FromSeconds(15)).TotalSeconds);
ctx.Response.Headers.RetryAfter = seconds.ToString();
return Results.Problem("Payment service is temporarily unavailable.", statusCode: 503);
}

Angular 22 side: an interceptor turns 503 + Retry-After into UI state held in a signal, so the “Pay” button disables itself and re-enables when the window ends.

import { Injectable, computed, signal, inject } from '@angular/core';
import { HttpErrorResponse, HttpInterceptorFn } from '@angular/common/http';
import { catchError, throwError } from 'rxjs';
@Injectable({ providedIn: 'root' })
export class ServiceStatusStore {
private readonly degradedUntil = signal(0);
readonly isDegraded = computed(() => this.degradedUntil() > 0);
markDegraded(seconds: number): void {
this.degradedUntil.set(Date.now() + seconds * 1000);
setTimeout(() => this.degradedUntil.set(0), seconds * 1000);
}
}
export const degradedInterceptor: HttpInterceptorFn = (req, next) => {
const status = inject(ServiceStatusStore);
return next(req).pipe(
catchError((err: HttpErrorResponse) => {
if (err.status === 503) {
const retryAfter = Number(err.headers.get('Retry-After') ?? 15);
status.markDegraded(Number.isFinite(retryAfter) ? retryAfter : 15);
}
return throwError(() => err);
})
);
};
// app.config.ts
// provideHttpClient(withInterceptors([degradedInterceptor]))
<button (click)="pay()" [disabled]="status.isDegraded()">Pay claim</button>
@if (status.isDegraded()) {
<p class="notice">Payments are temporarily unavailable. Please try again shortly.</p>
}

If the API is on a different origin, the browser only exposes Retry-After to JavaScript when the API sends Access-Control-Expose-Headers: Retry-After.

Level 3: Advanced

Performance and scalability

  • The breaker is per process. With 10 replicas, each keeps its own state and trips on its own sample. That is fine and even desirable (each instance protects itself), but a low-traffic replica may take longer to trip. Do not try to share state through Redis; the extra hop defeats the purpose.
  • Configure one breaker per dependency (per named/typed HttpClient), never one global breaker. A single pipeline shared across unrelated hosts lets one bad host trip calls to healthy ones.
  • Sampling maths: MinimumThroughput must be reachable within SamplingDuration. The defaults (100 calls in 30 s, 10% failure ratio) suit a busy service; a service doing 20 calls a minute will never trip. Size these from real traffic. For the standard resilience handler, SamplingDuration must be at least double the attempt timeout, and options validation enforces that.
  • Count slow calls as failures by placing the attempt timeout inside the breaker: a timeout surfaces as TimeoutRejectedException, which the breaker counts. Otherwise a “slow but 200 OK” dependency never trips it.

Strategy order (outer to inner) for HTTP: total timeout, retry, circuit breaker, attempt timeout. Retry outside the breaker means each attempt is counted, and once the circuit opens the remaining attempts fail instantly. Make sure your retry predicate does not retry BrokenCircuitException, otherwise you spend your backoff delays waiting for an open circuit.

Security

  • Only count server-side failures (5xx, 408, 429, timeouts, network errors). If 4xx counts, an attacker or a buggy client sending bad requests can trip the breaker and deny service to everyone (a self-inflicted denial of service).
  • Any operational endpoint that isolates or closes a circuit manually (CircuitBreakerManualControl.IsolateAsync() / CloseAsync()) must be authenticated and authorised; treat it like a kill switch.
  • Do not put claim data or identifiers in breaker callbacks’ log messages beyond the dependency name.

Failure modes and common mistakes

  • Breaker without fallback: users see errors faster, but availability does not improve. Design the fallback first (queue, cache, default, partial page).
  • Fallback that lies: returning “Approved” when fraud scoring is unavailable is a business risk, not a resilience win. Fallbacks must be business-approved (PendingFraudCheck, not Approved).
  • Thundering herd on recovery: in half-open state Polly allows a limited probe, but when the circuit closes all callers rush in. Use jittered break durations (BreakDurationGenerator) so replicas do not all probe in the same second, and combine with a rate limiter or bulkhead.
  • Flapping: a break duration that is too short (the 5 s default) against a dependency that needs 60 s to warm up causes open/half-open/open loops. Increase the break duration or use a generator that grows on repeated trips.
  • Tripping on business exceptions (e.g., “claim not found” modelled as an exception). Configure ShouldHandle explicitly.
  • No visibility: nobody knows the circuit opened until customers complain. Polly’s built-in telemetry emits OnCircuitOpened / OnCircuitHalfOpened / OnCircuitClosed events to ILogger and metrics via System.Diagnostics.Metrics; export them with OpenTelemetry and alert on opens.
  • Tests: use CircuitBreakerManualControl or a TimeProvider (ResiliencePipelineBuilder.TimeProvider) to make breaker tests deterministic instead of sleeping.

Level 4: Expert and Architect view

Related patterns compared (they are complementary, not competing):

PatternQuestion it answersReacts toWhere it livesWeakness alone
TimeoutHow long do I wait for one call?A single slow callClient libraryStill lets thousands of slow calls pile up
Retry with backoffWas that failure transient?One failed attemptClient libraryAmplifies load on a struggling dependency
Circuit BreakerIs this dependency currently unhealthy?A trend across many callsClient library, sidecar, gatewayNeeds a fallback to add availability
BulkheadHow much of my capacity can this dependency consume?ConcurrencyClient library / infraDoes not stop wasted calls to a dead service
Rate limiterHow much load do I accept or send?VolumeGateway / server / clientSays nothing about dependency health
FallbackWhat do I return instead?Any failure or open circuitCallerWrong fallback can be worse than an error
Mesh outlier detectionWhich instance is misbehaving?Consecutive 5xx per instanceEnvoy / IstioEjects instances, not whole-dependency awareness
Health-check-based load balancingIs the instance alive?Periodic probeLB / orchestratorSlow to react; probes can pass while real calls fail

Placement options for the breaker: in-code with Polly (most control, per-call fallback logic, language specific); in a sidecar or mesh (Dapr, Istio; language-agnostic, but no business fallback); at the gateway (APIM backend circuit breaker; protects backends from all clients, but the fallback is generic 503). In practice combine them: mesh/gateway for coarse protection, in-code for business-aware fallbacks.

Combines with: Retry & Backoff (Day 28), Bulkhead (Day 27), Fallback (Day 29), Rate Limiter (Day 30), Health Check API (Day 36), Distributed Tracing (Day 31), API Gateway (Day 19), and Transactional Outbox (Day 12) as the durable buffer behind the fallback.

ADR (architecture review style)

  • Title: ADR-026 Use circuit breakers on synchronous outbound calls from Claims API.
  • Status: Proposed.
  • Context: Claims API synchronously calls Fraud Scoring (third party, 99.5% SLA), Policy, and Payment. In the last two quarters, two fraud vendor slowdowns caused Claims API thread-pool starvation and a full claims-intake outage of about 40 minutes each.
  • Decision: Every outbound HTTP dependency gets its own typed HttpClient with a resilience pipeline of retry (max 2, jittered), circuit breaker (50% failure ratio over 30 s, minimum throughput sized per dependency, 20-30 s break), and a 3 s attempt timeout, using Microsoft.Extensions.Http.Resilience. Fraud Scoring falls back to PendingFraudCheck and asynchronous re-scoring; Payment returns 503 with Retry-After. Breaker state changes are exported as OpenTelemetry metrics and alerted on.
  • Consequences (positive): bounded blast radius, faster recovery, clear dependency-health telemetry. (Negative): more configuration to tune per dependency, a business-approved degraded mode must be defined and tested, per-instance state means slightly inconsistent behaviour across replicas.
  • Alternatives considered: gateway-only breaker (rejected: no business fallback), Istio-only (rejected: not all environments run a mesh; complementary later), no breaker with just timeouts (rejected: does not prevent retry storms or resource exhaustion).

Azure implementation

Azure services that implement or support the pattern

LayerAzure optionNotes
In-codePolly + Microsoft.Extensions.Http.Resilience on App Service, Container Apps, AKS, FunctionsFree (open source). You pay only for the compute. Most control.
GatewayAzure API Management backend circuit breakerSupported on all tiers except Consumption. One rule per backend.
PlatformAzure Container Apps service-discovery resiliency (circuit breaker policy)Preview at the time of writing. Ejects failing replicas from load balancing.
SidecarDapr resiliency policies (Container Apps managed Dapr or AKS Dapr extension)Timeouts, retries, and circuit breakers for service invocation, declared in YAML.
MeshIstio-based service mesh add-on for AKS (outlier detection)Ejects unhealthy pods; combine with connection-pool limits.
ObservabilityAzure Monitor / Application Insights via OpenTelemetryAlert when circuits open.

Configuration

  1. API Management backend circuit breaker (Bicep). When it trips, APIM stops sending requests to the backend and returns 503. Rules are evaluated per gateway instance, so behaviour is approximate across a scaled-out gateway. Use the backend API version your tooling supports; the documented example uses a preview version.
resource backend 'Microsoft.ApiManagement/service/backends@2025-03-01-preview' = {
name: 'apim-claims/fraud-scoring'
properties: {
url: 'https://fraud.example.com'
protocol: 'http'
circuitBreaker: {
rules: [
{
name: 'tripOn5xx'
failureCondition: {
count: 5
interval: 'PT1M'
statusCodeRanges: [ { min: 500, max: 599 } ]
}
tripDuration: 'PT30S'
acceptRetryAfter: true
}
]
}
}
}

Route to it in a policy with <set-backend-service backend-id="fraud-scoring" />. Keep acceptRetryAfter on so a dependency that answers 429 with Retry-After (Azure OpenAI style) is honoured.

  1. Container Apps circuit breaker policy (preview). Properties of the policy:
circuitBreakerPolicy: {
consecutiveErrors: 5 // consecutive errors before a replica is ejected
intervalInSeconds: 10 // how often ejection/restoration is evaluated
maxEjectionPercent: 50 // never eject more than half the replicas (at least one can be ejected)
}
  1. Dapr resiliency (YAML, applied to the Claims API app):
apiVersion: dapr.io/v1alpha1
kind: Resiliency
metadata:
name: claims-resiliency
scopes:
- claims-api
spec:
policies:
timeouts:
fraudTimeout: 3s
retries:
fraudRetry:
policy: exponential
maxInterval: 5s
maxRetries: 2
circuitBreakers:
fraudCB:
maxRequests: 1 # probes allowed in half-open
interval: 30s # counter reset interval in closed state
timeout: 20s # time to stay open
trip: consecutiveFailures > 5
targets:
apps:
fraud-scoring:
timeout: fraudTimeout
retry: fraudRetry
circuitBreaker: fraudCB
  1. Istio on AKS (DestinationRule):
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata:
name: fraud-scoring
spec:
host: fraud-scoring.claims.svc.cluster.local
trafficPolicy:
connectionPool:
http:
http1MaxPendingRequests: 50
http2MaxRequests: 200
outlierDetection:
consecutive5xxErrors: 5
interval: 10s
baseEjectionTime: 30s
maxEjectionPercent: 50
  1. Monitoring. Use Azure Monitor OpenTelemetry (UseAzureMonitor()), keep Polly telemetry on, and alert on circuit-open events. Example KQL for a log alert:
traces
| where timestamp > ago(1h)
| where message has "OnCircuitOpened"
| summarize opens = count() by cloud_RoleName, bin(timestamp, 5m)
| where opens > 0

Pricing and tier considerations

  • In-code (Polly): no licence cost; costs are the compute you already pay for.
  • API Management: the circuit breaker is not available on the Consumption tier, so it needs one of the dedicated tiers. Tier prices differ by region and change over time; this lesson does not quote figures, so check the Azure pricing calculator before choosing. Note that APIM is a shared edge component, so a rule there affects all consumers of that backend.
  • Container Apps: the resiliency policy is a platform setting on the app; charges follow the normal Container Apps plan (consumption or dedicated). It is in preview, so avoid depending on it for production-critical protection without an SLA review.
  • AKS with Istio or Dapr: the add-ons consume extra CPU and memory per pod for sidecars; budget for that in node sizing.
  • Application Insights / Log Analytics: ingestion is billed per GB. Alert on circuit events rather than logging every rejected call.

Reference architecture (text)

Angular 22 SPA (Static Web Apps or Front Door) calls the Claims API through Azure API Management. The Claims API runs on Azure Container Apps (or AKS) with a typed HttpClient per dependency, each with a Polly pipeline (retry, circuit breaker, timeout). Fraud Scoring is an external vendor reached through an APIM backend that also has a circuit breaker rule for coarse protection. Claims and the “PendingFraudCheck” queue live in Azure SQL Database or Azure Database for PostgreSQL; a background worker (Container Apps job or hosted service) re-scores parked claims. Secrets and endpoints come from Key Vault and App Configuration. OpenTelemetry exports traces, metrics, and Polly events to Application Insights; an Azure Monitor alert pages the on-call engineer when a circuit opens, and a dashboard shows breaker state per dependency next to the deployment markers.

Teaching guide for my team

2-minute explanation for a beginner

“Imagine you keep calling a friend who never picks up. After a few missed calls you stop calling for a while, so you do not waste your day and you do not annoy them. Later you try once. If they answer, you go back to normal. A circuit breaker is that rule for code. It has three states: closed (calling normally), open (not calling, failing instantly), half-open (trying one call). It protects our service from waiting on a broken dependency.”

5-minute explanation for an intermediate developer

Start from the fraud-scoring outage story: slow dependency, thread-pool starvation, retries multiplying load. Show the pipeline order (retry, breaker, attempt timeout) and why the timeout is inside the breaker. Explain the four numbers: failure ratio, minimum throughput, sampling duration, break duration, and why 4xx must not count. Show the fallback decision: which business outcomes are safe when the circuit is open. Close with observability: how you learn a circuit opened and how it maps to the dashboard.

Hands-on exercise

Build a tiny “Fraud API” minimal API that returns 200 normally and 503 for 30 seconds after POST /chaos/start. Build a Claims API endpoint that calls it through a typed HttpClient with the Level 2 pipeline. Drive load with hey/bombardier or a for loop hitting /submit.

Expected outcome: before chaos, all calls succeed. After chaos starts you see a burst of failures, then Polly logs OnCircuitOpened, and response times for /submit drop to a few milliseconds with status PendingFraudCheck. Fraud API request logs show near-zero traffic during the open window, then a single probe, then normal traffic once chaos ends and OnCircuitClosed appears. Bonus: point the Angular button at the payment endpoint and watch it disable itself for the Retry-After period.

Interview-style questions

  1. Why is a slow dependency worse than a dead one? A dead one fails immediately, freeing resources; a slow one holds threads and connections until the caller is exhausted.
  2. Where does the circuit breaker go relative to retry and timeout? Retry outside, breaker in the middle, attempt timeout inside, so timeouts count as failures and open-circuit rejections are not retried.
  3. Which responses should not count as failures? Client errors like 400, 404, 422, and cancellations by the caller; they say nothing about dependency health and could be used to trip the circuit deliberately.

Mastery checklist

  • You can draw the closed, open, and half-open transitions and say what triggers each.
  • You can explain failure ratio, minimum throughput, sampling duration, and break duration, and size them from a service’s actual request rate.
  • You can place retry, breaker, and timeout in the right order and justify it.
  • You can define a business-approved fallback for at least two dependencies and say when “fail fast with 503” is the right answer instead.
  • You can implement it on a typed HttpClient in .NET 10 and test it deterministically with manual control or a fake TimeProvider.
  • You can surface circuit state to Angular users through Retry-After and a signal-based store.
  • You can compare in-code, APIM, Dapr, and Istio circuit breaking and pick one for a given team and platform.
  • You can set up an alert that fires when a circuit opens and explain what the on-call engineer should check first.

Key takeaway

A circuit breaker turns a slow, cascading failure into a fast, contained one by refusing to call a dependency that is clearly struggling. It only pays off when paired with a sensible fallback, correctly sized thresholds, and an alert that tells you it opened.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 30: Rate Limiter / Throttling

A rate limiter decides, per caller, how many requests the system will accept in a given time (or how many it will process at once) and rejects or delays the rest, usually with HTTP 429 and a Retry-After header.

Manikandan
Manikandan·20 min read
Microservices

Day 29: Fallback

A Fallback is the "plan B" your service runs when a dependency call fails or is rejected (timeout, open circuit, bulkhead full, 5xx).

Manikandan
Manikandan·21 min read
Microservices

Day 28: Retry & Backoff

Retry & Backoff means that when a call to another service or resource fails with a *transient* error (a network blip, a 503, a throttling response, a deadlock victim), the caller waits a short, growing, randomised...

Manikandan
Manikandan·18 min read
Microservices

Day 27: Bulkhead

The Bulkhead pattern splits a service's shared resources (threads, connections, memory, queue slots, CPU) into isolated compartments, so that when one dependency or one traffic flow becomes slow or overloaded, only...

Manikandan
Manikandan·19 min read