Manikandan — Manikandan
Microservices

Day 27: Bulkhead

ManikandanManikandan
19 min read·Updated Sep 5, 2022

The Bulkhead pattern splits a service's shared resources (threads, connections, memory, queue slots, CPU) into isolated compartments, so that when one dependency or one traffic flow becomes slow or overloaded, only...

Intro

The Bulkhead pattern splits a service’s shared resources (threads, connections, memory, queue slots, CPU) into isolated compartments, so that when one dependency or one traffic flow becomes slow or overloaded, only its own compartment fills up. The name comes from ship hulls: watertight walls stop one breach from sinking the whole vessel. In an insurance claims system, it means a slow fraud-scoring vendor can exhaust its 20 permits without starving claim-status lookups that customers are hitting at the same time.

Why we need this

A microservice process has finite shared resources: the .NET thread pool, the HttpClient connection pool, database connections, memory, and CPU. Every outgoing call and every incoming request competes for them. Without partitioning, the “loudest” or “slowest” flow wins and everybody else loses.

  • Business reason: Claim submission and payment are revenue- and compliance-critical; “check third-party garage estimate” is not. A slow garage-estimate API must never stop a policyholder from filing a claim.
  • Technical reason: Latency, not errors, is the usual killer. A dependency that answers in 30 seconds instead of 200 ms holds a thread/connection for 150x longer, so concurrency needed to serve the same traffic rises 150x, and the shared pool is drained.
  • Blast-radius control: Bulkheads convert a system-wide outage into a partial degradation (one feature unavailable), which is far easier to operate and explain to the business.
  • It complements the Circuit Breaker (Day 26): the breaker reacts after failures are observed, while the bulkhead bounds the damage while you are still finding out.

What problem it solves

Problem from the topic list: One overloaded flow consumes shared resources and crashes others.

Concrete scenario: ClaimsApi calls three dependencies through one shared HttpClient and one shared pool: Policy Service, Fraud Scoring (vendor), and Document Service. The fraud vendor degrades and responds in 25 s. Requests to it pile up, each holding a pooled connection and a request thread. Within a minute, the connection pool and thread pool are saturated. Now calls to Policy Service (normally 50 ms) also queue or time out, health checks fail, the orchestrator restarts pods, and the whole claims portal is down, although only one non-essential vendor was broken.

Without a bulkhead:

  • Failure propagates horizontally across unrelated features (cascading failure).
  • One noisy tenant (e.g., a broker doing a bulk import of 50,000 claims) starves all other tenants.
  • Restarts don’t help: the moment the pod is back, the same pile-up recurs.

With a bulkhead: each dependency/flow has a fixed budget. When the budget is full, extra work is rejected immediately (or queued briefly) and the caller gets a fast, handled failure, while every other flow keeps its own resources.

When it is needed (and when it is NOT)

Use it when:

  • A service calls several downstream dependencies of different criticality (core vs. optional vendor).
  • You are multi-tenant and one tenant’s burst must not degrade others (noisy neighbour).
  • A dependency has known limits (DB connection cap, vendor licence of N concurrent calls) that you must never exceed.
  • Mixed workloads share a process: interactive API calls next to batch/import work.
  • The service is a fan-in point, e.g. an API Gateway or BFF, that talks to many backends.

Do NOT use it (or don’t bother) when:

  • The service has a single dependency and a single traffic type: there is nothing to isolate; a timeout and circuit breaker are enough.
  • Traffic is tiny and dependencies are fast and reliable; the extra configuration and tuning is overhead.
  • You cannot size the compartments with data. A guessed limit of 5 can cause an outage by rejecting healthy traffic. Measure first (p95 latency x peak RPS).
  • You expect it to replace timeouts. A bulkhead without timeouts just holds permits for a long time; you need both.
  • Your problem is total capacity (the whole system is under-provisioned). Bulkheads shed load fairly; they do not create capacity. Scale out instead.

How to identify the problem (key signals)

  1. One slow dependency, many unrelated failures. Alerts fire on endpoints that don’t touch the slow dependency.
  2. Thread pool starvation signals: ThreadPool.ThreadCount climbing, ThreadPool.PendingWorkItemCount (or the dotnet-counters “ThreadPool Queue Length”) rising, request latency increasing across all endpoints.
  3. Connection pool exhaustion: SqlException: Timeout expired. The timeout period elapsed prior to obtaining a connection from the pool, or HttpClient requests queuing on MaxConnectionsPerServer.
  4. Recovery only after restart, and the problem recurs minutes later once traffic returns.
  5. Noisy-neighbour complaints: “The portal is slow whenever Broker X uploads.”
  6. Code smell: a single static HttpClient/DbContext pool used for everything, no per-dependency limits, and Task.Run/.Result blocking on remote calls.
  7. Metrics shape: in-flight requests per dependency are unbounded and correlate with p99 latency of the whole service; CPU is low while latency is high (waiting, not working).

Flow Diagram

Each dependency gets its own bounded compartment, so one slow vendor cannot drain the others.

flowchart LR
REQ["Incoming requests"] --> API["ClaimsApi"]
API --> B1["Policy bulkhead - 100 permits"]
API --> B2["Documents bulkhead - 40 permits"]
API --> B3["Fraud bulkhead - 20 permits, no queue"]
B1 --> P["Policy Service"]
B2 --> D["Document Service"]
B3 --> F["Fraud vendor - slow"]
B3 -. "full: reject fast" .-> DEF["Defer to manual review"]

Level 1: Beginner

Analogy: A restaurant has three cashier lanes: “regular”, “express (under 5 items)”, and “online pickup”. If a customer with a huge, complicated order blocks the regular lane, express and pickup still move. One shared queue would stall everyone.

Core idea: give each flow a maximum number of concurrent operations. Anything above the limit is rejected (or waits briefly). In .NET the simplest bulkhead is a SemaphoreSlim per dependency.

using System;
using System.Threading;
using System.Threading.Tasks;
public sealed class SimpleBulkhead
{
private readonly SemaphoreSlim _permits;
public SimpleBulkhead(int maxConcurrent) =>
_permits = new SemaphoreSlim(maxConcurrent, maxConcurrent);
public async Task<T> RunAsync<T>(Func<Task<T>> work, CancellationToken ct = default)
{
// Wait(0): do NOT queue; reject immediately when the compartment is full.
if (!await _permits.WaitAsync(TimeSpan.Zero, ct))
throw new BulkheadRejectedException("Compartment is full.");
try { return await work(); }
finally { _permits.Release(); }
}
}
public sealed class BulkheadRejectedException(string message) : Exception(message);

Usage: one bulkhead instance per dependency, not one shared instance.

var fraudBulkhead = new SimpleBulkhead(maxConcurrent: 20);
var policyBulkhead = new SimpleBulkhead(maxConcurrent: 100);
// Fraud vendor gets 20 slots. Even if all 20 hang, policy lookups still have 100.
var score = await fraudBulkhead.RunAsync(() => fraudClient.ScoreAsync(claim));

Beginner rules: (1) one compartment per dependency, (2) always release in finally, (3) pair with a timeout so permits come back.

Level 2: Intermediate

Real applications use a resilience library rather than hand-written semaphores. In current .NET (.NET 10 LTS) that is Polly v8 through Microsoft.Extensions.Http.Resilience (which is built on Polly and Microsoft.Extensions.Resilience). Two places to put bulkheads: outbound (calls to dependencies) and inbound (requests to your API).

6.1 Outbound: one pipeline per dependency (ClaimsApi, .NET 10)

Each named/typed HttpClient gets its own handler pipeline and therefore its own concurrency limiter: that is the compartment.

// Program.cs (packages: Microsoft.Extensions.Http.Resilience, Polly.RateLimiting)
using Polly;
using Polly.RateLimiting;
var builder = WebApplication.CreateBuilder(args);
// Optional fraud vendor: small compartment, no queue, short timeout.
builder.Services.AddHttpClient<IFraudScoringClient, FraudScoringClient>(c =>
c.BaseAddress = new Uri(builder.Configuration["Fraud:BaseUrl"]!))
.AddResilienceHandler("fraud-bulkhead", pipeline =>
{
// Outermost: reject when 20 calls are already in flight; queue nothing.
pipeline.AddConcurrencyLimiter(permitLimit: 20, queueLimit: 0);
pipeline.AddTimeout(TimeSpan.FromSeconds(3));
});
// Critical policy service: big compartment, small queue.
builder.Services.AddHttpClient<IPolicyClient, PolicyClient>(c =>
c.BaseAddress = new Uri(builder.Configuration["Policy:BaseUrl"]!))
.AddResilienceHandler("policy-bulkhead", pipeline =>
{
pipeline.AddConcurrencyLimiter(permitLimit: 100, queueLimit: 25);
pipeline.AddTimeout(TimeSpan.FromSeconds(2));
});
var app = builder.Build();
app.Run();

Handling a rejection: Polly throws RateLimiterRejectedException (it has a RetryAfter property when available). Translate it into a handled business outcome:

public sealed class FraudScoringClient(HttpClient http) : IFraudScoringClient
{
public async Task<FraudResult> ScoreAsync(ClaimDto claim, CancellationToken ct)
{
try
{
var res = await http.PostAsJsonAsync("/score", claim, ct);
res.EnsureSuccessStatusCode();
return (await res.Content.ReadFromJsonAsync<FraudResult>(ct))!;
}
catch (RateLimiterRejectedException) // bulkhead full
{
// Degrade: mark for manual/async review instead of failing the claim.
return FraudResult.DeferredForManualReview;
}
}
}

Note: AddStandardResilienceHandler() also includes a concurrency rate limiter (default permit limit is 1000, queue 0) plus total timeout, retry, circuit breaker and attempt timeout. It is a sensible baseline, but its default limit is far too high to act as a real bulkhead for a fragile dependency, so tune options.RateLimiter.DefaultRateLimiterOptions.PermitLimit per client.

.AddStandardResilienceHandler(o =>
{
o.RateLimiter.DefaultRateLimiterOptions.PermitLimit = 20;
o.RateLimiter.DefaultRateLimiterOptions.QueueLimit = 0;
});

6.2 Inbound: isolate expensive endpoints (ASP.NET Core rate-limiting middleware)

The built-in middleware (Microsoft.AspNetCore.RateLimiting, .NET 7+) has a ConcurrencyLimiter policy that limits in-flight requests, which is a per-endpoint bulkhead:

using Microsoft.AspNetCore.RateLimiting;
using System.Threading.RateLimiting;
builder.Services.AddRateLimiter(o =>
{
o.RejectionStatusCode = StatusCodes.Status503ServiceUnavailable;
o.AddConcurrencyLimiter("bulk-import", c =>
{
c.PermitLimit = 4; // only 4 imports run at once
c.QueueLimit = 10;
c.QueueProcessingOrder = QueueProcessingOrder.OldestFirst;
});
});
var app = builder.Build();
app.UseRateLimiter();
app.MapPost("/api/claims/import", ImportClaimsAsync)
.RequireRateLimiting("bulk-import"); // heavy endpoint gets its own compartment
app.MapGet("/api/claims/{id}", GetClaimAsync); // unaffected

6.3 Angular side

The UI should treat 503/429 from a bulkhead as an expected, recoverable outcome: show a friendly message for the affected widget only, not a global error page.

// fraud-score.service.ts (Angular 20/21, standalone, functional style)
import { inject, Injectable } from '@angular/core';
import { HttpClient, HttpErrorResponse } from '@angular/common/http';
import { catchError, of, throwError } from 'rxjs';
@Injectable({ providedIn: 'root' })
export class FraudScoreService {
private http = inject(HttpClient);
getScore(claimId: string) {
return this.http.get<{ score: number | null; deferred: boolean }>(`/api/claims/${claimId}/fraud-score`).pipe(
catchError((e: HttpErrorResponse) =>
e.status === 503 || e.status === 429
? of({ score: null, deferred: true }) // widget shows "Review pending"
: throwError(() => e)
)
);
}
}

6.4 Database compartments

Separate connection pools per workload by using a different connection string (pool key). For SQL Server / PostgreSQL, a distinct Application Name (SQL Server) or distinct Max Pool Size per connection string creates separate pools:

{
"ConnectionStrings": {
"ClaimsOltp": "Server=...;Database=Claims;Application Name=claims-oltp;Max Pool Size=80",
"ClaimsReport": "Server=...;Database=Claims;Application Name=claims-report;Max Pool Size=10"
}
}

Now a runaway report query can consume at most 10 connections, not the 100 the OLTP path needs. (For PostgreSQL with Npgsql, different connection strings likewise get different pools; set Maximum Pool Size on each.)

Level 3: Advanced

Sizing (Little’s Law): concurrency needed = arrival rate x latency. If Fraud Scoring gets 40 req/s at p95 300 ms, steady-state concurrency is about 40 x 0.3 = 12. Set the limit with headroom (for example 20 to 25), not at 12 and not at 1000. Revisit when traffic or latency changes.

Semaphore vs. thread-pool isolation: In Java (Hystrix) thread-pool isolation was common. In .NET with async/await, calls do not hold a thread while waiting for I/O, so semaphore-style (concurrency-limit) isolation is the right, cheap choice. Dedicated threads are only relevant for blocking or CPU-bound legacy code, where a dedicated worker service/process is the better answer.

Queue or reject? A queue hides overload as latency. For interactive traffic prefer queueLimit: 0 or small; for background work a bounded queue is fine. Never use an unbounded queue: it converts a bulkhead into a memory leak.

Order in the pipeline matters. Typical order (outermost to innermost): concurrency limiter, total timeout, retry, circuit breaker, per-attempt timeout. Putting the bulkhead inside the retry means each retry attempt takes a permit, so retries can multiply load on a struggling dependency. Decide deliberately.

Failure modes and common mistakes:

  • Limits set by guess: too low rejects healthy traffic; too high gives no protection. Load test and use metrics.
  • No timeout: hung calls never release permits, so the compartment fills and stays full.
  • Shared compartment by accident: a singleton HttpClient reused for many dependencies, or one AddResilienceHandler pipeline instance shared across services, defeats the isolation.
  • Ignoring the rejection path: rejections surface as unhandled exceptions and 500s. Map them to 503/429 with Retry-After, or to a degraded response.
  • Retry storms: clients that retry immediately on 503 re-create the overload. Use exponential backoff with jitter (Day 28).
  • Scale-out illusion: limits are per instance. 10 pods x 20 permits = up to 200 concurrent calls to the vendor. If the vendor licence says 50, set the per-pod limit to about 5 or enforce it centrally (e.g. API Management limit-concurrency, or a queue).
  • Leaks: forgetting Release in a custom semaphore on exception paths permanently shrinks the compartment.

Security angle: bulkheads are a cheap defence against resource-exhaustion abuse (denial of service by slow or expensive requests) and against a tenant monopolising capacity. They are not a substitute for authentication, quotas, or WAF rules.

Observability: emit and alert on rejected count, permits in use, and queue depth per compartment. Polly v8 publishes resilience telemetry through Microsoft.Extensions.Telemetry/OpenTelemetry (for example resilience.polly.strategy.events with RateLimiterRejected); wire it to your metrics pipeline. A rising rejection rate on one compartment is an early warning of a dependency problem.

Level 4: Expert and Architect view

8.1 Isolation options compared

ApproachIsolation strengthCost / complexityBest forWeakness
In-process concurrency limit (Polly / SemaphoreSlim)Logical; same process, same memory/CPUVery lowPer-dependency and per-endpoint limitsA bug or memory blow-up still affects the whole process
Separate connection pools per workloadMedium (DB-level)LowOLTP vs. reporting on the same DBSame DB server still shared (CPU, IO)
Separate deployments / consumer groups per tenant or workload classStrong (process-level)MediumNoisy tenants, batch vs. interactiveMore units to operate and deploy
Separate Kubernetes namespaces / node pools / resource quotasStrong (infrastructure-level)Medium-HighCritical vs. best-effort servicesCapacity fragmentation, higher cost
Cell-based architecture (independent stacks per customer group)StrongestHighVery large multi-tenant SaaSSignificant engineering and cost overhead

8.2 Patterns it combines with

  • Circuit Breaker (Day 26): bulkhead caps concurrency; breaker stops calling a failing dependency altogether.
  • Timeout + Retry & Backoff (Day 28): timeouts free permits; backoff avoids re-overloading.
  • Fallback (Day 29): what to do when the bulkhead rejects (cached data, deferred processing).
  • Rate Limiter (Day 30): limits rate over time; bulkhead limits concurrency at a moment. They protect against different shapes of overload.
  • API Gateway / BFF (Days 19-20): natural places for per-route bulkheads.
  • Messaging (Day 16): a queue with bounded consumers is itself a bulkhead between producer and consumer.

8.3 ADR (architecture review style)

ADR-027: Isolate downstream dependencies in ClaimsApi with per-client concurrency bulkheads

  • Status: Proposed
  • Context: ClaimsApi calls Policy Service (critical), Document Service (important) and a third-party Fraud Scoring vendor (optional, occasionally slow, licence limited). A vendor slowdown last quarter saturated the shared connection/thread pool and made the claim portal unavailable for 18 minutes.
  • Decision: Give each downstream dependency its own HttpClient and its own Polly resilience pipeline with a concurrency limiter sized using Little’s Law plus 50% headroom (Policy 100/25 queue, Documents 40/10, Fraud 20/0), each combined with a timeout. Give the bulk-import endpoint its own ASP.NET Core concurrency policy. Use separate DB connection pools for OLTP and reporting. Map rejections to a degraded response (fraud: defer to manual review) or 503 with Retry-After.
  • Alternatives considered: (a) rely only on circuit breakers: reacts after damage, does not stop pile-up during the detection window; (b) separate deployments per dependency: stronger isolation but disproportionate operational cost now; (c) do nothing and scale out: does not address the root cause, cost grows.
  • Consequences: (+) partial degradation instead of full outage, predictable vendor load; (-) limits need tuning and review each quarter, rejected requests need UX and monitoring, per-instance limits multiply with pod count.
  • Review trigger: re-evaluate limits when p95 latency or peak RPS of any dependency changes by more than 30%.

Azure implementation

Bulkheading on Azure is a layered set of controls; Azure has no single “Bulkhead” service. Verify current limits and prices on the linked Microsoft pages before committing to numbers.

Services and how they support the pattern

  • Azure API Management (APIM): the limit-concurrency policy caps the number of concurrent requests entering a policy scope (for example per backend or per operation) and rejects extra requests; combine with rate-limit-by-key/quota-by-key for rate limits per subscription or tenant. Use separate backends and policy scopes per downstream so that one slow backend cannot consume all gateway capacity. Tier note: available across classic tiers; check the policy reference for per-tier and v2-tier support and the APIM service-limits page for capacity limits.
  • Azure Container Apps / AKS (compute isolation): set CPU and memory requests/limits per container so one service cannot starve its neighbours; use separate node pools (AKS) or separate container apps / environments for critical vs. best-effort workloads; configure KEDA/HTTP scaling rules with a max replicas value to protect databases from scale-out storms. In AKS, ResourceQuota and LimitRange per namespace give team-level compartments.
  • Azure App Service: use separate App Service plans for critical and best-effort apps; apps in the same plan share the same VM resources.
  • Azure Service Bus: queues/topics decouple and buffer; bound consumer concurrency (MaxConcurrentCalls in ServiceBusProcessorOptions) so a burst of messages can’t overwhelm downstream. Separate queues per workload or tenant class give queue-level bulkheads. Premium tier offers dedicated resources (messaging units) and more predictable isolation than Standard, which is multi-tenant (throttling is possible).
  • Azure SQL Database / PostgreSQL Flexible Server: separate pools in the app (Section 6.4); use Elastic Pools or separate databases to isolate tenants; on Azure SQL Database, isolation is typically done with separate databases, elastic pools, or read replicas for reporting (Resource Governor is a SQL Server / Managed Instance feature).
  • Azure Front Door / Application Gateway (WAF): edge rate limiting and bot rules protect the whole system before traffic reaches your compartments.
  • Azure Monitor / Application Insights / OpenTelemetry: capture rejection counts, in-flight requests, thread pool and connection pool metrics; alert when RateLimiterRejected spikes.

How to configure (example: APIM concurrency limit on the fraud backend, illustrative policy XML)

<inbound>
<base />
<limit-concurrency key="fraud-backend" max-count="20">
<set-backend-service backend-id="fraud-vendor" />
<!-- requests beyond 20 concurrent get 429 from the gateway -->
</limit-concurrency>
</inbound>

Pricing and tier considerations

  • Isolation costs money: separate plans, node pools, or Premium Service Bus each add fixed cost. Start with in-process bulkheads (free), add infrastructure isolation only for the highest-blast-radius workloads.
  • APIM: pick the tier by throughput, network (VNet) and SLA needs; consumption-style tiers scale differently from dedicated ones, so check the current tier comparison and pricing page.
  • Service Bus: Standard is pay per operation with shared capacity; Premium is priced per messaging unit with dedicated capacity. Choose Premium when throttling or noisy neighbours are a real risk.
  • Container Apps: consumption vs. dedicated workload profiles trade cost against guaranteed capacity.

Reference architecture (text)

Angular SPA on Static Web Apps / Front Door (WAF, edge rate limits) -> APIM (per-route limit-concurrency, per-tenant rate-limit-by-key) -> ClaimsApi on Container Apps (own CPU/memory limits, max replicas) -> inside ClaimsApi, one resilience pipeline per dependency: Policy Service (100), Documents (40), Fraud vendor (20, no queue) -> bulk imports are accepted at /import (concurrency 4) and pushed to a Service Bus queue consumed by a separate Container App with MaxConcurrentCalls tuned to DB capacity -> Azure SQL with separate connection pools for OLTP and reporting (reporting on a read replica) -> telemetry via OpenTelemetry to Application Insights with alerts on rejections and pool saturation.

Sources for verification: Bulkhead pattern (Azure Architecture Center), APIM limit-concurrency policy, Resilient app development in .NET, Microsoft.Extensions.Http.Resilience on NuGet.

Teaching guide for my team

10.1 Beginner explanation (2 minutes)

“Imagine our claims API is a call centre with 100 phone lines. Three departments share those lines: policies, documents and fraud checks. Fraud checks are handled by an outside vendor who sometimes takes half a minute to answer. If all 100 lines end up waiting on the vendor, nobody can call policies either, and the whole call centre looks dead. A bulkhead means we reserve lines: 20 for fraud, 40 for documents, 100 (in a bigger office) for policies. If the fraud lines are all busy, the 21st fraud call gets an instant ‘busy, try later’, and the other departments keep working. In code: one counter per dependency, take a permit before calling, give it back in finally, and always add a timeout.”

10.2 Intermediate explanation (5 minutes)

  • Failure mode: latency, not exceptions, exhausts shared pools (Little’s Law: concurrency = rate x latency).
  • Mechanism in .NET: AddResilienceHandler with AddConcurrencyLimiter(permitLimit, queueLimit) per typed HttpClient; RateLimiterRejectedException is the rejection signal; ASP.NET Core AddConcurrencyLimiter policy + RequireRateLimiting for inbound endpoints; separate connection strings for separate DB pools.
  • Sizing: measure p95 latency and peak RPS, add headroom, remember limits are per instance.
  • Pair with: timeouts (release permits), circuit breaker (stop calling a dead dependency), fallback (what the user sees on rejection), retry with jitter (do not re-amplify).
  • Observability: rejected count, permits in use, queue depth per compartment.
  • Show the ADR-027 numbers and ask the team what they would set for Document Service and why.

10.3 Hands-on exercise

Task: Build a small ASP.NET Core minimal API ClaimsApi with two endpoints, /policy/{id} (fast, 50 ms) and /fraud/{id} (calls a stub that sleeps 10 s). First implement both through one shared HttpClient limited by SocketsHttpHandler.MaxConnectionsPerServer = 10. Load test with 30 concurrent users hitting /fraud/{id} plus 5 users hitting /policy/{id} (use bombardier, k6 or dotnet-counters + a simple script). Then add a per-client AddConcurrencyLimiter(5, 0) for fraud and a 2 s timeout, with a separate client for policy.

Expected outcome: Before: /policy/{id} latency rises to seconds and times out while the fraud stub is slow. After: /policy/{id} stays near 50 ms; /fraud/{id} returns the deferred/503 response quickly for calls over the limit; the rejected-request counter is visible. Bonus: log the permit count and explain why setting the limit to 1000 gives no protection.

10.4 Interview-style questions

  1. What is the difference between a bulkhead and a circuit breaker? A bulkhead limits how many concurrent calls one dependency/flow may use (isolation, proactive); a circuit breaker stops calling a dependency after it has been failing (reactive, fail fast). They are used together.
  2. Why is a bulkhead useless without a timeout? Hung calls never release their permits, so the compartment stays full forever; the timeout guarantees permits return.
  3. You run 12 pods, each with a fraud bulkhead of 20, and the vendor allows 100 concurrent calls. What is wrong? Limits are per instance: 12 x 20 = 240 potential concurrent calls. Lower the per-pod limit (about 8) or enforce the cap centrally (APIM limit-concurrency or a queue with bounded consumers).

Mastery checklist

  • I can explain, with Little’s Law, why a slow dependency exhausts pools without a single error being thrown.
  • I can identify shared-resource pile-up from metrics (thread pool queue length, connection pool timeouts, in-flight requests per dependency).
  • I can implement per-dependency compartments in .NET with Polly v8 / Microsoft.Extensions.Http.Resilience and handle RateLimiterRejectedException.
  • I can size a compartment from p95 latency and peak RPS, and adjust for the number of instances.
  • I can order timeout, retry, circuit breaker and bulkhead in a pipeline and explain the consequences of the order.
  • I can choose between in-process, connection-pool, deployment-level and infrastructure-level isolation and justify the cost.
  • I can map rejections to a user-friendly degraded experience in Angular (503/429 handling per widget).
  • I can write an ADR for introducing bulkheads and name its review trigger.

Key takeaway

A bulkhead does not make a failing dependency healthy; it makes sure that dependency can only hurt its own compartment. Give every dependency and workload class its own bounded budget, always with a timeout, and decide in advance what “rejected” looks like to the user.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 30: Rate Limiter / Throttling

A rate limiter decides, per caller, how many requests the system will accept in a given time (or how many it will process at once) and rejects or delays the rest, usually with HTTP 429 and a Retry-After header.

Manikandan
Manikandan·20 min read
Microservices

Day 29: Fallback

A Fallback is the "plan B" your service runs when a dependency call fails or is rejected (timeout, open circuit, bulkhead full, 5xx).

Manikandan
Manikandan·21 min read
Microservices

Day 28: Retry & Backoff

Retry & Backoff means that when a call to another service or resource fails with a *transient* error (a network blip, a 503, a throttling response, a deadlock victim), the caller waits a short, growing, randomised...

Manikandan
Manikandan·18 min read
Microservices

Day 26: Circuit Breaker

A circuit breaker sits between your service and a dependency it calls (a fraud-scoring API, a payment gateway, another microservice).

Manikandan
Manikandan·20 min read