Manikandan — Manikandan
Microservices

Day 30: Rate Limiter / Throttling

ManikandanManikandan
20 min read·Updated Sep 8, 2022

A rate limiter decides, per caller, how many requests the system will accept in a given time (or how many it will process at once) and rejects or delays the rest, usually with HTTP 429 and a Retry-After header.

Intro

A rate limiter decides, per caller, how many requests the system will accept in a given time (or how many it will process at once) and rejects or delays the rest, usually with HTTP 429 and a Retry-After header. Throttling is the same idea viewed from the server’s side: slow callers down so that one noisy client, a bot, or a traffic spike cannot use up capacity that other clients need. In our insurance claims system, it is what stops one broker’s buggy integration from taking down claim submission for everyone else.

Version baseline for this lesson: .NET 10 (current LTS, ASP.NET Core built-in rate limiting middleware, available since .NET 7) and Angular 21 or later (standalone components, functional HTTP interceptors). Azure details were checked against the Microsoft Learn pages for API Management rate-limit-by-key / quota-by-key and Azure Front Door WAF rate limiting. Prices are deliberately not quoted as numbers; check the Azure pricing calculator for your region.

Why we need this

Every service has finite capacity: threads, database connections, CPU, and downstream quotas (a fraud-scoring vendor that charges per call, an OCR service with a hard limit). Demand is not finite and is not under our control. Without a limit, the caller who sends the most traffic wins, and everybody else loses.

The business reasons are fairness (one tenant must not degrade another), cost protection (pay-per-call downstream services and autoscaling bills), security (credential stuffing, scraping, brute force), and commercial plans (Basic plan: 60 calls/min, Premium plan: 600 calls/min). The technical reason is stability: a system that is over capacity does not degrade gracefully, it collapses, and a rate limiter is the cheapest way to turn “collapse” into “some callers get a fast, clear 429”.

What problem it solves

The problem from the topic list: malicious traffic, noisy neighbours, or traffic spikes overwhelm capacity.

Without it, a typical failure in the claims system looks like this. A broker’s integration has a bug and retries POST /api/claims in a tight loop, 500 requests per second. The Claims API thread pool and SQL connection pool fill up. Requests from all other brokers and from the Angular adjuster portal start timing out. Timeouts cause client retries, which add more load. The autoscaler adds instances, which opens more SQL connections, which makes SQL Server slower. One misbehaving client has caused an outage for all tenants and a large Azure bill.

With a rate limiter, the buggy broker gets 429 after its allowance, everyone else is unaffected, and the broker’s logs show a clear error with a Retry-After value.

When it is needed (and when it is NOT)

Needed when:

  • The API is exposed to external parties (brokers, partners, mobile apps, the public internet).
  • The system is multi-tenant and tenants share compute and database.
  • Endpoints are expensive (report export, full-text claim search, PDF generation) or call pay-per-use services.
  • Endpoints are attractive to attackers (login, password reset, OTP, claim status lookup by claim number).
  • You publish plans or SLAs with different limits per customer.

Not needed, or the wrong tool, when:

  • A purely internal service called by one known caller at a predictable rate. A bulkhead or a queue is usually a better protection there.
  • You are trying to protect against a slow dependency. That is a job for circuit breaker, timeout, and bulkhead (Days 26, 27). Rate limiting protects you from callers, not from your own dependencies.
  • You want to protect against a large volumetric DDoS. Application-level limiters run too late; use Azure DDoS Protection and edge WAF for that.
  • You have not decided what the correct limit is and have no metrics. A guessed limit that is too low becomes a self-inflicted outage. Measure first, start in log-only mode, then enforce.

How to identify the problem (key signals)

  • One client ID, API key, tenant or IP accounts for a disproportionate share of requests (top-1 caller above, say, 30 to 50 percent of traffic) in the gateway logs.
  • Latency and error rate rise for all tenants at the same moment that one tenant’s request rate spikes.
  • SQL Server or PostgreSQL shows connection pool exhaustion or many identical parameterised queries from the same API endpoint in a short window.
  • Repeated 401 responses for many different usernames from one IP or a small IP range (credential stuffing), or sequential claim-number lookups (enumeration).
  • Downstream vendor invoices or quota alerts spike, or the vendor starts returning 429 to you.
  • Autoscaling scales out to the maximum on a traffic pattern that has no business explanation.
  • Support tickets say “the portal is slow every morning at 9:00” and the cause is a partner’s batch job started with no concurrency control.

Flow Diagram

Layered rate limiting from edge to application, with 429 and Retry-After on rejection.

flowchart LR
C["Brokers and portal"] --> WAF["Front Door WAF - per IP"]
WAF --> APIM["APIM - rate-limit-by-key + quota per subscription"]
APIM --> APP["Claims API - token bucket per tenant"]
APP --> SVC["Business logic"]
WAF -. "block" .-> R1["429"]
APIM -. "exceeded" .-> R2["429 + Retry-After"]
APP -. "no tokens" .-> R3["429 + Retry-After"]

Level 1: Beginner

Analogy: a nightclub door. The club holds 200 people (capacity). The doorman lets in a fixed number per minute (rate) and, once the club is full, says “wait outside” (reject or queue). Nobody is thrown out; the door just controls how fast people enter, so the people inside still get served.

The four common algorithms:

AlgorithmIdeaGood forWatch out
Fixed windowN requests per calendar window (for example per minute); counter resets at the window edgeSimple limits, cheapBoundary burst: N at 12:00:59 plus N at 12:01:00 gives 2N in two seconds
Sliding windowWindow moves with time (implemented with segments)Smoother, fairer limitsSlightly more memory and CPU
Token bucketBucket refills at a steady rate; each request takes a token; allows bursts up to bucket sizeAPIs that should allow short bursts but cap the averageTwo parameters to tune (size and refill rate)
Concurrency limiterLimits requests in flight, not per timeExpensive endpoints (exports, PDF generation)Requests that hang hold permits; needs timeouts

Minimal working example (.NET 10, Program.cs, Web SDK, no extra packages):

using System.Threading.RateLimiting;
using Microsoft.AspNetCore.RateLimiting;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddRateLimiter(options =>
{
options.RejectionStatusCode = StatusCodes.Status429TooManyRequests;
// Named policy: 10 requests per 10 seconds, no queueing.
options.AddFixedWindowLimiter("claims-basic", limiter =>
{
limiter.PermitLimit = 10;
limiter.Window = TimeSpan.FromSeconds(10);
limiter.QueueLimit = 0;
});
});
var app = builder.Build();
app.UseRateLimiter();
app.MapGet("/api/claims/{id:guid}", (Guid id) => Results.Ok(new { id, status = "Open" }))
.RequireRateLimiting("claims-basic");
app.Run();

Note that this policy is global for the endpoint: all callers share the same 10 permits. That is the first mistake beginners make. Level 2 fixes it by partitioning per caller.

Level 2: Intermediate

Real applications need three things: limits per caller (partitioning by tenant or user, not global), a proper 429 response with Retry-After, and different policies for different endpoints. Below is a claims API with a token bucket for submissions, a sliding window for search, and a concurrency limit for exports.

Packages: Microsoft.AspNetCore.Authentication.JwtBearer (authority and audience configuration omitted for brevity).

using System.Globalization;
using System.Threading.RateLimiting;
using Microsoft.AspNetCore.Mvc;
using Microsoft.AspNetCore.RateLimiting;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddAuthentication().AddJwtBearer();
builder.Services.AddAuthorization();
builder.Services.AddRateLimiter(options =>
{
options.RejectionStatusCode = StatusCodes.Status429TooManyRequests;
options.OnRejected = async (context, ct) =>
{
var response = context.HttpContext.Response;
if (context.Lease.TryGetMetadata(MetadataName.RetryAfter, out var retryAfter))
{
response.Headers.RetryAfter =
((int)Math.Ceiling(retryAfter.TotalSeconds)).ToString(CultureInfo.InvariantCulture);
}
await response.WriteAsJsonAsync(new ProblemDetails
{
Status = StatusCodes.Status429TooManyRequests,
Title = "Too many requests",
Detail = "Request rate exceeded for this tenant. Retry after the time in the Retry-After header."
}, cancellationToken: ct);
};
// Claim submission: burst of 20, refill 5 per second, per tenant.
options.AddPolicy("claims-submit", httpContext =>
RateLimitPartition.GetTokenBucketLimiter(PartitionKey(httpContext), _ =>
new TokenBucketRateLimiterOptions
{
TokenLimit = 20,
TokensPerPeriod = 5,
ReplenishmentPeriod = TimeSpan.FromSeconds(1),
QueueLimit = 0,
AutoReplenishment = true
}));
// Search: 100 per minute per tenant, smoothed with 6 segments.
options.AddPolicy("claims-search", httpContext =>
RateLimitPartition.GetSlidingWindowLimiter(PartitionKey(httpContext), _ =>
new SlidingWindowRateLimiterOptions
{
PermitLimit = 100,
Window = TimeSpan.FromMinutes(1),
SegmentsPerWindow = 6,
QueueLimit = 0
}));
// Export: at most 2 concurrent exports per tenant; up to 5 wait in a queue.
options.AddPolicy("claims-export", httpContext =>
RateLimitPartition.GetConcurrencyLimiter(PartitionKey(httpContext), _ =>
new ConcurrencyLimiterOptions
{
PermitLimit = 2,
QueueLimit = 5,
QueueProcessingOrder = QueueProcessingOrder.OldestFirst
}));
});
var app = builder.Build();
app.UseAuthentication(); // must run before the limiter so the tenant claim is available
app.UseRateLimiter();
app.UseAuthorization();
app.MapPost("/api/claims", (ClaimRequest request) => Results.Accepted($"/api/claims/{Guid.NewGuid()}"))
.RequireAuthorization()
.RequireRateLimiting("claims-submit");
app.MapGet("/api/claims", (string? q) => Results.Ok(Array.Empty<string>()))
.RequireAuthorization()
.RequireRateLimiting("claims-search");
app.MapGet("/api/claims/export", () => Results.File(Array.Empty<byte>(), "text/csv"))
.RequireAuthorization()
.RequireRateLimiting("claims-export");
app.Run();
static string PartitionKey(HttpContext ctx) =>
ctx.User.FindFirst("tenant_id")?.Value
?? ctx.Connection.RemoteIpAddress?.ToString()
?? "anonymous";
public sealed record ClaimRequest(string PolicyNumber, decimal Amount, string Description);

Angular 21 client: honour Retry-After for safe, idempotent GET requests only, and give the user a clear message.

import { HttpErrorResponse, HttpInterceptorFn } from '@angular/common/http';
import { retry, throwError, timer } from 'rxjs';
export const rateLimitInterceptor: HttpInterceptorFn = (req, next) =>
next(req).pipe(
retry({
count: 2,
delay: (error: unknown) => {
if (
req.method !== 'GET' ||
!(error instanceof HttpErrorResponse) ||
error.status !== 429
) {
return throwError(() => error);
}
const seconds = Number(error.headers.get('Retry-After') ?? '1');
return timer(Math.min(Math.max(seconds, 1), 10) * 1000);
}
})
);
// app.config.ts
// provideHttpClient(withInterceptors([rateLimitInterceptor]))

Database side: keep plan limits in data, not code, so support can change a tenant’s limit without a deployment. SQL Server example:

CREATE TABLE dbo.TenantRateLimit
(
TenantId UNIQUEIDENTIFIER NOT NULL PRIMARY KEY,
Plan NVARCHAR(20) NOT NULL, -- Basic, Standard, Premium
SubmitBurst INT NOT NULL,
SubmitPerSecond INT NOT NULL,
SearchPerMinute INT NOT NULL,
UpdatedAtUtc DATETIME2(0) NOT NULL DEFAULT SYSUTCDATETIME()
);

Load these rows into IMemoryCache (refresh every few minutes) and read them inside the policy factory. Never query the database on every request just to decide whether to allow a request; that turns the protection into the bottleneck.

Level 3: Advanced

Performance and scalability

  • The built-in limiters are in-memory and per process. With 4 API instances and a policy of 100/min, the real tenant limit is up to 400/min, and it changes as the autoscaler adds instances. Options: divide the limit by the expected instance count (approximate), enforce the precise limit at a shared point (API Management or a Redis-backed limiter), or accept approximate limits at the app layer and use the gateway for the contractual limit.
  • The built-in limiter is cheap (lock-free or short lock per partition). A Redis round trip per request adds roughly a network hop of latency and a new dependency; use it only for limits that must be exact across instances.
  • Token bucket is usually the best default for APIs: it allows a sensible burst and caps the sustained rate. Fixed window is the cheapest but allows a 2x burst at window boundaries.

Distributed limiter sketch with Redis (package StackExchange.Redis). The Lua script makes increment-and-expire atomic:

using StackExchange.Redis;
public sealed class RedisFixedWindowLimiter(IConnectionMultiplexer redis)
{
private const string Script = @"
local current = redis.call('INCR', KEYS[1])
if current == 1 then redis.call('PEXPIRE', KEYS[1], ARGV[1]) end
return current";
public async Task<bool> TryAcquireAsync(string tenantId, int limit, TimeSpan window)
{
var db = redis.GetDatabase();
long windowMs = (long)window.TotalMilliseconds;
long windowId = DateTimeOffset.UtcNow.ToUnixTimeMilliseconds() / windowMs;
RedisKey key = $"rl:{tenantId}:{windowId}";
var count = (long)await db.ScriptEvaluateAsync(
Script, new[] { key }, new RedisValue[] { windowMs * 2 });
return count <= limit;
}
}

Decide in advance what happens when Redis is down: fail open (allow traffic, keep serving, risk overload) or fail closed (reject, protect the backend, cause an outage). For claim submission we usually fail open with a fall back to the local in-memory limiter; for login we usually fail closed.

Security

  • Behind a load balancer, RemoteIpAddress is the proxy’s address unless you configure UseForwardedHeaders with KnownProxies or KnownNetworks. Without that, either every client shares one bucket (accidental outage) or, if you trust X-Forwarded-For blindly, attackers rotate a spoofed header to get unlimited buckets.
  • Prefer authenticated identity (tenant, client ID) as the partition key. Use IP only for anonymous endpoints, and be aware of NAT (many users behind one corporate IP) and IPv6 (one user can own a whole /64, so partition IPv6 by prefix).
  • Apply a strict, separate policy to login, password reset, and OTP endpoints, keyed by both account name and IP.
  • Do not reveal internals in the 429 body; return Retry-After and a stable error shape.

Failure modes and common mistakes

  • A limit that is too low for legitimate peaks (month-end claim uploads) creates a self-inflicted outage. Roll out in log-only mode first, compare against real traffic percentiles.
  • Clients that retry a 429 immediately in a loop make it worse. Clients must honour Retry-After and use exponential backoff with jitter (Day 28).
  • Queueing (QueueLimit > 0) hides overload by turning it into latency. Keep queues short and bounded, and only for work where waiting is acceptable.
  • Limiting after expensive work (after deserialising a 20 MB body or after authentication against a slow identity provider) protects less than limiting early. Put a cheap coarse limit at the edge and finer limits inside.
  • Global limits with no per-tenant partition let one tenant consume everyone’s allowance.
  • Health check and metrics endpoints accidentally rate limited, so the orchestrator restarts healthy pods. Use DisableRateLimiting on them.
  • Unbounded number of partition keys (for example keyed on a random header) is a memory attack. Keep keys to a bounded, validated set.

Level 4: Expert and Architect view

Where to enforce the limit, compared:

LocationPrecision across instancesCost of a rejected requestPer-tenant plansTypical use
Edge WAF (Azure Front Door WAF)Approximate (counted per edge node)Lowest, never reaches originLimited (IP or header-based rules)Bots, abusive IPs, coarse protection
API gateway (Azure API Management)Good within one gateway deployment; approximate across regions or self-hosted gatewaysLowStrong (subscription key, product, claim)Contractual limits, quotas, partner plans
Application middleware (ASP.NET Core)Per instance onlyMedium (request reaches the app)Flexible (any claim, endpoint)Endpoint-specific and cost-aware limits
Distributed limiter (Redis)ExactMedium plus a Redis callFlexibleExact limits that the app itself must enforce
Service mesh or ingress (Envoy, NGINX)Per proxy (or global with a rate limit service)LowLimitedCluster-level protection on Kubernetes

Combines with:

  • API Gateway and BFF (Days 19, 20): the natural place for the coarse per-client limit.
  • Circuit Breaker, Bulkhead (Days 26, 27): limiter protects you from callers; breaker and bulkhead protect you from your dependencies.
  • Retry & Backoff (Day 28): the client-side partner of the limiter. A 429 with Retry-After is the signal the client’s retry policy needs.
  • Fallback (Day 29): degrade search results to cached data when the caller is throttled on the expensive path.
  • Application Metrics and Audit Logging (Days 33, 34): every rejection is a metric and, for security endpoints, an audit record.

Short ADR for an architecture review:

ADR-030: Layered rate limiting for the Claims platform
Status: Proposed
Context:
The Claims API is used by 40 broker integrations and the adjuster portal.
A single broker retry loop has already exhausted the SQL connection pool
once, degrading all tenants. Plans with different limits are being sold.
Decision:
1. Azure Front Door WAF rate limit rule per client IP as the coarse outer layer.
2. Azure API Management rate-limit-by-key (short window) and quota-by-key
(monthly) keyed on subscription, enforcing plan limits.
3. ASP.NET Core rate limiting middleware in the Claims API, partitioned by
tenant, for endpoint-specific limits (token bucket for submit,
concurrency for export).
4. All 429 responses carry Retry-After; clients must back off with jitter.
5. Plan limits are stored in SQL Server and cached; no per-request DB call.
6. New limits ship in log-only mode for one week before enforcement.
Consequences:
+ One tenant cannot degrade others; clear, machine-readable rejection.
+ Defence in depth: an outage of one layer does not remove all protection.
- Three places to configure; limits must be documented and kept consistent.
- App-layer limits are per instance and therefore approximate.
- Wrongly tuned limits can block legitimate month-end peaks.
Alternatives rejected:
- Only app-layer limits: not precise across scale-out, rejected requests
still consume app resources.
- Only gateway limits: no visibility into endpoint cost or per-user claims.
- Redis limiter everywhere: extra dependency and latency on the hot path
for limits that do not need exactness.

Azure implementation

Services that implement or support the pattern:

  • Azure API Management (APIM): rate-limit-by-key (short-term rate) and quota-by-key (long-term calls or bandwidth), returning 429 for rate limit and 403 for exceeded quota. Also rate-limit and quota scoped to subscription. Check the policy page for tier availability before choosing a tier, because not every policy is available in every tier (notably the Consumption tier has restrictions), and check the maximum renewal-period allowed for rate-limit-by-key (currently documented as a short maximum, in the range of minutes) so you do not design a monthly limit with it; use quota-by-key for that.
  • Azure Front Door (Standard or Premium) with WAF custom rules of type rate limit: threshold per duration (1 or 5 minutes), grouped by client IP or socket address, with actions such as Block, Log, or Allow. Counts are per Front Door edge node, so they are approximate.
  • Azure Application Gateway WAF v2 also supports rate limit custom rules if you use Application Gateway instead of Front Door.
  • Azure Managed Redis (or the Redis offering you already run) for exact, distributed counters.
  • Azure Monitor, Application Insights, Log Analytics for rejection metrics and alerts.
  • Azure Container Apps or AKS ingress: Container Apps can run Envoy-based ingress; on AKS use your ingress controller’s rate limit annotations or an Envoy/Istio rate limit service.

How to configure APIM (policy, inbound section):

<inbound>
<base />
<!-- Short-term rate: 100 calls per 60 seconds per subscription -->
<rate-limit-by-key calls="100"
renewal-period="60"
counter-key="@(context.Subscription?.Id ?? context.Request.IpAddress)"
remaining-calls-header-name="X-RateLimit-Remaining"
retry-after-header-name="Retry-After" />
<!-- Long-term quota: 100000 calls per 30 days (2592000 s) per subscription -->
<quota-by-key calls="100000"
renewal-period="2592000"
counter-key="@(context.Subscription?.Id ?? context.Request.IpAddress)" />
</inbound>

Front Door WAF rate limit rule (conceptual configuration, set in portal, Bicep, or CLI): rule type RateLimitRule, duration 1 minute, threshold 300 requests, group by client address, match condition on the request URI beginning with /api/, action Block (start with Log to observe first). Attach the WAF policy to the Front Door security policy that covers your endpoint.

Pricing and tier considerations (no figures quoted; verify in the Azure pricing calculator):

  • APIM is billed by tier and units per hour (Developer for non-production, Basic, Standard, Premium, and the v2 tiers), and Consumption is billed per call. Rate limit policies come with the tier; there is no separate charge, but the tier decides SLA, scale-out, multi-region, and virtual network options.
  • Front Door Standard and Premium have a base fee plus per-request and data transfer charges. Premium adds managed rule sets and bot protection; custom rules such as rate limiting are available in Standard as well as Premium at the time of writing. Confirm this in the current feature comparison before committing.
  • Redis is billed by capacity tier; size for operations per second (one to two operations per checked request) and choose a tier with high availability if the limiter fails closed.
  • The cheapest protection per request is the layer closest to the edge, because a rejected request never consumes origin compute.

Reference architecture (text):

  1. Clients (brokers, Angular adjuster portal) reach api.claims.example through Azure Front Door with a WAF policy. A rate limit rule blocks abusive IPs (coarse, per edge).
  2. Front Door forwards to Azure API Management, which authenticates the subscription key or validates the JWT, then applies rate-limit-by-key and quota-by-key per subscription. 429 with Retry-After is returned from here for plan violations.
  3. APIM routes to the Claims API running on Azure Container Apps (or AKS) with several replicas. ASP.NET Core rate limiting applies endpoint-specific policies partitioned by tenant claim. Plan data is read from SQL Server (or PostgreSQL) through an in-memory cache.
  4. Where an exact cross-instance limit is required (for example OTP attempts), the API calls a Redis-backed limiter in Azure Managed Redis.
  5. All layers emit rejection counts to Application Insights and Log Analytics. An Azure Monitor alert fires when 429 rate for a tenant exceeds a threshold, and a workbook shows top callers by rejected requests.
  6. Secrets (Redis connection, APIM keys) come from Azure Key Vault via managed identity.

Teaching guide for my team

Explain to a beginner in 2 minutes:

“Our API is like a shop counter with one cashier. If one customer sends a hundred friends to queue at once, everyone else waits forever. A rate limiter is a rule at the door: each customer gets a fair number of turns per minute. If you go over, we don’t crash, we politely say ‘too many requests, try again in 5 seconds’ (HTTP 429 and a Retry-After header). Your code, when it gets a 429, must wait and try again later, not hammer the door.”

Explain to an intermediate developer in 5 minutes:

Start with the failure story (one broker’s retry loop exhausts the SQL connection pool for all tenants). Then show the four algorithms and pick token bucket for submission (allows bursts, caps the average), sliding window for search, concurrency for exports. Show AddRateLimiter, a partitioned policy keyed by tenant, and the OnRejected handler that sets Retry-After. Stress three points: partition by identity not globally, the built-in limiter is per instance (so precise limits belong in APIM or Redis), and clients must back off with jitter on 429. Finish by showing where each layer lives on Azure (Front Door WAF, APIM, app middleware).

Hands-on exercise:

Task: Add per-tenant limiting to a sample Claims API and prove it works.

  1. Create the token bucket policy claims-submit from Level 2 (burst 20, refill 5 per second) and apply it to POST /api/claims.
  2. Use two different tenant_id claims (or, for a quick test, two different X-Test-Tenant header values mapped in PartitionKey).
  3. Send 50 requests in a burst from tenant A with a load tool (for example bombardier or hey), and at the same time a steady 1 request per second from tenant B.

Expected outcome: tenant A gets roughly 20 successful responses and then 429 responses with a Retry-After header; tenant B receives 200 on every request throughout. Then change the policy to a single global (non-partitioned) limiter and observe that tenant B now also gets 429, which demonstrates the noisy neighbour problem.

Interview-style questions:

  1. Q: Why is a fixed window limiter sometimes unfair, and what replaces it? A: A client can send the full allowance at the end of one window and again at the start of the next, giving double the intended rate in a short time. A sliding window or token bucket smooths this.
  2. Q: The built-in ASP.NET Core rate limiter is configured for 100 requests per minute and you run 5 instances. What is the effective limit and how do you fix it? A: Up to about 500 per minute because state is per instance. Enforce the exact limit in API Management or a Redis-backed limiter, or divide the local limit by the instance count as an approximation.
  3. Q: What should a client do when it receives a 429? A: Stop sending, wait for the Retry-After value (or use exponential backoff with jitter if absent), retry only idempotent requests automatically, and surface a clear message otherwise. Immediate retries make the overload worse.

Mastery checklist

  • I can explain the difference between fixed window, sliding window, token bucket, and concurrency limiters and choose one for a given endpoint.
  • I can implement a partitioned, per-tenant policy in ASP.NET Core (.NET 10) with a custom OnRejected that returns Retry-After and a ProblemDetails body.
  • I know that the built-in limiter is per instance and can explain three ways to get exact cross-instance limits.
  • I can choose the partition key correctly (tenant or user for authenticated endpoints, IP for anonymous ones) and handle proxies, NAT, and IPv6.
  • I can configure rate-limit-by-key and quota-by-key in Azure API Management and explain the difference between them.
  • I can describe where WAF, gateway, and application limits each belong and why a layered approach is used.
  • I can write a client (Angular interceptor or .NET resilience pipeline) that respects 429 and Retry-After without creating a retry storm.
  • I can roll out a new limit safely: log-only first, measure, enforce, alert on rejections.

Key takeaway

A rate limiter protects shared capacity by giving each caller a fair, measurable allowance and rejecting the excess quickly with a clear 429 and Retry-After. Partition by identity, layer it (edge, gateway, app), and pair it with client-side backoff so that rejection calms the system down instead of feeding a retry storm.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 29: Fallback

A Fallback is the "plan B" your service runs when a dependency call fails or is rejected (timeout, open circuit, bulkhead full, 5xx).

Manikandan
Manikandan·21 min read
Microservices

Day 28: Retry & Backoff

Retry & Backoff means that when a call to another service or resource fails with a *transient* error (a network blip, a 503, a throttling response, a deadlock victim), the caller waits a short, growing, randomised...

Manikandan
Manikandan·18 min read
Microservices

Day 27: Bulkhead

The Bulkhead pattern splits a service's shared resources (threads, connections, memory, queue slots, CPU) into isolated compartments, so that when one dependency or one traffic flow becomes slow or overloaded, only...

Manikandan
Manikandan·19 min read
Microservices

Day 26: Circuit Breaker

A circuit breaker sits between your service and a dependency it calls (a fraud-scoring API, a payment gateway, another microservice).

Manikandan
Manikandan·20 min read