Retry & Backoff means that when a call to another service or resource fails with a *transient* error (a network blip, a 503, a throttling response, a deadlock victim), the caller waits a short, growing, randomised...
Intro
Retry & Backoff means that when a call to another service or resource fails with a transient error (a network blip, a 503, a throttling response, a deadlock victim), the caller waits a short, growing, randomised delay and tries again a bounded number of times instead of failing the user’s operation immediately. Done well it turns many brief glitches into invisible non-events. Done badly (no limit, no delay, no jitter, retrying non-idempotent calls) it turns a small outage into a large one.
Why we need this
In a distributed system almost every call crosses a network, a load balancer, a DNS lookup, a TLS handshake, and a remote process that may be restarting, scaling or being throttled. Cloud platforms explicitly document that transient faults are normal: Azure SQL Database, Azure Service Bus, Azure Storage and Cosmos DB all return brief errors during failover, scale operations, or when throttling. A single-attempt call converts each of those brief moments into a user-visible failure.
Business reasons: fewer failed claim submissions, fewer support tickets, and a higher effective availability without buying more hardware. Technical reasons: transient faults typically clear in milliseconds to seconds, so a second attempt usually succeeds. The backoff and jitter parts exist because naive retries are dangerous: if 500 clients all retry every 100 ms against a struggling service, they keep it down (a “retry storm” or “thundering herd”).
What problem it solves
Problem (from the topic list): transient blips fail operations unnecessarily.
Insurance example: ClaimsService calls PolicyService to check that a policy is active before accepting a claim. During a rolling deployment of PolicyService, one pod is terminated and 1 in 20 requests gets a connection reset or a 503 for about two seconds.
Without retry: about 5% of claim submissions return HTTP 500 to the customer during every deployment, even though nothing is actually broken. Users resubmit, creating duplicate claims. Support sees “random” failures that cannot be reproduced.
With retry and backoff: the failed call is retried after roughly 1-2 seconds (with jitter) and lands on a healthy pod. The customer never notices. Backoff and jitter ensure that if the service is genuinely overloaded, the retries spread out over time rather than arriving in synchronized waves.
When it is needed (and when it is NOT)
Use it when:
- The call crosses a process or network boundary (HTTP, gRPC, SQL, Redis, Service Bus, Blob Storage).
- The failure is transient: timeouts, connection resets, HTTP 408, 429, 500, 502, 503, 504, SQL error codes documented as transient (for example deadlock victim 1205, Azure SQL failover errors).
- The operation is idempotent, or you can make it idempotent with an idempotency key (see Day 17).
- The caller can afford the added latency (the total retry budget fits within the user’s or upstream caller’s timeout).
Do NOT use it (or use it very carefully) when:
- The error is permanent: 400, 401, 403, 404, 409, 422, validation failures, business rule violations. Retrying just repeats the failure.
- The operation is not idempotent and has no idempotency key (for example
POST /paymentsthat charges a card). A retry after a timeout may double-charge, because the first request may have succeeded. - A downstream is clearly down. Retries there should be combined with a Circuit Breaker (Day 26) so you stop hammering it.
- You already retry at another layer. Retries multiply: 3 attempts at the gateway x 3 in the service x 3 in the SDK = 27 calls per user request. Retry at one layer, usually the one closest to the failing call.
- The user is waiting on an interactive request with a tight deadline. Prefer fewer retries with a short total budget, or fail fast and let the client retry.
How to identify the problem (key signals)
- Error rates spike briefly and self-heal: dashboards show 1-3 second bursts of 5xx or timeouts, especially during deployments, scale-outs, or node upgrades.
- Logs contain transient markers:
SocketException: Connection reset,TaskCanceledExceptiononHttpClient,SqlExceptionnumbers 40613, 40197, 49918 or 1205,ServiceBusExceptionwithIsTransient = true, HTTP 429 with aRetry-Afterheader. - “Works if I click again”: users or support say the second attempt always works.
- Failures correlate with infrastructure events: pod restarts, Azure SQL failover, cache node patching, throttling on Storage or Cosmos DB.
- Existing retry code is a smell: a hand-written
for (int i = 0; i < 3; i++)withThread.Sleep(1000), a fixed delay, no jitter, retrying on anyException. - Retry storms (the opposite signal): request rate to a dependency jumps 3x-10x exactly when it starts failing, and recovery is slow or never happens.
- Duplicate side effects: two claims, two emails, or two payments for one user action, which is a sign that retries hit a non-idempotent operation.
Flow Diagram
Retry only transient failures, with a cap, exponential backoff and jitter.
flowchart TD CALL["Call PolicyService"] --> RES{"Outcome"} RES -- "200" --> OK(["Success"]) RES -- "400/401/404/422" --> FAIL(["Fail - permanent"]) RES -- "408/429/5xx/timeout" --> LEFT{"Attempts left?"} LEFT -- "no" --> FB(["Give up - fallback"]) LEFT -- "yes" --> WAIT["Wait 2^n x base + jitter or Retry-After"] WAIT --> CALLLevel 1: Beginner
Analogy: You call a friend and the line is busy. You do not redial 50 times per second. You wait a bit, try again, wait a bit longer, try again, and after a few tries you leave a message (fail gracefully). And if the number does not exist, you do not keep dialling; you check the number.
The three knobs:
- Max attempts: how many tries (for example 3 total).
- Delay / backoff: how long to wait between tries. Fixed = same each time. Exponential = doubles each time (1s, 2s, 4s).
- Jitter: a random amount added or subtracted so clients do not all retry at the same moment.
Minimal working example (plain C#, no library), showing the concept only. In production use a library (Level 2):
using System.Net;
public static class SimpleRetry{ private static readonly Random Rng = new();
public static async Task<HttpResponseMessage> GetWithRetryAsync( HttpClient client, string url, int maxAttempts = 3) { for (var attempt = 1; ; attempt++) { try { var response = await client.GetAsync(url);
// Retry only on transient status codes. var transient = response.StatusCode is HttpStatusCode.RequestTimeout or HttpStatusCode.TooManyRequests or HttpStatusCode.BadGateway or HttpStatusCode.ServiceUnavailable or HttpStatusCode.GatewayTimeout;
if (!transient || attempt == maxAttempts) return response;
response.Dispose(); } catch (HttpRequestException) when (attempt < maxAttempts) { // network-level failure: fall through and retry }
// Exponential backoff with jitter: ~1s, ~2s, ~4s ... var baseDelay = TimeSpan.FromSeconds(Math.Pow(2, attempt - 1)); var jitter = TimeSpan.FromMilliseconds(Rng.Next(0, 500)); await Task.Delay(baseDelay + jitter); } }}What to notice: it retries only transient outcomes, it has a hard cap on attempts, the delay grows, and there is randomness. It does not retry a 404 or 400.
Level 2: Intermediate
Backend: .NET (current LTS: .NET 10) with Microsoft.Extensions.Http.Resilience
The recommended approach for HttpClient is the Microsoft.Extensions.Http.Resilience package, which is built on Polly v8. Do not use the older Microsoft.Extensions.Http.Polly package for new code; it targets Polly v7.
// Program.cs (ClaimsService)using Microsoft.Extensions.Http.Resilience;using Polly;
var builder = WebApplication.CreateBuilder(args);
builder.Services .AddHttpClient<IPolicyClient, PolicyClient>(c => c.BaseAddress = new Uri("https://policy-service.internal/")) .AddResilienceHandler("policy-retry", pipeline => { pipeline.AddRetry(new HttpRetryStrategyOptions { MaxRetryAttempts = 3, BackoffType = DelayBackoffType.Exponential, UseJitter = true, Delay = TimeSpan.FromMilliseconds(500), // Honour Retry-After from 429/503 responses when present ShouldRetryAfterHeader = true });
// Per-attempt timeout so a hung call does not eat the whole budget pipeline.AddTimeout(TimeSpan.FromSeconds(3)); });
var app = builder.Build();app.Run();Notes:
HttpRetryStrategyOptionsby default already retries on HTTP 408, 429, 5xx and onHttpRequestException/ timeouts. You rarely need to write the predicate yourself.- If you just want sensible defaults,
AddStandardResilienceHandler()adds a pipeline of rate limiter, total request timeout, retry, circuit breaker and per-attempt timeout. Its retry default is 3 retries with exponential backoff, jitter and a 2 second base delay. Caveat: by default it retries all HTTP methods, including POST. For non-idempotent calls, either disable retries for those methods (DisableForUnsafeHttpMethods()on the retry options) or use idempotency keys. - The order of strategies matters: a total-timeout outside the retry caps the whole operation, an attempt-timeout inside the retry caps each try.
Idempotent-safe POST with an idempotency key:
public sealed class PolicyClient(HttpClient http) : IPolicyClient{ public async Task<bool> ReserveCoverageAsync(Guid claimId, decimal amount, CancellationToken ct) { using var req = new HttpRequestMessage(HttpMethod.Post, "coverage/reservations") { Content = JsonContent.Create(new { claimId, amount }) }; // Same key on every retry of the same logical operation req.Headers.Add("Idempotency-Key", claimId.ToString());
using var resp = await http.SendAsync(req, ct); resp.EnsureSuccessStatusCode(); return true; }}Database: EF Core with SQL Server / Azure SQL / PostgreSQL
EF Core has built-in execution strategies for connection resiliency. They retry the whole unit of work when a transient error occurs.
// SQL Server / Azure SQLbuilder.Services.AddDbContext<ClaimsDbContext>(o => o.UseSqlServer(connectionString, sql => sql.EnableRetryOnFailure( maxRetryCount: 5, maxRetryDelay: TimeSpan.FromSeconds(30), errorNumbersToAdd: null)));
// PostgreSQL (Npgsql provider)builder.Services.AddDbContext<ClaimsDbContext>(o => o.UseNpgsql(pgConnectionString, npg => npg.EnableRetryOnFailure(5, TimeSpan.FromSeconds(30), null)));Important gotcha: when a retrying execution strategy is enabled, a user-initiated transaction must be wrapped so the whole block can be replayed:
var strategy = db.Database.CreateExecutionStrategy();await strategy.ExecuteAsync(async () =>{ await using var tx = await db.Database.BeginTransactionAsync(); db.Claims.Add(claim); await db.SaveChangesAsync(); await tx.CommitAsync();});If you skip this you get InvalidOperationException: The configured execution strategy does not support user-initiated transactions.
Frontend: Angular (current major version)
For browser-to-API calls, keep retries small and only for idempotent GETs. RxJS 7’s retry accepts a config with a delay function:
import { Injectable, inject } from '@angular/core';import { HttpClient, HttpErrorResponse } from '@angular/common/http';import { Observable, retry, timer } from 'rxjs';
@Injectable({ providedIn: 'root' })export class ClaimsService { private http = inject(HttpClient);
getClaim(id: string): Observable<Claim> { return this.http.get<Claim>(`/api/claims/${id}`).pipe( retry({ count: 2, delay: (error: HttpErrorResponse, retryCount: number) => { // Do not retry client errors if (error.status >= 400 && error.status < 500 && error.status !== 429) { throw error; } const base = 500 * 2 ** (retryCount - 1); // 500ms, 1000ms const jitter = Math.random() * 250; return timer(base + jitter); }, }) ); }}Do not auto-retry POST/PUT/DELETE in the browser unless the API supports an idempotency key. Show the user a clear “Try again” button instead.
Level 3: Advanced
Performance and scalability
- Retry budget / amplification: total load on a dependency = original requests x (1 + retries). With 3 retries and everything failing, load is 4x. Cap retries with a retry budget (for example allow retries only up to 10% of recent request volume) or combine with a circuit breaker. Once the breaker opens, calls fail fast instead of consuming more attempts against a dependency that is down.
- Latency budget: with 500 ms base exponential backoff and 3 retries, worst-case waiting is about 0.5 + 1 + 2 = 3.5 s plus attempt time. Make sure
total timeout >= sum of attempts + delays, and that it is less than what your upstream caller will wait. - Jitter types: “full jitter” (random between 0 and the exponential cap) spreads load best; “equal jitter” keeps half the delay fixed. Polly’s
UseJitter = trueadds randomness to each computed delay; do not implement your own unless you must. - Honour
Retry-After: 429 and 503 responses often say how long to wait. Ignoring it wastes attempts.
Security
- Retried requests carry the same auth token. A token that expires between attempts leads to 401, which is not transient; refresh the token rather than retrying blindly.
- Do not retry on responses that could indicate an attack or a policy block (403). Retrying login endpoints can trip lockouts.
- Logs must not include request bodies containing PII on every retry attempt.
Failure modes
- Retry storm / thundering herd: no jitter or too many retries. Symptom: dependency never recovers.
- Retry amplification across layers: gateway, service, SDK, and database driver each retrying. Retry in one place; disable SDK retries if you own the policy (for example set
ServiceBusRetryOptions.MaxRetries = 0only if you deliberately handle it yourself). - Duplicate side effects: timeout on a request that actually succeeded, then retry. Fix with idempotency keys or an Idempotent Consumer.
- Hidden latency: retries make p99 much worse than p50 and hide health problems. Always emit metrics for retry count.
- Retrying long, slow operations: a 30 s call retried three times blocks a thread or connection for 90+ s. Set per-attempt timeouts.
- Cancellation ignored: retry loops that do not honour
CancellationTokenkeep running after the client has gone.
Common mistakes
- Retrying on every exception (including
ArgumentExceptionorNullReferenceException). - Fixed delay with no jitter.
- Unlimited retries (“retry forever”).
- Retrying non-idempotent POSTs.
- Not logging or measuring retries (Polly v8 exposes telemetry via
Microsoft.Extensions.Http.Resilience; enable OpenTelemetry metrics). - Retrying inside a database transaction without an execution strategy.
Level 4: Expert and Architect view
Alternatives and companions compared
| Approach | What it does | Strengths | Weaknesses | Use when |
|---|---|---|---|---|
| Retry with exponential backoff + jitter | Re-attempts after growing random delay | Handles brief blips, simple | Adds latency, can amplify load, needs idempotency | Transient faults on idempotent calls |
| Immediate retry (no delay) | Retry at once | Fast for one-off packet loss | Useless for real outages, causes storms | At most once, for connection reset |
| Circuit Breaker (Day 26) | Stops calls after failure threshold | Protects dependency and caller | Needs tuning, fails fast for all callers | Dependency is down, not blipping |
| Fallback (Day 29) | Returns default or cached data | Keeps UX working | Data may be stale | Non-critical dependency |
| Timeout | Bounds a single attempt | Prevents hung threads | Does not recover by itself | Always, paired with retry |
| Queue-based retry / DLQ (Service Bus) | Message redelivered later, then dead-lettered | Durable, survives restarts, long delays | Async only, more infrastructure | Background processing, minutes-to-hours delays |
| Client-side hedging | Sends a second request if the first is slow | Cuts tail latency | Doubles load, needs idempotency | Latency-critical reads |
Patterns it combines with
- Timeout (per attempt) and Circuit Breaker: the standard trio. Order in a Polly pipeline (outer to inner): total timeout, retry, circuit breaker, attempt timeout.
- Idempotent Consumer (Day 17) and idempotency keys: required to make retries safe.
- Bulkhead (Day 27): limits how many concurrent retries can run so they cannot exhaust the pool.
- Fallback (Day 29): the last step when retries are exhausted.
- Transactional Outbox / Polling Publisher (Days 12, 14): the publisher itself retries delivery to the broker.
- Rate Limiter (Day 30): servers use it to shed load; clients see 429 and retry with backoff.
ADR-style justification
ADR-028: Standardise HTTP and database retry policy in the claims platform
- Status: Proposed
- Context:
ClaimsServicecallsPolicyService,FraudServiceand Azure SQL. During deployments and Azure maintenance we see 1-3% transient failures for 1-3 seconds. Teams have written ad-hoc retry loops with fixed delays and no jitter; one incident (retry storm againstPolicyService) extended an outage by 20 minutes. - Decision: Use
Microsoft.Extensions.Http.Resiliencefor all outboundHttpClientcalls (max 3 retries, exponential backoff, jitter, honourRetry-After, per-attempt timeout 3 s, total timeout 10 s, retries disabled for unsafe methods unless anIdempotency-Keyis sent). Use EF CoreEnableRetryOnFailurefor SQL. Retries live at the calling service only; the API gateway does not retry. Retry counts are exported as OpenTelemetry metrics and alerted on. - Consequences: (+) Fewer user-visible transient errors, one policy to reason about. (+) Consistent telemetry. (-) Worst-case latency rises by a few seconds. (-) Every unsafe endpoint must implement idempotency keys. (-) Teams must not add their own retries on top.
- Alternatives rejected: hand-rolled loops (inconsistent), retrying at the gateway (cannot know idempotency), unlimited retry queues (hide outages).
Azure implementation
Retry is mostly a client-side pattern implemented in your code, but many Azure services have built-in retry behaviour you should know and configure rather than duplicate. Verify current defaults in the service docs before relying on numbers.
Azure SDK clients (Azure.* libraries)
Azure SDK clients share a common RetryOptions model with Mode (Fixed or Exponential), MaxRetries, Delay, MaxDelay. Service Bus is a good example (ServiceBusRetryOptions); its documented defaults are exponential mode, 3 retries, 0.8 s base delay, 60 s max delay and 60 s try timeout.
var client = new ServiceBusClient(fqNamespace, new DefaultAzureCredential(), new ServiceBusClientOptions { RetryOptions = new ServiceBusRetryOptions { Mode = ServiceBusRetryMode.Exponential, MaxRetries = 5, Delay = TimeSpan.FromSeconds(0.8), MaxDelay = TimeSpan.FromSeconds(30), TryTimeout = TimeSpan.FromSeconds(30) } });Azure SQL Database / SQL Managed Instance
- Use EF Core
EnableRetryOnFailureorSqlClientconfigurable retry logic. Azure SQL can return transient errors such as 40613 (database unavailable), 40197, 49918 during failover and scaling. - Tier consideration: Serverless (General Purpose) tier can auto-pause; the first connection after a pause may take tens of seconds and fail, so use enough retries and a longer connect timeout for dev and test databases that use auto-pause.
Azure Functions
- Built-in retry policies exist for supported triggers (fixed delay and exponential backoff strategies configured via
host.jsonor attributes). Service Bus and Storage Queue triggers instead use the queue’s own delivery count and dead-letter behaviour; do not stack a function-level retry on top without thinking about duplicates. - Pricing: Consumption and Flex Consumption bill per execution and execution time, so retried executions cost money and each retry is a billable execution.
Azure Service Bus (queue-based retry)
- For asynchronous work, let the broker do the retrying:
MaxDeliveryCount(default 10) redelivers an abandoned message, then it moves to the dead-letter queue. Use scheduled messages or a delayed re-enqueue for long backoffs (minutes to hours). - Tiers: Standard and Premium support the features needed; Premium gives dedicated capacity and predictable latency (priced per messaging unit). Basic does not support topics or scheduled delivery. Check current pricing on the Azure pricing page.
Azure API Management (APIM)
- A
retrypolicy exists for the backend call inside a policy pipeline (condition,count,interval,max-interval,delta,first-fast-retry). Use it only for idempotent backends and only if you are not already retrying inside the service.
Azure Cosmos DB and Storage
- SDKs handle 429 (throttled) using the
x-ms-retry-after-msvalue; configureMaxRetryAttemptsOnRateLimitedRequestsandMaxRetryWaitTimeOnRateLimitedRequestsinCosmosClientOptions. Persistent 429s mean you need more provisioned RU/s or autoscale, not more retries.
Monitoring on Azure
- Send OpenTelemetry to Azure Monitor / Application Insights (
Azure.Monitor.OpenTelemetry.AspNetCore). In the Application Insights Failures and Dependencies views, look at dependency call counts per operation: a rising ratio of dependency calls to requests is a retry signal. Alert on retry rate and on dependency failure rate.
Reference architecture (text)
The Angular SPA calls Azure API Management or Application Gateway (no retries at this layer, rate limiting only). APIM routes to ClaimsService on Azure Container Apps or AKS. ClaimsService calls PolicyService using HttpClient with the resilience handler (retry + circuit breaker + timeouts) and writes to Azure SQL via EF Core with EnableRetryOnFailure. Claim-submitted events go to Azure Service Bus using the transactional outbox; consumers rely on Service Bus redelivery, MaxDeliveryCount, and a dead-letter queue, plus idempotent handlers. All services export traces and metrics through OpenTelemetry to Application Insights, with alerts on retry rate, dead-letter queue depth, and p99 latency.
Teaching guide for my team
2-minute beginner explanation
“Sometimes a call fails for a moment, like a phone line that is busy. If we give up immediately, the user sees an error even though a second try would work. So we try again a few times. But we wait a little longer each time (backoff) and add a random bit (jitter) so that hundreds of callers do not all retry at the same instant. We only retry problems that are temporary: timeouts and 503, not 400 or 404. And we only retry actions that are safe to repeat.”
5-minute intermediate explanation
Walk through the AddResilienceHandler snippet: max attempts, exponential delay, jitter, Retry-After, per-attempt timeout, total timeout. Then explain three risks: retry amplification (3 layers x 3 retries), duplicate side effects (need idempotency keys), and storms (need jitter and a circuit breaker). Show EF Core EnableRetryOnFailure and explain why user-managed transactions need CreateExecutionStrategy. Finish with where to retry (one layer only) and what to measure (retry count, p99 latency).
Hands-on exercise
Build a tiny flaky PolicyService (minimal API) that returns 503 for the first two requests of each minute and 200 otherwise, and a ClaimsService client.
- First call it with a plain
HttpClient: observe failures. - Add
AddResilienceHandlerwith 3 retries, exponential backoff and jitter. Log each attempt withOnRetry. - Change the flaky service to return 400 always: confirm there are no retries.
- Change it to fail 100% with 503 and add 50 parallel callers: measure how many calls hit the service with and without jitter.
Expected outcome: step 1 fails, step 2 succeeds on the third attempt with visible growing delays, step 3 shows exactly one call, step 4 shows that jitter spreads arrival times and that retries multiply load (up to 4x), motivating a circuit breaker.
Interview-style questions
- Why is jitter needed with exponential backoff? Without it, all clients that failed together retry together, producing synchronized spikes that keep the dependency overloaded. Randomising delays spreads the load.
- Which failures should not be retried? Permanent client-side errors (400, 401, 403, 404, 409, 422) and business rule violations; also non-idempotent operations without an idempotency key.
- What happens if the gateway, the service and the SDK each retry 3 times? Up to 27 calls for one user request (3 x 3 x 3), which can overwhelm the dependency. Retry at one layer only.
Mastery checklist
- I can explain the difference between transient and permanent errors and list at least five transient signals for HTTP and SQL.
- I can configure exponential backoff with jitter using
Microsoft.Extensions.Http.Resilienceand explain each option. - I can make a POST retry-safe with an idempotency key and explain why a timeout does not mean the request failed.
- I can use EF Core
EnableRetryOnFailurecorrectly, including wrapping user transactions in an execution strategy. - I can calculate worst-case latency and load amplification for a given retry configuration.
- I can explain how retry, timeout, circuit breaker, bulkhead and fallback are ordered in a pipeline.
- I can decide between client-side retry and broker-based retry (Service Bus redelivery and dead-letter queue).
- I can identify and fix retry storms and multi-layer retry amplification in an existing system using metrics.
Key takeaway
Retry only transient failures, only safe operations, only a few times, with exponential backoff and jitter, and only at one layer. Everything else turns a blip into an outage.
