An API Gateway is a single entry point that sits between external clients (browser, mobile app, partner systems) and your internal microservices.
Intro
An API Gateway is a single entry point that sits between external clients (browser, mobile app, partner systems) and your internal microservices. Instead of every client knowing the address of every service, clients talk to one URL. The gateway then handles cross-cutting edge concerns such as routing, authentication, rate limiting, TLS termination, and response aggregation, and forwards each request to the right service. In our insurance claims system, the Angular claims portal calls https://api.claimsco.example and the gateway decides whether the request goes to Claims, Policy, Payments, or Documents services.
Why we need this
- Clients should not know the internal topology. Services move, scale, split, and get renamed. The gateway gives clients one stable contract.
- Edge concerns belong in one place. Token validation, CORS, TLS, rate limiting, request size limits, and IP filtering are needed for every public route. Implementing them in 20 services means 20 chances to get it wrong.
- Reduce chattiness. A claim details screen needs claim, policy holder, documents, and payment status. Without a gateway, the browser makes four calls over a slow mobile network.
- Security perimeter. Only the gateway is exposed to the internet. Internal services live in a private network.
- Evolution. You can version APIs, canary a new service, or strangle a legacy monolith behind the same public URL (see Day 49).
What problem it solves
Problem (from the topic list): direct client-to-service calls cause chatter, security exposure, and CORS overhead.
Without a gateway in the claims system:
- The Angular app holds five base URLs in
environment.ts. Renaming or moving the Payments service breaks the released frontend. - Each of the five services must be internet-facing, has its own public TLS certificate, its own CORS configuration, and its own JWT validation code. One service forgets to validate the
audclaim and becomes an open door. - The mobile app makes 6 sequential HTTP calls to render one screen, and the screen takes 3 seconds on 4G.
- A partner integration hammers
POST /claimsand there is no single place to throttle it. - Security team asks “who accessed what and from where?” and logs are in five formats in five places.
When it is needed (and when it is NOT)
Needed when:
- You have more than a few services exposed to external clients.
- You have several client types (web, mobile, partners) and need central authentication and throttling.
- You need a public API product (keys, subscriptions, quotas, developer portal).
- You are migrating a monolith and need a routing layer in front of it.
Not needed, or wrong choice, when:
- You have one deployable (a modular monolith). A gateway adds a hop and an extra component to operate for no gain.
- Only internal service-to-service traffic is involved. A gateway is an edge pattern; internal calls should use service discovery, a mesh, or direct calls (Days 21-25, 48).
- You are tempted to put business logic in it. A gateway that contains claim-approval rules becomes a distributed monolith’s new god component.
- Every screen needs a different aggregated shape per client. Prefer one BFF per client type (Day 20) instead of one bloated gateway.
How to identify the problem (key signals)
- The frontend config lists many backend base URLs, and every service move requires a frontend release.
- Browser dev tools show 5+ API calls per screen load, with a waterfall of dependent calls.
- Multiple services contain copy-pasted JWT validation, CORS, and rate-limit code, with subtle differences.
- Security scan reports many internet-exposed endpoints and ports, each with a different TLS posture.
- A single abusive client (or a bot) can take down one service, and nobody can throttle it without redeploying that service.
- Support cannot follow one user request across services because there is no common entry point or correlation ID.
- Mobile users complain of slow screens while server-side latency per service is low (the delay is round trips).
Flow Diagram
One front door validates, throttles and routes; services stay private.
flowchart LR C1["Angular portal"] --> FD["Front Door + WAF"] C2["Mobile app"] --> FD C3["Partners"] --> FD FD --> GW["API Gateway: JWT, rate limit, CORS, routing"] GW -- "/claims/*" --> CS["Claims service"] GW -- "/policies/*" --> PS["Policy service"] GW -- "/payments/*" --> PY["Payments service"] GW -- "/documents/*" --> DS["Documents service"]Level 1: Beginner
Analogy: A hotel reception desk. Guests do not walk into the kitchen, laundry, and accounts office. They talk to reception, who checks their identity, and routes the request to the correct department.
Minimal example: a YARP gateway in .NET. YARP (Yet Another Reverse Proxy) is Microsoft’s reverse proxy library and ships as the Yarp.ReverseProxy NuGet package. It works on the current .NET LTS (.NET 10).
// Program.cs (dotnet new web; dotnet add package Yarp.ReverseProxy)var builder = WebApplication.CreateBuilder(args);
builder.Services .AddReverseProxy() .LoadFromConfig(builder.Configuration.GetSection("ReverseProxy"));
var app = builder.Build();app.MapReverseProxy();app.Run();{ "ReverseProxy": { "Routes": { "claims-route": { "ClusterId": "claims-cluster", "Match": { "Path": "/claims/{**catch-all}" } }, "policies-route": { "ClusterId": "policy-cluster", "Match": { "Path": "/policies/{**catch-all}" } } }, "Clusters": { "claims-cluster": { "Destinations": { "d1": { "Address": "http://claims-service:8080/" } } }, "policy-cluster": { "Destinations": { "d1": { "Address": "http://policy-service:8080/" } } } } }}A request to GET /claims/CLM-1001 on the gateway is forwarded to http://claims-service:8080/claims/CLM-1001. The client only knows the gateway.
Level 2: Intermediate
Real applications add authentication, rate limiting, and CORS at the gateway, and aggregate a few calls for the UI.
Gateway: JWT validation, rate limiting, CORS, and per-route policies (.NET):
using System.Threading.RateLimiting;using Microsoft.AspNetCore.RateLimiting;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddAuthentication("Bearer") .AddJwtBearer("Bearer", o => { o.Authority = builder.Configuration["Auth:Authority"]; // Entra ID / Entra External ID o.Audience = builder.Configuration["Auth:Audience"]; });
builder.Services.AddAuthorization(o =>{ o.AddPolicy("ClaimsHandler", p => p.RequireRole("claims.handler"));});
builder.Services.AddRateLimiter(o =>{ o.RejectionStatusCode = StatusCodes.Status429TooManyRequests; o.AddPolicy("per-user", ctx => RateLimitPartition.GetFixedWindowLimiter( ctx.User.Identity?.Name ?? ctx.Connection.RemoteIpAddress?.ToString() ?? "anon", _ => new FixedWindowRateLimiterOptions { PermitLimit = 100, Window = TimeSpan.FromMinutes(1) }));});
builder.Services.AddCors(o => o.AddPolicy("portal", p => p .WithOrigins("https://portal.claimsco.example") .AllowAnyHeader().AllowAnyMethod()));
builder.Services.AddReverseProxy() .LoadFromConfig(builder.Configuration.GetSection("ReverseProxy"));
var app = builder.Build();app.UseCors("portal");app.UseAuthentication();app.UseAuthorization();app.UseRateLimiter();app.MapReverseProxy();app.Run();Attach the policies to routes in configuration:
"claims-route": { "ClusterId": "claims-cluster", "AuthorizationPolicy": "ClaimsHandler", "RateLimiterPolicy": "per-user", "CorsPolicy": "portal", "Match": { "Path": "/claims/{**catch-all}" }}Aggregation endpoint (kept thin, no business rules):
app.MapGet("/api/claim-summary/{id}", async (string id, IHttpClientFactory f, CancellationToken ct) =>{ var claims = f.CreateClient("claims"); var payments = f.CreateClient("payments");
var claimTask = claims.GetFromJsonAsync<ClaimDto>($"claims/{id}", ct); var paymentTask = payments.GetFromJsonAsync<PaymentStatusDto>($"payments/by-claim/{id}", ct); await Task.WhenAll(claimTask, paymentTask);
return Results.Ok(new ClaimSummaryDto(claimTask.Result!, paymentTask.Result!));}).RequireAuthorization();
public record ClaimDto(string Id, string Status, decimal Amount);public record PaymentStatusDto(string ClaimId, string State);public record ClaimSummaryDto(ClaimDto Claim, PaymentStatusDto Payment);Angular client: one base URL, one interceptor (functional interceptors, Angular 15+; standalone bootstrap, Angular 17+).
import { HttpInterceptorFn } from '@angular/common/http';import { inject } from '@angular/core';import { TokenService } from './token.service';
export const authInterceptor: HttpInterceptorFn = (req, next) => { const token = inject(TokenService).accessToken(); return next(token ? req.clone({ setHeaders: { Authorization: `Bearer ${token}` } }) : req);};
// app.config.tsimport { ApplicationConfig } from '@angular/core';import { provideHttpClient, withInterceptors } from '@angular/common/http';
export const appConfig: ApplicationConfig = { providers: [provideHttpClient(withInterceptors([authInterceptor]))]};Database angle: the gateway itself should be stateless. If you need rate-limit counters shared across gateway instances, use a distributed store such as Redis, not SQL Server or PostgreSQL. Keep the claims and policy databases private behind their services.
Level 3: Advanced
Performance
- Every request pays one extra network hop. Keep the gateway in the same region and network as the services, and reuse HTTP connections (YARP pools them via
SocketsHttpHandler). - Do not buffer large uploads (claim photos, PDFs) in the gateway. Stream them and set explicit max body sizes per route.
- Cache safe, read-heavy, non-personalised responses at the gateway (for example, reference data such as claim types).
Scalability
- Run at least 2 instances across availability zones, stateless, autoscaled on CPU and request rate.
- Distributed rate limiting needs shared state (Redis) or a managed gateway that provides it.
Security
- Terminate TLS at the gateway and use TLS or private networking between gateway and services. Do not assume “internal means trusted”.
- Validate the token at the edge (issuer, audience, expiry, signature), then forward identity to services (Day 38). Services should still authorise what the caller may do; the gateway is not the only line of defence.
- Strip inbound headers you set yourself (such as
X-User-Id) so clients cannot spoof them. - Restrict request size, headers size, and timeouts to limit abuse.
Failure modes
- Single point of failure. If the gateway is down, everything is down. Run multiple instances and health-probe them.
- Timeout cascades. Set timeouts and retries deliberately (Days 26-29). Retrying non-idempotent
POST /claimsat the gateway can create duplicate claims unless the service is idempotent (Day 17). - Gateway becomes a bottleneck for both traffic and for team change velocity if every route change needs one central team.
Common mistakes
- Putting business logic or data transformation rules in the gateway.
- One gateway to serve all clients with client-specific branches (use BFFs, Day 20).
- Deep aggregation of many services synchronously, making the gateway as slow as its slowest dependency.
- Forgetting to propagate a correlation ID /
traceparentheader (Day 31). Both YARP and Azure API Management can forward it, but check your configuration. - Treating the gateway as a substitute for authorisation inside services.
Level 4: Expert and Architect view
Options compared
| Option | Strengths | Weaknesses | Best fit |
|---|---|---|---|
| YARP (code-first, self-hosted) | Full control in C#, testable, fits .NET teams, free | You operate and scale it; no built-in portal or subscriptions | Internal platform gateway for a .NET shop |
| Azure API Management | Managed, policies, subscriptions/keys, developer portal, analytics | Cost, policy expression learning curve, tier constraints | Public/partner APIs, governance |
| Azure Application Gateway (WAF) | L7 load balancing, WAF, TLS offload in a VNet | Not an API management layer; limited API-level policy | Regional WAF in front of AKS/App Service |
| Azure Front Door | Global edge, CDN, WAF, failover | Not designed for per-API policies | Global entry, static content, multi-region routing |
| Ocelot | Simple .NET gateway with aggregation | Smaller ecosystem than YARP for new work; verify maintenance status before adopting | Existing Ocelot estates |
| Envoy / Istio ingress / NGINX | High performance, proven | Different skill set; config complexity | Kubernetes-centric platforms |
Combines with: BFF (Day 20) for per-client shapes; Access Token (38) for edge auth; Rate Limiter (30); Circuit Breaker and Retry (26, 28); Distributed Tracing (31); Health Check API (36); Strangler Fig (49); Service Discovery (21-25).
ADR (for architecture review)
- Title: ADR-019 Introduce an edge API gateway for the claims platform.
- Status: Proposed.
- Context: The claims portal and partner API call 6 services directly. Auth, CORS, and throttling are implemented inconsistently. The frontend hard-codes service URLs.
- Decision: Route all external traffic through Azure API Management (public and partner APIs) in front of a YARP-based internal gateway on Azure Container Apps for routing and thin aggregation. Token validation and rate limiting happen at the edge. Services stay private.
- Consequences: (+) one public contract, consistent security, faster mobile screens. (-) extra hop and cost, new component to run and monitor, risk of gateway logic creep.
- Guardrails: no business rules in the gateway; each service team owns its route config in a shared repo with review; gateway SLO of 99.95%.
- Alternatives rejected: direct calls (security, coupling), single custom gateway without APIM (no partner subscription management).
Azure implementation
Services
- Azure API Management (APIM): the managed gateway. Provides policies (validate-jwt, rate-limit-by-key, rewrite-uri, cache-lookup), subscriptions and keys, and a developer portal.
- Azure Application Gateway with WAF, or Azure Front Door with WAF: OWASP-rule protection and TLS in front of the gateway.
- Azure Container Apps or AKS: hosts a self-managed YARP gateway.
- Microsoft Entra ID (or Entra External ID for customers): issues the tokens the gateway validates.
- Azure Key Vault: certificates and secrets. Application Insights / Azure Monitor: telemetry and traces.
Example APIM inbound policy
<policies> <inbound> <base /> <validate-jwt header-name="Authorization" failed-validation-httpcode="401"> <openid-config url="https://login.microsoftonline.com/{tenant-id}/v2.0/.well-known/openid-configuration" /> <required-claims> <claim name="aud"><value>api://claims-api</value></claim> </required-claims> </validate-jwt> <rate-limit-by-key calls="100" renewal-period="60" counter-key="@(context.Request.Headers.GetValueOrDefault("Authorization","").AsJwt()?.Subject)" /> <set-header name="X-Correlation-Id" exists-action="skip"> <value>@(Guid.NewGuid().ToString())</value> </set-header> </inbound></policies>Configuration notes
- Import the OpenAPI definition of each service as an API in APIM, then apply policies at product, API, or operation scope.
- Put the backend services in a VNet and expose them only via private endpoints or internal ingress. Deploy APIM in internal or external VNet mode as your network design requires (VNet features vary by tier).
- Use managed identity from the gateway to Key Vault, never stored secrets.
- Send APIM diagnostics to Log Analytics and Application Insights.
Pricing and tier considerations (verify current prices on the Azure pricing page before committing; they change by region and over time)
- Consumption: serverless, pay per call, scales to zero. Good for low or spiky traffic and for learning. It has feature limits (for example, no VNet injection).
- Developer: non-production only, no SLA.
- Basic / Standard / Premium (classic): dedicated capacity with a monthly unit fee; Premium adds multi-region, availability zones, and full VNet support.
- Basic v2 / Standard v2 / Premium v2: newer generation with faster provisioning and scaling and simpler VNet integration; check the Microsoft Learn tier comparison for exactly which features each v2 tier supports today.
- Self-hosted YARP has no licence cost; you pay for the compute (Container Apps or AKS nodes) and your team’s operating effort.
Reference architecture (text)
Internet clients (Angular portal, mobile app, partners) -> Azure Front Door with WAF -> Azure API Management (JWT validation, throttling, subscription keys) -> private link into the VNet -> YARP gateway on Azure Container Apps (routing, thin aggregation) -> Claims, Policy, Payments, and Documents services (each with its own database: Azure SQL / PostgreSQL Flexible Server). Entra ID issues tokens; Key Vault holds certificates; Application Insights receives traces from every hop using W3C trace context.
Teaching guide for my team
2-minute beginner explanation
“Imagine our claims system has five back-end services. Without a gateway, the website has to know all five addresses and log in to each one. With an API Gateway, the website talks to one door. The door checks who you are, slows down anyone hammering it, and sends you to the right service. If we move a service, the website does not change.”
5-minute intermediate explanation
Explain the three jobs: routing (path to cluster), cross-cutting policy (JWT, CORS, rate limit, headers), and light aggregation. Show the YARP config and the AuthorizationPolicy and RateLimiterPolicy route settings. Discuss the risks: single point of failure, an extra hop, and logic creep. Contrast with BFF: one gateway for the platform edge, one BFF per client type if their needs diverge. Finish with the rule “the gateway routes and protects; the services decide.”
Hands-on exercise
Build a YARP gateway in front of two stub minimal APIs (claims-service, policy-service).
- Route
/claims/**and/policies/**. - Add JWT bearer auth using a local test issuer (or
dotnet user-jwts create). - Add a
per-userfixed-window limit of 5 requests per minute. - Call the endpoint 6 times with a valid token.
Expected outcome: calls 1-5 return 200 from the right service, call 6 returns 429; a call without a token returns 401; the services are not reachable directly from outside the Docker network.
Interview questions
- How is an API Gateway different from a load balancer? A load balancer spreads traffic across identical instances, usually at L4 or basic L7. A gateway understands APIs: it routes by path, validates tokens, throttles per client, transforms, and can aggregate.
- What is the biggest risk of an API Gateway? It is a single point of failure and a magnet for business logic. Mitigate with multiple instances across zones, strict “no business rules” guardrails, and monitoring.
- Should the gateway validate tokens if services also validate them? Yes at the edge, to reject bad traffic early and centralise policy. Services still verify authorisation (defence in depth), typically using the forwarded signed token.
Mastery checklist
- Can explain what belongs in a gateway (routing, authN, throttling, CORS, TLS) and what does not (business rules).
- Can configure a YARP route and cluster with an authorisation and rate-limit policy.
- Can write an APIM policy using
validate-jwtandrate-limit-by-key. - Can distinguish API Gateway, BFF, load balancer, Front Door, and Application Gateway.
- Can explain why retrying
POSTat the gateway is dangerous without idempotency. - Can design a gateway deployment that removes it as a single point of failure.
- Can propagate correlation and trace headers end to end.
- Can justify the choice between APIM, YARP, or both in an ADR.
Key takeaway
An API Gateway gives clients one stable, secure front door and keeps edge concerns in one place, but it must stay thin: it routes and protects, the services own the business logic.
