Manikandan — Manikandan
Microservices

Day 15: Remote Procedure Invocation (RPI)

ManikandanManikandan
12 min read·Updated Aug 24, 2022

Remote Procedure Invocation (RPI) is the simplest way for one service to use another: the caller sends a request over the network (REST/HTTP, gRPC, or GraphQL), waits, and gets a response, as if it had called a local...

Intro

Remote Procedure Invocation (RPI) is the simplest way for one service to use another: the caller sends a request over the network (REST/HTTP, gRPC, or GraphQL), waits, and gets a response, as if it had called a local method. It is easy to understand, easy to debug, and gives an immediate answer — but it couples the caller to the callee’s availability and latency, so it must be used with timeouts, retries, and circuit breakers.

Why we need this

  • Many business questions need an answer now: “Is this policy active?”, “What is the claim’s current status?”, “Is this garage on the approved list?” A user is waiting on the screen.
  • Request-response is the mental model every developer already has. REST over HTTP is supported by every language, proxy, gateway, and browser.
  • It gives strong, immediate feedback: success, validation error, or failure, in one round trip.
  • Technically it is the default for client-to-service (Angular to API) and for many service-to-service queries. Even in event-driven systems, RPI is usually still used for queries.

What problem it solves

Problem: Services need simple synchronous request-response between them.

In our insurance claims system, ClaimsService must check with PolicyService that a policy is active and covers the loss type before accepting a First Notice of Loss (FNOL).

Without a defined RPI approach:

  • Each team invents its own ad-hoc HTTP calls, formats, error handling, and timeouts.
  • Callers copy database tables from other services “to avoid calling them”, breaking service boundaries.
  • Calls have no timeout, so a slow PolicyService hangs every ClaimsService thread.
  • Contracts are undocumented, so a field rename breaks callers silently.

When it is needed (and when it is NOT)

Fits:

  • Queries where the caller needs the answer to continue (policy validation, price quote, fraud score for a decision on screen).
  • Client-to-backend calls (Angular to API gateway/BFF).
  • Commands where the caller needs a definite accept/reject result immediately (create claim, return claim id).
  • Low-latency internal calls with typed contracts (gRPC).
  • Flexible client-driven queries across a graph of data (GraphQL).

Not a good fit:

  • Long-running work (claim settlement, document OCR taking minutes): use messaging or async request-reply with a status endpoint.
  • Broadcasting “something happened” to many consumers: use events (Day 11, Day 16).
  • Chains of 4+ synchronous hops: availability multiplies down (0.999^5 ≈ 99.5%). See Self-contained Service (Day 3).
  • Operations that must survive the callee being down: use a queue and retry later.

How to identify the problem (key signals)

  1. Timeouts and 5xx errors in ClaimsService whenever PolicyService is slow or deploying.
  2. Thread-pool starvation or many pending outbound requests in a service with low CPU.
  3. HttpClient created with new HttpClient() per call; socket exhaustion (SocketException, TIME_WAIT pile-up).
  4. Different services use different retry, timeout, and error formats; no shared conventions.
  5. Distributed traces show long serial waterfalls (A → B → C → D) with total latency equal to the sum.
  6. Breaking API changes discovered in production; callers deserialise fields that no longer exist.
  7. A retry storm: one failing dependency doubles or triples traffic during an incident.

Flow Diagram

A synchronous request-response call protected by timeouts, retry and a circuit breaker.

flowchart LR
U["Angular FNOL form"] --> C["ClaimsService"]
C --> H["Typed HttpClient + resilience handler"]
H --> T{"Circuit open?"}
T -- "yes" --> FF["Fail fast - 503 ProblemDetails"]
T -- "no" --> P["PolicyService - REST or gRPC"]
P -- "200 coverage" --> C
P -- "timeout / 5xx" --> R["Retry with backoff, idempotent calls only"]
R --> P

Level 1: Beginner

Analogy: A phone call. You dial, wait for the other person to answer, ask your question, and get an answer while you stay on the line. If they do not pick up, you must decide how long to wait and whether to call again.

Three styles:

StyleTransportContractTypical use
RESTHTTP/1.1 or HTTP/2 + JSONOpenAPIPublic and most internal APIs, browsers
gRPCHTTP/2 + Protobuf.proto fileInternal, low-latency, streaming
GraphQLHTTP + JSONSchemaClient-driven queries (often at a BFF)

Minimal example: PolicyService (ASP.NET Core minimal API, .NET 10):

PolicyService/Program.cs
var builder = WebApplication.CreateBuilder(args);
var app = builder.Build();
app.MapGet("/policies/{policyNumber}/coverage", (string policyNumber) =>
policyNumber == "POL-1001"
? Results.Ok(new CoverageResponse(policyNumber, true, 50_000m))
: Results.NotFound());
app.Run();
public record CoverageResponse(string PolicyNumber, bool IsActive, decimal CoverageLimit);

Caller: ClaimsService using a typed HttpClient:

ClaimsService/PolicyClient.cs
public record CoverageResponse(string PolicyNumber, bool IsActive, decimal CoverageLimit);
public class PolicyClient(HttpClient http)
{
public async Task<CoverageResponse?> GetCoverageAsync(string policyNumber, CancellationToken ct)
{
var response = await http.GetAsync($"policies/{policyNumber}/coverage", ct);
if (response.StatusCode == System.Net.HttpStatusCode.NotFound) return null;
response.EnsureSuccessStatusCode();
return await response.Content.ReadFromJsonAsync<CoverageResponse>(ct);
}
}
ClaimsService/Program.cs
builder.Services.AddHttpClient<PolicyClient>(c =>
c.BaseAddress = new Uri("https://policy-service.internal/"));

Key beginner rules: never new HttpClient() in a loop; always pass a CancellationToken; always handle “not found” and “service down” explicitly.

Level 2: Intermediate

6.1 Resilient REST client (.NET)

Add the standard resilience handler (package Microsoft.Extensions.Http.Resilience), which combines rate limiter, total timeout, retry, circuit breaker, and attempt timeout:

builder.Services
.AddHttpClient<PolicyClient>(c => c.BaseAddress = new Uri("https://policy-service.internal/"))
.AddStandardResilienceHandler(o =>
{
o.AttemptTimeout.Timeout = TimeSpan.FromSeconds(2);
o.TotalRequestTimeout.Timeout = TimeSpan.FromSeconds(6);
o.Retry.MaxRetryAttempts = 2;
// Only GET is safe to retry by default in your design; disable for POST unless idempotent:
o.Retry.DisableForUnsafeHttpMethods();
});

Note: CircuitBreaker.SamplingDuration must be at least twice AttemptTimeout; validation fails at startup otherwise.

6.2 gRPC contract and service

policy.proto
syntax = "proto3";
option csharp_namespace = "PolicyService.Grpc";
service PolicyLookup {
rpc GetCoverage (CoverageRequest) returns (CoverageReply);
}
message CoverageRequest { string policy_number = 1; }
message CoverageReply {
string policy_number = 1;
bool is_active = 2;
string coverage_limit = 3; // decimal as string to avoid float rounding
}
// Server (Grpc.AspNetCore)
public class PolicyLookupService : PolicyLookup.PolicyLookupBase
{
public override Task<CoverageReply> GetCoverage(CoverageRequest request, ServerCallContext context)
=> Task.FromResult(new CoverageReply
{
PolicyNumber = request.PolicyNumber,
IsActive = true,
CoverageLimit = "50000.00"
});
}
// Program.cs: builder.Services.AddGrpc(); app.MapGrpcService<PolicyLookupService>();
// Client (Grpc.Net.ClientFactory): retries via gRPC service config, not the HTTP resilience handler
builder.Services.AddGrpcClient<PolicyLookup.PolicyLookupClient>(o =>
o.Address = new Uri("https://policy-service.internal"))
.ConfigureChannel(o => o.ServiceConfig = new Grpc.Net.Client.Configuration.ServiceConfig
{
MethodConfigs =
{
new Grpc.Net.Client.Configuration.MethodConfig
{
Names = { Grpc.Net.Client.Configuration.MethodName.Default },
RetryPolicy = new Grpc.Net.Client.Configuration.RetryPolicy
{
MaxAttempts = 3,
InitialBackoff = TimeSpan.FromMilliseconds(200),
MaxBackoff = TimeSpan.FromSeconds(2),
BackoffMultiplier = 2,
RetryableStatusCodes = { Grpc.Core.StatusCode.Unavailable }
}
}
}
});

Use a deadline on every call: client.GetCoverageAsync(req, deadline: DateTime.UtcNow.AddSeconds(2)).

6.3 Angular calling the backend

// claims.service.ts (Angular, standalone, functional style)
import { inject, Injectable } from '@angular/core';
import { HttpClient } from '@angular/common/http';
import { Observable } from 'rxjs';
export interface Claim { id: string; policyNumber: string; status: string; }
@Injectable({ providedIn: 'root' })
export class ClaimsApi {
private http = inject(HttpClient);
getClaim(id: string): Observable<Claim> {
return this.http.get<Claim>(`/api/claims/${id}`);
}
}

Register with provideHttpClient(withInterceptors([authInterceptor])) in app.config.ts.

6.4 Error contract

Return RFC 9457 Problem Details so all services fail the same way:

builder.Services.AddProblemDetails();
app.UseExceptionHandler();
app.UseStatusCodePages();

6.5 Database angle

Do not let ClaimsService read PolicyService’s SQL Server tables. Store only the PolicyNumber (a reference) in the Claims database and fetch details by RPI, or cache a minimal snapshot (see Day 3 and Day 10).

Level 3: Advanced

Performance

  • Reuse connections: IHttpClientFactory / typed clients, or a single long-lived GrpcChannel. Set PooledConnectionLifetime (e.g., 2 minutes) so DNS changes are honoured.
  • Prefer HTTP/2 for many concurrent calls; gRPC’s binary Protobuf is typically smaller and faster to serialise than JSON, though the real gain must be measured for your payloads.
  • Avoid chatty calls (N+1 across services). Batch endpoints (POST /policies/coverage:batch) or API Composition (Day 10).
  • Propagate the deadline: if the user request has 5 s left, downstream calls must not use a fresh 30 s timeout.

Scalability

  • Stateless services behind a load balancer. For gRPC over HTTP/2, an L4 balancer pins all calls of a connection to one instance; use an L7 balancer (Application Gateway, Envoy, Kubernetes ingress) or client-side balancing.

Security

  • Validate a JWT (Day 38) at every service; forward the token or use token exchange, do not trust “internal network”.
  • TLS everywhere (mTLS via mesh where required). Use Managed Identity to call Azure resources instead of secrets.
  • Input validation and payload size limits; never expose internal DTOs to public clients.

Failure modes

  • Slow is worse than down: threads and connections pile up. Set timeouts short and explicit.
  • Retry amplification: 3 layers each retrying 3 times = 27 attempts. Retry at one layer only, with backoff and jitter (Day 28).
  • Non-idempotent retries create duplicate claims. Use an Idempotency-Key header for POST (Day 17).
  • Partial failure: the server processed the request but the response was lost. The caller cannot know; design for idempotency.
  • Cascading failure: add circuit breaker (Day 26) and bulkhead (Day 27).

Common mistakes

  1. new HttpClient() per request.
  2. No timeout (default HttpClient.Timeout is 100 s).
  3. Blocking on async (.Result) inside request handlers.
  4. Sharing DTO libraries across services so every change forces a lockstep release.
  5. Treating every failure as retryable (400/404 are not).
  6. Long synchronous call chains.
  7. Versioning by breaking changes instead of additive evolution.

Level 4: Expert and Architect view

Comparison of RPI styles

CriteriaREST/JSONgRPCGraphQL
Browser supportNativeNeeds gRPC-Web + proxyNative
ContractOpenAPI (optional, recommended).proto (mandatory, codegen)Schema (mandatory)
Payload efficiencyText, largerBinary, compactText; client selects fields
StreamingSSE/WebSocket separatelyBuilt-in (server, client, bidi)Subscriptions (extra transport)
CachingHTTP caching, CDN friendlyManualHarder (POST, per-query)
Tooling/debuggingExcellent (curl, browser)Needs grpcurl, PostmanGood (GraphiQL)
Best forPublic APIs, CRUD, most servicesInternal, high-throughput, polyglotAggregating for UIs (BFF)

RPI vs Messaging

AspectRPIMessaging (Day 16)
Coupling in timeBoth must be upBroker decouples
Latency to answerImmediateEventual
ComplexityLowHigher (broker, idempotency)
BackpressureWeakNatural (queue)

Combines with: API Gateway (19), BFF (20), Service Registry/Discovery (21–25), Circuit Breaker (26), Bulkhead (27), Retry & Backoff (28), Fallback (29), Rate Limiter (30), Distributed Tracing (31), Access Token (38), Consumer-Driven Contract Test (51).

ADR: Use synchronous RPI (REST for external and most internal calls, gRPC for hot internal paths)

  • Status: Proposed
  • Context: ClaimsService needs immediate policy coverage validation during FNOL, with a user waiting. Teams are .NET-heavy; a mobile and Angular client exist.
  • Decision: Use REST/JSON with OpenAPI for client-facing and default service-to-service calls. Use gRPC only for internal, high-volume, latency-sensitive paths (e.g., fraud scoring) after measurement. All outbound calls use typed clients, explicit timeouts, one retry layer, and a circuit breaker. Errors use Problem Details. No shared database access.
  • Consequences: (+) Simple, debuggable, immediate answers. (−) Runtime coupling and availability dependency on PolicyService; mitigated by resilience policies, caching of stable data, and moving long-running flows to messaging. Two protocols increase tooling surface.
  • Alternatives rejected: GraphQL between services (adds gateway complexity without clear need); messaging for coverage validation (user needs a synchronous answer).

Azure implementation

Services

  • Azure Container Apps / AKS / App Service: host the REST and gRPC services. Container Apps supports HTTP/2 and gRPC ingress; internal ingress keeps services private to the environment.
  • Azure API Management (APIM): front door for REST APIs (auth, rate limits, versioning, transformation). gRPC support in APIM is limited to specific tiers/gateways, so verify support for your tier before planning gRPC through APIM.
  • Azure Application Gateway / Front Door: L7 routing, WAF, TLS. Application Gateway supports HTTP/2 to backends for gRPC scenarios (verify current constraints).
  • Microsoft Entra ID + Managed Identity: token issuance and service identities.
  • Azure Monitor / Application Insights (OpenTelemetry): dependency calls, failures, and distributed traces across RPI hops.
  • Azure Private Link / VNet integration: keep service-to-service traffic private.

Configuration highlights

  1. Deploy PolicyService in Container Apps with internal ingress; set transport to http2 for gRPC.
  2. Publish the public REST facade via APIM with a JWT validation policy (validate-jwt) and rate-limit-by-key.
  3. Enable Application Insights with the Azure Monitor OpenTelemetry distro so each outbound call appears as a dependency.
  4. Set scale rules (HTTP concurrent requests) so slow callees scale out before timeouts occur.

Pricing and tier considerations

  • APIM offers Consumption, Developer, Basic, Standard, Premium, and the newer v2 tiers (Basic v2, Standard v2, Premium v2). Consumption is pay-per-call (good for spiky/low traffic); dedicated tiers charge per unit-hour with capacity and features (VNet, multi-region) rising by tier. Prices and feature matrices change; confirm on the Azure pricing page and the APIM feature comparison before choosing.
  • Container Apps has a Consumption plan (billed per vCPU-second/GiB-second/request, with a monthly free grant) and Dedicated workload profiles. Internal RPI traffic inside one environment does not incur APIM charges.
  • Application Insights and Log Analytics are billed by ingested GB: sample dependency telemetry on high-volume RPI paths.

Reference architecture (text) Angular SPA (Static Web Apps) → Front Door + WAF → APIM (JWT validation, rate limit) → Container Apps environment (VNet): ClaimsService → (REST/gRPC, internal ingress) → PolicyService and FraudService. Each service owns its Azure SQL / PostgreSQL Flexible Server database. All services emit OpenTelemetry to Application Insights; secrets and config come from Key Vault and App Configuration via Managed Identity.

Teaching guide for my team

2-minute beginner explanation “RPI is a phone call between services. ClaimsService calls PolicyService, asks ‘is this policy active?’, waits, and gets an answer. It is simple, but if the other side is slow or down, you are stuck waiting. So we always set a timeout, and we decide what to do if there is no answer.”

5-minute intermediate explanation Walk through: REST vs gRPC vs GraphQL (table in Level 1); why we use IHttpClientFactory; the standard resilience handler (timeout, retry, circuit breaker); why retries need idempotency; why deadlines must propagate; why we never share databases. Show a trace where three serial calls add up to the total latency, and show the effect of a circuit breaker during a PolicyService outage.

Hands-on exercise Build PolicyService (minimal API) and ClaimsService (typed client). Steps: (1) call succeeds; (2) add a 5 s Task.Delay to PolicyService and set AttemptTimeout to 2 s; (3) stop PolicyService and observe retries and the circuit opening; (4) return Problem Details for failures. Expected outcome: Step 2 fails in about 2 s per attempt and stops at the total timeout; step 3 shows retry attempts in logs, then BrokenCircuitException failing fast; ClaimsService returns HTTP 503 with a Problem Details body instead of hanging.

Interview questions

  1. Why not create a new HttpClient per request? It exhausts sockets and ignores DNS changes; use IHttpClientFactory or a long-lived client.
  2. When would you choose gRPC over REST? Internal, high-throughput, low-latency, typed contracts or streaming; not for browsers or public APIs without gRPC-Web.
  3. A retry created two claims. Why, and how do you fix it? The first request succeeded but the response was lost; make the operation idempotent with an idempotency key or natural key.

Mastery checklist

  • Can explain when RPI is right and when messaging is better, with an insurance example.
  • Uses IHttpClientFactory and never creates raw HttpClient per call.
  • Sets explicit timeouts and deadlines and propagates them downstream.
  • Applies retry only to safe/idempotent calls, at a single layer, with backoff and jitter.
  • Can write and version a .proto contract and an OpenAPI contract additively.
  • Returns consistent Problem Details errors across services.
  • Reads a distributed trace to find serial-call latency and dependency failures.
  • Can justify REST vs gRPC vs GraphQL in an ADR.

Key takeaway

RPI is the simplest way for services to talk, but every call is a bet on someone else’s availability: always use timeouts, one disciplined retry layer, and a circuit breaker, and keep long-running work off the synchronous path.

Interactive Architectural Roadmaps

Explore Complete Roadmaps & Pattern Checklists

Track your learning with interactive checklists for all 23 Gang of Four patterns and modern Microservice architecture patterns.

Share:
Back to Blog

Related Posts

View All Posts
Microservices

Day 18: Domain-specific Protocol

A domain-specific protocol means choosing a wire protocol that fits the *shape of the traffic* instead of defaulting to HTTP/JSON for everything.

Manikandan
Manikandan·22 min read
Microservices

Day 17: Idempotent Consumer

An idempotent consumer is a message handler that produces the same end result whether a message is processed once or several times.

Manikandan
Manikandan·15 min read
Microservices

Day 16: Messaging

Messaging means services talk to each other by putting messages on a durable broker (Azure Service Bus, Kafka, RabbitMQ) instead of calling each other directly and waiting.

Manikandan
Manikandan·10 min read