A domain-specific protocol means choosing a wire protocol that fits the *shape of the traffic* instead of defaulting to HTTP/JSON for everything.
Intro
A domain-specific protocol means choosing a wire protocol that fits the shape of the traffic instead of defaulting to HTTP/JSON for everything. Internal service-to-service calls with strict contracts fit gRPC (binary Protobuf over HTTP/2). Live screens that must update without refreshing fit WebSockets (via SignalR). Thousands of small, battery-powered devices on flaky networks fit MQTT (a lightweight publish/subscribe protocol). Using the right protocol cuts latency, bandwidth and server cost, but it also adds tooling, debugging and operational complexity, so it must be earned by a real need.
Running example for all lessons: an insurance claims system (Claims API, Fraud Scoring, Policy, Payments, Adjuster web app in Angular, field devices such as dashcams and drone/sensor kits used at loss sites).
Why we need this
HTTP/1.1 + JSON is a great default: it is human-readable, cacheable, and every tool understands it. But it was designed for document retrieval, not for every communication pattern:
- Text serialization is expensive. JSON repeats field names in every message, encodes numbers as text, and must be parsed. At tens of thousands of calls per second between internal services, CPU and bandwidth costs become measurable.
- Request/response only. Plain HTTP cannot push. To show “claim approved” on an adjuster’s screen, the browser has to poll, which wastes calls and still feels laggy.
- Heavy for tiny devices. A dashcam on a 3G/4G link sending a 200-byte telemetry reading every second cannot afford HTTP headers, TLS handshakes per request, and no built-in delivery guarantees.
- Weak contracts. JSON has no built-in schema. Two teams can disagree on a field type and only discover it in production.
Domain-specific protocols exist because different traffic shapes have different best-fit answers: RPC with contracts (gRPC), server push / bidirectional real-time (WebSocket/SignalR), and constrained-device pub/sub (MQTT).
What problem it solves
The problem from the topic list: HTTP/JSON adds serialization latency and bandwidth overhead for specialized flows.
Concrete claims-system examples of what goes wrong without it:
- The Claims service calls Fraud Scoring 4,000 times per second at month-end. Every call sends ~1.5 KB of JSON and burns CPU on serialization. p99 latency climbs, and you scale out servers purely to parse text.
- Adjusters keep a claim screen open. The Angular app polls
GET /claims/123/statusevery 5 seconds. With 800 adjusters that is 9,600 requests per minute, 99% of which return “no change”. - 20,000 field devices send telemetry over HTTPS POST. Each request pays header overhead; on weak networks retries create duplicate readings; devices drain batteries on reconnects. There is no “last will” to tell you a device died.
The fix is not “replace REST everywhere”. It is to use the specialised protocol only on the specific flows that justify it, and keep REST/JSON at the public edge.
When it is needed (and when it is NOT)
It fits when:
- Internal, high-volume, latency-sensitive calls with stable contracts (gRPC), such as Claims to Fraud Scoring or Policy lookup.
- Streaming data: server streaming (progress of a large claim-document batch), client streaming (uploading photo chunks), or bidirectional streaming (gRPC).
- The user must see changes immediately without polling (SignalR/WebSocket): live claim status, adjuster assignment boards, collaborative review.
- Very many constrained devices or unreliable networks (MQTT): dashcams, sensors, drones at catastrophe sites.
- Polyglot teams that want generated, type-safe clients from a single contract file (
.proto).
It is overkill or wrong when:
- The API is public or partner-facing and must be easy to call from curl, Postman, and any language. Keep REST/JSON (or GraphQL).
- Traffic is low (say under a few hundred requests per second). The JSON cost is noise; you will pay more in tooling complexity than you save.
- You need HTTP caching, CDN caching, or simple browser debugging. gRPC responses are not cacheable by standard HTTP caches.
- The team has no capacity to operate it (HTTP/2 load balancing, broker operations, certificate management for devices).
- Browsers must call the service directly and you cannot host a gRPC-Web-capable endpoint. Native gRPC does not work from browsers.
- Updates are rare (once a minute or slower) where simple polling or Server-Sent Events is enough.
How to identify the problem (key signals)
- CPU profile shows JSON serialization on top. In dotnet-trace / Application Insights profiler,
System.Text.Jsonserializers or deserializers appear near the top for an internal hot path. - Polling noise. Access logs show the same
GET .../statusendpoint hit every few seconds by every open browser tab, mostly returning 304/unchanged payloads. - Bandwidth bills or egress spikes between services or from devices, with small payloads dominated by headers and repeated field names.
- Chained internal HTTP calls with rising p95/p99. Each hop adds connection and parsing overhead; request fan-out multiplies it.
- Device complaints: battery drain, dropped readings, duplicate readings after reconnects, “we don’t know if the device is offline or just quiet”.
- Contract drift incidents: “the field was a string, now it’s a number” bugs between teams, with no compile-time safety.
- Users say the screen is stale (“I had to hit F5 to see the payment status”).
- Thread/connection exhaustion on the server caused by long-polling clients holding requests open.
Flow Diagram
Pick the protocol from the traffic shape and keep REST at the edge.
flowchart TD Q{"What is the traffic shape?"} Q -- "Public / CRUD from browser" --> REST["REST / JSON via gateway"] Q -- "Internal high-volume RPC" --> GRPC["gRPC over HTTP/2"] Q -- "Live UI updates" --> SR["SignalR / WebSocket"] Q -- "Constrained devices" --> MQTT["MQTT broker"] GRPC --> F["Claims to Fraud Scoring"] SR --> AD["Adjuster browser groups"] MQTT --> DEV["Dashcams and sensors"] REST --> SPA["Angular CRUD API"]Level 1: Beginner
Analogy. Sending a parcel: sometimes a plain envelope (JSON over HTTP) is fine. But for a fragile part you use a purpose-built box (gRPC); for a live phone call you need an open line (WebSocket); for thousands of small sensors you use a postal system built for tiny, frequent messages (MQTT). Same goal, different vehicle.
Minimal working example: gRPC unary call (Claims service asks Fraud Scoring for a score).
fraud.proto (shared contract file):
syntax = "proto3";
option csharp_namespace = "Claims.Fraud.Grpc";package fraud;
service FraudScoring { rpc Score (ScoreRequest) returns (ScoreReply);}
message ScoreRequest { string claim_id = 1; string policy_id = 2; int64 amount_minor = 3; // amount in minor units (cents) to avoid floating point}
message ScoreReply { double score = 1; // 0.0 - 1.0 string band = 2; // LOW | MEDIUM | HIGH}Server project (Grpc.AspNetCore package, .NET 10 LTS):
using Claims.Fraud.Grpc;using Grpc.Core;
var builder = WebApplication.CreateBuilder(args);builder.Services.AddGrpc();
var app = builder.Build();app.MapGrpcService<FraudScoringService>();app.Run();
public sealed class FraudScoringService : FraudScoring.FraudScoringBase{ public override Task<ScoreReply> Score(ScoreRequest request, ServerCallContext context) { // Toy rule for the lesson; real scoring would call a model. var score = request.AmountMinor > 5_000_000 ? 0.9 : 0.1; return Task.FromResult(new ScoreReply { Score = score, Band = score >= 0.8 ? "HIGH" : "LOW" }); }}The .proto file is added to the server project with <Protobuf Include="fraud.proto" GrpcServices="Server" /> in the .csproj, and to the client project with GrpcServices="Client". The C# base class and message types are generated at build time.
Client (Claims service, packages Grpc.Net.ClientFactory, Google.Protobuf, Grpc.Tools):
builder.Services.AddGrpcClient<FraudScoring.FraudScoringClient>(o => o.Address = new Uri("https://fraud-scoring"));
// usagepublic sealed class ClaimSubmitter(FraudScoring.FraudScoringClient fraud){ public async Task<string> CheckAsync(string claimId, string policyId, long amountMinor, CancellationToken ct) { var reply = await fraud.ScoreAsync( new ScoreRequest { ClaimId = claimId, PolicyId = policyId, AmountMinor = amountMinor }, deadline: DateTime.UtcNow.AddSeconds(2), cancellationToken: ct); return reply.Band; }}Beginner rules to teach: always set a deadline, never reuse or renumber a Protobuf field number, and keep amounts as integers.
Level 2: Intermediate
In a real .NET + Angular + database application, you typically use three protocols side by side, each on the flow where it fits:
| Flow | Protocol | Why |
|---|---|---|
| Angular app to Claims API (CRUD) | REST/JSON | Easy, cacheable, debuggable |
| Claims to Fraud Scoring / Policy (internal) | gRPC | Fast, typed contract |
| Live claim status to adjuster browser | SignalR (WebSocket) | Server push |
| Dashcam/sensor telemetry | MQTT | Tiny, unreliable networks |
6.1 Live status push with SignalR
Server: a strongly-typed hub. The hub is only the connection endpoint; business code pushes through IHubContext.
using Microsoft.AspNetCore.Authorization;using Microsoft.AspNetCore.SignalR;
public sealed record ClaimStatusDto(string ClaimId, string Status, DateTimeOffset At);
public interface IClaimsClient{ Task ClaimStatusChanged(ClaimStatusDto dto);}
[Authorize]public sealed class ClaimsHub : Hub<IClaimsClient>{ public Task Subscribe(string claimId) => Groups.AddToGroupAsync(Context.ConnectionId, $"claim-{claimId}");
public Task Unsubscribe(string claimId) => Groups.RemoveFromGroupAsync(Context.ConnectionId, $"claim-{claimId}");}Program.cs:
builder.Services.AddSignalR();// ... authentication configured with JWT bearer (see security note in Level 3)app.MapHub<ClaimsHub>("/hubs/claims");Pushing from a domain event handler (see Day 11, Domain Event):
public sealed class ClaimApprovedHandler(IHubContext<ClaimsHub, IClaimsClient> hub){ public Task Handle(string claimId, CancellationToken ct) => hub.Clients.Group($"claim-{claimId}") .ClaimStatusChanged(new ClaimStatusDto(claimId, "Approved", DateTimeOffset.UtcNow));}Authorization gap to teach: [Authorize] on the hub only proves who is connected. Subscribe must also check that this adjuster is allowed to see this claim, otherwise any logged-in user can subscribe to any claim ID.
Angular client (standalone service using signals, package @microsoft/signalr):
import { Injectable, signal } from '@angular/core';import * as signalR from '@microsoft/signalr';
export interface ClaimStatus { claimId: string; status: string; at: string; }
@Injectable({ providedIn: 'root' })export class ClaimLiveService { private connection = new signalR.HubConnectionBuilder() .withUrl('/hubs/claims', { accessTokenFactory: () => localStorage.getItem('access_token') ?? '' }) .withAutomaticReconnect() .build();
readonly latest = signal<ClaimStatus | null>(null);
async start(claimId: string): Promise<void> { this.connection.on('ClaimStatusChanged', (s: ClaimStatus) => this.latest.set(s)); // Re-subscribe after every automatic reconnect: group membership is lost with the old connection. this.connection.onreconnected(() => this.connection.invoke('Subscribe', claimId)); await this.connection.start(); await this.connection.invoke('Subscribe', claimId); }
stop(): Promise<void> { return this.connection.stop(); }}(Storing tokens in localStorage is shown for brevity; many teams prefer in-memory tokens or cookies, see Level 3.)
6.2 gRPC from the browser
Browsers cannot speak native gRPC (they lack control over HTTP/2 framing and trailers). Options: expose a gRPC-Web endpoint or, more simply, keep a REST/JSON façade for the browser and gRPC only between services. To enable gRPC-Web in ASP.NET Core:
app.UseGrpcWeb(new GrpcWebOptions { DefaultEnabled = true });app.MapGrpcService<FraudScoringService>();gRPC-Web supports unary and server-streaming calls only, not client or bidirectional streaming.
6.3 Device telemetry with MQTT
Topic design is the “schema” of MQTT. For the claims system: claims/{claimId}/devices/{deviceId}/telemetry. The example below uses the MQTTnet 4.x API (MqttFactory); MQTTnet 5 renamed some types (for example MqttClientFactory), so check the version you install.
using MQTTnet;using MQTTnet.Client;using MQTTnet.Protocol;
var factory = new MqttFactory();using var client = factory.CreateMqttClient();
var options = new MqttClientOptionsBuilder() .WithTcpServer("broker.contoso.example", 8883) .WithTlsOptions(o => o.UseTls()) .WithClientId("dashcam-7781") .WithCleanSession(false) // keep the session so queued QoS1 messages survive short disconnects .Build();
await client.ConnectAsync(options);
await client.PublishStringAsync( topic: "claims/C-1001/devices/dashcam-7781/telemetry", payload: "{\"speedKph\":42,\"ts\":\"2026-09-29T10:15:00Z\"}", qualityOfServiceLevel: MqttQualityOfServiceLevel.AtLeastOnce);6.4 Database interplay
Protocols do not change your data rules. Use the Transactional Outbox (Day 12) so that “claim approved” is committed together with the message that triggers the SignalR push or MQTT command. Store device readings in a time-series-friendly table (PostgreSQL partitioned by day, or SQL Server with a clustered index on (DeviceId, ReadingTime)), and make the consumer idempotent (Day 17), because MQTT QoS 1 and SignalR reconnects can both deliver duplicates.
Level 3: Advanced
Performance
- Reuse channels. A
GrpcChannelowns HTTP/2 connections; create it once (the client factory does this) rather than per call. - Protobuf design: use scalar types, avoid deeply nested optional messages on hot paths, and prefer
repeatedfields with packed encoding for numeric lists. - Streaming vs many unary calls: for batches (scoring 500 claims), one server-stream or client-stream call avoids per-call overhead.
- SignalR payload: consider the MessagePack protocol for high-volume hubs. Keep messages small; send IDs and let the client refetch details if payloads are large.
- Message size defaults: ASP.NET Core gRPC limits incoming messages to 4 MB by default. Raise it deliberately, or better, stream large payloads such as photos in chunks.
Scalability
- HTTP/2 and L4 load balancers. A layer-4 balancer pins all calls of a long-lived HTTP/2 connection to one backend, so scaling out does not spread load. Use an L7 balancer that understands HTTP/2 (Envoy, Azure Application Gateway v2, Kubernetes ingress with gRPC support) or client-side load balancing.
- SignalR scale-out. With more than one server instance, a message sent from one instance must reach connections on others. Use a backplane, either Redis or Azure SignalR Service, which also offloads connection handling.
- WebSocket connections are stateful and long-lived. Plan for connection counts (memory per connection), graceful draining on deploy, and reconnect storms after a restart (use randomized reconnect delays).
- MQTT broker capacity: plan connections, publish rate and subscription fan-out. Shared subscriptions allow multiple consumers to share a topic’s load (MQTT 5).
Security
- Transport: TLS everywhere. gRPC over HTTP/2 in production requires TLS (plain HTTP/2 without TLS is possible only in controlled internal setups).
- Browser WebSocket auth: browsers cannot set an
Authorizationheader on a WebSocket handshake, so SignalR sends the JWT in theaccess_tokenquery string. That token can end up in proxy and access logs: use short-lived tokens, scrub query strings from logs, or use cookie auth. - Authorize per resource, not only per connection (see the
Subscribeexample above). Also validate theOriginheader/CORS for browser clients. - Devices: use per-device identity (X.509 client certificates or per-device tokens), never one shared password. Enforce topic-level ACLs so
dashcam-7781can only publish to its own topic and never subscribe to others. - gRPC: validate every field server-side; a typed contract does not mean trusted input.
Failure modes
- No deadline on a gRPC call means a stuck downstream holds the caller’s resources forever. Always set deadlines and propagate the caller’s cancellation token.
- Retries on non-idempotent calls double-charge or double-approve. Configure gRPC retry policy only for idempotent methods, and pair with idempotency keys.
- Lost group membership after a SignalR reconnect (handled in the Angular example above).
- MQTT QoS: 0 = at most once (may lose), 1 = at least once (may duplicate), 2 = exactly once at the protocol level but slowest and rarely worth it end to end. Pick 1 plus idempotent consumers in most cases.
- Retained messages and Last Will: retained messages give new subscribers the last known state; Last Will publishes “offline” when a device disconnects ungracefully.
- Slow consumers on WebSockets: an unread client buffer can grow on the server. Set limits and drop or coalesce stale updates (only the latest claim status matters).
Common mistakes
- Using gRPC for a public API because “it is faster”, then finding partners cannot call it easily.
- Renumbering or reusing Protobuf field numbers (breaks older clients silently). Reserve removed numbers with
reserved. - Sharing one giant
.protoand generated code as a versioned NuGet without a compatibility policy. - Putting business logic inside the SignalR hub instead of using
IHubContextfrom application services. - Assuming WebSocket messages are ordered and delivered across reconnects (they are not, after a reconnect).
- Skipping observability: binary protocols cannot be inspected in the browser network tab the way JSON can. Add logging interceptors, OpenTelemetry instrumentation for gRPC, and broker metrics from day one.
Level 4: Expert and Architect view
Trade-off comparison
| Option | Strengths | Weaknesses | Best fit in claims system |
|---|---|---|---|
| REST / JSON (HTTP/1.1 or 2) | Universal, cacheable, easy debugging | Verbose, no built-in schema or streaming | Public/partner API, Angular CRUD |
| gRPC (Protobuf over HTTP/2) | Compact, fast, strict contracts, streaming, code generation | Not browser-native, harder debugging, needs HTTP/2-aware infrastructure | Claims to Fraud/Policy/Payments internal calls |
| SignalR / WebSocket | Server push, bidirectional, .NET and Angular clients, automatic fallback transports | Stateful connections, scale-out needs backplane | Live claim status, adjuster boards |
| Server-Sent Events | Simple one-way push over HTTP | Server to client only, browser connection limits on HTTP/1.1 | Simple notifications where SignalR is too much |
| MQTT | Tiny overhead, QoS levels, retained messages, Last Will, huge device counts | Needs a broker, topic and ACL design, not a request/response model | Dashcams, drones, sensors |
| Async messaging (Service Bus / Kafka) | Durable, decoupled, replayable | Not for interactive request/response | Domain events between services (Day 16) |
Patterns it combines with
- API Gateway / BFF (Days 19-20): the edge stays REST/JSON (or gRPC-Web) while internal calls use gRPC.
- Transactional Outbox (Day 12) + Idempotent Consumer (Day 17): reliable delivery to whichever protocol pushes the message.
- Circuit Breaker, Retry, Bulkhead (Days 26-28): wrap gRPC clients (Polly /
Microsoft.Extensions.Http.Resilienceon the underlying HttpClient, or gRPC retry policy). - Distributed Tracing (Day 31): OpenTelemetry context propagates over gRPC metadata and can be added to MQTT 5 user properties.
- Service Mesh (Day 48): mesh sidecars understand HTTP/2 and gRPC for load balancing, mTLS and metrics.
ADR (architecture review style)
ADR-018: Use gRPC internally, SignalR for live UI, MQTT for field devices; keep REST at the edge
- Status: Proposed
- Context: Month-end load sends about 4,000 scoring calls per second between Claims and Fraud Scoring. Adjusters poll claim status. Twenty thousand field devices report telemetry over HTTPS POST with duplicate and battery problems.
- Decision: Adopt gRPC for Claims to Fraud Scoring and Claims to Policy calls. Adopt SignalR (hosted on Azure SignalR Service) for adjuster live status. Adopt MQTT (Azure Event Grid MQTT broker, or IoT Hub if device management is required) for device telemetry. Keep REST/JSON for the Angular CRUD API and all partner APIs.
- Consequences (positive): lower CPU and bandwidth on hot internal paths, strict contracts checked at build time, no polling, better device reliability.
- Consequences (negative): three more protocols to operate, HTTP/2-aware load balancing required, extra observability work,
.protocontract governance needed. - Alternatives rejected: REST everywhere (does not solve push or device constraints); Kafka directly to devices (too heavy for constrained clients).
- Review trigger: revisit if measured JSON cost on the scoring path is under 5% of CPU after profiling, since then gRPC may not be justified.
Azure implementation
Service names, tiers and limits change, so treat the numbers below as a guide and confirm on the Azure pricing pages and the Azure Pricing Calculator before budgeting.
Services
- gRPC hosting. Azure Container Apps (set ingress transport to
http2for gRPC), AKS with an ingress controller that supports gRPC, or Azure App Service on Linux with HTTP/2 enabled. Azure Application Gateway v2 supports gRPC traffic when end-to-end HTTP/2 with TLS is configured. Verify current support before putting a specific gateway in front of gRPC. - SignalR. Azure SignalR Service (tiers: Free, Standard, Premium) handles connection scale-out for ASP.NET Core SignalR. It has Default mode (your app server hosts the hub, the service proxies connections) and Serverless mode (used with Azure Functions). Each Standard unit supports about 1,000 concurrent connections.
- Raw WebSocket / pub-sub. Azure Web PubSub for plain WebSocket clients and pub/sub with your own protocol (not SignalR). Per Microsoft’s billing documentation, each unit supports up to 1,000 concurrent connections, instances can be sized at 1, 2, 5, 10, 20, 50 or 100 units, and the Standard tier includes 1,000,000 messages per unit per day (counted in 2 KB increments of outbound traffic). Microsoft recommends staying at or below about 80% unit utilization before scaling up. Premium adds features such as replicas across regions.
- MQTT. Azure Event Grid namespaces (Standard tier) provide a managed MQTT broker supporting MQTT v3.1.1 and v5.0, with topic spaces, client certificates and routing of MQTT messages to Event Grid subscriptions. The Basic tier has no MQTT. Alternatively Azure IoT Hub supports MQTT (plus AMQP and HTTPS) with per-device identity, device twins, cloud-to-device messages and device management. Choose IoT Hub when you need fleet management; choose Event Grid MQTT for general pub/sub with flexible routing.
- Security and secrets. Microsoft Entra ID for user tokens, Azure Key Vault for certificates and secrets, Private Endpoints for SignalR/Web PubSub/Event Grid where required.
- Monitoring. Application Insights and OpenTelemetry (Azure Monitor OpenTelemetry distro) for gRPC client/server spans; Azure Monitor metrics for connection counts and message counts on SignalR, Web PubSub and Event Grid; Log Analytics alerts on connection drops and throttling.
Configuration highlights
- gRPC on Container Apps: in the container app ingress settings choose transport
HTTP/2; keep TLS on; set min replicas at 2 for the fraud service so a deploy does not drop capacity. - Azure SignalR Service: create the resource, add its connection string to Key Vault, then in the app call
AddSignalR().AddAzureSignalR()(packageMicrosoft.Azure.SignalR). Use Managed Identity instead of access keys where supported. - Event Grid MQTT: enable the MQTT broker on the namespace, create a client per device with certificate-based authentication, group clients with client groups, and define topic spaces plus permission bindings so each device can publish only to
claims/+/devices/{clientId}/telemetry(use the${client.authenticationName}variable in the topic template). - Alerts: connection count at or above 80% of capacity, message throttling events, and MQTT disconnect rate.
Pricing and tier considerations
| Service | How it is billed (summary) | Tier notes |
|---|---|---|
| Azure SignalR Service | Per unit per day (each unit about 1,000 connections) plus messages beyond the included daily quota | Free tier for dev/test with small limits; Standard for production; Premium adds higher scale and resilience options |
| Azure Web PubSub | Per unit per day plus outbound messages beyond the included quota (1M messages per Standard unit per day, per Microsoft docs) | Use Standard for production; Premium for replicas and advanced features |
| Event Grid namespace (MQTT) | Throughput-unit based charges plus per-operation charges for MQTT publishes/deliveries | MQTT requires the Standard tier; check the current pricing page for the exact rates |
| Azure IoT Hub | Per unit per day by tier and message allowance | Free (dev/test), Basic (no cloud-to-device, no twins) and Standard tiers |
| Container Apps / AKS | Compute (vCPU/memory seconds or node VMs) | gRPC itself has no extra charge; you pay for the compute and networking |
Cost tips: gRPC reduces internal egress and CPU; SignalR messages are billed per 2 KB so send small notifications and let the client refetch details; batch MQTT telemetry (for example one message per 5 seconds carrying several readings) to cut operation counts.
Reference architecture (text)
- Adjusters use the Angular app, served from Azure Static Web Apps or Blob Storage behind Azure Front Door. CRUD calls go over HTTPS/REST to Azure API Management and then to the Claims API on Container Apps.
- Claims API calls Fraud Scoring and Policy services over gRPC (HTTP/2, TLS) inside the Container Apps environment/VNet. Each call carries a deadline and OpenTelemetry trace context.
- When a claim changes state, the Claims API writes the state change and an outbox row in one Azure SQL / PostgreSQL transaction. An outbox publisher sends the event to Azure Service Bus; a notifier consumer (idempotent) pushes the status through Azure SignalR Service to the adjuster’s browser group
claim-{id}. - Field devices (dashcams, sensors) connect to the Event Grid MQTT broker using client certificates. Event Grid routes telemetry to an Azure Function or Event Hubs for ingestion into the telemetry store (PostgreSQL partitioned table or Azure Data Explorer).
- Observability: Application Insights and Log Analytics collect traces, metrics and logs; dashboards show gRPC latency, SignalR connections and MQTT disconnect rates. Secrets and certificates live in Key Vault; access uses Managed Identity.
Teaching guide for my team
Explain to a beginner in 2 minutes
“JSON over HTTP is like writing a letter in full sentences every time. It is easy to read, but slow and bulky when you send thousands per second. gRPC is like a pre-agreed short form: both sides have the same form template (the .proto file), so messages are tiny and fast. SignalR is a phone line that stays open so the server can tell the browser when something changes, instead of the browser asking every 5 seconds. MQTT is a tiny-message post office for gadgets with weak connections. We do not replace everything; we use each one where it clearly wins, and keep normal REST at the edge.”
Explain to an intermediate developer in 5 minutes
- Start from the traffic shape: internal RPC, live UI push, or device telemetry.
- gRPC: define the contract in
.proto, generate client and server, always set deadlines, never reuse field numbers, use an HTTP/2-aware load balancer, and use gRPC-Web or a REST façade for browsers. - SignalR: hub for connections only,
IHubContextfor pushing, groups for targeting, authorize per resource, re-subscribe on reconnect, and use Azure SignalR Service to scale out. - MQTT: topic design, QoS 1 plus idempotent consumers, Last Will for presence, per-device identity and topic ACLs.
- Combine with outbox and idempotent consumer so pushes are reliable, and add tracing since binary protocols are harder to inspect.
- Decide with evidence: profile first, and write an ADR.
Hands-on exercise
Task: Build a small “Claims Fraud Score + live status” slice.
- Create a
Fraud.GrpcASP.NET Core project with thefraud.protofrom Level 1 and implementScore. - Create a
Claims.Apiproject that calls it throughAddGrpcClientwith a 2-second deadline. - Add a
ClaimsHub(SignalR) and an endpointPOST /claims/{id}/approvethat pushesClaimStatusChangedto groupclaim-{id}usingIHubContext. - In an Angular standalone component, use the
ClaimLiveServicefrom Level 2, subscribe to claimC-1001, and display the latest status using a signal. - Use a load script (for example
ghzfor gRPC or a simple loop) to compare 1,000 unary calls against an equivalent REST/JSON endpoint and record payload size and p95 latency.
Expected outcome: approving a claim in one browser tab updates a second open tab within about a second without polling or refresh; the gRPC payload is several times smaller than the JSON equivalent; the team can explain why the Angular app still talks REST for CRUD and why a hub subscription needs a resource-level authorization check.
Interview-style questions
- Why can’t a browser call a normal gRPC service directly, and what are the options? Browsers do not expose the HTTP/2 framing and trailer control gRPC requires. Use gRPC-Web (unary and server streaming only) through a proxy or ASP.NET Core’s gRPC-Web middleware, or keep a REST/JSON façade for browsers.
- A SignalR app works with one server but loses messages with three. Why? Connections are spread across instances and a message published on one instance does not reach connections held by others. Add a backplane (Redis or Azure SignalR Service) and, if needed, make sure clients re-join groups after reconnecting.
- When would you choose MQTT over HTTP for device telemetry, and what QoS would you pick? For many constrained devices on unreliable networks needing low overhead, presence (Last Will) and queued delivery. QoS 1 with idempotent consumers is the usual choice; QoS 0 for disposable readings; QoS 2 is rarely worth its cost.
Mastery checklist
- I can explain, with measurements, why a given flow should or should not leave HTTP/JSON.
- I can write a
.protocontract, generate client and server, and follow safe schema evolution rules (no reused or renumbered fields, usereserved). - I always set gRPC deadlines, propagate cancellation, and only retry idempotent calls.
- I can build a SignalR hub with groups, resource-level authorization, reconnect handling and an Angular client.
- I understand HTTP/2 load balancing pitfalls and SignalR scale-out (backplane or Azure SignalR Service).
- I can design MQTT topics, choose QoS, and apply per-device identity and topic ACLs.
- I can choose between Azure SignalR Service, Web PubSub, Event Grid MQTT and IoT Hub and justify it.
- I can write an ADR that names the trade-offs and a review trigger.
Key takeaway
Pick the protocol from the traffic shape (RPC, live push, or device telemetry), use it only where measurements justify it, and keep REST/JSON at the public edge.
