Serverless deployment means you ship only your code and its triggers (an HTTP call, a queue message, a timer) and let the cloud platform decide where it runs, how many copies run, and when they are switched off.
Intro
Serverless deployment means you ship only your code and its triggers (an HTTP call, a queue message, a timer) and let the cloud platform decide where it runs, how many copies run, and when they are switched off. You are billed for what executes, not for servers that sit idle. In Azure the flagship service for this is Azure Functions, and for new .NET work the recommended hosting plan is Flex Consumption. In our insurance claims system, this fits spiky, event-driven work such as “a claim photo was uploaded, extract its metadata” or “a nightly reserve recalculation timer fired”.
Why we need this
- Idle cost. A claims system has very uneven traffic: a storm or a holiday weekend produces thousands of First Notice of Loss (FNOL) submissions, while 3 a.m. on a Tuesday produces almost none. Always-on VMs or containers are sized for the peak and paid for around the clock.
- Capacity planning is guesswork. Picking VM sizes, autoscale rules and node counts for dozens of small background jobs is real engineering effort that produces no business value.
- Operational load. Patching the OS, rotating instances and tuning scale rules are chores a small team would rather not own for small, simple workloads.
- Fit for event-driven work. Many claims tasks are naturally “when X happens, do Y”: a document lands in Blob Storage, a message arrives on Service Bus, a timer fires. A function bound to that trigger is the smallest unit of deployment that matches the requirement.
- Team speed. One function, one trigger, one deploy. Teams can release a small piece of behaviour without touching a larger application.
What problem it solves
Problem (from the topic list): idle services still cost money, and scaling needs heavy planning.
Without serverless, a low-traffic service such as ClaimDocumentThumbnailer gets its own App Service plan or container that runs 24/7 at, say, 2% utilization. Multiply that by 30 small services and you pay for 30 mostly idle hosts. When a burst arrives (10,000 photos after a hailstorm), the fixed capacity queues up work and claims processing slows, so someone has to write, test and tune autoscale rules. Serverless moves both problems to the platform: capacity follows events, and when there are no events the compute cost is zero (on-demand mode).
When it is needed (and when it is NOT)
Good fit
- Event-driven glue: Blob upload, Service Bus message, Event Grid event, timer, webhook.
- Spiky or unpredictable traffic, and long idle periods (nightly jobs, seasonal load).
- Short-running, stateless units of work that finish in seconds to a few minutes.
- Small teams that want minimal infrastructure ownership.
- Fan-out processing where many independent items are handled in parallel (for example, each claim photo).
Poor fit / overkill
- Constant, high, predictable load. A reserved App Service or AKS node is usually cheaper per request at steady 24/7 utilization.
- Latency-critical APIs that cannot tolerate any cold start (unless you pay for always-ready instances, which removes much of the cost advantage).
- Long-running or stateful work with tight in-memory state (use Durable Functions, Container Apps jobs, or AKS instead).
- Workloads that need custom OS packages, GPUs, or unusual runtimes (containers are a better match).
- A large, tightly coupled application: splitting it into hundreds of functions can produce a distributed-monolith mess that is harder to run than the original.
- Teams with no monitoring in place. Many small functions without tracing are hard to debug.
How to identify the problem (key signals)
- Low utilization on always-on hosts. CPU below ~10-15% for most of the day on a service that only reacts to messages or files.
- Bill vs. traffic mismatch. Monthly compute cost is flat regardless of traffic; a service costs the same in August as in a storm month.
- Queue backlogs during bursts. Service Bus queue depth or oldest-message age climbs after events, and someone manually scales out.
- Hand-tuned autoscale rules. A pile of scale-out/scale-in rules on CPU or queue length that nobody wants to touch.
- Cron jobs on VMs. A VM (or a container kept alive) whose only job is to run a script once a night.
- Tiny services with heavy hosting. A 40-line service that needs a Dockerfile, a Helm chart and a node pool.
- Ops tickets for chores. Repeated tickets to restart, patch or resize hosts running small background workers.
Flow Diagram
Events trigger functions that scale out on demand and back to zero.
flowchart LR UI["Angular"] --> API["HTTP function - Claims API"] API --> SQL[("Azure SQL + outbox")] API -- "202 Accepted" --> UI TMR["Timer function - polling publisher"] --> SQL TMR --> Q{{"Service Bus queue"}} Q --> FN["FraudPrecheck function - scales 0..N"] FN --> SQL BLOB["Blob upload"] -- "Event Grid" --> TH["Thumbnail function"]Level 1: Beginner
Analogy. A taxi versus owning a car. Owning a car (a VM) means paying for insurance, parking and maintenance whether you drive or not. A taxi (serverless) costs money only for the trips you take, and if fifty people need a ride at once, more taxis show up.
Core vocabulary
- Function: one piece of code that runs in response to a trigger.
- Trigger: what starts the function (HTTP request, queue message, timer, blob).
- Binding: a declarative way to read from or write to a service (for example, output to a queue) without writing client code.
- Cold start: the delay when the platform must start a new instance because none is running.
- Scale to zero: when nothing is happening, no instance runs and (in on-demand mode) nothing is billed for compute.
Minimal example: an HTTP-triggered function (C#, .NET 10, isolated worker model)
using Microsoft.Azure.Functions.Worker;using Microsoft.AspNetCore.Http;using Microsoft.AspNetCore.Mvc;
namespace Claims.Functions;
public class HelloFunction{ [Function("Ping")] public IActionResult Run( [HttpTrigger(AuthorizationLevel.Function, "get", Route = "ping")] HttpRequest req) => new OkObjectResult(new { status = "claims-functions alive", utc = DateTime.UtcNow });}Program.cs for the same project:
using Microsoft.Azure.Functions.Worker.Builder;using Microsoft.Extensions.Hosting;
var builder = FunctionsApplication.CreateBuilder(args);builder.ConfigureFunctionsWebApplication(); // ASP.NET Core integration for HTTP triggersbuilder.Build().Run();Required NuGet packages: Microsoft.Azure.Functions.Worker, Microsoft.Azure.Functions.Worker.Sdk, and Microsoft.Azure.Functions.Worker.Extensions.Http.AspNetCore. The project targets net10.0 with <AzureFunctionsVersion>v4</AzureFunctionsVersion>.
Level 2: Intermediate
Scenario. A policyholder submits a claim through our Angular app. The API stores the claim and puts a ClaimSubmitted message on a Service Bus queue. A function picks up each message, runs fraud pre-checks, and writes the result to SQL Server. Because messages arrive in bursts, the function scales out automatically.
using Azure.Messaging.ServiceBus;using Microsoft.Azure.Functions.Worker;using Microsoft.Extensions.Logging;using System.Text.Json;
namespace Claims.Functions;
public record ClaimSubmitted(Guid ClaimId, string PolicyNumber, decimal EstimatedAmount);
public class FraudPrecheckFunction(ILogger<FraudPrecheckFunction> logger, IFraudPrecheckService service){ [Function("FraudPrecheck")] public async Task Run( [ServiceBusTrigger("claim-submitted", Connection = "ServiceBus")] ServiceBusReceivedMessage message, ServiceBusMessageActions actions, CancellationToken ct) { var evt = JsonSerializer.Deserialize<ClaimSubmitted>(message.Body) ?? throw new InvalidOperationException("Empty message body");
logger.LogInformation("Fraud precheck for claim {ClaimId}", evt.ClaimId);
// Must be idempotent: Service Bus is at-least-once (see Day 17). await service.RunAsync(evt, message.MessageId, ct);
await actions.CompleteMessageAsync(message, ct); }}
public interface IFraudPrecheckService{ Task RunAsync(ClaimSubmitted evt, string messageId, CancellationToken ct);}Registering dependencies (EF Core with SQL Server, and Application Insights) in Program.cs:
using Microsoft.Azure.Functions.Worker.Builder;using Microsoft.Azure.Functions.Worker;using Microsoft.EntityFrameworkCore;using Microsoft.Extensions.DependencyInjection;using Microsoft.Extensions.Hosting;
var builder = FunctionsApplication.CreateBuilder(args);builder.ConfigureFunctionsWebApplication();
builder.Services .AddApplicationInsightsTelemetryWorkerService() .ConfigureFunctionsApplicationInsights();
builder.Services.AddDbContext<ClaimsDbContext>(o => o.UseSqlServer(builder.Configuration["SqlConnection"], sql => sql.EnableRetryOnFailure())); // transient faults are normal in the cloud
builder.Services.AddScoped<IFraudPrecheckService, FraudPrecheckService>();
builder.Build().Run();Angular side. Angular does not call the queue function directly. The claim API (an App Service or another HTTP function) accepts the request and returns 202 Accepted; Angular then polls or subscribes for status:
@Injectable({ providedIn: 'root' })export class ClaimsService { private http = inject(HttpClient);
submit(claim: NewClaim) { return this.http.post<{ claimId: string }>('/api/claims', claim); // returns 202 + id }
status(claimId: string) { return this.http.get<{ state: 'Pending' | 'Cleared' | 'Review' }>(`/api/claims/${claimId}/status`); }}Database notes. Functions can scale out to many instances, and each opens its own SQL connections. Keep the connection pool small, use EnableRetryOnFailure, and be aware that a thousand concurrent instances can exhaust SQL connection limits (see Level 3).
Level 3: Advanced
Performance and cold starts
- Cold start grows with package size, dependency count and startup work. Trim dependencies, avoid heavy work in
Program.cs, and avoid loading big configuration at start. - On Flex Consumption you can configure always-ready instances per trigger group to keep a number of instances warm. Always-ready instances are billed continuously and have no free grant, so use them only for latency-sensitive paths.
- Choose instance memory to match the workload. Flex Consumption offers 512 MB, 2,048 MB (default) and 4,096 MB instance sizes. Bigger instances cost more per second but allow more concurrency per instance.
- Tune per-instance concurrency for HTTP so fewer instances are needed when work is I/O-bound.
Scalability
- Flex Consumption scales functions in groups: all HTTP triggers together, all Blob triggers together, all Durable triggers together, and every other trigger on its own. One noisy queue does not force the HTTP endpoints to scale.
- Maximum instance count is configurable up to 1,000. There is also a regional per-subscription memory quota (default 250 cores per region); always-ready instances count against it.
- Downstream limits are the real ceiling. Scale-out will happily overwhelm SQL Server or a third-party API. Cap with
maximumInstanceCount, usemaxConcurrentCallson the Service Bus trigger, and consider a queue as a buffer. Pair with Bulkhead and Rate Limiter (Days 27 and 30).
Reliability and failure modes
- At-least-once delivery. Triggers like Service Bus and Event Grid may deliver duplicates; every function must be idempotent (Day 17).
- Poison messages. After the max delivery count the message goes to the dead-letter queue. Monitor it and have a replay process.
- Timeouts. Default execution timeout on Flex Consumption is 30 minutes, but HTTP-triggered requests are cut off at 230 seconds by the load balancer. Return
202and process asynchronously for long work. - Startup timeout. The host must start within 30 seconds on Flex Consumption.
- Statelessness. Local disk and memory do not survive between invocations. Keep state in SQL, Blob, or Cosmos DB, or use Durable Functions for workflows.
Security
- Prefer managed identity over connection strings. Use identity-based connections for Service Bus, Storage and Key Vault, and Microsoft Entra authentication for Azure SQL.
- Do not rely on function keys as the only protection for public HTTP endpoints; put API Management or Front Door with Entra ID validation in front.
- Use VNet integration and private endpoints when the function must reach private SQL or Service Bus. Flex Consumption supports both.
Common mistakes
- Creating a new
HttpClientper invocation instead of usingIHttpClientFactory(socket exhaustion). - Running a long orchestration inside one function instead of using Durable Functions.
- Using the in-process C# model. It is not supported on Flex Consumption and is being retired; use the isolated worker model.
- Ignoring downstream limits and letting scale-out take down the database.
- Skipping Application Insights, then being unable to trace a message across functions.
- Assuming the legacy Consumption plan is the default: Linux Consumption is scheduled for retirement on September 30, 2028, so new apps should target Flex Consumption.
Level 4: Expert and Architect view
Hosting options compared for a claims background workload
| Option | Scale to zero | Cold start | Ops effort | Cost profile | Best for |
|---|---|---|---|---|---|
| Azure Functions, Flex Consumption | Yes (on-demand mode) | Low, removable with always-ready | Very low | Pay per execution and GB-second; always-ready billed continuously | Spiky event-driven work, new serverless .NET apps |
| Azure Functions, Premium | No (min 1 warm instance) | Effectively none | Low | Continuous vCPU/memory billing | Steady load with low-latency needs, long runtimes |
| Azure Functions, Consumption (legacy) | Yes | Noticeable | Very low | Cheapest for very small usage | Existing apps only; Linux retires Sept 30, 2028 |
| Azure Container Apps (or ACA jobs) | Yes (KEDA) | Depends on replicas | Low-medium | vCPU/memory per second | Container-based services, custom runtimes |
| App Service (Dedicated) | No | None | Low | Fixed monthly per plan | Predictable, steady web workloads |
| AKS | Node-level only | None once running | High | Nodes 24/7 | Large platform teams, complex mesh needs |
Patterns it combines with
- Messaging and Idempotent Consumer (Days 16, 17): functions are natural consumers.
- Transactional Outbox / Polling Publisher (Days 12, 14): a timer-triggered function can be the polling publisher.
- Saga (Day 7): Durable Functions can host orchestration-style sagas.
- Circuit Breaker, Retry, Bulkhead (Days 26-28): protect downstream systems from scale-out.
- Distributed Tracing and Health Checks (Days 31, 36): mandatory when there are many small units.
- Strangler Fig (Day 49): peel functionality off a monolith into functions behind a router.
ADR (architecture review style)
ADR-045: Host claim background processors on Azure Functions Flex Consumption Status: Proposed Context: Claim intake produces bursty background work (fraud pre-checks, document processing, notifications). Load can rise 50x during weather events and is near zero overnight. The team is small and does not want to manage nodes or scale rules. The stack is .NET (current LTS) with Azure SQL and Service Bus. Decision: Host stateless, queue- and blob-triggered processors on Azure Functions (isolated worker model) using the Flex Consumption plan, with managed identity, VNet integration to reach private SQL, and Application Insights. Use 2 always-ready instances only on the HTTP intake endpoint if measured cold-start latency breaches the SLO. Consequences: (+) Near-zero idle cost, automatic scale-out, less infrastructure to own. (-) Cold starts on rarely used functions, scale-out must be capped to protect SQL, no deployment slots (use Flex site update strategies), one app per plan, and regional availability must be checked. Long-running or stateful workflows will use Durable Functions or Container Apps instead. Alternatives rejected: AKS (operational overhead too high for this team), Premium plan (always-on cost not justified by traffic pattern), legacy Consumption (Linux plan retirement).
Azure implementation
Services
- Azure Functions (Flex Consumption plan) for compute.
- Azure Service Bus, Event Grid, Blob Storage as triggers and event sources.
- Azure Storage account for the function host and the deployment package container.
- Azure SQL Database for claim data, accessed with Microsoft Entra authentication.
- Azure Key Vault for secrets that cannot be avoided.
- Application Insights + Log Analytics for telemetry.
- API Management or Front Door in front of public HTTP functions.
- Durable Functions (Azure Storage or Durable Task Scheduler backend) for orchestrations.
Configuration (Azure CLI sketch)
# Storage account for the hostaz storage account create -n stclaimsfunc01 -g rg-claims -l westeurope --sku Standard_LRS
# Flex Consumption function app (Linux, .NET 10 isolated)az functionapp create \ -n func-claims-fraud -g rg-claims \ --flexconsumption-location westeurope \ --runtime dotnet-isolated --runtime-version 10.0 \ --storage-account stclaimsfunc01 \ --instance-memory 2048 \ --maximum-instance-count 40
# Managed identity so no connection strings are neededaz functionapp identity assign -n func-claims-fraud -g rg-claims
# Identity-based Service Bus connection (setting name "ServiceBus")az functionapp config appsettings set -n func-claims-fraud -g rg-claims --settings \ "ServiceBus__fullyQualifiedNamespace=sb-claims.servicebus.windows.net"
# Keep 1 instance warm for the HTTP trigger group (billed continuously)az functionapp scale config always-ready set -n func-claims-fraud -g rg-claims \ --settings http=1Then grant the function’s managed identity the Azure Service Bus Data Receiver role on the namespace, and create a SQL user for it with CREATE USER [func-claims-fraud] FROM EXTERNAL PROVIDER. Check the exact CLI parameter names against your installed az version, as Flex Consumption commands have evolved.
Deployment. Build and zip the app in CI (GitHub Actions or Azure DevOps) and deploy with az functionapp deployment source config-zip or the Azure/functions-action. In Flex Consumption the package is stored in a blob container and pulled by the instances at startup. Deployment slots are not available; use the platform’s site update strategies for rolling updates.
Pricing and tier considerations (confirm current rates on the Azure Functions pricing page before quoting numbers to stakeholders)
- Flex Consumption, On Demand mode: billed per execution count and GB-seconds of memory while executing, with a monthly free grant. Minimum billable execution period is 1,000 ms, then rounded to the nearest 100 ms.
- Flex Consumption, Always Ready: a baseline charge for each always-ready instance plus execution charges; no free grant in this mode.
- Premium: continuous vCPU and memory billing with at least one warm instance; choose it when you need no cold start and long-running work at steady volume.
- Legacy Consumption: Linux plan retires on September 30, 2028; plan migration to Flex Consumption. There is no in-place migration; you create a new Flex app and redeploy.
- Hidden costs: storage transactions, Application Insights ingestion (sample or cap it), Service Bus tier (Standard or Premium), VNet-related networking, and SQL compute.
Reference architecture (text)
Angular SPA on Azure Static Web Apps calls Azure Front Door, which routes /api/* to API Management. API Management forwards to the Claims API (an HTTP-triggered function app on Flex Consumption with one always-ready instance). The API writes the claim and an outbox row to Azure SQL (private endpoint) and returns 202. A timer-triggered function (polling publisher) publishes outbox rows to a Service Bus queue. A separate Flex Consumption app hosts queue-triggered functions (fraud pre-check, notification) that read the queue with managed identity, write results to SQL, and emit events to Event Grid. Uploaded claim photos in Blob Storage raise Event Grid events that trigger a thumbnail and metadata function. All apps send telemetry to Application Insights with a shared Log Analytics workspace, and alerts fire on dead-letter queue depth, function failure rate and p95 duration. Each function app has its own system-assigned managed identity with least-privilege roles.
Teaching guide for my team
2-minute beginner explanation
“Normally we rent a server, install our app and pay for it every hour, even at night. With serverless we give Azure a small piece of code and say: run this when a claim message arrives. Azure starts it when needed, starts more copies if many messages arrive, and stops everything when there is nothing to do. We pay for the time the code actually runs. The trade-offs: the first call after a quiet period can be a bit slower (cold start), the code must not keep data in memory between calls, and we have to be careful that many copies running together do not overwhelm our database.”
5-minute intermediate explanation
Walk through the FraudPrecheck function: trigger, binding to Service Bus, DI and EF Core in Program.cs, and managed identity. Explain that scale is decided per trigger group on Flex Consumption, that maximumInstanceCount and maxConcurrentCalls protect SQL, that delivery is at-least-once so the handler must be idempotent, and that HTTP has a 230-second ceiling so long work should return 202. Show the Application Insights view of one invocation. Finish with the cost model: On Demand versus Always Ready, and why steady 24/7 load can favour Premium or App Service.
Hands-on exercise
Task: Build a Service Bus-triggered function ClaimAcknowledgement that reads a ClaimSubmitted message and writes a row to a ClaimAcks table with MessageId as a unique key.
- Create the project with
func init --worker-runtime dotnet-isolated --target-framework net10.0and add the function. - Run locally with the Azurite emulator and a Service Bus namespace (or the Service Bus emulator).
- Send the same message twice.
- Deploy to a Flex Consumption app with a managed identity.
- Send 500 messages in a loop and watch the instance count in Application Insights Live Metrics.
Expected outcome: Only one row per MessageId exists (idempotency works), the second delivery is completed without error, instance count grows during the burst and drops to zero afterwards, and setting maximumInstanceCount to 5 visibly caps the parallelism.
Interview-style questions
- Why is idempotency mandatory for queue-triggered functions? Because Service Bus delivers at least once; a retry after a crash or lock timeout will run the function again for the same message, and without idempotency you would double-process the claim.
- What is a cold start and how do you reduce it on Azure Functions? It is the delay while the platform starts a new instance. Reduce it with smaller packages, lighter startup code, and always-ready instances (Flex Consumption) or a Premium plan with pre-warmed instances.
- When would you not choose serverless for a claims service? When load is constant and high (cheaper on reserved compute), when latency must be consistently very low without paying for warm instances, when work is long-running and stateful, or when you need a custom OS or runtime.
Mastery checklist
- I can explain triggers, bindings, cold start and scale to zero without notes.
- I can build and deploy a .NET isolated-worker function with Service Bus trigger and DI.
- I can explain why the in-process model is not an option and how to migrate to isolated.
- I can make a handler idempotent and explain the dead-letter path.
- I can cap scale-out to protect SQL Server and justify the numbers.
- I can compare Flex Consumption, Premium, Container Apps and App Service for a given workload, including cost shape.
- I can configure managed identity for Service Bus, Storage and Azure SQL with no secrets in config.
- I can write an ADR that justifies (or rejects) serverless for a specific service.
Key takeaway
Serverless deployment trades server management and idle cost for platform-driven scaling, so use it for spiky, stateless, event-driven work. Always cap scale-out to protect downstream systems, and design every handler to be idempotent.
