Serverless Architecture — Patterns and Anti-Patterns
A practical guide to serverless architecture: function design, cold starts, event-driven patterns, state management, and common pitfalls with AWS Lambda, Azure Functions, and GCP Cloud Functions.
Overview
Serverless architecture lets you run code without provisioning or managing servers. The cloud provider handles infrastructure, scaling, and patching; you provide functions that execute in response to events. While serverless eliminates server management, it introduces new constraints: execution time limits, cold starts, statelessness, and distributed debugging. The following guide covers patterns that work and anti-patterns that cause pain.
When to Use
-
For alternatives, see Complete Guide to Serverless Architecture.
-
Variable or unpredictable traffic (pay-per-execution saves money)
-
Event-driven workflows (file uploads, database changes, scheduled tasks)
-
Microservices with independent deployment cycles
-
Prototypes and MVPs where speed matters more than optimization
-
Processing pipelines that can be broken into discrete steps
Core Patterns
| Pattern | Use Case | Example |
|---|---|---|
| Function-per-HTTP-endpoint | REST APIs | API Gateway → Lambda |
| Event-driven function | Async processing | S3 upload → Lambda thumbnail generator |
| Scheduled function | Cron jobs | CloudWatch Events → nightly report Lambda |
| Queue-triggered function | Decoupled workloads | SQS → Lambda order processor |
| Stream-triggered function | Real-time data | DynamoDB Streams → Lambda cache updater |
Function Design — What Works
# AWS Lambda handler — keep initialization outside handler for reuse
import boto3
import json
# Initialized once per container lifecycle
s3_client = boto3.client('s3')
dynamodb = boto3.resource('dynamodb')
table = dynamodb.Table('ProcessedFiles')
def lambda_handler(event, context):
# Handler runs on every invocation
bucket = event['Records'][0]['s3']['bucket']['name']
key = event['Records'][0]['s3']['object']['key']
# Process the file
response = s3_client.get_object(Bucket=bucket, Key=key)
content = response['Body'].read()
# Store metadata
table.put_item(Item={
'fileId': key,
'bucket': bucket,
'size': len(content),
'processedAt': context.aws_request_id
})
return {'statusCode': 200, 'body': json.dumps({'processed': key})}
// Azure Function with HTTP trigger and dependency injection
const { app } = require('@azure/functions');
class OrderService {
async createOrder(orderData) {
// Business logic here
return { id: crypto.randomUUID(), ...orderData };
}
}
app.http('createOrder', {
methods: ['POST'],
authLevel: 'anonymous',
handler: async (request, context) => {
const orderData = await request.json();
const service = new OrderService();
const order = await service.createOrder(orderData);
return { jsonBody: order, status: 201 };
}
});
Managing Cold Starts
| Strategy | Impact | Implementation |
|---|---|---|
| Keep-alive (ping) | Eliminates cold start | CloudWatch cron every 5 minutes |
| Provisioned concurrency | Pre-warmed instances | Lambda provisioned concurrency |
| Minimize dependencies | Faster initialization | Remove unused packages, tree-shake |
| Use compiled languages | Faster startup | Go, Rust, or .NET AOT compilation |
| Connection pooling | Reuse DB connections outside handler | Initialize clients globally |
State Management
Serverless functions are stateless. Persist state externally:
# AWS Step Functions — orchestrate stateful workflows
Comment: Order Processing Workflow
StartAt: ValidateOrder
States:
ValidateOrder:
Type: Task
Resource: arn:aws:lambda:...:validate-order
Next: ProcessPayment
ProcessPayment:
Type: Task
Resource: arn:aws:lambda:...:process-payment
Catch:
- ErrorEquals: ["PaymentFailed"]
ResultPath: "$.error"
Next: NotifyFailure
Next: NotifySuccess
Common Mistakes
- Monolithic Lambda — putting an entire application in one function; break into single-purpose functions
- Synchronous waiting — calling slow services synchronously inside a function; use async patterns
- Ignoring timeout limits — Lambda has a 15-minute max; long jobs need ECS or Batch
- Treating functions like servers — storing state in memory or local disk
- No retry strategy — transient failures should be handled with dead-letter queues
- Over-provisioned memory — memory controls CPU; test to find the sweet spot
Troubleshooting
- High latency between services: trace the request path. Look for synchronous chains, missing caching, and oversized payloads that cross network boundaries.
- Single point of failure: identify components without redundancy. Add replicas, failover, or circuit breakers before scaling traffic.
- Unexpected coupling between services: review shared databases, libraries, and schemas. Bound contexts should own their data and expose stable interfaces.
- Cost spikes after scaling: right-size instances and use autoscaling with limits. Reserved capacity or spot instances can reduce steady-state spend.
- Difficult to reason about the system: maintain architecture decision records and service dependency maps.
Production Notes
- Deploy gradually using canary or blue-green to catch regressions early.
- Configure alerts for error rate, p99 latency, and failure rate before enabling in production.
- Document the rollback in the runbook; test the procedure in staging at least once per quarter.
- Review structured logs with correlation IDs to trace requests end-to-end during incidents.
Key Takeaways
- Apply serverless architecture — patterns and anti-patterns when you need a practical solution for your use case.
- Monitor performance after implementation; measure latency, errors, and resource usage before and after.
- Check the Troubleshooting section for common failures; most have documented root causes with fixes.
- Keep dependencies updated and run tests in CI to prevent production regressions.
Advanced Topics
Detailed Scenario: Image Processing Pipeline
Architecture: S3 upload -> Lambda -> SQS -> Lambda -> DynamoDB
Platform: AWS
Volume: 10k images/day
Step 1: Image upload
User uploads image to S3 bucket: uploads/originals/
S3:ObjectCreated event triggers Lambda resize-function
Step 2: Lambda resize-function
- Read image from S3
- Generate 3 sizes: thumbnail (100x100), medium (500x500), large (1200x1200)
- Upload versions to S3: uploads/resized/
- Publish message to SQS: image-processed-queue
Message: { "originalKey": "...", "thumbnailKey": "...", "mediumKey": "...", "largeKey": "..." }
- Timeout: 30s (large image resize can take 10-15s)
- Memory: 1024MB (needed for image processing)
Step 3: SQS -> Lambda metadata-function
- Read message from SQS
- Extract metadata: dimensions, format, size, hash
- Store in DynamoDB table: ImageMetadata
PK: imageId (generated uuid), SK: version (thumbnail|medium|large)
- Delete message from SQS (successful processing)
- If fails 3 times -> DLQ: image-processed-dlq
IaC configuration (Terraform):
resource "aws_lambda_function" "resize" {
function_name = "image-resize"
handler = "handler.lambda_handler"
runtime = "python3.11"
memory_size = 1024
timeout = 30
reserved_concurrent_executions = 50
environment {
variables = {
OUTPUT_BUCKET = "uploads-resized"
QUEUE_URL = aws_sqs_queue.image_processed.id
}
}
}
resource "aws_sqs_queue" "image_processed" {
name = "image-processed-queue"
visibility_timeout = 60 # > lambda timeout
redrive_policy = jsonencode({
deadLetterTargetArn = aws_sqs_queue.dlq.arn
maxReceiveCount = 3
})
}
Monitoring:
- CloudWatch Alarms: errors > 1%, duration p95 > 20s
- DLQ alert: message in DLQ -> SNS -> Slack
- X-Ray tracing to see latency per step
Estimated cost (10k images/day):
Lambda: ~$15/month (1.5M invocations)
S3: ~$2/month (storage)
SQS: ~$0.40/month
DynamoDB: ~$1.25/month (on-demand)
Total: ~$19/month
How do I handle distributed transactions in serverless?
Do not use distributed ACID transactions. Use the saga pattern: each function does its part and publishes an event. If a step fails, a compensating function undoes the previous work. On AWS, Step Functions orchestrates sagas with compensation states. Alternatively, use the outbox pattern: the function writes to the DB and publishes the event in the same transaction. A separate process reads the outbox and publishes events.
How do I test serverless functions locally?
Use SAM CLI for AWS Lambda: sam local invoke -e event.json. For Azure, Azure Functions Core Tools: func start. Create test events in JSON that mimic real events (S3, SQS, API Gateway). For integration testing, use LocalStack which emulates AWS services locally. For CI, run tests in GitHub Actions with SAM CLI installed. Keep handler tests separate from business logic tests.
End of document. Review and update quarterly.
Common Production Pitfalls
- Treating the guide as a checklist to complete once rather than a practice to evolve.
- Adopting every recommendation at once instead of starting with one measured change.
- Skipping the maturity assessment and forcing advanced practices on an unprepared team.
- Not updating runbooks and on-call expectations as new practices are introduced.
- Ignoring real incident data when prioritizing which parts of the guide to apply first.
- Failing to assign an owner who reviews decisions quarterly.
- Copying examples without adapting them to the team’s actual tooling and constraints.
- Forgetting to measure outcomes before adding the next improvement.
Frequently Asked Questions
How do I get started with this in an existing project?
Start with a small, isolated part of your codebase. Apply the concepts from this guide to one module or service. Measure the impact, then expand to other areas.
What tools do I need?
The tools mentioned throughout this guide are listed in each section. Most are open-source and widely adopted. Check the related resources for setup instructions.
How do I measure success after implementing this?
Define clear metrics before starting: performance benchmarks, error rates, or maintainability indicators. Compare before and after. Iterate based on the data, not on assumptions.
Related Resources
AWS Basics — Core Services for Developers
A practical guide to AWS core services for developers: compute, storage, databases, networking, and security fundamentals with hands-on examples.
GuideAzure Basics — Core Services for Developers
A practical guide to Microsoft Azure core services for developers: compute, storage, databases, networking, and identity with hands-on examples.
GuideGCP Basics: Core Services for Developers
A practical guide to Google Cloud Platform core services for developers: compute, storage, databases, networking, and data analytics with hands-on examples.
RecipeReduce AWS Lambda Cold Start with Provisioned Concurrency
Minimize Lambda cold start latency using provisioned concurrency, ARM64 Graviton, lighter dependencies, and initialization code optimization.
RecipePackage Python Dependencies for AWS Lambda with Layers
Package Python dependencies for AWS Lambda using Lambda Layers, Docker builds for native extensions, and SAM/Serverless Framework integration.
RecipeBuild HTTP-Triggered Azure Functions with Python
Create HTTP-triggered Azure Functions in Python with binding configuration, async handlers, dependency injection, and deployment via Azure CLI.