The basic request-response Lambda is simple. Getting serverless right at scale is not.
After years running production serverless workloads that process millions of events daily, I've seen where the textbook patterns break down. Cold starts spike during traffic surges. State management becomes a mess across function boundaries. Costs balloon because every tutorial assumes infinite budget. The patterns below are what actually worked when response times mattered and the bill came due.
Event choreography vs orchestration
Most teams start with direct Lambda-to-Lambda calls because it feels natural. Function A invokes Function B, which calls Function C. Clean until it isn't.
Choreography flips this. Each function publishes events to a bus and subscribes only to what it needs. When an order is placed, the order service emits OrderCreated. The inventory service listens for that event and reserves stock independently. Payment service does the same for charging. No function knows the others exist.
The trade-off is visibility. With orchestration you have a single function coordinating everything, so the flow is explicit in code. Choreography spreads that logic across event subscriptions. Debugging means tracing events through logs, not reading a single file.
I prefer choreography for high-throughput paths where failures in one domain shouldn't block others. An inventory check can fail without stopping payment processing if you design the events right. Use orchestration when you need strong guarantees about execution order or when a process has a clear owner—think a Step Functions state machine handling a multi-hour approval workflow.
CQRS with event sourcing at the edge
Command Query Responsibility Segregation separates writes from reads. In serverless, this usually means one Lambda writes to DynamoDB and another reads from a denormalized view.
The real win is event sourcing the write side. Instead of updating a record in place, append every state change as an immutable event. Your order isn't a row that gets modified; it's a stream of OrderPlaced, PaymentProcessed, Shipped events. Rebuild current state by replaying the log.
Why bother? Two reasons. Audit trails come free—you have every state transition with a timestamp. More importantly, you can project multiple read models from the same event stream without touching the write path. Your customer service dashboard queries one DynamoDB table optimized for lookups by customer ID. Your analytics pipeline reads the same events into another table keyed by product SKU.
Push this pattern to the edge by placing the query side in Lambda@Edge or CloudFront Functions. The command side stays in regional Lambdas writing to DynamoDB Streams. Your read model sits in DynamoDB Global Tables so every edge location has a local copy. Response times drop below 50ms because you're serving cached projections from the nearest CloudFront POP.
The gotcha is eventual consistency. A customer places an order and immediately refreshes the page. If the read model hasn't caught up, the order won't appear. Handle this with optimistic UI updates on the client or a short-lived cache of recent writes that the query function checks first.
Saga pattern for distributed transactions
Serverless architectures split work across functions. What happens when a business process spans multiple services and one step fails halfway through?
A saga breaks a transaction into local transactions, each with a compensating action that undoes it. Book a flight, reserve a hotel, charge the card. If the card declines, you trigger compensations: cancel the hotel, cancel the flight. Each service handles its own rollback.
Implement this with Step Functions or a custom Lambda that maintains saga state in DynamoDB. The orchestrator calls each service in sequence. On failure, it walks backward through completed steps invoking compensations.
# Step Functions state machine (simplified)
States:
BookFlight:
Type: Task
Resource: arn:aws:lambda:function:bookFlight
Catch:
- ErrorEquals: ["States.ALL"]
ResultPath: $.error
Next: CancelFlight
Next: ReserveHotel
ReserveHotel:
Type: Task
Resource: arn:aws:lambda:function:reserveHotel
Catch:
- ErrorEquals: ["States.ALL"]
Next: CancelHotel
Next: ChargeCard
CancelHotel:
Type: Task
Resource: arn:aws:lambda:function:cancelHotel
Next: CancelFlight
The hard part is idempotency. Network issues mean a compensation might run twice. Your cancelFlight function must check if the flight is already canceled before trying again. Store a unique saga ID and step state in your database. Check it on every call.
Sagas add latency because each step is sequential. If your process can tolerate partial failures and fix them asynchronously, event choreography is faster. Use sagas when consistency guarantees matter more than speed—payment processing, account creation, anything involving money or contracts.
Eliminating cold starts with provisioned concurrency
Cold starts are the tax you pay for pay-per-use pricing. The first request to a new Lambda instance takes longer because AWS has to allocate a container, load your code, and run initialization.
Provisioned concurrency keeps instances warm. You tell Lambda to maintain a pool of pre-initialized containers ready to handle requests. Response time becomes predictable because there's no cold start penalty.
The cost is literal—you pay for those instances even when idle. I've seen teams enable provisioned concurrency on every function and watch their bill triple. Target it strategically.
Enable it on functions behind API Gateway that users hit directly. A 2-second cold start on your homepage Lambda is unacceptable. Background job processors can tolerate cold starts because no human is waiting.
Set concurrency based on observed traffic patterns, not theoretical max load. If your API serves 100 requests per second at peak with an average duration of 200ms, you need about 20 concurrent executions. Provision 25 to handle bursts. Monitor ProvisionedConcurrencySpilloverInvocations to see when you're undersized.
For functions with heavy initialization (large libraries, ML models), move that work into a Lambda Layer or container image cached in ECR. Initialization time drops when dependencies are already on disk.
Stream processing with partial batch failure
Lambda can read from Kinesis or DynamoDB Streams and process records in batches. The default behavior is all-or-nothing: if any record fails, the entire batch is retried.
This creates poison pills. One malformed record blocks the entire shard until you manually skip it or fix the bug. Messages pile up behind the bad record. Latency spikes.
Partial batch failure handling lets you report which specific records failed. Lambda retries only those and moves forward with the rest.
def handler(event, context):
records = event['Records']
failed_records = []
for record in records:
try:
process_record(record)
except Exception as e:
failed_records.append({
'itemIdentifier': record['kinesis']['sequenceNumber']
})
log_error(e, record)
return {
'batchItemFailures': failed_records
}
Enable ReportBatchItemFailures on your event source mapping. Your function returns a list of sequence numbers that failed. Lambda retries those specific records while processing newer ones.
You still need a dead-letter queue for records that fail repeatedly. After a few retries, send them to SQS or S3 for manual inspection. Don't let them block the stream forever.
Watch your batch size and window settings. A batch size of 1000 with a 10-second window means Lambda waits up to 10 seconds or until 1000 records arrive before invoking. For low-latency streams, drop the batch size and window. For high-throughput batch jobs, increase them to reduce invocation overhead.
State machine as a function
Step Functions is the obvious choice for orchestrating long-running workflows. But the pricing model—per state transition—can surprise you. A workflow with 100 state transitions costs more than one with 10, even if they do the same work.
For workflows with many small steps or high execution volumes, implement the state machine directly in a Lambda function using a state table pattern.
Each workflow instance has a row in DynamoDB tracking current state, pending actions, and history. A single Lambda function reads the state, executes the next action, updates the row, and triggers itself again if more work remains.
def handler(event, context):
workflow_id = event['workflow_id']
state = dynamo.get_item(Key={'id': workflow_id})
if state['current'] == 'PENDING_APPROVAL':
if check_approval(state):
state['current'] = 'APPROVED'
dynamo.update_item(Item=state)
lambda_client.invoke(
FunctionName=context.function_name,
InvocationType='Event',
Payload=json.dumps({'workflow_id': workflow_id})
)
elif state['current'] == 'APPROVED':
provision_resources(state)
state['current'] = 'COMPLETE'
dynamo.update_item(Item=state)
This cuts costs when you have thousands of workflows with simple state logic. The downside is you lose Step Functions' visual workflow editor and built-in retry/error handling. You're writing that logic by hand.
Use this for high-volume, straightforward workflows where the state transitions are predictable and you need fine-grained control over retry behavior. Keep Step Functions for complex branching, human-in-the-loop approvals, or anything that needs to run for hours.
Fan-out with SNS and selective filtering
A common pattern is one event triggering multiple downstream actions. An order confirmation sends an email, updates inventory, logs to analytics, and notifies a Slack channel.
The naive approach is one Lambda calling four services. Better is publishing to SNS and having four subscribers—one per action. Each subscriber is a Lambda that does one thing. They run in parallel, and if one fails it doesn't affect the others.
SNS subscription filters let you route messages without invoking every function. Each subscriber declares what it cares about using JSON filter policies.
{
"eventType": ["ORDER_PLACED"],
"orderTotal": [{"numeric": [">", 1000]}]
}
Your high-value order handler only gets invoked when orderTotal exceeds 1000. Low-value orders don't trigger that function at all. You skip unnecessary invocations and reduce cost.
SNS topic payloads are limited to 256 KB. For large messages, store the payload in S3 and publish a reference. Subscribers fetch the full object if needed.
One edge case that bit me: SNS retries failed deliveries with exponential backoff, but if a Lambda subscriber stays broken for too long, SNS moves the message to a dead-letter queue. Set up CloudWatch Alarms on your DLQ depth metric. Messages sitting in a DLQ mean a subscriber is down and events are getting lost.
How do I choose between Step Functions and a custom state machine?
If your workflow has branching logic, waits for external input, or needs a visual representation for non-engineers, use Step Functions. If you're running thousands of executions per second with simple linear state changes and cost matters, a custom Lambda state machine is cheaper. Step Functions bills per state transition; a Lambda state machine bills per invocation and compute time.
When should I use provisioned concurrency?
When cold starts directly impact user experience and you have predictable traffic patterns. Don't enable it globally. Profile each function, measure actual cold start frequency and duration, then provision only for user-facing endpoints where latency SLAs are tight. Background jobs can cold start.
How do I handle poison pill messages in stream processing?
Enable partial batch failure handling and implement retries with exponential backoff in your function code. After a configured number of retries (usually 3-5), send failed records to a dead-letter queue for manual review. Never let one bad record block the entire stream indefinitely.
What's the biggest mistake teams make with event-driven serverless?
Not designing for idempotency. Functions will be retried. Events will be delivered more than once. If your function charges a credit card or sends an email, check if you already processed that event ID before doing it again. Store processed event IDs in DynamoDB with a TTL to automatically clean up old entries.
What actually matters
Patterns don't scale systems. Understanding trade-offs does.
Choreography decouples services but makes debugging harder. CQRS gives you flexibility at the cost of eventual consistency. Provisioned concurrency buys speed but costs money even at idle. Every pattern has a price.
The serverless architectures that survive production are the ones designed around actual constraints: traffic patterns, budget limits, team size, acceptable latency. Start with the simplest thing that meets your requirements. Add complexity only when measurements prove you need it.
