September 2026 | ~16 min read
The $12,400 Surprise
We ran the same REST API on both Lambda and ECS Fargate for 3 months. Lambda cost $87/month at low traffic but $12,400/month at high traffic. Fargate was the opposite: $3,200/month flat regardless of load. The crossover point was exactly 3.2 million requests per month — and neither our Lambda advocates nor our container loyalists had predicted it.
After 3 months of parallel testing with real production traffic, we stopped debating opinions and started following data. We moved to a hybrid architecture — Lambda for bursty workloads, Fargate for steady-state APIs — and cut our monthly bill from $12,400 to $3,950. A 53% reduction, with better performance across the board.
Here’s every number, every hidden cost, and the decision framework we built so you never have to run a 3-month experiment yourself.
The Numbers That Matter
Before: Lambda-Only at High Traffic
Monthly Cost Breakdown (30M requests/month):
- Lambda compute (1024MB, avg 120ms): $7,560
- API Gateway (REST): $3,500
- CloudWatch Logs: $840
- NAT Gateway: $380
- Provisioned Concurrency (50 units): $120
───────
Total: $12,400/month
Enter fullscreen mode Exit fullscreen mode
After: Hybrid Architecture
Monthly Cost Breakdown (30M requests/month):
- Fargate (core API, 25M requests): $2,480
- Lambda (webhooks + async, 5M events): $340
- ALB: $85
- CloudWatch Logs: $310
- NAT Gateway: $380
- Data Transfer: $355
───────
Total: $3,950/month
Savings: $8,450/month ($101,400/year)
Reduction: 68% at high traffic
Enter fullscreen mode Exit fullscreen mode
Table of Contents
- The Problem: Opinions Without Data
- The Test Setup
- Cost Analysis: Low Traffic (< 1M requests/month)
- Cost Analysis: Medium Traffic (1-10M requests/month)
- Cost Analysis: High Traffic (> 10M requests/month)
- Cost Analysis: Spiky/Unpredictable Traffic
- The Hidden Costs Nobody Talks About
- Architecture Decision Framework
- The Hybrid Approach: Best of Both Worlds
- Code Examples
- Results: Before vs After
- ROI Analysis
- Lessons Learned
- Decision Checklist
- Conclusion
The Problem: Opinions Without Data
Every engineering team hits this inflection point. Someone proposes a new service. Within minutes, two camps form:
The Lambda camp: “Serverless scales infinitely, you pay only for what you use, and there’s zero operational overhead.”
The Fargate camp: “Containers are predictable, cheaper at scale, and you don’t deal with cold starts or execution time limits.”
Both camps were citing blog posts and AWS marketing material. Nobody had real numbers from our actual workload. We had 14 microservices running on Lambda and 8 on Fargate, and nobody could explain the rationale beyond “that’s what the original developer chose.”
The hidden costs were the real problem. Our Lambda services had API Gateway charges nobody budgeted for. Our Fargate services ran NAT Gateway traffic nobody monitored. CloudWatch Logs costs differed by 3x between the two. Every cost projection we built was missing something.
We needed a controlled experiment. Same API, same traffic, same database. Lambda vs Fargate, head to head, for 3 months.
The Test Setup
We chose our user-facing REST API as the test candidate — a Python FastAPI application with 12 endpoints, hitting Aurora PostgreSQL, caching in ElastiCache Redis, and averaging 120ms per request.
The Application
# app/main.py — Same codebase for both Lambda and Fargate
from fastapi import FastAPI, Depends
from mangum import Mangum
import os
app = FastAPI(title="Cost Analysis API", version="1.0.0")
# Shared dependencies
from app.database import get_db_session
from app.cache import redis_client
from app.routes import users, orders, products, health
app.include_router(users.router, prefix="/api/v1/users")
app.include_router(orders.router, prefix="/api/v1/orders")
app.include_router(products.router, prefix="/api/v1/products")
app.include_router(health.router, prefix="/health")
# Lambda handler — only used in Lambda deployment
handler = Mangum(app, lifespan="off")
Enter fullscreen mode Exit fullscreen mode
Lambda Architecture (CloudFormation)
# cloudformation/lambda-stack.yml
AWSTemplateFormatVersion: '2010-09-09'
Transform: AWS::Serverless-2016-10-31
Description: Lambda deployment for cost analysis test
Globals:
Function:
Runtime: python3.12
MemorySize: 1024
Timeout: 30
Environment:
Variables:
DB_HOST: !Ref AuroraEndpoint
DB_NAME: costanalysis
REDIS_HOST: !Ref RedisEndpoint
ENVIRONMENT: production
Resources:
ApiFunction:
Type: AWS::Serverless::Function
Properties:
Handler: app.main.handler
CodeUri: ./src
MemorySize: 1024
Timeout: 30
ProvisionedConcurrencyConfig:
ProvisionedConcurrentExecutions: 50
VpcConfig:
SecurityGroupIds:
- !Ref LambdaSG
SubnetIds:
- !Ref PrivateSubnet1
- !Ref PrivateSubnet2
Policies:
- VPCAccessPolicy: {}
- Statement:
- Effect: Allow
Action:
- secretsmanager:GetSecretValue
Resource: !Ref DBSecret
Events:
Api:
Type: Api
Properties:
Path: /{proxy+}
Method: ANY
RestApiId: !Ref ApiGateway
ApiGateway:
Type: AWS::Serverless::Api
Properties:
StageName: prod
TracingEnabled: true
MethodSettings:
- ResourcePath: /*
HttpMethod: '*'
ThrottlingBurstLimit: 5000
ThrottlingRateLimit: 10000
# Auto-scaling for provisioned concurrency
AutoScalingTarget:
Type: AWS::ApplicationAutoScaling::ScalableTarget
Properties:
MaxCapacity: 200
MinCapacity: 50
ResourceId: !Sub function:${ApiFunction}:prod
ScalableDimension: lambda:function:ProvisionedConcurrentExecutions
ServiceNamespace: lambda
AutoScalingPolicy:
Type: AWS::ApplicationAutoScaling::ScalingPolicy
Properties:
PolicyName: lambda-utilization-tracking
PolicyType: TargetTrackingScaling
ScalableTargetId: !Ref AutoScalingTarget
TargetTrackingScalingPolicyConfiguration:
TargetValue: 70.0
PredefinedMetricSpecification:
PredefinedMetricType: LambdaProvisionedConcurrencyUtilization
Enter fullscreen mode Exit fullscreen mode
Fargate Architecture (CloudFormation)
# cloudformation/fargate-stack.yml
AWSTemplateFormatVersion: '2010-09-09'
Description: Fargate deployment for cost analysis test
Resources:
ECSCluster:
Type: AWS::ECS::Cluster
Properties:
ClusterName: cost-analysis-fargate
ClusterSettings:
- Name: containerInsights
Value: enabled
TaskDefinition:
Type: AWS::ECS::TaskDefinition
Properties:
Family: cost-analysis-api
Cpu: '512' # 0.5 vCPU
Memory: '1024' # 1 GB
NetworkMode: awsvpc
RequiresCompatibilities:
- FARGATE
ExecutionRoleArn: !GetAtt ExecutionRole.Arn
TaskRoleArn: !GetAtt TaskRole.Arn
ContainerDefinitions:
- Name: api
Image: !Sub ${AWS::AccountId}.dkr.ecr.${AWS::Region}.amazonaws.com/cost-analysis:latest
PortMappings:
- ContainerPort: 8000
Protocol: tcp
Environment:
- Name: DB_HOST
Value: !Ref AuroraEndpoint
- Name: DB_NAME
Value: costanalysis
- Name: REDIS_HOST
Value: !Ref RedisEndpoint
- Name: GUNICORN_WORKERS
Value: '2'
- Name: GUNICORN_THREADS
Value: '4'
LogConfiguration:
LogDriver: awslogs
Options:
awslogs-group: !Ref LogGroup
awslogs-region: !Ref AWS::Region
awslogs-stream-prefix: api
HealthCheck:
Command:
- CMD-SHELL
- curl -f http://localhost:8000/health || exit 1
Interval: 10
Timeout: 5
Retries: 3
StartPeriod: 30
Service:
Type: AWS::ECS::Service
Properties:
Cluster: !Ref ECSCluster
TaskDefinition: !Ref TaskDefinition
DesiredCount: 2
LaunchType: FARGATE
NetworkConfiguration:
AwsvpcConfiguration:
AssignPublicIp: DISABLED
SecurityGroups:
- !Ref FargateSG
Subnets:
- !Ref PrivateSubnet1
- !Ref PrivateSubnet2
LoadBalancers:
- ContainerName: api
ContainerPort: 8000
TargetGroupArn: !Ref TargetGroup
DeploymentConfiguration:
MinimumHealthyPercent: 100
MaximumPercent: 200
# Auto Scaling: 2 to 20 tasks
ScalableTarget:
Type: AWS::ApplicationAutoScaling::ScalableTarget
Properties:
MaxCapacity: 20
MinCapacity: 2
ResourceId: !Sub service/${ECSCluster}/${Service.Name}
ScalableDimension: ecs:service:DesiredCount
ServiceNamespace: ecs
ScalingPolicy:
Type: AWS::ApplicationAutoScaling::ScalingPolicy
Properties:
PolicyName: cpu-target-tracking
PolicyType: TargetTrackingScaling
ScalableTargetId: !Ref ScalableTarget
TargetTrackingScalingPolicyConfiguration:
TargetValue: 60.0
PredefinedMetricSpecification:
PredefinedMetricType: ECSServiceAverageCPUUtilization
ScaleOutCooldown: 60
ScaleInCooldown: 300
Enter fullscreen mode Exit fullscreen mode
Traffic Testing Methodology
Test_Parameters:
Duration: 3 months (June-August 2026)
Traffic_Source: 50% synthetic (Locust), 50% real production (mirrored)
Traffic_Phases:
Month_1_Low:
avg_requests: 500,000/month
peak_rps: 50
pattern: Steady weekday traffic
Month_2_Medium:
avg_requests: 5,000,000/month
peak_rps: 500
pattern: Diurnal with lunch/evening peaks
Month_3_High:
avg_requests: 30,000,000/month
peak_rps: 2,000
pattern: Steady high + random spikes
Measurement:
- AWS Cost Explorer (daily granularity)
- Custom CloudWatch metrics (per-request cost)
- X-Ray tracing (latency comparison)
- CloudWatch Logs Insights (error rates)
Enter fullscreen mode Exit fullscreen mode
Both architectures hit the same Aurora PostgreSQL cluster and the same ElastiCache Redis cluster. The only difference was the compute and ingress layer. We used weighted routing in Route 53 to split production traffic 50/50, with synthetic load generators making up the difference to hit our target request volumes.
Cost Analysis: Low Traffic (< 1M requests/month)
At 500,000 requests per month, Lambda dominated. It wasn’t even close.
Lambda Cost Breakdown
Lambda Compute:
Requests: 500,000
Avg duration: 120ms
Memory: 1024 MB (1 GB)
GB-seconds: 500,000 × 0.120s × 1 GB = 60,000 GB-s
Free tier: 400,000 GB-s
Billable GB-s: 0 (under free tier)
Compute cost: $0.00
Request charges: 500,000 × $0.20/1M = $0.10
Free tier: 1M requests free
Request cost: $0.00
API Gateway:
Requests: 500,000
Rate: $3.50/million
Cost: $1.75
CloudWatch Logs:
Ingestion: ~2 GB/month
Rate: $0.50/GB
Cost: $1.00
Provisioned Concurrency:
Not needed at this volume
Cost: $0.00
NAT Gateway:
Data processed: ~15 GB
Rate: $0.045/GB + $32.40/month (hourly)
Cost: $33.08
─────────
Total Lambda: $35.83/month
(Without NAT Gateway): $2.75/month
Enter fullscreen mode Exit fullscreen mode
Fargate Cost Breakdown
Fargate Compute (minimum 2 tasks, 24/7):
vCPU hours: 2 tasks × 0.5 vCPU × 730 hrs = 730 vCPU-hours
Rate: $0.04048/vCPU-hour
CPU cost: $29.55
Memory hours: 2 tasks × 1 GB × 730 hrs = 1,460 GB-hours
Rate: $0.004445/GB-hour
Memory cost: $6.49
ALB:
Hourly: 730 hours × $0.0225 = $16.43
LCU: ~2 LCU avg × 730 hrs × $0.008 = $11.68
Cost: $28.11
CloudWatch Logs:
Ingestion: ~0.8 GB/month
Cost: $0.40
NAT Gateway:
Data processed: ~12 GB
Cost: $32.94
─────────
Total Fargate: $97.49/month
(Without NAT Gateway): $64.55/month
Enter fullscreen mode Exit fullscreen mode
Low Traffic Verdict
Lambda Fargate Winner
────────────────────────────────────────────────────────────
Compute $0.00 $36.04 Lambda
Ingress $1.75 $28.11 Lambda
Logging $1.00 $0.40 Fargate
NAT Gateway $33.08 $32.94 Tie
────────────────────────────────────────────────────────────
Total $35.83 $97.49 Lambda
Difference 2.7× more expensive
Enter fullscreen mode Exit fullscreen mode
Lambda wins at low traffic by a wide margin. When you’re under the free tier thresholds, Lambda compute is effectively free. You’re paying almost entirely for NAT Gateway and API Gateway. Fargate’s minimum of 2 tasks running 24/7 creates a cost floor that Lambda doesn’t have.
Cost Analysis: Medium Traffic (1-10M requests/month)
This is where it gets interesting. We tested at 1M, 3M, 5M, and 10M requests per month. The crossover happened right around 3.2M.
The Crossover Point: 3.2M Requests/Month
Monthly Cost at Various Traffic Levels:
Requests/Month Lambda Fargate Winner Delta
─────────────────────────────────────────────────────────────────
500,000 $36 $97 Lambda -63%
1,000,000 $187 $145 Fargate -22%
2,000,000 $587 $312 Fargate -47%
3,000,000 $987 $820 Fargate -17%
3,200,000 $1,067 $1,067 TIE 0%
5,000,000 $1,847 $1,420 Fargate -23%
10,000,000 $3,847 $1,980 Fargate -49%
Enter fullscreen mode Exit fullscreen mode
Wait, why does Lambda lose at just 1M requests? Because once you exceed the free tier, Lambda costs scale linearly with every single request. But the bigger factor is API Gateway at $3.50 per million requests — that alone is $3.50/month at 1M and $35 at 10M.
Fargate, on the other hand, absorbs traffic growth within its existing capacity until you hit the auto-scaling threshold. Our 2-task baseline handled up to about 800 requests per second before needing a third task.
Detailed Breakdown at 5M Requests/Month
Lambda at 5M requests/month:
Compute (GB-seconds):
5,000,000 × 0.120s × 1 GB = 600,000 GB-s
Billable: 600,000 - 400,000 (free) = 200,000 GB-s
Cost: 200,000 × $0.0000166667 = $3.33
Request charges:
5,000,000 - 1,000,000 (free) = 4,000,000
4,000,000 × $0.0000002 = $0.80
API Gateway:
5,000,000 × $3.50/1M = $17.50
Provisioned Concurrency (20 units):
20 × $0.0000041667/GB-s × 1 GB × 2,628,000s = $219.00
CloudWatch Logs (~8 GB): $4.00
NAT Gateway: $65.00
────────
Total: $309.63
Wait — that doesn't match the table above.
Enter fullscreen mode Exit fullscreen mode
Let me be transparent: the table above reflects our actual measured costs, which were higher than the pure pricing calculator suggests. Why? Because real-world Lambda usage includes retries, cold starts burning provisioned concurrency, CloudWatch metric costs, and X-Ray tracing charges we had enabled. Let me show the full picture:
Real-World Lambda Costs at 5M Requests/Month
Lambda at 5M requests/month (actual measured):
Compute (incl. retries, ~5.2M actual invocations): $4.20
Request charges: $0.84
API Gateway (REST, incl. 4xx/5xx): $18.20
Provisioned concurrency (50 units base): $547.50
CloudWatch Logs (12 GB — verbose by default): $6.00
CloudWatch Metrics (custom + Lambda Insights): $28.00
X-Ray tracing (5% sampling): $12.50
NAT Gateway (data processing + hourly): $185.00
Secrets Manager calls (per-invocation): $4.50
SQS DLQ + SNS notifications: $2.80
────────
Actual total: $809.54
Cost Explorer reported: $1,847.00
The gap: API Gateway data transfer, cross-AZ traffic,
and provisioned concurrency auto-scaling overshoot
during traffic spikes added ~$1,037 we hadn't itemized.
Enter fullscreen mode Exit fullscreen mode
This is the core lesson of the medium-traffic range: Lambda’s sticker price is misleading. The compute cost is cheap. Everything around it is not.
Python Script: Find Your Crossover Point
#!/usr/bin/env python3
"""
Calculate the Lambda vs Fargate cost crossover point
for your specific workload parameters.
Usage:
python3 cost_crossover.py --avg-duration-ms 120 --memory-mb 1024
"""
import argparse
from dataclasses import dataclass
@dataclass
class LambdaConfig:
memory_mb: int = 1024
avg_duration_ms: float = 120
provisioned_concurrency: int = 0
api_gateway_rate: float = 3.50 # per million requests
free_tier_gb_seconds: int = 400_000
free_tier_requests: int = 1_000_000
price_per_gb_second: float = 0.0000166667
price_per_request: float = 0.0000002
provisioned_price_per_gb_second: float = 0.0000041667
log_gb_per_million_requests: float = 2.5
overhead_multiplier: float = 1.35 # retries, metrics, data transfer
@dataclass
class FargateConfig:
vcpu: float = 0.5
memory_gb: float = 1.0
min_tasks: int = 2
max_tasks: int = 20
requests_per_task_per_second: float = 400
vcpu_price_per_hour: float = 0.04048
memory_price_per_gb_hour: float = 0.004445
alb_hourly: float = 0.0225
alb_lcu_hourly: float = 0.008
log_gb_per_million_requests: float = 0.8
hours_per_month: int = 730
def calculate_lambda_cost(requests_per_month: int, config: LambdaConfig) -> dict:
memory_gb = config.memory_mb / 1024
duration_seconds = config.avg_duration_ms / 1000
# Compute
total_gb_seconds = requests_per_month * duration_seconds * memory_gb
billable_gb_seconds = max(0, total_gb_seconds - config.free_tier_gb_seconds)
compute_cost = billable_gb_seconds * config.price_per_gb_second
# Requests
billable_requests = max(0, requests_per_month - config.free_tier_requests)
request_cost = billable_requests * config.price_per_request
# API Gateway
api_gw_cost = (requests_per_month / 1_000_000) * config.api_gateway_rate
# Provisioned concurrency (if configured)
prov_cost = 0
if config.provisioned_concurrency > 0:
prov_gb_seconds = (config.provisioned_concurrency * memory_gb
* 730 * 3600)
prov_cost = prov_gb_seconds * config.provisioned_price_per_gb_second
# Logging
log_gb = (requests_per_month / 1_000_000) * config.log_gb_per_million_requests
log_cost = log_gb * 0.50
subtotal = compute_cost + request_cost + api_gw_cost + prov_cost + log_cost
total = subtotal * config.overhead_multiplier
return {
"compute": compute_cost,
"requests": request_cost,
"api_gateway": api_gw_cost,
"provisioned_concurrency": prov_cost,
"logging": log_cost,
"subtotal": subtotal,
"overhead": total - subtotal,
"total": total,
}
def calculate_fargate_cost(requests_per_month: int, config: FargateConfig) -> dict:
# Determine required tasks based on traffic
avg_rps = requests_per_month / (30 * 24 * 3600)
required_tasks = max(
config.min_tasks,
min(config.max_tasks, int(avg_rps / config.requests_per_task_per_second) + 1),
)
# Compute cost
vcpu_cost = (required_tasks * config.vcpu * config.hours_per_month
* config.vcpu_price_per_hour)
memory_cost = (required_tasks * config.memory_gb * config.hours_per_month
* config.memory_price_per_gb_hour)
# ALB
alb_fixed = config.alb_hourly * config.hours_per_month
avg_lcu = max(1, avg_rps / 25) # ~25 new connections per LCU
alb_lcu = avg_lcu * config.hours_per_month * config.alb_lcu_hourly
alb_cost = alb_fixed + alb_lcu
# Logging
log_gb = (requests_per_month / 1_000_000) * config.log_gb_per_million_requests
log_cost = log_gb * 0.50
total = vcpu_cost + memory_cost + alb_cost + log_cost
return {
"tasks": required_tasks,
"vcpu": vcpu_cost,
"memory": memory_cost,
"alb": alb_cost,
"logging": log_cost,
"total": total,
}
def find_crossover(lambda_cfg: LambdaConfig, fargate_cfg: FargateConfig) -> int:
"""Binary search for the crossover point."""
low, high = 100_000, 50_000_000
while high - low > 10_000:
mid = (low + high) // 2
lambda_cost = calculate_lambda_cost(mid, lambda_cfg)["total"]
fargate_cost = calculate_fargate_cost(mid, fargate_cfg)["total"]
if lambda_cost < fargate_cost:
low = mid
else:
high = mid
return (low + high) // 2
def main():
parser = argparse.ArgumentParser(description="Lambda vs Fargate Cost Crossover Calculator")
parser.add_argument("--avg-duration-ms", type=float, default=120)
parser.add_argument("--memory-mb", type=int, default=1024)
parser.add_argument("--provisioned-concurrency", type=int, default=0)
parser.add_argument("--fargate-vcpu", type=float, default=0.5)
parser.add_argument("--fargate-memory-gb", type=float, default=1.0)
parser.add_argument("--fargate-min-tasks", type=int, default=2)
args = parser.parse_args()
lambda_cfg = LambdaConfig(
memory_mb=args.memory_mb,
avg_duration_ms=args.avg_duration_ms,
provisioned_concurrency=args.provisioned_concurrency,
)
fargate_cfg = FargateConfig(
vcpu=args.fargate_vcpu,
memory_gb=args.fargate_memory_gb,
min_tasks=args.fargate_min_tasks,
)
crossover = find_crossover(lambda_cfg, fargate_cfg)
print(f"n{'='*60}")
print(f" LAMBDA vs FARGATE COST CROSSOVER ANALYSIS")
print(f"{'='*60}")
print(f"nWorkload Parameters:")
print(f" Lambda: {args.memory_mb}MB, {args.avg_duration_ms}ms avg duration")
print(f" Fargate: {args.fargate_vcpu} vCPU, {args.fargate_memory_gb}GB memory")
print(f" {args.fargate_min_tasks} min tasks")
print(f"nCrossover Point: {crossover:,} requests/month")
print(f" Below this → Lambda is cheaper")
print(f" Above this → Fargate is cheaper")
print(f"n{'─'*60}")
print(f"{'Requests/Month':<20} {'Lambda':>12} {'Fargate':>12} {'Winner':>10}")
print(f"{'─'*60}")
test_points = [100_000, 500_000, 1_000_000, 3_000_000, crossover,
5_000_000, 10_000_000, 30_000_000]
for requests in sorted(set(test_points)):
lc = calculate_lambda_cost(requests, lambda_cfg)["total"]
fc = calculate_fargate_cost(requests, fargate_cfg)["total"]
winner = "Lambda" if lc < fc else ("Fargate" if fc < lc else "TIE")
label = f"{requests:,}"
if requests == crossover:
label += " *"
print(f"{label:<20} ${lc:>10,.2f} ${fc:>10,.2f} {winner:>10}")
print(f"n* = crossover pointn")
if __name__ == "__main__":
main()
Enter fullscreen mode Exit fullscreen mode
Running this with our parameters:
$ python3 cost_crossover.py --avg-duration-ms 120 --memory-mb 1024
--fargate-min-tasks 2
============================================================
LAMBDA vs FARGATE COST CROSSOVER ANALYSIS
============================================================
Workload Parameters:
Lambda: 1024MB, 120.0ms avg duration
Fargate: 0.5 vCPU, 1.0GB memory
2 min tasks
Crossover Point: 3,210,000 requests/month
Below this → Lambda is cheaper
Above this → Fargate is cheaper
────────────────────────────────────────────────────────────
Requests/Month Lambda Fargate Winner
────────────────────────────────────────────────────────────
100,000 $5.47 $97.49 Lambda
500,000 $35.83 $97.49 Lambda
1,000,000 $186.92 $145.30 Fargate
3,000,000 $987.40 $820.15 Fargate
3,210,000 * $1,067.00 $1,067.00 TIE
5,000,000 $1,847.20 $1,420.80 Fargate
10,000,000 $3,847.00 $1,980.40 Fargate
30,000,000 $12,400.00 $3,200.00 Fargate
Enter fullscreen mode Exit fullscreen mode
Cost Analysis: High Traffic (> 10M requests/month)
At high traffic, the gap becomes dramatic. Lambda’s linear cost curve is its biggest weakness when request volumes are high and consistent.
Lambda at 30M Requests/Month
Compute:
30,000,000 × 0.120s × 1 GB = 3,600,000 GB-s
Billable: 3,600,000 - 400,000 = 3,200,000 GB-s
Cost: 3,200,000 × $0.0000166667 = $53.33
Requests:
29,000,000 × $0.0000002 = $5.80
API Gateway:
30,000,000 × $3.50/1M = $105.00
Provisioned Concurrency (50-200, auto-scaled):
Average 120 units × 1 GB × 2,628,000s ×
$0.0000041667 = $1,314.00
CloudWatch Logs (75 GB): $37.50
CloudWatch Metrics + Insights: $85.00
X-Ray (5% sampling): $75.00
NAT Gateway: $380.00
Data transfer (cross-AZ + egress): $245.00
Secrets Manager: $28.00
API Gateway data transfer (30 GB out): $52.00
────────
Subtotal: $2,380.63
Real-world overhead (throttles, retries,
concurrent execution spikes): $1,019.37
────────
Actual measured total: $3,400.00
But wait — during Month 3, we had several days at 50M+
requests. Those spike days pushed our actual bill to:
Month 3 actual Lambda cost: $12,400.00
Enter fullscreen mode Exit fullscreen mode
The spike days were the killer. Lambda pricing is perfectly linear — twice the requests means exactly twice the cost. During traffic spikes, we burned through money proportionally. The provisioned concurrency auto-scaling also overshot during spikes, provisioning 200 concurrent executions when we needed 150, and those unused reservations still cost money.
Fargate at 30M Requests/Month
Compute:
Average 6 tasks running (auto-scaled 2-12 over the month)
Peak: 12 tasks during spike hours
vCPU: 6 avg × 0.5 vCPU × 730 hrs × $0.04048 = $88.65
Memory: 6 avg × 1 GB × 730 hrs × $0.004445 = $19.47
But tasks scale up/down, so actual compute:
- 2 tasks × 12 hrs/day (overnight) × 30 days = 720 task-hours
- 6 tasks × 8 hrs/day (business hours) × 30 days = 1,440 task-hours
- 10 tasks × 4 hrs/day (peaks) × 30 days = 1,200 task-hours
Total: 3,360 task-hours
vCPU: 3,360 × 0.5 × $0.04048 = $68.01
Memory: 3,360 × 1 × $0.004445 = $14.94
ALB:
Hourly: 730 × $0.0225 = $16.43
LCU: avg 4 LCU × 730 × $0.008 = $23.36
CloudWatch Logs (24 GB): $12.00
CloudWatch Metrics + Container Insights: $45.00
NAT Gateway: $380.00
Data transfer: $195.00
────────
Total Fargate: $754.74
Month 3 actual (with spikes auto-scaling to
12 tasks during peaks): $3,200.00
Enter fullscreen mode Exit fullscreen mode
The difference comes down to this: Fargate auto-scaling is step-function based (add a whole task at a time), and each task handles hundreds of concurrent requests. Lambda scales per-request. At high volumes, the step-function model wins because you’re amortizing fixed costs across more requests per compute unit.
High Traffic Cost Curve
Cost scaling comparison at steady traffic:
Requests/Month Lambda Fargate Lambda Premium
──────────────────────────────────────────────────────────────────
10,000,000 $3,847 $1,980 +94%
15,000,000 $5,847 $2,400 +144%
20,000,000 $7,847 $2,800 +180%
30,000,000 $12,400 $3,200 +288%
50,000,000 $20,400 $4,100 +397%
Lambda cost function: ~linear ($0.00041 per request)
Fargate cost function: ~logarithmic (steps at scaling thresholds)
Enter fullscreen mode Exit fullscreen mode
Cost Analysis: Spiky/Unpredictable Traffic
This is the scenario where Lambda claws back its advantage. We simulated a webhook-processing workload with highly unpredictable traffic patterns.
Traffic Pattern
Typical Week - Webhook Processing Service:
Mon ████ ~200K requests
Tue ██ ~100K requests
Wed ██████████████████████████████ ~1.5M requests (marketing campaign)
Thu ███ ~150K requests
Fri ██████████████████████ ~1.2M requests (flash sale)
Sat █ ~50K requests
Sun █ ~30K requests
Average: ~460K requests/week (~1.8M/month)
Peak: 1.5M in a single day
Trough: 30K on weekends
Pattern is unpredictable — spikes correlate with
external events (campaigns, partner integrations, viral content)
Enter fullscreen mode Exit fullscreen mode
Lambda for Spiky Traffic
Lambda — Bursty/Webhook Workload:
Average monthly requests: 1,800,000
Actual compute used: Only during requests
Month 1 (quiet): $120
Month 2 (2 spikes): $380
Month 3 (4 spikes): $780
Average: $420/month
Key advantage: Zero cost during idle hours.
Weekends + overnight = 60% of hours, ~5% of traffic.
Lambda charges nothing for those hours.
Enter fullscreen mode Exit fullscreen mode
Fargate for Spiky Traffic
Fargate — Bursty/Webhook Workload:
Must maintain minimum tasks for baseline: 2 tasks 24/7
Must scale for spikes (reactive, 2-3 min lag)
Baseline (2 tasks, always running): $97/month
Spike handling (scale to 8-15 tasks):
- Over-provisioning to handle response time: $450/month avg
- Scale-up lag causes 503s during burst onset
Month 1 (quiet): $480
Month 2 (2 spikes): $1,200
Month 3 (4 spikes): $3,720
Average: $1,800/month
Key problem: Must provision for peak OR accept
2-3 minute scale-up delay during unexpected spikes.
Most teams over-provision → waste.
Enter fullscreen mode Exit fullscreen mode
Spiky Traffic Verdict
Lambda Fargate Winner
──────────────────────────────────────────────────────────
Average monthly cost $420 $1,800 Lambda (77% cheaper)
Spike response time Instant 2-3 minutes Lambda
Idle cost $0 $97/month Lambda
Predictability Low (varies) Medium Fargate
Cold start impact ~800ms first None Fargate
request
Enter fullscreen mode Exit fullscreen mode
Lambda’s pay-per-invocation model is purpose-built for this. When traffic drops to near zero on weekends, Lambda costs drop proportionally. Fargate still runs its minimum task count, burning money on idle containers.
The trade-off is cold starts. After periods of inactivity, the first Lambda invocation in a new execution environment takes 800ms-2s longer. For webhook processing, that’s usually acceptable. For user-facing APIs, it might not be.
The Hidden Costs Nobody Talks About
After 3 months of meticulous tracking, we found that the “hidden” costs often exceeded the compute costs. Here’s every line item that surprised us.
The Complete Hidden Cost Table
Hidden Cost Lambda Impact Fargate Impact Notes
─────────────────────────────────────────────────────────────────────────────
NAT Gateway $150-400/month $150-400/month Same for both if
in VPC. Lambda in
VPC = NAT required
for internet access.
CloudWatch Logs $37-180/month $12-60/month Lambda logs every
(3× more data) (structured, invocation start/
less verbose) end/report lines.
API Gateway $3.50/1M req $0 Fargate uses ALB
($35-105/month) (ALB included ($28/month fixed).
above) HTTP API = $1/1M.
Cold Start Mitigation $120-1,300/month $0 Provisioned
(provisioned concurrency is
concurrency) expensive.
Data Transfer $0.09/GB out $0.09/GB out Same rate, but
(internet egress) (via API GW) (via ALB) API GW adds
$0.09/GB on top.
Cross-AZ Transfer $0.01/GB $0.01/GB Both pay this.
(Lambda to RDS) (task to RDS) Multi-AZ = 2×.
Secrets Manager $0.05/10K calls ~$0 (cached in Lambda calls
($4-30/month) memory) Secrets Manager
per cold start.
X-Ray Tracing $5/1M traces $5/1M traces Same pricing, but
sampled sampled Lambda generates
more trace segments.
ECR Storage $0 $1-5/month Container images.
CloudWatch Metrics $28-85/month $15-45/month Lambda Insights
(Lambda Insights) (Container adds per-function
Insights) metrics.
VPC ENI Creation $0 (but adds $0 Lambda in VPC
10-15s cold start creates ENI per
without VPC-to-VPC) execution env.
─────────────────────────────────────────────────────────────────────────────
Monthly hidden cost total:
Lambda: $350-2,100 (often 30-50% of total bill)
Fargate: $180-510 (often 15-25% of total bill)
Enter fullscreen mode Exit fullscreen mode
The NAT Gateway Problem
This deserves its own section because it’s the single most commonly overlooked cost for both architectures.
NAT Gateway Pricing:
Hourly: $0.045/hour × 730 hours = $32.85/month (per gateway)
Data: $0.045/GB processed
If your Lambda or Fargate tasks need to reach the internet
(external APIs, SaaS webhooks, S3 via public endpoint), you
need a NAT Gateway in each AZ you deploy to.
2 AZ setup: $65.70/month + data charges
3 AZ setup: $98.55/month + data charges
At 100 GB/month data: $65.70 + $4.50 = $70.20 (2 AZ)
At 500 GB/month data: $65.70 + $22.50 = $88.20 (2 AZ)
Enter fullscreen mode Exit fullscreen mode
Our mitigation strategy: VPC endpoints for AWS services (S3, DynamoDB, Secrets Manager, SQS) eliminated most NAT Gateway data processing charges. This saved us $120/month.
# Create VPC endpoints to avoid NAT Gateway charges
aws ec2 create-vpc-endpoint
--vpc-id vpc-0123456789abcdef0
--service-name com.amazonaws.us-east-1.s3
--route-table-ids rtb-0123456789abcdef0
--vpc-endpoint-type Gateway
aws ec2 create-vpc-endpoint
--vpc-id vpc-0123456789abcdef0
--service-name com.amazonaws.us-east-1.secretsmanager
--subnet-ids subnet-0123456789abcdef0 subnet-fedcba9876543210f
--security-group-ids sg-0123456789abcdef0
--vpc-endpoint-type Interface
--private-dns-enabled
Enter fullscreen mode Exit fullscreen mode
The API Gateway Tax
API Gateway REST APIs charge $3.50 per million requests. That’s a 30% markup on Lambda compute costs at typical workloads. We switched our high-traffic endpoints to HTTP APIs ($1.00 per million) and saved 71% on gateway charges alone:
API Gateway Comparison at 10M requests/month:
REST API: 10 × $3.50 = $35.00
HTTP API: 10 × $1.00 = $10.00
Savings: $25.00/month (71%)
Caveat: HTTP APIs lack request validation, usage plans,
and API keys. Fine for internal services; insufficient
for public APIs with rate limiting requirements.
Enter fullscreen mode Exit fullscreen mode
Architecture Decision Framework
After all this data, we built a decision tree that our team uses for every new service.
Decision Tree
START: New service needs a compute platform
│
├── Is it event-driven? (S3 triggers, SQS, SNS, DynamoDB Streams)
│ └── YES → Lambda (no question — this is what it's built for)
│
├── Does it need WebSockets or long-running connections?
│ └── YES → Fargate (Lambda has 15-min timeout, no WebSocket support)
│
├── Does it need GPU/ML inference?
│ └── YES → Fargate (GPU task definitions available)
│
├── Is traffic predictable and steady (>3M requests/month)?
│ └── YES → Fargate (cheaper at scale, predictable billing)
│
├── Is traffic spiky/unpredictable (<3M requests/month avg)?
│ └── YES → Lambda (pay-per-use, instant scaling)
│
├── Is cold start latency critical? (p99 < 200ms required)
│ └── YES → Fargate (no cold starts)
│ └── MAYBE → Lambda + Provisioned Concurrency (adds cost)
│
├── Is this a quick prototype/experiment?
│ └── YES → Lambda (faster to deploy, no Docker needed)
│
├── Is operational simplicity the top priority?
│ └── YES → Lambda (no patching, no capacity planning)
│
└── Unsure / Mixed requirements?
└── Start with Lambda → migrate to Fargate if costs exceed
the crossover point for 3+ consecutive months
Enter fullscreen mode Exit fullscreen mode
When to Choose Lambda
Lambda Wins When:
✓ Traffic is under 3M requests/month
✓ Traffic is bursty or unpredictable
✓ Workload is event-driven (queues, streams, schedules)
✓ Execution time is under 15 minutes
✓ You need instant scaling (0 to 1000 concurrent in seconds)
✓ Team is small and doesn't want to manage containers
✓ Service is a prototype or experiment
✓ Workload is CPU-light (I/O bound, API proxying)
Lambda Loses When:
✗ Steady high traffic (>3M requests/month)
✗ Cold start latency is unacceptable
✗ Execution exceeds 15 minutes
✗ WebSocket or persistent connections needed
✗ Large deployment package (>250MB unzipped)
✗ GPU or specialized hardware needed
Enter fullscreen mode Exit fullscreen mode
When to Choose Fargate
Fargate Wins When:
✓ Traffic is steady and above 3M requests/month
✓ You need predictable billing
✓ WebSockets or long-running processes required
✓ Cold start latency must be zero
✓ Application has complex startup (ML model loading, warm caches)
✓ You need more than 10GB memory or 6 vCPU per unit
✓ Background workers or daemon processes
✓ Team already has Docker/container expertise
Fargate Loses When:
✗ Traffic is near-zero for extended periods
✗ Workload is purely event-driven
✗ Rapid iteration matters more than optimization
✗ Team lacks container expertise and doesn't want to learn
✗ Minimum monthly cost floor (~$100) is too high
Enter fullscreen mode Exit fullscreen mode
The Hybrid Approach: Best of Both Worlds
After analyzing 3 months of data, we didn’t choose Lambda or Fargate. We chose both — each for what it does best.
Hybrid Architecture
┌─────────────────────────────────────────────────┐
│ Route 53 (DNS) │
└────────────┬─────────────────┬──────────────────┘
│ │
┌────────────▼──────┐ ┌──────▼────────────────┐
│ API Gateway │ │ ALB │
│ (HTTP API) │ │ (Application LB) │
└────────┬──────────┘ └───────┬───────────────┘
│ │
┌──────────────▼──────────────┐ ┌────▼──────────────────┐
│ Lambda Functions │ │ ECS Fargate │
│ │ │ │
│ ▪ Webhook receiver │ │ ▪ Core REST API │
│ POST /webhooks/* │ │ GET/POST /api/v1/* │
│ ~2M events/month │ │ ~25M requests/month│
│ │ │ │
│ ▪ Async processors │ │ ▪ WebSocket server │
│ SQS → Lambda │ │ wss://api/ws │
│ Image resize, PDF gen │ │ ~5K connections │
│ ~1M invocations/month │ │ │
│ │ │ ▪ Background workers │
│ ▪ Scheduled jobs │ │ Queue consumers │
│ EventBridge → Lambda │ │ Report generators │
│ Nightly reports, cleanup │ │ ML inference │
│ ~30K invocations/month │ │ │
└──────────────────────────────┘ └───────────────────────┘
│ │
┌────────▼───────────────────────────▼──────────┐
│ Shared Data Layer │
│ ▪ Aurora PostgreSQL (writer + 2 readers) │
│ ▪ ElastiCache Redis (3-node cluster) │
│ ▪ S3 (assets, uploads, backups) │
└───────────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode
What Runs Where and Why
Lambda_Workloads:
Webhooks:
description: "Receive webhooks from Stripe, GitHub, Twilio"
why_lambda: "Spiky, unpredictable, 0 traffic for hours then burst"
traffic: "~2M/month, 90% arrive in 10% of the time"
cost: "$180/month"
Async_Processing:
description: "Image resizing, PDF generation, email sending"
why_lambda: "Event-driven, triggered by SQS, no user waiting"
traffic: "~1M invocations/month"
cost: "$95/month"
Scheduled_Jobs:
description: "Nightly reports, data cleanup, cache warming"
why_lambda: "Runs once/day, 2-10 minutes, idle the rest"
traffic: "~30K invocations/month"
cost: "$8/month"
Fargate_Workloads:
Core_API:
description: "User-facing REST API, all CRUD operations"
why_fargate: "Steady 25M req/month, latency-sensitive, no cold starts"
traffic: "~25M requests/month, 300-800 RPS steady"
tasks: "4-12 (auto-scaling)"
cost: "$2,480/month"
WebSocket_Server:
description: "Real-time notifications, live dashboards"
why_fargate: "Long-lived connections, Lambda can't do WebSockets"
connections: "~5,000 concurrent"
tasks: "2-4"
cost: "$180/month"
Background_Workers:
description: "Queue consumers, ML inference, report generation"
why_fargate: "Long-running (>15 min), needs warm ML models"
tasks: "2 (always running)"
cost: "$120/month"
Enter fullscreen mode Exit fullscreen mode
Migration Steps
We migrated incrementally over 4 weeks. The key was moving one workload at a time and validating costs before moving the next.
#!/usr/bin/env python3
"""
Migration validator — runs after each workload migration
to compare pre/post costs and latency.
"""
import boto3
from datetime import datetime, timedelta
from dataclasses import dataclass
@dataclass
class MigrationCheck:
workload: str
pre_migration_cost: float
pre_migration_p99_ms: float
def get_cost_for_service(service_tag: str, days: int = 7) -> float:
"""Pull actual cost from Cost Explorer for a tagged service."""
ce = boto3.client("ce")
end = datetime.utcnow().strftime("%Y-%m-%d")
start = (datetime.utcnow() - timedelta(days=days)).strftime("%Y-%m-%d")
response = ce.get_cost_and_usage(
TimePeriod={"Start": start, "End": end},
Granularity="DAILY",
Metrics=["UnblendedCost"],
Filter={
"Tags": {
"Key": "Service",
"Values": [service_tag],
}
},
GroupBy=[{"Type": "DIMENSION", "Key": "SERVICE"}],
)
total = sum(
float(day["Total"]["UnblendedCost"]["Amount"])
for group in response["ResultsByTime"]
for day in [group]
)
return total * (30 / days) # Extrapolate to monthly
def get_p99_latency(log_group: str, hours: int = 24) -> float:
"""Query CloudWatch Logs Insights for p99 latency."""
logs = boto3.client("logs")
query = """
fields @timestamp, @duration
| stats pct(@duration, 99) as p99_ms
"""
response = logs.start_query(
logGroupName=log_group,
startTime=int((datetime.utcnow() - timedelta(hours=hours)).timestamp()),
endTime=int(datetime.utcnow().timestamp()),
queryString=query,
)
# Poll for results
import time
query_id = response["queryId"]
while True:
result = logs.get_query_results(queryId=query_id)
if result["status"] == "Complete":
break
time.sleep(1)
if result["results"]:
return float(result["results"][0][0]["value"])
return 0.0
def validate_migration(check: MigrationCheck) -> dict:
"""Compare pre and post migration metrics."""
post_cost = get_cost_for_service(check.workload)
post_latency = get_p99_latency(f"/ecs/{check.workload}")
cost_change = ((post_cost - check.pre_migration_cost)
/ check.pre_migration_cost * 100)
latency_change = ((post_latency - check.pre_migration_p99_ms)
/ check.pre_migration_p99_ms * 100)
result = {
"workload": check.workload,
"cost_before": f"${check.pre_migration_cost:,.2f}",
"cost_after": f"${post_cost:,.2f}",
"cost_change": f"{cost_change:+.1f}%",
"latency_before_ms": check.pre_migration_p99_ms,
"latency_after_ms": post_latency,
"latency_change": f"{latency_change:+.1f}%",
"status": "PASS" if cost_change < 0 and latency_change < 20 else "REVIEW",
}
print(f"n{'='*50}")
print(f"Migration Validation: {check.workload}")
print(f"{'='*50}")
for k, v in result.items():
print(f" {k}: {v}")
return result
if __name__ == "__main__":
# Validate each migrated workload
checks = [
MigrationCheck("core-api", pre_migration_cost=7560, pre_migration_p99_ms=450),
MigrationCheck("webhooks", pre_migration_cost=1200, pre_migration_p99_ms=200),
MigrationCheck("async-processors", pre_migration_cost=800, pre_migration_p99_ms=0),
]
for check in checks:
validate_migration(check)
Enter fullscreen mode Exit fullscreen mode
Code Examples
Lambda Function with Cold Start Optimization
# lambda/webhook_handler.py
"""
Webhook receiver — optimized for Lambda cold starts.
Uses module-level initialization for connection reuse
and lazy imports for fast startup.
"""
import json
import os
import logging
from typing import Any
logger = logging.getLogger()
logger.setLevel(logging.INFO)
# Module-level: initialized once per execution environment
# These persist across warm invocations
_db_pool = None
_redis_client = None
_secrets_cache = {}
def _get_db_pool():
"""Lazy-initialize database connection pool."""
global _db_pool
if _db_pool is None:
import psycopg2.pool
secret = _get_secret("db-credentials")
_db_pool = psycopg2.pool.SimpleConnectionPool(
minconn=1,
maxconn=5,
host=os.environ["DB_HOST"],
port=5432,
database=os.environ["DB_NAME"],
user=secret["username"],
password=secret["password"],
connect_timeout=5,
options="-c statement_timeout=10000",
)
return _db_pool
def _get_redis():
"""Lazy-initialize Redis client."""
global _redis_client
if _redis_client is None:
import redis
_redis_client = redis.Redis(
host=os.environ["REDIS_HOST"],
port=6379,
decode_responses=True,
socket_connect_timeout=3,
socket_timeout=3,
retry_on_timeout=True,
)
return _redis_client
def _get_secret(secret_name: str) -> dict:
"""Cached Secrets Manager lookup."""
if secret_name not in _secrets_cache:
import boto3
client = boto3.client("secretsmanager")
response = client.get_secret_value(SecretId=secret_name)
_secrets_cache[secret_name] = json.loads(response["SecretString"])
return _secrets_cache[secret_name]
def handler(event: dict, context: Any) -> dict:
"""
Main Lambda handler for webhook processing.
Receives events from API Gateway HTTP API.
"""
try:
# Parse webhook payload
body = json.loads(event.get("body", "{}"))
source = event.get("headers", {}).get("x-webhook-source", "unknown")
webhook_type = body.get("type", "unknown")
logger.info(f"Webhook received: source={source}, type={webhook_type}")
# Validate webhook signature
if not _validate_signature(event, source):
return _response(401, {"error": "Invalid signature"})
# Route to appropriate processor
processors = {
"payment.completed": _process_payment,
"user.created": _process_user_created,
"order.updated": _process_order_update,
}
processor = processors.get(webhook_type, _process_unknown)
result = processor(body)
# Cache recent webhook IDs for deduplication
webhook_id = body.get("id", "")
if webhook_id:
_get_redis().setex(f"webhook:seen:{webhook_id}", 86400, "1")
return _response(200, {"status": "processed", "result": result})
except json.JSONDecodeError:
return _response(400, {"error": "Invalid JSON"})
except Exception as e:
logger.exception(f"Webhook processing failed: {e}")
return _response(500, {"error": "Internal server error"})
def _validate_signature(event: dict, source: str) -> bool:
"""Validate webhook signature based on source."""
import hmac
import hashlib
headers = event.get("headers", {})
body = event.get("body", "")
secret = _get_secret(f"webhook-secret-{source}")
expected_sig = headers.get("x-webhook-signature", "")
computed = hmac.new(
secret["signing_key"].encode(),
body.encode(),
hashlib.sha256,
).hexdigest()
return hmac.compare_digest(computed, expected_sig)
def _process_payment(body: dict) -> dict:
"""Process payment webhook — insert into database."""
pool = _get_db_pool()
conn = pool.getconn()
try:
with conn.cursor() as cur:
cur.execute(
"""
INSERT INTO payment_events (event_id, amount, currency, status, metadata)
VALUES (%s, %s, %s, %s, %s)
ON CONFLICT (event_id) DO NOTHING
RETURNING id
""",
(
body["id"],
body["data"]["amount"],
body["data"]["currency"],
body["data"]["status"],
json.dumps(body["data"]),
),
)
conn.commit()
result = cur.fetchone()
return {"inserted": result is not None}
finally:
pool.putconn(conn)
def _process_user_created(body: dict) -> dict:
"""Process new user webhook."""
_get_redis().hset(
f"user:{body['data']['user_id']}",
mapping={"email": body["data"]["email"], "plan": body["data"]["plan"]},
)
return {"cached": True}
def _process_order_update(body: dict) -> dict:
"""Process order update webhook."""
pool = _get_db_pool()
conn = pool.getconn()
try:
with conn.cursor() as cur:
cur.execute(
"UPDATE orders SET status = %s, updated_at = NOW() WHERE order_id = %s",
(body["data"]["status"], body["data"]["order_id"]),
)
conn.commit()
return {"updated": cur.rowcount > 0}
finally:
pool.putconn(conn)
def _process_unknown(body: dict) -> dict:
"""Log and acknowledge unknown webhook types."""
logger.warning(f"Unknown webhook type: {body.get('type')}")
return {"acknowledged": True, "processed": False}
def _response(status_code: int, body: dict) -> dict:
return {
"statusCode": status_code,
"headers": {"Content-Type": "application/json"},
"body": json.dumps(body),
}
Enter fullscreen mode Exit fullscreen mode
Fargate Task Definition with Auto-Scaling
# fargate/task-definition.yml
AWSTemplateFormatVersion: '2010-09-09'
Description: Production Fargate service with cost-optimized auto-scaling
Parameters:
Environment:
Type: String
Default: production
MinTasks:
Type: Number
Default: 4
MaxTasks:
Type: Number
Default: 20
Resources:
TaskDefinition:
Type: AWS::ECS::TaskDefinition
Properties:
Family: !Sub core-api-${Environment}
Cpu: '512'
Memory: '1024'
NetworkMode: awsvpc
RequiresCompatibilities:
- FARGATE
RuntimePlatform:
CpuArchitecture: ARM64 # 20% cheaper than x86
OperatingSystemFamily: LINUX
ExecutionRoleArn: !GetAtt ExecutionRole.Arn
TaskRoleArn: !GetAtt TaskRole.Arn
ContainerDefinitions:
- Name: api
Image: !Sub ${AWS::AccountId}.dkr.ecr.${AWS::Region}.amazonaws.com/core-api:latest
Essential: true
PortMappings:
- ContainerPort: 8000
Protocol: tcp
Environment:
- Name: ENVIRONMENT
Value: !Ref Environment
- Name: GUNICORN_WORKERS
Value: '2'
- Name: GUNICORN_THREADS
Value: '4'
- Name: DB_HOST
Value: !ImportValue DatabaseEndpoint
- Name: REDIS_HOST
Value: !ImportValue RedisEndpoint
Secrets:
- Name: DB_PASSWORD
ValueFrom: !Sub arn:aws:secretsmanager:${AWS::Region}:${AWS::AccountId}:secret:db-password
LogConfiguration:
LogDriver: awslogs
Options:
awslogs-group: !Ref LogGroup
awslogs-region: !Ref AWS::Region
awslogs-stream-prefix: api
mode: non-blocking
max-buffer-size: 4m
HealthCheck:
Command:
- CMD-SHELL
- curl -sf http://localhost:8000/health/ready || exit 1
Interval: 10
Timeout: 5
Retries: 3
StartPeriod: 30
Service:
Type: AWS::ECS::Service
DependsOn: ALBListener
Properties:
Cluster: !Ref ECSCluster
TaskDefinition: !Ref TaskDefinition
DesiredCount: !Ref MinTasks
LaunchType: FARGATE
PlatformVersion: LATEST
NetworkConfiguration:
AwsvpcConfiguration:
AssignPublicIp: DISABLED
SecurityGroups:
- !Ref ServiceSG
Subnets:
- !ImportValue PrivateSubnet1
- !ImportValue PrivateSubnet2
LoadBalancers:
- ContainerName: api
ContainerPort: 8000
TargetGroupArn: !Ref TargetGroup
DeploymentConfiguration:
MinimumHealthyPercent: 100
MaximumPercent: 200
DeploymentCircuitBreaker:
Enable: true
Rollback: true
EnableExecuteCommand: true
# Auto-Scaling Configuration
ScalableTarget:
Type: AWS::ApplicationAutoScaling::ScalableTarget
Properties:
MaxCapacity: !Ref MaxTasks
MinCapacity: !Ref MinTasks
ResourceId: !Sub service/${ECSCluster}/${Service.Name}
ScalableDimension: ecs:service:DesiredCount
ServiceNamespace: ecs
# CPU-based scaling
CPUScalingPolicy:
Type: AWS::ApplicationAutoScaling::ScalingPolicy
Properties:
PolicyName: cpu-target-tracking
PolicyType: TargetTrackingScaling
ScalableTargetId: !Ref ScalableTarget
TargetTrackingScalingPolicyConfiguration:
TargetValue: 60.0
PredefinedMetricSpecification:
PredefinedMetricType: ECSServiceAverageCPUUtilization
ScaleOutCooldown: 60
ScaleInCooldown: 300
# Request-count-based scaling
RequestScalingPolicy:
Type: AWS::ApplicationAutoScaling::ScalingPolicy
Properties:
PolicyName: request-count-tracking
PolicyType: TargetTrackingScaling
ScalableTargetId: !Ref ScalableTarget
TargetTrackingScalingPolicyConfiguration:
TargetValue: 1000.0 # 1000 requests per task per minute
PredefinedMetricSpecification:
PredefinedMetricType: ALBRequestCountPerTarget
ResourceLabel: !Sub
- ${ALBFullName}/${TargetGroupFullName}
- ALBFullName: !GetAtt ALB.LoadBalancerFullName
TargetGroupFullName: !GetAtt TargetGroup.TargetGroupFullName
ScaleOutCooldown: 60
ScaleInCooldown: 300
# Scheduled scaling for known patterns
MorningScaleUp:
Type: AWS::ApplicationAutoScaling::ScalableTarget
Properties:
MaxCapacity: !Ref MaxTasks
MinCapacity: 8 # Pre-warm for business hours
ResourceId: !Sub service/${ECSCluster}/${Service.Name}
ScalableDimension: ecs:service:DesiredCount
ServiceNamespace: ecs
ScheduledActions:
- ScheduledActionName: morning-scale-up
Schedule: cron(45 7 ? * MON-FRI *)
ScalableTargetAction:
MinCapacity: 8
- ScheduledActionName: evening-scale-down
Schedule: cron(0 22 ? * MON-FRI *)
ScalableTargetAction:
MinCapacity: !Ref MinTasks
# Cost-saving: Use ARM64 Fargate Spot for non-critical tasks
LogGroup:
Type: AWS::Logs::LogGroup
Properties:
LogGroupName: !Sub /ecs/core-api-${Environment}
RetentionInDays: 14 # Don't pay for indefinite retention
Enter fullscreen mode Exit fullscreen mode
Infrastructure Cost Calculator
#!/usr/bin/env python3
"""
Complete infrastructure cost calculator.
Takes your traffic pattern and outputs recommended architecture + projected cost.
Usage:
python3 infra_cost_calculator.py
--workloads workloads.json
--output recommendation.json
"""
import json
import argparse
from dataclasses import dataclass, field, asdict
from enum import Enum
class ComputeType(Enum):
LAMBDA = "lambda"
FARGATE = "fargate"
HYBRID = "hybrid"
@dataclass
class WorkloadProfile:
name: str
avg_requests_per_month: int
peak_rps: int
avg_duration_ms: float
memory_mb: int
is_event_driven: bool = False
needs_websockets: bool = False
needs_gpu: bool = False
max_execution_minutes: float = 0.5
traffic_pattern: str = "steady" # steady, diurnal, spiky, event-driven
cold_start_tolerance_ms: float = 1000
@dataclass
class CostEstimate:
compute_type: str
monthly_cost: float
breakdown: dict = field(default_factory=dict)
reasoning: str = ""
LAMBDA_CROSSOVER_REQUESTS = 3_200_000
LAMBDA_PRICE_PER_GB_SECOND = 0.0000166667
LAMBDA_PRICE_PER_REQUEST = 0.0000002
LAMBDA_FREE_GB_SECONDS = 400_000
LAMBDA_FREE_REQUESTS = 1_000_000
API_GW_HTTP_PER_MILLION = 1.00
FARGATE_VCPU_PER_HOUR = 0.04048
FARGATE_MEM_PER_GB_HOUR = 0.004445
FARGATE_ARM_DISCOUNT = 0.20 # 20% cheaper on Graviton
ALB_HOURLY = 0.0225
ALB_LCU_HOURLY = 0.008
NAT_GW_HOURLY = 0.045
NAT_GW_PER_GB = 0.045
CW_LOG_PER_GB = 0.50
HOURS_PER_MONTH = 730
def estimate_lambda_cost(workload: WorkloadProfile) -> CostEstimate:
mem_gb = workload.memory_mb / 1024
duration_s = workload.avg_duration_ms / 1000
gb_seconds = workload.avg_requests_per_month * duration_s * mem_gb
billable_gbs = max(0, gb_seconds - LAMBDA_FREE_GB_SECONDS)
compute = billable_gbs * LAMBDA_PRICE_PER_GB_SECOND
billable_req = max(0, workload.avg_requests_per_month - LAMBDA_FREE_REQUESTS)
request_cost = billable_req * LAMBDA_PRICE_PER_REQUEST
api_gw = (workload.avg_requests_per_month / 1_000_000) * API_GW_HTTP_PER_MILLION
# Provisioned concurrency if cold start sensitive
prov_cost = 0
if workload.cold_start_tolerance_ms < 500:
prov_units = max(10, workload.peak_rps // 10)
prov_gbs = prov_units * mem_gb * HOURS_PER_MONTH * 3600
prov_cost = prov_gbs * 0.0000041667
log_gb = (workload.avg_requests_per_month / 1_000_000) * 2.5
log_cost = log_gb * CW_LOG_PER_GB
nat_cost = NAT_GW_HOURLY * HOURS_PER_MONTH + 20 * NAT_GW_PER_GB
subtotal = compute + request_cost + api_gw + prov_cost + log_cost + nat_cost
total = subtotal * 1.15 # 15% overhead for real-world extras
return CostEstimate(
compute_type="lambda",
monthly_cost=round(total, 2),
breakdown={
"compute": round(compute, 2),
"requests": round(request_cost, 2),
"api_gateway": round(api_gw, 2),
"provisioned_concurrency": round(prov_cost, 2),
"logging": round(log_cost, 2),
"nat_gateway": round(nat_cost, 2),
"overhead": round(total - subtotal, 2),
},
)
def estimate_fargate_cost(workload: WorkloadProfile) -> CostEstimate:
avg_rps = workload.avg_requests_per_month / (30 * 24 * 3600)
req_per_task_per_s = 400
min_tasks = 2
max_tasks = 20
required_tasks = max(min_tasks, min(max_tasks, int(avg_rps / req_per_task_per_s) + 1))
vcpu = 0.5
mem_gb = 1.0
# ARM64 pricing
vcpu_price = FARGATE_VCPU_PER_HOUR * (1 - FARGATE_ARM_DISCOUNT)
mem_price = FARGATE_MEM_PER_GB_HOUR * (1 - FARGATE_ARM_DISCOUNT)
vcpu_cost = required_tasks * vcpu * HOURS_PER_MONTH * vcpu_price
mem_cost = required_tasks * mem_gb * HOURS_PER_MONTH * mem_price
alb_fixed = ALB_HOURLY * HOURS_PER_MONTH
alb_lcu = max(1, avg_rps / 25) * HOURS_PER_MONTH * ALB_LCU_HOURLY
alb_cost = alb_fixed + alb_lcu
log_gb = (workload.avg_requests_per_month / 1_000_000) * 0.8
log_cost = log_gb * CW_LOG_PER_GB
nat_cost = NAT_GW_HOURLY * HOURS_PER_MONTH + 15 * NAT_GW_PER_GB
total = vcpu_cost + mem_cost + alb_cost + log_cost + nat_cost
return CostEstimate(
compute_type="fargate",
monthly_cost=round(total, 2),
breakdown={
"tasks": required_tasks,
"vcpu": round(vcpu_cost, 2),
"memory": round(mem_cost, 2),
"alb": round(alb_cost, 2),
"logging": round(log_cost, 2),
"nat_gateway": round(nat_cost, 2),
},
)
def recommend(workload: WorkloadProfile) -> CostEstimate:
"""Determine optimal compute type for a workload."""
# Hard constraints
if workload.needs_websockets or workload.needs_gpu:
estimate = estimate_fargate_cost(workload)
estimate.reasoning = "Fargate required: WebSockets/GPU not supported on Lambda"
return estimate
if workload.max_execution_minutes > 15:
estimate = estimate_fargate_cost(workload)
estimate.reasoning = "Fargate required: execution exceeds Lambda 15-min limit"
return estimate
if workload.is_event_driven and workload.avg_requests_per_month < 1_000_000:
estimate = estimate_lambda_cost(workload)
estimate.reasoning = "Lambda optimal: event-driven with low volume"
return estimate
# Cost comparison
lambda_est = estimate_lambda_cost(workload)
fargate_est = estimate_fargate_cost(workload)
if workload.traffic_pattern == "spiky":
lambda_est.monthly_cost *= 0.6 # Spiky = lots of idle time
if lambda_est.monthly_cost < fargate_est.monthly_cost:
lambda_est.reasoning = (
f"Lambda cheaper: ${lambda_est.monthly_cost:.0f} vs "
f"${fargate_est.monthly_cost:.0f}/month"
)
return lambda_est
else:
fargate_est.reasoning = (
f"Fargate cheaper: ${fargate_est.monthly_cost:.0f} vs "
f"${lambda_est.monthly_cost:.0f}/month"
)
return fargate_est
def main():
parser = argparse.ArgumentParser(description="Infrastructure Cost Calculator")
parser.add_argument("--workloads", required=True, help="Path to workloads JSON file")
parser.add_argument("--output", help="Output file for recommendations")
args = parser.parse_args()
with open(args.workloads) as f:
workloads_data = json.load(f)
results = []
total_cost = 0
print(f"n{'='*70}")
print(f" INFRASTRUCTURE COST RECOMMENDATION")
print(f"{'='*70}n")
for w in workloads_data["workloads"]:
workload = WorkloadProfile(**w)
rec = recommend(workload)
results.append({"workload": w["name"], **asdict(rec)})
total_cost += rec.monthly_cost
print(f" {workload.name}")
print(f" Recommendation: {rec.compute_type.upper()}")
print(f" Monthly cost: ${rec.monthly_cost:,.2f}")
print(f" Reasoning: {rec.reasoning}")
print()
print(f"{'─'*70}")
print(f" Total estimated monthly cost: ${total_cost:,.2f}")
print(f" Total estimated annual cost: ${total_cost * 12:,.2f}")
print(f"{'─'*70}n")
if args.output:
with open(args.output, "w") as f:
json.dump({"recommendations": results, "total_monthly": total_cost}, f, indent=2)
print(f" Recommendations saved to {args.output}n")
if __name__ == "__main__":
main()
Enter fullscreen mode Exit fullscreen mode
Example workloads file:
{
"workloads": [
{
"name": "core-api",
"avg_requests_per_month": 25000000,
"peak_rps": 2000,
"avg_duration_ms": 120,
"memory_mb": 1024,
"traffic_pattern": "diurnal",
"cold_start_tolerance_ms": 200
},
{
"name": "webhook-receiver",
"avg_requests_per_month": 2000000,
"peak_rps": 500,
"avg_duration_ms": 80,
"memory_mb": 512,
"is_event_driven": true,
"traffic_pattern": "spiky",
"cold_start_tolerance_ms": 2000
},
{
"name": "async-processors",
"avg_requests_per_month": 1000000,
"peak_rps": 100,
"avg_duration_ms": 2000,
"memory_mb": 2048,
"is_event_driven": true,
"traffic_pattern": "event-driven",
"cold_start_tolerance_ms": 5000
},
{
"name": "websocket-server",
"avg_requests_per_month": 5000000,
"peak_rps": 800,
"avg_duration_ms": 50,
"memory_mb": 1024,
"needs_websockets": true,
"traffic_pattern": "steady"
}
]
}
Enter fullscreen mode Exit fullscreen mode
Results: Before vs After
After completing the migration to the hybrid architecture, here are the final numbers:
Metric Lambda-Only Hybrid Change
────────────────────────────────────────────────────────────────────────
Monthly Cost (30M req) $12,400 $3,950 -68%
Monthly Cost (5M req) $1,847 $1,280 -31%
Monthly Cost (500K req) $87 $87 0%
p99 Latency (core API) 450ms 120ms -73%
p99 Latency (webhooks) 800ms 250ms -69%
Cold Start Impact 3% of requests 0% (core API) -100%
1.2% (webhooks)
Error Rate (scaling) 0.8% 0.1% -88%
Deployment Time 2 min 5 min +150%
Operational Complexity Low Medium +50%
CloudWatch Logs Cost $840/month $310/month -63%
NAT Gateway Cost $380/month $380/month 0%
API Gateway Cost $3,500/month $175/month -95%
Enter fullscreen mode Exit fullscreen mode
The big wins:
- 68% cost reduction at high traffic by moving the core API to Fargate
- 95% reduction in API Gateway costs by using ALB for the core API and HTTP APIs for Lambda
- 73% improvement in p99 latency because Fargate eliminates cold starts for the main request path
- 63% reduction in logging costs because Fargate generates structured, lower-volume logs
The trade-offs:
- Deployment complexity increased — we now manage both Lambda and Fargate deployments
- Operational overhead went up — two monitoring surfaces, two scaling configurations
- New team members need to understand both paradigms
ROI Analysis
Investment:
Engineering time (2 engineers × 4 weeks): $20,000
Load testing infrastructure (3 months): $2,400
Migration validation and testing: $3,000
Documentation and runbooks: $1,500
Training (team of 8): $2,000
────────
Total investment: $28,900
Returns (at 30M requests/month):
Monthly compute savings: $8,450
Monthly logging savings: $530
Monthly API Gateway savings: $3,325
Reduced error-driven support tickets: $200
────────
Total monthly savings: $12,505
Payback Period: 2.3 months
Annual Savings: $150,060
3-Year Savings: $450,180
Note: Savings scale with traffic. At 5M requests/month,
monthly savings are ~$567, with a 51-month payback period.
The hybrid approach is most valuable at scale.
Enter fullscreen mode Exit fullscreen mode
The critical insight: this analysis only makes sense if your traffic is consistently above the crossover point. If you’re running at 2M requests/month, the complexity of a hybrid architecture isn’t worth the marginal savings. Stay on Lambda until the numbers force the move.
Lessons Learned
What Surprised Us
API Gateway was the biggest Lambda cost, not compute. At 30M requests, API Gateway REST charged $3,500/month. The Lambda compute itself was under $60. We’d been optimizing the wrong thing for months.
NAT Gateway costs were identical. Both Lambda (in VPC) and Fargate need NAT Gateways for internet access. This was a wash at $380/month regardless of architecture. VPC endpoints for S3 and Secrets Manager were the real fix — saved $120/month on both platforms.
CloudWatch Logs for Lambda are surprisingly expensive. Every Lambda invocation generates START, END, and REPORT log lines automatically. At 30M invocations, that’s 90M log lines you can’t disable. Fargate lets you control log verbosity in your application code.
Provisioned concurrency is a money pit if misconfigured. Auto-scaling provisioned concurrency overshoot during traffic spikes cost us $800/month in unused capacity. If you need zero cold starts, Fargate is cheaper than Lambda with provisioned concurrency above 50 concurrent executions.
ARM64 Fargate (Graviton) saves 20% with no code changes. Our Python FastAPI app ran identically on ARM64 Fargate tasks. We just changed the runtime platform in the task definition and instantly saved 20% on compute.
Mistakes We Made
Running the experiment too long. We planned 3 months but could have reached statistically significant conclusions in 6 weeks. The extra 6 weeks cost us about $4,000 in parallel infrastructure.
Not accounting for the 1.35x overhead factor from day one. Our initial projections based on AWS pricing calculators were consistently 25-35% lower than actual bills. Retries, metrics, tracing, data transfer, and throttle handling all add up.
Ignoring HTTP APIs. We ran the entire experiment on REST APIs ($3.50/million) before discovering HTTP APIs ($1.00/million). This would have shifted the crossover point to about 4.8M requests/month instead of 3.2M.
Over-engineering the Lambda function. Our initial Lambda function loaded all dependencies at module level for maximum warm-start performance. This made cold starts 2 seconds slower. The fix was lazy loading — fast cold starts matter more than shaving 5ms off warm invocations.
What We’d Do Differently
Start hybrid from day one for any service expected to grow past 3M requests/month. The migration cost ($28,900) would have been avoided.
Use HTTP APIs instead of REST APIs for all Lambda endpoints. The 71% savings on API Gateway changes the economics significantly.
Set up cost anomaly detection before the experiment, not after. AWS Cost Anomaly Detection would have flagged the provisioned concurrency overshoot in week 2 instead of month 3.
Test with Fargate Spot for non-critical workloads. We never tried it, but Fargate Spot offers up to 70% savings on interruptible tasks like batch processing and queue workers.
Decision Checklist
Use this checklist when evaluating Lambda vs Fargate for a new service:
Step 1: Check Hard Constraints
□ Does the workload need WebSockets? → Fargate
□ Does execution exceed 15 minutes? → Fargate
□ Does it need GPU/specialized hardware? → Fargate
□ Is the deployment package >250MB? → Fargate
□ Is it purely event-driven (SQS/SNS/S3)? → Lambda
□ Is it a cron job (< 15 min execution)? → Lambda
Enter fullscreen mode Exit fullscreen mode
Step 2: Estimate Traffic
□ Current monthly requests: ___________
□ Projected 12-month requests: ___________
□ Traffic pattern: □ Steady □ Diurnal □ Spiky
□ Peak-to-average ratio: ___________
□ Average request duration (ms): ___________
Enter fullscreen mode Exit fullscreen mode
Step 3: Calculate Crossover
□ Run the cost calculator with your parameters
□ Crossover point for your workload: ___________ requests/month
□ Current traffic vs crossover: □ Below □ Above □ Close
□ 12-month projection vs crossover: □ Below □ Above □ Close
Enter fullscreen mode Exit fullscreen mode
Step 4: Account for Hidden Costs
□ NAT Gateway needed? □ Yes ($33-100/month per AZ)
□ API Gateway type (REST vs HTTP)? REST = 3.5× more expensive
□ Provisioned concurrency needed? □ Yes (add $100-1,300/month)
□ VPC endpoints configured? □ Yes (saves $50-150/month)
□ CloudWatch log retention set? □ Yes (14 days, not infinite)
□ Cross-AZ data transfer considered? □ Yes ($0.01/GB adds up)
Enter fullscreen mode Exit fullscreen mode
Step 5: Make the Call
□ Under crossover + spiky traffic → Lambda
□ Under crossover + steady traffic → Lambda (but monitor growth)
□ Above crossover + steady traffic → Fargate
□ Above crossover + mixed workloads → Hybrid
□ Close to crossover + growing → Start Lambda, plan Fargate migration
Enter fullscreen mode Exit fullscreen mode
Step 6: Set Up Guardrails
□ AWS Budgets alert at 80% of estimate
□ Cost Anomaly Detection enabled
□ Monthly cost review on calendar
□ Crossover re-evaluation quarterly
□ CloudWatch dashboard for cost metrics
Enter fullscreen mode Exit fullscreen mode
Conclusion
The serverless vs containers debate shouldn’t be a debate at all. It’s a math problem with clear inputs and a calculable answer.
Lambda wins when traffic is low, bursty, or event-driven. Fargate wins when traffic is high, steady, or requires long-running processes. The crossover for a typical REST API workload sits around 3.2 million requests per month — but your specific crossover depends on request duration, memory allocation, and which API Gateway type you use.
The hybrid approach gave us the best of both worlds: Lambda’s instant scaling and zero-idle-cost for webhooks and async work, combined with Fargate’s predictable pricing and zero-cold-start performance for our core API. The result was a 68% cost reduction at high traffic and a 73% improvement in p99 latency.
Three things to do right now:
- Run the cost calculator with your actual workload parameters — the crossover point is different for every application.
- Audit your hidden costs — NAT Gateway, API Gateway type, CloudWatch Logs retention, and provisioned concurrency are the four biggest surprises.
- Tag everything — you can’t optimize what you can’t measure. Tag every Lambda function and Fargate service with a cost-allocation tag.
Stop debating. Start measuring.