FinOps
How We Cut a SaaS Startup's AWS Bill by 40% in 7 Weeks
A Series B SaaS company came to us with a problem: their AWS bill had crept from $18k to $31k per month over 6 months, and they had no idea why. Seven weeks later, we got it down to $18.5k — a 40% reduction — with zero performance impact. Here's exactly what we did.
The Starting Point: $31,000/Month and Climbing
When we started the engagement, the client's AWS Cost Explorer looked like a hockey stick going the wrong direction. The company had scaled from 10 to 50 engineers in a year, and cloud costs had scaled faster than revenue.
The symptoms were familiar:
- Nobody could explain why the bill doubled in 6 months
- Cost Explorer showed everything, which meant it showed nothing useful
- Engineers spun up resources but never cleaned them up
- No tagging strategy, so costs couldn't be attributed to teams
- Everything ran on-demand — no savings plans, no reserved instances
Phase 1: The Audit (Week 1-2)
Before optimizing anything, we needed to understand where the money was going. We spent two weeks building a complete picture:
What we found:
| Category | Monthly Spend | % of Bill |
|---|---|---|
| EC2 (compute) | $14,200 | 46% |
| RDS (databases) | $6,800 | 22% |
| Data Transfer | $3,100 | 10% |
| S3 + EBS | $2,400 | 8% |
| NAT Gateways | $1,800 | 6% |
| Other | $2,700 | 8% |
The big insight: EC2 and RDS alone were 68% of the bill. If we could optimize compute and databases, everything else was rounding error.
Phase 2: The Quick Wins (Week 2-3)
We started with changes that required zero architectural changes — pure cleanup and configuration.
1. Zombie Resource Cleanup: $2,400/month saved
We found:
- 47 orphaned EBS volumes — databases that were deleted but volumes weren't
- 12 unused Elastic IPs — $50/month each for doing nothing
- 3 forgotten development environments — full copies of production, running 24/7
- 8 idle load balancers — pointing to nothing
Total cleanup: $2,400/month. No changes to production, no risk, pure waste elimination.
2. Dev/Staging Scheduling: $1,800/month saved
Development and staging environments don't need to run at 3 AM on weekends. We implemented:
# Lambda function triggered by EventBridge
# Stops dev instances at 8 PM, starts at 8 AM
# Weekends: stays off
def lambda_handler(event, context):
action = event.get('action')
instances = get_tagged_instances('Environment', ['dev', 'staging'])
if action == 'stop':
ec2.stop_instances(InstanceIds=instances)
elif action == 'start':
ec2.start_instances(InstanceIds=instances)
Running dev/staging 12 hours/day, 5 days/week instead of 24/7 cut those costs by 65%.
3. NAT Gateway Optimization: $900/month saved
They had NAT Gateways in every AZ (best practice for HA), but traffic was heavily skewed. 80% of traffic went through one AZ. We:
- Kept multi-AZ NAT for production
- Consolidated dev/staging to single NAT Gateway
- Set up S3 and DynamoDB VPC endpoints (free data transfer)
Phase 3: Right-Sizing (Week 3-5)
This is where the real savings came from. Most instances were over-provisioned because engineers picked sizes based on "what if" rather than actual usage.
The Right-Sizing Process
- Collect 2 weeks of CloudWatch metrics — CPU, memory (via CloudWatch agent), network, disk
- Identify over-provisioned instances — average utilization < 40% with no spikes above 70%
- Recommend new sizes — target 60-70% average utilization
- Test in staging first — run for a week, validate performance
- Rolling production changes — one service at a time, with rollback plan
What we changed:
| Service | Before | After | Monthly Savings |
|---|---|---|---|
| API servers (8x) | m5.2xlarge | m6g.xlarge | $1,920 |
| Worker nodes (6x) | c5.2xlarge | c6g.xlarge | $1,440 |
| Primary RDS | db.r5.2xlarge | db.r6g.xlarge | $840 |
| Redis | cache.r5.xlarge | cache.r6g.large | $320 |
Total right-sizing savings: $4,520/month
The Graviton Factor
Notice we moved from Intel (m5, c5, r5) to Graviton (m6g, c6g, r6g). Graviton instances are:
- ~20% cheaper for equivalent performance
- Often 10-20% faster for the same price
- ARM-based, so Docker images need multi-arch builds
The migration was straightforward because they were already containerized. We updated their CI/CD to build multi-arch images:
# GitHub Actions multi-arch build
- name: Build and push
uses: docker/build-push-action@v5
with:
platforms: linux/amd64,linux/arm64
push: true
tags: ${{ env.ECR_REGISTRY }}/${{ env.IMAGE_NAME }}:${{ github.sha }}
Phase 4: Savings Plans (Week 5-6)
With right-sizing done, we had a stable baseline to commit against. The client had been running 100% on-demand — leaving 30-40% savings on the table.
Our recommendation:
- Compute Savings Plan (1-year, no upfront): $8,000/month commitment for ~$5,600 effective cost
- Covered ~70% of their compute (the predictable portion)
- Remaining 30% stayed on-demand for flexibility
Monthly savings from Savings Plans: $2,400
We recommended 1-year no-upfront plans because:
- 3-year plans save more but lock you in during rapid growth
- No-upfront preserves cash for a startup
- 1-year is long enough for meaningful savings, short enough to adjust
Phase 5: Ongoing Optimization (Week 6-7)
The final phase was setting up systems to prevent cost creep:
1. Tagging Policy
Every resource must have:
Environment: production, staging, developmentTeam: platform, product, dataService: api, worker, database
AWS Config rule blocks untagged resource creation.
2. Cost Anomaly Detection
AWS Cost Anomaly Detection alerts when daily spend jumps unexpectedly. Slack notification for anything >$100 above normal.
3. Monthly Cost Reviews
30-minute monthly review with engineering leads. Dashboard shows cost by team, trends, and recommendations.
The Results
| Optimization | Monthly Savings |
|---|---|
| Zombie cleanup | $2,400 |
| Dev/staging scheduling | $1,800 |
| NAT Gateway optimization | $900 |
| Right-sizing + Graviton | $4,520 |
| Savings Plans | $2,400 |
| Total | $12,020 |
Final bill: $18,980/month (down from $31,000)
Reduction: 39%
Annualized savings: $144,240
What We Didn't Do
Some common cost optimization tactics we intentionally skipped:
- Spot instances — Their workloads weren't spot-tolerant yet. Future optimization.
- Reserved Instances — Savings Plans are more flexible for a growing company.
- Architecture changes — Focused on quick wins first. Serverless migration is a future project.
Lessons Learned
- Start with visibility. You can't optimize what you can't measure. Tagging and cost allocation come first.
- Right-size before committing. Don't buy Savings Plans based on bloated infrastructure.
- Graviton is free money. If you're containerized, the migration is straightforward and the savings are real.
- Zombies multiply. Without cleanup automation, orphaned resources will come back.
- Make costs visible. When teams see their costs, they care about optimization.
Could We Do This for You?
If your AWS bill has been climbing and you're not sure why, we can help. Our free cloud audit includes a cost analysis that identifies quick wins and estimates potential savings.
Most clients see 25-45% savings. The engagement often pays for itself in the first month.