FinOps

How We Cut a SaaS Startup's AWS Bill by 40% in 7 Weeks

· 12 min read

A Series B SaaS company came to us with a problem: their AWS bill had crept from $18k to $31k per month over 6 months, and they had no idea why. Seven weeks later, we got it down to $18.5k — a 40% reduction — with zero performance impact. Here's exactly what we did.

The Starting Point: $31,000/Month and Climbing

When we started the engagement, the client's AWS Cost Explorer looked like a hockey stick going the wrong direction. The company had scaled from 10 to 50 engineers in a year, and cloud costs had scaled faster than revenue.

The symptoms were familiar:

Phase 1: The Audit (Week 1-2)

Before optimizing anything, we needed to understand where the money was going. We spent two weeks building a complete picture:

What we found:

Category Monthly Spend % of Bill
EC2 (compute) $14,200 46%
RDS (databases) $6,800 22%
Data Transfer $3,100 10%
S3 + EBS $2,400 8%
NAT Gateways $1,800 6%
Other $2,700 8%

The big insight: EC2 and RDS alone were 68% of the bill. If we could optimize compute and databases, everything else was rounding error.

Phase 2: The Quick Wins (Week 2-3)

We started with changes that required zero architectural changes — pure cleanup and configuration.

1. Zombie Resource Cleanup: $2,400/month saved

We found:

Total cleanup: $2,400/month. No changes to production, no risk, pure waste elimination.

2. Dev/Staging Scheduling: $1,800/month saved

Development and staging environments don't need to run at 3 AM on weekends. We implemented:

# Lambda function triggered by EventBridge
# Stops dev instances at 8 PM, starts at 8 AM
# Weekends: stays off

def lambda_handler(event, context):
    action = event.get('action')
    instances = get_tagged_instances('Environment', ['dev', 'staging'])
    
    if action == 'stop':
        ec2.stop_instances(InstanceIds=instances)
    elif action == 'start':
        ec2.start_instances(InstanceIds=instances)

Running dev/staging 12 hours/day, 5 days/week instead of 24/7 cut those costs by 65%.

3. NAT Gateway Optimization: $900/month saved

They had NAT Gateways in every AZ (best practice for HA), but traffic was heavily skewed. 80% of traffic went through one AZ. We:

Phase 3: Right-Sizing (Week 3-5)

This is where the real savings came from. Most instances were over-provisioned because engineers picked sizes based on "what if" rather than actual usage.

The Right-Sizing Process

  1. Collect 2 weeks of CloudWatch metrics — CPU, memory (via CloudWatch agent), network, disk
  2. Identify over-provisioned instances — average utilization < 40% with no spikes above 70%
  3. Recommend new sizes — target 60-70% average utilization
  4. Test in staging first — run for a week, validate performance
  5. Rolling production changes — one service at a time, with rollback plan

What we changed:

Service Before After Monthly Savings
API servers (8x) m5.2xlarge m6g.xlarge $1,920
Worker nodes (6x) c5.2xlarge c6g.xlarge $1,440
Primary RDS db.r5.2xlarge db.r6g.xlarge $840
Redis cache.r5.xlarge cache.r6g.large $320

Total right-sizing savings: $4,520/month

The Graviton Factor

Notice we moved from Intel (m5, c5, r5) to Graviton (m6g, c6g, r6g). Graviton instances are:

The migration was straightforward because they were already containerized. We updated their CI/CD to build multi-arch images:

# GitHub Actions multi-arch build
- name: Build and push
  uses: docker/build-push-action@v5
  with:
    platforms: linux/amd64,linux/arm64
    push: true
    tags: ${{ env.ECR_REGISTRY }}/${{ env.IMAGE_NAME }}:${{ github.sha }}

Phase 4: Savings Plans (Week 5-6)

With right-sizing done, we had a stable baseline to commit against. The client had been running 100% on-demand — leaving 30-40% savings on the table.

Our recommendation:

Monthly savings from Savings Plans: $2,400

We recommended 1-year no-upfront plans because:

Phase 5: Ongoing Optimization (Week 6-7)

The final phase was setting up systems to prevent cost creep:

1. Tagging Policy

Every resource must have:

AWS Config rule blocks untagged resource creation.

2. Cost Anomaly Detection

AWS Cost Anomaly Detection alerts when daily spend jumps unexpectedly. Slack notification for anything >$100 above normal.

3. Monthly Cost Reviews

30-minute monthly review with engineering leads. Dashboard shows cost by team, trends, and recommendations.

The Results

Optimization Monthly Savings
Zombie cleanup $2,400
Dev/staging scheduling $1,800
NAT Gateway optimization $900
Right-sizing + Graviton $4,520
Savings Plans $2,400
Total $12,020

Final bill: $18,980/month (down from $31,000)

Reduction: 39%

Annualized savings: $144,240

What We Didn't Do

Some common cost optimization tactics we intentionally skipped:

Lessons Learned

  1. Start with visibility. You can't optimize what you can't measure. Tagging and cost allocation come first.
  2. Right-size before committing. Don't buy Savings Plans based on bloated infrastructure.
  3. Graviton is free money. If you're containerized, the migration is straightforward and the savings are real.
  4. Zombies multiply. Without cleanup automation, orphaned resources will come back.
  5. Make costs visible. When teams see their costs, they care about optimization.

Could We Do This for You?

If your AWS bill has been climbing and you're not sure why, we can help. Our free cloud audit includes a cost analysis that identifies quick wins and estimates potential savings.

Most clients see 25-45% savings. The engagement often pays for itself in the first month.

Ready to cut your cloud bill?
Start with a free cost audit.

We'll analyze your AWS/GCP/Azure spend and show you exactly where the savings are. No obligation.

Get Your Free Cost Audit