The Cloud Sprawl Problem
On average, organisations waste 32 % of their cloud spend (Gartner, 2024). Forgotten EC2 instances, unattached EBS volumes, month-old snapshots, data in S3 Standard that should have moved to Glacier after 30 days, and reservations that no longer match workloads after a refactor — these waste sources accumulate silently. FinOps applies to cloud spending the same rigour that DevOps applies to code: continuous measurement, team-level accountability, and systematic improvement cycles.
Step 0: Establish Visibility
You cannot optimise what you cannot measure. Before any action, activate these tools:
- AWS Cost Explorer: spend visualisation by service, tag, region, account. Free.
- Cost and Usage Report (CUR): raw export to S3, queryable via Athena. Hourly granularity per resource ID.
- AWS Cost Anomaly Detection: automatic alerts on abnormal deviations. Set a 10 % threshold with SNS → Slack notification.
- Resource tagging: without consistent tags (Environment, Team, Project, CostCenter), you will never know which team is generating which spend. Enforce tagging via AWS Config rules.
Lever 1: EC2 Right-Sizing
AWS Compute Optimizer analyses 14 days of CloudWatch metrics (CPU, memory via CloudWatch Agent, network, disk I/O) and recommends the optimal instance type. In every FinOps audit we run, we systematically find 30–50 % of EC2 instances significantly overprovisioned — often m5.xlarge running at 12 % CPU with 2 GB RAM used out of 16 GB available.
# Audit Compute Optimizer via CLI
aws compute-optimizer get-ec2-instance-recommendations --region eu-west-1 --filters name=Finding,values=Overprovisioned --query 'instanceRecommendations[].{Instance:instanceArn,CurrentType:currentInstanceType,RecommendedType:recommendationOptions[0].instanceType,Savings:recommendationOptions[0].estimatedMonthlySavings.value}' --output table
Golden rule: test right-sizing in non-production for a week, check metrics, then apply to production. Never right-size a database instance blindly.
Lever 2: Spot Instances and Savings Plans
| Option | Discount | Flexibility | Best for |
|---|---|---|---|
| Compute Savings Plans | up to 66 % | Full (instance, region, OS) | Changing workloads |
| EC2 Instance SP | up to 72 % | Family + region fixed | Stable workloads |
| Reserved Instances (1yr) | up to 40 % | Exact type fixed | Databases, very stable |
| Spot Instances | up to 90 % | Interruption possible | Batch, CI/CD workers |
Our recommended strategy: cover 60–70 % of base EC2 consumption with 1-year Compute Savings Plans, keep 20 % On-Demand for flexibility, and use Spot for the rest.
# Terraform — EKS nodegroup with Spot/On-Demand mix
resource "aws_eks_node_group" "workers" {
capacity_type = "SPOT"
instance_types = ["m6i.xlarge", "m6a.xlarge", "m5.xlarge", "m5a.xlarge"]
# Multiple types = better Spot availability
}
Lever 3: S3 Storage Optimisation
- S3 Intelligent-Tiering: automatically moves objects between Standard, IA, and Glacier tiers based on access patterns. ROI-positive from day 30 for unpredictably accessed data.
- Lifecycle policies: archive to Glacier Instant Retrieval after 30 days, to Glacier Deep Archive after 90 days. Reduces storage cost by 80 %.
- Multipart upload cleanup: incomplete uploads consume space silently. Set a rule to delete incomplete parts after 7 days.
Lever 4: Automated Orphan Resource Cleanup
Orphan resources (unattached EBS volumes, unassociated Elastic IPs, old snapshots, load balancers without targets) often represent 5–10 % of the AWS bill. Automate their detection and removal with a scheduled Lambda function.
import boto3
from datetime import datetime, timezone
def cleanup_orphaned_resources():
ec2 = boto3.client('ec2', region_name='eu-west-1')
# Unattached EBS volumes older than 7 days
volumes = ec2.describe_volumes(
Filters=[{'Name': 'status', 'Values': ['available']}]
)
for vol in volumes['Volumes']:
age = (datetime.now(timezone.utc) - vol['CreateTime']).days
tags = [t['Key'] for t in vol.get('Tags', [])]
if age > 7 and 'DoNotDelete' not in tags:
ec2.delete_volume(VolumeId=vol['VolumeId'])
# Unassociated Elastic IPs
for eip in ec2.describe_addresses()['Addresses']:
if 'AssociationId' not in eip:
ec2.release_address(AllocationId=eip['AllocationId'])
Client Results
- Fintech client (€500k/year AWS): 38 % saving over 6 months — EC2 right-sizing + Savings Plans
- E-commerce client (€1.2M/year AWS): 44 % saving — Spot Instances for workers + S3 Intelligent-Tiering
- SaaS startup (€80k/year AWS): 52 % saving — aggressive right-sizing + orphan resource cleanup
Conclusion
FinOps is not a one-time project — it is an ongoing practice that integrates into your DevOps rituals. Start by activating AWS Cost Explorer and Cost Anomaly Detection (both free), tag all your resources consistently, then iterate every sprint. A monthly 2-hour FinOps review can generate tens to hundreds of thousands of euros in annual savings. The key: make cloud spend visible at the team level, and give teams the tools to optimise it themselves.
