AWS Well-Architected Framework: The 6 Pillars Applied in Practice
Cloud

AWS Well-Architected Framework: The 6 Pillars Applied in Practice

May 5, 202511 min readAWSWell-ArchitectedArchitecture

The AWS Well-Architected Framework is more than a checklist — it is a methodology for building cloud systems that are secure, reliable, and cost-efficient from day one.

What Is the Well-Architected Framework?

Created by AWS in 2015, the Well-Architected Framework (WAF) is a structured set of best practices for designing cloud architectures that are secure, performant, resilient, efficient, and sustainable. At its creation it rested on 5 pillars; in 2021, AWS added a sixth: Sustainability, reflecting the growing awareness of the environmental impact of digital infrastructure.

Since its launch, more than 400,000 Well-Architected Reviews (WARs) have been conducted worldwide, identifying thousands of architectural risks before they become production incidents. The WAF is today the industry reference for assessing cloud architecture maturity.

Pillar 1 — Operational Excellence

Operational excellence is the ability to run and monitor systems to deliver business value, and to continuously improve processes and procedures. In 2025, this translates into four foundational practices.

Mandatory Infrastructure as Code

Every cloud resource must be provisioned via IaC (Terraform, AWS CDK, CloudFormation). No exceptions: resources created manually via the console are a source of configuration drift and an obstacle to security reviews. IaC code lives in Git, is reviewed via Pull Requests, and triggers CI/CD pipelines.

Runbooks for Every Incident Type

A runbook is a step-by-step procedure for responding to a specific type of incident (database unavailability, disk saturation, unexpected traffic spike). They are stored in a versioned wiki or documentation system and regularly tested during Game Days.

Blameless Post-Mortems

After every major incident (SEV1 or SEV2), a post-mortem is written within 48 hours. The objective is not to find a culprit but to identify systemic root causes and define corrective actions. This blameless post-mortem culture, popularised by Google SRE, is one of the most differentiating factors between mature and immature technology organisations.

Measure with DORA Metrics

DORA (DevOps Research and Assessment) metrics are the reference indicators for measuring engineering team performance:

  • Deployment Frequency: how often you deploy to production (elite target: multiple times per day)
  • Lead Time for Changes: time from commit to production deployment (elite target: under one hour)
  • Mean Time to Recovery (MTTR): average time to restore service after an incident (elite target: under one hour)
  • Change Failure Rate: percentage of deployments causing an incident (elite target: under 5%)

Pillar 2 — Security

Security in the cloud is a shared responsibility: AWS secures the physical infrastructure, you secure what you deploy on it. This pillar covers identity, detection, data protection, and incident response.

IAM: Principle of Least Privilege

No IAM policy in production should contain a wildcard * on actions or resources. Each IAM role is scoped precisely to the actions the service needs to perform, and nothing more. Human access uses IAM Identity Center (SSO) with mandatory MFA.

GuardDuty + Security Hub: Automated Detection

# Enable GuardDuty and Security Hub via Terraform
resource "aws_guardduty_detector" "main" {
  enable = true

  datasources {
    s3_logs { enable = true }
    kubernetes { audit_logs { enable = true } }
    malware_protection {
      scan_ec2_instance_with_findings {
        ebs_volumes { enable = true }
      }
    }
  }
}

resource "aws_securityhub_account" "main" {}

resource "aws_securityhub_standards_subscription" "cis" {
  standards_arn = "arn:aws:securityhub:::ruleset/cis-aws-foundations-benchmark/v/1.4.0"
  depends_on    = [aws_securityhub_account.main]
}

resource "aws_securityhub_standards_subscription" "aws_foundational" {
  standards_arn = "arn:aws:securityhub:eu-west-1::standards/aws-foundational-security-best-practices/v/1.0.0"
  depends_on    = [aws_securityhub_account.main]
}

CloudTrail, Config, and VPC Flow Logs

CloudTrail must be enabled in all regions (not just the primary region) and logs must be sent to an S3 bucket in a dedicated Log Archive account — inaccessible to application teams. AWS Config records the history of all resource configuration changes and detects drift. VPC Flow Logs capture all network flows and are essential for security investigations.

Pillar 3 — Reliability

Reliability is the ability to recover from disruptions and dynamically acquire resources to meet demand. Reliable architectures are designed to fail gracefully.

Multi-AZ and Multi-Region

Every critical production application must be deployed across at least 2 availability zones. For applications with near-zero RTO (financial services, e-commerce), an active-active or active-passive multi-region architecture is necessary.

Circuit Breaker with API Gateway

# Circuit Breaker: strict timeout on API Gateway integrations
resource "aws_api_gateway_integration" "backend" {
  rest_api_id             = aws_api_gateway_rest_api.main.id
  resource_id             = aws_api_gateway_resource.proxy.id
  http_method             = "ANY"
  type                    = "HTTP_PROXY"
  integration_http_method = "ANY"
  uri                     = "http://${aws_lb.backend.dns_name}/{proxy}"

  # Strict timeout: prevents backend slowness from propagating
  timeout_milliseconds = 3000
}

# AWS Backup for controlled RTO/RPO
resource "aws_backup_plan" "daily" {
  name = "daily-backup-plan"

  rule {
    rule_name         = "daily-backup"
    target_vault_name = aws_backup_vault.main.name
    schedule          = "cron(0 1 * * ? *)"  # 1am every day

    lifecycle {
      delete_after = 30  # 30-day retention
    }
  }
}

Chaos Engineering with AWS Fault Injection Service

AWS Fault Injection Service (FIS) enables you to inject controlled failures into your infrastructure (EC2 instance termination, artificial network latency, API errors) to test system resilience. Schedule monthly Game Days to validate that your runbooks work under real-world conditions.

Pillar 4 — Performance Efficiency

This pillar concerns the ability to use computing resources efficiently to meet system requirements, and to maintain this efficiency as demand and technologies evolve.

The Right Service for the Right Task

  • AWS Lambda: event-driven processing, short functions (< 15 min), no server management
  • Amazon ECS Fargate: long-running services, containers, without EC2 management
  • Amazon EC2 Graviton: CPU-intensive workloads, self-managed databases
  • AWS Batch: large-scale batch jobs, deferred processing
  • Amazon SageMaker: ML model training and inference

Multi-Layer Caching

A performant architecture uses multiple cache levels to minimise calls to costly services:

  • CloudFront: edge cache for static assets and public APIs — reduces latency to < 10 ms for users near a PoP
  • ElastiCache (Redis): session cache, frequent query results, real-time counters
  • DAX (DynamoDB Accelerator): in-memory cache for DynamoDB, reduces latency from milliseconds to microseconds

Pillar 5 — Cost Optimisation

Cost optimisation is not a one-time activity but a continuous process. This pillar covers visibility, governance, and active optimisation of cloud spending.

Mandatory Tagging Strategy

# Mandatory tagging policy via AWS Config Rule
resource "aws_config_config_rule" "required_tags" {
  name = "required-tags"

  source {
    owner             = "AWS"
    source_identifier = "REQUIRED_TAGS"
  }

  input_parameters = jsonencode({
    tag1Key   = "Environment"
    tag1Value = "production,staging,development"
    tag2Key   = "Team"
    tag3Key   = "CostCenter"
    tag4Key   = "Application"
  })
}

AWS Budgets: Alerts at 80% and 100%

resource "aws_budgets_budget" "monthly" {
  name         = "monthly-budget"
  budget_type  = "COST"
  limit_amount = "5000"
  limit_unit   = "USD"
  time_unit    = "MONTHLY"

  notification {
    comparison_operator        = "GREATER_THAN"
    threshold                  = 80
    threshold_type             = "PERCENTAGE"
    notification_type          = "ACTUAL"
    subscriber_email_addresses = ["finops@company.com"]
  }

  notification {
    comparison_operator        = "GREATER_THAN"
    threshold                  = 100
    threshold_type             = "PERCENTAGE"
    notification_type          = "FORECASTED"
    subscriber_email_addresses = ["finops@company.com", "cto@company.com"]
  }
}

Savings Plans + Spot Strategy

For predictable workloads (base traffic of a web application), commit to Compute Savings Plans (1 year: -30%, 3 years: -50%). For interruption-tolerant workloads (ML jobs, video rendering, integration tests), use Spot Instances with a multi-type multi-AZ strategy to maximise availability.

Pillar 6 — Sustainability

Introduced in 2021, this pillar reflects the environmental responsibility of cloud architects. Reducing the carbon footprint of a cloud architecture is both a growing obligation (CSRD) and a cost optimisation opportunity.

Low-Carbon Regions

Prioritise eu-north-1 (Stockholm, ~12 gCO2eq/kWh via hydropower) and eu-west-1 (Ireland, ~240 gCO2eq/kWh via wind) for your batch workloads and archives. Avoid us-east-1 (~415 gCO2eq/kWh) for workloads not sensitive to localisation.

Graviton3 and Customer Carbon Footprint Tool

Graviton3 instances deliver 60% better performance per watt than equivalent x86 instances. Combine this migration with the AWS Customer Carbon Footprint Tool to measure your progress and feed your CSRD reporting.

How to Run a Well-Architected Review (WAR)

A WAR is a structured review of your architecture based on the Well-Architected Tool questions. It typically runs over 2 days with a certified AWS Partner:

  • Day 1: pillar-by-pillar workshops with technical teams — identification of High Risk Issues (HRI) and Medium Risk Issues (MRI)
  • Day 2: risk prioritisation, remediation plan construction, effort estimation and projected savings

The deliverable is a detailed report with all identified risks, prioritised and linked to concrete recommendations. Companies that follow WAR recommendations save an average of 30% on their AWS costs in the following 6 months.

Move2Cloud is an AWS Partner and offers free WARs for eligible architectures. Contact us to have your architecture assessed.

Conclusion

The Well-Architected Framework is not a state to be reached once — it is a process of continuous improvement. Start with the Security pillar (most immediate risk), then Reliability, then the other four according to your priorities. Schedule an annual WAR to adapt your architecture as your requirements and AWS services evolve. In 2025, the Sustainability pillar is emerging as a regulatory priority for all companies subject to CSRD.

← Back to blog