A review is not a questionnaire
Most failed Well-Architected Reviews share the same root cause: someone opened the AWS Well-Architected Tool, clicked through 60 questions in two hours with a single architect in the room, exported the PDF, and filed it. Six months later nothing has changed, and the exercise is remembered as compliance theatre.
A review that actually produces value is a facilitated audit: evidence-based, multi-role, time-boxed, and — critically — ending with a costed, owned, scheduled remediation plan. The framework's questions are a checklist for coverage, not the deliverable. What you ship is a decision document for the people who fund and operate the workload.
Scoping: one workload, not "the AWS account"
The unit of review is a workload: a set of components delivering business value, with an identifiable owner and a defined blast radius. "Our AWS estate" is not a workload. "The customer-facing order API and its data stores across prod and pre-prod" is.
Before booking any workshop, nail down:
- Boundaries: accounts, regions, VPCs, repositories, pipelines, third-party SaaS dependencies.
- Business context: criticality tier, RTO/RPO targets, regulatory scope (DORA, NIS2, PCI DSS, health data hosting), seasonality.
- Lens selection: the base Well-Architected lens, plus Serverless, SaaS, Data Analytics, Machine Learning or a custom lens encoding your own platform standards.
- Participants: the workload architect, an SRE/ops engineer, a developer who ships on it weekly, a security representative, and ideally the product owner. Missing the ops voice is the single fastest way to produce a fictional maturity score.
The workshop, step by step
A realistic timeline for a mid-sized workload is three weeks end to end. Compressing everything into a single day is possible but produces shallow findings.
| Phase | Timing | Who | Output |
|---|---|---|---|
| Kick-off & scoping | D-15, 1h | Architect, product owner, facilitator | Workload definition, lenses, participant list |
| Evidence collection | D-14 to D-3 | Facilitator + platform team | Diagrams, IaC repos, Trusted Advisor / Security Hub / Compute Optimizer exports, incident history, cost breakdown |
| Pre-fill | D-2, 2h | Facilitator | Draft answers for factual questions, list of open points |
| Workshop | D0, 2 × 3h | Full panel | Validated answers, identified risks, quick wins |
| Consolidation | D+3 | Facilitator | Report, risk register, remediation roadmap |
| Restitution | D+7, 1h30 | Panel + engineering management | Arbitrated, funded plan |
The pre-fill phase is what separates a good facilitator from a note-taker. Questions like "do you have automated deployments?" or "are your backups tested?" should arrive in the room already answered with evidence attached, so the group spends its time on the genuinely debatable items: trade-offs, deliberate technical debt, disagreement between dev and ops on what actually happens in production.
Running the six pillars without losing the room
Six pillars in six hours means roughly 50 minutes each — and they are not equally interesting to everyone. A few facilitation rules that consistently work:
- Start with Operational Excellence, not Security. It warms the room up on concrete topics (runbooks, deployments, on-call) and surfaces the real operating model, which then colours every other pillar.
- Ask for evidence, not opinions. "We monitor everything" becomes "show me the alert that fired during the last incident and who acknowledged it".
- Time-box debates to three minutes and park the rest. A parking lot with 15 items is a healthy sign, not a failure.
- Separate the risk from the fix. The workshop identifies gaps; solution design happens afterwards with the right people. Otherwise you spend 40 minutes arguing about Aurora vs. RDS Multi-AZ and never reach the Sustainability pillar.
Use the tool's risk classification honestly. A High Risk Issue (HRI) should mean "this can cause a serious outage, breach or cost overrun in the next 12 months". If everything is an HRI, nothing is.
The deliverables of an AWS architecture audit
The exported PDF from the Well-Architected Tool is a raw artefact, not a deliverable. Here is the package that actually gets acted on:
| Deliverable | Content | Audience |
|---|---|---|
| Executive summary (3–5 slides) | Maturity per pillar, top 5 risks in business terms, budget and effort required | CTO, engineering management |
| Target architecture diagram | Current state and target state, with deltas highlighted | Architects, tech leads |
| Risk register | One line per gap: pillar, question, description, impact, likelihood, severity, owner | Architecture & security governance |
| Remediation plan | Waves, dependencies, estimates, acceptance criteria per item | Delivery teams, PMO |
| Technical appendix | Tool exports, IaC extracts, configuration evidence, screenshots | Auditors, next reviewer |
| Milestone in the WA Tool | Frozen state of answers on review date | Everyone, for the next review |
Keep the risk register machine-readable. A spreadsheet or a YAML file in the workload's repository beats a paragraph in a Word document, because it can be diffed, imported into Jira, and re-evaluated at the next review.
- id: REL-03
pillar: Reliability
question: How do you back up data?
finding: RDS automated backups enabled (7 days) but restore never tested;
no cross-region copy for a workload with a 4h RTO.
severity: high
impact: RTO unachievable in a regional incident; contractual penalty risk.
remediation:
- Enable cross-region automated backups to eu-west-3
- Add a quarterly restore drill to the on-call runbook
- Alert on backup age > 26h via EventBridge + SNS
effort: M
owner: team-payments
wave: 1
acceptance: Documented restore drill executed, RTO measured < 4h
Turning gaps into a remediation plan that survives contact with the backlog
The classic failure mode is a list of 47 recommendations handed to a team that has no capacity. Prioritisation must be explicit and defensible. An impact/effort matrix works well, with three buckets:
- Wave 1 — 0 to 30 days: high-severity, low-effort items. Enabling MFA delete, restricting a security group open to 0.0.0.0/0, turning on GuardDuty in every region, setting a budget alarm. These fund the credibility of the whole exercise.
- Wave 2 — 1 to 3 months: structural items requiring design and testing. Multi-AZ, network segmentation, SLO-based alerting, replacing long-lived IAM users with OIDC-based federation in the CI pipeline.
- Wave 3 — 3 to 12 months: items requiring architectural change or budget arbitration. Migrating to a managed data store, splitting a monolith, adopting a multi-account landing zone.
Two non-negotiable rules. First, every item has a single named owner — a team, not a person who might leave. Second, every item has an acceptance criterion that is observable: "restore drill executed and documented", not "improve backups". If a remediation item cannot be verified, it will be declared done without being done.
Wherever possible, the remediation should land in Infrastructure as Code and, better still, in a policy that prevents regression. Fixing one open security group is a task; adding a Service Control Policy or an OPA/Conftest rule to the pipeline is a fix.
Tooling: automate the factual half
Roughly half of the Well-Architected questions can be answered from data rather than from memory. Pull the evidence before the workshop:
- AWS Trusted Advisor — natively integrated with the WA Tool, surfaces checks directly against relevant questions.
- AWS Security Hub with the AWS Foundational Security Best Practices and CIS standards for the Security pillar.
- AWS Compute Optimizer and Cost Explorer (rightsizing, Savings Plans coverage, untagged spend) for Cost Optimization.
- Prowler or Steampipe for a deeper open-source scan across accounts.
- AWS Resilience Hub to confront declared RTO/RPO with the actual architecture.
- Custom lenses to encode your own platform standards (mandatory tags, approved regions, logging baseline) and review them alongside the AWS questions.
The WA Tool API lets you script the boring parts — creating the workload, applying a review template, exporting the report into your documentation repository:
aws wellarchitected create-workload \
--workload-name "order-api-prod" \
--description "Customer-facing order API" \
--environment PRODUCTION \
--aws-regions eu-west-1 eu-west-3 \
--lenses wellarchitected serverless \
--review-owner "platform-team@example.com"
# List remaining improvement items after the workshop
aws wellarchitected list-lens-review-improvements \
--workload-id $WID --lens-alias wellarchitected \
--query 'ImprovementSummaries[?Risk==`HIGH`].[QuestionTitle,ImprovementPlanUrl]' \
--output table
# Archive the signed-off report
aws wellarchitected get-lens-review-report \
--workload-id $WID --lens-alias wellarchitected \
--query Base64String --output text | base64 -d > review-2025Q2.pdf
Making the review a cadence, not an event
The value of a Well-Architected Review compounds only if it repeats. Create a milestone in the tool at the end of each review: you then get a diff between two dates, which is far more persuasive to management than an absolute score. A workable cadence is a full review annually per critical workload, a lightweight re-scoring quarterly, and an ad-hoc review triggered by any major architectural change or serious incident.
Plug the remediation items into the normal delivery backlog with the same ticket format as features, and reserve a fixed capacity slice — 10 to 20 % of a sprint — for architecture debt. Items that live in a separate "audit plan" spreadsheet never get done.
Common failure modes
- Reviewing everything at once. Three workloads reviewed properly beat fifteen reviewed superficially.
- The self-congratulatory review. If the team facilitates its own review with no external challenge, expect a flattering score. Bring in someone from another team or an outside consultant.
- The report with no owner. A recommendation without a named team and a target date is a wish.
- Ignoring Sustainability and Operational Excellence. They tend to be rushed at the end of the day, yet they are where the cheapest wins usually hide.
- Confusing the tool score with reality. The number of HRIs is a conversation starter, not a KPI to optimise.
Done properly, a Well-Architected Review is one of the cheapest risk-reduction exercises available on AWS: a few days of facilitation, a shared vocabulary between dev, ops and security, and a prioritised plan that stops the same incident from happening a third time.
