Framing a Cloud Migration: Business Case, PMO Governance and Wave Planning That Hold Up
Cloud

Framing a Cloud Migration: Business Case, PMO Governance and Wave Planning That Hold Up

September 20, 20266 min readMigration CloudBusiness CasePMO

How to frame a cloud migration programme so it survives execution: defensible business case, lightweight PMO governance, wave-based milestones and a risk register that actually predicts failures.

Cloud programmes rarely fail on technology

When a cloud migration programme derails, the post-mortem almost never points at Terraform, Kubernetes or the landing zone. It points at scope that was never bounded, a business case nobody owned after signature, a steering committee that turned into a status-reporting theatre, and a risk register that was updated the week before an audit.

Framing — the two to twelve weeks before the first workload moves — is where a migration programme is actually won or lost. It is also the phase organisations compress hardest, because it produces slides instead of running systems. That trade is a bad one: every ambiguity left in the framing phase gets paid back with interest during execution, usually in the form of scope creep, dual-run costs and stalled decommissioning.

Framing: turning intent into a measurable perimeter

"Migrate to the cloud" is not a perimeter. A usable framing output answers four questions precisely.

What exactly is in scope? An application inventory with, for each item: business owner, technical owner, criticality tier, data classification, runtime dependencies, upstream/downstream integrations, licensing constraints and current hosting cost. Discovery tooling (agentless network flow analysis, CMDB extraction, APM topology) gets you 70% of the way; the remaining 30% comes from interviews, and that is where the surprises live — the undocumented FTP job, the Oracle feature nobody can replace, the appliance whose vendor has no cloud version.

What is the target? Not just "AWS" or "OVHcloud", but the disposition per application. The 7R model (retire, retain, rehost, relocate, replatform, repurchase, refactor) is useful precisely because it forces a decision per app rather than a blanket strategy. Expect a mix: a portfolio that is 100% refactor is a fantasy budget, and one that is 100% rehost is a cost problem deferred by 18 months.

What is the sequencing constraint? Dependencies, freeze periods, contract end dates, datacentre exit dates, regulatory deadlines. The datacentre lease expiry is usually the only truly hard date in the programme — everything else negotiates.

What does "done" mean? Migration is not complete when the workload runs in the cloud. It is complete when the source is decommissioned, the licence is terminated, the runbook is updated and the on-call rotation has taken ownership. Define this at framing or you will carry zombie infrastructure for years.

A business case the CFO cannot dismantle

The weakest business cases compare a cloud invoice to a hardware invoice. That comparison loses, always, because it ignores everything that makes on-premises expensive and everything that makes cloud cheap beyond compute.

Build the model on total cost of ownership over a defined horizon — typically three to five years, aligned with your hardware refresh cycle — and include the cost lines below on both sides.

Cost lineOn-premisesCloud target
Compute & storageCAPEX amortised + refreshOPEX, elastic, commitment-discountable
DatacentreSpace, power, cooling, network transitIncluded in unit price
Software licensingPer-socket/per-core, often over-provisionedBYOL or managed-service pricing; can go up or down
Infrastructure opsInternal FTEs + hardware support contractsReduced infra ops, increased platform/FinOps skills
DR & backupSecond site, often idleCross-AZ/region, pay-per-use
Migration—One-off: tooling, integration, testing, dual run
Dual run—Both platforms paid simultaneously, per wave

Two lines deserve special attention because they are systematically underestimated. Dual run is the period where you pay for both the source and the target. Its cost is a direct function of migration velocity — halve your wave duration and you halve the dual-run bill. That is the strongest financial argument for industrialising migration factories rather than treating each app as a project. Egress and data transfer during the migration itself, and structurally afterwards for hybrid integrations, can be material for data-heavy workloads.

On the benefit side, resist vague "agility" claims. Quantify what you can defend: avoided hardware refresh, floor space returned, licence consolidation, reduced environment provisioning lead time (measured in days before, hours after), capacity headroom no longer pre-purchased. Put a confidence level on each benefit and let the sceptical numbers carry the case. A business case that is credible at 60% of its optimistic scenario survives contact with a CFO; one that needs 100% of its benefits does not.

PMO governance: three bodies, not ten

Over-governed programmes burn their best engineers in meetings. A workable structure has three decision layers, each with a distinct mandate.

  • Steering committee (monthly, executive): budget, scope arbitration, escalated risks, go/no-go on waves. Sponsor is a business executive, not the CTO alone — a purely IT-sponsored migration loses priority the first time a product roadmap conflicts.
  • Programme review (weekly, delivery): wave progress, blockers, resourcing, dependency conflicts. Chaired by the programme manager, attended by workstream leads.
  • Architecture & security board (bi-weekly or on-demand): target patterns, exceptions, landing zone changes, derogations. Its power is the right to say "no, use the standard pattern" — and the obligation to publish the pattern.

Underneath, a Cloud Centre of Excellence or platform team owns the reusable assets: landing zone, IaC modules, CI/CD pipelines, guardrails, FinOps tagging policy. The CCoE is not a governance body that reviews; it is a product team that ships. The distinction matters. When the CCoE only reviews, application teams route around it.

Write a RACI once, for the decisions that actually recur: choice of target disposition per app, acceptance of a cutover, approval of a security derogation, ownership of post-migration cost. Ambiguity on the last one is the single most common cause of FinOps failure: nobody is accountable for a bill that arrives after the project team has been dissolved.

Waves, milestones and entry/exit criteria

Sequence the portfolio into waves that each have a coherent theme — a shared dependency cluster, a common target pattern, one business domain. The first wave should be deliberately boring: low criticality, few dependencies, a disposition your teams already know. Its purpose is to validate the factory, not to prove ambition.

Every wave needs explicit entry and exit criteria, version-controlled alongside the rest of the programme artefacts:

wave: W03-retail-backoffice
window: 2026-03-02 .. 2026-04-24
applications: [ORD-142, INV-078, PRC-330]
disposition: replatform (RDS PostgreSQL + ECS Fargate)
entry_criteria:
  - dependency_map_validated: true
  - landing_zone_account_provisioned: true
  - iac_modules_available: [vpc, rds, ecs-service, alb]
  - runbook_drafted: true
  - rollback_plan_tested_in_preprod: true
  - business_freeze_window_confirmed: true
exit_criteria:
  - functional_tests_passed: 100%
  - perf_within_slo: p95 < 400ms
  - observability: dashboards + alerts wired to on-call
  - source_environment_decommissioned: true
  - licences_terminated: true
  - cost_tags_applied: [app, env, owner, cost-center]
hypercare: 10 business days

The two criteria teams skip under pressure — rollback_plan_tested_in_preprod and source_environment_decommissioned — are exactly the ones that determine whether the programme delivers its business case. Make them non-negotiable exit gates and report the decommissioning rate in every steering committee.

The risks that actually materialise

A risk register that lists "project delay" is decoration. Useful entries name a mechanism, an early-warning signal and an owner.

RiskEarly signalMitigation
Hidden dependencies discovered at cutoverDiscovery coverage below ~90% of network flowsExtend flow analysis; shadow-traffic testing before cutover
Dual run extends, budget erodesWave duration slipping past plan by >20%Hard decommissioning gate; reduce wave size rather than delay
Post-migration cost overrunUntagged resources; no committed-use coverageTagging policy enforced in IaC; FinOps review per wave
Key-person dependencyOne engineer on every critical pathPair on runbooks; codify patterns in modules
Business teams unavailable for testingUAT slots unconfirmed 3 weeks aheadContract test windows at framing, in the wave plan
Regulatory/sovereignty blockerData classification incompleteClassify before disposition; qualify sovereign options early

Steering with numbers, not narratives

Four indicators are enough to run a monthly steering committee: percentage of the portfolio migrated and decommissioned (two separate curves — the gap between them is your dual-run exposure), actual versus forecast cloud spend per wave, cutover incident rate during hypercare, and lead time per application from entry criteria to exit. Track velocity as a trend: if the factory is not getting faster by wave three, the problem is industrialisation, not effort.

The final governance point is temporal. A migration programme has an end date; cloud operations do not. Plan the transfer of ownership — cost, security, reliability — to permanent teams from the first wave, not as a closing formality. The programmes that deliver their business case are the ones where, six months after the last cutover, someone is still accountable for the bill.

← Back to blog