Cloud migrations rarely fail for technical reasons
Landing zones get built. Terraform modules get written. Workloads do move. And yet a large share of migration programmes still end up over budget, behind schedule, or quietly abandoned halfway through — with two production estates to run instead of one. The failure mode is almost never "we couldn't make it work on AWS/Azure/GCP/OVHcloud". It is scope that was never framed, a business case nobody owned, a governance model that couldn't arbitrate, and application teams that were informed rather than engaged.
That is precisely the territory of business analysis (what French practice calls AMOA), the PMO, and change management. This article lays out how we structure and run cloud migration programmes: framing, business case, roles, governance cadence, the frameworks worth borrowing from (PMP, SAFe, ITIL 4), and the metrics that actually tell you whether the programme is on track.
Framing: turning intent into a defensible scope
"We're going to the cloud" is a direction, not a scope. The framing phase — typically four to eight weeks for a mid-sized estate — has one job: convert that direction into a set of decisions everyone can be held to. Four deliverables matter.
1. Application inventory and dependency map. Not a spreadsheet of VMs, an inventory of applications with their owners, criticality, data classification, technical debt, contractual constraints, and — most importantly — their inbound and outbound dependencies. Discovery tooling (Azure Migrate, AWS Application Discovery Service, CMDB extraction, network flow analysis) gets you 70% of the way; interviews with the people who actually operate the app get you the rest.
2. Disposition per application (the 7 Rs). Retire, retain, rehost, replatform, repurchase, refactor, relocate. This decision must be traceable and scored, not the result of a corridor conversation. A simple weighted model works well:
applications:
- name: billing-core
owner: finance-it
criticality: tier-1
scores: # 1 (low) to 5 (high)
business_value: 5
change_appetite: 2 # team's capacity to absorb refactoring
technical_debt: 4
cloud_readiness: 2 # stateless? containerisable? licensing?
regulatory_weight: 4
disposition: replatform
wave: 3
rationale: >
Tier-1 revenue path, low change appetite this year.
Move to managed PostgreSQL + containers, defer refactor to FY+1.
3. Target architecture and landing zone principles. Multi-account/subscription structure, network topology, identity model, security guardrails, tagging taxonomy, observability standard. Decide once, centrally, then let teams consume it — every deviation negotiated later costs a multiple of what it costs to fix now.
4. Wave plan. Sequence by dependency cluster and risk, not by alphabetical order or by which team shouts loudest. Wave 0 should be a genuinely representative pilot: an application with real users, real data, real on-call — not the internal wiki.
The business case: escaping the "lift and shift is cheaper" trap
The most damaging business case is the one that promises infrastructure savings from a pure rehost. Moving a fleet of over-provisioned VMs to on-demand instances usually costs more in month one, and the delta only closes with rightsizing, commitment coverage, scheduling of non-production, and storage tiering — all of which are FinOps work that has to be explicitly staffed and planned.
A credible business case has four value pools, and you should model each separately:
| Value pool | What it contains | When it materialises | Confidence |
|---|---|---|---|
| Infrastructure run cost | Compute, storage, network, licences, datacentre, hardware refresh avoidance | After rightsizing + commitments, typically 6–12 months post-migration | High — measurable |
| Operational efficiency | Reduced patching, backup, hardware ops; managed services replacing self-hosted middleware | Progressive, tied to decommissioning | Medium |
| Delivery velocity | Environment provisioning time, deployment frequency, lead time for change | Requires platform + CI/CD investment, 12+ months | Medium — track via DORA metrics |
| Business capability | Elasticity for seasonal peaks, data/AI platform access, faster market entry | Long-term, business-owned | Low — model as scenarios, not a single number |
Two disciplines keep the business case honest. First, model the double-run cost: the period where you pay for both estates. It is the single largest cost line in most migration programmes and the reason decommissioning must be a tracked deliverable with a date, not an aspiration. Second, treat the exit from the datacentre as a hard milestone with a contractual date — a migration without a decommissioning deadline drifts indefinitely.
Express the outcome as a cash-flow curve over three to five years, with a break-even point, not as a single percentage. And be explicit about the cost of the programme itself: cloud engineering, the PMO, business analysis, training, licence renegotiation, and the productivity dip during the transition.
AMOA, PMO, Product Owner: who does what
The most common organisational failure is conflating these roles, or assigning them all to an overloaded infrastructure manager.
| Role | Owns | Key deliverables | Anti-pattern |
|---|---|---|---|
| Business analysis / AMOA | The why and the what: business requirements, disposition decisions, acceptance criteria | Application inventory, requirement specs, UAT strategy, business readiness | A pure documentation function with no decision authority |
| Technical delivery / AMOE | The how: landing zone, IaC, migration factory, cutover execution | Architecture decisions, runbooks, automated pipelines | Deciding scope unilaterally because "business isn't available" |
| PMO | The when and the how much: planning, risk, budget, dependencies, reporting | Wave plan, RAID log, budget tracking, steering pack | Status-report factory that surfaces problems after they've happened |
| Platform Product Owner | The landing zone as an internal product with real users | Platform backlog, golden paths, adoption metrics | Building a platform nobody asked for and nobody adopts |
The PMO in a cloud programme is not administrative. Its real value lies in dependency management — the cross-application, cross-team, cross-vendor sequencing constraints that no single squad can see — and in making risk visible early enough to be acted on. A PMO that only reports RAG status is overhead; a PMO that unblocks the Wave 3 network dependency six weeks before it bites is the highest-leverage role in the programme.
Governance: fewer meetings, sharper decisions
Governance exists to make decisions at the right level, at the right speed. Three tiers usually suffice:
- Steering committee (monthly) — sponsor plus business and IT executives. Agenda: budget versus business case, milestone and decommissioning trajectory, top risks, and arbitrations only. If the steerco is being briefed rather than deciding, it is theatre.
- Programme committee (weekly) — PMO, lead architect, business analysis lead, wave leads. Agenda: wave progress, cross-team blockers, scope change requests, RAID review.
- Migration factory standup (daily during cutover windows) — operational, 15 minutes, blockers only.
Add two artefacts that punch well above their weight. An architecture decision record (ADR) repository in Git, versioned alongside the IaC, so that six months later nobody re-litigates why you chose a transit gateway topology or a specific database engine. And a change-request process with a cost tag: every scope change is priced in euros and days before it is accepted. Nothing disciplines scope creep faster than a visible price.
PMP, SAFe, ITIL 4, PRINCE2: borrow, don't import wholesale
No single framework fits a cloud migration end to end. Migrations are a hybrid: predictive at the programme level (dependencies, cutover dates, contractual deadlines are not negotiable with a backlog), adaptive at the delivery level (each application reveals surprises).
| Framework | What to take | Where it hurts |
|---|---|---|
| PMP / PMBOK | Rigorous scope, risk, stakeholder and procurement management; earned value for long predictable waves | Heavyweight change control slows down teams that need to iterate weekly |
| PRINCE2 | Clear business case ownership, stage gates, tolerance-based escalation | Document-driven ceremony if applied literally |
| SAFe | PI planning to synchronise many teams, cross-team dependency visualisation, a dedicated platform/enabler stream | Full SAFe on a 4-team programme is pure overhead; take the planning cadence, skip the org chart |
| ITIL 4 | Service transition, change enablement, service continuity — essential for the day-2 operating model and the run/build handover | Applied rigidly, its change advisory board becomes the bottleneck a cloud platform was meant to remove |
The pragmatic pattern: PRINCE2-style stage gates at wave boundaries, a SAFe-style quarterly planning event to align application squads and the platform team, ITIL 4 for change enablement and the target service model, and PMP discipline on risk and procurement. Certifications matter less than whether the people running the programme can tell you why they chose each mechanism.
Change management is the deliverable, not the communication plan
Migrating an application without migrating the operating model produces a virtual machine in someone else's datacentre and a larger bill. Change management here means three concrete workstreams.
Skills. Map the target operating model role by role: who handles incidents on a managed database, who owns IAM policy reviews, who approves a Terraform module change. Then build the training and certification path against that map, not against a generic catalogue. Run a genuine internal community of practice — a weekly clinic where teams bring real problems beats a one-off training week.
Processes. Incident, change, capacity and continuity processes all shift. Backup becomes snapshot policy plus restore testing. Capacity planning becomes autoscaling policy plus commitment planning. Write the new runbooks during the migration, with the run team in the room — not after go-live.
Adoption. Measure it. Percentage of workloads deployed through the golden path, number of teams onboarded to the platform, tickets raised against self-service. A landing zone with a 20% adoption rate is a failed product, whatever its technical quality.
The metrics that actually steer
Replace the RAG-status pack with a small set of indicators that can trigger a decision:
- Migration velocity — applications or VMs migrated per wave versus plan, with a cumulative burn-up.
- Decommissioning rate — the only metric that proves value is being captured. Track the gap between migrated and decommissioned; it is your double-run exposure.
- Actual versus business-case spend, broken down by wave, with unit economics (cost per application, per environment).
- Post-migration stability — incidents in the first 30 days after cutover per application; it tells you whether your pre-migration testing is good enough before wave N+1 amplifies the problem.
- DORA metrics on migrated applications — deployment frequency and lead time, to evidence the velocity value pool.
- Platform adoption — golden-path usage rate.
Five failure patterns worth naming
The pilot that proves nothing. A stateless internal tool with no users validates neither your cutover process nor your run model. Pick a pilot that hurts a little.
Decommissioning as a later problem. Without a contractual datacentre exit date and a named owner per decommissioning, the old estate survives for years and the business case evaporates.
FinOps bolted on at the end. Tagging taxonomy, budget alerts, showback and commitment strategy belong in the landing zone design, not in a post-migration cost-reduction project.
A platform team with no product mandate. If application teams can bypass the golden path with no consequence and no dialogue, you will end up with a central platform and a shadow estate.
Governance that reports instead of arbitrating. If no decision was taken in the last steering committee, cancel the next one and fix the agenda.
Cloud migration is an organisational transformation with a technical component, not the other way round. The engineering is broadly a solved problem; the framing, the business case discipline, the governance cadence and the change work are where programmes are won or lost.
