Cloud Migration TCO and ROI: Building a Business Case That Survives Year One
FinOps & Coûts

Cloud Migration TCO and ROI: Building a Business Case That Survives Year One

October 10, 20267 min readFinOpsTCOROI

Building a cloud migration business case that survives contact with reality: full on-premise baseline, honest target modelling, one-off costs, NPV/IRR, and post-migration measurement.

Almost every cloud migration starts with a slide that says "we will save 30%". Eighteen months later, the finance team is looking at a cloud bill that is higher than the old data centre, and nobody can explain where the business case went wrong. The problem is rarely the cloud itself — it is the model. A credible TCO and ROI analysis is not a spreadsheet exercise you do once to unlock the budget; it is an engineering artefact that you version, challenge and re-measure.

Why most cloud business cases are wrong

Three systematic errors show up again and again when we audit a migration business case:

  • An incomplete on-premise baseline. Hardware depreciation is counted, but not floor space, power, the storage team's time, the DR site that runs at 5% utilisation, or the licence true-ups.
  • A cloud target priced as a like-for-like lift-and-shift. You size the cloud VMs on the peak specs of physical servers bought five years ago, then wonder why the result is expensive. Of course it is: you are paying on-demand rates for a machine sized for 2019 peak load.
  • Migration costs treated as a rounding error. Project teams, dual-running both environments, data transfer, application remediation, training and the productivity dip during cutover can easily dominate year one.

A fourth, subtler error: counting only costs. A migration that costs the same but cuts lead time for change from six weeks to two days has an enormous ROI — it just does not appear in the infrastructure line.

Building a defensible on-premise baseline

The baseline must cover a full economic cycle, typically five years, because that is the refresh horizon of the hardware you are avoiding. Anything shorter flatters the status quo. Include:

CategoryWhat to includeCommon omission
Compute & storage hardwareAmortised purchase cost, maintenance contracts, spare capacityOver-provisioning headroom (often 50–70% idle)
Data centreRack space, power, cooling, physical securityInternal cross-charges that hide real cost
NetworkWAN links, load balancers, firewalls, MPLSLinks that disappear post-migration
Software licencesHypervisor, OS, databases, backup, monitoringPer-socket licensing that changes model in cloud
PeopleInfrastructure ops, patching, hardware incidents, capacity planningTime spent by app teams waiting for environments
Risk & continuityDR site, insurance, downtime costCost of an outage, expressed per hour

The one number worth fighting for is real utilisation. Pull it from vCenter, Prometheus, or your hypervisor API over at least 90 days — p95 CPU and memory per workload, not averages. This single dataset drives both the credibility of the baseline and the rightsizing of the target.

Modelling the target cloud cost

Model the target as a set of scenarios, not a single number. At minimum: rehost as-is, rehost with rightsizing and commitments, and replatform (managed databases, containers, serverless where it fits). Present all three; the comparison is what makes the recommendation credible.

Lines that are routinely forgotten in the cloud column:

  • Egress and inter-AZ traffic. Chatty microservices spread across availability zones generate real, recurring cost. Sovereign European providers often have more favourable egress terms than hyperscalers — worth modelling explicitly if you are comparing OVHcloud or Scaleway against AWS or Azure.
  • Observability. Log and metric ingestion priced per GB can become one of the largest line items on a chatty platform. Budget it from day one.
  • Backup, snapshots and cross-region replication. Cheap per GB, expensive at retention scale.
  • Support plans, usually a percentage of spend.
  • Managed service premium. A managed Postgres costs more than a VM running Postgres — and replaces part of a DBA's workload. Model both sides of that trade.
  • Non-production. Dev/test is where elasticity pays off the most; model scheduled shutdown and count the saving.

On the savings side, be explicit about which levers you are actually committing to: rightsizing from measured p95, Savings Plans or Reserved Instances on the stable base, Spot/preemptible for batch and CI, storage tiering, and shutting down non-production outside working hours. A business case that assumes commitments but does not name an owner for purchasing them is fiction.

Migration costs: the one-off column nobody likes

These are the costs that make year one negative and the payback period meaningful:

  • Discovery and dependency mapping
  • Application remediation and refactoring effort (by wave)
  • Dual-run: both environments live, sometimes for months
  • Data transfer and initial replication
  • Landing zone build: networking, IAM, IaC, CI/CD, observability
  • Training and certification, external consulting
  • Decommissioning and asset disposal — and the contractual exit cost of leases or hosting contracts

That last point is decisive. If your data centre contract has three years left and cannot be terminated, the savings do not start when you migrate, they start when the contract ends. Model it honestly; the migration may still be justified on agility grounds.

From TCO to ROI: the arithmetic that convinces a CFO

TCO is a cost comparison. ROI is an investment decision. Finance will expect net present value, internal rate of return, and payback period — discounting future cash flows at the company's cost of capital.

from itertools import count

# Illustrative only — replace with your own figures
DISCOUNT_RATE = 0.08

# Year 0 = migration investment, years 1-5 = net annual benefit
# (on-prem avoided cost - cloud run cost - residual contracts)
cash_flows = [-1_800_000, 250_000, 620_000, 760_000, 790_000, 810_000]

def npv(rate, flows):
    return sum(cf / (1 + rate) ** t for t, cf in enumerate(flows))

def irr(flows, lo=-0.9, hi=3.0, tol=1e-6):
    for _ in range(200):
        mid = (lo + hi) / 2
        if npv(mid, flows) > 0:
            lo = mid
        else:
            hi = mid
    return (lo + hi) / 2

def payback(flows):
    cumulative = 0.0
    for year, cf in enumerate(flows):
        previous = cumulative
        cumulative += cf
        if previous < 0 <= cumulative:
            return year - 1 + abs(previous) / cf
    return None

print(f"NPV  : {npv(DISCOUNT_RATE, cash_flows):,.0f}")
print(f"IRR  : {irr(cash_flows):.1%}")
print(f"Payback: {payback(cash_flows):.1f} years")

Run this as a sensitivity analysis, not a point estimate. Vary three parameters: achieved rightsizing (how much of the theoretical saving you actually capture), workload growth, and migration slippage. If NPV stays positive when you capture only half the rightsizing and the project runs six months late, you have a robust case. If it only works in the optimistic scenario, you have a sales pitch.

The value nobody puts in the spreadsheet

Infrastructure cost is the least interesting part of a cloud business case. The components that usually dominate NPV are harder to quantify but very real:

  • Time to market. Environment provisioning falling from weeks to minutes; measure it with DORA lead time before and after.
  • Avoided capex and capacity risk. No more buying for a peak that may never come.
  • Availability. Multi-AZ and automated failover against a single-site DR plan that has never been fully tested.
  • Team focus. Hours reallocated from firmware updates to product work.
  • Compliance posture. For regulated or sovereignty-sensitive data, a European or SecNumCloud-qualified provider can remove an entire risk and audit workload — or, conversely, add constraints you must price in.

Quantify these conservatively and state the assumption in plain language next to the number. A CFO will accept "we assume 20% of the three FTEs currently on hardware maintenance are redeployed" far more readily than an unexplained productivity multiplier.

Measuring the ROI after the migration

The business case is worthless if nobody checks it. Set up, from the first wave:

  • A tagging and showback policy enforced in IaC — application, environment, cost centre, owner. Untagged resources should fail the pipeline, not generate a quarterly email.
  • A unit economics metric — cost per order, per customer, per 1000 API calls. Absolute spend grows with the business; unit cost is the only honest efficiency signal.
  • A quarterly variance review comparing actuals against the business case model, with written explanations for gaps.
  • Commitment coverage and utilisation tracked as an explicit KPI with a named owner.

In practice the gap between modelled and actual savings almost always comes from the same places: rightsizing never executed after the lift-and-shift, non-production left running 24/7, orphaned volumes and snapshots, and log ingestion nobody budgeted. All four are detectable within a month of go-live if you instrumented the platform properly.

A pragmatic approach

Build the model in three passes. A two-week order-of-magnitude estimate to decide whether the question is worth asking. A detailed per-wave model once discovery data is available, with scenarios and sensitivity. Then a living model, updated quarterly with real consumption, that becomes your FinOps baseline rather than a document archived after the steering committee.

The deliverable that matters is not the saving percentage. It is a model whose assumptions are explicit, testable and owned — one that tells you, six months in, whether you are on track and which lever to pull if you are not.

← Back to blog