Why cloud programmes stall at the governance layer
Most large cloud programmes do not fail on technology. The landing zone gets built, the Kubernetes clusters run, Terraform modules are published. What breaks is the connective tissue: who decides what gets funded, who authorises a production change, how incidents are handled when half the estate is managed by a platform team and the other half by a legacy ops function, and how the whole thing is reported upwards without turning into a monthly PowerPoint theatre.
Two frameworks dominate the conversation in large European enterprises: SAFe for the delivery side (portfolio, Agile Release Trains, PI Planning) and ITIL 4 for the service side (change enablement, incident, problem, service configuration management). They are usually owned by different parts of the organisation, speak different vocabularies, and are often set against each other — "agility versus control". That framing is wrong, and expensive. Used deliberately, SAFe answers what should we build and in what order, ITIL answers how do we run it safely and recover when it breaks. A cloud programme needs both answers.
Two grammars, one value stream
The first practical step is to stop treating the two frameworks as competing operating models and map them onto a single value stream. SAFe describes flow from strategy to deployment; ITIL 4 describes the Service Value System from demand to value, including the run-time practices. They overlap in the middle — and that overlap is exactly where governance disputes happen.
| Governance question | SAFe artefact | ITIL 4 practice | Who should own it |
|---|---|---|---|
| Funding and prioritisation | Lean Portfolio Management, Epics, guardrails | Portfolio & demand management | LPM function, with FinOps input |
| Scope commitment | PI Planning, PI Objectives | Service level management | ART / product management |
| Production change authorisation | Continuous Delivery Pipeline, DoD | Change enablement | Platform team + change authority |
| Environment and asset truth | — | Service configuration management (CMDB) | Platform, fed automatically from IaC |
| Outage response | — (DevOps on-call) | Incident / major incident management | Product team on-call, ITSM coordination |
| Recurrence elimination | Inspect & Adapt, improvement backlog | Problem management | Shared, single backlog |
The pattern that works: ITIL practices define the control objectives, SAFe cadence defines where the work to satisfy them lives. Problem management does not get its own parallel backlog — its outputs become enabler stories in the team backlog, visible at PI Planning, competing for capacity like everything else. If a problem record cannot win a place in the backlog, that is a prioritisation conversation, not a process failure.
Change enablement is where the integration lives or dies
Nothing kills cloud velocity faster than running a weekly CAB over Terraform plans. ITIL 4 explicitly moved away from the monolithic Change Advisory Board towards the notion of a change authority appropriate to the change type, and it defines three types: standard (pre-authorised, low risk, documented procedure), normal (assessed and authorised), and emergency.
The entire goal of a cloud programme's governance design should be to push as much deployment traffic as possible into the standard change category, with the pipeline itself acting as the change authority. That is a legitimate ITIL 4 interpretation, not a workaround — but it only holds if the pipeline genuinely enforces the controls a human reviewer would have applied.
Concretely, a change is eligible for the standard path when it can demonstrate: an approved pull request with peer review, a successful policy-as-code evaluation, automated tests, a known blast radius, and a reversible deployment strategy. Encode that as policy rather than as a Confluence page.
package change.standard
# A deployment qualifies as a pre-authorised standard change
default eligible := false
eligible if {
input.pr.approvals >= 1
input.pr.author != input.pr.approvers[_]
input.tests.unit == "passed"
input.policy.scan.critical == 0
input.deployment.strategy in {"canary", "blue-green", "rolling"}
input.target.criticality != "tier-0"
not touches_shared_network
}
touches_shared_network if {
some r in input.plan.resource_changes
startswith(r.type, "aws_transit_gateway")
}
# Everything else falls back to normal change with human authority
requires_cab if { not eligible }The pipeline then writes the change record automatically — after the fact for standard changes, before deployment for normal ones. The ITSM tool stops being a gate that engineers resent and becomes a ledger that auditors can read.
- name: Register change record
if: always()
run: |
curl -sS -X POST "$SNOW_URL/api/now/table/change_request" \
-u "$SNOW_USER:$SNOW_PASS" \
-H 'Content-Type: application/json' \
-d "{
\"type\": \"standard\",
\"short_description\": \"Deploy ${SERVICE} ${GIT_SHA:0:7} to prod\",
\"cmdb_ci\": \"${CI_SYS_ID}\",
\"assignment_group\": \"${ART_NAME}\",
\"implementation_plan\": \"${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}/actions/runs/${GITHUB_RUN_ID}\",
\"backout_plan\": \"argocd app rollback ${SERVICE}\",
\"close_code\": \"${JOB_STATUS}\"
}"The CMDB problem, and how IaC solves it
Service configuration management is the ITIL practice that most consistently collapses in cloud environments. A CMDB maintained by humans describing auto-scaling, ephemeral infrastructure is wrong within hours. Yet without it, incident management has no impact analysis, change enablement has no risk assessment, and FinOps has no cost allocation.
The fix is to invert the flow: the CMDB must be a consumer of infrastructure truth, not its source. Terraform state, cloud provider inventory APIs and Kubernetes resources are the system of record; a reconciliation job projects them into configuration items. Mandatory tagging becomes a governance control with teeth, because tags are what carry the service, ART, criticality and cost-centre relationships into both the CMDB and the FinOps allocation model.
variable "governance_tags" {
type = object({
service_id = string # CI identifier in the CMDB
art = string # Owning Agile Release Train
criticality = string # tier-0 | tier-1 | tier-2
cost_center = string
data_class = string # public | internal | confidential | regulated
})
}
# Enforced at plan time, not discovered at audit timeA single tagging contract serves three governance functions at once: incident impact analysis, change risk scoring, and showback. That is the kind of leverage worth fighting for in the first increments of a programme.
Funding: LPM guardrails meet FinOps reality
SAFe's Lean Portfolio Management replaces project-based funding with funding of value streams under guardrails. In cloud, the guardrails must include consumption, not just headcount. An ART that is funded for twelve engineers and whose cloud spend doubles unnoticed is not operating within guardrails, whatever the burn-down chart says.
Practically, this means cloud cost becomes a first-class item in the quarterly portfolio review and in the ART's own cadence: unit economics (cost per transaction, per tenant, per model inference) reported alongside PI Objectives; anomaly detection routed to the team, not to a central FinOps mailbox; optimisation work that exceeds a threshold entering the backlog as an enabler epic with a business case. The mistake to avoid is creating a central FinOps team that produces reports nobody acts on because the teams producing the spend have no budget accountability.
Metrics that survive both audiences
Executives want assurance; engineers want signal. A governance dashboard that mixes SAFe flow metrics with ITIL service metrics and DORA delivery metrics gives both, provided it is small.
| Dimension | Metric | Why it matters for governance |
|---|---|---|
| Flow | Flow time, flow load per ART | Detects overcommitment before PI objectives slip |
| Delivery | Deployment frequency, lead time | Shows whether change enablement is a brake |
| Stability | Change failure rate, MTTR | Evidence that the standard-change path is safe |
| Service | SLO attainment, major incident count | Connects delivery choices to customer experience |
| Economics | Cost per unit of value, waste ratio | Keeps portfolio guardrails honest |
| Compliance | % deployments via standard path, policy violations | Proves the control model works at scale |
One metric deserves special attention: the share of production changes going through the pre-authorised standard path. If it rises while change failure rate stays flat, the integration is working. If it rises while failures climb, the policy gate is too permissive. If it stays low, the organisation has adopted SAFe vocabulary on top of unchanged ITIL bureaucracy.
Anti-patterns worth naming
Three failure modes recur. The first is the parallel governance stack: an ITSM process for the legacy estate and an informal one for cloud, with no bridge. Audit eventually forces reconciliation, usually at the worst moment. The second is SAFe as a reporting layer, where PI Planning produces a plan that project managers then track in a separate tool, with ITIL change processes untouched — all the ceremony, none of the flow. The third is the platform team as a ticket queue: every environment request, IAM role and DNS record flows through a service desk, which guarantees that the platform becomes the constraint on every ART.
The counter to all three is the same: treat the internal platform as a product with an interface (self-service APIs, golden paths, a service catalogue) and encode governance into that interface. Governance consumed as a paved road is adopted; governance consumed as a gate is circumvented.
A pragmatic sequencing
Do not attempt a big-bang operating model redesign. In the first quarter, agree the value stream map and the shared vocabulary, and define which services are tier-0. In the second, build the standard change path for one ART and one non-critical service, with policy-as-code and automated change records. In the third, wire the CMDB to IaC and introduce the tagging contract. Only then extend to tier-1 services and the rest of the portfolio, and fold FinOps guardrails into the LPM cadence.
The target state is unglamorous and highly effective: engineers who deploy many times a day without filing a ticket, auditors who get a complete and queryable change ledger, and a portfolio board that can see where money and capacity are actually going. SAFe and ITIL do not need to be reconciled philosophically. They need to be wired into the same pipeline.
