Why run a Well-Architected review on OVHcloud at all?
The Well-Architected Framework was popularised by AWS and later echoed by Azure and Google, but the pillars themselves — operational excellence, security, reliability, performance efficiency, cost optimisation, sustainability — are not vendor property. They are a review grid. What changes from one provider to another is the set of building blocks you can lean on, and where the shared-responsibility line actually sits.
On OVHcloud, two structural differences drive the whole exercise. First, the managed catalogue is deliberately narrower than a hyperscaler's: there is no managed WAF-as-a-service bolted onto every load balancer, no twenty-service observability suite, no drag-and-drop event bus. More of the platform is yours to build and operate, which means the operational excellence and reliability pillars carry more weight. Second, the economics and the sovereignty posture are genuinely different: outbound bandwidth is included in most regions, the vRack private network spans datacentres and regions, and a SecNumCloud-qualified perimeter exists for workloads that need it. Architectural choices that would be financially painful at a hyperscaler — cross-region replication, chatty hybrid topologies, keeping a copy of your data in a second site — become natural here.
A Well-Architected review on OVHcloud is therefore not a copy-paste of the AWS questionnaire. It is the same questions, answered with a different map.
Operational excellence: everything through the API, nothing through the panel
The OVHcloud control panel is fine for discovery and terrible for reproducibility. The baseline for any serious platform is Terraform with two providers side by side: the ovh provider for account-level and managed-service resources (Public Cloud projects, Managed Kubernetes, managed databases, IAM policies, Private Registry), and the openstack provider for the lower-level primitives exposed by Public Cloud (instances, volumes, security groups, Octavia load balancers).
A practical detail that trips most teams up: there is no DynamoDB equivalent for state locking. Since Terraform supports native S3 lock files, you can host the state in OVHcloud Object Storage (S3-compatible) and get locking without any external component:
terraform {
backend "s3" {
bucket = "tfstate-platform-prod"
key = "platform/terraform.tfstate"
region = "gra"
endpoints = { s3 = "https://s3.gra.io.cloud.ovh.net" }
use_lockfile = true
skip_credentials_validation = true
skip_requesting_account_id = true
skip_region_validation = true
skip_s3_checksum = true
}
}The other operational-excellence levers worth auditing:
- One Public Cloud project per environment. Projects are a hard isolation boundary (separate Keystone tenants, separate quotas, separate invoicing lines). They are a far better blast-radius control than any policy you will write.
- Managed Kubernetes upgrade policy. MKS exposes an update strategy per cluster; pick it deliberately and document it, rather than discovering a control-plane upgrade during a Friday incident. Same for node pools: use rolling updates with a surge node rather than in-place reboots.
- GitOps for everything above the cluster. Argo CD or Flux on MKS, with container images in Managed Private Registry (Harbor) so that image scanning and retention policies live in the same place as your supply chain.
- Quota management as a first-class task. Public Cloud quotas (instances, vCPU, volumes, floating IPs, load balancers) are per project and per region, and they are low by default. A Well-Architected review should include "can this project actually scale to its target size today?" — the answer is often no, and quota increases take time.
Security and sovereignty: coarse IAM, strong network isolation
OVHcloud IAM has improved substantially — identities, resource groups, policies with actions and conditions — but it is still less granular than AWS IAM. Do not architect as if you can express every least-privilege nuance in a policy document. Architect around isolation instead: separate Public Cloud projects, separate S3 credentials per bucket and per workload, separate API application keys per automation pipeline, each restricted and rotated.
On the network side, OVHcloud gives you more than most people use. The vRack is a layer-2 private network that spans datacentres and regions, and it is the glue for hybrid designs: bare-metal database servers or GPU nodes on one side, elastic Public Cloud instances and MKS worker nodes on the other, all on private addressing. Combined with a Public Cloud private network, a gateway for egress SNAT, and MKS clusters deployed without public IPs on the workers, you get a topology where nothing is exposed except the load balancers you explicitly declare.
Checklist items that come up in nearly every review: anti-DDoS is included but it is not a WAF — put a real WAF (Ingress-level or a dedicated appliance) in front of public HTTP; encryption at rest should be explicit (managed database encryption, KMS-backed keys, or a Vault instance you operate on MKS); MFA must be enforced on every control-panel identity; and if you are in a regulated sector, be precise about which offer actually carries the qualification you need — SecNumCloud and HDS apply to specific perimeters, not to the whole catalogue by default.
Reliability: know which of your regions are multi-AZ
This is the single biggest source of false assumptions on OVHcloud. Most historical regions are single-datacentre; the Paris region introduced true availability zones for Public Cloud. If your production runs in a single-AZ region, "high availability" means three instances in the same building, and your real resilience story is a second region.
Two viable patterns, and you should choose consciously:
- Multi-AZ inside the 3-AZ region. One node pool per zone, anti-affinity via topology spread constraints, managed databases with replicas across zones, Octavia load balancers in front. Low latency, simple failover.
- Active/passive across two regions over vRack. Asynchronous database replication, object storage replicated with a scheduled job, infrastructure described once in Terraform and instantiated twice with a different region variable. Slower RTO, but it survives the loss of an entire site — and because egress is included, continuous replication does not wreck the budget.
resource "ovh_cloud_project_kube" "prod" {
service_name = var.project_id
name = "prod-par"
region = "EU-WEST-PAR"
version = "1.31"
}
resource "ovh_cloud_project_kube_nodepool" "workers" {
for_each = toset(["eu-west-par-a", "eu-west-par-b", "eu-west-par-c"])
service_name = var.project_id
kube_id = ovh_cloud_project_kube.prod.id
name = "workers-${substr(each.key, -1, 1)}"
flavor_name = "b3-16"
availability_zones = [each.key]
autoscale = true
desired_nodes = 2
min_nodes = 2
max_nodes = 8
}And the eternal reminder: a volume snapshot stored in the same region is not a backup. Export to object storage, ideally in another region, and run a restore drill on a schedule. A Well-Architected review that does not ask for the date of the last successful restore is a documentation exercise, not an audit.
Performance efficiency: pick the right shape, not the biggest one
Public Cloud instance families map cleanly to workload profiles — balanced (b3), compute-optimised (c3), memory-optimised (r3), local-NVMe for IOPS-hungry databases, GPU families for training and inference. Block storage comes in several performance tiers, and object storage distinguishes a high-performance NVMe-backed S3 tier from the standard tier and a cold archive tier. Most over-spending and most latency complaints come from mismatches here: an application sitting on classic block storage when it needs NVMe, or a Prometheus/Loki stack writing to the wrong object storage class.
The hybrid card is worth playing. A steady-state PostgreSQL or an always-busy GPU training fleet is often dramatically cheaper and faster on dedicated bare metal connected through vRack, while the stateless tier stays elastic on MKS. This is an architecture that hyperscaler-trained teams rarely consider, and it is one of OVHcloud's genuine strengths.
For observability, you have a real choice: Logs Data Platform (managed Graylog/OpenSearch) and Metrics Data Platform when you want a managed, sovereign, retention-compliant store with no operational burden; or a self-hosted Prometheus/Mimir/Loki/Grafana stack on MKS with object storage as the long-term backend when you want maximum flexibility and control over cost. Our default recommendation: self-host the metrics and traces stack on MKS with S3-compatible object storage, and use the managed log platform when compliance requires a tamper-evident, long-retention log store you do not want to babysit.
Cost optimisation: different levers from the hyperscalers
| Lever | OVHcloud specificity | What to audit |
|---|---|---|
| Commitment | Monthly pricing and Savings Plans on Public Cloud instances | Any instance running 24/7 on hourly billing is money left on the table |
| Egress | Bandwidth included in most regions | Re-evaluate designs that were contorted to avoid egress fees elsewhere |
| Baseline vs burst | Bare metal over vRack for steady load | Workloads with flat utilisation curves still on elastic instances |
| Storage tiering | High-performance / standard / cold archive object storage | Backups and logs sitting on the premium tier |
| Chargeback | Billing granularity at the project level | Shared "everything" projects that make showback impossible |
| Waste | Orphan volumes, snapshots, floating IPs, idle load balancers | Monthly sweep driven by the consumption API |
Because tagging is less central than on AWS, project-per-application-per-environment is also your FinOps model. Design it that way from day one; retrofitting cost allocation later is painful.
Sustainability: a pillar OVHcloud actually helps with
OVHcloud's industrial model — in-house water cooling, long hardware reuse cycles, heat recovery on some sites, and European datacentres connected to relatively low-carbon grids — gives you a better starting point than most. But the infrastructure-side gains are the provider's; your pillar score depends on what you run. Scale non-production to zero outside working hours, use the cluster autoscaler aggressively, right-size from real metrics rather than from the ticket that asked for "8 vCPU to be safe", and prefer a smaller number of well-utilised nodes over a sprawl of half-idle ones. Choose the region with the cleaner grid when latency allows.
Running the review in practice
A useful Well-Architected session on OVHcloud lasts half a day per application domain and produces a prioritised backlog, not a slide deck. Our working checklist: Is every resource in Terraform, with a locked remote state? Is there one project per environment? Are MKS workers free of public IPs and behind a gateway? Is the region single-AZ, and does the team know it? When was the last restore test? Are quotas sized for the next twelve months? Are API keys rotated and scoped? Which offer carries the compliance qualification the contract assumes? Which workloads are running 24/7 without a commitment? What is the cost per environment, and can you even answer that question today?
Nine out of ten findings fall into three buckets: resilience assumed but never tested, isolation assumed but implemented through a single shared project, and cost assumed to be optimised because the provider is cheap. The framework's value is not the pillars — it is the discipline of asking the awkward questions before the incident does.
