AWS Loses Customer Data in Bahrain: What Every Cloud Architecture Should Change
Cloud

AWS Loses Customer Data in Bahrain: What Every Cloud Architecture Should Change

September 20, 20267 min readAWSAWS BackupS3

AWS confirmed in September 2026 the permanent loss of certain customer data in its Bahrain region. A practical review of durability myths, failure domains, immutable backups and the residency-versus-resilience trade-off.

What actually happened

In mid-September 2026, AWS confirmed to affected customers that an incident in its Bahrain region (me-south-1) resulted in the permanent, unrecoverable loss of certain customer data. Not degraded performance. Not an extended outage. Data that no longer exists and that AWS cannot restore.

Details remain scoped to the customers concerned, and it would be irresponsible to speculate on root cause, blast radius or the exact resource types involved before a full post-event summary is published. What matters for everyone else is far more actionable: the failure mode that most architectures quietly assume away — the cloud provider losing your data — is not theoretical.

If your reaction was "that can't happen, S3 has eleven nines of durability", this article is written for you.

Durability is a statistical model, not a guarantee

The famous durability figures published for object storage describe the expected annual probability of losing an object under a specific replication model, within a specific scope. Three things are routinely misread:

  • Scope. Durability claims apply to a service's own redundancy design — typically across multiple Availability Zones within one region. They say nothing about correlated failures affecting an entire region, control-plane bugs, or a lifecycle rule you wrote yourself.
  • Failure class. Durability models cover hardware failure. They do not cover software defects, operator error (theirs or yours), malicious deletion, ransomware, account closure, or KMS key deletion.
  • Remedy. AWS's customer agreement and service level agreements provide service credits for availability. They are not a data-loss insurance policy. Credits do not rebuild your production database.

Durability is an engineering input. It is not a backup strategy.

The shared responsibility model, read literally

AWS is responsible for the resilience of the cloud. You are responsible for resilience in the cloud — which explicitly includes your backup and recovery strategy, your replication topology, and your ability to restore. That sentence appears in every AWS whitepaper. It is also the sentence most frequently skipped during a migration under deadline pressure.

Concretely, the question to ask of every stateful component in your platform is: if this region became permanently unavailable tomorrow, what would I have left, where does it live, and how long would it take to serve traffic again? If the answer requires a meeting to find out, you do not have a DR plan — you have an intention.

Know your failure domains

Most silent single points of failure in AWS architectures come from resources whose default failure domain is narrower than engineers assume.

ResourceDefault failure domainWhat actually protects you
EBS volumeSingle AZSnapshots + cross-region snapshot copy
EC2 instance storeSingle host — data lost on stopNever store state here
RDS / Aurora (single-AZ)Single AZMulti-AZ + automated backups replicated cross-region
S3 StandardMulti-AZ, single regionVersioning + Object Lock + Cross-Region Replication
S3 One Zone-IASingle AZReplication, or simply do not use it for anything you need
EFS One ZoneSingle AZAWS Backup with cross-region copy
DynamoDBMulti-AZ, single regionGlobal Tables + PITR + on-demand backups
ECR imagesSingle regionCross-region replication rules
KMS CMKSingle regionMulti-Region Keys — otherwise your cross-region copies are unreadable

That last row deserves emphasis. A perfectly replicated encrypted snapshot in a second region is worthless if the only key able to decrypt it lived in the lost region. Multi-Region KMS keys, or re-encryption at copy time, are part of the recovery path — not a detail.

Build a recovery path that does not trust the provider

The classic 3-2-1 rule translates cleanly to cloud: three copies of the data, on two distinct storage technologies or fault domains, with at least one copy outside the primary blast radius. In practice, three progressive levels:

Level 1 — cross-region copies inside AWS. Cheap, automatable, covers regional loss. Insufficient against account compromise or an organisation-wide misconfiguration.

Level 2 — a separate AWS account with isolated credentials. Backup vaults in a dedicated account, written to by a service role, with Vault Lock in compliance mode so that even an account administrator cannot shorten retention. This is what turns a backup into an immutable backup.

Level 3 — a copy off AWS entirely. For your genuinely irreplaceable datasets (billing ledger, source of truth for customers, regulatory archives), an export to a second provider or to on-premises storage. Expensive, slow, and the only thing that survives a provider-level catastrophe or a commercial dispute that locks your account.

Here is the level-2 baseline as Terraform, with cross-region copy and Vault Lock:

resource "aws_backup_vault" "primary" {
  name        = "prod-me-south-1"
  kms_key_arn = aws_kms_key.backup_mrk.arn
}

resource "aws_backup_vault_lock_configuration" "primary" {
  backup_vault_name   = aws_backup_vault.primary.name
  changeable_for_days = 3      # cooling-off window before the lock is immutable
  min_retention_days  = 30
  max_retention_days  = 365
}

resource "aws_backup_plan" "prod" {
  name = "prod-daily-with-dr-copy"

  rule {
    rule_name         = "daily-0200"
    target_vault_name = aws_backup_vault.primary.name
    schedule          = "cron(0 2 * * ? *)"
    start_window      = 60
    completion_window = 360

    lifecycle {
      delete_after = 35
    }

    copy_action {
      destination_vault_arn = aws_backup_vault.dr.arn  # aliased provider, second region
      lifecycle {
        cold_storage_after = 30
        delete_after       = 365
      }
    }
  }
}

resource "aws_backup_selection" "tagged" {
  name         = "tag-backup-true"
  plan_id      = aws_backup_plan.prod.id
  iam_role_arn = aws_iam_role.backup.arn

  selection_tag {
    type  = "STRINGEQUALS"
    key   = "Backup"
    value = "true"
  }
}

Tag-based selection matters: it makes backup coverage a property enforced by policy (deny resource creation without a Backup tag) rather than something a team remembers to configure.

The uncomfortable part: data residency versus resilience

Bahrain is a region chosen overwhelmingly for regulatory reasons — Gulf data residency requirements, financial-sector rules, public-sector mandates. Which creates a real tension: the obvious answer to regional data loss is "copy it to another region", and the obvious answer to residency is "don't".

There are workable paths, and they all require a decision rather than a default:

  • Map residency obligations to data classes, not to the whole platform. Very often, only a subset of personal or regulated data must remain in-country. Aggregates, operational metadata, artefacts and infrastructure state frequently can be copied elsewhere.
  • Use a second region within the same jurisdictional bloc where the regulator accepts it — this is exactly the argument for multi-region topologies in the Middle East rather than single-region deployments.
  • Keep an in-country copy outside AWS. A local provider or your own datacentre for the regulated subset. This is the same argument European teams make about sovereignty: a single provider, however large, is a single point of failure both technically and contractually.
  • Encrypt with customer-managed keys you control, so that an off-provider copy remains legally defensible as "not readable by the host".

The worst outcome is the implicit one: residency used as an excuse to have no second copy at all.

A backup you have not restored is a hypothesis

Every serious incident review reaches the same conclusion — the backups existed, and the restore did not work. Snapshots of a database whose engine version is no longer available. An encrypted copy without its key. A 4 TB restore that takes eleven hours when the stated RTO was two. A runbook referencing a person who left.

Make restoration a scheduled, measured, boring operation:

  • Quarterly game day: restore the most critical dataset into an isolated account, from the DR region, with no access to the primary region. Measure wall-clock time to first successful query.
  • Automate a weekly partial restore in CI — a single table, a single volume — and alert on failure. A backup job that reports success proves the write path, not the read path.
  • Record actual RTO and RPO measured, not declared, and put them next to the business expectation. The gap is your roadmap.
  • Keep infrastructure-as-code and container images reproducible outside the primary region. Recovering data is pointless if you cannot recreate the platform that serves it.

What to do this week

You do not need a twelve-month programme to materially reduce this risk. In order of effort-to-value:

  1. Inventory every stateful resource and its failure domain. A tag-based query plus an AWS Config rule will surface the single-AZ surprises in an afternoon.
  2. Turn on S3 versioning everywhere, and Object Lock on the buckets that matter. Enable PITR on DynamoDB tables.
  3. Create a dedicated backup account, move vaults into it, and apply Vault Lock. Revoke the ability of production roles to delete recovery points.
  4. Add cross-region copy actions to your backup plans, and verify that KMS keys are multi-region or that copies are re-encrypted.
  5. Schedule a restore drill with a date, an owner, and a measured outcome.

The Bahrain incident will eventually be explained, remediated, and written up. The durable lesson is independent of the root cause: cloud providers are extraordinarily reliable, and extraordinary reliability is not the same as infallibility. Design as though your provider will, one day, lose your data — because for a handful of customers in September 2026, it did.

← Back to blog