Cloud & Kubernetes

Upgrading Kubernetes

An operational guide to absorbing the pace of three minor releases a year without a maintenance window: compatibility, cutover strategies, controlled draining and automation.

September 2026

Kubernetes has released three minor versions per year since 2021, and each version is maintained for only 14 months (12 months of active support, 2 months of maintenance). A platform that upgrades only once a year therefore lives structurally at end of support, with no upstream security fixes and often outside the guarantees of its managed provider.

The constraint is hardening on the provider side: AWS bills EKS extended support at $0.60 per cluster per hour, roughly six times the standard rate, before an automatic upgrade at the end of the line; AKS and GKE apply equivalent mechanisms. On top of this comes API debt — removal of Ingress extensions/v1beta1 in 1.22, of PodSecurityPolicy in 1.25 — and the independent schedules of CNIs, CSIs and operators, which turn a deferred version jump into a remediation project.

The good news is that Kubernetes' compatibility policy makes the operation predictable: the control plane moves up one minor version at a time and kubelets can stay up to 3 minor versions behind the API server, which opens a genuine window for progressive migration. This document describes how to exploit that margin: detect obsolete APIs before the switch, choose between in-place and blue/green node pools, calibrate PodDisruptionBudgets and draining, then automate and verify the whole thing.

The tempo set by upstream

The question is no longer whether to upgrade, but how often and under what conditions. Three dynamics are converging today to turn the Kubernetes upgrade from a one-off project into a permanent operational routine.

What remains is to translate this calendar constraint into engineering practice: compatibilities to check, migration strategy, draining and automation.

Version skew and deprecated APIs

Kubernetes allows a bounded skew between its components, and that window — not the team's calendar — determines the order and pace of the upgrades.

**Key takeaway:** Don't just validate the manifests in your Git repository: enable audit logs and the `apiserver_requested_deprecated_apis` metric for at least one full cycle (including monthly batch jobs) to capture the actual calls made by operators, Helm charts and CI scripts before locking in the upgrade date.

Upgrade strategies compared

Three families of strategies coexist in production, and they differ not in elegance but in how long it takes to roll back.

**Common pitfall:** A minor-version downgrade of the control plane is not supported: etcd schemas only migrate forward. Any in-place strategy must therefore rely on a roll-forward plan, a verified etcd snapshot and tested application backups — not on the illusion of a rollback button.

Draining, PDBs and awkward workloads

Draining is the only moment in an upgrade when disruption becomes visible: that is where the 502s and the pods stuck for hours are concentrated.

Automate, test, validate continuously

Three minor versions a year means one upgrade per quarter: at that pace, only automation keeps the exercise from turning back into a project.

**Key takeaway:** An upgrade repeated four times a year is an incident; a node rotation executed every week is a routine. Automate rotation before automating the upgrade: it is rotation that reveals misconfigured PDBs and unclean shutdowns, cold, outside the critical window.

Key takeaways

  • Do you know the End of Life date of the minor version currently running in production on each of your clusters?
  • Have you instrumented `apiserver_requested_deprecated_apis` and audit logs over a full cycle, monthly batches included, rather than relying solely on the contents of your Git repositories?
  • Does every exposed application have a PodDisruptionBudget consistent with its replica count, and a `terminationGracePeriodSeconds` aligned with the real duration of its requests?
  • Have you restored — not merely produced — an etcd snapshot within the last six months?
  • Are you able to create an ephemeral cluster on version n+1 from your CI and deploy your entire stack onto it without manual intervention?
  • Are your nodes immutable and replaced regularly, rather than updated in place and kept for several months?

YOUR GUIDE TO KEEP

Upgrading Kubernetes

Get the complete guide to explore the topic further and share best practices with your team.

Download the PDF

Free PDF · Direct access

Ready to put it into practice?

Our experts help you move your cloud projects forward.

Go further

AWS EKS Auto Mode: Simplifying Kubernetes Management in ProductionCloud

AWS EKS Auto Mode: Simplifying Kubernetes Management in Production

EKS Auto Mode, launched in late 2024, fully delegates node group management, networking and storage to AWS. Here is what that changes for DevOps teams.

Kubernetes Gateway API: Migrating from Nginx IngressKubernetes & Conteneurs

Kubernetes Gateway API: Migrating from Nginx Ingress

The Gateway API became GA in Kubernetes 1.31 and is gradually replacing the Ingress. More expressive, multi-tenant and extensible — here is how to migrate your workloads.

Europcar Mobility GroupCase Study

Europcar Mobility Group

Run several brands' services on one platform able to absorb seasonal peaks, with little in-house Kubernetes expertise.