AI-Assisted FinOps: From Cost Anomaly Detection to Cost per Customer
FinOps & Coûts

AI-Assisted FinOps: From Cost Anomaly Detection to Cost per Customer

July 31, 20269 min readFinOpsKubernetesOpenCost

How to build an AI-assisted FinOps loop: statistically sound cost anomaly detection, honest ML forecasting, Kubernetes namespace allocation with OpenCost, showback and per-customer unit economics.

FinOps has outgrown the monthly spreadsheet

For years, cloud cost management meant a monthly export, a pivot table, and a tense meeting where someone asked why the bill went up 12%. That model breaks the moment your platform becomes elastic: autoscaling node pools, spot interruptions, serverless functions billed per millisecond, GPU nodes rented for a training run, and a dozen teams sharing a handful of Kubernetes clusters. The billing signal becomes high-cardinality, noisy, and delayed — exactly the conditions where humans stop noticing drift.

Two things changed the game. First, the FinOps Foundation's FOCUS specification gave us a vendor-neutral billing schema, so anomaly detection and forecasting no longer need per-provider parsers. Second, the maturing of managed anomaly detection (AWS Cost Anomaly Detection, Azure Cost Management anomaly alerts, GCP cost anomaly detection) plus open-source allocation engines (OpenCost, Kubecost) made it realistic to close the loop: detect, attribute, forecast, and bill back — automatically.

This article is about building that loop pragmatically, and about being honest on where machine learning genuinely helps versus where a well-chosen statistical baseline is better.

Cost anomaly detection: what ML actually buys you

A cost anomaly is not "spend went up". It is "spend went up in a way that is not explained by known seasonality, known growth, or a known deployment". That distinction is the whole engineering problem.

Static thresholds fail because cloud spend is non-stationary: it has weekly seasonality (CI runners idle on weekends), monthly seasonality (batch closing jobs), and structural growth. Percentage-over-last-week rules generate so many false positives that teams mute the channel within a month — the classic alert fatigue death spiral.

ApproachGood atWeak atWhen to use
Static budget thresholdHard financial guardrails, prepaid commitmentsAnything elastic; no seasonality awarenessAccount-level circuit breakers only
Robust z-score / MAD on daily spendCheap, explainable, few dependenciesStrong seasonality, low-volume servicesFirst iteration, per service × account
Seasonal decomposition (STL, ETS, Prophet-style)Weekly/monthly patterns, trend separationAbrupt regime changes, sparse seriesMature FinOps with 6+ months history
Managed cloud anomaly detectionZero infra, provider-native granularityMulti-cloud correlation, Kubernetes-level attributionBaseline coverage on every account
Gradient boosting with exogenous featuresCorrelating cost with deploys, traffic, tenantsExplainability, cold start, maintenanceLarge platforms with a data team

In practice the biggest win is not a fancier model — it is granularity plus context. Detecting a 30% jump on a whole AWS account is useless; detecting that EC2-Other / NAT Gateway data processing in one region for one cost-allocation tag tripled at 14:00 UTC yesterday is actionable. Run detection per (service, region, tag) tuple, not on the aggregate.

A robust, explainable starting point on daily FOCUS data:

import pandas as pd, numpy as np

# df: FOCUS-like columns ChargePeriodStart, ServiceName, x_Team, BilledCost
daily = (df.groupby([pd.Grouper(key="ChargePeriodStart", freq="D"),
                     "ServiceName", "x_Team"])["BilledCost"]
           .sum().reset_index())

def flag(group, window=28, k=4.0, min_delta=50.0):
    s = group.set_index("ChargePeriodStart")["BilledCost"].asfreq("D", fill_value=0)
    # same weekday baseline removes weekly seasonality cheaply
    base = s.shift(1).rolling(window).apply(lambda w: np.median(w[-28::7]), raw=True)
    mad  = s.shift(1).rolling(window).apply(
        lambda w: np.median(np.abs(w - np.median(w))), raw=True)
    score = (s - base) / (1.4826 * mad + 1e-9)
    group = group.assign(baseline=base.values, score=score.values)
    return group[(group.score > k) & (group.BilledCost - group.baseline > min_delta)]

anomalies = daily.groupby(["ServiceName", "x_Team"], group_keys=False).apply(flag)

Note the min_delta guard: a service going from $2 to $12 is a 6x spike and financially irrelevant. Every anomaly alert must carry an estimated monthly impact if unfixed, otherwise engineers cannot triage.

The second must-have is correlation with change events. Join your anomaly timestamps against deployment records (Argo CD application sync history, GitHub deployment events, Terraform apply logs). "Spend on nat-gateway in eu-west-3 rose 4σ two hours after payments-api v2.14.0 rolled out" turns a finance alert into an engineering ticket with an owner.

ML forecasting: pick the right horizon and the right question

Forecasting cloud spend is easier than forecasting revenue, because a large share of the bill is inertial: committed instances, storage that only grows, baseline compute. The useful questions are narrower than "what will we spend next year":

  • End-of-month landing (7–25 day horizon): mostly extrapolation of a partially observed month. A simple model conditioned on day-of-month and weekday beats a generic LSTM almost every time.
  • Commitment coverage (1–12 months): forecast the stable floor of on-demand usage per instance family, not total spend. This directly drives Savings Plans / Reserved Instance / committed-use discount decisions.
  • Capacity and unit cost (quarterly): forecast cost per business unit (per tenant, per order, per inference) to feed pricing and margin models.

Two disciplines separate credible forecasts from decoration. First, backtest: rolling-origin evaluation with MAPE or better, weighted absolute percentage error, reported per service. If you cannot state your forecast error, no CFO should use it. Second, decompose: separate committed spend (deterministic), baseline usage (predictable), and elastic usage (the only part that really needs a model). Forecasting the aggregate hides where the uncertainty lives.

Also expose prediction intervals, not a single line. A budget conversation goes very differently when engineering says "P50 is X, P90 is Y, and the gap is driven by the batch ML training schedule".

Kubernetes cost allocation: the hard part

Kubernetes is where FinOps attribution breaks, because the cloud bill stops at the node. The provider invoices you for an m6i.4xlarge; it has no idea that 40% of it runs the search team's indexers. You need an allocation engine that joins Prometheus workload metrics with node pricing.

OpenCost (CNCF incubating, and the engine underlying Kubecost) is the de facto standard. Its model is straightforward: for each pod, compute CPU, memory, GPU, storage and network allocation over time, price it against the node's hourly cost, and roll it up by namespace, label, controller or annotation.

The single most important design decision is the allocation basis:

BasisEffectRisk
max(requests, usage)Default; charges teams for reserved capacity they blockTeams over-request and pay for it — which is the point
Usage only"Fair" to bursty workloadsOver-requesting becomes free; cluster idle explodes
Requests onlyPerfectly predictable, drives rightsizingPenalises legitimately bursty jobs

Use max(requests, usage). It creates the correct incentive: reserving capacity costs money whether you use it or not — exactly how the underlying node billing works.

The second decision is idle and shared cost. A cluster is never 100% packed; there is headroom, DaemonSets, control plane, ingress, service mesh, observability agents. You can either leave idle unallocated (honest, but nobody owns it, so nobody fixes it) or redistribute it proportionally to each namespace's allocated share (creates pressure to improve bin-packing). We recommend: show both. Charge back proportionally, but display the idle line explicitly so the platform team owns the bin-packing KPI.

A minimal namespace cost query against OpenCost's exported metrics:

sum by (namespace) (
  avg_over_time(container_cpu_allocation[1h])
    * on (node) group_left()
      avg_over_time(node_cpu_hourly_cost[1h])
)
+
sum by (namespace) (
  avg_over_time(container_memory_allocation_bytes[1h]) / 1024^3
    * on (node) group_left()
      avg_over_time(node_ram_hourly_cost[1h])
)

And the governance that makes it work — allocation is only as good as your labels. Enforce them at admission time rather than chasing owners later:

apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
  name: require-cost-labels
spec:
  validationFailureAction: Enforce
  rules:
    - name: check-owner-labels
      match:
        any:
          - resources:
              kinds: [Namespace]
      validate:
        message: "Namespaces must carry cost-center, team and env labels"
        pattern:
          metadata:
            labels:
              cost-center: "?*"
              team: "?*"
              env: "?*"

Showback before chargeback

Chargeback — actually moving money between internal budgets — is an organisational decision, not a technical one, and it fails when the data is not yet trusted. Start with showback: every team sees its monthly allocated cost, its trend, its idle share, and its top three optimisation opportunities. No invoice, no argument about 3% attribution error.

Move to chargeback only when three conditions hold: label coverage above a level you are willing to defend publicly, a documented and stable allocation methodology (including how idle and shared services are split), and a dispute process. Changing the allocation formula mid-year without a migration note destroys credibility faster than any billing surprise.

Unit economics: cost per customer, per request, per model call

This is where FinOps stops being a cost-cutting function and becomes a product function. Absolute spend is a poor signal — a bill that doubles while you triple your customer base is good news. What leadership actually needs is a denominator.

Common unit metrics, in rough order of usefulness:

  • Cost per tenant / per account: essential for SaaS gross margin and for spotting customers you are losing money on.
  • Cost per transaction (order, API call, document processed): the engineering efficiency metric; it should decline as you scale.
  • Cost per 1k inferences or per token: now the dominant metric for AI features, where a single heavy user can distort margins.
  • Cost per environment: the fastest source of quick wins, because non-production spend is rarely optimised.

For a multi-tenant platform where tenants share the same pods, allocation needs an application-level signal. Emit a per-tenant usage counter from the application, then split the namespace cost proportionally:

-- dbt model: cost_per_tenant.sql
with ns_cost as (
  select date_day, namespace, allocated_cost
  from {{ ref('k8s_namespace_cost_daily') }}
),
usage as (
  select date_day, namespace, tenant_id,
         sum(weighted_requests) as units       -- e.g. requests * cpu_ms
  from {{ ref('app_tenant_usage_daily') }}
  group by 1,2,3
)
select u.date_day, u.tenant_id,
       sum(c.allocated_cost * u.units
           / sum(u.units) over (partition by u.date_day, u.namespace)) as tenant_cost
from usage u join ns_cost c using (date_day, namespace)
group by 1,2

Weight the driver properly: raw request counts are misleading if one endpoint is 100x heavier. Use CPU-milliseconds, bytes stored, or tokens consumed — whatever actually drives the resource being billed. Then join tenant cost with contract revenue and you have per-customer gross margin, refreshed daily.

A reference architecture that stays maintainable

The pattern that holds up across clients is deliberately boring:

  1. Ingest: provider cost exports (CUR 2.0 / FOCUS, Azure cost exports, GCP BigQuery billing export) into object storage; OpenCost allocation data exported daily to the same lake.
  2. Normalise: one FOCUS-shaped table across providers, plus a dimension table for teams, cost centres and tenants sourced from Git, not from a spreadsheet.
  3. Model: dbt models for daily allocated cost, unit metrics and forecasts, all version-controlled and tested.
  4. Detect: anomaly job running on the normalised table, plus native provider detectors as a safety net.
  5. Deliver: alerts into the owning team's Slack channel with impact and probable cause; dashboards in the same tool engineers already use.

Resist building a bespoke FinOps portal. Cost data that lives where engineers already work — Grafana, Backstage, their PR checks — gets acted upon; a separate portal gets bookmarked once.

Anti-patterns to avoid

Sending anomaly alerts to a central FinOps channel instead of the owning team. Optimising the tail of the bill while ignoring the top three services. Treating idle cluster capacity as a team problem rather than a platform KPI. Publishing a forecast without an error metric. And the most common one: rolling out chargeback before label coverage is trustworthy, which converts every cost conversation into an argument about the data instead of about the architecture.

Done right, AI-assisted FinOps is not about a model that magically cuts your bill. It is about compressing the time between a cost event and the engineer who can fix it — from a month to a few hours — and giving product leadership a per-customer margin they can actually act on.

← Back to blog