The Three Pillars of Kubernetes Observability
A Kubernetes cluster without observability is a ticking time bomb. Failures happen — a pod stuck in CrashLoopBackOff, a node under MemoryPressure, an ingress that silently stops responding. Without monitoring, you learn about incidents from end users. With a complete observability stack, you are alerted before the impact ever becomes visible.
Kubernetes observability rests on three complementary pillars:
- Metrics (Prometheus): numeric time series — CPU, memory, request count, latency, error rate. These tell you what is happening and how much.
- Logs (Loki): structured or unstructured text events from applications and the system. These tell you why something happened.
- Traces (Tempo / Jaeger): end-to-end tracking of individual requests across multiple microservices. These tell you where latency or errors originate in a distributed system.
Grafana serves as the unified visualisation layer across all three sources, enabling cross-correlation between metrics and logs in the same dashboard panel. When an alert fires, you can jump directly from a metric spike to the relevant log lines — cutting mean time to resolution (MTTR) dramatically.
This guide focuses on the most battle-tested open-source stack: Prometheus + Alertmanager for metrics and alerting, Loki + Promtail for logs, and Grafana as the frontend. The whole stack is installable via a single Helm chart.
Installing kube-prometheus-stack
kube-prometheus-stack is the community Helm chart that bundles Prometheus Operator, Grafana, Alertmanager, kube-state-metrics, node-exporter, and a full set of pre-built alerting rules for Kubernetes internals. It is the recommended starting point for any new cluster — it avoids hours of manual wiring.
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
helm install kube-prometheus-stack prometheus-community/kube-prometheus-stack --namespace monitoring --create-namespace --version 61.0.0 --values values.yaml
The critical part is a well-crafted values.yaml. The defaults are intentionally conservative; a production deployment needs explicit retention, persistence, and Alertmanager configuration:
# values.yaml — production-ready configuration
grafana:
adminPassword: "replace-with-a-vault-secret"
persistence:
enabled: true
storageClassName: gp3
size: 10Gi
ingress:
enabled: true
ingressClassName: nginx
annotations:
cert-manager.io/cluster-issuer: letsencrypt-prod
hosts: ["grafana.myapp.example.com"]
prometheus:
prometheusSpec:
retention: 30d
retentionSize: "40GB"
storageSpec:
volumeClaimTemplate:
spec:
storageClassName: gp3
resources:
requests:
storage: 50Gi
resources:
requests:
cpu: 500m
memory: 2Gi
limits:
memory: 4Gi
alertmanager:
alertmanagerSpec:
storage:
volumeClaimTemplate:
spec:
storageClassName: gp3
resources:
requests:
storage: 5Gi
config:
route:
group_by: [alertname, cluster, namespace]
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
receiver: slack-critical
receivers:
- name: slack-critical
slack_configs:
- channel: '#alerts-prod'
api_url: "YOUR_SLACK_WEBHOOK_URL"
title: '[{{ .Status }}] {{ .GroupLabels.alertname }}'
text: >-
{{ range .Alerts }}
*Alert:* {{ .Annotations.summary }}
*Severity:* {{ .Labels.severity }}
*Namespace:* {{ .Labels.namespace }}
{{ end }}
Key decisions baked into this configuration: 30-day retention ensures you can investigate incidents that happened last month; the 40 GB size cap prevents Prometheus from consuming unbounded storage; Alertmanager groups by alertname, cluster, and namespace to suppress noisy floods during cluster-wide outages.
Custom Application Metrics with prom-client
Kubernetes system metrics (CPU, memory, pod restarts) are necessary but not sufficient. You must also instrument your applications with business-level metrics: request throughput, error rates by endpoint, payment processing latency, queue depth. These are the metrics that actually define the user experience.
Here is a complete Node.js instrumentation example using the official prom-client library:
// Node.js — instrumentation with prom-client
const promClient = require('prom-client');
// Enable default metrics (event loop lag, GC, heap size)
promClient.collectDefaultMetrics({ prefix: 'myapp_' });
const httpRequestsTotal = new promClient.Counter({
name: 'http_requests_total',
help: 'Total number of HTTP requests',
labelNames: ['method', 'route', 'status'],
});
const httpRequestDuration = new promClient.Histogram({
name: 'http_request_duration_seconds',
help: 'HTTP request duration in seconds',
labelNames: ['method', 'route'],
buckets: [0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5],
});
const activeConnections = new promClient.Gauge({
name: 'http_active_connections',
help: 'Number of active HTTP connections',
});
// Express middleware
app.use((req, res, next) => {
const end = httpRequestDuration.startTimer();
activeConnections.inc();
res.on('finish', () => {
httpRequestsTotal.inc({
method: req.method,
route: req.route?.path || req.path,
status: res.statusCode,
});
end({ method: req.method, route: req.route?.path || req.path });
activeConnections.dec();
});
next();
});
// Metrics endpoint scraped by Prometheus
app.get('/metrics', async (req, res) => {
res.set('Content-Type', promClient.register.contentType);
res.end(await promClient.register.metrics());
});
For Prometheus to discover and scrape this endpoint, add a ServiceMonitor custom resource (the Prometheus Operator's way of configuring scrape targets):
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: myapp-metrics
namespace: production
labels:
release: kube-prometheus-stack # must match Prometheus selector
spec:
selector:
matchLabels:
app: myapp
endpoints:
- port: http
path: /metrics
interval: 15s
Critical Alerting Rules
The kube-prometheus-stack ships with a comprehensive set of default alert rules. You should extend these with rules tailored to your own applications. Below are four rules that every production cluster needs:
groups:
- name: kubernetes-critical
rules:
- alert: PodCrashLooping
expr: rate(kube_pod_container_status_restarts_total[5m]) * 60 * 5 > 0
for: 5m
labels:
severity: critical
annotations:
summary: "Pod {{ $labels.namespace }}/{{ $labels.pod }} is crash-looping"
description: "{{ $value }} restarts in the last 5 minutes — check logs immediately"
- alert: NodeDiskPressure
expr: kube_node_status_condition{condition="DiskPressure",status="true"} == 1
for: 2m
labels:
severity: critical
annotations:
summary: "Node {{ $labels.node }} is under disk pressure"
description: "The node's disk is nearly full — pods may be evicted"
- alert: PodMemoryUsageHigh
expr: |
container_memory_working_set_bytes{container!=""}
/ on(pod, namespace) kube_pod_container_resource_limits{resource="memory"}
> 0.9
for: 5m
labels:
severity: warning
annotations:
summary: "{{ $labels.namespace }}/{{ $labels.pod }} is using >90% of its memory limit"
description: "Current usage: {{ $value | humanizePercentage }} — OOMKill risk is high"
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
/ sum(rate(http_requests_total[5m])) by (service) > 0.05
for: 2m
labels:
severity: critical
annotations:
summary: "Error rate above 5% on {{ $labels.service }}"
description: "Current error rate: {{ $value | humanizePercentage }} — investigate immediately"
These rules follow the RED method (Rate, Errors, Duration) — the three metrics that best describe the health of any service from a user-facing perspective. For infrastructure, complement them with USE (Utilization, Saturation, Errors) metrics on nodes and persistent volumes.
Log Aggregation with Loki and Promtail
Loki is Prometheus for logs. Rather than indexing the full log content (which is expensive), Loki only indexes the labels (namespace, pod name, container) and stores compressed log chunks. This makes it dramatically cheaper than Elasticsearch for the same retention period, while integrating seamlessly with Grafana.
helm install loki grafana/loki-stack --namespace monitoring --set loki.persistence.enabled=true --set loki.persistence.storageClassName=gp3 --set loki.persistence.size=20Gi --set promtail.enabled=true
Promtail runs as a DaemonSet, tailing the container logs from every node and forwarding them to Loki with the correct namespace/pod/container labels. Once deployed, you can run LogQL queries in Grafana:
# Show all ERROR lines from the production API, parsed as JSON
{namespace="production", app="api"} |= "ERROR" | json | line_format "{{.level}} {{.message}}"
# Count error log lines per minute over the last hour
sum(rate({namespace="production"} |= "ERROR" [1m])) by (app)
# Find slow database queries (>500ms) in the last 15 minutes
{namespace="production", app="api"} | json | duration > 500ms | logfmt
The real power of Grafana + Loki is correlation: when a Prometheus alert fires, you can click through to the exact log lines from the offending pod during the alert window — no manual timestamp hunting across multiple tools.
Essential Grafana Dashboard IDs
Grafana's dashboard marketplace (grafana.com/grafana/dashboards) hosts hundreds of community dashboards. These five are the most valuable for a Kubernetes environment:
- ID 15760 — Kubernetes Cluster Overview: top-level view of CPU, memory, pod counts, and namespace-level resource consumption. This is the first dashboard you open when something is wrong.
- ID 13770 — Node Exporter Full: detailed host-level metrics for every node (per-CPU core utilisation, memory pressure zones, disk I/O, network saturation).
- ID 14205 — Kubernetes API Server: API server request latency by verb and resource, error rates, and etcd round-trip times. Critical for diagnosing kubectl slowness or failed deployments.
- ID 15358 — Kubernetes Nginx Ingress: per-ingress latency percentiles, HTTP status code breakdown, request throughput, and upstream response times.
- ID 12559 — Loki Logs Explorer: a pre-built panel layout for exploring Loki logs alongside Prometheus metrics, enabling side-by-side correlation.
Import these dashboards directly from the Grafana UI (Dashboards → Import → enter the ID). They automatically bind to your Prometheus and Loki data sources if the data source names match the defaults.
Defining SLOs and Error Budgets
Dashboards and alerts have real value only when anchored to explicit Service Level Objectives (SLOs). An SLO is a target for a service level indicator (SLI) measured over a rolling time window. Without SLOs, every alert is either ignored or treated as equally critical — both failure modes erode team trust in the monitoring system.
A practical SLO definition for a production API:
- Availability: 99.9% over 30 days — this corresponds to a maximum of 43.8 minutes of total downtime per month. Any Prometheus
upmetric or successful health check can serve as the SLI. - Latency P99: less than 500 ms for 99% of requests, measured over any 5-minute window. Use
histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m])). - Error rate: less than 0.1% of requests returning 5xx status over any 5-minute window.
The error budget is the inverse of the SLO: if availability is 99.9%, you have a budget of 0.1% failures per month. Configure burn-rate alerts on this budget rather than raw thresholds:
- If you burn 10% of the monthly budget in 1 hour → page the on-call engineer immediately (14.4x burn rate)
- If you burn 5% of the monthly budget in 6 hours → create a high-priority ticket (5x burn rate)
- If the budget is more than 50% consumed with 10 days left → freeze risky releases and focus on reliability
This approach prevents alert fatigue: low-severity issues that consume little budget are silently tracked, while rapid budget consumption triggers immediate action proportional to the risk.
Conclusion
A complete Prometheus, Grafana, and Loki stack can be operational on any Kubernetes cluster in under one hour using kube-prometheus-stack and the Loki Helm chart. But tooling alone is not observability. Real observability requires three things working in concert: the right instrumentation in your applications (custom metrics at the business layer), well-defined SLOs that give alerts meaning, and clear runbooks so that when an alert fires, whoever picks it up knows exactly what to investigate and what actions to take.
Start with the kube-prometheus-stack defaults, add your application's ServiceMonitor, import the five recommended dashboard IDs, and define SLOs for your most critical services. Within a week, your cluster will go from a black box to a system you genuinely understand — and incidents will shift from unpleasant surprises to managed, rehearsed procedures.
