Why OpenTelemetry?
Before OpenTelemetry, every observability tool had its own proprietary SDK: Jaeger SDK, Zipkin SDK, Datadog Agent, New Relic SDK. Instrumenting an application meant choosing a vendor from the start, with tight coupling and prohibitive migration costs. OpenTelemetry (OTel) solves this fundamental problem by providing a single, vendor-neutral SDK that has become the CNCF de-facto standard.
OTel unifies the three observability signals in a single API:
- Traces: end-to-end tracking of a request across multiple services (spans)
- Metrics: counters, gauges, histograms exposed in OTLP or Prometheus format
- Logs: structured logs with automatic correlation to traces via Trace ID
OTel Architecture
The OTel collection chain follows this flow:
Application (OTel SDK)
│ OTLP gRPC/HTTP
▼
OTel Collector (pipeline)
├── Receivers : OTLP, Prometheus, Jaeger, Zipkin, Fluentd
├── Processors : batch, memory_limiter, resource, sampling
└── Exporters : Tempo (traces), Prometheus (metrics), Loki (logs)
│
├── Grafana Tempo → Traces
├── Prometheus → Metrics
└── Grafana Loki → Logs
Node.js Auto-Instrumentation
OTel provides auto-instrumentation libraries that automatically patch popular modules (Express, HTTP, gRPC, Redis, PostgreSQL) without modifying application code:
// tracing.ts — load BEFORE any other module
import { NodeSDK } from '@opentelemetry/sdk-node';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-grpc';
import { OTLPMetricExporter } from '@opentelemetry/exporter-metrics-otlp-grpc';
import { PeriodicExportingMetricReader } from '@opentelemetry/sdk-metrics';
import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node';
import { Resource } from '@opentelemetry/resources';
import { SemanticResourceAttributes } from '@opentelemetry/semantic-conventions';
const sdk = new NodeSDK({
resource: new Resource({
[SemanticResourceAttributes.SERVICE_NAME]: 'payment-service',
[SemanticResourceAttributes.SERVICE_VERSION]: process.env.APP_VERSION,
[SemanticResourceAttributes.DEPLOYMENT_ENVIRONMENT]: process.env.NODE_ENV,
}),
traceExporter: new OTLPTraceExporter({
url: 'http://otel-collector:4317',
}),
metricReader: new PeriodicExportingMetricReader({
exporter: new OTLPMetricExporter({ url: 'http://otel-collector:4317' }),
exportIntervalMillis: 15000,
}),
instrumentations: [getNodeAutoInstrumentations()],
});
sdk.start();
process.on('SIGTERM', () => sdk.shutdown());
OTel Collector Configuration
The Collector is the central component that receives, transforms, and exports signals to backends:
# otel-collector-config.yaml
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
prometheus:
config:
scrape_configs:
- job_name: 'kubernetes-pods'
kubernetes_sd_configs:
- role: pod
processors:
batch:
timeout: 5s
send_batch_size: 1024
memory_limiter:
limit_mib: 512
spike_limit_mib: 128
resource:
attributes:
- key: deployment.environment
value: production
action: upsert
tail_sampling:
decision_wait: 10s
policies:
- name: errors-policy
type: status_code
status_code: {status_codes: [ERROR]}
- name: slow-traces
type: latency
latency: {threshold_ms: 500}
- name: probabilistic
type: probabilistic
probabilistic: {sampling_percentage: 10}
exporters:
otlp/tempo:
endpoint: http://tempo:4317
tls:
insecure: true
prometheusremotewrite:
endpoint: http://prometheus:9090/api/v1/write
loki:
endpoint: http://loki:3100/loki/api/v1/push
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch, resource, tail_sampling]
exporters: [otlp/tempo]
metrics:
receivers: [otlp, prometheus]
processors: [memory_limiter, batch]
exporters: [prometheusremotewrite]
logs:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [loki]
Grafana Tempo as Trace Backend
Tempo is Grafana Labs' distributed trace backend, designed for massive, low-cost trace storage using S3/GCS as object storage:
# values.yaml for the Tempo Helm chart
tempo:
storage:
trace:
backend: s3
s3:
bucket: my-tempo-traces
region: eu-west-1
access_key: ${AWS_ACCESS_KEY_ID}
secret_key: ${AWS_SECRET_ACCESS_KEY}
retention: 168h # 7 days
tempoQuery:
enabled: true
helm repo add grafana https://grafana.github.io/helm-charts
helm upgrade --install tempo grafana/tempo -f values.yaml --namespace monitoring
Traces ↔ Logs ↔ Metrics Correlation in Grafana
The power of OTel lies in the automatic correlation of all three signals. In Grafana, configure datasources to enable navigation between signals:
- Trace to Logs: from a Tempo trace, open Loki logs filtered by
traceID - Trace to Metrics: from a trace, view Prometheus metrics for that service at that exact moment
- Exemplars: from a Prometheus metric, jump directly to a representative trace
W3C Trace Context Propagation
For distributed cross-service traces, OTel uses the W3C Trace Context standard (RFC 7230). The traceparent and tracestate headers are automatically propagated by SDKs between HTTP and gRPC services.
Sampling Strategies
- Head-based sampling: decision made at the entry of the trace (before having complete data). Simple, low overhead. Configurable by rate (10 % of traces).
- Tail-based sampling: decision made after the trace ends (access to all data). Allows keeping 100 % of error traces while sampling 5 % of normal traces. Requires more memory in the Collector.
OTel Operator for Kubernetes
The OTel Operator enables auto-instrumentation of applications via a simple Kubernetes annotation, without modifying code or Docker images:
apiVersion: apps/v1
kind: Deployment
metadata:
name: payment-service
spec:
template:
metadata:
annotations:
instrumentation.opentelemetry.io/inject-nodejs: "true"
# Also available: inject-java, inject-python, inject-dotnet
spec:
containers:
- name: app
image: payment-service:1.0.0
# Instrumentation CRD configuration
apiVersion: opentelemetry.io/v1alpha1
kind: Instrumentation
metadata:
name: default-instrumentation
namespace: production
spec:
exporter:
endpoint: http://otel-collector:4317
propagators:
- tracecontext
- baggage
sampler:
type: parentbased_traceidratio
argument: "0.1" # 10 % sampling
Conclusion
OpenTelemetry has solved the vendor lock-in problem in observability. With a single SDK, you can change backends (Datadog to Grafana Stack, or NewRelic to Jaeger) without touching application code. The open-source stack OTel Collector + Grafana Tempo + Prometheus + Loki delivers enterprise-grade observability at controlled cost on Kubernetes.
