AniSri
Observability Platform engineering June 25 2026 14 min read

Stop being scared of OpenTelemetry

OTEL looks like a vendor deck until you see the actual pipeline: apps talk to a node agent, the agent talks to a cluster gateway, traces land in Datadog, and logs land in Splunk. One instrumentation path, two exporter mappings. This is the version that skips the acronyms and ends with a pipeline you can build and break yourself.

When something broke, you were the correlation layer

Before OpenTelemetry, collecting telemetry was a local agreement. Every vendor had a proprietary format. Traces in one tool, metrics in another, logs in a third. Correlation was a manual exercise. When something broke at 2am, you had three browser tabs open, three query languages running, and zero shared context tying them together.

Rolling it out is a platform problem, not a mandate. Product teams will not re-instrument because you published a standard. Give them language-specific auto-instrumentation for HTTP, gRPC, and database calls, plus a Collector pipeline that continues sending APM data to Datadog and logs to Splunk. The Collector receives and processes telemetry; it does not instrument the application for you. Hand teams a working path and a sane default config. Let them break it. Adoption friction matters more than the specification.

OpenTelemetry is the technical piece. It's a CNCF-graduated project, the same graduation tier as Kubernetes, with common APIs, SDKs, semantic conventions, data models, and a transport protocol called OTLP. Instrument once, then send to any backend your Collector distribution has a compatible exporter for. Switching or splitting backends should become a pipeline change instead of an application rewrite. The platform team's job is to make that path the easy one.

The one-liner

OpenTelemetry is the USB-C of observability. One standard connection on the application side, then OTLP or a compatible exporter behind the Collector. Google, Microsoft, AWS, Datadog, and Grafana all support it, but the fidelity of that support still needs testing.

Monitoring tells you error rate is up. Observability lets you ask why and get a real answer. That gap shows up at 2am when the failure mode isn't in any runbook. Monitoring answers questions you wrote down ahead of time. Observability is for the ones you didn't. OTEL supplies the shared resource context, log trace IDs, and metric exemplars that a capable backend can use to move from a spike to a representative trace and then to its logs.

Three stable signals, one emerging signal

Traces, metrics, and logs have stable OpenTelemetry specifications. Profiles reached alpha in 2026 and are not ready for critical production workloads. SDK maturity still varies by language, especially for logs, profiles, and browser instrumentation. The signals share resource attributes such as service.name and deployment.environment.name, but they do not all carry a trace_id: spans do, correlated logs can, and metric data points link to traces only through optional exemplars. Baggage is also part of OpenTelemetry, but it is propagated application context rather than a telemetry stream you store and query like the other signals.

Signal 01
Logs
The diary. What happened, at what timestamp. An OTEL log body can be a plain string or a structured value. It carries trace_id and span_id only when a log bridge injects active context or your filelog pipeline parses those fields.
Stable
Signal 02
Metrics
Aggregated numbers over time. Counters, histograms, gauges. Low storage cost, always on. Powers dashboards and alerts. Optional exemplars preserve a few representative trace links. Plays well with Prometheus via scrape, remote write, or OTLP ingest.
Stable
Signal 03
Traces
One request, broken into spans across service boundaries. Shows where time went. Best signal for debugging latency in distributed systems.
Stable
Signal 04
Profiles
CPU flame graphs at the code level. Where your service actually spends cycles. Traces tell you a function was slow. Profiles tell you which line made it slow.
Alpha
Not a signal
RUM / browser
Real user monitoring is a use case, not a fifth signal type. Browser instrumentation is still experimental. Browser traces, and supported metrics, use OTLP/HTTP because browsers cannot use the OTLP gRPC transport; browser log support is not yet stable.
OTLP/HTTP
Why it matters
Shared context
Semantic conventions define standard attribute and metric names across signals. http.server.request.duration can mean the same thing in every service, but no Collector enforces that your teams used the convention correctly.
By design

The Collector: where the pipeline work happens

An SDK can export directly to a backend, but the production architecture in this article sends to the OTel Collector, a standalone service between applications and backends. That extra hop gives the platform team one place to enrich, filter, sample, retry, and reroute telemetry without redeploying every service.

Collector pipeline model
Receivers
Data in
OTLP gRPC/HTTP, Prometheus scrape, Jaeger, Fluent Forward, Kafka, cloud integrations. This is how telemetry enters the pipeline.
→
Processors
Transform
Batch, sample, enrich with K8s attributes, redact PII, filter health check noise, add region labels. All without touching app code.
→
Exporters
Data out
Datadog, Splunk HEC, Prometheus, Tempo, Loki, OTLP. Fan out to multiple backends from one pipeline. Traces and metrics one way, logs another.

Put pipeline-wide enrichment in the Collector, not in application code. Applications should still identify themselves with OTEL_SERVICE_NAME; otherwise SDKs fall back to an unknown_service value. The Collector can then add k8s.pod.name, deployment.environment.name, and k8s.namespace.name as resource attributes. Developers do not have to repeat those values on every span.

The Kubernetes shape: agent, then gateway, then vendors

In this design, apps send telemetry to a Collector on the same node rather than directly to Datadog or Splunk. That agent enriches and forwards to a cluster gateway. The gateway is the only component that holds vendor credentials, samples traces, and splits signals: traces and metrics to Datadog, logs to Splunk. Two Collector configs can replace separate vendor agents for these application signals, though you still need to account for any infrastructure or security features those agents used to provide.

This design uses both OTLP transports. OTLP/gRPC on :4317 is a common choice for server SDKs and Collector-to-Collector hops. OTLP/HTTP on :4318 works for browsers, some serverless runtimes, curl, and HTTP-only ingress paths. Both use the same protobuf data model, and language defaults differ, so configure the SDK protocol rather than assuming it. The gateway and node agent listen on both; this article chooses gRPC from agent to gateway.

OTLP/gRPC :4317 OTLP/HTTP :4318 stdout / filelog
Kubernetes OpenTelemetry architecture: node agents, cluster gateway, Datadog APM, Splunk logs Browser telemetry sends OTLP over HTTP through Ingress to the cluster gateway. Application pods send OTLP over gRPC to a DaemonSet node agent on the same node, which also tails container logs. Node agents forward all signals over gRPC to a cluster gateway Deployment. The gateway exports traces and metrics to Datadog and logs to Splunk. Correlated logs preserve the trace ID across both backends. Browser RUM outside the cluster OTLP/HTTP :4318 KUBERNETES CLUSTER Ingress HTTP only · terminates TLS NODE A checkout-service SDK · OTLP/gRPC inventory-service SDK · OTLP/gRPC hostIP:4317 · stdout Node Agent DaemonSet · one pod per node recv :4317 gRPC · :4318 HTTP · filelog k8sattributes · batch · forward NODE B payment-service SDK · OTLP/gRPC postgres client spans, not OTLP hostIP:4317 stdout Node Agent DaemonSet · same image tails /var/log/pods on this node k8sattributes · batch · forward OTLP/gRPC :4317 · otel-gateway.observability.svc CLUSTER GATEWAY · DEPLOYMENT RECEIVE OTLP/gRPC :4317 from agents OTLP/HTTP :4318 from Ingress no kubelet access required PROCESS memory_limiter · batch APM stats · sample traces route by signal, not by team EXPORT datadog · traces + metrics splunk_hec · logs vendor creds live only here HTTPS HTTPS Datadog APM · traces · metrics datadog exporter · site + API key Splunk logs · HTTP Event Collector splunk_hec · token + index shared trace_id when logs carry it

Figure 1. Two Collector tiers. SDKs and container logs stay on the node. The gateway is the only hop that talks to vendors. Solid lines are OTLP/gRPC :4317. Dashed lines are OTLP/HTTP :4318. Dotted is stdout tailed by filelog, not an OTLP call from the app.

Hop Protocol Why this one
App SDK → node agent OTLP/gRPC :4317 Chosen here for server workloads. Set the SDK protocol explicitly and point OTEL_EXPORTER_OTLP_ENDPOINT at the node's hostIP. The DaemonSet manifest must expose that port with hostPort or host networking. Exporter queues provide in-memory buffering; durable buffering needs a file-backed queue.
Browser RUM → Ingress → gateway OTLP/HTTP :4318 Browsers cannot use the OTLP gRPC transport. Put a reverse proxy in front of the public receiver, restrict CORS, allow the endpoint in CSP, rate-limit it, and terminate TLS there. This path never hits a node agent.
Container logs → node agent filelog (local files) Apps write stdout. The agent tails /var/log/pods on that node. Do not make every service emit logs over OTLP unless you have a reason.
Node agent → gateway OTLP/gRPC :4317 The choice in this design, not a protocol requirement. High volume, in-cluster. otel-gateway.observability.svc:4317.
Gateway → Datadog datadog exporter, HTTPS Traces and application metrics go to Datadog. The Datadog connector computes APM trace metrics before sampling. Vendor credentials live only on the gateway, not on every node or in app env vars.
Gateway → Splunk splunk_hec, HTTPS Logs only. A log bridge or parser must populate the OTEL LogRecord trace field; the Splunk HEC exporter then writes it as trace_id.

Node agent: enrich and forward

One DaemonSet, one pod per node, with hostPort 4317 and 4318 so local pods can reach it through the downward API status.hostIP. k8sattributes watches the Kubernetes API and associates incoming telemetry with a pod, by connection IP here. Running it on the node is convenient because the agent receives the application's connection directly; physical proximity to the kubelet is not what makes enrichment work. resourcedetection with the k8snode detector only sees the node. It will not stamp k8s.pod.name or k8s.namespace.name for you.

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: "0.0.0.0:4317"
      http:
        endpoint: "0.0.0.0:4318"
  filelog:
    include: ["/var/log/pods/*/*/*.log"]
    exclude: ["/var/log/pods/*/otel-agent/*.log"]
    include_file_path: true
    operators:
      - type: container

processors:
  memory_limiter:
    check_interval: 1s
    limit_percentage: 75
    spike_limit_percentage: 15
  k8sattributes:
    auth_type: serviceAccount
    extract:
      metadata: [k8s.namespace.name, k8s.pod.name, k8s.deployment.name, k8s.node.name]
    pod_association:
      - sources:
          - from: connection
  filter:
    error_mode: ignore
    traces:
      span:
        - 'attributes["url.path"] == "/healthz" or attributes["url.path"] == "/readyz"'
    logs:
      log_record:
        - 'IsMatch(body, "healthz") or IsMatch(body, "readyz")'
  batch:
    timeout: 5s

exporters:
  otlp:
    endpoint: "otel-gateway.observability.svc:4317"
    tls:
      insecure: true   # in-cluster; mTLS in production

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, k8sattributes, filter, batch]
      exporters: [otlp]
    metrics:
      receivers: [otlp]
      processors: [memory_limiter, k8sattributes, batch]
      exporters: [otlp]
    logs:
      receivers: [otlp, filelog]
      processors: [memory_limiter, k8sattributes, filter, batch]
      exporters: [otlp]

This is the Collector config, not the whole DaemonSet. The pod still needs the host ports, log directory mounts, RBAC for Kubernetes metadata, resource limits for memory_limiter, and a downward-API value that applications can use as their endpoint. The OTLP exporter retries and usually queues in memory. If telemetry must survive an agent restart, add the file_storage extension and point the exporter's sending queue at a persistent volume.

Gateway: sample, route, export

The gateway is a Deployment. It does not scrape nodes and it does not talk to the kubelet. It receives already-enriched OTLP, computes Datadog APM statistics from the complete trace stream, samples the traces it stores, and fans out by signal. Datadog gets sampled traces, APM statistics, and application metrics. Splunk gets logs. Separate pipelines isolate ordinary backend failures, though queue exhaustion or memory pressure in the shared Collector process can still affect every signal.

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: "0.0.0.0:4317"
      http:
        endpoint: "0.0.0.0:4318"

processors:
  memory_limiter:
    check_interval: 1s
    limit_percentage: 75
    spike_limit_percentage: 15
  probabilistic_sampler:
    sampling_percentage: 10
  batch:
    timeout: 5s

connectors:
  datadog/connector:
    traces:
      compute_stats_by_span_kind: true

exporters:
  datadog:
    api:
      site: datadoghq.com
      key: ${env:DD_API_KEY}
  splunk_hec:
    token: ${env:SPLUNK_HEC_TOKEN}
    endpoint: https://http-inputs.example.splunkcloud.com/services/collector
    source: otel
    sourcetype: otel
    index: k8s_logs

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [datadog/connector]
    traces/sampled:
      receivers: [datadog/connector]
      processors: [probabilistic_sampler, batch]
      exporters: [datadog]
    metrics:
      receivers: [otlp, datadog/connector]
      processors: [memory_limiter, batch]
      exporters: [datadog]
    logs:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [splunk_hec]

The connector placement is deliberate. Datadog APM statistics must be computed from the full trace stream before the probabilistic sampler drops spans. Since Collector Contrib 0.95.0, the exporter no longer computes those statistics by default; without the connector, services and trace metrics can be missing from Datadog APM.

Do not skip the agent

A single gateway Deployment that every pod hits directly can work, especially at smaller scale. The agent tier earns its keep when you need node-scoped log tails, host-local receive, and a buffer on each node. The gateway is for centralized sampling, routing, and vendor credentials. Keep those jobs apart once their operational value justifies the second tier.

How signals actually connect

Context propagation is what makes distributed tracing work. When a request crosses a service boundary, W3C trace context travels in HTTP headers, gRPC metadata, or message properties for systems such as Kafka and SQS. Each service extracts the incoming context and links its own spans to the same trace. Drop that context and your trace breaks at the boundary. What you see is a flat list of disconnected spans instead of an end-to-end waterfall.

Signal correlation - how logs, metrics and traces share context Three boxes showing service A calling service B calling a database, with trace spans connecting them and a metric and log linked via shared trace_id checkout-service service A traceparent payment-service service B client span postgres db layer TRACE SPANS span: POST /checkout 320ms span: charge_card 240ms span: SELECT 185ms METRIC EXEMPLAR http.server.request.duration trace_id=a3c1f1e2b6c74c7d value=0.32s | sampled observation LOG WARN: slow db query detected trace_id=a3c1f1e2b6c74c7d span_id=5b9d6f0c4a1d3e92 same trace_id

Figure 2. Spans carry trace context. Correlated logs can carry the same trace and span IDs. Aggregated metrics do not have one trace ID, but an exemplar can preserve a representative measurement's IDs so a chart links to a trace. From that trace, use the same ID to find its Splunk events.

Two backends is normal. Dual instrumentation is not.

Datadog and Splunk are not the same kind of backend. The Datadog exporter sends traces and metrics over Datadog's APIs. The Splunk HEC exporter sends log events through HEC, which is not OTLP. The Collector speaks both, which is how this split works without instrumenting the application twice. Each exporter still maps the OTEL data model into a vendor-specific store. Your job is to make sure trace_id, service.name, and the semantic conventions remain searchable after both mappings.

"OTEL-native" and "OTLP compatible" are marketing labels, not conformance levels. Every backend stores data in its own model, and accepting OTLP does not prove that exemplars, resource attributes, schema URLs, or trace IDs survive ingestion. Most enterprises do not get a greenfield stack. They get Datadog for APM and Splunk for logs. Fine. Use one application instrumentation path and test the two exporter mappings. Keep a vendor agent only when it still provides infrastructure, security, or product features your Collector setup does not replace.

Test What must survive Failure you are looking for
Trace lookup Full trace ID and span relationships Truncated, reformatted, or unsearchable IDs
Log correlation LogRecord trace_id and span_id IDs left inside an unparsed body or renamed unexpectedly
Metric correlation Exemplars and their trace links Aggregation retained but exemplars dropped
Resource identity service.name, environment, Kubernetes attributes Attributes flattened, renamed, or indexed at unusable cardinality
Backend swap The same OTLP fixture through a second exporter Application code or vendor-specific SDK required to preserve key features

Accepting OTLP is a checkbox. Keeping trace_id intact through the Datadog exporter and the Splunk HEC exporter is the actual requirement. Prove it with one known trace before you call the migration done.

OTEL and Prometheus are not a competition

If you already run Prometheus, you're not replacing it. Prometheus is a metrics database and scraper. OTEL is a collection and forwarding framework. They're complementary. Three ways they work together in practice:

01
Collector scrapes existing endpoints. Already exposing /metrics? The OTel Collector can use its Prometheus receiver to scrape them, add resource attributes, and forward the converted metrics. Watch cardinality when resource attributes become backend labels.
02
Push metrics into Prometheus. The Collector can use Prometheus Remote Write, or Prometheus can receive OTLP metrics at /api/v1/otlp/v1/metrics. OTLP ingest arrived experimentally in v2.47 and remains disabled until you set --web.enable-otlp-receiver. Existing dashboards survive only if your translation strategy preserves the metric names they query.
03
Use an SDK Prometheus exporter where your language supports one. Instrument with the OTEL SDK, expose a /metrics endpoint in Prometheus format, and let Prometheus keep scraping. You gain standardized instruments and attributes, subject to the OTEL-to-Prometheus translation.

Standardized metric names are the practical win. One team calls it latency_s, another calls it response_time. Same measurement, different names, broken federation queries. The HTTP semantic conventions define http.server.request.duration as a histogram with unit s (seconds, not milliseconds). Prometheus still translates OTLP names by default, so configure its OTLP translation_strategy deliberately if dashboards need the dotted name unchanged.

Ten things that matter in production

01
Always set service.name. Without it, SDKs fall back to unknown_service, sometimes with the executable appended. Set OTEL_SERVICE_NAME in the workload or configure the SDK resource explicitly. Do not try to reconstruct service identity later from pod names.
02
Start with auto-instrumentation where the language supports it. Java and .NET agents, Python and Node packages, and the OpenTelemetry Operator can cover common HTTP, DB, and gRPC libraries with little or no application code. An SDK by itself does not auto-instrument anything. Get the supported baseline first, then add manual spans for business logic.
03
Use two Collector tiers when the jobs require it. A DaemonSet agent handles node-local logs and host-local receive. A gateway centralizes sampling, routing, and vendor credentials. Smaller clusters may not need both, but do not put the Datadog API key in every application.
04
Know what head sampling cannot do. A 10% probabilistic sampler makes its decision before the request outcome is known, so it also drops about 90% of errors. Keeping every error requires tail sampling or an application decision made with outcome information. Tail sampling needs every span for a trace routed to the same gateway, usually with the load-balancing exporter keyed by traceID.
05
Use histograms for latency, not gauges. An average latency gauge hides the tail that's hurting users. Histograms preserve bucketed or exponential distributions from which a backend estimates percentiles. Pick boundaries and aggregation settings that retain the resolution your SLO needs.
06
Verify context across async boundaries. Instrumentation for Kafka, SQS, or Redis may inject and extract context automatically, but library and language coverage varies. Confirm that W3C trace context reaches message properties or attributes. Add it manually only where instrumentation does not.
07
Use semantic conventions before inventing attributes. If a stable convention exists for it, such as db.system.name, http.request.method, or k8s.namespace.name, use the standard. Your dashboards are more likely to survive library and team changes.
08
Drop health check noise at the agent. Kubernetes liveness probes can generate thousands of low-value spans and logs. The agent config above filters /healthz and /readyz before they cross the cluster network. Keep probes whose failures carry diagnostic value.
09
Version your Collector config in Git. Treat it like application code. PRs for changes, CI validation, rollback capability. The Collector is infrastructure. Teams that don't version it end up with undocumented config drift that takes down the pipeline.
10
Two vendors, one trace_id. Datadog for APM and Splunk for logs is a normal split. First make sure the application log bridge or filelog parser puts the ID into the OTEL LogRecord trace field. Then confirm both exporters preserve a searchable representation of it.

Build your first OTEL pipeline

Reading about pipelines only gets you so far. The interactive builder below walks you through five steps: pick how telemetry gets in, what happens in transit, where it goes, then read the collector config you built and fire live metrics through it. No Docker, no cluster, no YAML files. Each component has a plain-english label so you're not guessing what a receiver does.

Interactive

The guided pipeline builder is embedded below. Build a pipeline, break it, watch metrics flow through your receivers, processors, and exporters, and see them land on a live dashboard.

GitHub repo

Want to deploy the real thing? The anipublik/OTEL repo on GitHub has working configurations, Collector setups, and deployment resources you can use to instrument and ship your own workloads.

Loading pipeline builder…

A few years ago you had to choose between vendor lock-in and fragmented telemetry. With OTEL you can instrument once, own the pipeline, and split backends without rewriting every service. The agent and gateway configs above show one practical Kubernetes shape: gRPC for server and Collector hops, HTTP for the browser path, Datadog for traces and metrics, Splunk for logs. Start with supported auto-instrumentation and add the second Collector tier when its node-local and centralized jobs earn the complexity. The core trace, metric, and log specifications are stable; SDKs, Collector components, browser support, and profiles mature independently. The platform team's job is still making adoption easy enough that teams actually ship it.