When something broke, you were the correlation layer
Before OpenTelemetry, collecting telemetry was a local agreement. Every vendor had a proprietary format. Traces in one tool, metrics in another, logs in a third. Correlation was a manual exercise. When something broke at 2am, you had three browser tabs open, three query languages running, and zero shared context tying them together.
Rolling it out is a platform problem, not a mandate. Product teams will not re-instrument because you published a standard. Give them language-specific auto-instrumentation for HTTP, gRPC, and database calls, plus a Collector pipeline that continues sending APM data to Datadog and logs to Splunk. The Collector receives and processes telemetry; it does not instrument the application for you. Hand teams a working path and a sane default config. Let them break it. Adoption friction matters more than the specification.
OpenTelemetry is the technical piece. It's a CNCF-graduated project, the same graduation tier as Kubernetes, with common APIs, SDKs, semantic conventions, data models, and a transport protocol called OTLP. Instrument once, then send to any backend your Collector distribution has a compatible exporter for. Switching or splitting backends should become a pipeline change instead of an application rewrite. The platform team's job is to make that path the easy one.
OpenTelemetry is the USB-C of observability. One standard connection on the application side, then OTLP or a compatible exporter behind the Collector. Google, Microsoft, AWS, Datadog, and Grafana all support it, but the fidelity of that support still needs testing.
Monitoring tells you error rate is up. Observability lets you ask why and get a real answer. That gap shows up at 2am when the failure mode isn't in any runbook. Monitoring answers questions you wrote down ahead of time. Observability is for the ones you didn't. OTEL supplies the shared resource context, log trace IDs, and metric exemplars that a capable backend can use to move from a spike to a representative trace and then to its logs.
Three stable signals, one emerging signal
Traces, metrics, and logs have stable OpenTelemetry specifications. Profiles reached alpha in 2026 and are not ready for critical production workloads. SDK maturity still varies by language, especially for logs, profiles, and browser instrumentation. The signals share resource attributes such as service.name and deployment.environment.name, but they do not all carry a trace_id: spans do, correlated logs can, and metric data points link to traces only through optional exemplars. Baggage is also part of OpenTelemetry, but it is propagated application context rather than a telemetry stream you store and query like the other signals.
trace_id and span_id only when a log bridge injects active context or your filelog pipeline parses those fields.http.server.request.duration can mean the same thing in every service, but no Collector enforces that your teams used the convention correctly.The Collector: where the pipeline work happens
An SDK can export directly to a backend, but the production architecture in this article sends to the OTel Collector, a standalone service between applications and backends. That extra hop gives the platform team one place to enrich, filter, sample, retry, and reroute telemetry without redeploying every service.
Put pipeline-wide enrichment in the Collector, not in application code. Applications should still identify themselves with OTEL_SERVICE_NAME; otherwise SDKs fall back to an unknown_service value. The Collector can then add k8s.pod.name, deployment.environment.name, and k8s.namespace.name as resource attributes. Developers do not have to repeat those values on every span.
The Kubernetes shape: agent, then gateway, then vendors
In this design, apps send telemetry to a Collector on the same node rather than directly to Datadog or Splunk. That agent enriches and forwards to a cluster gateway. The gateway is the only component that holds vendor credentials, samples traces, and splits signals: traces and metrics to Datadog, logs to Splunk. Two Collector configs can replace separate vendor agents for these application signals, though you still need to account for any infrastructure or security features those agents used to provide.
This design uses both OTLP transports. OTLP/gRPC on :4317 is a common choice for server SDKs and Collector-to-Collector hops. OTLP/HTTP on :4318 works for browsers, some serverless runtimes, curl, and HTTP-only ingress paths. Both use the same protobuf data model, and language defaults differ, so configure the SDK protocol rather than assuming it. The gateway and node agent listen on both; this article chooses gRPC from agent to gateway.
Figure 1. Two Collector tiers. SDKs and container logs stay on the node. The gateway is the only hop that talks to vendors. Solid lines are OTLP/gRPC :4317. Dashed lines are OTLP/HTTP :4318. Dotted is stdout tailed by filelog, not an OTLP call from the app.
| Hop | Protocol | Why this one |
|---|---|---|
| App SDK → node agent | OTLP/gRPC :4317 | Chosen here for server workloads. Set the SDK protocol explicitly and point OTEL_EXPORTER_OTLP_ENDPOINT at the node's hostIP. The DaemonSet manifest must expose that port with hostPort or host networking. Exporter queues provide in-memory buffering; durable buffering needs a file-backed queue. |
| Browser RUM → Ingress → gateway | OTLP/HTTP :4318 | Browsers cannot use the OTLP gRPC transport. Put a reverse proxy in front of the public receiver, restrict CORS, allow the endpoint in CSP, rate-limit it, and terminate TLS there. This path never hits a node agent. |
| Container logs → node agent | filelog (local files) | Apps write stdout. The agent tails /var/log/pods on that node. Do not make every service emit logs over OTLP unless you have a reason. |
| Node agent → gateway | OTLP/gRPC :4317 | The choice in this design, not a protocol requirement. High volume, in-cluster. otel-gateway.observability.svc:4317. |
| Gateway → Datadog | datadog exporter, HTTPS | Traces and application metrics go to Datadog. The Datadog connector computes APM trace metrics before sampling. Vendor credentials live only on the gateway, not on every node or in app env vars. |
| Gateway → Splunk | splunk_hec, HTTPS | Logs only. A log bridge or parser must populate the OTEL LogRecord trace field; the Splunk HEC exporter then writes it as trace_id. |
Node agent: enrich and forward
One DaemonSet, one pod per node, with hostPort 4317 and 4318 so local pods can reach it through the downward API status.hostIP. k8sattributes watches the Kubernetes API and associates incoming telemetry with a pod, by connection IP here. Running it on the node is convenient because the agent receives the application's connection directly; physical proximity to the kubelet is not what makes enrichment work. resourcedetection with the k8snode detector only sees the node. It will not stamp k8s.pod.name or k8s.namespace.name for you.
receivers:
otlp:
protocols:
grpc:
endpoint: "0.0.0.0:4317"
http:
endpoint: "0.0.0.0:4318"
filelog:
include: ["/var/log/pods/*/*/*.log"]
exclude: ["/var/log/pods/*/otel-agent/*.log"]
include_file_path: true
operators:
- type: container
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 75
spike_limit_percentage: 15
k8sattributes:
auth_type: serviceAccount
extract:
metadata: [k8s.namespace.name, k8s.pod.name, k8s.deployment.name, k8s.node.name]
pod_association:
- sources:
- from: connection
filter:
error_mode: ignore
traces:
span:
- 'attributes["url.path"] == "/healthz" or attributes["url.path"] == "/readyz"'
logs:
log_record:
- 'IsMatch(body, "healthz") or IsMatch(body, "readyz")'
batch:
timeout: 5s
exporters:
otlp:
endpoint: "otel-gateway.observability.svc:4317"
tls:
insecure: true # in-cluster; mTLS in production
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, k8sattributes, filter, batch]
exporters: [otlp]
metrics:
receivers: [otlp]
processors: [memory_limiter, k8sattributes, batch]
exporters: [otlp]
logs:
receivers: [otlp, filelog]
processors: [memory_limiter, k8sattributes, filter, batch]
exporters: [otlp]
This is the Collector config, not the whole DaemonSet. The pod still needs the host ports, log directory mounts, RBAC for Kubernetes metadata, resource limits for memory_limiter, and a downward-API value that applications can use as their endpoint. The OTLP exporter retries and usually queues in memory. If telemetry must survive an agent restart, add the file_storage extension and point the exporter's sending queue at a persistent volume.
Gateway: sample, route, export
The gateway is a Deployment. It does not scrape nodes and it does not talk to the kubelet. It receives already-enriched OTLP, computes Datadog APM statistics from the complete trace stream, samples the traces it stores, and fans out by signal. Datadog gets sampled traces, APM statistics, and application metrics. Splunk gets logs. Separate pipelines isolate ordinary backend failures, though queue exhaustion or memory pressure in the shared Collector process can still affect every signal.
receivers:
otlp:
protocols:
grpc:
endpoint: "0.0.0.0:4317"
http:
endpoint: "0.0.0.0:4318"
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 75
spike_limit_percentage: 15
probabilistic_sampler:
sampling_percentage: 10
batch:
timeout: 5s
connectors:
datadog/connector:
traces:
compute_stats_by_span_kind: true
exporters:
datadog:
api:
site: datadoghq.com
key: ${env:DD_API_KEY}
splunk_hec:
token: ${env:SPLUNK_HEC_TOKEN}
endpoint: https://http-inputs.example.splunkcloud.com/services/collector
source: otel
sourcetype: otel
index: k8s_logs
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [datadog/connector]
traces/sampled:
receivers: [datadog/connector]
processors: [probabilistic_sampler, batch]
exporters: [datadog]
metrics:
receivers: [otlp, datadog/connector]
processors: [memory_limiter, batch]
exporters: [datadog]
logs:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [splunk_hec]
The connector placement is deliberate. Datadog APM statistics must be computed from the full trace stream before the probabilistic sampler drops spans. Since Collector Contrib 0.95.0, the exporter no longer computes those statistics by default; without the connector, services and trace metrics can be missing from Datadog APM.
A single gateway Deployment that every pod hits directly can work, especially at smaller scale. The agent tier earns its keep when you need node-scoped log tails, host-local receive, and a buffer on each node. The gateway is for centralized sampling, routing, and vendor credentials. Keep those jobs apart once their operational value justifies the second tier.
How signals actually connect
Context propagation is what makes distributed tracing work. When a request crosses a service boundary, W3C trace context travels in HTTP headers, gRPC metadata, or message properties for systems such as Kafka and SQS. Each service extracts the incoming context and links its own spans to the same trace. Drop that context and your trace breaks at the boundary. What you see is a flat list of disconnected spans instead of an end-to-end waterfall.
Figure 2. Spans carry trace context. Correlated logs can carry the same trace and span IDs. Aggregated metrics do not have one trace ID, but an exemplar can preserve a representative measurement's IDs so a chart links to a trace. From that trace, use the same ID to find its Splunk events.
Two backends is normal. Dual instrumentation is not.
Datadog and Splunk are not the same kind of backend. The Datadog exporter sends traces and metrics over Datadog's APIs. The Splunk HEC exporter sends log events through HEC, which is not OTLP. The Collector speaks both, which is how this split works without instrumenting the application twice. Each exporter still maps the OTEL data model into a vendor-specific store. Your job is to make sure trace_id, service.name, and the semantic conventions remain searchable after both mappings.
"OTEL-native" and "OTLP compatible" are marketing labels, not conformance levels. Every backend stores data in its own model, and accepting OTLP does not prove that exemplars, resource attributes, schema URLs, or trace IDs survive ingestion. Most enterprises do not get a greenfield stack. They get Datadog for APM and Splunk for logs. Fine. Use one application instrumentation path and test the two exporter mappings. Keep a vendor agent only when it still provides infrastructure, security, or product features your Collector setup does not replace.
| Test | What must survive | Failure you are looking for |
|---|---|---|
| Trace lookup | Full trace ID and span relationships | Truncated, reformatted, or unsearchable IDs |
| Log correlation | LogRecord trace_id and span_id |
IDs left inside an unparsed body or renamed unexpectedly |
| Metric correlation | Exemplars and their trace links | Aggregation retained but exemplars dropped |
| Resource identity | service.name, environment, Kubernetes attributes |
Attributes flattened, renamed, or indexed at unusable cardinality |
| Backend swap | The same OTLP fixture through a second exporter | Application code or vendor-specific SDK required to preserve key features |
Accepting OTLP is a checkbox. Keeping trace_id intact through the Datadog exporter and the Splunk HEC exporter is the actual requirement. Prove it with one known trace before you call the migration done.
OTEL and Prometheus are not a competition
If you already run Prometheus, you're not replacing it. Prometheus is a metrics database and scraper. OTEL is a collection and forwarding framework. They're complementary. Three ways they work together in practice:
/metrics? The OTel Collector can use its Prometheus receiver to scrape them, add resource attributes, and forward the converted metrics. Watch cardinality when resource attributes become backend labels./api/v1/otlp/v1/metrics. OTLP ingest arrived experimentally in v2.47 and remains disabled until you set --web.enable-otlp-receiver. Existing dashboards survive only if your translation strategy preserves the metric names they query./metrics endpoint in Prometheus format, and let Prometheus keep scraping. You gain standardized instruments and attributes, subject to the OTEL-to-Prometheus translation.
Standardized metric names are the practical win. One team calls it latency_s, another calls it response_time. Same measurement, different names, broken federation queries. The HTTP semantic conventions define http.server.request.duration as a histogram with unit s (seconds, not milliseconds). Prometheus still translates OTLP names by default, so configure its OTLP translation_strategy deliberately if dashboards need the dotted name unchanged.
Ten things that matter in production
service.name. Without it, SDKs fall back to unknown_service, sometimes with the executable appended. Set OTEL_SERVICE_NAME in the workload or configure the SDK resource explicitly. Do not try to reconstruct service identity later from pod names.traceID.db.system.name, http.request.method, or k8s.namespace.name, use the standard. Your dashboards are more likely to survive library and team changes./healthz and /readyz before they cross the cluster network. Keep probes whose failures carry diagnostic value.trace_id. Datadog for APM and Splunk for logs is a normal split. First make sure the application log bridge or filelog parser puts the ID into the OTEL LogRecord trace field. Then confirm both exporters preserve a searchable representation of it.Build your first OTEL pipeline
Reading about pipelines only gets you so far. The interactive builder below walks you through five steps: pick how telemetry gets in, what happens in transit, where it goes, then read the collector config you built and fire live metrics through it. No Docker, no cluster, no YAML files. Each component has a plain-english label so you're not guessing what a receiver does.
The guided pipeline builder is embedded below. Build a pipeline, break it, watch metrics flow through your receivers, processors, and exporters, and see them land on a live dashboard.
Want to deploy the real thing? The anipublik/OTEL repo on GitHub has working configurations, Collector setups, and deployment resources you can use to instrument and ship your own workloads.
A few years ago you had to choose between vendor lock-in and fragmented telemetry. With OTEL you can instrument once, own the pipeline, and split backends without rewriting every service. The agent and gateway configs above show one practical Kubernetes shape: gRPC for server and Collector hops, HTTP for the browser path, Datadog for traces and metrics, Splunk for logs. Start with supported auto-instrumentation and add the second Collector tier when its node-local and centralized jobs earn the complexity. The core trace, metric, and log specifications are stable; SDKs, Collector components, browser support, and profiles mature independently. The platform team's job is still making adoption easy enough that teams actually ship it.