AniSri
Case Study: Platform Economics

My take on Build vs Buy after a $13M observability migration

We replaced Datadog and Splunk with a self-hosted LGTM stack, moved more than 100 engineering teams, and cut annual spend by over $13 million. Then we applied the same ownership model to on-call paging and incident communication.

Published: Aug 25, 2026, Edits by AniBot · Author: Ani Sridharan
$13M+
annual savings after migration
100+
engineering teams moved
8 mo
development through production rollout
3
operational platforms brought in-house
The short version
  • The old contracts made growth expensive. More hosts, custom metric series, and log ingest meant a larger bill, whether the extra data helped anyone or not.
  • We moved metrics, logs, traces, and dashboards to Mimir, Loki, Tempo, and Grafana, with object storage underneath and OpenTelemetry at the collection layer.
  • The tooling was not the hard part. Cardinality controls, data tiering, structured logs, team training, ITSM approvals, and ongoing ownership did the real work.
  • AI made prototypes cheaper and faster. It did not remove the cost of maintenance, upgrades, incident response, governance, or staffing.
  • Build won here because the capability was core, the spend was large, and leadership funded a platform team to own it. Those conditions do not apply to every decision.
01

Buy is not the automatic answer anymore

For most of my career, build vs. buy was settled before the spreadsheet opened. A vendor could deliver a mature product in days. Rebuilding observability, paging, or a status page took a team months just to reach basic feature parity. Once salaries, delivery risk, and support were included, buying was usually the sensible answer.

Two parts of that calculation changed. Open source observability platforms became credible at enterprise scale, and AI-assisted development shortened the path from an idea to a working prototype. A platform team can now test an ingestion pipeline, write an OpenTelemetry processor, or model an escalation policy without committing a quarter to discovery.

What changed

AI did not make software free. It made the first serious experiment cheaper. The final decision still depends on maintenance, reliability, adoption, and whether the organization is willing to own the result.

That distinction matters. A prototype can tell you whether a design works. It cannot carry the pager for five years, plan upgrades, write migration guides, or explain the system to the next engineer. The initial build stopped being the main barrier. Long-term ownership did not.

02

What the old model charged us for

Before the migration, we used Datadog for metrics and APM, and Splunk for logs. Both products worked. The problem was how our contracts turned normal engineering growth into recurring cost.

Metrics and APM

Hosts and custom metric series

Every new workload added hosts. High-cardinality labels multiplied custom metric series. A useful tag added by one team could become a large shared cost after broad adoption.

Logs

Ingest volume

Verbose errors, temporary debug logging, and rising traffic all increased ingest. The bill grew even when the extra bytes did not improve detection or diagnosis.

Neither failure came from bad intent. Teams added labels because they wanted better filtering. Engineers raised log levels because they needed evidence during incidents. At enterprise scale, small decisions made across hundreds of services compound quickly.

The deeper issue was that cost followed raw activity rather than useful signal. More services, traffic, and engineers meant higher spend by default. Renewal then started from that larger baseline.

Vendor friction was part of the cost

Pricing comparisons miss the time spent in support queues, waiting on roadmap items, and working around features the team cannot change. We kept hearing the same sentence in postmortems: “If this were ours, we could fix it.” Eventually that became part of the business case.

03

The platform we moved to

We replaced the commercial stack with Loki for logs, Grafana for dashboards, Tempo for traces, and Mimir for metrics. OpenTelemetry collectors handled ingestion and policy enforcement. Object storage carried the durable data instead of a proprietary indexing tier.

The product swap alone would not have produced the result. Savings came from changing the data itself and controlling it before storage.

01

Control metric cardinality

We reviewed labels across the fleet and removed request IDs, customer identifiers, pod names, and other high-cardinality values that did not belong in metrics. Collector rules dropped or aggregated risky labels before ingestion.

02

Match retention to actual use

Recent data stayed hot for fast investigation. Older data moved to lower-cost object storage for compliance and postmortems. We stopped paying the same storage and query cost for every day of retention.

03

Sample traces with intent

Errors and slow requests kept full fidelity. High-volume healthy traffic was sampled more aggressively. This preserved the cases engineers needed without treating every successful request as equally valuable.

04

Standardize structured logs

Services moved from inconsistent free text to JSON with common field names, severity levels, and correlation IDs. Queries became simpler, incident filtering improved, and teams could govern noisy fields at the collector.

Those four changes, combined with leaving per-unit commercial licensing, produced more than $13 million in annual savings. There was no single trick. The result came from many boring controls applied consistently.

04

Why an eight-month migration was still fast

The migration covered more than 100 engineering teams, and moving from development to full production took eight months.

The technology was not the bottleneck. Pipelines, dashboards, and backend services moved faster than expected. Most of the schedule went into change approvals, migration sequencing, team education, and a knowledge base that had to work just as well for the last team as it did for the first.

Discover

Map signals and cost drivers

Inventory collectors, dashboards, labels, retention, ingest paths, compliance needs, and team-specific dependencies.

Prove

Run representative workloads

Test real query patterns, cardinality, failure handling, storage behavior, and operating cost before choosing the design.

Migrate

Move teams in controlled waves

Translate dashboards, validate alerts, train engineers, and keep rollback paths until each team signs off.

Operate

Treat the platform as a product

Publish standards, assign owners, track service health, plan upgrades, and support teams after cutover.

The adoption lesson

If more than one team must change how it works, engineering effort is only part of the estimate. Training, documentation, approvals, support, and trust are delivery work too.

05

What we built next

Once the organization could operate the observability platform, two adjacent decisions became easier: on-call paging and the enterprise status page. Both were recurring subscriptions attached to growth. Both were also narrow enough to own if we respected their reliability requirements.

On-call

A small system with a hard independence rule

Schedules, rotations, escalation policies, deduplication, acknowledgement, and multi-channel delivery were straightforward. The hard requirement was isolation. Paging could not depend on the same platform it monitored, so it used a separate delivery path with few dependencies and regular live tests.

Incident communication

One status surface for the enterprise

The self-hosted status page consumed incident state from the alerting workflow and published the right view to internal or external audiences. We owned the data flow and avoided pricing tied to pages, subscribers, or headcount.

These systems were not chosen because every vendor product was bad. They were chosen because the required workflows were well understood, the organization already had a capable platform team, and recurring licensing exceeded the cost of disciplined ownership.

06

Why we did not buy a cheaper replacement

Replacing one commercial platform with another would have lowered the starting price. It would not have changed the basic model. Usage would grow, the renewal baseline would rise, and another migration would be waiting when the new price stopped working.

The useful question was not, “Which vendor is cheapest this year?” It was, “Is this capability core to how we operate, or is it context we are comfortable renting?” At this scale, observability, paging, and incident communication affected detection time, response quality, sensitive telemetry, and every production team. We treated them as core operational infrastructure.

Decision factorBuy another platformOwn the capability
Cost modelLower entry price, then usage and renewal exposureCompute, storage, delivery, and platform staffing we could model directly
RoadmapVendor priorities and support processChanges ordered by our operational needs
Data controlThird-party processing and contract controlsTelemetry stays inside the architecture we operate
OperationsVendor carries most product maintenanceOur team owns reliability, upgrades, and support
Exit costAnother product and data migrationOpen components, internal expertise, and infrastructure portability

Ownership also changed vendor negotiations elsewhere. A credible in-house option reduces switching fear. That leverage matters, even for products the organization continues to buy.

07

The costs that do not disappear

Build is not free because AI can produce code and configuration quickly. Fast generation can make the wrong system arrive sooner. Without a clear design, the team ends up with several versions of the same escalation logic, inconsistent relabeling rules, or generated configuration that nobody can explain during an incident.

The governance requirement

AI-assisted code needs the same review, tests, ownership, and change control as any other production code. The team also needs a firm scope boundary. Easy generation is not a reason to add every requested feature.

The platform team now owned upgrades, backward compatibility, capacity planning, security patches, documentation, support, and the problem of monitoring the monitoring system. This could not survive as spare-time work. Leadership funded it as a product with an operating model, not as a migration project that ended at cutover.

That is the real threshold. The build case works only when current spend is large enough to fund infrastructure and a durable team, when ownership solves a problem beyond price, and when leadership accepts the ongoing responsibility.

08

A practical build vs. buy test

AI widened the set of options worth evaluating. It did not make build the default. Use the shape of the problem and the operating model to decide.

Build is worth testing when

  • The capability is core to operations or data control.
  • The required feature set is bounded and well understood.
  • The system can run on infrastructure the team already knows.
  • Vendor spend can fund both hosting and a permanent owner.
  • The organization needs control a vendor roadmap cannot provide.
  • A prototype can exercise real workloads, failure modes, and cost.

Buy is usually better when

  • The vendor has years of hard-to-reproduce edge cases.
  • The system is a complex infrastructure ecosystem, not a bounded workflow.
  • Internal demand or spend is too small to support a platform team.
  • Reliability depends on capabilities the organization cannot staff.
  • The product is useful context but not a source of control or advantage.
  • Leadership will fund delivery but not long-term maintenance.
Prototype the decision, not just the happy path

Test upgrades, rollback, backup and restore, noisy tenants, high-cardinality data, dependency loss, support workflow, and the pager. A demo proves that software runs. It does not prove that a team can operate it.

09

What I would carry into the next decision

  • 01Start with core versus context. Price matters, but control, data, roadmap, and operating responsibility decide whether ownership is useful.
  • 02Use AI to reduce discovery cost. Build a real prototype quickly, then judge it with production constraints instead of treating generated code as the finished product.
  • 03Model the team, not only the infrastructure. Hosting can be cheap while support, upgrades, adoption, and incident response remain expensive.
  • 04Fix the data before changing the backend. Cardinality, retention, sampling, and logging standards did more than a product swap could have done alone.
  • 05Count organizational work as delivery work. More than 100 teams moved because training, documentation, approvals, and migration support were planned from the start.
  • 06Do not build a distributed database to prove a point. Bounded operational workflows were good candidates here. Deep infrastructure products with hidden edge cases often are not.

Build vs. buy is still a question about ownership. The change is that teams can now test the build option before a long planning cycle makes the decision for them. For this organization, observability, on-call paging, and incident communication were core. We built them, staffed them, and accepted the pager that came with them.