AniSri
Anirudh Sridharan

Anirudh Sridharan (Ani)

Truck under starry winter sky

Hi, I'm Ani

I've spent 10+ years leading platform and reliability engineering at Fortune 500 companies and high-growth startups: 30+ platform transformations, legacy stacks to modern cloud. Now I build self-healing infrastructure and AI reliability agents that prevent the 2am page instead of heroically answering it.

Beyond the Code

When I'm not architecting systems, I'm chasing horizons in one of the 48 states, often by backroads, camera in hand, exploring trails, and sampling craft beers. The road teaches patience, perspective, and the value of the scenic route.

Exploring the hidden gems in America's backroads and scenic bywaysPhotographyCraft Beer

Ask me about cloud, AI, road adventures, or how to make complex systems feel effortless.

30+
Platform TransformationsOld school to modern tech
$8.0M+
Verified FinOps Savings
~65%
Incident Reduction
10+
Years at Fortune 500s

Capabilities

All capabilities

Engineering Leadership

Built high performing teams from scratch and scaled multi team initiatives across infrastructure operations, observability, and systems reliability.

Architected SLO driven frameworks and predictive reliability practices that balanced delivery velocity with system resilience, establishing scalable automation and incident management standards that reduced operational toil and improved platform stability.

Hands-on Engineer

Designing and shipping infrastructure as code, Kubernetes platforms, and backend services; debugging production systems.

I work with Terraform and Helm, CI/CD, and TypeScript/Go to turn runbooks into automated, platformized workflows.

Thought Leadership

Practicing customer obsessed engineering, using real user journeys to shape reliability, CX, and operational guardrails.

I bridge non technical needs to engineering decisions with clear narratives and decision memos, aligning teams on measurable outcomes.

Platform Engineering

Architecting and building scalable internal developer platforms with golden paths, self service infrastructure, and CI/CD at scale.

Design platform abstractions that accelerate delivery, enforce guardrails, and reduce cognitive load for product teams.

Systems Reliability & Operations

Defining SLIs/SLOs, capacity planning, and chaos/performance testing to make quiet oncall a first class outcome.

Actionable observability with OpenTelemetry and service mesh visibility, with policy backed SLOs and guardrails.

Operations & Incident Management

Leading incident command, tuning escalation policies, and using correlationID tracing to accelerate root cause analysis.

Reduce alert noise, codify communication templates, and turn retrospectives into durable engineering improvements.

Reliability at Scale

Delivering multi region architectures and high throughput telemetry pipelines with four nines availability targets.

I manage large fleets, forecast capacity with data, and design safe failover and graceful degradation strategies.

AI-Native Operations

Leading teams in building production grade AI agents for reliability workflows, MCP servers for observability platforms, and LLM powered incident analysis that turns tribal knowledge into instant insights.

YODA (AI ops agent), secure MCP servers for New Relic, Datadog, and PagerDuty, plus automated RCA that compresses days of investigation into minutes. AI-first observability correlates signals across millions of metrics; autonomous remediation resolves 70%+ of recurring incidents; intelligent capacity planning optimizes spend while ensuring headroom; and ChatOps agents reduce MTTR by 50%+ through automated runbook execution. Led an 8-person observability team adopting AI-native practices, and established guardrails for safe, compliant agent deployment in production. Delivered $3M+ annualized impact through intelligent automation roadmaps.

AI-First Observability

Using LLMs and ML models to correlate signals across millions of metrics, traces, and logs to surface the few things that actually matter.

Connects telemetry to service ownership, change events, and runbooks so the right engineer sees the right context instantly.

Autonomous Remediation

Building self-healing automation that resolves 70%+ of recurring incidents safely and consistently.

Designed for safety: guardrails, approvals, blast-radius controls, and post-action validation.

Intelligent Capacity Planning

Forecasting demand and tuning capacity with ML-driven models to reduce waste while protecting availability.

Aligns performance headroom with cost objectives so reliability and efficiency improve together.

ChatOps & AI Agents

Context-aware assistants that reduce MTTR by 50%+ via automated triage, summaries, and runbook execution.

Turns tribal knowledge into action: guided troubleshooting, safe remediations, and compliant incident notes.

Tech I Get My Hands Dirty With

The platforms and tools I actually use, not just talk about.

AI / ML
OpenAILangChainHugging FacePyTorchAzure Cognitive ServicesCustom LLMsEval harnessesMLOps
Use cases
  • Ship standing objectives, not one-shot prompts: triage, routing, and retrieval that keep working while you are offline.
  • Run evaluation harnesses with wired release gates so model and tool changes do not silently regress.
  • Treat production AI like any other system: observability, blast-radius limits, and rollback paths.
Patterns
  • Evaluation-first: golden datasets, regression checks, and gates that halt bad releases automatically.
  • Prompt and config as code: versioned, reviewed, and deployed through CI like any other change.
  • Guardrails the runtime enforces: quality floors, latency budgets, and stop conditions that do not trust the model to self-report.
Outcomes
  • AI features that behave predictably because done means a check passed, not a confident summary.
  • Lower regression risk when models, prompts, or tools change because release gates catch drift before users do.
  • Runtime behavior you can explain: what it may do, what it may not, and what stops it.

References

What colleagues and managers say!

View All

My Products

AI-native products I'm building, from concept to production.

View All

Interactive Tools

25+ diagnostics, converters & calculators

View All Tools

Latest Insights

Experiments, playbooks & 2AM thoughts

View All