
Anirudh Sridharan (Ani)

Hi, I'm Ani
I've spent 10+ years leading platform and reliability engineering at Fortune 500 companies and high-growth startups: 30+ platform transformations, legacy stacks to modern cloud. Now I build self-healing infrastructure and AI reliability agents that prevent the 2am page instead of heroically answering it.
Beyond the Code
When I'm not architecting systems, I'm chasing horizons in one of the 48 states, often by backroads, camera in hand, exploring trails, and sampling craft beers. The road teaches patience, perspective, and the value of the scenic route.
Ask me about cloud, AI, road adventures, or how to make complex systems feel effortless.
Capabilities
All capabilitiesEngineering Leadership
Built high performing teams from scratch and scaled multi team initiatives across infrastructure operations, observability, and systems reliability.
Architected SLO driven frameworks and predictive reliability practices that balanced delivery velocity with system resilience, establishing scalable automation and incident management standards that reduced operational toil and improved platform stability.
Hands-on Engineer
Designing and shipping infrastructure as code, Kubernetes platforms, and backend services; debugging production systems.
I work with Terraform and Helm, CI/CD, and TypeScript/Go to turn runbooks into automated, platformized workflows.
Thought Leadership
Practicing customer obsessed engineering, using real user journeys to shape reliability, CX, and operational guardrails.
I bridge non technical needs to engineering decisions with clear narratives and decision memos, aligning teams on measurable outcomes.
Platform Engineering
Architecting and building scalable internal developer platforms with golden paths, self service infrastructure, and CI/CD at scale.
Design platform abstractions that accelerate delivery, enforce guardrails, and reduce cognitive load for product teams.
Systems Reliability & Operations
Defining SLIs/SLOs, capacity planning, and chaos/performance testing to make quiet oncall a first class outcome.
Actionable observability with OpenTelemetry and service mesh visibility, with policy backed SLOs and guardrails.
Operations & Incident Management
Leading incident command, tuning escalation policies, and using correlationID tracing to accelerate root cause analysis.
Reduce alert noise, codify communication templates, and turn retrospectives into durable engineering improvements.
Reliability at Scale
Delivering multi region architectures and high throughput telemetry pipelines with four nines availability targets.
I manage large fleets, forecast capacity with data, and design safe failover and graceful degradation strategies.
AI-Native Operations
Leading teams in building production grade AI agents for reliability workflows, MCP servers for observability platforms, and LLM powered incident analysis that turns tribal knowledge into instant insights.
YODA (AI ops agent), secure MCP servers for New Relic, Datadog, and PagerDuty, plus automated RCA that compresses days of investigation into minutes. AI-first observability correlates signals across millions of metrics; autonomous remediation resolves 70%+ of recurring incidents; intelligent capacity planning optimizes spend while ensuring headroom; and ChatOps agents reduce MTTR by 50%+ through automated runbook execution. Led an 8-person observability team adopting AI-native practices, and established guardrails for safe, compliant agent deployment in production. Delivered $3M+ annualized impact through intelligent automation roadmaps.
AI-First Observability
Using LLMs and ML models to correlate signals across millions of metrics, traces, and logs to surface the few things that actually matter.
Connects telemetry to service ownership, change events, and runbooks so the right engineer sees the right context instantly.
Autonomous Remediation
Building self-healing automation that resolves 70%+ of recurring incidents safely and consistently.
Designed for safety: guardrails, approvals, blast-radius controls, and post-action validation.
Intelligent Capacity Planning
Forecasting demand and tuning capacity with ML-driven models to reduce waste while protecting availability.
Aligns performance headroom with cost objectives so reliability and efficiency improve together.
ChatOps & AI Agents
Context-aware assistants that reduce MTTR by 50%+ via automated triage, summaries, and runbook execution.
Turns tribal knowledge into action: guided troubleshooting, safe remediations, and compliant incident notes.
Tech I Get My Hands Dirty With
The platforms and tools I actually use, not just talk about.
- Ship standing objectives, not one-shot prompts: triage, routing, and retrieval that keep working while you are offline.
- Run evaluation harnesses with wired release gates so model and tool changes do not silently regress.
- Treat production AI like any other system: observability, blast-radius limits, and rollback paths.
- Evaluation-first: golden datasets, regression checks, and gates that halt bad releases automatically.
- Prompt and config as code: versioned, reviewed, and deployed through CI like any other change.
- Guardrails the runtime enforces: quality floors, latency budgets, and stop conditions that do not trust the model to self-report.
- AI features that behave predictably because done means a check passed, not a confident summary.
- Lower regression risk when models, prompts, or tools change because release gates catch drift before users do.
- Runtime behavior you can explain: what it may do, what it may not, and what stops it.
References
What colleagues and managers say!
My Products
AI-native products I'm building, from concept to production.
Interactive Tools
25+ diagnostics, converters & calculators
Latest Insights
Experiments, playbooks & 2AM thoughts