In-place pod resize: a CPU fix, not a memory one
Kubernetes 1.35 made CPU and memory requests mutable on a running pod. CPU can use that. A JVM heap cannot. The resize policy, VPA mode, and HPA collision that follow from that.
Read articleArticles on reliability, leadership, and platform engineering. Lessons from the field and late-night debugging sessions.
9 of 23 articles tagged Platform Engineering
Interactive simulations to practice incident response and observability skills.
Experience a real 2AM incident response: golden signals, distributed traces, log correlation.
Respond to incident →Black Friday checkout outage scenario. Make decisions, get scored on your SRE skills.
Start simulation →Interactive modules on observability and resilience. Percentiles, SLOs, chaos engineering, and more.
Start learning →Filter by tag
ClearKubernetes 1.35 made CPU and memory requests mutable on a running pod. CPU can use that. A JVM heap cannot. The resize policy, VPA mode, and HPA collision that follow from that.
Read articleOTEL looks like a vendor deck until you see the actual pipeline: apps talk to a node agent, the agent talks to a cluster gateway, traces land in Datadog, and logs land in Splunk. One instrumentation path, two exporter mappings.
Read articleA prompt is one instruction you have to be awake to send. A loop is a standing agent that holds your context, chases a goal, and keeps going while you sleep. You already build control loops for everything that matters in production. Here is how to point that instinct at your agents: loop anatomy, shared state files, drift prevention, hard gates, the autonomy ladder, and five platform fleets worth standing up this week.
Read articleStop confusing SLIs, SLOs, and SLAs. This confusion is burning out your engineers. Here's how to get it right.
Read articleA raw, practitioner take on one of the most powerful - and most misused - observability stacks in enterprise. The good, the bad, the ugly, the MLTK deep dive, and the SPL patterns that actually keep platforms healthy.
Read articleTechnical debt compounds faster than you think. Here's how to spot the warning signs and what actually works to pay it down.
Read articleMost requests blocks are day-one guesses that never get revisited. VPA turns usage into a recommendation, and now Kubernetes can apply it without always killing the pod. Where it fits, where it fights HPA, and which update mode to use.
Read articleProtected main and develop, disposable feature and release branches, and a non-SNAPSHOT version string that doubles as the environment gate. How this model works when a change window, not a merge, is what ships code.
Read articleA B2B SaaS journey from chaos to predictable reliability. How we went from 47 P1 incidents per month to 3, and made on-call a non-event.
Read article