For six years, the answer to "this container needs more CPU" was the same regardless of the question: kill the pod, schedule a new one, hope the new number is closer to correct. For a stateless web tier behind a deployment that's a non-event. For a batch job twenty-five minutes into an hour-long run, or a multiplayer server carrying live player state in memory, that cost is very real, and it always lands on the workload, never on whoever guessed wrong about how much CPU it would need.
That changed on December 17, 2025, when Kubernetes 1.35 shipped and in-place pod resize graduated to stable after starting as alpha in 1.27 and beta in 1.33. Kubernetes 1.35 also introduced pod-level resource resize as alpha, and 1.36 promoted it to beta. Kubernetes 1.37, released August 26, 2026, added an alpha mechanism that lets the scheduler preempt lower-priority pods to clear room for a deferred resize.
None of that is the interesting part. The interesting part is that the feature closes the CPU problem almost completely and barely touches the memory problem, and the gap between those two outcomes is exactly where a naive VPA rollout will hurt you.
Where this doesn't apply at all
Before any of the mechanics: does this even apply to your fleet? These read like the fine print at the bottom of the docs, and that's exactly why they get skipped, but they're the most load-bearing part of the whole feature. Check the list against your own workloads before you read another word about how it works.
- Only CPU and memory. No other resource type is resizable this way.
- QoS class can't be touched, ever. Whatever class a pod was born with, it dies with. Guaranteed has to keep requests and limits identical on every resize, forever. Burstable is barred from the one move that would accidentally turn it into Guaranteed, namely requests catching up to limits on both CPU and memory at the same time. BestEffort can never grow a request or limit out of nothing. And no field gets deleted once it exists, you can only swap its value.
- Shrinking a memory limit is a request, not an order. The kubelet checks live usage first, and if usage is already above where you want the new ceiling, it just leaves the resize hanging rather than risk an OOM kill. There's still a narrow window where usage climbs right after that check clears, so best-effort really does mean best-effort here.
- Two container types are locked out entirely: regular init containers and ephemeral debug containers. A sidecar, meaning a restartable init container, is the one exception that gets to resize.
- No Windows support, full stop.
- Pin a pod to a static CPU or memory manager policy and this feature stops seeing it. That's most of your NUMA-sensitive, latency-critical fleet, gone from the eligible list.
- Running on swap blocks memory request changes, unless you've already set memory's resize policy to
RestartContainer. - A memory-backed emptyDir will only resize on cgroup v2 plus its own alpha gate; cgroup v1 nodes reject the request outright. Anything disk-backed, or a persistent volume, doesn't go through this mechanism at all, that's a different resize path entirely.
Roughly half of that list applies even if you're on a single node with no scheduler in the loop. Cleared all of it? Then the mechanics below are worth learning.
What's actually mutable, and who's watching
There are three different answers to "how big is this container," and they disagree on purpose during a resize. spec.containers[*].resources is the ask, now editable for CPU and memory. status.containerStatuses[*].resources is reality, what the container is running with at this second. A quieter third field, status.containerStatuses[*].allocatedResources, holds the values the kubelet has confirmed, and it exists mainly to feed internal scheduling logic. It's not a dashboard field. Reach for it when you're debugging allocation accounting rather than a container's actual footprint.
Node capacity math takes the pessimistic path on purpose: while a resize is still landing, whichever of those three numbers is largest is the one the scheduler bills against that pod. There's no window where you can start a resize and squeeze a second pod onto a node before the first one's new size gets counted.
The only door in is the /resize subresource. An ordinary patch to a pod's spec still leaves resources untouched, on purpose, per KEP-1287. In practice:
kubectl patch pod checkout-worker -n orders --subresource resize --patch \
'{"spec":{"containers":[{"name":"worker","resources":{"requests":{"cpu":"1200m"},"limits":{"cpu":"1200m"}}}]}}'Point an old client at that flag and it'll tell you the subresource doesn't exist, you need v1.32 or newer. And the door only swings on three hinges: a container's own resources, the resources on a restartable init container (a sidecar, specifically, not an ordinary init container), and the resize policy field.
Most RBAC audits I've seen never separate this out, and they should. pods/resize is its own resource string in RBAC, separate from plain pods, because a subresource gets its own line in a Role. Granting patch on pods does not carry patch on pods/resize with it. A grant that lets someone patch a deployment's pod spec does not automatically let them reshape a live node's resource math, unless you've bundled the two together carelessly. Split them, and audit pods/resize the way you'd audit anything that can move capacity around underneath a neighbor.
And getting past the kubelet isn't the finish line. The request still has to clear admission first, and ResourceQuota is explicit about it: the KEP documents quota tallying a resize into the namespace's running total, so a namespace sitting close to its limit can bounce a resize the node itself would've happily granted. The same admission chain also runs LimitRange, which is where you'd expect a per-container floor or ceiling to get enforced too, though the docs don't spell that case out as directly as they do for quota. Either way, that's a reads-as-a-broken-feature moment for whoever's paged if nobody warned them quota was in the loop at all.
You decide the blast radius, per resource
Every container declares its own tolerance for change through resizePolicy, one entry per resource, choosing between NotRequired (apply live, default) and RestartContainer. Two rules worth internalizing before you write policy for real workloads:
- If the pod's own
restartPolicyisNever, no resource is allowed to declareRestartContainer. A pod that can't be restarted can't have a resize policy that requires one, so every resource falls back toNotRequiredwhether you like it or not. - Resize requests that touch two resources at once inherit the stricter policy. If CPU is
NotRequiredand memory isRestartContainer, a request that changes only CPU applies live. A request that changes memory, or both together, restarts the container.
Setting memory to RestartContainer reads, at a glance, like you've disabled the feature for the resource that matters most. For most JVM and CPython services it's actually the honest setting, not a hedge, and the next section is about why that's a structural fact rather than a matter of taste.
The wall that no automation gets around
The most honest line in the entire feature's rollout is a small one buried in the future-work notes: Java and Python runtimes don't support resizing their own memory ceiling without a restart, and it's still an open conversation with the people who maintain those runtimes.
There are actually three separate ceilings stacked on top of each other here, and it's worth pulling them apart instead of treating "memory" as one number. The bottom one is the cgroup limit, the OS-level ceiling the kubelet controls, and that one genuinely moves the moment a resize lands. Above that sits whatever the JVM decided about its own heap at boot, usually an explicit -Xmx or a container-aware calculation done once at startup, and nothing about a live resize goes back and reruns that math. If -Xmx was set at 2Gi, it's still 2Gi after the cgroup limit moves to 4Gi, full stop. Native and off-heap allocations are the murkier third layer, direct buffers, JNI code, whatever a Python C-extension is holding outside the interpreter's own object heap, and those can sometimes actually use the new cgroup headroom, since they're not gated by a JVM-style fixed ceiling the way heap is. CPython specifically doesn't have a clean -Xmx equivalent at all, its allocator and RSS behavior are messier and more workload-dependent than the JVM's, so don't treat "CPython has a fixed ceiling too" as a clean parallel. The safe operating assumption either way: unless you've actually measured that a given process can use extra cgroup room without a restart, treat a memory increase as restart-required.
This is why a memory-focused VPA policy that leans on NotRequired for a JVM service isn't buying you a restart-free resize. It's buying you a slower path to the exact restart you were trying to avoid, minus the part where you got to choose when it happened. RestartContainer for memory isn't a compromise. For these runtimes it's the accurate policy.
What a stuck resize actually looks like
The two pending reasons sit next to each other in a dashboard and mean opposite things. Infeasible is about the node's total capacity: ask a four-core node for eight cores and no amount of patience or evicted neighbors will ever satisfy it. The condition lands on PodResizePending with reason: Infeasible and stays there. That's not a retry loop, it's a dead end, and the fix is to shrink the ask or move the workload somewhere bigger. A human makes that call.
Deferred is about what's free right now. The node could hold the new size, but today's neighbors are in the way, so the kubelet keeps retrying on its own schedule. When several resizes are queued, higher PriorityClass goes first, Guaranteed outranks Burstable at equal priority, and the longest wait breaks the tie. A deferred high-priority resize doesn't block the ones behind it; they all keep getting retried. This one can and often does resolve itself. The alert you actually want isn't "a pod is Deferred," that's normal background noise, it's "a pod has been Deferred for longer than your retry window would explain," which is a judgment call you have to tune per cluster, but a few minutes is a reasonable starting point to page on.
Once the kubelet accepts a request, the condition becomes PodResizeInProgress, and if something goes wrong while it's landing, for example the memory-shrink usage check discussed above, the failure shows up as reason: Error in that same condition's message field, not as a separate event you have to go hunting for.
One field makes the whole sequence legible, and it went stable in the same release: status.observedGeneration reports the metadata.generation the kubelet has actually processed, and each resize condition carries its own observedGeneration as well. If that number lags the pod's current generation, the kubelet simply hasn't caught up to your patch yet. That's a different problem from a resize it has seen and can't satisfy, and before 1.35 there was no reliable way to tell the two apart.
The pod-level envelope has its own trap
Kubernetes 1.35 introduced pod-level resize for spec.resources, an aggregate budget for the whole pod rather than per-container, and 1.36 promoted it to beta. It needs cgroup v2, a handful of feature gates, Linux only, and a CRI runtime new enough to speak the extra call this needs: container-level resize already relied on UpdateContainerResources, and pod-level resize added a second call, UpdatePodSandboxResources, so the runtime has to support both (containerd 2.0+ or CRI-O). The sequencing avoids overshoot on either direction: growing the budget means the outer pod cgroup opens up before any container gets a bigger number, and shrinking flips that order, containers get throttled down first and the outer boundary only narrows once they've already backed off.
Here's the part that bites people: any container that skipped setting its own limit is quietly borrowing the pod's shared ceiling as its own. Nudge the pod-level number and you've just changed that container's real limit too, whether you meant to touch that container or not, and the kubelet logs it exactly like any other resize. So if that particular container happens to have memory set to RestartContainer, turning the pod-level dial restarts a container you never named in your patch. "Pod-level resize skips restarts" only holds up for containers whose own policy happens to agree, and that's a per-container fact you have to go check, not something you get to assume.
Turning this into a VPA policy, not just a feature
In-place resize is the mechanism. The Vertical Pod Autoscaler is what decides when to use it, and it has more modes than most teams bother reading past the first one.
| mode | what it does |
|---|---|
| Off | Computes a recommendation. Applies nothing. |
| Initial | Applies the recommendation only at pod creation. Never touches a running pod. |
| Recreate | Evicts and recreates the pod when the recommendation drifts far enough from the current value. |
| InPlaceOrRecreate | Tries an in-place resize first, falls back to eviction when in-place isn't possible. Beta and on by default since VPA 1.5.0, GA in 1.6.0, and its feature gate was removed in 1.7.0. |
| InPlace | In-place only, never falls back to eviction; defers and retries instead. Alpha as of VPA 1.7.0, needs Kubernetes 1.33+ and gates on both the updater and admission controller. |
| Auto | Deprecated since VPA 1.4.0. Behaves identically to Recreate. Don't design around it. |
Start on InPlaceOrRecreate, not InPlace. InPlace never evicting sounds like the gentler option, but a node that's genuinely full will leave the recommendation sitting in Deferred indefinitely, and a pod that's actually under-provisioned just starves quietly with no forcing function. An eviction under Recreate or InPlaceOrRecreate is noisy on purpose. That noise is the thing that gets someone to look at the node. Reserve InPlace for workloads you've specifically decided must never be evicted, and build alerting on stuck Deferred and Infeasible conditions before you turn it on.
Put together, a JVM service running this policy looks like this. The Deployment carries the resizePolicy split from earlier, the VPA object just points at it and picks the mode:
# Deployment, relevant fields only spec: template: spec: containers: - name: worker resources: requests: { cpu: 500m, memory: 1Gi } limits: { cpu: "1", memory: 2Gi } resizePolicy: - resourceName: cpu restartPolicy: NotRequired - resourceName: memory restartPolicy: RestartContainer --- # VerticalPodAutoscaler apiVersion: autoscaling.k8s.io/v1 kind: VerticalPodAutoscaler metadata: name: worker-vpa spec: targetRef: apiVersion: apps/v1 kind: Deployment name: worker updatePolicy: updateMode: "InPlaceOrRecreate"
Field names and available modes depend on your exact VPA release and cluster feature gates, check both against the versions running in your own cluster before you ship this.
Don't enable InPlacePodVerticalScalingSchedulerPreemption yet. It's alpha as of 1.37 and lets the scheduler evict a lower-priority pod's neighbor specifically to make room for someone else's deferred resize. That's exactly the kind of alpha-stage blast radius you want hardened by someone else's cluster first.
The math problem nobody draws a picture of
HPA's CPU trigger is a fraction, current usage divided by whatever the pod's request field currently says, not divided by some number fixed at deploy time. VPA's whole job, meanwhile, is rewriting that exact denominator on a live pod. Here's what that does concretely: a pod is genuinely using 400m of CPU, sitting on a 500m request, so HPA reads 80% utilization and is getting ready to scale out. VPA looks at the same pod, decides 500m was an underestimate, and bumps the request to 1000m. Nothing about real traffic changed, that 400m of actual usage is identical, but HPA's percentage just dropped to 40% and the scale-out signal vanishes. Run the sequence the other direction and it's worse: VPA shrinking a request makes utilization look artificially high and can trigger a scale-out that real demand never asked for. Put both controllers on the same workload watching the same resource and you've wired one autoscaler's output into the other one's input. Either put HPA on a metric VPA never rewrites, or sit down and trace the loop by hand before you let both run unattended on the same dimension.
In-place resize removes a restart from CPU right-sizing. It does not remove the operational choices around node capacity, runtime memory behavior, quota, or controller ownership. Use InPlaceOrRecreate as the default, require restarts where the process needs one, and keep HPA and VPA from rewriting the same signal.
For the broader rollout path, recommendation model, and production patterns, see Vertical pod autoscaler: stop guessing your resource requests.