Vertical pod autoscaler: stop guessing your resource requests
Most resources.requests blocks are day-one guesses nobody revisited. VPA turns measured usage into a recommendation, and newer Kubernetes releases can often apply that without killing the pod. Where it fits, where it fights HPA, and which update mode to use.
Someone typed 500m and 1Gi because it felt safe, shipped it, and never looked again. Multiply that across a few hundred deployments and you are not running a cluster, you are running a museum of day-one anxiety. The vertical pod autoscaler replaces that guess with measurement and keeps re-measuring as the app changes. Tutorials usually stop at a YAML sample and skip the parts that matter in production: the HPA collision, the half hour of silence while the recommender gathers samples, and the 2am eviction.
This is the practical path: what VPA does, where it breaks, and what changed once Kubernetes stopped requiring a restart for every resize.
What VPA actually does
VPA is three controllers shipped as a CRD addon from kubernetes/autoscaler, not a built-in kubelet feature. It watches container CPU and memory over time, computes a percentile-based recommendation, and depending on mode either shows you the number or writes it into the pod spec.
The distinction that trips people up: VPA changes how big each pod is. HPA changes how many pods you have. They solve different problems. Done wrong, they undermine each other. That collision is covered below.
On one web tier I worked with, pods requested 500m CPU and 1Gi memory while consuming roughly 50m and 256Mi. That is about 450m CPU and 768Mi memory reserved and unused, per pod, indefinitely, because nothing was measuring the gap.
When it's worth deploying
VPA is not a default-on tool. It's a fit for a specific shape of problem, and a liability for another.
- You genuinely don't know the right requests for a workload
- Usage drifts over weeks or months as the app or its traffic changes
- You're paying for reserved capacity nobody uses
- The workload is stateless and tolerates a restart or a brief resize event
- You want continuous right-sizing without a human re-running load tests every quarter
- HPA is already scaling the same resource VPA would control
- The workload has hard uptime requirements and you're not on in-place resize yet
- Usage is genuinely flat and predictable, there's nothing to learn
- It's a bare pod with no owning controller, which VPA won't target
The mode landscape, and why old tutorials are stale
Most VPA write-ups still describe three modes: Off, Initial, Auto. That model is out of date. Auto has been deprecated since VPA 1.4.0. It still works as an alias for Recreate, and newer VPA releases return an API warning when you use it, so set an explicit mode. A VPA created without one now defaults to Recreate.
| Mode | Behavior | Restarts pods? | Maturity | Use it for |
|---|---|---|---|---|
| Off | Computes and stores recommendations only, touches nothing | Never | Stable | Every VPA you deploy, on day one |
| Initial | Applies the current recommendation at pod creation only, through the admission webhook | Only if you replace the pod | Stable | Production workloads not yet on in-place resize |
| Recreate | Actively evicts running pods to apply new recommendations | Always, by eviction | Stable | Stateless apps with real PodDisruptionBudgets |
| Auto | Deprecated alias for Recreate, kept for backward compatibility | Always, by eviction | Deprecated since 1.4.0 | Nothing, migrate off it |
| InPlaceOrRecreate | Patches the running pod's resources directly; falls back to eviction only when an in-place resize isn't possible | Only on fallback | GA since VPA 1.6.0 | Stateful or long-lived workloads on Kubernetes 1.33 or newer |
| InPlace | Patches in place and never evicts; if a resize can't be applied yet it waits and retries | Never | Alpha in VPA 1.7.0 | Experimentation, not production, yet |
In-place pod resize (KEP-1287) went GA in Kubernetes 1.35 (December 2025) and the feature gate is now locked on. Before that, applying a VPA recommendation to a running pod meant killing it. The kubelet can now mutate a container's cgroup limits through the /resize subresource, often with no new pod and no restart, subject to each container's resizePolicy. VPA's InPlaceOrRecreate mode rides on that and reached GA in VPA 1.6.0. It is real progress for StatefulSets and other workloads that used to be too disruptive to automate, but it is not magic: it falls back to eviction when the node lacks room, only CPU and memory can be resized at all, and it's prohibited in combination with the static CPU manager and the static memory manager. A pod's QoS class also can't change across a resize.
One asymmetry complicates the split this piece recommends everywhere else. CPU resizes cleanly in both directions. Memory really only resizes upward: a limit decrease is best-effort, and the kubelet skips it and leaves the resize pending when current usage sits above the new ceiling. A process that fixed its heap at startup also can't use headroom that arrives later. So if HPA owns CPU and VPA owns memory, in-place resize helps least with exactly the direction VPA usually drives, which is trimming memory down. For the resize policy rules and what the runtime actually sees, see In-place pod resize: a CPU fix, not a memory one.
Practical order is unchanged even though the mode names are not: start on Off, move to Initial once you trust the numbers, then reach for Recreate or InPlaceOrRecreate only after you have tested eviction behavior against your PodDisruptionBudgets.
Standing one up
The whole point of starting on Off mode is that it's a free, zero-risk capacity audit. It will not touch a single running pod. Point it at a deployment and let it watch.
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: my-app-vpa
namespace: my-namespace
spec:
# what this VPA watches
targetRef:
apiVersion: apps/v1
kind: Deployment
name: my-app
# start here, always. recommendations only, nothing is touched
updatePolicy:
updateMode: "Off"
# guardrails so a future switch to Initial/Recreate can't do something insane
resourcePolicy:
containerPolicies:
- containerName: "*"
minAllowed:
cpu: 50m
memory: 64Mi
maxAllowed:
cpu: 2000m
memory: 4Gi
kubectl apply -f my-app-vpa.yaml
# give it real time before judging anything, see the convergence table below
kubectl describe vpa my-app-vpa -n my-namespace
The output you're looking for is the Recommendation block:
Recommendation:
Container Recommendations:
Container Name: my-app
Lower Bound:
Cpu: 100m
Memory: 256Mi
Target: ← this is what VPA recommends
Cpu: 150m
Memory: 384Mi
Upper Bound:
Cpu: 500m
Memory: 1Gi
Target is what actually gets written into your pod spec. Lower bound is the smallest request VPA thinks might still work, and the API is explicit that it's not a guarantee: run below it and expect performance or availability to suffer. Upper bound is the largest request worth giving the container, so anything above it is likely waste. The two bounds exist for the updater, which compares a running pod's requests against them to decide whether it has drifted far enough to be worth acting on. The admission controller doesn't use them; it writes target.
There's a fourth field worth knowing about: uncappedTarget, the recommendation as computed from usage alone, before minAllowed and maxAllowed clip it. Compare it against target when you suspect your own guardrails are the thing shaping the number.
Three patterns that hold up in production
01Recommendation-only baseline›
When: you want visibility into what every workload actually needs, with zero chance of disrupting anything. This should be the default state for every deployment in a cluster you care about, whether or not you ever flip it further.
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: my-app
updatePolicy:
updateMode: "Off" # recommendation only
Run this fleet-wide with kubectl get vpa -A -o wide as a monthly ritual and you'll find your worst offenders in about five minutes.
02Memory-only VPA alongside CPU-based HPA›
When: HPA already scales replica count off CPU utilization and you want vertical optimization on top, without the two systems touching the same number.
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: my-app
updatePolicy:
updateMode: "Initial"
resourcePolicy:
containerPolicies:
- containerName: "*"
controlledResources: ["memory"] # CPU stays untouched, HPA owns it
This is the safe combination and the default I would use in any cluster running both controllers. VPA owns memory. HPA owns replica count off CPU. Neither writes a field the other reads.
03Bounded automatic sizing›
When: you trust VPA's recommendations for a specific workload and want them applied automatically, but you're not willing to let it pick truly arbitrary values.
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: my-app
updatePolicy:
updateMode: "Initial"
resourcePolicy:
containerPolicies:
- containerName: "my-app"
controlledResources: ["cpu", "memory"]
minAllowed:
cpu: 100m # never go below this
memory: 256Mi # never go below this
maxAllowed:
cpu: 1000m # never go above this
memory: 2Gi # never go above this
minAllowed and maxAllowed are the seatbelt. Without them, a bad metrics window, a memory leak, or a bug in your own app can produce a recommendation nobody would have signed off on manually.
Six ways this breaks in production
1. VPA and HPA fight when they're pointed at the same metric
HPA replica math is roughly desiredReplicas = ceil(currentReplicas × (currentMetricValue / desiredMetricValue)). For CPU utilization scaling, that ratio is measured against the pod's resource request, not an absolute number. The moment VPA rewrites the CPU request, HPA's utilization percentage moves even though load did not. HPA reacts to that phantom signal, scales, changes aggregate usage, and feeds the next VPA recommendation. Two control loops disagree about reality and take turns overcorrecting: replica oscillation, unnecessary churn, and an on-call page that looks like a launch on a quiet Tuesday.
Split the resources. Let VPA control memory. Let HPA scale on CPU, or better, on a custom metric such as request rate that VPA never touches.
# VPA manages memory only
resourcePolicy:
containerPolicies:
- containerName: "*"
controlledResources: ["memory"]
---
# HPA scales on CPU only, and never on the field VPA writes to
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
spec:
metrics:
- type: Resource
resource:
name: cpu
2. VPA needs real time before its recommendation means anything
The recommender ships default settings that keep about 8 days of history per container via VerticalPodAutoscalerCheckpoint objects, but you don't need to wait that long to get a usable number. You do need to wait longer than most people's patience allows.
| Time since deploy | Recommendation quality |
|---|---|
| 15–30 minutes | Initial only, based on a handful of samples, treat as rough |
| ~1 hour | Usable directionally, still noisy at the edges |
| 24 hours | Good, captures a full daily traffic cycle |
| 48+ hours | Best, spans multiple cycles and smooths one-off spikes |
Judging VPA in the first hour and calling it wrong is the most common false start.
3. Whether it restarts your pods depends entirely on the mode you picked
| Mode | When pods change | Disruption |
|---|---|---|
| Off | Never | None |
| Initial | Only on a restart you triggered | Low |
| Recreate | VPA evicts the pod to apply the change | High |
| InPlaceOrRecreate | Patched live in most cases, evicted only when in-place isn't possible | Low, occasionally high |
Use Initial for production workloads that are not yet on in-place resize. You get automation without a controller that can wake you up.
4. The updater quietly refuses to act on single-replica workloads
By default the updater won't evict a pod unless at least two replicas are live, so resizing a workload never takes it fully offline. On a single-replica deployment in Recreate mode, that safety rail is indistinguishable from VPA being broken: a sensible recommendation, PROVIDED: True, and pods that never change. The cluster-wide default comes from the updater's --min-replicas flag, and since it isn't always reachable on managed control planes, you can override it per object:
spec:
updatePolicy:
updateMode: "Recreate"
minReplicas: 1 # accept a brief gap in availability
Set that only where you've genuinely accepted the gap. The in-place modes are the better answer for single-replica workloads, because a resize that doesn't restart the container never creates a gap to accept.
5. Bare pods aren't supported, and DaemonSets come with a catch
VPA needs an owning controller. A bare Pod with nothing above it has no selector for VPA to resolve and nothing to recreate it after an eviction, so VPA won't touch it. Everything in the standard controller zoo works, DaemonSets included:
- Deployments
- StatefulSets
- ReplicaSets
- ReplicationControllers
- DaemonSets
- Jobs and CronJobs
- Bare pods with no owning controller
- Anything with neither a scale subresource nor a well-known controller kind
The DaemonSet catch is topology, not support. A DaemonSet pod evicted for a resize comes back on the same node, so if that node can't fit the larger request the pod sits pending and you've lost coverage there rather than relocating it. The updater does cap concurrent evictions using the DaemonSet's ready count, but you steer that with its eviction rate flags rather than a PodDisruptionBudget. InPlaceOrRecreate is a much better fit for DaemonSets than Recreate.
6. StatefulSets need a longer look before you automate them
For StatefulSets running databases, queues, or anything with quorum, stick to Initial unless you've specifically validated InPlaceOrRecreate against your failure domain. An eviction-based update on a quorum member at the wrong moment is how a routine resize becomes an incident. Set a real PodDisruptionBudget, and test the restart path in staging before you ever let this run against production data stores.
Verifying it's actually doing something
kubectl get vpa -n my-namespace
# NAME MODE CPU MEM PROVIDED AGE
# my-app-vpa Initial 150m 384Mi True 2d
PROVIDED: True means the recommender has computed a recommendation. PROVIDED: False means it's still gathering history or something upstream is broken, see the troubleshooting section below.
# full detail
kubectl describe vpa my-app-vpa -n my-namespace
# or as structured JSON
kubectl get vpa my-app-vpa -n my-namespace -o jsonpath='{.status.recommendation}' | jq .
# confirm the running pod actually matches the recommendation
kubectl get vpa my-app-vpa -n my-namespace -o jsonpath='{.status.recommendation.containerRecommendations[0].target}'
kubectl get pod <pod-name> -n my-namespace -o jsonpath='{.spec.containers[0].resources.requests}'
If those two numbers match, the recommendation has been applied. If they don't, and you're in Initial mode, the pod hasn't restarted since the recommendation last changed, which is expected, not broken.
Best practices
| Guidance | |
|---|---|
| Do | Start every VPA on Off. Review recommendations for a few days before switching anything. |
| Do | Always set minAllowed and maxAllowed. A recommendation with no ceiling is a recommendation you haven't actually reviewed. |
| Do | Split responsibilities with HPA: VPA for memory, HPA for CPU or a custom metric. |
| Do | Put kubectl get vpa -A -o wide on a recurring cadence. Recommendations drift as usage patterns drift, this isn't a set-and-forget system. |
| Don't | Reach for Recreate or full automation on day one. Earn it with a week of clean Off-mode data first. |
| Don't | Let VPA and HPA read and write the same resource. That's the oscillation bug described above, and it's entirely avoidable. |
| Don't | Judge a recommendation inside the first hour. Give it at least 24. |
A full walkthrough, with the actual math
Where you start: an over-provisioned deployment
resources:
requests:
cpu: 500m # guessed on day one
memory: 1Gi # also guessed
limits:
cpu: 2000m
memory: 2Gi
$ kubectl top pods -n my-namespace
NAME CPU MEMORY
my-app-5d4b9c8f7b-abc12 50m 256Mi
my-app-5d4b9c8f7b-def34 45m 240Mi
my-app-5d4b9c8f7b-ghi56 55m 270Mi
Average usage sits around 50m CPU and 256Mi memory against a 500m/1Gi request. That's roughly 450m CPU and 768Mi memory wasted per pod. Across three replicas: 1,350m CPU (1.35 cores) and 2,304Mi (~2.25Gi) of memory reserved and doing nothing.
Step 1 · deploy in Off mode, wait 24 hours
Target:
Cpu: 100m # 80% reduction from the original 500m request
Memory: 384Mi # 62% reduction from the original 1Gi request
That 384Mi includes a real safety margin over the observed 256Mi average, this isn't VPA trimming to the bone, it's leaving room for normal variance.
Step 2 · switch to Initial and trigger a restart
kubectl patch vpa my-app-vpa -n my-namespace --type merge \
-p '{"spec":{"updatePolicy":{"updateMode":"Initial"}}}'
kubectl rollout restart deployment my-app -n my-namespace
Step 3 · confirm it took
$ kubectl get pod <new-pod-name> -n my-namespace -o jsonpath='{.spec.containers[0].resources.requests}'
{"cpu":"100m","memory":"384Mi"}
Before: 1,500m CPU and 3Gi memory reserved across three pods. After: 300m CPU and 1,152Mi memory, same app, same replica count, measured numbers instead of a day-one guess. What you recover is slightly less than the waste figures earlier in this section, because those measured requests against observed usage while the new requests still carry a deliberate margin above it. Repeat this across a few hundred deployments and you recover real headroom without buying another node pool.
FAQ
Will VPA restart my pods?›
Depends entirely on mode: never in Off, only when you replace the pod in Initial, always by eviction in Recreate, and mostly not at all in InPlaceOrRecreate, which still falls back to eviction when the node can't fit the larger request.
Can I run VPA alongside HPA?›
Yes, as long as they never manage the same resource. VPA on memory plus HPA on CPU (or a custom metric) is safe. VPA and HPA both touching CPU is the oscillation bug covered in the failure modes above, and upstream lists it as a known limitation.
Can I run VPA alongside KEDA?›
Yes, cleanly. VPA adjusts resource requests, KEDA scales replica count off event sources like queue depth or HTTP concurrency. They react to different signals entirely, there's no shared field to fight over.
What if a recommendation looks wrong?›
Set minAllowed and maxAllowed in the resource policy rather than fighting the recommender directly. That's what the boundaries are for.
Can I exempt one container, like a sidecar, from VPA?›
Yes. Give that container its own entry in containerPolicies with mode: "Off", and it's left alone while the rest of the pod is still managed.
Does VPA work for Jobs and CronJobs?›
Yes. Use Initial mode so each new job run picks up the latest recommendation, since there's no long-running pod to evict and re-apply against.
What happens after an OOMKill?›
The recommender treats the OOM event as a signal and raises its memory target on the next reconciliation, though upstream is careful to say it reacts to most out-of-memory events rather than all of them. Either way it won't help the container that just died. A container restarting inside an existing pod keeps that pod's original requests, because the admission webhook only fires when a pod is created. Relief arrives when the pod object is actually replaced, or when an in-place mode pushes the new value into the running pod.
Does VPA need Prometheus?›
No. By default the recommender only requires metrics-server and keeps its own history through VerticalPodAutoscalerCheckpoint objects in-cluster. You can optionally point it at a Prometheus history provider so a fresh install converges faster instead of rebuilding history from scratch.
Does this work the same on EKS, GKE, and AKS?›
Mostly. On EKS and AKS you install the standard open-source VPA addon yourself, same YAML as everywhere else. GKE ships a managed, cluster-integrated version you can turn on per-cluster or per-workload without installing anything, worth checking before you deploy your own.
Is VPA safe in a multi-tenant cluster?›
Yes, it's namespace-scoped like any other CRD. Each team can own VPAs for their own workloads without touching anyone else's.
Troubleshooting
Check that the components are actually up
Where they run depends on how you installed VPA. Upstream vpa-up.sh creates three deployments in kube-system with one replica each. Several managed distributions run two of each for leader-election HA, so six pods is equally normal, and Helm charts may put them in a namespace of their own.
kubectl get pods -n kube-system | grep vpa
# expect recommender, updater, admission-controller
# two of each is HA on some distros, not a bug
PROVIDED: False and it won't go away
- Too new: the recommender hasn't finished its first reconciliation cycle. Give it 15–30 minutes minimum.
- metrics-server unreachable: check
kubectl top podsworks at all before blaming VPA. - Selector mismatch: confirm
targetRefactually resolves to a live Deployment/StatefulSet with running pods. - RBAC: the recommender's service account needs read access to metrics APIs and the target workload; check its logs for permission errors.
Recommendations exist but pods never pick them up
- Confirm
updateModeisn't still"Off". - In
Initialmode, the pod has to be replaced. A container restarting inside the existing pod doesn't go through admission and keeps the old requests. - In
Recreatemode on a single-replica workload, the updater is refusing to evict by design. See theminReplicasfailure mode above. - Check the admission webhook is registered:
kubectl get mutatingwebhookconfigurations | grep vpa. - A frequently missed cause: the admission controller's TLS certificate rotated or expired. If cert rotation silently failed, the webhook stops mutating pods with no obvious error in the pod's own events, only in the admission-controller's logs.
A recommendation looks unreasonably high or low
- Check whether
minAllowed/maxAllowedare clipping it. ComparetargetagainstuncappedTargetto confirm. - Check
controlledResourcesisn't silently excluding the resource you're looking at. - Not enough history yet. See the convergence table earlier in this piece.
- A memory target that jumped after an OOMKill isn't a bug, it's the recommender doing exactly what it's supposed to.
Too many pods evicting at once in Recreate mode
VPA respects PodDisruptionBudgets when they exist. If you haven't defined one, there's nothing stopping it from evicting more of a deployment at once than you'd like. Set a real PodDisruptionBudget with a sane maxUnavailable before you ever run Recreate against something that matters.
Logs worth reading
# these labels match the upstream manifests
# Helm charts label with app.kubernetes.io/component instead
# recommendation math
kubectl logs -n kube-system -l app=vpa-recommender --tail=50
# eviction and in-place decisions
kubectl logs -n kube-system -l app=vpa-updater --tail=50
# pod mutation at creation time
kubectl logs -n kube-system -l app=vpa-admission-controller --tail=50
Right-sizing is a process problem with a YAML costume. Deploying the controller is easy. The hard part is that most teams have no mechanism to learn their day-one guess was wrong, so the waste compounds for years. Run VPA in Off mode across the fleet even if you never flip it to automatic. The recommendation alone beats most dashboards built to watch the same gap.
If you already treat saturation as a reliability signal, this is what that signal is for: not whether you paged, but whether you were quietly reserving capacity you never used and calling it safety margin. For the broader signal contract, see We have 400 dashboards and still don't know if we're healthy.