AniSri

Vertical pod autoscaler: stop guessing your resource requests

Most resources.requests blocks are day-one guesses nobody revisited. VPA turns measured usage into a recommendation, and newer Kubernetes releases can often apply that without killing the pod. Where it fits, where it fights HPA, and which update mode to use.

Insights · Kubernetes · Published June 2024 · Edited August 2025 · Edit Assist: AniBot powered by Claude

Someone typed 500m and 1Gi because it felt safe, shipped it, and never looked again. Multiply that across a few hundred deployments and you are not running a cluster, you are running a museum of day-one anxiety. The vertical pod autoscaler replaces that guess with measurement and keeps re-measuring as the app changes. Tutorials usually stop at a YAML sample and skip the parts that matter in production: the HPA collision, the half hour of silence while the recommender gathers samples, and the 2am eviction.

This is the practical path: what VPA does, where it breaks, and what changed once Kubernetes stopped requiring a restart for every resize.

What VPA actually does

VPA is three controllers shipped as a CRD addon from kubernetes/autoscaler, not a built-in kubelet feature. It watches container CPU and memory over time, computes a percentile-based recommendation, and depending on mode either shows you the number or writes it into the pod spec.

Recommender
Pulls usage from metrics-server, computes a target with headroom above observed peaks
→
Admission controller
Mutates new pods at creation time to carry the current recommendation
→
Updater
Applies recommendations to already-running pods, by eviction or in place, depending on mode

The distinction that trips people up: VPA changes how big each pod is. HPA changes how many pods you have. They solve different problems. Done wrong, they undermine each other. That collision is covered below.

What over-requesting actually costs

On one web tier I worked with, pods requested 500m CPU and 1Gi memory while consuming roughly 50m and 256Mi. That is about 450m CPU and 768Mi memory reserved and unused, per pod, indefinitely, because nothing was measuring the gap.

When it's worth deploying

VPA is not a default-on tool. It's a fit for a specific shape of problem, and a liability for another.

Good fit
  • You genuinely don't know the right requests for a workload
  • Usage drifts over weeks or months as the app or its traffic changes
  • You're paying for reserved capacity nobody uses
  • The workload is stateless and tolerates a restart or a brief resize event
  • You want continuous right-sizing without a human re-running load tests every quarter
Poor fit
  • HPA is already scaling the same resource VPA would control
  • The workload has hard uptime requirements and you're not on in-place resize yet
  • Usage is genuinely flat and predictable, there's nothing to learn
  • It's a bare pod with no owning controller, which VPA won't target

The mode landscape, and why old tutorials are stale

Most VPA write-ups still describe three modes: Off, Initial, Auto. That model is out of date. Auto has been deprecated since VPA 1.4.0. It still works as an alias for Recreate, and newer VPA releases return an API warning when you use it, so set an explicit mode. A VPA created without one now defaults to Recreate.

ModeBehaviorRestarts pods?MaturityUse it for
OffComputes and stores recommendations only, touches nothingNeverStableEvery VPA you deploy, on day one
InitialApplies the current recommendation at pod creation only, through the admission webhookOnly if you replace the podStableProduction workloads not yet on in-place resize
RecreateActively evicts running pods to apply new recommendationsAlways, by evictionStableStateless apps with real PodDisruptionBudgets
AutoDeprecated alias for Recreate, kept for backward compatibilityAlways, by evictionDeprecated since 1.4.0Nothing, migrate off it
InPlaceOrRecreatePatches the running pod's resources directly; falls back to eviction only when an in-place resize isn't possibleOnly on fallbackGA since VPA 1.6.0Stateful or long-lived workloads on Kubernetes 1.33 or newer
InPlacePatches in place and never evicts; if a resize can't be applied yet it waits and retriesNeverAlpha in VPA 1.7.0Experimentation, not production, yet
What changed in Kubernetes 1.35

In-place pod resize (KEP-1287) went GA in Kubernetes 1.35 (December 2025) and the feature gate is now locked on. Before that, applying a VPA recommendation to a running pod meant killing it. The kubelet can now mutate a container's cgroup limits through the /resize subresource, often with no new pod and no restart, subject to each container's resizePolicy. VPA's InPlaceOrRecreate mode rides on that and reached GA in VPA 1.6.0. It is real progress for StatefulSets and other workloads that used to be too disruptive to automate, but it is not magic: it falls back to eviction when the node lacks room, only CPU and memory can be resized at all, and it's prohibited in combination with the static CPU manager and the static memory manager. A pod's QoS class also can't change across a resize.

One asymmetry complicates the split this piece recommends everywhere else. CPU resizes cleanly in both directions. Memory really only resizes upward: a limit decrease is best-effort, and the kubelet skips it and leaves the resize pending when current usage sits above the new ceiling. A process that fixed its heap at startup also can't use headroom that arrives later. So if HPA owns CPU and VPA owns memory, in-place resize helps least with exactly the direction VPA usually drives, which is trimming memory down. For the resize policy rules and what the runtime actually sees, see In-place pod resize: a CPU fix, not a memory one.

Practical order is unchanged even though the mode names are not: start on Off, move to Initial once you trust the numbers, then reach for Recreate or InPlaceOrRecreate only after you have tested eviction behavior against your PodDisruptionBudgets.

Standing one up

The whole point of starting on Off mode is that it's a free, zero-risk capacity audit. It will not touch a single running pod. Point it at a deployment and let it watch.

my-app-vpa.yamlautoscaling.k8s.io/v1
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: my-app-vpa
  namespace: my-namespace
spec:
  # what this VPA watches
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: my-app

  # start here, always. recommendations only, nothing is touched
  updatePolicy:
    updateMode: "Off"

  # guardrails so a future switch to Initial/Recreate can't do something insane
  resourcePolicy:
    containerPolicies:
      - containerName: "*"
        minAllowed:
          cpu: 50m
          memory: 64Mi
        maxAllowed:
          cpu: 2000m
          memory: 4Gi
apply and inspectbash
kubectl apply -f my-app-vpa.yaml

# give it real time before judging anything, see the convergence table below
kubectl describe vpa my-app-vpa -n my-namespace

The output you're looking for is the Recommendation block:

kubectl describe vpa output
Recommendation:
  Container Recommendations:
    Container Name:  my-app
    Lower Bound:
      Cpu:     100m
      Memory:  256Mi
    Target:                  ← this is what VPA recommends
      Cpu:     150m
      Memory:  384Mi
    Upper Bound:
      Cpu:     500m
      Memory:  1Gi

Target is what actually gets written into your pod spec. Lower bound is the smallest request VPA thinks might still work, and the API is explicit that it's not a guarantee: run below it and expect performance or availability to suffer. Upper bound is the largest request worth giving the container, so anything above it is likely waste. The two bounds exist for the updater, which compares a running pod's requests against them to decide whether it has drifted far enough to be worth acting on. The admission controller doesn't use them; it writes target.

There's a fourth field worth knowing about: uncappedTarget, the recommendation as computed from usage alone, before minAllowed and maxAllowed clip it. Compare it against target when you suspect your own guardrails are the thing shaping the number.

Three patterns that hold up in production

01Recommendation-only baseline›

When: you want visibility into what every workload actually needs, with zero chance of disrupting anything. This should be the default state for every deployment in a cluster you care about, whether or not you ever flip it further.

vpa-recommendation-only.yaml
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: my-app
  updatePolicy:
    updateMode: "Off"  # recommendation only

Run this fleet-wide with kubectl get vpa -A -o wide as a monthly ritual and you'll find your worst offenders in about five minutes.

02Memory-only VPA alongside CPU-based HPA›

When: HPA already scales replica count off CPU utilization and you want vertical optimization on top, without the two systems touching the same number.

vpa-memory-only.yaml
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: my-app
  updatePolicy:
    updateMode: "Initial"
  resourcePolicy:
    containerPolicies:
      - containerName: "*"
        controlledResources: ["memory"]  # CPU stays untouched, HPA owns it

This is the safe combination and the default I would use in any cluster running both controllers. VPA owns memory. HPA owns replica count off CPU. Neither writes a field the other reads.

03Bounded automatic sizing›

When: you trust VPA's recommendations for a specific workload and want them applied automatically, but you're not willing to let it pick truly arbitrary values.

vpa-with-boundaries.yaml
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: my-app
  updatePolicy:
    updateMode: "Initial"
  resourcePolicy:
    containerPolicies:
      - containerName: "my-app"
        controlledResources: ["cpu", "memory"]
        minAllowed:
          cpu: 100m    # never go below this
          memory: 256Mi # never go below this
        maxAllowed:
          cpu: 1000m  # never go above this
          memory: 2Gi   # never go above this

minAllowed and maxAllowed are the seatbelt. Without them, a bad metrics window, a memory leak, or a bug in your own app can produce a recommendation nobody would have signed off on manually.

Six ways this breaks in production

1. VPA and HPA fight when they're pointed at the same metric

HPA replica math is roughly desiredReplicas = ceil(currentReplicas × (currentMetricValue / desiredMetricValue)). For CPU utilization scaling, that ratio is measured against the pod's resource request, not an absolute number. The moment VPA rewrites the CPU request, HPA's utilization percentage moves even though load did not. HPA reacts to that phantom signal, scales, changes aggregate usage, and feeds the next VPA recommendation. Two control loops disagree about reality and take turns overcorrecting: replica oscillation, unnecessary churn, and an on-call page that looks like a launch on a quiet Tuesday.

Safe pattern

Split the resources. Let VPA control memory. Let HPA scale on CPU, or better, on a custom metric such as request rate that VPA never touches.

safe split
# VPA manages memory only
resourcePolicy:
  containerPolicies:
    - containerName: "*"
      controlledResources: ["memory"]
---
# HPA scales on CPU only, and never on the field VPA writes to
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
spec:
  metrics:
    - type: Resource
      resource:
        name: cpu

2. VPA needs real time before its recommendation means anything

The recommender ships default settings that keep about 8 days of history per container via VerticalPodAutoscalerCheckpoint objects, but you don't need to wait that long to get a usable number. You do need to wait longer than most people's patience allows.

Time since deployRecommendation quality
15–30 minutesInitial only, based on a handful of samples, treat as rough
~1 hourUsable directionally, still noisy at the edges
24 hoursGood, captures a full daily traffic cycle
48+ hoursBest, spans multiple cycles and smooths one-off spikes

Judging VPA in the first hour and calling it wrong is the most common false start.

3. Whether it restarts your pods depends entirely on the mode you picked

ModeWhen pods changeDisruption
OffNeverNone
InitialOnly on a restart you triggeredLow
RecreateVPA evicts the pod to apply the changeHigh
InPlaceOrRecreatePatched live in most cases, evicted only when in-place isn't possibleLow, occasionally high

Use Initial for production workloads that are not yet on in-place resize. You get automation without a controller that can wake you up.

4. The updater quietly refuses to act on single-replica workloads

By default the updater won't evict a pod unless at least two replicas are live, so resizing a workload never takes it fully offline. On a single-replica deployment in Recreate mode, that safety rail is indistinguishable from VPA being broken: a sensible recommendation, PROVIDED: True, and pods that never change. The cluster-wide default comes from the updater's --min-replicas flag, and since it isn't always reachable on managed control planes, you can override it per object:

vpa-single-replica.yaml
spec:
  updatePolicy:
    updateMode: "Recreate"
    minReplicas: 1  # accept a brief gap in availability

Set that only where you've genuinely accepted the gap. The in-place modes are the better answer for single-replica workloads, because a resize that doesn't restart the container never creates a gap to accept.

5. Bare pods aren't supported, and DaemonSets come with a catch

VPA needs an owning controller. A bare Pod with nothing above it has no selector for VPA to resolve and nothing to recreate it after an eviction, so VPA won't touch it. Everything in the standard controller zoo works, DaemonSets included:

Supported
  • Deployments
  • StatefulSets
  • ReplicaSets
  • ReplicationControllers
  • DaemonSets
  • Jobs and CronJobs
Not supported
  • Bare pods with no owning controller
  • Anything with neither a scale subresource nor a well-known controller kind

The DaemonSet catch is topology, not support. A DaemonSet pod evicted for a resize comes back on the same node, so if that node can't fit the larger request the pod sits pending and you've lost coverage there rather than relocating it. The updater does cap concurrent evictions using the DaemonSet's ready count, but you steer that with its eviction rate flags rather than a PodDisruptionBudget. InPlaceOrRecreate is a much better fit for DaemonSets than Recreate.

6. StatefulSets need a longer look before you automate them

Read this before touching a database

For StatefulSets running databases, queues, or anything with quorum, stick to Initial unless you've specifically validated InPlaceOrRecreate against your failure domain. An eviction-based update on a quorum member at the wrong moment is how a routine resize becomes an incident. Set a real PodDisruptionBudget, and test the restart path in staging before you ever let this run against production data stores.

Verifying it's actually doing something

is it runningbash
kubectl get vpa -n my-namespace

# NAME         MODE      CPU    MEM     PROVIDED   AGE
# my-app-vpa   Initial   150m   384Mi   True       2d

PROVIDED: True means the recommender has computed a recommendation. PROVIDED: False means it's still gathering history or something upstream is broken, see the troubleshooting section below.

full recommendation and confirming it appliedbash
# full detail
kubectl describe vpa my-app-vpa -n my-namespace

# or as structured JSON
kubectl get vpa my-app-vpa -n my-namespace -o jsonpath='{.status.recommendation}' | jq .

# confirm the running pod actually matches the recommendation
kubectl get vpa my-app-vpa -n my-namespace -o jsonpath='{.status.recommendation.containerRecommendations[0].target}'
kubectl get pod <pod-name> -n my-namespace -o jsonpath='{.spec.containers[0].resources.requests}'

If those two numbers match, the recommendation has been applied. If they don't, and you're in Initial mode, the pod hasn't restarted since the recommendation last changed, which is expected, not broken.

Best practices

Guidance
DoStart every VPA on Off. Review recommendations for a few days before switching anything.
DoAlways set minAllowed and maxAllowed. A recommendation with no ceiling is a recommendation you haven't actually reviewed.
DoSplit responsibilities with HPA: VPA for memory, HPA for CPU or a custom metric.
DoPut kubectl get vpa -A -o wide on a recurring cadence. Recommendations drift as usage patterns drift, this isn't a set-and-forget system.
Don'tReach for Recreate or full automation on day one. Earn it with a week of clean Off-mode data first.
Don'tLet VPA and HPA read and write the same resource. That's the oscillation bug described above, and it's entirely avoidable.
Don'tJudge a recommendation inside the first hour. Give it at least 24.

A full walkthrough, with the actual math

Where you start: an over-provisioned deployment

my-app deployment (current)
resources:
  requests:
    cpu: 500m     # guessed on day one
    memory: 1Gi   # also guessed
  limits:
    cpu: 2000m
    memory: 2Gi
what it actually usesbash
$ kubectl top pods -n my-namespace
NAME                          CPU    MEMORY
my-app-5d4b9c8f7b-abc12       50m    256Mi
my-app-5d4b9c8f7b-def34       45m    240Mi
my-app-5d4b9c8f7b-ghi56       55m    270Mi

Average usage sits around 50m CPU and 256Mi memory against a 500m/1Gi request. That's roughly 450m CPU and 768Mi memory wasted per pod. Across three replicas: 1,350m CPU (1.35 cores) and 2,304Mi (~2.25Gi) of memory reserved and doing nothing.

Step 1 · deploy in Off mode, wait 24 hours

kubectl describe vpa my-app-vpa -n my-namespace
Target:
  Cpu:     100m    # 80% reduction from the original 500m request
  Memory:  384Mi   # 62% reduction from the original 1Gi request

That 384Mi includes a real safety margin over the observed 256Mi average, this isn't VPA trimming to the bone, it's leaving room for normal variance.

Step 2 · switch to Initial and trigger a restart

bash
kubectl patch vpa my-app-vpa -n my-namespace --type merge \
  -p '{"spec":{"updatePolicy":{"updateMode":"Initial"}}}'

kubectl rollout restart deployment my-app -n my-namespace

Step 3 · confirm it took

bash
$ kubectl get pod <new-pod-name> -n my-namespace -o jsonpath='{.spec.containers[0].resources.requests}'
{"cpu":"100m","memory":"384Mi"}
-80%
CPU request
-62%
Memory request
1.2 cores
Freed across 3 replicas
~1.9Gi
Memory freed across 3

Before: 1,500m CPU and 3Gi memory reserved across three pods. After: 300m CPU and 1,152Mi memory, same app, same replica count, measured numbers instead of a day-one guess. What you recover is slightly less than the waste figures earlier in this section, because those measured requests against observed usage while the new requests still carry a deliberate margin above it. Repeat this across a few hundred deployments and you recover real headroom without buying another node pool.

FAQ

Will VPA restart my pods?›

Depends entirely on mode: never in Off, only when you replace the pod in Initial, always by eviction in Recreate, and mostly not at all in InPlaceOrRecreate, which still falls back to eviction when the node can't fit the larger request.

Can I run VPA alongside HPA?›

Yes, as long as they never manage the same resource. VPA on memory plus HPA on CPU (or a custom metric) is safe. VPA and HPA both touching CPU is the oscillation bug covered in the failure modes above, and upstream lists it as a known limitation.

Can I run VPA alongside KEDA?›

Yes, cleanly. VPA adjusts resource requests, KEDA scales replica count off event sources like queue depth or HTTP concurrency. They react to different signals entirely, there's no shared field to fight over.

What if a recommendation looks wrong?›

Set minAllowed and maxAllowed in the resource policy rather than fighting the recommender directly. That's what the boundaries are for.

Can I exempt one container, like a sidecar, from VPA?›

Yes. Give that container its own entry in containerPolicies with mode: "Off", and it's left alone while the rest of the pod is still managed.

Does VPA work for Jobs and CronJobs?›

Yes. Use Initial mode so each new job run picks up the latest recommendation, since there's no long-running pod to evict and re-apply against.

What happens after an OOMKill?›

The recommender treats the OOM event as a signal and raises its memory target on the next reconciliation, though upstream is careful to say it reacts to most out-of-memory events rather than all of them. Either way it won't help the container that just died. A container restarting inside an existing pod keeps that pod's original requests, because the admission webhook only fires when a pod is created. Relief arrives when the pod object is actually replaced, or when an in-place mode pushes the new value into the running pod.

Does VPA need Prometheus?›

No. By default the recommender only requires metrics-server and keeps its own history through VerticalPodAutoscalerCheckpoint objects in-cluster. You can optionally point it at a Prometheus history provider so a fresh install converges faster instead of rebuilding history from scratch.

Does this work the same on EKS, GKE, and AKS?›

Mostly. On EKS and AKS you install the standard open-source VPA addon yourself, same YAML as everywhere else. GKE ships a managed, cluster-integrated version you can turn on per-cluster or per-workload without installing anything, worth checking before you deploy your own.

Is VPA safe in a multi-tenant cluster?›

Yes, it's namespace-scoped like any other CRD. Each team can own VPAs for their own workloads without touching anyone else's.

Troubleshooting

Check that the components are actually up

Where they run depends on how you installed VPA. Upstream vpa-up.sh creates three deployments in kube-system with one replica each. Several managed distributions run two of each for leader-election HA, so six pods is equally normal, and Helm charts may put them in a namespace of their own.

bash
kubectl get pods -n kube-system | grep vpa
# expect recommender, updater, admission-controller
# two of each is HA on some distros, not a bug

PROVIDED: False and it won't go away

  • Too new: the recommender hasn't finished its first reconciliation cycle. Give it 15–30 minutes minimum.
  • metrics-server unreachable: check kubectl top pods works at all before blaming VPA.
  • Selector mismatch: confirm targetRef actually resolves to a live Deployment/StatefulSet with running pods.
  • RBAC: the recommender's service account needs read access to metrics APIs and the target workload; check its logs for permission errors.

Recommendations exist but pods never pick them up

  • Confirm updateMode isn't still "Off".
  • In Initial mode, the pod has to be replaced. A container restarting inside the existing pod doesn't go through admission and keeps the old requests.
  • In Recreate mode on a single-replica workload, the updater is refusing to evict by design. See the minReplicas failure mode above.
  • Check the admission webhook is registered: kubectl get mutatingwebhookconfigurations | grep vpa.
  • A frequently missed cause: the admission controller's TLS certificate rotated or expired. If cert rotation silently failed, the webhook stops mutating pods with no obvious error in the pod's own events, only in the admission-controller's logs.

A recommendation looks unreasonably high or low

  • Check whether minAllowed/maxAllowed are clipping it. Compare target against uncappedTarget to confirm.
  • Check controlledResources isn't silently excluding the resource you're looking at.
  • Not enough history yet. See the convergence table earlier in this piece.
  • A memory target that jumped after an OOMKill isn't a bug, it's the recommender doing exactly what it's supposed to.

Too many pods evicting at once in Recreate mode

VPA respects PodDisruptionBudgets when they exist. If you haven't defined one, there's nothing stopping it from evicting more of a deployment at once than you'd like. Set a real PodDisruptionBudget with a sane maxUnavailable before you ever run Recreate against something that matters.

Logs worth reading

bash
# these labels match the upstream manifests
# Helm charts label with app.kubernetes.io/component instead

# recommendation math
kubectl logs -n kube-system -l app=vpa-recommender --tail=50

# eviction and in-place decisions
kubectl logs -n kube-system -l app=vpa-updater --tail=50

# pod mutation at creation time
kubectl logs -n kube-system -l app=vpa-admission-controller --tail=50

Right-sizing is a process problem with a YAML costume. Deploying the controller is easy. The hard part is that most teams have no mechanism to learn their day-one guess was wrong, so the waste compounds for years. Run VPA in Off mode across the fleet even if you never flip it to automatic. The recommendation alone beats most dashboards built to watch the same gap.

If you already treat saturation as a reliability signal, this is what that signal is for: not whether you paged, but whether you were quietly reserving capacity you never used and calling it safety margin. For the broader signal contract, see We have 400 dashboards and still don't know if we're healthy.