Gardener Shoot Autoscaling: Where Worker Pool Bounds Meet Cluster-Autoscaler Reality
Gardener shoot autoscaling hinges on worker-pool bounds, expander strategy, and conservative scale-down defaults. This post shows how to configure priority-based pool selection, tune scale-down delays, and verify behavior before production surprises hit the bill.
30 Aug 2026, 16:14 UTC

The problem: nodes that won't scale down (or up) when you expect
You've defined a Gardener Shoot with two worker pools—one for general workloads, one for GPU jobs. The GPU pool has min: 0, max: 5. A training job lands, pods go Pending, but the pool stays at zero. Meanwhile the general pool scales up instead, and the GPU pods never schedule. Or the reverse: the job finishes, but nodes linger for hours because a single PodDisruptionBudget blocks drain. Both scenarios stem from the same root: the interaction between Gardener's worker-pool bounds and the upstream cluster-autoscaler flags that Gardener surfaces.
Worker-pool bounds are your primary cost lever
Each workerPool in the Shoot spec declares minimum and maximum node counts. The cluster-autoscaler—deployed per shoot control plane—will never exceed those bounds. That makes pool sizing the first engineering decision: minimum = committed spend, maximum = risk ceiling.
A common pattern is a "burst" pool with minimum: 0 for spot or GPU instances. But minimum: 0 only works if the underlying cloud provider and machine type support scaling from zero. On some providers (or with certain instance families) the MachineDeployment controller refuses to create the first node, leaving pods Pending indefinitely. Verify this in a test shoot before committing to production.
Expander strategy decides which pool grows
When multiple pools can satisfy a pending pod, the autoscaler's expander flag picks the winner. Gardener exposes this via .spec.kubernetes.clusterAutoscaler.expander. The default is random; least-waste and priority are common alternatives.
- least-waste: chooses the node group that leaves the least unused CPU/memory after scheduling the pod. Good for homogeneous pools.
- priority: uses the
priorityfield on each worker pool (higher wins). Use this when you must steer GPU pods to the GPU pool.
If you rely on random (the default), the GPU pool in the opening example loses the coin flip. Set expander: priority and give the GPU pool a higher priority value.
Scale-down is conservative by design
The autoscaler only considers a node for removal after it has been underutilized for scale-down-delay-after-add (default 10m) and scale-down-unneeded-time (default 10m). Even then, a pod with a restrictive PodDisruptionBudget (e.g., minAvailable: 1 on a single-replica deployment) or a pod using emptyDir or hostPath volumes blocks eviction. The node stays, the pool stays at max, and the bill grows.
Gardener lets you tune these delays in the same clusterAutoscaler section:
kubernetes:
clusterAutoscaler:
scaleDownDelayAfterAdd: 5m
scaleDownUnneededTime: 5m
scaleDownUtilizationThreshold: 0.5
expander: priority
Shorter delays reduce waste but increase churn. The utilization threshold (default 0.5) means a node must be below 50% requested resources to be considered unneeded.
Worked example: two-pool shoot with priority expander
The following Shoot fragment shows a general pool (on-demand, min 3) and a GPU pool (spot, min 0, priority 10). Apply it to the garden cluster where the shoot is managed (requires gardener.cloud/shoots write permission).
apiVersion: core.gardener.cloud/v1beta1
kind: Shoot
metadata:
name: ml-training
namespace: garden-dev
spec:
provider:
type: aws
workers:
- name: general
machineType: m6i.xlarge
minimum: 3
maximum: 20
priority: 1
- name: gpu
machineType: g5.xlarge
minimum: 0
maximum: 5
priority: 10
kubernetes:
version: "1.29"
clusterAutoscaler:
expander: priority
scaleDownUnneededTime: 5m
scaleDownDelayAfterAdd: 5m
scaleDownUtilizationThreshold: 0.4
Expected behavior: a pod with nodeSelector: accelerator: nvidia-gpu triggers scale-up of the gpu pool because its priority (10) beats general (1). When the pod completes, the GPU node becomes unneeded; after 5 minutes of <40% utilization it is drained—provided no PodDisruptionBudget or local storage blocks it.
Verification: after deploying a test workload, check the shoot's control-plane namespace in the seed cluster for the cluster-autoscaler pod logs. Look for lines like scale-up: adding node group gpu or scale-down: node ip-xxx unneeded. Also inspect the MachineDeployment replicas in the shoot namespace to confirm the pool's actual size.
Trade-offs and version drift
- Flag drift: Gardener pins a specific
cluster-autoscalerversion per Kubernetes release. Upgrading the shoot's Kubernetes version can change default flags or add new ones. Always diff theclusterAutoscalersection against the release notes. - Control-plane scaling is separate: The shoot's API server and etcd run in the seed cluster. Gardener supports vertical pod autoscaling for those components, but it does not react to worker-node pressure. A chatty 500-node shoot may need manual control-plane resource bumps.
- Minimum-zero surprise: As noted,
minimum: 0can silently fail on some cloud/machine-type combinations. Test each pool type in a non-production shoot.
Actionable next step
Create a test shoot mirroring your production pool topology. Deploy a workload that exceeds current capacity in each pool, then remove it. Confirm:
1. Scale-up targets the intended pool (check autoscaler logs).
2. Scale-down occurs within your configured delay.
3. A pod with a strict PodDisruptionBudget blocks scale-down as expected.
Repeat after any Gardener or Kubernetes upgrade. This five-minute experiment prevents the "why are we paying for 20 idle GPU nodes?" conversation later.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.