Choosing a Scaling Signal for IBM Cloud Kubernetes Service Node Pools
Node pool autoscaling in IBM Cloud Kubernetes Service is a control loop, not an event handler. The signal you choose — CPU threshold, custom metric, or schedule — determines which failure mode you get.
02 Nov 2025, 23:51 UTC

A node pool that scales up at 09:00 and back down at 09:05
A familiar pattern in IBM Cloud Kubernetes Service (IKS): a node pool with autoscaling enabled sits at its minimum overnight. At the start of the working day, a burst of batch jobs pushes average CPU above the configured threshold. The pool scales out. Ten minutes later the batch finishes, CPU falls, and the pool scales back in. The next morning the same thing happens, except now the scale-out has to wait for a zone to have capacity.
The threshold is rarely the problem. The signal driving the decision usually is.
What IKS is actually reconciling
Node pool autoscaling in IKS is a control loop, not an event handler. The controller periodically compares the desired worker count against the actual count and acts on the difference. IBM's documentation has described that reconcile interval as being on the order of minutes — commonly cited as roughly five. Confirm the current value for your cluster version before designing around it, because two consequences follow:
- A load spike shorter than the reconcile interval can be over before a scale-out decision is made. Autoscaling is a capacity tool, not a latency tool.
- Scale-down is not instantaneous either. Nodes are drained, and pod disruption budgets (PDBs — rules stating how many replicas of a workload may be unavailable at once) can hold a drain open.
Scale-out is also bounded by things outside the autoscaler's control: the minimum and maximum worker counts per zone, whether the chosen instance type is available in that zone, and the regional quota for that instance family. Quota exhaustion shows up as a pool that wants to grow but does not.
Three scaling signals and what each one gets wrong
| Signal | Good fit | Characteristic failure mode |
|---|---|---|
| CPU utilization threshold | Steady, CPU-bound services where load and CPU move together | I/O-bound or memory-bound workloads: CPU stays flat while latency climbs |
| Custom metric from IBM Cloud Monitoring | Queues, request latency, or business metrics that lead CPU | Metric ingestion lag — the controller can act on a value that is already minutes old |
| Scheduled (time-based) rule | Predictable peaks: business hours, nightly batch, known event windows | Anything that does not follow the calendar; an over-generous schedule is just a permanently larger bill |
Most production pools end up combining two: a scheduled floor that covers the predictable part of the load, plus a metric-based rule to absorb the rest.
Worked example: read the pool before you change it
Run these from a workstation with the IBM Cloud CLI and kubectl installed, authenticated to the account that owns the cluster. Reading node pool configuration requires at least the Viewer platform role on the Kubernetes Service instance; changing it requires Operator or Editor. kubectl access comes from the cluster's own RBAC, not from IAM.
# Replace with your cluster name or ID and node pool name
CLUSTER="my-cluster"
POOL="default"
# Current worker count, per-zone size, and autoscaling bounds
ibmcloud ks node-pool get --cluster "$CLUSTER" --name "$POOL"
# Cluster-level view: zones, add-ons, and Kubernetes version
ibmcloud ks cluster get --cluster "$CLUSTER"
# Recent scaling-related events emitted by the controller
kubectl get events -n kube-system --sort-by=.lastTimestamp | grep -i scal
What to look for: the node-pool get output should show the autoscaling bounds and the current workers per zone. If the pool is already at its maximum and pods are still pending, the problem is the bound or the quota, not the signal. The event stream is the fastest way to see whether the controller tried to act and was refused.
Event reason strings differ between IKS versions and add-on releases, so treat the grep as a filter that narrows output rather than a fixed contract. If it returns nothing, drop the filter and read the raw list.
Changing the bounds
To change autoscaling bounds or enable autoscaling on a pool, use the node pool create or update command for your CLI version. Flag names have changed across releases, so confirm them first:
ibmcloud ks node-pool create --help
ibmcloud ks node-pool update --help
Capture the current minimum, maximum, and per-zone size before changing anything; that output is your rollback target. Applying a new bound is a state change, so restoring the previous values is the way back if the new setting causes unexpected scale-out.
The trade-off: oscillation costs more than the spike
Aggressive scale-down is where the money leaks. A pool that scales in quickly after every dip pays for repeated node provisioning, and workloads rescheduled onto a shrinking pool can see transient latency while pods restart. The usual guard is a stabilization or cooldown window on scale-down that is longer than the one on scale-out — slow to shrink, quick to grow.
The second limitation is specific to the custom-metric path. Metrics exported to IBM Cloud Monitoring are not instantaneous. If a scaling rule keys off a metric with a multi-minute ingestion delay, the controller may scale out after the queue has already drained, then scale back in. For bursty workloads, a scheduled floor is often cheaper and more predictable than a reactive rule.
Checking that the decision worked
After a change, verify in this order:
- Pool configuration matches intent:
ibmcloud ks node-pool getshows the expected minimum, maximum, and per-zone size. - The controller is acting: scaling-related events appear in
kube-systemduring a known load window. - Capacity is genuinely available: no pending pods with insufficient-capacity reasons, and no quota errors in the cluster's activity log.
- Cost tracks the load curve: worker count over a week should follow the shape of the workload, not the shape of the reconcile interval.
If step 2 succeeds and step 3 fails, the signal is fine and the constraint is quota or instance availability. If step 2 never fires, the signal is the thing to change.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.