Diagnosing Gardener Shoot API Server Unavailability
Step‑by‑step diagnostic guide for troubleshooting Gardener Shoot API server unavailability, covering etcd quorum, pod resources, network policies, node taints, and cloud load balancer checks.
26 Oct 2025, 10:21 UTC

Recognizable condition
When a Gardener Shoot cluster’s API server becomes unavailable, kubectl commands return connection timeouts or errors such as Unable to connect to the server: dial tcp. Workloads cannot be scheduled, and the Gardener dashboard shows the shoot control plane in a degraded state.
Short cause/diagnostic table
| Observed symptom | Likely cause | Key indicator |
|---|---|---|
API server returns no response or service unavailable | Etcd quorum loss | Etcd members report unhealthy or leader election failed in logs |
| API server pod restarts repeatedly | Insufficient CPU/Memory limits or OOMKilled | Pod events show OOMKilled or BackOff with restart count increasing |
| Worker nodes cannot reach API server on port 6443 | Network policy or security group blocks ingress | Network policy logs show dropped packets; cloud security groups lack rule for 6443 |
Control plane pods stuck in Pending or CrashLoopBackOff | Node taints or resource pressure prevent scheduling | Node describes show taints like node-role.kubernetes.io/control-plane:NoSchedule or high memory/CPU usage |
| Intermittent connectivity, occasional timeouts | Cloud load balancer misconfiguration | Load balancer health checks report OutOfService or backend pool missing API server IPs |
Ordered checks
- Verify API server health endpoint
Run from a machine with
kubectlaccess to the shoot’s kubeconfig:kubectl get --raw='/healthz' --server=<apiserver-address>Expected output:
ok. Any other response or timeout indicates the API server is not reachable.Required permission: Any authenticated user able to read the shoot kubeconfig.
Risk: None; read‑only.
- Check etcd cluster health
From a bastion or any node with etcdctl installed and the shoot’s etcd certificates:
etcdctl endpoint health --cluster --cert=<etcd-client-cert> --key=<etcd-client-key> --cacert=<etcd-ca>Expected output: each endpoint shows
healthy. If any endpoint showsunhealthyor the cluster reports loss of quorum, proceed to etcd recovery steps.Required permission: Access to etcd certificates (usually held by the Gardener admin).
Risk: Read‑only; no data modification.
- Inspect API server pod state and logs
List pods in the shoot’s control plane namespace:
kubectl get pods -n garden shoot-<shoot-name>Look for
kube-apiserverpod status. If it isCrashLoopBackOfforOOMKilled, fetch logs:kubectl logs -n garden shoot-<shoot-name> -l k8s-app=kube-apiserver --tail=50Search for
OOMKilled, insufficient resource messages, or failed leader election.Required permission: Ability to read pods in the shoot namespace.
Risk: None.
- Review network policies and security groups
Check for any NetworkPolicy that denies ingress to port 6443:
kubectl get networkpolicy -n garden shoot-<shoot-name> -o yamlAlso verify the cloud provider’s load balancer/security group rules allow traffic from worker node CIDRs and the Gardener seed’s control plane to the shoot’s API server IP on TCP 6443.
Required permission: Read access to NetworkPolicy and cloud provider networking console/CLI.
Risk: None.
- Validate cloud load balancer backend and health checks
Using the provider’s CLI or console, list the load balancer backing the shoot’s API server and confirm:
- Backend targets include the API server node IPs.
- Health check protocol is TCP on port 6443 (or HTTP if using a health‑check path).
- All targets report
InServiceor equivalent.
Example (AWS):
aws elbv2 describe-target-health --target-group-arn <tg-arn>Required permission: Read access to the load balancer service.
Risk: None.
Fixes tied to findings
Etcd quorum loss
- Take a etcd snapshot before any changes (if possible):
ETCDCTL_API=3 etcdctl --endpoints=<etcd-ip>:2379 \
--cert=<cert> --key=<key> --cacert=<ca> snapshot save /var/etcd/snapshot.db
- If a member is permanently offline, remove it only after confirming the snapshot is valid and you have a backup:
etcdctl member remove <member-id> \
--endpoints=<remaining-etcd-ip>:2379 \
--cert=<cert> --key=<key> --cacert=<ca>
- Restart the API server pod to re‑elect a leader:
kubectl rollout restart -n garden shoot-<shoot-name> deploy/kube-apiserver
Limitation: If etcd data is corrupted, the cluster may be unrecoverable; always verify the snapshot can be restored.
Verification: After the pod is Running, rerun the healthz check and etcdctl endpoint health.
API server pod crashloop due to resources
- Edit the shoot’s control plane component settings (via Gardener config) to increase limits. Example patch for the API server:
apiVersion: core.gardener.cloud/v1beta1
kind: Shoot
metadata:
name: <shoot-name>
spec:
provider:
workers:
- name: worker
# ...
kubernetes:
kubeAPIServer:
resources:
limits:
cpu: "2"
memory: "4Gi"
requests:
cpu: "500m"
memory: "2Gi"
- Apply the change; Gardener will roll out a new API server pod with the updated limits.
Risk: During the rolling update there is a brief window where the old pod is terminated before the new one starts; if the new limits are still insufficient, the crashloop may persist.
Verification: Watch the pod’s status (kubectl get pods -w -n garden shoot-<shoot-name> -l k8s-app=kube-apiserver) and ensure it stays Running with no restarts. Then rerun the healthz endpoint.
Network policy or security group blocking ingress
- Identify the offending NetworkPolicy:
kubectl get networkpolicy -n garden shoot-<shoot-name> -o yaml | grep -A5 -B5 "6443"
- Either delete the policy (if it is unnecessary) or add an ingress rule allowing TCP 6443 from the worker node CIDR and the seed’s control plane:
kubectl apply -f - <
- If using a cloud provider security group, add an inbound rule for TCP 6443 from the same sources.
Required permission: Ability to modify NetworkPolicy (cluster admin) and cloud security group (cloud admin).
Verification: From a worker node, run nc -vz <apiserver-ip> 6443; expect a successful connection. Then confirm healthz returns ok.
Control plane node taints or resource pressure
- Describe nodes hosting the control plane:
kubectl get nodes -l node-role.kubernetes.io/control-plane=true -o wide
- If a node is tainted (
NoSchedule) or shows high usage, either:
- Remove the taint:
kubectl taint nodes <node> node-role.kubernetes.io/control-plane:NoSchedule- - Or drain and replace the node via Gardener’s machine‑deployment scaling.
- If resource pressure is the cause, increase the machine type or add more workers via the Shoot spec.
Required permission: Ability to edit nodes (cluster admin) and to modify Shoot machine deployments (Gardener admin).
Risk: Removing a taint may allow workloads to schedule on control plane nodes, which is generally discouraged; prefer scaling the machine pool.
Verification: After node changes, ensure kube-apiserver, kube-controller-manager, and kube-scheduler pods are Running on healthy nodes. Then rerun healthz.
Cloud load balancer misconfiguration
- Confirm the load balancer’s backend pool contains the API server node IPs.
- Ensure the health check matches the API server’s listening port (TCP 6443) and interval.
- If health checks are failing, inspect the security group/network ACL attached to the load balancer to verify it allows traffic from the health‑check source.
Required permission: Read/write access to the load balancer service.
Risk: Changing the backend or health‑check parameters may cause a brief interruption; schedule during a maintenance window if possible.
Verification: After fixing, run the provider’s health‑check description (e.g., aws elbv2 describe-target-health) and confirm all targets are InService. Then validate the API server healthz endpoint.
Escalation criteria
- If etcd quorum loss persists after attempting member replacement and you lack a recent etcd snapshot, escalate to Gardener support or your etcd vendor for potential disaster‑recovery procedures.
- If API server pods continue to crashloop despite raising resource limits and you observe OOMKilled events at the node level, consider a deeper node‑capacity review and involve your infrastructure team.
- When network‑policy or security‑group changes do not restore connectivity and cloud provider logs show dropped packets with no obvious rule, open a ticket with the cloud provider’s networking support.
- If the load balancer consistently reports unhealthy targets despite correct backend registration, verify with the cloud provider that there are no regional service incidents; otherwise, escalate to their support.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.