Minimal GKE Autopilot Design for Stateless Web Services
A practical architecture note for deploying a stateless web service on GKE Autopilot, covering requirements, smallest design, trust boundaries, operational checks, and failure modes.
23 Dec 2025, 15:38 UTC

Requirements for a Minimal GKE Autopilot Deployment
You have a stateless web service that must be reachable on Google Cloud without the overhead of managing virtual machines. GKE Autopilot removes node provisioning and OS patching, but it imposes boundaries you must respect before committing.
Cloud project and quota readiness
Ensure your Google Cloud project has the gke API enabled and that your quota allows the intended pod count, CPU, and memory. Autopilot charges based on resource usage rather than per-VM, but exceeding quotas results in scheduling failures.
- Project ID must be set or replace <PROJECT_ID> in commands.
- Run gcloud services enable container.googleapis.com if the API is new.
- Check current quotas with gcloud compute quotas list --filter=resource.type=GKE_CLUSTER.
Service characteristics
Autopilot is optimized for workloads that do not require host-level access, custom kernels, or persistent local storage. Stateless HTTP services, APIs, and background workers fit the model. Workloads that need hostNetwork, privileged containers, or emptyDir on local SSD are unsupported.
Smallest Suitable Design
The most compact production-ready setup uses a single Autopilot cluster, a Deployment describing your container, and a Service exposing it internally or via an external LoadBalancer.
gcloud container clusters create-autopilot
--project <PROJECT_ID>
--region <REGION>
--cluster-name <CLUSTER_NAME>
Run the above in Cloud Shell or a local workstation with the Google Cloud SDK authenticated to the project.
Deployment manifest
apiVersion: apps/v1
kind: Deployment
metadata:
name: <APP_NAME>
namespace: <NAMESPACE>
spec:
replicas: 2
selector:
matchLabels:
app: <APP_NAME>
template:
metadata:
labels:
app: <APP_NAME>
spec:
containers:
- name: <CONTAINER_NAME>
image: <IMAGE_URL>
ports:
- containerPort: 8080
resources:
limits:
cpu: 500m
memory: 256Mi
requests:
cpu: 100m
memory: 128Mi
Placeholders such as <APP_NAME>, <IMAGE_URL>, and <NAMESPACE> should be replaced with your values. The resource limits above are typical for a lightweight web service and stay within Autopilot default pod-level allocations.
Service manifest
apiVersion: v1
kind: Service
metadata:
name: <APP_NAME>-svc
namespace: <NAMESPACE>
spec:
type: LoadBalancer
selector:
app: <APP_NAME>
ports:
- protocol: TCP
port: 80
targetPort: 8080
Setting the Service type to LoadBalancer provisions an external IP address. If you only need internal traffic, use type: ClusterIP and omit the external load balancer.
Trust and Data Boundaries
GKE Autopilot runs the control plane on Google-managed Kubernetes masters. You do not patch the control plane or select node sizes, but you are responsible for container image integrity and pod security.
Image validation
Pull only images from registries you control or verified sources. Untagged or untracked images increase the risk of supply-chain attacks. Scan images for vulnerabilities before pushing them to Artifact Registry or Docker Hub.
Pod security policies and restrictions
Autopilot does not allow privileged containers, hostNetwork, or access to host directories. If your service requires any of these, you must redesign it for Standard mode or refactor to use sidecar patterns that run within the container boundary.
Operational Checks
After deployment, verify that pods are scheduled and running as expected. Use kubectl commands in the cluster context.
Verify pod scheduling
kubectl describe pod <POD_NAME> -n <NAMESPACE>
Run the command after kubectl get pods shows your pods. Expected output includes a Status: Running line, no WarningScheduling events, and Node: default-pool/auto-generated-name indicating Autopilot assignment. If you see 0/1 nodes available, quota or resource constraints are likely the cause.
Monitoring and health checks
Open the Google Cloud Console Cloud Monitoring dashboard for the GKE Pod Health metric. Look for eviction events, restart counts, and CPU/memory usage trends. A healthy Autopilot deployment shows steady usage within the limits you set in the Deployment manifest.
Failure Modes and Conditions That Would Change the Design
| Failure Mode | Impact | Practical Check |
|---|---|---|
| GCE VM shortage in region | Pods remain Pending; service unavailable | Run gcloud compute regions describe <REGION> and check available capacity. |
| NetworkPolicy blocking pod traffic | Service responds with connection refused or timeout | Run kubectl get networkpolicy and kubectl describe pod <POD_NAME> -n <NAMESPACE> to inspect events. |
| Privileged container request | Pod rejected during admission; immediate eviction | Attempt to create a pod with privileged true; admission will fail on Autopilot. |
| Exceeded pod or CPU quota | Scheduling failures, repeated restarts | Monitor gke/autopilot/pod_count and gke/autopilot/cpu_usage in Cloud Monitoring. |
If any of these conditions appear and cannot be resolved by adjusting quotas or redesigning the workload, the architecture note conditions that would change the design clause triggers a move to GKE Standard mode or a multi-cluster approach.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.