Architecture Note: Minimal Stateless Web Service on DigitalOcean Managed Kubernetes
A concise architecture note for running a stateless web service on the smallest DigitalOcean Managed Kubernetes cluster, covering requirements, minimal design, trust boundaries, ops checks, failure modes, and when to redesign.
28 Jun 2026, 01:36 UTC

Requirements
The service must accept HTTP traffic, scale horizontally, keep any session state outside the pod, and meet a 99.9 % availability SLA. Expected resource usage per replica is modest: less than 2 vCPU and under 4 GB of RAM. The design should be operable with the smallest allowed DigitalOcean Kubernetes (DOKS) cluster while still providing clear trust boundaries and observable failure modes.
Minimal Suitable Design
Start with a single‑node DOKS cluster (the smallest size permitted by DigitalOcean). Use one node pool consisting of a basic s‑1vcpu‑2gb droplet. Deploy the web service as a Deployment with a replica count of at least two to enable horizontal scaling within the node’s capacity. Expose the service through a DigitalOcean Load Balancer configured for TCP/443 with TLS termination; the Load Balancer forwards traffic to the pod’s ClusterIP service inside the VPC. For any persistent state that the application might need (e.g., file uploads), create a PersistentVolumeClaim (PVC) backed by Block Storage. All secrets are injected via Kubernetes Secrets and never written to the node’s local disk. Egress traffic to external APIs is allowed only through the node’s outbound‑enabled security settings.
Example Manifests
# Deployment (stateless web service)
apiVersion: apps/v1
kind: Deployment
metadata:
name: web-service
spec:
replicas: 2
selector:
matchLabels:
app: web
template:
metadata:
labels:
app: web
spec:
containers:
- name: web
image: your-registry/web-service:latest
ports:
- containerPort: 8080
envFrom:
- secretRef:
name: web-secrets
readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 5
periodSeconds: 10
---
# Service exposed via LoadBalancer
apiVersion: v1
kind: Service
metadata:
name: web-lb
annotations:
service.beta.kubernetes.io/do-loadbalancer-enable-proxy-protocol: "true"
spec:
type: LoadBalancer
selector:
app: web
ports:
- port: 443
targetPort: 8080
protocol: TCP
---
# PVC for optional persistent storage
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: web-storage
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 10Gi
storageClassName: do-block-storage
Replace your-registry/web-service:latest with the actual image reference and web-secrets with a Kubernetes Secret that contains any required API keys or TLS credentials.
Trust and Data Boundaries
- Ingress boundary: Traffic terminates at the DigitalOcean Load Balancer; TLS is offloaded there, so the pod sees plain HTTP inside the VPC.
- Network boundary: The Load Balancer, nodes, and pods all reside in the same VPC; no public IPs are assigned to pods.
- Secret boundary: Kubernetes Secrets are stored in etcd (encrypted at rest by DOKS) and injected as environment variables or mounted files; they never touch the node’s ephemeral disk.
- Storage boundary: The PVC maps to a Block Storage volume that lives outside the node’s local disk; snapshots of this volume are managed independently by DigitalOcean.
Operational Checks
- Verify node size: After creating the cluster, run
doctl compute droplet list --tag(requires a Personal Access Token withreadandwritescope). Confirm that the droplet size showss-1vcpu-2gb. - Check Load Balancer TLS: In the DigitalOcean Control Panel, navigate to Networking → Load Balancers, select the LB created for
web-lb, and view the certificate details under the “Settings” tab. Ensure the certificate matches the domain you intend to serve. - Validate PVC snapshots: Execute
doctl compute snapshot listand look for snapshots with names prefixed by the PVC’s UID or the cluster name. A successful snapshot will appear with a status ofavailable. To test restore, usedoctl compute volume-snapshot restore --volume-nameand then attach the new volume to a test pod to confirm data integrity. - Monitor health: Enable the DigitalOcean Monitoring agent on the node (installed by default on DOKS). Create an alert policy that triggers when CPU usage exceeds 80 % for five minutes or when memory usage exceeds 85 % of available RAM.
Failure Modes
- Node failure: DOKS automatically marks the node as unhealthy; the Kubernetes scheduler evicts pods to the remaining node(s). With a single‑node cluster, pods are rescheduled after the node is replaced, causing a brief interruption proportional to the node‑replacement time (typically a few minutes).
- Load Balancer failure: If the LB loses connectivity, DNS resolution continues to point to the LB’s IP, but traffic drops until DigitalOcean fails over to a healthy LB instance (usually under 60 seconds). During this window, clients see connection timeouts.
- Resource exhaustion: Sustained CPU >80 % or memory pressure on the
s‑1vcpu‑2gbnode can lead to throttling or OOMKilled pods. The design does not automatically add capacity; manual intervention or a redesign is required. - Storage snapshot failure: If Block Storage snapshot service is degraded, daily PVC backups may fail, leaving the most recent snapshot older than the intended RPO. This does not affect runtime but impacts recovery point objectives.
Conditions That Would Change the Design
- Higher per‑replica resource demand: If the service regularly needs >2 vCPU or >4 GB RAM, move to a larger node type (e.g.,
s‑2vcpu‑4gbors‑4vcpu‑8gb) or enable the cluster autoscaler with a mixed node pool. - Strict data‑residency or encryption requirements: Enable encryption at rest for Block Storage volumes (available as a toggle when creating the PVC) or migrate the PVC to a volume in a specific region that satisfies regulatory constraints.
- Increased traffic beyond a single LB’s capacity: Although DigitalOcean LBs scale automatically, if you anticipate bursts that exceed the LB’s connection‑per‑second limits, consider adding a second LB and using DNS‑based load sharing or a global load balancer.
- Need for multi‑zone resilience: Expand the node pool to span multiple availability zones within the region; this changes the trust boundary because node failure in one zone no longer impacts the entire service.
Practical Verification Summary
To confirm that the architecture behaves as described:
- Create the cluster:
doctl k8s cluster create web-cluster --region nyc1 --node-pool s-1vcpu-2gb:1 --tag web-cluster - Apply the manifests above with
kubectl apply -f web-service.yaml. - Run the verification steps in the “Operational Checks” section.
- Observe that the Load Balancer shows a valid TLS certificate, that the PVC snapshot appears in
doctl compute snapshot list, and that node health metrics stay within alert thresholds.
If any check fails, adjust the node size, enable storage encryption, or revisit the LB configuration before promoting the design to production.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.