Guaranteed vs Burstable QoS: Choosing Memory Requests for Production Stability
27K reputation · 08 Dec 2023, 14:24 UTC
Resource Allocation Trade‑offs
In Kubernetes, the relationship between resource requests and limits defines a pod’s Quality of Service (QoS) class. Setting requests equal to limits yields a Guaranteed pod, protecting it from eviction during node memory pressure. Setting requests lower than limits creates a Burstable pod, allowing it to use spare node capacity but increasing the likelihood of OOMKills when the node is overcommitted.
Constraints and Uncertainty
Local development environments typically lack strict resource quotas, so Burstable pods often run without contention. In contrast, production clusters experience heavy resource competition, making the choice between Guaranteed and Burstable a critical decision for stability versus utilization. Over‑provisioning for Guaranteed status can waste resources, while Burstable configurations may lead to unpredictable cascading failures during traffic spikes.
- What is the acceptable overcommit ratio for Burstable pods before the risk of node‑level instability outweighs utilization benefits?
- How does the risk of CFS throttling in Burstable CPU configurations compare to the stability gains of Guaranteed memory limits?
- Can a hybrid strategy—setting requests close to limits for critical services while allowing lower requests for non‑critical ones—balance cost and reliability in production?