Diagnosing and Resolving Kubernetes CrashLoopBackOff States
A diagnostic guide for resolving Kubernetes CrashLoopBackOff states. Learn how to use exit codes, previous logs, and cluster events to identify OOMKilled processes and configuration errors.
28 Feb 2026, 11:40 UTC

The CrashLoopBackOff Problem
A CrashLoopBackOff is not an error in itself, but a status indicating that a container is repeatedly crashing after startup. When a container fails, the kubelet (the agent running on each node) attempts to restart it. To prevent the node from being overwhelmed by a failing process, Kubernetes implements an exponential back-off delay, increasing the time between restarts until it reaches a maximum of five minutes.
The primary challenge is that by the time you run a diagnostic command, the container that actually crashed is gone, and you are looking at a new, fresh instance that is about to crash again.
Quick Diagnostic Reference
| Symptom | Likely Cause | Key Indicator |
|---|---|---|
| Immediate crash upon start | Config/Env Error | Exit Code 1 or 255 |
| Crash after several minutes | Liveness Probe Failure | "Liveness probe failed" in Events |
| Random crashes under load | Memory Exhaustion | Reason: OOMKilled |
| Crash during file access | Permission/Volume Issue | "Permission denied" in logs |
Step-by-Step Diagnostic Workflow
1. Inspect the Termination Reason
Before looking at logs, determine why the process stopped. The current pod status only shows the current attempt; you must look at the previous state.
# Run on a machine with kubectl configured for the cluster
# Permission: Needs 'get' and 'list' permissions for pods
kubectl describe pod <pod-name>
Locate the Containers: section and look for Last State. Specifically, check the Reason and Exit Code:
- OOMKilled: The container exceeded its memory limit. The Linux kernel terminated the process.
- Exit Code 0: The application finished its task and exited normally. Kubernetes expects services to run indefinitely; if it exits 0, it is still treated as a crash loop.
- Exit Code 1 or 137: General application error or a SIGKILL signal.
2. Retrieve Logs from the Crashed Instance
Standard kubectl logs shows the current (likely starting) container. To see why the previous one died, use the --previous flag.
# Retrieve logs from the failed container instance
kubectl logs <pod-name> --previous
Risk: If the container crashed so quickly that it couldn't flush its stdout buffer, this command may return nothing. In such cases, check the cluster events.
3. Check Cluster Events
If logs are empty, the issue is likely occurring at the Kubernetes orchestration level (e.g., failing probes or scheduling issues) rather than inside the application code.
# List events in the namespace, sorted by time
kubectl get events --sort-by='.lastTimestamp'
Look for Unhealthy warnings. If a liveness probe fails repeatedly, Kubernetes kills the container, which triggers the CrashLoopBackOff cycle.
Fixes Based on Findings
Scenario A: OOMKilled
If the reason is OOMKilled, the application is attempting to use more memory than the limits defined in the manifest.
- Fix: Increase the
resources.limits.memoryin the Deployment spec. - Caution: Do not simply double the limit without profiling. If the memory usage grows linearly over time, you have a memory leak that needs a code fix, not a larger limit.
Scenario B: Configuration or Secret Missing
If logs show "File not found" or "Environment variable X is required," the application is crashing during initialization.
- Fix: Verify that the
ConfigMaporSecretreferenced in the pod spec exists in the same namespace. Check for typos in the environment variable keys.
Scenario C: Liveness Probe Misconfiguration
If the application takes 30 seconds to start but the liveness probe starts checking after 5 seconds and fails, Kubernetes will kill the app before it ever becomes ready.
- Fix: Implement a
startupProbe. This allows the container a grace period to initialize before the liveness probe takes over.
Verification and Rollback
To verify the fix, delete the pod to force a fresh restart without the back-off delay:
kubectl delete pod <pod-name>
Monitor the status using kubectl get pods -w. The pod should transition from ContainerCreating to Running and stay there for several minutes.
Rollback: If the change to resource limits or probes causes node instability (e.g., causing other pods to be evicted), revert the manifest to the previous version using:
kubectl rollout undo deployment <deployment-name>
Escalation Criteria
If the following conditions are met, escalate to the infrastructure or security team:
- The pod is
OOMKilleddespite limits being set to the maximum available node memory. - Logs indicate
Permission Deniedon a volume that is correctly mounted (potentialSecurityContextor RBAC issue). - The pod is stuck in
ContainerCreatingbefore it even reaches theCrashLoopBackOffstate (potential CNI or ImagePullBackOff issue).
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.