Diagnosing HTTP 503 and 504 Errors in Azure App Service
Learn how to identify the root cause of intermittent 503 and constant 504 errors in Azure App Service, run diagnostic checks, apply fixes, and know when to escalate.
30 Sept 2025, 06:33 UTC

Problem: Intermittent 503 and constant 504 errors in Azure App Service
When users see HTTP 503 (Service Unavailable) or HTTP 504 (Gateway Timeout) responses, the application appears down or slow, but the underlying cause can be either an App Service process crash or a request that exceeds the platform’s idle timeout.
Useful takeaway: By following a structured set of checks—logs, metrics, and configuration—you can pinpoint whether the error stems from memory exhaustion or a long‑running request, apply the appropriate fix, and know when to involve deeper support.
Recognizable condition and cause/diagnostic table
| Error | Typical Cause | Diagnostic Signal |
|---|---|---|
| HTTP 503 | App Service worker process crash (often OOM) | System.OutOfMemoryException in logs; "Web App Down" event in Diagnose and solve problems |
| HTTP 504 | Request exceeds Azure Load Balancer idle timeout (default 4 min) | No response after >240 s; logs show request started but no completion; Metrics show high "Average Response Time" |
Ordered checks
- Open the App Service resource in the Azure Portal, select Diagnose and solve problems, and look for Web App Down or High CPU/Memory incidents.
- Stream logs via Kudu (
https://<app-name>.scm.azurewebsites.net/DebugConsole) or enable Log Stream and search for System.OutOfMemoryException or StackOverflowException. - In Azure Monitor > Metrics, chart Http5xx, Average Response Time, and Memory Working Set for the App Service plan.
- If a 504 is suspected, enable detailed request logging (
DIAGNOSTICSLOGS) or use Application Insights to capture request duration; look for requests >240 s. - Verify the Always On setting under Configuration > General settings; it should be On for apps that need to stay warm.
Fixes tied to findings
- For 503 caused by memory pressure
- Scale up the App Service Plan:
az appservice plan update --resource-group <rg> --name <plan> --number-of-workers 1 --sku P2v2(requires Contributor role). - Profile the application with a memory profiler or add
GC.Collect()at safe points; consider reducing large in‑memory caches. - Create an auto‑heal rule:
az webapp config set --resource-group <rg> --name <app> --auto-heal-enabled true --auto-heal-actions actionType=Recycleand a trigger based on memory > 80 %.
- Scale up the App Service Plan:
- For 504 caused by long‑running requests
- Refactor to asynchronous pattern: push work to an Azure Queue Storage message and have an Azure Function process it.
- If synchronous work is unavoidable, place the App Service behind Azure Front Door or Application Gateway and set the backend timeout to >240 s; keep Always On enabled.
- Adjust the idle timeout via app setting:
az webapp config appsettings set --resource-group <rg> --name <app> --settings WEBSITE_IDLE_TIMEOUT=240(maximum 240 s).
Escalation criteria
- Memory‑related 503 persists after scaling up two tiers and after applying auto‑heal.
- 504 errors continue after request refactoring and timeout adjustments, suggesting backend latency or dependency issues.
- Logs show repeated StackOverflowException or native crashes that require a support ticket with Azure.
Rollback and verification
If scaling up the plan did not resolve the 503, you can roll back to the original SKU: az appservice plan update --resource-group <rg> --name <plan> --sku <original-sku>. Verify the fix by deploying a test workload:
- For 504: add a delay (
Thread.Sleep(260000)) in an endpoint and confirm a 504 response. - For 503: allocate a large byte array (
new byte[800_000_000]) to provoke OOM and check for 503.
Check results in the Azure Portal metrics (Http5xx should drop) and in the logs (no exception traces).
Limitations
- Scaling up masks the underlying memory leak; the leak will reappear under load.
- Enabling Always On increases hourly cost and may affect cold‑start latency for consumption‑based plans.
- TCP keep‑alive or other OS‑level timeout tweaks are not available in the App Service PaaS environment.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.