Diagnosing Intermittent 504 Gateway Timeout Errors from API Gateway with Lambda Integration
A step‑by‑step diagnostic guide for identifying why Amazon API Gateway returns occasional 504 errors when invoking Lambda functions, with concrete checks, fixes, and escalation criteria.
06 Nov 2025, 12:28 UTC

Recognizable condition
You observe sporadic 504 Gateway Timeout responses from Amazon API Gateway when clients call an API that is backed by an AWS Lambda function. The errors appear intermittently, often under load, and the Lambda function itself may succeed when retried.
Short cause/diagnostic table
| Possible cause | What to look for |
|---|---|
| Lambda execution exceeds the default API Gateway integration timeout (29 s for REST APIs) | CloudWatch Logs show Lambda Duration close to or greater than 29 s; API Gateway IntegrationLatency spikes. |
| Concurrency throttling (function or account limit) | API Gateway metrics show 429 errors before 504; Lambda Throttles metric rises. |
| VPC‑related networking problems (missing NAT, blocked SG, DNS failure) | Lambda logs contain timeout or connection‑refused errors; VPC Flow Logs show dropped traffic. |
| Custom domain or edge‑optimized configuration adds DNS/TLS latency | API Gateway Access Logs show prolonged TLS handshake times; CloudFront logs (if used) show high latency. |
Ordered checks
- Inspect Lambda execution time
- Open the Lambda function in the console → Monitor → Logs or run:
aws logs filter-log-events --log-group-name /aws/lambda/<function-name> --filter-pattern "REPORT" --limit 20- Look for the
Durationfield. If values frequently approach 29000 ms, the function is near the timeout.
Permissions needed:
logs:FilterLogEventson the log group. - Review API Gateway metrics
- In CloudWatch, select the API Gateway namespace and view:
IntegrationLatency(Average, p95)5XXErrorcount4XXError(specifically 429) count
- Permissions needed:
cloudwatch:GetMetricStatistics. - Check for throttling
- If
4XXErrorshows a rise in 429 before 504, concurrency is likely the cause. - Run:
aws lambda get-function-concurrency --function-name <function-name>- Compare the reserved concurrency with observed concurrent invocations (CloudWatch
ConcurrentExecutionsmetric).
Permissions needed:
lambda:GetFunctionConcurrencyandcloudwatch:GetMetricStatistics. - If
- Validate VPC networking (if the Lambda is VPC‑enabled)
- Ensure the Lambda’s security group allows outbound traffic to the required endpoints (e.g., RDS port, S3 via VPC endpoint, or public internet via NAT).
- Check that a NAT gateway exists in the subnet’s route table if the function needs to reach public AWS services.
- Use VPC Flow Logs to look for REJECT entries on the Lambda’s elastic network interface:
aws ec2 describe-flow-logs --filter "name=resource-id,values=<eni-id>"
Permissions needed:
ec2:DescribeFlowLogs,ec2:DescribeSubnets,ec2:DescribeRouteTables. - Test custom domain latency
- If you use a custom domain name (edge‑optimized or regional), enable detailed CloudFront logging (if applicable) or API Gateway access logging.
- Look for fields like
tlsHandshakeTimeordnsLookupTimethat add to overall latency.
Permissions needed:
apigateway:GETon the stage,cloudfront:GetDistributionConfigif using CloudFront. - Perform a low‑load manual test
- In the API Gateway console, choose the method → Test → provide a sample payload and invoke.
- Alternatively, run:
aws apigateway test-invoke-method --rest-api-id <api-id> --resource-id <resource-id> --http-method POST --path-with-query-string '/' --body '{}'- Measure the response time. If it succeeds quickly, the problem likely appears only under higher concurrency or network load.
Permissions needed:
apigateway:TestInvokeMethod.
Fixes tied to findings
If Lambda duration exceeds the timeout
- Increase the function timeout (maximum 30 s for REST API integration). Use:
aws lambda update-function-configuration --function-name <function-name> --timeout 30- If the function regularly needs >30 s, consider switching to an HTTP API with payload version 2.0 (supports up to 50 s) or refactor to asynchronous processing (e.g., SNS → Lambda).
- Risk: Higher timeout can increase concurrent execution count and cost; monitor
ConcurrentExecutions.
If concurrency throttling is observed
- Raise the reserved concurrency for the function:
aws lambda put-function-concurrency --function-name <function-name> --reserved-concurrency <new-value>- Verify the account‑level concurrency limit via Service Quotas (
aws service-quotas list-service-quotas --service-code lambda) and request an increase if needed. - Risk: Allocating too much concurrency to one function can starve others; set alarms on
Throttlesmetric.
If VPC networking is the issue
- Add missing outbound rules to the Lambda’s security group (e.g., allow TCP 443 to S3 endpoints, or TCP 5432 to RDS).
- Ensure a NAT gateway is present in the subnet’s route table for internet access.
- Confirm DNS resolution: the Lambda’s VPC should have DHCP options set to
AmazonProvidedDNSor a working custom DNS. - Verification: After changes, re‑run the low‑load test and check CloudWatch Logs for successful outbound connections.
If custom domain latency pushes total time over the limit
- Regional custom domains avoid the extra CloudFront edge hop; consider migrating from edge‑optimized to regional.
- Enable HTTP keep‑alive and use a recent TLS version (TLS 1.2 or 1.3) to reduce handshake time.
- If using CloudFront, increase the
TTLor adjust the origin request policy to reduce round‑trips. - Verification: Compare
IntegrationLatencybefore and after the change; look for a reduction in the TLS handshake component.
Escalation criteria
- If after applying the relevant fix the 504 error rate remains above an acceptable threshold (e.g., >0.1 % of requests) for more than 15 minutes, proceed to:
- Enable API Gateway detailed metrics (
execute-api:Stagelevel) to capture per‑request latency breakdown. - Capture a sample of the failing requests with AWS X‑Ray (if enabled) to see where time is spent.
- Consider moving the backend to an asynchronous pattern (API Gateway → SQS → Lambda) to decouple request latency from function execution time.
- If the problem persists and you suspect a platform‑level issue, open a support case with the API Gateway and Lambda teams, providing the CloudWatch Logs timestamps, request IDs, and the metric screenshots.
Limitations and practical verification
The REST API integration timeout is a hard limit of 30 seconds; you cannot exceed it without changing the API type or redesigning the workflow. After any change, verify the fix by:
- Setting a CloudWatch alarm on
5XXErrorfor the API stage with a threshold of, for example, 0.05 % over 5 minutes. - Confirming that the alarm stays in OK state for at least two consecutive evaluation periods.
- Checking that the Lambda
Durationmetric’s 95th percentile is comfortably below the timeout (e.g., < 24 s for a 30 s limit).
If the alarm transitions to ALARM, revisit the checks above; otherwise, consider the issue resolved.
Example scenario
A Lambda function that processes uploaded PDFs occasionally takes 32 seconds because of a large image‑extraction step. The API Gateway is a REST API, so the integration timeout of 29 seconds is exceeded, producing a 504. The following steps resolved the issue:
- Logs showed
Durationvalues of 31000‑34000 ms. - The function timeout was increased to 30 seconds (the maximum for REST API).
- Because the function still regularly exceeded 30 seconds under peak load, the team switched to an HTTP API with payload version 2.0, raising the allowed integration timeout to 50 seconds.
- After the switch, the 504 error rate dropped from 0.3 % to <0.01 % and the Lambda Duration 95th percentile settled at 27 seconds.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.