Diagnosing Hanging Requests and Timeouts in Fastify
Learn how to diagnose and fix hanging requests in Fastify, from identifying event loop blockage and missing next() calls to resolving database connection pool exhaustion.
17 Jul 2025, 13:52 UTC

The Symptom: The 'Pending' Request
A Fastify application is healthy when it processes requests asynchronously. However, you may encounter a scenario where specific endpoints—or the entire server—stop responding. The client sees the request stay in a pending state until a gateway (like Nginx or AWS ALB) returns a 504 Gateway Timeout, or the TCP socket closes. This usually indicates that the request has entered the Fastify lifecycle but never reached the final response phase.
Diagnostic Matrix: Identifying the Cause
| Condition | Likely Cause | Primary Indicator |
|---|---|---|
| Only specific routes hang | Missing await or next() |
Logs show request entry but no exit/response. |
| All routes hang under load | Event Loop Blockage | High CPU usage; /health route becomes unresponsive. |
| Intermittent hangs during DB calls | Connection Pool Exhaustion | DB driver logs show queueing or timeout errors. |
| Hangs during large data transfers | JSON Serialization Block | CPU spikes during reply.send() with large arrays. |
Step-by-Step Resolution Path
1. Establish a Baseline with a Health Check
To determine if the entire Node.js process is frozen or if a specific route is leaking, implement a minimal health check route. This route must avoid all external dependencies (DBs, APIs).
// Run this in your main server file
fastify.get('/health', async (request, reply) => {
return { status: 'ok' };
});
Verification: While the problematic route is hanging, attempt to curl /health. If /health responds instantly, the event loop is free, and the issue is likely a logic error (missing await) in the specific route. If /health also hangs, the event loop is blocked.
2. Audit Async Handlers and Hooks
Fastify supports both async/await and callback-style handlers. A common failure occurs when a non-async function is used in a hook or route but fails to call the next() callback.
Check for this pattern:
// RISK: This will hang the request forever if the if-statement is false
fastify.addHook('preHandler', (request, reply, next) => {
if (request.headers['x-api-key']) {
next();
}
// Missing 'else { next() }' or 'return reply.send()' here
});
Fix: Ensure every possible execution path in a callback-based handler calls next() or sends a response via reply.send().
3. Detect Event Loop Blockage
Node.js is single-threaded. If you perform a CPU-intensive task (e.g., fs.readFileSync, heavy cryptography, or parsing a 50MB JSON string) inside a handler, no other requests can be processed.
Diagnostic: Use the native perf_hooks module to monitor lag. If the lag exceeds 100ms, the loop is blocked.
const { monitorEventLoopDelay } = require('perf_hooks');
const h = monitorEventLoopDelay();
h.enable();
// Check this value periodically or on a timer
console.log(`Event loop delay: ${h.mean / 1e6} ms`);
Fix: Offload CPU-intensive tasks to worker_threads or break large synchronous loops into smaller chunks using setImmediate().
4. Verify Database Connection Limits
If you use a database driver (e.g., pg for PostgreSQL), the internal connection pool has a maximum limit. If all connections are checked out and not returned, subsequent requests will wait indefinitely for a connection.
Check: Review your pool configuration. If max: 10 and you have 11 concurrent long-running queries, the 11th request will hang.
Fix: Increase the pool size or, more importantly, ensure all clients are released back to the pool using try...finally blocks.
Escalation Criteria
If the following conditions are met, move from code auditing to infrastructure profiling:
- The
/healthroute hangs, but CPU usage is low (indicates a deadlock or external resource wait). - Memory usage climbs steadily during hangs (indicates a memory leak causing excessive Garbage Collection pauses).
- The issue only occurs under specific concurrency levels (indicates a race condition or pool exhaustion).
Rollback and Recovery
If these changes are deployed to production and cause instability:
- Revert the connection pool size to the previous known-stable value to avoid overloading the database server.
- Remove custom
perf_hooksmonitoring if it introduces unexpected overhead in high-throughput environments.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.