Direct Answer
First, try to shrink the transaction scope. A smaller scope returns connections to the pool faster and usually resolves intermittent exhaustion. Only enlarge the pool after confirming that transaction length cannot be reduced further.
Confirmed Facts
- HikariCP reports
ActiveConnections and PendingRequests via JMX or its log output. - The pool blocks new requests when
ActiveConnections == maximumPoolSize and PendingRequests starts to grow. - Connection acquisition timeouts appear when
PendingRequests wait longer than connectionTimeout. - Database‑side
wait_timeout can close idle sockets, but it does not cause the pool to report pending requests.
Likely Explanation for the Intermittent Exhaustion
When a Hibernate session (or @Transactional method) holds a connection for the duration of the transaction, the connection stays in ActiveConnections. If those transactions occasionally run long (e.g., due to lazy loading, batch processing, or OSIV keeping the session open during view rendering), the pool can exhaust even though the configured maximumPoolSize is adequate for short‑lived work.
Steps to Diagnose and Act for This Case
- Enable HikariCP leak detection and log stack traces:
HikariConfig config = new HikariConfig();
config.setLeakDetectionThreshold(3000); // ms
config.setLogConnectionState(true);
- Monitor
ActiveConnections and PendingRequests during a peak load window. - If
PendingRequests rises while ActiveConnections stays near maximumPoolSize, the bottleneck is transaction length. - If
ActiveConnections stays well below the limit but PendingRequests still spikes, the pool may be too small for the concurrent request rate. - Profile the application (e.g., with VisualVM or YourKit) to measure the average duration of @Transactional methods under load.
Missing Diagnostic Detail That Would Change the Recommendation
What is the average duration of @Transactional methods (or Hibernate sessions) during the periods when timeouts occur? If the average is on the order of seconds or more, focus on reducing transaction scope; if it remains sub‑second and the pool still exhausts, consider increasing maximumPoolSize after verifying the database can handle the extra load.
Verification Actions
- Check HikariCP logs for messages like "Pool exhausted" and associated stack traces.
- Correlate spikes in
PendingRequests with application‑level transaction timestamps. - After adjusting transaction boundaries (e.g., removing OSIV, shortening service methods, or using read‑only transactions), re‑run the load test and confirm that
PendingRequests no longer grows. - If you increase the pool size, observe database CPU and connection count to ensure the server is not saturated.