Diagnosing Sudden Query Latency Spikes in SurrealDB
When query response times jump above 200 ms or timeouts appear while CPU and memory stay normal, follow this step‑by‑step guide to isolate the cause, apply the appropriate fix, and know when to escalate.
17 Nov 2025, 03:47 UTC

Recognizable Condition
Average query response time exceeds ~200 ms or you see frequent timeout errors in client logs, while server CPU and memory utilization remain within normal ranges. This pattern indicates that the bottleneck is not raw resource saturation but something affecting query execution or communication.
Cause & Diagnostic Table
| Possible Cause | What to Look For |
|---|---|
| A) Insufficient query‑planning cache size | Low query_cache_hit_ratio; repeated planning of the same statement. |
| B) Lock contention on heavily updated tables | Elevated lock_wait_time; many concurrent writes on the same table. |
| C) Network packet loss or latency between client and SurrealDB | Increased round‑trip time, retransmits, or jitter observed from the client side. |
| D) Outdated index statistics leading to suboptimal plans | Query plans show SeqScan where an IndexScan was previously used; statistics timestamp stale. |
Ordered Verification Steps
Enable SurrealDB query logging (if not already active) and inspect recent entries for plan details.
# Example: set log level via HTTP (requires admin token) curl -X POST http://localhost:8000/admin/log-level \ -H "Authorization: Bearer " \ -d '{"level":"debug"}'Look for the
planfield in each log line; note whether it containsIndexScanorSeqScan.Query the built‑in HTTP admin endpoint for key metrics.
curl -s http://localhost:8000/admin/metrics | grep -E 'query_cache_hit_ratio|lock_wait_time'Record the values; a hit ratio significantly below your baseline (e.g., < 0.7) or a lock_wait_time trending upward points to causes A or B.
From the client host, measure network health to the SurrealDB node.
ping -c 20 # or for loss detection mtr --reportLook for >1 % packet loss or average RTT > 10 ms higher than usual.
Run EXPLAIN on a representative query that participates in the latency spike.
EXPLAIN SELECT * FROM person WHERE email = '[contact removed]';Compare the output to a known‑good baseline; note any change from index usage to a full scan.
Fixes Tied to Findings
If logs show missing indexes or frequent SeqScans (Cause D):
- Create the appropriate index:
CREATE INDEX ON TABLE person (email); - Or refresh statistics:
UPDATE STATISTICS ON TABLE person; - After applying, re‑run EXPLAIN to verify the plan now uses
IndexScan.
- Create the appropriate index:
If query_cache_hit_ratio is low (Cause A):
- Increase the planning cache size via configuration (e.g., set
query_planning_cache_sizeinsurreal.cfgor via CLI flag). - Restart the SurrealDB instance during an approved maintenance window.
- Confirm the ratio rises above the baseline after restart.
- Increase the planning cache size via configuration (e.g., set
If lock_wait_time is high (Cause B):
- Identify hot tables from the lock metrics or by enabling
lock_log. - Reduce write hotspots: batch updates, shard the table, or increase concurrent transaction limit (
max_concurrent_txns). - Restart the server to apply the new limit, then monitor lock_wait_time drop.
- Identify hot tables from the lock metrics or by enabling
If network loss or latency is detected (Cause C):
- Check NIC errors, switch port statistics, or cable integrity.
- Consider enabling TLS session resumption or using a persistent connection pool to reduce handshake overhead.
- After remediation, re‑run the ping/mtr test to confirm loss < 0.1 % and RTT stable.
Escalation Criteria
Escalate to SurrealDB support when:
- Latency remains > 500 ms after applying the relevant fix, or
- Error rates (timeouts or failed requests) exceed 1 % of total traffic despite corrective actions.
Provide the support team with:
- Recent query logs (including plan fields).
- Metric snapshots from
/admin/metricsbefore and after changes. - Network diagnostics (ping/mtr output) and any observed packet loss.
- Details of the workload (query patterns, write frequency).
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.