Diagnosing Latency and State Loss in Cloudflare Workers Durable Objects
A practical diagnostic flow for Cloudflare Workers Durable Objects when you see unpredictable latency or missing state. Follow the cause‑diagnostic table, ordered checks, and targeted fixes to get your application back to predictable performance.
25 Jul 2025, 10:13 UTC

Problem Statement
When a Cloudflare Worker uses a Durable Object (DO) you may notice two common symptoms:
- Requests to the DO return with unpredictable, often high, latency.
- State that was written in a prior request is missing or incomplete in subsequent requests.
Both symptoms can stem from subtle misconfigurations or runtime limitations. This guide gives you a concise diagnostic flow that maps observed conditions to likely causes, ordered checks, and fixes. It also tells you when to involve the Cloudflare support team.
Cause‑Diagnostic Table
| Observed Condition | Common Cause | Key Diagnostic Check |
|---|---|---|
| High latency on concurrent writes | Serialization bottleneck on the same DO instance | Measure per‑instance write latency in the dashboard |
| Missing or stale state after write | Async calls not awaited or storage API misuse | Verify that all state.storage.put calls are awaited |
| Partial data returned or silent failures | 10 MB per‑instance storage limit exceeded | Check state.storage.get return size and error logs |
| 404 or 500 errors when binding the DO | Misconfigured binding name or missing env var | Confirm DO_BINDING is set in wrangler.toml and env |
| Repeated failures from different origins | CORS not configured for the DO endpoint | Inspect response headers for Access-Control-Allow-Origin |
| Throttling or timeouts after many reads | Rapid reads saturating request quota | Check request quota usage in the dashboard |
| Data corruption or inconsistent reads | Race condition due to missing withState callback | Ensure withState is defined and used correctly |
Ordered Diagnostic Checks
- Verify Binding Configuration
Open
wrangler.tomland ensure the binding matches the code reference.[durable_objects] bindings = [{ name = "MY_DO", class_name = "MyObject", script_name = "my-worker" }]On the Cloudflare dashboard, confirm the
MY_DObinding is listed under the Workers > Durable Objects section. - Check Environment Variables
Run
wrangler kv:namespace listto ensure any KV namespaces used by the DO exist and are bound. - Measure Per‑Instance Latency
In the Workers dashboard, navigate to the Durable Objects tab, select the relevant namespace, and review the “Latency” chart. A spike that appears only when multiple requests target the same instance indicates serialization bottlenecks.
- Audit Async Storage Calls
Search your DO code for
state.storage.putandstate.storage.get. Ensure every call is awaited:async handle(request, env, ctx) { await env.MY_DO.idFromName("session").then(id => { const obj = env.MY_DO.get(id); return obj.fetch(request); }); }Missing
awaitor returning a promise without awaiting can cause state to appear lost. - Validate Storage Size
Insert a diagnostic endpoint in the DO that reports the current storage size:
async fetch(request) { if (request.url.includes("/debug/size")) { const keys = await this.state.storage.list(); let size = 0; for await (const { key } of keys) { const value = await this.state.storage.get(key); size += JSON.stringify(value).length; } return new Response(`Size: ${size} bytes`); } }If the size approaches 10 MB, you will see partial reads or silent failures.
- Inspect CORS Headers
Make a fetch from a browser console to the DO endpoint and check the response headers for
Access-Control-Allow-Origin. If missing, configure the DO to set the header:async fetch(request) { const res = await this.doSomething(); return new Response(res.body, { status: res.status, headers: { "Access-Control-Allow-Origin": "*", ...res.headers } }); } - Monitor Request Quota
Navigate to the Worker’s “Usage” tab. A high
Requests per secondcounter with a corresponding error rate suggests throttling. - Check
withStateUsageWhen defining a DO class, you must supply a
withStatecallback to access the internal storage:class MyObject { constructor(state, env) { this.state = state; this.env = env; } static async withState(state, env) { return new MyObject(state, env); } }Omitting this callback will cause the DO to operate without persistence, leading to race conditions.
Targeted Fixes
- Serialization Bottleneck
- Distribute writes across multiple DO instances by using
state.storage.get('instanceId')to decide which instance to target. - Batch writes: collect multiple updates in memory and commit them in a single
putcall. - Use the
state.storage.setWithTTLAPI to reduce write frequency.
- Distribute writes across multiple DO instances by using
- Async Call Misuse
- Wrap all storage interactions in
awaitstatements. - Return a resolved promise only after the storage operation completes.
- Wrap all storage interactions in
- Storage Limit Exceeded
- Compress data before storing (e.g.,
JSON.stringify(data).length / 2after gzip). - Move large blobs to Cloudflare R2 or KV and store only a reference in the DO.
- Compress data before storing (e.g.,
- Binding Errors
- Re‑deploy the Worker after correcting the binding name in
wrangler.toml. - Run
wrangler publishwith the--dry-runflag to catch misconfigurations.
- Re‑deploy the Worker after correcting the binding name in
- CORS Issues
- Add a global middleware that sets
Access-Control-Allow-Originfor all responses. - Test with
curl -I https://example.com/your-do-endpointto confirm the header.
- Add a global middleware that sets
- Quota Saturation
- Implement client‑side caching using
Cache APIto reduce read frequency. - Use
Cache-Control: max-age=60for idempotent reads.
- Implement client‑side caching using
- Race Conditions
- Ensure
withStateis defined and returns a fully constructed object. - Wrap critical sections in
this.state.storage.getfollowed byputinside anawaitchain.
- Ensure
Escalation Criteria
If after completing the above checks the symptoms persist, consider escalating to Cloudflare support:
- Latency remains >2 s for >10 % of requests after 30 min of mitigation.
- State loss occurs even when all async calls are awaited and storage size is <5 MB.
- You observe error logs containing
DurableObjectErrororKV::Error::StorageLimitExceededwithout explicit error messages. - Quota usage spikes without a corresponding increase in traffic volume.
When contacting support, provide:
- Worker and DO namespace IDs.
- Sample request logs showing latency spikes.
- Snapshot of the
wrangler.tomlbindings section. - Results of the storage size diagnostic endpoint.
Practical Verification Checklist
- Run
wrangler devlocally and use--watchto observe real‑time latency. - Use
wrangler tailto stream logs and look forDurableObjectErrorentries. - Execute a minimal concurrency test harness:
async function testConcurrentWrites() {
const id = env.MY_DO.idFromName("test");
const obj = env.MY_DO.get(id);
const promises = [];
for (let i = 0; i < 10; i++) {
promises.push(obj.fetch(new Request("/update", { method: "POST", body: `{"i":${i}}` } )));
}
await Promise.all(promises);
}
/debug/size endpoint to confirm data persisted.Conclusion
Unpredictable latency and state loss in Durable Objects usually trace back to a handful of well‑understood causes. By following the ordered checks and applying the targeted fixes above, most teams can restore consistent performance without involving external support. Use the practical verification steps to confirm that each fix has taken effect before escalating.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.