Per‑User Rate Limiting with Cloudflare Workers KV
Learn how to use Cloudflare Workers KV to enforce per‑user request quotas with sliding‑window counters, including a runnable example and verification steps.
09 Dec 2025, 02:46 UTC

Problem: uncontrolled per‑user traffic can exhaust backend resources
When an API endpoint is open to the internet, a single authenticated user can generate enough requests to overload the origin server, degrade latency for others, and increase cost. A simple global rate limit does not protect against a burst from one user, while per‑IP limits are easily evaded. The goal is to enforce a request quota that is tied to each user’s identity (e.g., a JWT sub claim) and reset after a sliding time window.
Thesis: Cloudflare Workers KV offers a low‑latency, globally distributed store that can implement sliding‑window counters with only a few lines of middleware code
Workers run at the edge, so the decision to allow or block a request happens close to the client, adding minimal round‑trip time. KV provides eventually consistent key‑value storage with sub‑second propagation across Cloudflare’s network, which is sufficient for most rate‑limiting use cases when a small safety margin is applied.
Section 1: KV data model for sliding‑window counters
We store a counter per user and per window using a key that encodes the user identifier and the start of the current window. The value is an integer representing the number of requests seen in that window. When the window expires, the key is automatically deleted by setting an expiration time (TTL).
// Example key format: \"rate:{userId}:{windowStartUnixSec}\"\n// TTL = window length in seconds (e.g., 60 for a 1‑minute window)\n
On each request we:
- Extract the user ID from the request (e.g., from a verified JWT).
- Compute the current window start:
Math.floor(Date.now() / 1000) / windowSize * windowSize. - Form the KV key.
- Increment the counter atomically with
KV.put(key, newValue, { expiration: ttl })(if the key does not exist we start at 1). - Read the updated value; if it exceeds the allowed limit, return HTTP 429.
Section 2: Worker middleware implementation
The following Worker shows the core logic. It assumes a KV namespace binding named RATE_LIMIT_KV and expects the user ID in a header X-User-ID (in practice you would validate a JWT and extract the sub claim).
addEventListener('fetch', event => {\n event.respondWith(handleRequest(event.request))\n})\n\nconst WINDOW_SIZE = 60 // seconds\nconst LIMIT = 100 // max requests per window\n\nasync function handleRequest(request) {\n const userId = request.headers.get('X-User-ID')\n if (!userId) {\n return new Response('Missing user ID', { status: 400 })\n }\n\n const nowSec = Math.floor(Date.now() / 1000)\n const windowStart = nowSec - (nowSec % WINDOW_SIZE)\n const key = `rate:${userId}:${windowStart}`\n\n // Get current count (may be undefined if key expired)\n let count = await RATE_LIMIT_KV.get(key)\n count = count === null ? 0 : parseInt(count, 10)\n\n if (count >= LIMIT) {\n return new Response('Too Many Requests', { status: 429 })\n }\n\n // Increment and store with TTL equal to window size\n await RATE_LIMIT_KV.put(key, String(count + 1), { expiration: WINDOW_SIZE })\n\n // Forward request to origin (or return a mock response)\n return fetch(request)\n}\n
Notes:
- The
expirationoption tells KV to automatically delete the key after the window ends, eliminating the need for manual cleanup. - Because KV is eventually consistent, under extremely high concurrency a request might see a stale count and briefly exceed
LIMIT. Adding a small safety margin (e.g., settingLIMIT = 90when the true quota is 100) mitigates this risk. - Cold starts add a few milliseconds; for traffic patterns with long idle periods you can keep the Worker warm via scheduled triggers or place the logic in a Service Binding if sub‑millisecond latency is required.
Worked example: testing locally with wrangler dev
- Create a KV namespace:
wrangler kv:namespace create RATE_LIMIT_KV. - Add the binding to
wrangler.toml:\n
\n[[kv_namespaces]]\nbinding = \"RATE_LIMIT_KV\"\nid = \"<your‑kv‑id>\"\n - Start the dev server:
wrangler dev. - Send test requests (replace
<user>with any string):\n
\nfor i in {1..10}; do\n curl -i -H \"X-User-ID: <user>\" http://127.0.0.1:8787/\ndone\n - Observe that the first 6 requests return 200 (or the origin response) and the 7th returns 429 Too Many Requests when
LIMIT=5andWINDOW_SIZE=10seconds. - Wait longer than the window, then repeat the loop; the counter resets and requests succeed again.
To verify behavior in a preview deployment:
- Deploy:
wrangler publish --env preview. - Check the KV namespace via the Cloudflare dashboard; you should see keys matching the pattern
rate:{userId}:{windowStart}with a TTL equal to the window size. - Inspect response headers (e.g.,
Retry-After) if you choose to add them.
Trade‑offs and monitoring
Using KV gives you global durability and low operational overhead, but the eventual consistency model means you cannot guarantee a hard ceiling under massive burst traffic. For strict guarantees you would need Cloudflare Durable Objects, which provide strong consistency at the cost of higher complexity and per‑object latency.
Monitoring can be done by exporting KV metrics via Cloudflare GraphQL API or by logging the counter value before the decision. Setting an alert when the hit rate approaches the limit helps you detect abusive patterns early.
Actionable closing
Implement the middleware above, tune WINDOW_SIZE and LIMIT to match your service’s SLA, and test with wrangler dev as shown. Remember to add a small safety margin to absorb KV’s eventual consistency lag, and consider Durable Objects only if you observe frequent limit violations under peak load. With these steps you get a lightweight, edge‑based per‑user rate limiter that protects your origin without adding significant latency.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.