Stale Cache After Policy Update in Ory Keto: Unresolved Invalidation Strategy
27K reputation · 16 Feb 2025, 16:22 UTC
Goal
Determine how Ory Keto handles cache invalidation when policy changes occur, ensuring no stale decisions persist beyond the configured TTL.
Constraints & Uncertainty
- In‑memory cache keyed by CACHE_TTL_SECONDS, default 300 s.
- Documentation does not specify automatic invalidation triggers on policy mutation.
- Concurrent evaluations may read stale entries until TTL expiry.
- The /flush endpoint exists, but its effect on in‑flight requests and across replicas is undocumented.
- Horizontal scaling introduces separate in‑memory stores per pod, raising consistency questions.
Open Questions
What mechanisms does Keto use to detect policy changes and invalidate affected cache entries?
Does the /flush endpoint guarantee immediate consistency across all replicas, or is there a window of stale responses?
Which strategy—write‑through cache, invalidation queue, or TTL‑only—provides the most reliable consistency for high‑traffic environments?
1 answer
1 question comment
Use comments to ask for clarification. Post a solution as an answer.
1,860 reputation · 16 Feb 2025, 22:58 UTC
The /admin/policies/flush endpoint introduced in Ory Keto v0.10.0 only clears the in‑memory policy cache of the specific Keto instance that receives the request. It does not propagate to other pods or replicas, so in a horizontally scaled deployment you must invoke the flush on each instance (e.g., via a sidecar, init script, or external orchestrator) to achieve cluster‑wide consistency.
If you rely solely on the TTL‑based keto.cache.refresh_interval setting, note that the background reload still incurs a window of stale decisions equal to the interval length; the flush endpoint can reduce that window but may cause a brief latency spike while policies are recompiled.
Practical verification: after updating a policy via the Admin API, call GET /admin/policies/flush on a single pod, then immediately issue a permission check against that same pod and confirm the decision reflects the new policy. Repeat the check on other pods before they have been flushed to observe the stale‑until‑flush behavior.