Short answer
Keto does not, as of the versions where this question typically arises, document a general-purpose cache-invalidation API or a /flush endpoint with guaranteed cross-replica semantics. If you are seeing stale check results after writing relation tuples, the most likely cause is not a broken invalidation mechanism inside Keto — it is almost always one of three things: a read replica between Keto and its database, an application-side cache in front of Keto, or in-process state that a restart will clear. Treat the premise of a "Keto cache invalidation strategy" with suspicion until you verify which layer is actually serving the stale answer.
Separating confirmed behavior from assumption
Confirmed: Keto separates writes (relation tuples) from reads (check/expand APIs). Its consistency guarantees come from the backing SQL store (PostgreSQL, MySQL, or CockroachDB). With a single primary and default settings, a check issued after a committed write should see that write — the database provides read-your-writes, not Keto.
Likely explanation, not confirmed: When operators report "stale cache after policy update," the usual culprits are:
- Keto's DSN points at a lagging read replica, so checks read pre-write state.
- The application (or a gateway/CDN) caches check responses; Keto has no way to invalidate a cache it doesn't own.
- Connection pooling or in-process state in a specific pod, which is why the symptom appears on some replicas and not others.
Because caching behavior and configuration keys vary across Keto releases, verify against the changelog and config reference for your version before designing around any assumed TTL or flush setting.
Diagnosing your deployment
- Write a relation tuple, then immediately call the check API against the same Keto endpoint. Log the result.
- Repeat with Keto pointed directly at the database primary (bypassing any replica or proxy). If staleness disappears, the replica lag is your answer.
- Restart one Keto pod and rerun the failing check sequence. If the stale result vanishes after restart, in-memory state or pooled connections are involved, not the database.
- Audit everything between your caller and Keto: HTTP caches, service-mesh layers, and any application-level memoization of check results.
What to do about each cause
Replica lag: route both writes and checks to the primary, or accept a bounded retry window for checks that must reflect a just-completed write. Do not add fixed sleeps — they mask the problem and fail under load.
Application-side cache: this is the one case where you genuinely need an invalidation strategy, and it is yours to build. On every relation-tuple write, publish an invalidation event (e.g., a pub/sub message keyed by the affected subject/object) so each application's cache entry is dropped. A short TTL alone is a fallback, not a guarantee.
Multi-replica Keto: since each pod is stateless with respect to the database, consistency across replicas is a database concern, not a Keto concern — provided all pods talk to the same primary. If you believe a /flush endpoint exists in your version, confirm it in that version's docs before relying on it; do not assume it propagates across pods.
Answering the strategy question
For high-traffic environments, the reliable pattern is: no cache in front of Keto's check API, primary-only reads, and — if latency forces a client-side cache — event-driven invalidation keyed to tuple writes with a short TTL as backstop. A write-through cache adds complexity without fixing replica lag, and TTL-only accepts a window of wrong authorization decisions that is rarely acceptable for permission systems.
One detail that would change this advice
Which Keto version are you running, and does your check path pass through any proxy or cache you control? If staleness persists with primary-only reads, no client cache, and after a restart, that points to something version-specific worth escalating with a minimal reproduction.