The Cause of Stale Results
The consistency level that explains stale CheckPermission results is minimize_latency (bounded staleness). With this level, SpiceDB may serve the request from the dispatch cache or a replica snapshot that has not yet caught up with the latest writes. That creates a window where a successfully written relationship is not yet visible on the read path — exactly what shows up after a deployment, when caches are cold, revisions are quantized, and clients are not yet propagating fresh tokens.
Likely Explanation vs. Confirmed Facts
Likely explanation: your clients are not capturing and passing the zedtoken returned by WriteRelationships, so every read is effectively a blind read at whatever snapshot the serving node happens to hold. Under minimize_latency, revision quantization and the dispatch cache make that snapshot even older than strictly necessary.
Confirmed facts:
minimize_latency does not guarantee the most recent write is visible.at_least_as_fresh requires a stored zedtoken and guarantees the read is at least as current as that token.at_exact_snapshot pins the read to one revision; fully_consistent is linearizable but costs latency and leader load.- Flag names and cache defaults vary by version, so verify against your release's
spicedb serve --help output.
Diagnostic Signals
- A/B the consistency level: repeat the same
CheckPermission with fully_consistent. If it returns the correct answer while minimize_latency does not, you have a staleness problem, not a schema or tuple problem. A schema/tuple bug would be wrong at every consistency level. - Compare revisions to head: check the revision reported for the read against the datastore head revision (visible in metrics or the datastore's head-revision endpoint, depending on backend). A persistent gap points to lag or cache staleness; matching revisions with a wrong result points back to the schema or the written tuples.
- Dispatch cache hits: a high hit rate on the affected permission path right after the write, combined with a revision gap, confirms the cached negative result is being replayed.
The GC Window Factor
The datastore GC window bounds how old a snapshot can be. A zedtoken older than the window produces a hard error, not a stale answer — so if you see errors rather than outdated ALLOW/DENY responses, suspect token age, especially for long-lived sessions or tokens persisted across deployments. If you see silently stale results, the token is either absent or within the window but behind head. Ensure the GC window comfortably exceeds your longest client token lifetime.
Recommendation
at_least_as_fresh with stored zedtokens is the right default fix. It gives causal consistency (users see their own writes) without the cost of fully_consistent on every call. Reserve fully_consistent for rare paths that must be linearizable, and use minimize_latency only where staleness is genuinely acceptable.
One missing detail that changes the recommendation: are the stale reads hitting all replicas or only some? If only specific nodes lag, the fix is node synchronization or health-checking, not the client consistency level.