Understanding ZedToken Stale Reads
A stale read failure occurs when the ZedToken provided in a request specifies a consistency timestamp that the underlying PostgreSQL datastore cannot yet satisfy or has already surpassed in a way that violates the requested consistency level. In SpiceDB, this typically manifests when a client requests a read at a specific snapshot (via the token), but the database node handling the request is lagging behind that timestamp or the token itself is outdated relative to the required consistency point.
Conditions Triggering Failures with PostgreSQL
When using PostgreSQL as the backend, stale read errors are generally triggered by the following conditions:
- Replication Lag: In a distributed PostgreSQL setup, if the read request hits a replica that has not yet ingested the WAL (Write Ahead Log) entries corresponding to the ZedToken's timestamp, the system cannot guarantee the requested consistency and returns a stale read error.
- Token Mismatch: If a client reuses a cached ZedToken after a subsequent write has occurred, and the new request requires a consistency level that must include that latest write, the old token is considered stale.
- Clock Skew: Significant drift between the SpiceDB nodes and the PostgreSQL primary can lead to timestamps that are logically impossible to satisfy immediately.
Synchronization Lag vs. Permanent Inability
SpiceDB distinguishes between temporary lag and permanent failure primarily through timeouts and retry logic rather than a distinct "permanent failure" status code.
A temporary synchronization lag is treated as a transient state; the system will typically wait for the datastore to catch up to the requested timestamp until the request context deadline is reached. If the timeout expires before the datastore reaches the token's timestamp, the request fails. A "permanent" inability is usually an indicator of a configuration error, such as a disconnected replica or a corrupted transaction log, which manifests as a persistent connection error or a timeout that never resolves regardless of the token used.
Resolution Steps
- Refresh the Token: Immediately obtain a new ZedToken from the most recent write operation before performing the read.
- Avoid Long-term Caching: Do not cache ZedTokens across unrelated sessions; treat them as short-lived pointers to a specific database state.
- Verify Replication Health: Check the PostgreSQL replication lag metrics to ensure replicas are keeping pace with the primary.
- Update Client Libraries: Ensure you are using the latest version of the SpiceDB client to avoid known bugs in token encoding or handling.
Diagnostic Note: To refine this recommendation, please provide the consistency level being requested (e.g., at_least_as_fresh_as) and whether you are using a single-node or clustered PostgreSQL deployment.