Architecture Note: Google Cloud Spanner External Consistency
An architecture note on achieving external consistency with Google Cloud Spanner's TrueTime: minimal multi-region design, trust boundaries, operational signals, failure modes, and design-change triggers.
15 Feb 2026, 10:55 UTC

Requirements
Applications that need globally distributed, strongly consistent reads and writes without implementing application-level conflict resolution face a fundamental tension: linearizable ordering across data centers requires a global clock, but physical clocks drift. Spanner solves this by exposing external consistency—the guarantee that transactions appear to execute at a single instant between their start and commit—using TrueTime, Google's globally synchronized clock service. The requirement set is narrow: transactional workloads that cannot tolerate stale reads or write conflicts, bounded latency (typically tens to low hundreds of milliseconds), and no offline or eventually consistent write paths.
Smallest Suitable Design
The minimal viable topology is a single Spanner instance configured for multi-region deployment with one primary region and one or more read‑only replica regions. A typical starter configuration:
instance-config: regional-us-central1 (primary)
replicas:
- region: us-east1
type: READ_ONLY
- region: europe-west1
type: READ_ONLY
Writes route to the primary region; reads can be served locally from any replica. The Spanner client library handles transaction routing automatically—read‑write transactions go to the primary, read‑only transactions can specify a region via TransactionOptions.ReadOnly with exact_staleness or strong (default). No additional coordination layer, consensus service, or application-level timestamp assignment is needed.
Trust and Data Boundaries
Encryption at rest (AES‑256) and in transit (TLS 1.2+) is managed entirely within Spanner. The trust boundary ends at the Spanner API frontier: clients authenticate via a service account with IAM roles (e.g., roles/spanner.databaseUser), and TrueTime itself is not exposed—clients never see uncertainty intervals or clock synchronization internals. Replication traffic stays on Google's private network. If your threat model requires customer‑managed encryption keys (CMEK), enable them at instance creation; rotation and access logging are then handled through Cloud KMS.
Operational Checks
Monitor these signals continuously; they are the observable surface of TrueTime health:
- Commit latency (p50, p99) — sustained increases often correlate with TrueTime uncertainty growth.
- Transaction abort rate — spikes indicate contention or clock uncertainty forcing retries.
- TrueTime uncertainty (internal metric, visible via Cloud Monitoring) — not tunable, but alerts on sustained >10 ms uncertainty are warranted.
- Replication lag — seconds‑level lag on read‑only replicas is normal; minutes indicates regional stress.
- Node/replica health — CPU, storage, and split points per node.
Example alerting query (PromQL‑style):
rate(spanner_transaction_abort_total[5m]) > 0.05
AND
spanner_commit_latency_seconds_bucket{le="0.5"} < 0.9
Investigate when both conditions hold: high aborts with elevated latency suggest TrueTime pressure, not just application contention.
Failure Modes
| Scenario | Behavior | Client Impact |
|---|---|---|
| Primary region outage | Automatic failover to a designated secondary (if configured as witness or read‑write); write unavailability window typically 10–60 s. | Read‑write transactions fail with UNAVAILABLE; read‑only transactions on healthy replicas continue. |
| TrueTime uncertainty spike | Commit wait increases to absorb uncertainty; aborts rise for contended keys. | Latency jumps (2–10× baseline); exponential backoff with jitter on retries is mandatory. |
| Network partition between regions | Spanner chooses consistency over availability for the affected partition; minority side rejects writes. | Same as primary outage for the partitioned region; no split‑brain because TrueTime provides a total order. |
Spanner avoids split‑brain by construction: TrueTime timestamps impose a global order, so two partitions cannot both believe they hold the latest write.
Conditions That Would Change the Design
- Multi-region write requirement — If writes must originate from multiple regions with low latency, the single-primary model breaks. Consider dual-region instance configs (e.g.,
dual-region-us-central1-us-east1) or accept higher latency with a single primary. - Offline or edge write needs — Spanner does not support disconnected operation. Layer a local queue (e.g., Cloud Pub/Sub + Dataflow) for ingestion, then reconcile.
- Cost sensitivity at scale — Node-hour pricing and inter-region replication traffic can dominate. Evaluate read-only replica count versus read latency SLOs; each additional replica adds ~30% storage cost.
- Version/configuration drift — Instance configs and consistency guarantees evolve (e.g.,
REGIONALvsMULTI_REGIONALvsDUAL_REGIONplacement options). Re-validate placement against current documentation before each major deployment.
Verification Checklist
- Inspect
gcloud spanner instances describe INSTANCE_ID --format=jsonto confirm region placement and replica types. - Run a load test (e.g., 1k QPS mixed read/write) and capture commit timestamps and abort rates from Cloud Monitoring; verify aborts stay <2% under steady state.
- Review IAM bindings:
gcloud spanner databases get-iam-policy DATABASE_ID --instance=INSTANCE_ID— ensure the client service account has onlyspanner.databases.select,spanner.databases.update, andspanner.databases.beginOrRollbackTransactionas needed.
Limitations
TrueTime uncertainty is internal and non-tunable; design must tolerate latency excursions during clock skew events (e.g., leap seconds, GPS anomalies). Spanner does not replace application caching for read-heavy workloads—strong consistency carries a latency floor. Failover timing and durability guarantees depend on the specific instance configuration chosen; consult the SLA for your config type before committing to RPO/RTO targets.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.