DigitalOcean Read Replicas: When Offloading Reads Beats a Bigger Node
DigitalOcean read replicas can take read load off your primary — if your queries tolerate slightly stale data. How to set one up, measure lag, and decide what stays put.
27 Nov 2025, 22:27 UTC

Your managed PostgreSQL cluster on DigitalOcean has been fine for months. Then the analytics dashboard — a few heavy, read-only aggregate queries — starts running while customers are checking out, and writes on the primary slow to a crawl.
The reflexive fix is a bigger node. The targeted fix, if your workload fits, is a read replica: a second node that continuously copies the primary's data and serves read traffic from its own endpoint, leaving the primary free to handle writes. The decision hinges on one question: can the reads you want to offload tolerate data that is a few seconds old? If yes, a replica is often the cheapest scaling lever available. If no, replicas will not help — they are eventually consistent by design, meaning reads can briefly reflect older data.
What a read replica gives you
DigitalOcean's managed database platform supports read replicas for PostgreSQL and MySQL; this post uses PostgreSQL, so confirm availability before planning around other engines. The mechanics are asynchronous streaming replication: the primary ships every change to the replica, which applies it with a small delay. Same-region setups are low-latency in practice, but the number that matters is what you measure on your own cluster.
A few properties worth knowing up front:
- Each cluster supports up to five read replicas — confirm the current limit in the control panel before designing around it.
- Every replica gets its own connection string, so routing reads to it is an application configuration change, not a DNS project.
- Replicas inherit the cluster's engine version and configuration; you pick each replica's plan size, and each is billed at the standard hourly rate for that plan.
- Connections pass through the same trusted-sources firewall (the allowed client IP list) that guards the primary, and TLS is required.
Worked example: a reporting replica with a lag check
Say the dashboard queries belong to a separate reporting service. The goal: give that service its own replica and prove it stays within a few seconds of the primary before trusting it with real traffic.
Create and route the replica
Run these from your workstation or CI, with doctl authenticated using a read-write API token (generate one in the control panel's API section). Replace <cluster-id> with your cluster's ID and the size slug with the plan you want:
# Find your cluster ID
doctl databases list
# Create a replica; check flags with: doctl databases replica create --help
doctl databases replica create <cluster-id> reporting-replica --size db-s-2vcpu-4gb
# Watch it come online
doctl databases replica list <cluster-id>Expect the replica to reach online status within a few minutes, and confirm it appears under the cluster's Read Replicas tab in the control panel. Get the replica's connection details from the panel (or doctl databases replica get): a dedicated host on port 25061 for PostgreSQL, with sslmode=require in the connection string. Point the reporting service's connection string at the replica host and leave everything else untouched.
Prove it is keeping up
The check that matters is replication lag. Run this on the replica, for example with psql using the replica's connection string:
SELECT now() - pg_last_xact_replay_timestamp() AS replay_delay;pg_last_xact_replay_timestamp() returns the transaction time of the last change the replica replayed, so the result is how far behind it is. Two caveats keep the check honest: run it during normal write traffic, and remember that on a quiet database the number grows even when replication is healthy, because no new transactions are arriving to replay. A same-region replica should sit well under a few seconds under load. If the delay climbs steadily during your busiest hour, the replica is falling behind and needs a bigger plan — or a smaller read set.
One risk: the replica is billable from the moment it is created. If the experiment does not earn its keep, delete it with doctl databases replica delete <cluster-id> <replica-id> and the cluster returns to its previous shape.
The trade-off: define a staleness budget
Replicas are eventually consistent, and the classic failure is read-after-write. A user changes their email, the write lands on the primary, and the settings page — served from the replica — shows the old address for a few seconds. That is not a bug; it is the design.
The workable rule: route only staleness-tolerant reads to the replica — analytics, search, feeds, public pages, internal dashboards. Keep anything where users expect to see their own write on the primary. If a page mixes both, read it from the primary or accept the staleness explicitly rather than by accident.
| Situation | Better fit |
|---|---|
| Heavy read-only queries, slightly stale data acceptable | Read replica |
| Write-bound, or running out of RAM and storage | Bigger primary plan |
| Automatic failover when the primary dies | HA standby nodes, not replicas |
| The same small queries over and over | A cache, before any of this |
Where replicas stop helping: failover
It is tempting to treat a replica as a hot spare. It is not one. If the primary fails, promoting the replica is a manual step — via the control panel or doctl databases replica promote (check doctl databases replica --help for the exact arguments in your doctl version) — and promotion is one-way: the replica detaches and becomes a standalone cluster that cannot be turned back into a replica. Your original primary keeps running until you decommission it, which is why promotion works as a migration path but not as automatic failover.
If failover is the actual requirement, DigitalOcean's high-availability option for managed databases — standby nodes with automatic failover — is the feature built for it. Replicas and HA standbys solve different problems, and paying for the wrong one is an easy mistake. Cross-region replicas, where available for your engine and region, add replication delay and data-transfer costs and may not fit data-residency requirements, so verify both before planning one.
A practical way to act on this
Pick the read-heaviest path in your application this week. Create one replica at the smallest plan that handles its queries, measure replay delay under real traffic, and set an explicit staleness budget — for example, dashboard data may be up to 30 seconds old. Check the lag query against that budget during your busiest hour. If the replica holds it, move the next read path. If not, you have learned that for the cost of one small node for a few days, and you can delete it.
The primary keeps doing what it was doing the whole time. That is the point: replicas let you scale the reads that can wait, without touching the writes that cannot.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.