Using Facebook's TAO for Low‑Latency Social‑Graph Queries
Learn how TAO’s cache‑backed MySQL store serves billions of graph reads with sub‑millisecond latency, and how to handle its eventual consistency in practice.
21 Jan 2026, 09:22 UTC

Problem: needing fast, flexible reads from a massive social graph
When building features like News Feed or Graph Search, a service must retrieve thousands of edges (friend, like, comment) per request with sub‑millisecond latency while the underlying data store continues to accept writes from billions of users. A traditional relational schema would require costly joins or frequent migrations for new edge types.
Thesis: TAO’s cache‑backed MySQL store gives low‑latency, schema‑free reads, but applications must tolerate or mitigate its eventual consistency.
How TAO works
TAO stores each object (user, page, photo) and each edge (friendship, like) as a row in a MySQL table. A geographically distributed cache layer sits in front of MySQL and serves read requests directly from memory when the data is hot. Writes are first applied to the MySQL master and then propagated asynchronously to replicas; the cache is updated via a invalidation‑or‑refresh mechanism after the write commits. This split lets reads achieve the speed of an in‑memory store while writes retain the durability of MySQL.
Worked example: querying a friend edge
Assume a service wants to know whether user A (ID 12345) follows user B (ID 67890). Using the TAO client library (available in the open‑source facebook/tao repo), the call looks like:
// Java‑style pseudocode; run inside a service that has TAO client access
TaoAsyncClient tao = TaoAsyncClient.create(); // configured with cluster endpoints
Long src = 12345L;
Long dst = 67890L;
String edgeType = "friend"; // arbitrary edge type, no schema change needed
CompletableFuture exists = tao.get(src, edgeType, dst);
exists.thenApply(present -> {
if (present) {
// edge exists – proceed with downstream logic
} else {
// edge absent – handle accordingly
}
return null;
});
The get method first checks the local cache; a hit returns immediately (typically <1 ms). On a miss, TAO forwards the request to the MySQL backend, populates the cache, and returns the result. No schema migration is required to add a new edgeType.
Mitigating stale reads
Because the cache may hold a version of an edge that precedes the latest write, a read can temporarily return stale data. Applications that require read‑after‑write consistency (e.g., showing a newly added friend immediately) should:
- Perform the write via TAO’s
mutateAPI. - Implement a short‑lived read‑after‑write barrier: after the mutate completes, issue a read with the
forceFreshflag (or bypass the cache) for the specific object/edge pair. - Monitor the latency of the forced‑fresh read; if it exceeds a threshold, consider falling back to eventual consistency for non‑critical paths.
This pattern adds a small latency penalty (usually a few milliseconds) only for the critical path, while the majority of reads continue to benefit from the cache.
Trade‑offs and operational considerations
- Eventual consistency: Stale reads are inevitable; design UI to tolerate brief inconsistencies or use read‑after‑write where needed.
- Cache warm‑up: After a cluster restart or scaling event, the cache is cold and read latency can spike until hot data is reloaded. Pre‑warming with known hot keys mitigates this.
- Capacity planning: Mis‑sized cache memory leads to thrashing, increasing load on MySQL masters. Monitor cache hit ratio (target > 90 %) and MySQL QPS to detect overload.
- Operational complexity: Operating TAO requires custom tooling for cache invalidation, replication lag monitoring, and failure detection.
Actionable closing
If you are evaluating a graph store for a read‑heavy social‑graph workload, TAO demonstrates how a simple MySQL backend paired with a smart cache can deliver sub‑millisecond reads at massive scale. Start by instrumenting read latency and cache hit ratio in your environment, adopt the read‑after‑write pattern for any path that demands fresh data, and plan cache warm‑up procedures for deployments or scaling events. These steps let you harness TAO’s performance strengths while keeping its consistency limitations under control.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.