Architecting High-Concurrency Feeds: The Fan-Out Pattern
Learn how the fan-out pattern solves the read-latency problem in social feeds and why a hybrid push/pull model is necessary to handle celebrity accounts and write amplification.
19 Jul 2025, 23:11 UTC

The Latency Trade-off in Feed Generation
The primary challenge in building a social feed is the conflict between write-time efficiency and read-time latency. If a system calculates a user's feed at the moment they open the app (Read-time Aggregation), the database must perform complex joins across millions of rows to find every person the user follows and sort their latest posts. This creates unacceptable latency for the end user.
The solution is Fan-out: shifting the computational burden from the reader to the writer. Instead of calculating the feed on demand, the system pre-computes the feed and stores it in a fast-access cache. When a user posts, the system \"fans out\" that post to the caches of all their followers.
The Smallest Suitable Design
A basic fan-out architecture requires three core components:
- Tweet Store: A durable database (e.g., NoSQL or Distributed SQL) that holds the permanent record of the tweet.
- Timeline Cache: An in-memory store (like Redis) that holds a list of Tweet IDs for each active user. This is a transient, ordered list, not the full tweet content.
- Fan-out Worker: An asynchronous process that reads the follower list of the author and updates the Timeline Caches of those followers.
Trust and Data Boundaries
To maintain system stability, the architecture must strictly separate the Source of Truth from the Delivery Mechanism.
| Boundary | Responsibility | Persistence | Consistency Requirement |
|---|---|---|---|
| Tweet Store | Permanent storage of content | Disk-based | Strong/Eventual |
| Timeline Cache | Fast retrieval of IDs | Memory-based | Ephemeral |
By treating the Timeline Cache as disposable, the system can recover from cache crashes by re-populating the list from the Tweet Store without losing user data.
Handling the 'Celebrity' Problem (Hybrid Model)
Pure fan-out fails when a user has millions of followers. If an account with 50 million followers posts, the system must perform 50 million write operations to different caches. This creates write amplification, which can clog the message queue and delay delivery for all other users.
To solve this, the system implements a Hybrid Model based on user profile:
- Standard Users: Use the Push model. Their tweets are fanned out to followers' caches.
- High-Profile Users: Use a Pull model. Their tweets are not fanned out. Instead, when a follower loads their feed, the system pulls the latest tweets from the high-profile user and merges them into the cached feed in real-time.
Operational Checks and Failure Modes
Monitoring a fan-out system requires tracking the Fan-out Lag. This is the time delta between when a tweet is written to the Tweet Store and when it appears in the last follower's cache.
Common Failure Modes:
- Queue Backlog: A surge in activity can overwhelm fan-out workers. If the queue grows too large, users experience \"stale feeds\" where tweets appear minutes or hours late.
- Cache Eviction: To save memory, inactive users' caches are evicted. The first request after a long absence will trigger a \"cache miss,\" forcing the system to rebuild the timeline from the Tweet Store, causing a temporary spike in read latency.
Verification and Testing
To verify the health of the fan-out pipeline, engineers can use a canary account to track delivery speed. Run the following conceptual check on the distribution queue:
# Check the depth of the fan-out processing queue
# Required permissions: Queue Administrator
# Expected result: Queue depth should be within 10% of baseline during peak hours
queue-cli get-depth --queue fanout_distribution_queue
If the depth increases while CPU utilization on workers remains low, the bottleneck is likely I/O or lock contention in the cache layer.
Conditions for Design Change
The hybrid fan-out design remains effective until the following conditions are met:
- Extreme Read Volume: If the number of \"high-profile\" users grows so large that the read-time merge becomes the primary bottleneck.
- Real-time Requirements: If the business requires sub-millisecond consistency across all followers, regardless of account size, necessitating a move toward more complex streaming architectures like Apache Flink.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.