BigPipe: How Facebook Streamed Pages Before Streaming SSR Was Cool
Facebook's BigPipe streamed pages in chunks instead of waiting for the slowest backend service. Here's how pagelets worked, what they cost, and why the idea still matters.
30 Dec 2025, 08:32 UTC

Imagine a page that needs data from five backend services. Four respond in 50ms; the fifth takes 800ms. Under classic server-side rendering, the server waits for all five before sending a single byte of HTML. The user stares at a blank tab for the full 800ms, even though most of the page was ready almost immediately. That was the problem Facebook's engineering team attacked with BigPipe around 2009–2010, and the idea they shipped still shapes how streaming rendering works today.
The thesis: don't wait for the slowest dependency
BigPipe's core decision was to stop treating a page as one indivisible response. Instead, the server sends a page skeleton immediately — the layout, CSS, and JavaScript — and then streams each section of content as its data becomes available. The browser starts rendering and the user starts reading while slower fragments are still being computed on the server.
The important part of the thesis is what it did not require: backend services didn't need to change their interfaces or become faster. The improvement came purely from overlapping server work with network transfer and browser rendering, which is why it was deployable across a large existing codebase.
How pagelets actually work
BigPipe decomposes a page into independent chunks called pagelets. The flow looks like this:
- The server receives the request and quickly renders the page skeleton: the document head, the overall layout, and an empty placeholder
<div>for each pagelet. It flushes this to the browser right away, keeping the HTTP connection open. - As each backend dependency resolves, the server renders that pagelet's HTML and flushes it down the same connection as a small chunk of JavaScript — essentially a call like
onPageletArrive({id: "pagelet_ads", content: "..."}). - A tiny client-side script receives each chunk, finds the matching placeholder, and swaps in the real markup.
Because the connection stays open and chunks arrive in dependency order, the browser progressively fills in the page rather than waiting for one monolithic response. If this sounds familiar, it should: modern streaming SSR in React (renderToPipeableStream with <Suspense>) is conceptually the same pattern — shell first, then streamed fragments that hydrate placeholders.
A worked example: the News Feed
Consider a simplified News Feed page with three pagelets:
- Header and navigation — cheap to render, available almost instantly.
- First feed stories — moderate cost, ready a few hundred milliseconds in.
- Right-rail ads — depends on an ad-ranking service, the slowest dependency.
Under classic rendering, the ads service holds the whole page hostage. With BigPipe, the skeleton and header flush first, the stories arrive next, and the user is already scrolling when the ads pagelet finally lands in its placeholder. Perceived load time drops dramatically even though total server work is unchanged — the win comes from reordering when bytes reach the browser, not from making anything faster.
The trade-offs nobody mentions in the demo
Streaming partial pages is not free. The engineering costs are real:
- Ordering and dependencies. Some pagelets depend on others (a comment box needs its parent story). You need an explicit dependency graph and a scheduler, not just independent flushes.
- Error handling mid-stream. Once the skeleton is flushed, you can no longer return a 500 or redirect. A failed pagelet has to be handled in-band — render an error state into the placeholder and log aggressively, because partial failures are now a normal case.
- Client state complexity. JavaScript that assumes the whole DOM exists at load time breaks. Pagelet-scoped initialization becomes necessary.
- Caching and SEO. A partially streamed response is harder to cache at the edge, and crawlers historically handled progressively injected markup poorly. Facebook's audience was logged-in users, so SEO mattered little — your mileage may differ sharply.
What to take from it
BigPipe is largely historical — it was built for Facebook's PHP/Hack stack, and the company's modern frontend uses different rendering strategies. But the engineering decision behind it is durable: identify your slowest dependency, and ask whether it should block the first byte. You can apply this today without any Facebook-era machinery: most modern frameworks support streaming responses, and even a hand-rolled version — flush the head and shell early, then stream fragments — is a few lines in Node:
// Express-style sketch; run on your server, Node 14+
res.writeHead(200, { 'Content-Type': 'text/html; charset=utf-8' });
res.write(renderShell()); // skeleton + placeholders, flushed now
const stories = await fetchStories();
res.write(`<template data-pagelet="stories">${stories}</template>`);
const ads = await fetchAds(); // slow dependency arrives last
res.write(`<template data-pagelet="ads">${ads}</template>`);
res.end();A small client script moves each <template> into its placeholder. To verify the benefit, open your browser's network panel and check the "Timing" tab: you should see a low Time to First Byte with content arriving in chunks, rather than one long wait followed by a single large download. If TTFB still equals total server time, your framework or proxy is buffering — check for gzip middleware or an nginx proxy_buffering on setting, both of which silently defeat streaming.
The lesson BigPipe leaves behind is less about pagelets and more about a posture: perceived performance is a scheduling problem. When you can't make the work faster, change the order in which its results reach the user.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.