Paginated Unary RPCs vs Server-Side Streaming in gRPC for Large Ordered Results
Compare paginated unary RPCs and server-side streaming for large ordered gRPC result sets, including latency, memory, retry behavior, and client compatibility.
06 Nov 2025, 04:31 UTC

Decision and constraints
When a gRPC service must return a large or unbounded ordered collection, the transport pattern is a design decision, not a default. The choice between a unary paginated RPC and a server-side streaming RPC changes latency, memory, retry behavior, and client compatibility. The constraints that usually decide the case are:
- Time to first item: how soon the client can process the first result.
- Server memory per request: how much state the server holds while serving one call.
- Retryability: whether a transient failure can resume without duplicates or gaps.
- Client and proxy support: whether gRPC-Web, load balancers, or language runtimes handle streaming as expected.
Also consider HTTP/2 stream limits, connection occupancy, and observability. The exact behavior of flow control, keepalive, and cancellation is version-sensitive, so verify against your gRPC implementation and runtime before committing.
Option comparison
| Pattern | First item | Server memory | Retry | State |
|---|---|---|---|---|
| Unary paginated | Waits for a full page | Low: page buffer per call | Easy: opaque page token | Stateless per RPC |
| Server-side streaming | Low: first item as ready | Higher: stream held open | Hard: restart stream | Stateful for stream lifetime |
Trade-offs
Unary pagination is simpler to operate. Each call is independent, so horizontal scaling is straightforward, error handling is ordinary unary status handling, and gRPC-Web or buffering proxies usually work without special configuration. The cost is more round trips and a higher time to first item when page size is small. Pagination tokens must be opaque and stable; do not expose raw offsets, and make sure tokens remain valid across deployments if clients can resume later.
Server-side streaming reduces round trips and improves perceived latency because the server can send each item as soon as it is available. The trade-off is resource occupancy: the server holds an HTTP/2 stream for the lifetime of the call, which can exhaust MAX_CONCURRENT_STREAMS under high fan-out. Partial failure recovery is harder because a broken stream usually means restarting from the beginning unless the application adds its own resume token. Clients must also handle cancellation and backpressure correctly.
Implementation sketch
The following proto shows both patterns for a hypothetical ItemService. It is illustrative, not tested output. Compile it with your normal protoc toolchain and inspect the generated stubs to confirm the method signatures.
Proto definition
syntax = "proto3";
package inventory;
service ItemService {
// Unary paginated RPC
rpc ListItemsPage(ListItemsRequest) returns (ListItemsResponse);
// Server-side streaming RPC
rpc ListItemsStream(ListItemsRequest) returns (stream Item);
}
message ListItemsRequest {
string parent = 1; // e.g., project or folder identifier
int32 page_size = 2; // ignored by streaming variant
string page_token = 3; // ignored by streaming variant
}
message ListItemsResponse {
repeated Item items = 1;
string next_page_token = 2;
}
message Item {
string id = 1;
string name = 2;
}
Server-side handling
For pagination, the server reads one page from the backing store and returns it. For streaming, the server iterates the source and sends each item, checking the stream context between sends so cancellation and deadlines are observed.
// Pseudocode: paginated unary handler
func (s *itemService) ListItemsPage(ctx context.Context, req *ListItemsRequest) (*ListItemsResponse, error) {
items, nextTok, err := s.store.Page(req.Parent, req.PageSize, req.PageToken)
if err != nil {
return nil, err
}
return &ListItemsResponse{Items: items, NextPageToken: nextTok}, nil
}
// Pseudocode: server-side streaming handler
func (s *itemService) ListItemsStream(req *ListItemsRequest, stream ItemService_ListItemsStreamServer) error {
iter := s.store.Iterate(req.Parent)
defer iter.Close()
for iter.Next() {
if err := stream.Send(iter.Item()); err != nil {
return err
}
if err := stream.Context().Err(); err != nil {
return err
}
}
return iter.Err()
}
Client behavior
A paginated client loops until next_page_token is empty. A streaming client reads until EOF and treats a mid-stream error as a restart unless the application adds a resume mechanism.
// Pseudocode: paginated client
pageToken := ""
for {
resp, err := client.ListItemsPage(ctx, &ListItemsRequest{
Parent: proj, PageSize: 200, PageToken: pageToken,
})
if err != nil {
// apply retry policy; token makes this idempotent
}
for _, item := range resp.Items {
process(item)
}
if resp.NextPageToken == "" {
break
}
pageToken = resp.NextPageToken
}
// Pseudocode: streaming client
stream, err := client.ListItemsStream(ctx, &ListItemsRequest{Parent: proj})
if err != nil {
// handle
}
for {
item, err := stream.Recv()
if err == io.EOF {
break
}
if err != nil {
// transient error: restart stream or use resume token
}
process(item)
}
Validation approach
To decide with evidence, deploy both variants against the same data source and drive them with a fixed dataset and concurrent clients. Measure time to first item, total completion time, server resident memory, and active stream count. Then inject a transient failure after N items and observe recovery cost. Pagination should resume from the last token without duplication; streaming will normally require a restart unless you built a resume protocol.
Practical checks:
- Compile the proto and inspect generated stubs: unary methods return a response, streaming methods return a stream type.
- Run a local client and observe first message arrival and stream termination on cancellation.
- Monitor active streams versus request count and memory growth under sustained load.
- Confirm that pagination tokens are opaque and stable across deployments.
Limitations and when to choose which
Choose unary pagination when retryability, proxy compatibility, simple operations, and bounded server memory matter more than first-item latency. Choose server-side streaming when the client can handle streams, the result set is large or unbounded, and low time to first item is the dominant requirement. Many systems use pagination for public or browser-facing APIs and streaming for internal, high-throughput consumers.
Streaming support is not uniform. gRPC-Web historically limits full duplex streaming without a transcoding proxy, and edge proxies may buffer responses. Flow control and keepalive behavior are version-sensitive, and cancellation propagation depends on correct context handling in both client and server libraries. Treat the decision as reversible only at the API design stage; changing the RPC type later is a breaking change for generated clients.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.