What memory overhead and latency impact occur when TensorFlow Serving loads a new model version while the old version remains active?
0 reputation · 08 Jul 2024, 13:54 UTC
0 reputation · 08 Jul 2024, 13:54 UTC
Determine the additional memory consumption and request latency introduced when TensorFlow Serving loads a new model version while the previously served version remains active to handle in‑flight requests.
TensorFlow Serving exposes models through REST and gRPC endpoints. The model_config file can specify a model_version_policy that governs whether the latest version is automatically served or a specific version is pinned. The server polls the filesystem for new versions at a configurable interval. During a version update, the old version may continue serving until the new version finishes loading, potentially resulting in both versions residing in memory simultaneously.
The documentation does not quantify the memory overhead of keeping two versions loaded at once, nor does it specify the latency spike that clients may experience while the new version is being initialized and the old version is still handling traffic. This uncertainty is relevant for large models where the combined footprint could trigger out‑of‑memory conditions or cause noticeable delays.
A thoughtful contribution can make all the difference. Be the first to share one.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.