Streamlit Caching in Production: Choosing Between cache_data and cache_resource
Streamlit 1.18 split caching into st.cache_data (serializable returns) and st.cache_resource (singleton objects). Covers when to use each, TTL/LRU eviction, widget pitfalls, mutation risks, and distributed deployment patterns with a semantic search example.
07 Nov 2025, 07:41 UTC

The Problem: Cache Confusion in Production
You deploy a Streamlit app that loads a 500 MB model and queries a PostgreSQL database. After the first user session, memory usage climbs. A second user hits the same endpoint and gets a stale DataFrame. The logs show CachedWidgetWarning spam. You added @st.cache everywhere because the docs said it speeds things up — but now the app behaves unpredictably.
Since Streamlit 1.18, the legacy @st.cache decorator has been split into two distinct primitives: st.cache_data for serializable return values and st.cache_resource for unserializable objects like database connections, ML models, and HTTP clients. Understanding this distinction isn't optional for production apps — it's the difference between a stable deployment and a memory leak that crashes your container.
When to Use Each Decorator
The rule of thumb: if the return value can be pickled and you want a fresh copy on each cache hit, use st.cache_data. If the return value holds open file descriptors, GPU memory, or network sockets — things that cannot be meaningfully copied — use st.cache_resource.
st.cache_data: Serializable Data
- Pandas DataFrames, NumPy arrays, dicts, lists, primitives
- Stores a pickled copy per argument signature
- Returns a deep copy by default (since 1.32,
copy=Falseopts out) - Ideal for query results, computed features, downloaded CSVs
st.cache_resource: Singleton Objects
- SQLAlchemy engines, Redis clients, transformers pipelines,
requests.Session - Stores the actual object instance in a global map keyed by function identity
- Same object returned across all sessions and reruns within the process
- Not garbage-collected when a user session ends
# cache_data: returns a fresh DataFrame copy each hit
@st.cache_data(ttl=300, max_entries=10)
def load_sales_data(region: str) -> pd.DataFrame:
return pd.read_sql(f"SELECT * FROM sales WHERE region='{region}'", engine)
# cache_resource: returns the SAME engine instance every time
@st.cache_resource
def get_engine() -> sqlalchemy.Engine:
return sqlalchemy.create_engine(
"postgresql://user:pass@db:5432/app",
pool_size=5, max_overflow=10
)
engine = get_engine() # singleton
sales_df = load_sales_data("EMEA") # cached per region, TTL 5 min
Invalidation Strategies That Actually Work
Both decorators accept ttl (seconds or timedelta) and max_entries (LRU eviction). These operate independently: ttl expires entries by age, max_entries bounds memory by discarding least-recently-used entries.
Time-Based Expiration
Use ttl when data has a known freshness window — dashboard metrics, exchange rates, feature flags. The cache key includes the function's qualified name, source file hash, and argument values (via inspect.signature + pickle). Changing default argument values or the function body invalidates existing entries automatically.
Manual Invalidation
Call func.clear() to purge all entries for that function across all sessions. There is no per-session or per-user granularity without namespacing arguments yourself (e.g., include user_id in the cache key).
@st.cache_data(ttl=600, max_entries=50)
def fetch_user_features(user_id: int) -> dict:
return api_client.get(f"/users/{user_id}/features").json()
# Invalidate for one user after a write
if st.button("Refresh features"):
fetch_user_features.clear() # clears ALL users — consider namespacing
st.rerun()
Common Pitfalls and How to Avoid Them
Widget Calls Inside Cached Functions
Any st.* widget call inside a cached function raises CachedWidgetWarning and is ignored. Move widget calls outside and pass values as plain arguments:
# WRONG — widget inside cached function
@st.cache_data
def bad_example():
threshold = st.slider("Threshold", 0, 100, 50) # warning, ignored
return query_db(threshold)
# RIGHT — widget outside, value passed in
@st.cache_data
def good_example(threshold: int):
return query_db(threshold)
threshold = st.slider("Threshold", 0, 100, 50)
result = good_example(threshold)
Mutable Returns with copy=True (Default)
With st.cache_data, the default copy=True deep-copies the return value on every hit. Mutating the returned DataFrame in-place wastes CPU and memory without affecting the cache. For large DataFrames, use copy=False — but treat the returned object as read-only:
@st.cache_data(copy=False)
def large_dataframe() -> pd.DataFrame:
return pd.read_parquet("s3://bucket/large.parquet")
# SAFE: read-only operations
df = large_dataframe()
st.dataframe(df.head())
# DANGEROUS: mutation persists in cache, affects other sessions
df["new_col"] = 1 # mutates the cached object!
Distributed Deployments: No Shared Cache
The in-process cache is not shared across replicas. Two containers behind a load balancer maintain independent caches. For cross-instance coherence, attach a persistent backend (Redis, Memcached, or a database) via a custom wrapper or packages like streamlit-cache-redis. This is not built into Streamlit core.
Worked Example: A Production-Ready Pattern
Here's a pattern used in a real-time analytics dashboard serving 200+ concurrent users. The app loads a 200 MB embedding model and queries a read-replica PostgreSQL cluster.
import streamlit as st
import sqlalchemy
from sentence_transformers import SentenceTransformer
import pandas as pd
# 1. Singleton resources — initialized once per process
@st.cache_resource
def get_db_engine() -> sqlalchemy.Engine:
return sqlalchemy.create_engine(
st.secrets["postgres_dsn"],
pool_size=10, max_overflow=20,
pool_pre_ping=True # handles stale connections
)
@st.cache_resource
def get_embedder() -> SentenceTransformer:
# Loads ~200 MB model weights into GPU/CPU memory once
return SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
# 2. Cached data — TTL for freshness, LRU for memory bound
@st.cache_data(ttl=300, max_entries=200, copy=False)
def search_similar(query_embedding: list[float], top_k: int = 10) -> pd.DataFrame:
engine = get_db_engine()
# pgvector similarity search
sql = """
SELECT id, text, 1 - (embedding <=> %s) AS similarity
FROM documents
ORDER BY embedding <=> %s
LIMIT %s
"""
return pd.read_sql(sql, engine, params=(query_embedding, query_embedding, top_k))
# 3. UI layer — widgets outside cached functions
st.title("Semantic Search")
query = st.text_input("Search query")
top_k = st.slider("Results", 5, 50, 10)
if query:
embedder = get_embedder()
embedding = embedder.encode(query).tolist()
results = search_similar(embedding, top_k)
st.dataframe(results)
# 4. Admin: manual cache clear for deployments
with st.sidebar:
if st.button("Clear all caches (admin)"):
st.cache_data.clear()
st.cache_resource.clear()
st.success("Caches cleared")
st.rerun()
Key decisions in this pattern:
copy=Falseonsearch_similaravoids copying the result DataFrame (read-only downstream)pool_pre_ping=Trueprevents stale connection errors after database failover- Embedding model loaded once per process via
cache_resource, not per session - Cache keys include
query_embedding(list of floats) — hashable, stable - Admin clear button calls both
clear()methods for full reset after model updates
Trade-offs and Limitations
| Concern | cache_data | cache_resource |
|---|---|---|
| Cross-replica sharing | No (in-process only) | No (in-process only) |
| Per-session granularity | Only via argument namespacing | Not supported |
| Memory growth | Bounded by max_entries | Unbounded — persists for process lifetime |
| Mutation risk | High if copy=False | High — same object everywhere |
| Serialization cost | Pickle on miss/hit (unless copy=False) | None — object stored directly |
The biggest operational gap: no built-in distributed cache. If you run multiple replicas (Kubernetes, Cloud Run, Heroku), each replica warms its own cache independently. For read-heavy workloads, this means redundant computation. The workaround is a wrapper around Redis or a shared database table, but that adds infrastructure complexity.
Actionable Checklist for Your Next Deploy
- Audit existing
@st.cachedecorators. Split each intocache_dataorcache_resourcebased on return type. Removeallow_output_mutationandsuppress_st_warning— they don't exist in the new API. - Add
ttlandmax_entriesto everycache_datafunction. Start withttl=300(5 min) andmax_entries=100; tune based on memory profiling. - Set
copy=Falseon large DataFrame functions only after confirming downstream code treats them as read-only. Add a comment:# READ ONLY — cached object shared. - Move all
st.*widget calls outside cached functions. Pass widget values as arguments. This eliminatesCachedWidgetWarningand makes cache keys explicit. - Test invalidation. Run the app, trigger a cache hit, call
func.clear()from a button, verify fresh data loads. Confirmstreamlit cache clearCLI (added in 1.28) purges on-disk artifacts. - Plan for distributed deployment. If you'll run >1 replica, evaluate
streamlit-cache-redisor a custom Redis wrapper before you need it.
Run the verification steps from the Streamlit docs: create a minimal app with @st.cache_data(ttl=10) returning a timestamp DataFrame, confirm updates only after TTL expiry. Wrap sqlalchemy.create_engine in @st.cache_resource and verify id(engine) is stable across reruns. These five-minute tests catch 90% of caching bugs before they reach production.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.