Optimizing Streamlit Apps with st.cache_data: A Practical Guide
Streamlit’s @st.cache_data memoizes heavy data functions by hashing arguments, reducing recomputation on reruns. This guide shows how to configure TTL, size limits, and avoid stale data with a concrete CSV‑loading example.
26 Jul 2026, 08:30 UTC

Problem: Heavy Data Loads Slow Down Every Rerun
In a typical Streamlit dashboard you might read a 200 MB CSV, run a faceted group‑by, or query a remote API. Every time the user clicks a widget, the entire script re‑executes, pulling the data again and re‑calculating. The result is a sluggish experience and wasted compute.
Thesis: @st.cache_data Memoizes Results by Argument Hash
The @st.cache_data decorator stores a function’s output in memory the first time it runs. Subsequent calls with the same arguments fetch the cached result instead of re‑executing the function. Internally Streamlit hashes the function arguments; the hash becomes the cache key. If the arguments change, the hash changes and a new computation occurs.
Key Terms Explained
- Hash – a fixed‑size string derived from input data. If two inputs produce the same hash, they are considered equal for caching purposes.
- TTL (Time‑to‑Live) – the maximum age a cached entry remains valid before it is automatically invalidated.
- max_entries – the maximum number of distinct cache entries allowed. When exceeded, the least‑recently used entry is evicted.
Configuring the Cache
Typical usage is a single decorator line:
import streamlit as st
@st.cache_data(ttl=3600, max_entries=10)
def load_large_csv(path):
return pd.read_csv(path)
Here the cache expires after one hour and only ten different CSV paths can be cached at once. If you need a persistent cache across app restarts, you must store the data externally (e.g., in a database or a file) and load from there.
Worked Example: Loading a 200 MB CSV
Below is a minimal Streamlit app that reads a large CSV and shows a simple histogram. The first run takes ~4 s. After caching, the same run is instant because the data is pulled from memory.
import streamlit as st
import pandas as pd
import time
@st.cache_data(ttl=600, max_entries=5)
def load_data():
start = time.time()
df = pd.read_csv("/data/huge_dataset.csv")
st.success(f"Loaded in {time.time() - start:.2f}s")
return df
st.title("Large CSV Demo")
# Button forces a rerun; without the button the app would rerun on any widget change
if st.button("Reload data"):
df = load_data()
st.line_chart(df["value_column"]) # simple visual
How to verify the caching works:
- Run the app. Observe the “Loaded in Xs” message.
- Click the button again. The message should disappear and the chart appears instantly.
- Open the Streamlit sidebar, click the Cache tab, and confirm that the function appears with a memory footprint count.
Trade‑offs and Pitfalls
- Memory Usage: Caching large dataframes can consume gigabytes of RAM. Monitor the Cache panel and set
max_entriesorttlto bound growth. - Stale Data: Without a TTL, cached results may stay forever, showing outdated information. Always set a reasonable
ttlfor data that changes. - Mutable Arguments: Lists, dicts, or pandas objects are not hashable by default. Pass immutable types or convert them (e.g.,
tuple(my_list)) to avoid cache misses. - Non‑Deterministic Functions: Functions that rely on external state (like random numbers) should not be cached, or you must seed them deterministically.
- App Restarts: The in‑memory cache is cleared when the Streamlit process restarts. For production, consider using
@st.cache_resourcefor long‑lived objects or external persistence.
Actionable Checklist for Production
- Identify expensive functions (IO, heavy calculations).
- Wrap them with
@st.cache_dataand setttlto match data freshness needs. - Set
max_entriesto prevent runaway memory. - Test cache hits by measuring execution time with
time.time()or the built‑inst.cache_datadiagnostics. - Monitor the Cache panel during load tests; adjust parameters if memory spikes.
- For critical data, add a fallback to re‑load from disk or a database if the cache is missing.
By following these steps, you can turn a slow, re‑computing Streamlit app into a responsive, efficient dashboard that scales with user demand.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.