Speeding Up Streamlit Apps with st.cache_data: When and How to Use It
Learn how to use Streamlit’s st.cache_data decorator to avoid re‑loading large datasets on every interaction, with a concrete CSV‑loading example, version notes, and cache‑staleness mitigation.
29 Mar 2026, 04:11 UTC

Problem: UI feels sluggish after every interaction
Streamlit re‑executes the entire script whenever a user clicks a button, moves a slider, or types in a text box. If your app loads a large CSV, queries a database, or runs a heavy computation inside the main script, each interaction can cause a noticeable delay.
Thesis: Use st.cache_data to memoize expensive, pure functions
By wrapping a data‑loading or computation function with st.cache_data, Streamlit stores the return value keyed by the function’s arguments. On subsequent runs with the same arguments, the cached result is returned instantly, avoiding redundant work.
How caching works in Streamlit
When the script runs, Streamlit evaluates the decorated function. If the arguments have not been seen before, the function executes and its result is saved. If the arguments match a previous call, the saved result is returned without re‑entering the function body. The cache lives for the duration of the server session and is shared across all users of that session.
Version note
Starting with Streamlit 1.20, st.cache_data (and its companion st.cache_resource) replaced the experimental st.experimental_memo. If you are on an older version, use st.experimental_memo but plan to upgrade.
Worked example: caching a large CSV read
Suppose you have a 200 MB file data/big_dataset.csv and you want to show a summary table that updates when the user selects a column.
import streamlit as st
import pandas as pd
@st.cache_data(show_spinner=False)
def load_csv(path: str) -> pd.DataFrame:
# This print helps us see when the function actually runs
print(f"Loading {path}...")
return pd.read_csv(path)
# UI
st.title("CSV Explorer")
file_path = st.text_input("CSV path", "data/big_dataset.csv")
if file_path:
df = load_csv(file_path)
column = st.selectbox("Select column", df.columns)
st.write(f"You chose {column}")
st.dataframe(df[[column]].head())
Place the script in app.py and run it with:
streamlit run app.py
When you first load the page, you will see the print statement in the console, indicating the CSV is read. After that, changing the selector or re‑typing the path (as long as the string stays identical) will not trigger the print again, because the cached DataFrame is returned.
Trade‑off: cache staleness
The cached result is valid only as long as the underlying data stays unchanged. If the CSV is overwritten while the app is running, users will still see the old version until the cache is cleared. To avoid stale data:
- Include a version string or file‑modification timestamp as an argument to the cached function.
- Call
st.cache_data.clear()manually (e.g., via a “Refresh data” button) when you know the source has changed. - Use
st.cache_resourcefor objects that should persist across reruns but still need explicit invalidation.
Practical way to verify freshness
Add a mutable argument that reflects the source state, for example the file’s modification time:
import os
import time
@st.cache_data(show_spinner=False)
def load_csv_with_mtime(path: str, mtime: float) -> pd.DataFrame:
print(f"Loading {path} (mtime={mtime})")
return pd.read_csv(path)
# In the UI
if file_path and os.path.exists(file_path):
mtime = os.path.getmtime(file_path)
df = load_csv_with_mtime(file_path, mtime)
# … rest of UI as before
Now, if the file is updated, mtime changes, causing a cache miss and a fresh read.
Actionable closing
- Check your Streamlit version:
pip show streamlitorstreamlit --version. Ensure you are on ≥ 1.20 forst.cache_data. - Identify the expensive, pure function in your script (data load, API call, heavy computation).
- Decorate it with
@st.cache_data. Useshow_spinner=Falseif you want to hide the default spinner. - Add a print or logging statement inside the function to confirm it runs only when expected.
- Run the app, interact with the UI, and observe the console: the print should appear once per unique set of arguments.
- If your data can change, incorporate a versioning argument (file hash, timestamp, or manual version number) or provide a button that calls
st.cache_data.clear().
By caching effectively, you keep your Streamlit app responsive without sacrificing correctness—just remember to invalidate the cache when the underlying data evolves.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.