Streamlit Caching: @st.cache_data vs @st.cache_resource
Learn when to use @st.cache_data and @st.cache_resource in Streamlit, see a worked example, and understand the trade‑offs of cached values.
01 Jul 2025, 16:27 UTC

The problem: full script reruns on every interaction
When a user moves a slider, clicks a button, or changes any widget, Streamlit reruns the entire script from top to bottom. If the script loads a large CSV, trains a model, or opens a database connection on each run, the app feels sluggish.
Choosing the right caching decorator
@st.cache_data– for pure functions that return data (e.g., filtering, aggregation). The output depends only on the function's arguments.@st.cache_resource– for heavy objects that should be created once, such as ML models, database clients, or API sessions.
Worked example: CSV filter with a cached model
import streamlit as st
import pandas as pd
# Dummy model placeholder – replace with your actual model loader
def load_model():
# Simulate an expensive load
return {'weights': [0.1, 0.2, 0.3]}
@st.cache_data
def load_data(path: str) -> pd.DataFrame:
return pd.read_csv(path)
@st.cache_data
def filter_data(df: pd.DataFrame, threshold: float) -> pd.DataFrame:
return df[df['value'] > threshold]
@st.cache_resource
def get_model():
return load_model()
def main():
st.title('Cached Streamlit Demo')
# File path – adjust to your environment
data = load_data('sample.csv')
threshold = st.slider('Value threshold', 0.0, 10.0, 5.0)
filtered = filter_data(data, threshold)
model = get_model()
st.write('Rows after filter:', len(filtered))
st.write('Model weights:', model['weights'])
if st.button('Clear cache'):
st.cache_data.clear()
st.cache_resource.clear()
if __name__ == '__main__':
main()
Save the script as app.py and run streamlit run app.py. The first execution loads the CSV and creates the model; subsequent slider moves reuse the cached data and model, so the UI updates instantly.
Trade‑off and practical check
- Cached objects occupy memory for the lifetime of the Streamlit session. Monitor memory with your OS task manager or
psutilto ensure the cache does not exceed available RAM. - If the underlying CSV changes, the cached DataFrame becomes stale. To avoid serving outdated results, either add a version string to the cache key (e.g.,
@st.cache_data(show_spinner=False)with a custom hash) or manually clear the cache when you know the source updated. - You can verify the cache is working by adding a temporary
st.write('Loading data…')insideload_data; the message appears only on the first run and after you click Clear cache.
Start by annotating the slowest function with @st.cache_data, measure the difference with st.time (or a simple time.time() wrapper), then promote expensive objects to @st.cache_resource. Finally, expose a Clear cache button as shown to let users reset when needed.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.