Designing Streamlit Apps with st.cache_data: Requirements, Boundaries, and Operational Checks
Learn when and how to use Streamlit's st.cache_data decorator, including purity and hashability requirements, memory boundaries, operational checks, failure modes, and triggers for redesign.
03 Dec 2025, 08:06 UTC

Requirements for @st.cache_data
The st.cache_data decorator memoizes the return value of a pure function. For it to work, two conditions must be satisfied:
- Purity – the function must return the same output every time it receives the same input arguments, without relying on external state that can change between calls.
- Hashable arguments – each argument must be hashable so Streamlit can compute a deterministic cache key. Built‑in immutable types (int, float, str, tuple of immutables) are hashable; lists, dicts, sets, and custom objects without a
__hash__method are not.
If either condition fails, Streamlit raises a ValueError at runtime.
Smallest Suitable Design
The minimal implementation is to place the decorator directly above the function definition:
import streamlit as st
import pandas as pd
@st.cache_data
def load_csv(path: str) -> pd.DataFrame:
"""Read a CSV file and return a DataFrame."""
return pd.read_csv(path)
# In the app
st.write(load_csv("data/sales.csv"))
On the first call with a given path, Streamlit executes the function, stores the returned DataFrame in an in‑memory dictionary keyed by the argument hash, and returns the stored object on subsequent calls with the same path. No additional boilerplate is required.
Trust and Data Boundaries
The cache lives in the memory of the individual Streamlit worker process that serves a user session. By default:
- Each worker maintains its own cache; data is not automatically shared across workers.
- If the same worker handles multiple sessions (e.g., in a multi‑user deployment), the cached return value is accessible to those sessions unless the function is marked with
allow_output_mutation=False(the default) and the return value is immutable. - Sensitive information should not be cached unless the app enforces access controls at the route or session level, because a cached object could be read by another user sharing the same worker.
When allow_output_mutation=True is used, the caller may mutate the returned object; Streamlit then assumes the caller will not rely on the cached value being unchanged across calls.
Operational Checks
Developers can observe and manage cache behaviour through a few built‑in mechanisms:
- Cache inspection – call
st.cache_data.clear()to evict all entries; after clearing, the next invocation with the same arguments will re‑execute the function. - Spinner control – set
show_spinner=Falsein the decorator to suppress the default "Running..." UI while the function executes. - Runtime warnings – Streamlit emits a warning if an unhashable argument is encountered, helping developers catch misuse early.
- Hit/miss visibility – although Streamlit does not expose hit/miss counters by default, community plugins or custom logging around the cached function can be added to track usage.
Failure Modes
Several scenarios can cause the cache to behave unexpectedly or exhaust resources:
- Unhashable arguments – passing a list, dict, or unhashable custom object raises
ValueError: Unhashable type: .... - Memory pressure – each cached return value occupies memory proportional to its size. Repeatedly caching large objects (e.g., multi‑megabyte DataFrames) can lead to OOM kills of the worker.
- Hidden side effects – if the cached function mutates external state (writes a file, updates a database), those side effects occur only on the first execution; later calls skip them, potentially causing inconsistent application behaviour.
- Stale data across deployments – Streamlit clears the in‑memory cache when the main script changes, but long‑running workers may retain old values until a restart, leading to temporary mismatches after a code update.
When the Design Should Change
The basic st.cache_data approach is appropriate for many prototypes and internal tools, but certain requirements necessitate a different caching strategy:
- Cross‑worker persistence – if cached data must survive worker restarts or be shared across multiple Streamlit instances, replace the in‑memory store with an external backend (e.g., Redis, Memcached) and implement a wrapper that reads/writes to that store.
- Time‑based eviction – when entries should expire after a fixed interval, add a TTL layer (e.g., using
functools.lru_cachewith a custom timestamp check or a RedisEXPIREcommand). - Safe caching of mutable objects – to cache objects that may be mutated after retrieval, serialize them (e.g.,
pickle.dumps) before storing and deserialize on retrieval, or useallow_output_mutation=Falseand return immutable equivalents (e.g., tuples, frozensets). - Integration with validation pipelines – if the function’s purity is broken by validation steps that depend on external state, consider moving validation outside the cached core and caching only the pure data‑loading portion.
Verification Steps (No Claim of Prior Testing)
To confirm that the cache behaves as described, you can perform the following checks in a local Streamlit environment:
- Timestamp test – create a cached function that returns the current time via
time.time()(note: this is impure, so the timestamp will only change when the arguments differ). Observe that repeated calls with identical arguments return the same timestamp. - Clear cache – after a few calls, invoke
st.cache_data.clear()and call the function again; the timestamp should update, indicating a cache miss. - Memory observation – run Docker (
docker stats <container>) or a system monitor while repeatedly invoking a cached function that returns a large NumPy array; memory usage should rise on the first call and then plateau. - Unhashable argument – pass a list as an argument to a cached function; Streamlit should raise a
ValueErrorabout unhashable types. - Deploy‑time check – if using Streamlit Community Cloud, enable any available cache‑stats plugin and verify that hit/miss counters increase as expected during normal interaction.
These steps help you validate assumptions without asserting that they have been executed in a specific environment.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.