Cut Pandas Memory with Categorical Columns: When Low‑Cardinality Strings Pay Off
Learn how converting low‑cardinality string columns to pandas categorical dtype can cut memory usage by up to 90 % and boost performance. Follow a quick checklist to decide when the conversion is worth it.
04 Apr 2026, 11:01 UTC

The Problem: Large String Columns Drain RAM
In many analytics pipelines a DataFrame ends up with columns that hold a handful of repeating text values—status codes, department names, product categories, etc. Those columns are stored as object dtype, which means each string is an independent Python object. Even if the strings are short, the per‑row overhead is heavy: the pointer to the string, the string object itself, and the hash table entry. On a million‑row table, a single low‑cardinality column can consume 30–50 MB of RAM.
How Categorical Dtype Helps
When you cast a column to category, Pandas replaces the per‑row string with a 2‑byte integer code that indexes a categories table holding the unique values. Memory usage becomes:
- Integer codes:
N * 2 bytes(or 4 bytes if the cardinality > 65535) - Categories table:
cardinality * avg_string_len+ small overhead
For a column with 10 unique values, the code array is 2 bytes × 1 000 000 = 2 MB, while the categories table is only a few kilobytes. That’s a 90 % reduction** in memory compared to a 45 MB object column.
Practical Example: A 1‑Million Row Department Table
import pandas as pd
# 1 M rows, 10 unique department names
dept_names = ["HR", "Finance", "Engineering", "Sales", "Marketing", "Support", "Legal", "Ops", "R&D", "IT"]
rows = 1_000_000
df = pd.DataFrame({
"employee_id": range(rows),
"department": [dept_names[i % len(dept_names)] for i in range(rows)],
})
# Measure memory before conversion
mem_before = df.memory_usage(deep=True).sum() / (1024 ** 2)
print(f"Memory before: {mem_before:.2f} MB")
# Convert to categorical
df["department"] = df["department"].astype("category")
mem_after = df.memory_usage(deep=True).sum() / (1024 ** 2)
print(f"Memory after : {mem_after:.2f} MB")
Typical output (on a recent Pandas 2.1 release):
- Memory before: ~45 MB
- Memory after : ~5 MB
Notice the ~90 % cut in RAM. The same conversion also speeds up groupby and value_counts because the operations run on the integer codes.
When Categorical Is NOT a Win
Not every string column is a candidate:
- High cardinality – If a column has nearly as many unique values as rows (e.g., user IDs or UUIDs), the categories table grows almost linearly and can exceed the original object size.
- Some
straccessor methods (.str.contains(),.str.split()) returnobjectdtype. If you keep the column categorical, you’ll need to cast back tostror use the.cataccessor for category‑specific operations. - Serializing to CSV or JSON strips categorical information. If you need to preserve the dtype across I/O, consider
to_parquetorfeatherwhich support categories.
To decide, run a quick cardinality check:
# cardinality threshold heuristic
card = df["department"].nunique()
if card < 1000:
print("Categorical likely beneficial")
else:
print("Consider keeping as object")
Takeaway: Quick Checklist
- Measure current memory:
df.memory_usage(deep=True).sum() - Check cardinality:
df[col].nunique() - Convert if
cardinality < 1000and you’re not using heavy string methods. - Validate performance: run a
groupbyormergeand compare timings. - When writing to disk, use a format that preserves categories or re‑convert after reading.
By applying this simple cast to low‑cardinality string columns, you can shave hundreds of megabytes off your DataFrame’s footprint and often gain speed for downstream analytics. It’s a low‑risk, high‑reward tweak that fits into any Pandas workflow.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.