MATLAB Tall Arrays: Processing Data Larger Than RAM Without Rewriting Your Code
MATLAB tall arrays defer execution until gather(), letting you run familiar table operations on data larger than RAM. See a grouped-statistics example on 50 GB of sensor logs, plus the memory, debugging, and performance trade-offs before adopting.
11 Sept 2026, 06:49 UTC

The Problem: 50 GB of Sensor Logs, 16 GB of RAM
You have a folder of CSV files from a month of sensor readings — 50 GB total. Your laptop has 16 GB of RAM. Loading everything with readtable crashes the session. The classic workaround is manual chunking: write a loop that reads a file, processes it, saves partial results, and stitches them together. That works, but it turns a one-line analysis into a fragile script you have to maintain.
MATLAB's tall arrays offer a different path. They let you write the same table operations you already know — filtering, grouping, statistics — on data that never fully enters memory. The catch: you must understand when the computation actually runs and what fits in RAM at the end.
How Tall Arrays Defer Work
A tall array is a lazy wrapper around a datastore — an object that knows how to read pieces of your data from disk (or a database, or HDFS). When you call tall(ds), MATLAB builds an execution plan, not a result. Every subsequent operation (mean, join, rowfun) adds a step to that plan. Nothing executes until you call gather() or write output with write().
This means errors surface at gather() time, not where you wrote the operation. A typo in a variable name inside rowfun won't appear until the final materialization, which can make debugging feel disconnected from the source.
Worked Example: Grouped Statistics Without Loading Everything
Suppose each CSV has columns device_id, timestamp, temperature, vibration. You want the mean temperature and vibration per device.
% 1. Point a datastore at the folder
folder = 'sensor_logs';
ds = datastore(folder, 'FileExtensions', '.csv', 'ReadSize', 'file');
% 2. Create the tall table
t = tall(ds);
% 3. Express the analysis — no data moves yet
stats = groupsummary(t, 'device_id', 'mean', {'temperature', 'vibration'});
% 4. Materialize only the small aggregated result
result = gather(stats); % result fits in memory: one row per device
The key insight: result is a regular in-memory table with one row per device. The 50 GB never loaded at once. ReadSize controls chunk granularity; 'file' processes one CSV per worker, which is often a good starting point.
Where Tall Arrays Shine — and Where They Don't
- Good fit: Reductions (grouped stats, histograms, tallies), filtering that shrinks the dataset, feature extraction that outputs fewer rows than input.
- Poor fit: Row-by-row algorithms needing random access (e.g., iterative solvers, custom recursions), operations that expand data (certain joins), or any step where the final
gather()would still exceed RAM.
Unsupported functions error at gather(). Check the "Extended Capabilities" section of each function's documentation for your MATLAB release — coverage grows each version but is not universal.
Trade-offs Worth Knowing
- Memory safety:
gather()on an unreduced tall array pulls the full dataset into RAM and can crash the session. Always aggregate or filter first. - Performance: For pure reductions, tall arrays often match or beat manual chunking because MATLAB optimizes the execution plan. For complex row-wise logic, explicit
transformedDatastorepipelines ormemmapfilecan be faster and more predictable. - Scaling: The same code runs on a laptop, a multi-core workstation (with Parallel Computing Toolbox), or a Spark cluster — just change the datastore or enable a parallel pool.
Quick Verification Before You Commit
- Generate 3–5 small CSV files (a few MB each) with the same schema.
- Run the pattern above and confirm
gather()returns the expected grouped means. - Open the MATLAB Profiler (
profile on; ... profile viewer) and compare peak memory and runtime against a naivereadtable+groupsummaryon a subset that fits in RAM. - Verify each function used (e.g.,
groupsummary,mean) lists tall-array support in the documentation for your release.
If the tall-array version is simpler and no slower, you've bought yourself the ability to scale without rewriting. If it's slower or hits unsupported functions, a transformedDatastore with a custom read function is often the lighter-weight alternative.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.