Speeding Up Line Processing in Bash with mapfile (readarray)
Bash’s mapfile builtin reads a file into an array in one go, cutting runtime and simplifying code compared to the classic while‑read loop. This guide shows how to use mapfile, benchmark the speedup, and when it’s appropriate.
01 Aug 2025, 14:21 UTC

Problem: The Classic While‑Read Loop Is Slow
When you need to read a text file line by line in Bash, the idiomatic pattern is a while IFS= read -r line; do …; done < file loop. Each iteration spawns a subshell, and the shell has to perform word‑splitting, globbing, and variable assignment on every line. If the loop body is minimal or you only need the data later, this overhead becomes noticeable, especially with large files.
Thesis: Use mapfile (or readarray) for a Cleaner, Faster Solution
Bash 4.0 introduced the mapfile builtin, which reads an entire file straight into an indexed array in a single system call. The result is a concise one‑liner, no subshells, and a measurable speed improvement for many use‑cases. This article shows how to replace a while‑read loop with mapfile, benchmarks the difference, and explains the trade‑offs.
What mapfile Actually Does
- Reads the file descriptor line by line and stores each line as an array element.
- When the
-toption is used, trailing newlines are stripped from each element. - Does not perform IFS‑based word splitting unless
-tis omitted. - Requires Bash 4.0 or newer; older releases lack this builtin.
Performance Comparison
Below is a quick benchmark that compares the classic loop with mapfile on a 1‑million‑line file. The timing is illustrative; your mileage may vary.
# Create a test file
seq 1 1000000 > test.txt
# Classic loop
/usr/bin/time -p bash -c 'while IFS= read -r line; do :; done < test.txt'
# mapfile
/usr/bin/time -p bash -c 'mapfile -t arr < test.txt'
Typical output on a modern laptop:
real 0.35
user 0.28
sys 0.07
real 0.08
user 0.07
sys 0.01
Notice the real time drops from ~0.35 s to ~0.08 s, a ~75 % improvement. The difference grows with file size.
Practical Usage Patterns
1. Loading a File into an Array
#!/usr/bin/env bash
# Ensure we run with Bash 4.0+
if (( ${BASH_VERSINFO[0]} < 4 )); then
echo 'This script requires Bash 4.0 or newer' >&2
exit 1
fi
# Load file
mapfile -t lines < "$1"
# Iterate with line numbers
for i in "${!lines[@]}"; do
printf '%d: %s\n' $((i+1)) "${lines[i]}"
done
Run with ./script.sh data.txt. The -t option ensures each element is a clean line without a trailing newline.
2. Processing Only a Subset of Lines
# Read first 10 lines
mapfile -t first10 -s 0 -n 10 < data.txt
for line in "${first10[@]}"; do
echo "$line"
done
The -s (skip) and -n (max lines) options give fine‑grained control without a loop.
3. Avoiding IFS Side‑Effects
Because mapfile respects the current IFS unless you use -t, it’s safe to keep IFS at its default when you don’t need splitting. If you set IFS to a custom value elsewhere, remember to use -t or reset IFS before calling mapfile.
Limitations & Trade‑offs
- Memory Footprint:
mapfileloads the entire file into RAM. For gigabyte‑scale files or systems with limited memory, this can cause swapping or OOM kills. In such cases, fall back to a streamingwhile‑readloop. - NUL Bytes:
mapfiletreats NUL bytes as regular characters. Many downstream Unix tools (e.g.,grep,sed) cannot handle NUL, so if your data may contain them, you must sanitize or avoidmapfile. - Compatibility: Bash 3.x and earlier lack
mapfile. On those systems, use the traditional loop or upgrade Bash. - Line Splitting: Omitting
-tcan cause unexpected word splitting ifIFShas been altered. Always use-twhen you need literal lines.
Verification Checklist
- Confirm Bash version:
bash --versionshould show 4.0 or newer. - Verify
mapfileis builtin:help mapfileshould display usage. - Create a small test file:
seq 1 1000 > test.txt. - Run both approaches and compare
timeoutput. - Inspect the array:
declare -p linesshould list each line as an element. - Check that
${lines[0]}does not end with a newline when-tis used.
Actionable Takeaway
If your script needs to read a file that comfortably fits in memory and you want cleaner, faster code, replace the classic while‑read loop with mapfile -t. Always benchmark on your target data set, and guard against large memory usage or NUL bytes. For legacy Bash environments or streaming requirements, keep the traditional loop.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.