Using HDFS Erasure Coding to Cut Storage Costs in Hadoop 3.x
Learn how to enable HDFS erasure coding in Hadoop 3.x to cut storage costs for cold data, with a step‑by‑step example and trade‑offs.
30 Oct 2025, 05:10 UTC

Problem: Storage overhead from three‑way replication
In a default HDFS setup each block is stored three times, so a 1 TB dataset occupies roughly 3 TB of disk space. For archives, logs or other cold data this overhead can be a significant cost factor.
Thesis: Enable HDFS erasure coding for cold files to cut storage to ~1.5× while accepting higher read latency and CPU use
How erasure coding works in HDFS
HDFS 3.x can replace replication with an erasure‑coding (EC) policy such as RS‑6‑3. The original file is split into six data blocks and three parity blocks. Any three of the nine blocks are sufficient to reconstruct the original data, so the total storage is (6+3)/6 = 1.5× the original size.
EC policies are configured via dfs.namenode.erasure-code.enable and a set of policy definitions. The NameNode must be restarted after changing the flag.
When to enable erasure coding
- Data that is rarely read (e.g., historical logs, backups) – the extra CPU for decoding is amortized over infrequent accesses.
- Workloads that do not rely on HDFS append or concurrent writers (e.g., not HBase tables, not Flume sinks that append).
- Clusters running Hadoop 3.0+ with a supported codec (Intel ISA‑L, built‑in Java codec, or a native library).
Worked example: creating and verifying an RS‑6‑3 policy
- Enable the feature (requires NameNode restart). Add to
hdfs-site.xml:
Then restart the NameNode:<property> <name>dfs.namenode.erasure-code.enable</name> <value>true</value> </property>sudo -u hdfs hdfs --daemon stop namenodefollowed bysudo -u hdfs hdfs --daemon start namenode. - Create an EC policy (run as HDFS superuser):
Thehdfs ec -addPolicy -p RS-6-3 -rsLegacy-rsLegacyflag selects the Reed‑Solomon codec compatible with ISA‑L. - Apply the policy to a directory (example:
/data/cold):
You need write permission on the directory; typically the hdfs user or a user withhdfs ec -setPolicy -p RS-6-3 /data/colddfs.permissions.enabledfalse or appropriate ACL. - Verify the policy is active:
Look forhdfs ec -listPoliciesRS-6-3in the output. Then check a file placed under the directory:
In the output each block should showhdfs fsck /data/cold -files -blocks -locationsECinstead ofREPLICATED(e.g.,BP-... blk_1073741824_1234 len=134217728 repl=1 [EC]).
Trade‑offs and limitations
- Read latency: Reading an EC block requires decoding parity, which adds CPU overhead and can increase latency compared to a local replica read.
- CPU usage: Decoding is more intensive; monitor NameNode and DataNode CPU when many EC reads occur.
- Incompatible workloads: Append‑only workloads (HBase, Flume, Kafka Connect HDFS sink with append) may fail or suffer severe performance drops because EC does not support concurrent writers.
- Recovery cost: Reconstructing a missing block involves reading all parity and data blocks and computing the missing piece, which is slower than simple replica copy.
- Codec requirement: Without a fast native codec (e.g., ISA‑L) the software fallback can be prohibitively slow.
Actionable closing
If you have cold data that is rarely read, enable erasure coding after confirming your cluster runs Hadoop 3.x and has a supported codec. Start with a test directory, verify the EC block type with hdfs fsck, and monitor read latency and CPU for a week before rolling the policy out to larger datasets. Keep the traditional three‑way replication for hot or append‑only workloads to avoid latency penalties.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.