Short answer
Yes — the most likely cause is that your configured replication factor (3) exceeds the number of live DataNodes (2). When a write cannot be placed on the required number of replicas, the client fails with an error like could only be replicated to N nodes instead of minReplication. The heartbeat settings are, at most, a secondary factor: they control how fast the NameNode declares the dead node, not whether writes succeed.
Separate the likely cause from confirmed facts
One nuance matters here: in HDFS, a write does not strictly require the full dfs.replication count to succeed. It requires dfs.namenode.replication.min (default 1). So a complete write block with 2 live nodes usually means one of:
- The minimum replication was raised (e.g., set to 3), which cannot be satisfied with 2 nodes.
- Fewer nodes are actually live than you think — the second node may be dead, decommissioning, or the NameNode may be in safe mode.
Confirm before changing anything:
hdfs dfsadmin -report
hdfs dfsadmin -safemode get
hdfs dfs -put small-test-file /tmp/rf-check
Compare the Live datanodes count against dfs.replication and dfs.namenode.replication.min in hdfs-site.xml. If the test write's error mentions insufficient replicas, the hypothesis is confirmed. Also check the NameNode UI or logs for UnderReplicatedBlocks.
Fixing it
In order of preference:
- Restore capacity. Bring the failed node back, or decommission it cleanly (
dfs.hosts.exclude plus hdfs dfsadmin -refreshNodes) and add a replacement. Avoid simply yanking the node — abrupt removal triggers a large re-replication storm on the survivors.
- Temporarily lower the replication factor so writes proceed while you repair:
# for new files
hdfs dfs -setrep -w 2 /path/to/data
# and set dfs.replication=2 in hdfs-site.xml, then restart or refresh
This reduces durability — one more node loss means data loss — so treat it as a stopgap and restore 3 once a third node is healthy.
- Verify recovery. After capacity returns, run
hdfs fsck / or just watch the Under replicated blocks metric; the NameNode re-replicates automatically until it reaches zero.
About heartbeat settings
dfs.heartbeat.interval and the dead-node detection timeout only affect how quickly the NameNode stops considering the failed node a write candidate. A slow detection can cause short-lived write retries right after the failure, but it cannot cause a sustained block once the node is marked dead. Tune it only if you see long stalls immediately after failures, not as the fix for this problem.
The same pattern elsewhere
This failure mode generalizes: Kafka producers fail when min.insync.replicas exceeds the in-sync replica count, and Cassandra writes fail when the requested consistency level exceeds available replicas. In every case the rule is the same — the durability guarantee you demand must not exceed the replicas you actually have, and the correct fix is restoring capacity first, relaxing the guarantee only temporarily.
Assumes a recent Hadoop 3.x release; property names and defaults can differ in older versions or vendor distributions, so verify against your deployed configuration.