Using Nomad's spread stanza to balance workloads across failure domains
Learn how Nomad’s spread stanza balances allocations across zones, racks, or custom attributes to improve service resilience, with a concrete three‑node example and verification steps.
07 Jul 2026, 21:22 UTC

Problem: allocations cluster on the same nodes
When you run a service in Nomad, you might notice that several allocations of the same task group end up on nodes that share the same rack or availability zone. If that zone fails, the whole service can go down, even though you have other healthy nodes in the cluster.
Thesis: the spread stanza lets Nomad distribute allocations by node attributes
Nomad’s spread stanza, available since version 0.9.0, tells the scheduler to prefer nodes with the lowest current count for a given attribute (e.g., zone, rack, or a custom metadata key). By referencing one or more attributes, you can force Nomad to balance workloads across failure domains.
Worked example: three‑zone cluster
Assume a small Nomad cluster with three nodes, each labeled with a different zone:
- node1:
zone=us-east-1a - node2:
zone=us-east-1b - node3:
zone=us-east-1c
We want a job that runs three allocations of a simple web service, ideally one per zone.
# example.hcl
job \"web\" {
datacenters = [\"dc1\"]
group \"app\" {
count = 3
spread {
attribute = \"zone\"
}
task \"server\" {
driver = \"docker\"
config {
image = \"nginx:latest\"
ports = [\"http\"]
}
}
}
}
To run the job, you need a Nomad token with job:submit and job:read capabilities.
- Check Nomad version:
nomad version(must be ≥0.9.0). - Submit the job:
nomad job run example.hcl. - Watch allocations:
nomad alloc status -json.
After a few placement cycles, you can verify the spread by counting allocations per zone:
nomad alloc status -json | jq -r '.[].NodeAttributes.zone' | sort | uniq -c
You should see roughly 1 allocation for each zone value. If a zone label is missing on any node, Nomad falls back to its default scoring and may place multiple allocations on the same node—this is the key limitation.
Trade‑offs and limitations
- Node diversity required: spread only works when there are enough distinct attribute values. With fewer nodes than spread values, some allocations will share a node.
- Scheduling overhead: each spread attribute adds a scoring step. Using many attributes or high‑cardinality metadata (e.g., per‑host UUID) can increase evaluation time in large clusters.
- Fallback behavior: when a referenced attribute value is absent, Nomad ignores that spread constraint for the offending allocation, which can lead to uneven distribution.
Actionable closing
Start by identifying the failure domain you want to protect against (rack, zone, hardware generation). Add a minimal spread stanza to your job, run it, and verify the allocation counts with the jq command above. If you see imbalance, check that each attribute value exists on at least as many nodes as the job’s count. Adjust the spread attributes or add more nodes to the domain as needed.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.