Guide
Diagnosing Nomad Job Pending Issues with the Autoscaler Feature
Learn how to diagnose Nomad jobs stuck in pending when the autoscaler is involved, check logs, policy, node status, and quota, then apply targeted fixes.
Published by Tasadduq Burney
17 May 2026, 13:05 UTC
3 min134K views0

Recognizable Condition
Jobs stay in the pending state longer than expected while Nomad client nodes report sufficient CPU, memory, and disk resources and are not drained.
Cause/Diagnostic Table
| Possible Cause | What to Look For |
|---|---|
| Autoscaler plugin not evaluating | No autoscaler events in server logs; policy not triggered. |
| Misconfigured scaling policy (min/max limits) | Policy file shows min/max values higher than current node count or target metric thresholds too high. |
| Autoscaler plugin disabled | Plugin stanza missing or set to disabled in agent config. |
| Quota limits preventing new allocations | Quota usage shows allocated resources at or near limit for the job’s namespace. |
| Insufficient nominal node count for trigger | Number of ready nodes is below the autoscaler’s min_ready_nodes or evaluation window requires more nodes. |
Ordered Checks
- Check Nomad server logs for autoscaler events:
journalctl -u nomad | grep -i autoscaler(requires read access to system logs or Nomad agent log file). - Verify client node attributes and drain status:
nomad node status -selfornomad node status(Nomad token with node read permission). - Inspect the autoscaler policy file: look for
policy.hclin the plugin directory; confirmtarget_metric,threshold,min,maxvalues. - Confirm Nomad client registration and health:
nomad node status(requires node read token). - Review quota usage:
nomad quota usage -namespace(requires quota read permission).
Fixes Tied to Findings
- If autoscaler logs show no evaluation: ensure the plugin is enabled in the agent config (
plugin "nomad-autoscaler" { config = { ... } }) and restart the Nomad agent (sudo systemctl restart nomad). - If policy thresholds are too high: lower the
thresholdor adjustmin/maxin the policy file, then reload the autoscaler (nomad autoscaler policy apply -policy policy.hcl) which requires a plugin restart or agent reload. - If client nodes are drained: mark them ready with
nomad node eligibility -enable(node write permission). - If quota exceeded: increase quota limits via
nomad quota apply -namespace quota.hclor request an infrastructure increase from the platform team. - If node count insufficient: add healthy client nodes to the cluster or adjust the autoscaler’s
min_ready_nodesto reflect the current capacity.
Escalation Criteria
- The autoscaler plugin fails to start after agent restart (check logs for repeated plugin load errors).
- Node health checks continue to fail despite correct eligibility and resource availability.
- Quota adjustments require changes to underlying infrastructure (e.g., requesting new host machines) that are beyond the operator’s scope.
Verification Steps
- Run
nomad node statusand confirm the expected number of nodes showreadyandavailable. - List active autoscaler policies:
nomad autoscaler policy listand verify the updated policy appears. - Submit a test job with known resource requests (e.g., a simple
rediscontainer) and watchnomad job statustransition frompendingtorunningwithin the expected timeframe. - Enable debug logging temporarily (
nomad agent -log-level=debug) and check for autoscaler evaluation messages matching the policy thresholds.
Limitations
- The Nomad Autoscaler only works with drivers that expose the required task metrics (e.g.,
docker,exec,java). - Metric collection must be enabled (
telemetryblock) and the autoscaler must have access to the metric source (Prometheus, StatsD, etc.). - Changes to the autoscaler plugin or policy require a Nomad agent restart; test modifications in a staging cluster before applying to production.
- Avoid setting min/max values that could cause thrashing; monitor the evaluation window to prevent rapid scale‑up/down cycles.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.