Nomad health check flapping and event stream noise
26K reputation · 02 Oct 2022, 02:14 UTC
HashiCorp Nomad utilizes check stanzas to monitor task health via HTTP, TCP, or scripts. When integrated with external monitoring tools via the Event Stream API or Prometheus endpoints, these checks provide visibility into task stability.
A challenge arises when tasks experience "flapping," where a health check rapidly oscillates between healthy and unhealthy states. Because Nomad triggers events for these state transitions, a flapping task can generate a high volume of noise in the event stream, potentially overwhelming external alerting systems before the restart policy can stabilize the task.
Given the current behavior of the scheduler and the Event Stream API, what mechanisms exist to dampen these notifications? Is there a native way to implement a cooldown period for health check state changes to prevent alert storms?