Podman Quadlet and systemd restart policy conflict during healthcheck failures
26.5K reputation · 17 Feb 2022, 21:24 UTC
When using Podman Quadlet to manage container lifecycles, the generator translates .container files into systemd unit files. While resource limits and basic restart directives map to cgroups and systemd directives, the interaction between Podman's internal healthcheck and systemd's restart logic remains ambiguous.
nIf a container fails a defined healthcheck, Podman may attempt a restart based on internal logic. However, if the generated systemd unit also has a Restart= policy defined, it is unclear how the two managers arbitrate a restart triggered by a healthcheck failure versus a process-level exit. This overlap can lead to unpredictable restart loops or suppressed recovery if the KillMode or unit type are not perfectly aligned between the two layers.
- How does systemd prioritize a restart request when a Podman healthcheck fails but the container process remains running? Is there a specific Quadlet-supported configuration to ensure systemd handles healthcheck-driven restarts without conflicting with internal Podman policies?
1 answer
1 question comment
Use comments to ask for clarification. Post a solution as an answer.
26,525 reputation · 18 Feb 2022, 07:22 UTC
To build on the previous point, it is important to clarify that Podman's HealthCmd is purely diagnostic; it changes the container status to unhealthy but does not send a termination signal to the process. Since systemd only reacts to process exits or signals, a container can remain unhealthy indefinitely while systemd reports it as active (running).
If you require a restart specifically triggered by a healthcheck failure, you cannot rely on Restart=on-failure alone. You must implement an external trigger. A common pattern is using a sidecar container or a simple systemd timer that runs a script to check the status:
# Example logic for a monitoring script
if [ "$(podman inspect --format='{{.State.Health.Status}}' container_name)" == "unhealthy" ]; then
systemctl restart container-name.service
fi
When verifying this behavior, check systemctl status to see if the Main PID has changed. If the PID remains the same despite an unhealthy status, systemd is correctly ignoring the application-level failure.