NixOS switch-to-configuration: Rollback Policy and Journal Access Boundaries
25.5K reputation · 18 Dec 2024, 19:59 UTC
NixOS deploys by building a new system closure with nixos-rebuild switch and then activating it through switch-to-configuration. A failure can surface at build time, during activation, or later at service runtime, and an interrupted activation can leave a partially applied generation rather than an atomic switch.
The unresolved decision is policy: should a failed activation trigger automatic rollback, or should the operator inspect the failing generation first? NixOS exposes nixos-rebuild switch --rollback and boot-time generation selection, but defines no policy for when either is appropriate. Automatic rollback can also obscure the original cause. Option names and rollback semantics are release-sensitive and should be confirmed against the target channel.
A second boundary is log access. Reading the full systemd journal typically requires root or membership in a journal-reading group such as systemd-journal or adm, while build logs retrieved through nix log are a separate source and do not include runtime output. Remote or CI deployments may not retain activation messages where an operator expects them.
Which source should be authoritative for activation failures when terminal output is discarded? Should automatic rollback be gated on a detectable activation failure, and how should a partially activated generation be represented? What journal permissions let a non-root operator diagnose a failed unit without full root access?