Yunohost App Sandbox Architecture Using systemd-nspawn
An architecture note on how Yunohost uses systemd-nspawn, overlayfs, and cgroups to create secure, resource-limited app sandboxes with minimal host privileges.
24 Oct 2025, 23:41 UTC

Requirements
Yunohost must run multiple independent web applications on a single host while ensuring that a compromised or misbehaving application cannot affect the host or other tenants. To achieve this, the isolation mechanism must satisfy these specific requirements:
- Multi-tenant security: Each application must be restricted to its own filesystem view and process space.
- Enforceable resource limits: CPU, memory, and I/O must be cappable per application to prevent noisy-neighbor effects.
- Minimal privilege execution: Applications must run with the smallest possible set of Linux capabilities.
- Admin integration: Lifecycle management (start/stop) and observability (logs) must integrate with the Yunohost admin panel.
- Stack agnostic: The sandbox must support various runtimes (PHP, Node.js, Python) without requiring host-wide dependency installations.
Minimal Suitable Design
The architecture implements a per-app sandbox using systemd-nspawn, a lightweight containerization tool built into systemd. This design avoids the overhead of a full virtual machine while providing stronger isolation than a simple chroot.
- Read-only base image: The application source code is mounted as read-only using
--bind-ro, preventing the app from modifying its own binaries or core logic. - Writable overlay: An
overlayfslayer (via--overlay) provides a persistent directory for user uploads, databases, and configuration files. - Stripped capability set: All kernel capabilities are dropped via
--capability-drop=all. Only essential capabilities (e.g.,CAP_NET_BIND_SERVICE) are added back. - Private network namespace: The
--private-networkflag isolates the network stack, preventing an app from sniffing traffic on other containers' interfaces. - UID/GID mapping: User namespaces are used to map the container's root user to an unprivileged user on the host, ensuring that a container escape does not grant root access to the host.
To verify the isolation settings of a running app (e.g., wallabag), run the following command as root on the host:
systemctl cat wallabag.serviceCheck the ExecStart line for the following flags: --private-network, --capability-drop=all, --bind-ro, and --uid-map. If these are missing, the application is not running in the intended hardened sandbox.
Trust and Data Boundaries
Isolation is enforced across three primary boundaries:
Filesystem and Device Boundaries
The container is restricted to its own root filesystem. Host directories such as /dev, /sys, and the data directories of other applications are invisible. The --private-devices setting prevents direct access to host hardware.
System Call Filtering
A seccomp (secure computing mode) filter is applied to block dangerous system calls—such as keyctl or open_by_handle_at—that could be used to bypass filesystem restrictions or attack the kernel.
Secret Handling
Sensitive data, such as LDAP passwords, are passed into the container via environment variables using --setenv. These secrets exist only within the container's process space and are not written to the persistent disk image.
To verify that the overlayfs is active and providing the necessary write layer, run:
mount | grep wallabagThe expected output should show an overlay type mount with lowerdir (read-only source) and upperdir (writable data) paths.
Operational Checks
Yunohost leverages standard Linux tooling for monitoring the sandbox health:
- Service State:
systemctl is-active wallabag.serviceconfirms if the container is running. - Resource Accounting: Memory and CPU usage are tracked via cgroups. You can check current memory usage with:
systemctl show wallabag.service -p MemoryCurrent - Log Aggregation: All
stdoutandstderrstreams from the container are captured byjournald. Access them viajournalctl -u wallabag.service.
Failure Modes
The following scenarios represent the primary risks to this architecture:
- Container Escape: If a capability is accidentally retained (e.g.,
CAP_SYS_ADMIN), a process inside the container may exploit kernel interfaces to break out of the namespace. - I/O Exhaustion: While CPU and memory are capped, heavy write workloads on
overlayfscan cause high I/O wait times, potentially slowing down the entire host. - Kernel Incompatibility: The design relies on user namespaces and seccomp. On older kernels that lack these features,
systemd-nspawnwill fail to launch the container.
Conditions for Design Change
The current systemd-nspawn approach would be replaced if the following requirements emerge:
- Hardware Acceleration: If apps require direct GPU or specialized device access (e.g., for transcoding), the architecture would move toward device whitelisting or lightweight VMs (KVM).
- Standardized Image Formats: If the need for OCI-compliant images (Docker/Podman) outweighs the simplicity of the current overlay setup, the runtime would shift to a Podman-based backend.
- Extreme I/O Demands: If
overlayfslatency becomes a bottleneck for high-traffic apps, Yunohost would transition to dedicated block devices or Btrfs subvolumes per app.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.