Glitch Free‑Tier Sleep Mechanism: An Architecture Note (Historical)
An architectural overview of Glitch’s free‑tier sleep/wake mechanism, covering requirements, minimal design, trust boundaries, operational checks, failure modes, and factors that would alter the design. Note: the service was discontinued in July 2025.
13 Oct 2025, 20:59 UTC

Overview
Glitch offered a free‑tier hosting model where projects ran in isolated Linux containers. To conserve platform resources, containers were automatically put to sleep after a period of inactivity and instantly revived on the next HTTP request. This note examines the sleeping mechanism as an architectural case study, highlighting the requirements, minimal design, trust/data boundaries, operational checks, failure modes, and conditions that would have altered the design. The description reflects the system while the service was operational (up to July 2025). As of the supplied date (2026‑10‑08) Glitch hosting has been discontinued, so the mechanism cannot be exercised today.
Requirements
- Resource efficiency: idle projects must consume minimal CPU, memory, and network bandwidth.
- Developer transparency: waking a project should require no code changes and preserve the file system and environment variables.
- Fast wake‑up: latency from sleep to serving a request should be low enough for typical web workloads.
- State safety: any in‑memory state may be lost; persistent state must be stored outside the container.
- Isolation: each project must remain isolated from others, even during sleep/wake cycles.
Smallest Suitable Design
The design that satisfies the above requirements with the fewest moving parts consists of three core components:
- Idle timer – a platform‑level watchdog that measures the time since the last incoming HTTP request; when the threshold (5 minutes) is exceeded, it signals the container runtime to pause.
- Snapshot‑and‑pause – the container’s memory state is discarded, while its root filesystem is snapshotted to persistent storage. The container process is then stopped, freeing CPU and memory.
- Wake‑on‑request – the platform’s edge router detects an incoming HTTP request for a sleeping project, provisions a fresh container from the latest snapshot, re‑attaches the persisted disk, and forwards the request. The container boots, runs the user’s start script, and begins serving traffic.
No additional components (e.g., separate scheduler, external state store) are required for the basic sleep/wake cycle.
Trust and Data Boundaries
- Project isolation – each free project runs in its own container with a unique filesystem snapshot; the platform enforces Linux namespace separation, preventing cross‑project file or process access.
- Persistent data boundary** – the snapshotted disk is the only trusted persistent layer. Environment variables are stored in the platform’s metadata service and are reinjected on wake‑up.
- Control plane trust** – the idle timer, snapshot mechanism, and wake‑on‑request logic reside in the Glitch control plane, which is trusted to manage container lifecycle safely.
Operational Checks
Developers could verify the sleep/wake behavior using the following steps (while the service was active):
- Create a free Glitch project with a simple Express server that logs the request timestamp and a global in‑memory counter.
- Leave the project idle for > 6 minutes to ensure the idle timer fires.
- Send an HTTP request (via browser or
curl) and measure the response time; this is the wake‑up latency. - Send a second request a few seconds later; the response time should reflect the warm container (typically < 50 ms).
- Inspect the project logs: the first request after sleep shows a fresh startup log and the counter reset to zero, confirming loss of in‑memory state.
Operational metrics exposed by the platform included container start‑up events, sleep timestamps, and request latency histograms, which could be viewed in the project dashboard.
Failure Modes
- Snapshot corruption – if the persistent disk snapshot cannot be read, wake‑up fails and the project returns an error (typically 502). Recovery requires manual re‑creation of the project from the latest backup.
- Routing mis‑direction – the edge router might fail to map the hostname to a sleeping project, causing the request to be dropped or sent to the wrong container.
- Extended wake latency** – under heavy platform load, provisioning a new container could exceed the typical 200‑800 ms window, leading to user‑perceived delays.
- Background process loss** – any long‑running process (e.g., WebSocket server, cron loop) terminates on sleep and must be restarted in the request handler after wake‑up; failure to do so results in silent loss of functionality.
- Disk attach failure** – if the persisted volume cannot be re‑attached to the new container, the project starts with an empty file system, losing all user data.
Conditions That Would Change the Design
- Paid tiers that allow disabling sleep (as Glitch offered) would remove the idle timer and snapshot‑pause steps, trading resource savings for predictable latency.
- If the target workload required sub‑50 ms first‑request latency (e.g., real‑time games), the platform might replace sleep with a warm‑pool of idle containers or a different scheduling policy.
- Changes in cost model (e.g., cheaper persistent storage) could make it viable to keep containers running but freeze their CPU cycles via cgroups, eliminating the need for full snapshotting.
- A shift from container‑based isolation to lightweight VMs or sandboxed runtimes would affect the snapshot mechanism; the design would need to adopt VM image diffs or memory‑page sharing techniques.
- Regulatory or security requirements mandating encryption of idle state would necessitate encrypting snapshots at rest and managing decryption keys on wake‑up.
Historical Note and Verification Limitations
Glitch announced the termination of its hosting service in May 2025, with the shutdown completed on July 8 2025. Consequently, as of the current date (2026‑10‑08) the free‑tier sleeping mechanism described above is no longer operational, and the verification steps cannot be performed. The architecture remains a useful illustration of how a platform‑managed sleep/wake cycle can be built around container snapshotting, idle timers, and transparent request‑driven revival, but any attempt to apply it today would require a different hosting provider or a self‑hosted implementation.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.