Error: Failed to connect to database – NetBox migration strategy for zero‑downtime deployment
0 reputation · 10 Sept 2023, 13:49 UTC
0 reputation · 10 Sept 2023, 13:49 UTC
To upgrade a small NetBox instance to a newer release while keeping the application online for clients.
NetBox automatically applies pending Django migrations during startup. This process acquires schema locks that can block read/write traffic, potentially exposing a brief outage. The documented option to suppress automatic migrations is the environment variable DJANGO_MIGRATE, but its exact effect on the web server startup remains unclear. Additionally, the Failed to connect to database error is logged when NetBox cannot establish a connection during this migration phase, which may cause the service to halt entirely. A key decision is whether to rely on the automatic migration mechanism or to perform migrations manually in a blue‑green deployment pattern, and how to guarantee the database remains reachable throughout the switch.
DJANGO_MIGRATE=0 fully disable automatic migrations during NetBox startup, or does it only affect the python manage.py migrate command?To achieve a zero-downtime deployment for NetBox, you must decouple the database schema migration from the application startup process. Relying on automatic migrations during a rolling update often leads to Failed to connect to database errors because the application attempts to boot while the database is locked by a migration process, or multiple pods attempt to run the same migration simultaneously.
Setting DJANGO_MIGRATE=0 (or omitting it, depending on the specific NetBox version's wrapper) is intended to prevent the application container from automatically executing python manage.py migrate upon startup. This does not affect the manual execution of the command via a shell; it specifically instructs the entrypoint script to skip the migration phase before launching the web server (Gunicorn/uWSGI).
When manage.py migrate is executed, Django acquires Access Exclusive locks on tables being modified. For small instances, these locks are typically brief (milliseconds to seconds). However, if a migration involves adding a column with a default value to a large table or altering a constraint, the lock persists until the transaction completes. During this window, any API request attempting to read from or write to that specific table will hang, eventually resulting in a 504 Gateway Timeout or a database connection error if the connection pool is exhausted.
To maintain availability, follow this sequence to ensure the database is ready before the new application version attempts to connect:
python manage.py migrate
DJANGO_MIGRATE=0. This ensures the new pods boot immediately without attempting to re-run migrations or fight for locks.To verify lock contention during a migration, you can monitor active locks in PostgreSQL using:
SELECT pid, locktype, mode, granted FROM pg_locks WHERE granted = false;
Missing Detail: Are you utilizing a container orchestrator (like Kubernetes) or a standard systemd installation? This determines whether you should use a Job/InitContainer or a manual CLI sequence.
Use comments to ask for clarification. Post a solution as an answer.
26,525 reputation · 10 Sept 2023, 17:36 UTC
Setting DJANGO_MIGRATE=0 in the NetBox container entrypoint only skips the automatic python manage.py migrate step that the default startup script runs. It does not prevent you from invoking migrations manually via netbox migrate or python manage.py migrate; those commands will still execute if called explicitly. Therefore, the variable is a convenience flag to decouple migration execution from web‑process start‑up, not a hard disable of Django’s migration framework.
When a migration is run manually on a live NetBox instance, Django acquires an exclusive ACCESS EXCLUSIVE lock on each table being altered for the duration of the ALTER statement. For simple column additions or index creations this lock typically lasts only a few seconds, but for operations that rebuild a table (e.g., changing a column type) the lock can persist for the entire rebuild time, blocking all write operations (INSERT/UPDATE/DELETE) while still allowing read‑only queries. Consequently, ongoing API write requests will fail or be delayed until the lock is released.
For a blue‑green zero‑downtime deployment the recommended sequence is:
DJANGO_MIGRATE=0 so the web pods start without triggering migrations.netbox migrate against the primary database. Verify connectivity and required privileges beforehand.This approach ensures the database remains reachable throughout the switch and limits the window where exclusive locks could affect client traffic.