Diagnosing Apache Airflow Scheduler Deadlocks that Keep Tasks in Queued State
Identify why Airflow tasks stay queued due to scheduler deadlocks, check logs, DB transactions, and process counts, then apply targeted fixes and know when to escalate.
30 Aug 2025, 02:26 UTC

Recognizable Condition
Tasks remain in the queued state for longer than the scheduler’s loop interval (default 5 seconds). The scheduler logs show no recent heartbeat messages or contain repeated "Database lock timeout" warnings.
Cause/Diagnostic Table
| Possible Cause | What to Look For | Initial Check |
|---|---|---|
| Long‑running DB transaction holding the DagModel lock | Scheduler logs lack activity; DB shows a session in idle in transaction state | Query the DB’s process list for idle‑in‑transaction sessions |
| Too many parallel scheduler jobs (> CPU cores) | High number of SchedulerJob processes; DB lock contention messages | Count scheduler processes and compare to core count |
Misconfigured max_threads or worker_concurrency | Scheduler spawns more threads than DB connection pool allows; lock timeouts | Review airflow.cfg values and DB max_connections |
Ordered Checks
- Verify scheduler heartbeat
Run on the scheduler host:
# Replace with your Airflow base directory cd airflow scheduler --help 2>&1 | grep -i heartbeatIf the command fails or returns no output, the scheduler process may not be running. Ensure you have read access to the
airflowuser’s home directory.Risk: Mis‑typing the path can start a second scheduler instance, worsening contention.
- Inspect scheduler logs for lock warnings
Locate the log file (default
$AIRFLOW_HOME/logs/scheduler/*.log) and grep:grep -i "lock timeout\|deadlock" $AIRFLOW_HOME/logs/scheduler/*.log | head -20Look for timestamps; repeated entries indicate an ongoing issue.
Permission: Read access to the log directory (usually the
airflowuser). - Check for blocking database transactions
Depending on your metadata database, run the appropriate query. For PostgreSQL:
# Connect as a user with pg_stat_activity access (e.g., airflow) psql -U airflow -d airflow -c " SELECT pid, state, query, now() - pg_stat_activity.xact_start AS xact_duration FROM pg_stat_activity WHERE datname = 'airflow' AND state = 'idle in transaction' ORDER BY xact_duration DESC; "For MySQL/MariaDB:
mysql -u airflow -p -e " SELECT id, user, host, db, command, state, info, time FROM information_schema.innodb_lock_waits JOIN information_schema.processlist ON blocking_engine_transaction_id = blocking_trx_id; "If a session shows a long‑running query or is idle in a transaction for more than a few seconds, it may be holding the DagModel lock.
Permission: Database user with access to
pg_stat_activityorinformation_schematables. - Count scheduler processes and compare to CPU cores
On the scheduler host:
ps -ef | grep "[a]irflow scheduler" | wc -lThen check core count:
nprocIf the scheduler count exceeds the core count significantly (e.g., > 2 × cores), you may be over‑provisioning.
Permission: Ability to run
psandnprocas theairflowuser or root. - Review scheduler configuration values
Examine
$AIRFLOW_HOME/airflow.cfg(or environment overrides):grep -E "^scheduler\.max_threads|^worker_concurrency" $AIRFLOW_HOME/airflow.cfgEnsure
scheduler.max_threadsdoes not exceed the DB’smax_connectionssetting.Permission: Read access to the configuration file.
Fixes Tied to Findings
- If a blocking transaction is found:
- Confirm it is safe to terminate (e.g., it is an idle import or a stuck sensor).
- Kill the session:
# PostgreSQL example SELECT pg_terminate_backend(pid) FROM pg_stat_activity WHERE pid = ;Replace
with the PID from the query.After killing, restart the scheduler:
airflow scheduler restartRisk: Killing a transaction that is actively writing can leave the metadata DB in an inconsistent state; only terminate sessions confirmed to be idle or stuck.
- If scheduler process count is too high:
Reduce parallelism by lowering
scheduler.max_threads(or the number of scheduler replicas if you run multiple). Example change inairflow.cfg:[scheduler] max_threads = 2 # set to number of CPU cores or lessThen restart the scheduler.
- If
max_threadsorworker_concurrencyexceed DB limits:Either increase the database’s
max_connections (requires DB admin) or lower the Airflow values so that:scheduler.max_threads + worker_concurrency ≤ DB max_connections – reserved_connectionsApply the change and restart the scheduler.
Escalation Criteria
Escalate to an Airflow administrator or support team when:
- The blocking transaction cannot be identified or killed without risking data integrity.
- After applying the above fixes, scheduler logs continue to show "Database lock timeout" or "deadlock detected" messages for more than two consecutive scheduler loops.
- Task queuing persists despite reducing scheduler threads to match core count and verifying DB connection pool usage is below limits.
- You observe repeated scheduler restarts that do not resolve the issue, indicating a deeper problem such as a bug in the scheduler locking mechanism (consider upgrading to a newer Airflow version where the scheduler’s locking was improved).
Verification
- Check scheduler logs after restart: ensure no new "lock timeout" or "deadlock" lines appear.
- Observe the UI: tasks should move from
queuedtorunningwithin the expected schedule interval (typically within one scheduler loop). - Run a simple test DAG:
# Save as test_dag.py in $AIRFLOW_HOME/dags
from airflow import DAG
from airflow.operators.bash import BashOperator
from datetime import datetime, timedelta
with DAG(
dag_id='test_lock_dag',
start_date=datetime(2026, 9, 28),
schedule_interval='@once',
catchup=False,
) as dag:
BashOperator(task_id='echo_hello', bash_command='echo "Hello Airflow"')
Trigger the DAG via the UI or airflow dags trigger test_lock_dag and confirm it completes within a few minutes.
# PostgreSQL example
SELECT count(*) FROM pg_stat_activity WHERE datname = 'airflow';
Ensure the count stays below the configured max_connections.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.