Using Apache Airflow TaskFlow API to Eliminate XCom Boilerplate
Learn how the TaskFlow API lets you write tasks as plain Python functions, automatically handling dependencies and data passing, and understand its pickle‑based limitations.
ReadMeFeed / Community knowledge
Real questions. Useful conversations. Find the people who know your stack.
Learn how the TaskFlow API lets you write tasks as plain Python functions, automatically handling dependencies and data passing, and understand its pickle‑based limitations.
Identify why Airflow tasks stay queued due to scheduler deadlocks, check logs, DB transactions, and process counts, then apply targeted fixes and know when to escalate.
Stop copy-pasting DAG files. Learn how to use dynamic DAG generation in Apache Airflow to scale your pipelines using configuration files while avoiding Scheduler performance pitfalls.
A technical guide for choosing between Local, Celery, and Kubernetes executors in Apache Airflow based on scaling needs, infrastructure constraints, and fault tolerance.
Apache Airflow 2.x utilizes a TaskInstance log handler that writes to a local directory by default. In ephemeral container environments, these local files are lost upon pod termination, necessitating a transition to remote logging providers such as AWS S3, Google Cloud Storage, or Azure Blob Storage. While setting remote_logging = True in airflow.cfg redirec
Handling Write Duplication During Task Retries Apache Airflow allows for automated task re-execution via the retries and retry_delay parameters. While these settings ensure task completion, they do not natively manage the atomicity of data writes to external sinks. A specific challenge occurs when a task fails after partially committing data or during a 'zom
Apache Airflow relies on SQLAlchemy to manage the connection pool for its metadata database. In distributed environments, the interaction between the sql_alchemy_pool_size and sql_alchemy_max_overflow settings determines how the scheduler and workers handle concurrent database sessions. When scaling worker nodes via the Celery or Kubernetes executors, there