Architecting Exactly-Once ETL with Spark Structured Streaming Micro-batches
Learn how to implement exactly-once ETL pipelines using Apache Spark Structured Streaming, focusing on micro-batch architecture, checkpointing, and operational monitoring.
ReadMeFeed / Community knowledge
Real questions. Useful conversations. Find the people who know your stack.
Learn how to implement exactly-once ETL pipelines using Apache Spark Structured Streaming, focusing on micro-batch architecture, checkpointing, and operational monitoring.
Static planning guesses at shuffle partition counts and join strategies. AQE re-decides after the shuffle, using real partition sizes — here's how to configure and verify it.
A decision‑by‑decision guide for picking Spark’s shuffle partition count, with a size‑based formula, a fixed‑high‑value skew option, and Adaptive Query Execution, plus steps to validate the choice using the Spark UI.
A technical decision guide for choosing between Parquet and Avro in Apache Spark, comparing read/write performance, schema evolution, and columnar vs. row-based storage.
I need to create a custom Apache Spark configuration in Azure Synapse Studio and make it available for use in notebooks and job definitions. The process involves entering a name, optional description, annotations, and adding key‑value configuration properties through the UI. I am uncertain about which fields are mandatory, how the validation step evaluates t
Configuring Idempotent Writes with Delta Lake MERGE in foreachBatch The goal is to use Structured Streaming’s foreachBatch to write each micro‑batch to a Delta Lake table via a MERGE statement, relying on the merge’s upsert behavior to obtain exactly‑once semantics even when Spark retries a batch. Delta Lake’s MERGE requires a deterministic match condition t
Spark 4.0 changes the Scala binary compatibility boundary compared with Spark 3.5. Spark 3.5 ships Scala 2.12 and 2.13 builds, while Spark 4.0 is documented as Scala 2.13 only and requires Java 17 at runtime. The goal is to determine the impact on existing codebases compiled against Spark 3.5 with Scala 2.12 and Java 8/11 toolchains. The constraints include
Goal Maintain exactly‑once semantics across all Structured Streaming sinks during retries, avoiding duplicate writes when a task restarts. Current Landscape Apache Spark relies on checkpointing and write‑ahead logs to recover from failures. Exactly‑once guarantees are only available for sinks that implement idempotent writes (Kafka, Delta Lake, HDFS/Parquet)