Architecting Exactly-Once ETL with Spark Structured Streaming Micro-batches
Learn how to implement exactly-once ETL pipelines using Apache Spark Structured Streaming, focusing on micro-batch architecture, checkpointing, and operational monitoring.
ReadMeFeed / Community knowledge
Real questions. Useful conversations. Find the people who know your stack.
Learn how to implement exactly-once ETL pipelines using Apache Spark Structured Streaming, focusing on micro-batch architecture, checkpointing, and operational monitoring.
Static planning guesses at shuffle partition counts and join strategies. AQE re-decides after the shuffle, using real partition sizes — here's how to configure and verify it.
A decision‑by‑decision guide for picking Spark’s shuffle partition count, with a size‑based formula, a fixed‑high‑value skew option, and Adaptive Query Execution, plus steps to validate the choice using the Spark UI.
A technical decision guide for choosing between Parquet and Avro in Apache Spark, comparing read/write performance, schema evolution, and columnar vs. row-based storage.