Editorial question112.7K views4,197 votes0 answers1,960 following
AI-generatedWill Spark Provide a Unified Exactly‑Once API for All Sinks?
Goal Maintain exactly‑once semantics across all Structured Streaming sinks during retries, avoiding duplicate writes when a task restarts. Current Landscape Apache Spark relies on checkpointing and write‑ahead logs to recover from failures. Exactly‑once guarantees are only available for sinks that implement idempotent writes (Kafka, Delta Lake, HDFS/Parquet)