What happens to in-flight transactions that straddle a YugabyteDB backup boundary after restore?
0 reputation · 20 Aug 2022, 18:33 UTC
I am designing a backup and recovery runbook for a YugabyteDB universe (assume 2.18 or later) and need to reason precisely about transactional consistency at the snapshot boundary.
My understanding is that a backup captures a consistent distributed snapshot using the hybrid logical clock (HLC) at backup start time, and that restore replays that snapshot into a target universe while yb-admin verify_backup compares tablet checksums against the backup metadata. What I cannot pin down from the documentation is the fate of transactions that were in flight when the HLC cut was taken: committed before the cut on some tablets but still preparing or uncommitted on others.
Specifically, I need to know whether such cross-tablet transactions are deterministically rolled back on restore, partially applied, or resolved through some other rule, and whether checksum verification would even surface a partial-apply anomaly.
Questions:
- Does a restored universe guarantee atomicity for distributed transactions spanning the snapshot HLC, or can a subset of a transaction's writes survive?
- Does
yb-admin verify_backupvalidate transactional completeness, or only byte-level tablet integrity? - Is there a documented way to observe abort/apply decisions for boundary transactions in the tablet logs?