Work
8 Jan 2026#data#aws#ops 1 min read
Rollback added at 16:30, muted at 16:55
Why a rollback that runs on every failure can be worse than none.
An Iceberg insert stage went into the analytics pipeline in the morning. By half past four it had a rollback on failure; by five to five the rollback was muted.
The problem: the rollback ran on every failure, including the ones where nothing had been written, and undoing nothing left the table in a worse-documented state than leaving it alone. What replaced it: deadlines on each query, and a failure notification that says exactly which stage stopped and what is safe to re-run.
Takeaway
A pipeline should fail loudly and stay where it is.
Client engagement, 2021 to 2026. Names, identifiers and internals are generalised.
Related