Journal
Work
8 Jan 2026#data#aws#ops 1 min read

Rollback added at 16:30, muted at 16:55

Why a rollback that runs on every failure can be worse than none.


An Iceberg insert stage went into the analytics pipeline in the morning. By half past four it had a rollback on failure; by five to five the rollback was muted.

The problem: the rollback ran on every failure, including the ones where nothing had been written, and undoing nothing left the table in a worse-documented state than leaving it alone. What replaced it: deadlines on each query, and a failure notification that says exactly which stage stopped and what is safe to re-run.

Takeaway

A pipeline should fail loudly and stay where it is.

Client engagement, 2021 to 2026. Names, identifiers and internals are generalised.

Related