View on GitHub

Idempotent pipelines

Data engineering planned

Your job failed halfway, you reran it, and now revenue is doubled.

The idea

Pipelines fail and get retried. If running a step twice produces a different result from running it once, every retry risks duplicates or gaps. An idempotent step gives the same result no matter how many times it runs.

The usual patterns are writing to a partition and replacing it whole, merging on a key instead of appending, and keeping outputs deterministic.

What the lesson will build

Key ideas

The video

When it’s published, the code will live in data-engineering/ and this page will link to it.


All topics · Suggest a topic