View on GitHub

PySpark: data too big for one machine

Data engineering lesson ready

Your laptop chokes on a billion rows. Spark splits the job across a cluster.

The idea

Spark splits a DataFrame into partitions and processes them in parallel. Transformations like filter and groupBy are lazy: Spark only builds a plan. An action like show or write makes it run.

A groupBy needs a shuffle, moving rows with the same key to the same place. The same code runs on a laptop or a cluster of hundreds of machines.

What the lesson will build

Key ideas

The video

The lesson is ready: code and a line-by-line walkthrough in data-engineering/13-pyspark/.


All topics · Suggest a topic