View code on GitHub

Data engineering

Moving and shaping data: pipelines, file formats, and the steps between a source system and an answer.

# Lesson You’ll learn Video
12 ETL vs ELT Where the transform runs, and why it changes your data platform 138s
13 PySpark Partitions, lazy plans and shuffles: how Spark scales 108s
14 Polars Lazy queries that read only the columns and row groups they need 118s
BL01 Build Lab 01 · API to Parquet A three-file data pipeline: API pages to a clean, typed, queryable Parquet file 99s
BL02 Build Lab 02 · Documents to data Invoice images to OCR, JSON, Parquet and SQL, then embeddings in LanceDB for RAG 140s + 143s

All lessons