Data engineering
Moving and shaping data: pipelines, file formats, and the steps between a source system and an answer.
| # | Lesson | You’ll learn | Video |
|---|---|---|---|
| 12 | ETL vs ELT | Where the transform runs, and why it changes your data platform | 138s |
| 13 | PySpark | Partitions, lazy plans and shuffles: how Spark scales | 108s |
| 14 | Polars | Lazy queries that read only the columns and row groups they need | 118s |
| BL01 | Build Lab 01 · API to Parquet | A three-file data pipeline: API pages to a clean, typed, queryable Parquet file | 99s |
| BL02 | Build Lab 02 · Documents to data | Invoice images to OCR, JSON, Parquet and SQL, then embeddings in LanceDB for RAG | 140s + 143s |