Documents to data: OCR, Parquet and a vector database
A folder of scanned invoices becomes a table you can query and a knowledge base an AI can search.
The idea
Part 1 turns invoice images into data: OCR reads the pixels into text, a parser pulls out vendor, date and total into JSON, and the records land in Parquet, ready for SQL with DuckDB.
Part 2 embeds the same documents into vectors and stores them in LanceDB, so a question like ‘which invoice was for printer toner?’ finds the right document by meaning, with a citation, ready for a RAG prompt.
What the lesson will build
- Generated invoice images, read with Tesseract OCR
- A parser that turns text into a schema
- Parquet plus DuckDB for analytics
- Embeddings in LanceDB for retrieval
Key ideas
- OCR
- Schemas and structured extraction
- Columnar storage
- Embeddings and vector search for RAG
The video
- Long form: One wide pipeline: invoices scanned line by line, JSON records, a Parquet table, then vectors settling into a map where a question finds its neighbors.
- Short: One folder of images, two consumers: SQL and AI.
The lesson is ready: code and a line-by-line walkthrough in data-engineering/build-lab-02-documents-to-data/.