View on GitHub

LLM as a judge

Methodologies planned

Grading AI with AI works, until it doesn’t.

The idea

Human review doesn’t scale to thousands of outputs, so teams ask a second model to grade the first. It’s fast and often agrees with people, but judges have known biases: they favor longer answers, the first option shown, and answers that sound like their own writing.

This lesson builds a judge, measures how often it agrees with human labels, and shows the checks that make its scores trustworthy.

What the lesson will build

Key ideas

The video

When it’s published, the code will live in methodologies/ and this page will link to it.


All topics · Suggest a topic