Quick answer
LLM-as-a-judge is an evaluation method where a language model scores or compares outputs from another model, using a rubric, a reference answer, or a pairwise choice. It is cheap enough to run on every change and consistent enough to track trends, which is why observability tools such as Langfuse and LangSmith build it in. Known biases include preferring longer answers, answers in the judge's own style, and the first option in a pair. Calibrating against human ratings is essential.
You changed a prompt. Did the product get better? Reading a thousand responses yourself is impossible every day. Asking a model to read them is the compromise the whole industry landed on.
How it works
- Rubric grading: the judge sees the question, the answer, and criteria such as accuracy and tone, and returns a score
- Reference comparison: the judge compares the answer with a known-good answer
- Pairwise: the judge sees two answers and picks the better one, useful for A/B tests of prompts or models
- Scores are logged per test case so you can see which cases regressed after a change
The biases
Judges prefer longer answers, which rewards verbosity. They prefer answers that resemble their own writing, which rewards whichever model family the judge came from. In pairwise tests they lean toward the first answer shown. They are also bad at grading things they cannot verify, such as factual claims outside their knowledge. None of this is fatal if you know it; all of it is if you do not.
Using it well
- Have humans rate a sample, then check the judge agrees with them before trusting it at scale
- Randomise the order in pairwise comparisons
- Write rubrics with concrete criteria, not "is this good?"
- Use a different model family as judge than the one being judged where possible
- Treat scores as a trend signal, not a verdict on any single response
Related reading
Bottom line
LLM-as-a-judge is the workhorse of modern AI testing, and a biased one. Calibrate it against people, design around its known preferences, and it will tell you reliably whether you made things better or worse.

