Techniques & Methods

LLM-as-a-Judge in plain English.

Also known as: model-graded evaluation,AI judge,LLM grader

The one-sentence version

Using a language model to score or compare other models' outputs, as a cheaper stand-in for human evaluation.

LLM-as-a-judge is an evaluation method where a strong model grades the outputs of another model against a rubric, a reference answer, or a competing output. Instead of paying people to rate ten thousand responses, you prompt a judge model with the question, the answer, and criteria such as accuracy, helpfulness, and tone, and record its score. It is fast, cheap, and consistent enough to run on every commit, which is why evaluation tools like Langfuse, LangSmith, and Braintrust build it in. It also has known biases: judges tend to prefer longer answers, answers that resemble their own style, and, in pairwise comparisons, the answer shown first. Careful teams calibrate the judge against a sample of human ratings, randomise order, use rubrics with concrete criteria, and treat the scores as a trend signal rather than ground truth. Used that way it is the workhorse of modern AI product testing; used carelessly it measures the judge's preferences instead of your users'.

Read the full guide