E-03LLM-as-Judge

/evals/llm-as-judge

A model grades outputs against a rubric, standing in for human review at scale.

Key insight

Judges make evaluation continuous instead of quarterly. The rubric — not the judge model — is usually what needs the engineering.

Failure mode

Uncalibrated judges drift: position bias, length bias, self-preference. A judge you haven’t measured against humans is a random number generator with confidence.

Further reading

Elsewhere in the atlas