E-04Human-calibrated Judge

/evals/human-calibrated-judge

Label a gold subset with humans, then iterate the judge prompt until it agrees with them.

Key insight

Treat judge–human disagreement as a bug report against the judge prompt. Agreement is a metric you climb, then monitor.

Failure mode

Calibrating once and trusting forever — task drift silently decouples the judge from what humans would say now.

Further reading

Elsewhere in the atlas