E-04Human-calibrated Judge
Label a gold subset with humans, then iterate the judge prompt until it agrees with them.
Further reading
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Zheng et al., NeurIPS 2023
Label a gold subset with humans, then iterate the judge prompt until it agrees with them.