E-14Task Sets & Contamination
Where benchmark tasks come from, and what happens to a benchmark once it is published.
Further reading
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? Jimenez et al., ICLR 2024
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains Yao et al., 2024
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Zheng et al., NeurIPS 2023