E-01Outcome Eval
Run the task end to end; check only the final result.
Further reading
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? Jimenez et al., ICLR 2024
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains Yao et al., 2024