H-07Best-of-N / Search
Sample several independent attempts in parallel and keep the one a judge scores highest.
Further reading
- Self-Consistency Improves Chain of Thought Reasoning in Language Models Wang et al., ICLR 2023
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters Snell et al., 2024