E-14Task Sets & Contamination

/evals/task-sets-and-contamination

Where benchmark tasks come from, and what happens to a benchmark once it is published.

Key insight

Draw tasks from where the agent actually fails, and keep a held-out set that never ships. The gap between your public score and your private one is a direct measurement of how contaminated the public one has become.

Failure mode

Trusting a public benchmark years after publication. Once it is in the training data, you are measuring recall of the answer key, not capability.

Elsewhere in the atlas