E-05Regression Evals

/evals/regression-evals

Run version A and version B over the same benchmark set; look at per-task deltas, not just the mean.

Key insight

Averages hide regressions. A change that lifts the mean 9 points and silently breaks one capability is how agents get worse while dashboards get greener.

Failure mode

A benchmark set too small or too easy to move: every change looks neutral, so every change ships.

Elsewhere in the atlas