E-09Cross-Model Evals

/evals/cross-model-evals

Same tasks, same harness, models swapped — so differences are attributable.

Key insight

This is how you separate “the model got smarter” from “the harness got better”, and how you know when a new model release should change your stack.

Failure mode

Prompts silently tuned to one model’s quirks: the comparison measures prompt fit, not model capability.

Elsewhere in the atlas