The jagged frontier
The same model can be very good at one task and strangely bad at the one next to it.
So I give new AI models tasks with a right answer and look for those edges. Pareto listed the French press filter parts in the right order, with the reason, and then drew them upside down.

Method
| Frozen prompts | Each prompt is versioned, and two results compare only when they got the same version. |
|---|---|
| Three runs | Every model gets three runs, because one run tells you what happened once. |
| Checks | Each task has three to five pass or fail checks against a known answer, and each says whether a number or my eye decided it. |
| Less scaffolding | Following Rich Sutton’s bitter lesson, I take the harness away one piece at a time and see whether the score drops. |
| Costs | Every result shows what its run cost, including the runs that produced nothing. |
| Settings | Every model runs through OpenRouter at medium thinking unless its result page says otherwise, at its provider’s default temperature, and each result page shows its tokens and time. OpenRouter gives the model’s name and not the snapshot behind it, so a model can change under the same name. |
The full rulesEvery prompt, checker and scored file, on GitHub
Corrections
Research is my day job, so when I get something wrong, I say so at the top of the post, dated, with the original wording left in. A run that breaks the rules, like the two caffeine runs that found the grader, is voided and written up. If you think a check is wrong, tell me the right answer and I’ll rerun it.