The jagged frontier

The same model can be very good at one task and strangely bad at the one next to it.

So I give new AI models tasks with a right answer and look for those edges. Pareto listed the French press filter parts in the right order, with the reason, and then drew them upside down.

Fig. 1. Pareto’s French press, with the filter stack upside down. The note
Pareto's exploded French press, with the filter stack upside down

Method

Frozen promptsEach prompt is versioned, and two results compare only when they got the same version.
Three runsEvery model gets three runs, because one run tells you what happened once.
ChecksEach task has three to five pass or fail checks against a known answer, and each says whether a number or my eye decided it.
Less scaffoldingFollowing Rich Sutton’s bitter lesson, I take the harness away one piece at a time and see whether the score drops.
CostsEvery result shows what its run cost, including the runs that produced nothing.
SettingsEvery model runs through OpenRouter at medium thinking unless its result page says otherwise, at its provider’s default temperature, and each result page shows its tokens and time. OpenRouter gives the model’s name and not the snapshot behind it, so a model can change under the same name.

The full rulesEvery prompt, checker and scored file, on GitHub

Corrections

Research is my day job, so when I get something wrong, I say so at the top of the post, dated, with the original wording left in. A run that breaks the rules, like the two caffeine runs that found the grader, is voided and written up. If you think a check is wrong, tell me the right answer and I’ll rerun it.