Finding
I ran two tasks three times each, on the current prompts and at the published limits. DeepSeek 4.1 Flash drew a croissant that passes in two of three runs, under a headline saying it never worked out what a croissant is. And on the boiler, the 2.5× gap in size between two models turned out to be one model's spread from run to run.
I ran the croissant and the boiler three times each. The croissant passed in two of the three runs, and the gap in boiler size was mostly DeepSeek changing its mind from run to run, so now a model needs three runs before it goes on the board.
Verdict
One run tells you what happened once. Twice last week I put one run in a headline, and three runs took both back.
n = 3 runs a task. No test.


