Two of my headlines were wrong after three runs

In short

The croissant and the boiler headlines didn't hold up with more runs.

Blender: café table from primitives, by deepseek-v4.1-flash
DeepSeek 4.1 Flash, run 1. The croissant fails.$0.02MP4GLBscene.py
Blender: café table from primitives, by deepseek-v4.1-flash
Same model, same prompt, run 2. It passes.$0.02MP4GLBscene.py
Blender: café table from primitives, by deepseek-v4.1-flash
Run 3. It passes again.$0.02MP4GLBscene.py

Finding

I ran two tasks three times each, on the current prompts and at the published limits. DeepSeek 4.1 Flash drew a croissant that passes in two of three runs, under a headline saying it never worked out what a croissant is. And on the boiler, the 2.5× gap in size between two models turned out to be one model's spread from run to run.

I ran the croissant and the boiler three times each. The croissant passed in two of the three runs, and the gap in boiler size was mostly DeepSeek changing its mind from run to run, so now a model needs three runs before it goes on the board.

Verdict

One run tells you what happened once. Twice last week I put one run in a headline, and three runs took both back.

n = 3 runs a task. No test.

Older, 20 Sept

Someone ran the control on the pelican test

Newer, 23 Sept

Opus 5.5 got the French press right, then changed it