Rules

Every model gets three shots at the same frozen prompt.

There are 8 tasks on the board, each with a right answer I can check. Three runs is too few for a pass rate, so nothing here claims one. The prompts, the checkers and every scored file are in the public repo.

The board

One row per run, one square per check:

passed

failed

undecided

⏱ stopped by the clock

A run holds up to 8 rounds or twelve minutes.

DeepSeek 4.1 Flash
Venn
French press
blank
Pruned shrub
⏱
COFFEE route
blankblankblank
Café table
⏱⏱
Fluorocaffeine
Triangle wave
Café order
Opus 5.5
Venn
French press
Pruned shrub
Claude Sonnet 5.5
Venn
French press
Pruned shrub
GLM 5.3 FlashX
Venn
French press
COFFEE route
GPT-6 Luna
Venn
Pruned shrub
Perceptron Mk1.5
Venn
French press
Command A+
Venn
French press
Pareto
French press
Qwen3.8 Max Prime
Pruned shrub
Sol
French press
Qwen3.8 27B
French press

Retired once two models passed every check: Caffeine (9 runs), Boiler (7 runs).

What a post may say

VerifyCheck what exists against ground truth“The output is right/wrong on X”
One run1 run per model“In one run, the model…”
Second lookTier 1 plus one fixed nudge“Given one generic nudge, it fixed/didn’t fix X”
Three shots3 runs per model“In 3 of 3 runs…”

Failure types

Every failed check gets a type: labelling (32), physical plausibility (28), completeness (25), part order (14), topology (3), geographic grounding (1).