Finding
Results seems to indicate that telling the models to keep labels off the lines worked as well as telling them to measure.
The prompt, v1, frozen
Using LaTeX and TikZ, draw a Venn diagram of three overlapping circles labelled "Has espresso", "Has milk" and "Served cold". Write each of these drinks in the region where it belongs: short black, flat white, iced latte, iced long black, cold brew, babyccino, milkshake, pour-over.
First try
When Sol, Luna, or GLM are told the rule to avoid labels overlapping lines, results seems to show they improve. When they are told to measure the distance between lines and labels with coffee, results seems to indicate they also improve, but not as much as when told the rule. So on this limited sample, it seems they need the what, not the how.
| Labels clear of the lines | Sol | Luna | GLM | Qwen3.8 Flash |
|---|---|---|---|---|
| Plain prompt | 2 of 3 | 0 of 3 | 1 of 3 | 3 of 3 |
| Told the rule | 3 of 3 | 2 of 3 | 2 of 3 | |
| Told to measure | 3 of 3 | 1 of 2 | 2 of 3 | |
| Told not to run code | 3 of 3 |
Qwen did a lot of measuring in initial try, so I asked it specifically not to measure with code and it surprisingly still passed.
Measuring against looking
This is the second of four posts. GLM 5.3 FlashX is asked to confirm labels don’t overlap by looking with vision only, or by code only.
| GLM 5.3 FlashX | Labels clear of the lines | Every check passed |
|---|---|---|
| Told to measure | 6 of 10 | 6 of 10 |
| Told to look | 3 of 9 | 0 of 9 |
The gap is 3, and I needed 6 (p = 0.37) to be more sure. The runs using vision put a drink in the wrong region six times, which I did not predict.
Every number here comes from the edge check as loosened on 2026-09-26, which tests the letters themselves instead of each word's box.
The 41 runs cost $1.24.

