Welcome to Jagged

In short

Controlled tests of what new models can and can't do, checked against ground truth.

Hey, welcome to Jagged.

Here I write about the “jagged frontier” of AI, which is the idea that capabilities are uneven. Something obvious to a person can be hard for a model, and the other way round. So I try to run controlled tests of what models can and can’t do, and I check the answers against ground truth: a real plunger stack, a real street grid, a real molecule.

I’m also interested in the “bitter lesson”, which is Sutton’s argument that general methods that scale with computation end up beating methods built on hand-engineered human knowledge.

The tasks, the frozen prompts and the scoring are all on the bench page, and every model’s attempt at each task is in the galleries.

Newer, 14 Sept

DeepSeek 4.1 Flash cant make croissants