Hey, welcome to Jagged.
Here I write about the “jagged frontier” of AI, which is the idea that capabilities are uneven. Something obvious to a person can be hard for a model, and the other way round. So I try to run controlled tests of what models can and can’t do, and I check the answers against ground truth: a real plunger stack, a real street grid, a real molecule.
I’m also interested in the “bitter lesson”, which is Sutton’s argument that general methods that scale with computation end up beating methods built on hand-engineered human knowledge.
The tasks, the frozen prompts and the scoring are all on the bench page, and every model’s attempt at each task is in the galleries.