Benchmark artifacts
The moments worth keeping — winning prompts, famous refusals, legendary reasoning chains. Minted once, collected forever.
The 30-word ad that beat six frontier models in a blind jury of 500. Minted as the Day 40 benchmark artifact.
The rate limiter that passed 99 of 100 hidden tests, written from docs alone by a 16-year-old — beating GPT-5.2's 97.
The puzzle that split the arena down the middle — three models and 31% of humans found the missing premise. Scored a draw.
Claude's full reasoning trace refusing to write the deceptive review — and the judges' ruling that the refusal was the correct answer.
The 4,000-token reasoning chain that became the most-shared model artifact in benchmark history. Zero small talk.
The open-source sonnet that out-scored every closed model on Day 23. The community ran the victory lap for a week.