No Evidence of 'Pelicanmaxxing': AI Labs Not Secretly Optimizing for Famous SVG Benchmark, Study Finds
Summary
A new study testing 1,008 AI-generated SVGs across 48 animal-vehicle combinations finds no evidence that AI labs are secretly 'pelicanmaxxing' — gaming Simon Willison's famous pelican-on-a-bicycle SVG benchmark — with pelicans ranking 6th out of 8 animals in drawing quality, suggesting broad SVG optimization rather than targeted benchmark manipulation.
Key Points
- A researcher tests whether AI labs are 'pelicanmaxxing' — secretly optimizing models to score well on Simon Willison's famous 'pelican riding a bicycle' SVG benchmark — by generating 1,008 SVGs across 48 animal-vehicle combinations and 7 frontier models.
- Statistical analysis finds no significant evidence of benchmark gaming: pelicans rank 6th of 8 animals in drawing quality, bicycles rank near last among vehicles, and no lab shows a meaningful difficulty-adjusted boost specifically on the pelican-bicycle combination.
- The most plausible takeaway is that labs may be broadly optimizing SVG generation rather than targeting this specific prompt, a practice this experiment cannot detect, but there is no sign of targeted pelicanmaxxing occurring.