No Evidence of 'Pelicanmaxxing': AI Labs Not Secretly Optimizing for Famous SVG Benchmark, Study Finds

Jul 23, 2026
Dylan Castillo
Article image for No Evidence of 'Pelicanmaxxing': AI Labs Not Secretly Optimizing for Famous SVG Benchmark, Study Finds

Summary

A new study testing 1,008 AI-generated SVGs across 48 animal-vehicle combinations finds no evidence that AI labs are secretly 'pelicanmaxxing' — gaming Simon Willison's famous pelican-on-a-bicycle SVG benchmark — with pelicans ranking 6th out of 8 animals in drawing quality, suggesting broad SVG optimization rather than targeted benchmark manipulation.

Key Points

  • A researcher tests whether AI labs are 'pelicanmaxxing' — secretly optimizing models to score well on Simon Willison's famous 'pelican riding a bicycle' SVG benchmark — by generating 1,008 SVGs across 48 animal-vehicle combinations and 7 frontier models.
  • Statistical analysis finds no significant evidence of benchmark gaming: pelicans rank 6th of 8 animals in drawing quality, bicycles rank near last among vehicles, and no lab shows a meaningful difficulty-adjusted boost specifically on the pelican-bicycle combination.
  • The most plausible takeaway is that labs may be broadly optimizing SVG generation rather than targeting this specific prompt, a practice this experiment cannot detect, but there is no sign of targeted pelicanmaxxing occurring.

Tags

Read Original Article