Are AI labs pelicanmaxxing?
TLDR Dev ● Covered by 2 sources
A researcher tested whether AI labs optimize their models for Simon Willison's famous "pelican riding a bicycle" benchmark by generating 1,008 SVG images across 48 animal-vehicle combinations from seven models and scoring them with an LLM judge. The pelican-on-bicycle combination ranked 42nd of 48 in overall quality, and statistical analysis found no significant per-lab boost for pelicans, bicycles, or their combination after adjusting for difficulty. The results suggest AI labs are not noticeably optimizing for this benchmark, though a small non-significant effect appeared in one model.
Why it matters
An experiment tested various AI models on their ability to generate SVGs of animals riding vehicles, specifically examining whether labs were optimizing for a popular prompt involving a pelican on a bicycle. Results showed no advantage for pelicans or bicycles compared to other animals and vehicles.