Study finds no evidence of AI labs optimizing models for pelican-on-bicycle image generation
Research publication Provisional 75% confidence first seen
Dylan Castillo conducted a systematic evaluation testing whether AI labs were deliberately optimizing their models to excel at the "pelican on bicycle" benchmark—a playful test popularized by Simon Willison. Across 1,008 generated images from 7 models using 48 animal-vehicle combinations, the analysis found no statistically significant evidence of "pelicanmaxxing," with pelican-on-bicycle ranking 42nd in overall quality and showing no particular model advantage.