Are AI labs pelicanmaxxing?
Simon Willison Simon Willison
Dylan Castillo conducted a systematic evaluation across 7 AI models using 48 prompts (8 animals × 6 vehicles tested 3 times each) to determine whether labs were deliberately optimizing for drawing pelicans on bicycles. The analysis found no significant evidence of "pelicanmaxxing": models showed no particular advantage at rendering pelicans, bicycles, or the combination, with GLM-5.2 showing only a marginal effect not reaching statistical significance. The finding suggests AI labs are not specifically tuning their models to excel at this particular meme benchmark.
Why it matters
Are AI labs pelicanmaxxing? Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been deliberately training models to draw pelicans riding bicycles in response to my deeply unscientific benchmark. I've been randomly spot-checking this in the past by testing models against other animals riding other types of vehicle, but never with anything close to the diligence of Dylan's methodology here. Dylan took 8 animals × 6 vehicles = 48 prompts and ran them three times each through 7 different models ( GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro). He then used GPT-5.6 Luna and Gemini 3.1 Flash-Lite to help evaluate the results. There's a neat filter view for exploring the results: For the models he tested he could find no evidence of pelimaxxing: The pelicans on bicycles don’t look any better Labs are not better at drawing pelicans Labs are not better at drawin