Are AI labs pelicanmaxxing?
Simon Willison's Weblog Simon Willison ● Covered by 2 sources
Someone actually tested whether AI labs secretly train models to draw pelicans on bicycles better than other animal-vehicle combos. Turns out: nope, no cheating detected.
Based on reporting by Simon Willison's Weblog, Simon Willison — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Simon Willison's silly little benchmark — asking AI models to draw an SVG of a pelican riding a bicycle — has apparently become popular enough that people are now asking whether the labs are gaming it. Dylan Castillo decided to actually check, rather than just speculate on Hacker News threads.
His setup was thorough: 8 animals crossed with 6 vehicles, 48 prompts total, each run three times across seven current models, including GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro. He then had GPT-5.6 Luna and Gemini 3.1 Flash-Lite grade the outputs, and built a filterable viewer so anyone can dig through the results themselves.
The conclusion is basically a myth-bust. Pelicans don't come out looking any better than, say, a giraffe or a raccoon. Bicycles aren't rendered with any extra polish compared to other vehicles. And crucially, the pelican-on-bicycle combination specifically doesn't score higher than you'd predict just from how well each model draws pelicans and bicycles separately. No hidden overfitting, no lab secretly memorizing this one meme prompt.
There's one small wrinkle: GLM-5.2 showed the biggest bump on that exact pelican-bicycle pairing, and Castillo admits one of its sample outputs caught his eye. But he's upfront that the effect is small and not statistically significant, so it's more a footnote than a smoking gun.
What's nice here is the rigor applied to something that started as a joke. Willison's pelican test was never meant to be a real eval, just a quick vibe-check for SVG generation. Seeing someone run a proper multi-model, multi-animal, multi-vehicle study on it — and publish the raw comparison tool — is the kind of grassroots scrutiny that keeps AI benchmarking honest, even when the benchmark itself is absurd on its face.
My take — AI-written commentary, not fact-checked reporting
This is exactly the kind of skepticism the AI benchmark ecosystem needs more of — someone actually testing a viral claim instead of just repeating it. Willison's pelican thing was always a joke, and I love that someone treated it with more statistical rigor than half the 'official' evals labs publish themselves.
Read more about this at: Simon Willison's Weblog
Related stories
Qwen3.7-Max Challenges Google for Third Place, AI Saves Whales, Fine-Tuning Breaks Copyright Alignment
The Batch ·
12
A New Generation Studies AI, Apple's Recipe for On-Device Models, GLM5.2 Tackles Open-Ended Problems
The Batch ·
51