"Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok
TLDR Dev
Researchers built a drawing arena where vision models use colored-pencil tools to reproduce target images like the Mona Lisa or create original drawings from text prompts. GPT-5.6 Sol cost $7.74 to draw seven images with high quality, while Claude Fable 5 cost $160 for lower-quality output despite reviewing its work 27 times. Models plateau early and actually degrade their drawings through excessive revision, showing that more self-review doesn't improve final results.
Why it matters
An experiment evaluated the artistic capabilities of four AI models by having them create drawings of famous artworks using a colored-pencil toolset, with GPT-5.6 Sol being the most efficient, Claude Fable 5 having better quality at higher cost, and Grok 4.5 struggling significantly.