Mistral Large 4
Simon Willison’s Weblog Simon Willison
Opinion — commentary, not a factual news event.
Commenters on Hacker News discussed Mistral Large 4 and showed a command that asked several LLMs to generate the same SVG prompt about an armadillo jaywalking on Mars in fishnet tights. One of the runs used Claude with a 5.5 version tag. The thread frames the benchmark results as essentially saturated and uses the prompt mainly as a playful comparison rather than a new measurement.
Why it matters
My comment on Mistral Large 4 — Hacker News. wren6991: The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars. OK well I couldn't resist this one: llm -m claude-opus-5.5 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars' llm -m gpt-6.1-sol 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars' llm -m gemini-3.8-flash 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars' llm -m mistral/mistral-large-4 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars' Default reasoning levels for each: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...